The NVIDIA A10, A30 and A40 solve different buying problems despite sharing the Ampere generation. A10 combines a single-slot package with graphics and inference capabilities. A30 adds hardware partitioning through MIG and HBM2 memory. A40 offers 48GB for larger working sets and graphics workloads. Choose the capability your application actually needs, then validate the server and software combination.
The numbers are identifiers, not a ranking
A40 is not automatically a better purchase than A30, and A30 is not simply an A10 with a larger number. Begin with a short list of requirements that can rule a device in or out: usable memory for the intended application, hardware partitioning, rendering features, video codec support, physical clearance and deployment software.
The table describes manufacturer reference cards. It does not establish an offered unit's condition, current availability or performance in your application. Memory bandwidth figures are published peaks, not measured application throughput.
| Reference characteristic | A10 | A30 | A40 |
|---|---|---|---|
| Memory per accelerator | 24GB GDDR6 | 24GB HBM2 | 48GB GDDR6 |
| Peak memory bandwidth | 600 GB/s | 933 GB/s | 696 GB/s |
| Maximum board power | 150W | 165W | 300W |
| Reference form factor | Full-height, single-slot | Full-height, dual-slot | Full-height, dual-slot |
| Cooling | Passive | Passive | Passive |
| MIG | Not supported | Up to four instances | Not supported |
| Auxiliary power type | PCIe eight-pin | CPU eight-pin | CPU eight-pin |
Sources: NVIDIA A10 product brief, NVIDIA A30 product brief and NVIDIA A40 datasheet. The connector names are deliberately different; treat the cable specification as part of compatibility.
Separate memory capacity from memory bandwidth
A10 and A30 both advertise 24GB, but use different memory systems. Capacity answers whether the planned working set can fit. Bandwidth describes a maximum rate of movement through a particular memory interface. Neither number alone predicts tokens per second, batch capacity or the time to complete a rendering job.
For inference, create a memory budget that includes model weights, the KV cache and runtime allocations. Precision, context length and simultaneous sequences change that budget. Use the intended model configuration and serving engine rather than copying a memory estimate from a different deployment. Hugging Face's cache documentation explains why cache implementation and offloading choices also matter. Source: Hugging Face KV cache guide
For a deliberately simplified example, ten billion parameters stored at two bytes each occupy 20 billion bytes before cache and other allocations. That arithmetic is not evidence that the model will operate safely inside a 24GB card. Nor does it establish a useful concurrency level. The inference memory guide walks through a fuller planning approach.
A40's larger capacity can make it a candidate when a 24GB budget is insufficient. That is a reason to investigate it, not a speed prediction. Ask for a representative workload run with the actual model, precision, software versions and latency requirement. Keep configuration and results together so a later offer can be compared on the same basis.
A30's MIG support changes the deployment question
Among these three devices, A30 supports Multi-Instance GPU. NVIDIA lists a maximum of four instances for A30. A10 and A40 are not interchangeable MIG alternatives. MIG availability must be checked against the specific product rather than inferred from the word Ampere. Source: NVIDIA supported MIG GPUs
This matters when the requirement is several separately allocated workloads rather than one application using the whole card. However, “up to four” does not mean four independent applications each receive the entire card's resources. Select the supported profile and verify that its memory and compute allocation fit each workload. NVIDIA documents the supported profile configurations separately. Source: NVIDIA MIG profiles
Scroll across the diagram to read every label
Download diagramWrite the isolation requirement explicitly in the buying brief. “Several users” is not precise enough: determine whether they need dedicated hardware partitions, an approved virtual GPU configuration, or simply a scheduler controlling access to a whole device. Have the platform owner confirm the deployment mode before requesting hardware.
Graphics, video encoding and video decoding are separate checks
A graphics application, a video encoder and a video decoder can exercise different capabilities. Do not use successful CUDA inference as evidence that a card will also satisfy a rendering or transcoding requirement. The A10 brief describes graphics and video workloads, while the A40 datasheet lists RT cores and display functionality. Those specifications are relevant starting points for application qualification, not universal software certification.
For video, first name the input and output codecs, resolution, bit depth, chroma format, frame rate and required simultaneous streams. Check encoding and decoding independently in NVIDIA's current support matrix. A “supported” codec entry does not establish your application's sustained throughput, quality settings or license entitlement. Source: NVIDIA video encode and decode support matrix
Then reproduce the intended pipeline in the application you will deploy. Record whether processing falls back to the CPU, which software build is used and what measurement defines success. This is especially useful when a workload includes both inference and video processing: a nominally suitable accelerator can still leave another stage as the practical constraint.
For remote graphics, ask which application, guest operating system and virtualization product are involved. Do not make physical display connectors the only selection criterion for a remotely accessed workload. Conversely, if local display outputs matter, confirm their operating mode and the actual offered card instead of assuming that every datacenter accelerator behaves like a workstation graphics card.
A PCIe slot is only the beginning of compatibility
All three reference cards are passive. They depend on a suitable server cooling arrangement rather than an onboard fan solving the installation problem. A10's thinner package may help a physical layout, but it does not establish how many cards a particular server can support. Read the OEM's GPU population, thermal and power restrictions for the exact chassis configuration.
The different auxiliary connector types in the table deserve attention. Obtain the server manufacturer's approved cable and installation configuration. Do not select an adapter because the connector appears to have eight positions. Have qualified personnel check the intended installation; this guide is not a wiring procedure.
Keep the board-power rating separate from wall power. CPUs, memory, storage, networking, fans and conversion losses also contribute to a server's demand. A simple sum of GPU ratings is an accelerator budget, not a complete rack allocation. The power and cooling guide provides a planning worksheet and explains that boundary.
Include the server model, chassis revision, risers, GPU kit and intended card count in the request. Ask the supplier to identify which accessories are included. Where a multi-GPU bridge is required, confirm the supported bridge and slot spacing separately. A loose card and a complete qualified installation are different offer scopes.
Treat historical minimum versions as history
Manufacturer product briefs can list the driver or CUDA release available when the hardware launched. Those tables are useful provenance, but they are not a recommendation to deploy an old release today. Validate the current application, framework, driver, operating system and any virtualization manager as one configuration.
NVIDIA's CUDA compatibility documentation explains the relationship between toolkit and driver compatibility. Use it alongside the application's requirements; the presence of an accelerator on a hardware list does not replace that software review. Source: NVIDIA CUDA compatibility documentation
Where virtual GPU software is involved, consult the support matrix for the intended software release and hypervisor. Obtain written clarification of required software entitlements and what the proposed transaction includes. Do not assume that possession of a used card transfers a previous owner's subscription or support agreement. Source: NVIDIA virtual GPU software documentation
For a used offer, ask for unit identifiers and available seller diagnostic records. Keep evidence of hardware condition separate from evidence that your application stack is supported. A successful device query is useful identification evidence; it is not a complete application acceptance test.
Turn the comparison into an exact request
Start with the relevant variant: A10 PCIe 24GB, A30 PCIe 24GB or A40 PCIe 48GB. Include quantity, destination, timing, server configuration and the capability that makes this model necessary.
A useful A30 brief might state that supported MIG profiles are required and ask the platform team to supply the intended profile layout. An A40 brief might identify a measured memory requirement exceeding the available budget on a smaller card. An A10 brief might identify the qualified single-slot layout and application stack. These are hypothetical examples of clearer requirements, not claims that any configuration is already validated.
Ask for the exact part number, seller-described condition, included accessories, available diagnostic evidence and offer-specific warranty terms. Cardinal can use that brief to source hardware and review seller documentation. The comparison itself does not certify a particular unit or establish a price.
Common questions
Is A30 better than A10 because both have 24GB?
No single ranking follows from capacity. A30 has a different memory system and MIG support; A10 has a single-slot reference package. Select the required features and compare representative application measurements.
Does A40 support MIG?
No. Among A10, A30 and A40, A30 is the MIG option, with up to four instances. Verify the supported profiles and software configuration before deployment.
Does a 48GB A40 guarantee that an application will fit?
No. Budget all relevant allocations and validate the actual runtime. The advertised capacity alone does not establish usable workload headroom or a concurrency target.
Can these cards go into any desktop with a PCIe slot?
Do not assume compatibility. These reference cards use passive cooling and require appropriate airflow, power connections, clearance and software support. Have the complete system configuration checked before buying.