H100 PCIe, H100 SXM and H100 NVL are different buying decisions. Start with the server you can operate and the memory your workload needs. Then compare the exact hardware, rather than choosing from the H100 name alone.

Manufacturer specifications and original buying analysis · No measured performance claims

What changes between the three H100 variants

The standard H100 PCIe card has 80 GB of HBM2e. H100 SXM5 has 80 GB of HBM3. H100 NVL has 94 GB of HBM3 per card. Those memory differences sit alongside different cooling, power and interconnect requirements.

Published specifications per accelerator
VariantMemoryMemory bandwidthPower ratingPhysical integration
H100 PCIe 80GB80 GB HBM2e2,000 GB/sUp to 350 WDual-slot PCIe card
H100 SXM5 80GB80 GB HBM33,350 GB/sUp to 700 WSXM5 module
H100 NVL 94GB94 GB HBM33,900 GB/s350–400 WDual-slot PCIe card

The figures above come from NVIDIA's H100 specification table, PCIe product brief and Hopper architecture description. Power is an accelerator rating, not the complete server's consumption. Confirm the offered OEM configuration before ordering.

For procurement, write the variant before the quantity. “Eight H100 GPUs” leaves substantial ambiguity. “Eight H100 SXM5 80GB modules installed in a specified server” tells the supplier what must match. Add the server manufacturer and part number when you have them.

Do the same for an NVL request. State whether the quantity counts individual cards or pairs. State whether the bridges belong in the offer. A seller's photograph of two connected cards should not decide how you interpret the line-item price.

An NVL pair is not a single 188 GB GPU

Two H100 NVL cards provide 188 GB in aggregate, with 94 GB physically attached to each GPU. Your application still needs an appropriate distribution strategy. Adding the capacities does not establish that every tensor can reside wherever the application expects.

H100 NVL memory belongs to two separate acceleratorsOne card has 94 GB. A second card has 94 GB. An NVLink connection joins the cards, while the application determines how work and memory are distributed.H100 NVL card94 GBH100 NVL card94 GBNVLink188 GB aggregate across two devicesApplication distribution still matterscardinalsilicon.com

Scroll across the diagram to read every label

Download diagram
Original schematic of per-card and aggregate capacity, not a bandwidth or performance measurement

Consider a hypothetical model with 70 billion parameters stored at two bytes per parameter. Its weights alone occupy 140 decimal GB. A single 80 GB or 94 GB H100 cannot hold that idealized weight allocation. Two cards have more aggregate capacity, but weights are only one part of the memory budget.

Reserve space for the KV cache, activations and runtime allocations. Record the actual model artifact, because quantization may change weight storage and introduce metadata. Longer prompts or more simultaneous requests can change the peak without changing the parameter count. See the worked memory explanation before selecting a card count.

A practical buying test is to run the intended request distribution, including your longest supported context. Measure peak allocated and reserved memory where the framework exposes both. Repeat the test at the desired concurrency. A successful short prompt is weak evidence for a production service with longer requests.

The server platform can decide the shortlist

If you already have a PCIe server

Ask the server OEM whether the exact H100 PCIe or NVL part number is supported. Provide the chassis, motherboard and riser identifiers. Include the current power supplies and cooling configuration. An available PCIe slot alone does not answer the installation question.

NVIDIA's PCIe brief describes passive cooling and a 16-pin auxiliary power connector. It also distinguishes supported power configurations. Treat a missing cable or unqualified adapter as an unresolved integration item, rather than an inexpensive detail to fix later. Source: H100 PCIe product brief, power and thermal sections

For two cards, ask the seller to identify the slot spacing and bridge assembly. For a larger installation, request a diagram of the PCIe switches and CPU connections. Keep this diagram with the quote so the delivered topology can be compared with the proposed system.

If you are buying a complete SXM system

Evaluate the server as a complete system. A low module price is not a complete deployment price. Ask whether the baseboard, heatsinks, host CPUs, memory, networking and storage are included. Keep the accelerator quantity separate from the number of servers.

NVIDIA's DGX H100 documentation illustrates the distinction: its eight H100 GPUs provide 640 GB of aggregate GPU memory, within a much larger system configuration. That is a system example, not a statement that every SXM offer has the same components. Source: DGX H100/H200 system introduction

Have the facility operator review the proposed rack power and cooling requirements. Use the OEM system documentation, not eight times the GPU power rating as the entire budget. Confirm management access and service procedures before planning a maintenance window.

Compare performance under the same conditions

Memory bandwidth, tensor throughput and application throughput describe different things. Higher published bandwidth can matter when data movement is limiting execution. It does not establish a fixed speedup across applications. A buyer needs evidence for the workload that creates business value.

Write a benchmark contract before comparing offers. For inference, specify the model revision, numerical format, serving engine, prompt lengths, output lengths and concurrency. Choose a latency target alongside throughput. A high aggregate token rate may be unsuitable if individual requests wait too long.

For training, hold the model, optimizer, batch configuration and convergence target constant. Compare elapsed time to the same useful outcome. A shorter iteration is not sufficient if the training settings changed the work performed or the quality achieved.

Mark theoretical figures as theoretical. NVIDIA's H100 table identifies sparse tensor results with a footnote. Do not compare a sparse figure from one column with dense performance from another source. Keep the numerical precision visible beside every compute figure. Source: H100 specifications and sparsity footnote

Resolve the interconnect specification carefully

NVIDIA lists 900 GB/s NVLink for SXM and 600 GB/s for NVL. The standard PCIe brief contains a discrepancy: its introduction says 900 GB/s, while its dedicated NVLink table gives 600 GB/s. Cardinal uses the explicit table value and records the discrepancy rather than silently choosing the larger number.

Ask for the actual topology and a relevant communication test when distributed performance matters. NVIDIA's NCCL tests provide collective-communication checks. Their results still need the test configuration and software version. A communication benchmark is useful evidence, but it does not replace the full application benchmark.

Check the software before treating the hardware as interchangeable

Record the operating system, driver, CUDA runtime and application build. Add the container image digest if you deploy containers. This creates a reproducible target for the seller's demonstration and your own acceptance work.

CUDA compatibility has conditions. NVIDIA documents minimum driver requirements and limitations for minor-version compatibility. A statement that a server “supports CUDA” is too broad to establish that your application image will run correctly. Source: CUDA minor-version compatibility

Separate compatibility from optimization. An application can execute without using the hardware efficiently. Ask whether the intended framework and kernels support your selected precision. Keep an existing working image available while evaluating changes, so hardware and software differences do not become impossible to separate.

Software subscriptions need their own evidence in a used transaction. Do not infer transferable entitlement from a product family's original bundle description. Ask the seller to state the license included, the remaining term and the basis for transfer. If nothing is established, compare the hardware offer without assuming that entitlement.

Evaluate a used H100 offer as a documented configuration

Request photographs of the actual units and readable part-number labels. Ask for a unit list linked to the seller's diagnostic records. The list should make it possible to identify which evidence belongs to which offered card or module.

NVIDIA's management tool exposes identifiers and status information, but a pasted screenshot is only a fragment of evidence. Request the date and machine context alongside the output. Ask why any expected field is unavailable. Source: NVIDIA System Management Interface documentation

Review diagnostic coverage rather than accepting “tested” as a complete condition statement. NVIDIA explicitly distinguishes DCGM checks from comprehensive hardware diagnostics. A passing result does not establish future reliability or a transferable warranty. Cardinal reviews seller-supplied documents; it does not claim to have tested the hardware. Source: DCGM diagnostic scope

Evidence to requestQuestion it should answer
Exact part numbers and unit identifiersWhich configuration and physical units are offered?
Dated seller test recordsWhat was checked, on which units, under which conditions?
Included-component listAre bridges, cooling assemblies and required mounting parts included?
OEM compatibility confirmationCan the proposed server support this configuration?
Written commercial termsWhat happens if the delivered configuration differs from the agreed offer?

Set the acceptance process before shipment. Specify who receives the units, checks identifiers and reviews the seller's evidence. Agree how a mismatch will be reported and resolved. Keep the promised condition distinct from the manufacturer's specification.

Questions H100 buyers ask

Is H100 SXM always the better purchase?

No. It may suit a workload and system design that benefits from its capabilities, but the decision includes the required server platform. Compare a deployable configuration with another deployable configuration, including integration work and application evidence.

Can I put an H100 SXM module in a PCIe slot?

No. The SXM module requires its compatible baseboard and system. A PCIe H100 is a different physical configuration. Use the exact variant reference when requesting either one.

Does H100 NVL mean two cards are included?

Do not assume so. NVL is commonly discussed as a pair, but an offer may count individual cards. Require the card quantity and bridge quantity in writing. Each card has 94 GB.

Should I choose NVL because it has more memory than SXM?

Only after checking the full workload and server requirements. Capacity per card is one constraint. The distribution strategy, cooling, topology and software can change which configuration is appropriate.

What should I send with a quote request?

Send the exact variant, card or server quantity, destination, desired condition and required date. Add the server part number and a short workload description. If the variant is undecided, say which platform and memory constraints are already known.

Compare the H100 variant references, or use the used GPU buying guide to prepare an offer checklist.