H200 is worth evaluating when memory capacity or memory movement constrains an H100 deployment. The useful comparison is an exact H200 configuration against an exact H100 configuration, running the same workload. A larger memory figure alone does not determine the better purchase.
Compare SXM with SXM, then compare NVL separately
H100 and H200 both include SXM and PCIe-based configurations in the reference catalog. Avoid a comparison that silently changes the form factor. If your starting point is an existing PCIe server, the SXM specification is not the relevant installation target.
| Configuration | Memory capacity | Memory bandwidth | Maximum accelerator power |
|---|---|---|---|
| H100 SXM5 80GB | 80 GB HBM3 | 3.35 TB/s | 700 W |
| H200 SXM 141GB | 141 GB HBM3e | 4.8 TB/s | 700 W |
| H100 NVL 94GB | 94 GB HBM3 | 3.9 TB/s | 400 W |
| H200 NVL 141GB | 141 GB HBM3e | 4.8 TB/s | 600 W |
NVIDIA's H100 page and H200 page provide the comparison values. The H200 specification table is marked preliminary by NVIDIA. Confirm current OEM documentation for a purchase. Power limits are configurable and do not represent complete server consumption.
The standard H100 PCIe 80GB is another configuration, distinct from H100 NVL. Keep it on the shortlist when that is the hardware your server supports. The H100 variant guide explains those differences in more detail.
More memory does not mean a higher peak arithmetic rate
The current manufacturer tables list the same 67 TFLOPS of FP32 vector compute for H100 SXM and H200 SXM. Their FP8 Tensor Core figures also match at 3,958 TFLOPS, explicitly with sparsity. These examples show why H200 should not be described as a universal compute upgrade: an application can benefit from its memory system without gaining a higher published peak arithmetic rate. Source: H100 specifications; source: H200 specifications and sparsity footnote
Keep vector and Tensor Core operations in separate columns. A dense model does not automatically achieve a structured-sparse figure, and the numerical format of a checkpoint does not prove which arithmetic its kernels execute. Use the inference precision guide to qualify those assumptions. Then determine whether the measured service spends its time on arithmetic, memory movement, communication or another stage.
The memory increase can change how a workload is distributed
Using the published capacities, H200 has 61 GB more memory per accelerator than H100 SXM. That is a 76.25% increase in stated capacity. Relative to H100 NVL's 94 GB, the increase is 50%. These are arithmetic comparisons, not measured application gains.
Scroll across the diagram to read every label
Download diagramExtra capacity can be valuable when it removes a constraint in the actual deployment. It may create room for a larger model artifact, more cached tokens or a different batch configuration. It may also leave unused memory if the existing workload already fits comfortably. Measure the constraint before paying to remove it.
For a hypothetical 70-billion-parameter model, BF16 weights alone require 140 decimal GB at two bytes per parameter. H200's stated 141 GB should not be interpreted as a comfortable fit for the complete inference process. Cache and runtime allocations still need space. Even the weight artifact can differ from this simplified calculation.
A quantized artifact changes the calculation, but it also changes the validation work. Check the delivered quality at the chosen quantization settings. Record any higher-precision tensors and metadata. Do not select a GPU from a parameter-count label without inspecting the intended model format.
Cache behavior also depends on the serving engine and model architecture. Hugging Face documents multiple cache strategies, including implementations that trade memory use against other behavior. A full-attention KV estimate does not describe every sliding-window or offloaded configuration. Source: Transformers cache strategies
Fewer GPUs is a hypothesis to test
Suppose a measured service requires more memory than one H100 can provide, but appears to fit within one H200's operating budget. The H200 may allow a different distribution strategy. That possibility is useful, but the comparison must include throughput and latency, not just memory totals.
One larger GPU may still be insufficient for the desired request rate. Conversely, several smaller replicas may suit independent jobs even when a larger device has more capacity. Test the proposed arrangement rather than assuming that the smallest device count minimizes the deployed cost.
Three situations that lead to different decisions
Your current H100 service runs out of memory at long context
First isolate what grows. Record memory at model load, after a short request and at the intended long-context concurrency. Check whether the serving engine reserves a large static cache. Establish whether the memory failure comes from the desired workload or an avoidable configuration choice.
If the working configuration still exceeds the H100 budget, evaluate H200 with the same model and request mix. Require the longest supported input and the desired output length in the test. Compare the result with an H100 configuration that also meets those requirements. An H100 run that fails is not a useful denominator for a speedup claim.
Your H100 workload already fits and meets its latency target
Start by defining the improvement that would justify a change. It could be higher throughput per server, fewer servers for a fixed service level or additional workload capacity. Without that target, a larger memory specification can become an expensive feature with little operational value.
Keep the existing deployment as a measured baseline. Include migration effort and any changes to scheduling, monitoring or spare hardware. If a new configuration requires different operating procedures, account for that work explicitly rather than hiding it inside a hardware comparison.
You are buying a first cluster
Compare complete proposals with the same acceptance criteria. Ask each supplier to specify accelerator variants, server components, networking and installation requirements. Use the workload as the common reference. A proposal with more GPU memory may still differ in host memory, storage performance or support arrangements.
Preserve a record of what is included and what is optional. Ask for separate prices where the scope differs. This lets you compare the hardware and the integration work without inventing a universal market price for either GPU family.
A fair H200 versus H100 benchmark needs a written workload
Describe the model revision and numerical format, then freeze the application image for the first comparison. Record the driver and framework versions. Hold the request distribution constant. If either platform needs a different optimized image, run that as a second comparison and explain the change.
For interactive inference, collect time to first token, inter-token latency and completed throughput. Evaluate more than the average if your users care about slow requests. State whether token counts refer to input, output or both. A benchmark label without those definitions can conceal a different workload.
For batch processing, define the completion objective and dataset. Count successful outputs and record any retries. For training, compare the same effective batch and target quality. Throughput is useful only when the resulting work remains equivalent.
Keep utilization and memory observations beside the application results. If the accelerator waits on input preparation or external storage, replacing it may leave the bottleneck unchanged. Profile the system to determine whether a hardware difference actually addresses the limiting stage.
NVIDIA publishes workload-specific comparisons with stated configurations. Those results can help form a test hypothesis, but they do not establish a universal H200 multiplier. Treat the model, batch size and server arrangement as part of every benchmark claim. Source: H200 performance examples and test notes
The platform review is part of the purchase
For SXM, request the qualified server configuration rather than assuming that a module with a related product name can replace another module. The server OEM should confirm mechanical, electrical, firmware and thermal support. Get that answer for the exact part numbers.
For NVL, the published power ratings differ materially between H100 and H200. Have the OEM confirm the power delivery, cooling and bridge arrangement. A chassis that supports one NVL configuration is not automatically evidence that it supports the other.
NVIDIA's DGX documentation describes separate H100 and H200 system configurations, including 640 GB and 1,128 GB aggregate GPU memory respectively. Use such documentation to understand a complete platform. Do not treat the shared manual as a blanket field-upgrade procedure. Source: DGX H100/H200 system introduction
Keep software compatibility separate from licensing
Confirm that the intended software image supports the offered hardware and driver. NVIDIA's CUDA compatibility documentation describes restrictions as well as compatibility paths. Record the tested combination so the deployment does not depend on a vague promise of CUDA support. Source: CUDA compatibility requirements
Then ask what software entitlement is included in the transaction. A hardware specification does not prove the transferability of a subscription. Require the seller to identify any included license and its remaining term. If that evidence is unavailable, keep the license outside the assumed value of the offer.
A used H200 or H100 offer needs unit-level evidence
Ask for part numbers and identifiers before comparing condition labels. A family-level stock photograph cannot establish which accelerator will be shipped. Request dated photographs of the offered units and a component list for any complete server.
Request seller diagnostic records with their tool version, test configuration and date. Match the reported identifiers to the unit list. If the records describe a different host platform, ask how the seller established the proposed configuration's behavior.
Do not treat one passing report as comprehensive assurance. DCGM documentation explains the limits of its diagnostic scope. Use the report as one piece of evidence, together with the seller's written condition statement and the agreed acceptance process. Source: DCGM diagnostics
| Before accepting an offer | Record |
|---|---|
| Variant identity | Exact card or module configuration and OEM part number |
| Deployment scope | Bare accelerator, populated baseboard or complete server |
| Condition evidence | Seller statement and dated records linked to units |
| Operational fit | OEM confirmation and intended workload evidence |
| Acceptance | Checks, responsibilities and mismatch resolution agreed before shipment |
Cardinal checks the documents supplied by the seller. It does not claim to test the offered hardware. A sourcing request should state the evidence your team needs so that incomplete offers can be identified early.
Questions buyers ask about H200 and H100
How much more memory does H200 have?
Published capacity is 141 GB per GPU. That is 61 GB more than H100 SXM's 80 GB and 47 GB more than H100 NVL's 94 GB. The differences do not measure application performance.
Can a 70B BF16 model fit on one H200?
Do not assume it from the weight calculation. Idealized weights alone need 140 decimal GB, leaving little stated capacity for other allocations. Validate the actual model artifact and serving configuration.
Is H200 a drop-in replacement for H100?
Do not assume a field replacement is supported. Ask the server OEM to confirm the exact part-number combination and any required changes. Shared architecture or similar naming is not an installation approval.
Is H200 always faster than H100?
No. The result depends on the exact variants and application. More capacity can enable a different batch or distribution strategy, while more memory bandwidth can help a memory-constrained stage. Neither establishes a universal speedup. Compare equivalent outputs at the same latency and quality requirements.
Is the H200 premium justified?
That depends on the offered configuration and your measured constraint. Compare current offers that meet the same workload target. Include integration work and the value of any verified operational improvement.
Should I compare H200 with a used H100 system?
Yes, when both are realistic deployment options. Keep condition, included components and acceptance terms visible. A cheaper accelerator line item is not necessarily a cheaper working system.
Use the H200 references and H100 references to prepare exact variant requests. For offer documentation, continue with the used GPU buying guide.