L40S and H100 overlap in AI workloads, but they are not substitutes for every job. L40S includes graphics and video capabilities alongside compute. H100 variants provide different memory and interconnect options for datacenter compute. Choose around the application and server requirements.

Manufacturer capabilities and original buying analysis · No universal performance ranking

Keep the H100 variant visible

A comparison labeled only “L40S versus H100” leaves the H100 side ambiguous. Standard PCIe, SXM5 and NVL configurations differ. Begin with the one your proposed server can support.

Selected specifications per accelerator
ConfigurationMemoryMemory bandwidthMaximum powerNVLink
L40S PCIe 48GB48 GB GDDR6 ECC864 GB/s350 WNo
H100 PCIe 80GB80 GB HBM2e2,000 GB/s350 WSupported pairing
H100 SXM5 80GB80 GB HBM33,350 GB/s700 WPlatform dependent topology
H100 NVL 94GB94 GB HBM33,900 GB/s400 WSupported pairing

Sources: NVIDIA's L40S specifications, H100 specifications and H100 PCIe brief. Accelerator power is not the complete server load. These figures do not describe measured application throughput.

The table should help eliminate unsuitable configurations, not create a winner by adding numbers. A required feature may matter more than memory bandwidth. A larger memory allocation may matter more than peak compute. The workload establishes which comparison is useful.

Graphics and video change the shortlist

When the application needs hardware graphics or video features, inspect those requirements before looking at AI tensor figures. L40S includes ray-tracing hardware and dedicated video encode/decode capabilities. NVIDIA's product specification also lists display outputs.

For a video pipeline, specify the codec, resolution, pixel format and concurrent streams. Then check the applicable hardware support matrix and software implementation. A statement that a GPU supports video does not establish every codec or processing path. Source: NVIDIA video encode and decode support matrix

For rendering, use a representative scene and the intended renderer. Include assets large enough to expose memory limits. Measure time to an accepted output and record quality settings. A benchmark for a different render engine may say little about your application.

For AI-only work, L40S remains a candidate when the complete workload fits and the software supports the required operations. H100 may offer capabilities that matter for a different workload. Compare the actual service rather than ranking all tasks under one label.

A workload-first shortlistStart with required application features. Validate graphics and video requirements separately from AI memory and distribution requirements. Both paths lead to a complete server and application test.Required application behaviorGraphics or video featuresCheck exact engine and codecAI memory and distributionCheck the full working setValidate the complete server and applicationcardinalsilicon.com

Scroll across the diagram to read every label

Download diagram
Original decision diagram. The branches identify checks, not automatic product recommendations

Separate local memory from multi-GPU behavior

L40S has 48 GB per card. That capacity belongs to the individual accelerator. Installing two cards gives two devices with 48 GB each, not one device with a single 96 GB local allocation.

For a hypothetical model with 20 billion parameters at two bytes per parameter, idealized weights occupy 40 decimal GB. That leaves limited stated capacity on a 48 GB device for cache and runtime allocations. It is a reason to calculate further, not a fit guarantee.

Measure the actual model artifact under the desired context and concurrency. Record the cache format separately from the weight format. Include model-loading peaks and any framework reservations. The GPU specification guide explains the main memory components.

No NVLink does not mean no multi-GPU software

NVIDIA explicitly lists no NVLink support for L40S. That tells you which hardware link is absent. It does not, by itself, establish whether a multi-GPU application is possible or whether it will meet a performance target.

Ask how the intended application communicates between devices on the proposed server. Inspect the PCIe topology and host placement. Then benchmark the actual arrangement. A workload consisting of independent replicas may have different communication needs from one model distributed across devices.

For H100, confirm which links and bridges are present in the offer. The family name does not document the installed topology. The H100 variant guide covers the distinction between standard PCIe, SXM and NVL.

A tensor number is not an application result

Before comparing compute figures, keep precision and sparsity visible. A sparse tensor figure should not be compared with a dense result as if both described the same operation. Keep ordinary floating-point execution separate from tensor operations.

Then ask whether the application uses the relevant path. Support for a numerical format in hardware is not proof that the selected model, framework and kernel use it efficiently. It is also not evidence that the model quality remains acceptable after a precision change.

For inference, define the model revision and request distribution. Compare successful throughput while meeting a latency target. Record input and output lengths. A result with very short prompts should not be presented as evidence for a long-context service.

For a mixed graphics and AI service, test the mixture you expect to operate. Separate feature validation from contention testing. A card may support both tasks, while simultaneous execution still changes their completion times. That operating pattern belongs in the evaluation.

For training, compare the same useful outcome. Record the optimizer, precision and effective batch. If one proposal depends on different software settings, disclose them. Do not infer training value from inference results or from the memory capacity alone.

Use a benchmark worksheet

RecordWhy it belongs in the comparison
Application and model revisionEstablishes which work was executed
Numerical format and quality targetPrevents a precision change from hiding a different result
Batch, context and concurrencyDefines memory pressure and the request pattern
Server topology and power settingIdentifies the tested hardware arrangement
Latency and completed throughputConnects performance with the service objective

Keep the raw result and command or configuration file. Ask for an explanation of failures and retries. A selected chart without its test definition is weaker evidence than a reproducible run that answers your actual buying question.

The same power rating does not establish the same installation

L40S and standard H100 PCIe both have a published 350 W maximum in the comparison table. That does not make them mechanically or operationally interchangeable. The server OEM should confirm the exact card, riser, cable and cooling configuration.

NVIDIA lists L40S as a passive dual-slot card with a 16-pin power connector. Passive cooling relies on system airflow. A conventional desktop enclosure should not be treated as qualified simply because the card can be physically inserted.

Check clearance around the card and any adjacent components. Ask whether the quoted mounting hardware fits the chassis. For a larger server, confirm how populated slots affect the supported configuration. Use the OEM documentation for the complete server's power and thermal requirements.

For H100 SXM, the installation category changes entirely. A bare SXM module needs its qualified baseboard and thermal assembly. Compare a complete SXM proposal with a complete PCIe proposal, rather than comparing module prices without the required platform.

Validate the deployment image and any required license

Record the operating system, driver, CUDA runtime and application version. Add a container digest where applicable. This creates a target that can be reproduced during evaluation and after delivery.

NVIDIA's compatibility documentation describes minimum driver requirements and feature limitations. A broad claim of CUDA support is not sufficient to validate a specific application image. Review the version combination and any custom extensions. Source: CUDA minor-version compatibility

For virtualized graphics or compute, identify the required software edition and entitlement. Do not infer a transferable license from a used card's hardware capabilities. Request the seller's written description of what is included, then verify the entitlement through the appropriate channel.

L40S does not support MIG in NVIDIA's specification. If your operating plan depends on MIG profiles, that is a concrete requirement to resolve before shortlisting the card. Do not confuse hardware partitioning with every other way software may schedule work.

Compare used offers on matching scope

Ask for exact part numbers and photographs of the offered units. Confirm the memory configuration and included accessories. If a listing uses a stock image, request unit-specific evidence before relying on its apparent condition.

Ask the seller for dated test records and the host configuration used. Match the unit identifiers to the offered inventory. Request a clear description of what the seller checked, rather than accepting “tested” as a complete answer.

A diagnostic pass has a defined scope. NVIDIA's DCGM documentation does not describe its checks as comprehensive hardware diagnosis. Use the seller's records alongside the written condition statement and agreed acceptance process. Source: DCGM diagnostic limits

Compare the complete deployed scope: accelerator hardware, required server components and explicitly included services. Keep shipping, taxes and currency treatment visible. No current price ratio between L40S and H100 is assumed in this guide.

Cardinal reviews seller-supplied documents. It does not claim to test the offered cards. Include your application and acceptance requirements in the sourcing request so that a technically incomplete offer can be identified early.

Questions buyers ask about L40S and H100

Is L40S only for graphics?

No. NVIDIA describes compute capabilities alongside graphics and video features. Evaluate it for an AI workload when the memory, software and measured performance meet the requirement.

Can two L40S cards share one 96 GB memory allocation?

They provide 96 GB in aggregate across two devices. That is not one 96 GB local memory device. The application needs a suitable distribution strategy.

Does L40S have NVLink or MIG?

NVIDIA lists neither feature as supported. If your proposed deployment requires either, resolve that mismatch before purchasing.

Is H100 always faster for AI?

A universal answer would ignore the application and tested configuration. Compare exact variants under the same workload and service target. Theoretical specifications are useful inputs, not a complete result.

What should I send for a useful quote?

Specify the exact accelerator, quantity, destination and desired condition. Include the server part number and the required graphics, video or AI behavior. State any software entitlement that must be included.

Review the L40S reference and H100 variants to prepare a configuration-specific request.