Buying Intel Gaudi 3 means selecting a hardware platform and validating a software path. Distinguish the HL-325L OAM module from the HL-338 PCIe card. Then review the server, Ethernet arrangement and application evidence as one proposed deployment.
Two configurations need separate references
Intel's product overview separates the HL-325L OAM accelerator from the HL-338 add-in card. HLB-325 identifies a baseboard, not an individual accelerator. Preserve those distinctions when reading an offer or preparing a request.
| Property | HL-325L OAM 128GB | HL-338 PCIe 128GB |
|---|---|---|
| Memory | 128 GB HBM2e | 128 GB HBM2e |
| Memory bandwidth | 3.7 TB/s | Not independently verified for this variant |
| Accelerator power | 900 W | 600 W |
| Integrated networking | 24 × 200 GbE ports | 18 × 200 GbE ports |
| Physical form | OAM 2.0 module | Dual-slot, full-height add-in card |
The Intel Gaudi 3 product overview provides these distinctions. Its dedicated HL-338 slide does not give a memory bandwidth figure. This guide leaves that field unverified instead of copying the OAM value into the PCIe record.
Ask the seller for the full part number and revision. Intel's software support matrix also uses an HL-338A designation for a listed platform component. Do not silently treat every similar label as identical. Match the actual offered hardware to the appropriate documentation.
Specify whether the offer is a module, a populated baseboard or a complete server. A baseboard photograph may include components that are not in the commercial line item. Require the accelerator count and included-component list in writing.
Integrated Ethernet still needs a topology
The integrated port count is an architectural capability, not a count of ready-to-use external server sockets. The server or baseboard determines how connections are allocated and exposed. Ask for its wiring and network diagram before specifying switches or cables.
Keep three parts of the installation distinct: communication among accelerators, communication between servers and the management network. They have different roles. A successful management login does not establish that the accelerator fabric is configured or performing correctly.
Scroll across the diagram to read every label
Download diagramFor a complete proposal, ask which external fabric ports are included and which transceivers or cables are required. Record the switch model and configuration assumptions. Keep optional networking items separate from included hardware so that the proposal can be priced and installed consistently.
Intel's readiness documentation distinguishes operating-system access, BMC access and optional accelerator-fabric connections for its described systems. It directs buyers to system-vendor power and thermal requirements. Use the exact server documentation rather than applying one reference layout to every Gaudi configuration. Source: hardware and network readiness
If the application will span nodes, include a multi-node demonstration in the evaluation. Test the intended message pattern and workload scale. A single-server result should not be multiplied by the server count to predict cluster throughput.
A software plan belongs before the purchase
Inventory the application first. Record the framework, model revision and custom operations. Identify any CUDA-specific extensions or third-party libraries. Establish which pieces have a supported Gaudi path and which require engineering work.
Intel publishes a support matrix that ties together software versions, operating systems and related components. Use that matrix for the proposed release. A supported framework name alone does not establish every feature or package combination. Source: Intel Gaudi support matrix
Record the driver, firmware and container image alongside the application. Ask the supplier to reproduce the proposed stack, rather than demonstrate an unrelated sample image. Preserve the configuration so your team can compare the delivered system with the evaluated system.
Migration assistance is not application validation
Intel's GPU Migration Toolkit provides a documented route for adapting supported PyTorch code. Its documentation includes a support matrix and limitations. Treat it as tooling within a migration project, not proof that an arbitrary CUDA application will work unchanged. Source: GPU Migration Toolkit
Split evaluation into correctness and operating performance. First, confirm that outputs meet your quality requirements. Include difficult inputs and the intended numerical format. Then measure whether the application meets throughput, latency and stability targets.
Assign an owner for unsupported operations before accepting a hardware proposal. Record the expected work and how it will be tested. If the project depends on a future software change, make that dependency visible instead of assuming it will be available at deployment.
Review the maintenance path as well as the initial port. Ask how the team will update frameworks, monitor failures and rebuild custom components. A successful demonstration is valuable, but a repeatable operating procedure makes it useful beyond the demonstration.
Benchmark a useful outcome
For inference, specify the model artifact, precision and request distribution. Set a latency target and measure completed throughput while meeting it. Include the longest supported context and the desired concurrency. Count failures rather than removing them from the result.
For training, define the effective batch, optimizer and quality target. Record changes to checkpointing or distributed execution. Compare useful progress under those settings, not only a peak hardware figure.
Measure warmup separately from steady execution when it affects your workload. A long-running service and a short-lived batch task may assign different importance to startup behavior. State which phase a reported number covers.
Keep comparisons between architectures explicit about software differences. Each system may use its own optimized stack, but the task and acceptance criteria should remain comparable. Document any model or precision changes that alter the work.
Use the memory budget to eliminate impossible configurations, then measure the remaining candidates. The 128 GB capacity does not by itself establish that a given model and serving configuration fit. Weights, cache and runtime allocations share the device's memory budget.
Do not turn port counts into throughput claims
A port rate is not an application benchmark. It does not include every effect of topology, communication scheduling or congestion. Request measurements on the proposed network rather than a multiplication of the integrated interface specifications.
Keep units visible. A 200 GbE port uses gigabits in its name, while memory bandwidth is commonly reported in gigabytes per second. Those numbers should not be placed beside each other without explaining the units and measurement boundary.
Confirm that the complete platform is ready
For OAM, request the qualified baseboard and server configuration. Include thermal assemblies, host components and management access in the scope. A bare accelerator does not establish that the buyer has a usable deployment platform.
For PCIe, confirm the exact HL-338 part and revision with the server OEM. Review space, power delivery and cooling together. A dual-slot description alone is not enough to approve a chassis installation.
Use the OEM's complete system power requirements for facility planning. The table's accelerator ratings exclude other server components. Confirm the intended power configuration and any service requirements before arranging delivery.
Ask for a documented firmware and software state. Do not update an unfamiliar system simply to make version numbers look newer. Establish the supported combination and the maintenance procedure first. The same discipline applies when evaluating a seller's existing installation.
Evaluate a used Gaudi offer with unit-level records
Request part numbers, revisions and unit identifiers from the actual hardware. Ask for photographs and a complete component list. If several modules are offered, keep their identifiers in a table that can be matched to the seller's evidence.
Intel documents the hl-smi management interface for device information and monitoring. Ask the seller to include the relevant output with the software version and collection date. A cropped status image is less useful than a complete record with context. Source: hl-smi documentation
Request the seller's test plan and results for the offered units. Establish whether tests covered memory, connectivity and the intended workload. A statement that the server boots should not be interpreted as an application acceptance test.
| Offer component | Question to resolve |
|---|---|
| Accelerator identity | Which exact part numbers, revisions and units will be delivered? |
| Baseboard or server | Is the offer deployable with the buyer's proposed platform? |
| Network arrangement | Which connections and external components are included? |
| Software state | Which supported driver, firmware and application image were demonstrated? |
| Condition evidence | What did the seller check, when and on which units? |
Agree how your team will verify the delivered configuration and report discrepancies. Keep the acceptance process and any stated remedy in the commercial terms. Do not infer a warranty or remaining service life from a manufacturer's product brief.
Cardinal reviews seller-supplied documents and sources offers against the requested configuration. It does not claim to test the hardware or complete an application port. Include your validation requirements in the request so the supplier can respond to them directly.
Questions buyers ask about Intel Gaudi 3
Is HL-325L the same thing as HLB-325?
No. HL-325L identifies an OAM accelerator. HLB-325 identifies a baseboard in the checked Intel documentation. State both where the offer includes a populated baseboard.
Does HL-338 have the same specifications as the OAM module?
Do not assume that. The checked overview shows different power and networking counts. It does not provide a separate HL-338 memory bandwidth figure, so that field remains unverified here.
Can I use my existing Ethernet switches?
That requires a review of the exact system and fabric requirements. Ask for a topology and supported switch configuration. An Ethernet label alone does not establish the intended performance.
Will an existing PyTorch application run without changes?
It depends on the application and supported feature set. Inventory dependencies, follow the documented migration path and validate correctness. Then measure operating performance on the proposed configuration.
What makes a useful Gaudi quote request?
Specify HL-325L or HL-338, quantity, destination and desired condition. State whether a complete server is required. Include the application stack, networking plan and evidence needed for acceptance.
Review the Gaudi 3 variant references and the used hardware buying guide before comparing offers.