Contact Us
Build Your DGX B300 SuperPOD with Dilight Deployment Support and Products
As LLM parameter sizes, context lengths, and inference concurrency continue to expand, AI infrastructure is shifting from merely boosting GPU compute to enhancing overall system efficiency. Powered by the Blackwell Ultra architecture, the B300 delivers higher compute performance and larger HBM3e capacity—establishing a stronger foundation for large-scale AI training, inference, and AI Factories.
The B300 is equipped with 288GB HBM3e memory and 8TB/s memory bandwidth, with FP4 Tensor Core performance up to 15 PFLOPS. Compared with the B200, the B300 features increased memory capacity, higher FP4 performance and elevated TDP, enabling larger models to run with fewer GPUs.
NVIDIA DGX
B200
FP4 PERFORMANCE
9PFLOPS
HBM3e
192GB
STANDARD NIC
ConnectX-7 400Gb/s
TECHNICAL COMPATIBILITY
Base / MoE Models (≤200B)Context Length ≤ 128K Tokens
BEST FOR WORKLOADS
Large-scale Cluster Training
NVIDIA DGX
B300
FP4 PERFORMANCE
15PFLOPS+67%
HBM3e
288GB+50%
STANDARD NIC
ConnectX-8 800Gb/s
TECHNICAL COMPATIBILITY
Trillion-Parameter / MoE ModelsContext Length ≥ 256K / Million Tokens
BEST FOR WORKLOADS
Ultra-High Concurrency Inference & Extreme Training

Data chart courtesy of SemiAnalysis
Based on the latest SemiAnalysis InferenceX™ benchmark for DeepSeek V4 Pro 1.6T (FP4, Agentic using vLLM), B300 demonstrates a fundamental shift in LLM serving performance. At low latency regimes (P90 E2E Latency < 60s), B300 delivers up to 2x higher token throughput, proving that its 15 PFLOPS FP4 compute effectively resolves the compute-bound bottleneck during high-concurrency Prefill/Decode stages. As batch size and latency limits expand, B200 hits a plateau at ~60k tok/s/chip due to HBM capacity bounds, whereas B300’s 288GB HBM3e seamlessly handles ultra-large KV-Cache retention, unlocking an 83% higher throughput ceiling at ~110k tok/s/chip.
This trajectory confirms B300’s superior compute-to-memory architecture for complex, multi-turn Agentic interactions.
The B300 GPU is not a standalone compute system. Within the DGX B300, eight B300 GPUs are interconnected via 5th-Gen NVLink and NVSwitch to form an ultra-high-bandwidth Scale-Up network inside the chassis. Simultaneously, each node integrates eight ConnectX-8 SuperNICs over InfiniBand or Ethernet for inter-node Scale-Out communication. As the primary compute unit bridging single GPUs to SuperPOD clusters, the value of the DGX B300 extends beyond raw compute power—it lies in seamlessly unifying GPUs, SuperNICs, switching fabrics, and physical interconnects into a low-latency, highly scalable, and integrated node architecture.

Image courtesy of NVIDIA

Image courtesy of NVIDIA
To support AI workloads exceeding single-node limits, multiple DGX B300 systems are aggregated via high-speed compute networks into standardized Scalable Units (SU). This modular compute unit integrates compute nodes, network switching and management infrastructure, supporting up to 72 DGX B300 systems for XDR or 64 for Spectrum‑X, with both high-density and low-density deployment options available.
Four air-cooled DGX B300 systems per 48U MGX rack, equipped with Active Rear Door Heat Exchangers.
Two traditional air-cooled DGX B300 systems per 48U/52U rack.
At the data center scale, multiple SUs are interconnected via a two-tier, non-blocking Scale-Out switching fabric to form a complete DGX SuperPOD. The underlying architecture seamlessly integrates high-speed intra-node GPU interconnects (powered by NVLink/NVSwitch) with robust inter-node Scale-Out networking (enabled by ConnectX-8 and high-speed switches). This layered architecture ensures ultra-low latency, ample non-blocking bandwidth, and minimal additional communication overhead introduced by increased network hierarchy, even as the cluster scales up to tens of thousands of GPUs.
When multiple DGX B300 systems are combined into a SuperPOD, the network is no longer merely an infrastructure connecting servers, but the core communication backbone of the entire AI cluster. Based on different workload responsibilities, a SuperPOD typically requires structured traffic planning across compute, storage, and management planes—ensuring high-performance GPU communication, data access, and system management operate stably on isolated network planes.
The compute fabric is the primary backbone for GPU clusters, handling collective communication, parameter synchronization, and inter-node data exchange. Logically, intra-rail nodes connect to the same Leaf switch for single-hop forwarding, while cross-rail or cross-SU traffic is routed through the Spine layer. Powered by ConnectX-8 (supporting both InfiniBand and Ethernet), the B300 compute fabric delivers full line-rate, ultra-low latency, and non-blocking forwarding to meet the high-concurrency demands of AI workloads.

B300 SuperPOD supports InfiniBand or Ethernet as its scale-out compute network. Both architectures work with ConnectX-8 and enable large-scale GPU cluster expansion via high-bandwidth networks. However, they differ in networking technology, switch ecosystem, and the integration approach for data center infrastructure. Selection should be based on GPU cluster scale, workload types, network operation & maintenance framework, and existing data center facilities.

InfiniBand is a dedicated network architecture for HPC and AI clusters, delivering mature end-to-end high-performance fabric and a robust AI/HPC networking ecosystem.
Single-plane networking via ConnectX-8 adapters and Q3400 IB switches: ConnectX-8 supports 800G single-port or 400G dual-port; Q3400 provides 144×800G ports.
Deployed in two-tier Spine-Leaf, this solution scales up to 10,368 GPUs. Clusters beyond 10,368 GPUs can expand using dual-plane networking or three-tier architecture.

RoCE operates on standard Ethernet infrastructure, suited for deployments integrating AI networks with existing data center Ethernet.
In Ethernet mode, ConnectX-8 adapters do not support single-port 800G, maxing out at dual-port 400G. High-performance Ethernet enables scale-out networks with high scalability, favorable TCO and open ecosystem.
High-density 51.2Tbps switches can be adopted, offering 64×800GbE or 128×400GbE ports. Paired with Rail-Optimized networking and multi-plane architecture, it delivers high-performance networking for large-scale GPU clusters.
The Rail-Optimized topology works as follows: within each system, GPU 1 pairs with NIC 1 and connects to Switch 1, GPU 2 pairs with NIC 2 and connects to Switch 2, and so on. This design enables single-hop communication between GPUs with the same index across different systems in the cluster, ultimately building a non-blocking all-to-all network fabric across all servers in a large-scale AI network.

The core feature of the B300 solution is a dual-plane Leaf-Spine design (Plane A and Plane B). Based on a Rail-Optimized network, GPUs are arranged along rails, as shown below: SU1 (GPU-001 to 064) and SU16 (GPU-961 to 1024). Each GPU connects to two independent planes via two NVIDIA ConnectX-8 SuperNICs, with each link running at 400GbE, delivering a total of 2×400Gbps per GPU.
The two planes are physically isolated, with a small number of escape paths established between spine-layer switches. The multi-plane design not only delivers high performance and low latency for AI training but also significantly improves fault tolerance: when a single switch, optical module, or cable fails, jobs can continue running on independent paths at half of the original bandwidth. Moreover, because the two planes are independent at the core layer, only two layers of Leaf-Spine are needed to connect twice as many nodes, lowering build costs while maintaining high performance. The architecture also automates traffic optimization through the plane load balancer and software built into the ConnectX-8 NICs, while providing real-time network telemetry and monitoring to ensure the efficient and reliable operation of AI workloads.

Each plane is a physically independent network. A device failure or link outage in Plane A never impacts Plane B, eliminating the risk of a single point of failure bringing down the entire network.
Traffic automatically fails over to the healthy plane, so AI training tasks are virtually unaffected by network faults — significantly improving overall cluster availability.
Shuffle is a structured optical connectivity scheme for B300 high-density Multi-Plane and Rail-Optimized designs. It pre-defines cross-connection mapping in the factory, routing different lanes of a high-speed port to distinct planes—e.g., Shuffle cables, Shuffle box. By hardening network topology into physical interconnects before deployment, it minimizes on-site patching and precisely resolves port breakout, rail alignment, and multi-plane optical organization, and drastically reduces manual on-site cabling work.

The evolution of AI infrastructure relies fundamentally on the deep synergy between compute density and network topology. By integrating modular SUs, physically isolated storage fabrics, and Rail-Optimized structured cabling, the B300 architecture achieves linear scaling of compute and bandwidth while maintaining strict topological and engineering consistency—establishing a deterministic framework for hyper-scale AI clusters.
From a single node to a multi-SU SuperPOD, unlocking peak system performance depends not only on raw compute density, but also on the precise implementation of physical interconnects. Under an 800G high-density and Rail-Optimized architecture, structured optical shuffle systems, high-reliability optical transceivers, and low-latency cabling assemblies serve as the essential physical foundation—ensuring non-blocking throughput and zero-loss performance for next-generation AI infrastructure.
| InfiniBand | RoCE | |
|---|---|---|
| Spine & Leaf Switch | ||
| Transceiver (Switch-side) | ||
| Transceiver (NIC-side) | ||
| DAC/AOC/AEC | ||
| Fiber Cable |
Dilight delivers full-stack turnkey construction capabilities for SuperPOD clusters, independently managing end-to-end delivery spanning network architecture design, topology optimization, equipment deployment, high-speed interconnection cabling, optical system integration and system commissioning.

AI infrastructure centers on tangible business value. With high-performance scale-out networking and professional managed services, Dilight delivers industry-specific AI solutions to accelerate industrial intelligence.
Build Your DGX B300 SuperPOD with Dilight Deployment Support and Products