As AI workloads grow exponentially, high-bandwidth, low-latency, and energy-efficient networking solutions have become central to modern data center architectures. NVIDIA ConnectX-8 SuperNIC as the world's first 800G high-speed SmartNIC, redefines how servers, GPUs, and broader network fabrics connect.This article will explore the core technical breakthroughs of ConnectX-8 SuperNIC, examine how it becomes a key acceleration engine for AI factories and AI cluster network, and analyze its profound impact on server architecture and AI computing deployment.
ConnectX-8 SuperNIC is NVIDIA’s 7th-gen smart network interface card, designed for next-generation AI clusters, ultra-large-scale data centers, and high-performance computing (HPC) scenarios. It deeply integrates network acceleration with compute offload capabilities, supports ultra-high-speed 400GbE/800GbE connectivity, and significantly lowers network latency while boosting throughput efficiency through hardware-level protocol offloading and GPU-NIC co-optimization. This enables ultra-low latency and lossless network transmission for scenarios such as AI training, inference, and distributed storage.

Maximum total bandwidth: 800Gb/s
Host interface: PCIe Gen6-up to 48 lanes
Network interface: Supports flexible configurations of single-port 800G OSFP112 or dual-port 400G QSF112. Backward compatible with 200G/100G rates, adapting to existing infrastructure.
Portfolio: A variety of hardware packaging forms to adapt to different server/motherboard specifications.
PCle HHHL 1P x OSFP: PAM4 200G/Lane, half-height, half-length card with one OSFP port, suitable for high-density chassis environments.

PCle HHHL 2P x QSFP112: PAM4 100G/Lane, half-height, half-length card with two QSFP112 ports, suitable for standard rack server deployment.

Dual Connect X-8 Mezzanine: PAM4 200G/Lane, dual network card module, suitable for custom server motherboards or OCP platforms.


Through telemetry-based congestion control and intelligent routing, it maximizes network and AI workload efficiency. The InfiniBand version also supports the SHARP protocol, enabling in-network computing to accelerate distributed training and inference in HPC scenarios. The overall architecture offers strong scalability and future adaptability, making it an ideal choice for building next-generation AI infrastructure. This series of features is also explained in detail in the official introduction released by NVIDIA, click here for more information.
The NVIDIA ConnectX-8 SuperNIC is the industry’s first SuperNIC to integrate a PCIe 6.0 switch and ultra-high-speed networking into a single device. Featuring a built-in PCIe 6.0 switch with 48 lanes of PCIe 6.0 connectivity, it overcomes the limitations of traditional designs that rely on a discrete PCIe switch.
PCIe (Peripheral Component Interconnect Express) is the most widely used peripheral interconnect protocol today. It is an interface standard for connecting high-speed components. In modern computer systems, devices such as graphics cards, storage, and network cards are connected through PCIe. It provides reliable and efficient communication speed. In the PCIe architecture, the PCIe Switch functions like a network switch. It connects multiple PCIe devices, enables data forwarding between them, and addresses communication and bandwidth allocation issues.

Architecture: Independent PCIe switches are used to connect CPUs, GPUs, NICs, and other devices. For a configuration with eight GPUs and multiple NICs, two to four independent PCIe switches are required for GPU-to-GPU and GPU-to-NIC connectivity. This architecture not only increases the number of components and overall complexity, but also increases component cost and cooling costs.
Performance: GPU-to-GPU communication spans two CPU sockets, and this path may encounter bottlenecks within the host CPU and internal sockets. Depending on inter-CPU link utilization, speeds may be limited to 25 GB/s or less. In a 2:1 GPU-to-NIC configuration, GPU-to-NIC communication has limited bandwidth and is affected by the PCIe version.
Architecture: NVIDIA ConnectX-8 is redefining the possibilities of PCIe-based systems. By integrating a built-in PCIe 6.0 switch that provides 48 lanes of PCIe 6.0 connectivity, it addresses the limitations of traditional standalone PCIe switch designs. Consolidating GPU-to-GPU and GPU-to-NIC communication into a single high-performance device eliminates the need for separate PCIe switches, reduces component count, and simplifies motherboard design—creating a more cost-effective and scalable architecture for AI infrastructure.
Performance: In NVIDIA's optimized design, the ConnectX-8 SuperNIC replaces dedicated PCIe switches, integrating PCIe 6.0 switching and 800 Gb/s networking into a single network card. This doubles the network bandwidth per GPU, helping to eliminate IO bottlenecks and speeding up data movement between GPUs, NICs, and storage. As a result, this NVIDIA RTX PRO server platform delivers up to 2x the NCCL all-to-all performance, accelerating collective communications critical for multi-GPU and multi-node workloads and improving the scalability of AI factories.


The ConnectX-8 SuperNIC provides 800Gb/s of ultra-high bandwidth and supports one Infiniband XDR(800Gb/s) interface or two 400Gb/s Ethernet interfaces. While doubling the speed (compared to the previous generation ConnectX-7), the ConnectX-8 SuperNIC also strives to improve network performance and optimize network efficiency.


Positioned as a "SuperNIC," the ConnectX-8 SuperNIC not only supports 800G bandwidth, which is double that of the ConnectX-7, but also integrates a programmable acceleration engine and rich network functions to flexibly offload host network loads.

The ConnectX-8 SuperNIC utilizes a radically new server design architecture, breaking free from the limitations of traditional PCIe switches and supporting higher-bandwidth, lower-latency direct connections. This shift places even higher demands on data transmission capabilities between servers and high-performance computing nodes. As single-server network bandwidth leaps to 800G, traditional 400G optical tranceivers are no longer sufficient to meet the growing demands of cluster interconnection. This will directly drive the rapid development of 800G and 1.6T optical transceivers.
For example, the ConnectX-8 optimized design enables NVIDIA RTX PRO servers to achieve 800GB/s of omnidirectional bandwidth, doubling the network bandwidth per GPU to 400Gb/s (based on a 2:1 GPU-to-NIC ratio). To match this performance, servers require 800G optical transceivers to directly carry full-rate transmission on a single port, or to achieve 800G bandwidth by aggregating two 400G ports. At the same time, switches also face the challenge of bandwidth upgrades. To fully utilize the performance of 800G interfaces on the server side, switches require higher-capacity uplink ports, which is the core application scenario for 1.6T optical transceivers. Using 1.6T transceivers enables larger-scale GPU cluster interconnection, optimizing overall system power consumption and network topology complexity.
As a solution expert focused on AI networking, Dilight offers comprehensive 800G and 1.6T optical transceiver that support Ethernet and InfiniBand networks. These solutions feature ultra-low BER, high reliability, and broad compatibility. Dilight's latest generation of 1.6T optical transceivers have already been delivered and deployed worldwide. ( See the Dilight 1.6T transceiver test reports: Compatibility with Quantum-X800 Switch; Interoperability with NVIDIA MMS4A00 ) Currently, 1.6T optical transceivers and 1.6T DAC cables are primarily used for interconnects between QuantumX-800 switches and connection between QuantumX-800 switches and 800G ConnectX-8 NICs.

Whether it's current high-speed GPU-to-NIC interconnects or the ultra-high-bandwidth deployments of next-generation AI data centers, Dilight's high-performance interconnect solutions provide you with solid support. Feel free to connect with Dilight experts and explore what’s possible in AI networking.