In the ICT(Information and Communication Technology) domain, the concept of "Scale" is frequently encountered. It typically refers to enhancing the capabilities of an existing system to handle a greater workload. There are two primary scaling methods: Scale-Up (Vertical Scaling) and Scale-Out (Horizontal Scaling).
Scale-Up (Vertical Scaling): This involves boosting the performance of a single system by adding more resources, such as increasing processor speed, memory, or storage capacity. Essentially, it makes one system more powerful.
Scale-Out (Horizontal Scaling): This entails adding more systems with the same or similar configurations to distribute the workload. It means scaling by increasing the number of independent systems working in parallel. Using a three-tier/two-tier Clos architecture, tens of thousands or even hundreds of thousands of GPUs can be connected to a unified cluster to support parallel training scenarios like DP/PP. Its core goal is to prioritize scale and manage costs.

Example in Networking:
In many cases, Scale-Up and Scale-Out can be combined to build larger and more efficient networks.

In today’s AI Computing Networks, both Scale-Up and Scale-Out networks coexist:
The synergy between Scale-Up and Scale-Out networks is what powers today’s AIGC large models.
While both aim to enable memory-level data transfer between GPUs, their design purposes and application scenarios differ significantly. With the rise of large AI models, the scale continues to grow. A single GPU server is no longer sufficient. Parallel computing becomes necessary, which introduces communication overhead, complexity in partitioning, and programming challenges. Transformers, with their attention mechanisms and feedforward layers, place high demands on memory and compute resources.
Ideally, one would have a super GPU chip capable of processing an entire large model alone. Since this is unfeasible, the model is split:
While RDMA simulates memory access to some extent, it is not ideal for frequent small memory read/writes and is not a true memory-semantic network. This dual-network architecture balances performance and cost: Scale-Up focuses on ultimate performance and Scale-Out prioritizes flexibility and affordability.
In large-scale model training, both network types support GPU-to-GPU data transfers, but they differ greatly in latency.
Network latency is the time data takes to travel through a network. It comprises:
This is a bus-domain network enabling direct GPU memory access. Since modern GPUs can run at over 1GHz (less than 1ns per cycle), ultra-low latency is critical. Scale-Up networks must achieve sub-microsecond latency or even lower. To achieve this, Scale-up network designs must be tightly coupled with specific application requirements, eliminating traditional transport and network layers while employing credit-based flow control and link-layer retransmission to ensure reliability. Simultaneously, addressing challenges from high-speed SerDes technologies—such as PAM4 signaling and 112Gbps/224Gbps implementations with DSP architectures—makes deterministic latency control critically important. Existing RS(544, 514) forward error correction (FEC) schemes may become inadequate at these speeds, necessitating exploration of new FEC approaches to further reduce latency.
In contrast, Scale-Out networks are inherently more flexible and diverse. They draw inspiration from traditional layered network architectures such as the OSI model, which allows them to support a wide variety of communication and data transmission needs. While this flexibility does come with compromises in latency performance, it also ensures the network can adapt to a broader range of application scenarios.
In a Scale-Out network, end-to-end latency is typically maintained between 1 to 10 milliseconds, ensuring users perceive the system as responsive. For compute-intensive tasks in AI and HPC environments, although ultra-low latency isn’t mandatory, stable low latency is still a key factor in achieving high performance. Scale-Out networks leverage the existing industrial ecosystem—including switches and optical modules—and apply performance enhancements such as UEC (Ultra Ethernet Consortium) and GSE (Generic Streaming Extensions) to reduce dynamic latency. However, due to the inherent complexity of the architecture, static latency remains relatively high.
The latency objectives of Scale-Up and Scale-Out networks diverge significantly. Scale-Up networks are engineered to drive round-trip times (RTT) from sub-millisecond down to sub-microsecond levels, emphasizing extreme low-latency performance. Scale-Out networks prioritize flexibility and cost efficiency, delivering millisecond-level latency suitable for a broad range of workload types. This distinction in latency performance is precisely what defines their different roles in AI, HPC, and other demanding computing environments.
To address the challenges of rapidly growing AI workloads, modern scale-out networks must support massive node expansion and unpredictable traffic bursts. Designed with these demands in mind, the Dilight 51.2T switch N9570-128QC delivers exceptional performance and scalability. With 128 high-density 400GE QSFP112 ports, it supports the interconnection of up to 128 GPU servers and can easily scale to network architectures with tens of thousands of nodes. Powered by the NVIDIA Spectrum-4 chip, this 400G AIDC switch uses Priority Flow Control (PFC) and Explicit Congestion Notification (ECN) to effectively manage congestion and minimize dynamic latency. Additionally, it supports Adaptive Routing (AR), enabling intelligent packet-level load balancing across multi-path networks. It resolves the issue of uneven static load distribution inherent in ECMP and effectively improves link utilization.

Merging Scale-Up and Scale-Out networks is impractical due to fundamental differences in their design, objectives, and implementation.
Scale-out networks, rooted in traditional data centers, connect geographically dispersed nodes for efficient, long-distance communication. They excel at remote transmissions, heterogeneous device interconnections, and diverse business communications. Scale-Up networks are a newer paradigm that enhances system capabilities by boosting single-device performance. These tightly integrated networks consolidate resources within limited physical space for significant performance gains, and they are deeply coupled with business logic.
In the AI and AGI era, intelligent computing network demands are rising. Simply enhancing traditional data center network's load-store capabilities or attempting to expand networks using load-store technology won't meet scale-up network needs. Their initial design premises differ, leading to significant distinctions in technical realization, performance, and cost-effectiveness.
From a business logic perspective, Scale-Up networks (like NVLink) align with load-store semantics, emphasizing direct, high-speed memory access. In contrast, Scale-Out networks (like InfiniBand) are based on message semantics, prioritizing flexibility and scalability. While some technical specifications might appear similar, this is merely a coincidence and doesn't indicate their potential for convergence or interchangeability.
Therefore, Scale-Out and Scale-Up networks should not be combined due to their inherent differences in technical philosophy, application goals, and business logic. Each plays an indispensable role in its specific domain, collectively advancing computing network technology.
In March 2024, NVIDIA introduced the GB200 NVL72 SuperNode, integrating 36 Grace CPUs and 72 Blackwell GPUs into a single liquid-cooled cabinet. This system delivers up to 720 PFLOPs of AI training performance or 1440 PFLOPs for inference. Its architecture not only overcomes the inter-node bandwidth bottleneck seen in earlier generations like H100/GH200, but also combines "GPU-GPU NVLink Scale-Up" and "Node-to-Node RDMA Scale-Out" to provide scalable infrastructure for exabyte-scale data processing and trillion-parameter model training.

Within the SuperNode cabinet, 72 B200 GPUs housed across 18 Compute Trays are fully interconnected via NVLink 5 and copper cables, linking to 18 NVSwitch chips distributed across 9 Switch Trays.
Theoretical Bandwidth:

Physical Cabling:
Cabling Medium:
The system uses a cable cartridge approach (copper-based interconnect). In short-range transmission scenarios, copper offers higher reliability and lower cost compared to optical modules, and simplifies cabling. As such, direct copper connections have become a mainstream solution for Scale-Up interconnects.

In short, NVL72 creates a massive Scale-Up network that enables high-bandwidth, low-latency communication between GPUs.
Scale-Out allows the integration of eight DGX GB200 NVL72 units into a SuperPOD with 576 B200 GPUs. Each Compute Tray equips each of its four GPUs with a CX8 800Gbps RNIC (RDMA NIC), connecting them to the InfiniBand RDMA-based Scale-Out network.

NVL72’s Scale-Up interconnect offers:
On the Scale-Up side, the industry was traditionally dominated by Chassis SuperNodes, such as NVIDIA NVL72 and Huawei CloudMatrix384. These architectures emphasize reliability and build a stable interconnection domain through a "Cable Backplane + L1/L2 Switch" design. Electrical interconnects are used within the chassis to achieve high bandwidth, while optical interconnects are used between cabinets.
However, the emergence of Box SuperNodes after 2025 disrupted this pattern. By packaging multiple GPUs into independent boxes and directly interconnecting them through a first-level Clos switch, Box SuperNodes significantly reduce the number of switching layers and lower interconnect costs. At the same time, this architecture places higher requirements on first-hop optical connectivity and system reliability.
On the Scale-Out side, the dominant trend is simplification. Traditional three-layer Clos architectures are gradually being replaced by two-layer Clos networks through a flatter design approach. Advances in switch silicon have enabled Radix=512, allowing two-layer Clos networks to directly support over 130,000 GPUs without requiring a dedicated core layer. At the same time, multi-plane networking is gaining traction. DeepSeek proposed a "Mutli-plane architecture" in its ISCA paper: the AI-NIC accesses four Clos planes through four 200G ports, and the data packets of each flow are evenly distributed to different planes through Round-Robin. The receiving end uses DDP out-of-order write technology to reorganize the data, increasing the single-GPU Scale-Out bandwidth utilization to more than 95%.
Together, these trends show that Scale-Up architectures are becoming more flexible, while Scale-Out networks are becoming flatter and more scalable.
Scale-Up pursues ultra-low latency and low power consumption, making the LPO/NPO transceiver that removes the DSP chip the preferred solution. 800G LPO optical transceivers consume only 6W, 60% less than 15W optical transceivers with DSPs. By eliminating the DSP's signal processing, they reduce one-way latency by 60ns and round-trip latency by 240ns. For example, Dilight's 800G LPO transceivers are optimized for AI data centers and meet the low power and latency requirements of Scale-Up architectures.
Scale-Out's demand for cross-vendor interoperability and long-distance transmission makes DSP-based optical transceivers difficult to replace.
BER Advantage: DPO transceivers with DSP support complex FEC (forward error correction) algorithms, achieving a corrected BER of up to 1E-15. However, the BER of LPO is only 1E-12~1E-13, which cannot meet the reliability requirements of Scale-Out long-distance (10km+) transmission. Whether it is DPO or LPO modules, when deployed in real production environments, reliability depends not only on lab-tested BER performance but also on system-level stability and long-term operational robustness, which is why “passing tests” does not necessarily mean the infrastructure is production-ready.
Cross-vendor interoperability is a must: Scale-Out clusters often use equipment from multiple vendors. LPO/NPO interoperability has yet to reach a unified standard. However, the DPO transceivers, based on the IEEE 802.3 protocol, enables seamless interoperability.
At this point, the technical divisions of optical transceivers are clear: Scale-Up chooses LPO/NPO, while Scale-Out chooses DPO/LRO. While these two technologies differ, they both share the core goal of achieving "224G optical speeds." In line with this trend, the industry has already begun developing related products. For example, Dilight's OSFP-1.6T-2xDR4 and OSFP-1.6T-2xFR4 support the high bandwidth and low latency requirements of future supernode architectures, laying the foundation for data centers to move toward the next stage of Scale-Up.
Large-scale AI models continue to grow, seemingly without limit, and place unprecedented demands on computing infrastructure. Building ultra-powerful supernodes using a Scale-Up strategy, and then extending them across clusters through Scale-Out, has become a common and efficient practice. This layered approach not only meets the performance and scalability needs of modern AI, but also lays a practical foundation for future innovation.
As a dedicated provider of AI networking solutions, Dilight leverages deep technical expertise and premium-quality products (including optical transceivers, AOC/DAC cables, switches and NICs, etc...) to help customers build cadvanced AI computing infrastructure. From scalable network architecture to precise optical interconnects, we deliver the demanding performance, reliability, and scalability required in AI era. Reach out to explore how we can support your intelligent computing networks.