Exploring Network Architecture Options for Reduced Latency and Enhanced Performance

Kevin Di

Hatched by Kevin Di

Mar 28, 2024

5 min read

0

Exploring Network Architecture Options for Reduced Latency and Enhanced Performance

Introduction:
In today's digital age, where data transfer and communication play a crucial role in various industries, optimizing network architecture becomes essential. This article delves into the key technologies and options available to reduce latency and improve performance in multi-machine, multi-card communication scenarios. We will explore RDMA technology, specifically the InfiniBand and RoCEv2 solutions, and discuss their advantages over traditional TCP/IP networks. Additionally, we will touch upon the components and suppliers in the InfiniBand market. Furthermore, we will discuss the RoCE solution, its versatility, and the challenges associated with its implementation. Lastly, we will highlight the significance of A100-SXM4-80G Multi-Query Attention and the improvements it brings to the ChatGLM2-6B language model.

Reducing Latency with RDMA Technology:
RDMA (Remote Direct Memory Access) technology is a key factor in reducing end-to-end communication latency between multiple machines and cards. It bypasses the operating system kernel, allowing direct access to another host's memory. There are four methods for implementing RDMA: InfiniBand, RoCEv1, RoCEv2, and iWARP. RoCEv1 is now considered obsolete, and iWARP has limited usage. Currently, the most widely adopted RDMA solutions are InfiniBand and RoCEv2. By bypassing the kernel protocol stack, these solutions offer significant latency improvements compared to traditional TCP/IP networks. In lab tests for scenarios where one hop is reachable within the same cluster, the end-to-end latency at the application layer can be reduced from 50us (TCP/IP) to 5us (RoCE) or 2us (InfiniBand) when bypassing the kernel protocol stack.

Understanding InfiniBand Network Architecture:
The InfiniBand network architecture consists of several key components, including the Subnet Manager (SM), InfiniBand network cards, InfiniBand switches, and InfiniBand cables. NVIDIA introduced the Quantum-2 series switch with a speed of 400Gbps in 2021. This switch features 32 800G OSFP ports, which can be converted to 64 400G QSFP ports using cables. InfiniBand switches do not run any routing protocols. The forwarding table for the entire network is computed and distributed by a centralized Subnet Manager (SM). Additionally, the SM manages the configuration of partitions, QoS, and other aspects of the InfiniBand subnet. Specialized cables and optical modules are required for interconnecting switches and network cards in an InfiniBand network. Adaptive Routing in InfiniBand is based on per-packet dynamic routing, ensuring optimal network utilization in large-scale deployments. InfiniBand networks have found extensive use in the industry, particularly in large GPU clusters such as Baidu Intelligent Cloud and Microsoft Azure.

Major InfiniBand Solution Providers:
Several suppliers offer InfiniBand network solutions in the market, with NVIDIA being the dominant player, occupying over 70% of the market share. Here are some key players:

  1. NVIDIA: NVIDIA is one of the major suppliers of InfiniBand technology. They provide various InfiniBand adapters, switches, and related products.

  2. Intel Corporation: Intel is another important InfiniBand supplier, offering a range of InfiniBand network products and solutions.

  3. Cisco Systems: Cisco, a renowned network equipment manufacturer, also offers InfiniBand switches and related products.

  4. Hewlett Packard Enterprise (HPE): HPE, a large IT company, provides various InfiniBand network solutions and products, including adapters, switches, and servers.

Exploring RoCE Solution and its Advantages:
The RoCE (RDMA over Converged Ethernet) solution offers greater versatility and comparatively lower prices than the InfiniBand solution. Apart from high-performance RDMA networks, RoCE can also be used in traditional Ethernet networks. However, configuring parameters such as Headroom, PFC, and ECN on switches can be complex. In large-scale deployments like those involving thousands of cards, the throughput performance of RoCE networks is slightly weaker than InfiniBand networks. Numerous switch manufacturers support RoCE, with prominent names like Huawei and H3C (New H3C) leading the market. In terms of network cards, NVIDIA's ConnectX series holds a significant market share for supporting RoCE.

Enhancing Language Models with A100-SXM4-80G and ChatGLM2-6B:
In addition to network architecture, advancements in language models also contribute to improved performance. The A100-SXM4-80G Multi-Query Attention technology plays a crucial role in reducing the memory usage of the KV Cache during the generation process. Furthermore, ChatGLM2-6B, as an open bilingual chat Language Model (LLM), utilizes Causal Mask during training to enable reusability of the KV Cache from previous rounds in continuous conversations. These optimizations result in enhanced memory utilization. For instance, while the initial ChatGLM-6B model with 6GB of GPU memory could generate a maximum of 1119 characters before running out of memory, ChatGLM2-6B can generate at least 8192 characters under the same memory constraints.

Actionable Advice:

  1. Evaluate the communication requirements of your multi-machine, multi-card systems and consider implementing RDMA technology, such as InfiniBand or RoCEv2, to significantly reduce latency and improve performance.

  2. When deploying large GPU clusters or high-performance RDMA networks, carefully select the appropriate network architecture solution based on factors like scalability, cost, and ease of configuration. Consider InfiniBand for ultra-large-scale deployments and RoCE for versatile usage scenarios.

  3. Stay updated with advancements in language models, such as the A100-SXM4-80G Multi-Query Attention technology, to optimize memory utilization and enhance the performance of AI-driven applications.

Conclusion:
In the quest for reduced latency and improved performance, network architecture plays a vital role. RDMA technology, specifically the InfiniBand and RoCEv2 solutions, offers significant advantages over traditional TCP/IP networks. While InfiniBand excels in ultra-large-scale GPU clusters, RoCE provides versatility and cost-effectiveness. Understanding the key components and major suppliers in the InfiniBand market is crucial for making informed decisions. Additionally, advancements in language models, such as the A100-SXM4-80G Multi-Query Attention and ChatGLM2-6B, contribute to enhanced performance in AI-driven applications. By embracing these technologies and considering the actionable advice provided, organizations can optimize their network architecture and unlock new possibilities for efficient and high-performance communication.

Sources

← Back to Library

Hatch New Ideas with Glasp AI 🐣

Glasp AI allows you to hatch new ideas based on your curated content. Let's curate and create with Glasp AI :)

Start Hatching 🐣