"The Synergistic Development of CXL and RDMA in Memory Architecture Evolution"

Kevin Di

Hatched by Kevin Di

Jul 03, 2024

4 min read

0

"The Synergistic Development of CXL and RDMA in Memory Architecture Evolution"

Introduction:
The evolution of memory architecture has paved the way for the development of new capacity layers based on DRAM. One such advancement is the concept of "remote memory," where DRAM can exist in another server or memory device, creating opportunities for increased capacity. This article explores the collaborative development of two key technologies in memory architecture - CXL and RDMA - and their implications for server designers and the industry as a whole.

CXL and its Industry Support:
Despite being a relatively young technology, CXL has garnered significant industry support, surpassing previous coherent interconnect standards like CCIX, OpenCAPI, and HyperTransport. Notably, while Intel was the original developer of the CXL specification, AMD has also embraced and implemented CXL. For server designers, adding CXL support only requires the latest EPYC or Xeon processors and careful allocation of PCIe channels. The CXL specification defines three device types and three protocols required for different use cases, with CXL.mem protocol facilitating cache-coherent memory access. CXL.io protocol is necessary for all three device types, but Type 3 devices primarily use it for configuration and control purposes. The key difference is that CXL.mem (and CXL.cache) employs fixed-length messages, while CXL.io utilizes variable-length data packets similar to PCIe. In CXL 1.1 and 2.0 versions, CXL.mem uses 68-byte flow control units (flits) to handle 64-byte cache lines. CXL 3.0, however, introduces 256-byte flits from PCIe 6.0 for forward error correction (FEC) while optimizing delay with a split of error check (CRC) into two 128-byte blocks.

The Technological and Business Logic of AI Factories and Clouds:
In the realm of AI clusters, three network setups are prevalent. One is the Scale-Up network, such as NVLink, while another is an East-West Scale-Out network based on RDMA. The third framework encompasses the existing Front-End storage/control and North-South traffic. Gilad's speech at GTC titled "Entering A New Frontier of AI Networking Innovation" also delves into this distinction. One solution involves direct CPU-to-GPU interconnects like C2C, while another approach merges Front-End and Scale-Out networks. Google's deployment on its A3 H100 instance exemplifies this merging, wherein any general-purpose CPU computation VM can connect directly to the Scale-Out network through the Front-End. This integration of SCALE-OUT and FRONT-END networks also affirms the compatibility of protocols like GPUDirectTCPX or future options like Falcon to accommodate the lossy FRONT-END network, as opposed to using ROCEv2. AWS's NVL32 and NVL72 systems, on the other hand, continue to utilize EFASRD built with Nitro architecture.

Common Ground and Natural Connection:
While the concepts of CXL and RDMA may seem distinct, they share a common goal of optimizing memory performance and capacity. CXL's focus on cache-coherent memory access and RDMA's emphasis on efficient data transfer and remote memory access converge in their pursuit of improved memory architecture. Both technologies aim to minimize latency, enhance scalability, and enable seamless integration between components, ultimately contributing to more efficient and powerful computing systems. The collaboration and integration of CXL and RDMA can unlock new possibilities in memory architecture, fostering innovation and advancements in various domains, including AI, cloud computing, and data-intensive applications.

Actionable Advice:

  1. Embrace CXL: As CXL gains industry support and becomes more prevalent, server designers should consider incorporating CXL support into their designs by utilizing the latest EPYC or Xeon processors. This will ensure compatibility with emerging memory architectures and enable efficient cache-coherent memory access.
  2. Optimize RDMA Integration: When implementing RDMA in AI clusters or cloud environments, explore merging the Front-End and Scale-Out networks for improved efficiency and reduced latency. This integration can streamline data transfer and enhance overall performance, especially in scenarios where CPU-GPU interconnects are essential.
  3. Stay Informed and Adapt: As memory architecture continues to evolve, it is crucial to stay updated with the latest advancements, research, and industry trends. By actively seeking knowledge and adapting to emerging technologies like CXL and RDMA, server designers can stay ahead of the curve and leverage these innovations to optimize their systems' performance and capabilities.

Conclusion:
The collaborative development of CXL and RDMA in memory architecture presents exciting opportunities for the industry. By combining cache-coherent memory access with efficient data transfer and remote memory capabilities, CXL and RDMA aim to revolutionize memory performance, scalability, and integration. Implementing CXL support, optimizing RDMA integration, and staying informed about evolving memory architectures will enable server designers to unlock the full potential of these technologies and drive innovation in the AI, cloud computing, and data-intensive realms.

Sources

← Back to Library

Hatch New Ideas with Glasp AI 🐣

Glasp AI allows you to hatch new ideas based on your curated content. Let's curate and create with Glasp AI :)

Start Hatching 🐣