The Evolution of Memory Architecture: CXL and RDMA's Collaborative Development

Kevin Di

Hatched by Kevin Di

Apr 12, 2024

3 min read

0

The Evolution of Memory Architecture: CXL and RDMA's Collaborative Development

Introduction:
In the world of technology, advancements in memory architecture play a crucial role in enhancing performance and efficiency. Two significant developments in this area are Speculative decoding in the 8xH100 inference unit and the collaborative development of CXL (Compute Express Link) and RDMA (Remote Direct Memory Access). This article explores the potential of speculative decoding, the opportunities created by the DRAM-based capacity layer, and the industry support for CXL. Additionally, it delves into the different device types and protocols defined by the CXL specification and their implications for server designers.

Speculative Decoding and the 8xH100 Inference Unit:
Speculative decoding is a technique that can significantly enhance the throughput of the 8xH100 inference unit. With this approach, the unit can achieve impressive throughputs of up to 420 tokens per second per user. While this figure is already remarkable, it's worth noting that implementing speculative decoding on MoE (Mixture of Experts) models poses certain challenges. Despite these challenges, speculative decoding holds immense potential for accelerating inference processes and improving overall performance.

The Opportunities of DRAM-based Capacity Layer:
The concept of swapping memory pages to solid-state drives (SSDs) often leads to substantial performance losses. However, this drawback has created an opportunity for a new capacity layer based on DRAM. Referred to as "remote memory," this DRAM can exist in a separate server or memory device. CXL, with its relatively short history of just over three years, has garnered significant industry support, surpassing previous coherent interconnect standards like CCIX, OpenCAPI, and HyperTransport. Notably, even though Intel was the original developer of the CXL specification, AMD has also embraced and implemented it. For server designers, adding CXL support only requires the latest EPYC or Xeon processors and careful consideration of PCIe channel allocation.

Understanding CXL Device Types and Protocols:
The CXL specification defines three device types and three protocols required for different use cases. The first protocol is CXL.mem, which enables cache-coherent memory access. The second protocol, CXL.io, is required for all three device types but is primarily used for configuration and control purposes in Type 3 devices. One key distinction is that CXL.mem (and CXL.cache) utilize fixed-length messages, while CXL.io employs variable-length data packets similar to PCIe. In versions 1.1 and 2.0, CXL.mem utilizes a 68-byte flow control unit (flit) to handle 64-byte cache lines. However, CXL 3.0 introduces a 256-byte flit, as introduced in PCIe 6.0, to accommodate forward error correction (FEC) while optimizing latency. It also divides the error check (CRC) into two 128-byte blocks for improved efficiency.

Insights and Actionable Advice:

  1. Embrace Speculative Decoding: To enhance the throughput of inference units, consider implementing speculative decoding techniques. While challenges may arise, exploring this approach can lead to significant performance improvements.

  2. Leverage DRAM-based Capacity Layer: Instead of swapping memory pages to SSDs, explore the possibilities offered by a DRAM-based capacity layer. By utilizing remote memory, servers can achieve higher performance without suffering from the performance losses associated with traditional swapping methods.

  3. Adopt CXL for Enhanced Connectivity: With industry-wide support and compatibility with both Intel and AMD processors, integrating CXL into server designs can provide improved connectivity and cache-coherent memory access. Ensure the appropriate allocation of PCIe channels to fully leverage the benefits of CXL.

Conclusion:
Advancements in memory architecture continue to shape the technological landscape, enabling faster and more efficient computing. Speculative decoding and the potential it holds for accelerating inference processes, along with the opportunities presented by a DRAM-based capacity layer, showcase the continuous evolution in this field. Furthermore, the collaborative development of CXL and RDMA signifies the industry's recognition of the importance of enhanced connectivity and cache-coherent memory access. By embracing speculative decoding, leveraging DRAM-based capacity, and adopting CXL, technology professionals can unlock new levels of performance and efficiency in their systems.

Sources

← Back to Library

Hatch New Ideas with Glasp AI 🐣

Glasp AI allows you to hatch new ideas based on your curated content. Let's curate and create with Glasp AI :)

Start Hatching 🐣