The Rise of AI Data Centers: Unveiling the Secrets Behind the "智算中心"

Kevin Di

Hatched by Kevin Di

Jan 25, 2024

4 min read

0

The Rise of AI Data Centers: Unveiling the Secrets Behind the "智算中心"

In recent years, there has been a surge in the development of AI data centers, commonly known as "智算中心" in China. As of the end of 2023, there were a staggering 128 projects with "智算中心" designation, out of which 83 projects had disclosed their scales, totaling over 77,000P in computational power. These centers vary in standards and sizes, with compute capabilities ranging from 50P to a staggering 12,000P and beyond.

One of the key considerations in optimizing AI data centers is the interconnection of parameters, particularly the use of OXC (Optical Cross-Connect) systems. To achieve the best solution, 1024 A100 GPUs are employed to train GPT models. The GPUs are grouped into 8P nodes, with model parallelism within each node. This is followed by an 8-level pipelining parallelism using 8 groups of 8x8, resulting in a batch size of 16 for data parallelism. The key lies in utilizing the parameter plane network transmission and pipelining parallelism, which demands significantly higher bandwidth than storage network requirements.

But what exactly is OXC? OXC stands for Optical Cross-Connect, which is a system of optical interconnections. Unlike traditional switches that rely on electrical exchanges, OXC leverages optical technology for data transmission. While optical fibers are abundantly available, they are primarily used for data transmission rather than data exchange. In the current data exchange process within data centers, data goes through multiple conversions - electrical to optical, optical transmission, optical to electrical, and so on. This conversion process incurs significant latency, akin to the difference between railway and road transportation systems. Optical exchange, on the other hand, faces the challenge of slower switching speeds due to the need for changing routes, resulting in delays in the millisecond range. It is unable to handle the frequency and complexity of interactions found in urban road networks.

So, how does optical exchange work? It relies on a series of reflective mirrors to direct an incoming light beam to a specific output port. While there are numerous online resources explaining the principles in detail, the management of optical exchange systems can be complex, as multiple light beams must not be directed to a single output port simultaneously.

When it comes to AI data centers, the requirements for interconnectivity are substantial. For instance, a system with 2 CPUs and 8 GPUs would only require a bandwidth of 200Gbps for the CPUs. However, with 8 GPUs, the requirement skyrockets, necessitating eight 400Gbps cards, resulting in a total bandwidth of 3,200Gbps. With a hypothetical scenario of 20,000 GPUs, each with a 400Gbps bandwidth, the total exchange bandwidth would reach a staggering 8,000,000Gbps, or 8Pbps. This specialized network, which needs to provide 16 times the performance of a standard Data Center Network (DCN) and achieve non-converging binary bandwidth, comes at a significant cost. Implementing an InfiniBand (IB) network would require the purchase of mlx cards, switches, and optical modules, making the construction cost of this network account for up to 20% of the entire data center budget. However, in comparison to the cost of GPU cards, this expense may be justifiable for those with substantial financial resources.

In conclusion, the rise of AI data centers, represented by the "智算中心," is transforming the landscape of computational power in China. The optimization of interconnectivity through OXC systems is crucial for achieving the best performance. With the demand for AI-specific networks surpassing standard DCNs by 16 times, the cost of constructing such networks becomes a significant consideration. To navigate this evolving landscape, here are three actionable pieces of advice:

  1. Consider the scale and compute requirements of your AI project carefully. Understanding the computational power needed and the potential network bottlenecks will help in making informed decisions about infrastructure investments.

  2. Stay updated with the latest advancements in interconnectivity technologies, such as OXC systems. Keeping a pulse on emerging trends will allow you to leverage cutting-edge solutions to optimize your AI data center's performance.

  3. Work closely with experts and industry leaders who have experience in building and managing AI data centers. Collaborating with knowledgeable professionals can help ensure that your data center design and implementation align with best practices and industry standards.

By following these recommendations, you can navigate the complexities of AI data center development and harness the power of computational resources to drive innovation and advancements in AI technology.

Sources

← Back to Library

Hatch New Ideas with Glasp AI 🐣

Glasp AI allows you to hatch new ideas based on your curated content. Let's curate and create with Glasp AI :)

Start Hatching 🐣