The Future Development Trends of Computing Chips: Insights from the Rise of Large Models like ChatGPT and OAI Server Design Guidelines
Hatched by Kevin Di
Mar 05, 2024
4 min read
6 views
The Future Development Trends of Computing Chips: Insights from the Rise of Large Models like ChatGPT and OAI Server Design Guidelines
In recent years, the development of computing chips has been driven by the rise of large models such as ChatGPT and the need for high-performance networks. These advancements have paved the way for new possibilities and trends in the field of computing. In this article, we will explore the common points between the emergence of large models and the design guidelines for open AI servers, shedding light on the future development trends of computing chips.
One common point between the rise of large models and the development of computing chips is the increasing demand for high-performance networks. Large models like ChatGPT require extensive data transmission between multiple nodes, resulting in a significant amount of east-west network traffic. Statistics show that east-west network traffic accounts for over 85% of the total data flow in large data centers. Similarly, AI training clusters, with their large number of nodes exceeding 1000, experience east-west traffic estimated to be over 90%. To meet these demands, high-performance network functionalities are optimized, including congestion control, multi-path load balancing ECMP, out-of-order delivery, high scalability, fast fault recovery, and Incast optimization. These optimizations aim to achieve superior network performance and efficiency.
Another common point lies in the integration of high-bandwidth networks in computing chips. One notable example is the Gaudi chip, which integrates a high-performance network with ultra-high bandwidth. This integration significantly enhances the efficiency of east-west traffic interaction between cluster nodes and enables the design of even larger and more scalable clusters. The seamless integration of high-performance networks into computing chips is crucial for meeting the growing demands of large models and advancing the field of computing.
Looking at the design guidelines for open AI servers, such as the OAI-UBB 1.0 specification released by OCP, we can identify further insights into the future development trends of computing chips. The OAI-UBB 1.0 specification was introduced to support a variety of AI acceleration cards from different vendors without the need for hardware modifications. It defines the physical and electrical forms of AI acceleration cards to accommodate higher power consumption and larger interconnect bandwidth required for ultra-large-scale deep learning training. The OAI group within OCP unified the design specification of AI acceleration card baseboards, known as OAI-UBB, based on the OAI-UBB 1.0 specification. The OAI-UBB specification further defines the host interface, power supply method, cooling method, management interface, inter-card interconnect topology, and scale-out approach for the baseboard, with 8 OAI-UBB cards forming a cohesive unit.
To achieve external expansion between nodes and create interconnected clusters, the UBB links can be split into ×8 links. However, if all 7 ports are configured as ×16, it would not be possible to achieve external expansion. Therefore, the UBB baseboard limits the interconnect links to ×8 and designates the latter half of port 1 (×8), commonly referred to as the 1H port, for external expansion.
These design guidelines highlight the need for standardized AI acceleration card forms and interfaces to address the challenges arising from diverse forms and interfaces. By unifying the design specifications, computing chips can be developed to meet the demands of ultra-large-scale deep learning training and enable seamless integration between different vendors' products.
In conclusion, the rise of large models and the design guidelines for open AI servers provide valuable insights into the future development trends of computing chips. The increasing demand for high-performance networks and the integration of high-bandwidth networks within computing chips are two prominent trends shaping the field. To stay ahead in this rapidly evolving landscape, here are three actionable pieces of advice:
- Prioritize the optimization of high-performance network functionalities, such as congestion control and load balancing, to enhance the efficiency of east-west traffic in large-scale computing clusters.
- Embrace the integration of high-bandwidth networks within computing chips to enable more scalable and efficient cluster designs, accommodating the growing demands of large models and AI training.
- Adopt standardized design guidelines for AI acceleration cards and interfaces to promote compatibility and interoperability among different vendors' products, facilitating the creation of cohesive and scalable computing systems.
By following these recommendations, chip manufacturers and system designers can navigate the future development trends of computing chips and contribute to the advancement of the field in an increasingly interconnected and data-driven world.
Sources
Hatch New Ideas with Glasp AI 🐣
Glasp AI allows you to hatch new ideas based on your curated content. Let's curate and create with Glasp AI :)
Start Hatching 🐣