The Battle for HBM Supremacy and the Future of Computing Chips

Kevin Di

Hatched by Kevin Di

Feb 13, 2024

5 min read

0

The Battle for HBM Supremacy and the Future of Computing Chips

In recent years, there has been fierce competition among storage giants for the development of High Bandwidth Memory (HBM) technology. HBM is a technique that involves stacking multiple DRAM chips on top of each other using through-silicon vias (TSVs) - thousands of tiny holes that connect the upper and lower chips with vertical electrodes. This technology allows for the stacking of several DRAM chips on a buffer chip, enabling the transmission of signals, instructions, and current through columnar channels that pass through all the chip layers. Compared to traditional packaging methods, HBM technology can reduce the volume by 30% and lower energy consumption by 50%.

HBM offers several advantages over traditional memory solutions. It provides higher bandwidth, more I/O (input/output) counts, lower power consumption, and a smaller form factor. HBM1, for example, operates at a frequency of around 1600 Mbps, has a drain voltage of 1.2V, and a chip density of 2Gb (4-hi). Its bandwidth exceeds that of DDR4 and GDDR5 products while consuming lower power in a smaller form factor, making it ideal for processors with high bandwidth requirements such as GPUs.

To support increased bandwidth and capacity, JEDEC introduced the HBM2E specification at the end of 2018. With a data transfer rate of 3.6Gbps per pin, HBM2E can achieve a memory bandwidth of 461GB/s per stack. Furthermore, HBM2E supports up to 12 stacked DRAMs, with a memory capacity of up to 24GB per stack. Compared to HBM2, HBM2E offers more advanced technology, broader application range, faster speed, and larger capacity.

Samsung, one of the leading players in the storage industry, introduced its 16GB HBM2E Flashbolt, which vertically stacks eight layers of 10-nanometer 16GB DRAM chips. This configuration provides a memory bandwidth level of up to 410GB/s and a data transfer speed of 3.2GB/s per pin.

In January 2022, JEDEC officially released the standard specification for the next-generation high-bandwidth memory, HBM3. The HBM3 specification aims to expand and upgrade storage density, bandwidth, channels, reliability, and energy efficiency. Some of the key features include:

  1. The use of a 0.4V low swing interface in the main interface, reducing operating voltage to 1.1V for improved energy efficiency.
  2. Doubling the data transfer rate compared to HBM2, with a transfer rate of 6.4Gbps per pin and a maximum bandwidth of 819GB/s per chip, thanks to a 1024-bit interface.
  3. Doubling the number of independent channels from 8 to 16, with support for up to 32 channels per chip, including virtual channels.
  4. Support for 4-layer, 8-layer, and 12-layer TSV stacks, with preparation for future expansion to 16-layer TSV stacks.
  5. Each storage layer can have a capacity of 8/16/32Gb, with a starting chip capacity of 4GB (8Gb 4-high) and a maximum capacity of 64GB (32Gb 16-high).
  6. Integration of platform-level RAS (Reliability, Availability, and Serviceability) features, including ECC (Error-Correcting Code) error detection and correction, real-time error reporting, and transparency.

SK Hynix, a major player in the memory market, offers two HBM3 products. One is a 24GB (196Gb) chip stack using the 12-layer silicon through-hole technology, while the other is a 16GB (128Gb) chip stack with 8 layers. Both options provide a bandwidth of 819GB/s, with the former having a chip height of only 30 micrometers. Compared to the previous generation HBM2E with 460GB/s bandwidth, HBM3 offers a 78% increase in bandwidth. Additionally, HBM3 memory incorporates on-chip error correction technology, enhancing product reliability. HBM3 technology has already entered mass production in 2022, with a single chip interface width of 1024 bits and a transfer rate of 6.4Gbps, representing a 1.8x improvement over the previous generation. With 6-layer stacking, it can achieve a total bandwidth of 4.8TB/s.

The rise of large-scale models like ChatGPT has shed light on the future trends in computing chips. With N nodes connected by lines, the total amount of data transmitted through these connections is N*(N-1)/2. In current large data centers, east-west network traffic accounts for over 85% of the total. For AI large model training clusters with over 1000 nodes, east-west traffic likely exceeds 90%. By optimizing high-performance networking capabilities such as congestion control, ECMP load balancing, out-of-order delivery, scalability, fast fault recovery, and Incast optimization, the efficiency of east-west traffic interaction between cluster nodes can be improved. Gaudi, for example, integrates a high-bandwidth, high-performance network, enhancing the efficiency of east-west traffic interaction between cluster nodes and enabling the design of even larger clusters.

Traditional architectures face a "power wall" due to the significant power consumption associated with data movement between memory units and compute units. The energy consumed in data transfer far exceeds that used for computations, resulting in a low percentage of energy and time dedicated to actual computation. The frequent migration of data between memory and processors introduces significant power consumption challenges. New memory technologies aim to alleviate the burden of data movement between memory and processors.

In conclusion, the battle for HBM supremacy among storage giants has led to significant advancements in memory technology. HBM offers higher bandwidth, lower power consumption, and smaller form factors compared to traditional memory solutions. With the introduction of HBM3, the industry is set to witness further improvements in storage density, bandwidth, channels, reliability, and energy efficiency. These advancements will contribute to the development of more powerful and energy-efficient computing systems.

Actionable Advice:

  1. Stay updated on the latest advancements in memory technology, such as HBM3. Understanding the capabilities and benefits of these new technologies can help you make informed decisions when upgrading your computing systems.

  2. Consider the specific requirements of your workload when choosing a memory solution. If you have high bandwidth requirements, such as in AI or GPU-intensive applications, HBM technology may be the ideal choice. However, for other workloads, alternative memory solutions may offer better cost-effectiveness.

  3. Collaborate with technology partners or vendors who have expertise in memory technologies. They can assist in evaluating your workload requirements and recommending the most suitable memory solutions for your specific use case.

Sources

← Back to Library

Hatch New Ideas with Glasp AI 🐣

Glasp AI allows you to hatch new ideas based on your curated content. Let's curate and create with Glasp AI :)

Start Hatching 🐣