The Incredible Power of Large Language Models and H100 GPUs

Kevin Di

Hatched by Kevin Di

May 23, 2024

5 min read

0

The Incredible Power of Large Language Models and H100 GPUs

With the advancements in technology, we are witnessing the rise of large language models and powerful GPUs that are revolutionizing various industries. In this article, we will explore the capabilities and impact of these cutting-edge technologies.

Let's start by delving into the world of large language models. These models have become increasingly popular due to their ability to process and understand vast amounts of text data. One such example is the model with 52 billion parameters, which is a staggering feat. To put it into perspective, the weight size of this model can be calculated by multiplying its parameter count by 2. This means that the model with 52 billion parameters has a weight size of an enormous magnitude.

On the other hand, we have the H100 GPUs, which have taken the GPU market by storm. Compared to its predecessor, the A100, the H100 offers a significant boost in both inference and training speeds. In fact, the H100 has improved inference speed by 3.5 times and training speed by 2.3 times on a single card. When used in server clusters, the training speed can be further enhanced up to 9 times. This means that what used to take a week to accomplish can now be done in just 20 hours.

Although the H100 comes with a higher price tag compared to the A100, its efficiency in training large models has increased by a staggering 200%. This results in a higher "performance per dollar" ratio, making it an attractive choice for many customers. Major players in the industry, such as Microsoft Azure, Google, Oracle, Tesla, and Amazon, have already acquired tens of thousands of H100 GPUs. CoreWeave, a prominent company in this field, reportedly has an allocation promise of 35,000 GPUs, with an actual delivery of around 10,000. Other companies have relatively smaller demands, with only a few thousand GPUs.

According to a prediction by GPU Utils, the current demand for H100 GPUs is approximately 432,000 units. OpenAI requires 50,000 GPUs for training GPT-5, Inflection needs 22,000 GPUs, and Meta has a demand of either 25,000 or 100,000 GPUs, depending on different sources. Each of the four major public cloud providers requires at least 30,000 GPUs, while the private cloud industry demands 100,000 GPUs. Additionally, smaller model manufacturers also require 100,000 GPUs. With NVIDIA expecting to ship around 500,000 H100 GPUs in 2023, the scarcity of H100 GPUs will likely be alleviated by the end of the year.

Looking ahead, the demand for H100 GPUs is projected to soar in 2024. According to the Financial Times, the shipment volume of H100 GPUs is expected to reach a staggering 1.5 to 2 million units, a three to four-fold increase compared to this year's 500,000 units. This exponential growth indicates the widespread adoption and recognition of the power of large language models and the need for more powerful GPUs.

Now, let's take a closer look at the technical details behind the H100 GPUs. The core logic chip has a size of 814mm^2 and is manufactured in TSMC's state-of-the-art Fab 18 in Tainan. The process node used for the chip is referred to as "4N," which stands for 5nm+. Despite the name starting with "4," it is actually a 5nm+ technology. Each 12-inch wafer, with an area of 70,695mm^2, can ideally yield 86 chips. However, considering an 80% yield rate and losses during cutting, the final yield is reduced to 65 core logic chips per wafer.

But what about the cost of producing these core logic chips? According to TSMC's 2023 pricing, a single 12-inch wafer is priced at $13,400. Taking this into account, the cost of each core logic chip is approximately $200.

Moving on to the HBM (High Bandwidth Memory), the specific pricing remains undisclosed by the manufacturers. However, according to reports from South Korean media, the price of HBM is currently around 5 to 6 times higher than existing DRAM products. Considering that the price of GDDR6 VRAM is approximately $3 per GB, the estimated price of HBM is around $15 per GB. Therefore, the cost of HBM for one H100 SXM is around $1,500.

Additionally, we have the CoWoS (Chip-on-Wafer-on-Substrate) packaging technology, which plays a crucial role in the production of AI chips. TSMC's 2022 financial report revealed that the CoWoS process accounted for 7% of the company's total revenue. Based on the production capacity and the size of the bare die, analysts estimate that packaging one AI chip brings approximately $723 in revenue to TSMC.

Taking all these factors into consideration, the total cost of producing an H100 GPU is around $2,500. TSMC's share of this cost, including the core logic chip and CoWoS packaging, is approximately $1,000, while SK Hynix accounts for $1,500. When factoring in other materials like PCB, the overall material cost does not exceed $3,000.

In conclusion, the combination of large language models with billions of parameters and powerful GPUs like the H100 is reshaping the landscape of various industries. From natural language processing to machine learning, these technologies offer unprecedented capabilities and efficiencies. As we move forward, it is crucial to harness their power responsibly and explore new possibilities for innovation.

Actionable Advice:

  1. Embrace large language models: Consider integrating large language models into your business processes to unlock their potential for understanding and generating human-like text.
  2. Leverage powerful GPUs: Invest in high-performance GPUs like the H100 to accelerate training and inference speeds, driving efficiency and productivity in your AI projects.
  3. Stay updated with technological advancements: Keep a close eye on the latest developments in large language models and GPU technologies to stay ahead of the curve and capitalize on emerging opportunities.

Sources

← Back to Library

Hatch New Ideas with Glasp AI 🐣

Glasp AI allows you to hatch new ideas based on your curated content. Let's curate and create with Glasp AI :)

Start Hatching 🐣