"Nvidia H100 GPUs: Supply and Demand" - Startup Co-Founded by YouTuber Airrack Sells to VidIQ: The Impact on GPU Demand
Hatched by David Tao
Jan 31, 2024
5 min read
16 views
"Nvidia H100 GPUs: Supply and Demand" - Startup Co-Founded by YouTuber Airrack Sells to VidIQ: The Impact on GPU Demand
In recent years, the demand for high-end GPUs, particularly Nvidia's H100 GPUs, has been on the rise. Startups, especially those involved in fine-tuning large open-source models, have been seeking these powerful GPUs to drive their operations. Companies using private clouds, such as CoreWeave and Lambda, have also been investing in hundreds or even thousands of H100s for their machine learning models.
So, what exactly are these high-end GPUs being used for? For companies with extensive H100 deployments, the primary use case is fine-tuning large language models (LLMs) and some diffusion model work. These startups, some of which are building new models from scratch, are securing contracts worth millions of dollars over several years. They require a significant number of GPUs, ranging from a few hundred to a few thousand, to meet their training needs.
When it comes to GPU preferences, the H100 is the go-to choice for most companies. It offers the fastest performance for both training and inference tasks, making it ideal for LLMs. Additionally, the H100 provides an excellent price-performance ratio for inference, further solidifying its position as the preferred GPU for startups.
In the realm of LLM training and inference, several factors play a crucial role in determining the GPU of choice. These factors include memory bandwidth, FLOPS (tensor cores or equivalent matrix multiplication units), caches and cache latencies, additional features like FP8 compute, compute performance (related to the number of CUDA cores), and interconnect speed (e.g., InfiniBand). The H100 outshines the A100 in these areas, making it the preferred option due to lower cache latencies and better FP8 compute capabilities.
One might wonder why LLM companies do not opt for AMD GPUs instead. While it is theoretically possible for a company to purchase AMD GPUs, the process of getting everything to work seamlessly takes time. This development time, even if just a couple of months, can mean being late to market compared to competitors. Thus, Nvidia's CUDA framework serves as a significant advantage for the company, acting as a moat against AMD's GPUs.
The comparison between H100s and A100s reveals that the former offers approximately 3.5 times faster performance for 16-bit inference tasks and about 2.3 times faster for 16-bit training. Given these metrics, it is expected that most companies will opt for 8-GPU HGX H100s over DGX H100s or 4-GPU HGX H100 servers.
Of course, the cost of these GPUs is a significant consideration for startups and companies. A single DGX H100 with 8 H100 GPUs costs around $460,000, including the required support. However, startups can avail the Inception discount, which offers approximately $50,000 off, making it more affordable. Startups can use this discount for up to 8 DGX H100 boxes, resulting in a total of 64 H100s.
Considering the number of GPUs needed for training large models, it is estimated that GPT-4 was likely trained on somewhere between 10,000 to 25,000 A100s. Meta, Tesla, and Stability AI have already invested in thousands of A100s for their operations. Inflection, on the other hand, utilized 3,500 H100s for their GPT-3.5 equivalent model.
Based on these figures, it is evident that the demand for H100s is substantial. OpenAI, Inflection, Meta, big clouds like Azure, Google Cloud, and AWS, as well as private clouds like Lambda and CoreWeave, are all potential buyers of H100s. These numbers add up to an estimated 432,000 H100s, translating to approximately $15 billion worth of GPUs. Additionally, Chinese companies like ByteDance (TikTok), Baidu, and Tencent are also expected to invest heavily in H800s.
TSMC is the manufacturer responsible for producing the H100 GPUs. The production, packaging, and testing process takes around 6 months before the GPUs are ready to be sold to customers. However, the bottleneck lies in the CoWoS packaging, a 3D stacking technology, which TSMC needs to address to meet the growing demand.
In terms of H100 previews and availability, CoreWeave was the first to launch their H100 preview. Other major cloud providers like Azure, Oracle, Lambda Labs, AWS, and Google Cloud followed suit, each announcing their availability at different times. These previews allowed customers to get a glimpse of the H100's capabilities and performance.
Nvidia's allocation strategy plays a crucial role in determining which customers receive a certain number of GPUs. While they have a standard allocation per customer, the specific end customer matters to Nvidia. For instance, if a cloud provider like Azure requests 10,000 H100s exclusively for a specific end customer, Nvidia considers the request differently compared to a general request for 10,000 H100s for Azure's cloud. Nvidia prefers customers with strong brand names or startups with impressive pedigrees. They are also cautious about providing large allocations to companies that directly compete with them in the GPU market.
In conclusion, the demand for Nvidia's H100 GPUs, especially among startups and companies involved in fine-tuning large language models, is high. The H100 offers superior performance for training and inference tasks, making it the preferred choice for most companies. However, the availability and production timeline of these GPUs, along with the growing competition and demand, pose challenges for both Nvidia and manufacturers like TSMC. Despite these challenges, the market for high-end GPUs continues to expand, driven by the increasing demand for advanced machine learning models and applications.
Actionable Advice:
-
Startups and companies looking to invest in high-end GPUs for machine learning should carefully consider the performance, price, and support offered by different GPU options. The H100, with its superior performance and price-performance ratio, is a popular choice for training large models.
-
When planning for large-scale GPU deployments, it is crucial to account for the production and packaging timelines. The demand for GPUs often exceeds the manufacturing capacity, leading to delays in availability. Startups should plan accordingly and consider potential alternatives if there are significant delays.
-
Building strong partnerships and having a reputable brand name can influence Nvidia's allocation decisions. Startups should focus on establishing their credentials and demonstrating their potential to drive innovation in the GPU market. This can increase their chances of receiving larger allocations and gaining access to the latest GPU technologies.
In summary, the demand for Nvidia's H100 GPUs continues to rise, driven by startups and companies involved in fine-tuning large language models. The H100's superior performance and price-performance ratio make it the preferred choice for training and inference tasks. However, challenges related to production capacity, competition, and the need for strong partnerships persist in this rapidly evolving market. Startups must navigate these challenges strategically to secure the necessary GPU resources for their machine learning endeavors.
Sources
Hatch New Ideas with Glasp AI 🐣
Glasp AI allows you to hatch new ideas based on your curated content. Let's curate and create with Glasp AI :)
Start Hatching 🐣