"Nvidia H100 GPUs: Supply and Demand" - Escaping State Controls in 2022

David Tao

Hatched by David Tao

Aug 26, 2023

5 min read

0

"Nvidia H100 GPUs: Supply and Demand" - Escaping State Controls in 2022

In 2022, the demand for Nvidia H100 GPUs has been on the rise due to various factors. Startups, in particular, are utilizing these high-end GPUs for fine-tuning large open source models. Many of these companies are involved in building new models from scratch, often securing contracts worth millions of dollars over a span of several years. Their need for H100 GPUs stems from the fact that these GPUs offer the fastest performance for both training and inference for Language and Learning Models (LLMs).

When it comes to LLM training and inference, startups prioritize factors such as memory bandwidth, FLOPS, caches and cache latencies, additional features like FP8 compute, compute performance, and interconnect speed. The H100 GPUs are preferred over the A100 GPUs due to their lower cache latencies and FP8 compute capabilities. These factors contribute to faster training times and quicker time-to-market for startups.

Despite the availability of AMD GPUs, many LLM companies hesitate to adopt them due to the additional time required for development and integration. The compatibility and familiarity of Nvidia's CUDA platform make it the preferred choice for most companies, as it allows them to avoid potential delays in product launches. However, it's worth noting that the time difference between adopting AMD GPUs and Nvidia GPUs may not be as significant as initially perceived.

The cost of these GPUs varies depending on the configuration. For example, a DGX H100 with 8 H100 GPUs is priced at $460,000, including the required support. However, startups can benefit from the Inception discount, which offers a reduction of approximately $50,000 per DGX H100 box, making it a more affordable option.

In terms of the number of GPUs needed, the scale varies based on the specific project. GPT-4, for instance, was likely trained using anywhere between 10,000 to 25,000 A100 GPUs. Companies like Meta, Tesla, Stability AI, and Inflection have invested in thousands of A100 GPUs for their models. Inflection, for example, used 3,500 H100 GPUs for their GPT-3.5 equivalent model.

Considering the demand for H100 GPUs, it is estimated that companies like OpenAI, Inflection, Meta, and various big clouds could potentially require hundreds of thousands of H100 GPUs. This demand translates to a substantial investment of approximately $15 billion in GPUs alone. Additionally, Chinese companies like ByteDance, Baidu, and Tencent are also expected to require a significant number of H800 GPUs.

The production of H100 GPUs is handled by TSMC, with a production cycle of approximately six months from start to finish. The packaging process, specifically CoWoS (3D stacking) packaging, is the bottleneck at TSMC. Despite this, startups tend to choose suppliers based on a combination of factors such as support, price, and capacity.

The big clouds, including Azure, Oracle, Lambda Labs, AWS, and Google Cloud, have started offering previews and limited availability of H100 GPUs to their customers. Nvidia allocates GPUs based on the end customer, taking into consideration factors such as brand name and startup pedigree. They prefer not to give large allocations to companies that directly compete with them in the GPU market.

In a different context, in 2022, the author recounts their experience of escaping state controls by taking the final flight out of Shanghai to Yunnan, a province in China's farthest southwest. Yunnan boasts diverse geography, ranging from historic Tibet in the north to a landscape reminiscent of Thailand in the south. The province is known for its natural beauty, including rainforests, rice terraces, fast rivers, and snowy mountains. It is also home to numerous ethnic groups that have historically resisted Han rule, making it a destination for tourists seeking ethnic exoticism.

In conclusion, the demand for Nvidia H100 GPUs is driven by the growing needs of startups and companies engaged in fine-tuning large open source models. The preference for H100 GPUs over A100 GPUs is due to their superior performance in terms of training and inference for LLMs. While Nvidia's CUDA platform remains the dominant choice, the potential for AMD GPUs exists, despite the initial development hurdles. The production and allocation of H100 GPUs are key factors influencing their availability and adoption. Overall, the demand for these GPUs, coupled with their high price, reflects the significant investment companies are willing to make to stay competitive in the field of artificial intelligence.

Three actionable advice for companies seeking to leverage GPUs for LLM development and training:

  1. Consider the trade-offs between performance and cost: While the H100 GPUs offer superior performance, it's essential to evaluate the cost-effectiveness of these GPUs for your specific project. Assess whether the performance gains justify the price difference compared to alternative options.

  2. Stay up-to-date with GPU developments: As technology evolves rapidly, it's crucial to stay informed about the latest advancements in GPU technology. Keep an eye on both Nvidia and AMD GPUs, as well as other emerging players in the market. Evaluate their capabilities and compatibility with your infrastructure to make informed decisions.

  3. Build strong partnerships and leverage cloud services: Collaborate with reputable GPU suppliers and cloud providers to ensure access to the latest hardware and resources. Establishing strong partnerships can help secure priority allocations and access to cutting-edge technologies, enabling your organization to stay at the forefront of AI development.

By considering these recommendations and understanding the dynamics of the GPU market, companies can maximize their potential for LLM training and inference, ultimately driving innovation and success in the field of artificial intelligence.

Sources

← Back to Library

Hatch New Ideas with Glasp AI 🐣

Glasp AI allows you to hatch new ideas based on your curated content. Let's curate and create with Glasp AI :)

Start Hatching 🐣