Nvidia: The Next Inflection Point
Hatched by David Tao
Sep 07, 2023
4 min read
11 views
Nvidia: The Next Inflection Point
In the world of deep learning and GPU computing, Nvidia has emerged as a dominant player. Their CUDA platform allows developers to easily deploy deep learning models on Nvidia GPUs, making it the go-to choice for many developers. However, this compatibility is limited to Nvidia GPUs, creating a barrier for those who want to use other GPU brands.
One of the main reasons why Nvidia GPUs are in high demand is their performance and speed. Startups that are heavily involved in fine-tuning large open-source models require powerful GPUs like the Nvidia H100. These startups are often working on multi-million dollar contracts over several years, using hundreds or even thousands of GPUs. For them, the H100 is the preferred choice due to its fast training and inference capabilities.
When it comes to training deep learning models, there are several factors that companies consider. These include memory bandwidth, FLOPS (tensor cores or equivalent matrix multiplication units), caches and cache latencies, additional features like FP8 compute, compute performance, and interconnect speed. The H100 excels in many of these areas, which is why it is favored over the A100. Lower cache latencies and FP8 compute make the H100 a better option for startups looking to compress time and improve their models quickly.
Despite the availability of AMD GPUs, most companies stick to Nvidia due to the time and effort required to make AMD GPUs work seamlessly. The development time alone can set a company back by months, putting them at a disadvantage compared to their competitors. Nvidia's CUDA platform acts as a moat, keeping competitors at bay and solidifying their dominance in the market.
The cost of these high-end GPUs is another important consideration. An 8-GPU HGX H100 setup, which is popular among most companies, costs around $460,000, including required support. Startups can avail the Inception discount, which provides a $50,000 reduction and can be used on up to 8 DGX H100 boxes, totaling 64 H100s. The high cost of these GPUs reflects their value and the demand for them in the market.
The number of GPUs needed for training deep learning models can vary significantly. GPT-4, for example, was likely trained on anywhere between 10,000 to 25,000 A100s. Companies like Meta, Tesla, and Stability AI have thousands of A100s in their possession. Inflection, a company known for its GPT-3.5 equivalent model, used 3,500 H100s for training. These numbers give us an idea of the scale at which these models are being trained and the demand for GPUs.
Predicting the number of H100s that companies may want is challenging but essential. OpenAI may require around 50,000 H100s, while Inflection wants 22,000. Big clouds like Azure, Google Cloud, and AWS may want 30,000 each, and private clouds like Lambda and CoreWeave may want a total of 100,000. These numbers are estimates, but they highlight the massive demand for H100s, amounting to billions of dollars worth of GPUs.
TSMC is the manufacturer of the H100 GPUs, and their production process takes around 6 months. The bottleneck in production is not wafer starts but CoWoS packaging, which is a 3D stacking technique. This information sheds light on the time and effort required to meet the demand for these GPUs.
The big clouds, including Azure, Oracle, Lambda Labs, AWS, and Google Cloud, have all launched their H100 previews at different times. Nvidia allocates GPUs per customer, but the end customer matters. Nvidia prefers customers with strong brand names or startups with a strong pedigree. They also want to avoid giving large allocations to companies that directly compete with them. This allocation process ensures that Nvidia maintains control over their GPUs and the reputation associated with their brand.
In conclusion, Nvidia is at an inflection point in the GPU market. Their CUDA platform and high-performance GPUs like the H100 have solidified their position as the go-to choice for deep learning models. The demand for these GPUs is staggering, with companies requiring tens of thousands of them. The production process and allocation strategy play a crucial role in meeting this demand.
Actionable Advice:
- Consider the specific requirements of your deep learning models and choose the GPU that best suits your needs. Nvidia's H100 is often the preferred choice due to its performance and compatibility with CUDA.
- If you're a startup looking to fine-tune large open-source models, prioritize speed and efficiency in your GPU selection. The H100's fast training and inference capabilities make it a valuable asset.
- Take into account the time and effort required to incorporate AMD GPUs into your workflow. Nvidia's CUDA platform provides a seamless experience and eliminates the need for extensive development time.
By understanding the dynamics of the GPU market and the factors that drive demand, companies can make informed decisions when it comes to selecting the right GPUs for their deep learning projects. Nvidia's position as a market leader is a testament to their commitment to innovation and their ability to meet the evolving needs of the industry.
Sources
Hatch New Ideas with Glasp AI 🐣
Glasp AI allows you to hatch new ideas based on your curated content. Let's curate and create with Glasp AI :)
Start Hatching 🐣