And make allowance for their doubting too;

David Tao

Hatched by David Tao

Aug 12, 2023

4 min read

0

And make allowance for their doubting too;
If you can wait and not be tired by waiting,
Or being lied about, don't deal in lies,
Or being hated, don't give way to hating,

If you can dream—and not make dreams your master;
If you can think—and not make thoughts your aim;
If you can meet with Triumph and Disaster
And treat those two impostors just the same;
If you can bear to hear the truth you've spoken
Twisted by knaves to make a trap for fools,

Or watch the things you gave your life to, broken,
And stoop and build 'em up with worn-out tools:

If you can make one heap of all your winnings
And risk it on one turn of pitch-and-toss,
And lose, and start again at your beginnings
And never breathe a word about your loss;
If you can force your heart and nerve and sinew
To serve your turn long after they are gone,
And so hold on when there is nothing in you
Except the Will which says to them: 'Hold on!'

If you can talk with crowds and keep your virtue,
Or walk with Kings—nor lose the common touch,
If neither foes nor loving friends can hurt you,
If all men count with you, but none too much;
If you can fill the unforgiving minute
With sixty seconds' worth of distance run,
Yours is the Earth and everything that's in it,
And—which is more—you'll be a Man, my son!

Nvidia H100 GPUs: Supply and Demand

In the world of startups and companies utilizing large open-source models (LLMs), the demand for high-end GPUs like the Nvidia H100 is on the rise. These GPUs are being used for tasks such as fine-tuning existing models and building new models from scratch. With contracts ranging from $10 million to $50 million over three years, it's no wonder that companies are investing in hundreds or even thousands of H100s.

The reason why H100s are the preferred choice for many companies is because of their speed and performance. They are the fastest GPUs for both training and inference when it comes to LLMs. Additionally, the H100 has better scalability with higher numbers of GPUs, which is crucial for startups looking to quickly launch or improve their models.

When it comes to LLM training, there are several factors that companies prioritize. These include memory bandwidth, FLOPS (tensor cores or equivalent matrix multiplication units), caches and cache latencies, additional features like FP8 compute, compute performance, and interconnect speed. The H100 is favored over the A100 because of its lower cache latencies and FP8 compute capabilities.

While theoretically, companies can purchase AMD GPUs, the time it takes to get everything to work is a major deterrent. The development time required to make AMD GPUs compatible with existing systems can result in delays and put companies at a disadvantage compared to their competitors. This is why Nvidia's CUDA platform has become the industry standard.

In terms of cost, the 1x DGX H100 with 8x H100 GPUs is priced at $460,000, including the required support. Startups can avail the Inception discount, which offers $50,000 off and can be used for up to 8x DGX H100 boxes, totaling 64 H100s.

The demand for GPUs is staggering. GPT-4, for example, was likely trained on anywhere between 10,000 to 25,000 A100s. Companies like Meta, Tesla, and Stability AI have thousands of A100s in their possession. Inflection used 3,500 H100s for their GPT-3.5 equivalent model. With these numbers in mind, it's estimated that companies may require hundreds of thousands of H100s, amounting to billions of dollars worth of GPUs.

TSMC is the manufacturer behind the H100 GPUs, and the production process, including packaging and testing, takes approximately six months. The bottleneck in the production cycle lies in the CoWoS packaging, a 3D stacking technology.

The big cloud providers, such as Azure, Oracle, Lambda Labs, AWS, and Google Cloud, have all launched previews of the H100. Nvidia allocates a specific number of GPUs per customer, taking into consideration the end customer's reputation and brand name. However, they are cautious about providing large allocations to companies that directly compete with them.

In conclusion, the demand for Nvidia H100 GPUs is driven by the need for high-performance computing in the field of LLMs. Startups and companies require GPUs that can handle large-scale training and inference tasks. While there is potential for AMD GPUs to enter the market, the time required to make them compatible with existing systems poses a significant barrier. Nvidia's CUDA platform and the speed and performance of their H100 GPUs have solidified their position as the industry leader.

Actionable Advice:

  1. For startups or companies looking to venture into the field of large open-source models, consider investing in Nvidia H100 GPUs. Their speed and scalability make them the ideal choice for training and inference tasks.

  2. When considering GPU options, prioritize factors such as memory bandwidth, FLOPS, caches and cache latencies, additional features, compute performance, and interconnect speed. A comprehensive understanding of these factors will help you make informed decisions.

  3. Stay updated with the latest developments in the GPU market, including new releases and advancements in technology. This will ensure that you are aware of the most efficient and cost-effective options available.

By following these actionable advice, startups and companies can optimize their GPU usage and stay ahead in the rapidly evolving field of large open-source models.

Sources

← Back to Library

Hatch New Ideas with Glasp AI 🐣

Glasp AI allows you to hatch new ideas based on your curated content. Let's curate and create with Glasp AI :)

Start Hatching 🐣