How to Build a Data Center for Startups: owning GPUs vs renting

18.7K views
•
September 5, 2026
by
20VC with Harry Stebbings
YouTube video player
How to Build a Data Center for Startups: owning GPUs vs renting

TL;DR

Owning GPUs can be cheaper over time than renting when used at scale, since yearly costs can undercut hourly cloud rates. The approach involves balancing training and inference needs, memory locality, and capacity to run many experiments simultaneously. Startups can leverage owned hardware and spot or long-term cloud rentals to optimize costs and speed.

Transcript

You said 11 raps kind of leaprogged you. Is that on you? Yeah, 100% it's on me. It's the biggest strategic mistake I made in the history of Speedify. How do you reflect on that? So today is a real freaking discussion. Cliff Whitesman, founder and CEO at Speechify, one of the fastest growing text to speech startups in the world on the show. The best... Read More

Key Insights

  • Owning GPUs often reduces annual costs versus renting if utilization is high enough.
  • Memory locality across a large GPU cluster is essential for large scale training.
  • Inference can run on older GPUs while training uses the newest hardware, optimizing cost and performance.
  • A mix of owned hardware and cloud rentals can balance capacity and flexibility.
  • Forecasting demand involves analyzing peak months and diversifying procurement through long term and spot rentals.
  • Hardware depreciation is expected but can be offset by uptime and utilization across teams.
  • Data center logistics include energy supply, networking, and skilled on-site engineers for maintenance.
  • Open source models can be run on owned hardware to reduce token costs compared with branded cloud options.

Explore YouTube Video Summarizer or Get YouTube Transcript Extractor

Questions & Answers

Q: What is the main reason Speechify buys GPUs instead of renting them?

The primary reason Speechify buys GPUs is to reduce ongoing costs for large scale training and to ensure immediate access to a large memory cluster. They compare the annual renting cost of a high end GPU, which can be 3.5 to 5 dollars per hour, to a one time purchase around the price of a single card, arguing that owning the hardware is financially advantageous when usage is heavy and sustained. Owning also allows them to run many experiments simultaneously, which is critical for rapid iteration and model development.

Q: How does Speechify differentiate between training and inference in hardware strategy?

Speechify treats training and inference as two separate needs with different hardware requirements. Training requires fast, high end GPUs to test hypotheses quickly, while inference can run adequately on older or less powerful GPUs and still respond within acceptable latency. This allows them to allocate new hardware for training while continuing to use legacy GPUs for inference, optimizing cost without compromising user experience.

Q: What role do memory and data locality play in their compute strategy?

Memory locality is crucial because large scale training needs a cluster where all GPUs can access the same data efficiently. Speechify emphasizes that renting through hyperscalers often cannot provide the same level of memory locality for a training run of their size, so owning hardware enables faster, more controlled data access and reduces bottlenecks during training.

Q: How does the economics of owning GPUs compare to renting for a year?

The economics hinge on a simple comparison: if a single GPU card costs around 30,000 dollars and cloud rental costs are 3.5 to 5 dollars per hour, renting for a year can total 35,000 to 50,000 dollars. Therefore owning can be cheaper by about 1.5x to 2x over a year, especially when you need persistent capacity and multiple GPUs running in parallel.

Q: Why do they still rent GPUs from hyperscalers if they own many GPUs?

They still rent because cloud rental provides flexibility to scale up during peak demand, run spot instances to reduce costs, and access newer hardware types when needed. The combination of owned GPUs with cloud rentals allows Speechify to balance long term capacity with short term spikes in demand, ensuring they can meet user needs without overprovisioning.

Q: What is their approach to forecasting GPU demand and purchases?

Forecasting involves estimating the average monthly demand for training and inference, then layering in anticipated peaks for specific months. They plan 100 percent base capacity, add 20 percent for future needs, and allocate another 25 percent via long term contracts with hyperscalers, reserving remaining capacity for spot instances to stay flexible.

Q: How does the team size influence the compute strategy?

Team growth directly affects compute demand because more engineers mean more experiments and models to train. They project expanding from 45 engineers to about 150, which increases the need for parallel training capacity and reliable access to GPUs. Planning aligns with hiring to ensure hardware capacity matches development needs.

Q: What considerations are there for deploying hardware in data centers?

Deploying hardware requires access to data center space, power, networking, and on site engineers who can install and maintain equipment. They emphasize the importance of energy availability as a major constraint, plus the need for efficient cooling and robust networking to support a multi rack GPU cluster that runs large scale AI workloads.

Summary & Key Takeaways

  • A data center strategy centers on owning GPU hardware to reduce long term costs and enable large scale AI training and inference.

  • A blended approach uses owned hardware, long term cloud contracts, and spot rentals to meet variable demand across months.

  • A successful setup requires planning around memory locality, power and cooling, and team growth to maximize return on investment.


Read in Other Languages (beta)

Share This Summary 📚

Explore More Summaries from 20VC with Harry Stebbings 📚