How Does AI Hardware-Software Co-Design Work?

103.0K views
•
June 30, 2026
by
Sequoia Capital
YouTube video player
How Does AI Hardware-Software Co-Design Work?

TL;DR

AI efficiency improves most when models, kernels, interconnects, and silicon are optimized together, allowing several incremental gains to compound into improvements approaching 100x. SemiAnalysis measures this progress through InferenceX, a continuously updated benchmark that runs current models daily on more than $50 million of donated hardware and tracks cost per unit of quality.

Transcript

I think it's really fun inside of semi analysis because we have 90 people and like a big chunk of them are technologists engineers across the whole supply chain. Um, and then a big chunk is people who are formerly at hedge funds. And you see these arguments like people are like, "Oh, well that doesn't matter." And it's like, then someone's like, "W... Read More

Key Insights

  • Hardware-software co-design is the process of optimizing the model, kernels, interconnects, and silicon together. Several improvements of roughly 2x can compound into gains approaching 100x, making coordinated system design more consequential than evaluating faster chips in isolation.
  • InferenceX is a living benchmark that runs the latest models every day on more than $50 million of donated hardware. It evaluates cost relative to model quality, giving SemiAnalysis a continuously updated view of practical inference efficiency across competing systems.
  • Inference cost per unit of quality is falling at roughly 60x annually according to InferenceX. This metric combines economic cost with useful model quality, so it captures improvements arising from models, software, and hardware rather than attributing progress to silicon alone.
  • Model architecture is shaped by the hardware and interconnect on which it is developed. DeepSeek's experts were designed around Nvidia's Hopper platform, and that specialization helps explain why the same architecture is difficult to run effectively on TPUs.
  • Sparse and dense models create different hardware incentives. OpenAI's sparser models and Anthropic's denser models pull the companies toward different system designs, showing that no single accelerator architecture is automatically optimal for every model development strategy.
  • The CUDA moat is shifting from a narrow programming-interface advantage toward a broader ecosystem and co-design advantage. The relevant competitive strength includes software, hardware, model optimization, and the surrounding tools that enable those components to work efficiently together.
  • Cerebras can deliver notable inference speed, but speed alone does not establish universal superiority. Its limitations illustrate why accelerator comparisons must consider model compatibility, economics, system constraints, and quality-adjusted cost instead of relying on a single throughput measurement.
  • The compute crunch persists because models expand the value of useful computational work faster than available compute grows. Patel consequently expects inference to become a market larger than oil and describes Jensen Huang's financing of neoclouds as an effort to create a multipolar infrastructure market.

Install to Summarize YouTube Videos and Get Transcripts

Explore YouTube Video Summarizer or Get YouTube Transcript Extractor

Questions & Answers

Q: What is hardware-software co-design in AI?

Hardware-software co-design means optimizing the AI model, computational kernels, interconnects, software ecosystem, and silicon as a coordinated system. The argument is that the largest efficiency gains do not come from improving one chip specification alone. Multiple gains of roughly 2x across different layers can compound, potentially producing an overall improvement approaching 100x in useful performance or economics.

Q: Why can AI co-design produce gains approaching 100x?

AI co-design can approach 100x because improvements at several system layers multiply rather than merely add. Better model structure, more efficient kernels, suitable interconnects, optimized software, and specialized silicon can each contribute an improvement. When several roughly 2x gains are combined, their cumulative effect becomes much larger than any individual advance, which is why coordinated optimization matters so much.

Q: What is the InferenceX benchmark?

InferenceX is SemiAnalysis's living benchmark for evaluating current AI inference systems. It runs the latest models daily on more than $50 million of donated hardware and measures cost per unit of quality. Because it is continually updated, it is intended to follow rapid changes in models, software, and accelerators instead of presenting a one-time comparison that may quickly become outdated.

Q: How quickly is AI inference cost improving?

InferenceX tracks a roughly 60x annual decline in cost per unit of quality. That measurement is broader than raw chip speed because it relates the price of inference to the quality produced by a model. The reported improvement therefore reflects combined progress in model design, kernels, software, systems, and hardware rather than assigning the entire change to faster silicon.

Q: Why does DeepSeek run differently on Hopper and TPUs?

DeepSeek's expert-based architecture was shaped for Nvidia's Hopper hardware and its interconnect characteristics. That co-design influences how computation and communication are organized across the system. TPUs struggle to run it because an architecture optimized around one platform's properties does not necessarily map efficiently onto another, demonstrating that model design and hardware design cannot always be separated.

Q: How do sparse and dense models affect hardware choices?

Sparse and dense model strategies place different demands on computing systems. OpenAI's sparser models and Anthropic's denser models therefore pull the organizations toward different hardware designs. The contrast shows that accelerator selection depends on model architecture, communication patterns, and software optimization, so a platform that works well for one company's models may be less suitable for another's.

Q: What is the real source of Nvidia's CUDA moat?

The described CUDA moat is not simply the CUDA interface by itself. Its strength increasingly comes from the larger ecosystem connecting models, kernels, software tools, interconnects, and Nvidia hardware. This integrated environment supports extensive co-design and optimization. Competitors consequently need more than compatible programming features, because effective AI systems depend on how well the entire stack operates together.

Q: Why does the AI compute crunch continue despite efficiency gains?

The compute crunch continues because improving models increase the value and range of useful computational work faster than the supply of compute expands. Even as InferenceX records steep reductions in quality-adjusted inference cost, demand can grow more quickly. Patel connects this dynamic to the prediction that inference will become a market larger than oil and to investment in competing neocloud providers.

Summary & Key Takeaways

  • Dylan Patel traces his technical curiosity to repairing an Xbox 360 hardware defect and participating heavily in online technology forums. Years spent comparing processors, GPUs, smartphones, power efficiency, margins, and price-performance taught him to examine semiconductor technology and business economics together rather than treating them as separate subjects.

  • SemiAnalysis grew from Patel's long-running habit of researching, debating, and publishing about technology. After personal and professional disruptions in early 2020, he concentrated more heavily on online analysis. Being doxed prompted him to publish under his real identity, and two carefully prepared posts on his 24th birthday gained significant traction.

  • The central technical argument is that AI's largest efficiency gains emerge from co-design across models, kernels, interconnects, software ecosystems, and chips. InferenceX tests current models daily across donated hardware, while comparisons involving DeepSeek, OpenAI, Anthropic, Nvidia, TPUs, and Cerebras illustrate how model architecture and hardware suitability shape performance and cost.


Read in Other Languages (beta)

Share This Summary 📚

Explore More Summaries from Sequoia Capital 📚