How Does DeepSeek Deliver AI at Lower Cost?

TL;DR
DeepSeek achieves competitive AI performance at far lower reported cost by combining efficient engineering, constrained hardware, better algorithms, and a sparse model architecture that activates only part of its capacity at once. Its R1 reasoning model also gained attention because it displays its problem-solving process, is openly available, and has smaller distilled versions that people can run on laptops.
Transcript
Imad deep seek no surprise for you was this something expected or was this something like wow I think it was actually expected like last February I said deep seek one of my favorite AI companies out there um like they took the original ethos that we had at stability another ex hedge fund manager and they released amazing models open I think when th... Read More
Key Insights
- DeepSeek's rise was expected by some AI builders because the company had already released DeepSeek Coder, reached the top of code rankings, and later introduced V3. Its development moved from replicating Meta's Llama toward algorithms and engineering techniques that produced stronger, more independent models.
- DeepSeek V3 reportedly matched GPT-4o while carrying a stated training cost of $6 million. The discussion treats that result as evidence that competitive base models can be trained at a fraction of the spending associated with leading American AI laboratories.
- DeepSeek R1 gained attention because it displays the reasoning used to break down a problem. That visible process, combined with longer thinking and stronger output quality, made interactions feel more like working with another person than receiving an unexplained answer from a closed system.
- R1's rapid public impact came from a combination of usability, benchmark performance, visible reasoning, and open availability. Smaller distilled versions could be run on laptops, allowing people to experiment directly instead of relying exclusively on access to a centrally hosted model.
- DeepSeek's reported query cost was 96 percent lower than the comparison discussed for OpenAI's reasoning system. The participants identify this order-of-magnitude reduction, rather than model quality alone, as a central reason the release challenged assumptions about AI economics and infrastructure requirements.
- DeepSeek reportedly used about 2,000 H800 chips for the cited training run, although the company did not claim that this was its total chip inventory. The speakers say accusations about a much larger hidden GPU supply do not invalidate the stated requirements of that particular run.
- DeepSeek's model uses 640 billion parameters but activates only about 30 billion at one time. This sparse design scales capacity through memory rather than depending only on the fastest silicon, helping the company work around hardware and interconnect constraints described in the discussion.
- Hardware restrictions can create pressure for greater efficiency because teams unable to scale compute freely must improve data, algorithms, memory use, and low-level code. The speakers frame DeepSeek as an example of engineering innovation emerging from the need to accomplish more with constrained resources.
Install to Summarize YouTube Videos and Get Transcripts
Explore YouTube Video Summarizer or Get YouTube Transcript Extractor
Questions & Answers
Q: Why did DeepSeek R1 attract so much attention?
DeepSeek R1 attracted attention because it combined several features rather than relying on a single benchmark result. It showed its reasoning as it broke down problems, produced high-quality answers after thinking longer, was openly available, and had smaller distilled versions that people could run on laptops. Its reported cost advantage also challenged assumptions about how much expensive infrastructure advanced reasoning models require.
Q: How does DeepSeek compare with GPT-4o and OpenAI o1?
The discussion says DeepSeek V3, the base model released in December, matched GPT-4o but did not initially match OpenAI o1. DeepSeek later released R1, a reasoning model that could think longer and compete with o1-style performance. The speakers emphasize that R1 was reportedly 96 percent cheaper and displayed its reasoning, while OpenAI hid the chain-of-thought process behind a thinking message.
Q: How much did DeepSeek V3 reportedly cost to train?
DeepSeek V3 was described as a $6 million training-cost model that matched GPT-4o and other base models. The speakers contrast that figure with the claim that OpenAI spent $3 billion on model training during the referenced year. They also distinguish V3's full training figure from the later R1 evolution, which one speaker estimates may have required only $200,000 beyond the original trained model.
Q: How did DeepSeek work with fewer or slower AI chips?
DeepSeek reportedly used about 2,000 H800 chips for the relevant training run. These chips had reduced interconnect speed, meaning communication between chips was slower. The team compensated through intensive engineering, including low-level PTX code, while emphasizing better data, algorithms, memory use, and efficiency. The speakers argue that the reported training requirements are plausible based on their own experience building models under similar interconnect constraints.
Q: What makes DeepSeek's model architecture efficient?
DeepSeek's architecture contains 640 billion parameters but activates only about 30 billion parameters at one time. The discussion contrasts this with dense models such as a 70 billion parameter Llama model, where the architecture is characterized differently. By scaling through memory and selective activation, DeepSeek reduced its dependence on extremely fast silicon and adapted its system to the hardware constraints it faced.
Q: Can DeepSeek R1 run locally on a laptop?
Smaller distilled versions of DeepSeek R1 can run locally on laptops, according to the discussion. However, those versions are not the full primary model, and the main version is described as difficult to operate locally. Local execution can address some concerns about sending information to a Chinese-hosted service, but the speakers caution that most users are unlikely to deploy the model themselves.
Q: Why did hardware restrictions potentially help DeepSeek innovate?
The speakers argue that restrictions on access to top Nvidia chips created evolutionary pressure for Chinese teams to accomplish more with less. Without the option to solve every problem by adding large numbers of faster GPUs, DeepSeek focused on data quality, algorithms, memory capacity, sparse activation, and low-level engineering. The discussion presents these constraints as a practical example of necessity encouraging technical invention and efficiency.
Q: What privacy concerns apply to using DeepSeek?
The discussion raises uncertainty about what happens to information submitted to DeepSeek, comparing the concern with earlier corporate restrictions on ChatGPT and public questions surrounding TikTok. Users can reduce exposure by running smaller distilled versions locally, but these are not the full model and most people may not use them. The transcript does not resolve the data-handling question or specify DeepSeek's privacy practices.
Summary & Key Takeaways
-
DeepSeek progressed from replicating Meta's Llama to releasing DeepSeek Coder, V3, and the R1 reasoning model. V3 reportedly matched GPT-4o at a $6 million training cost, while R1 added longer reasoning and competitive performance. The guests argue that this progression was expected by people following the company closely.
-
R1 attracted broad attention because several qualities arrived together: visible reasoning, strong benchmark results, lower pricing, open availability, and smaller versions that could run on laptops. The visible problem-solving process made the system feel more immediate and understandable than a model that simply displays a thinking notice before returning its final response.
-
DeepSeek reportedly trained with about 2,000 H800 chips for the relevant run and compensated for slower interconnects through low-level engineering. Its 640 billion parameter architecture activates only about 30 billion parameters at once. The discussion presents hardware constraints as pressure that encouraged improvements in memory use, data, algorithms, and overall efficiency.
Read in Other Languages (beta)
Share This Summary 📚
Summarize YouTube Videos and Get Video Transcripts with 1-Click
Try YouTube Summary with ChatGPT & Claude or YouTube Transcript Generator
Explore More Summaries from Peter H. Diamandis 📚






Summarize YouTube Videos and Get Video Transcripts with 1-Click
Try YouTube Summary with ChatGPT & Claude or YouTube Transcript Generator