Why Meta Open-Sourced Llama 3.1 405B for Free

TL;DR
Meta released Llama 3.1 405B as an open model primarily to serve as a 'massive teacher' that improves smaller models through distillation and synthetic data generation. Meta's business model does not depend on selling the model directly, so open-sourcing drives maximal adoption, lets the community red-team and harden it, and turns Llama into an industry standard.
Transcript
you know if I was a Founder right now I would absolutely adopt open source um it forces me though to look at the engineering complexion of my or right and think like I I'm going to need people doing L Ops and and uh and you know things like you know F data fine tuning and and how to build Rag and things and apis there's plenty of apis that allow yo... Read More
Key Insights
- Llama 3.1 405B is described as a 'monster' model whose biggest value is acting as a 'massive teacher' for other models, since a big model can be used to improve small models through distillation, which is how the 8B and 70B became strong models.
- Meta changed the Llama license so developers can use the model's outputs and data, resolving a long-standing community pain point where closed models forbade using outputs. Meta now actively encourages people to train on Llama-generated data.
- Zero-shot tool use is one of the most exciting capabilities, letting the model call tools like Wolfram, Brave search, or Google search, run code through a code interpreter, and support custom plugins for uses like Rag, all without task-specific training.
- Meta's rationale for open source is that its business model does not depend on selling the model directly; Meta has never been a cloud company and instead works through a partner ecosystem, so giving the model away costs it little strategically.
- Open sourcing improves security and quality because transparency lets academia and companies red-team and jailbreak the models, letting Meta find and fix issues faster, similar to how Linux and its open kernel become more secure when bugs surface publicly.
- The PyTorch experience shaped Meta's open ethos: publishing PyTorch created a bridge to outside innovation, letting Meta pull community architectures and open-sourced models internally, evaluate them, and see capabilities improve week over week and month over month.
- Partners including Nvidia and AWS began building distillation recipes and synthetic data generation services on top of the release, enabling developers to create specialized models from Llama's high-quality data, which Meta knows is good because it improves its own smaller models.
- Meta believes there is room for both open and closed models, comparing the split to Linux versus Windows where people choose based on their needs, and expects a world of open models and closed models to coexist without one having to win.
Install to Summarize YouTube Videos and Get Transcripts
Explore YouTube Video Summarizer or Get YouTube Transcript Extractor
Questions & Answers
Q: What makes Llama 3.1 405B unique compared to other state-of-the-art models?
According to Spisak, the 405B is a 'monster' and a great model, but its biggest differentiator is being a 'massive teacher' for other models. Because it is so large, it can be used to improve small models through distillation, which is exactly how the 8B and 70B models became strong. It also delivers long context, expanded multilingual support, and zero-shot tool use, all capabilities the community requested.
Q: Why did Meta open source Llama 3.1 405B for free?
Spisak explains that Meta's business model does not depend on this model to make money directly, since Meta has never been a cloud company or sold a cloud service. Open sourcing follows the same ethos as PyTorch, building a bridge to outside innovation so the world builds on Meta's technology. It also drives maximal adoption, helps make Llama a standard, and lets the community harden the models.
Q: How does the 405B model help improve smaller models?
The 405B acts as a large teacher model. Spisak says that when you have a big model you can use it for distillation to improve small models, and that this is how the 8B and 70B became the great models they are. Additionally, partners like Nvidia and AWS built distillation recipes and synthetic data generation services, letting developers create specialized models from Llama's high-quality data, which Meta knows is good because it improves its own smaller models.
Q: What is zero-shot tool use in Llama 3.1 and why does it matter?
Zero-shot tool use lets the model call external tools without task-specific training. Spisak shows examples like calling Wolfram, Brave search, or Google search, and says it works great. It also enables calling a code interpreter to actually run code and building custom plugins for things like Rag. He calls it a potential game changer for the community, making these capabilities state-of-the-art and broadly accessible.
Q: Why did Meta change the Llama license, and what did it unlock?
Meta changed the license so developers can use the model's outputs and data, which had been a pain point since closed models forbade using outputs or only allowed it unscrupulously. Spisak says this was a big decision discussed in many meetings with Mark. The goal was to unlock new capabilities, remove artificial barriers, drive maximal adoption, and enable partners to build distillation and synthetic data generation services on top of the release.
Q: How does open sourcing improve the security and quality of Meta's models?
Spisak argues transparency makes models better and more secure, comparing it to Linux and its open kernel where bugs can be found and pushed faster. When academia and companies red-team or try to jailbreak the models, Meta wants them to do so because it lets Meta improve. This mirrors PyTorch, where Meta watched the community improve things week over week and month over month and pulled those improvements internally.
Q: Should founders adopt open source models according to Joe Spisak?
Spisak says if he were a founder right now he would absolutely adopt open source. He notes it forces founders to look at the engineering complexion of their organization, needing people for LLM Ops, fine-tuning, building Rag, and APIs. Ultimately he argues you want control, because your moat is your data and your interaction with users, which open source lets you own rather than depend on closed APIs.
Q: How does Meta view the coexistence of open and closed AI models?
Spisak believes there is room for both open and closed models. He compares the situation to Linux and Windows, where people use whichever fits their needs and applications. He expects a world of open models and a world of closed models, and considers that outcome totally fine. Meta does not want the environment to become completely closed, which is part of the motivation for releasing Llama openly.
Summary & Key Takeaways
-
Speaking two days after the Llama 3.1 405B launch, Meta's Joe Spisak calls the 405B a 'monster' model. Its biggest strength is serving as a massive teacher for smaller models through distillation, the same process that made the 8B and 70B models strong performers.
-
Meta prioritized capabilities the community asked for: much longer context, expanded multilingual support with more languages coming, heavy post-training and safety work rather than just pre-training on data, and zero-shot tool use that Spisak expects to be a game changer for calling search, code, and plugins.
-
Meta changed the license so outputs and data can be freely used, a decision debated in meetings with Mark Zuckerberg. The goal is maximal adoption and making Llama a standard, removing artificial barriers, and enabling partners like Nvidia and AWS to build distillation and synthetic data services.
Read in Other Languages (beta)
Share This Summary 📚
Summarize YouTube Videos and Get Video Transcripts with 1-Click
Try YouTube Summary with ChatGPT & Claude or YouTube Transcript Generator
Explore More Summaries from Sequoia Capital 📚






Summarize YouTube Videos and Get Video Transcripts with 1-Click
Try YouTube Summary with ChatGPT & Claude or YouTube Transcript Generator