What Changed in Granite 4, Claude 4.5, and Sora 2?

TL;DR
Granite 4 delivers smaller language models that use less memory, support long contexts, and fit on a single GPU while outperforming the largest Granite 3 model in the comparison discussed. The episode also presents Claude Sonnet 4.5 as a coding-focused release with longer run times and sharper reasoning, Sora 2 as a video-production model, and Buy in ChatGPT as a new commerce feature.
Transcript
Yeah, Sora 2 is, I think the Vibe video producing app. I mean the Claude is the Vibe coding we have. I mean vibe thinking, Vibe everything going on. I'm just Vibe living at this point. Exactly. Exactly right. All that and more on today's Mixture of Experts. I'm Tim Huang and welcome to Mixture of Experts. Each week MOE brings together a panel of cu... Read More
Key Insights
- Granite 4 is a family of efficient smaller language models designed for developer experimentation and enterprise deployment. The models were released on Hugging Face and can fit on a single GPU, including less expensive options such as an L40S or A100, according to the panel.
- Granite 4's hybrid architecture is designed to improve memory efficiency. The smallest model reportedly occupies roughly four gigabytes while running a 128k context length, yet it outperforms the largest Granite 3 model in the comparison presented during the discussion.
- Granite 4 is positioned around smarter use of compute rather than a race toward ever-larger models. The panel reports almost 72 percent less memory, half the model size or less, better performance, and contexts that can be four times longer or more.
- ISO 42001 certification is presented as evidence of the governance, safety, and security processes used to develop Granite 4. The panel describes Granite as one of the first, if not the first, open source model families on Hugging Face with that certification.
- Cryptographic signing is used to make Granite model development more verifiable. The described mechanism signs checkpoints created during training, then releases corresponding signature information so another party can check whether the published training history matches the claimed process.
- Enterprise efficiency is important because compute costs are rising, regulatory pressure is increasing, and customers care about total cost of ownership. The panel also argues that environmental costs make continued dependence on increasingly expensive model scaling difficult to sustain.
- Claude Sonnet 4.5 is presented as a more specialized release than earlier general-purpose model announcements. Its launch materials emphasize coding, while the episode description highlights longer run times and sharper reasoning as notable improvements associated with the model.
- Sora 2 is characterized as a vibe video-production model, extending the informal concept of vibe coding into visual creation. The episode places it alongside OpenAI's Buy in ChatGPT feature, which introduces an e-commerce dimension to interactions conducted through ChatGPT.
Install to Summarize YouTube Videos and Get Transcripts
Explore YouTube Video Summarizer or Get YouTube Transcript Extractor
Questions & Answers
Q: How is Granite 4 different from Granite 3?
Granite 4 reduces the memory footprint while improving performance in the comparison discussed by the panel. Even its smallest model, which takes roughly four gigabytes while running a 128k context length, reportedly outperforms the largest Granite 3 model. The new family also uses a hybrid architecture, fits on a single GPU, and places greater emphasis on governance and verifiable development practices.
Q: What hardware is required to run Granite 4?
Granite 4 is designed so every model in the announced family can fit on a single GPU. The panel specifically mentions options including an L40S and an A100, contrasting them with deployments that require eight H100 GPUs. This reduced hardware requirement is attributed to the models' smaller sizes and their hybrid architecture, which provides memory efficiencies during deployment.
Q: Why does Granite 4 use a hybrid architecture?
Granite 4 uses a hybrid architecture to gain memory efficiency while preserving strong model capability. The panel connects this design to Bamba and state-space models, arguing that transformer-based systems scale effectively but expensively. The architectural goal is therefore to deliver comparable results with less memory and compute, enabling long contexts and practical operation on consumer or lower-cost single-GPU systems.
Q: What governance measures are associated with Granite 4?
Granite 4 received ISO 42001 certification before its release, which the panel presents as evidence of governance, safety, and security throughout the AI model development system. IBM has also engaged with the Stanford Transparency Index team and introduced cryptographic signing for models. Together, these measures are intended to improve transparency, process accountability, and independent verification of model development.
Q: How does cryptographic signing help verify an AI model?
Cryptographic signing creates a way to verify checkpoints produced during model training. As described in the discussion, signatures are attached throughout the training process, and the relevant signature information is released with the model. Another party can then use it to check whether the training proceeded as represented, adding a layer of verification beyond simply trusting the final published model.
Q: Why are smaller AI models valuable to enterprises?
Smaller AI models can lower computing requirements, reduce deployment costs, and run on more accessible hardware. The panel argues that enterprises care about total cost of ownership more than leaderboard position alone. It also identifies rising compute costs, mounting regulatory pressure, and environmental costs as reasons to prioritize models that remain capable while using memory and processing resources more efficiently.
Q: What is notable about Claude Sonnet 4.5?
Claude Sonnet 4.5 is presented as a release with a particularly strong coding focus rather than a broad claim to excel equally across every task. The episode description also attributes longer run times and sharper reasoning to the model. This positioning reflects a shift toward describing a foundation model through a specific practical strength, in this case sustained and reasoning-intensive coding work.
Q: What topics beyond language models does the episode cover?
The episode covers Sora 2 as a vibe video-production model and Buy in ChatGPT as a new e-commerce feature. Its news roundup also mentions Meta using AI-assistant conversations to inform Facebook and Instagram ads, Microsoft's vibe working agents, DoorDash's Dot delivery robot, and Tilly, an AI-generated actress. A bonus cybersecurity segment asks whether users can trust their AI.
Summary & Key Takeaways
-
Granite 4 focuses on efficient, enterprise-ready language models that developers can download from Hugging Face, test, and deploy without extremely large computing systems. Its hybrid architecture reduces memory requirements, enables operation on a single GPU, and supports a 128k context length even in a model occupying roughly four gigabytes.
-
IBM positions Granite 4 around capability per unit of compute rather than model size alone. The panel says the smallest model outperforms the largest Granite 3 model while using a smaller memory footprint. Governance measures include ISO 42001 certification, model transparency work, and cryptographic signatures intended to support verification of training checkpoints.
-
The broader discussion covers several shifts in applied AI. Claude Sonnet 4.5 is presented as a coding-focused model with longer run times and sharper reasoning, while Sora 2 extends the idea of vibe creation into video production. The episode also examines Buy in ChatGPT, AI-assisted work, advertising, delivery robots, and AI security.
Read in Other Languages (beta)
Share This Summary 📚
Summarize YouTube Videos and Get Video Transcripts with 1-Click
Try YouTube Summary with ChatGPT & Claude or YouTube Transcript Generator
Explore More Summaries from IBM Technology 📚






Summarize YouTube Videos and Get Video Transcripts with 1-Click
Try YouTube Summary with ChatGPT & Claude or YouTube Transcript Generator