What's New in AI This Week: Google, OpenAI & Open Source

TL;DR
Google shipped the most this week, launching Firebase Studio for AI coding, its sixth-gen Ironwood TPU with 192GB RAM per chip for inferencing, and public VEO 2 access via Gemini and API. Meta's Llama 4 Scout and Maverick landed to a lukewarm reception since neither runs on consumer hardware, while open-source image models HiDream and a Flux-based stylizer arrived free on Hugging Face.
Transcript
This week has been absolutely colossal with AI news. I mean, so many different releases from big names in the tech space, little open- source tidbits, cool demos, you name it, we got it today. These are the kinds of weeks that I like to see in the AI world. So, let's just jump right in. First up, quick little recap. Llama dropped from Meta. I did a... Read More
Key Insights
- Meta's Llama 4 released two models, Scout and Maverick, both with 16 billion active parameters, but Scout uses 16 experts while Maverick uses 128 experts. Neither runs on consumer-grade hardware, as both are aimed at corporations and businesses.
- HiDream AI is a brand new MIT-licensed image generation model with full, dev, and fast variants. Its benchmarks reportedly beat Dolly 3, SDXL, and Flux, though not by much, and it requires significant VRAM to run.
- Google's Ironwood is its sixth-generation TPU built for AI inferencing, offering 192GB of RAM per chip and 4.5 times faster data access, positioned as a cheaper alternative to Nvidia GPUs.
- Firebase Studio is Google's AI vibe-coding platform that automates typical coding processes, but it runs on a weaker Gemini model rather than 2.5 Pro, leading to less than satisfactory results in early testing.
- Google's VEO 2 is now public and usable directly in Gemini, though image uploading is not supported there. The API adds inpainting, outpainting, camera presets like panning, and first and last frame control.
- Gen 4 Turbo is now available at five times faster and half the cost of the original Gen 4. It trades some quality and prompt coherence for rapid ideation and struggles with large non-human motion and complex prompts.
- A new research paper, one-minute video generation with test-time training, produces coherent full one-minute Tom and Jerry cartoons with characters interacting logically, and could potentially extend to five or ten minutes.
- 11 Labs launched a new MCP server that gives Claude and Cursor access to its entire AI audio platform through simple text prompts, enabling use cases like spinning up voice agents to perform outbound calls.
Install to Summarize YouTube Videos and Get Transcripts
Explore YouTube Video Summarizer or Get YouTube Transcript Extractor
Questions & Answers
Q: What are the differences between Llama 4 Scout and Maverick?
Llama 4 Scout and Maverick both have 16 billion active parameters, but they differ in their mixture of experts. Scout only has 16 experts, while Maverick comprises 128 experts. Neither model, including the smaller Scout, runs on consumer-grade hardware. Both are meant for corporations and businesses to run. They remain open source but do not have the greatest license, and the overall release was considered pretty meh compared to previous Llama releases.
Q: What is HiDream AI and where can you test it?
HiDream AI is a brand new open-source image generation model that comes in three variants: full quality, dev quality, and fast quality. The model itself is MIT licensed, though its text encoder is Llama 3 base, which still maintains the Llama 3 license. Its benchmarks reportedly beat Dolly 3, SDXL, and Flux, though not by much. It requires a significant amount of VRAM to run, but you can test all variants for free on Hugging Face, where a quantized version is used.
Q: What is Google's Ironwood TPU?
Ironwood is Google's sixth-generation TPU built specifically for AI inferencing. It features 192GB of RAM per chip and offers 4.5 times faster data access. Big AI companies are looking for alternative, cheaper methods to Nvidia GPUs for AI inferencing, and Ironwood looks like a pretty solid competitor in that space. It reflects the broader industry push to reduce reliance on Nvidia hardware for running AI models at scale.
Q: How well does Google Firebase Studio work?
Firebase Studio is Google's AI vibe-coding style platform that aims to automate a lot of typical coding processes. It uses Gemini under the hood, but not the strong 2.5 Pro coding model, instead relying on a weaker model, which has led to less than satisfactory examples. Viewer comments reported it isn't holding its own too well, with one user's generated app not going well and another noting environment issues at launch. It looks promising but appears to be in early stages with issues to iron out.
Q: What can you do with VEO 2 now that it is public?
Google officially dropped VEO 2 publicly, and you can use it right inside Gemini, though image uploading to Gemini is not supported. In the API, VEO 2 offers inpainting and outpainting features, camera presets such as panning to the right, and first and last frame control. These capabilities are aimed more at professional team-level work, like creating commercials or stories, where larger groups manage the systems at a higher level. VEO 2 is regarded as a top-charting model that most others do not compete with.
Q: What is Gen 4 Turbo and how does it compare to Gen 4?
Gen 4 Turbo is a new AI video generation model that is five times faster and half the cost of the original Gen 4, which had its full release the prior week. It does not have the best quality or prompt coherence compared to Gen 4, but it allows for rapid ideation while keeping still-good quality. It struggles with large motion that isn't human and with more complex prompts, but it kept the creator impressively consistent across shots for being a turbo model.
Q: What does the one-minute video generation paper accomplish?
The paper, titled one-minute video generation with test-time training, generates full Tom and Jerry cartoons that are one minute long, with the possibility of extending to five or ten minutes. The stories are coherent and the characters interact in ways that make sense, such as Jerry unplugging a computer, the screen going fuzzy, and Tom getting angry and investigating, just like a classic episode. It was described as a really promising and underrated paper from the week.
Q: What updates did 11 Labs and other video tools announce?
11 Labs launched a new MCP server that gives Claude and Cursor access to its entire AI audio platform through simple text prompts, enabling use cases like spinning up voice agents for outbound calls. Higgsfield AI added the ability to combine multiple motion controls in a single shot, including moves impossible with real cameras, plus 10 new motion controls for speed and cinematic impact. LTX Studio added custom AI actors, letting users train characters from reference images to keep faces and styles consistent across shots.
Summary & Key Takeaways
-
This was a colossal week for AI news across major companies and open source. Meta's Llama 4 Scout and Maverick launched with 16 billion active parameters each but need corporate hardware, drawing a meh reception. Ply had already jailbroken Llama 4 shortly after release.
-
Google dominated the week, launching Firebase Studio for AI coding on a weaker Gemini model, the Ironwood sixth-gen inferencing TPU with 192GB RAM per chip, updated Imagen generation, Gemini 2.5 Flash, and public VEO 2 access through Gemini and a feature-rich API.
-
AI video generation surged with Gen 4 Turbo running five times faster at half the cost, Higgsfield AI's new combinable camera motion controls, and LTX Studio adding consistent custom AI actors. Open image models HiDream and a Flux-based stylizer arrived free, and 11 Labs shipped an MCP server.
Read in Other Languages (beta)
Share This Summary 📚
Summarize YouTube Videos and Get Video Transcripts with 1-Click
Try YouTube Summary with ChatGPT & Claude or YouTube Transcript Generator
Explore More Summaries from MattVidPro 📚






Summarize YouTube Videos and Get Video Transcripts with 1-Click
Try YouTube Summary with ChatGPT & Claude or YouTube Transcript Generator