How Google's Nano Banana Achieves Character Consistency

367.7K views
•
November 11, 2025
by
Sequoia Capital
YouTube video player
How Google's Nano Banana Achieves Character Consistency

TL;DR

Google’s Nano Banana achieves character consistency by combining model architecture, high-quality data, long multimodal context windows, and disciplined human evaluation rather than relying on scale alone. The capability was a deliberate goal because earlier image models struggled to preserve identity and untouched regions. Team members tested it with their own faces, while users extended it to 3D figurines, video workflows, and sketch notes. Read on to see how the breakthrough was developed and evaluated.

Transcript

There's something about like visual media that really excites people that it's like the fun thing, but it's not just fun. It's exciting. It's intuitive. The visual space is so much of how we as humans experience uh experience life that I think I've loved how much it's moved people. I think we're really now making it possible to like tell stories th... Read More

Key Insights

  • Character consistency was an explicit goal from the beginning, not an emergent property of scale, because the team knew it was a gap in the image models they had released in the past.
  • Human evaluations are a foundational game changer for image generation, since qualities like faces, aesthetic quality, and subtle details are very hard to quantify with automated metrics alone.
  • Judging face consistency reliably requires people to evaluate outputs against their own faces, so the team built evals using team members' own faces because you can only truly judge whether an image looks like you on yourself.
  • The right recipe for character consistency combined both model architecture and data quality, and the team was surprised by just how good the model turned out once training finished and they actually used it.
  • Preserving parts of an image you are not editing is shockingly technically difficult, even though lay users expect editing to leave untouched areas unchanged like Photoshop or phone apps.
  • Nano Banana started as a 2 a.m. code name and became a cultural phenomenon, with a common use being turning yourself into a 3D figurine complete with a toy box and computer.
  • Users hacked the model in unanticipated ways, including creating coherent sketch notes from technical lectures despite text rendering not being where the team wants it to be.
  • Advertisers demanded character consistency for years because products placed in lifestyle shots must look 100% like the real product, or they cannot be used in an ad.

Explore YouTube Video Summarizer or Get YouTube Transcript Extractor

Questions & Answers

Q: How does Google’s Nano Banana achieve character consistency?

Nano Banana combines model architecture, high-quality data, long multimodal context windows, and disciplined human evaluations. Character consistency was an explicit development goal because the team recognized it as a gap in earlier image models, rather than treating it as an automatic result of scale.

Q: Is Nano Banana’s character consistency simply an emergent property of scale?

No. The team says character consistency was intentionally targeted from the beginning because previous image models had struggled with it. They credit the right combination of model design and data quality, with craft and infrastructure mattering as much as scale.

Q: Why are human evaluations important for Nano Banana?

Faces, aesthetic quality, and subtle visual details are difficult to quantify with automated metrics alone. Disciplined human evaluations therefore helped the team judge whether the model preserved identity and produced convincing results.

Q: How did the Nano Banana team evaluate facial consistency?

Team members evaluated outputs made from their own faces and the faces of familiar colleagues. A generated stranger may appear convincing to an observer, but the person depicted is more likely to notice when specific facial details do not look right.

Q: What was the breakthrough moment during Nano Banana’s development?

One team member submitted a photo and asked the model to place her on a red carpet in full glam. The result actually looked like her, unlike outputs from the earlier models she compared it with. It took a couple of weeks for other colleagues to test their own photos and recognize how significant that identity preservation was.

Q: Why is preserving unedited parts of an image difficult?

People expect an image editor to leave everything outside the requested change untouched, as familiar editing tools do. For image models, however, preserving those regions can be shockingly difficult depending on the model’s design and the decisions behind it.

Q: How are people using Nano Banana with video models?

Users combine Nano Banana with video models from different sources to preserve characters and scenes across cuts. The workflow is not necessarily fluid, but the team observed videos becoming much smoother, with scene transitions that felt more natural after Nano Banana launched.

Q: What unexpected uses have people found for Nano Banana?

Users have created coherent sketch notes from technical material, even though text rendering is not yet where the team wants it to be. One person fed university chemistry lectures into Gemini with Nano Banana to better understand his father’s work. Another popular use turns a person into a 3D figurine alongside a toy box and computer.

Summary & Key Takeaways

  • Nano Banana, Google's image model built by Nicole Brichtova and Hansa Srinivasan, achieved breakthrough single-image character consistency. It started as a 2 a.m. code name and became a cultural phenomenon, letting people see themselves in AI-generated worlds and tell visual stories they could not before.

  • Character consistency was a deliberate goal from the start because earlier models had a clear gap in preserving identity and untouched image regions. The team credits a combination of model architecture and high-quality data, and was surprised by how well it worked once the model was actually built and used.

  • Human evaluation is central to the model's quality, since faces and aesthetic quality resist quantification. Team members eval outputs against their own faces because identity is only truly judgeable on yourself. Users have creatively extended the model into video workflows, learning, and personal storytelling.


Read in Other Languages (beta)

Share This Summary 📚

Explore More Summaries from Sequoia Capital 📚