How Google's Nano Banana Achieves Character Consistency

367.7K views
•
November 11, 2025
by
Sequoia Capital
YouTube video player
How Google's Nano Banana Achieves Character Consistency

TL;DR

Nano Banana achieves single-image character consistency through the right combination of model architecture and high-quality data, not scale alone. Its team relied heavily on human evaluations, including team members judging outputs against their own faces, because identity preservation is subtle and hard to quantify. The goal was set from the start to close a known gap in earlier image models.

Transcript

There's something about like visual media that really excites people that it's like the fun thing, but it's not just fun. It's exciting. It's intuitive. The visual space is so much of how we as humans experience uh experience life that I think I've loved how much it's moved people. I think we're really now making it possible to like tell stories th... Read More

Key Insights

  • Character consistency was an explicit goal from the beginning, not an emergent property of scale, because the team knew it was a gap in the image models they had released in the past.
  • Human evaluations are a foundational game changer for image generation, since qualities like faces, aesthetic quality, and subtle details are very hard to quantify with automated metrics alone.
  • Judging face consistency reliably requires people to evaluate outputs against their own faces, so the team built evals using team members' own faces because you can only truly judge whether an image looks like you on yourself.
  • The right recipe for character consistency combined both model architecture and data quality, and the team was surprised by just how good the model turned out once training finished and they actually used it.
  • Preserving parts of an image you are not editing is shockingly technically difficult, even though lay users expect editing to leave untouched areas unchanged like Photoshop or phone apps.
  • Nano Banana started as a 2 a.m. code name and became a cultural phenomenon, with a common use being turning yourself into a 3D figurine complete with a toy box and computer.
  • Users hacked the model in unanticipated ways, including creating coherent sketch notes from technical lectures despite text rendering not being where the team wants it to be.
  • Advertisers demanded character consistency for years because products placed in lifestyle shots must look 100% like the real product, or they cannot be used in an ad.

Install to Summarize YouTube Videos and Get Transcripts

Explore YouTube Video Summarizer or Get YouTube Transcript Extractor

Questions & Answers

Q: How did Google achieve character consistency in Nano Banana?

The team achieved it through the right recipe combining model architecture and data, not scale alone. There are different genres of ways to do image generation, and that choice plays a part in how good it is. Character consistency was an explicit goal from the beginning because it was a known gap in earlier models, and the team was surprised by how good it turned out once the model was actually trained and used.

Q: Is character consistency an emergent property of scale?

No. According to the team, character consistency is not just an emergent property of scale. It was an explicit goal set from the beginning of the model's development because they knew it was a gap in the image models they had previously released. Achieving it required the right combination of model architecture and data quality, which they described as having the right recipe both in terms of model design and the underlying data.

Q: Why are human evaluations important for image generation models?

Human evaluations are foundational because qualities like faces, aesthetic quality, and subtle visual details are very hard to quantify with automated metrics. The team has a dedicated group that builds tooling and practices for having humans evaluate these subtle outputs. Human evals were described as a big game changer, especially for judging things like facial character consistency that are difficult to measure numerically.

Q: How does the team evaluate whether a generated face looks correct?

The team runs evals using their own team members' faces, because you can really only judge character consistency on faces you know well. An AI version of a stranger might look acceptable to you, but the person themselves would notice parts of their face are not quite right. By evaluating the model's output on familiar faces like colleagues they see often, the team can reliably judge identity preservation.

Q: What was the aha moment during Nano Banana's development?

One team member took a photo of himself and prompted the model to put him on the red carpet in full glam, a vanity prompt. The result actually looked like him, which no previous model had achieved. Comparing it to earlier models made the leap clear. It then took a couple of weeks for other colleagues to try their own photos and realize how magical accurate identity preservation was.

Q: Why is preserving unedited parts of an image so difficult?

When people edit images in phone apps or Photoshop, they expect a high degree of preservation of the things they are not touching. Depending on how models are designed and the decisions behind them, keeping untouched areas unchanged is shockingly technically difficult. It is something a lay person assumes is basic to editing, that you do not mess with things you do not want changed, but it is actually very hard for models to do well.

Q: How are people using Nano Banana in unexpected ways?

Users have combined it with video models to get consistent characters and scenes across scene cuts, producing smoother, more natural video workflows. One person created coherent sketch notes from technical university chemistry lectures, feeding material into Gemini with Nano Banana to finally understand and discuss his father's work. Others turn themselves into 3D figurines with toy boxes, using it to express and enhance their own identity.

Q: Why do advertisers need character consistency in AI image models?

The team heard demand from advertisers for years because products placed in lifestyle shots must look exactly like the real product, 100%, or they cannot be used in an advertisement. Prior models were not good at preserving parts of an image while changing others, which made them unusable in professional workflows. This gap in consistency confirmed there was real demand and motivated the team to finally close it.

Summary & Key Takeaways

  • Nano Banana, Google's image model built by Nicole Brichtova and Hansa Srinivasan, achieved breakthrough single-image character consistency. It started as a 2 a.m. code name and became a cultural phenomenon, letting people see themselves in AI-generated worlds and tell visual stories they could not before.

  • Character consistency was a deliberate goal from the start because earlier models had a clear gap in preserving identity and untouched image regions. The team credits a combination of model architecture and high-quality data, and was surprised by how well it worked once the model was actually built and used.

  • Human evaluation is central to the model's quality, since faces and aesthetic quality resist quantification. Team members eval outputs against their own faces because identity is only truly judgeable on yourself. Users have creatively extended the model into video workflows, learning, and personal storytelling.


Read in Other Languages (beta)

Share This Summary 📚

Summarize YouTube Videos and Get Video Transcripts with 1-Click

Download browser extensions on:

Try YouTube Summary with ChatGPT & Claude or YouTube Transcript Generator

Explore More Summaries from Sequoia Capital 📚

Summarize YouTube Videos and Get Video Transcripts with 1-Click

Download browser extensions on:

Try YouTube Summary with ChatGPT & Claude or YouTube Transcript Generator