How to scale multimodal photo enhancement at Uber Eats?

TL;DR
The video explains a production workflow for multimodal image editing at Uber Eats that preserves authenticity while improving quality. It uses an image understanding and routing stage to decide whether to enhance and a multi-step editing loop with QA and logging. Continuous tuning uses human labels as gold and a closed feedback loop to handle drift.
Transcript
[music] >> My name is Jay and I'm here with Sonya. We are part of the computer vision team at Aruba. We're going to talk to you about a real world production use. Oh, my son done. Okay. Try again. Okay. Don't worry. I'll I'll manage. You hear me now? Okay, so we're going to talk to you today about a real world production use case and specifically w... Read More
Key Insights
- Image understanding plus routing determine whether to enhance or keep the original image.
- The router uses a rubric and a pass/fail decision to guide downstream processing.
- A structured JSON logging approach makes end-to-end auditing and diagnosis possible.
- Guardrails focus on safety, trust, and marketplace diversity to avoid homogenization.
- Human labels are used as the gold standard to align models before shipping.
- Drift handling is essential; offline models must adapt online to avoid stale outputs.
- A feedback loop includes an umbrella diagnosis agent that triggers auto-tuning when mismatches occur.
- Continuous learning loops require regular sampling of production data and re-training based on golden data comparisons.
Install to Summarize YouTube Videos and Get Transcripts
Explore YouTube Video Summarizer or Get YouTube Transcript Extractor
Questions & Answers
Q: What is the primary goal of the routing and image editing pipeline in Uber Eats' multimodal system?
The primary goal is to preserve authenticity and trust while improving quality selectively and safely at scale. The routing step uses multimodal input to decide whether an image should be enhanced or left as is, based on a rubric and recall guardrails. This prevents deterioration of a merchant’s brand and avoids adding unnecessary edits that could misrepresent the dish.
Q: How does the system handle drift and keep models aligned with real world data?
The system continuously samples production data, labels it with objective guidelines, and compares agent outputs to the golden labels. If a mismatch is detected, an umbrella diagnosis agent localizes the issue and triggers an auto-tuning pipeline. The tuned model is benchmarked against the golden dataset before being shipped again, ensuring alignment with real world data.
Q: Why is logging emphasized so strongly in this Uber Eats workflow?
Logging is emphasized because it provides a transparent, auditable record of all agent actions and decisions. A flat JSON structure across the end-to-end orchestration enables anyone from engineers to non technical staff to diagnose cases and roll up insights for aggregates. It also creates a foundation for the self learning loop.
Q: What challenges arise when enhancing high quality images, and how are they addressed?
Enhancing high quality images can waste compute or degrade the image if prompts introduce hallucinations. The system uses guardrails to avoid unnecessary enhancement when it would not improve quality. It also accounts for fidelity by ensuring the output remains faithful to the actual dish and branding, thus preventing misleading edits.
Q: What role do human labelers play in this multimodal setup?
Human labelers provide the golden truth for model alignment. They label a representative dataset across geographies, dish types, and image quality. This labeled data is used to tune agents, assess performance against guardrails such as recall, and guide improvements in the routing and editing stages until production metrics meet thresholds.
Q: How is risk of homogenization avoided in a diverse marketplace like Uber Eats?
The approach avoids homogenization by preserving merchant branding and by not using a single generic prompt for all edits. The routing and editing pipeline are designed to maintain diversity in the marketplace, leveraging model constraints and evaluation criteria that safeguard authenticity while enabling selective quality improvements.
Q: What constitutes a successful production cycle for model deployment in this context?
A successful cycle means the model outputs meet guardrail metrics on the golden dataset, the system can handle drift through online tuning, and the final publish step passes post processing QA. It also requires comprehensive logging and a scalable loop for continuous improvement across multiple merchants and regions.
Q: What is the importance of the end-to-end JSON logging structure mentioned in the talk?
The end-to-end JSON logging structure is important because it enables clear traceability, diagnostics, and rollups for the entire agent pipeline. It supports non technical stakeholders to understand cases, supports aggregates analysis, and underpins the self learning loop by providing the data foundation for audits, retraining, and governance.
Summary & Key Takeaways
-
The talk presents a production design for end-to-end multimodal image enhancement in Uber Eats, emphasizing authenticity, brand preservation, and safety guardrails. It outlines staged processing from understanding to routing to editing and post-processing, plus logging for observability and learning. The system uses a feedback loop to improve over time while controlling costs.
-
A key point is that human labels are treated as the golden truth to align models, with a repeatable data collection and evaluation pipeline that guides model tuning and gatekeeping before publishing results.
-
Another takeaway is the emphasis on drift handling and continuous learning, where production data is periodically re-labeled and used to auto-tune components, ensuring the pipeline remains reliable as the marketplace image quality varies across geographies and merchants.
Read in Other Languages (beta)
Share This Summary 📚
Summarize YouTube Videos and Get Video Transcripts with 1-Click
Try YouTube Summary with ChatGPT & Claude or YouTube Transcript Generator
Explore More Summaries from AI Engineer 📚






Summarize YouTube Videos and Get Video Transcripts with 1-Click
Try YouTube Summary with ChatGPT & Claude or YouTube Transcript Generator