Ronny Kohavi on A/B Testing and Experimentation Culture

76.2K views
•
July 27, 2023
by
Lenny's Podcast
YouTube video player
Ronny Kohavi on A/B Testing and Experimentation Culture

TL;DR

Any code change or feature should run inside an experiment, because even small bug fixes can have surprising, outsized impact. At Bing, simply moving an ad's second line to the first line lifted revenue about 12% (worth roughly $100 million) without hurting user metrics. Big wins are rare, so expect roughly 80% of bold ideas to fail.

Transcript

I'm very clear that I'm a big fan of test everything which is any code change that you make any feature that you introduce has to be in some experiment because again I've observed this sort of surprising result that even small bug fixes even small changes can sometimes have surprising unexpected impact and so I don't think it's possible to experime... Read More

Key Insights

  • Testing everything is the safest default: every code change or new feature should sit inside an experiment, because even small bug fixes and minor tweaks can produce surprising, unexpected impact that intuition fails to predict.
  • The largest revenue win in Bing's history came from a trivial change: moving an ad's second line up to the first line, enlarging the title font. It raised revenue about 12%, worth around $100 million at the time.
  • Increasing revenue without harming the user experience is the real test. Displaying more ads is a trivial 'theatrics' way to raise revenue, but the Bing title change was a home run because it did not significantly hurt guardrail metrics.
  • High-risk, high-reward ideas deserve allocation even though most fail. Kohavi cautions teams pursuing radical new designs to be ready to fail roughly 80% of the time while chasing occasional home runs.
  • Institutional memory is fragile: a winning experiment can be forgotten over time. Kohavi reintroduced the 'open listings in a new tab' idea at Airbnb years after it had been done and semi-forgotten, again producing big improvements.
  • Opening search results in a new tab was tested around 2008 for Hotmail and MSN, well before Airbnb, and proved highly beneficial despite heavy pushback from designers who argued users never asked for it.
  • Humans are consistently bad at predicting experiment outcomes. Ideas sit low on backlogs for months while trivial-to-implement changes, cheap in engineering time, turn out to deliver the biggest return on investment.
  • Most gains come inch by inch, not from gold nuggets. Bing ads improved revenue per thousand searches through many small monthly improvements, with occasional degradations, rather than rare massive one-off wins.

Install to Summarize YouTube Videos and Get Transcripts

Explore YouTube Video Summarizer or Get YouTube Transcript Extractor

Questions & Answers

Q: What was the most surprising A/B test result Ronny Kohavi has seen?

The most surprising public example was at Bing, where someone proposed moving the second line of an ad up to become the first line, making the title line larger. This trivial change, which sat on the backlog for months rated below other ideas, increased revenue by about 12%, worth roughly $100 million at the time, and did not significantly hurt user guardrail metrics.

Q: Why should companies test every code change and feature?

Kohavi is a strong advocate of testing everything, meaning any code change or new feature should run inside an experiment. He has repeatedly observed that even small bug fixes and minor changes can sometimes have surprising, unexpected impact. Because outcomes are so hard to predict, he does not believe it is possible to experiment too much, so experimentation guards against unintended effects.

Q: How often do teams find quick, high-impact experiment wins?

These 'gold nugget' wins, where an hour or a few days of work produces a massive result, are very rare and should not be expected. Kohavi shows several in chapter one of his book, but stresses most winnings are made inch by inch. Bing ads improved revenue per thousand searches through many small monthly gains, sometimes with degradations, rather than rare breakthroughs.

Q: Why did the Bing revenue experiment trigger a false alarm?

When the experiment launched, an alarm fired warning that something was wrong with the revenue metric. This same alarm had previously caught real mistakes like revenue being logged twice. The team's first reaction was that a 12% revenue jump was too good to be true, so they searched for a bug several times and replicated the experiment, but found nothing wrong.

Q: How do you increase revenue without hurting the user experience?

Kohavi notes it is very easy to increase revenue through 'theatrics' like displaying more ads, but that hurts the user experience, which their experiments have shown. The Bing title-line change was considered a home run precisely because it improved revenue while not significantly hurting the guardrail metrics, making it a genuine win rather than a short-term revenue grab.

Q: What is the experiment about opening search results in a new tab?

Kohavi ran this experiment around 2008, first for opening Hotmail in the UK and then MSN search, well before Airbnb existed. Opening results in a new tab was highly beneficial despite heavy pushback from designers who noted users never asked for it. At Airbnb, listings had long opened in a new tab, and reintroducing the idea to newer designs produced big improvements.

Q: Why is institutional memory important for experimentation?

Winning experiments can be forgotten over time, a lesson Kohavi learned about institutional memory. Although opening listings in a new tab had been done at Airbnb for a long time, it was semi-forgotten and not applied to newly designed features. When he reintroduced it to the team, they saw big improvements again, showing why teams should document and remember their winners.

Q: How well can teams predict the outcomes of their experiments?

Teams are often humbled by how bad they are at predicting experiment outcomes. Ideas can sit low on a backlog for months while trivial-to-implement changes deliver the biggest return on investment. The Bing title change was rated below many other backlog items yet became the biggest revenue win in Bing's history, showing intuition frequently fails to rank ideas correctly.

Summary & Key Takeaways

  • Ronny Kohavi, an experimentation expert who led teams at Airbnb, Microsoft, and Amazon, argues that every code change or feature should run inside an experiment, because even small bug fixes can produce surprising and unexpected impact that is impossible to predict in advance.

  • The biggest revenue win in Bing's history was a trivial change: moving an ad's second line to the first line, enlarging the title. It raised revenue about 12%, worth roughly $100 million, and triggered a false alarm because the result looked too good to be true.

  • Crucially, that Bing change did not significantly hurt user guardrail metrics, unlike showing more ads. Similar surprising wins, like opening search results in a new tab tested around 2008, show teams are humbled by how poorly they predict outcomes, with most gains earned inch by inch.


Read in Other Languages (beta)

Share This Summary 📚

Summarize YouTube Videos and Get Video Transcripts with 1-Click

Download browser extensions on:

Try YouTube Summary with ChatGPT & Claude or YouTube Transcript Generator

Explore More Summaries from Lenny's Podcast 📚

Summarize YouTube Videos and Get Video Transcripts with 1-Click

Download browser extensions on:

Try YouTube Summary with ChatGPT & Claude or YouTube Transcript Generator