How to Build an AI Research Lab on a Budget

28.9K views
•
August 11, 2026
by
Sequoia Capital
YouTube video player
How to Build an AI Research Lab on a Budget

TL;DR

Application companies can compete with frontier labs by using the broader model ecosystem instead of reproducing every capability internally. The practical sequence is to build reliable domain benchmarks, generate synthetic data under expert guidance, post-train competitive open-source models with specialized partners, and establish evaluation, routing, fallback, and monitoring infrastructure before deploying those models in production.

Transcript

Next up, we are honored to have Gabe with us. Gabe is co-founder and president of Harvey. Um you were a research scientist at DeepMind ages ago, um and then at Meta I think before you showed your college roommates uh what GPT-3 could do. And this duo uh then became Harvey. Um we're really excited to have you here. I think everybody here in the audi... Read More

Key Insights

  • Benchmarks are the foundation of domain model development because they determine whether training produces meaningful improvement. Harvey begins with taxonomies and datasets representing real legal work, including drafting fund formation documents, researching case law, negotiating contracts, and conducting diligence across large data rooms.
  • Sensitive customer information is not required to begin specialized model training. Harvey cannot train generic or internal models on privileged legal data from law firms and enterprises, so its domain experts guide synthetic data generation to create realistic material for product evaluation and model development.
  • Domain experts can guide coding models to produce realistic synthetic datasets. Harvey's lawyers developed skill in directing these systems and trained other lawyers to follow the approach, combining professional judgment with model generation before outside providers helped scale the process into larger training sets.
  • Synthetic data is a starting point rather than a complete solution. Harvey uses expert guidance to establish realistic datasets, then works with companies such as Mercor and Snorkel to expand the process, particularly when larger quantities of training data are required.
  • Efficient reinforcement-learning environments are necessary when evaluation becomes large and expensive. Harvey's diligence dataset includes more than 1,000 unit tests that grade model outputs with an LLM judge, so reducing evaluation costs is important when running repeated rollouts over complex, long-context tasks.
  • Open-sourcing benchmarks can improve their quality by exposing them to broad training and evaluation. Harvey receives pull requests and suggestions that reveal issues, while model labs increasingly report performance on its datasets, making the benchmarks more useful for comparing new systems.
  • Multiple specialized model labs expand research capacity and expose an application company to different technical bets. Harvey collaborates with several providers because it has more projects than internal teams or one partner can handle, and each lab brings distinct infrastructure, models, recipes, and research perspectives.
  • Production readiness depends on layered evaluation and monitoring rather than benchmark scores alone. Harvey combines generic evaluations, human comparisons, product-specific tests, cost, latency, and regional availability before release, then tracks experiments, engagement, uptime, token efficiency, and customer feedback after deployment.

Install to Summarize YouTube Videos and Get Transcripts

Explore YouTube Video Summarizer or Get YouTube Transcript Extractor

Questions & Answers

Q: How can an application company compete with frontier AI labs?

An application company can compete by using the frontier ecosystem rather than attempting to reproduce every advantage of a frontier lab. Harvey builds specialized benchmarks and expert-guided synthetic data, collaborates with multiple labs for post-training, and uses open-source and closed-source models together. It then applies product-specific evaluation, routing, fallbacks, and production monitoring to deliver strong intelligence for focused legal tasks.

Q: Why should an AI research lab build benchmarks before training models?

A benchmark provides the measurement needed to determine whether training improves a model on the intended tasks. Harvey treats this as the first step because training cannot be directed or assessed reliably without it. Its benchmarks represent complex legal workflows across practice areas, including drafting, research, negotiation, and diligence, creating a concrete target for model development and product evaluation.

Q: How can companies train models without using sensitive customer data?

Companies can ask domain experts to guide the creation of realistic synthetic data. Harvey cannot train on privileged legal information belonging to major law firms and enterprises, including within its own models. Its lawyers instead direct coding models to generate representative datasets for evaluation and training. External providers can then help scale that expert-led process when larger training sets are needed.

Q: Why did Harvey open-source some of its legal datasets?

Harvey open-sourced datasets because broad use helps reveal whether a benchmark is genuinely useful and where it contains problems. Researchers and developers can train against the data, submit pull requests, and recommend improvements. Model labs also increasingly evaluate new releases on Harvey's benchmarks, which strengthens their role as shared measures and provides more evidence about model performance on legal tasks.

Q: When is post-training an open-source model worthwhile?

Post-training becomes worthwhile when the base model is strong enough that specialized training can make it competitive on a defined task. Previously, rapid improvements from pre-training could absorb the value of application-level post-training. Harvey argues that current open-source base models can now reach frontier-level performance in narrow domains, even if they do not achieve general frontier intelligence across every task.

Q: Why does Harvey work with multiple specialized model labs?

Harvey works with multiple labs because its research portfolio exceeds the bandwidth of its internal team or any single partner. Different labs pursue different research ideas, support different open-source models, and provide distinct infrastructure and training recipes. Working across these partners lets Harvey run more projects, learn from varied approaches, validate its datasets, and gradually strengthen its own post-training capabilities.

Q: How does Harvey decide whether a new model enters production?

Harvey combines generic and product-specific evidence before deployment. Automated legal benchmarks indicate overall strength and relevant practice areas, while human side-by-side testing compares the candidate with other models. Each product surface also has critical user journeys, automated tests, and human testing. Cost, latency, and regional availability are added to these signals before a production decision is made.

Q: How does Harvey monitor AI models after deployment?

Harvey continues evaluating models after release because pre-production performance does not guarantee reliable user outcomes. For major changes, it uses A/B testing and examines engagement to determine whether behavior matches expectations. Operational and efficiency signals include uptime and token efficiency, while direct product feedback, including dissatisfied customer messages, helps expose failures that automated evaluations or aggregate metrics may miss.

Summary & Key Takeaways

  • Harvey's research strategy begins with domain-specific benchmarks because models cannot be improved systematically without reliable measurements. Its datasets cover law-firm tasks, contract negotiation, and diligence. Since customer legal information is sensitive and privileged, Harvey asks lawyers to guide synthetic data generation, then works with outside companies to scale the resulting training sets.

  • Open-source models have become strong enough for task-specific post-training to produce competitive intelligence. Harvey works with multiple specialized labs because each offers different research approaches, infrastructure, recipes, and model expertise. These partnerships help validate training data, expand research capacity, and transfer knowledge that supports Harvey's growing ability to conduct post-training internally.

  • Production deployment requires more than a successful training run. Harvey serves multiple model families across 60 countries, product surfaces, regional requirements, and customer preferences. It evaluates candidate models with automated benchmarks, human comparisons, critical user journeys, product tests, cost, latency, and availability, then monitors live performance through experiments, engagement, uptime, token efficiency, and customer feedback.


Read in Other Languages (beta)

Share This Summary 📚

Explore More Summaries from Sequoia Capital 📚