What are RL environments and how do they train AI to use real apps

52.6K views
•
August 12, 2026
by
Sequoia Capital
YouTube video player
What are RL environments and how do they train AI to use real apps

TL;DR

RL environments are built from three components: worlds, apps, and tasks. Humans remain essential for measuring the frontier with rubrics and verifiers, enabling high quality data. Post-training results show meaningful gains on domain specific tasks when using expert-built environments and data from Mercor’s ecosystem.

Transcript

Um, Record, I think you guys grew from a 1 to a 2 billion dollar revenue run rate in the last 4 months or so. Um, so this company's off to the races and I think you were just so front and center to how companies are thinking about uh post training their own models, uh building their own intelligence. So, thank you for joining us for this conversati... Read More

Key Insights

  • RL environments consist of three parts: worlds, apps, and tasks, which together simulate real project contexts.
  • The worlds include messages, docs, and files that mirror a real company environment for agents to interact with.
  • Apps are high fidelity clones of tools like Salesforce and Google Workspace that agents use via CLI or APIs.
  • Tasks are prompts paired with rubrics or unit tests that serve as evaluators for training and evaluation.
  • Humans are essential to define rubrics and verify results, acting like professors grading essays or TAs evaluating slides.
  • Verifiers and leaderboards enable objective comparison of model trajectories across tasks and domains.
  • Post-training results on Apex Agents show significant improvements in specific domains with modest compute investments.
  • The data strategy shifts from crowdsourcing to expert-built environments and frontier lab collaboration to scale intelligence.

Explore YouTube Video Summarizer or Get YouTube Transcript Extractor

Questions & Answers

Q: What is a RL environment and what are its three core parts

A reinforcement learning environment in this context is a framework where AI agents interact with structured worlds, apps, and tasks to learn and be evaluated. The three core parts are worlds, which contain the data and context of a project; apps, which are high fidelity clones of common tools; and tasks, which are prompts paired with verifiers to drive training and assessment. This structure lets agents practice real-world workflows and receive rubric-based feedback to improve performance across complex domains.

Q: Why are humans still needed to measure the frontier in AI training

Humans are needed because many domains involve nuanced judgment and complex reasoning that are difficult to simulate. Humans create rubrics and verifiers that capture a full solution space, including multiple potential paths and common mistakes. This human oversight prevents reward hacking and ensures that model evaluations reflect meaningful, real-world competencies beyond simple calculations or surface-level accuracy.

Q: What role do verifiers play in RL environments

Verifiers serve as objective evaluators for model outputs. They can be rubrics or unit tests that define which behaviors are correct and how to score trajectories. By codifying what constitutes a high-quality response, verifiers enable consistent scoring across many tasks and models, reducing subjective bias and helping to align model training with real-world requirements.

Q: What is post-training performance like on Apex Agents

Post-training performance on Apex Agents shows notable gains across tasks when using the data and prompts derived from the RL environment. The example cited demonstrates improvements in domain-specific capabilities, such as corporate law tasks, after training with this framework. The gains are achieved with a feasible amount of compute and by leveraging the scalably produced task data.

Q: How do companies buy data for RL environments

There are three common data buying models: task-based, where companies pay per complex task that may take hours to days; off-the-shelf data, where repositories of ready-to-use data sets are sold to multiple customers; and custom per-task pricing for highly specialized or large-scale needs. This mix supports scalable data acquisition while tailoring to specific domains.

Q: What is meant by the shift from crowdsourcing to agentic data

The shift from crowdsourcing to agentic data means moving away from low-skilled, generalized labeling to obtaining data from highly skilled experts who can collaborate in teams. These experts create frontier evaluations, RL environments, and domain-specific data that better measure and improve model capabilities, enabling more reliable and targeted AI advancement.

Q: How do RL environments help with real-world apps like Google Workspace

RL environments include high fidelity app clones such as Google Workspace that agents can interact with via dedicated interfaces. This enables agents to practice real workflows, such as document handling or data room interactions, in a controlled but realistic setting. By training with these authentic apps, models learn to perform complex tasks more effectively in real-world work contexts.

Q: What is meant by frontier intelligence in this talk

Frontier intelligence refers to advancing model capabilities toward the leading edge of what is possible in AI. It involves creating and evaluating highly capable models through sophisticated environments, expert data, and rigorous verifiers. The goal is to enable organizations to own more of their intelligence and apply it directly to their own domains and workflows.

Summary & Key Takeaways

  • RL environments are three-part systems that include worlds, apps, and tasks to simulate real work contexts.

  • The human-in-the-loop is critical for creating robust rubrics and verifiers that guide model evaluation and training.

  • Post-training data and leaderboards demonstrate improving performance across complex domains using expert-built environments.


Read in Other Languages (beta)

Share This Summary 📚

Explore More Summaries from Sequoia Capital 📚