How Can AI Agents Learn From Every Interaction?

TL;DR
AI agents improve through continual learning when their real interactions, corrections, tool calls, and sub-agent traces become reusable learning signals. A practical system connects complete traceability to behavioral specifications, production-based evaluations, model training, and harness improvements, while deciding carefully whether each lesson belongs in model weights or in contextual tools and infrastructure.
Transcript
Okay, last talk to bring us home, we have Arjun and Ronak, co-founders of Trajectory. They're doing a bunch of interesting research around what I think is one of the hotter topics, continual learning. Please welcome Arjun to the stage. >> Cool. What's up? How's it going everyone? Thank you so much. I'm really excited. I know I'm close to the last o... Read More
Key Insights
- AI intelligence has two distinct axes: general capability and accumulated experience. Models may keep improving in capability while still acting like new employees because they do not reliably retain lessons from the real work, feedback, and corrections produced during previous uses.
- Agent interactions are potential learning signals rather than disposable outputs. The work agents perform, the actions people take afterward, and the corrections users make can support systems that become faster, better, and cheaper while improving repeatedly through continued use.
- Complete traceability requires capturing the entire execution tree. Recording only the main action while discarding tool calls or sub-agent activity removes information needed to understand results and learn from every component involved in producing them.
- Corrective behavior is often more informative than simple ratings. Thumbs-up and thumbs-down feedback can be noisy because users may accept an output initially, then discover several commits later that it caused a problem and respond through edits, retries, undos, or other corrections.
- Production activity should shape evaluation and training. The product experience, evaluation environment, and training environment should be as similar as possible, with tasks drawn from current traffic and frontier requests, replayed when possible, and graded through the actual production harness.
- Agent harnesses should expose useful primitives instead of prescribing rigid workflows. Search tools, private information, and other product capabilities can be made available for the agent to orchestrate, reflecting a shift from primarily preventing mistakes toward allowing stronger models greater flexibility.
- Agent tools should mirror the user interface as closely as possible. When every action available to a person through the interface also has a corresponding tool call, training and continual learning become easier because agents can operate through the same product capabilities.
- Tool responses should report what actually happened rather than merely saying an action finished. Informative responses show what was written, read, or returned, giving the agent and the learning system a useful signal for understanding outcomes and improving future behavior.
Install to Summarize YouTube Videos and Get Transcripts
Explore YouTube Video Summarizer or Get YouTube Transcript Extractor
Questions & Answers
Q: What is the experience gap in AI agents?
The experience gap is the difference between increasing model capability and the practical experience an agent accumulates from doing real work. Models are getting smarter, but conversations can still feel like their first day on the job. The talk argues that intelligence alone is insufficient, just as a highly intelligent person may need workplace experience before becoming effective in a specialized role.
Q: How can AI agents learn from real user interactions?
AI agents can learn by preserving complete interaction data and turning it into usable feedback rather than discarding it after each task. Relevant signals include agent traces, tool calls, sub-agent work, user edits, retries, undos, and subsequent actions. These records can help define desired behavior, evaluate performance, improve models, and adjust the harness that supplies tools and context.
Q: What should an AI agent trace for continual learning?
An agent should trace the entire execution tree, including the primary action, tool calls, and work delegated to sub-agents. It should also record meaningful user responses, especially corrective behavior such as edits, retries, and undos. Capturing only the main result leaves gaps that make it difficult to identify why a task succeeded or failed and what should change.
Q: Why are thumbs-up and thumbs-down ratings insufficient for agent feedback?
Simple ratings can be noisy because users may approve agent work without closely examining every consequence. A problem might become visible only several commits later, when the user realizes that an earlier action broke something and reverses it. Edits, retries, undos, and similar corrective behaviors therefore provide important evidence that a basic approval signal may fail to capture.
Q: How should production traffic be used to evaluate AI agents?
Evaluations should be drawn from the ways people actually use the product, including current requests and frontier requests that may not yet be fully achievable. The product, evaluation, and training environments should remain as similar as possible. User tasks should be replayable when infrastructure permits, and performance should be graded through the real production harness rather than a separate approximation.
Q: What is the role of an agent harness in continual learning?
The harness supplies the capabilities and context through which an agent works, including search tools, private information, and other product primitives. Instead of enforcing a fixed sequence of actions, it can allow the model to orchestrate those primitives. The harness can also hold contextual or changing information that should remain accessible without being trained permanently into the model itself.
Q: When should information go into the model versus the harness?
Behavior learned from full interaction traces may be appropriate for improving the model, including through reinforcement learning approaches discussed in the talk. Contextual or changing facts may belong in the harness instead. The example given is that a company has been delisted: rather than training that fact into model weights, the system can make it available as context through the harness.
Q: Why should tool responses contain detailed information?
Detailed tool responses give agents a clear signal about the outcome of an action. A response that merely says an operation is done or finished does not reveal what was written, read, or returned. That ambiguity makes the result confusing for the agent and provides little useful evidence for training, evaluation, debugging, or learning from the completed interaction.
Summary & Key Takeaways
-
Current models are becoming more capable, but they often behave as if every interaction were their first day at work. Trajectory calls this the experience gap. Continual learning addresses it by preserving agent activity and human responses, then converting that accumulated experience into improvements that can compound as people continue using the system.
-
The proposed platform begins with complete interaction traces and derives specifications describing what the agent should do. Those specifications guide improvements to both models and harnesses. Reinforcement learning methods can train on long traces, while changing facts or contextual information can remain available through the harness instead of being embedded into model weights.
-
Companies pursuing continual learning should trace complete agent trees, capture corrective behavior, build evaluations from production traffic, and replay user tasks when possible. Their harnesses should expose product capabilities as flexible primitives, align agent tools with the user interface, return informative results, and let increasingly capable agents orchestrate actions without unnecessarily rigid flows.
Read in Other Languages (beta)
Share This Summary 📚
Summarize YouTube Videos and Get Video Transcripts with 1-Click
Try YouTube Summary with ChatGPT & Claude or YouTube Transcript Generator
Explore More Summaries from Sequoia Capital 📚






Summarize YouTube Videos and Get Video Transcripts with 1-Click
Try YouTube Summary with ChatGPT & Claude or YouTube Transcript Generator