How Can Trace Mining Continuously Improve Agents?

15.9K views
•
August 12, 2026
by
AI Engineer
YouTube video player
How Can Trace Mining Continuously Improve Agents?

TL;DR

Agent improvement starts by shipping an agent, collecting its traces, mining those traces for behavioral patterns, and testing changes against evidence from prior runs. Because autonomous behavior is difficult to predict from code alone, trace-level analysis can reveal failures, user reactions, compaction effects, model differences, and the domain-specific feedback needed for evaluation, prompt refinement, and fine-tuning.

Transcript

[music] >> Hey everyone. I'm Vic and I lead applied research at LangChain and I'm going to talk about something that I think is sexy, which is data mining, but it's not as sexy as LLM, so we're going to try to like make it sexy together. And the problem that we're going to talk about today is how do we continuously improve agents, but how do we do ... Read More

Key Insights

  • Agent improvement is a four-stage loop: ship the agent, collect extensive traces, mine those traces, and run data-driven experiments. The experiments determine whether changes to prompts, tools, orchestration, or agent loops actually improve performance on behaviors found in previous runs.
  • Observability is tightly coupled with continual learning because an agent's environmental activity produces the record needed for reflection and updating. Traces preserve tool calls, output messages, API interactions, CLI usage, and other evidence that can guide changes to an agent's knowledge or operating instructions.
  • Agent behavior is difficult to infer from code because modern systems combine prompts, tools, skills, hooks, middleware, nested agents, and swarm orchestration. As systems trade determinism for autonomy, teams need trace-based methods to understand how prompt changes affect real behavior across different domains.
  • Trace mining can answer concrete behavioral questions that aggregate metrics cannot fully resolve. Agents can search other agents' traces for positive or negative user interactions, performance degradation after context compaction, and the fine-grained behavioral differences produced by substituting one model for another.
  • Large-scale trace analysis is expensive because its input cost grows with both trace volume and trace length. Some coding-agent interactions are also too large to fit inside another agent's context, requiring systems that treat long traces as external objects that can be queried selectively.
  • Open models can perform specialized trace judging at substantially lower cost when their harnesses are engineered from observed evidence. In LangChain's work with Harvey's legal benchmark, an open model approximately matched a frontier model's judging capability at one to two orders of magnitude lower cost.
  • Harness engineering is the fastest first step for improving a model on an evaluation because it provides feedback in about two minutes. Once repeated prompt and guidance changes reach an intelligence ceiling, fine-tuning becomes the method for breaking through that limit on a narrow domain-specific task.
  • Dense feedback is more actionable than a simple pass-or-fail benchmark result because it identifies what happened inside an agent's trajectory. Traces contain fine-grained signals that can support evaluations, human-readable feedback, issue discovery, and training data preparation for continual improvement.

Explore YouTube Video Summarizer or Get YouTube Transcript Extractor

Questions & Answers

Q: How can trace mining continuously improve AI agents?

Trace mining supports a continuous improvement loop that begins with deploying an agent and recording its activity in real environments. Teams then search tool calls, messages, API usage, CLI activity, and outcomes for recurring successes and failures. Those findings become evaluations and experiments that test whether revised prompts, tools, orchestration, or execution loops produce better behavior on previously observed cases.

Q: Why are traces necessary for understanding agent behavior?

Traces are necessary because an agent's behavior cannot reliably be inferred from source code alone. Its actions emerge from prompts, tools, skills, hooks, middleware, orchestration, and interactions with its environment. A trace captures the actual sequence users experienced, allowing teams to inspect fine-grained behavior, identify failure points, and assess the effects of changes in a specific domain.

Q: What information should an agent trace contain?

An agent trace should preserve the activity generated while the agent operates in its environment. The transcript identifies tool calls, output messages, API calls, and CLI usage as important examples. Centralizing this information by agent or across multiple agents gives analysis systems a record they can search for mistakes, successful interactions, user reactions, compaction effects, and model-specific behavioral differences.

Q: How can agents analyze traces produced by other agents?

A reviewing agent can search centralized traces with targeted questions instead of reading every interaction indiscriminately. It can look for cases where users became upset or especially satisfied, examine whether performance declined after the first or second context compaction, and investigate what a different model might have done. This agentic search approach helps make large collections of behavioral data more manageable.

Q: Why is large-scale trace analysis difficult and expensive?

Large-scale analysis becomes expensive when teams have millions of traces and each trace may contain millions of tokens. Input cost grows with the number of traces and their average size. Some long coding-agent interactions cannot fit inside another agent's context at all, so the trace must be treated as an external object that an analysis system can query selectively rather than load completely.

Q: How can open models reduce the cost of trace judging?

Open models can reduce trace-judging costs when teams first establish that a task is possible with a frontier model, inspect its traces, and then engineer a stronger harness for a cheaper model. In LangChain's work with Harvey's legal benchmark, this process allowed an open model to approximately match a frontier model's judging capability at one to two orders of magnitude lower cost.

Q: When should teams move from prompt engineering to fine-tuning?

Teams should begin with harness and prompt engineering because it provides feedback in about two minutes and can quickly reveal whether added guidance improves evaluation results. Repeated adjustments eventually reach a ceiling where further prompt changes yield little benefit. At that point, fine-tuning a base model on the narrow, domain-specific tasks customers actually need can reach or exceed frontier performance on those tasks.

Q: Why is dense feedback better than pass-or-fail evaluation?

Dense feedback is better because it tells an agent or development team what happened within the trajectory, while a pass-or-fail result provides little information about what should change. Traces already contain fine-grained evidence about decisions, tool usage, intermediate behavior, and outcomes. Mining that evidence can generate evaluations, identify issues, prepare training data, and produce feedback that humans can review.

Summary & Key Takeaways

  • Improving an agent is a continuous loop: deploy it into real environments, retain its tool calls, messages, API activity, and other traces, then mine that evidence for useful patterns. Teams can turn the resulting findings into evaluations and experiments that test whether changes to prompts, tools, orchestration, or execution loops improve observed behavior.

  • Agent traces are too numerous, costly, and sometimes too long for straightforward human or model review. LangChain therefore uses agents to search traces produced by other agents, looking for successful and unsuccessful interactions, user reactions, degradation after context compaction, and counterfactual evidence about how another model might perform on the same task.

  • Trace analysis also informs economic and technical choices. LangChain found that an open model could approximately match a frontier model's legal trace-judging capability at one to two orders of magnitude lower cost after harness engineering. When prompt improvements reach a ceiling, narrow domain-specific fine-tuning can reach or exceed frontier performance on targeted tasks.


Read in Other Languages (beta)

Share This Summary 📚

Explore More Summaries from AI Engineer 📚