How to Build Reliable Evals for AI Agents

20.7K views
•
May 14, 2026
by
AI Engineer
YouTube video player
How to Build Reliable Evals for AI Agents

TL;DR

Reliable agent evaluation starts by tracing actual runs, reading those traces, and categorizing failures before choosing any metric. Combine cheap deterministic code checks with LLM judges, validate the judges through meta-evaluation, and run datasets as experiments so prompt or model changes can be measured for regressions instead of approved by intuition.

Transcript

Hi everybody. Uh my name is Lori Voss. I am head of developer experience at AriseAI. Uh in a former life, I co-founded npm Inc. So some of you may remember me from when I used to talk incessantly about JavaScript. Now I talk incessantly about AI. Uh what I think about mostly is how to test AI systems, how to make AI systems that actually work. Uh s... Read More

Key Insights

  • Evals are tests for AI applications, while traces are the runtime logs that supply those tests with evidence. A trace records the agent's execution as nested spans, including agent turns, tool calls, LLM invocations, inputs, outputs, timing, token counts, and related metadata.
  • The vibes problem is the practice of testing an AI feature with a few queries and deciding whether the responses look right. It misses unfamiliar wording, edge cases, adversarial inputs, and regressions, while offering no repeatable test suite that can run in continuous integration.
  • Agent outputs cannot usually be validated through exact string matching because the same prompt can produce different text across runs, and several different responses may all be correct. Effective evaluation therefore needs checks suited to both structured requirements and flexible semantic judgments.
  • Code evals are deterministic functions written in languages such as Python or TypeScript. They run quickly, cost almost nothing, and can reproducibly check properties such as valid JSON, output length, token limits, or whether a required subject appears in the response.
  • LLM-as-a-judge evals assess semantic qualities that simple functions cannot reliably capture. Built-in evaluators provide ready-made judgments, while custom evaluators can apply a task-specific rubric and labeled examples to decide whether an agent's behavior succeeds or fails.
  • Choosing the right evaluator is more important than merely tuning one. A correctness evaluator scored the financial agent 0 out of 13, while a faithfulness evaluator scored the same agent 13 out of 13, because the judge could not verify forward-looking financial data from its own knowledge.
  • Meta-evaluation is the process of testing whether an evaluator's judgments are themselves correct. Applying it to an LLM judge helps reveal whether the rubric and labeled examples produce trustworthy classifications instead of treating the judge's output as unquestionable ground truth.
  • Datasets and experiments transform evaluation from failure reporting into an iteration loop. Teams can run repeatable cases before and after a prompt or model change, compare results, detect unintended regressions, and demonstrate that the proposed modification improved the agent rather than merely looking better.

Explore YouTube Video Summarizer or Get YouTube Transcript Extractor

Questions & Answers

Q: What is an eval for an AI agent?

An eval is a test that measures whether an AI application behaved as intended. Its evidence commonly comes from traces, which record agent calls, tool calls, LLM invocations, inputs, outputs, timing, token counts, and other metadata. Because valid AI responses can vary in wording, an eval may use deterministic code, an LLM judge, or both.

Q: Why is checking a few agent responses not enough?

Checking a few responses by eye creates the vibes problem. It can miss edge cases, adversarial inputs, unexpected user vocabulary, and failures caused by later prompt changes. Human review also does not scale, catch regressions consistently, or run in continuous integration. A repeatable eval suite tests broader behavior and shows whether a change broke something that previously worked.

Q: What is the difference between a trace and a span?

A trace records what an AI application did during execution, including its inputs, outputs, timing, token counts, and intermediate activity. A span represents one step within that trace, such as an LLM invocation, a tool call, or a complete agent turn. Spans can contain other spans, creating a nested structure that exposes the agent's execution path.

Q: Why should teams read traces before writing evals?

Teams should read traces first because the observed failures reveal what actually needs to be measured. Running representative queries and categorizing problems by root cause prevents the team from selecting an evaluator based on assumptions. The resulting categories can guide the choice of deterministic checks, built-in judges, or custom rubrics that correspond to genuine agent failure modes.

Q: When should a code eval be used for an AI agent?

A code eval should be used when success can be expressed as a clear deterministic rule. Examples from the workshop include checking whether an output is valid JSON, whether it stays under a token limit, and whether it mentions the requested subject. These evaluations run in milliseconds, cost almost nothing, and return reproducible results across repeated executions.

Q: How does an LLM-as-a-judge eval work?

An LLM-as-a-judge eval asks a language model to assess the semantic content of an agent's output and classify whether it succeeded or failed. Teams can use built-in evaluators or create a custom evaluator with a task-specific rubric and labeled examples. This approach handles flexible qualities that cannot be reliably measured through basic string matching or structural code checks.

Q: Why can the wrong evaluator produce misleading results?

An evaluator can fail when its judgment requires knowledge it does not possess. In the workshop, a correctness evaluator gave the financial agent 0 out of 13, while a faithfulness evaluator gave it 13 out of 13. The correctness judge could not verify forward-looking financial information because the model did not know the relevant year, whereas faithfulness tested use of supplied sources.

Q: How do datasets and experiments improve an AI agent?

Datasets provide repeatable test cases, and experiments run an agent or a proposed change against those cases. This creates an iteration loop in which teams can compare prompt or model versions, quantify improvement, and detect regressions. Instead of eyeballing a modified response and declaring success, the team can show whether the change improved measured behavior without breaking existing capabilities.

Summary & Key Takeaways

  • AI evals are tests built from trace data. Traces capture agent calls, tool calls, LLM invocations, inputs, outputs, timing, token counts, and other metadata as nested spans. This evidence replaces the common practice of trying a few queries and deciding by intuition whether an agent appears ready to ship.

  • The practical workflow begins with instrumenting an agent through Phoenix, running multiple representative queries, and reading the resulting traces. Failures should be categorized by root cause before metrics are selected. Evaluation can then combine deterministic code checks, built-in LLM judges, and custom rubric-based judges supported by labeled examples and meta-evaluation.

  • Datasets and experiments turn evals into an improvement system. They let teams compare prompt or model changes against repeatable cases, detect regressions, and verify that an intended fix actually helped. The workshop also introduces the impact hierarchy, the data flywheel, pairwise evaluation, and reliability scoring as techniques for further development.


Read in Other Languages (beta)

Share This Summary 📚

Explore More Summaries from AI Engineer 📚