Why Tracing Your Agent Isn’t the Same as Testing It
Here’s the gap nobody on your AI team wants to admit: 89% of organizations have implemented agent observability, but only 52% run offline evals, and just 37% run online evals against production traffic, according to LangChain’s survey of 1,340 practitioners fielded in late 2025 (LangChain, 2025). That’s a 37-point hole between seeing what your agent does and knowing whether it did it well.
I’ve been building agentic systems for the past year: a multi-agent code reviewer, a few internal automation agents, an evaluation pipeline for one of them. The pattern repeats: teams ship the agent, wire up tracing, stare at LangSmith dashboards, and then quietly hope nothing’s broken. Tracing tells you what happened. Evals tell you whether what happened was right. Most teams have one and not the other.
This piece walks through the three-grader system that production teams actually use (code grader, model grader, human grader): what each one catches, where each one fails, and how to write your first eval suite without buying a platform on day one.
what agentic AI is and how multi-step agent systems work
Key Takeaways
- 89% of teams have agent observability but only 52% run offline evals, a 37-point gap between seeing what an agent does and knowing whether it was right (LangChain, 2025 survey).
- Quality (accuracy, relevance, consistency, tone) is the #1 blocker keeping agents out of production, cited by 32% of respondents; latency is second at 20%.
- The three-grader system layers code graders (deterministic, cheap), model graders (LLM-as-judge, scalable), and human graders (gold standard, expensive). No single grader covers everything.
- Even GPT-4 Turbo, a top judge in a major 2025 survey’s own test, agreed with human labels only 61.5% of the time (arXiv, 2025). Model graders need human calibration.
- Start with 30 hand-annotated traces and a code grader before buying a platform. Most teams overbuy tooling and underwrite tests.
What Is Eval-Driven Agent Development?
Eval-driven agent development means you write the test before you ship the prompt. Instead of vibing through prompt iterations and hoping production users will tell you when things break, you maintain a versioned suite of input/output pairs scored by automated graders, and you treat regressions in those scores the same way you’d treat a failing unit test in a backend service. LangChain’s 2025 survey of 1,340 practitioners found 57% already have agents in production, so for most teams this is no longer a prototype problem (LangChain, 2025). My read, after building a few of these: the teams shipping reliable agents treat evals as deployment gates, not analytics dashboards.
The discipline borrows directly from test-driven development. You define the behavior you want, encode it as a graded eval, run the eval against every prompt change, and block deploys that drop the score below a threshold. The shift is cultural more than technical. Most engineering teams already know how to write tests. They just haven’t accepted that LLM apps need them.
When I was building a multi-agent code review skill, I made every classic mistake first. I shipped without evals. I tweaked the system prompt based on whichever output looked good in my latest run. I changed three things at once and couldn’t tell which change broke which behavior. Two weeks in, I had no idea if the new version was better than the old one. Only that the demos still looked impressive. That’s the trap eval-driven development is designed to break.
how I built the multi-agent code review skill that pushed me toward evals
The 37-Point Gap: Why Observability Isn’t Enough
Observability without evals is a smoke detector with no thermostat. According to LangChain’s State of Agent Engineering 2025, 89% of teams have implemented some form of agent observability (traces, span timing, tool call logs), but only 52.4% run offline evals on those traces, and just 37.3% run online evals against live production traffic (LangChain, 2025). The instrumentation is there; the scoring isn’t.

Why does this gap exist? Observability ships first because vendors push it first. Datadog, Langfuse, LangSmith, Arize Phoenix: every platform leads with traces because traces are easy to demo. You add a wrapper, you get a waterfall view, you feel productive. Evals require you to define what “correct” means for your agent, and that’s the hard part. It’s product work, not platform work.
Datadog’s State of AI Engineering report (July 2026) adds a complication. Across thousands of its customers, more than 70% of organizations now use three or more models, and the share using more than six nearly doubled (Datadog, 2026). Every extra model is another thing that can quietly change behavior under you. Without evals, you can see each one in a dashboard, but you can’t tell whether swapping or upgrading it made your agent better or worse.
Citation Capsule: 89% of AI teams have implemented agent observability, but only 52% run offline evals and 37% run online evals, a 37-point gap that explains why most agent projects stall after launch (LangChain, 2025). Tracing tells you what happened. Evals tell you whether it should have.
how AI agents audit codebases at the eval boundary
How Do Production Agents Actually Fail?
Production agents fail in predictable ways, and most of those failures aren’t the model being dumb. Arize’s January 2026 breakdown of common agent failures lists eight recurring modes: retrieval noise and context-window overload, hallucinated arguments in tool calls, recursive loops, guardrail failures around sensitive data, pre-training bias overriding retrieved context, unhandled API schema changes, instruction drift in long sessions, and unsafe code generation (Arize, 2026). Read that list again and count how many live in the scaffolding (retrieval, tools, APIs, session length) rather than in the model weights. It’s most of them.
For eval purposes, I group those modes into three buckets, because each bucket needs a different grader:
- Context failures. The agent didn’t have the right inputs at the right step, or its training-data priors beat the documents you retrieved. Long sessions make it worse as the original instructions lose weight. These show up as confident, wrong answers, which is a model-grader problem.
- Action failures. The agent called a tool with parameters that “feel” right but don’t exist, or took a destructive action it was told not to take. Arize cites the Replit incident, where an agent ran
DROP TABLEdespite explicit instructions. Most of these are catchable with code graders that check tool-call arguments and block state-changing calls. - Execution failures. Recursive polling loops that fire hundreds of API calls for one task, or an agent that “hallucinates success” after a 400, 401, 403 or 429 error instead of reporting it. Code graders on trajectory length and error handling catch these cheaply.
None of these throw a clean exception that your observability stack will page you about. That’s the whole problem.
Then there’s silent degradation, the one that should keep you up at night. Your agent’s response time looks fine. Error rates look fine. Token usage looks fine. But the answers are getting subtly worse: a slightly wrong format here, a hallucinated tool call there, a fact that used to be right and now isn’t. Observability won’t catch this. Only an eval suite scoring outputs against a ground-truth set will.
Why Are Agents Stuck in Pilot?
The reason most agents never escape pilot mode isn’t cost or latency. It’s quality. LangChain’s 2025 survey found a third of respondents named quality (accuracy, relevance, consistency, tone and guideline adherence) as their primary blocker, 32% in total, with latency a distant second at 20%. Cost came up less often than in previous years as model prices fell. Among enterprises with 2,000+ employees, security jumps to the number-two spot at 24.9% (LangChain, 2025). In my experience, when stakeholders kill an agent project, it’s because they don’t trust the output.

Here’s the inversion that matters: cost and latency are measurable. You can graph them. You can SLA them. Quality, in the absence of evals, is vibes: a debate between a PM, an engineer, and three demo screenshots. Without an eval suite producing a numeric score that everyone agrees on, “is the agent good enough?” becomes a political question, not a technical one. That’s why projects stall.
What Is the Three-Grader System?
The three-grader system is a layered approach where every agent output gets scored by three different mechanisms (code, model, and human), and each one catches what the others miss. Anthropic’s engineering team lays out exactly these three grader types in its agent-evals guide: code-based graders are fast, cheap and reproducible but brittle; model-based graders are flexible and scalable but non-deterministic and need human calibration; human graders are the gold standard but slow and expensive (Anthropic, 2026). No single grader is cheap, scalable, and aligned with human judgment at once. You combine them deliberately, and the combination is the system.

Code graders are deterministic. They run in milliseconds, cost nothing, and never disagree with themselves. They’re perfect for format checks, schema validation, and any rule you can express as code. They’re useless for “does this answer feel correct?” Model graders (LLM-as-judge) use a separate LLM call to score outputs against a rubric. They scale, they handle nuance, and in my runs they cost cents per eval, not dollars. But they drift from human judgment in subtle ways. Human graders are the gold standard you calibrate everything else against, but they’re slow and expensive. You can’t run them on every commit.
Citation Capsule: In the meta-evaluation run by A Survey on LLM-as-a-Judge, GPT-4 Turbo outperformed the other standard judges tested, yet it agreed with human labels on only 61.5% of 5,106 samples, and reasoning models such as o3-mini landed in the same range (61.7%). GPT-4 Turbo did keep the same verdict 80.3% of the time when the two candidate answers were swapped (arXiv: A Survey on LLM-as-a-Judge, 2025). The lesson: LLM judges are consistent enough to scale, but not aligned enough to trust uncalibrated. Use them for volume and calibrate them against humans monthly.
how AI coding agents are graded across platforms
How Do You Build a Code Grader (Layer 1)?
A code grader is just a function that takes the agent’s output and returns a pass/fail or numeric score using deterministic rules. Start here because it’s free to run, impossible to misinterpret, and catches more bugs than people expect. Anthropic’s engineering guidance is to choose “deterministic graders where possible, LLM graders where necessary,” and code graders are the deterministic layer: format, schema, and tool-call correctness (Anthropic, 2026).
Here’s a minimal example for an agent that’s supposed to return JSON-formatted recommendations:
def code_grader(output: str, expected_keys: set[str]) -> dict:
"""Score an agent output on schema correctness."""
try:
parsed = json.loads(output)
except json.JSONDecodeError:
return {"score": 0, "reason": "Output is not valid JSON"}
missing = expected_keys - set(parsed.keys())
if missing:
return {"score": 0.5, "reason": f"Missing keys: {missing}"}
if not isinstance(parsed.get("recommendations"), list):
return {"score": 0.5, "reason": "recommendations must be a list"}
if len(parsed["recommendations"]) < 3:
return {"score": 0.7, "reason": "Need at least 3 recommendations"}
return {"score": 1.0, "reason": "Schema valid"}
That’s it. Run that against every output the agent produces, and you’ve already caught most “the model returned weird formatting” failures before they reach a user. The open-source DeepEval framework (Apache 2.0, 18.6K+ GitHub stars as of October 2026, Confident AI / DeepEval) is built around this same pytest-style pattern and pulled roughly 2.6 million PyPI downloads in the month to early October 2026 (pypistats, 2026). Its pre-built metrics for faithfulness, hallucination, and answer relevancy are mostly LLM-judged, though, which is really layer two.
The trap with code graders is overconfidence. They’ll happily tell you the schema is valid even when the content is garbage. A weather agent that returns {"forecast": "purple lasagna"} passes every JSON validator on Earth. Code graders catch shape, not substance. That’s why you need layer two.
When Should You Use a Model Grader (Layer 2)?
Use a model grader (LLM-as-judge) whenever you need to score something subjective at scale: factual correctness, tone, helpfulness, faithfulness to a source. You write a rubric in a prompt, point a separate LLM at the output, and ask it to score on a 1-5 scale with reasoning. The 2025 survey on LLM-as-a-judge notes that with appropriate prompt design, a judge’s agreement with human judgment “can be promising,” but also lists the biases you’re signing up for, from position bias to a preference for longer answers (arXiv, 2025). A tight, structured rubric is the cheapest prompt-design win available.
A model grader prompt looks roughly like this:
You are evaluating whether an agent's response is FAITHFUL to the source.
Source document: {source}
Agent response: {response}
Rate faithfulness on a 1-5 scale:
1 = Response contradicts the source
2 = Response includes claims not in the source
3 = Response is partially supported
4 = Response is mostly supported
5 = Every claim is directly supported by the source
Respond as JSON: {"score": int, "unsupported_claims": [list], "reason": str}
The thing nobody tells you about LLM-as-judge: the judge model needs its own eval suite. If you’re using GPT-6.1 Sol to grade your Claude Sonnet 5.5 agent, you’ve now got two models in the loop, and you need to verify the judge agrees with humans before you trust it. Treat your LLM judge as a piece of production infrastructure that needs its own calibration set of 50-100 human-scored examples. Recompute agreement monthly. If it drops, your judge has drifted, and every eval score downstream is suspect.
how solo founders can keep eval costs under control
Model graders shine on factual accuracy and tone, both areas where code graders are useless. They’re cheap (cents per eval on a mid-tier model like Sonnet 5.5 at $2/$10 per million tokens), fast enough to run in CI, and version-controllable as part of your prompt repo. Just don’t outsource your conscience to them. They drift, they have biases (they prefer longer responses, they prefer their own family of models), and they will silently over-credit outputs that sound right.
Why Is the Human Grader (Layer 3) Still Non-Negotiable?
Human graders are the calibration anchor for the entire eval system, and you cannot eliminate them no matter how good your model judge gets. Even GPT-4 Turbo, the strongest standard judge in the LLM-as-a-judge survey’s own meta-evaluation, matched human labels just 61.5% of the time (arXiv, 2025). That’s why practitioners haven’t dropped humans: LangChain’s survey found that among organizations running evals, 59.8% still use human review, slightly more than the 53.3% using LLM-as-judge (LangChain, 2025). Humans disagree with each other too, but they disagree in ways you can debug. Model judges disagree in ways you often can’t detect until production damage is done.
The pattern that works: maintain a “golden set” of 50-200 examples that domain experts have hand-scored. Rerun your model judge against this set once a month. If model-human agreement on the golden set drops below your threshold (say, 75%), your judge has drifted and you need to either retune the rubric, swap models, or expand the calibration set. Anthropic’s open-source Bloom framework, released in December 2025, shows the same loop at research scale: a four-stage agentic eval generator (Understanding, Ideation, Rollout, Judgment) whose judge stage was validated against 40 hand-labeled transcripts, with Claude Opus 4.1 reaching a 0.86 Spearman correlation with human scores (Anthropic Alignment, 2025).
When I added human review to my own agent project, the surprise wasn’t that humans caught things the model judge missed. It was that humans flagged a category of failure the model judge couldn’t even see. The agent was producing technically-correct outputs that no human would actually use, because the phrasing was robotic. The model judge scored these 5/5 on factual accuracy. A real human reviewer wrote one comment: “I’d never read past the first sentence.” That’s the kind of feedback that changes a product. You can’t get it from an LLM.

https://www.youtube.com/watch?v=uiza7wp1KrE
Which Eval Platform Should You Actually Use?
There’s no single right answer, but the eval platform market has resolved into clear archetypes. Braintrust positions evals as deployment gates and prices on usage, with unlimited users even on its free Starter tier and a $249/month Pro plan (Braintrust pricing, 2026). Langfuse is the open-source Swiss Army knife: observability plus evals plus prompt management, with an MIT-licensed, self-hostable core. LangSmith offers the deepest LangChain/LangGraph integration, and its free plan includes 5,000 base traces a month for a single seat (LangChain pricing, 2026).
Here’s how the major options stack up for a production deployment as of October 2026:
| Platform | OSS or Commercial | Best For |
|---|---|---|
| Braintrust | Commercial | Eval-as-CI/CD; engineering teams that want regression-blocking deploys |
| Galileo | Commercial | Production guardrails and runtime evaluators |
| DeepEval / Confident AI | OSS (Apache 2.0) + commercial cloud | Pytest-style unit testing for LLM apps; teams that already think in test suites |
| Langfuse | OSS (MIT core) + commercial cloud | OpenTelemetry-native tracing, evals, prompt mgmt without lock-in |
| LangSmith | Commercial | LangChain/LangGraph apps where zero-config tracing matters |
| Arize Phoenix | Source-available (Elastic License 2.0), self-hostable + Arize cloud | ML platform teams who want OTel-based tracing and evals they can run themselves |
| Latitude | OSS (MIT) + commercial cloud | Prompt management and evals in one open-source platform |
| Promptfoo | OSS (MIT) | Config-driven evals and red-teaming from the CLI, easy to drop into CI |
| Anthropic Workbench / Bloom | Console + OSS Bloom | Behavioral and alignment evals on Claude |
| OpenAI Evals | OSS (MIT) | Model-centric scaffolding; baseline framework |
My recommendation: don’t pick a platform first. Pick a grader pattern first. Write your code grader and your model grader as plain Python functions. Hand-label 30-50 examples in a CSV. Run your suite. Then decide if you need a platform. And once you do, the choice usually becomes obvious from your stack (LangChain → LangSmith, OTel-everything → Langfuse, regression-blocking deploys → Braintrust).
How Do You Write Your First Eval Suite?
Start tiny. The single biggest mistake I see is teams shopping for an eval platform before they’ve written their first eval. You don’t need infrastructure for 10 examples. You need a CSV and an afternoon. Hamel Husain’s evals FAQ, which he keeps updating (most recently in September 2026), puts it plainly: start with error analysis, not infrastructure. He recommends annotating at least 30 traces yourself and reviewing at least 100 during error analysis, continuing until new traces stop revealing new failure modes (Hamel Husain, 2026). The discipline is in the labeling, not the tooling.
Concrete starting steps:
- Pull 30 real production traces. Not synthetic data, but real inputs and real outputs your agent produced. If you don’t have production traces yet, run the agent against a sample of expected user inputs and capture them.
- Hand-score each one as pass/fail. No 1-5 scales yet. Just: did this output do what the user needed?
- Cluster the failures. Read the failed examples. Group them. You’ll find 3-5 categories: schema bugs, hallucinations, missing context, wrong tool, wrong tone. Those are your eval dimensions.
- Write a code grader for the easy categories. Schema and format checks first. On my projects this alone caught roughly a third of failures, with no LLM call.
- Write a model grader rubric for the hard categories. Factual accuracy, tone, helpfulness. Use the structured 1-5 rubric pattern from the section above.
- Run the suite on every prompt change. Block any change that drops aggregate score by more than 5% without explicit override.
- Add 10 new examples a week. Pull from production. Especially pull from edge cases and recent failures.
That’s it. By month three you’ll have a 150-example suite, two graders, and a rough model-human agreement number. You’ll know, numerically, whether your agent is getting better or worse. You’ll catch silent degradations before users do. And you’ll be in the 52% of teams running offline evals, not the 48% still flying blind.
how MCP architecture changes which eval boundaries matter
https://www.youtube.com/watch?v=BsWxPI9UM4c
Frequently Asked Questions
What’s the difference between offline evals and online evals?
Offline evals run against a fixed dataset of input/output pairs, usually in CI before a deploy. Online evals run against live production traffic, scoring real user interactions in real time. According to LangChain’s 2025 survey, 52% of teams run offline evals while only 37% run online evals (rising to 44.8% among teams with agents in production), and the online layer is where silent degradation actually gets caught (LangChain, 2025).
How many evals do I actually need to start?
Start with 30 hand-labeled examples and one code grader. Hamel Husain’s evals FAQ recommends annotating at least 30 traces yourself before automating anything, and reviewing 100+ during error analysis (Hamel Husain, 2026). Thirty is enough to detect obvious regressions and small enough to label in an afternoon. Grow to 150-300 examples over the first quarter, weighted toward production failure cases.
Can I use a GPT model to grade my Claude agent?
Yes, but calibrate it. In the LLM-as-a-judge survey’s meta-evaluation, even GPT-4 Turbo, one of the strongest judges tested, agreed with human labels only 61.5% of the time (arXiv: A Survey on LLM-as-a-Judge, 2025). Newer judges like GPT-6.1 Sol or Claude Sonnet 5.5 haven’t been run through that same test, so don’t assume they’re better aligned without checking. Maintain a 50-example human-scored golden set and recompute model-human agreement monthly to catch judge drift.
Do I need an eval platform or can I roll my own?
Roll your own first. A Python script + a CSV will get you to 100 evals before you outgrow it. Once you need versioned datasets, regression-blocking CI, or LLM-judge dashboards, switch to Braintrust, Langfuse, DeepEval, or Promptfoo depending on your stack. Don’t pre-buy infrastructure for problems you don’t have yet.
Why does observability cover 89% of teams but only 52% have evals?
Observability is bought; evals are built. Vendors ship trace dashboards as a one-line install, but evals require teams to define what “good” means for their agent, and that’s product work, not platform work. The 37-point gap exists because the harder layer doesn’t have a checkbox in a procurement form.
Conclusion
If your agent has observability but no evals, you’ve built half of a quality system. You can see the failure happening; you just can’t tell it’s a failure. That’s the gap behind the 89% / 52% number, and it’s the same gap behind the 32% of teams who name quality as the reason their agent is stuck in pilot.
The three-grader system isn’t novel. It’s just the discipline that production teams converge on once they’ve shipped enough agents to know what breaks. Code graders catch shape. Model graders catch substance. Human graders catch what neither of the others can see, and they keep the whole system honest.
You don’t need a platform to start. You need 30 examples, a Python file, and an afternoon. Once that’s in place, you’re already ahead of the 48% of teams who haven’t even tried.
If you want to go deeper on the architectural side of agent design (how MCP, multi-agent coordination, and tool integration shape what your eval boundaries should look like) start with our guide to agentic AI fundamentals. From there, walk through our multi-agent code review walkthrough.