Blog
What to collect when you’re evaluating an AI agent
Your eval score moved seven points and nobody can say why. What you have to record for a number about an AI agent to still mean something six weeks later.
Your AI agent works in the demo. Now someone asks whether it works in general, and you need a number.
So you turn on tracing. Your framework gives you a span tree: the model calls, the tool calls, the retrieval step, token counts, latency. It looks thorough. You run your test cases, you get 71% pass rate, you write it in the doc.
Two weeks later the number is 64%. Nobody can say why.

That seven-point drop is the problem. Not because 64% is necessarily bad, but because you can’t explain why it changed. The question is: what do you need to record so that, six weeks later, the number still means something?
Two problems, and the second one is worse
We’ve seen this go wrong in two different ways.
Sometimes the problem is obvious: you simply didn’t record the thing you later needed. You know you should log “the traces”, so you enable whatever your framework emits and hope it covers the important parts. That’s a hard list to design before you’ve seen the failures it needs to explain.
The more frustrating version is when you recorded everything your framework offered and still can’t answer the question. That’s the one that tends to burn a week. You have every span from every run, but “did the model get worse, or did we change something?” is still unanswerable because the trace never separated the model from the code around it.
Take a refund agent. It calls a refund tool. The tool returns HTTP 200. Your trace records a successful tool call, your agent says “done, refund issued”, and your eval marks the case passed. No refund was issued. The tool accepted the call and silently did nothing, because the record it was pointed at didn’t exist.
Nothing in the span tree is wrong. The span tree faithfully recorded what the agent did. It just never recorded whether anything happened as a result, and that’s a different field.
The refund case is only one version of this:
A tool call that the model asked for but that never ran, because a guardrail blocked it or the parser dropped it. If you only log calls that executed, that failure looks like the model not trying.
A run that stopped because it hit your step limit. Averaged in with genuine wrong answers, it makes a budget problem look like a capability problem.
A model whose output format drifted so your extractor stopped finding the answer. Scored as a reasoning failure, fixed by prompt engineering, when the bug was in the regex.
A refusal rate that looks reassuringly high because the agent has quietly stopped calling tools at all. Doing nothing is very safe.
These aren’t edge cases. They happen in ordinary systems, and they fail in the same annoying way: the trace looks complete while the conclusion you draw from it is wrong.
Why this is genuinely hard
This isn’t mainly a discipline problem.
No single standard covers this. OpenTelemetry’s GenAI semantic conventions give you an excellent vocabulary for what the agent did, and significant parts of them are still marked Development. NIST’s audit controls tell you how to keep a record trustworthy but leave retention and sampling as decisions you have to make. W3C PROV-O handles lineage and isn’t an evaluation schema. The EU AI Act tells you to keep logs without telling you which fields make a score defensible. They solve different pieces of the problem. None tells you, end to end, what has to be captured for an eval result to remain explainable.
And the parts that matter most are exactly the parts no instrumentation library can give you. An auto-instrumentor sees inside your process. It cannot tell you which tools you offered the model but it never picked, whether the refund actually landed in the database, what the judge was shown before it scored, or how many cases you quietly dropped from the denominator. Your instrumentation can’t infer those facts. You have to design for them.

Where we got the answers
We wanted the planner to be grounded in more than our own experience, so we used two kinds of evidence.
The first is a research corpus of our own. Evaluating AI Agents: A Practitioner’s Guide is assembled from every arXiv paper submitted between January 2025 and September 2026 across cs.AI, cs.CL, cs.LG, cs.MA, cs.SE, cs.HC and stat.ML. That’s 174,437 papers. After dedupe and two screening gates, 89,677 were about evaluating AI systems, and 40,959 carried both an agent-or-judge subject and an assessment predicate. The top candidates were scored against a fixed rubric, and 1,507 papers were read in full, each with an evidence record quoting the passage and the number behind every claim.
The corpus is large, but two choices in how we built it matter more than the paper count.
Every claim carries its source as an arXiv id, and a verifier re-checks the guide against the papers on disk: every number must appear verbatim in the extracted text of the paper it cites. Rounding a figure fails the build. Computing a difference between two of a paper’s own numbers fails the build. Both happened during writing and both were repaired.
Claims are also marked by evidence strength. An unhedged sentence is something a paper measured. Wording like “the authors recommend” or “proposed but not validated” means nobody tested it, and you should treat it as a hypothesis worth trying rather than a practice to adopt. We kept that distinction in the skill. In practice, it matters more than having another long list of “best practices.”
The second body of work covers what practitioners and standards bodies already publish. We worked through the trace and span models from OpenTelemetry’s GenAI conventions and OpenInference, the per-framework capture behaviour of LangSmith, Braintrust, Arize Phoenix, MLflow, W&B Weave, Microsoft Foundry and Semantic Kernel, the OpenAI Agents SDK, CrewAI and Google ADK, plus the UK AI Security Institute’s Inspect log format, Anthropic’s task-trial-grader-outcome vocabulary, NIST SP 800-53 and the AI Risk Management Framework, W3C PROV-O, EU AI Act Article 12, and the Model Cards and Datasheets documentation proposals.
The research tells us what can change the conclusion. The implementation docs tell us what you’re likely already collecting and what each field is called. We needed both.
What we built
agent-trace-planner is an open agent skill. It asks you about your agent, then writes the trace-collection plan for it.
Behind it is a taxonomy of 1,728 recordable fields, organised into ten layers by granularity: what pins the run, what describes the task, what each step and tool call emits, what’s derived across the whole trajectory, where the verdict came from, what the environment holds, what the judge and human labellers produced, what it cost, and what the safety arm observed. Cross-cutting that are nine agent types, because a coding agent and a voice agent need different subsets.
Eighteen of those rows are the floor. The rule for getting on that list is narrow: without this row, a published score cannot be attributed to the agent at all. Everything else sharpens a claim you can already attribute.
There are also 329 warnings about how those signals can mislead you. Some are traps, where a signal looks like evidence and isn’t. Some are contested, where two credible results point opposite ways and the only honest move is to log both sides and measure it yourself. Some were asserted by authors without being tested. And 70 entries are questions the research doesn’t answer, which we list rather than paper over. How many repeats you need. What your noise floor is. How often to re-validate a judge. We leave those questions open on purpose. A made-up default would look helpful and create false confidence.
Using it
If you work in Claude Code, clone the repo and copy one directory:
git clone --depth 1 https://github.com/kayzn-io/kayzn-skills.git /tmp/kayzn-skills
cp -r /tmp/kayzn-skills/.claude/skills/agent-trace-planner ~/.claude/skills/
Start a new session afterwards, since skills load at session start. Then ask in your own words:
What traces should I collect to evaluate my support agent?
Or clone the repo and open it directly, in which case the skill is discovered without installing anything.
The skill follows the Agent Skills format, so it isn’t tied to one tool. OpenCode reads .claude/skills/ as well. Codex looks in .agents/skills/. It’s a directory of markdown and one Python script, with no dependencies beyond Python 3.8.
It starts with questions about your system. There are 47 in total, but the interview branches, so most users see far fewer. It asks what your agent does, what it calls, what cuts a run short, whether anything persists between runs, and who or what decides that a run passed. Every question comes with a concrete example and guidance for when you don’t know the answer, because we watched people get stuck on the first draft of those questions.

Your answers become capability flags, and the flags select rows. That mapping is deterministic, in a script you can run yourself:
scripts/recommend.py --flags calls_tools,uses_llm_judge,in_production --format markdown
The mapping lives in code deliberately. The same rules will eventually drive a web version, and two implementations that disagree about what you need would be worse than none.
What comes out
A markdown document you can hand to an engineer. It opens with what the skill understood about your system, so you can correct it before reading further. Then the eighteen floor items, then the rows for your setup grouped by layer, then anything your specific answers demanded.
The most useful parts come after the tables.
The instrumentation mapping tells you which rows your framework already emits, so you verify rather than build, and which ones nothing will give you for free.
The traps section lists the warnings that apply to your particular setup, phrased as what not to conclude.
Then the declared gaps: everything the skill could not recommend, and why. If you said you have no way to log outside the agent’s reach, it doesn’t quietly hand you environment-snapshot rows and let you discover later that they were impossible. It says so.
We think the declared gaps are the most important output. A plan that admits four holes is more useful than one that pretends to have none, because now you know what to close.

How this helps in practice
You find out what you’re missing before you need it, rather than during an incident. The gap between “we log tool calls” and “we log whether the call did anything” is invisible until it costs you a week.
You stop confusing your own bugs with the model’s. A large share of what teams read as model regression turns out to be a parser change, a moved cap, an environment that drifted, or a test set with defects in it. Separating those is most of what the floor rows are for.
You also get a boundary around the claim. If a finding was measured on one model and one benchmark, the planner keeps that limitation attached to it instead of letting it harden into a rule you cite in a design review.
What it doesn’t do
The skill won’t tell you how many tasks to run, how many repeats you need, where to set a threshold, how to weight process against outcome, how long to keep your traces, or how often to sample production.
Those are decisions you still have to make for your system. The planner’s job is narrower: make sure you collect enough evidence to understand what actually caused the score you report.
And if your setup makes some evidence impossible to collect, say for example, you can’t independently verify what happened after a tool call, capture the relevant content, or replay a run, then it calls that out as a gap rather than pretending the trace plan is complete.
Get it
The repo is at github.com/kayzn-io/kayzn-skills, free for personal and commercial use under Apache-2.0. The guides work as reading material on their own if you’d rather not install anything; start at references/00-INDEX.md.
If you run agent evals, the most useful thing you can send us is the question the interview gets wrong for your system, or the trace you needed that isn’t in there. If you send us either, we’ll use it to improve the next version.
