Free
Send us traces. We will show you what an evaluation would find.
You send traces your AI agent has already produced. We read them, code the ways it actually fails, draft criteria against those failures, and label a first pass ourselves. What you get back is not a document. It is a view you log into, with your own traces in it while we are working on them.
What comes back
- The failure categories we found, coded from your traces rather than from a list of failure modes in general, with how often each one turned up in what you sent.
- Criteria drafted against those failures, with pass criteria and fail triggers written out, in the form a judge could run.
- An agreement figure on the ones we labelled: how often an automated judge scoring those criteria matched our own label. Including the criteria where it did badly, because those are the ones that tell you something.
- A straight statement of what this could not establish, which on a sample this size is most things. See the limits below.
All of it in a running view your team logs into, with the runs behind every number. If you want something to forward to whoever asked, it exports.
What survives, and what does not. The failure categories, the criteria and the labels are derived from your traces rather than being them, so they are yours and they stay yours. The traces themselves follow the same rule as every other engagement: deleted when this one ends, or sooner if you ask. So export anything you want to keep before then, and the data handling page is the rule in full.
What we need from you
- Between twenty and a hundred traces from one AI agent that is already live. More is not better here: a hundred read properly beats a thousand skimmed.
- Whatever the agent had in front of it when it answered, if you capture it. The retrieved material is what makes it possible to tell a grounded answer from a fluent guess, and without it several of the most useful criteria cannot be scored at all.
- One sentence on what you would call a wrong answer, in whatever words you already use. Not a rubric. We are going to write that part.
Send the least that answers the question. If a trace can be redacted and still show what the agent did and what evidence it had, redact it. What happens to the traces after they arrive, how long they are kept, how they are separated and who else touches them is written down rather than promised on the call: the data handling page is the whole answer, and it is worth reading before you send anything.
The honest limits
A hundred traces cannot tell you how often your agent is right. The sample is too small and it is whatever you happened to send rather than a representative draw, so nothing in this establishes a rate you could take to a release meeting. Any number in it is an indication of where to look, and it is labelled that way.
The judge in it is not calibrated. Calibration means measuring an automated judge against labels your own experts produced, and your experts have not been in the room yet. The agreement figure here is against our labels, not theirs, and the two are not the same thing. Getting a criterion to a bar worth gating on is the Calibrated Release Gate Pilot, and this is not a small version of it.
What it does establish is whether there is anything here worth measuring properly, and what it would take. That is the decision this is for.
How many we can read at a time
You watch the method run on your own data instead of reading a claim about it. The limit is that one of us reads every trace, and reading does not get faster. We take a small number at a time. If yours has to wait we will tell you when you ask rather than after you have sent the traces.
Not ready to send production data to anyone? Send us one criterion instead. No data, no NDA, and it answers a narrower version of the same question.
Start a Trace Review
Tell us what your AI agent does.
Write to hello@kayzn.io and say what the agent does and what you would call a wrong answer. We reply within two business days with the NDA and where to put the traces. Do not attach them to that first email.
A call is the faster route if you are not sure this is the right agent to start with. Also free, and it ends in the same recommendation.