Blog

We put four AI judges on trial. Here is who got it right, and what it cost.

Before you let a model grade your AI agent, someone has to grade the model. We ran four judges, including Jev, TypeSafe’s new System One model, over 230 support conversations with a hidden answer key.

Your AI agent talks to a thousand customers a day. You want a number for how often it got things right. So you hand each conversation to a second model and ask it: pass or fail?

That second model is a judge. And nobody has checked whether the judge is any good.

We hear this a lot. A team picks the smartest model they can afford, writes a rubric, and starts trusting the score. Then the score moves, and nobody can say whether the agent changed or the judge did.

So we ran the trial ourselves. Four judges, 230 real support conversations, and an answer key none of the judges could see. This post is about what we found, what it cost, and a design decision we made early that the results forced us to reverse. It changed the outcome more than any choice of model did.

How the trial worked

The six steps of the study: 230 conversations, a hidden answer key, a reading copy, four judges, comparison to the key, harder tests

The conversations come from a public benchmark called tau-bench. Each one is a customer asking a retail support agent for something: return these two items, exchange this kettle for a different colour, change the address on a pending order. The agent is an AI. It looks things up, makes changes, and talks to the customer, who is also simulated.

What makes this benchmark useful is the answer key. Every task has a known correct end state: which orders should have changed, to what, and paid how. After the conversation ends, a script compares the database to that expected state. Match and it’s a pass; anything else is a fail. No opinion involved. That is the ground truth the judges are measured against.

We ran the same 115 tasks twice. Once with a careful agent that follows the store’s policy and confirms every change with the customer before making it. Once with a rushed agent, told to act on the first thing the customer says. That gave us 230 conversations, and a second question to ask later: can a judge tell these two agents apart?

Then each conversation was written out as a transcript. Every message, every tool the agent called, every result it got back. The answer key was stripped out. This transcript is all a judge sees.

Four judges read every transcript five times each, so we could tell whether a verdict was a considered opinion or a coin toss that happened to land one way.

Two of the four are text models. gpt-5 is the strongest we tested; it reads the rubric and the transcript, thinks at length, and writes out a pass or fail with its reasoning. gpt-4o-mini is a fast, cheap model doing the same job.

The third is Jev, from a company called TypeSafe. It is the first of what they call a System One model, and the name is a nod to Kahneman: fast, intuitive System 1 thinking, as against the slow, deliberate System 2 that gpt-5 is doing when it writes out 1,700 words of reasoning. Jev does not write anything. You give it a fixed statement (“the agent completed the request in line with the policy”) and it returns a probability that the statement is true. That number is the whole output. TypeSafe says it is trained to produce calibrated probabilities over a fixed set of typed answers, all in one pass rather than one token at a time, which is where the speed comes from.

The pitch is that giving up text buys three things: an answer in well under a second, at a small fraction of what a text model costs, and a probability that means what it says. The announcement is two weeks old. Part of the reason Jev is in this trial is to see how those three claims hold up on a judging task with an answer key. Two of them did, and we come back to the third.

The fourth is a rule-based check: plain code that compares the agent’s tool calls against the expected actions in the answer key. Nobody could deploy it, because it needs the answer key. It is here to show what a check with perfect information looks like.

The scoreboard

How often each judge agreed with the answer key, with fail recall, pass rate, consistency, and ranking quality

Start with the first column. gpt-5 agreed with the answer key on 60% of conversations. Jev and gpt-4o-mini on 57%. A judge that said “pass” to everything would score 49%, because about half of these conversations passed.

So the buzzing headline is: the text and decision models are better than guessing, but not by much. gpt-5’s lead over always-pass is 11 points. The other two are about 7 points ahead. The rule-based check, which reads the answer key, scores 79%.

That would be a short and sad post if the first column were the whole story, so read across the row.

gpt-4o-mini reaches 57% by saying pass to 88% of everything. It caught 19% of the failures. That is not a judge so much as a rubber stamp: it approves nearly every conversation it sees, and the few it marks as fail are the ones where the agent visibly did nothing.

Jev reaches the same 57% a different way. It caught a third of the failures, and the last column shows what it is good at: if you hand it one conversation that passed and one that failed, its score puts the passing one higher 72% of the time. That is the best ranking of the three. It is also almost perfectly consistent. Ask it five times and you get the same answer on 99% of conversations.

gpt-5 sits between them. Best at catching failures (nearly half), the closest of the three to saying pass as often as the answer key does, and the least consistent: one conversation in five got a different answer on at least one of its five readings.

Three judges with nearly the same accuracy, and three different ways of getting there. Anyone choosing between them on accuracy alone would have learned nothing useful.

Jev and gpt-5 agree with each other more than with the truth

Jev and gpt-5 agree with each other far more often than either agrees with the answer key. In statistical terms their agreement is kappa 0.50, where their agreement with the truth is 0.14 and 0.21. Two very different systems, one a small model that returns a number and one a large model that writes an essay, are converging on the same verdicts. And those verdicts are not the ones in the answer key.

What are they agreeing on? The second agent tells us. The figure below shows, for each judge, what share of the careful agent’s conversations it marked as pass and what share of the rushed agent’s.

The rushed agent stopped asking for confirmation; the answer key drops 4 points, gpt-4o-mini 8, gpt-5 17, Jev 37

The rushed agent skips one thing: asking the customer to confirm before making a change. By the answer key, that costs it 4 points. It still gets the right end state most of the time, because customers usually meant what they first said.

The share of conversations Jev marks as pass falls 37 percentage points. gpt-5’s falls 17. Both are marking the rushed agent’s conversations as fail for skipping confirmation, whether or not the final state was right. They are grading behaviour, because the rubric told them to. Our rubric says “confirmed with the user before any order-changing write” is part of passing. The answer key does not care about that. The judges read the rubric and did what it said.

This matters for anyone choosing a judge. There are two different things you might mean by “did the agent do a good job”: did it get the right result, or did it behave the way policy requires? A judge that works from the transcript, as all three models here do, can see the behaviour directly. Every confirmation asked for, every identity check, every policy rule followed or skipped is right there on the page. The result is not on the page. Whether the right order was changed in the right way lives in the database, and the transcript only shows what the agent said it did and what the tools said back. A judge reading the transcript has to piece the result together from those clues, and we will see below how often it pieces it together wrong. So when a rubric names behaviour, which ours did, a transcript judge grades behaviour, and grades it well. If what you want is the result, you need something that checks the database, and a transcript judge will look worse than it is. If what you want is the behaviour, the transcript judges here are the only ones that noticed the policy had been broken, and they noticed it consistently. Decide which of those two you want before you pick the judge, because the judge will not decide it for you.

The decision we reversed

When we designed the transcript, we chose to cut every tool result at 600 characters.

The reasoning was sound on paper. The tool results in these conversations are raw JSON. A single product lookup returns about 2,200 characters of variant IDs, option lists, and prices, and one conversation can contain a dozen lookups. Put all of that in the transcript and the customer’s words and the agent’s replies end up separated by walls of data. A judge is supposed to be grading a conversation, not reading a database dump. The agent’s own summary of what it found is right there in the next message, in plain language, which is what the customer heard and what we thought the judge should be weighing. Keeping the first 600 characters of each result seemed like enough to show which tool was called and roughly what came back, without drowning the conversation. It also made every transcript shorter and every verdict cheaper.

On the first run, gpt-5 scored 54%, exactly chance on the careful agent, and was the least consistent judge of the four. The finding was going to be that the most expensive model was no better than a coin. That seemed wrong enough to make us read its reasoning before writing it up.

Over and over, on conversations that had passed the answer key, gpt-5 was marking the agent as fail for giving “a specific tracking number without any visible tool output supporting that number”, or “a specific refund amount without any tool output supporting that calculation”. We pulled 14 of the numbers it called unsupported and went looking for them in the raw data. All 14 were in the tool results. The agent had not made anything up.

The judge had never seen them. Across the 230 conversations, 59% of tool results had been cut, and the numbers the agent quoted to the customer were often in the part we removed. Then we handed that transcript to a judge with a rubric that said “fail the agent if it states things unsupported by the tools”, and it did exactly what we asked.

The argument for cutting had a hole in it. The agent’s summary is not a substitute for the tool result, because checking the summary against the result is one of the judge’s jobs. Take the result away and a careful judge cannot do that job, and a strict one will fail the agent for the gap.

We reversed the decision, wrote the transcripts out whole, and ran all three model judges again.

Ranking quality before and after the fix: gpt-5 0.58 to 0.69, Jev 0.65 to 0.73, gpt-4o-mini 0.58 to 0.60

gpt-5 went from 54% to 60%, from the least consistent judge to a usable one, and from chance to the most accurate model in the study. Jev improved on ranking and became almost perfectly consistent. gpt-4o-mini didn’t move, which tells you how closely it was reading the tool output in the first place. The whole transcripts were about 30% longer, so the cost saving we had been protecting turned out to be small.

The transcript is part of the judge. Change what the judge can see and you change the verdict, sometimes more than switching models would. Before comparing judges, check what your pipeline leaves out, and whether your rubric asks the judge to verify things that are no longer on the page.

The expensive judge’s essay also earned its keep once. We would not have found this without gpt-5’s written rationales. Jev gives a probability and nothing else, which is a virtue when you want speed and a liability when you want to know why.

The failures nobody caught

44 conversations failed the answer key and were marked as pass by all three model judges. We read them.

In 21, the agent did the right kind of thing to the wrong item. A customer asked to exchange two identical kettles for two different variants; the agent confirmed the details, got a yes, and sent the same new variant twice. gpt-5’s reasoning walks through every step approvingly and marks it as pass. In another, the agent processed one of two requested changes and charged the wrong card, and all three judges praised its confirmation flow.

These are not reasoning failures. They are identifier lookups: does the item ID the agent sent match the one the customer described? A model reading the transcript does that badly. A line of code does it in one comparison, and the rule-based check caught 40 of the 44.

TypeSafe describes Jev as a model that cannot hallucinate, and in the sense they mean it that is true: it cannot return an answer outside the set you gave it. It can still return the wrong one. On the kettle, it did, along with both text models.

If your agent’s correctness turns on getting IDs, amounts, and variants exactly right, a model reading the transcript is the wrong tool for that part. Have code do it instead: write down what a correct outcome looks like in concrete fields (this order, these item IDs, this card) and compare the agent’s actions, or the record it left behind, against that list. That is what our rule-based check does, and it is why it caught the kettle. Keep the model for the questions code cannot answer, such as whether the agent explained itself clearly or followed policy along the way.

What it costs

Cost per thousand verdicts and seconds per verdict, log scale

Jev and gpt-5 rank conversations equally well. We could not tell them apart on that measure with 230 conversations. One costs 19 cents per thousand verdicts and answers in under a quarter of a second. The other costs $28 per thousand and takes half a minute, because it writes about 1,700 words of reasoning every time.

Of the $33 we spent judging the whole set of transcripts, gpt-5 accounted for $32.

Two things not to trust

The confidence number. Every judge here says how sure it is of each verdict, as a percentage. You might hope that when a judge says 90%, it is right about nine times in ten. None of them is. gpt-4o-mini says about 90% whether it turns out to be right or wrong. gpt-5 and Jev are closer to honest but still off by 20 points or more on average. There is a standard repair for this: take a set of conversations you have already graded, see how often the judge was right at each confidence level, and use that to translate its future numbers. We tried it. It fixes the percentages, so 90% starts meaning roughly 90%, but the judge still cannot tell you which of its verdicts are the wrong ones. Treat the confidence as a hint, not a number to act on.

That includes Jev, and it is the one part of TypeSafe’s pitch we could not confirm. Jev is trained, in their words, for calibrated decisions, so that higher confidence means higher accuracy. On these transcripts, scored against the answer key, its raw probabilities were off by about as much as gpt-5’s. The repair above brings them into line, but you have to do the repair. Speed, cost and consistency were as advertised. Calibration, out of the box, was not.

The score change between agent versions. A common use for a judge is to compare your agent before and after a change: run both versions over the same conversations, score each with the judge, and see whether the share it marks as pass went up or down. We tried that here with the careful agent and the rushed agent. By the answer key, the careful agent passed 51% of its conversations and the rushed agent passed 47%, a drop of 4 percentage points. Jev marked 90% of the careful agent’s conversations as pass and 53% of the rushed agent’s, a drop of 37 points. gpt-5 went from 72% to 55%, a drop of 17. Both judges got the direction right, the rushed agent was worse, but both made the drop look several times bigger than it was.

That is not the judges inventing a difference. We checked by splitting the careful agent’s own conversations into two random halves and asking each judge to compare them. There is no real difference between the halves, and no judge reported one. The judges only see a big gap when there is a real change in how the agent behaves, and in this case the behaviour did change a lot, even though the results changed only a little. So if you use a judge this way, trust the sign of the change and be careful with the number. A 15-point drop in the judge’s score may be a 2-point drop in what your customers get.

Let’s get a bit more technical

The accuracy figures are the modal verdict of five readings per conversation, scored against tau-bench’s state match, n = 230. Some intervals, all paired bootstrap on the same conversations:

  • gpt-5 beats always-pass by 3.5 to 19 points. Jev and gpt-4o-mini beat it by 0.4 to 14 and 3 to 12.
  • gpt-5 and Jev cannot be separated on accuracy (difference −2 to +10 points) or on AUROC of the “completed” score (−0.09 to +0.03). Both beat gpt-4o-mini on AUROC (+0.01 to +0.17 and +0.04 to +0.21).
  • Removing the 600-character cap raised AUROC by +0.06 to +0.18 for gpt-5 and +0.04 to +0.11 for Jev. gpt-4o-mini’s change (−0.03 to +0.08) is noise.
  • Raw Brier on pass probability is 0.27 to 0.38 for the model judges; a coin scores 0.25. Cross-fitted isotonic correction brings them to 0.22 to 0.26.
  • A confidence-gated cascade from gpt-4o-mini to gpt-5 peaks at 62% for $0.008 per conversation, one point over gpt-5 alone and one over Jev alone.
  • Jev’s default 0.5 threshold is lenient on whole transcripts (it marks 72% as pass; 49% truly passed). A threshold chosen by 5-fold cross-validation brings it to 68% agreement, against 66% for gpt-5 treated the same way.

Full tables, confusion matrices, the miss taxonomy, and rationale themes are in docs/judge-comparison.md in the repository.

What we would tell a client

If you can write down what the correct end state looks like, have code compare the agent’s work against it. Nothing else here came close.

If you cannot, and you want a model that reads the conversation, know that it will grade behaviour more than result, because behaviour is what it can see. That is fine if process is what you care about. Jev does that grading as well as gpt-5 for a small fraction of the cost, and gives you the same answer every time. gpt-5 gives you a written reason, which you will want the day the score moves and nobody knows why. gpt-4o-mini marks nearly everything as pass and should not be a gate.

Whichever you pick: give it the whole transcript, set its threshold on conversations you have labelled yourself, and do not read its confidence as a probability.

And test the judge before you trust it. Ours were fine. One of our own design decisions was not, and we only found out because we read the verdicts instead of just counting them.

The code

Everything here reproduces from saved data. The 230 conversations, every verdict from every judge, and the scripts that turn them into these numbers are in the repository at github.com/kayzn-io/decision-models-as-judges, under the Apache 2.0 license. Two commands install it and open an app that walks through the steps. Rerunning the judges yourself costs about $35, almost all of it gpt-5; replaying from the cache costs nothing.

One thing we have not written up yet: the repository also carries a fourth model judge, Laya, an open-source decision model you can run on your own machine for free. It only reads 512 tokens, which is too short for a full transcript, so it works from a compressed summary and we tested it separately. That is a post for another day.

If you are running an agent and want to know how often it is right, our Trace Review is free. Send us the conversations and we will tell you what we find, and how sure we are of it.

← All posts