Services

What we do, and how an engagement starts.

Evaluation Fit CallFree

A working conversation about the AI agent, not a sales call.

We look at what the agent does or is meant to do, who decides what a right answer is, and what you need to know before it reaches users.

Bring one AI agent at any stage, from an early design to production. We will tell you honestly what can be evaluated now, what has to exist first, and what an evaluation would and would not establish.

Trace ReviewFree

You send traces your AI agent has already produced. We read them and show you what an evaluation of it would actually find.

Not a document. You get a view you log into, with your own traces in it, the failure categories we found while reading them, criteria drafted against those failures, and an agreement figure on the ones we labelled. It stays there afterwards.

We take a small number of these at a time, because the reading is the part that does not automate. The traces come under an NDA and the data handling answer is published rather than promised on the call.

What we need from you

  • Between twenty and a hundred traces from one AI agent that is already live.
  • A sentence on what you would call a wrong answer, in whatever words you already use.
Evaluation Readiness Audit

We read your production traces and tell you what you can and cannot currently measure.

What you get

  • A failure taxonomy coded from your own traces, with how often each failure actually happens.
  • A first criteria set drafted against those failures, with pass criteria and fail triggers written out.
  • Agreement numbers on a first pass: how often an automated judge scoring those criteria matches a human label.
  • A straight statement of which criteria are nowhere near good enough to gate anything yet, and why.
  • Written agreement on who on your side settles a close call.

What we need from you

  • Production traces, and whatever you already capture alongside them.
  • Time with the person whose judgement counts as correct.
Calibrated Release Gate Pilot

One AI agent, one calibrated gate, running in your pipeline. Not a slide deck.

We score the agent against a few hundred of your own past cases, then measure the judge itself against human labels and keep revising until each criterion either clears the bar we agreed or is marked as not yet trusted to decide anything.

This is the step that turns an opinion into a number, and then tells you what that number is worth. Afterwards you know how often the agent is right, how often the judge is right about the agent, and which criteria are allowed to hold a release.

The gate ships as a running system your team logs into, not a repo you inherit: run history, trends, the difference between one version and the next, and the runs behind every number.

What we need from you

  • A few hundred real past cases we can score against.
  • Someone who can settle the close calls on what counts as correct, and time from them for labelling.
  • A place in your pipeline for the gate to run.
Evaluation System Build

One evaluation engine your other teams can adopt by configuration rather than by fork.

Criteria, thresholds, pre-checks and judge framing live outside the engine, so a second product brings its own rubric instead of a second pipeline. Keeping the engine ignorant of any one domain is what makes that possible.

What you get

  • The engine running in your stack, with a second product onboarded by configuration to prove it transfers.
  • The gate in your release process, so a change that makes things worse cannot reach your customers.
  • Run history per criterion, with runs that are not comparable marked as not comparable rather than drawn through.
  • A written runbook for the team who will operate it.
  • Training for the people who will do the labelling and maintain the criteria.
Continuous Quality GateOngoing

We keep it accurate as your data and your business change.

An agent does not stay right on its own. The data moves, the model changes underneath it, and the day it drifts you hear it from a customer, not a dashboard.

What we do each month

  • Re-run the criteria against traffic the AI agent has actually seen, and send you a dated result you can forward to whoever asked.
  • Refresh the cases so they still reflect the work as it is now.
  • Re-check that the automated scoring still agrees with your experts, and re-calibrate against fresh labels when it does not.
  • Review what moved, what we changed, and what to watch.
  • A named person who answers when something looks wrong.

After the Handover the number is yours. The Continuous Quality Gate does not take it back: it checks it against labels your team did not produce, and a named person who is not on that team reads the result before you act on it.

Cancel any month, with notice.

Three more ways to work with us

Independent Agent Evaluation

We do not build the AI agents we evaluate.

An evaluation of an AI agent somebody else built, whether that was your own team or another vendor. What you get is the same running dashboard and evaluation record every engagement here ends in, plus a written summary of what we found and what we could not cover.

  • We read what the agents actually did in production.
  • We catalogue every way they fail, and how often each one really happens.
  • We build a set of tests against your definition of a right answer.
  • We check the automatic scoring against your own experts and report how closely they agree.
  • We state in writing what our review could not cover.

Before we start, you write down what you already believe is wrong with it, and that list goes in the engagement letter, so what we find can be told apart from what you already knew.

Useful before a renewal, a handover, or any decision where somebody is going to ask you how you know.

The Calibration Handover

Your own team taught to keep the number true: the labelling protocol including the close-call rule, a gold set they produced themselves, one criterion taken to the agreement bar in the room by them, and a re-calibration procedure with the triggers that should set it off.

The Calibration Handover hands over the maintenance. When it works, your team owns the labelling, the criteria and the re-calibration, and you do not need us in the room to keep the number true. What a trained team still cannot supply itself is a re-calibration against labels drawn from outside it, and someone who is not on it reading the result. That is what the Continuous Quality Gate is for.

A client who takes the Handover and never renews is the Handover working as sold.

Executive Working Session

A working session in a room with the person accountable for the AI agent and their technical counterpart, working on your actual system rather than a slide deck.

Useful when several people need to reach the same understanding at the same time, or when you want to pressure-test an idea before committing to it.

You leave with a written view on what can and cannot be measured on this AI agent today, and a proposal for the Audit or the Pilot.

Questions we get asked

The things worth asking before you book.

Do you build AI agents?

We build evaluation systems, and that is most of what we do. What we do not build is the AI agent under evaluation, for anybody we are evaluating for. The moment we build it, we cannot be the ones who tell you honestly how often it is wrong. If it helps, we will tell you what to look for in whoever does build it, and we will evaluate what they ship.

What happens to our traces?

They are held only for the engagement they were sent for, and deleted when it ends, or sooner if you ask. The criteria, the rubric and the labels your experts produce are yours and stay with you whatever happens to the traces. The data handling page is the full answer, including everyone else who touches them.

How we handle your data →

How is this different from the eval tool we already pay for?

A tool gives you somewhere to put criteria. It does not tell you whether your criteria are the right ones, it does not know what your experts mean by correct, and it does not measure its own judge against them. That is the work. If the tool you own is where the criteria should live, we will build them there rather than move you off it.

Who owns the evaluation when you leave?

You do. It runs in your pipeline, on criteria your experts signed, and the runs behind every number stay in it. The Continuous Quality Gate is a decision about whether you want a re-calibration against labels your own team did not produce, and somebody who is not on that team reading the result. It is not a decision about who holds the system.

Will our security review accept your evaluation?

We do not know, and we would rather find that out with you than promise it. What you get is your own evaluation record: the criteria, the runs, which criteria cleared the agreement bar, which are still marked indicative, and what the review could not cover, in a dashboard your team keeps. It is yours to show whoever is asking. It is not a certificate and we will not dress it up as one. Tell us what your reviewer needs to see and we will write the record to answer it.

Start with a conversation

Book an Evaluation Fit Call.

Free, and useful whether or not we work together. Bring one AI agent at any stage, from an early design to production. We will tell you honestly what can be evaluated now, what has to exist first, and what an evaluation would and would not establish.

Not ready for a call? Send us the one criterion you trust least and we will write back on what it would take to make it decide something.

Pick a time

No deck. No discovery workshop. One AI agent, what it does or is meant to do, and a partner on the call.