Skip to content
Kayzn
What we take onCase studiesTeamServicesWhat's newBlog
Menu
What we take onCase studiesTeamServicesWhat's new (research digest)Blog (field notes)

Start

Send us one criterionTrace ReviewBook a Fit Call
Book a Fit Call

One analytics cookie, to count visits. What we collect.

Blog / calibration

Tagged “calibration”

  • 29 September 2026

    We put four AI judges on trial. Here is who got it right, and what it cost.

    Before you let a model grade your AI agent, someone has to grade the model. We ran four judges, including Jev, TypeSafe’s new System One model, over 230 support conversations with a hidden answer key.

    →
  • 8 August 2026

    Evals, part 1: when “it looks fine” stops working

    GenAI output is probabilistic, so the old question, did it match the label does not work. Why AI products fail silently, and why spot-checking cannot catch it.

    →
  • 8 August 2026

    Evals, part 2: choosing the right ones for your product

    A stack, seven product archetypes, and the one question that decides most of your eval design: is there a right answer to compare against?

    →
  • 8 August 2026

    Evals, part 3: what a year of building them taught us

    Four principles, how to build an LLM-as-a-judge forwards rather than backwards, the five metrics to start with, and why robustness comes from stacking imperfect layers.

    →
  • 1 August 2026

    Evaluating LLM agents in production: a source-aware, calibrated approach

    How to move from subjective spot checks to a repeatable, evidence-based evaluation platform that can gate releases, and serve more than one product.

    →

← All posts

Kayzn

We measure how often your AI agents are right, and we show you how accurate that measurement is.

Work

What we take onServicesSend us one criterionTrace ReviewBook a Fit Call

Reading

Case studiesWhat's newBlog

Industries

Financial servicesHealthcareInsuranceLegal

Contact

hello@kayzn.io

Los Angeles · Montréal

© 2026 KayznPrivacyData handling