Evals as CI: release gates for LLM systems

Versioned datasets, deterministic scorers, a calibrated LLM judge and a paired statistical gate that blocks a merge when quality drops. With tested code.

Level
Advanced
Stack
TypeScript, Vitest, Any LLM provider

In short

  • Run every change to prompts, models, retrieval or post-processing against a versioned dataset of real, anonymised cases.
  • Score with code wherever possible, and only use an LLM judge after measuring its agreement with humans using Cohen's kappa.
  • Gate releases on paired results: block critical regressions, significant drops (exact McNemar) and drops beyond an agreed margin.
  • Report pass rates with Wilson intervals, and keep the eval check separate from unit tests so outages read as "could not run".
In this article · 11 sections
  1. 01What counts as a change
  2. 02The dataset is code
  3. 03Scorers: deterministic first
  4. 04The LLM judge, and how to earn the right to use it
  5. 05Running the suite
  6. 06The gate: paired comparison, three rules
  7. 07Reporting uncertainty
  8. 08Wiring it into CI
  9. 09Testing the eval code itself
  10. 10What to build first
  11. 11References

Every change to an LLM system is a behaviour change you cannot fully predict. A reworded instruction fixes the case that prompted it and quietly breaks three others. A new model version is better on average and worse on the one task your largest customer depends on. A retrieval tweak improves recall and floods the context with near-duplicates.

Unit tests do not catch this, because the code is the same. What changed is the behaviour of a component you do not control, steered by text. The only reliable way to know whether a change is safe is to run it against a fixed set of cases and compare the outcome with what you have today.

That is what an eval suite is, and it belongs in CI, not in a notebook someone runs before a big launch. This article builds the pieces: a versioned dataset, scorers that are deterministic wherever possible, an LLM judge that has been calibrated before it is trusted, and a release gate that compares a candidate with the baseline using a paired statistical test. As in the rest of the series, every snippet is type-checked and tested in CI before this page is built.

A pull request triggers two runs of the same versioned dataset, baseline and candidate. Scorers grade each output, and a gate compares the paired results and blocks the merge on a critical regression, a significant drop or a drop larger than the margin.
Figure 1. Evals as a release gate. The dataset and the gate policy live in the repository; the decision is made by code, not by someone reading a dashboard.

What counts as a change

The trigger for an eval run is wider than most teams first assume. Anything that can change what the system outputs for the same input is a change:

  • the prompt template, including the system prompt and examples,
  • the model, the model version or the provider route,
  • sampling parameters such as temperature or maximum output length,
  • retrieval: the index, the chunking, the embedding model, the number of results and any reranking,
  • tool definitions and their descriptions, which the model reads as instructions,
  • post-processing: parsers, validators and repair logic.

In the gateway from the first article, prompts are registered with a version and a content hash, and routes are configuration. That makes it possible to say precisely what differs between two runs, which is the precondition for comparing them.

The dataset is code

The dataset is the most valuable part of the suite and the part most often treated carelessly. We keep it in the repository as JSONL, one case per line, reviewed in pull requests like any other change.

evals/dataset.ts
export interface EvalCase {
  /** Stable id. Never reuse an id for a different case: history depends on it. */
  id: string;
  input: string;
  /** Reference answer or structured expectation, depending on the scorer. */
  expected?: unknown;
  /** e.g. ["billing", "swedish", "critical"]. Drives per-slice reporting. */
  tags: string[];
}

Each case has a stable id, the input, an optional expectation and tags. Tags drive per-slice reporting, so you can see that a change improved billing questions and hurt Swedish-language ones, instead of a single number that hides both. One tag is special: critical marks behaviour that must never regress, and the gate treats it separately.

evals/dataset.ts
/**
 * Datasets live in the repository as JSONL, one case per line, reviewed in
 * pull requests like code. Loading is strict: a malformed line fails the
 * run instead of silently shrinking the suite.
 */
export function parseJsonl(text: string): EvalCase[] {
  const cases: EvalCase[] = [];
  const seen = new Set<string>();
  text.split('\n').forEach((line, i) => {
    if (!line.trim()) return;
    let raw: unknown;
    try { raw = JSON.parse(line); } catch { throw new Error(`line ${i + 1}: invalid JSON`); }
    const c = raw as Partial<EvalCase>;
    if (typeof c.id !== 'string' || !c.id) throw new Error(`line ${i + 1}: missing id`);
    if (typeof c.input !== 'string') throw new Error(`line ${i + 1} (${c.id}): missing input`);
    if (c.tags !== undefined && !(Array.isArray(c.tags) && c.tags.every((t) => typeof t === 'string'))) {
      throw new Error(`line ${i + 1} (${c.id}): tags must be strings`);
    }
    if (seen.has(c.id)) throw new Error(`line ${i + 1}: duplicate id ${c.id}`);
    seen.add(c.id);
    cases.push({ id: c.id, input: c.input, expected: c.expected, tags: c.tags ?? [] });
  });
  return cases;
}

Loading is strict. A malformed line, a missing input or a duplicate id fails the run. A loader that skips bad lines is worse than no loader, because the suite shrinks silently and the pass rate goes up for the wrong reason.

Three rules keep a dataset useful over time.

An id never changes meaning. If the correct answer for a case changes, because the product changed, retire the old id and add a new one. Reusing an id with a new expectation breaks every historical comparison that includes it.

Cases come from production. The best cases are real inputs that went wrong: support tickets, thumbs-down feedback, outputs flagged in review. The gateway's redacted samples are a steady source. Synthetic cases are useful for coverage of edge cases you have not seen yet, but a suite made only of cases someone imagined tests what the team expected, not what users do.

Personal data stays out. Production inputs must be anonymised before they become cases, because the dataset lives in the repository and is sent to model providers on every run. Replace names, addresses and identifiers with realistic placeholders, and review that as carefully as you review the expectation.

Scorers: deterministic first

A scorer takes an output and a case and returns pass or fail with a reason. The reason matters as much as the verdict, because it is what an engineer reads in the CI report when the gate blocks their change.

evals/scorers.ts
const normalise = (s: string) => s.normalize('NFKC').trim().toLowerCase().replace(/\s+/g, ' ');

/** For classification-style tasks with a single correct label. */
export const exactLabel: SyncScorer = (output, c) => {
  const ok = normalise(output) === normalise(String(c.expected));
  return { pass: ok, reason: ok ? 'label matches' : `expected "${String(c.expected)}", got "${output.slice(0, 80)}"` };
};

/**
 * For extraction: the output must be JSON, and every expected field must
 * match. Extra fields are allowed; missing or wrong ones are named.
 */
export const jsonFields: SyncScorer = (output, c) => {
  let parsed: Record<string, unknown>;
  try {
    const v: unknown = JSON.parse(output);
    if (typeof v !== 'object' || v === null || Array.isArray(v)) return { pass: false, reason: 'output is not a JSON object' };
    parsed = v as Record<string, unknown>;
  } catch {
    return { pass: false, reason: 'output is not valid JSON' };
  }
  const expected = (c.expected ?? {}) as Record<string, unknown>;
  const wrong = Object.entries(expected)
    .filter(([k, v]) => JSON.stringify(parsed[k]) !== JSON.stringify(v))
    .map(([k]) => k);
  return wrong.length ? { pass: false, reason: `wrong or missing: ${wrong.join(', ')}` } : { pass: true, reason: 'all fields match' };
};

/** Hard constraints that must hold whatever the content, e.g. "never mention a price". */
export function forbids(pattern: RegExp, label: string): SyncScorer {
  return (output) => (pattern.test(output) ? { pass: false, reason: `contains ${label}` } : { pass: true, reason: `no ${label}` });
}

The order of preference is simple: if code can check it, code should check it.

Exact labels for classification: routing a ticket, detecting a language, choosing an intent. Normalise whitespace, case and Unicode form, and nothing more. If the model returns "Billing." with a full stop, that is a format problem worth knowing about, not something to hide with a fuzzy comparison.

Structured fields for extraction. Parse the output as JSON and compare the fields you care about. Extra fields are allowed, so the expectation does not have to change every time the schema grows. The reason names the fields that are wrong, which is far more useful than "output did not match".

Hard constraints that must hold whatever the content: never quote a price, never include an internal URL, never answer in English to a Swedish question. These are cheap, precise and catch exactly the regressions that reach customers.

evals/scorers.ts
/** A case passes only if every scorer passes. Reasons are kept for the report. */
export async function scoreAll(output: string, c: EvalCase, scorers: Scorer[]): Promise<Score> {
  const results = await Promise.all(scorers.map((s) => s(output, c)));
  const failed = results.filter((r) => !r.pass);
  return failed.length ? { pass: false, reason: failed.map((r) => r.reason).join('; ') } : { pass: true, reason: 'ok' };
}

A case passes only if every scorer passes, and the report keeps all the reasons. Deterministic scorers are synchronous by type, which is a small but useful signal in review: anything returning a promise involves a model or a network, and deserves a second look.

The LLM judge, and how to earn the right to use it

Some criteria cannot be checked by code. Is the answer polite? Does it stay within the refund policy? Does it actually answer the question? For these, a second model grading the first is the practical option, and it is also the easiest part of an eval suite to get wrong.

evals/judge.ts
/**
 * LLM-as-judge for criteria that code cannot check ("is the tone polite",
 * "does the answer stay within the policy"). The judge answers in strict
 * JSON with a verdict and a reason; anything else counts as a failed
 * judgement, never as a pass.
 */
export function llmJudge(complete: Complete, rubric: string): Scorer {
  const system = [
    'You are grading the output of another system against a rubric.',
    'Respond with JSON only: {"reason": string, "verdict": "pass" | "fail"}.',
    'Write the reason first, then the verdict. Grade strictly; if unsure, fail.',
  ].join('\n');
  return async (output: string, c: EvalCase): Promise<Score> => {
    const user = `Rubric:\n${rubric}\n\nInput:\n${c.input}\n\nOutput to grade:\n${output}`;
    const raw = await complete(system, user);
    return parseVerdict(raw);
  };
}

export function parseVerdict(raw: string): Score {
  const json = raw.trim().replace(/^```(?:json)?\s*|\s*```$/g, '');
  try {
    const v = JSON.parse(json) as { verdict?: unknown; reason?: unknown };
    if ((v.verdict === 'pass' || v.verdict === 'fail') && typeof v.reason === 'string') {
      return { pass: v.verdict === 'pass', reason: `judge: ${v.reason}` };
    }
  } catch { /* fall through */ }
  return { pass: false, reason: 'judge returned an unparseable verdict' };
}

The judge prompt asks for JSON with the reason before the verdict, so the model commits to an argument before a conclusion. The parser fails closed: anything other than a well-formed verdict counts as a failure. A judge that says "PASS" in plain text, wraps its JSON in prose or invents a third verdict has not passed the case.

Judges have documented biases. Zheng and colleagues found, in their study of LLM-as-a-judge, that judge models can prefer the first of two answers in pairwise comparisons, favour longer answers and favour output from their own model family. Three mitigations follow:

  • Prefer single-output grading against a rubric to pairwise comparison where you can. It removes position bias entirely.
  • When you must compare two outputs, run both orders and only count a preference if both runs agree.
  • Use a different model family for the judge than for the system under test, or at least measure whether it matters.

None of this tells you whether the judge agrees with your team. That requires measurement.

evals/judge.ts
/**
 * Before a judge is allowed to gate releases, compare it with human labels
 * on the same outputs. Raw agreement is misleading when most outputs pass,
 * so we also report Cohen's kappa, which corrects for chance agreement.
 */
export function agreement(human: boolean[], judge: boolean[]): { accuracy: number; kappa: number; n: number } {
  if (human.length !== judge.length || human.length === 0) throw new Error('label arrays must be non-empty and equal length');
  const n = human.length;
  let both = 0, neither = 0, h = 0, j = 0;
  for (let i = 0; i < n; i++) {
    if (human[i]) h++;
    if (judge[i]) j++;
    if (human[i] && judge[i]) both++;
    if (!human[i] && !judge[i]) neither++;
  }
  const po = (both + neither) / n;
  const pe = (h / n) * (j / n) + ((n - h) / n) * ((n - j) / n);
  const kappa = pe === 1 ? 1 : (po - pe) / (1 - pe);
  return { accuracy: po, kappa, n };
}

Before a judge is allowed to gate releases, label a sample of real outputs by hand, a hundred or two is a reasonable start, run the judge on the same outputs and compare. Report two numbers.

Raw agreement is the share of outputs where the judge and the humans agree. It is intuitive and misleading: if 90 % of outputs pass, a judge that always says "pass" agrees 90 % of the time while detecting nothing.

Cohen's kappa corrects for the agreement you would expect by chance, given how often each side says pass. In the test for this code, a judge that agrees with humans on 8 out of 10 outputs scores a kappa of 0.375, because most of that agreement is what two raters who mostly say "pass" would reach by luck. A judge that will block merges should reach substantially higher agreement than that on your own data. Where you set the bar is a product decision; write it down next to the rubric.

Running the suite

The runner is deliberately boring.

evals/runner.ts
/**
 * Runs every case with bounded concurrency. Bounded, because the eval run
 * shares rate limits with production traffic on the same API key.
 */
export async function runSuite(cases: EvalCase[], system: System, scorers: Scorer[], concurrency = 4): Promise<CaseResult[]> {
  const results: CaseResult[] = new Array(cases.length);
  let next = 0;
  async function worker() {
    while (next < cases.length) {
      const i = next++;
      const c = cases[i]!;
      let output = '';
      let score: Score;
      try {
        output = await system(c.input);
        score = await scoreAll(output, c, scorers);
      } catch (err) {
        // A crash is a failure of the system under test, not of the harness.
        score = { pass: false, reason: `error: ${err instanceof Error ? err.message : String(err)}` };
      }
      results[i] = { id: c.id, tags: c.tags, output, ...score };
    }
  }
  await Promise.all(Array.from({ length: Math.min(concurrency, cases.length) }, worker));
  return results;
}

Concurrency is bounded, because an eval run shares rate limits with production traffic if it uses the same account. Route eval traffic through the gateway with its own tenant id and budget, so a large run can never starve customers.

A crash is a failure of the system under test. If the system throws on a case, because of a timeout, a refused request or an unparseable response, that case fails with the error as its reason. The harness itself does not crash halfway and leave you with half a report.

Temperature is not your friend here. If the system samples at a non-zero temperature, the same case can pass on one run and fail on the next. Either evaluate at the temperature you run in production and repeat each case several times, or pin the temperature to zero for evaluation and accept that you are testing a slightly different system. Both are defensible. Not deciding is not.

The gate: paired comparison, three rules

The naive gate compares the candidate's pass rate with a fixed threshold, or with the baseline's pass rate. Both are blunt instruments. A fixed threshold blocks changes that improve things when the suite is hard, and lets real regressions through when it is easy. Comparing two pass rates as if they were independent samples ignores the most useful fact you have: both versions ran on the same cases.

evals/runner.ts
export interface GatePolicy {
  /** Cases with this tag must pass on the candidate if they passed on the baseline. */
  criticalTag: string;
  /** Block if the candidate is significantly worse at this level. */
  alpha: number;
  /** Block if the overall pass rate drops by more than this, significant or not. */
  maxDrop: number;
}

export interface GateDecision {
  pass: boolean;
  reasons: string[];
  baseline: { rate: number; low: number; high: number };
  candidate: { rate: number; low: number; high: number };
  regressions: string[];
  fixes: string[];
  pValue: number;
}

/**
 * Compares a candidate run with the baseline run on the same cases.
 * Three independent reasons to block a release:
 *   1. any critical case regressed,
 *   2. the candidate is significantly worse (paired McNemar test),
 *   3. the pass rate fell more than the allowed margin.
 */
export function gate(baseline: CaseResult[], candidate: CaseResult[], policy: GatePolicy): GateDecision {
  const before = new Map(baseline.map((r) => [r.id, r]));
  const regressions: string[] = [], fixes: string[] = [], criticalRegressions: string[] = [];
  for (const r of candidate) {
    const b = before.get(r.id);
    if (!b) continue; // new case: no baseline to compare with
    if (b.pass && !r.pass) {
      regressions.push(r.id);
      if (r.tags.includes(policy.criticalTag)) criticalRegressions.push(r.id);
    }
    if (!b.pass && r.pass) fixes.push(r.id);
  }
  const rate = (rs: CaseResult[]) => {
    const passes = rs.filter((r) => r.pass).length;
    return { rate: rs.length ? passes / rs.length : 0, ...wilson(passes, rs.length) };
  };
  const base = rate(baseline), cand = rate(candidate);
  const pValue = mcnemarWorse(regressions.length, fixes.length);

  const reasons: string[] = [];
  if (criticalRegressions.length) reasons.push(`critical cases regressed: ${criticalRegressions.join(', ')}`);
  if (pValue < policy.alpha) reasons.push(`significantly worse (McNemar p = ${pValue.toFixed(4)})`);
  if (base.rate - cand.rate > policy.maxDrop) reasons.push(`pass rate fell ${((base.rate - cand.rate) * 100).toFixed(1)} points`);
  return { pass: reasons.length === 0, reasons, baseline: base, candidate: cand, regressions, fixes, pValue };
}

The gate blocks a merge for three independent reasons.

1. A critical case regressed. If a case tagged critical passed on the baseline and fails on the candidate, the change is blocked, whatever happened to the overall rate. Critical cases encode things like "never disclose another customer's data" or "always escalate a threat of self-harm". An average that improves elsewhere does not compensate for them.

2. The candidate is significantly worse. The gate uses the exact McNemar test on paired results. Only cases that changed outcome carry information: b regressions that passed before and fail now, and c fixes that failed before and pass now. If the candidate were no worse, a case that changed would be equally likely to go either way, so b follows a binomial distribution with n = b + c and probability one half. The gate blocks when the one-sided p-value for "worse" falls below the chosen significance level.

3. The pass rate fell more than the agreed margin. Significance tests answer "is this real?", not "does it matter?". A tiny suite can hide a large drop behind a high p-value. The margin is a product decision: how much worse are you willing to let a change make things, for the sake of whatever else it improves?

evals/stats.ts
/**
 * Exact McNemar test for paired results: the same cases, run on the baseline
 * and on the candidate. Only discordant pairs carry information:
 *   b = passed before, fails now (regressions)
 *   c = failed before, passes now (fixes)
 * Under "no difference", b ~ Binomial(b + c, 0.5). Returns the one-sided
 * p-value for "the candidate is worse".
 */
export function mcnemarWorse(b: number, c: number): number {
  const n = b + c;
  if (n === 0) return 1;
  // P(X >= b) for X ~ Bin(n, 0.5), summed in log space to avoid overflow.
  let p = 0;
  for (let k = b; k <= n; k++) p += Math.exp(logChoose(n, k) - n * Math.LN2);
  return Math.min(1, p);
}

function logChoose(n: number, k: number): number {
  let s = 0;
  for (let i = 1; i <= k; i++) s += Math.log(n - k + i) - Math.log(i);
  return s;
}

The test is easy to compute exactly for the sizes an eval suite produces. The sum runs in log space so that large n does not overflow. Concretely: if a candidate regresses eight cases and fixes none, p = 1/256 ≈ 0.004, and the gate blocks. If it regresses three and fixes three, p ≈ 0.66, and the gate lets it through, because trading three cases for three others is not evidence of anything.

That second example is the strength of a paired test. On a suite of 200 cases where 150 pass on both versions and 44 fail on both, those 194 cases say nothing about the difference between the versions. Only the six that changed do.

Reporting uncertainty

The gate makes a decision; the report explains it. Every pass rate in the report should come with an interval, because a suite is a sample of the inputs your system will see.

evals/stats.ts
/**
 * Wilson score interval for a pass rate. Unlike the naive p ± 1.96·√(p(1-p)/n)
 * it behaves at small n and near 0 or 1, which is exactly where eval suites
 * live (n = 200, pass rate 0.97).
 */
export function wilson(passes: number, n: number, z = 1.96): { low: number; high: number } {
  if (n === 0) return { low: 0, high: 1 };
  const p = passes / n;
  const denom = 1 + (z * z) / n;
  const centre = (p + (z * z) / (2 * n)) / denom;
  const half = (z * Math.sqrt((p * (1 - p)) / n + (z * z) / (4 * n * n))) / denom;
  return { low: Math.max(0, centre - half), high: Math.min(1, centre + half) };
}

Use the Wilson score interval rather than the textbook normal approximation. The normal approximation behaves badly with small samples and near 0 % or 100 %, which is exactly where eval suites live. For 97 passes out of 100, the Wilson interval is roughly 91.5 % to 99.0 %. A dashboard that shows "97 %" without that range invites people to celebrate or panic over moves that are noise.

Two refinements matter as suites mature. If cases are clustered, for example several questions about the same document, they are not independent, and intervals should be computed with clustered standard errors; Evan Miller's paper on error bars for evals, in the references, covers this. And if you repeat each case several times to handle sampling randomness, analyse the repetitions as repeated measurements of one case, not as additional cases.

Wiring it into CI

The pieces fit into an ordinary CI job. On every pull request that touches prompts, routes, retrieval or post-processing:

  1. Load the dataset at the pull request's commit and record its hash.
  2. Get the baseline results. Run the suite on the main branch, or reuse cached results for main's commit and dataset hash, so the baseline is not re-run on every pull request.
  3. Run the candidate from the branch, through the gateway, with an eval tenant and budget.
  4. Score and gate. Apply the policy and post the decision as a comment on the pull request: the three rule outcomes, the pass rates with intervals, every regression with its reason and every fix.
  5. Fail the check if the gate fails. Overrides are allowed, but explicit: a label on the pull request that someone has to add, and that shows up in review.

Run the full suite on a schedule as well, against production configuration. Providers update models behind stable names more often than their changelogs suggest, and a scheduled run is how you find out on a Tuesday morning instead of from a customer.

Testing the eval code itself

An eval harness that is wrong is worse than none, because it gives confidence it has not earned. The code in this article has thirteen tests. A few of them are worth copying into any harness:

  • The statistics against known values. The Wilson interval for 97 of 100, and the McNemar p-value for eight regressions and one fix, which is exactly 10/512.
  • The judge fails closed. Plain-text verdicts, unknown verdicts and malformed JSON all count as failures.
  • Kappa on a skewed example. 80 % raw agreement, kappa 0.375. If this test ever passes with a higher kappa, the formula is wrong.
  • A critical regression blocks even when the overall rate improves. This is the rule people are most tempted to weaken, so it deserves a test that fails loudly.

What to build first

If you have no evals today, do not start with a judge or with statistics. Start with fifty real cases from production, anonymised, with the deterministic scorers that fit your task, and run them on every change to a prompt. The first time the suite stops a change that looked obviously better, the rest of this article will be easier to justify.

Then add the pieces in this order: tags and a critical slice, the paired gate, intervals in the report, and finally a judge for the criteria code cannot check, once you have measured it against your team.

The first article in this series, a reference architecture for LLM systems in production, covers the gateway that makes every call traceable to a prompt version and a provider. The whole series is on the Luniat Engineering page.

References

  1. Lianmin Zheng et al., "Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena", NeurIPS Datasets and Benchmarks, 2023.
  2. Evan Miller, "Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations", 2024.
  3. Thomas G. Dietterich, "Approximate Statistical Tests for Comparing Supervised Classification Learning Algorithms", Neural Computation 10(7), 1998.
  4. Quinn McNemar, "Note on the sampling error of the difference between correlated proportions or percentages", Psychometrika 12(2), 1947.
  5. Lawrence D. Brown, T. Tony Cai and Anirban DasGupta, "Interval Estimation for a Binomial Proportion", Statistical Science 16(2), 2001.
  6. Jacob Cohen, "A Coefficient of Agreement for Nominal Scales", Educational and Psychological Measurement 20(1), 1960.
Read next →Retrieval that holds up: hybrid search, chunking and citations in PostgresEngineering · No. 03 · 14 min

Frequently asked questions

What is an eval in an LLM system?
A test case for behaviour that is not fully deterministic. You run a fixed set of inputs through the system, score each output against an expectation, and compare the results between versions before you ship a change.
How many eval cases do I need?
Enough to detect the regressions you care about. With 50 cases a single regression moves the pass rate by two points and is hard to tell from noise; a few hundred cases, tagged by slice, is a practical starting point for a gate. Critical behaviours need dedicated cases regardless of size.
Can I trust an LLM as a judge?
Only after you have measured it. Label a sample of outputs by hand, compare with the judge and report chance-corrected agreement such as Cohen's kappa. Use the judge only for criteria code cannot check, and fail closed when its answer cannot be parsed.
Why compare against a baseline instead of a fixed threshold?
Because the same cases run on two versions are paired. A paired test uses only the cases that changed outcome, which makes it far more sensitive than comparing two pass rates, and it ignores cases that were hard for both versions.