Structured output and tool calls that do not break production

Validate model output against a schema, repair it with exact errors, and run tool calls with permissions, validation and idempotency keys. Tested code.

Level
Advanced
Stack
TypeScript, JSON Schema, Any LLM with tool use

In short

  • Constrained decoding guarantees shape, not correctness: validate against the schema and business rules every time.
  • Repair with the model's own output and every error with a path, then stop after a small limit with a typed error.
  • Treat tool calls as requests: check existence, permission, arguments and tenant on every call, and return errors as data.
  • Make every write idempotent on tenant, tool and call id, so a retried turn can never repeat a side effect.
In this article · 8 sections
  1. 01Two different guarantees
  2. 02A schema you can validate against
  3. 03Generate, validate, repair
  4. 04Tool calls are requests, not commands
  5. 05Idempotency: at most once, even when retried
  6. 06Testing output and tools
  7. 07A checklist for structured output and tools
  8. 08References

The moment an LLM's output is consumed by code instead of read by a person, the rules change. A human reader tolerates "Sure! Here's the JSON you asked for:" and a missing field. A parser does not. And the moment a model can call tools, its output is no longer just text: it is a request to change something in the world, made by a component that can be steered by any text it reads.

This article covers both: getting output you can trust from a model, and executing the actions it proposes without letting it do more than it should. As in the rest of the series, the code is type-checked and tested in CI before this page is built.

Two different guarantees

It helps to separate two properties that are often blurred together.

Syntactic validity: the output parses as JSON and has the right shape. Fields exist, types match, enums hold. Many providers can now guarantee this with constrained decoding, where the model is only allowed to produce tokens that keep the output valid against a schema.

Semantic validity: the content is correct. The invoice total equals the sum of the lines. The due date is after the invoice date. The customer id belongs to this tenant. No decoder can guarantee this, because it depends on facts outside the output.

Constrained decoding is worth using where it is available, because it removes a whole class of failures cheaply. It does not remove the need to validate. The validator is also what protects you when you switch provider or model, when a setting is changed, or when the output is produced by a fallback route that does not support the same feature.

A schema you can validate against

We use a small subset of JSON Schema: objects with required and forbidden properties, arrays with bounds, strings with enums, patterns and lengths, numbers with ranges, and anyOf for nullable fields. It is enough for almost every extraction or classification task, and small enough to read in one sitting.

structured/schema.ts
/** The subset of JSON Schema that model output schemas actually need. */
export type Schema =
  | { type: 'object'; properties: Record<string, Schema>; required?: string[]; additionalProperties?: false; description?: string }
  | { type: 'array'; items: Schema; minItems?: number; maxItems?: number; description?: string }
  | { type: 'string'; enum?: string[]; pattern?: string; maxLength?: number; description?: string }
  | { type: 'number' | 'integer'; minimum?: number; maximum?: number; description?: string }
  | { type: 'boolean'; description?: string }
  | { type: 'null' }
  | { anyOf: Schema[]; description?: string };

export interface ValidationError {
  /** JSON Pointer to the offending value, e.g. /lines/2/amount */
  path: string;
  message: string;
}
structured/schema.ts
/**
 * Returns every error, not just the first: the list is fed back to the model
 * in a repair attempt, and fixing all problems at once saves a round trip.
 */
export function validate(value: unknown, schema: Schema, path = ''): ValidationError[] {
  const err = (message: string): ValidationError[] => [{ path: path || '/', message }];
  if ('anyOf' in schema) {
    const results = schema.anyOf.map((s) => validate(value, s, path));
    return results.some((r) => r.length === 0) ? [] : err(`does not match any allowed shape`);
  }
  switch (schema.type) {
    case 'null':
      return value === null ? [] : err('expected null');
    case 'boolean':
      return typeof value === 'boolean' ? [] : err('expected boolean');
    case 'number':
    case 'integer': {
      if (typeof value !== 'number' || !Number.isFinite(value)) return err(`expected ${schema.type}`);
      if (schema.type === 'integer' && !Number.isInteger(value)) return err('expected integer');
      if (schema.minimum !== undefined && value < schema.minimum) return err(`must be >= ${schema.minimum}`);
      if (schema.maximum !== undefined && value > schema.maximum) return err(`must be <= ${schema.maximum}`);
      return [];
    }
    case 'string': {
      if (typeof value !== 'string') return err('expected string');
      if (schema.enum && !schema.enum.includes(value)) return err(`must be one of ${schema.enum.join(', ')}`);
      if (schema.maxLength !== undefined && value.length > schema.maxLength) return err(`longer than ${schema.maxLength}`);
      if (schema.pattern && !new RegExp(schema.pattern, 'u').test(value)) return err(`does not match ${schema.pattern}`);
      return [];
    }
    case 'array': {
      if (!Array.isArray(value)) return err('expected array');
      const out: ValidationError[] = [];
      if (schema.minItems !== undefined && value.length < schema.minItems) out.push(...err(`fewer than ${schema.minItems} items`));
      if (schema.maxItems !== undefined && value.length > schema.maxItems) out.push(...err(`more than ${schema.maxItems} items`));
      value.forEach((v, i) => out.push(...validate(v, schema.items, `${path}/${i}`)));
      return out;
    }
    case 'object': {
      if (typeof value !== 'object' || value === null || Array.isArray(value)) return err('expected object');
      const obj = value as Record<string, unknown>;
      const out: ValidationError[] = [];
      for (const k of schema.required ?? []) if (!(k in obj)) out.push({ path: `${path}/${k}`, message: 'is required' });
      for (const [k, v] of Object.entries(obj)) {
        const sub = schema.properties[k];
        if (sub) out.push(...validate(v, sub, `${path}/${k}`));
        else if (schema.additionalProperties === false) out.push({ path: `${path}/${k}`, message: 'is not allowed' });
      }
      return out;
    }
  }
}

Two properties of this validator matter more than its completeness.

It returns every error, with a path. /lines/1/amount: is required tells both a developer and a model exactly what to fix. A validator that stops at the first error turns a single repair into several round trips.

It is strict about extra properties when the schema says so. additionalProperties: false catches a model inventing a field, which is often the first sign that it misunderstood the task. Allow extras only where you have a reason to.

Generate, validate, repair

The generation loop is short, and every line is there for a reason.

structured/generate.ts
/**
 * Models wrap JSON in prose or code fences more often than anyone would like.
 * Take the outermost object; anything else is a parse failure, not a guess.
 */
export function extractJson(raw: string): unknown {
  const fenced = raw.match(/```(?:json)?\s*([\s\S]*?)```/);
  const text = (fenced ? fenced[1]! : raw).trim();
  const start = text.indexOf('{');
  const end = text.lastIndexOf('}');
  if (start === -1 || end < start) throw new SyntaxError('no JSON object in output');
  return JSON.parse(text.slice(start, end + 1));
}

Extract, do not guess. Models wrap JSON in prose or code fences more often than anyone would like, even when told not to. Extraction takes the outermost object and nothing else. If there is no object, that is a failure, not a reason to try a regular expression on the prose.

structured/generate.ts
/**
 * Generate → parse → validate → repair. On failure the model sees its own
 * output and the exact list of errors, which fixes most problems in one
 * extra call. Business rules that a schema cannot express go in `check`.
 */
export async function generateStructured<T>(o: {
  complete: Complete;
  system: string;
  input: string;
  schema: Schema;
  check?: (value: T) => ValidationError[];
  maxAttempts?: number;
}): Promise<{ value: T; attempts: number }> {
  const max = o.maxAttempts ?? 2;
  const system = `${o.system}\n\nRespond with a single JSON object that matches this JSON Schema, and nothing else:\n${JSON.stringify(o.schema)}`;
  const messages: Message[] = [{ role: 'user', content: o.input }];
  let errors: ValidationError[] = [];

  for (let attempt = 1; attempt <= max; attempt++) {
    const raw = await o.complete(system, messages);
    let value: unknown;
    try {
      value = extractJson(raw);
      errors = validate(value, o.schema);
      if (!errors.length && o.check) errors = o.check(value as T);
    } catch (e) {
      errors = [{ path: '/', message: `invalid JSON: ${(e as Error).message}` }];
    }
    if (!errors.length) return { value: value as T, attempts: attempt };
    messages.push(
      { role: 'assistant', content: raw },
      { role: 'user', content: `Your output failed validation:\n${errors.map((e) => `- ${e.path}: ${e.message}`).join('\n')}\nReturn the corrected JSON object only.` },
    );
  }
  throw new StructuredOutputError(`no valid output after ${max} attempts`, errors, max);
}

The schema is in the system prompt. Even with constrained decoding, the model writes better values when it can see the schema's descriptions and constraints.

Validation includes business rules. The check function takes the parsed value and returns errors in the same format. Rules that a schema cannot express, such as a due date after the invoice date or line amounts that sum to the total, live here, and the model sees their errors just like schema errors.

Repair with the exact errors. On failure, the conversation is extended with the model's own output and a list of errors with paths. Structural problems, such as a missing field or a wrong format, are usually fixed in one repair attempt. The tests assert that the second call receives /invoiceNumber: does not match ^[A-Z]-\d+$, which is the whole point.

Then stop. After the limit the loop throws a StructuredOutputError with the last errors and the number of attempts. Retrying indefinitely turns a bad input, such as a scanned document with no invoice in it, into an unbounded cost. The caller decides what to do: route to a person, ask the user, or mark the document as unprocessable.

Where repair belongs

The loop sits above the gateway from the first article, not inside it. The gateway treats output that fails validation as a permanent error at its layer, because whether and how to repair depends on the feature. An extraction pipeline repairs; a chat feature might show the raw text; a classification step might fall back to a default label. Keeping that decision in the caller keeps the gateway simple and predictable.

Tool calls are requests, not commands

When a model has tools, it does not execute anything. It produces a structured request: call this tool, with these arguments. Your code decides whether to honour it. That decision point is the most important security boundary in an agentic system, and it deserves the same rigour as an API endpoint, because that is what it is.

The model proposes a tool call. The registry checks that the tool exists, that it is permitted for this caller, that the arguments are valid, and for writes claims an idempotency key, before the tool runs with the tenant from the session. Errors return to the model as structured data.
Figure 1. Five checks between a proposed call and a side effect. Each failure is reported back to the model as data it can act on.
structured/tools.ts
export interface Tool<A = unknown, R = unknown> {
  name: string;
  /** Read by the model. Treat it as part of the prompt: versioned and reviewed. */
  description: string;
  input: Schema;
  /** Writes change the world and get stricter handling than reads. */
  effect: 'read' | 'write';
  run(args: A, ctx: ToolContext): Promise<R>;
}

export interface ToolContext {
  /** From the authenticated session, never from the model's arguments. */
  tenantId: string;
  signal: AbortSignal;
}

/** What the model gets back. Errors are data the model can act on, not exceptions. */
export type ToolResult =
  | { ok: true; content: unknown; replayed?: boolean }
  | { ok: false; error: 'unknown_tool' | 'invalid_arguments' | 'not_permitted' | 'in_progress' | 'conflict' | 'failed'; detail: string };

A tool has a name, a description, an input schema, an effect and a function. Two fields deserve attention.

The description is part of the prompt. The model decides when and how to call a tool largely from its description. Changing it changes behaviour, so version it and run it through your evals like any other prompt change.

The effect is explicit. A tool is either a read or a write. Writes change the world: they send messages, create records, move money. They get stricter handling, and the distinction is visible in review.

structured/tools.ts
export class ToolRegistry {
  private readonly tools = new Map<string, Tool<never, unknown>>();

  constructor(private readonly idempotency: IdempotencyStore, private readonly timeoutMs = 10_000) {}

  register<A, R>(tool: Tool<A, R>): void {
    this.tools.set(tool.name, tool as unknown as Tool<never, unknown>);
  }

  /** Tool definitions for the provider API, filtered by what this caller may use. */
  definitions(allowWrites: boolean) {
    return [...this.tools.values()]
      .filter((t) => allowWrites || t.effect === 'read')
      .map((t) => ({ name: t.name, description: t.description, input_schema: t.input }));
  }

  async execute(call: ToolCall, ctx: ToolContext & { allowWrites: boolean }): Promise<ToolResult> {
    const tool = this.tools.get(call.name);
    if (!tool) return { ok: false, error: 'unknown_tool', detail: `no tool named ${call.name}` };
    // Enforced here, not only by hiding the definition: models can and do call tools they were not offered.
    if (tool.effect === 'write' && !ctx.allowWrites) return { ok: false, error: 'not_permitted', detail: `${call.name} is not allowed here` };

    const errors = validate(call.arguments, tool.input);
    if (errors.length) return { ok: false, error: 'invalid_arguments', detail: errors.map((e) => `${e.path}: ${e.message}`).join('; ') };

    // Writes run at most once per tool call id, even if the turn is retried.
    if (tool.effect === 'write') {
      const begin = this.idempotency.begin(ctx.tenantId, `${call.name}:${call.id}`, call.arguments);
      if (begin.kind === 'replay') return { ok: true, content: begin.result, replayed: true };
      if (begin.kind === 'in-progress') return { ok: false, error: 'in_progress', detail: 'this call is already running' };
      if (begin.kind === 'conflict') return { ok: false, error: 'conflict', detail: 'call id reused with different arguments' };
    }

    const signal = AbortSignal.any([ctx.signal, AbortSignal.timeout(this.timeoutMs)]);
    try {
      const content = await tool.run(call.arguments as never, { tenantId: ctx.tenantId, signal });
      if (tool.effect === 'write') this.idempotency.complete(ctx.tenantId, `${call.name}:${call.id}`, content);
      return { ok: true, content };
    } catch (e) {
      if (tool.effect === 'write') this.idempotency.abandon(ctx.tenantId, `${call.name}:${call.id}`);
      return { ok: false, error: 'failed', detail: e instanceof Error ? e.message : String(e) };
    }
  }
}

The registry runs five checks, in order, on every call.

  1. The tool exists. Models occasionally call tools that were never defined, sometimes with plausible names. The answer is an error message the model can read, not an exception.
  2. The caller may use it. A read-only context, such as answering a question in a help widget, must not be able to send email. Hiding the tool from the definitions is not enough: a retrieved document or a user message can name a tool the model was not offered, and some models will call it. Permission is enforced where the tool runs.
  3. The arguments are valid, against the tool's schema, with path-level errors returned to the model so it can correct itself.
  4. Writes claim an idempotency key before running, covered in the next section.
  5. The tool runs with the tenant from the session, with a timeout and the caller's abort signal. The model's arguments never decide which tenant's data a tool touches. The test for this passes a tenantId in the arguments and asserts that the tool ignored it.

Errors as data

Every failure in the registry returns a structured ToolResult with an error code and a detail, instead of throwing. That is deliberate. The model is the one that made the call, and it is often able to fix it: correct an argument, choose a different tool, or tell the user it cannot complete the request. An exception that aborts the whole turn throws that ability away, and a generic "something went wrong" teaches the model nothing.

The error codes are a closed set: unknown_tool, not_permitted, invalid_arguments, in_progress, conflict and failed. A closed set is easy to count in telemetry, and a spike in one of them usually points straight at the cause: invalid_arguments after a prompt change, not_permitted after something started steering the model towards tools it should not use.

Idempotency: at most once, even when retried

Retries are everywhere in an LLM system. The gateway retries provider errors, a client retries a timed-out request, an orchestrator replays a turn after a crash. For text generation that is harmless. For a tool that sends an email or creates an order, a retry is a duplicate.

The fix is old and well understood from payment APIs: an idempotency key.

structured/idempotency.ts
/**
 * At-most-once execution for side effects. The key is scoped to the tenant,
 * and the request body is hashed: the same key with a different body is a
 * client bug and must fail loudly instead of returning someone else's result.
 * In production this lives in Redis or a database with a unique constraint;
 * the protocol is the same.
 */
export class IdempotencyStore {
  private readonly entries = new Map<string, Entry>();

  constructor(private readonly ttlMs: number, private readonly now: () => number = Date.now) {}

  static hash(body: unknown): string {
    return createHash('sha256').update(JSON.stringify(body)).digest('hex');
  }

  begin(tenantId: string, key: string, body: unknown): Begin {
    const id = `${tenantId}:${key}`;
    const requestHash = IdempotencyStore.hash(body);
    const e = this.entries.get(id);
    if (e && e.expiresAt > this.now()) {
      if (e.requestHash !== requestHash) return { kind: 'conflict' };
      return e.state === 'done' ? { kind: 'replay', result: e.result } : { kind: 'in-progress' };
    }
    this.entries.set(id, { requestHash, state: 'running', expiresAt: this.now() + this.ttlMs });
    return { kind: 'new' };
  }

  complete(tenantId: string, key: string, result: unknown): void {
    const e = this.entries.get(`${tenantId}:${key}`);
    if (e) Object.assign(e, { state: 'done', result });
  }

  /** On failure, release the key so a retry can run the operation again. */
  abandon(tenantId: string, key: string): void {
    this.entries.delete(`${tenantId}:${key}`);
  }
}

The protocol has four outcomes:

Outcome When What the registry does
new First time this key is seen Runs the tool and stores the result
replay Same key, same arguments, finished Returns the stored result without running
in-progress Same key, still running Returns an error; do not start a second run
conflict Same key, different arguments Returns an error; a client bug, never ignored

The key is derived, not invented. For tool calls we use the tenant, the tool name and the call id the provider assigned. When a turn is retried with the same model output, the call ids are the same and the side effect does not repeat. When the model is called again and proposes a new call, it gets a new id, which is correct: it is a new decision.

The arguments are hashed. The same key with different arguments means something is wrong upstream. Returning the old result would hide the bug; running the new request would break the guarantee. Failing loudly is the only safe option.

Failure releases the key. If the tool throws, the key is abandoned so that a retry can run it again. A failed side effect that blocked all retries would turn a transient error into a permanent one.

Idempotency in your own code does not make the external system idempotent. If the tool calls a third-party API, pass the same key on to that API if it supports one, so a retry after a network timeout between you and them is also safe.

Testing output and tools

Both halves of this article are easy to test without a model, because the model is behind a function.

  • Schema errors with paths. One invalid object with five different problems, asserted as five specific errors. If the validator ever stops reporting all of them, the repair loop gets slower without anyone noticing.
  • Repair feeds back the exact errors. A scripted model returns an invalid object, then a valid one; the test asserts that the second call saw the error path in its prompt.
  • Business rules fail closed. A check that rejects a past due date, with the loop giving up after the limit and reporting the error.
  • Writes run once. The same call executed twice runs the tool once and replays the result; the same id with different arguments is a conflict.
  • Permissions hold against a model that ignores them. A write tool called in a read-only context fails, even though it was never in the definitions.
  • The tenant comes from context. Arguments that name a different tenant are ignored.

Add your tools' real schemas to your eval suite too. A model upgrade that changes how often a tool is called with invalid arguments shows up there long before it shows up in support tickets.

A checklist for structured output and tools

Question What good looks like
Is output validated even with constrained decoding? Yes, against the schema and business rules.
What does the model see when output is invalid? Its own output and every error with a path, once or twice.
What happens after the repair limit? A typed error with the last errors; the caller decides.
Where are tool permissions enforced? Where the tool runs, on every call, not only in the tool list.
Which tenant does a tool act for? The one from the session. Never an argument.
Can a retried turn repeat a side effect? No: writes are keyed on tenant, tool and call id.
What does the model get when a tool fails? A structured error from a closed set, not an exception.

The retrieval layer that feeds many tool-using systems is covered in the RAG article, and the whole series is on the Luniat Engineering page.

References

  1. JSON Schema specification, JSON Schema organisation.
  2. Brandon T. Willard and Rémi Louf, "Efficient Guided Generation for Large Language Models", 2023.
  3. IETF HTTPAPI working group, The Idempotency-Key HTTP Header Field, Internet-Draft.
  4. Stripe, Idempotent requests, API reference.
  5. OWASP, Top 10 for Large Language Model Applications: Improper Output Handling and Excessive Agency.
Read next →Multi-tenant AI without leaks: isolation from the database to the promptEngineering · No. 05 · 11 min

Frequently asked questions

If the provider supports structured outputs, do I still need to validate?
Yes. Constrained decoding guarantees that the output parses and matches the shape, not that it is correct. You still need business rules, such as a due date after the invoice date, and a validator is the only thing that protects you when a provider, model or setting changes.
Should failed validation be retried?
Once or twice, with the exact errors fed back to the model. Structural errors are usually fixed in one repair attempt. After that, fail with a clear error instead of retrying indefinitely.
What is an idempotency key for tool calls?
A key that makes a side effect run at most once. If a turn is retried after a timeout, the same tool call id maps to the same key, and the stored result is returned instead of sending a second email or creating a second order.
Where should tool permissions be enforced?
In the code that executes the tool, on every call. Hiding a tool from the model's tool list is not enough, because models can call tools they were not offered when a prompt or retrieved document tells them to.