A reference architecture for LLM systems in production
One gateway between your services and every model provider, with per-tenant limits, versioned prompts, fallbacks, retries and telemetry. Tested code.
- Level
- Advanced
- Stack
- TypeScript, Node.js 22, Any LLM provider
In short
- Put every model call behind one gateway that owns keys, limits, prompts, routing, retries, breakers and telemetry.
- Limit per tenant in tokens, reserve the worst case and refund the difference, and reject instead of queueing.
- Treat prompts as versioned, hashed artifacts, and only fall back to providers that pass the same evals.
- Classify errors explicitly, retry in one layer only within a shared deadline, and emit one redacted event per call.
In this article · 14 sections
- 01What the gateway owns
- 02The interface
- 03Errors are a classification problem
- 04Stage 1: admission per tenant
- 05Stage 2: prompts as versioned artifacts
- 06Stage 3: routing and fallbacks
- 07Stage 4: circuit breakers
- 08Stage 5: retries within a deadline
- 09Putting it together
- 10Telemetry without leaking data
- 11Testing it without a network
- 12What we left out on purpose
- 13A checklist for your own gateway
- 14References
The first LLM feature in most codebases is a vendor SDK call inside a request handler. It works. By the fifth feature there are five slightly different retry loops, provider keys in three services, no way to tell what a tenant costs, and no record of which prompt produced the output a customer is complaining about.
The fix is not a framework. It is one internal service with a narrow interface that every caller goes through: an LLM gateway. This article describes the gateway we use as a reference, stage by stage, with the code. Every snippet comes from a small TypeScript implementation that is type-checked and tested in CI before this page is built, so what you read is what runs.
The design goals are the usual ones for a shared dependency that is slow, expensive, rate limited and occasionally down:
- No caller can hurt another caller. One tenant's batch job must not exhaust the quota that everyone shares.
- Failures are classified, not guessed. Some errors should be retried, some never, and the difference must be explicit.
- Every call is traceable. For any output you can say which tenant, which prompt version, which provider, how many attempts and what it cost.
- Nothing sensitive leaves through telemetry. Prompts contain whatever users typed.
What the gateway owns
It is easier to agree on a gateway when you list what moves into it, and what each concern looks like when it is scattered across callers instead.
| Concern | Without a gateway | In the gateway |
|---|---|---|
| Provider keys | In every service that calls a model | In one service, never in callers |
| Rate limits | Discovered as 429s in production | Admission per tenant, before the call |
| Prompts | String literals next to the call | Versioned templates with a content hash |
| Provider choice | Hard-coded SDK and model name | Logical model names, routed with fallbacks |
| Outages | Each caller retries on its own schedule | One breaker per provider, one retry policy |
| Cost | Monthly invoice, not attributable | Per call, per tenant, per prompt version |
| Telemetry | Ad hoc logging, often of raw prompts | One redacted event per call |
The rest of this article walks through the stages in the order a request meets them.
The interface
The interface is deliberately small. Callers say who they are acting for, which logical model they want, which rendered prompt to send and how much they are willing to wait and spend. They never name a provider.
/** One request to a model, as the rest of the system sees it. */
export interface ModelRequest {
/** Tenant the request is made on behalf of. Never taken from user input. */
tenantId: string;
/** Logical model name, resolved to a provider by the gateway. */
model: string;
/** Rendered prompt, produced by the prompt registry. */
prompt: RenderedPrompt;
/** Upper bound on output tokens. Also used for rate limiting and cost. */
maxOutputTokens: number;
/** End-to-end deadline for this call, including retries. */
timeoutMs: number;
}
export interface RenderedPrompt {
id: string;
version: number;
/** sha256 of the template, so logs can prove which text was sent. */
hash: string;
system: string;
user: string;
}
export interface Usage {
inputTokens: number;
outputTokens: number;
}
export interface ModelResponse {
text: string;
usage: Usage;
/** Which provider actually served the call (after fallback). */
provider: string;
}
/** A provider adapter wraps one vendor SDK behind a common interface. */
export interface Provider {
readonly name: string;
complete(req: ModelRequest, signal: AbortSignal): Promise<Omit<ModelResponse, 'provider'>>;
}Two details matter more than they look.
tenantId is never taken from user input. It comes from the authenticated session in the product API, or from the job record in a background worker. The gateway uses it to choose a rate-limit bucket and to attribute cost. If a client could set it, one customer could spend another customer's budget and pollute their cost reports. In a multi-tenant system this is the same rule as for database access: the tenant is derived server-side, always.
model is a logical name. Callers ask for default, extraction or long-context, and the gateway maps that to a list of providers. Moving a feature to a different vendor, region or model version becomes a configuration change that you can run through your evals, instead of a code change in every caller.
Errors are a classification problem
Many production incidents in LLM integrations come down to retrying the wrong thing, or not retrying the right one. The gateway makes the classification explicit with three error types.
/** Safe to retry: 429, 5xx, timeouts, connection resets. */
export class RetryableError extends Error {
override readonly name = 'RetryableError';
constructor(message: string, readonly retryAfterMs?: number) {
super(message);
}
}
/** Never retry: 4xx validation errors, auth errors, content policy refusals. */
export class PermanentError extends Error {
override readonly name = 'PermanentError';
}
/** Raised by the gateway itself before a provider is called. */
export class RejectedError extends Error {
override readonly name = 'RejectedError';
constructor(message: string, readonly reason: 'rate_limited' | 'circuit_open' | 'invalid_input') {
super(message);
}
}The mapping from provider responses to these classes is the most important table in the adapter layer. Check it against your provider's documentation; status codes and their meaning vary.
| Response | Class | Why |
|---|---|---|
400, 404, 413, 422 |
Permanent | The request is wrong. Retrying sends the same wrong request. |
401, 403 |
Permanent, and alert | A key is revoked or misconfigured. Retrying hides an incident. |
408, connection reset |
Retryable | Transient. But see the warning below. |
429 |
Retryable, with Retry-After |
The provider is telling you when capacity returns. |
500, 502, 503, 504 |
Retryable | Transient server-side failures. Some providers also use non-standard codes for overload. |
| Content policy refusal | Permanent | The same input will be refused again. Handle it as a product case. |
| Output fails validation | Permanent at this layer | Repairing output is a caller decision, covered in a later article. |
RejectedError is the gateway's own refusal: the tenant is over budget, all circuits are open or the model name is unknown. It is raised before any network call, carries a machine-readable reason, and maps cleanly to a 429 or 503 in the caller's own API.
Stage 1: admission per tenant
Every provider limits your account, not your customers. Your tier might allow a few hundred thousand input tokens per minute, shared by every tenant and every feature. Without admission control, one tenant importing a large document archive will consume the whole budget and every other tenant will see 429s for something they did not do.
The gateway gives each tenant a token bucket, measured in model tokens rather than requests.
/**
* Token bucket measured in model tokens, not requests: one request with a
* 200k-token context costs the provider far more than ten short ones, and
* the provider's own limits are in tokens per minute.
*/
export class TokenBucket {
private tokens: number;
private updatedAt: number;
constructor(private readonly opts: BucketOptions, private readonly clock: Clock) {
this.tokens = opts.capacity;
this.updatedAt = clock.now();
}
private refill(): void {
const now = this.clock.now();
const added = ((now - this.updatedAt) / 1000) * this.opts.refillPerSecond;
this.tokens = Math.min(this.opts.capacity, this.tokens + added);
this.updatedAt = now;
}
/** Takes `cost` tokens if available. Never blocks. */
tryTake(cost: number): boolean {
this.refill();
if (cost > this.tokens) return false;
this.tokens -= cost;
return true;
}
/** Returns tokens when the real usage was lower than the estimate. */
refund(amount: number): void {
this.refill();
this.tokens = Math.min(this.opts.capacity, this.tokens + amount);
}
/** Milliseconds until `cost` tokens are available, for a Retry-After header. */
waitTimeMs(cost: number): number {
this.refill();
if (cost <= this.tokens) return 0;
return Math.ceil(((cost - this.tokens) / this.opts.refillPerSecond) * 1000);
}
}/** One bucket per tenant, so a single tenant can never drain the shared quota. */
export class TenantLimiter {
private readonly buckets = new Map<string, TokenBucket>();
constructor(
private readonly optsFor: (tenantId: string) => BucketOptions,
private readonly clock: Clock,
) {}
bucket(tenantId: string): TokenBucket {
let b = this.buckets.get(tenantId);
if (!b) {
b = new TokenBucket(this.optsFor(tenantId), this.clock);
this.buckets.set(tenantId, b);
}
return b;
}
}Three decisions in this code are worth defending.
Tokens, not requests. A request counter treats a 300-token classification and a 150,000-token document summary as equal. The provider does not. Counting tokens aligns your limit with the provider's limit and with your bill.
Reserve the worst case, then reconcile. Before the call you do not know how many tokens it will use, but you do know the upper bound: the estimated input plus maxOutputTokens. The gateway reserves that, and refunds the difference once the provider reports real usage. Reserving less invites a burst of concurrent calls that each looked affordable. Reserving the bound and refunding keeps the guarantee and costs nothing in the common case. Input estimation does not need a tokenizer per model; a conservative characters-per-token ratio is enough, because the refund corrects it.
Reject, do not queue. When a tenant is over budget the gateway rejects immediately with the time until capacity returns. Queueing inside the gateway looks friendlier but turns overload into latency for everyone, and requests that wait in a queue often time out upstream anyway. Rejecting early with a precise Retry-After lets the caller decide: an interactive feature shows a message, a batch job sleeps exactly as long as it needs to.
Per-tenant budgets also need a global ceiling. The sum of all tenant capacities will usually exceed your provider limit, because not every tenant is active at once. Size the global limit at roughly 80–90 % of the provider's limit and let the per-tenant buckets decide who gets it, so the provider's own 429s become rare instead of routine.
Stage 2: prompts as versioned artifacts
A prompt is code. It changes the behaviour of the system as much as any function, and it deserves the same discipline: review, versioning and an immutable record of what was deployed.
/**
* Prompts are code: versioned, reviewed and immutable once published.
* Every rendered prompt carries id, version and a content hash, so a log
* line or an eval result can always be traced to the exact text sent.
*/
export class PromptRegistry {
private readonly templates = new Map<string, PromptTemplate & { hash: string }>();
register(t: PromptTemplate): void {
const key = `${t.id}@${t.version}`;
const hash = createHash('sha256').update(`${t.system}\u0000${t.user}`).digest('hex');
const existing = this.templates.get(key);
if (existing && existing.hash !== hash) {
throw new Error(`${key} is already registered with different content; bump the version`);
}
this.templates.set(key, { ...t, hash });
}
render(id: string, version: number, vars: Record<string, string>): RenderedPrompt {
const t = this.templates.get(`${id}@${version}`);
if (!t) throw new Error(`unknown prompt ${id}@${version}`);
const needed = new Set([...t.user.matchAll(VAR)].map((m) => m[1]!));
const missing = [...needed].filter((k) => !(k in vars));
const extra = Object.keys(vars).filter((k) => !needed.has(k));
if (missing.length || extra.length) {
throw new Error(`prompt ${id}@${version}: missing [${missing.join(', ')}], unexpected [${extra.join(', ')}]`);
}
const user = t.user.replace(VAR, (_, k: string) => vars[k]!);
return { id, version, hash: t.hash, system: t.system, user };
}
}The registry enforces three rules.
- A version is immutable. Registering the same
id@versionwith different text fails. If you change the text, you bump the version, which shows up in review and in every telemetry event after the deploy. - Every rendered prompt carries a content hash. When a customer reports a bad answer, the event for that call tells you the exact template, not "the prompt as it looked around that time".
- Variables are checked both ways. A missing variable is an obvious bug. An unexpected one usually means the caller and the template have drifted apart, for example after someone renamed
{{ticket}}to{{message}}in the template but not in the call site.
The hash is computed over the system and user templates with a separator, so moving a sentence from one to the other changes the hash, as it should.
Stage 3: routing and fallbacks
A route maps a logical model name to an ordered list of providers and a price table. The first provider is preferred; later ones are fallbacks.
Fallbacks are where well-intentioned gateways create silent quality regressions. Two models that both "support JSON output" can differ in how strictly they follow a schema, how they handle an ambiguous instruction and what they refuse. If the fallback route activates only during an outage, you discover those differences in production, at the worst moment, without any record that a different model answered.
The gateway in this article reports which provider served each call, so the switch is visible. That is necessary but not sufficient. The rule we follow is stricter:
Routes are also where cost becomes visible. The price table is per route, in USD per million input and output tokens, and the gateway computes cost from the usage the provider actually reports, not from estimates.
Stage 4: circuit breakers
When a provider is down, every call to it waits until its timeout and then fails. Without a breaker, those calls tie up connections and worker capacity, and if you have a fallback, every request pays the full timeout before it gets there. A circuit breaker turns a slow failure into a fast one.
/**
* One breaker per provider. While open, calls fail immediately instead of
* queueing behind a provider that is down, which protects both latency and
* the shared rate limit. After the cooldown exactly one probe is allowed;
* its outcome closes or re-opens the circuit.
*/
export class CircuitBreaker {
private failures: number[] = [];
private openedAt: number | null = null;
private probing = false;
constructor(private readonly opts: BreakerOptions, private readonly clock: Clock) {}
state(): BreakerState {
if (this.openedAt === null) return 'closed';
return this.clock.now() - this.openedAt >= this.opts.cooldownMs ? 'half-open' : 'open';
}
/** Returns false if the call must not be attempted. */
tryAcquire(): boolean {
const s = this.state();
if (s === 'closed') return true;
if (s === 'open' || this.probing) return false;
this.probing = true; // half-open: let a single probe through
return true;
}
onSuccess(): void {
this.failures = [];
this.openedAt = null;
this.probing = false;
}
onFailure(): void {
const now = this.clock.now();
if (this.probing) {
this.probing = false;
this.openedAt = now; // probe failed: back to open, restart cooldown
return;
}
this.failures = this.failures.filter((t) => now - t < this.opts.windowMs);
this.failures.push(now);
if (this.failures.length >= this.opts.failureThreshold) this.openedAt = now;
}
}The states are the classic ones. Closed: calls go through and failures are counted in a sliding window. Open: once failures reach the threshold within the window, calls fail immediately for the cooldown period. Half-open: after the cooldown, exactly one probe call is let through; success closes the circuit, failure reopens it and restarts the cooldown.
Two choices are specific to LLM traffic.
One breaker per provider, not per tenant or per caller. The thing that fails is the provider. A per-tenant breaker would need every tenant to discover the outage separately.
Permanent errors do not count. A 400 means the request was wrong, not that the provider is unhealthy. If a bug in one caller produces malformed requests, it must not open the circuit for everyone else. The gateway only reports a failure to the breaker after a retryable error has exhausted its retries on that provider.
Stage 5: retries within a deadline
Retries are necessary and dangerous in equal measure. The policy has three parts: how long to wait, which errors qualify and when to stop.
/**
* "Full jitter" exponential backoff: a random delay between 0 and
* min(cap, base * 2^attempt). Spreads retries from many clients so they
* do not hit the provider in synchronised waves after an outage.
*/
export function backoffDelay(attempt: number, p: RetryPolicy, random: () => number = Math.random): number {
const ceiling = Math.min(p.maxDelayMs, p.baseDelayMs * 2 ** attempt);
return Math.floor(random() * ceiling);
}The wait uses full jitter: a random delay between zero and an exponentially growing ceiling. Without jitter, every client that failed at the same moment retries at the same moment, and the provider sees synchronised waves exactly as it is trying to recover. Full jitter spreads them out. The AWS Architecture Blog post on backoff and jitter, listed in the references, compares the variants with simulations.
/**
* Retries only RetryableError. A Retry-After from the provider wins over
* our own backoff, because the provider knows when capacity returns.
* Gives up early if the remaining deadline cannot fit the next wait.
*/
export async function withRetry<T>(
fn: (attempt: number) => Promise<T>,
policy: RetryPolicy,
clock: Clock,
deadline: number,
signal: AbortSignal,
random: () => number = Math.random,
): Promise<T> {
for (let attempt = 0; ; attempt++) {
try {
return await fn(attempt);
} catch (err) {
const last = attempt + 1 >= policy.maxAttempts;
if (!(err instanceof RetryableError) || last) throw err;
const wait = err.retryAfterMs ?? backoffDelay(attempt, policy, random);
if (clock.now() + wait >= deadline) throw err;
await clock.sleep(wait, signal);
}
}
}A Retry-After from the provider wins. The provider knows when capacity returns; your backoff formula is a guess.
Only RetryableError is retried. Everything else propagates immediately.
The deadline is shared by all attempts. The caller sets timeoutMs once, for the whole operation. The gateway does not start a wait that would end after the deadline, and the same AbortSignal is passed to every attempt so an in-flight request is cancelled when time runs out. Without a shared deadline, three attempts with a 30-second timeout each become a 90-second request that the user abandoned after ten.
Putting it together
Here is the gateway itself. It is short because each stage is a separate, separately tested unit.
export class Gateway {
private readonly breakers = new Map<string, CircuitBreaker>();
private readonly limiter: TenantLimiter;
constructor(private readonly o: GatewayOptions) {
this.limiter = new TenantLimiter(o.limits, o.clock);
}
async complete(req: ModelRequest): Promise<ModelResponse> {
const started = this.o.clock.now();
const deadline = started + req.timeoutMs;
const base = {
tenantId: req.tenantId, model: req.model,
promptId: req.prompt.id, promptVersion: req.prompt.version, promptHash: req.prompt.hash,
};
let attempts = 0;
let provider: string | null = null;
const route = this.o.routes[req.model];
if (!route) throw this.reject(base, started, new RejectedError(`unknown model ${req.model}`, 'invalid_input'));
// 1. Admission: reserve the worst case for this tenant before any network call.
const estimate = this.o.estimateInputTokens(req.prompt.system + req.prompt.user) + req.maxOutputTokens;
const bucket = this.limiter.bucket(req.tenantId);
if (!bucket.tryTake(estimate)) {
const wait = bucket.waitTimeMs(estimate);
throw this.reject(base, started, new RejectedError(`tenant over budget, retry in ${wait} ms`, 'rate_limited'));
}
const signal = AbortSignal.timeout(req.timeoutMs);
try {
// 2. Try providers in order; a provider with an open circuit is skipped.
for (const p of route.providers) {
const breaker = this.breaker(p.name);
if (!breaker.tryAcquire()) continue;
provider = p.name;
try {
const res = await withRetry(
() => { attempts++; return p.complete(req, signal); },
this.o.retry, this.o.clock, deadline, signal, this.o.random,
);
breaker.onSuccess();
// 3. Reconcile: refund what we reserved but did not use.
const used = res.usage.inputTokens + res.usage.outputTokens;
if (used < estimate) bucket.refund(estimate - used);
const costUsd = (res.usage.inputTokens * route.price.input + res.usage.outputTokens * route.price.output) / 1e6;
this.o.onEvent({
...base, provider, outcome: 'ok', attempts, latencyMs: this.o.clock.now() - started,
usage: res.usage, costUsd, sample: redact(res.text).slice(0, 200),
});
return { ...res, provider: p.name };
} catch (err) {
// Permanent errors are the caller's problem, not the provider's health.
if (err instanceof PermanentError) throw err;
breaker.onFailure();
if (signal.aborted) throw err;
// Retryable and exhausted on this provider: fall through to the next one.
}
}
throw new RejectedError('no provider available (all circuits open or failing)', 'circuit_open');
} catch (err) {
bucket.refund(estimate); // nothing was served
const e = err instanceof Error ? err : new Error(String(err));
this.o.onEvent({
...base, provider, outcome: e instanceof RejectedError ? 'rejected' : 'error',
error: `${e.name}: ${redact(e.message)}`, attempts, latencyMs: this.o.clock.now() - started,
});
throw e;
}
}
private breaker(name: string): CircuitBreaker {
let b = this.breakers.get(name);
if (!b) { b = new CircuitBreaker(this.o.breaker, this.o.clock); this.breakers.set(name, b); }
return b;
}
private reject(base: Omit<GatewayEvent, 'provider' | 'outcome' | 'attempts' | 'latencyMs'>, started: number, e: RejectedError): RejectedError {
this.o.onEvent({ ...base, provider: null, outcome: 'rejected', error: `${e.name}: ${e.message}`, attempts: 0, latencyMs: this.o.clock.now() - started });
return e;
}
}Walk through a call that meets an outage:
- Admission. The tenant's bucket reserves the estimated input plus
maxOutputTokens. If that fails, the caller gets aRejectedErrorwith the wait time and nothing else happens. - Primary provider. Its breaker is closed, so the call goes out. It returns
503, which is retryable. The retry policy waits with jitter and tries again, twice, all within the deadline. - Fallback. After the retries are exhausted, the primary's breaker records a failure and the gateway moves to the next provider in the route, which succeeds.
- Reconcile. Real usage is lower than the reservation, so the difference is refunded to the tenant's bucket. Cost is computed from the real usage and the route's price table.
- Telemetry. One event records the tenant, prompt id, version and hash, the provider that served the call, the number of attempts, latency, usage, cost and a short redacted sample.
After three such requests the primary's breaker opens, and later requests go straight to the fallback without paying the timeout. After the cooldown, one probe tests the primary again.
If anything throws on the way out, the reservation is refunded in full, because nothing was served, and an error or rejected event is still emitted. A call that leaves no trace is the hardest kind to debug.
Telemetry without leaking data
The gateway emits exactly one event per call. It is the single source for dashboards, cost reports, alerts and the samples that seed your eval datasets, so it needs to be complete and safe at the same time.
Map the fields to the OpenTelemetry semantic conventions for generative AI where you can, for example gen_ai.request.model, gen_ai.usage.input_tokens and gen_ai.usage.output_tokens. You get dashboards and tooling that already understand them, and you can change observability vendor without rewriting the gateway. Add your own attributes for what the conventions do not cover: tenant, prompt id, version and hash, attempts and computed cost.
Prompts and completions are a different matter. They contain whatever users typed: names, email addresses, order numbers, sometimes national identity numbers or card numbers pasted into a support chat. Shipping them in full to a logging vendor turns that vendor into a processor of personal data, with everything that implies under GDPR.
/**
* Applied to everything that leaves the gateway as telemetry. Prompts and
* completions contain whatever users typed, so treat them as personal data
* by default. Patterns are deliberately broad: a false positive costs a
* slightly less useful log line, a false negative costs an incident.
*/
const PATTERNS: Array<[RegExp, string]> = [
[/\b[\w.+-]+@[\w-]+(?:\.[\w-]+)+\b/g, '[email]'],
[/\b(?:sk|pk|rk)[-_][A-Za-z0-9_-]{16,}\b/g, '[secret]'],
[/\bBearer\s+[A-Za-z0-9._~+/-]+=*/g, 'Bearer [secret]'],
[/\b(?:\d[ -]?){13,19}\b/g, '[card]'],
[/\b(?:19|20)?\d{6}[-+]?\d{4}\b/g, '[national-id]'],
[/\+?\d[\d\s().-]{7,}\d/g, '[phone]'],
];
export function redact(text: string): string {
return PATTERNS.reduce((s, [re, replacement]) => s.replace(re, replacement), text);
}The gateway therefore logs only a short, redacted sample. The patterns are broad on purpose: a false positive costs a slightly less useful log line, a false negative costs an incident report.
Testing it without a network
A gateway is concurrency, time and failure handling, which are the three hardest things to test against a real provider. The implementation therefore takes a Clock and talks to providers through an interface, and the tests replace both.
it('falls back when the primary keeps failing, then skips it while its circuit is open', async () => {
const primary = scripted('primary', [new RetryableError('503')]);
const fallback = scripted('fallback', ['ok']);
const gw = gateway([primary, fallback]);
for (let i = 0; i < 3; i++) expect((await gw.complete(request())).provider).toBe('fallback');
expect(primary.calls).toBe(9); // 3 requests × 3 attempts, then the circuit opens
await gw.complete(request());
expect(primary.calls).toBe(9); // open circuit: not called at all
});The fake clock makes sleep advance time instantly and records every wait, so a test can assert that the gateway honoured a Retry-After of exactly 1,000 ms without waiting a second. A scripted provider plays back a sequence of outcomes. Fifteen tests cover the policy end to end: backoff bounds, Retry-After, deadlines, breaker transitions, token refill, fallback, tenant isolation while calls are in flight, reservation refunds, prompt immutability and redaction. They run in well under a second on every pull request that touches this article.
Two properties are worth testing in any gateway, whatever its implementation:
- Tenant isolation under concurrency. Start a call for tenant A that has not finished, then show that a second call for A is rejected while a call for B succeeds. A sequential test cannot catch a reservation bug, because the refund happens before the next call starts.
- Permanent errors bypass everything. A
400from the primary must not be retried, must not touch the fallback and must not move the breaker.
What we left out on purpose
A reference architecture is useful partly for what it does not include. These are real concerns, each with its own trade-offs, and each is covered by a later article in this series.
Streaming. Users expect tokens to appear as they are generated. Streaming changes the failure model: once you have forwarded half an answer you cannot transparently retry or fall back, and your latency metric splits into time to first token and total time. The admission and telemetry stages stay the same; retries and fallbacks only apply before the first byte.
Caching. Provider-side prompt caching reduces cost and latency for long, repeated prefixes and needs no extra infrastructure. Response caching in your own system is riskier. Cache keys must include the tenant, the prompt hash and every input variable, or one tenant can be served another tenant's answer. Both are covered in cost and latency and multi-tenant AI.
Idempotency. For generations that trigger side effects, the caller should pass an idempotency key and the gateway should store the result for a period, so a retried request returns the original result instead of running twice. See structured output and tool calls.
Structured output. Validating output against a schema, and deciding whether to repair, retry or fail, belongs in a layer above the gateway, because the right answer depends on the feature.
Evaluation. The gateway makes every call traceable to a prompt version and a provider. That is what makes the next step possible: running the same cases against two versions and deciding, with evidence, whether a change can ship. The next article in the series, Evals as CI, builds that release gate.
A checklist for your own gateway
If you are building or reviewing a gateway, these are the questions we would ask first.
| Question | What good looks like |
|---|---|
| Where do provider keys live? | Only in the gateway's secret store. Callers cannot reach providers directly. |
Where does tenantId come from? |
The authenticated session or the job record. Never request input. |
| What is rate limited? | Tokens per tenant, with a global ceiling below the provider's limit. |
| What happens over budget? | An immediate rejection with a precise retry time. |
| Which errors are retried? | An explicit table per provider, with permanent errors never retried. |
| Where are retries? | In the gateway only, within one deadline, with full jitter. |
| What happens during an outage? | A breaker per provider fails fast and a tested fallback serves. |
| Can you trace an output? | Tenant, prompt id, version, hash, provider, attempts, usage and cost on every call. |
| What reaches the logs? | A redacted sample, with short retention and strict access. |
You can find the whole series on the Luniat Engineering page, and the rest of the magazine on the Luniat Magazine front page.
References
- Marc Brooker, "Exponential Backoff And Jitter", AWS Architecture Blog, 2015.
- Martin Fowler, "CircuitBreaker", 2014.
- Google, Site Reliability Engineering, chapters "Handling Overload" and "Addressing Cascading Failures", O'Reilly, 2016.
- IETF, RFC 9110 HTTP Semantics, section 10.2.3: Retry-After, 2022.
- OpenTelemetry, Semantic conventions for generative AI systems.
- OWASP, Top 10 for Large Language Model Applications.