Cost and latency: caching, batching, routing and budgets
Where the money and the milliseconds go in an LLM system, and four levers that move both, with the arithmetic and tested code for each.
- Level
- Intermediate to advanced
- Stack
- TypeScript, Any LLM provider
In short
- Output tokens drive both cost and latency; ask for shorter outputs and enforce them.
- Order prompts from stable to volatile and keep timestamps out of the prefix, so provider caching can cut repeated input cost by most of its price.
- Send work that can wait to batch interfaces, and route each task to the cheapest model that passes its evals on the lower bound.
- Track spend per tenant, project it from the burn rate, and decide in advance what degrade and stop mean.
In this article · 8 sections
LLM systems have an unusual cost profile. The marginal cost of a request is not close to zero, as it is for most web traffic, and it varies by two orders of magnitude depending on the model, the length of the input and the length of the output. Latency behaves the same way: the same endpoint can answer in 400 milliseconds or in 40 seconds.
Both are controllable, but only if you know where they come from. This article starts with the arithmetic, then covers four levers that move cost and latency together: prompt caching, batching, routing per task and budgets per tenant. As in the rest of the series, the code is type-checked and tested in CI. The prices in the code and tests are illustrative round numbers, not any provider's current price list, which changes too often to print.
Where the money goes
A model call has two meters: input tokens and output tokens. Output tokens are typically priced several times higher than input tokens, and they are also what makes a call slow, because the model produces them one at a time.
/**
* Prices per million tokens, per route. The numbers in tests are illustrative;
* real prices change and belong in configuration, not in code. Cached input is
* split into writes (first time a prefix is cached) and reads (later hits),
* because providers price them differently.
*/
export interface Price {
input: number;
output: number;
cacheWrite?: number;
cacheRead?: number;
/** Multiplier for asynchronous batch processing, e.g. 0.5. */
batchFactor?: number;
}
export interface Usage {
/** Uncached input tokens. */
input: number;
output: number;
cacheWrite?: number;
cacheRead?: number;
}
export function costUsd(u: Usage, p: Price, opts: { batch?: boolean } = {}): number {
const raw =
u.input * p.input +
u.output * p.output +
(u.cacheWrite ?? 0) * (p.cacheWrite ?? p.input) +
(u.cacheRead ?? 0) * (p.cacheRead ?? p.input);
return (raw / 1e6) * (opts.batch ? p.batchFactor ?? 1 : 1);
}The cost model has four kinds of token, because providers with prompt caching price them differently:
| Token type | What it is | Relative price (typical) |
|---|---|---|
| Input | Prompt tokens processed normally | 1× |
| Cache write | Prompt prefix processed and stored for reuse | Somewhat more than input |
| Cache read | Prompt prefix served from the cache | A small fraction of input |
| Output | Generated tokens | Several times input |
Two consequences shape everything else in this article.
Long prompts are cheap to repeat if they are cached, and expensive if they are not. A support assistant with a 20,000-token system prompt of policies and tool definitions pays for those 20,000 tokens on every call unless the prefix is cached.
Long outputs are expensive in both money and time. An instruction like "answer in at most three sentences" is one of the most effective cost and latency optimisations available, and it costs nothing to try.
Before optimising anything, measure. The gateway from the first article records usage and cost per call, per tenant and per prompt version. Sort by total cost per prompt id, and the first optimisation target is usually obvious.
Lever 1: prompt caching
Provider-side prompt caching stores the processed state of a prompt prefix, so that a later request starting with the same prefix skips that work. It reduces both cost and time to first token. It is also easy to defeat by accident, because the match is on an exact prefix.
/**
* Provider prompt caches match on an exact prefix. Order the prompt from the
* most stable part to the most volatile, so the long stable part is shared by
* every request: instructions, then tool definitions, then reference material
* that changes daily, then the conversation, then the new question.
*/
export interface PromptParts {
instructions: string;
tools: string;
reference: string;
history: string;
question: string;
}
export function layout(p: PromptParts): { prefix: string; suffix: string; prefixHash: string } {
const prefix = [p.instructions, p.tools, p.reference].join('\n\n');
const suffix = [p.history, p.question].filter(Boolean).join('\n\n');
return { prefix, suffix, prefixHash: createHash('sha256').update(prefix).digest('hex').slice(0, 16) };
}
/**
* Finds values that change on every request and silently defeat the cache
* when they end up in the stable prefix: timestamps, dates, UUIDs, request ids.
*/
export function volatileTokens(prefix: string): string[] {
const patterns = [
/\b\d{4}-\d{2}-\d{2}[T ]\d{2}:\d{2}(?::\d{2})?/g,
/\b[0-9a-f]{8}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{12}\b/gi,
/\b(?:req|trace|request)[_-]?id[:=]\s*\S+/gi,
];
return patterns.flatMap((re) => prefix.match(re) ?? []);
}Order the prompt from stable to volatile. Instructions first, then tool definitions, then reference material that changes daily, then the conversation, then the new question. Everything up to the first part that differs between requests can be cached; everything after it cannot.
Keep volatile values out of the prefix. The most common way to break caching is a timestamp in the system prompt: "Today is 2 October 2026, 09:14". It changes every request, and every token after it misses the cache. The volatileTokens check catches timestamps, UUIDs and request ids in the prefix; run it in a test against your real prompt templates. If the model needs the date, put the date, not the time, at the end of the prefix, or in the user message.
The arithmetic
The test for this section uses a 20,000-token stable prefix, 500 new input tokens and 300 output tokens per request, with illustrative prices of 3 per million input tokens, 15 per million output tokens, 3.75 per million for cache writes and 0.30 for cache reads.
| Request | Cost |
|---|---|
| Uncached | 0.066 |
| First request, writing the cache | 0.081 |
| Later requests, reading the cache | 0.012 |
The first request is more expensive than not caching at all. Every later request within the cache lifetime is about 80 % cheaper than uncached. With these numbers the write premium is recovered on the first cache hit. The general rule holds across realistic price ratios: caching pays off for any prefix that is reused at least once or twice within the cache's lifetime, and it costs a little extra for prefixes that are never reused.
Lever 2: batching
Many providers offer an asynchronous batch interface: you submit many requests, and results arrive within hours instead of seconds, at a substantial discount, often around half price. The trade is latency for cost, and for a large share of LLM workloads latency does not matter at all.
Good candidates for batch processing:
- Nightly summaries, reports and digests.
- Re-classifying or re-extracting a corpus after a prompt change.
- Embedding a large document set at ingestion.
- Running the full eval suite on a schedule, as opposed to the gating subset on every pull request.
Bad candidates are anything a user is waiting for, and anything whose result feeds the next step of an interactive flow.
The fair scheduling article describes the queue for batch work that runs through the normal API. A provider batch interface is a third path next to it: work that can wait a day goes there, work that must finish within minutes goes through the fair queue, and interactive work goes straight through. The cost model takes a batch flag so that dashboards and budgets show the real price of each path.
Lever 3: routing per task
Not every task needs the largest model. Classifying a support ticket into one of eight categories is a different problem from drafting a careful reply to an angry customer, and a smaller, faster, cheaper model may do the first as well as a large one.
The question is how to know. The answer is the eval suite from the evals article: run each task's suite against each candidate route, record the pass rate with its confidence interval, the latency and the cost, and choose from measurements instead of intuition.
/** Measured per task and route by the eval suite, not estimated. */
export interface RouteStats {
route: string;
task: string;
passRate: number;
/** Lower bound of the pass rate's confidence interval. */
passRateLow: number;
p95LatencyMs: number;
costPerCallUsd: number;
}
export interface TaskPolicy {
minPassRate: number;
maxP95LatencyMs: number;
}
/**
* Picks the cheapest route that meets the task's quality and latency bars.
* Uses the lower bound of the pass rate, so a route that scored well on a
* handful of cases does not win on luck. Returns the reason when nothing
* qualifies, instead of silently picking the "best available".
*/
export function chooseRoute(stats: RouteStats[], task: string, policy: TaskPolicy):
{ route: string; costPerCallUsd: number } | { route: null; reason: string } {
const candidates = stats.filter((s) => s.task === task);
if (!candidates.length) return { route: null, reason: `no measurements for task ${task}` };
const ok = candidates
.filter((s) => s.passRateLow >= policy.minPassRate && s.p95LatencyMs <= policy.maxP95LatencyMs)
.sort((a, b) => a.costPerCallUsd - b.costPerCallUsd);
if (!ok.length) return { route: null, reason: `no route meets pass rate ≥ ${policy.minPassRate} and p95 ≤ ${policy.maxP95LatencyMs} ms` };
return { route: ok[0]!.route, costPerCallUsd: ok[0]!.costPerCallUsd };
}Use the lower bound of the pass rate. A small model that passed 19 out of 20 cases has a pass rate of 95 % and a Wilson lower bound below 80 %. It has not shown that it is as good as a large model with 97 % on 400 cases; it has shown that you need more cases. Choosing on the lower bound makes the router conservative where the evidence is thin.
Respect a latency bar per task. An interactive classification step might need a p95 under one second; a background draft can take ten. A route that is cheap but too slow does not qualify.
Refuse to guess. When no route meets the bar, the router says so with a reason. Silently picking the "best available" route hides a quality or latency problem until a customer finds it.
In the tests, ticket classification goes to the small route, which is a tenth of the cost of the large one and well within the bars, while reply drafting stays on the large route because the small one does not reach the required pass rate. That pattern is common: a system that uses one model for everything is usually paying large-model prices for its simplest steps.
Lever 4: budgets per tenant
Cost per call is an engineering metric. Cost per tenant per month is a business one, and in a multi-tenant product it needs a control loop, not just a dashboard.
export type BudgetState = 'ok' | 'warn' | 'degrade' | 'stop';
export interface BudgetPolicy {
monthlyUsd: number;
/** Fractions of the monthly budget. */
warnAt: number;
degradeAt: number;
}
/**
* Monthly spend per tenant with three thresholds. Projecting from the burn
* rate warns early in the month, when there is still time to act, instead of
* on the day the budget runs out.
*/
export function budgetState(spentUsd: number, policy: BudgetPolicy, dayOfMonth: number, daysInMonth: number): {
state: BudgetState;
projectedUsd: number;
} {
const projectedUsd = dayOfMonth > 0 ? (spentUsd / dayOfMonth) * daysInMonth : spentUsd;
const share = spentUsd / policy.monthlyUsd;
const state: BudgetState =
share >= 1 ? 'stop'
: share >= policy.degradeAt ? 'degrade'
: share >= policy.warnAt || projectedUsd > policy.monthlyUsd ? 'warn'
: 'ok';
return { state, projectedUsd };
}The budget has three thresholds and one projection.
- Warn when spend reaches a share of the budget, or when the burn rate projects past it. Spending 120 out of 300 by day 10 projects to 360 by month end; that is worth knowing on day 10, not day 25.
- Degrade near the limit: route the tenant's non-critical tasks to cheaper models, shorten outputs, move background work to batch. The product keeps working, a little less well.
- Stop at the limit: reject new non-essential work with a clear message, or switch the tenant to a pay-as-you-go tier if that is the commercial model.
What "degrade" and "stop" mean is a product decision, and it should be made before the first tenant reaches the limit, not during the incident.
Latency: what users actually feel
Cost and latency move together for most of these levers, but latency has a few specific techniques of its own.
Stream the output. Time to first token is what users perceive as responsiveness. Streaming shows text as it is generated and makes a ten-second answer feel interactive. It complicates retries and validation, as noted in the gateway article, so stream where a human reads the output and buffer where code parses it.
Shorten the output. Generation time is roughly proportional to output length. Ask for the length you need, and set maxOutputTokens to enforce it.
Cache the prefix. A cached prefix reduces time to first token as well as cost, because the provider skips processing it.
Run independent calls in parallel. If a flow classifies a ticket, extracts the order number and checks the sentiment, those are three independent calls. Run them concurrently and the latency is the slowest one, not the sum.
Use a smaller model where it qualifies. Smaller models are usually faster per token as well as cheaper.
Measure latency as a distribution, not an average. The p95 and p99 of an LLM endpoint are often several times the median, because output length varies, and those tails are what users remember. The observability article covers how to record them.
A cost review checklist
| Question | What good looks like |
|---|---|
| Do you know cost per prompt version and per tenant? | Yes, from the gateway's per-call events. |
| Is the stable prefix first, and free of timestamps? | Yes, checked by a test against the real templates. |
| What is the cache hit rate? | Measured from provider responses, per prompt. |
| Is work that can wait a day sent to a batch interface? | Yes, with the batch price in the cost model. |
| Is each task routed to the cheapest route that passes its evals? | Yes, on the lower bound of the pass rate. |
| Are outputs as short as the task allows? | Instructed and enforced with maxOutputTokens. |
| Can one tenant run up the bill unnoticed? | No: projected spend warns early, then degrades, then stops. |
The whole series is on the Luniat Engineering page.
References
- Anthropic, Prompt caching, documentation.
- OpenAI, Prompt caching, documentation.
- Lingjiao Chen, Matei Zaharia and James Zou, "FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance", 2023.
- Isaac Ong et al., "RouteLLM: Learning to Route LLMs with Preference Data", 2024.
- Lawrence D. Brown, T. Tony Cai and Anirban DasGupta, "Interval Estimation for a Binomial Proportion", Statistical Science 16(2), 2001.