Observability for LLM systems: traces, metrics and samples without leaking data

Spans with OpenTelemetry's generative AI attributes, tail sampling that keeps every interesting trace, latency histograms, and no prompts in the logs.

Level
Intermediate to advanced
Stack
TypeScript, OpenTelemetry, Any tracing back end

In short

  • Record metadata on every call with OpenTelemetry's generative AI attributes, your own under app.*, and no content.
  • Keep every error, slow and flagged trace with tail sampling; sample the rest deterministically by trace id.
  • Track latency as histograms with time to first token separate, and alert on cache share, cut-offs and policy denials.
  • Capture content only through a sampled, redacted, short-lived and access-controlled path that also feeds your evals.
In this article · 8 sections
  1. 01What to record, and what not to
  2. 02Spans with standard attributes
  3. 03Sampling that keeps what matters
  4. 04Latency as a distribution
  5. 05Metrics worth an alert
  6. 06Content samples for debugging and evals
  7. 07A checklist for LLM observability
  8. 08References

When an LLM feature misbehaves, the questions are always the same. Which tenant? Which prompt version? Which model actually answered? How long did it take before the first token, and how long in total? Was it retried? What did it cost? Did anything in the chain fail quietly?

If the system cannot answer those questions from its telemetry, every incident becomes an investigation that starts with "can you reproduce it?", and the answer is usually no, because the model will not produce the same output twice.

This article describes the telemetry we would want for any LLM system: spans with standard attributes, a sampling policy that keeps every interesting trace, latency histograms that show the tails, and a firm line between metadata, which is recorded always, and content, which is not. As in the rest of the series, the code is type-checked and tested in CI.

What to record, and what not to

The gateway already emits one event per call. Observability is about turning those events into something you can query, alert on and follow through a whole request.

Record on every call Record only as a sample Do not record
Operation, provider, requested and response model Short, redacted excerpts of input and output Full prompts and completions in general logs
Tokens in, out and from cache; max tokens Retrieved chunk ids for flagged answers Secrets, credentials or keys in any form
Latency: time to first token and total Tool arguments for failed calls, redacted Raw personal data from user input
Finish reason, error type, attempts
Tenant, prompt id and version, cost

The first column is cheap, contains no personal data beyond identifiers, and answers most operational questions. The second is valuable for debugging and for building eval datasets, and must be handled like the user data it is. The third column is where incidents come from.

Spans with standard attributes

OpenTelemetry has semantic conventions for generative AI: a shared vocabulary for model calls, so that tracing back ends can show tokens, models and errors without custom configuration. Using it costs nothing and keeps you portable between observability vendors.

observability/genai-span.ts
/** What the gateway knows about one model call. */
export interface CallRecord {
  operation: 'chat' | 'embeddings' | 'execute_tool';
  provider: string;
  requestModel: string;
  responseModel?: string;
  maxTokens?: number;
  usage?: { input: number; output: number; cacheRead?: number };
  finishReasons?: string[];
  error?: string;
  /** Our own context, kept in an app.* namespace. */
  tenantId: string;
  promptId: string;
  promptVersion: number;
  attempts: number;
  costUsd?: number;
  ttftMs?: number;
}

type Attr = string | number | boolean | string[];

/**
 * Maps a call to span attributes. Standard names follow the OpenTelemetry
 * semantic conventions for generative AI, so tracing back ends understand
 * them; everything specific to us lives under app.*. Prompt and completion
 * text are deliberately absent: content capture is opt-in, sampled and
 * redacted, and never part of the default span.
 */
export function genAiAttributes(c: CallRecord): Record<string, Attr> {
  const a: Record<string, Attr> = {
    'gen_ai.operation.name': c.operation,
    'gen_ai.provider.name': c.provider,
    'gen_ai.request.model': c.requestModel,
    'app.tenant_id': c.tenantId,
    'app.prompt.id': c.promptId,
    'app.prompt.version': c.promptVersion,
    'app.attempts': c.attempts,
  };
  if (c.responseModel) a['gen_ai.response.model'] = c.responseModel;
  if (c.maxTokens !== undefined) a['gen_ai.request.max_tokens'] = c.maxTokens;
  if (c.usage) {
    a['gen_ai.usage.input_tokens'] = c.usage.input;
    a['gen_ai.usage.output_tokens'] = c.usage.output;
    if (c.usage.cacheRead !== undefined) a['app.usage.cache_read_tokens'] = c.usage.cacheRead;
  }
  if (c.finishReasons?.length) a['gen_ai.response.finish_reasons'] = c.finishReasons;
  if (c.error) a['error.type'] = c.error;
  if (c.costUsd !== undefined) a['app.cost_usd'] = Number(c.costUsd.toFixed(6));
  if (c.ttftMs !== undefined) a['app.ttft_ms'] = c.ttftMs;
  return a;
}

/** Span name convention: "{operation} {model}". Low cardinality, readable in a trace view. */
export const spanName = (c: CallRecord) => `${c.operation} ${c.requestModel}`;

Standard names for standard facts. gen_ai.operation.name, gen_ai.provider.name, gen_ai.request.model, gen_ai.response.model, gen_ai.usage.input_tokens, gen_ai.usage.output_tokens, gen_ai.request.max_tokens, gen_ai.response.finish_reasons and error.type cover what every call has in common.

Request model and response model are both recorded. You ask for a logical model name or a model alias; the provider reports which concrete version answered. When behaviour changes overnight without a deploy, the response model is often where the explanation is.

Everything specific to you lives under app.*. Tenant, prompt id and version, attempts, cost, cache reads and time to first token. A separate namespace keeps your attributes from colliding with future versions of the standard.

No content in the span. The test asserts that no attribute name looks like a prompt, completion or message field. The conventions define optional ways to capture content; leave them off by default, and capture content through the sampling path described below instead.

What a trace should show

A single user request in an LLM feature is rarely a single model call. A support answer might involve retrieval, a classification call, a generation call, a citation check and perhaps a repair. Each should be a span, nested under the request:

  • retrieve with candidate counts and the fused scores of the top results,
  • chat default for the classification, with its usage and latency,
  • chat default for the answer, with time to first token,
  • execute_tool lookup_order, with the tool's outcome code,
  • the citation check, with the number of invalid citations.

With that structure, a slow answer in a trace view shows immediately whether the time went to retrieval, to a retry, to a long generation or to a tool waiting on a third-party API.

Sampling that keeps what matters

Recording every trace is expensive at volume, and most of them are identical successes. Sampling at the start of a request, with a fixed probability, throws away errors and slow requests at the same rate as everything else. That is the wrong trade for LLM systems, where the interesting traces are rare.

observability/sampling.ts
export interface TraceSummary {
  traceId: string;
  durationMs: number;
  error: boolean;
  /** Set by product code: thumbs-down, policy denial, failed citation check. */
  flagged: boolean;
  tenantId: string;
}

/**
 * Tail-based sampling: decide after the trace is complete, so every
 * interesting trace is kept and the boring majority is sampled. The random
 * part is a hash of the trace id, so every service that sees the same trace
 * makes the same decision without coordinating.
 */
export function keepTrace(t: TraceSummary, o: { slowMs: number; baseRate: number; tenantRate?: (tenant: string) => number }):
  { keep: boolean; reason: 'error' | 'flagged' | 'slow' | 'sampled' | 'dropped' } {
  if (t.error) return { keep: true, reason: 'error' };
  if (t.flagged) return { keep: true, reason: 'flagged' };
  if (t.durationMs >= o.slowMs) return { keep: true, reason: 'slow' };
  const rate = o.tenantRate?.(t.tenantId) ?? o.baseRate;
  const bucket = createHash('sha256').update(t.traceId).digest().readUInt32BE(0) / 0xffffffff;
  return bucket < rate ? { keep: true, reason: 'sampled' } : { keep: false, reason: 'dropped' };
}

Tail-based sampling decides when the trace is complete, and keeps:

  • every trace with an error,
  • every trace that product code flagged: a thumbs-down, a policy denial from the security layer, a failed citation check,
  • every slow trace, above a threshold set from your latency targets,
  • a fixed share of the rest, chosen by hashing the trace id.

The hash makes the decision deterministic: every service that sees the same trace id makes the same choice, without coordinating, so a sampled trace is complete rather than missing the spans from one service. The test checks that the sampled share stays close to the configured rate over ten thousand traces, and that the same trace always gets the same decision.

The per-tenant rate is there for a practical reason. A tenant with a hundred times more traffic than the others will dominate a uniform sample, and the small tenants' traces, which you need when they report a problem, will be rare. Raising the rate for small tenants keeps their traces visible.

Latency as a distribution

LLM latency has long tails. Most answers might take under a second; a few take thirty, because they are long, because they were retried, or because the provider was slow. An average blends these into a number that describes no actual request.

observability/histogram.ts
/**
 * Fixed-bucket histogram for latency. Buckets are chosen for LLM traffic,
 * where p99 is often ten times the median: fine-grained under a second for
 * time to first token, coarse up to a minute for full generations.
 */
export const LATENCY_BUCKETS_MS = [50, 100, 250, 500, 750, 1_000, 1_500, 2_500, 5_000, 10_000, 20_000, 40_000, 60_000];

export class Histogram {
  readonly counts: number[];
  count = 0;
  sum = 0;

  constructor(readonly bounds: number[] = LATENCY_BUCKETS_MS) {
    this.counts = new Array(bounds.length + 1).fill(0);
  }

  record(v: number): void {
    let i = this.bounds.findIndex((b) => v <= b);
    if (i === -1) i = this.bounds.length; // overflow bucket
    this.counts[i]!++;
    this.count++;
    this.sum += v;
  }

  /**
   * Quantile estimate by linear interpolation inside the bucket, the same
   * method Prometheus uses for histogram_quantile. Accurate to the bucket
   * width, which is why the bucket layout matters.
   */
  quantile(q: number): number {
    if (!this.count) return NaN;
    const rank = q * this.count;
    let seen = 0;
    for (let i = 0; i < this.counts.length; i++) {
      const c = this.counts[i]!;
      if (seen + c >= rank && c > 0) {
        const lo = i === 0 ? 0 : this.bounds[i - 1]!;
        const hi = this.bounds[i] ?? this.bounds[this.bounds.length - 1]!;
        return lo + ((hi - lo) * (rank - seen)) / c;
      }
      seen += c;
    }
    return this.bounds[this.bounds.length - 1]!;
  }
}

Use histograms, not averages. The test records ninety typical answers at 600 milliseconds, nine long ones at four seconds and one at thirty. The mean is 1.2 seconds, which no user experienced. The median is under 750 milliseconds and the 95th percentile is between 2.5 and 5 seconds, which describes what people actually saw.

Choose buckets for the traffic. The default buckets are fine-grained under a second, where time to first token lives, and coarse up to a minute, where full generations end. Quantiles from a histogram are only as precise as the bucket they fall in, so the layout matters more than it seems.

Measure time to first token and total duration separately. For streamed answers, time to first token is what the user perceives as responsiveness; total duration is what your capacity planning needs. A change that improves one often worsens the other, for example a longer cached prefix.

Metrics worth an alert

Spans are for investigating; metrics are for noticing. A small set covers most of what goes wrong:

Metric Alert when
Error rate by provider and error type Rises above its usual level for several minutes
p95 time to first token and total duration Exceeds the latency target
Retries per call Climbs: the provider is degrading before it fails
Breaker state per provider Opens
Cost per hour, per tenant and in total Exceeds the projected budget, as in the cost article
Cache read share of input tokens Drops suddenly: a prompt change broke the cached prefix
Finish reason length share Rises: outputs are being cut off
Policy denials and removed URLs Spike: possible injection or exfiltration attempt

Keep metric labels low-cardinality. Provider, model, operation, error type and prompt id are fine. Tenant id is fine for a few hundred tenants and becomes a problem for a few hundred thousand; at that scale, aggregate per tenant in a separate store rather than as a metric label. Never use free text, user ids or trace ids as labels.

Content samples for debugging and evals

Sometimes you need the text. A user reports a bad answer; an eval regression needs real examples; a new failure mode has to be understood. The answer is a separate, controlled path, not the general log stream.

  • Opt-in per tenant, where your agreements require it.
  • Sampled, from the traces the sampler already kept for being interesting.
  • Redacted with the same patterns as the gateway, before storage.
  • Short-lived: days, not months.
  • Access-controlled like the database the text came from, with access logged.
  • Linked by trace id, so an engineer can go from a trace to its sample without searching text.

This is also the path that feeds the eval suite. A flagged answer, anonymised, becomes a new test case, and the next change to the prompt has to pass it.

A checklist for LLM observability

Question What good looks like
Can you say which tenant, prompt version and model produced an output? Yes, from attributes on every span.
Do spans use the standard generative AI attributes? Yes, mapped in one function, with your own under app.*.
Are prompts and completions in the general logs? No: only redacted, sampled, short-lived samples.
Are errors, slow and flagged traces always kept? Yes, by tail sampling; the rest is sampled by trace id hash.
Is latency tracked as percentiles? Yes, from histograms, with time to first token separate.
Are cache hit share and cut-off outputs monitored? Yes, with alerts on sudden changes.
Are metric labels low-cardinality? Yes: no free text, user ids or trace ids.

The whole series is on the Luniat Engineering page.

References

  1. OpenTelemetry, Semantic conventions for generative AI systems.
  2. OpenTelemetry, Gen AI attribute registry.
  3. OpenTelemetry, Sampling, including tail sampling.
  4. Prometheus, Histograms and summaries.
  5. Google, Site Reliability Engineering, chapter "Monitoring Distributed Systems", O'Reilly, 2016.
Read next →Agents in production: state machines, budgets and humans in the loopEngineering · No. 10 · 11 min

Frequently asked questions

Should prompts and completions be logged?
Not by default. They contain whatever users typed and are personal data. Record metadata on every call, and capture content only as an opt-in, sampled, redacted and short-lived sample, stored with the same access controls as the source data.
Which OpenTelemetry attributes should an LLM span have?
At least the operation, provider, requested and response model, token usage, maximum tokens, finish reasons and error type, using the generative AI semantic conventions. Add your own attributes, such as tenant and prompt version, under an application namespace.
Why use tail-based sampling for LLM traces?
Because the traces you need are rare: errors, slow generations, answers a user flagged. Tail-based sampling decides after the trace is complete, so it can keep all of those and sample only the ordinary ones.
Why are averages misleading for LLM latency?
Output length varies a lot, so latency distributions have long tails. The mean hides both the typical case and the slow one. Track percentiles from histograms, and time to first token separately from total duration.