Luniat Engineering
Technical deep dives for engineers and technical leaders. Every code sample is type-checked and tested in CI before it is published.
10 articles · 118 min read · All the code is on GitHub →
Start readingAI in production
How to build, evaluate and operate LLM systems that have to work every day, not just in the demo.
- 01 A reference architecture for LLM systems in productionOne gateway between your services and every model provider, with per-tenant limits, versioned prompts, fallbacks, retries and telemetry. Tested code. 17 min
- 02 Evals as CI: release gates for LLM systemsVersioned datasets, deterministic scorers, a calibrated LLM judge and a paired statistical gate that blocks a merge when quality drops. With tested code. 14 min
- 03 Retrieval that holds up: hybrid search, chunking and citations in PostgresStructure-aware chunking, vector and full-text search fused with RRF in one SQL query, tenant-scoped by construction, with citations you can verify. 14 min
- 04 Structured output and tool calls that do not break productionValidate model output against a schema, repair it with exact errors, and run tool calls with permissions, validation and idempotency keys. Tested code. 11 min
- 05 Multi-tenant AI without leaks: isolation from the database to the promptForced row-level security tested in CI, a request-scoped tenant context, tenant-bound caches and secrets, and why the prompt is a new place to leak data. 11 min
- 06 One rate limit, many tenants: fair queues and adaptive throttlingShare one provider rate limit between tenants with deficit round robin, pace calls from quota headers, and keep a reserve for interactive traffic. 10 min
- 07 Cost and latency: caching, batching, routing and budgetsWhere the money and the milliseconds go in an LLM system, and four levers that move both, with the arithmetic and tested code for each. 10 min
- 08 Security for LLM applications: injection, exfiltration and excessive agencyWhy prompt injection cannot be filtered away, and the controls that hold anyway: spotlighting, taint-aware tool policies, output sanitising. 11 min
- 09 Observability for LLM systems: traces, metrics and samples without leaking dataSpans with OpenTelemetry's generative AI attributes, tail sampling that keeps every interesting trace, latency histograms, and no prompts in the logs. 9 min
- 10 Agents in production: state machines, budgets and humans in the loopRun an agent as an explicit state machine with a step budget, a deadline and an error cap, saved after every step and paused for human approval. 11 min