Multi-tenant AI without leaks: isolation from the database to the prompt
Forced row-level security tested in CI, a request-scoped tenant context, tenant-bound caches and secrets, and why the prompt is a new place to leak data.
- Level
- Advanced
- Stack
- TypeScript, PostgreSQL row-level security, Node.js AsyncLocalStorage
In short
- Derive the tenant from a verified credential once, keep it in a request context, and make code without a tenant throw.
- Use forced row-level security with policies for every command, a transaction-scoped tenant setting and an audit for open policies.
- Scope retrieval, history, tools, caches and telemetry to the tenant before anything reaches the prompt or leaves the system.
- Bind per-tenant secrets to the tenant as authenticated data, and test every layer with two tenants in a blocking CI suite.
In this article · 10 sections
- 01Why leaks happen
- 02Layer 1: the tenant comes from the credential
- 03Layer 2: row-level security that fails closed
- 04Layer 3: retrieval and the prompt
- 05Layer 4: caches and derived data
- 06Layer 5: secrets bound to the tenant
- 07Layer 6: telemetry is tenant data too
- 08The cross-tenant test suite
- 09A checklist for tenant isolation
- 10References
In a multi-tenant system, a cross-tenant leak is the incident that ends customer relationships. It is also the one that is most often found by a customer rather than by a test, because the code that causes it usually works perfectly for every tenant on its own. The bug only exists in the relationship between tenants, and most test suites never create two.
LLM features make this worse in two ways. They add new places where data is stored and reused: embeddings, retrieval indexes, response caches, conversation memory, tool results. And they add a component, the model, that will repeat anything in its context to whoever is asking, with no concept of whose data it was.
This article builds tenant isolation as a set of independent layers, from the database to the prompt, and tests each of them. As in the rest of the series, the code is type-checked and tested in CI. The row-level security policies here are tested against a real Postgres instance on every pull request.
Why leaks happen
Cross-tenant leaks are rarely exotic. They come from a short list of patterns, and it is worth naming them, because each layer in this article exists to stop one of them.
| Pattern | Typical example |
|---|---|
| Tenant taken from the request | An endpoint reads tenantId from the body or a query parameter and trusts it |
| A query without a filter | A new endpoint or a quick fix forgets where tenant_id = $1 |
| An open policy | A permissive policy with true added during debugging, which overrides every other policy on the table |
| Shared state keyed wrongly | A cache keyed on the question text, not the tenant, serves one customer's answer to another |
| Filtering after retrieval | Search returns the global top 20, then drops other tenants' results |
| Background jobs | A job that infers its tenant from the data it processes instead of carrying it explicitly |
| Telemetry | Prompts with customer data shipped to a shared log index everyone can search |
The common cause is relying on application code to get the tenant right every time, with no second line of defence and no test that would notice when it does not.
Layer 1: the tenant comes from the credential
Everything starts with one rule: the tenant is derived from a verified credential on the server, and from nowhere else. Not a request parameter, not a body field, not a header the client controls. Authorisation flaws of this kind are the first entry on the OWASP API Security Top 10, under the name broken object level authorization, and they are the easiest to introduce in a hurry.
export interface Principal {
userId: string;
tenantId: string;
}
const storage = new AsyncLocalStorage<Principal>();
/**
* The tenant is established once, at the edge, from a verified credential,
* and is then available to every function in the request without being
* passed through (or overridden by) request parameters.
*/
export async function runAuthenticated<T>(
token: string | undefined,
verify: (token: string) => Promise<Principal | null>,
fn: () => Promise<T>,
): Promise<T> {
const principal = token ? await verify(token) : null;
if (!principal) throw Object.assign(new Error('unauthenticated'), { status: 401 });
return storage.run(Object.freeze({ ...principal }), fn);
}
/** Throws instead of returning undefined: code that needs a tenant must never run without one. */
export function currentTenant(): string {
const p = storage.getStore();
if (!p) throw new Error('no tenant in context');
return p.tenantId;
}The tenant is established once, at the edge, and stored in an AsyncLocalStorage context that follows the request through every await. Code deep in the call stack asks for currentTenant() instead of receiving a tenant id as a parameter, which removes the temptation to pass one along from the request.
currentTenant() throws when there is no tenant. That is deliberate: a function that silently returns undefined invites where tenant_id = undefined, which in some query builders means "no filter". Background jobs establish the context the same way, from the tenant id stored on the job record when it was enqueued by an authenticated request.
Layer 2: row-level security that fails closed
Application code will eventually forget a filter. Row-level security in Postgres makes that mistake return nothing instead of everything, by moving the tenant check into the database itself.
-- The application connects as a role that owns nothing and bypasses nothing.
create role app_user nologin;
create table conversations (
id bigint generated always as identity primary key,
tenant_id uuid not null default nullif(current_setting('app.tenant_id', true), '')::uuid,
title text not null,
created_at timestamptz not null default now()
);
create index conversations_tenant_idx on conversations (tenant_id);
alter table conversations enable row level security;
-- Also applies RLS to the table owner, so a migration or admin script that
-- forgets to switch role does not silently see every tenant.
alter table conversations force row level security;
-- One policy per command, all keyed on the same setting. If the setting is
-- missing the comparison is NULL and every row is filtered out: the failure
-- mode is "nothing", never "everything". nullif is needed because once a
-- connection has set the variable, it reads as '' (not NULL) after the
-- transaction ends, and ''::uuid is an error rather than an empty result.
create policy tenant_select on conversations for select to app_user
using (tenant_id = nullif(current_setting('app.tenant_id', true), '')::uuid);
create policy tenant_insert on conversations for insert to app_user
with check (tenant_id = nullif(current_setting('app.tenant_id', true), '')::uuid);
create policy tenant_update on conversations for update to app_user
using (tenant_id = nullif(current_setting('app.tenant_id', true), '')::uuid)
with check (tenant_id = nullif(current_setting('app.tenant_id', true), '')::uuid);
create policy tenant_delete on conversations for delete to app_user
using (tenant_id = nullif(current_setting('app.tenant_id', true), '')::uuid);
grant select, insert, update, delete on conversations to app_user;Every line of this has a reason.
The application connects as a dedicated role that owns nothing and does not bypass row-level security. Policies apply to it on every statement.
force row level security applies the policies to the table owner too. Without it, a migration script or an admin tool running as the owner sees every tenant, and so does anyone who copies one of its queries into application code.
One policy per command, with both using for rows that can be seen or changed and with check for rows that can be written. A table with only a select policy lets a tenant insert rows into another tenant's space, as long as it does not try to read them back. The tests try exactly that.
The policies read a setting, not a function argument. app.tenant_id is set at the start of each transaction. If it is missing, the comparison is null and no rows match.
Setting the tenant per transaction
Connections are pooled. A setting made with plain SET on a connection stays on that connection when it goes back to the pool, and the next request that borrows it inherits the previous request's tenant. That is a cross-tenant leak caused by an optimisation.
/**
* Every tenant-scoped query runs inside this function. The tenant id and the
* role are set with transaction scope (set_config(..., true) and SET LOCAL),
* so they vanish at commit or rollback and can never leak to the next request
* that borrows the same pooled connection.
*/
export async function withTenant<T>(db: TxDb, tenantId: string, fn: (tx: Tx) => Promise<T>): Promise<T> {
if (!UUID.test(tenantId)) throw new Error('invalid tenant id');
return db.transaction(async (tx) => {
await tx.query(`select set_config('app.tenant_id', $1, true)`, [tenantId]);
await tx.query('set local role app_user');
return fn(tx);
});
}withTenant opens a transaction, sets the tenant with set_config(..., true) and switches role with SET LOCAL. Both are transaction-scoped and vanish at commit or rollback. The tenant id is validated as a UUID before it goes anywhere near the database, and it is passed as a bound parameter, never interpolated.
Auditing the policies themselves
Permissive policies in Postgres are combined with OR. One policy with using (true) for a client role turns every carefully written policy on the same table into decoration. Such policies get created: during debugging, by a template, by an administration interface that offers to "enable access".
-- Run in CI: no policy on a tenant table may be unconditionally true for
-- client roles. A policy created with "true" in a hurry turns every other
-- policy on the table into decoration, because permissive policies are ORed.
create view open_policies as
select schemaname, tablename, policyname, roles, cmd
from pg_policies
where (qual = 'true' or with_check = 'true')
and roles && array['public', 'app_user']::name[];The audit is a view over pg_policies that lists any policy which is unconditionally true for a client role. The test suite asserts that it is empty, and also creates an open policy on purpose to prove that the audit catches it. A check that has never been seen to fail is not yet a check.
Layer 3: retrieval and the prompt
The model is a new kind of component in a multi-tenant system: it has no tenant, and it will repeat anything in its context. Isolation therefore has to be complete before anything reaches the prompt.
Retrieval filters on the tenant in every candidate query, as covered in the RAG article, and row-level security applies to the chunk table as well. Filtering results after a global search is not equivalent: it leaks through ranking and timing, and it returns fewer results for small tenants.
Conversation history and memory are tenant data, stored under the same policies, never in a shared in-memory structure keyed by conversation id alone.
Tool calls act for the session's tenant, as in the tools article. A tool that takes a tenant id as an argument lets any prompt that mentions another tenant's id reach their data.
Shared content is a shared injection surface. If tenants can contribute content that other tenants' assistants read, such as a public knowledge base, a marketplace listing or a shared document, one tenant can plant instructions that run in another tenant's context. Treat any content that crosses a tenant boundary as untrusted, and keep it out of contexts where the model can call write tools.
Layer 4: caches and derived data
Every cache is a copy of tenant data under a different key. A response cache keyed on the prompt text alone is the textbook cross-tenant leak in LLM systems: two tenants ask the same question, and the second gets the first tenant's answer, complete with their order numbers.
/** JSON with sorted keys, so {a, b} and {b, a} produce the same key. */
export function canonicalJson(v: unknown): string {
if (Array.isArray(v)) return `[${v.map(canonicalJson).join(',')}]`;
if (v && typeof v === 'object') {
return `{${Object.keys(v).sort().map((k) => `${JSON.stringify(k)}:${canonicalJson((v as Record<string, unknown>)[k])}`).join(',')}}`;
}
return JSON.stringify(v);
}
/**
* A cached model response may only be served to the tenant that produced it.
* The tenant is a prefix (so a tenant's entries can be listed and purged),
* and everything that can change the output is inside the hash.
*/
export function responseCacheKey(k: { tenantId: string; model: string; promptHash: string; inputs: Record<string, unknown> }): string {
const digest = createHash('sha256').update(canonicalJson({ model: k.model, promptHash: k.promptHash, inputs: k.inputs })).digest('hex');
return `llm:v1:${k.tenantId}:${digest}`;
}The key starts with the tenant, so a tenant's entries can be listed and purged when they leave. Everything else that can change the output, the model, the prompt hash and every input variable, is inside a SHA-256 digest of canonical JSON, so key order in the inputs does not create accidental misses.
The same rule applies to every derived store: embeddings carry the tenant id, evaluation samples are tagged by tenant, and anything exported for analysis is either aggregated across tenants or stays inside one.
Layer 5: secrets bound to the tenant
Many AI products store secrets on behalf of tenants: their own provider keys, OAuth tokens for the systems the assistant connects to, webhook signing secrets. These deserve encryption in the application, with the key held outside the database, and one more property that is easy to add and often missed.
/**
* Per-tenant secrets (a tenant's own provider key, OAuth tokens) encrypted in
* the application with AES-256-GCM. The tenant id is bound as additional
* authenticated data: a ciphertext copied into another tenant's row fails to
* decrypt instead of quietly working for the wrong customer.
*
* The key comes from a secret manager or KMS, never from the same database
* as the ciphertext. Format: v1.<iv>.<tag>.<ciphertext>, base64url.
*/
export function encryptForTenant(plaintext: string, tenantId: string, key: Buffer): string {
if (key.length !== 32) throw new Error('key must be 32 bytes');
const iv = randomBytes(12);
const cipher = createCipheriv('aes-256-gcm', key, iv);
cipher.setAAD(Buffer.from(`tenant:${tenantId}`));
const ct = Buffer.concat([cipher.update(plaintext, 'utf8'), cipher.final()]);
return ['v1', iv, cipher.getAuthTag(), ct].map((p) => (typeof p === 'string' ? p : p.toString('base64url'))).join('.');
}
export function decryptForTenant(token: string, tenantId: string, key: Buffer): string {
const [v, iv, tag, ct] = token.split('.');
if (v !== 'v1' || !iv || !tag || !ct) throw new Error('unsupported format');
const decipher = createDecipheriv('aes-256-gcm', key, Buffer.from(iv, 'base64url'));
decipher.setAAD(Buffer.from(`tenant:${tenantId}`));
decipher.setAuthTag(Buffer.from(tag, 'base64url'));
return Buffer.concat([decipher.update(Buffer.from(ct, 'base64url')), decipher.final()]).toString('utf8');
}The tenant id is bound to the ciphertext as additional authenticated data in AES-256-GCM. The ciphertext is useless with any other tenant id: if a bug or an attacker copies an encrypted token from one tenant's row to another's, decryption fails instead of quietly working for the wrong customer. The test copies a token between tenants and expects the authentication tag to reject it.
The encryption key does not live in the same database. Encryption with a key stored next to the ciphertext protects against very little. Keep the key in a secret manager or a KMS, and use envelope encryption with per-tenant data keys if you need to rotate or revoke per tenant.
Layer 6: telemetry is tenant data too
Logs and traces contain prompts, outputs and identifiers. They are as sensitive as the database they came from, and usually far less protected: a shared log index that every engineer can search, with longer retention than the data itself.
The rules from the gateway article apply: one event per call, tagged with the tenant, with only a short redacted sample of the text. Add access control that matches the data, and a retention period no longer than the shortest one you promise customers.
The cross-tenant test suite
None of these layers is trustworthy until it has been tested against two tenants. The suite for this article does what an attacker would try, as tenant A against tenant B:
- Read: select from a table with rows for both tenants and get only A's.
- Write: insert a row with B's tenant id and get a row-level security violation.
- Modify: update and delete with a filter that matches B's rows and affect zero rows.
- Forget: query with no tenant set and get zero rows, not an error and not everything.
- Reuse: run a tenant transaction, then check that the setting is gone on the same connection.
- Open policy: add a permissive policy and see the audit report it.
- Inject: pass a tenant id with SQL in it and have it rejected before the database.
- Cache: build cache keys for the same input as two tenants and get two different keys.
- Decrypt: move an encrypted secret from A to B and fail to decrypt it.
Make the suite blocking in CI, and add a case for every new table, endpoint, cache or store that holds tenant data. When a leak is found anyway, write the test that would have caught it before writing the fix, and watch it fail.
A checklist for tenant isolation
| Question | What good looks like |
|---|---|
| Where does the tenant come from? | A verified credential, once, at the edge. Never request input. |
| What happens if code forgets the tenant? | It throws, and the database returns zero rows. |
| Is row-level security forced and complete? | Forced, with policies for select, insert, update and delete. |
| Can a pooled connection leak a tenant? | No: the setting is transaction-scoped. |
| Can a permissive policy slip in? | An audit in CI fails if one exists. |
| Is retrieval filtered before ranking? | Yes, in every candidate query. |
| Do cache keys include the tenant? | As a prefix, with all output-affecting inputs hashed. |
| Can a secret be used by another tenant? | No: the tenant is bound as authenticated data. |
| Is all of this tested with two tenants? | A blocking cross-tenant suite in CI. |
The whole series is on the Luniat Engineering page.
References
- PostgreSQL documentation, Row Security Policies.
- OWASP, API Security Top 10 2023, API1: Broken Object Level Authorization.
- OWASP, Top 10 for Large Language Model Applications: Sensitive Information Disclosure.
- AWS, SaaS Tenant Isolation Strategies, whitepaper.
- NIST, SP 800-38D: Recommendation for Block Cipher Modes of Operation: Galois/Counter Mode (GCM) and GMAC, 2007.
- Node.js documentation, Asynchronous context tracking (AsyncLocalStorage).