Multi-Tenant SaaS Architecture: The Complete Guide for 2026

How to design multi-tenant SaaS architecture — tenant isolation models, identification, row-level scoping, cross-tenant analytics, migrations — anchored in 9 production apps.

Huifer
Huifer
August 14, 20269 min read

Written by Huifer, solo developer and maintainer of TanStack Ship. Across nine production multi-tenant SaaS apps — including a multi-tenant analytics tool at ~140k requests/day, two B2B billing platforms, a workspace collab tool, and five smaller apps — I have shipped, broken, and re-shipped the patterns that decide whether a tenant boundary holds up under real traffic. This guide consolidates those patterns: isolation models, tenant identification, row-level scoping, cross-tenant concerns, and migration paths. Every pattern below is in shipped code. The shorter version is in the multi-tenant architecture guide.

Verified sources: Cloudflare D1 documentation · Cloudflare Workers documentation · SQLite query planner · TanStack Router documentation · TanStack Ship GitHub Organization · TanStack Ship multi-tenant reference repo

Last updated: 2026-07-14 · Changelog


TL;DR: Multi-tenant SaaS architecture is the discipline of isolating customer data while sharing infrastructure. For ~95% of SaaS products, the right model is row-level isolation: one database, one schema, a tenant_id column on every table, and a tenant context injected at the edge. This guide walks through the three isolation models, tenant identification, the row-level discipline that prevents data leaks, cross-tenant analytics, and migration paths. If you are choosing a stack, see the TanStack Ship features page; for broader context, see the SaaS architecture 2026 guide.


The three isolation models and when each one earns its complexity

The honest recommendation: row-level for most apps

Every multi-tenant decision begins with the same three options: separate databases per tenant, separate schemas per tenant, or shared tables with a tenant_id column. The temptation is to start with the most isolated option because it feels safer. Resist it. The most isolated option is also the most expensive to operate, the slowest to migrate, the hardest to reason about at scale. According to the Cloudflare D1 documentation, D1 supports multiple databases per account; the question is whether you want that complexity on day one.

The model I ship for ~95% of SaaS products is row-level isolation: one database, one schema, every domain table starts with tenant_id, every query is scoped by tenant, and the tenant context is injected at the edge before any server function runs. One database means one connection, one migration, one backup. One schema means EXPLAIN works the same way for every tenant. One set of tables means cross-tenant analytics is a single SQL query. The other two models have legitimate use cases — regulated industries, enterprise customers who demand physical isolation, very large tenants whose data volume dwarfs everyone else. The path to those models is well-trodden; the path back is painful. Start simple.

When database-per-tenant earns its complexity

Database-per-tenant gives each tenant its own D1 database; the application routes requests based on tenant context. Advantages: physical isolation, per-tenant backup and restore, the ability to give large tenants their own read replica topology. Disadvantages: every additional database is another binding, another migration, another billing line. Honest use cases I have shipped for: a healthcare platform where one customer required HIPAA-aligned per-tenant data residency; a B2B platform with one customer generating ~40% of write volume (splitting that customer dropped write-lock contention for everyone else); a white-label platform where each instance was sold as a separate product. If none apply, do not reach for it — overhead grows linearly with tenant count.

Schema-per-tenant as the middle ground that rarely earns its keep

Separate schemas per tenant in a single database sounds appealing: physical isolation, one connection, one migration. In practice, schema-per-tenant in SQLite (or D1) does not give you what Postgres users expect. The query planner does not optimize across schemas the way it does across tables; cross-schema joins are slow; partial indexes do not transfer. Schema-per-tenant earns its keep in Postgres-backed apps where search_path switching is well-understood — a different stack. The rest of this guide assumes row-level isolation; the patterns transfer directly to schema-per-tenant (replace tenant_id with schema_name) and database-per-tenant (replace tenant_id with database_name).

Tenant identification: three patterns, one rule

The single rule: resolve the tenant before any business logic runs

The first architectural decision is how the request identifies its tenant. The second — and more important — is when that identification happens. The single rule I ship with: resolve the tenant before any business logic runs. The tenant context flows through the request as ambient state; every server function, every loader, every webhook handler reads it without re-deriving. Three identification patterns cover ~99% of SaaS apps: subdomain (acme.example.com — host header carries the slug), path prefix (example.com/t/acme/dashboard — first path segment carries the slug, TanStack Router's nested layout makes the context available to every nested route), and auth token (the tenant lives in the session, useful for B2B apps where the user belongs to one workspace). Subdomain is the cleanest because every request is a tenant request. Path prefix is the most flexible because it works on a single domain. Auth token is the fallback for white-label apps.

Subdomain resolution at the edge

The pattern on Cloudflare Workers resolves the host header against a tenants lookup and returns the tenant as ambient context:

typescript
// src/middleware/tenant.ts
export const tenantMiddleware = defineMiddleware({
  before: async ({ request }) => {
    const host = request.headers.get('host') ?? ''
    const slug = host.split('.')[0]
    if (['www', 'api', 'app'].includes(slug)) return
    const tenant = await db
      .prepare('SELECT id, plan, status FROM tenants WHERE slug = ?')
      .bind(slug).first<Tenant>()
    if (!tenant || tenant.status !== 'active') {
      return new Response('Tenant not found', { status: 404 })
    }
    return { tenantId: tenant.id, tenantPlan: tenant.plan }
  },
})

The middleware runs before every server function; the returned object becomes ambient context — every loader, mutation, and webhook handler reads context.tenantId without a second lookup. TanStack Router picks it up automatically.

Path-prefix resolution with route guards

When subdomains are not viable — local dev, a single staging environment, a path-first product — the path-prefix pattern uses TanStack Router's nested layouts. A _layout.tsx under routes/t/$tenantSlug/ resolves the slug, redirects on miss, and returns the tenant context. The $tenantSlug segment is a route parameter, not a query string — it survives page reloads and is shareable. The TanStack Router advanced patterns guide covers the conventions.

Row-level isolation: the four disciplines

Every domain table starts with tenant_id

The first discipline is mechanical: every domain table has a tenant_id column as its first column, every composite index has tenant_id as its first column, every foreign key references back to the tenant. The reason is the SQLite query planner's leftmost-prefix rule: a composite index (tenant_id, created_at) serves WHERE tenant_id = ? ORDER BY created_at DESC without a sort. According to the SQLite query planner documentation, the planner uses this index efficiently and ignores it for queries that omit tenant_id — exactly what you want for a database where every query must filter by tenant. The partial index on archived_at IS NULL is the SQLite-specific optimization that keeps the index small as the table grows. The full D1 schema-design playbook is in the D1 production guide.

Every query is scoped, and the type system enforces it

The second discipline actually prevents data leaks: every query is scoped by tenant, and the type system makes an unscoped query a compile error. The pattern is a thin wrapper that takes tenantId as a required argument and returns a query builder with tenant_id baked into every clause:

typescript
// src/lib/tenant-db.ts
export function tenantDb(tenantId: string) {
  return {
    projects: {
      list: () => db.prepare(
        'SELECT * FROM projects WHERE tenant_id = ? AND archived_at IS NULL ORDER BY created_at DESC'
      ).bind(tenantId),
      insert: (name: string) => db.prepare(
        'INSERT INTO projects (id, tenant_id, name) VALUES (?, ?, ?)'
      ).bind(crypto.randomUUID(), tenantId, name),
    },
  }
}
// Every server function calls tenantDb(context.tenantId).
// There is no `db.projects` global — only the scoped wrapper.

The compile-time benefit: no global db.projects to forget. The runtime benefit: every query that reaches the database already has the tenant filter. Type system + ambient context + no unscoped global — that combination is what makes row-level isolation safe at scale.

Soft delete with an audit trail

Hard deletes are a footgun: a user clicks "delete project," the row vanishes, the next invoice still references the project ID, and the foreign-key violation fires in production. The fix is soft delete with an audit table — every delete inserts into project_deletions and updates projects.deleted_at inside a single db.batch(); reads filter on deleted_at IS NULL; a nightly cron hard-deletes rows older than 90 days after exporting to R2. The full D1 migration playbook is in the D1 deep-dive guide.

Rate limits, quotas, and plan boundaries are tenant-scoped

The fourth discipline prevents the most embarrassing production incident: rate limits are scoped per tenant, not per IP. A single tenant can run a thousand concurrent syncs without tripping a global limit; a noisy neighbor cannot exhaust a shared bucket and starve everyone else. The rate-limit key is rl:${tenantId}:${route}, so every tenant gets its own bucket. The plan boundary — free 60 req/min, pro 600 req/min, enterprise unlimited — is read from tenantPlan in ambient context. Plan changes do not require code changes; the plan lives in the database. The full D1 operational playbook is in the D1 production guide.

Cross-tenant concerns: the work that is easier in one database

Analytics, billing, and the queries that need every tenant

The work that is genuinely easier in row-level isolation is cross-tenant analytics. A single SQL query joins across every tenant; no fanout, no per-tenant job, no aggregation pipeline. A monthly recurring revenue query is a GROUP BY plan against the same indexes the tenant-facing queries already use; no special analytics pipeline, no separate warehouse, no ETL job. The full observability shape is in the SaaS monitoring and observability guide.

The admin superuser pattern: a separate auth boundary

The work that requires care is the admin superuser pattern: the operator who logs in once and sees every tenant. Treat admin as a separate auth boundary, not a special tenant. Admin server functions bypass the tenantDb wrapper and call the raw db directly; every admin action is logged with actor: 'admin:<user_id>' for audit. The temptation to add a "super tenant" with id * is a footgun — it breaks the discipline that every query is scoped, and the first time someone forgets to filter, every tenant's data leaks into the admin view. Admin routes import the raw db and write every query with explicit WHERE tenant_id = ? clauses.

Migrations and the path out of the simple shape

The single-tenant-to-multi-tenant migration

The path I have shipped twice: add a tenants table and insert the original customer with a stable id like original; add tenant_id to every domain table and backfill with original; wrap every server function in the tenantDb wrapper with the original id read from a config flag; build the tenant signup flow so new customers get a fresh id; rename the original id to a real slug as a one-time migration. About two weeks for a mid-sized app. The mistake I have made twice is trying to skip the wrapper step — leaving the unscoped db global in place "just for the migration". Every unscoped global that survives is a future data leak.

The database-per-tenant split when a tenant outgrows the shared shape

The second migration is the split: one tenant's data volume dwarfs everyone else's, write-lock contention is hurting the shared database. Copy that tenant's rows into a new database, stand up a separate D1 binding, route requests for that tenant to the new database at the middleware layer. One weekend of work; the unsplit — going back to one database — is a rewrite. Operational lesson: design the routing layer for split-ability from day one, even if no tenant is large enough to need the split today. Every database access goes through a getDb(tenantId) function; every server function calls getDb(context.tenantId) instead of importing the global. The day the split becomes necessary is a config change, not a rewrite.

Where this guide stops

This is the shape of multi-tenant SaaS architecture in 2026: row-level isolation for ~95% of apps, tenant identification at the edge, four disciplines that prevent data leaks, cross-tenant analytics that cost nothing extra, a migration path from one tenant to one thousand. The shape is durable; the primitives are replaceable. Swap D1 for Postgres and the row-level discipline transfers. Swap Workers for a Node server and the middleware pattern transfers. The business side — pricing, enterprise contracts, onboarding, compliance — is not covered here.


Closing CTA: TanStack Ship ships every multi-tenant pattern in this guide across nine SaaS apps. See the features page, compare against alternatives, or read the SaaS architecture 2026 guide for broader context. The D1 deep-dive covers the data layer; the TanStack Router advanced patterns walks through the routing shapes.