Chapter 25: Multi-Tenant and Platform Architectures
How do I build a platform that serves multiple tenants securely and efficiently?
Multi-tenancy makes SaaS economics work but is also the architectural decision most likely to wake you at 3am. One codebase serves many customers, each isolated from others while sharing infrastructure; your month-one isolation strategy determines your year-three crises.
Get isolation right and you have a scalable business with predictable costs. Get it wrong and you face failures from noisy neighbour problems to data leaks that end your company; one missing WHERE clause means explaining to Customer A why they can see Customer B's data.
A missing tenant filter can expose data in shared tables. Separate databases contain that particular mistake, but incorrect routing can still select another tenant's database. Test authorisation and tenant-to-resource mapping whichever data model you choose.
Cloudflare's architecture naturally fits multi-tenant patterns. Many small databases rather than one large one, per-entity Durable Objects, Workers for Platforms: these primitives assume multi-tenancy as the default. But "naturally fits" doesn't mean "automatically correct". This chapter covers decision frameworks for isolation strategies, their operational implications, and common failure modes.
Choose the isolation boundaries
Compute, data and state isolation are independent choices. Begin with the boundary a failure must not cross, then choose the mechanism that enforces it.
| Requirement | Mechanism | Responsibility that remains |
|---|---|---|
| Shared application code with tenant-scoped data | Shared Worker and tenant-filtered tables | Establish tenant identity and enforce it on every query |
| Independently managed tenant data | Separate D1 databases | Authenticate tenant routing; automate provisioning and migrations |
| Tenant-contributed code | Separate user Workers in a Workers for Platforms dispatch namespace | Grant bindings and resource limits deliberately |
| Separate administrative ownership | Separate accounts where required | Operate identity, deployment and resource lifecycle across accounts |
A separate database is a data boundary, not dedicated hardware. A separate account is an administrative boundary, not a promise of physical infrastructure separation. Match contractual language to the platform guarantee before offering either as a product tier.
Choosing your data isolation strategy
Data isolation is where the hard decisions live. Two patterns dominate: row-level isolation in a shared database, and database-per-tenant. Each has profound implications for security, operations, and cost. Understand both thoroughly before committing.
Row-level isolation
Row-level isolation stores all tenant data in shared tables with a tenant identifier column, requiring every query to filter by tenant: add tenant_id to every table, include it in every relevant index, and enforce it in every WHERE clause.
The most compelling case for row-level isolation is cross-tenant analytics. If your product includes reports comparing tenant performance to benchmarks or aggregating usage across your customer base, row-level isolation makes these queries trivial (a single SQL statement computes averages, percentiles, or aggregations across everyone), whereas database-per-tenant requires iterating through every database and aggregating results in application code.
The cost structure also favours row-level isolation at scale. One D1 database versus ten thousand affects both your bill and operational surface area. Monitoring one database is simpler than monitoring ten thousand, even with excellent automation.
Database-per-tenant
Database-per-tenant gives each tenant a separately managed D1 database. A query cannot accidentally include rows held in another database, but the shared application may hold access to several databases and can still route a request incorrectly. Authenticate the tenant before selecting its database, and test that mapping as part of the security boundary.
This pattern suits independently managed retention, restores, schema changes and access boundaries. It can help meet contractual isolation requirements, provided those requirements accept separate managed databases rather than dedicated physical infrastructure.
Schema flexibility is another driver. If tenants need custom fields, different indexing strategies, or schema variations, database-per-tenant accommodates this naturally. Each database evolves independently. Row-level isolation forces a single schema on all tenants, with customisation limited to nullable columns or JSON fields.
The operational cost is substantial. Schema migrations become distributed systems problems. Adding a column means running ten thousand ALTER TABLEs, tracking which succeeded, handling failures, and managing the period where different tenants run different schema versions. This requires tooling, monitoring, and operational discipline that row-level isolation doesn't demand.
The economics
D1 does not charge a fixed instance price for each database. The relevant costs are aggregate rows read and written, stored data and the tooling needed to operate the fleet. A thousand mostly idle databases can therefore be economical, but their number alone is not a cost estimate.
Compare alternatives at the same boundary. Several PostgreSQL databases can share one server; database-per-tenant does not necessarily mean an RDS instance per tenant. Include the isolation guarantees, performance contention and administrative work of each arrangement, alongside the bill.
Capacity also has several limits. The per-database size and account database count do not supersede the account's total storage allowance or the Worker's binding limits. Plan tenant routing and account capacity together before committing to a database per customer.
Making the decision
Neither pattern is universally superior. Row-level isolation fits when data sensitivity is moderate, cross-tenant operations are common, and engineering resources for migration tooling are limited. If your product requires heavy cross-tenant analytics or you have thousands of tenants with limited operational capacity, row-level isolation is probably correct.
Database-per-tenant fits when separate lifecycle operations or a database boundary justify fleet management. Sensitive data still requires correct routing and authorisation in either model; the choice should follow a concrete requirement rather than an industry label.
Tenant count matters but isn't absolute. Managing schema migrations across fifty thousand databases differs qualitatively from managing fifty. The tooling investment scales stepwise as you cross operational thresholds. At a hundred tenants, manual intervention during migrations is annoying but feasible. At ten thousand, it's impossible.
Consider which mistake each design contains. A separate database contains a missing tenant filter; it does not contain selection of the wrong tenant binding. A shared database with enforced access policies addresses a different failure boundary. Review the likely failure and its consequence before paying for an isolation mechanism.
Tenant tiering
Many platforms use different isolation levels for different tenant tiers. Free tenants share a database; enterprise tenants get dedicated databases. This resolves the row-level versus database-per-tenant tension for applications where tenant value varies dramatically.
Tiering can concentrate the cost of independent provisioning, migrations and restores on customers who need those capabilities. Price the promised boundary and service level explicitly; annual spend alone does not determine an appropriate isolation mechanism.
Implementation requires routing logic that determines database binding based on tenant tier; shared database for standard tenants, dedicated database for enterprise tenants. This adds complexity but concentrates operational overhead on tenants generating revenue to justify it.
Make tiering explicit in your pricing and contracts. Describe a dedicated database as a separate managed database, including what is shared in compute, administration and support access. Promise physical separation only when the service contract actually provides it.
Migration between models
Can you change your mind? Moving from row-level to database-per-tenant requires significant effort: provisioning databases for each tenant, migrating data from shared tables to dedicated databases, updating all application code to use tenant-specific bindings, and handling the transition period where some tenants are migrated and others aren't. Expect weeks to months depending on data volume and application complexity.
Moving from database-per-tenant to row-level is theoretically possible but rarely done. The scenarios driving database-per-tenant adoption (regulatory requirements, enterprise contracts) don't typically reverse.
Start with the simplest boundary that satisfies the access-control and data-lifecycle requirements. Include future migration work in the decision, but do not assume either shared tables or separate databases is a universally safer default.
Tenant metadata architecture
Tenant metadata, the registry of tenants, their configuration, feature flags, and billing status, differs from tenant data in one crucial way: your application needs it before knowing which tenant database to use. The authentication flow checks credentials, identifies the tenant, retrieves configuration, then routes to their data.
A shared D1 database can hold the tenant registry, domain mappings, entitlements and billing status. This is an authorisation control plane: a wrong mapping or changed entitlement can expose data or grant access. Restrict its writes and protect tenant-visible reads just as carefully as the tenant databases.
Cache metadata according to the consequence of staleness. Branding and display preferences may tolerate minutes; access revocation, domain ownership and billing entitlements may not. Give each category an explicit expiry or invalidation strategy rather than applying one aggressive cache policy to the registry.
Define outage behaviour for each category too. Cached branding can keep a page usable, while an action requiring current entitlement or revocation state should remain blocked if that authority is unavailable. A previously authenticated identity does not establish current permission. Do not extend cached authorisation beyond its agreed freshness policy merely because the registry is down.
How Cloudflare handles noisy neighbours
A shared Worker does not receive a separate isolate for each tenant. Its 128 MB memory limit is per isolate and may be shared by concurrent requests. CPU limits apply to invocations, not to an authenticated customer's aggregate usage. Enforce tenant admission and usage limits in the application.
Separate D1 databases and Durable Objects separate data and coordination boundaries, reducing contention on a single database or object. They do not promise dedicated hardware or immunity from account quotas and shared infrastructure failures. An overloaded object can also become a bottleneck for every request assigned to it.
Workers for Platforms adds an execution boundary for tenant-contributed code. User Workers within a dispatch namespace receive the bindings and invocation limits the platform grants them. Keep those controls separate from the commercial tenant's aggregate quota, which may cover many invocations and services.
State isolation
Durable Objects provide natural tenant isolation through their naming scheme. Including the tenant identifier in the object name ensures each tenant's state lives in separate objects:
const id = env.SESSION.idFromName(JSON.stringify([tenantId, sessionId]));
Each object has private storage, but the application chooses which object a request reaches and which other services it may call. Derive the tenant identity from authenticated context and test cross-tenant attempts. A name is a routing key, not an authorisation check.
Separate bindings can restrict which namespace tenant-contributed code may address. Tenant-prefixed names within a shared namespace are a routing convention; code holding that namespace binding can construct other names. Match the binding grant to the trust boundary rather than describing it as physical compute separation.
Workers for Platforms
Workers for Platforms solves a specific problem: executing tenant-provided code safely. When your platform allows customers to write JavaScript running on your infrastructure (webhooks, custom transformations, integration logic), this is how you isolate their code from yours and from each other.
When configuration suffices
Before reaching for Workers for Platforms, ask whether tenant customisation can be expressed as configuration. Code execution is a capability tax you pay forever; configuration is a one-time design investment.
Most "we need custom code" requirements, examined closely, reduce to "we need more flexible configuration." Tenants want to transform data, apply conditional logic, or customise behaviour. These needs often fit well-designed configuration systems.
Consider field mappings. A tenant wants to transform webhook payloads: "rename field A to field B, extract nested field C.D to top level." JSONPath expressions, field mapping rules, or template strings handle this without code execution.
Consider conditional logic. A tenant wants to route data based on field values: "field X equals Y" or "field X contains Y." A rule engine with predefined operators handles it. You don't need arbitrary JavaScript to evaluate "status equals approved."
Consider formatting. A tenant wants to customise notification messages. Mustache, Handlebars, or simple string interpolation let tenants customise output without arbitrary code.
Workers for Platforms becomes necessary when customisation requires logic you can't anticipate: calling external APIs based on data values, implementing domain-specific business logic, transforming data in ways requiring loops or computation.
Integration platforms, workflow automation tools, and extensible applications fall into this category. The question isn't whether you can avoid code execution but whether your platform's value proposition requires it.
Architecture and implications
A dispatch namespace groups the tenant Workers your platform deploys. Your dispatch Worker authenticates the request and selects the tenant Worker through a binding:
Here authenticateAndResolveTenant is an application helper that authenticates the caller and authorises the selected tenant before returning its dispatch name.
export default {
async fetch(request: Request, env: Env) {
const tenantId = await authenticateAndResolveTenant(request);
const tenantWorker = env.TENANT_WORKERS.get(tenantId);
return tenantWorker.fetch(request);
}
};
This indirection is the security boundary. Authentication, rate limiting, and request validation happen in your dispatch Worker before tenant code executes. You control what data reaches tenant code and what tenant code can do with responses. But once you hand off, you're trusting the sandbox.
Tenant code runs in separate V8 isolates, subject to the runtime’s memory limit and configured CPU constraints. Granted bindings and application routing still determine which data that code can reach; review those capabilities rather than treating runtime isolation as the entire security boundary.
Trust modes
Workers for Platforms offers two isolation modes with different security implications:
Untrusted mode (the default) provides complete isolation between customer Workers, including separate cache namespaces. The caches.default API is disabled entirely, preventing tenants from polluting or reading each other's cached content. Use this when customers control deployed code (the normal case for integration platforms and extensibility scenarios).
Trusted mode allows shared cache access across the namespace. Only appropriate when you, the platform operator, control all Worker code deployed to the namespace. If you're deploying your own code variants rather than accepting customer code, trusted mode enables performance optimisations impossible with full isolation.
Outbound controls
An outbound Worker can inspect and restrict outgoing fetch() requests from user Workers. Use it to enforce destination policy, add credentials without exposing them to tenant code, and record usage.
Pass the tenant identity through the dispatch Worker's trusted outbound parameters. Do not derive it from an arbitrary request header supplied by tenant code. A header named CF-Worker-Tenant-ID is not a documented platform identity mechanism.
Enabling an outbound Worker disables the user Worker's outbound TCP connect() API, but interception does not cover fetches from Durable Objects or mTLS certificate bindings. Review every granted binding as an additional egress capability; one outbound policy is not complete control of every possible bound service.
Dynamic Workers for runtime code execution
Workers for Platforms assumes you know tenant code at deployment time. Dynamic Worker Loader handles the case where code is generated at runtime, whether by AI or by users through a platform's scripting interface. A Worker can instantiate a new Worker on the fly with code specified as a string, running it in its own isolate with precisely scoped capabilities passed through the env object and network access controlled via globalOutbound. The spawned Worker starts in milliseconds, uses a few megabytes of memory, and can be discarded after a single execution.
This is particularly relevant for platforms building AI-powered automation, where an LLM writes code on behalf of a user and that code needs to execute safely. Rather than managing warm container pools or accepting the latency of cold container starts, Dynamic Workers provide per-request isolation at isolate-level speed. The security model is additive: capabilities are explicitly granted rather than filtered from a default-allow posture, which is inherently safer when the code being executed is unpredictable. Chapter 19 covers Dynamic Workers in detail for agent use cases.
Durable Object Facets for per-tenant stateful code
Dynamic Workers solve dynamic code; Durable Object Facets (covered in Chapter 7) solve dynamic stateful code. A supervisor Durable Object you write can instantiate a child Durable Object whose class is loaded at runtime out of a Dynamic Worker, with each child getting its own isolated SQLite database stored alongside the parent's. For platforms running tenant-supplied or AI-generated code that needs persistent state, this collapses what previously required a per-tenant Durable Object namespace into a single namespace whose objects host arbitrary tenant code on demand. The supervisor is the right place to enforce per-tenant logging, billing, and resource limits before tenant code runs; the parent's database stays under platform control regardless of what the child does to its own SQLite. For Workers-for-Platforms-style multi-tenancy where each tenant needs a long-lived stateful database, Facets are the missing piece.
Dynamic Workflows for per-tenant durable execution
Dynamic Workers solve dynamic code and Facets solve dynamic stateful code; Dynamic Workflows solve dynamic durable execution. When each tenant defines their own multi-step automation, onboarding sequences, approval chains, or billing-retry logic, and those processes must survive restarts and long sleeps, you run the Workflow inside a Dynamic Worker rather than deploying a workflow class per tenant. The @cloudflare/dynamic-workflows library tags each instance with the tenant's identity so the correct tenant code reloads when the instance wakes, days later if it must, while the underlying Durable Object preserves the usual durability guarantees. This turns per-tenant durable workflows into a single namespace running tenant code on demand at near-zero idle cost. Chapter 8 covers the mechanism and the case where a plain Workflow taking per-tenant input is the simpler choice.
Per-tenant AI Search instances
If your platform exposes search-over-tenant-content as a feature, the ai_search_namespaces Workers binding (Chapter 18) lets you create a dedicated AI Search instance per tenant at runtime via create(), delete(), list(), and search() methods. Combined with the 5,000-instances-per-account ceiling on paid plans, this makes namespace-per-tenant a viable shape for B2B SaaS knowledge bases. Each tenant has a separate search instance, hybrid search and relevance boosting can be tuned per tenant, and you avoid the ranking contamination that comes from mixing customer corpora into a single index.
The documentation obligation
When you build a platform, you inherit documentation obligations. Cloudflare documents the guarantees, limits, and failure modes of Workers. You must do the same for your platform.
Your tenants can't reason about their code's behaviour without understanding what your platform promises. Document resource limits: CPU time, memory, execution duration. Document available bindings and APIs. Document failure modes: what happens when limits are exceeded, when outbound requests fail, when your infrastructure has issues. Document what you log and monitor.
When a tenant's code fails mysteriously, they need documentation to diagnose the issue. When a tenant wants to push limits, they need to know what limits exist. This documentation isn't optional; it's part of running a platform.
Operational reality
Running tenant code creates operational challenges that shared Workers avoid.
Debugging is harder. When a tenant reports their Worker is failing, you're investigating code you didn't write. Your observability is limited to logs, error messages, and resource consumption. Building good debugging tools for tenants reduces support burden.
Version management is complex. Tenants update their code independently. Platform changes (new APIs, deprecated features, security patches) may break tenant Workers. You need a strategy for communicating changes, providing migration periods, and handling tenants who don't update.
Resource limits require tuning. Too restrictive, and legitimate tenant code fails. Too generous, and one tenant's expensive computation affects platform stability. Start conservative and increase limits based on observed legitimate usage.
Support burden increases. Tenants will ask for help with code that doesn't work. Draw the line between platform support and coding assistance explicitly. "Your code timed out" is platform support. "Why does my JavaScript throw a TypeError" is coding assistance.
Custom domains
Custom domains create operational dependencies on infrastructure you don't control. Enterprise tenants expect their branding (your API accessible at api.tenant.com rather than api.yourplatform.com/tenants/tenant), and customers pay more for white-labelling, but this couples your platform's reliability to your tenants' DNS hygiene.
Cloudflare for SaaS handles certificate issuance, renewal, and routing, while your responsibility covers the tenant experience, operational edge cases, and resilience to failures outside your control.
The routing foundation
Your Worker must identify tenants by domain:
const hostname = new URL(request.url).hostname;
const tenantId = await getTenantByDomain(hostname, env);
if (!tenantId) {
return new Response("Unknown domain", { status: 404 });
}
Cache domain lookups only under an explicit stale-routing policy. Validate ownership and domain reassignment at an authoritative boundary, and check the resolved tenant against the authenticated request. A KV update or deletion does not immediately remove cached mappings everywhere.
Support fallback patterns. Subdomains of your platform (tenant.yourplatform.com) should coexist with custom domains, providing a working configuration while tenants sort out custom domain setup. When custom domain lookup fails, check whether the request arrived at a platform subdomain and route accordingly.
Verification strategy
Adding a custom domain requires the tenant to prove ownership. The verification method affects tenant experience significantly.
HTTP validation can simplify onboarding once the tenant's DNS points at the platform. Track validation and certificate status explicitly, and show what the tenant needs to correct when provisioning stalls.
DNS TXT validation works before DNS points to your platform. Tenants add a verification record, you confirm it, then they update their A or CNAME records. More complex but allows verification before go-live; suits tenants who want to validate before cutting over production traffic.
For less technical customers, consider the support implications. DNS configuration confuses people who don't do it regularly. Documentation, validation status dashboards, and proactive support for stuck domains improve the experience. Monitor tenants who start custom domain setup but never complete it; they may be stuck.
Resilience to external failures
DNS propagation takes time. Tenants will configure their DNS and immediately complain that custom domains don't work. Set expectations explicitly: propagation takes minutes to hours depending on DNS provider and TTL settings. Provide tools to check configuration status.
Certificate validation can fail for reasons outside your control: DNS misconfigurations, CAA records blocking issuance, propagation delays. Monitor certificate status across tenants and alert on persistent failures. Distinguish expected delays from genuine problems.
Tenants forget to renew their domains. When a tenant's domain expires, their registrar may redirect traffic to a parking page, making your platform look broken. Monitor for unexpected DNS changes, certificate validation failures on previously working domains, or sudden traffic drops. Notify tenants proactively when domain configuration appears broken.
Build your platform to handle custom domain failures gracefully. When a custom domain stops working, the tenant's platform subdomain should still function. Don't rely solely on custom domain routing; maintain subdomain routes as fallback.
Usage tracking and billing
Usage tracking serves three purposes: billing customers accurately, enforcing quotas, and understanding platform capacity, though the precision required differs by purpose.
Accuracy requirements
Decide what the customer is being charged for before choosing the meter. Operational trends can use sampled estimates. An invoice needs a documented unit, treatment of retries and failures, and enough retained evidence to explain the charge. Sampling can support a contract based on estimated usage, but its error must be acceptable at the individual tenant's volume.
Delayed counts and lost counts are different. Delayed events can be reconciled; increments lost when an isolate is evicted do not reappear at month end. Do not assume these losses average out.
The metering cost problem
Writing once per request can make metering more expensive than the work being measured. Batch durable usage events or counter updates, assign stable identifiers where retries can duplicate them, and reconcile the results against an independent measure. Keep the unflushed-loss window explicit if a cheaper, approximate meter is a deliberate product choice.
Analytics Engine provides weighted usage aggregates and supports usage-based billing designs. Match that sampled model to the billing contract; use a durable event ledger when disputes require exact event-level evidence. Chapter 21 covers the sampling arithmetic.
Quota enforcement patterns
The naive approach reads current usage, compares to limit, then proceeds or rejects. It adds latency to every request and doesn't work correctly under concurrent load.
Asynchronous metering lets requests proceed without waiting for a quota update. The overage depends on propagation delay, concurrent work and enforcement frequency; it is not automatically small. Use this when the product can tolerate that uncertainty, and define how to charge for or reject further work once the limit is observed.
For hard limits, a Durable Object can make the admission decision atomically for one tenant or quota key. The check adds a round trip to that object's location, which may be intercontinental. Decide whether the strict global limit justifies that latency, or whether local allocations with an explicitly bounded overage meet the requirement.
Reserve Durable Objects for quotas that must be hard: API rate limits where abuse is a concern, resource limits protecting your infrastructure, or contractual obligations requiring exact enforcement.
Database-per-tenant operations
If you choose database-per-tenant, several operational realities require planning before you have a thousand databases with no tools to manage them.
Migration as distributed systems
Schema migrations become distributed operations. Adding a column means executing the same change across every tenant database, tracking which succeeded, handling failures, and managing the transition period.
Track schema versions durably per tenant and make the runner safe to resume after an uncertain outcome. Check whether the database change committed before retrying it. An IF NOT EXISTS clause helps only for statements that support it; it is not a general migration protocol.
The fleet's progress record and each database's actual schema can diverge if an update fails between them. Reconcile those two facts instead of treating a separate KV status write as proof that the schema operation and its record committed together.
Build tooling that expects partial failure. When migration 4,293 fails, you need to know immediately and understand why: data constraint violation, timeout, infrastructure issue. Your tooling should support both stopping to investigate and continuing with remaining tenants.
Consider lazy migration for non-critical changes. Rather than migrating all databases immediately, check schema version on first access and migrate then. This spreads load over time and ensures you only migrate active tenants. Dormant tenants migrate if and when they return.
Cross-tenant analytics
You cannot query across D1 databases; no federation, no cross-database joins. If you need data from multiple tenants, iterate through databases and aggregate in application code.
For reporting and analytics, maintain aggregate tables separately. Process usage data from tenant databases into a shared analytics database periodically. Tenant data stays isolated, but aggregated, anonymised metrics live in a shared store optimised for cross-tenant queries.
Consider what aggregates you'll need before you need them. Adding aggregation after the fact requires backfilling from thousands of databases.
:::warning Subrequest Limits and Cross-Tenant Operations A high subrequest allowance does not make one invocation a safe fleet-wide migration or reporting job. Bound concurrency, track progress and checkpoint work so it can resume after interruption. Account for database capacity and the invocation's CPU and lifecycle limits as well as the number of calls. Use queued batches or durable orchestration when the operation must survive independently of an HTTP request. :::
Tenant lifecycle
Provisioning creates resources that must be tracked, and a robust flow handles failure at any step. Database creation might succeed but schema application might fail (empty database, partially onboarded tenant; detect and retry or clean up). Schema might apply but storing the tenant record might fail (functional database you can't find; include the database identifier in the database name so you can discover orphans through the API). Tenant record might store but subsequent setup might fail (registry points to valid database but default data, welcome emails, or integrations didn't complete; design onboarding as a state machine that can resume from any point).
Track resources created through every entry point. An agent created from a Slack message needs the same tenant ownership record as one created through the web interface; a list assembled only by the dashboard will miss it during export or deletion. Keep the tenant-to-resource mapping in a shared provisioning path, including Durable Objects and search instances created after onboarding.
Offboarding must close those creation paths before cleaning up. Record the tenant's closing state and enforce it when admitting work from integrations, scheduled jobs and queued messages. Account for work already in flight, which may still create resources or write data. Keep the tenant marked closed after resource deletion so a delayed webhook cannot recreate its state. Apply the agreed export, grace-period and retention rules, then reconcile the resource inventory before declaring offboarding complete. Chapter 23 covers stopping and reconciling work across these boundaries.
Build reconciliation processes. Periodically enumerate all resources and verify they belong to active tenants. Flag orphans for investigation. This catches provisioning failures, offboarding bugs, and edge cases you haven't imagined.
What goes wrong
Multi-tenant failures are predictable in kind if not in timing. The question isn't whether you'll face data isolation bugs, migration failures, and resource exhaustion; it's whether you've built detection and mitigation before they become incidents.
Data isolation failures
In row-level isolation, the most common failure is a query without tenant filtering: typically in new code paths, rarely-exercised error handlers, or debugging sessions where someone runs a query manually.
Prevent it through multiple layers. ORM configurations can add tenant filters automatically; configure this as default, not opt-in. Database views can include tenant filters, ensuring raw table access is never needed in application code. Code review checklists should specifically verify tenant filtering. Integration tests should attempt cross-tenant access and verify failure.
Detect it through audit logging. Log all database queries with tenant context. Anomaly detection can flag queries returning data for multiple tenants or query patterns that don't match expected access.
Migration failures
In database-per-tenant, migrations fail partially. Some tenants migrate successfully; others fail. Causes vary: data violating new constraints, queries timing out on large tables, transient infrastructure issues.
Schema drift is a subtle variant. With lazy migration, different tenants run different schema versions. Your application must handle this (either supporting multiple schema versions simultaneously or forcing migration before operations requiring newer schemas). Design this handling explicitly; discovering schema incompatibility at runtime creates debugging challenges.
Resource exhaustion
A shared Worker does not automatically enforce each commercial tenant's resource budget. Separate databases or objects can reduce contention, but account limits and shared application dependencies remain. Monitor and enforce tenant usage at the boundary that admits the work.
Monitor per-tenant resource consumption. Track CPU time, subrequests, and database operations by tenant. Rate limit tenants approaching concerning thresholds before they impact the platform.
With Workers for Platforms, configure resource limits conservatively. Tenant code running at maximum allowed time on every request consumes substantial resources. Start tight and increase based on observed legitimate usage.
Signs that the operating model needs to change
Tenant count is a rough proxy for operational load. The clearer signals are a growing provisioning backlog, migration failures that require repeated manual recovery, cross-tenant reports consuming production query capacity, or one customer's activity dominating a shared dependency.
Automate the operation whose failure or delay is becoming material. Measure completion time and recovery effort, not just the number of databases. A small fleet of heavily customised tenants can demand more tooling than thousands of uniform, mostly idle ones.
Reference architecture
The entry point is a dispatch Worker handling all incoming requests. It authenticates requests, identifies tenants from tokens or domains, retrieves tenant configuration from cached metadata, and routes to appropriate handlers. Tenant code never runs without authentication succeeding first.
Tenant metadata can live in shared D1, with caching chosen per field. Branding and display preferences may tolerate stale reads; domain ownership, revocation and billing entitlements require the freshness promised by the access policy. Apply the same distinction in every request path.
Tenant data lives according to your isolation strategy: shared D1 with tenant filtering for row-level isolation, dynamic binding resolution for database-per-tenant, or routing logic selecting shared or dedicated databases for tiered isolation.
Tenant state uses Durable Objects with tenant-prefixed names. Sessions, real-time features, and coordination all route through objects named with tenant identifiers.
For platforms with tenant-contributed code, a dispatch namespace contains tenant Workers. The dispatch Worker routes to tenant code after authentication, with outbound Workers controlling external access.
Usage tracking aggregates in a dedicated analytics database, populated by background processing. Quota enforcement uses Durable Objects for hard limits, optimistic checking for soft limits.
Custom domains route through Cloudflare for SaaS, with domain-to-tenant mappings in the metadata database. Fallback to platform subdomains ensures availability when custom domain issues arise.
Your specific platform will differ, but the pattern of centralised authentication, cached metadata, appropriate data isolation, and explicit resource boundaries applies broadly.
What comes next
Tenant boundaries make the platform's limits concrete: a single large customer, a residency obligation or a demanding workload can change the design. Chapter 26 gathers those reasons to choose a different runtime, storage service or platform.