Skip to main content

Chapter 24: Architectural Patterns and Reference Designs

What patterns should I follow for common architectural challenges?


Part VI established operational foundations: cost management, observability, secure deployment and recovery. Part VII applies everything from this book to real architectural challenges.

Patterns encode judgement. The code is the easy part. Knowing when to apply which pattern separates architecture from implementation. This chapter presents architectural patterns adapted for Cloudflare's edge platform, not as recipes to follow blindly, but as decision frameworks for your specific constraints.

Each pattern exists because it solves a real problem. But patterns have costs: complexity, latency, operational overhead. The goal isn't applying patterns because they're elegant. It's applying them because they're appropriate. Sometimes the right answer is no pattern at all.

The virtue of boring​

Before exploring patterns, internalise this principle: boring is good. Software systems should be predictable, doing what they're supposed to do without surprises, puzzles, or heroics. The lack of excitement and suspense is actually a desirable property of source code, unlike a detective story.

This principle guides pattern selection. The right pattern isn't the cleverest or most elegant; it's the one that makes your system boring. Predictable. Understandable. Debuggable by someone at 3am who's never seen the code.

When evaluating patterns, ask whether they make the system easier or harder to understand, introduce complexity that will confuse future maintainers, or solve problems you actually have.

Clever architectures become legacy nightmares while simple architectures endure. When in doubt, choose boring.

Establishing your latency budget​

Before selecting patterns, establish your latency budget. A checkout flow might allow two seconds total; a real-time game might require sub-100ms response. Work backwards from your budget to determine which patterns you can afford.

Start with measurements for the dependencies and user regions that matter. A colocated service binding avoids a public network round trip; a call to another location pays for the distance. The same applies to Durable Objects and external databases. Use the p95 or p99 path your product promises, including serial calls, rather than treating one typical hop time as a platform guarantee.

If your budget is 200ms and your backend database call takes 150ms, you have 50ms for everything else. A gateway Worker with rate limiting might consume 15ms, while a cross-continental Durable Object call for user session data would blow the budget entirely. These numbers should drive your architecture, not the other way around.

The patterns in this chapter have different latency profiles: gateway patterns add a hop but enable caching that can eliminate hops entirely; event-driven patterns trade immediate latency for eventual processing; collaboration patterns accept cross-continental latency as the cost of strong consistency. Know your budget before choosing your pattern.

Choosing your pattern​

Before diving into individual patterns, consider what problem you're solving.

ProblemPatternCloudflare PrimitiveWhen to Avoid
Centralised auth, routing, rate limitingAPI GatewayWorker + service bindingsSimple apps with single backend
Client-specific API shapesBackend for FrontendWorker per client typeSmall teams, stable API
Real-time synchronisationCollaborationDurable ObjectsLow-frequency updates, tolerance for polling
Multi-step reliable processesSagaWorkflowsSimple operations that can retry atomically
Decoupled async processingEvent-drivenQueuesSynchronous requirements, ordering guarantees

The patterns build on primitives covered earlier: Workers for compute, Durable Objects for coordination, Workflows for durability, Queues for decoupling. If you haven't internalised those primitives, the patterns will feel arbitrary. If you have, they'll feel inevitable.

Choose your primitive: sync, coordinated, ordered, or fan-out

Loading diagram…

Use this as a starting point for the dominant requirement. A design can combine primitives: a queue consumer may call a Durable Object, and a Workflow may coordinate queued work.

API gateway pattern​

Every Worker is already a gateway. It receives requests, processes them, returns responses. The pattern question isn't whether to have a gateway; it's whether to make it explicit.

Where the gateway earns its place​

A gateway can authenticate callers, apply admission policy and route to internal services. Make its authority explicit: which backends can it call, what caller context does it assert, and can another path reach those backends?

A backend can rely on authentication performed by a trusted gateway when that trust boundary is enforced. It still needs to authorise the requested operation and resource. Rechecking permissions does not make the gateway redundant; central authentication and resource-specific authorisation solve different parts of the request.

Service bindings avoid a public HTTP endpoint for each internal service and can keep calls on the same server thread. Measure the complete path, including authentication lookups and quota checks. A nearby gateway does not make a distant database or quota authority local.

The decision framework​

Make the gateway explicit when: multiple backend services need consistent authentication, rate limiting must be globally consistent, you need request/response transformation, or observability requires a single instrumentation point. Keep routing implicit when: single backend service, auth handled externally through Cloudflare Access, the gateway would be pass-through with no logic, or every millisecond matters and you've measured the gateway's cost.

Choose the boundary around shared policy and ownership, not a backend count. A single service may need a dedicated gateway for a strict ingress boundary; several services may already share suitable authentication infrastructure.

Implementation essence​

The gateway authenticates, checks admission policy and dispatches through a configured binding. Pass a narrow caller context derived from validated credentials; do not forward user-supplied identity headers as trusted claims.

The receiving service checks that the caller may act on the requested resource. Keep service entrypoints private where possible, and test alternative hostnames and internal routes for bypasses. That test is more valuable than another generic routing wrapper.

Trade-offs to accept​

Gateway centralisation means gateway outages affect everything. Your gateway Worker becomes critical infrastructure. Test thoroughly, monitor obsessively, accept that Cloudflare issues take down your gateway with everything else.

The gateway becomes a coordination point for deployments. Changing auth logic or rate limits requires gateway deployment, not backend deployment. For some teams, this enables centralised policy management; for others, it's friction slowing backend teams.

Rate limiting at the edge​

Rate limiting deserves special attention because it illustrates the edge advantage most clearly, and because getting it wrong creates either security vulnerabilities or latency problems.

Local abuse controls or a global quota​

A strict shared counter needs an authority that every relevant request consults, or a scheme that allocates limited capacity to local authorities. The choice is the same whether that authority is Redis, a database or a Durable Object. Distant callers pay for the coordination path.

Use local rate limits when the aim is to make abuse expensive and some distributed overage is acceptable. Define that scope explicitly; a per-location limit is not the same product promise as one global quota.

A Durable Object can own the admission decision for one quota key. Partitioning by user or API key spreads independent counters, and automatic routing removes the need to maintain an object directory. A single global key still concentrates requests on one object. Users who move or share an API key across continents may also be far from their object.

For strict quotas, measure that round trip and keep the update atomic. For high-volume distributed use, consider allocating bounded local budgets if the domain allows it. The useful distinction is the required precision and scope of the limit, not whether the counter's product is described as global.

Backend for frontend​

Different clients need different API shapes. Mobile apps on cellular connections need minimal payloads. Web applications need richer data. Internal tools need everything. The Backend for Frontend pattern provides client-specific facades that aggregate and transform data for each client's needs.

Why the pattern exists​

Generic APIs serve the lowest common denominator, returning too much data (wasting bandwidth for mobile clients) or too little (forcing multiple round-trips from web clients). BFF acknowledges that different clients have legitimately different needs and provides APIs shaped for each.

Give each BFF an owner who understands that client's release cycle and data needs. Separate client teams can make ownership natural, but a small team can also justify a BFF when it removes substantial client complexity. Count the extra interface and deployment work.

Placement and parallelism​

A BFF can combine calls and trim responses before they cross a slow client connection. That benefit comes from aggregation, and can exist whether the BFF runs near the user or near the backends.

If four independent backend calls each take an assumed 180ms, issuing them sequentially spends about 720ms waiting; issuing them together spends roughly one such wait, before processing and other network legs. Confirm the calls are independent. Neither a BFF nor Promise.all() can parallelise a query that needs the preceding result.

Place backend-heavy aggregation near its data, and lightweight client-specific transformation where measurement supports it. Compare the complete response from the client, including payload size and the slow tail.

The decision framework​

Use a BFF when clients need materially different aggregation, payloads or release cycles and an owner can maintain those contracts. Keep one API when field selection or a shared response already serves them well. GraphQL may provide the query flexibility required, but still needs authorisation and controls on expensive queries.

Implementation essence​

A BFF Worker is an aggregation point. It knows what its client needs and fetches exactly that.

These handler excerpts receive an authenticated user ID; the web query helpers are application code.

Client-specific dashboard handlers: application excerpt
async function getMobileDashboard(userId: string, env: Env) {
// Mobile needs summary only: single query, minimal payload
const summary = await env.DB.prepare(`
SELECT COUNT(*) as orders, SUM(total) as revenue
FROM orders WHERE user_id = ?
`).bind(userId).first();

return Response.json({ summary });
}

async function getWebDashboard(userId: string, env: Env) {
// Web needs detail: parallel fetches, rich payload
const [summary, recentOrders, topProducts, notifications] = await Promise.all([
getSummary(env, userId),
getRecentOrders(env, userId),
getTopProducts(env, userId),
getNotifications(env, userId)
]);

return Response.json({ summary, recentOrders, topProducts, notifications });
}

The mobile endpoint makes one database call and returns a small aggregate payload. The web endpoint makes four parallel calls and returns a larger detailed payload. The data is related, but each client receives a different shape. Each BFF knows its client intimately; when client needs change, the BFF changes with them.

A/B testing at the edge​

A/B testing traditionally requires client-side JavaScript that flickers as variants load, or server-side infrastructure that adds latency to every request. Edge A/B testing eliminates both problems: the Worker intercepts the request, assigns the user to a variant, and routes to the appropriate content before the response begins.

Why edge A/B testing​

The latency advantage is significant. Traditional server-side A/B testing requires a round-trip to a central assignment service before routing can occur. Edge A/B testing makes the assignment decision in the same location handling the request, adding sub-millisecond overhead rather than tens of milliseconds.

Same-URL testing becomes straightforward. Both control and test variants serve from identical URLs. The Worker routes to appropriate origins based on user assignment, eliminating client-side variant detection and the associated flicker.

The assignment pattern​

Keep assignment stable for the experiment's lifetime. A persisted cookie or a deterministic hash of an appropriate identifier can achieve that; choose according to the unit being tested and the privacy requirements. Validate an existing variant against the current experiment configuration before routing.

Separate assignment from authorisation. A user changing an experiment cookie must not gain access to protected features. Preserve path and query parameters when forwarding, and make the cache key distinguish variants so one participant's response does not become another's treatment.

Log exposure only when the treatment was actually served. Assignment, exposure and conversion are different events; counting the first as the second can bias results even when routing works correctly.

Configuration-code separation​

Store experiment configuration in KV, completely decoupled from Worker deployments. Traffic allocation changes, experiment activation, and targeting rule updates propagate globally without redeploying code.

This separation enables marketing and product teams to manage experiments without engineering involvement for routine changes. The Worker code handles mechanics; KV holds business logic.

When edge A/B testing fits​

Use edge A/B testing when every request must be assigned quickly and you control the serving infrastructure. It excels at page-level experiments (different landing pages), feature flag evaluation, and traffic splitting for canary deployments.

Consider alternatives when you need sophisticated analytics integration (most analytics platforms have their own A/B testing), when experiments span multiple domains or platforms, or when you're already invested in a feature flag service that handles assignment.

Cloudflare's own entry in that last category is Flagship, in public beta: feature flags with targeting rules and percentage rollouts, evaluated inside your Worker through a native binding with no outbound call, and flag changes propagating globally within seconds. It's built on the same KV and Durable Objects primitives this pattern uses, so evaluate it before building your own assignment machinery, with the usual beta caveats about evolving APIs and unannounced pricing.

Real-time collaboration​

A Durable Object gives a collaboration room one address and one owner of authoritative state. Clients can reach that owner over WebSockets without the application maintaining a directory of connection servers. This removes infrastructure work, but the product still has to decide what an edit means when its author is offline or looking at an old version.

Hibernation preserves eligible WebSocket connections while an object is inactive; application memory must be reconstructed when it wakes. It does not deliver messages across a disconnected network or reconcile changes saved on a device. Chapter 7 covers the connection lifecycle. The larger design question is what the user may safely believe before the server has accepted their work.

Choose authority from the invariant​

Three approaches serve different product promises. An application may use all three for different operations.

ApproachSuitable workMeaning while disconnected
Server accepts each commandReserving scarce capacity, approving a releaseThe operation remains pending until the authority accepts it
Client applies changes optimistically; server validates themForms, comments, ordinary record editsThe user sees a provisional result that may be rejected or rebased
Replicas use defined merge rulesCollaborative text or shared drawingsLocal edits can proceed; convergence depends on the data model and protocol

Choose from what must remain true. Two inspectors can add independent observations without agreeing on a global order. Two users cannot both reserve the last appointment merely because their local screens say it is available. A conflict-free replicated data type can make replicas converge under its merge rules; convergence alone does not enforce that booking invariant.

A server owner can serialise acceptance without requiring every keystroke to wait for a round trip. Keep typing local and reconcile provisional edits, while commands such as “approve inspection” require authoritative acceptance. This gives the application different consistency choices at the level where users experience them.

Follow an inspection through disconnection​

Consider an inspection application. A Durable Object owns each inspection's editable record and accepted operation history. R2 stores photographs. Two inspectors, Alice and Ben, open revision 40; Alice loses connectivity while Ben changes the equipment's condition from “serviceable” to “unsafe”. Alice edits a note, takes a photograph and marks the inspection complete on her device.

The interface must distinguish “saved on this device”, “waiting to submit” and “accepted by the service”. A local write may protect Alice from an application restart, subject to the device's storage behaviour, but it does not put her work on the server or on another device. In particular, the application must not present her offline completion as an approved inspection.

Store proposed operations durably on the client before reporting local save success. Give each one a stable operation identifier, its target inspection, schema version and the revision or field version it was based on. Derive the actor from authenticated identity at submission, not from a trusted-looking user ID in the payload. Keep the original proposal until its acceptance or rejection is durably recorded locally.

On reconnection, Alice resubmits pending operations with their existing IDs. The owner checks current permission, validates the proposal and commits the mutation, operation receipt and new sequence in one local transaction. A repeated ID with the same payload returns the recorded outcome; the same ID with different content is an error. Retain enough receipts to cover the supported offline and retry period. Record terminal rejections under the same operation ID too; resolving a conflict creates a new proposal and ID. A transport failure can be retried, but an already rejected operation must not quietly become accepted because unrelated state changed.

Output gating delays outbound communication behind pending storage writes under the default storage behaviour. That supports durability before acknowledgement, but the application must still commit the correct state and receipt together. A durable wrong decision is still wrong. Neither output gating nor single-threaded JavaScript authorises Alice's change or detects its conflict with Ben's edit.

Make conflicts a product decision​

Alice's new observation can be appended independently if that is the record's meaning. Replacing the entire inspection with her revision-40 copy would erase Ben's warning. Send a field-specific intention or a mergeable edit instead of uploading a stale document as the new truth.

For a conflicting condition field, return the current value and preserve Alice's proposed value for resolution. For completion, validate the latest record and required evidence at the owner. The server might accept the note, leave the photo pending and reject completion until an authorised inspector reviews the unsafe condition. Partial acceptance needs per-operation outcomes; a single “sync successful” banner cannot express this result.

Choose automatic merge only where both outcomes retain their meaning. Appending observations may work; combining two conflicting safety assessments into one arbitrary value does not. A last-write-wins rule can be appropriate for a disposable preference, but device clocks and arrival order are weak substitutes for a deliberate conflict policy on consequential records.

Presence belongs outside this durable history. A cursor position or “Ben is typing” indicator can expire and be replaced by a newer value. Persisting and replaying every cursor movement wastes storage and can display a user as active long after they have left. Define different delivery and retention rules for awareness and accepted work.

Recover the state, not just the socket​

A reconnecting client supplies the last server sequence it has applied. The owner returns later accepted operations, with a defined switch to a snapshot when the client's cursor is older than retained history. A snapshot must identify the sequence it includes; subsequent changes start after that boundary. Otherwise an edit racing snapshot creation can be applied twice or missed. Establish the live subscription against that catch-up boundary, buffering concurrent updates where necessary. The client suppresses duplicates and fills sequence gaps before advancing its cursor, so an edit cannot disappear between replay and live delivery.

The client keeps its unaccepted proposals separate from the authoritative snapshot. It replaces or updates the accepted base, then reapplies only still-valid provisional work. Blindly replaying every local mutation after downloading a snapshot can resurrect rejected changes or duplicate an already accepted observation whose acknowledgement was lost.

Keep a history identity as well as a sequence. If recovery replaces the server's history, sequence 500 in the restored history may mean something different from the client's sequence 500. A changed history identity forces a deliberate resynchronisation rather than treating matching numbers as matching state. Retain pending work for review instead of silently deleting it during that reset.

Receipt expiry limits retries as well as storage. Once an operation is older than the supported retry boundary, a missing receipt no longer establishes that it is new. Reject it for reconciliation instead of admitting it automatically. Fetching a fresh snapshot must not reset the old operation's identity or age; otherwise an observation accepted before a lost acknowledgement can be appended again. Enforce the retry boundary using server-recognised history or validity information, rather than trusting the device to report when it created the operation.

Deletion needs a similar boundary. A tombstone records that an entity was deleted so an old client cannot recreate it by replaying an edit. If tombstones expire, clients older than that retention boundary must fetch a fresh base and cannot submit stale operations as if the history were still complete. Support for months of offline work has a retention and migration cost even if opening the socket is cheap.

Recheck permission before exchanging data​

Now suppose Alice's access was revoked while she was offline. On reconnect, check access before returning the snapshot or accepting her pending changes. A valid cursor and a previously authorised device do not establish current permission. Keep the refusal distinct from a network failure so the client stops retrying an operation that cannot succeed.

For this application, revocation takes effect when the inspection owner records it. Every subsequent mutation checks that authoritative state, and connected subscribers lose access to subsequent updates. If permissions live elsewhere, define how that change reaches the owner and what delay the product permits. Caching permission at WebSocket creation can leave a revoked user able to read or write for the life of the connection.

The local proposal still deserves deliberate treatment. The product may retain it in a restricted local state for an authorised recovery process, or remove it according to policy. It must not bypass revocation to upload it through an alternate endpoint. Nor can a server guarantee deletion of plaintext already copied to a disconnected device. Offline access expands the data boundary to the device; choose its encryption, retention and account-switching behaviour accordingly.

Photographs expose another boundary. Uploading bytes to R2 does not attach them to an accepted inspection. Issue a scoped upload grant only after authorisation, store the file under an application-selected key and require a separate authorised command to associate it with the inspection. A previously issued presigned URL may remain usable until expiry, so quarantine uploads and recheck permission before attachment or publication. Finalise evidence under a protected immutable version key and bind the attachment to its version and digest; a reusable staging grant must not alter an accepted photograph. Coordinate abandoned-upload removal with acceptance as in Chapter 23.

Account for placement, scale and upgrades​

The inspection owner has a location. Clients elsewhere pay the distance to it when obtaining acceptance, even when their typing feels immediate. Default placement is near the initial request where possible; location hints are best-effort preferences. Test user-to-owner paths from the regions you serve rather than treating the first user's city as a placement guarantee.

Eligible hibernation avoids compute-duration charges while idle, but active processing and storage still cost money. Compare a bursty room with a continuously active one before projecting savings.

A busy room also concentrates validation and fan-out on one object. Measure message rate, payload size and recipient count together. A distribution tree can spread delivery work when needed, but it adds recovery and ordering boundaries and does not make a single authoritative mutation stream unlimited. Managed collaboration services remain credible choices when their client libraries, history and presence semantics remove more work than owning the protocol saves.

Disconnected clients also return with old code. Version operation envelopes and define a supported upgrade window. A server may translate an old proposal only when its meaning is preserved; otherwise ask the client to upgrade while protecting its unsent work. Test a deployment against an old client with pending edits, not only a newly opened browser.

Adopt a sync library for its data model​

A suitable library removes protocol work when its semantics fit. tldraw sync provides a concrete collaborative-canvas model and a Cloudflare template using Durable Objects for rooms and R2 for assets. Its production integration still needs application authentication, authorisation and upload controls.

LiveStore takes another approach, synchronising events and materialising application state locally. Its documentation describes a central ordered event log and client rebasing, with separate limitations around conflict handling and compaction. Evaluate the supported version against your required behaviour; do not infer that two libraries solve the same problem because both can use a Durable Object.

Compare data ownership, supported merges, schema upgrades, licence and the responsibilities left to your team. For a small form with occasional edits, optimistic updates and explicit conflict rejection may be enough. A shared rich-text editor deserves a proven editing model. In either case, rehearse Alice's sequence: disconnect, concurrent change, permission revocation, reconnection and an ambiguous acknowledgement. The protocol is ready when the user can understand what survived and the authoritative record explains why.

Event-driven architecture​

Events decouple producers from consumers. Services emit events when state changes; other services react asynchronously. The pattern enables independent scaling, fault isolation, and system evolution without tight coordination.

Why events matter​

Synchronous request-response creates coupling: when Service A calls Service B directly, A waits for B, fails if B fails, and must know where B lives. For simple systems, this coupling is fine; it's not inherently bad. But for complex systems with many services, coupling becomes brittle.

Events are for consequences, not requirements. If the caller needs to know whether an operation succeeded, that's a synchronous call. If the caller just needs to announce that something happened (order created, payment received, user signed up), that's an event. The caller moves on; interested parties react when ready.

What's different at the edge​

Queues move work out of the user's request path and give consumers up to 15 minutes of wall time per invocation. CPU time has a separate configurable limit; the wall-time allowance is not a CPU budget. Use a queue when work should survive the client disconnecting, absorb bursts, or retry independently; changing the trigger does not provide more memory or unlimited computation.

Cloudflare Queues provide at-least-once delivery. Design consumers so repeated delivery produces one intended business effect; the available mechanism depends on where that effect commits.

The idempotency imperative​

At-least-once delivery means the same event can be processed again after a timeout or crash. Checking for a receipt, performing the effect and then inserting the receipt leaves two gaps: concurrent consumers can both pass the check, and a crash can occur after the effect but before recording it.

When the effect and receipt share a database, commit them in one transaction with a unique event identifier. For an external effect, pass a stable idempotency key to a receiver that honours it. If the receiver offers no such guarantee, design reconciliation or compensation; a local receipt cannot make an external call atomic.

Retain receipts for the full retry and replay window. The useful guarantee is that repeated delivery produces one business effect, not that the transport invokes the handler only once.

Queues vs Workflows​

Both handle asynchronous work but solve different problems. Queues process independent items where order doesn't matter and items don't depend on each other. Workflows orchestrate dependent steps where step 2 needs step 1's result and failures require compensation.

Processing a batch of independent tasks (uploaded files, notification deliveries, data imports)? Use Queues. Coordinating a multi-step process (order fulfilment, account provisioning, data migration)? Use Workflows. The saga pattern shows how Workflows handle coordination.

Saga pattern for distributed transactions​

A saga records progress across several independently committed actions and defines compensation when a later action fails. A booking might reserve stock, take payment and confirm delivery; a failed delivery reservation may require releasing stock and refunding payment.

Workflows can persist this progress and retry individual steps, as Chapter 8 explains. The service cannot choose the business meaning of compensation. Decide which effects are reversible, which require a new action such as a refund, and who handles a compensation that also fails. Intermediate states remain visible, so expose a pending or recovery state rather than presenting the process as an atomic transaction.

Trade-offs to accept​

Event-driven systems are harder to debug. Tracing a request through synchronous calls is straightforward: follow the stack. Tracing events through queues requires correlation IDs and distributed tracing. Invest in observability before you need it.

Eventual consistency is the default. After publication, consumers may lag, fail or exhaust their retries. Track progress and unresolved outcomes rather than treating publication as completion. If downstream systems need immediate consistency, events are the wrong pattern.

Queues does not guarantee ordering. A Workflow can order steps within one instance, but it does not order independently delivered events. When entity updates must be applied in order, retain authoritative sequence information, detect gaps and replay missing updates; otherwise design events to be order-independent.

Service composition​

Complex applications comprise multiple services. How those services communicate affects latency, reliability, and operational complexity. Choices here ripple through your architecture.

The three-layer model​

When decomposing a system, the three-layer architecture provides useful vocabulary: the interaction layer handles user-facing concerns (API gateways, BFFs, authentication); the control layer contains business logic and orchestration, deciding what to do based on requests and coordinating components; the mechanism layer provides capabilities (storage, external APIs, notification services).

A gateway Worker is interaction; a Workflow coordinating an order is control; D1 and R2 are mechanism. This framing helps decide where new functionality belongs and prevents mixing concerns (a gateway that also contains business logic, or a storage abstraction that also makes policy decisions). Unsure where code belongs? Ask which layer it serves.

Service bindings​

Service bindings can keep calls between Workers on the same server thread, avoiding a public HTTP round trip. Placement and downstream state still determine the complete path. Split services where ownership, authority or independent deployment earns the interface; a cheap call does not make that interface free to maintain.

Service bindings aren't just faster than HTTP; they're a different kind of call. No public internet traversal. No DNS resolution, no connection pooling, no TLS negotiation. The cost is coupling: bound services must be Cloudflare Workers, sharing a deployment context in the sense that a binding targets a specific service name.

Configure bindings in wrangler.jsonc and use them like local services:

Calling a bound service and checking its response
const response = await env.USER_SERVICE.fetch(
new Request(`http://internal/users/${encodeURIComponent(userId)}`)
);
if (!response.ok) throw new Error(`User service returned ${response.status}`);
const user = await response.json();

The http://internal URL is convention; the request never hits the internet. The binding routes directly to the target Worker.

External APIs​

Calls to external APIs traverse the public internet and need different handling. Use aggressive timeouts; a Worker waiting on a slow external API still consumes resources. Have fallback strategies for when external services are slow or unavailable.

Circuit breakers add value when: external API has frequent partial outages, timeouts are expensive, you have a meaningful fallback (cached data, degraded functionality). If you can't do anything useful when the circuit is open, a circuit breaker just changes when you return an error, not whether you return one. Consider whether the complexity is worth it.

Composition patterns that combine​

Patterns rarely exist in isolation. Common combinations: gateway plus rate limiting (almost every production API); BFF plus caching (BFF knows what data its client needs and can cache accordingly); event-driven plus saga (operations spanning services needing reliable completion); collaboration plus events (persist changes and maintain audit trail alongside real-time updates).

When combining patterns, watch for complexity accumulation. Each pattern solves a problem but adds operational surface. A system with five patterns is harder to debug than one with two. Add patterns when they solve problems you actually have, not problems you might theoretically encounter.

Anti-patterns to avoid​

Patterns have failure modes. These anti-patterns show what happens when patterns are applied without judgement.

The everything gateway​

A gateway containing business logic becomes a monolith at the edge. Gateways should handle cross-cutting concerns: authentication, rate limiting, routing. When business logic creeps in (validation rules, data transformation, conditional workflows), the gateway becomes a development bottleneck. Every feature change requires gateway deployment. Teams can't work independently.

The test: could you replace this gateway with a different implementation handling the same cross-cutting concerns? If business logic is embedded, you can't.

Premature event-driven​

Using queues for operations that should be synchronous adds complexity without benefit. Not every operation needs decoupling. Caller needs to know whether operation succeeded? Operation is fast and reliable? No benefit to processing later? Use a synchronous call.

The smell: a queue consumer immediately processing every message with no batching benefit, no retry benefit, no decoupling benefit. You've added a queue for architectural purity rather than solving a real problem.

Durable Objects for everything​

Durable Objects are powerful but not the only tool. Using them when KV or D1 would suffice means paying for coordination you don't need. KV handles configuration and caching. D1 handles relational data. Durable Objects handle coordination.

Choose from the invariant and access path. KV suits configuration whose freshness policy permits cached reads. D1 can enforce atomic updates and constraints within a relational partition. Choose a Durable Object when the application benefits from an active owner coordinating requests, connections and state.

Saga without compensation​

Implementing a saga that reserves resources without defining how to release them is a recipe for leaked state. Every saga step that acquires something (reserves inventory, charges a card, allocates capacity) needs a corresponding compensation step that releases it.

danger

Design compensations before acquisitions. If you can't figure out how to undo a step, reconsider whether it belongs in a saga.

Architecture decision records​

Patterns represent judgement crystallised into reusable form. Architecture Decision Records capture the judgement behind specific choices: why you chose this pattern over that one, what trade-offs you accepted, what would make you reconsider.

ADRs matter because architectural decisions outlive the people who made them. New team member asks why you have separate databases per tenant? The ADR explains. Circumstances change and you wonder whether to revisit a decision? The ADR provides context.

ADR structure​

A useful ADR answers five questions: What context prompted this decision? What alternatives did we consider? What did we choose and why? What trade-offs did we accept? What would make us revisit this?

Keep ADRs concise. A page is usually enough. Capture judgement, not documentation.

ADR: multi-database vs single database for tenants​

Context: Multi-tenant SaaS application. Tenant data must be isolated; regulatory requirements prohibit any data leakage between tenants. We need to decide how to structure D1 databases.

Options considered:

Single database with tenant_id column: every table includes tenant identifier, every query filters by tenant. Application logic enforces isolation. Simpler operationally: one database to manage, cross-tenant queries possible for analytics. But query mistakes can leak data, noisy neighbours affect everyone, 10 GB limit applies to all tenants combined.

Separate database per tenant: each tenant gets a D1 database, selected from an authenticated tenant identity. A query in one database cannot join another tenant's rows, but incorrect routing can still select the wrong database. The 10 GB limit applies per database; provisioning, migrations and cross-tenant analytics require automation.

External PostgreSQL with row-level security: Hyperdrive connects Workers to PostgreSQL, where tested policies can enforce tenant access centrally. This retains relational tooling and larger database capacity, with an external dependency and a placement decision for latency. Policy design and application roles need deliberate testing.

Decision: Separate D1 databases per tenant.

Rationale: The application requires a separately managed database boundary per tenant, and its team can operate the provisioning and migration tooling. Separate databases reduce the consequences of a missing row filter. Tenant routing remains security-critical and must be tested independently.

The 10 GB per-database limit fits the expected tenant sizes: the largest are forecast at 5 GB, most below 1 GB. Forecast aggregate account storage as well, and reconsider the partitioning before either limit is approached.

Schema migration complexity manageable through automation. Build tooling to apply migrations across all tenant databases with rollback capability.

Consequences accepted: No cross-tenant analytics without aggregation layer. Tenant onboarding requires database provisioning. Schema changes require careful rollout coordination.

Reconsideration triggers: Tenant needs more than 10 GB: partition their data or move to external PostgreSQL. Cross-tenant analytics becomes critical: need aggregation strategy.

ADR: Durable Objects vs external Redis for rate limiting​

Context: The API requires one consistent quota per customer key. All requests sharing a key must consult the same admission authority. The team already operates Workers but no shared Redis service.

Options considered:

Durable Objects: one object per quota key makes the decision atomically and provides automatic routing. Distant callers still pay a round trip to its location. The application owns the counter algorithm and lifecycle.

External Redis: an atomic command or script can enforce the same decision at the Redis authority. This adds a dependency but may reuse existing operational expertise and data structures. Its location and replication mode determine latency and consistency.

Decision: Use Durable Objects because the platform already operates the surrounding application and the measured object round trip meets its budget. This avoids introducing another service solely for counters.

Consequences accepted: One heavily shared key can become a hotspot. The counter code is platform-specific, and global clients may be far from its object. Sliding windows and leaky buckets are possible in either design; neither removes the coordination cost.

Reconsideration triggers: The workload develops globally shared hot keys, measured latency exceeds the budget, or an existing Redis deployment becomes a simpler operational home for the counters.

ADR: Queues vs Workflows for background processing​

Context: Uploaded files pass through validation, transformation, storage and notification. Each file is independent, processing takes 10–30 seconds, and repeating the whole operation is safe with a stable operation identifier.

Decision: Use Queues. A consumer processes each file and retries failures; after retries are exhausted, a dead-letter queue retains the work for investigation and replay within its configured retention window. Preserve and reconcile accepted intent separately when the promised recovery window is longer, as Chapter 23 explains. Aggregate outcomes meet the initial reporting requirement.

Consequences accepted: The application must make output writes and notifications idempotent. Queues does not persist a separate checkpoint after each processing stage or supply a per-file progress view; build those only if needed. Processing order is not guaranteed.

Reconsideration triggers: Adopt Workflows when redoing completed stages becomes expensive, independent step retries or waits matter, or compensation and persisted progress become part of the product. Mere dependence between two sequential operations does not require a workflow engine.

What comes next​

Chapter 25 applies these patterns to systems serving many customers. Routing, database partitioning and workflow orchestration must also enforce tenant identity, limit one tenant's impact on others, and produce usage records you can defend.