Chapter 7: Durable Objects: Stateful Compute at the Edge
How do I build applications that need coordination, real-time updates, or consistent state?
Durable Objects brings entity identity, state and application logic together in a managed actor. The important distinction is how directly that model connects to Workers and live requests. Actor systems such as Orleans and Azure Durable Entities offer related ideas with different execution and integration trade-offs.
Durable Objects aren't servers you provision; they're database rows that run code. Think "one object per user" or "one object per document," not "how many objects do I need for capacity." This shift from vertical to horizontal thinking determines success with Durable Objects.
Durable Objects launched in 2020 to solve a problem Workers couldn't: coordinating state across distributed requests. Five years of production use across thousands of applications have proven the model. The SQLite storage backend, added in 2024, brought full relational capabilities to what was originally a key-value system.
A Durable Object is a Worker with a globally-unique identity and private, strongly-consistent storage. That simple description has profound implications: you can create an object representing a single user, document, game session, or chat room, and that object coordinates all operations on that entity from anywhere in the world with strong consistency guarantees, without you managing any infrastructure.
If you find yourself counting Durable Objects to minimise cost, you're approaching them wrong. Shift your thinking from "how do I minimise instances" to "one object per logical entity." This shift produces cleaner architecture and better performance.
The conceptual model
Understanding Durable Objects requires abandoning assumptions other platforms instil.
Single-threaded, globally-unique actors
Each Durable Object has a unique identity and a single-threaded execution context. Requests to the same object reach its active instance. Synchronous code cannot run simultaneously, but requests can interleave when a handler awaits external I/O. Storage gates protect specific state operations; they do not serialise every handler from start to finish.
This is the actor model at global scale. Carl Hewitt conceived the actor model at MIT in 1973, and it evolved into production systems like Erlang (powering much telecommunications), Microsoft Orleans (powering Xbox Live), and Akka (widely used in financial services). The core insight is simple: instead of shared memory protected by locks, each actor owns its state exclusively and communicates only through messages. What distinguishes Durable Objects is global uniqueness backed by automatic durable storage, capabilities other actor systems require significant infrastructure to achieve.
More precisely, Durable Objects are active objects: an entity owns state and the logic that changes it. That ownership makes an invariant easier to locate in the code. It still needs deliberate handling of external awaits, which can let another request enter before the first operation finishes.
The uniqueness is global. If you create a Durable Object named "user-123", every request targeting "user-123" reaches the same instance, whether originating in Tokyo, Toronto, or Tallinn. Routing is automatic. You address the object by name; the platform handles the rest.
Think database rows, not servers
Think of objects as small stateful entities, created on demand, rather than servers to provision. A chat application might use one per conversation; a game might use one per match. In-memory instances can be evicted, but stored data remains until the application deletes it. The domain defines the partitioning, and lifecycle policy defines how long the state should remain.
Coordination primitive first, storage second
Durable Objects combine state with an addressable owner that runs application logic. That can remove the need for a separate coordination service, provided the object's key matches the boundary over which the invariant must hold.
Consider rate limiting. One object per rate-limited entity gives each counter an owner. A synchronous increment can be decided locally; if the decision awaits an external service, recheck the state or protect the transition before treating it as atomic.
Consider document collaboration. Routing edits through one object gives the document an authoritative sequencer and a place to enforce its rules. Concurrent edits from disconnected clients may still need version checks, operational transformation or CRDTs; central coordination does not decide how conflicting user intent should merge.
Each object can store up to 10 GB in SQLite, but coordination guarantees often matter more than storage capacity.
When to use Durable Objects
Durable Objects aren't the only stateful option on Cloudflare. Knowing when to reach for them versus D1 or KV prevents over-engineering and unnecessary cost.
The decision framework
Ask what must be coordinated and who owns that decision. Use D1 when database transactions and relational queries express the invariant. Use KV when read-heavy key-value access can tolerate stale data. Use a Durable Object when an entity needs an active owner: ordering operations, holding live connections or coordinating state across requests.
Strong consistency alone does not require an actor. A shopping cart can use transactional database updates. A live auction with bid ordering, timers and connected participants may benefit from one object per auction. Choose the abstraction that makes the invariant visible with the least machinery.
The cost reality
Durable Objects bill requests, active duration and storage. The cost depends on how long objects remain active and how efficiently each object serves its entity. Hibernation can materially change a WebSocket application's bill; keeping an object active for occasional events can do the opposite.
Compare a representative entity lifecycle, including idle periods and reconnects, with the complete alternative. A price ratio without connection duration, message size and storage operations says little. Chapter 20 develops the full cost model.
If you're using Redis for coordination
Many readers use Redis for counters, locks or live state. The useful shift is from a shared data service to an entity that owns both state and logic. Durable Objects simplify local invariants, but code that awaits an external service still needs a deliberate concurrency strategy.
Redis requires you to manage connection pools, handle failover, and reason about consistency during network partitions. Durable Objects handle all this invisibly. But Redis offers portability across any cloud or on-premises environment, a massive ecosystem of client libraries and tooling, rich data structures like sorted sets and streams, and horizontal scaling through clustering without application-level sharding.
The trade-off is portability versus simplicity. If you might run on AWS, Azure, or on-premises, Redis keeps your options open. If you're committed to Cloudflare, Durable Objects eliminate operational complexity Redis requires. But there's no lift-and-shift path from Durable Objects to anything else.
Communicating with Durable Objects
Before examining storage and durability guarantees, understand how Workers talk to Durable Objects. The communication model shapes how you design your object's interface.
RPC: the modern approach
With compatibility date 2024-04-03 or later, define typed methods directly on your Durable Object class and call them from your Worker. This interface describes the application methods the object implements; the caller already has the target object ID:
interface ProfileMethods {
getProfile(): Promise<Profile | null>;
updateProfile(updates: Partial<Profile>): Promise<void>;
}
// In the Worker
const stub = env.USER_PROFILE.get(id);
const profile = await stub.getProfile();
await stub.updateProfile({ name: "New Name" });
RPC provides type safety, cleaner code, and better error messages than parsing requests manually in a fetch() handler.
Always await RPC calls. Unawaited calls create dangling promises that swallow errors silently. Your code appears to work while failing invisibly.
A Durable Object can read its identifier through this.ctx.id. Use that identifier for correlation; keep application names or tenant metadata separately when those meanings matter.
When fetch() still makes sense
The older pattern routes HTTP requests through a fetch() handler on the Durable Object. This remains appropriate in two scenarios.
First, when proxying HTTP requests where request/response semantics matter. If your Durable Object mediates access to an external API and clients expect standard HTTP responses with specific headers, status codes, and streaming bodies, fetch() preserves those semantics naturally. RPC would require reconstructing them.
Second, when migrating an existing codebase incrementally. A production system built on fetch() routing can migrate method by method. New functionality uses RPC; legacy paths continue through fetch() until converted.
For new code without these constraints, prefer RPC. Type safety catches errors at development time, and the calling code reads like what it is: method invocations on an object.
Storage backend
Each Durable Object contains a private SQLite database, accessed through the sql property on the storage API. You have full SQL capabilities: queries, joins, indexes, transactions, FTS5 full-text search, and JSON functions.
this.ctx.storage.sql.exec(
"INSERT INTO moves (player_id, move_data) VALUES (?, ?)",
playerId, JSON.stringify(moveData)
);
const leaders = this.ctx.storage.sql.exec(
"SELECT name, score FROM players ORDER BY score DESC LIMIT 10"
).toArray();
A legacy key-value API exists for compatibility but is backed by SQLite internally; for new code, use SQL directly, as it's more powerful and no slower. New Durable Object namespaces must use the SQLite backend, and accounts with no existing key-value namespaces cannot create one at all. The key-value backend is a compatibility path for existing deployments, not a choice for new designs.
Storage limits and their implications
Each Durable Object can store up to 10 GB in its SQLite database, with individual rows limited to 2 MB. These limits rarely constrain correct designs. If you're approaching 10 GB in a single object, reconsider your sharding strategy. The 2 MB row limit matters for document storage; chunk large documents or store them in R2 with references in SQLite.
Durable Object Storage Limits Reference
| Resource | Limit | Notes |
|---|---|---|
| SQLite database size | 10 GB | Per Durable Object |
| Row size | 2 MB | Use R2 for larger documents |
| SQLite key-value entry | 2 MB combined key and value | SQLite storage backend |
| Concurrent connections | 32,768 | WebSocket connections per object |
After roughly 10 seconds of inactivity, an eligible Durable Object hibernates and its in-memory state is discarded; the constructor runs again on the next request. Full eviction from the host follows later, after 70 to 140 seconds. Either way, anything not written to SQLite is gone. This isn't a bug; it's fundamental to the economic model, so make memory-versus-SQLite decisions explicit for every piece of state.
Memory, SQLite, and the eviction reality
Two timers govern an idle Durable Object, and conflating them causes confusion. After roughly 10 seconds of inactivity, an object that qualifies for hibernation hibernates: it is removed from memory, its in-memory state is discarded, and the constructor runs again on the next request. An object that cannot hibernate, because it holds a timer, an in-flight fetch(), a standard (non-hibernation) WebSocket, or an outbound connection, stays resident until it is evicted from the host entirely, which happens after 70 to 140 seconds of inactivity. The practical consequence is identical either way: state not persisted to SQLite is lost, and the constructor re-runs on the next request. The 10-second figure is simply the earliest point at which that can happen.
Use SQLite for state that must survive eviction. Use memory for reconstructible caches. Delaying persistence of a counter until a periodic flush is write-behind: a crash can lose increments since the last completed flush. Choose that loss window explicitly; do not describe it as durable merely because SQLite is used eventually.
The pattern that causes trouble is implicit reliance on memory. Storing state in class properties without consciously deciding whether it should survive eviction creates bugs that appear in production (where objects go idle between requests) but not in development (where objects stay warm during active testing).
Outbound connections qualify that idle rule. An active outbound TCP socket or WebSocket prevents eviction for up to 15 minutes per connection. After that, the connection stops preventing eviction; it is not automatically disconnected at the fifteen-minute mark. Plain fetch() response streams do not provide this protection. An upstream connection therefore cannot guarantee that an agent or subscription stays resident until its work finishes. Persist progress and design for reconnection.
Point-in-time recovery
SQLite-backed Durable Objects support point-in-time recovery, allowing restoration to any moment in the last 30 days. This matters most for disaster recovery and debugging: understanding what state an object held at a specific moment, or restoring an object corrupted by a bug. The mechanism uses bookmark strings you obtain for the current moment or any past timestamp, then schedule restoration on the object's next restart. Practise restoring an object with the code and dependencies that will use its recovered state. Restoration does not undo external effects, and pending work may refer to decisions made after the restore point. The recovery chapter in Part VI develops that application-level procedure.
How durability actually works: output gating
Output gating controls when outgoing communication may leave an object with pending writes. It lets application code rely on the runtime to hold a response behind the state it has just persisted.
With the default storage guarantees, outgoing communication waits for pending writes to become durable. The object must actually write the state: returning success for an in-memory mutation does not persist it. A lost response can still leave the caller uncertain about the outcome.
The mechanism
Storage work can proceed locally while durability is being established. Output gates hold outgoing communication behind pending writes: the caller should not observe success for state that has not become durable. If persistence fails, the runtime resets the object and fails the pending work rather than releasing a successful response.
Output gating holds responses until pending writes are durable
Loading diagram…
Why this matters for architects
Output gating prevents a successful response from getting ahead of the writes it depends on. It does not make external side effects transactional with storage, or remove ambiguity when a client loses the response. Retried business operations still need an idempotency strategy.
The architectural gain is that the runtime coordinates persistence with outgoing communication. You can write local state logic without manually delaying each response until replication finishes. Keep the distinction between a durable local decision and a completed operation in another system.
Measure the response, not only the local storage call. Durability and network transit remain on the externally observed path, especially when a request makes several sequential object calls.
Input gates: the boundary around state operations
Single-threaded JavaScript prevents simultaneous execution, but an external await can let another request run. Durable Object storage adds a stronger rule: while an asynchronous storage operation is pending, input gates normally prevent other events from entering the object.
That makes a read-modify-write sequence using storage safe when it does not yield to unrelated I/O. A fetch to a payment service between the read and write changes the situation: another request can change the state while the first waits. Separate the reservation from the external action, then validate the state when processing the result.
Use explicit concurrency control only where that invariant requires it. Holding every request behind a slow external call may preserve an invariant while creating a throughput bottleneck.
Write coalescing and its constraints
The runtime coalesces writes made without intervening awaits. That avoids a separate durable commit for every small update. Keep an explicit transaction boundary when failure partway through application logic must roll back all its writes:
const sql = this.ctx.storage.sql;
this.ctx.storage.transactionSync(() => {
sql.exec("UPDATE players SET name = ? WHERE id = ?", name, id);
sql.exec("UPDATE players SET score = ? WHERE id = ?", score, id);
sql.exec("INSERT INTO audit_log (player_id, action) VALUES (?, 'update')", id);
});
transactionSync() requires a synchronous callback. An external API call cannot join that transaction. Commit the local decision, perform the external action with an appropriate retry strategy, and reconcile its outcome.
Placement and routing
Durable Objects are placed geographically based on where they're first accessed. A simple algorithm that works well, with caveats worth understanding.
How placement works
On first use, Cloudflare chooses a location for the object, normally near the initiating request. Location hints can guide that decision, while jurisdictional restrictions express a stronger placement boundary. Request origin is an input to placement, not an exact data-centre guarantee.
Every caller then reaches the same object. An administrative job that first touches all tenant objects can therefore influence their placement. Consider who creates or first accesses them, and test the latency of the users who will actually use them.
The constructor can run again after hibernation or eviction. Reconstruct memory from storage and keep initialisation safe to repeat; restarting the in-memory instance is different from creating a new entity.
Geographic latency realities
A globally addressable object still has one active location. Users near it and users on another continent take different paths to the same state. That is the price of a single coordination point.
Measure the locations that matter to your application, including the first request after hibernation. A chat room shared across continents cannot place its authority next to every participant. Keep latency-sensitive work local where possible and send only the decisions that need shared coordination to the object.
Location hints and jurisdictional restrictions
When you know the optimal location before first access, provide a location hint when obtaining the object’s stub for the first time. Location hints are suggestions, not guarantees (Cloudflare may place objects elsewhere based on capacity), but they improve placement when pre-creating objects for users or placing objects near known backend infrastructure.
For compliance requirements, restrict Durable Objects to specific jurisdictions in your configuration. The object only runs in data centres within that jurisdiction and storage only persists there, but the object remains globally accessible. Requests from outside the jurisdiction route to it; processing happens within the jurisdiction. This can meet an object-specific residency requirement; it does not establish compliance for the application’s other data flows.
Three jurisdictions are available: eu, us, and fedramp. You select one when deriving the object's ID, and the object can read its own jurisdiction back through ctx.id.jurisdiction. The us jurisdiction pins both compute and storage inside the United States, which matters when a contract or regulator demands genuine US data residency rather than merely US-reachable infrastructure; fedramp restricts the object to FedRAMP-authorised facilities for the same reason in a higher-assurance context. The architectural point holds across all three: residency constrains where the object lives, not who can reach it, so a single global codebase can still serve a jurisdiction-pinned object to users anywhere.
Real-time capabilities: the economics of long-lived connections
Durable Objects excel at real-time features, but not because WebSocket handling is novel. The innovation is economic: hibernation makes long-lived connections viable at scale.
The cost problem with WebSockets
Traditional architectures charge for compute time while connections are open. A chat room with 100 idle users still consumes server resources, still incurs costs. This creates pressure to disconnect idle users, implement complex connection pooling, or accept high costs for real-time features.
Hibernation changes this calculus. When a Durable Object has no active work, just idle WebSocket connections, it hibernates. Connections remain open, but the object stops consuming compute resources. When a message arrives, the object wakes, processes the message, and can return to hibernation. Long-lived connections become a feature, not a cost centre.
When hibernation matters
The economics become significant above roughly 10,000 concurrent idle connections. Below that threshold, the cost difference is negligible, just a few dollars per month. Above it, hibernation can reduce costs by an order of magnitude.
If connections are short-lived (seconds rather than hours), hibernation rarely engages. If objects receive constant traffic with no idle periods, hibernation never triggers. The sweet spot: applications with many connections mostly idle but occasionally active. Collaborative documents where users have files open but aren't always editing, chat applications where conversations go quiet for hours, dashboards updating periodically rather than continuously.
What hibernation requires
Hibernation works when the Durable Object acts as a WebSocket server, accepting connections from clients. Outgoing WebSocket connections to external services cannot hibernate; the object must remain active to maintain them. If your Durable Object needs connections to external services, it cannot benefit from hibernation's cost savings.
Enable hibernation by using ctx.acceptWebSocket() rather than manually managing WebSocket pairs, and by implementing webSocketMessage() and webSocketClose() handlers rather than attaching event listeners.
Connection metadata survives hibernation through serializeAttachment(). Keep attachments small: user and session identifiers can locate richer state in SQLite when the object wakes. Attachments last only for the connection's lifetime, so information needed after reconnection belongs in durable storage.
Deployment resets connections
When an object changes to a different deployed version, its restart closes existing WebSocket connections. Design clients to reconnect and recover accepted state and pending operations. Chapter 24 explains why reopening the socket alone is insufficient.
A version change that restarts an object disconnects its WebSockets. Gradual deployment assigns versions per object, so a deployment need not restart every object at once. Client reconnection logic isn't optional; it's required for production systems. Test reconnection flows as thoroughly as initial connection flows.
The one pattern that matters
Every Durable Objects architecture reduces to one pattern: one object per logical entity that needs coordination. Rate limiting, counters, leader election, chat rooms, collaborative documents: not separate patterns but applications of the same principle.
Routing to the right object
Your Worker routes requests to objects based on entity identity:
const userId = getAuthenticatedUserId(request);
const id = env.USER.idFromName(userId);
const stub = env.USER.get(id);
return stub.fetch(request);
All operations on a user route to one object. All operations on a document route to one object. All operations on a game session route to one object. The pattern is always the same; only the entity changes.
This aligns the coordination boundary with the data boundary. You don't need distributed transactions because all related operations happen in one place. You don't need consensus protocols because one object makes all decisions for its entity.
Applied coordination: rate limiting and leader election
A rate limiter can assign one object to each limited resource and make the counter check and update one synchronous SQL operation or local transaction. An external await between the check and update would allow another request to interleave. The object supplies the owner; the mutation boundary enforces the limit.
A Durable Object can similarly own a lease, storing leader_id, expires_at and an increasing ownership token. Candidates acquire or renew it through an atomic local mutation. Lease expiry does not stop a delayed former leader from acting: the protected resource must reject stale ownership tokens, or consequential work must pass through the owner. A lease record without downstream enforcement cannot prevent two generations of leader from producing effects.
Control plane and data plane separation
For systems managing many resources (game sessions, chat rooms, tenant workspaces), separate coordination concerns from data concerns using the control plane / data plane pattern.
The control plane Durable Object handles administrative operations: creating resources, listing resources per user, tracking resource metadata. The data plane comprises individual Durable Objects for each resource, handling actual operations on that resource's state and logic.
The critical insight is that data plane operations bypass the control plane entirely. Once you know which game session to join, requests route directly to that session's Durable Object. The control plane isn't in the hot path. This prevents a single coordination point from becoming a bottleneck while maintaining the ability to manage and discover resources.
The sketch omits imports, schema creation and domain-method implementations. Callers supply an authenticated user ID.
// Control plane: manages resource registry
type ProjectRow = { id: string; user_id: string; name: string };
class WorkspaceRegistry extends DurableObject {
async createProject(userId: string, name: string): Promise<string> {
const projectId = crypto.randomUUID();
this.ctx.storage.sql.exec(
"INSERT INTO projects (id, user_id, name) VALUES (?, ?, ?)",
projectId, userId, name
);
return projectId;
}
async listProjects(userId: string): Promise<ProjectRow[]> {
return this.ctx.storage.sql.exec<ProjectRow>(
"SELECT id, user_id, name FROM projects WHERE user_id = ?", userId
).toArray();
}
}
// Data plane: handles actual project operations (separate DO per project)
class Project extends DurableObject {
async addDocument(doc: Document): Promise<void> { /* ... */ }
async updateSettings(settings: Settings): Promise<void> { /* ... */ }
}
The pattern scales horizontally. The control plane tracks thousands of resources; each operates independently. Single-threaded guarantees apply within each object, not across them.
When to split into multiple objects
Sometimes state that seems related should live in separate Durable Objects.
Split when entities have different access patterns. A user's profile (updated occasionally) and their real-time cursor position (updated constantly) probably belong in separate objects. The profile tolerates higher latency and benefits from simpler access patterns; the cursor position needs minimal latency and generates high write volume.
Split when entities have different consistency requirements. If some data needs strong consistency while related data tolerates eventual consistency, separating them lets you use the appropriate primitive for each.
Split when combining would create a throughput bottleneck. A single Durable Object handles roughly 1,000 requests per second, plenty for one user, one document, or one game session, but not for aggregate traffic. If you're routing all requests to one object, you've misunderstood the model.
A single Durable Object handles approximately 1,000 requests per second. This is plenty for per-user or per-entity coordination, but routing aggregate traffic to one object creates a bottleneck that limits scalability.
Combine when operations need atomicity across the data. If updating A and B must either both succeed or both fail, they belong in the same object. Splitting them means accepting eventual consistency or building your own coordination, defeating the purpose of using Durable Objects.
Throughput limits and what to do about them
The documented 1,000 requests per second figure is a soft limit; actual throughput depends on the work each request performs. One object has one execution thread, so expensive synchronous work limits throughput. Asynchronous I/O can overlap, subject to the concurrency controls protecting state. Measure the entity workload before deciding to split it.
First, verify you actually do. A single user won't generate 1,000 requests per second. A document with 100 concurrent editors, each making one edit per second, generates 100 requests per second, well within limits. Scenarios exceeding the limit are rare.
If you've verified the need, your options: client-side batching (aggregate multiple operations into single requests), read replicas via KV for read-heavy patterns (cache frequently-read data when staleness of 60 seconds or longer is acceptable), or sharding within the entity (split a high-traffic counter into 10 shard objects, aggregate reads). Each adds complexity. Often, hitting the throughput limit indicates a design that doesn't fit the model. Reconsider whether Durable Objects are right for that entity.
Hierarchical coordination
For complex systems, use parent-child relationships. A lobby object tracks active games (IDs, status, player counts) while individual game objects handle each game's state and logic. The parent coordinates and indexes; children hold the data. This scales horizontally: the lobby tracks thousands of games, each operates independently, and single-threaded guarantees apply within each object, not across them.
Facets: dynamically loading child Durable Objects
Hierarchical coordination assumes you know the child class at deploy time. Facets, in beta on the Workers Paid plan, lift that assumption. A parent Durable Object can call this.ctx.facets.get(name, () => loadCode()) to instantiate a child Durable Object whose code is loaded at runtime, typically out of a Dynamic Worker (Chapter 25). Each facet runs in its own isolate and gets its own SQLite database, separate from the parent's, but stored together as part of the same overall Durable Object. Requests to the facet flow through facet.fetch() or its RPC stub, exactly like any other DO.
Two consequences are worth understanding. First, you can run AI-generated application code with its own persistent database without provisioning a fresh DO namespace per app: the parent acts as a supervisor that loads the child on first request and routes traffic to it thereafter. The supervisor is the place to enforce per-tenant logging, observability, billing, and resource limits before the agent's code ever runs. Second, the parent's database remains under your control regardless of what the child does to its own SQLite, which gives platform builders a clean separation between platform-managed metadata and tenant-managed application state. For Workers for Platforms-style multi-tenancy, this is the missing piece that lets each tenant have its own long-lived stateful database without a one-namespace-per-tenant explosion. Chapter 25 returns to this in the multi-tenant context.
Lifecycle and scheduled work
Understanding when Durable Objects wake and sleep prevents subtle bugs and enables efficient designs.
Idle timeout and constructor semantics
A Durable Object with no pending operations enters an idle state. After roughly 10 seconds it hibernates if it qualifies, discarding in-memory state; otherwise it stays resident until full eviction at 70 to 140 seconds. The next request recreates the in-memory object either way, running the constructor again.
This means three things. First, don't rely on in-memory state persisting; always store important data to SQLite. Second, the constructor runs on every wake, so use it to initialise state, load from storage, and run migrations. Third, there is no destructor. You cannot run cleanup code on eviction because it might never happen, or might happen without warning.
Use blockConcurrencyWhile() in the constructor to ensure initialisation completes before handling requests. This blocks all requests until the promise resolves, appropriate for migrations and loading critical state, but a throughput bottleneck if overused.
Alarms for scheduled wake-up
Alarms wake a Durable Object at a scheduled time, even if no external requests arrive. Set an alarm, handle it in alarm(), and reschedule if work is recurring. Each object has one alarm slot; setting a new alarm overwrites the previous. Alarms provide at-least-once execution with up to six retries after failure. Exhaustion can leave the alarm unscheduled until it is set again, so consequential recurring work needs a reschedule or escalation path. Scheduled time is not a real-time deadline.
Use alarms when the schedule is per-object: each user's subscription renews on a different date, each document's cleanup runs 24 hours after last edit. Use Cron Triggers when the schedule is global: daily cleanup at midnight, hourly aggregation on the hour. Use Queues when work is triggered by events rather than time.
Testing Durable Objects
Test the invariant that motivated the object: concurrent reservations must not exceed capacity, a repeated command must not repeat its side effect, or a reconnect must recover the right session state. The Workers Vitest integration can exercise runtime bindings and storage locally, including eviction and reconstruction.
Add deployed tests for geographic routing, production limits and external dependencies. Keep an external await in concurrency tests when production code has one; omitting it can hide exactly the interleaving that matters. Chapter 5 sets out the testing layers.
Comparing to hyperscaler alternatives
Understanding what Durable Objects replace clarifies their value.
| Pattern | Cloudflare | AWS | Azure |
|---|---|---|---|
| Coordination with consistent state | Durable Objects | Step Functions + DynamoDB | Durable Functions |
| WebSocket state management | DO with hibernation | API Gateway + Lambda + DynamoDB | SignalR + Functions |
| Rate limiting | DO counter | ElastiCache or DynamoDB | Redis Cache |
| Actor model | Native | Custom with SQS + Lambda | Orleans (self-hosted) |
| Leader election | DO with lease | DynamoDB conditional writes | Blob leases |
Durable Objects can reduce the number of components needed for entity coordination. That is most useful when one object can own the invariant and the live connections. An existing actor platform or database transaction may already express the same requirement with less migration work.
Real-time presence: a concrete comparison
On AWS, building a real-time presence system (who's online, cursor positions, live updates) typically requires API Gateway WebSocket API for connection management, Lambda functions for connection and message handling, DynamoDB for storing connection state and presence data with TTL for cleanup, and often ElastiCache for pub/sub to broadcast updates across Lambda instances. You manage connection routing, handle eventual consistency between DynamoDB and Lambda state, and pay for idle connections through API Gateway. Four distinct services with their own pricing models, failure modes, and operational concerns.
On Cloudflare, the same system uses a Worker for initial routing and a Durable Object per presence scope. The Durable Object holds WebSocket connections directly, maintains presence state in SQLite, and broadcasts updates to connected clients. Hibernation means idle connections cost nothing. No external pub/sub needed because all connections for a scope route to the same object. No eventual consistency issues because all state lives in one place with single-threaded access.
Centralising a presence scope in one object makes its authoritative state easier to reason about. Clients can still disconnect, reconnect with stale views or lose messages. Define expiry and resynchronisation explicitly; fewer services does not remove those protocol boundaries.
The trade-off is lock-in. Moving away from Durable Objects means replacing their routing, lifecycle and storage integration, even if the actor model remains. If PostgreSQL already handles your coordination needs clearly, keep it. Choose Durable Objects when giving each entity its own execution and storage boundary simplifies the application enough to justify that dependency.
Failure modes worth naming
Certain failures recur in Durable Object systems. Naming them creates vocabulary for design discussions and post-incident analysis.
Placement latency mismatch occurs when a Durable Object is placed based on its first requester, but subsequent users are in distant regions. A game session object placed near its first request in London may later serve distant users. When players in Tokyo join later, every request crosses continents. Fix with location hints when you know optimal placement beforehand, or accept the latency. For objects serving globally distributed users equally, no placement is optimal. Consider whether the coordination requirement demands a single object.
Hibernation wake-up requires the constructor to reconstruct application state before handling the next event. Eligible WebSocket connections survive hibernation; the application recovers their metadata rather than asking every client to reconnect. Measure the first event with representative stored state and constructor work.
Alarm drift occurs when alarms fire late. Make handlers idempotent and enforce expiry on the authoritative access path when it matters; a delayed cleanup alarm must not extend a permission or reservation.
State eviction surprise catches developers relying on in-memory state. It manifests as intermittent bugs in production but not local testing, where objects rarely go idle long enough to hibernate or be evicted.
The god object anti-pattern
A team building a rate limiter creates one Durable Object for all users. It seems logical: centralised state, single source of truth, no coordination needed between objects. It works in staging.
As traffic grows, every user competes for the same object's execution time. Queueing raises latency and can lead to overload errors. The onset depends on the work per request, so measure saturation with the real handler rather than treating 1,000 requests per second as a guaranteed operating point.
The fix: embrace the model. One Durable Object per rate-limited entity. Each user's rate limiter lives in their own object. A million users means a million potential objects, but that's correct. Objects are database rows, not servers. Each handles only that user's traffic; the system scales horizontally without configuration.
Every Durable Objects anti-pattern shares this root cause: treating them like servers rather than database rows.
What comes next
A Durable Object gives an entity a place to coordinate. A business process also needs to remember which work has completed and where to resume after failure. Chapter 8 covers Workflows and the boundary between replaying a step and repeating its effects.