Skip to main content

Chapter 23: Designing for Failure and Recovery

How do I keep the application useful during a partial failure, and restore correct service afterwards?


A document has been released to a customer. The external delivery service has accepted it, the Workflow has recorded success, and D1 contains the receipt. Then a migration corrupts the database. Time Travel restores D1 to ten minutes before the release.

The database is healthy again. The application believes the document still needs sending.

Every primitive can fulfil its contract while the application reaches the wrong conclusion. Database recovery restored database history; it did not recall the document or rewind the Workflow. The recovery problem is deciding which facts remain true when their local records disagree.

The earlier chapters explain persistence, retries and deployment boundaries separately. This chapter connects them through a document-processing service: users upload documents, extraction prepares them for review, a person approves a particular version, and the service releases that version to an external recipient. The design is illustrative, not a claim about a measured production system. Its failures apply equally to provisioning, order fulfilment and agent-triggered actions.

Decide what the service promises to preserve​

An application rarely has one useful availability state. During an extraction outage, this service can display existing documents and accept new uploads. During a database outage, it may be unable to accept responsibility for new work even while the upload endpoint remains reachable. During a suspected approval bug, it should stop releases while investigation continues.

Give these distinctions a product meaning. “Uploaded” means bytes reached storage. “Accepted” means the application durably recorded the request and can find it again. “Ready for review” means extraction completed. “Released” means the receiving system confirmed the agreed handoff. None means the human recipient has read the document.

The receipt returned to the user should identify that promise: an operation ID, the accepted document version and a status endpoint. A timeout leaves the caller uncertain, so a repeat submission with the same operation ID must recover the existing result or safely continue it. Reject reuse of that ID for a different document or recipient. Silently treating different requests as duplicates hides a client defect.

Recovery objectives follow the same boundaries. The recovery time objective is how long an interruption may last; the recovery point objective is how much acknowledged state may be lost. Set them for a user journey, including reconciliation and backlog drain. A database restored in five minutes does not establish a five-minute recovery time if releases remain unsafe for two hours.

For this service, preserving an accepted document matters more than preserving an extracted preview. The preview can be regenerated from retained source bytes. An approval cannot be reconstructed by asking a model what the reviewer probably intended. Classifying those facts determines where durability investment belongs.

Choose the failure scope before the mechanism​

Durable acceptance needs a stated failure scope. A process crash, an accidental database rewind and loss of the account are different events. An outbox protects a handoff after a crash; it is rewound with its database. Choose the promise before adding another copy of state.

Required promiseDesign consequence
Resume after a process failure while authoritative storage survivesDurable business intent, repeatable dispatch and reconciliation may suffice
Recover from corruption with a declared possible loss windowKeep independently recoverable exports and business evidence; measure the gap from the last usable recovery point
Preserve every acknowledged decision across a database rewindRetain independent, ordered decision evidence before acknowledgement and resulting effects; its unavailability blocks that protected path
Recover from account compromise or loss of accessPut evidence and restoration capability across the relevant administrative boundary

These choices can differ within one service. A regenerated preview may tolerate loss; an acknowledged cancellation may not. A scheduled export does not establish its promised recovery point until the export is usable and restoration has been tested. The independent journal described later is for the stronger rewind promise, with an availability and operating cost to match.

Give each fact an authority​

A recovery design starts with a small map of facts, rather than a diagram of every product in the account.

FactAuthority in this designRecovery treatment
Exact uploaded bytesR2 object at an application-assigned immutable version keyPreserve; verify against the recorded digest
Accepted request, approval, cancellationTenant D1 databasePreserve or reconstruct from retained business evidence
Progress through extraction and reviewWorkflow checkpointsResume only after checking against current business state
Work awaiting a consumerQueueRedispatch from durable intent if delivery history is insufficient
External release acceptedReceiver's receipt for the stable release IDReconcile before repeating the effect
Search results, previews and dashboard countsDerived from authoritative recordsRebuild or invalidate

“Immutable” here is an application storage rule: a new document version gets a new key, and ordinary writers cannot overwrite accepted versions. It is not an assertion that naming an R2 object makes it undeletable. Configure retention and deletion authority to preserve the bytes for the agreed recovery period.

D1 owns the business decision because approval, document version and release intent fit in one relational partition. A Workflow owns the sequence and waits. Keeping a business record outside the Workflow is justified when the record must survive instance retention, support querying or establish the outcome of an external action. Avoid copying every step into another state machine merely to reproduce the Workflow dashboard.

There is no requirement to add a Durable Object to this design. Use one when a document needs an active owner for live collaboration or ordered decisions. Its private storage can then own those facts instead of D1. Giving both stores independent authority over approval would create a new reconciliation problem without solving an existing one.

Keep accepted work findable after a crash​

Persist the R2 object first, then commit the accepted operation and an outbox entry together in D1. The outbox describes work that must be dispatched. A dispatcher sends that work to a queue or starts a Workflow, then records the handoff. A periodic scan recovers entries left pending after a crash. The D1 batch is the local transaction; the R2 write and queue send remain outside it.

Protect the accepted bytes​

This ordering chooses an unreferenced object over an accepted request whose bytes never arrived. To make that choice safe, create an upload lifecycle record with separate staging and final keys. A direct-upload credential may write only to staging. The server finalises and verifies bytes under a key that credential cannot modify, then accepts that final version and its digest in the D1 transaction. Verify the final object itself: staging may change between an earlier check and the copy.

R2 presigned URLs can be reused until expiry. A unique key alone therefore does not protect an accepted version. Ordinary writers must also lack authority to replace or delete final bytes during their required retention.

R2 Bucket Locks can enforce that retention on the final-object prefix, taking precedence over lifecycle expiry. Keep staging cleanup separate from final-object retention. Chapter 14 covers both mechanisms, including their administrative limit: someone authorised to change bucket configuration can remove a lock rule.

Acceptance and abandonment compete through the same durable upload lifecycle. Cleanup first claims an eligible upload as abandoned, preventing later acceptance. Track candidate keys for each finalisation attempt so a delayed write can be found and removed without touching an accepted version. A grace period followed by a separate reference check leaves a race: acceptance can commit between that check and deletion.

Resume dispatch without creating new work​

If the dispatcher sends the work but crashes before recording the handoff, it may send the same outbox entry again. That is expected. Cloudflare Queues provides at-least-once delivery, so the consumer already needs a stable operation identity.

When the dispatcher creates a Workflow, retain the relationship between the business operation and instance ID. A lost create response must lead to lookup and reconciliation, not a new randomly named process. Workflows rejects creation with an ID belonging to a retained instance; that protection ends with instance retention. The business operation record must outlive it if older requests can be replayed.

A successful handoff does not mean the job finished. Reconcile accepted unfinished operations against execution state and receiver receipts as well as scanning outbox entries awaiting dispatch. A published message can expire before processing; redispatch only when the surviving evidence makes it safe, using the same business identity.

For a small system, a scheduled outbox dispatcher can start Workflows directly. Add a Queue when buffering, independent consumer scaling or downstream capacity warrants it. Reliability comes from the accepted intent and repeatable handoff, not from maximising the number of durable services on the path.

Preserve identity across retries and repairs​

The service needs different identities for the user's intention and its execution attempts. One release might be attempted by a Workflow, retried after a timeout, and later investigated by a repair job. They are all trying to perform the same release.

Use a stable release ID tied to the tenant, document version, recipient and approved action. An execution ID distinguishes a particular Workflow or repair run; an attempt number distinguishes calls within it. Neither should replace the release ID at the receiver. Otherwise, starting a replacement Workflow defeats the duplicate protection precisely when it matters most.

Before delivery, make cancellation and release commitment compete at the business authority. In this design, a conditional D1 transition checks the current approval and exact version, then changes the operation from approved to committed for release. Cancellation can instead change it to cancelled. Only one transition wins, and the dispatcher sends only committed releases. If the stronger journal promise applies, its evidence boundary must also be satisfied before sending.

Cancellation and release compete for one decision

Loading diagram…

Once release commitment wins, the product reports cancellation as a request whose outcome needs reconciliation or corrective action. It cannot promise to stop a request that may already be travelling to the receiver. The local transition establishes permission to attempt the effect; the receiver's contract establishes whether that attempt happened once.

The dangerous gap is between an external effect and recording its result locally. Suppose the receiver accepts the document, but the response is lost. A local “not sent” row proves only that the sender lacks confirmation. Checking that row again before retrying changes nothing.

There are three honest ways to handle the gap:

  • Use a receiver that atomically associates the release ID with the effect and returns the same outcome on repetition.
  • Query a receiver with a sufficiently strong operation-status contract, and reconcile uncertain outcomes before another attempt.
  • Accept the remaining ambiguity, stop automatic retries when necessary, and provide a human recovery path appropriate to the consequence.

A status endpoint whose “not found” response can race with an in-flight request is insufficient evidence to resend. Establish how the receiver reports pending operations, how long its duplicate protection lasts, and whether it rejects reuse of a key with different content. A client timeout does not cancel the receiver's work.

For the example, assume the delivery service supports a stable release ID, durable status lookup and duplicate suppression for an agreed window. If the actual provider lacks that contract, the architecture must expose an “outcome unknown” state. It cannot promise one release merely because the sender uses Workflows.

The same distinction applies to compensation. Revoking an accessible link may prevent future downloads; it cannot erase a copy already retrieved. Record a corrective action under its own identity and link it to the original release. An incident record should explain what happened, including any exposure that cannot be reversed.

Make the replay window explicit​

Suppose operators may replay a request up to fourteen days after acceptance. The source bytes, operation record and duplicate protection needed for that replay must remain usable for at least that period. If the receiver forgets release IDs after one day, a fourteen-day automatic replay policy is unsafe even though the application still has its records.

This example is a design requirement, not a default Cloudflare retention policy. Queue retention is configurable up to fourteen days on Paid, and fixed at twenty-four hours on Free; messages expire even while delivery is paused. A dead-letter queue is another finite buffer.

Choose the recovery window from business consequences, then make each dependency support it or shorten the permitted action. Sometimes the right decision is to retain document evidence for months but allow automatic resending only while receiver idempotency remains valid. Older cases go through reconciliation or a new, explicitly approved operation.

Control admission before recovery becomes overload​

A queue makes a temporary outage easier to tolerate by letting work wait. It also makes it easy to keep promising service after the application has lost the capacity to deliver it.

Consider an illustrative extraction service normally receiving 40 documents per second. Its downstream processor can sustainably complete 100 per second for the expected document mix. A thirty-minute outage admits 72,000 waiting documents. After recovery, 40 new documents per second leave only 60 per second for draining that backlog: another twenty minutes under those assumptions.

The useful estimate is:

Backlog recovery with continuing arrivals
drain time = waiting work / (sustainable completion rate - new arrival rate)

At 90 new documents per second, the same backlog needs two hours. At 100, it never drains. These calculations ignore variable document size and retries; production planning should measure work in pages, processing seconds or another unit that reflects the constrained resource. Queue depth alone can conceal a backlog dominated by unusually expensive files.

Admission control should protect a promise the application can keep. Bound outstanding work per tenant, cap total waiting work and stop accepting when predicted delay exceeds the product's tolerance or retention headroom. A useful response explains that processing is unavailable and whether the user should retry with the same operation ID. Persisting unlimited “accepted” requests merely moves the outage into tomorrow.

For this service, preserve retrieval and review of already prepared documents while restricting new extraction. Keep release delivery on a separate budget so a document-upload surge cannot consume all capacity for previously approved releases. Partition queues where priorities or downstream failure domains differ, rather than creating one queue per endpoint by habit.

Spend retries deliberately​

Retries consume the dependency that is already struggling. If a client, a service wrapper and a Workflow each make three total attempts, one operation can produce up to 27 downstream attempts. Backoff spaces them out; it does not remove them.

Choose one layer to own retries for each external operation, give it a bounded deadline and include the attempts in the capacity budget. Other layers should preserve the operation identity and report the failure. Add jitter to avoid synchronised retries, and honour a dependency's retry guidance when it fits the overall deadline. The operational reasons for bounded retries and jitter are developed in the Amazon Builders' Library.

A timeout is an observation about waiting, not a verdict on the remote outcome. For extraction, repeating computation may be acceptable if outputs are written under distinct attempt keys and only one result is accepted. For release delivery, the same timeout requires the receiver's idempotency or reconciliation contract.

Consumer concurrency is only an indirect rate control. Ten concurrent calls can become a much higher request rate when latency drops. Where a downstream contract is expressed in requests per second, enforce that budget explicitly across the callers sharing it; an in-memory counter in each global Worker is not a shared budget. A Durable Object per quota domain can own admission, with its latency and availability added to that path.

During an outage, stop futile processing without discarding the accepted intent. Pause queue delivery, suspend relevant Workflows and retain enough information to resume safely. A Workflow may be finishing current work before reaching its paused state. Neither pausing transport nor requesting a pause recalls an external request already sent.

Restore the application in a deliberate order​

Return to the failed migration. D1 is corrupt; R2 still has the uploaded bytes; some Workflows have advanced; a delivery service has accepted releases; queued messages describe an older view of the world. Restoring every component to the same clock time would not create a shared transaction boundary, even if every service offered that control.

The same release ID connects the restored database to the receiver's surviving record. Under the receiver contract described earlier, recovery can look up the outcome without sending the document again. This shows the receipt repair; the controls that stop other work during recovery follow below:

A database rewind does not rewind delivery

Loading diagram…

Prefer a bounded repair when you can identify the damaged records and derive their correct values from trustworthy evidence. It preserves legitimate writes that a whole-database restore would discard. Use restoration when the corruption is too broad or its extent is uncertain, with an explicit plan for accepted work after the restore point.

D1 Time Travel restores in place and cancels in-flight queries. SQLite-backed Durable Objects restore their own storage through a bookmark applied at restart. Neither mechanism restores another resource with it. Keep the pre-restore recovery bookmark and a separate record of the intervention.

Stop new effects before changing their history​

First put the affected tenant into recovery mode. Block new mutations and external releases, stop outbox dispatch, pause relevant queue delivery and request Workflow pauses. Identify the in-flight operations and wait for, cancel or reconcile them according to the receiving service's contract. “Pause requested” is not a safe point for restoration.

The control must cover every writer: HTTP handlers, consumers, scheduled dispatch, Workflows and administrative repair tools. Keep the recovery flag outside the database being rewound, or a restore can erase the instruction to remain stopped. A fail-closed path for consequential actions is preferable to assuming a missing flag means permission to continue.

A KV flag alone cannot provide an immediate stop: a writer may still read a cached permission to continue. Chapter 15 explains this freshness trade-off. Consequential actions need an authoritative control path, followed by the checks below to establish that old work can no longer write or release.

A recovery generation is a number that distinguishes current work from work started before the intervention. Increment it when beginning recovery and attach it to permitted execution. Install the new generation in the restored database while the external recovery control remains closed.

Suppose recovery moves the generation from 7 to 8. A delayed writer carrying 7 must be rejected, even if its code is otherwise valid. Compare the generation in the same transaction as the local write so recovery cannot slip between the check and the change. This is fencing: the storage boundary refuses work from the old generation.

External effects need a separate stop boundary. A check before sending still leaves a race; stop the controlled release dispatcher and account for its in-flight calls before rewinding state. Use the receiver's stable operation identity when resolving them. On resumption, an earlier Workflow step that read “approved” remains historical evidence; it cannot replace the current checks required by the release transition.

D1 Time Travel restores a database, so row-level tenant isolation does not provide a tenant-specific restore. The database-per-tenant pattern in Chapter 25 can keep that rewind within one customer’s data. A shared queue or release credential may still force a wider pause unless the application supports tenant-specific controls. Choose the data partition and the scope of recovery together.

When acknowledged decisions must survive a rewind​

Decide before an incident how to recover accepted requests made after the chosen restore point. For the stronger rewind promise, keep independent acceptance evidence. One design commits the D1 operation and a journal obligation together, then copies that evidence into a separate recovery journal before returning acceptance. Retries use the same operation ID and finish the incomplete handoff. The journal supports reconstruction; D1 remains the normal authority for current business state.

The journal must preserve business order even when writes arrive out of order:

An approval followed by cancellation
Business history: revision 41 approves; revision 42 cancels.
Journal arrival: 42 arrives; 41 is delayed.
Recovery: the gap keeps this history incomplete and release blocked.
After 41 arrives: reconstruct 41 then 42; the result is cancelled.

Each journal entry must answer two questions: what was decided, and where does that decision belong in the history?

For the decision, retain the operation, document version, requested action and the data needed to reconstruct it, or a reference to retained immutable data. Include a digest to verify the recovered data. The digest alone cannot reconstruct missing contents.

For the order, assign an entity revision with the business mutation and record its predecessor and history identity. The predecessor links the decision to the one before it. The history identity distinguishes revisions made before a restore from those made afterwards; start a new history identity on restoration so reused revision numbers cannot collide. Duplicate entries must agree on their content. Missing predecessors or conflicting successors keep the affected history incomplete.

This design waits for a continuous journalled history through the decision before acknowledging it or allowing the resulting release. Otherwise a dispatcher could send a document from a committed D1 row before its recovery evidence exists. Track that durable boundary explicitly and keep incomplete decisions pending. Apply the same rule to approvals, cancellations and revocations; preserving submission alone cannot recover the latest decision. Journal records whose responses were lost remain committed evidence even if the caller never saw success.

The stronger promise adds a write, ordering and a dependency to acceptance. An unavailable journal blocks the protected path, and its administrative boundary must match the failure being protected against. An asynchronous export offers a different trade-off: simpler acceptance with a possible loss window. State that window honestly. Uploaded bytes alone cannot establish that a user submitted them, approved release or subsequently cancelled it.

Reconcile surviving evidence before reopening​

Capture evidence before overwriting state: the current export where available, restore bookmarks, affected deployment identifiers, retained Workflow progress, outbox contents and receiver receipts. Store the recovery record separately from the data being restored. If the evidence needed to distinguish a completed release is already missing, preserve that uncertainty; do not convert it into “not released”.

Workers Logs and traces can help reconstruct attempts using the operation IDs discussed in Chapter 21. Their sampling and retention settings matter: they do not automatically provide the complete acceptance journal described above. An absent log entry cannot establish that no release occurred.

After restoration, keep effects disabled and reconcile each affected operation by identity:

Evidence after restorationAction before normal processing resumes
Receiver confirms the same release ID and document digestRecord its receipt; suppress another release
Receiver cannot establish whether a release happenedHold the operation for resolution
Accepted upload is absent from restored D1, but surviving evidence proves acceptance and identifies retained bytesReconstruct the business record and its intended work
Workflow expects an approval that restoration removedRecover the recorded approval or obtain a new one before release
Queue message refers to a cancelled or superseded versionPreserve the cancellation; do not execute the stale command
Search or preview refers to a document version absent from authoritative stateRemove or rebuild that projection

A receipt proves the receiver accepted something. It does not prove that the action was properly authorised. If the incident involved approval logic, reconcile authorisation separately and record wrongful releases as incidents even when technical delivery succeeded.

Restoration can also resurrect permissions and deleted records. Reapply subsequent revocations and deletion decisions before reopening access. Recovery evidence must preserve enough to recognise these changes without retaining document contents beyond the application's deletion policy. Where that evidence is unavailable, keep the affected scope closed until its permissions can be re-established.

Reconciliation changes should be repeatable and conditional on the expected state. Give each repair a reason, input evidence and outcome; rerunning the repair must not create a new release. Compare sets of operation IDs and their outcomes. Equal row counts can conceal one missing operation and one duplicate.

Reopen in stages, with a completion test​

Restore read access only when document references and current permissions agree. Then resume internal processing, followed by external releases, using a small, verified set of operations before increasing the rate. Do not let a repaired consumer's automatic scaling turn the entire backlog into a fresh outage.

Track oldest accepted unfinished work, unresolved external outcomes and sustainable completion rate. A shrinking queue is insufficient evidence: messages can disappear through expiry or acknowledgement without a successful business result. Each accepted operation should have a completed, cancelled, rejected or explicitly unresolved outcome, with someone responsible for unresolved cases.

The incident is operationally recovered when the agreed service is safe and usable. Record any remaining reconciliation separately, with an owner and a deadline. Declaring everything complete because the homepage loads invites the next on-call engineer to discover the unfinished work accidentally.

Keep old work compatible with new code​

A release contains more than a Worker bundle. Existing rows, queue messages, Workflow checkpoints, service interfaces and stored document formats remain after code changes. Compatibility lasts as long as the oldest work that may still execute or be replayed.

Suppose a new release separates approved into approved_for_internal_review and approved_for_external_release. Old consumers must not interpret either as permission to send. Add explicit release intent, migrate known records conservatively, and keep unknown states from gaining authority through a default branch. An additive schema change can still be a breaking change in meaning.

Use an expand-and-contract sequence: deploy readers that accept the transition, introduce the new writes, migrate existing records, drain or translate old work, then remove the old representation. Keep the intermediate states testable. The Workers version rollback covered in Chapter 22 restores earlier code; it does not rewind D1 data or undo a completed external release. Routing traffic back is safe only while that code can handle the data and messages produced since its deployment.

Durable Object gradual deployment assigns a version per object rather than independently per request, and callers and objects can still use different versions. Service bindings also join independently deployed code. Test those mixed interfaces.

For breaking Workflow changes, retain an explicitly versioned implementation and its dependencies for existing instances, directing new work to the replacement. Chapter 8 covers that boundary. Before promoting a release, pause an old instance after a consequential step, deploy the change, resume it and inspect actual effects. A successful test of a freshly created instance answers a different question.

Reprocessing a document with a different extraction model is also a versioned change. Preserve the source version, processor version and output identity. If the interpretation changes, the prior approval should not silently attach to the replacement output. Recovery should restore the approved decision, not opportunistically upgrade it.

Rehearse the failures your design claims to survive​

A recovery drill should start with business facts that can be independently checked. For the document service, create several accepted documents, approve one version, cancel another and arrange for the delivery receiver to accept a release while its response is lost. Introduce a known metadata error, restore a test database and reconcile the resulting disagreement.

Use an isolated receiving service that records every call, including duplicates, so a final “released” row cannot hide two external effects. Interrupt at the handoffs: after R2 persistence, after the acceptance transaction, after queue publication, after receiver acceptance and before the local receipt. Test a stale queued command and an old Workflow against the recovered state. Delay an approval journal entry until after its cancellation arrives; recovery must detect the gap and preserve decision order.

Define the cancellation cutoff in the drill. Cancellation accepted before the controlled release commitment must prevent dispatch; once a request may have reached the receiver, recovery must establish its outcome and any corrective action. The evidence should show that every accepted operation remains accountable, stale authority cannot start another release, and confirmed releases are not repeated. Measure the time until safe service and the time until the backlog clears.

Include the person following the procedure, permissions required to run it and the location of the recovery evidence. A drill that relies on its author's undocumented knowledge has found a defect.

Also test the recovery path when the management API is unavailable. An application-level stop can remain useful when deploying a fix or changing queue settings is impossible, provided its own control path is available. It should be installed and exercised before an incident. A runbook whose first step is “deploy the emergency control” has an undeclared dependency.

Independent backups serve a different failure scope from managed point-in-time recovery. An export in the same account may help with an accidental database change while remaining exposed to the same account compromise or loss of access. Choose a separate administrative boundary or provider when that is the requirement, then test restoring code, configuration, credentials and data together. Chapter 13 covers D1 export; Chapter 27 covers transferring authoritative writers.

A full second-cloud application is justified only when the outage scope and recovery target require it. It introduces another write-ownership and reconciliation problem. Many services gain more from a smaller set of explicit guarantees: durable acceptance, limited admission, a safe stop, retained evidence and a rehearsed way to bring the facts back into agreement.

What comes next​

Chapter 24 applies the runtime, data and recovery choices to complete application patterns. Carry the recovery boundary into those designs: identify who owns the decision, where accepted work survives and what prevents an old attempt from doing the wrong thing after the system changes.