Skip to main content

Chapter 8: Workflows: Durable Execution

How do I build processes that must complete even if individual steps fail?


Most Workers follow a simple pattern: handle a request, return a response, complete in milliseconds or seconds, and retry on failure. But some processes don't fit this model at all. Order fulfilment spans payment capture, inventory reservation, shipping label generation, and email notification; document approval waits days for human review; data synchronisation processes millions of records over hours. These long-running, multi-step processes must survive infrastructure failures, code deployments, and network outages while retaining enough progress to resume or reach an explicit failure outcome.

Workflows persists progress and coordinates retries and waits. It removes orchestration infrastructure, while leaving the application responsible for terminal failures and uncertain external effects.

What "durable" actually means​

The term "durable" gets used loosely in distributed systems, but for Workflows it means something specific: after each step completes, its result persists to storage before the next step begins. On replay, retained completed steps return their saved results; an incomplete step can run again. Code outside those checkpoints can also run again, so it must not hide unprotected side effects.

Durable vs. Retry

This differs from retry logic. A retry re-executes the entire operation. Durable execution checkpoints progress and resumes from where it left off. A workflow that completed steps 1 through 7 before failing resumes at step 8, not step 1.

Durable execution has over a decade of production history across systems like Amazon SWF, Step Functions, Uber's Cadence, Temporal, and Azure Durable Functions. Cloudflare Workflows applies these proven patterns to edge computing with characteristic simplicity: you write TypeScript, each instance is isolated, and the platform handles replay, checkpointing, and step persistence.

The infrastructure beneath​

Each workflow instance runs on its own SQLite-backed Durable Object, and understanding this architecture isn't merely academic; it explains Workflows' behaviour and constraints in practical terms.

Workflows uses Durable Objects for durable execution state, but its public contract is the step model. Parallel steps can overlap I/O, and step-result limits are Workflows limits rather than a direct restatement of SQLite row size. Design against the workflow contract instead of inferring it from the underlying implementation.

Waiting workflow instances do not occupy active concurrency slots. They still retain state, and a large population that resumes together must share the account's running capacity. Plan the resume burst as well as the quiet waiting period.

The Paid plan supports 50,000 concurrent running instances, with creation limited to 300 per second per account and 100 per second per workflow. Fan-out still needs backpressure: a burst can reach creation or running limits even when individual instances are small. Size the producer rate and recovery path around those limits.

You cannot access the underlying Durable Object directly, as Workflows provides an abstraction that handles replay logic and step isolation. If you need direct Durable Object access (custom storage patterns, WebSocket connections, real-time state queries), use Durable Objects rather than Workflows.

The cost of durability​

Durability isn't free; every checkpoint is a write, and every step carries overhead. A workflow with 50 steps processing 10,000 orders daily writes 500,000 checkpoints, meaning you're paying for orchestration machinery on every step.

Workflows pricing involves per-step charges alongside the usual CPU time and request billing. Steps on paid plans bill at $0.80 per 100,000 beyond a monthly included 500,000 (the free plan includes 3,000 steps per day), and every step operation counts, including sleeps and event waits. Persisted state beyond 1 GB-month bills at $0.20 per GB-month. A 50-step workflow therefore costs roughly $0.0004 per execution in step charges. At 10,000 daily executions, that's approximately $115 per month in orchestration overhead alone, before accounting for compute within each step. Processing 10,000 independent items through a Queue costs a fraction of that, with no step checkpointing, no replay machinery, and no orchestration state.

This overhead is justified when failure is expensive. An order that captures payment but fails to reserve inventory creates a customer service nightmare. A document approval that loses its place after three of five approvers signed wastes days of human effort. Checkpoint cost is negligible compared to recovery cost.

Not every background process needs this machinery, however. Processing 10,000 uploaded images where each is independent works better with a Queue and idempotent consumers, which is simpler, cheaper, and faster since the images don't depend on each other and you'd otherwise pay for orchestration you don't use.

Use Workflows when steps depend on each other, when failure mid-process creates inconsistent state, or when you need visibility into where a process is stuck. Workflows earns its overhead in these scenarios. For independent, idempotent tasks, simpler patterns suffice.

A checkpoint protects a completed result​

Once a step result is durably recorded, replay can reuse it instead of rerunning the callback. A callback may nevertheless run more than once before that checkpoint exists. An external service can accept an action just before the callback fails or its result cannot be persisted.

Keep external effects inside steps and make their retries safe at the destination. Code outside steps can run again as Workflows reconstructs progress, including code after a completed step. Returning from a callback is not a transaction across the workflow engine and every service it called.

The constraints that shape design​

Before examining step types, you need to understand the constraints governing workflow architecture, as these aren't limitations to work around but boundaries that actually define good design.

The 1 MiB barrier​

danger

Each step can return at most 1 MiB as a JSON-serialisable value. Keep ordinary checkpoint results small; exceeding the limit fails the step. Larger data needs a storage reference or the supported stream-return path.

This constraint shapes everything about workflow design. If your natural inclination is to pass rich objects between steps, you'll hit this wall quickly, so instead store large data externally (R2, KV, or D1) and pass references. A step processing a large dataset returns { resultPath: "results/workflow-123" } rather than the dataset itself, and later steps fetch from storage using the path.

Design for this from the start. Discovering the constraint after building means restructuring your entire workflow. Small step results also improve observability: you can inspect workflow state in the dashboard without wading through megabytes of payload.

For the specific case where a step naturally produces a larger blob (a model response, an exported document, a generated image), step.do() accepts a ReadableStream return value. The platform persists the stream to R2 and hands later steps a stream-shaped reference, so you can keep the "step returns its output" mental model without manually splitting data and reference. The 1 MiB JSON ceiling still applies to plain values; streams are the explicit escape hatch for larger payloads.

The 128 MB runtime memory limit​

The 1 MiB barrier governs what you can persist between steps, while a separate constraint governs what you can hold in memory during a step: Workflows share Workers' 128 MB memory limit. This distinction trips up developers who assume larger persisted state limits mean larger runtime capacity.

Memory vs Persisted State

Persisted state limits (1 MiB per step, 1 GB total) describe what survives between steps. The 128 MB memory limit describes what you can hold in memory while a step executes. These are independent constraints serving different purposes.

A step processing a 500 MB file cannot buffer it in memory regardless of how little state it persists, because the step crashes before completing and never reaches the point where persisted state limits matter. The 1 GB total persisted state limit doesn't help here because memory exhaustion happens during execution, not during persistence.

The streaming pattern solves this. Instead of fetching a file into memory, transforming it, and writing the result, stream data directly from source to destination:

This example assumes a synchronous, byte-length-preserving processChunk transform. R2 needs a known output length; FixedLengthStream enforces the declared byte count. A transform that changes length needs a different upload strategy, such as multipart upload. Keep the input immutable and the transformation deterministic so a retried step writes the same result.

Streaming a length-preserving transformation to R2
const processedPath = await step.do("process-large-file", async () => {
const sourceObject = await env.SOURCE_BUCKET.get(inputPath);
if (!sourceObject) throw new Error("Source file not found");

// Stream through a transform without buffering the entire file
const transformedStream = sourceObject.body
.pipeThrough(new TransformStream({
transform(chunk, controller) {
// Process chunk by chunk, never holding the full file
controller.enqueue(processChunk(chunk));
}
}))
.pipeThrough(new FixedLengthStream(sourceObject.size));

const outputPath = `processed/${instanceId}/${filename}`;
await env.DEST_BUCKET.put(outputPath, transformedStream);
return outputPath; // Return reference, not data
});

The file flows through bounded stream buffers rather than being loaded in full. A 500 MB file can therefore fit the memory model, provided the transform and execution budget fit too. Streaming does not remove R2’s 5 GiB single-upload limit; larger outputs require multipart upload.

For workloads requiring analysis of entire files (not just streaming transforms), chunk the work across multiple steps where each step processes a portion, persists intermediate results to R2, and returns a reference so the next step can pick up where the previous left off. This pattern trades step overhead for memory safety.

Determinism requirements​

Replay and Determinism

Workflows may replay, re-executing your workflow code to reconstruct state after failures. Non-deterministic operations outside steps cause problems. If you generate a random number outside a step and use it to decide whether to execute a subsequent step, replay generates a different number and makes a different decision.

Wrap anything that might differ between executions in a step. Timestamps, random numbers, external reads: if the value changes, checkpoint it. The step's result persists; replays use the persisted value rather than re-computing.

Idempotency for external calls​

Steps may retry, and if a step calls an external API, that API might receive the same request multiple times. Payment processors, email services, and inventory systems must handle duplicate requests gracefully.

This discipline isn't Workflows-specific; any distributed system needs it, but Workflows makes it explicit and unavoidable. Every step modifying external state needs an idempotency strategy. Most payment processors accept idempotency keys, where a retry with the same key returns the original result instead of charging again. When the effect and its receipt share a database, commit them atomically with a unique operation identifier. A local check followed by an external call leaves a gap: the remote effect may succeed before its receipt is stored. If the destination has no idempotency contract, reconcile an uncertain outcome or route it for intervention rather than assuming a retry is harmless.

Step types and when to use each​

Workflows provides four step types, and understanding when to use each matters far more than memorising their syntax.

step.do(): execute and persist​

The fundamental step wraps code execution with durability, persisting the result before the workflow continues; if the step fails, Workflows retries according to configuration.

Capturing payment with a stable operation key
const paymentResult = await step.do("capture-payment", async () => {
return await paymentService.capture({
paymentMethodId: order.paymentMethodId,
amount: order.total,
idempotencyKey: `capture:${order.id}`
});
});

The string identifier serves three purposes: observability in dashboards, debugging in logs, and replay identification. Step names must be deterministic. Don't include timestamps or random values; replay will fail to match steps with their persisted results.

step.sleep() and step.sleepUntil(): durable waiting​

These steps pause execution without holding active compute for the interval. A sleep counts as a step, and persisted state remains subject to storage billing. A longer wait does not require a polling loop.

Use a durable sleep when the process owns the next deadline: a reminder, an expiry or a scheduled follow-up. Use an event wait when another system determines when work can continue. Both keep that decision in the workflow rather than in a separate scheduler and status table.

AWS Step Functions Standard also supports long waits without duration-based execution charges; Express is limited to five minutes and bills duration. Choose between platforms on integration, recovery behaviour and the complete cost model.

step.waitForEvent(): pause for external signals​

This step type enables workflows that wait for external input such as human approvals, webhook callbacks, file uploads, or third-party notifications. The workflow pauses until a matching event arrives or a timeout expires.

Waiting for an approval event
const approval = await step.waitForEvent("manager-approval", {
type: "approval-decision",
timeout: "48 hours"
});

The workflow owns the suspended state, but the receiving system still needs to authenticate the event and route it to the correct instance. Keep the external correlation identifier durable so a webhook can be retried without losing its destination.

Set the timeout explicitly to the business deadline; the default is twenty-four hours. An event sent after instance creation can arrive before the matching wait and be buffered. Give approval events a stable identity and check their current business meaning when consumed: matching an event type does not establish that permission remains valid.

If the event never arrives, the timeout path needs a useful outcome: escalation, cancellation or a recorded incomplete state. Waiting for forty-eight hours is a product policy, not an error-handling strategy on its own.

The step granularity decision​

One fundamental design question is how much work belongs in a single step, and this decision has very real consequences for your workflow's efficiency and resilience.

Too granular creates overhead: a workflow with 100 steps for a simple operation writes 100 checkpoints, meaning you're paying for durability granularity you don't need.

Too coarse bundles operations that should fail independently: a single step making five API calls retries all five if the fifth fails, and if any call has side effects, you risk duplicates.

The principle is simple: one step per side effect. Bundle reads freely since failure just means re-reading, but isolate writes ruthlessly. Multiple queries gathering data can share a step, but the moment you write, send, or charge, you need a step boundary.

A step boundary cannot make updates across three services atomic. That requires a transaction mechanism covering all three. Without one, design a saga: checkpoint each action and define compensation for the effects already completed. Intermediate states remain visible and compensation can fail, so the later rollback section includes a path for intervention.

Parallel steps with Promise.all and Promise.race​

Steps can run concurrently using Promise.all, allowing independent operations that don't depend on each other's results to execute in parallel and reduce total workflow duration:

Parallel independent steps
const [userData, orderHistory, recommendations] = await Promise.all([
step.do("fetch-user", () => fetchUser(userId)),
step.do("fetch-orders", () => fetchOrders(userId)),
step.do("fetch-recommendations", () => fetchRecommendations(userId))
]);

Promise.race and Promise.any require more care because the workflow engine may restart between steps and steps are cached by name. If you use Promise.race with steps directly, the cached result after a restart might differ from the original race winner, causing subtle bugs.

Promise.race Caching Gotcha

Wrap Promise.race or Promise.any calls in a containing step.do to ensure consistent caching across workflow restarts. Without the wrapper, the race result may change if the workflow hibernates and resumes.

Persisting the result of a race
// Correct: wrap the race in a step for consistent caching
const winner = await step.do("race-for-response", async () => {
return await Promise.race([
step.do("fast-source", () => fetchFromFastSource()),
step.do("slow-source", () => fetchFromSlowSource())
]);
});

The outer step persists the race result, so on replay, the workflow retrieves the persisted winner rather than re-racing, ensuring deterministic behaviour regardless of which source happens to respond first on a given execution.

What goes wrong: non-idempotent steps​

A team builds an order workflow where the email notification step calls their email service, which sends successfully. But then the step fails, perhaps because a subsequent line throws or a timeout occurs after the send but before acknowledgment. The workflow retries, the email sends again, and the customer receives two identical order confirmations.

Worse still: a payment capture step charges the card, then fails on response parsing, so retry charges again and the customer pays twice.

The root cause is treating steps as atomic when they're not. A step making an external call and then doing more work can fail after the external effect but before completion; the effect happened, and retry makes it happen again.

The fix is a stable operation identity that the destination uses to suppress repeated effects. Where the service supports idempotency keys, derive one from the durable business operation and intended effect, reuse it across retries and replacement Workflows, and verify its retention period. Keep execution IDs separate so starting a repair run does not create a new external effect. A key passed to an API that ignores it provides no protection.

The Idempotency Imperative

Every step that modifies external state needs an idempotency strategy. This isn't paranoia; retries are normal operation in durable execution systems. If your step can't safely run twice, it's not production-ready.

Failure handling​

Steps fail, networks time out, and services return errors, so understanding how Workflows handles failure separates robust workflows from fragile ones.

Retry behaviour​

When a step throws an exception or times out, Workflows catches the error, records it, and decides whether to retry using a default policy of five retries after the initial attempt, with a ten-second initial delay and exponential backoff.

Step code may run multiple times, but successful results persist exactly once, and retry state persists to storage so infrastructure failure during backoff doesn't lose track of where the workflow was.

Retry configuration adapts your workflow to the reliability characteristics of what it calls; defaults assume well-behaved services while customisation acknowledges reality.

Rate-limited APIs need longer backoff, since aggressive retries make things worse when an API returns 429 errors. Increase initial delay and let exponential backoff create breathing room. The delay can also be computed per attempt: retry configuration accepts a function receiving the attempt number and the triggering error, so a rate-limit error can back off progressively while other failures retry quickly.

Slow-failing services need shorter timeouts, since waiting for the default timeout wastes time if a service hangs 30 seconds before failing. Set explicit step timeouts to fail fast and retry sooner; the default step timeout of ten minutes is appropriate for long-running operations but excessive for APIs that should respond in milliseconds.

Permanently failing operations shouldn't retry, since a payment declined for insufficient funds won't succeed on attempt six; validation failures, business rule violations, and malformed inputs are all permanent. Throw NonRetryableError to stop immediately:

Stop the workflow when retry can't help
import { NonRetryableError } from "cloudflare:workflows";

if (!validation.valid) {
throw new NonRetryableError(`Invalid order: ${validation.reason}`);
}

Flaky but critical services might need more attempts if they fail often but eventually succeed and you can't fix them, so increase retry limits. But treat this as technical debt rather than a proper solution.

Each step callback receives a context object exposing the current attempt number, the step's own name, an invocation counter, and the resolved step configuration. This is useful for logging, progressive degradation, or conditional logic based on how many retries have occurred:

Using step context for retry-aware logic
await step.do("look-up-product-information", async (ctx) => {
console.log(`Step ${ctx.step.name} (call ${ctx.step.count}), attempt ${ctx.attempt}`);

// A read-only lookup can try another source after two failures
const provider = ctx.attempt <= 2 ? primaryProvider : fallbackProvider;
return await provider.lookup(productId);
});

Do not apply this fallback to an uncertain payment: the first provider may have charged successfully before its response was lost. Reconcile with that provider before deciding whether another charge is permissible.

The ctx.attempt value is 1-indexed: 1 on the first try, 2 on the first retry, and so on. ctx.step.name lets shared helper functions log against the calling step without duplicating string literals, and ctx.step.count distinguishes successive invocations of the same step within a workflow (for example, when a step name appears inside a loop). Use these signals to add observability to retry-heavy workflows or to implement fallback strategies that escalate gracefully rather than repeating the same failing call.

When failure is acceptable​

Some operations are truly optional: enrichment services adding nice-to-have data, analytics calls that don't affect business logic, or notifications retried later through other means. Wrap these in try/catch and continue:

Handling optional step failures
try {
await step.do("optional-enrichment", () => enrichmentService.enhance(data));
} catch {
// Enrichment unavailable; proceed without it
}

The workflow continues despite failure, so log it and perhaps alert on high failure rates, but don't let optional operations block critical paths.

Compensation and rollback​

When a step fails after previous steps succeeded, you may need to undo earlier work: payment captured but shipping failed means refunding the payment; inventory reserved but payment failed means releasing the inventory.

The saga pattern pairs a sequence of local operations with compensating actions. A refund can correct a charge without erasing its history; an email or a downloaded document cannot be recalled. Define compensation where it exists and an explicit resolution path where it does not.

Saga pattern: compensate a completed effect

Loading diagram…

Workflows runs the saga for you. Attach a rollback handler to any step.do(), and when the workflow fails terminally Cloudflare invokes the registered handlers automatically, in reverse step-start order, through the same durable machinery that runs the forward path. Without it, you would track which operations succeeded and execute compensations yourself in a catch block.

The application helpers must make both payment capture and refund safe to retry, using stable operation identities at the payment provider.

Registering payment compensation
const payment = await step.do(
"capture-payment",
async () => capturePayment(order),
{
rollback: async ({ output }) => {
if (output?.chargeId) await refundCharge(output.chargeId);
},
rollbackConfig: {
retries: { limit: 5, delay: "10 seconds", backoff: "exponential" },
},
},
);

Each handler receives the triggering error, the step's persisted output (which may be undefined when the step never completed, so guard for it), and a ctx object carrying the step name, attempt count, and resolved configuration for logging and customised recovery. The rollbackConfig gives compensation its own retry policy and timeout, which matters because compensation that fails leaves you worse off than the original failure: a refund that never goes through is a worse state than a shipping error. Have a fallback for when compensation exhausts its retries, usually alerting humans for manual intervention.

One behaviour deserves emphasis: rollback runs only when the workflow fails terminally. Catch an error and continue, and the steps on that path are not rolled back. Compensation is the platform's response to a workflow giving up, not a general-purpose try/finally.

Step ordering also matters for compensation design. Place operations that are harder to reverse later in the workflow so they only execute after easier-to-compensate steps have already succeeded. Notifications deserve particular attention here: a payment refund is mechanical, but you cannot recall a sent email or SMS. If business rules allow it, defer customer-facing notifications until all preceding steps have confirmed success. This saves you from the awkward alternative of sending correction messages explaining that the previous message should be disregarded.

Workflows that succeed but are wrong​

Not all failures throw errors; a payment step might return success while charging the wrong amount, or a reservation step might confirm inventory that doesn't exist, and the workflow completes with green checkmarks and incorrect business state.

These silent failures are harder to catch than exceptions because they require validation steps verifying that outcomes match expectations, reconciliation processes comparing workflow results against source-of-truth systems, and monitoring that detects anomalies in business metrics even when technical metrics look healthy.

Build verification into critical workflows by verifying that captured amounts match order totals after payment capture, confirming that reservations exist after inventory reservation. External systems often lie, usually unintentionally.

When Workflows get stuck​

Workflows can sometimes get stuck, and these failure modes deserve explicit names so you can recognise them.

Poison step loop occurs when a step fails consistently but not with NonRetryableError, so the workflow retries according to configuration, exhausts attempts, and either fails or loops longer if you've configured more retries. A step calling a service that always returns 500 retries until something intervenes. The fix is recognising which failures are permanent and throwing NonRetryableError.

Event starvation happens when a waitForEvent step never receives its event because the external system fails, loses the event, or was never configured correctly, so the workflow reaches its configured timeout, or the twenty-four-hour default. Choose the deadline explicitly and design the escalation path.

Compensation cascade failure manifests when your rollback logic itself fails: step 3 failed, you try to compensate step 2, but compensation also fails, leaving you with partially captured payments and partially reserved inventory in a workflow stuck trying to clean up. Design compensation to be more reliable than original operations and have manual intervention as the ultimate fallback.

State size overflow can occur when an individual non-stream result exceeds 1 MiB or total retained instance state reaches its own limit. Use small checkpoint results; store large or long-lived artefacts externally and pass references. Streamed results also count towards total instance storage.

Your intervention playbook breaks down into clear steps:

Diagnosis first: The dashboard shows step-level progress, retry counts, and error messages, so before terminating, understand why the workflow is stuck (external service down, bug in your code, or event source failing to send).

Stop stuck instances deliberately. Termination stops further execution; terminate({ rollback: true }) requests compensation through registered rollback handlers. Deleting an instance stops execution and erases its stored state without running those handlers. Use deletion for disposal, after any required compensation and evidence capture, rather than as a substitute for rollback.

Send missing events for workflows waiting on external signals; if an approval event was lost, resend it programmatically, and if a webhook failed, replay it so the workflow resumes from where it paused.

Fix and redeploy for bugs, but remember that running instances continue with their persisted state and don't automatically pick up new code. See versioning below for how to handle this.

Versioning and deployments​

Treat workflow code, checkpoint data and external interfaces as a versioned contract. A deployment test should include an instance created before the change: pause it at a meaningful boundary, deploy, resume it, and verify the resulting steps and side effects. Testing only newly created instances misses the compatibility question.

For a breaking change, a separately named Workflow and deployment provides an explicit boundary. Direct new work to it, keep the old implementation and its dependencies available, and drain or deliberately migrate existing instances before retiring them. Give that transition an owner and a deadline.

Keep dependencies compatible​

A long-lived workflow can call a separately deployed Worker long after the process began. Maintain the RPC methods, event payloads and data formats that outstanding instances may still need. Add a versioned method or event type when a change cannot remain compatible.

A sleeping process is an outstanding dependency, even while it consumes no active compute. Code rollback alone cannot undo its earlier payments, reservations or writes. Include those effects in the migration and recovery plan.

Handling platform transients​

Classify known permanent failures inside the step callback, where NonRetryableError can prevent futile retries. Let transient failures use the step's retry policy. An error caught after await step.do() may mean that policy has already been exhausted; rethrowing it does not grant another automatic set of attempts.

Unknown failures deserve a bounded retry policy and useful diagnostic context. If retries finish without recovery, the workflow still needs an operational decision: resume through a supported path, compensate, or start a replacement operation whose side effects are safe to repeat.

Observability in production​

The dashboard shows workflow instances, step progress, and error messages, which is useful for debugging individual failures but insufficient for production operations. The dashboard also generates visual diagrams of each workflow by parsing the TypeScript source via abstract syntax trees, so a complex workflow with conditional branches, parallel sections, and waitForEvent calls renders as a flowchart you can read at a glance. Useful for code review, onboarding, and explaining a stuck instance to a stakeholder who does not read TypeScript.

Progress visibility and evidence retention are separate decisions. An instance subscription streams retained workflow and step events, then follows live execution, so a progress screen need not poll or maintain a second state machine. Filters select relevant events, and a cursor resumes a disconnected subscriber. Keep business actions idempotent if subscriptions trigger follow-up work.

Completed and errored instance state has a finite retention period. New Workflows on the Paid plan default to seven days, with up to 30 days configurable; existing Workflows retain their configured period. The Free plan retains state for three days. Set success and error retention explicitly when investigation or data-handling requirements matter, and persist business records outside the workflow before its state expires or is deleted. Durable execution is not a permanent audit archive.

Alerting on stuck workflows requires external monitoring. Export workflow metrics to your observability platform. Alert when workflows remain "running" beyond expected duration, when step retry counts exceed thresholds, when failed workflow rates increase.

Duration analysis identifies bottlenecks by answering questions like which step takes longest and whether it's consistently slow or shows variable latency. Track step durations over time since degradation often appears gradually before causing failures.

Business metrics matter more than technical metrics: a workflow completing successfully but taking 4 hours instead of 4 minutes might not trigger technical alerts, so monitor what actually matters: time from order placement to shipping label generation, approval workflow cycle time, and data synchronisation lag.

Log context aggressively: when a step fails at 3 AM, you'll want to know what inputs it received, what external calls it made, and what state preceded the failure. Include workflow instance ID in all logs for correlation.

Long-running process patterns​

Workflows excel at processes spanning extended time (hours, days, or weeks), and these patterns share common characteristics: steps depending on each other, external waits where polling cost would be prohibitive, and the need for visibility into where processes stand.

Sequential with external waits​

Order processing exemplifies this pattern: validate, capture payment, reserve inventory, wait for warehouse confirmation, generate shipping, wait for carrier pickup, notify customer, wait for delivery, and request review. The waits are where Workflows earns its keep because each waitForEvent hibernates the workflow without active compute while waiting. Event-wait steps and retained state still contribute to billing. Without this capability, you'd poll external systems, maintain status in a database, and run scheduled jobs to check for updates.

Approval chains with escalation​

Document approvals need timeouts and escalation paths like waiting 48 hours for manager approval, escalating to director with 24 hours if timeout occurs, and auto-rejecting with notification if still no response. The workflow encodes the escalation policy directly with no external state machine, no cron jobs checking approval status, and no separate escalation service. The policy is readable in code, and changes deploy like any other code change.

Chunked batch processing​

Large datasets benefit from per-chunk checkpointing because a data migration processing 1 million records can't be a single step; failure at record 999,000 would restart from record 1. Chunk the data and process each chunk as a separate step so failure at chunk 47 resumes at chunk 47, not chunk 1.

Make each chunk a step, and progress persists automatically while chunk size balances checkpoint overhead (smaller chunks mean more checkpoints) against recovery cost (larger chunks mean more re-processing on failure). For most workloads, 1,000 to 10,000 items per chunk balances well.

When Workflows versus when alternatives​

Workflows versus Queues​

Chapter 9 covers Queues in depth, but the architectural choice between them needs to happen now.

Workflows orchestrate dependent steps where step 3 needs output from step 2, you need visibility into where a process is stuck, processes span significant time, order matters, and compensation is required for partial failure.

Queues distribute independent tasks for processing 10,000 items where each is separate, order doesn't matter, fire-and-forget is acceptable, and you want parallel processing.

A common pattern combines both, where Workflows orchestrates the overall process while Queues handles parallel fan-out within steps. An order workflow might queue 100 notification sends to process concurrently, wait for completion, and then continue.

Workflows versus Durable Objects​

Workflows is built on Durable Objects, so when should you use the abstraction versus the primitive?

Use Workflows when orchestration fits: you have sequential steps with persistence between them, human wait times measured in hours or days, and standard retry and timeout needs. You want the framework to handle replay, checkpointing, and step isolation.

Use Durable Objects directly when you need patterns Workflows doesn't express: arbitrary mutable state that clients query and change outside step execution, bidirectional WebSocket sessions beyond workflow progress subscriptions, complex conditional logic that doesn't map to linear steps, custom storage patterns beyond step results, or receiving messages while processing.

The test is straightforward: if you're fighting Workflows' step model to express your logic, you probably want a Durable Object, but if your process naturally decomposes into "do this, then this, then wait, then this," Workflows saves you from building that machinery yourself.

A third option combines both: the AgentWorkflow class (Chapter 19) lets AI agents delegate durable execution to Workflows while maintaining real-time WebSocket connections with clients. The agent handles interactive decision-making and user communication while the Workflow handles reliable multi-step execution with checkpointing and retry. This composition is particularly effective for human-in-the-loop patterns where users interact with an agent that triggers and monitors long-running processes.

Comparing to external orchestrators​

If you're evaluating Workflows against AWS Step Functions or Temporal, you'll find the differences are architectural in nature.

AspectCloudflare WorkflowsAWS Step FunctionsTemporal
DefinitionTypeScript codeJSON (ASL) or SDKCode (multiple languages)
Execution ModelDurable Object per instanceManaged serviceSelf-hosted or Temporal Cloud
Max DurationNo fixed instance duration; step limits apply1 year (Standard)Unlimited
State Limit1 MiB per non-stream step result256 KiB per task, state or execution input/outputConfigurable
Step PricingPer step, plus CPU time and storagePer state transitionPer action (cloud)

Step Functions uses a JSON-based definition language (ASL) that constrains expression but integrates tightly with AWS services, with native connectors to Lambda, SNS, DynamoDB, and dozens of other services reducing integration code. Pricing per state transition can be expensive for step-heavy workflows, though if you're deep in AWS and need tight integration with Lambda and DynamoDB, Step Functions' native connectors save integration work despite ASL's constraints.

Temporal offers a mature orchestration model, with worker infrastructure you operate or Temporal Cloud. Evaluate it when its workflow tooling, integrations and operational model fit your process. Long duration or child processes alone do not distinguish it from Cloudflare Workflows; compare the specific recovery, versioning and dependency behaviour you need.

Workflows fits teams that want durable orchestration alongside Workers and Cloudflare bindings. Step Functions may reduce integration work for an AWS-centred system; Temporal may fit more elaborate orchestration requirements and existing operational expertise. Compare the failure and migration paths your process needs, rather than treating a smaller API as proof of a simpler production system.

Dynamic Workflows: durable execution that follows the tenant​

Everything so far assumes you know the workflow's code at deploy time. You write a WorkflowEntrypoint, deploy it, and every instance runs that same definition with different input. That assumption breaks for platforms where the shape of the process belongs to the tenant rather than to you: a SaaS product where each customer defines their own onboarding sequence, approval chain, or billing-retry logic, or an agent framework where the model generates a multi-step plan at runtime. You cannot deploy a workflow class per tenant when there are tens of thousands of them, and you certainly cannot redeploy every time a tenant edits their automation.

Dynamic Workflows, the MIT-licensed @cloudflare/dynamic-workflows library, closes this gap by running a Workflow inside a Dynamic Worker, the runtime code-loading mechanism Chapter 19 covers for agents. The workflow definition is loaded on demand rather than deployed ahead of time, so each tenant's code lives wherever you keep it and is instantiated only when needed.

The important problem is recovering the same process after its code has left memory. Dynamic loading makes code identity part of durable state: retain the definition an outstanding instance needs, including its dependencies. Waiting removes active compute consumption, but retained workflow state and code storage can still incur costs.

Reach for this when the orchestration itself is tenant- or agent-supplied. If all tenants run one process with different data, a deployed Workflow with configuration input is easier to operate. Dynamic loading earns its complexity when tenants supply different steps; preserve code identity and apply limits to that execution boundary. Chapters 19 and 25 cover the agent and multi-tenant implications.

Testing Workflows​

Test Workflows through the Workers Vitest integration and its introspection helpers. Inspect the steps and events relevant to the business outcome, including failures that should trigger compensation. The test should explain why the process is correct, not merely report that an instance completed.

The cloudflare:test helpers include introspectWorkflowInstance for a known instance and introspectWorkflow when the test needs to discover created instances. These support local tests without a deployed Workflow.

The introspection APIs let you mock step results, inject events at precise moments, and disable sleeps for time-dependent tests, so a workflow normally sleeping seven days can have sleeps disabled entirely, and a step calling an external payment API can return a mocked response. You control exactly what each step produces and verify what subsequent steps receive.

For time-dependent behaviour, call disableSleeps() on the modifier object rather than parameterising durations; for waitForEvent testing, use mockEvent() to inject events without querying instance status manually; and for compensation logic, mock step failures to trigger rollback paths on demand.

The principle is straightforward: don't test Workflows' machinery because Cloudflare tests that steps persist and retries happen; instead, test your business logic by verifying that given these inputs, the right steps execute, and given this mocked failure, compensation happens correctly. The introspection APIs make these assertions straightforward.

What comes next​

Not every background task needs a persistent sequence of steps. Chapter 9 covers Queues: buffering independent work, controlling consumer load and recovering from duplicate delivery. Combine the two when a workflow needs to fan work out and collect its results.