Skip to main content

Chapter 9: Queues: Asynchronous Processing

How do I handle work that doesn't need immediate response?


Not every operation belongs in a request/response cycle; sending a welcome email shouldn't block user registration, processing an uploaded image shouldn't delay the upload confirmation, and syncing data to a third-party system shouldn't slow down the operation that triggered it.

Queues decouple producers from consumers by allowing a Worker to write a message and return immediately while another Worker (or the same one in a different context) processes that message later. The user can receive a response before processing finishes; the application must still track failures, retry exhaustion and expired work.

Queues offer a simpler model than Workflows: no orchestration, no step dependencies, and no durable state spanning operations. Messages go in, messages come out, and work gets done. For many background processing needs, that simplicity is exactly right, though knowing where queues stop being the right abstraction matters more than mastering their features.

What's different from SQS and Service Bus​

If you're coming from AWS or Azure, the biggest difference isn't features but rather the operational model. SQS is a standalone service configured separately from Lambda, with IAM roles connecting them, visibility timeouts that interact badly with Lambda timeouts, and a separate deployment pipeline. Cloudflare Queues, by contrast, is a binding within your Worker project where queue and consumer share the same codebase, the same deployment, and the same configuration file.

Fewer moving parts means fewer failure modes: you won't debug production incidents caused by mismatched visibility timeout and Lambda timeout settings because those settings don't exist separately, and you won't trace IAM permission errors across services since there are no cross-service permissions to configure.

The trade-off is capability. SQS FIFO offers ordering by message group and submission deduplication; it still needs consumers that handle retries safely. Azure Service Bus adds sessions, routing and larger-message options. Choose those capabilities when they match the requirement, even if the rest of the application uses Workers.

The honest comparison is that Cloudflare Queues is simpler to operate and natively integrated with Workers, though this comes at the cost of advanced messaging features you may or may not need. If your requirements are simple (process this later, retry on failure, alert if it keeps failing), Cloudflare Queues handles that with less operational surface than hyperscaler alternatives, but if you need strict ordering, exactly-once semantics, or messages larger than 128 KB without indirection, you should look elsewhere.

The mental model​

Queues retain work for later delivery within a configured retention window. Batching and retries help consumers make progress; dead letters expose work that needs intervention. The application owns the promise made to the user.

Queues answer "do this later"; Workflows answer "do these things in order"; Durable Objects answer "coordinate this now". Choose based on what your problem actually is, not which tool you know best.

Think of a queue as a time-shifted function call. Instead of calling processOrder(order) synchronously and waiting, you write a message describing the call and let the system invoke it later. The queue handles the mechanics: persisting the intent, distributing it to available workers, retrying on failure, routing failures somewhere visible.

Queues defers work and absorbs bursts between independently progressing components. A status record can track an individual task. Choose Workflows when progress depends on durable checkpoints across dependent steps, or when recovery needs to resume a business process rather than repeat one task.

The producer-consumer relationship is deliberately loose because producers don't know which consumer will handle their message, when, or whether it will succeed on the first attempt, and consumers don't know how many producers exist or what rate messages will arrive. This loose coupling lets producers and consumers scale independently. The queue alone does not tell a user whether the business operation succeeded; maintain an application status record when that answer matters.

When to choose Queues​

Use queues deliberately, not by default.

Queues excel when tasks are independent (each message processes without knowledge of others), order doesn't matter (or can be handled at the consumer level through timestamps), fire-and-forget is acceptable (tracking completion through side effects rather than queue state), you need parallel processing (multiple consumers handling messages concurrently), and latency is tolerable (seconds or minutes of delay between send and processing).

Dependent steps, durable waits and compensation are reasons to consider Workflows. Progress visibility alone is not: an independent queued task can update a status record without becoming an orchestrated process.

Use waitUntil() when work is truly fire-and-forget: no retry needed, no confirmation needed, and failure acceptable. The work is nice-to-have, and you can't afford even minimal queue dispatch overhead.

Use Durable Objects when you need real-time coordination (rate limiting, presence detection, live updates) where state must be immediately consistent and multiple requests must see that consistent state.

Use Workflows when dependent steps must checkpoint progress and resume after failure, or when a multi-step process needs durable waits and compensation. Use Queues for independent tasks whose retry boundary is the task itself.

Look for orchestration when queue consumers accumulate dependent steps, durable waits, per-step retries and compensation. A business status ledger can be useful with either Queues or Workflows; it does not by itself justify moving the execution. Choose Workflows when its checkpointed sequence removes machinery you would otherwise maintain.

Delivery guarantees and their implications​

Queues provide specific delivery guarantees that shape how you design consumers, and understanding these guarantees prevents architectural mistakes that surface only under failure conditions.

At-least-once delivery​

Queues use at-least-once delivery, so consumers must handle redelivery after a failure or an ambiguous acknowledgement. Delivery is still bounded by retention and retry policy: exhausted messages go to a configured dead-letter queue or are deleted. Monitor those boundaries rather than assuming every accepted message eventually completes its business operation.

Queue delivery and business effects are separate guarantees. A broker can deduplicate submissions or preserve order within a group, while a consumer still repeats work after a timeout. Design the side effect to remain safe when the same message is processed again.

The idempotency requirement​

Idempotency Is Required

Queues provide at-least-once delivery. Messages will be redelivered after failures, timeouts, or infrastructure events. Design every consumer to handle duplicate messages safely.

Messages will be redelivered. A consumer crash, network failure, or acknowledgment timeout causes redelivery even when processing succeeded. Design every consumer to handle duplicate messages correctly from day one. Retrofitting idempotency after production incidents is expensive and error-prone.

Message acknowledgment transfers ownership. The queue keeps the message safe and redelivers if needed until the consumer calls ack(). That call transfers responsibility to your code. If your code crashes after acknowledgment, the message is gone. Acknowledge only after processing completes successfully.

Idempotent processing means repeating an operation has the same intended effect. For an external action such as sending an email or charging a card, use a destination that accepts a stable idempotency key and honours it across retries:

Keep the operation identity stable across deliveries
await sendEmail(task, { idempotencyKey: task.messageId });
message.ack();

Here sendEmail represents a provider integration that implements that guarantee; the key is not a Cloudflare Queues feature. Verify the provider's retention window and failure semantics.

A local record saying an email was sent is useful for tracking, but a check followed by a send is not atomic. Two consumers can pass the check, or one can crash after sending but before recording success. For database-only work, keep the deduplication record and mutation in one transaction. For external work without destination-supported idempotency, acknowledge the remaining duplicate risk and design reconciliation around it.

Ordering belongs to a defined entity​

Queues can deliver messages out of order. Use them for work whose correctness does not depend on delivery sequence, or whose updates carry enough version information to reject stale effects.

When order matters, identify its scope. A Workflow expresses dependencies within one process. A Durable Object can own ordering for one entity. Neither recovers the producer's original order merely because messages are forwarded to it: if messages arrive out of sequence, the application still needs a sequence or version rule.

A FIFO broker may be the simpler choice when ordered delivery by key is the requirement. Ordering one customer's events does not require serialising every customer's work. Decide which operations must wait for each other before choosing the transport.

Designing reliable producers​

Any Worker can write messages to a queue through a binding. Call send() with a JSON-serialisable object up to 128 KB. For larger payloads, store in R2 and send a reference. But mechanics aren't where production systems fail.

The architectural questions matter more: When should you batch sends versus sending individually? How do you handle queue unavailability? How do you evolve message schemas?

Batching sends reduces overhead when you have multiple messages ready simultaneously (up to 100 messages or 256 KB per batch). But batching changes failure semantics. If a batch send fails, do you retry the whole batch? If some messages are more critical than others, should they be batched together? For independent messages of equal importance, batch aggressively. For messages with different criticality, send individually or implement batch-level retry logic.

A queue send can fail or return an ambiguous result. Retry with a stable operation identifier. If the business operation must remain accepted while the queue is unavailable, persist an outbox record with the state change and dispatch it later. A buffer in Worker memory is not a durable fallback.

Message schema evolution is the concern most producers ignore until an incident. You add a field, but old messages in the queue lack it. You rename a field, but consumers see both names during transition. You remove a field that consumers still expect.

The safest approach is additive-only changes with consumer tolerance for missing fields. New fields get default values; removed fields are ignored rather than causing errors. For breaking changes, version your message schema explicitly and have consumers handle multiple versions during transitions.

Messages can be delayed up to 24 hours. Use delay for a simple deferral; use a Workflow when the wait belongs to a process with later dependent actions, cancellation or recovery decisions.

Configuring consumers​

A Worker consumer receives a batch. Successful return acknowledges messages unless they were explicitly marked for retry; explicit per-message acknowledgement lets completed items survive a later batch failure. HTTP pull consumers have a separate visibility-timeout model.

Choose batch size, batch timeout and concurrency together:

WorkloadBiasCost to watch
Throughput with flexible latencyLarger batches and longer fill timeQueue age and repeated work after failure
User-visible notificationsShort fill timeMore invocations at low volume
Rate-limited downstream serviceBounded concurrency and deliberate backoffBacklog growth during an outage

Concurrency should reflect downstream capacity, not how quickly Cloudflare can start consumers. A queue absorbs a burst; it does not make the database or external API faster. Increase throughput only while error rates, dependency saturation and oldest-message age remain acceptable.

Retry strategy​

The default retries a failed delivery three times. Configure delay or backoff to suit the downstream service; exponential backoff is an application policy, not the default retry schedule.

As Chapter 8 discussed, exponential backoff reflects hope that a transient condition will clear; immediate acknowledgment of permanent failures reflects certainty that retrying won't help. The critical decision is distinguishing between them:

Distinguishing transient from permanent failures
try {
await processTask(message.body, env);
message.ack();
} catch (error) {
if (isTransient(error)) {
message.retry({ delaySeconds: 60 });
} else {
await logPermanentFailure(message.body, error);
message.ack(); // Stop retrying
}
}

Transient failures (network timeouts, rate limits, temporary unavailability) should retry with increasing delays. Permanent failures (malformed data, business rule violations, missing dependencies) should not retry. Acknowledging a permanently failed message prevents infinite retry loops; logging ensures visibility.

The isTransient() function encodes your understanding of downstream failure modes. HTTP 429 (rate limited) is transient; HTTP 400 (bad request) is permanent; HTTP 500 might be either. Connection timeouts are usually transient; validation errors are always permanent. Getting this classification wrong wastes resources on hopeless retries or drops recoverable messages.

Per-message acknowledgement for ETL pipelines​

Individual acknowledgement prevents already acknowledged messages from being replayed because a later message in the batch fails. It reduces repeated work, but does not close the crash window between a side effect and its acknowledgement. Keep the operation idempotent where possible.

Per-message acknowledgement in ETL
export default {
async queue(batch: MessageBatch<ETLTask>, env: Env) {
for (const message of batch.messages) {
try {
await processRecord(message.body, env);
message.ack(); // Acknowledge immediately on success
} catch (error) {
if (isTransient(error)) {
const delaySeconds = Math.min(3600, 2 ** message.attempts * 10);
message.retry({ delaySeconds });
} else {
await logPermanentFailure(message.body, error, env);
message.ack(); // Stop retrying permanent failures
}
}
}
}
};

Use per-message acknowledgement when messages can succeed independently and retrying the whole batch would waste work. A crash before acknowledgement can still repeat a completed action. Choose batch acknowledgement for an atomic batch operation or when repeating the whole batch is safe; avoid treating acknowledgement granularity as a substitute for idempotency.

Pull-based consumers​

Everything discussed so far assumes push-based consumption: Cloudflare invokes your Worker when messages arrive. Queues also support pull-based consumption, where any HTTP client fetches messages on its own schedule. A queue supports one or the other, not both.

Pull consumers exist for when the processor cannot be a Worker: message processing runs on Kubernetes, your consumer is a Go binary that won't transpile to JavaScript, or your legacy system needs queue integration but can't adopt Workers yet. Pull consumers bridge Cloudflare's queue infrastructure to wherever your processing runs.

The mental model shifts from "Cloudflare calls you" to "you call Cloudflare." Your consumer polls the queue's HTTP endpoint, receives a batch of messages, processes them, then acknowledges through another HTTP call. The queue doesn't care what makes these calls: a Lambda function, a container in ECS, a cron job on a VM, or a developer's laptop during debugging.

Set the account, queue and API-token environment variables before running these commands. Replace the two lease placeholders with distinct lease IDs returned by the pull: one for a processed message and one for a message to retry.

Pull-based queue consumption via HTTP API
QUEUE_URL="https://api.cloudflare.com/client/v4/accounts/${ACCOUNT_ID}/queues/${QUEUE_ID}"

# Pull up to 100 messages with a 30-second visibility timeout
curl -X POST "${QUEUE_URL}/messages/pull" \
-H "Authorization: Bearer ${API_TOKEN}" \
-H "Content-Type: application/json" \
-d '{"batch_size": 100, "visibility_timeout": 30000}'

# Acknowledge processed messages; mark others for retry
curl -X POST "${QUEUE_URL}/messages/ack" \
-H "Authorization: Bearer ${API_TOKEN}" \
-H "Content-Type: application/json" \
-d '{"acks": [{"lease_id": "processed-lease-id"}], "retries": [{"lease_id": "retry-lease-id"}]}'

The visibility timeout matters more here than with push consumers. When you pull a message, other consumers can't see it until the timeout expires or you explicitly acknowledge or retry. Set the timeout longer than expected processing time but short enough that stuck messages don't block indefinitely. Thirty seconds works for most workloads.

Pull consumers use short polling. If no messages exist, the API returns immediately with an empty response. Your consumer must implement its own polling loop with appropriate backoff during quiet periods. Poll every 100 milliseconds when messages flow, back off to several seconds when the queue is empty.

When to use pull consumers​

Default to push consumers. They're simpler to operate, scale automatically, and integrate naturally with your Cloudflare architecture.

Choose pull consumers when Workers cannot handle your processing: the consumer needs capabilities Workers don't provide (arbitrary binaries, GPU access, massive memory, hours-long processing), must run in a specific location for compliance or latency, or integrates with existing infrastructure that isn't moving to Cloudflare. These are legitimate constraints, not preferences.

Pull consumers also suit controlled consumption rates. If your downstream handles only 10 requests per second and you need precise enforcement, pull consumers let you control exactly when and how fast you consume. Push consumers offer concurrency limits; pull consumers give complete control over consumption timing.

The hybrid pattern works well during migrations: produce to Cloudflare Queues from Workers, consume from existing infrastructure via pull, then migrate consumers to Workers once producers are stable. The queue decouples producer migration from consumer migration.

What pull consumers sacrifice​

Pull consumers give up automatic scaling. Push consumers scale with queue depth. Pull consumers scale only if you build that scaling yourself.

Pull consumers give up Cloudflare's retry orchestration. Push consumers use message.retry() with platform-managed delays. Pull consumers must implement their own retry logic or rely on the visibility timeout.

Pull consumers require API token management. Push consumers authenticate implicitly through bindings; pull consumers need tokens with queue read and write permissions, rotated and secured like any other credential.

If Workers can process your messages, push consumers' operational simplicity outweighs pull's theoretical flexibility.

Dead letter Queues​

Messages that fail repeatedly need somewhere to go. Without a dead letter queue, they're deleted after exhausting retries.

wrangler.jsonc: orders consumer and dead letter queue
{
"queues": {
"consumers": [
{
"queue": "orders",
"dead_letter_queue": "orders-dlq"
}
]
}
}

The dead letter queue is just another queue with its own consumer. That consumer might alert on-call engineers, log details for investigation, attempt processing with different logic, or queue for manual review.

Messages in a DLQ signal something is wrong with your consumer code, upstream data, or downstream dependencies. A growing DLQ indicates a problem that won't fix itself. Monitor DLQ depth and alert when messages accumulate; don't let it become a graveyard of ignored failures.

Failure modes​

Every queue failure mode traces back to the same root: the producer-consumer relationship is deliberately loose. That looseness enables scaling but prevents answering questions about individual messages.

Poison messages fail repeatedly while useful work waits behind them. Acknowledge completed items separately, bound retries, and configure a dead-letter queue for exhausted messages. Acknowledging a failed item removes it rather than sending it to the DLQ; preserve its payload and failure reason elsewhere first if it still needs investigation or correction.

Partial batch failure occurs when some messages succeed and others fail. Explicitly acknowledging completed messages prevents a later batch failure from replaying those acknowledged items. A crash between a side effect and acknowledgement can still repeat that effect, so combine acknowledgement with idempotent processing.

Consumer starvation happens when a slow message blocks batch processing. If your batch contains nine fast messages and one slow one, all nine complete quickly but the consumer invocation doesn't finish until the slow message completes. High concurrency helps (other invocations process other batches), but within a single batch, slow messages create head-of-line blocking.

Solutions depend on why messages are slow. If processing time varies predictably by message type, route slow types to a separate queue. If slowness is unpredictable, smaller batches reduce blocking impact. If slowness indicates a downstream problem, circuit breakers prevent one slow dependency from blocking all processing.

Backlog runaway during outages is expected since queues buffer work when consumers can't keep up. But unbounded accumulation indicates a fundamental mismatch between production rate and consumption capacity. Monitor backlog depth and alert on sustained growth. A backlog that grows during an outage and drains after recovery is healthy; indefinite growth indicates architectural problems.

Monitoring queue health​

Debugging queue problems requires visibility into the full pipeline.

Essential metrics: backlog depth (messages waiting), consumer error rate (failures per message), processing latency (time from send to acknowledgment), and DLQ inflow rate (messages failing to dead letter queue per period).

Backlog depth trending upward means insufficient consumer capacity: not enough concurrency, processing too slow, or too many retries. Error rate spikes indicate bad messages or downstream problems. Processing latency degradation suggests consumer slowness or batch configuration issues. DLQ inflow indicates persistent failures needing investigation.

Set alerts from the oldest accepted unfinished work, its business deadline and the retention time remaining. An hour to drain is already too late for a minute-sensitive task. Measure retry exhaustion and dead-letter inflow directly; a low aggregate error rate can hide a few permanently failing operations. A falling queue depth can also reflect expiry. Chapter 23 works through admission and drain capacity.

Cloudflare provides basic analytics through the dashboard. For production systems, supplement with your own logging: capture message types, processing duration, failure reasons, and retry counts.

Cost considerations​

Queues are available on both paid and free Workers plans. The free plan includes 10,000 operations per day across reads, writes, and deletes, with all features available including event subscriptions. The meaningful restriction is retention: free plan messages expire after 24 hours rather than the 14 days available on paid plans. For prototyping, learning, and low-volume production workloads, the free tier is usable rather than merely decorative.

Queue costs on paid plans scale with operations: writes, reads, and deletes. Each message typically incurs three operations minimum (write from producer, read by consumer, delete on acknowledgment). Retries add reads; DLQ routing adds writes and reads.

At low volumes, queue costs are negligible. Don't optimise.

At moderate volumes (millions of messages per day), costs become meaningful but are still dominated by what consumers do, not queue operations.

At high volumes, architecture matters more than configuration. Batch size doesn't reduce queue operations (10 messages still means 10 writes, 10 reads, 10 deletes), but batching reduces consumer invocations, lowering Workers costs if your consumer has significant per-invocation overhead. The optimisation that matters most is reducing message volume: aggregate events into single messages where semantics allow, use waitUntil() for fire-and-forget work that doesn't need retry guarantees, and question whether every event actually needs to be a separate message.

Scalability boundaries​

Queues have limits that matter for architecture decisions.

Message size caps at 128 KB. Many payloads exceed this. Store large data in R2 and send a reference. Design for indirection from the start if your payloads might grow.

Throughput caps at approximately 5,000 messages per second per queue. For higher throughput, shard across multiple queues. Most applications never approach this limit, but high-volume event streaming or IoT scenarios might require sharding.

Consumer wall time caps at 15 minutes; CPU time at 30 seconds by default, extendable to 5 minutes on paid plans. Most queue consumers complete in milliseconds or seconds. If yours routinely approaches these limits, you might need Containers, covered in Chapter 10.

When Queues aren't enough​

Signs to evaluate Workflows: implementing a durable multi-step sequence, resuming at a failed step, or coordinating waits and compensations. Keep Queues when independent work and buffering are the main requirements. Business status and duplicate protection may still need application records with either choice.

Choosing between Cloudflare Queues and hyperscaler alternatives​

Most background processing doesn't require advanced messaging features. If your requirements are "process this later, retry on failure, dead-letter if it keeps failing," Cloudflare Queues handles that with less operational complexity than SQS or Service Bus.

But some requirements do demand those features:

If ordered delivery by key and broker-side submission deduplication are requirements, evaluate SQS FIFO. A consumer can still repeat an external effect after a failure, so destination idempotency remains part of the design.

If you need messages larger than 128 KB without indirection, use Azure Service Bus (up to 100 MB on premium tiers). The R2 indirection pattern works, but native support is cleaner if large messages are common.

If you need sophisticated routing, dead-letter policies, or message sessions, Azure Service Bus offers capabilities Cloudflare Queues doesn't attempt.

If your requirements are simpler (independent tasks, at-least-once is fine, ordering doesn't matter, messages are small), Cloudflare Queues integrates natively with Workers, deploys as part of your Worker project, and operates with less configuration surface.

Hybrid approaches work. Use Cloudflare Queues for simple async tasks within your Cloudflare stack; use SQS or Service Bus for workloads needing advanced features. Design producers and consumers to be queue-agnostic where practical.

What comes next​

Queues change when work runs, but the consumer still needs a suitable runtime. Chapter 10 covers Containers for native binaries, larger working sets and processing that cannot fit a Worker invocation.