Skip to main content

Chapter 10: Containers: Beyond V8 Isolates

What do I do when Workers' constraints genuinely don't fit?


Containers extend the execution model when a workload needs a conventional runtime, more memory or a longer computation. The decision is whether to preserve that workload or restructure it to fit Workers. Cost matters, but so do startup latency, concurrency and the engineering effort of changing tested code.

Workers bills active CPU and requests. Containers also bills active CPU, with provisioned memory and disk charged while an instance runs. For I/O-heavy work, that makes instance utilisation and idle time part of the decision.

Concurrent requests can share those provisioned resources. Compare cost at the expected concurrency and sleep policy: a lightly used instance and a busy shared instance can have very different economics, even for the same I/O-heavy code.

Containers can preserve software whose runtime, memory or execution needs do not fit Workers. They introduce startup behaviour and provisioned resources alongside the existing Worker model. Choose them when that trade-off is better than restructuring the workload.

Containers: Durable Objects with bigger bodies​

Understanding Cloudflare's Container architecture reveals capabilities not obvious from documentation and explains why Containers don't need load balancers, service discovery, or orchestration complexity.

Every Container instance is backed by a Durable Object where the DO provides the brain (coordination, state, global addressability) and the Container provides the muscle (arbitrary runtimes, more memory, and longer computation). Requests flow through this chain:

How a request reaches a Container: Worker → DO → Container

Loading diagram…

This architecture inherits everything from Chapter 7 because when you derive an ID from a user identifier, every request worldwide for that user routes to the same Durable Object and therefore the same Container instance. The globally-unique routing that makes DOs powerful for coordination makes Containers globally addressable without load balancers.

The DO maintains state that survives container restarts using the same SQLite storage and strong consistency guarantees as any other Durable Object, so when a Container sleeps or crashes, the DO persists and provides whatever state the Container needs to resume when it wakes. The Container is ephemeral compute; the Durable Object is durable coordination.

Your sharding strategy (one container per user, per session, or shared pools) is expressed through how you derive DO IDs. The same decision you'd make for pure Durable Objects, with the same trade-offs. The routing code is trivial; the strategy decision is everything.

The real question: escape or restructure?​

Before choosing Containers, you must ask yourself whether you can restructure your workload to fit Workers instead.

Many workloads that seem to exceed Workers' limits can be refactored: image processing that loads entire images into memory can stream instead, data transformations that buffer complete datasets can chunk and process incrementally, and batch operations that accumulate results can write intermediate state to R2 or D1.

Refactoring takes engineering time and can add complexity. Containers preserves existing code at the cost of another execution model, startup behaviour and provisioned resources to manage. Compare those costs for the actual workload.

Estimate the refactoring cost, then compare measured running cost and latency over the expected workload lifetime. Restructure when it removes a recurring bottleneck at a reasonable maintenance cost. Keep the container when preserving tested code is worth more than the likely savings. Request count alone cannot establish that break-even point.

Restructure for Workers when:

  • Refactoring is straightforward (streaming instead of buffering, chunking instead of accumulating)
  • The workload runs frequently (cost savings compound rapidly)
  • Latency matters (Workers start in milliseconds; Containers in seconds)
  • You want architectural simplicity (one compute model, not two)

Accept Container complexity when:

  • Refactoring requires rewriting core logic or changing interfaces across services
  • You need a runtime Workers don't support (Go, Java, .NET, Rust without WASM)
  • The workload runs infrequently (Container overhead amortises over fewer invocations)
  • Existing containerised code works and rewriting provides no benefit

The worst outcome is reaching for Containers out of convenience, then discovering ongoing costs exceed what restructuring would have required, so do the analysis first.

Hard boundaries: when Containers are unavoidable​

Some constraints can't be engineered around, and these represent cases where Containers are necessary, not preferences but actual requirements.

Memory beyond 128 MB: If your workload must hold more than 128 MB in memory simultaneously and the problem requires it, Workers can't help. Machine learning models that don't fit in 128 MB, image processing of very large images where streaming isn't possible, and in-memory computation for specialised workloads all need Containers, which provide up to 12 GiB. If that's insufficient, use hyperscaler compute.

Non-JavaScript runtimes: Workers run JavaScript, TypeScript, Python, or WebAssembly, but Go, Rust without WASM compilation, Java, .NET, Ruby, and other runtimes require Containers. This includes existing containerised applications like internal tools, legacy services, and third-party software you can't or won't rewrite.

Long uninterrupted computation: HTTP Workers allow up to five minutes of CPU time per request. Some other triggers have different limits; dividing work into durable stages may fit. Use Containers or external compute when the job needs a longer uninterrupted execution, a native runtime or explicit CPU capacity.

For supported video resizing or frame and audio extraction, check the Media Transformations binding before adopting a container pipeline. Chapter 14 covers its limits and the R2 processing pattern.

Filesystem requirements: Workers offer bundled files and request-scoped temporary storage through their virtual filesystem. Containers provide up to 20 GB disk per instance for software needing a conventional filesystem or working data larger than Worker memory permits. That disk survives across requests within the instance, but it is ephemeral: keep durable data in external storage.

Hard boundaries: when Containers can't help​

Sometimes Workers don't fit but neither do Containers, so recognising these cases early saves considerable effort.

Resources beyond Container limits: Containers max out at 4 vCPU and 12 GiB memory, so if you need 32 GB RAM for a large ML model or 8 cores for parallel processing, use hyperscaler compute (EC2, GCE, Azure VMs) where larger instance types exist.

Inbound UDP, and inbound TCP only in private beta: a Container normally receives requests through a Worker or Durable Object over HTTP, not as a direct inbound socket. The exception is a private-beta TCP path: Spectrum forwards an inbound TCP socket to a Worker through a connect() handler, and the Worker passes it to a gRPC server in your Container, giving full-duplex gRPC in any language. It is TCP-only. Inbound UDP has no equivalent, so game servers expecting direct UDP, and IoT gateways speaking MQTT or CoAP over UDP, need traditional cloud infrastructure with load balancers and public IP addresses.

Nested containers: Docker-in-Docker is not possible, so CI/CD systems spawning containers and container-based testing frameworks cannot run on Cloudflare.

Choosing a routing strategy​

Your Durable Object ID strategy is your Container scaling strategy because per-user IDs mean per-user containers, pool IDs mean shared containers, and session IDs mean session-sticky containers. Each approach has different cost, isolation, and latency characteristics.

Per-user containers derive the DO ID from a user identifier, giving each user their own container instance and coordinating Durable Object with strong isolation and simplified state management since all requests route to the same instance. The cost is potentially many containers with low individual utilisation: 10,000 active users might mean 10,000 container instances, most sleeping. Sleep is free, but each request to a sleeping container incurs cold-start latency. This pattern fits when users need isolated resources, per-user state is substantial, or security requires separation.

Shared pools derive the DO ID from a pool identifier, routing multiple users to the same container instances so concentrated traffic keeps containers warm and reduces cold starts, with fewer containers meaning lower costs when traffic is moderate. The trade-off is noisy neighbours, where one user's expensive operation affects everyone sharing that container. This pattern fits stateless workloads or workloads where state lives in external storage and requires careful attention to timeouts and resource limits within your container application.

Session-sticky routing derives the DO ID from a session identifier so requests within a session route to the same container while different sessions may route to different instances. This provides in-session state without per-user cost and containers sleep when sessions end, so economics fall between per-user and pooled approaches. Watch session duration carefully: long sessions mean long-running containers while very short sessions mean frequent cold starts.

RequirementStrategyTrade-off
Strong user isolationPer-userHigher cost, more cold starts
Cost efficiencyShared poolNo isolation, noisy-neighbour risk
In-session stateSession-stickyModerate cost, session-duration sensitivity
Mixed workloadHybridComplexity, but optimises each case

For hybrid approaches, route different request types differently. Stateless API calls go to shared pools for cost efficiency; user-specific heavy processing goes to per-user containers for isolation. The routing logic lives in your Worker; the Container receives requests without knowing how they were routed.

Per-user containers separate application state, but they still rely on Cloudflare's isolation of shared compute and storage. Cloudflare's investigation of a remediated Containers vulnerability demonstrated that residual disk blocks could cross tenant boundaries; Sandboxes were affected too. Cloudflare reports completed fleet-wide cleanup, no required customer configuration changes, and no evidence of malicious exploitation within its available telemetry. The architectural lesson is to evaluate storage reuse alongside runtime isolation. An ephemeral disk describes how long your application can rely on its contents; it does not by itself establish how safely the underlying blocks are reassigned.

Instance sizing​

Choose the smallest predefined or custom instance that supports the workload at its intended concurrency. Cloudflare's predefined range runs from lite at 256 MiB and 1/16 vCPU to standard-4 at 12 GiB and 4 vCPU. Custom types have their own resource ratios and a minimum of one vCPU.

Measure memory, CPU and temporary-disk demand separately. Increasing concurrency can exhaust memory even when one request fits comfortably. More CPU helps a compute bottleneck only when the application can use it; a single-threaded process may need different parallelism rather than a larger instance.

An out-of-memory kill, sustained CPU queueing and a full disk are different failures. Diagnose which resource is exhausted before resizing. Clean up temporary files, bound concurrent jobs and rule out leaks, then repeat the representative load test.

The cost model in practice​

Containers charge for active CPU usage and for provisioned memory and disk while the instance runs. Waiting for an API does not consume CPU charges by itself, but the allocated memory and disk remain billable until the container sleeps. Workers, Durable Object coordination and network egress can add to the bill.

For Workers, estimate CPU per request and request volume. For Containers, also estimate how long instances remain awake, how many requests overlap, and how much memory each concurrent request needs. Multiplying each request's wall time by its count overstates running time when one instance handles overlapping requests; ignoring idle time understates it.

The routing strategy is part of this calculation. Shared pools can improve utilisation, while per-user instances buy separation with more provisioned capacity and more cold starts. Longer sleep timeouts trade an idle cost floor for response latency.

Compare a representative day, including quiet periods, bursts and failures. The right question is whether preserving the container's runtime and code saves enough engineering work to justify its measured operating cost.

Designing for cold starts​

You cannot eliminate cold starts; you can only choose who experiences them.

A request to a sleeping container first triggers startup through its Durable Object. Cloudflare describes cold starts often in the one-to-three-second range, with image size and application initialisation affecting the result. Measure your image and decide whether that delay fits the interactive path.

Cold Start Reality

Startup latency depends on the image and its initialisation. If the interactive deadline is shorter than measured cold startup, keep capacity warm or acknowledge the job and complete it asynchronously. Include the cost of warm capacity in that decision.

The solution is architectural rather than optimisation: route interactive traffic to Workers and heavy computation to Containers so the user gets millisecond responses while the Container wakes in the background.

Design for graceful degradation. When a request requires Container processing and the Container is cold, return immediately with a "processing" acknowledgement. Deliver results via webhook, polling, or WebSocket notification. The user experiences fast feedback; the Container processes without blocking them.

Use a durable handoff. Validate and authorise in the Worker, then hand accepted background work to a Queue or Workflow that invokes the Container. Return a job identifier once the handoff succeeds. Starting an unawaited container request and returning a response does not establish durable acceptance.

Pre-warm on predictable traffic. If you know traffic is coming (user logged in, batch job scheduled, webhook expected), send a lightweight request to wake the container before the real work arrives. Shift cold-start latency from user-facing requests to background preparation.

Extend sleep timeout selectively. Configure longer sleepAfter values for containers with frequent but irregular traffic. A container sleeping after 30 minutes instead of 10 stays warm for more requests, at the cost of paying for idle time. This trade-off makes sense when cold-start latency matters more than cost.

Measure and reduce startup work. Minimise image size and defer initialisation that the first request does not need. Re-measure cold startup before choosing a warm-pool or asynchronous design; an assumed five-second floor is not a platform guarantee.

For batch processing, scheduled jobs, and asynchronous workflows, cold starts don't matter because a workflow step taking 30 seconds of processing doesn't suffer noticeably from 5 seconds of startup, so route interactive traffic away from cold-start-sensitive paths.

Observability: what container failures actually look like​

Standard observability advice applies to Containers as to any compute, though what matters here is understanding Container-specific failure patterns and what they indicate.

Error rate spikes in the Container but not the DO indicate a Container application bug, not routing or coordination. The DO successfully received and forwarded requests; the Container failed to process them. Debug your application code, not your Cloudflare configuration.

Error rate spikes in both DO and Container simultaneously suggest resource exhaustion or infrastructure issues. The DO might be failing to start the Container, or the Container crashing on startup. Check resource limits, image validity, and Cloudflare status.

Latency increases without error rate changes have two common causes: if latency correlates with traffic, your Container is compute-bound and requests are queuing, and if latency increases randomly regardless of traffic, you're seeing cold starts. Distinguish by correlating request timestamps with container start events.

Intermittent failures under load usually indicate memory exhaustion because a Container running fine at low traffic OOMs when concurrent requests multiply memory usage. If memory climbs towards limits before failures, increase instance size or reduce concurrency.

Requests succeed but return wrong results require tracing the data as well as the transport. Check stale state, request correlation, dependency responses and application logic. A successful HTTP status identifies neither the faulty layer nor the correctness of the result.

Trace the request across the Worker, Durable Object and Container, then find the first divergence from expected behaviour. The layer reporting an error may be relaying a failure from another layer. Correlation IDs and lifecycle events make that distinction visible.

When logs and metrics are not enough, you can SSH into a running Container instance via Wrangler for live debugging. This is enabled by default and exposes no public ports: the SSH service is reachable only through wrangler containers ssh <INSTANCE_ID>, which authenticates against your Cloudflare account. You still add an ssh-ed25519 public key to the container's authorized_keys before anyone can connect, so the transport being on does not by itself grant access; set ssh.enabled to false to turn it off entirely. Find running instance IDs with wrangler containers instances <APPLICATION>, and run a single command without an interactive shell by appending it after --. This is useful for inspecting running processes, checking filesystem state, or executing one-off diagnostic commands, but treat it as a debugging tool rather than an operational workflow. If you find yourself SSHing routinely, that is a signal your logging and monitoring need improvement.

SSH is for humans; exec() is for code. From a class extending Container, or from another Durable Object holding the container through this.ctx.container, exec() starts a process inside the running container and streams its standard input and output back to you, with exit codes and per-process signals. This turns a container from a single HTTP server into something your Worker can drive directly: run a one-off database migration, invoke a CLI tool the image already ships, or spawn a helper process alongside the main one and coordinate several commands in a single round-trip. The command is an array that starts an executable directly, with no implicit shell, so reach for an explicit shell only when you need pipes, redirects, or variable expansion. The architectural gain is that orchestration logic stays in the Worker or Durable Object where you already handle routing and lifecycle, rather than being baked into an entrypoint script inside the image.

When to choose hyperscalers instead​

Cloudflare Containers optimise for global distribution and tight integration with Workers and Durable Objects, while hyperscaler alternatives optimise for different things that sometimes matter more.

Choose hyperscalers when you need deep VPC integration: Cloudflare Containers communicate via HTTP through Workers, and while Workers VPC Services (currently in beta) provides secure connectivity to private resources through Cloudflare Tunnel, hyperscaler containers offer native VPC attachment without intermediate layers. For workloads requiring extensive private network access or complex networking topologies, hyperscaler containers with native VPC integration remain simpler.

Choose hyperscalers when you need larger instances: Cloudflare Containers reaches 4 vCPU and 12 GiB memory. AWS Fargate supports larger tasks, including 16 vCPU and 120 GB memory. Compare the required instance envelope before considering operational convenience.

Choose hyperscalers when you need deep integration with hyperscaler-specific services: If your architecture depends on SQS, DynamoDB, BigQuery, or Cosmos DB, running containers on the same platform simplifies authentication, reduces latency, and consolidates billing.

Choose Cloudflare when global distribution matters more than raw instance size: Cloudflare Containers deploy globally by default while hyperscaler containers require explicit multi-region configuration. If your users are worldwide and latency matters, Cloudflare's automatic distribution is valuable.

Choose Cloudflare when your architecture already uses Workers and Durable Objects: The DO coordination model backing Containers is the same model you're already using, so adding Containers extends your existing architecture whereas adopting hyperscaler containers introduces a separate system with separate deployment, monitoring, and operational patterns.

Choose Cloudflare when the DO coordination model simplifies your design: Container routing can reuse the entity ownership, persistent state and lifecycle logic already expressed in Durable Objects. Include DO requests, duration and storage in the cost model; the integration does not make coordination free.

Deciding factorCloudflareHyperscaler
Global distributionAutomaticManual multi-region
Maximum resources4 vCPU, 12 GiBLarger profiles; limits vary by service
Private networkingVPC Services (beta)Native VPC integration
Coordination modelDurable ObjectsBuild your own
Ecosystem integrationWorkers, R2, D1Full hyperscaler suite

What comes next​

Chapter 11 turns to live audio and video. Durable Objects can coordinate a room and Containers can process media, but transporting interactive streams calls for WebRTC infrastructure. Cloudflare Realtime supplies that layer.