Chapter 3: Workers: The Core Compute Primitive
How do Workers actually work, and how do I use them effectively?
Chapter 1 introduced the V8 isolate model. This chapter makes it concrete: the request lifecycle, the hard constraints and their implications, and the patterns that separate effective Workers code from code that fights the platform.
Understanding Workers deeply matters because every other Cloudflare service builds on them. Durable Objects are Workers with persistent state. Workflows are Workers with durable execution. Queues deliver messages to Workers. Containers route through Workers. Master this primitive and you understand the foundation of everything else.
The execution model
Workers execute JavaScript in V8 isolates, the same sandboxing technology that isolates browser tabs from each other within the browser. This isn't a virtual machine or container; it's a lightweight sandbox within an existing runtime. Sharing the runtime avoids starting a separate process for each execution environment. Application initialisation and backend calls still contribute to the first response.
Sharing a runtime across isolates
Loading diagram…
Single-threaded by design
Each request executes in a single thread without parallelism. You cannot parallelise computation within a request: no worker threads API, no shared memory between concurrent operations. For parallel processing, you make multiple subrequests executing in potentially different isolates, or use Promise.all for concurrent I/O operations.
The design implication: compute-heavy operations benefiting from parallelism need decomposition across multiple Workers invocations or delegation to specialised services. Workers excel at orchestration and I/O coordination, not raw parallel computation.
Isolate lifetime and state
An isolate running your Worker may persist between requests as an optimisation for performance. This is how Cloudflare achieves near-instant response times for warm isolates. But you cannot rely on this persistence as a guarantee. Global variables set during one request may or may not exist during the next.
Always treat every request as if it's the first to a fresh isolate for correctness. That persistence might not exist is a fact; that it sometimes does is a performance optimisation, not a guarantee you can rely on.
This reality shapes how you use global scope strategically. Expensive initialisation producing immutable results (parsing configuration, compiling regular expressions, establishing reusable structures) belongs in global scope where it might persist between requests. State that matters (user sessions, request counts, accumulated data) belongs in Durable Objects or external storage where persistence is guaranteed.
The successful pattern emerges: use global variables for caches where loss is acceptable and reconstruction is cheap. A cache miss means slightly slower response, not incorrect behaviour. If losing data would cause errors or require complex recovery logic, it doesn't belong in isolate memory.
The request lifecycle
Workers respond to events through handlers: HTTP requests, scheduled events, queue messages and other triggers. The handler receives the event and its configured bindings; HTTP handlers also receive an execution context for work such as waitUntil().
Request lifetime and isolate lifetime are different. The runtime may reuse an isolate after a request ends, but that does not guarantee completion of work started by the earlier request. Await operations needed for the response, and explicitly register eligible background work. A streamed response can also keep request work active after the handler returns.
Background work with waitUntil()
waitUntil() allows work to continue after an HTTP response without delaying that response, for up to 30 seconds after the response completes or the client disconnects. It is a bounded opportunity to finish work, not durable job execution.
Use waitUntil() for operations that don't affect the response and where clients shouldn't wait: analytics tracking, logging to external systems, cache warming after serving stale content, cleanup operations. The client gets their response quickly; background work completes afterward.
Don't use waitUntil() for operations where failure must change the outcome to the user. If the background operation fails, you've already sent a success response. The client believes their action succeeded, but your system knows it didn't. If failure should change the response, await the operation directly rather than deferring it.
waitUntil() does not increase the CPU allowance. Its failures cannot change a response already sent. Use Queues or Workflows when completion must survive beyond the request context.
Resource constraints
Workers operate within hard constraints that shape application design. These aren't suggestions or soft limits that increase cost. They're boundaries that cause failures when exceeded.
Memory: 128 MB per isolate
Each isolate has 128 MB of memory total, encompassing the JavaScript heap, WebAssembly linear memory, and buffers. This memory is shared across concurrent requests handled by the same isolate, though in practice most isolates handle one request at a time.
128 MB sounds restrictive until you realise most web requests don't need to hold entire files in memory; they need to stream bytes from one place to another. The limit constrains buffering, not capability.
What fits comfortably: typical request/response handling, JSON parsing of documents up to a few megabytes, reasonable in-memory data structures, most API gateway patterns.
What doesn't fit within constraints: buffering large file uploads before processing (stream directly to R2), loading datasets for in-memory analysis (use D1 or process incrementally), high-resolution image manipulation without streaming (use R2's image transformations), large ML models (use Workers AI).
The key strategy is streaming. When handling large request or response bodies, pipe directly to destination rather than accumulating in memory. A Worker can proxy gigabytes through R2 while staying under the memory limit because it never holds more than a buffer's worth at once.
When you need more memory because the workload requires holding substantial data in memory simultaneously rather than streaming it through, Workers aren't the right tool for that workload. Containers offer up to 12 GiB at the cost of longer cold starts. Recognise this constraint early rather than fighting it with overly complex streaming gymnastics.
CPU time: the billing boundary
CPU time defaults to a 30-second limit on paid plans (extendable to 5 minutes). The free tier allows only 10 milliseconds, enough for simple request routing but not substantial computation.
Understanding what consumes CPU time matters profoundly for staying within limits and controlling costs. Active JavaScript execution, JSON serialisation, cryptographic operations, WebAssembly computation, and regular expression evaluation all consume CPU time. Waiting does not. A fetch() to an external API, a D1 query, a KV read: while your code awaits these operations, CPU time doesn't accumulate.
This distinction is fundamental to Workers economics and receives deeper treatment shortly. The configured limit applies to computation time, not elapsed time. A Worker making dozens of external calls over several seconds of wall time might consume only milliseconds of CPU time.
Subrequest limits
Paid Workers default to 10,000 subrequests per invocation, configurable up to ten million with limits.subrequests. Free plans allow 50 external subrequests and 1,000 calls to Cloudflare services.
Set the ceiling from the expected work per invocation and a bounded allowance for retries. A high limit can accommodate legitimate fan-out, but it can also let a loop amplify a small request into substantial downstream work. Lower limits are useful containment when inputs or loaded code are less trusted.
When constraints don't fit
Workers have a 128 MB per-isolate memory limit. Paid HTTP invocations can configure CPU time up to five minutes; other triggers have their own limits. Paid subrequests default to 10,000 and can be configured up to ten million. Check the trigger and plan before treating a limit as universal.
If your workload routinely needs more than the isolate's 128 MB and streaming won't help because you need data in memory simultaneously, you need Containers.
If a job exceeds its trigger's CPU allowance, split the work or choose a longer-running compute boundary. Paid Cron Triggers scheduled hourly or less often allow fifteen minutes of CPU; more frequent schedules allow thirty seconds. Containers can fit processes whose memory, runtime or lifecycle needs remain outside Workers' limits.
Before raising a subrequest ceiling, account for fan-out and retry amplification. Increase it for bounded legitimate work, and keep it low enough to contain a loop. Memory, CPU and subrequest budgets protect different resources; none replaces a workload design.
Trace the security path
An incoming request and a subrequest from your Worker do not necessarily pass through the same security checks. Behaviour depends on the route, target zone and feature configuration. For example, same-zone subrequests can affect rate limiting under some settings.
Rules can distinguish Worker-originated subrequests with cf.worker.upstream_zone, which records the source Worker's zone and is empty for a direct visitor request. Use it to scope deliberate exceptions; a request from one of your Workers still needs the appropriate user authorisation.
Draw the actual request path and identify where authentication, authorisation and rate limits run. A service binding establishes the caller's capability to reach a service; it does not establish that the original user may perform every operation that service exposes. Carry and validate the application context the operation requires.
Test a permitted request and a forbidden one through each entry path, including internal calls. Chapter 22 owns the deployment and security controls around those boundaries.
CPU time versus wall time
The separation of CPU time from wall time isn't billing detail. It's the architectural insight that determines whether Workers make economic sense for your workload.
Lambda's duration meter includes I/O wait; Workers' CPU meter excludes it. For I/O-heavy workloads, price both the orchestration and the services doing the work.
The economics of waiting
Consider a typical API endpoint authenticating a request, querying a database twice, calling an external service, and returning formatted results:
- JWT validation: 2ms CPU time
- D1 query 1: 5ms waiting, 3ms CPU for parsing
- D1 query 2: 5ms waiting, 2ms CPU for parsing
- External API call: 200ms waiting, 3ms CPU for processing
- Response serialisation: 5ms CPU time
- Total wall time: ~225ms
- Total CPU time: ~15ms
On Lambda with 512 MB allocated, you pay for 225ms x 0.5 GB = 112.5 GB-milliseconds. The function sat idle for 93% of that time, but you paid for all of it.
On Workers, you pay for 15ms of CPU time. The 210ms of waiting adds no Worker CPU charge.
For I/O-heavy workloads (most API endpoints, most web applications, most integration layers), this difference compounds dramatically. Workers charge for computation. Lambda charges for existence.
The inverse is also true. For compute-heavy workloads (data transformation, complex calculations, CPU-bound processing), billing models converge or even favour Lambda's larger memory allowances and longer time limits. Workers optimise for the common case of web workloads, not compute-intensive batch processing.
Designing for free I/O
The cost model rewards designs maximising I/O relative to computation.
Move computation to specialised services. Need image resizing? R2's image transformations run on Cloudflare's infrastructure, not yours. Need ML inference? Workers AI handles the heavy lifting. Don't reimplement computationally intensive operations when bindings provide them.
Prefer focused queries over local filtering. A database query returning exactly what you need consumes I/O time (free waiting) and minimal parsing (cheap computation). Fetching excessive data and filtering locally consumes the same I/O time plus significant computation for filtering. Push predicates to the data source.
Cache computed results aggressively. A KV read is I/O; the computation producing the cached value already happened. For values changing infrequently relative to read frequency, computing once and caching beats recomputing every request.
Use streaming for large payloads. Parsing a 10 MB JSON document consumes substantial CPU time. Streaming the same data through without parsing consumes almost none. When you don't need the full parsed structure (when you're proxying, storing, or passing data through) don't parse it.
Common CPU time sinks
Certain patterns consume more CPU time than developers expect.
Repeated JSON parsing/serialisation, string concatenation in loops, and heavy validation libraries accumulate CPU time quickly. Parse JSON once at the boundary. For hot paths, profile before adding heavyweight validation frameworks.
Repeated JSON operations accumulate quickly. Parse once at the boundary; pass the object through your code. Don't serialise and deserialise repeatedly as data moves between functions.
String manipulation in loops scales poorly. Concatenation in loops, repeated substring extraction, and applying regular expressions to the same content multiple times can be problematic. Each operation is cheap, but thousands add up.
Validation libraries vary enormously in efficiency. Schema validation frameworks designed for developer ergonomics may do far more work than manual validation for simple cases. Measure whether library overhead is acceptable.
Node.js dependencies may include abstractions costing CPU time without providing value in Workers. Workers-optimised alternatives often exist for common tasks.
Profiling and measurement
The Cloudflare dashboard shows CPU time distribution across requests. For development, measure synchronous code paths with timing calls. For operations without I/O waits, wall time approximates CPU time.
If P95 CPU time significantly exceeds P50, investigate outlier requests. They often hit edge cases: regular expressions catastrophically backtracking on certain inputs, loops iterating far more than expected, or unusually large payloads requiring more parsing.
Worker placement
Workers run globally by default: your code executes in whichever Cloudflare location is closest to the user. For applications serving static content or performing computation not depending on external data, this is optimal. Users everywhere get low latency.
But many Workers make backend calls: queries to databases, API calls, access to services living in specific regions. A Worker running in Sydney making ten calls to a database in Virginia adds 200ms of round-trip latency per call. The Worker is close to the user but far from the data.
Smart Placement and explicit hints optimise execution location for performance. Residency is a separate requirement with a different scope of controls.
Smart placement
Smart Placement analyses your Worker's traffic patterns, specifically where subrequests go. If a Worker consistently calls backends in a particular region, Smart Placement runs the Worker closer to those backends rather than closer to the user.
{
"placement": {
"mode": "smart"
}
}
The trade-off is explicit: users further from the backend region experience higher latency reaching the Worker, but the Worker experiences lower latency reaching the backend. For workloads making many backend calls, total latency decreases despite the longer initial hop.
Enable Smart Placement when your Worker makes multiple calls to backends in a specific region and those calls dominate total latency. An API endpoint making five database queries to a PostgreSQL instance in eu-west-1 benefits from running in Europe rather than wherever the user happens to be.
Don't enable Smart Placement for Workers primarily serving cached content, performing computation without external calls, or calling globally distributed services without a single location.
Explicit placement hints
Use a region hint when the dominant backend lives in a known cloud region. Use a host or hostname probe when its location is better identified by an endpoint. These options place the Worker near the target; they are not exact-machine or residency controls.
Endpoint probes need a single-homed target. They do not locate anycast or replicated services reliably, because different probes may reach different instances. Do not use a distributed front door as a placement hint for the database behind it.
A hint is useful when one backend dominates the critical path. If several dependencies compete, compare the complete request time rather than the latency of one call. Keep the configuration only if it improves the workload you actually serve.
Regional execution and residency
Regional Services can restrict processing for a configured custom hostname. This is distinct from a performance hint or a Durable Object jurisdiction. Its scope does not extend to Queue and Cron triggers or automatically constrain outgoing subrequests; Worker code and secrets remain globally deployed. Trace the complete path before treating regional entry as regional processing.
Chapter 22 examines those boundaries across storage, execution and dependencies. A residency requirement should determine the permitted path before placement optimisation begins.
Service bindings and composition
Workers can call other Workers through service bindings, enabling separate deployment and ownership with low transport overhead. Interface design and failure handling still cross that boundary.
Use service boundaries for independent ownership, deployment or capability control. Low transport overhead makes those boundaries cheaper; it does not justify splitting every function into a Worker.
How service bindings work
A service binding gives one Worker access to another through configuration. By default, both execute on the same server thread, avoiding a public HTTP round trip. Placement can move a callee elsewhere when that shortens its backend path.
This reduces transport overhead without removing the service boundary. Calls can fail, callees can exhaust limits, and independently deployed code can break an interface. Keep deadlines, error handling and compatibility rules appropriate to the operation. Measure the complete chain before treating decomposition as free.
Fetch forwarding versus RPC
Service bindings support two communication styles. Fetch forwarding passes an entire Request to the target Worker, which processes it and returns a Response. RPC-style bindings expose typed methods the calling Worker invokes directly.
Fetch forwarding suits scenarios where you're proxying requests or the target Worker needs full request context: headers, method, body, URL. An authentication Worker examining cookies and setting response headers benefits from receiving the complete Request.
RPC suits scenarios calling specific functionality with specific parameters. A rate-limiting Worker needing a user ID and action name doesn't need full request context, just those two values, returning a boolean. RPC provides type safety, clearer interfaces, and avoids the ceremony of constructing Request objects for simple operations.
Use fetch forwarding when the target Worker processes requests as requests. Use RPC when the target Worker provides functionality as functions.
When to split Workers
Decomposing into multiple Workers involves trade-offs that near-zero latency doesn't eliminate.
Start with a single Worker unless you have a specific reason to split. Service bindings make splitting cheap but not free. Each call adds latency even if sub-millisecond, and each additional Worker adds deployment complexity, versioning concerns, and cognitive overhead.
Split when deployment independence matters. If one part changes frequently while another is stable, deploying separately avoids unnecessary risk to the stable component.
Split when resource profiles differ significantly. A CPU-intensive Worker benefits from different optimisation strategies than an I/O-heavy Worker.
Split when team boundaries align with service boundaries. Different teams owning different functionality benefit from separate Workers with clearer ownership and reduced coordination overhead.
Split when sharing functionality across multiple entry points. A common authentication Worker called by several API Workers is cleaner than duplicating authentication logic: one place to update security policies, one place to fix vulnerabilities.
Don't split because microservices are architecturally fashionable. A design requiring twenty service binding calls per request performs worse than a monolith even with sub-millisecond latency, and complexity cost is substantial. Measure before distributing.
Error handling in edge systems
Error handling in Workers differs from traditional servers in ways affecting both implementation and architecture.
Named failure modes
Edge systems exhibit failure patterns worth naming explicitly.
Orphaned background work. A waitUntil() promise fails after the response was sent. The client believes success; your system recorded failure. Data inconsistency results, difficult to diagnose because no error surfaced to the user.
Regional backend outage. An external API or database fails in one region but remains healthy elsewhere. Aggregate error rate is 5%, but São Paulo users see 80% failures while London users see none. Aggregate metrics hide severity for affected users.
Isolate state assumption. Code assumes a global variable persists between requests. Works in testing where the same isolate handles sequential requests, then fails in production when requests hit fresh isolates.
Subrequest exhaustion. A fan-out pattern works at normal load, but unusual requests trigger more subrequests than expected, exceeding the configured limit and failing abruptly. Even with configurable limits (up to 10 million on paid plans), setting an appropriate ceiling and monitoring actual usage prevents surprises.
Timeout cascade. A slow upstream exhausts wall time budget. Other healthy upstreams in the same request path never get called because time ran out waiting for the slow one.
Cold cache stampede. A cached value expires. Many concurrent requests miss simultaneously and all attempt to regenerate the value, overwhelming the backend.
Naming these patterns gives your team vocabulary for design discussions and post-incident analysis. "We hit orphaned background work" communicates more precisely than "something went wrong with the async stuff."
The ephemeral context
Workers execute in ephemeral isolates. You cannot maintain in-memory circuit breaker state across requests since each request might execute in a different isolate with no knowledge of previous failures.
For circuit breaker behaviour, state must live outside the isolate. Durable Objects can track failure counts and provide consistent circuit breaker decisions. KV can store failure state with some eventual consistency tolerance. The pattern changes from "maintain state in memory" to "coordinate through external storage."
Global distribution implications
The same Worker code runs in over 300 locations. A regional backend outage affects requests routed to that region but not requests elsewhere. An external API having problems in Asia might cause failures for Asian users while European users experience normal operation.
Error rates can be geographically localised in ways aggregate metrics obscure. A 2% global error rate might represent 50% failures in one region and 0% everywhere else. Alerting and investigation should account for this distribution.
Retry economics
Waiting two seconds for a timeout adds little Worker CPU cost, but the attempted operation may still incur database, inference or external API charges. Repeating it can repeat those costs even when the Worker itself is inexpensive.
Retry only when the operation’s semantics permit it, use bounded backoff and account for retries in downstream clients and orchestration layers. Chapter 23 explains how those layers can amplify one user request.
Timeout decisions
Workers don't impose timeouts on subrequests by default. A slow upstream can consume your entire wall time budget waiting. Implement explicit timeouts using AbortController, setting limits appropriate to SLA requirements and expected upstream performance.
Five seconds is reasonable for most external API calls. Database queries through Hyperdrive might tolerate longer; calls to services with strict latency SLAs might require shorter. Choose timeouts reflecting "how long am I willing to wait before giving up" rather than "how long does this usually take."
Error response strategy
Handle errors at the handler level to return controlled responses. Unhandled exceptions produce generic 500 responses with no useful information. Wrap handler logic, catch exceptions, log with sufficient context for investigation, return responses helping clients understand what happened without leaking implementation details.
The categories that matter: client errors (4xx) for invalid requests the client can fix, server errors (5xx) for problems the client cannot fix. A missing resource is 404, not 500; a malformed request is 400, not 404. Precise status codes help clients respond appropriately and help you categorise errors in monitoring.
Language support
Workers execute JavaScript natively, with additional options for teams with different requirements or existing codebases.
JavaScript and TypeScript
JavaScript executes directly in the V8 runtime with access to modern language features and standard Web APIs. TypeScript compiles to JavaScript through Cloudflare's tooling with no additional configuration. For most new projects, TypeScript provides type safety benefits without meaningful downsides.
You have access to Web APIs you'd use in browsers: fetch, Request, Response, Headers, URL, crypto, TextEncoder, TextDecoder, streams. Code written against these APIs is portable across Workers, browsers, Deno, and other Web-API-compatible environments.
Node.js compatibility
Workers aren't Node.js, but a compatibility layer provides many Node.js APIs. It is enabled by default for compatibility dates of 2026-08-04 or later; Workers pinned to earlier dates can opt in with nodejs_compat. Pinning a compatibility date remains an architectural control: test a date upgrade against your dependencies before deploying it.
The filesystem distinction matters when assessing npm packages. node:fs exposes a virtual filesystem with read-only bundled files and a writable, request-scoped /tmp. Temporary data consumes memory; it provides neither persistent disk nor access to the host. Native binary addons and child processes remain outside the Workers model. Check what a dependency needs from Node.js rather than rejecting it merely because it imports a Node module.
Worker code has a 64 MiB uncompressed deployment limit on both Free and Paid plans. This gives dependency-heavy applications room to deploy, but the 128 MB isolate memory limit and startup constraints still apply. A bundle that fits at upload can still be unsuitable at runtime. Prefer Web APIs for new code where they meet the need, and use Node.js compatibility to preserve useful dependencies.
WebAssembly
WebAssembly allows running code compiled from Rust, C, C++, Go, and other languages: computationally intensive algorithms where JavaScript performance isn't adequate, existing compiled libraries you can't or won't rewrite, and code sharing across WASM-compatible environments.
WASM executes within the same constraints as JavaScript: 128 MB memory limit and CPU time limits apply identically. WebAssembly provides access to efficient compiled code, not resource limit evasion.
Choose WebAssembly when you have existing compiled code prohibitively expensive to rewrite, or when profiling shows JavaScript performance is inadequate. Don't choose WebAssembly because it seems faster; for I/O-heavy workloads, execution speed is rarely the bottleneck.
Python Workers
Python Workers are generally available and execute Python code at the edge through Pyodide, a Python runtime compiled to WebAssembly. This isn't a compatibility layer or transpilation. It's actual Python executing in the Workers runtime with access to the standard library and a growing ecosystem of packages.
How Python Workers execute
When you deploy a Python Worker, Cloudflare uploads your Python code and packages specified in pyproject.toml. The runtime creates a V8 isolate, injects Pyodide, scans your code for imports, executes them, then captures a memory snapshot of this initialised state.
This snapshot is key to Python Workers' performance. Cold starts in Python would normally require loading the runtime, importing packages, and executing top-level code on every new isolate. With memory snapshots, expensive initialisation happens once at deploy time. Subsequent cold starts load the snapshot directly, bypassing import cost.
Snapshotting removes repeated import work, but Python startup still depends on the runtime and packages your application loads. Benchmark cold and warm requests with your actual dependency set before choosing Python for a latency-sensitive path; JavaScript's isolate startup figures are not a Python latency guarantee.
Package compatibility
Python Workers use pywrangler with uv for dependency management and deployment. Package support covers pure Python packages from PyPI, compatible PyEmscripten wheels, and Pyodide's package set. Native extensions need a compatible WebAssembly build.
Test the dependency set before committing to the runtime, including optional imports and code paths used only under load. A framework starting successfully does not establish that its database driver, image library or background extension will work.
Accessing bindings from Python
Python Workers access Cloudflare services through the same binding model as JavaScript Workers, with Python-native syntax. This entry point assumes the python_workers compatibility flag, configured MY_KV, DB and BUCKET bindings, and an existing users table:
from workers import WorkerEntrypoint, Response
class Default(WorkerEntrypoint):
async def fetch(self, request):
# Access KV
value = await self.env.MY_KV.get("key")
# Fixed fixture ID for this binding example
user_id = "demo-user"
result = await self.env.DB.prepare(
"SELECT * FROM users WHERE id = ?"
).bind(user_id).first()
# Access R2
stored_object = await self.env.BUCKET.get("file.txt")
return Response.json({
"value": value, "user": result,
"file_found": stored_object is not None
})
Python's binding wrappers convert ordinary Python values for calls into Cloudflare services and return Python values for common results. Pyodide's foreign function interface remains underneath, but application code can use familiar Python collections rather than handling JavaScript proxies for every binding call.
Cron triggers, queue consumers, and other handler types work identically. Define the appropriate method on your WorkerEntrypoint class; the runtime invokes it when the trigger fires.
When to choose Python
Python Workers make sense when your team's expertise is Python and rewriting in JavaScript would slow development significantly. They're appropriate for I/O-heavy workloads where Python's execution speed isn't the bottleneck: API orchestration, data transformation, AI inference coordination. ASGI and WSGI adapters support frameworks such as FastAPI, Django, and Flask, while Hyperdrive connects Python Workers to PostgreSQL and MySQL. An existing Python application can therefore retain its framework and database when its dependencies fit Pyodide and its workload fits Workers' limits. Test the database driver and concurrency model as well as the framework; support for a framework does not imply support for every extension it can load.
Choose Python Workers when retaining Python code and expertise saves more work than adapting its dependencies creates. Validate package compatibility, memory use, and startup latency with a representative workload before committing.
Python Workers are inappropriate when raw performance matters (JavaScript and WebAssembly execute faster for CPU-intensive computation) or when you need native extensions without a compatible WebAssembly build.
The mental model is simple: Python Workers let Python developers build on Cloudflare without learning JavaScript. They're not faster or more capable than JavaScript Workers; they're an alternative for teams where Python is the better choice for human reasons.
Comparing to hyperscaler serverless
Coming from Lambda, Azure Functions, or Cloud Functions, certain differences affect how you design and operate applications.
Startup and application latency
Workers avoids starting a separate process for each fresh isolate. That reduces runtime startup overhead, but the response still includes application initialisation and the path to its dependencies. Compare cold and warm requests with representative code; the Python case earlier in this chapter makes that distinction especially clear.
Memory: a hard boundary
Lambda offers up to 10 GB of memory per function. Workers offer 128 MB. This is a hard boundary, not a tuning parameter.
The response isn't to avoid Workers but to understand which workloads fit each model. Request/response handling, API gateways, edge logic, and coordination tasks fit Workers' memory model. Data transformation, large file processing, and memory-intensive computation fit Lambda or Workers Containers.
The global model
Lambda functions deploy to regions. You choose a region, your function runs there, users far from that region experience latency. Multi-region deployment requires explicit configuration, additional infrastructure, and careful data synchronisation.
Workers deploy globally by default. You don't choose regions; your code runs everywhere. Users close to any Cloudflare location get low latency. The operational complexity of multi-region deployment doesn't exist because deployment is inherently global.
On Lambda, going global is a project. On Workers, you're global from the first deployment. Questions shift from "should we deploy to additional regions" to "are there reasons to constrain where we run."
| Aspect | Workers | Lambda |
|---|---|---|
| Initialisation | Shared V8 runtime plus application setup | Execution environment plus application setup |
| Maximum memory | 128 MB | 10 GB |
| HTTP execution limit | Up to 5 minutes CPU | Up to 15 minutes elapsed |
| Deployment scope | Global (automatic) | Regional |
| Billing model | CPU time | GB-seconds (wall time) |
| Multi-region | Default | Additional complexity |
Workers optimise for the common case of web workloads (I/O-heavy, latency-sensitive). Lambda optimises for flexibility (arbitrary memory, longer execution, regional control).
What comes next
Workers provide the execution boundary. Chapter 4 puts a full application around it: choosing where to render, serving static assets, and keeping personalised responses out of shared caches.