Skip to main content

Chapter 1: The Cloudflare Developer Platform

What makes Cloudflare architecturally different, and why does it matter?


A Worker is serverless code that handles requests and other events. If you've used AWS Lambda or Azure Functions, the starting point is familiar: write a handler, deploy it, and let the platform supply the execution environment.

The first time you deploy a Worker, something feels wrong. You write a function, run a command, and seconds later it is available across Cloudflare's network. No server to provision, no region selection dropdown. The instinct honed by years of working with AWS, Azure, or GCP insists that you must have missed a step.

You haven't. The absence of those steps is the point.

The work has moved. You spend less time arranging where processes run and more time choosing where state belongs, which operations must agree, and what each request is allowed to do. Three parts of the platform explain that shift: isolates, global routing and bindings.

The V8 isolate model​

Workers execute inside V8 isolates: separate JavaScript execution contexts within a running process. Each has its own heap and globals. Cloudflare supplies the surrounding runtime and mediates access to network and storage services.

What an isolate actually is​

A conventional serverless execution environment includes a language runtime inside a process. Workers share the cost of an existing V8 runtime across many isolates. Loading another application does not require booting another operating system or starting another Node.js process.

That lowers startup overhead, but isolate creation is only one part of request latency. Your code still has to initialise, fetch data and produce a response. A small handler and a dependency-heavy Python application should not inherit the same latency estimate simply because both are Workers.

What happens when a request arrives​

The runtime can route a request to an existing isolate or create one as needed. An isolate may serve several concurrent requests through its event loop and may be reused later. It can also be evicted; two requests from the same user need not reach the same isolate.

Use reusable globals for safe caches or immutable setup, never as the only copy of application state. A warm isolate is an optimisation you may benefit from, not an identity or persistence guarantee you can build on.

Measure startup with the dependencies and first backend call included. Scheduled warming requests do not establish that the next user's request will reach the isolate you warmed.

Memory: 128 MB and what that means​

Each isolate running your Worker can consume up to 128 MB of memory, encompassing the JavaScript heap and any WebAssembly linear memory. This is a hard, non-configurable limit that cannot be adjusted through pricing tiers or support requests. If your workload routinely requires more than 128 MB per request, Workers are the wrong tool for that workload.

The subtlety is that this limit applies per isolate, not per request. A single isolate can handle many concurrent requests, and those requests share the 128 MB pool. Memory-heavy workloads face compounding pressure under load, precisely when resource constraints hurt most.

Stream large payloads through R2, Cloudflare’s object store, and process records incrementally where possible. If the algorithm must hold a larger working set at once, choose Containers or external compute. The useful distinction is between a large input and a large working set: a multi-gigabyte file can pass through a Worker that never holds more than a small buffer.

CPU time versus wall time​

CPU time measures active computation; wall time includes waiting. A Worker awaiting a database response accumulates elapsed time without spending CPU on that wait. Parsing the returned rows and constructing a response do consume CPU.

The Free plan allows 10 milliseconds of CPU per request. Paid HTTP handlers default to 30 seconds and can be configured for up to five minutes. HTTP wall time has no hard limit while the client remains connected, but a long connection is not durable orchestration. Use Queues for buffered background work or Workflows for a persisted sequence of steps when work must survive the client leaving.

This distinction changes the compute bill for I/O-heavy applications. A handler that waits two seconds while using twenty milliseconds of CPU incurs the same Worker compute usage as one that uses twenty milliseconds without waiting. The external API, database, storage operations and request itself can still carry charges.

Do not replace cheap local computation with an expensive remote service merely to reduce Worker CPU. Compare the complete request path: cost, latency, failure handling and data movement. Chapter 20 develops that calculation.

Resource constraints beyond memory and CPU​

A subrequest is a call the Worker makes to another service, such as an external API or a bound Cloudflare resource. Paid Workers default to 10,000 subrequests per invocation, configurable up to ten million through limits.subrequests. Free plans allow 50 external subrequests and 1,000 calls to Cloudflare services.

A subrequest ceiling bounds how far one invocation can amplify work. Set it from the expected fan-out and retry allowance, including calls through bindings. Chapter 3 examines the runtime limits in more detail.

Single-threaded execution within requests​

JavaScript executes on one thread per isolate. Promise.all() can overlap independent I/O waits; it does not make CPU-heavy JavaScript run on several cores.

Service bindings let one Worker call another. They compose services without adding CPU cores to the calling invocation. By default, the caller and callee run on the same thread of the same server. For independent background computation, divide work into separate invocations and choose a runtime suited to each task. For one calculation requiring several cores, assess Containers or external compute.

Global by default​

A Worker deployment is available across Cloudflare's network without selecting a region for each copy. Request routing normally favours a nearby Cloudflare location, although network routes, capacity and placement configuration determine the actual path. A city's presence on the network map is not a per-request routing guarantee.

What proximity can remove​

Logic that needs only the incoming request can benefit directly: validating a signed token, selecting a redirect or transforming a response need not visit a distant origin. Logic that depends on a database still has to reach that database.

This distinction explains why moving an API handler closer to users can make it slower. If each handler performs a sequence of remote queries, the saved user-to-compute distance may be smaller than the added compute-to-data distance. Draw the round trips before deciding where the code should execute.

Placement control: when global isn't optimal​

Consider an Australian user whose request needs seven sequential queries to a database in Frankfurt. Executing the handler near the user pays the Australia–Europe round trip repeatedly. Executing near the database pays that distance for the user's overall request, while the query sequence stays close to its data.

Smart Placement uses observed traffic to choose execution locations that reduce request duration. Placement configuration can also target known backend locations. These are performance controls; neither should be mistaken for a contractual residency boundary.

Use backend-proximity placement when repeated calls to geographically concentrated services dominate the request. Keep the default when work is mostly local to the handler or its dependencies are already distributed. Measure the full response time from representative user locations, including tail latency: improving one backend leg may make another worse.

Jurisdictional restrictions​

Some contracts and regulatory requirements constrain where data may be processed or stored. Establish the requirement before selecting the product control; GDPR, for example, does not impose a blanket requirement to keep all personal data inside the EU.

A Durable Object combines one entity’s application logic with private durable storage. An object created in the eu jurisdiction keeps its execution and storage there. Workers request processing uses different controls, including Regional Services where applicable. Neither setting by itself settles the locations of logs, external API calls, backups or other bound storage.

Map those paths separately. A restricted Durable Object remains reachable from outside its jurisdiction, so users farther away still pay the network latency. Chapter 22 examines these boundaries in the context of application security and compliance.

The binding model​

A binding gives a Worker a named capability to use a configured resource. The runtime supplies it through the environment object:

Resource access through bindings
// DB: D1; STORAGE: R2; CACHE: KV
const result = await env.DB.prepare("SELECT id FROM users LIMIT 10").all();
const file = await env.STORAGE.get("config.json");
const cached = await env.CACHE.get("recent-queries");

Configuration, not code​

The code names env.DB; configuration determines which D1 relational database it means. Wrangler is Cloudflare’s development and deployment CLI. Its configuration (wrangler.jsonc or wrangler.toml) can bind production to one resource and staging to another. Local development can supply local storage or a configured remote resource.

This keeps resource credentials and client setup out of ordinary application code. It also moves an important review obligation into configuration. A staging Worker bound to production data still has production data access.

Authority does not imply locality​

Bindings let Cloudflare route and authorise resource calls without your code managing public service URLs and credentials. They do not make every call local. A Durable Object has a location, D1 has a primary, and a read from KV, Cloudflare’s distributed key-value store, may miss the nearby cache.

Batch work when it reduces round trips and respect each service's failure model. A convenient method call can still cross an ocean.

A narrower SSRF surface​

Server-side request forgery tricks application code into fetching an attacker-chosen destination. A private service exposed only through a binding has no public URL that an attacker can hand to the Worker's global fetch(). This removes one route to that service.

The handler still needs to validate user-selected URLs, resource identifiers and permitted actions. Code that accepts an arbitrary path and forwards it through a privileged binding can misuse that capability. A private-network binding may grant broader reach than a single-service binding; review what it actually exposes.

Service bindings for composition​

A Worker can expose methods to other Workers through RPC service bindings:

Composing internal services
const user = await env.AUTH.authenticate(request);
const quote = await env.PRICING.getQuote(user.tier, items);

This supports independently deployed components without requiring a public HTTP endpoint for each. By default the Workers share a server thread; placement can change the path. Use the boundary when it clarifies ownership, reuse or authority. Splitting a handler into five services still creates five interfaces to version and investigate.

Design around independent units​

Workers began with request handling, and the platform's stateful services extend that model with explicit units of ownership. A Durable Object coordinates one entity. A D1 database holds one relational partition. A queue distributes work that consumers can process independently.

A database per tenant fits when most transactions stay within that tenant. An object per document fits when collaborators need one authority for changes to that document. Neither choice removes the need to assess the largest tenant or busiest document. Partitioning distributes independent work; it cannot make a single hot partition unlimited.

Choose the unit before the size

Start with the smallest unit whose operations must agree. Then check whether that unit fits the storage and throughput limits. Splitting it further creates a coordination problem you will have to solve elsewhere.

The cost of composition appears between resources. A database updates its own indexes transactionally; a D1 write and a derived KV value require an explicit synchronisation strategy. A Queue can retry a message; it cannot undo an email already sent. The platform supplies useful boundaries, and the application owns what crosses them.

The work you still own​

Automatic scaling reduces fleet management, but the database, external API and account budget still have limits. Plan for overload at the narrowest dependency, and decide whether excess work should wait, fail or be rejected before admission.

Global code availability reduces deployment work, but shared dependencies can still produce a global outage. Use gradual releases, compatible data changes and a recovery plan that distinguishes reverting code from restoring state.

Managed services reduce infrastructure work, but somebody still owns schemas, authorisation, retention and on-call diagnosis. Those responsibilities should shape the choice of primitive as much as its API does.

Cloudflare is a strong candidate when request handling, independent partitions and explicit coordination match the application. Keep external services when they solve requirements the platform does not: a large relational partition, specialised runtime, established operational contract or model capability. A useful architecture can cross a provider boundary without being unfinished.

What comes next​

Chapter 2 turns this platform model into an adoption decision: which workloads fit, which organisational costs matter, and what evidence should persuade you to proceed or stop.