Skip to main content

Chapter 5: Local Development, Testing, and Debugging

How do I develop, test, and debug Workers effectively before deploying?


The promise of serverless is reduced operational burden. The reality is that the burden shifts from managing servers to managing the gap between local development and production behaviour. This gap exists on every serverless platform, but Cloudflare's architecture creates a distinctive version. Understanding where simulation ends and reality begins determines whether you develop with confidence or get surprised in production.

Use local tests for runtime behaviour and fast feedback, then deployed staging for configuration, geographic latency and service behaviour the local environment cannot reproduce. The two layers answer different questions.

The simulation boundary​

Miniflare runs Worker code through workerd, the runtime used in production, and supplies local implementations of bindings. That gives useful fidelity for application behaviour. It does not reproduce Cloudflare's network, placement or production capacity controls.

What Miniflare actually simulates​

Miniflare uses workerd rather than a Node.js imitation of the Workers API. Align compatibility dates and tool versions, and test against the runtime you intend to deploy. A passing local test establishes behaviour in that environment; it is not proof of production equivalence.

Binding interfaces match production. KV operations use local SQLite storage with the same API semantics. D1 queries run against actual SQLite with the same SQL dialect. R2 operations write to local files with the same interface. Durable Objects run with the same storage API and single-threaded execution guarantee.

Local workerd does not enforce the production CPU, memory and subrequest limits. Profile locally, but validate resource headroom in a deployed Worker under representative payloads and concurrency. A local run that completes is not evidence that it fits the production limits.

For business logic, data transformation, request handling, and business rules, local simulation provides high-fidelity accuracy.

What needs deployed validation​

Database round trips. Local D1 queries avoid the network path to a deployed database. Twenty sequential queries can look harmless locally and dominate production latency. Test realistic query shapes locally, then measure the deployed path from relevant user and backend locations.

Cache distribution. Local tests can verify keys, expiry decisions and response handling. Deployed traffic is needed to assess hit rates, per-location cold misses and purge behaviour across the network.

Placement and external services. A laptop does not reproduce Durable Object geography, Worker placement or the source addresses seen by an external API. Validate those paths in staging, including rate limits and access policies.

Resource headroom. Measure CPU and memory against production enforcement. Use realistic payloads and concurrency rather than treating a small successful request as a capacity test.

The practical implication​

Run the cheapest test that can expose the failure. Local runtime tests catch binding misuse, SQL errors and coordination mistakes. Remote bindings test selected service interactions. A deployed staging environment tests the configured application across its real network paths. None substitutes for the others.

Development modes and when to use them​

Worker execution and binding location are separate choices. The normal development loop runs code locally, with each supported binding connected to a local simulator or a remote resource. A deployed test adds the production execution location and configuration.

Local development​

wrangler dev runs your Worker locally with simulated bindings. D1 uses local SQLite files. KV and R2 use local storage. Durable Objects run in-process without network traversal. Changes reflect instantly; the feedback loop is sub-second.

Use local development when iterating on request handling, routing, or business logic; developing features that don't depend on specific data; running unit tests; working with sensitive production data you shouldn't touch; or when iteration speed matters more than integration accuracy.

The mental model: you're testing your code, not your infrastructure.

When you need someone, or something, else to reach your local server, Wrangler and the Vite plugin can put it behind a Cloudflare Tunnel on demand. Press t while wrangler dev is running (t then Enter in Vite) to get a public URL: either a throwaway *.trycloudflare.com hostname or a named tunnel you can place behind Cloudflare Access. This is the practical way to test an inbound webhook, share a preview with a colleague, or open your in-progress Worker on a phone, all without deploying it.

Remote development​

Mark supported bindings with remote: true to connect local code to deployed resources while other bindings remain local. Both Wrangler and the Vite plugin support this pattern. It is useful for isolating a particular integration, but the network path still begins on your development machine.

Use remote development when validating database queries against production schemas, testing with realistic data volumes, debugging integration issues that local simulation can't reproduce, verifying latency characteristics, or when accuracy matters more than speed.

The mental model: you're testing your infrastructure, not just your code.

The environment strategy​

Neither mode should touch production data. Configure separate environments, replacing these placeholder IDs with the IDs of the corresponding D1 databases:

Per-environment D1 bindings: dev, staging, production
{
"d1_databases": [
{
"binding": "DB",
"database_name": "myapp-dev",
"database_id": "dev-database-id"
}
],
"env": {
"staging": {
"d1_databases": [
{
"binding": "DB",
"database_name": "myapp-staging",
"database_id": "staging-database-id"
}
]
},
"production": {
"d1_databases": [
{
"binding": "DB",
"database_name": "myapp-production",
"database_id": "production-database-id"
}
]
}
}
}
Essential Safety Practice

Configure dev/staging/production environments before writing code. The five minutes spent prevents the production incident you'll otherwise have.

Local development uses default bindings that are isolated, safe, and disposable. Remote development marks the relevant bindings remote: true and runs wrangler dev --env staging: real services, test data. Production deployment uses wrangler deploy --env production: real services, real users.

Staging environment design​

Staging should mirror production structure without production data.

Separate resources, same schemas. Your staging D1 database should have identical schema to production. Your staging KV namespace should have the same key patterns. Structural differences defeat the purpose of staging.

Realistic data, not production data. Seed staging with synthetic data exercising your code paths. If production has millions of records, staging needs enough to reveal pagination bugs and query performance issues; thousands, not millions, but not empty.

Same binding names, different resources. Code references env.DB everywhere. Environment configuration determines which database that resolves to. Never branch on environment in application code; let configuration handle it.

Rehearse schema changes in staging. Apply and test the migration there before promoting it to production. Keep any schema difference intentional and temporary; unexplained drift gives false confidence when staging tests exercise a different structure.

The Vite alternative​

Wrangler is the standard development tool for Workers, but the Cloudflare Vite plugin provides an alternative for teams already using Vite.

Vite is a frontend build tool that's become the default for React, Vue, Svelte, and other modern frameworks. If your project already uses Vite (a React SPA with an API backend, a full-stack framework like React Router or SvelteKit), the Vite plugin integrates Workers development into your existing workflow rather than requiring a separate tool.

What the plugin provides​

The plugin runs your Worker code within Vite's development server, providing hot module replacement. Change your Worker code and see results immediately without restarting the server. For frontend-heavy projects with Workers backends, this unified experience is significantly smoother than running Wrangler and Vite simultaneously.

TypeScript types for your bindings generate automatically. The plugin reads your Wrangler configuration and produces type definitions, ensuring accurate types for KV, D1, R2, and other bindings without manual maintenance.

Local simulation works identically to Wrangler. The same Miniflare runtime powers both. The difference is developer experience, not simulation fidelity.

When to choose Vite over Wrangler​

Use the Vite plugin when building frontend-centric applications where Vite is already your build tool. React Router v7 officially supports the Vite plugin for full-stack SSR development. If your framework documentation recommends it, follow that guidance.

Use the Vite plugin when hot module replacement matters. Wrangler restarts on file changes; Vite updates modules in place. For rapid frontend iteration, HMR provides faster feedback.

Use Wrangler when building API-only Workers without frontend components; it's simpler for Workers that don't need Vite's frontend capabilities.

Remote bindings are available with both tools. Choose Wrangler or Vite for the project workflow; choose remote resources separately for the integration being tested.

Keep build and deployment commands explicit in the project. The Vite plugin participates in building the Worker as well as local development; deployment should publish the tested output with the intended environment configuration.

Configuration​

Configure the plugin in your Vite configuration file:

vite.config.ts
import { cloudflare } from "@cloudflare/vite-plugin";
import { defineConfig } from "vite";

export default defineConfig({
plugins: [cloudflare({ configPath: "./wrangler.jsonc" })],
});

The plugin reads your existing wrangler.jsonc for bindings, environment variables, and other configuration; you have one configuration file for both development and deployment.

The choice between Wrangler and Vite is about workflow fit, not capability. Both provide accurate local simulation. Teams using Vite for frontend development benefit from the unified experience; teams building backend-only Workers gain nothing from adding Vite to their toolchain.

Inspect local state directly​

Local Explorer exposes simulated KV, R2, D1, Durable Objects and Workflows through an inspector. Use it to inspect a failed test's state, seed a scenario or reset resources. This is useful evidence when application behaviour and your assumptions disagree.

Keep repeatable setup in scripts or fixtures so another developer can reproduce the same state. An inspector helps explain a failure; it should not become the only record of how the environment was prepared.

Testing stateful edge systems​

The testing pyramid applies to Workers, but edge-specific considerations change what matters at each level.

Test runtime behaviour before geography​

Workers removes much process-startup work, but application initialisation, dependency loading and cache warmth can still affect latency. Test the actual application, especially Python or dependency-heavy code, rather than treating isolate startup as the complete response time.

Global deployment also preserves geographic questions: where data lives, where a Durable Object is placed, and which backend a request reaches. Deployed tests should sample those paths. Local tests should exercise the application rules without waiting for a global environment.

Testing Durable Objects​

Use the Workers Vitest integration to exercise objects in the runtime with their storage and bindings. Send concurrent requests that compete for the same state, and check the resulting invariant. Include an external await in tests of code that performs one; single-threaded execution does not prevent interleaving there.

Test persistence by writing state, evicting the object with evictDurableObject() or resetting its in-memory instance, then reading through a new invocation. Stored state should survive; cache state should be reconstructed. These tests can run locally.

Keep deployed tests for routing, placement, resource limits and interactions with real services. The goal is to test your use of the platform's guarantees, including their boundaries, rather than to rebuild a proof of the runtime itself.

Choosing test granularity​

Every test has a cost (time to run, infrastructure to maintain) and a benefit (bugs it catches, confidence it provides). The ratio should guide your strategy.

Test TypeCostCatchesUse When
Unit tests (mocked bindings)MillisecondsLogic bugsBusiness logic, data transformation
Unit tests (Miniflare)SecondsBinding API misuseComplex binding interactions
Integration tests10+ secondsSchema mismatches, query bugsDatabase code, critical paths
E2E testsMinutesDeployment configurationCritical user journeys

Trivial binding interactions (simple get, put, delete operations): unit test with hand-rolled mocks. The binding API is stable; you're testing your logic, not the platform.

Complex binding interactions (SQL queries, transactions, Durable Object coordination): start with tests in the Workers runtime. Add deployed tests for behaviour affected by network placement, service configuration or production limits. Avoid mocks that merely encode the behaviour you hope the service provides.

Critical code (authentication, payment, data integrity): test failure paths at both layers. Fast runtime tests explore retries, malformed input and state transitions; deployed tests verify the configuration and dependencies that enforce the production boundary.

Writing more mock code than application code? Step back and reconsider. Either simplify your binding interactions or accept the cost of integration testing.

Binding mock mismatch deserves special attention because your mock behaves the way you think the real service behaves, which may differ from reality. D1's transaction semantics, KV's eventual consistency, R2's conditional operations all have subtle behaviour. Mocks often implement the happy path while omitting edge cases production surfaces. The more complex the interaction, the more likely your mock diverges in ways that matter.

Debugging distributed edge systems​

When something breaks in production, the debugging approach differs from traditional server applications because failure modes differ. Understanding how failures propagate helps you locate problems faster.

The debugging mental model​

A request arrives at Cloudflare's edge, executes your Worker, potentially calls bindings or external services, returns a response. Failures can occur at any point:

Worker execution failures crash the request with an exception. These appear in logs with stack traces; they're usually straightforward to diagnose and fix.

Binding failures include rejected queries, permission failures and unavailable services. Handle thrown errors separately from valid absent results: a missing KV key returns null. Treating both as “not found” can hide an outage or grant access on incomplete evidence.

External service failures occur when third-party APIs time out, return errors, or behave unexpectedly. Hardest to diagnose because the failure is outside Cloudflare's observability.

Coordination failures in Durable Objects manifest as unexpected state rather than crashes, resulting in wrong results. Two requests that should have been serialised weren't, or state that should have persisted didn't.

When debugging, use the symptom to choose where to investigate first. Stack traces point to Worker code or bindings; unexpected state suggests examining coordination; failures correlated with an external service suggest tracing that dependency. These are starting points, not conclusive diagnoses.

Use the symptom to choose where to investigate first

Loading diagram…

Tracing requests across services​

With tracing enabled, Workers record JavaScript RPC sessions and method calls across Workers and Durable Objects automatically; Chapter 21 covers that facility and its limits. Application correlation IDs remain useful for linking logs to business operations, particularly across Queues and external services. Generate an ID at the entry point and carry it through those paths:

Selecting an application correlation ID
const traceId = request.headers.get("x-trace-id") ?? crypto.randomUUID();
// Include traceId in all log statements
// Carry the ID into queue messages and external calls

Where native traces end, correlation IDs keep related logs searchable. Without them, investigating an asynchronous business process means matching timestamps and guessing which events belong together.

Log sampling and its implications​

Live streams, Workers Logs and exported logs have different sampling and delivery behaviour. A missing error may reflect the collection path rather than a successful request. Head sampling excludes every log from an unselected request, including errors.

For critical operations, specify the evidence that must survive and verify the pipeline against that requirement. An export destination alone does not guarantee complete delivery. Chapter 21 explains sampling, retention and durable evidence for business events.

The diagnostic toolkit​

wrangler tail streams logs in real-time with filtering. Use for active debugging: something is wrong now, you need to see what. Filter by status code, URL path, or IP.

Workers Logs persists your logs for seven days on paid plans, queryable through the dashboard. Use for recent investigations: what happened to that failing request an hour ago? Enable it with a single configuration flag and your console output, errors, and request metadata are retained and searchable without external infrastructure.

Dashboard analytics show aggregated patterns. Use for retrospective analysis: errors spiked yesterday at 3pm. What changed? Correlate error rate spikes with deployment times, traffic spikes, or geographic patterns.

Logpush exports logs to storage or an observability platform. Use it for retention and analysis needs beyond the built-in tools, with explicit sampling and delivery expectations. Chapter 21 covers those trade-offs.

Common failure patterns and their signatures​

Certain failures recur across Workers applications. Naming them creates shared vocabulary for design discussions and post-incident analysis.

Sequential database query latency​

Symptom: Requests work locally but time out or perform poorly in production. Wall time is high; CPU time is low.

Cause: Sequential queries repeatedly pay the Worker-to-database round trip. As an illustrative calculation, twenty calls each adding 20ms of network time contribute 400ms before query execution and response handling. The actual cost depends on placement and the service path.

Diagnosis: Look for loops containing await on database operations. High wall time with low CPU time indicates waiting on I/O.

Fix: Batch queries where possible. Use D1.batch() for multiple independent queries. Restructure code for fewer round trips. Consider denormalising data to reduce query count.

This timing pattern catches nearly every developer new to Workers at least once.

Subrequest limit exhaustion​

Symptom: Requests fail after many fetch calls. Error references subrequest limits.

Cause: Workers default to 10,000 subrequests per invocation on paid plans (50 external and 1,000 to Cloudflare services on free plans). Fan-out patterns, long-lived WebSocket connections, and extended Workflows can hit this limit.

Diagnosis: Count fetch calls, including bindings and external services. Each binding operation counts as a subrequest. Check your configured limit in wrangler.jsonc under limits.subrequests; if unset, the default of 10,000 applies.

Fix: Establish the intended fan-out and retry budget before raising the limit. Paid plans support up to 10 million subrequests per invocation through Wrangler configuration, but a higher ceiling does not increase a dependency's capacity or the invocation's memory. Raise it when the measured workload fits those other constraints. For workloads where the subrequest count is unpredictable or unbounded, batch operations where APIs support it, use continuation tokens rather than fetching all pages at once, or use Queues to distribute work across multiple Worker invocations. You can also set a lower limit alongside cpu_ms to protect against runaway code.

CPU time exhaustion​

Symptom: Requests fail with CPU time limit errors. High CPU time in analytics.

Cause: Computation-heavy operations exceed the 30-second default limit (or 10ms on free plans).

Diagnosis: Profile to identify expensive operations. JSON parsing of large payloads, complex string manipulation, and cryptographic operations are common culprits.

Fix: Stream instead of buffering; parse JSON incrementally and process data in chunks. Offload heavy computation to Containers. Use Queues to distribute work. For legitimate heavy computation that must happen synchronously, configure a higher CPU limit (up to 5 minutes on paid plans).

CPU Time Limits by Plan
PlanDefault CPU TimeMaximum with Configuration
Free10ms10ms
Paid (Standard)30 seconds5 minutes

Note: Legacy "Bundled" plans (no longer available to new accounts) have a 50ms CPU limit. If you're on a Bundled plan from before March 2024, this limit still applies.

Memory pressure​

Symptom: Requests fail mysteriously without clear error messages. This could be crashes or incomplete responses.

Cause: Large payloads or accumulated state exhaust the 128 MB isolate limit.

Diagnosis: Memory issues are hard to diagnose because the failure mode is a crash, not a catchable exception. Look for patterns: failures correlating with large request bodies, data accumulating in loops without releasing references, unbounded collection growth.

Fix: Stream large data instead of buffering; use Request and Response bodies as streams. Release references early. Prefer processing and discarding over collecting and batch-processing.

Unhandled promise rejections​

Symptom: Requests fail with unhandled rejection errors. Stack traces may be unhelpful.

Cause: Async operations fail without catch handlers.

Diagnosis: Look for await calls without try/catch, or .then() chains without .catch().

Fix: Await or deliberately manage every promise whose result matters. Catch errors at the boundary that can recover or return a useful failure; rethrow when recovery belongs elsewhere. A handler-level try/catch cannot catch an unawaited promise that rejects later, and a type check cannot establish that background work survives the request.

Edge-specific failures​

Some failures only manifest at the edge because they depend on properties that don't exist locally.

Geographic distribution mismatch occurs when code assumes consistent behaviour across locations. User A in London writes data; User B in Sydney reads immediately and gets stale results. This isn't wrong code per se; global distribution introduces propagation delays that don't exist locally. KV is particularly susceptible: cached values and misses can remain visible for 60 seconds or longer. Test the application response to a stale value, including revocation and missing-record cases.

External service IP filtering causes mysterious failures where code that worked locally fails in production. Many third-party APIs rate-limit or block requests from cloud provider IP ranges. Your local machine has a residential IP; Cloudflare's edge has well-known IP ranges that security systems treat differently.

Request routing variance surfaces when Smart Placement or custom placement hints route requests unexpectedly. A Durable Object placed near its first user in Tokyo serves that user well but adds latency for users in New York. This is working as designed, but "working as designed" and "meeting expectations" aren't the same.

Carrying a development practice across​

Teams moving from Lambda or Azure Functions can retain their testing and release discipline. The useful change is the runtime and resource boundary they exercise. Local workerd runs Worker code with platform APIs, while deployed tests establish network paths, service configuration and enforced limits. Neither global deployment nor a shared runtime removes the need for those tests.

Keep an existing CI system when it already provides the required checks, credentials and audit evidence. Use Wrangler or the Vite plugin for the local loop and Cloudflare's deployment tools for the target environment. Compare the time from a reported defect to a verified fix, including environment preparation and diagnosis. Deployment duration alone is a poor measure of the development experience.

Deployment and rollback​

A quick deployment shortens both the experiment and the route by which a defect reaches users. Put required checks before exposure, then observe the deployed behaviour. Chapter 22 develops the release policy; the development workflow must make that policy executable.

Bound the first exposure​

Upload a version and use gradual deployment when the change can be tested with a limited traffic allocation. Choose the allocation and observation period from the failure you need to detect: a quiet tenant may not exercise a migration during a brief canary. Version overrides can direct a test request to a particular version, but its bindings and effects still need the intended test boundary.

Define the response to a failed check before deployment. Routing back can be automated when the retained code is compatible with current state. A schema or permission change may require a different recovery path. Monitor correctness and business outcomes alongside errors and latency; a successful HTTP response can still describe the wrong result.

Rollback mechanics​

wrangler rollback can restore an eligible previous Worker version. Practise it with the same bindings and resource lifecycle changes the application uses. Rollback changes code and versioned configuration; it does not reverse stored data or an external action. Deleted bindings and Durable Object lifecycle changes can also restrict which versions remain deployable.

Keep a known recovery target and verify what it can process, including messages created by the new code. If a rollback would misinterpret that work, pause or isolate the affected path and use the prepared repair procedure. A deployment command is one step in restoring service.

Preview deployments​

Worker Previews let a branch run under the same Worker with its own code, configuration, URL, and observability. Deploy one with wrangler preview or create them through Workers Builds. This gives reviewers a live environment for testing changes while production continues to serve its deployed code.

State isolation depends on the resource. Cloudflare provisions separate Durable Object namespaces and Container instances for each Preview. KV, D1, R2, and other account-level resources need explicit bindings to separate resources. A preview URL alone does not make a production database safe to write to.

Use Previews for branch-level validation and keep a staging environment for integration tests that need a complete application boundary. Service bindings from a Preview target the other Worker's production deployment, and Workflow bindings reference existing deployed Workflows. Previews cannot consume Queues or execute cron triggers. Check the entire dependency path before running a destructive test; Chapter 22 covers how that boundary fits into deployment policy.

Working across accounts​

Consultancies and platform teams juggle multiple Cloudflare accounts, and re-running wrangler login to switch between them is both tedious and an incident waiting to happen. Wrangler's authentication profiles fix this: create a named OAuth login per client or environment, bind it to a directory, and Wrangler switches accounts automatically as you move between project folders. A --profile flag pins any single command to a specific profile, and in CI, CLOUDFLARE_API_TOKEN still takes precedence over every profile, so automation stays explicit.

What comes next​

Chapter 6 applies this development loop to coding agents: turning architectural questions into experiments, giving generated changes useful tests, and controlling which resources autonomous work can affect.