Chapter 2: Strategic Assessment
Should we adopt the Cloudflare Developer Platform, and what are the implications?
Platform adoption is never purely technical. The architecture you choose must fit team capability, budget cycles, vendor relationships, and incentives. A platform that is technically strong but organisationally incompatible will fail as reliably as one with fundamental technical flaws.
This chapter focuses on the Cloudflare Developer Platform: Workers, Durable Objects, D1, R2, Queues, Workflows, Containers, Realtime, and the primitives covered throughout this book. Cloudflare also offers networking and security products (WAF, DDoS protection, rate limiting, bot management, Zero Trust access), but those serve different architectural decisions and deserve separate evaluation.
The distinction matters because the Developer Platform delivers substantial networking and security benefits automatically. When you deploy a Worker to a custom domain, your application receives Cloudflare's DDoS mitigation, TLS termination, and network optimisation without configuration. Your traffic flows through infrastructure handling billions of attack requests daily. The platform you're evaluating for compute runs on one of the world's most capable security networks.
Understanding what you receive automatically versus what requires explicit enablement shapes both evaluation and architecture. This chapter helps you make that assessment.
The decision in brief
The core framework follows. The rest of the chapter substantiates these recommendations.
Choose Cloudflare when users are globally distributed and latency matters, when you need real-time coordination that would require significant custom infrastructure elsewhere, when egress costs dominate your cloud budget, or when operational simplicity matters more than maximum flexibility. Greenfield projects benefit most by designing for Cloudflare's model from the start.
Choose hyperscalers when you need specific managed services (SageMaker, BigQuery, Cosmos DB), when workloads are compute-intensive rather than I/O-heavy, when your team has deep hyperscaler expertise and time-to-market is critical, or when regulations mandate specific clouds. Significant reserved capacity commitments also shift the economics until they expire.
Choose hybrid when you're migrating incrementally, when different workloads have different needs, or when risk tolerance requires gradual adoption with rollback capability. Hybrid isn't a compromise; it's often optimal for complex organisations.
Some requirements represent genuine mismatches with Cloudflare's architecture rather than limitations to engineer around. These include: applications requiring more than 128 MB memory per request where Containers' cold starts are unacceptable, workloads needing inbound UDP, or GA inbound TCP today (an inbound-TCP path is in private beta via Spectrum), databases exceeding 10 GB that cannot be horizontally partitioned, or contractual requirements mandating a specific hyperscaler. If any of these apply today, the relevant sections of this book will help you understand why and what alternatives exist. If none apply today, understanding these boundaries helps you recognise when future requirements might push you towards hybrid architectures.
The following sections develop these recommendations with the detail needed to apply them.
The platform in context
Before evaluating workload fit and migration paths, understand what the Developer Platform includes and how it relates to Cloudflare's broader offerings.
What the developer platform comprises
The Developer Platform includes the compute and storage primitives you deploy and operate:
- Workers: Core compute executing JavaScript, TypeScript, Python, and WebAssembly at the edge
- Durable Objects: Stateful coordination through globally-unique, strongly-consistent actors
- D1, R2, KV: Storage for structured data, objects, and key-value pairs
- Queues and Workflows: Asynchronous processing and durable execution
- Containers: Extended compute for workloads exceeding isolate constraints
- Workers AI and Vectorize: Inference and vector storage for intelligent applications
- Workers for Platforms: Multi-tenant isolation for platform builders
These primitives share a deployment model, billing model, and operational philosophy. They compose through bindings and represent the compute layer you architect upon.
What comes bundled
Deploying on the Developer Platform automatically provides capabilities that would require separate products, configuration, and cost on other platforms.
DDoS protection operates continuously on all traffic reaching your Workers. Cloudflare's network absorbs volumetric attacks before they reach your code. The same systems protecting the world's largest websites protect your application from deployment. Enterprise customers gain additional controls and analytics, but core protection works for everyone.
TLS terminates automatically with certificates provisioned and renewed without intervention. Your Workers receive requests over HTTPS; certificate management complexity disappears.
Network routing is managed by Cloudflare. Anycast brings requests into the network without an application-managed global load balancer. The resulting latency still depends on routing, congestion and backend location; it is not guaranteed to beat every direct path to an origin.
Global distribution requires no configuration. Deploy once and your code runs in over 300 locations. The alternative (replicating infrastructure across AWS regions, configuring multi-region databases, implementing failover logic) represents weeks of engineering and ongoing operational burden. On Cloudflare, it's the default.
Bot controls require deliberate configuration. Bot Fight Mode and Bot Management have their own scope and settings. Do not infer a configured bot policy from the fact that a request reaches a Worker.
These capabilities aren't upsells positioned as standard; they're architectural consequences of building on Cloudflare's network. Traffic to your Workers traverses Cloudflare infrastructure, which includes DDoS mitigation, TLS termination, and network optimisation by design.
The strategic implication: evaluating the Developer Platform purely on compute and storage undersells the value. You're also acquiring networking and security infrastructure that would otherwise require separate evaluation, negotiation, and operational investment.
What requires additional products
Some enterprise requirements exceed what the Developer Platform provides automatically.
Web Application Firewall (WAF) provides application-layer protection beyond basic DDoS mitigation: OWASP ruleset enforcement, custom rules blocking specific attack patterns, and managed rules updated for emerging threats. Many applications operate safely with bundled protections; those handling sensitive data or facing sophisticated attackers benefit from explicit WAF deployment.
Advanced rate limiting operates at the Cloudflare edge with granular controls. Chapter 24 distinguishes approximate local abuse controls from strict quotas owned by a Durable Object. For edge-level rate limiting with complex rules, geographic conditions, and threat intelligence integration, Cloudflare's rate limiting products exceed what you'd build yourself.
Bot Management adds classification and policy controls for automated traffic. Select rules for the actual threat model and test legitimate API clients as well as browsers; automated traffic is not inherently unwanted.
Access and Zero Trust control who reaches your applications based on identity, device posture, and context. If your Workers serve internal tools, administrative interfaces, or B2B applications with controlled user populations, Access policies enforce authentication before requests reach your code.
Map the required controls to the services and settings that provide them. DDoS protection does not establish an application authorisation policy, and a managed certificate does not validate an incoming payload. Include additional products where that control map requires them, then test the complete request path.
Workload fit
Cloudflare's architecture favours globally distributed, I/O-heavy, coordination-intensive workloads. It's less suited to memory-hungry, compute-bound workloads or those tightly coupled to hyperscaler services.
Workloads that thrive
The platform's sweet spot becomes clear once you understand the underlying model.
API backends with global users can avoid managing separate regional compute deployments. The latency benefit depends on the complete path: a nearby Worker may still call a distant database. Stateless request handling benefits most directly; replicated data and coordinated state require their own decisions.
Latency-sensitive authentication can avoid a remote round trip when a Worker verifies a token locally. Current permissions, revocation and session state may still require an authoritative lookup. Separate signature verification from those checks before estimating the latency benefit.
Real-time coordination and collaboration can fit Durable Objects when a room, document or match is a useful ownership boundary. The object combines state, routing and live connections. Clients still need reconnect and conflict-handling behaviour, and an existing actor or collaboration service may already meet the requirement. Chapter 7 examines when central authority simplifies the design.
I/O-heavy orchestration exploits the billing model's deepest advantage. API aggregation, webhook processing, and service composition spend most of their time waiting on external calls. Cloudflare's pricing favours workloads that wait; you pay for waiting on hyperscalers but not on Cloudflare. A Worker making five API calls and waiting 500ms might consume only 10ms of billable CPU time.
Multi-tenant SaaS platforms benefit from horizontal scaling patterns matching the platform's design. The database-per-tenant pattern aligns naturally with D1's architecture. Rather than fighting a 10 GB limit on a single database, embrace thousands of isolated tenant databases, each well under the limit with natural isolation boundaries. Workers for Platforms extends this to tenant-specific code execution.
AI-powered applications with global users can combine Workers AI, Vectorize and storage through platform bindings. That simplifies integration without guaranteeing that every stage runs near the user. Measure retrieval, inference and tool calls together; Chapter 17 develops the inference decision.
Workloads that struggle
These constraints are hard, not soft. Understanding them prevents costly discoveries in production.
Memory-intensive processing hits the 128 MB isolate limit without recourse. Image manipulation libraries buffering entire images, data processing loading large datasets into memory, and ML inference with substantial model sizes cannot run in Workers. Containers offer up to 12 GiB but introduce cold-start latency measured in seconds. If your workload routinely needs more than 128 MB and cannot tolerate those cold starts, you need different infrastructure.
Long-running computation can exceed a paid HTTP Worker’s five-minute CPU ceiling. Use Containers, external compute or shorter durable stages when the job needs more. Waiting and computing have different limits: an HTTP request using four minutes of CPU over an hour still depends on the client connection, while other triggers have their own execution budgets.
Single large databases conflict with D1's architectural assumptions. The 10 GB-per-database limit exists by design, guiding you towards horizontal patterns. If your application requires a monolithic relational database with complex joins across hundreds of gigabytes, D1 won't work, but Cloudflare still can. Hyperdrive accelerates connections to external PostgreSQL or MySQL databases, providing connection pooling and query caching at the edge. Many production systems run permanently with Workers connecting to external databases via Hyperdrive. You're operating that external database with its associated complexity, but you're also gaining edge compute benefits for your application logic.
Inbound UDP connections don't fit, and inbound TCP only partly. Cloudflare's ingress is HTTP-centric; non-HTTP traffic reaches the network through Spectrum, its TCP and UDP proxy. Spectrum can hand an inbound TCP socket directly to a Worker or Container through a connect() handler, which is what lets you serve gRPC on Cloudflare, but that path is in private beta and TCP-only. Inbound UDP has no equivalent, so game servers expecting direct UDP connections and latency-sensitive custom protocols need traditional cloud infrastructure for ingress. Treat inbound TCP as a private beta, not something to stake a launch on.
Assessing your workloads
For each significant workload, assessment reduces to key questions.
Memory envelope? Measure peak memory with representative payloads and concurrency. Workers share a 128 MB limit across requests in an isolate, so a request that fits alone may still leave too little headroom under load. If required working data cannot fit after streaming or chunking, use Containers or external compute.
Compute profile? Under one second of CPU time per request is ideal; up to thirty seconds fits the paid default. HTTP handlers can be configured for up to five minutes of CPU time. Beyond that, use Containers, external compute, or divide the work into shorter invocations. Queues decouple processing from the request; they do not provide unlimited CPU.
Where are your users? Global or distributed users benefit from edge execution. Users concentrated in one region see less advantage, though operational simplicity may still justify adoption.
Where are your backends? A database in one location can dominate request latency wherever the Worker runs. Placement can bring compute closer to it. Draw the path to every required dependency; using a Cloudflare binding does not by itself make that dependency local.
What protocols? HTTP and WebSocket fit perfectly. Inbound UDP doesn't fit; inbound TCP is in private beta via Spectrum (which also lets you serve gRPC), so treat it as emerging rather than production-ready today.
Need real-time coordination? Durable Objects provide capabilities unavailable elsewhere without significant custom work. This alone can justify adoption even if other factors are neutral.
A workload doesn't need perfect scores to succeed on Cloudflare, but multiple poor-fit dimensions signal you'll fight the platform rather than benefit from it.
Which of your current workloads would benefit most from edge execution, and which suit centralised compute? The answer often reveals that hybrid architecture is optimal rather than a compromise.
Translating hyperscaler experience
Engineers from AWS, Azure, or GCP carry accumulated intuitions about cloud platforms. Some translate directly; others require deliberate unlearning. This section maps familiar hyperscaler concepts to Cloudflare equivalents, highlights where familiar assumptions fail, and identifies patterns that don't apply.
The concept map
| Hyperscaler Concept | Cloudflare Equivalent | Key Difference |
|---|---|---|
| Lambda / Azure Functions | Workers | Isolates avoid process startup; measure application initialisation |
| API Gateway | Workers Routes | Integrated into compute; no separate service or pricing |
| DynamoDB / Cosmos DB | D1 + Durable Objects | D1 for relational queries; Durable Objects for coordination |
| S3 / Blob Storage | R2 | Zero egress fees; S3-compatible API |
| ElastiCache / Redis | KV (caching) or Durable Objects (coordination) | KV for eventual consistency; DO for strong consistency |
| Step Functions / Durable Functions | Workflows | Simpler model; automatic retry and state persistence |
| SQS / Service Bus | Queues | Worker consumers or HTTP pull; at-least-once delivery |
| ECS / Container Apps | Containers | Must route through Workers; no direct inbound connections |
| CloudFront / Front Door | Built-in | Every Worker deployment includes global CDN automatically |
| WAF / Shield | Built-in (basic) or WAF product | DDoS protection included; advanced WAF separate |
| Secrets Manager | Workers Secrets | Encrypted at rest; accessed via environment bindings |
| CloudWatch / Monitor | Workers Logs + Analytics Engine | 3-day Free / 7-day Paid log retention; export for longer retention |
Startup work changes
Workers creates isolates within an existing runtime, avoiding the process startup common to regional serverless functions. That removes much of the reason to maintain warming infrastructure for ordinary JavaScript handlers.
Application startup still matters. Module initialisation, Python dependencies and empty caches can add work to a first request. Benchmark the application before importing a warming strategy from Lambda or declaring startup irrelevant.
Connection management: the disappeared problem
Traditional serverless architectures struggle with database connections. Lambda functions spin up independently, each wanting its own connection, quickly exhausting database connection limits. Solutions include RDS Proxy, PgBouncer, connection pooling layers, and careful timeout tuning.
Cloudflare's binding model eliminates this for platform-native storage. D1, KV, R2, and Durable Objects use bindings the platform manages, not connection pools. Your code references env.DB and the platform handles everything else: no connection string, no pool size to tune, no connection exhaustion to monitor.
For external databases, Hyperdrive provides connection pooling at the edge. You configure once; the platform manages globally. The mental shift: connections become someone else's problem for native storage, and a simpler problem for external databases.
Regional architecture: the inverted model
Hyperscaler architecture starts with a region decision: which region hosts primary infrastructure, which need replicas, how to handle failover. Multi-region deployment is an advanced pattern requiring explicit design.
Cloudflare inverts this. Global deployment is the default: deploy once and code runs everywhere. The advanced pattern is restricting deployment to specific jurisdictions for compliance, not expanding it to multiple regions for performance.
Global compute does not make data uniformly local. D1 places a primary database and can use read replicas; Durable Objects place each entity at one location. Choose consistency and placement deliberately, then measure the distance between users, compute and state.
Smart Placement further inverts expectations. Instead of running compute near users and accepting backend latency, Smart Placement runs compute near backends when that produces better total latency. You enable it with configuration, not architecture.
The services that don't exist
Some hyperscaler services have no Cloudflare equivalent because the need they address doesn't exist or manifests differently.
NAT Gateway: Workers have outbound internet access by default. No VPC to escape, no NAT to provision, no hourly charge accumulating invisibly.
Load Balancer: The Anycast network routes requests to the nearest available location automatically. Load balancing is inherent to the deployment model.
Auto Scaling Configuration: Workers distributes stateless invocations without instance-count settings. Capacity still needs attention where work converges on a database, one Durable Object or a rate-limited external service.
Container Orchestration: Containers exist but don't require orchestration in the Kubernetes sense. Deploy a container; the platform runs it. No cluster to manage, no node pools to size, no pod specifications.
Service Mesh: Worker-to-Worker communication uses service bindings with RPC semantics. No sidecar proxy, no service mesh control plane, no traffic policies. Bindings provide type-safe communication without network configuration.
These absences aren't gaps; they're consequences of the model. When the platform handles distribution, load balancing, and scaling inherently, services configuring those concerns become unnecessary.
Pricing model translation
Hyperscaler serverless pricing typically combines request charges, duration charges, and memory allocation. Lambda charges per GB-second (memory multiplied by duration). API Gateway adds per-request fees. Data transfer adds egress fees.
Cloudflare's model differs in ways affecting cost intuition.
CPU time, not wall time: Workers charge for CPU milliseconds, not elapsed time. A request waiting 500ms for external APIs but computing for 5ms pays for 5ms. On Lambda, you'd pay for 500ms. I/O-heavy orchestration becomes dramatically cheaper, though compute-heavy workloads don't benefit.
Zero egress: R2 charges nothing for data leaving the platform. If egress exceeds 20% of your current cloud bill, this single factor may dominate all other comparisons.
No API Gateway equivalent: Workers handle HTTP routing directly. No separate API Gateway service with its own per-request pricing; compute includes what hyperscalers charge separately for routing.
Usage-based storage queries: D1 charges per rows read and written, not per provisioned capacity unit. Costs align with actual usage but require different capacity planning: you're forecasting query patterns, not instance sizes.
Patterns that transfer
Not everything requires unlearning.
Presigned URLs for direct upload: R2 supports presigned URLs with S3-compatible APIs. Existing client-side upload patterns work unchanged.
Event-driven processing: Queues trigger Workers just as SQS triggers Lambda. The consumer pattern is familiar; only configuration syntax differs.
Scheduled execution: Cron Triggers work like CloudWatch Events or Azure Timer Triggers. Specify a cron expression; the platform invokes your code.
Environment-based configuration: Workers Secrets and environment variables work like Lambda environment configuration. Sensitive values encrypted, non-sensitive plaintext. The binding model provides type-safe access, but the concept is familiar.
Infrastructure as code: Wrangler configuration files serve the same purpose as CloudFormation or Terraform. You can also use Terraform with the Cloudflare provider if that's your team's standard.
What requires rethinking
Several patterns require fundamental rethinking rather than direct translation.
Actor-based coordination stands out as Cloudflare's most distinctive capability. Durable Objects have no direct hyperscaler equivalent. The closest patterns (Step Functions with DynamoDB, Durable Functions with entity functions, self-managed Orleans) require multiple services and significant custom code. If your application needs coordination, Durable Objects likely replace an entire architectural layer you'd otherwise build from components.
Horizontal database scaling: D1's many-small-databases model differs from RDS's scale-up approach or DynamoDB's horizontal sharding. If you're used to one database handling everything, shifting to database-per-tenant requires rethinking data architecture, not just translating configuration.
Edge-first AI: Workers AI runs inference at the edge, not in centralised GPU clusters. The latency profile differs from Bedrock or Azure OpenAI. Model selection differs too: Workers AI runs open-source models Cloudflare hosts, not proprietary models from OpenAI or Anthropic (though AI Gateway can proxy to external providers).
The economics
The sticker price of any cloud platform misleads. Real cost includes direct charges, indirect operational overhead, and savings through architectural efficiency.
How Cloudflare charges
Workers charge for requests and CPU time. The paid plan ($5/month base) includes 10 million requests and 30 million CPU-milliseconds monthly. Beyond that, you pay $0.30 per million requests and $0.02 per million CPU-milliseconds.
Storage services add costs. D1 charges per rows read and written plus storage. R2 charges for storage and operations but zero egress. KV charges for reads, writes, and storage. Durable Objects charge for requests, duration, and storage.
The free tier (100,000 requests daily, 10ms CPU time per request) suffices for development and low-traffic applications but doesn't represent production capabilities.
What hyperscalers actually cost
Hyperscaler pricing appears straightforward until you deploy. The compute cost you model is rarely what you pay.
Egress fees compound invisibly. AWS charges $0.09 per GB leaving a region. An API serving 1 TB monthly pays $90 in egress alone, before compute, storage, or anything else. Organisations often discover egress costs only when the bill arrives. R2's zero-egress model eliminates this category.
NAT Gateway fees accumulate for VPC-connected functions. Lambda functions requiring external connectivity through NAT Gateway pay $0.045 per hour plus $0.045 per GB processed. A moderately active architecture processing 100 GB monthly through NAT pays approximately $40/month, rarely appearing in initial estimates.
Count startup mitigation separately. Provisioned concurrency can add a standing charge to a Lambda deployment. Compare that configuration with measured Worker startup behaviour, including application initialisation, rather than assuming a fixed saving from isolate startup alone.
Include internal data transfer where it is billed. Charges depend on the services and paths involved. Trace the actual traffic between compute, databases and load balancers rather than applying one cross-zone rate to every interaction.
Reserved capacity commits regardless of usage. Reserved instances and savings plans reduce per-hour costs but lock you to capacity regardless of demand. Traffic dropping 50% still costs 100% of your reserved commitment.
A representative calculation
Consider 50 million requests a month at 20ms of CPU per request. That is one billion CPU milliseconds. On Workers Paid, after the included ten million requests and thirty million CPU milliseconds, the compute bill is $5 + $12 + $19.40 = $36.40. This covers Workers requests and CPU, not the application's database, logs or other services.
The useful comparison is with a measured alternative. Lambda bills execution duration and allocated memory, so CPU time alone does not establish its cost. API Gateway pricing depends on the API type, and database costs need explicit read and write volumes. Compare the same request path, traffic profile, caching assumptions and included allowances on both platforms.
For an I/O-heavy service, investigate duration billing and gateway charges first. For a media-heavy service, investigate egress. Those are hypotheses to test against your bill, not a universal savings percentage.
When Cloudflare isn't cheaper
Cost advantages disappear in predictable scenarios.
Compute-intensive workloads consuming hundreds of milliseconds of CPU time per request pay for that computation. Workers' per-CPU-millisecond pricing becomes expensive for sustained computation. Lambda's per-GB-second model may cost less for workloads needing significant memory but modest CPU.
Very low traffic may not justify the $5/month Workers Paid base cost. First check whether the workload fits the Free plan’s daily request, CPU and service limits; 10,000 monthly requests alone do not imply a paid bill. Compare the allowances and required features on both platforms.
Heavy hyperscaler service integration can make adding Cloudflare more expensive. If you're using AWS services charging for external access (RDS data transfer, cross-account SQS messaging, Kinesis streaming), Cloudflare as a front layer may increase total cost.
Existing reserved commitments represent sunk costs. Organisations with multi-year reserved instance or savings plan commitments don't recoup that investment by migrating. Hyperscaler compute is effectively cheaper until those commitments expire.
Modelling your own costs
The representative comparison illustrates dynamics but doesn't tell you your costs. To model your situation, gather these data points:
Monthly request volume drives base Workers cost directly.
Average CPU time per request: if unknown, instrument a sample. The ratio of CPU time to wall time determines whether Cloudflare's billing model advantages you.
Current egress volume is often buried in bills or unknown entirely. It's the largest hidden cost on hyperscalers and the largest potential saving with R2.
Cold start mitigation costs: count provisioned concurrency or warming infrastructure that the new workload would avoid. Retest application initialisation and dependency loading before assuming that entire cost disappears.
Operational overhead: multi-region deployment, capacity planning, scaling policy tuning, cold start debugging consume engineering time. This cost doesn't appear on cloud bills but is real.
With these numbers, construct a comparison meaningful for your situation rather than relying on representative examples.
Lock-in and portability
Switching costs exist with any platform. Understanding them enables informed commitments rather than accidental ones. Lock-in is proportional to differentiation: the features hardest to leave are the features you can't get elsewhere.
Accept lock-in proportional to the value it creates. The features you can replicate elsewhere impose low lock-in cost; those you cannot should deliver value justifying the switching costs if you ever need to leave.
What travels with you
Application logic ports readily. Workers use standard JavaScript and TypeScript with Web APIs. Business logic (request handling, data transformation, algorithmic code) runs elsewhere with modest adaptation.
R2 data extracts via standard S3 tools. The S3-compatible API means your data isn't trapped. You can migrate to any S3-compatible storage using tools you already know.
D1 data exports as SQLite. Schemas and data move to any SQLite-compatible database. The data model is standard; only the runtime is Cloudflare-specific.
Queue message formats are your own. The semantics (at-least-once delivery) are standard. Migrating means implementing equivalent consumers elsewhere, not reformatting messages.
What requires rewriting
Durable Objects creates a substantial execution-model commitment. A replacement must reproduce the routing, storage, concurrency and connection behaviour the application uses. Other actor systems provide related abstractions, but an exit plan must map those guarantees rather than assume an exported dataset is a portable application.
Binding semantics don't exist elsewhere. The env.RESOURCE pattern for resource access is Cloudflare-specific. Code heavily dependent on bindings needs modification, though changes are mechanical rather than architectural.
Workers-specific APIs require alternatives. HTMLRewriter for streaming HTML transformation, specific caching APIs, and other Cloudflare-native features need replacement. These are typically small codebase portions but require attention during migration.
The global deployment model doesn't port. Architectures assuming instantaneous global deployment need rethinking for hyperscalers' regional models. This is a design assumption affecting application structure.
Calibrating your tolerance
Lock-in correlates with feature depth. Stateless Workers for APIs create low lock-in with moderate migration effort. D1 for CRUD applications increases lock-in slightly. R2 for storage creates minimal lock-in due to S3 compatibility. Durable Objects for coordination create high lock-in with high migration effort but deliver capabilities unavailable elsewhere. Workers for Platforms for multi-tenant architectures creates very high lock-in.
The question isn't whether to accept lock-in but how much value justifies it. Durable Objects' lock-in reflects genuine differentiation: capabilities you cannot replicate without significant custom infrastructure. That lock-in buys something real.
Accept lock-in consciously, understanding what you're trading for what you're getting. Accidental lock-in (discovering dependencies you didn't know you'd created) is the risk to avoid. Intentional lock-in for genuine value is often the right choice.
Team readiness
A pilot should test whether the team can operate the application as well as build it. Familiar JavaScript and HTTP APIs shorten the first deployment; they do not establish that someone can diagnose a stale read, recover a stuck workflow or explain a cost spike.
Test the unfamiliar parts
Actor-model experience helps with Durable Objects. Database experience helps with partitioning and migrations. Existing Terraform and CI practices still matter: bindings describe what code can access, while infrastructure tooling manages resources and their lifecycle. Chapter 5 covers the local and deployed tests that exercise those boundaries.
Avoid promising a standard ramp-up period. Have the team build one representative slice, explain its failure modes to a colleague, and recover it from a deliberately introduced fault. The gaps in that exercise give you a more useful training plan than a forecast of “two months to proficiency”.
Put operations inside the pilot
Name the people who will investigate failures, review costs and maintain deployment tooling. Give them access to logs and budgets before the first production release. A prototype operated only by its enthusiastic author can conceal a substantial support burden.
Record the decisions that a new on-call engineer would need: which resource owns the state, which writes can be replayed, how to disable a failing path, and what a rollback restores. If a second engineer cannot follow those instructions, the pilot has more work to do.
Set an evidence threshold
Choose a bounded workload with a measurable problem: a slow request path, expensive data transfer or coordination that is difficult to operate. Include its hardest dependency. A disconnected demonstration may be easy to finish but prove little about the intended application.
Agree on the baseline, the acceptable result and the conditions for stopping. Share the result with engineering, operations and the budget owner, including the parts that became harder. A pilot that rules out a poor fit has earned its cost.
Keep unrelated product additions outside the migration's acceptance criteria. Changing the platform and expanding the feature set together makes regressions harder to explain and completion harder to recognise.
Migration realities
Migration is where theoretical benefits meet practical friction. Understanding what makes migrations hard (not just what steps to follow) determines whether yours succeeds. This section covers strategic thinking. Chapter 27 provides detailed playbooks for specific scenarios including S3 to R2, Lambda to Workers, and container migrations.
Interrogate your requirements first
Before mapping Lambda to Workers or DynamoDB to D1, ask: what problem was this system actually solving?
Systems accumulate features, workarounds, and cargo-cult patterns. Not all serve the original purpose. Migration is an opportunity to shed accumulated complexity, but only if you define success through outcomes rather than feature parity.
Instead of "migrate our user service from Lambda to Workers," try "achieve sub-50ms authentication response times globally." Instead of "replicate our DynamoDB schema in D1," try "support 10,000 concurrent users with consistent read-after-write semantics." These framings licence your team to use Cloudflare's primitives as designed rather than contorting them to match AWS patterns.
Ask your team: if we were building this today with no legacy constraints, would we build it this way? The answer is often no, and migration is your chance to act on that knowledge.
This reframing has practical benefits. When success is defined by outcomes rather than feature parity, teams make autonomous decisions about implementation without constant approval cycles. "Does this help us hit sub-50ms latency?" is a question engineers can answer themselves. "Does this exactly replicate what Lambda did?" requires archaeology into decisions nobody remembers making.
The fundamental challenge
Data migration is where timelines fail. Teams budget months for code and weeks for data; they should budget the reverse.
Code migration is largely mechanical: handlers rewrite to new APIs, logic transfers with syntax changes, and tests adapt to new mocking patterns. The work is tedious but predictable. Estimate by counting handlers and measuring complexity.
Data migration is unpredictable because your data contains surprises: edge cases code never handled, implicit assumptions never documented, format inconsistencies accumulated over years, referential integrity existing only in application code. Data migration surfaces these; code migration doesn't.
Running parallel systems with data synchronisation compounds complexity. Dual writes must handle failures on either side. Consistency windows create edge cases. Cutover timing affects users. The longer you run parallel, the more edge cases you discover, and the more you need to run parallel to handle them.
Map the overgrowth before you cut
Every system operating for more than a year has accumulated dependencies you don't know about. Some are documented in code; many aren't. Before beginning migration, discover what your system actually depends on, not just what you think it depends on.
Timing assumptions hide everywhere. Does your code assume specific latency ranges for database queries, queue processing, or external API calls? Cloudflare's global distribution may violate these. A process assuming "database queries return in under 10ms" may behave incorrectly when Smart Placement puts your Worker in Frankfurt while your backend remains in Virginia.
Consistency assumptions deserve particular attention. KV provides eventual consistency; cached values and misses can remain visible for 60 seconds or longer. Code reading immediately after writing may see stale data. D1 supports transactions within a single database, but not across databases or between D1 and R2. Chapter 12 maps these guarantees precisely; the strategic implication is simpler: audit what your code assumes and verify those assumptions hold in the new environment.
SDK behaviours don't transfer. The AWS SDK's automatic retry with exponential backoff doesn't exist in fetch calls to external services. You'll need to implement it yourself or accept different failure modes.
IAM patterns need consolidation. Authorisation logic distributed across IAM policies, resource policies, and application code must consolidate into Workers-native patterns or Cloudflare Access rules.
Filesystem assumptions need inspection. Workers expose a virtual filesystem through Node.js compatibility, including bundled files and temporary storage, but no host filesystem or persistent local disk. A dependency that reads packaged configuration may fit; one that expects durable local writes or operating-system access may still require Containers.
The goal isn't eliminating all dependencies before migration; that's rarely practical. The goal is knowing what they are, so surprises become planned transitions rather than emergency debugging sessions.
Migrating from Lambda and DynamoDB
Difficulty correlates directly with how deeply you've used DynamoDB's patterns.
Teams estimate migration time by counting Lambda functions, then discover data model translation takes 3x longer than expected. Assess DynamoDB access patterns first: GSIs, sparse indexes, and composite keys require schema redesign surfacing buried assumptions.
Simple key-value access translates cleanly to D1 or KV. Single-table design with GSIs, sparse indexes, and composite keys requires schema redesign. If your DynamoDB usage looks like SQL with a different API, migration is straightforward. If it exploits DynamoDB's unique capabilities, expect significant design work.
Assess DynamoDB access patterns exhaustively first. Identify every query pattern, every GSI, every scan operation. Understand which patterns translate to SQL and which require rethinking. The data model assessment, not the function count, determines migration difficulty.
The coordination question matters too. Lambda functions coordinating through DynamoDB conditional writes, transactions, or streams may be implementing patterns belonging in Durable Objects. Recognising coordination logic masquerading as database operations, and migrating it to purpose-built coordination primitives, often improves the resulting architecture significantly.
Migrating from S3 to R2
R2 supports a substantial subset of the S3 API. Check the operations, headers and lifecycle features your application uses before calling the migration straightforward; API compatibility does not imply every S3 capability is present.
Primary work is validating your S3 API usage falls within R2's supported operations. Most does. Some advanced features (certain analytics, inventory reports, cross-region replication, object lock compliance mode) have no R2 equivalent.
Event-driven architectures need attention. S3 events triggering Lambda functions don't have a direct equivalent; R2 offers event notifications triggering Workers, but the integration pattern differs. Architectures built around S3 events need rewiring, not reconfiguration.
The migration mechanics are simple: configure R2, update SDK endpoints and credentials, migrate data, switch traffic. Simple doesn't mean careless: run parallel reads comparing responses, verify data integrity, monitor error rates.
Migrating from ECS or Fargate
First question: should you migrate at all?
Many ECS workloads can run as Workers with acceptable constraints. If your container exists because "that's how we deploy" rather than because you need specific container capabilities, evaluate Workers first. Operational simplification often justifies modest code changes.
Containers make sense when you need more than 128 MB memory, arbitrary runtimes that won't compile to WebAssembly, existing container images whose rewrite cost exceeds migration benefit, or CPU time beyond Workers' limits.
If you need Containers, understand how differently they work from ECS: traffic routes through Workers or Durable Objects, not directly to containers; container instances tie to Durable Objects for coordination; maximum resources (4 vCPU, 12 GiB RAM, 20 GB disk) may be less than current instances; cold starts measure in seconds; no VPC equivalent, so network isolation patterns need rethinking.
This isn't lift-and-shift; it's restructure-and-shift. Teams expecting to move container images unchanged discover they're redesigning architecture.
Common failure patterns
The Big Bang fails because one component failure stalls everything. Teams get stuck running both platforms indefinitely, paying for both while neither works well. Migrate incrementally, one workload at a time. Prove each migration before starting the next.
Underestimating data migration leaves reconciliation and cutover until too late. Start early, keep one write authority, replicate changes through a recoverable path and reconcile through a known position before cutover. Chapter 27 develops that sequence; include rehearsals and repairs in the estimate.
Assuming parity fails because "it works the same way" is rarely true. Subtle differences in timing, consistency, error handling, and edge cases cause production issues testing missed. Run parallel implementations comparing results. Don't declare victory until production traffic has run for weeks.
Ignoring operational readiness fails when the new system works but nobody knows how to operate it. Alerts aren't configured, runbooks don't exist, 3am incidents find teams unprepared. Build operational capability alongside migration: train the team, set up monitoring, write runbooks.
Feature creep fails by turning migration into transformation. "While we're migrating, let's also add these features" expands scope, slips timelines, loses momentum. Migrate first, improve second.
Hybrid architectures
If you're uncertain, start hybrid. This path proves value with minimal commitment and builds expertise for future decisions.
Deploy Cloudflare at the edge while keeping existing backends. Workers handle authentication, rate limiting, and caching. Backends remain unchanged. You've added capability without removing anything that works.
Static assets move to R2 for immediate egress cost savings with no backend changes, just storage migration. The savings often fund further experimentation.
Edge personalisation (A/B testing, geographic customisation, response transformation) adds value without backend modification. Workers transform responses; backends serve content.
From this foundation, evaluate further migration with data rather than projections. You understand the development experience. You've measured operational impact. The decision to migrate more, or stay hybrid permanently, rests on evidence from your environment, not spreadsheet estimates.
Staying hybrid permanently is valid. Many architectures are optimal with Cloudflare handling edge concerns and hyperscalers handling backend compute. Hybrid isn't a transition state you must exit; it's an architecture that may be right indefinitely.
Hybrid security architectures
Hybrid doesn't only mean hybrid compute. You can deploy Cloudflare's networking and security products in front of existing infrastructure without migrating compute at all.
Deploy Cloudflare's WAF, DDoS protection, and CDN in front of existing AWS/Azure/GCP infrastructure without migrating compute. This delivers immediate value while building organisational familiarity with Cloudflare operations.
Consider an organisation with significant AWS investment. The Developer Platform evaluation might conclude migration is premature: team expertise, existing commitments, and risk tolerance all favour staying on AWS. But that doesn't preclude using Cloudflare's WAF, DDoS protection, and CDN in front of AWS origins.
This pattern delivers immediate value while building organisational familiarity. Traffic flows through Cloudflare's network for performance and security, then reaches AWS backends unchanged. The AWS investment remains protected, the organisation gains Cloudflare operational experience, and the path to Developer Platform adoption becomes clearer.
When you eventually deploy Workers, they integrate naturally with security products already in place. WAF rules configured for AWS origins work identically for Worker-based applications. Rate limiting policies transfer. Access controls apply. Each step adds capability without compromising what came before.
This is the strategic case for Cloudflare beyond any single product category: security products, networking products, and the Developer Platform share operational model, configuration interface, and deployment philosophy. Adopting any subset creates familiarity with the whole.
Control the release boundary
A Worker deployment can serve users worldwide. Regional isolation will not limit a faulty release unless you have deliberately created such a boundary. The release strategy must supply the containment.
Gradual deployments split traffic between versions, so global availability does not require exposing every request to new code at once. Promote against error rates and latency, keep a tested rollback path, and consider how a rollback interacts with writes already made. Rolling back code does not reverse a database migration or an external side effect.
The relevant question is how many users and which state a release can affect before you detect failure. Answer it in the deployment design, rather than assuming geography will answer it for you.
Organisational readiness
Before expanding beyond the pilot, settle three commitments. Finance needs a way to approve and monitor variable usage. Engineering needs time to learn the platform while maintaining the current system. Operations needs an owner for the additional deployment and incident-response surface.
Keep the decision record independent of its sponsor: the workload, measured benefit, remaining risks, expected cost and review date should survive a change of personnel. Adoption that depends on one advocate is difficult to sustain and difficult to question.
Evaluate the complete request path
Cloudflare's security and networking products may form part of the design, but their presence in the same account does not mean every request passes through them. Record the public hostname, Worker entry point, private bindings and outbound path; then identify which controls apply to each.
Price and validate additional products against that path. An Access policy on one hostname, for example, should not be treated as proof that every route to the application is authenticated. Chapter 22 covers identity verification, private connectivity and deployment authority; Chapter 21 covers operational evidence.
Making the decision
Proceed when the pilot demonstrates a useful improvement, the workload fits the limits, and a named team can support it. Record the switching costs you accept and the evidence that would make you reconsider.
Keep a hybrid architecture when the useful boundary lies between systems: Workers close to users, an existing database close to its dependent services, or R2 serving objects while compute stays elsewhere. There is no requirement to complete a migration that has already delivered its benefit.
If the evidence remains uncertain, test the uncertain assumption next. If it points away from Cloudflare, keep the result and stop the adoption work. Chapter 26 examines those limits in more detail. For workloads that fit, Chapter 3 begins with the execution model your production design will depend on.