Chapter 26: When Not to Use Cloudflare
When is Cloudflare the wrong choice?
Every platform has a shape. Previous chapters explored what fits within Cloudflare's shape; this chapter examines what doesn't. Understanding both separates useful guidance from marketing.
Cloudflare works best when its primitives match the domain. A room maps naturally to a coordination object; a large shared relational model may fit an external database better. Count the work the platform removes and the work its constraints introduce. That balance determines whether the design is worthwhile.
The shape of the platform
Shared isolates make small invocations economical. D1's small databases encourage independent data partitions. Those design choices matter more than any prediction about a future limit increase.
Treat the published limits as constraints on the system you can deploy. Their exact values can change; your architecture should not depend on an unannounced increase. If a required workload does not fit, partition it where the domain permits or choose another execution or storage service.
Separate a missing API from a mismatched data model. An API addition might remove a compatibility obstacle. A larger database allowance would not, by itself, make transactions across thousands of independent tenant databases easy. Assess the requirement that remains after the hoped-for change.
Hard limits and what they mean
Before evaluating architectural fit, check the binary filters: failing any doesn't mean Cloudflare is bad but rather your workload doesn't match the platform's assumptions.
| Constraint | Threshold | If You Exceed It |
|---|---|---|
| Memory per isolate | 128 MB | Containers (up to 12 GiB) or external compute |
| CPU time per paid HTTP invocation | 5 minutes maximum; 30 seconds by default | Partition the work, use Containers, or use external compute |
| Database size (single) | 10 GB | Multiple D1 databases or Hyperdrive to external DB |
| Container resources | 4 vCPU / 12 GiB RAM | Hyperscaler compute |
| Inbound protocols | HTTP and WebSocket; inbound TCP in private beta via Spectrum; UDP not available; WebRTC via Realtime | Traditional cloud with load balancers for UDP or generally available TCP ingress |
| Cross-partition transactions | None | Saga patterns via Workflows, or external database |
Need more than 12 GiB RAM per instance? Need inbound UDP connections (other than WebRTC)? Need true archival storage under $0.001/GB-month? Need cross-database atomic transactions? These aren't "hard but possible"; they're "Cloudflare cannot do this." Check binary filters before investing in detailed evaluation.
Memory: the isolate trade-off
Workers have 128 MB of memory per isolate, shared across concurrent requests, blocking image processing that loads large files, ML inference with substantial models in-process, and any operation building large in-memory data structures.
Workers aren't bad at memory-intensive work; they're simply not designed for it. The isolate model optimises for high concurrency with moderate per-request resources, not resource-intensive individual operations.
When you hit this limit, Containers offer up to 12 GiB per instance; stream processing avoids the problem by never loading complete data into memory; external services let you offload specific operations while keeping edge logic in Workers. The question isn't whether 128 MB is "enough" but whether your workload's memory profile matches the isolate model.
CPU time: compute has a clock
Workers have a maximum CPU time of five minutes on paid plans and thirty seconds by default. CPU time means actual compute (cycles spent processing, not wall time spent waiting), a distinction that matters because Workers are designed for I/O-heavy workloads.
The limit blocks video transcoding, complex simulations, and large data transformations; Containers have no CPU time limit within their wall-time constraints, and Queues can help for batch processing (consumers have fifteen minutes of wall time and can spread work across many messages).
But consider whether heavy compute at the edge makes sense. Edge execution's value is latency reduction through geographic proximity, while compute-intensive batch processing benefits from resource availability, not edge proximity. Running heavy compute on Cloudflare to "stay on platform" may optimise for the wrong thing.
Database size: horizontal by design
Each D1 database is limited to 10 GB. The paid account limit of 50,000 databases does not imply 500 TB of available storage: the default account storage limit is 1 TB. Evaluate database count, per-database size and account capacity separately. The architectural question is whether a transaction and query boundary of one small database fits the domain.
The question is whether horizontal partitioning fits your data model. Multi-tenant SaaS with database-per-tenant works beautifully, as do consumer applications with database-per-user. But a data warehouse aggregating across all tenants or analytical workloads requiring joins across partitions fight the model.
When 10 GB per queryable unit doesn't fit or horizontal partitioning doesn't align with your data model, Hyperdrive provides a production-ready alternative: connecting Workers to existing PostgreSQL or MySQL databases with connection pooling and query caching at the edge. This isn't a compromise or migration stepping-stone but rather a valid permanent choice; many production systems run with Hyperdrive as their data layer. Choose between D1's horizontal model and Hyperdrive's external database connectivity based on your data architecture, not on one being inherently superior.
Inbound UDP, and TCP only in beta: the proxy architecture
Workers and Containers receive end-user traffic through Cloudflare's HTTP proxy, not as raw inbound sockets. There is one narrow exception, in private beta: a connect() handler in the Workers runtime lets Spectrum hand an inbound TCP socket directly to a Worker, which can forward it to a gRPC server in a Container for full-duplex, any-language gRPC. It is TCP-only.
Inbound UDP has no equivalent, which blocks game servers requiring UDP for latency-sensitive state synchronisation and IoT protocols speaking UDP without an HTTP bridge. The exception for real-time media is Cloudflare Realtime, which handles WebRTC for audio and video; if your non-HTTP requirement is specifically real-time media, Chapter 11 covers it.
The honest position: HTTP and WebSocket are native, inbound TCP is a private beta through Spectrum, inbound UDP is unavailable, and WebRTC media has a home in Realtime. If you need generally available inbound TCP or any inbound UDP, that ingress lives in traditional cloud infrastructure.
No distributed transactions
D1 provides transactions within a single database. No transactions span multiple D1 databases, D1 and R2, or D1 and external systems.
The horizontal scaling patterns the platform encourages (database-per-tenant, database-per-user) create exactly the scenario where you might want cross-database transactions. When user A transfers funds to user B and each has their own database, no atomic operation spans both.
A saga can coordinate separately committed actions through Workflows. Define and register the compensations the business permits, and give failed compensation an operational owner. Intermediate states remain visible; durable progress does not provide transaction isolation.
Write down the invariant before partitioning. If a transfer must debit one balance and credit another atomically, keep both within a suitable transaction boundary. Repartitioning the data or retaining a shared database can be simpler than explaining why partially completed transfers are acceptable.
Red flags: when to look elsewhere
Beyond hard limits, certain project characteristics signal poor platform fit. Consider looking elsewhere if:
"We need to migrate our existing application without changes." Cloudflare rewards native design. Lift-and-shift migrations fight the model at every turn. If you're not willing to rearchitect, hyperscaler VMs or managed containers provide straightforward paths.
"Our data model assumes a single large database." D1's horizontal model deliberately rejects this assumption. However, this doesn't disqualify Cloudflare; Hyperdrive connects Workers to external PostgreSQL or MySQL databases, letting you keep your existing data architecture while gaining edge compute benefits. The question becomes whether you want D1's model or prefer connecting to external databases via Hyperdrive.
"We need atomic transactions across tenant boundaries." Saga patterns provide eventual consistency, but if regulatory or business requirements demand true atomicity across partitions, the complexity cost may exceed any platform benefit.
"Our protocol requires direct UDP connections." Non-negotiable, unless that protocol is WebRTC for audio/video (which Cloudflare Realtime handles).
"Our team has five years of AWS expertise and three months to deliver." Platform fit isn't just technical. If the team is productive on their current platform and the workload doesn't strongly favour edge execution, forcing a platform change may destroy more value than it creates.
"We're already deeply integrated with cloud-native services." A machine learning pipeline using SageMaker, S3, Lambda, and Step Functions has dependencies throughout the AWS ecosystem. Extracting this means replacing not just compute but the integration fabric connecting everything.
If several of these describe your situation, Cloudflare probably isn't your answer. That's not because it's inferior, but because the fit isn't there.
A large shared database with cross-tenant transactions is a poor candidate for arbitrary partitioning into D1 databases. Workers can still use that database through Hyperdrive. Evaluate the compute and data decisions separately, and reject a migration when the required changes cost more than the benefits they enable.
Different versus worse
Most platform mismatches aren't limitations but differences: Cloudflare doesn't do X poorly but rather does Y instead. The question is whether X or Y fits your workload.
If you're evaluating Cloudflare after years on AWS, Azure, or GCP, you carry accumulated intuitions about systems that are sometimes universal principles and sometimes adaptations to your current platform's constraints, constraints that may not apply here.
Separate the requirement from its existing implementation. Provisioned concurrency is one response to startup latency; on Workers, measure isolate and application initialisation plus backend calls before deciding what control remains necessary. Bindings simplify resource access and credentials, but data location, connection capacity and query behaviour still matter. Likewise, a requirement for one large transaction boundary deserves an explicit database choice rather than automatic partitioning.
None of these intuitions are wrong in their original context. But applying them uncritically leads to "Cloudflare can't handle our workload" when the accurate conclusion is "Cloudflare handles our workload differently."
The test is whether the requirement survives a change of platform. A user-visible latency target does; a particular provisioned-concurrency setting may not. Needing 2 GB of simultaneously resident data is a resource requirement that the Workers isolate model cannot satisfy.
The "almost fits" problem
Clear mismatches are easy (if you need 64 GB of RAM per instance, Cloudflare obviously doesn't fit), but harder cases are partial matches where workloads fit the platform 90% of the time but have components that don't.
An e-commerce platform might handle browsing, sessions, carts, and checkout beautifully on Workers with Durable Objects but require more memory than Workers provide for generating PDF invoices from complex templates. A real-time collaboration tool might handle presence and synchronisation perfectly in Durable Objects but exceed the memory limit when generating document thumbnails.
Three options exist:
Carve out to an external service. PDF generation on a hyperscaler Lambda, image processing on dedicated infrastructure. This adds operational complexity and latency for cross-service calls.
Use Cloudflare Containers. Keep everything on platform but accept Containers' different characteristics: slower startup, fewer locations, different pricing.
Redesign the feature. Perhaps thumbnails can be generated client-side, or PDF generation can be simplified to fit within constraints.
Choosing between them
| Factor | Favour External Service | Favour Containers | Favour Redesign |
|---|---|---|---|
| Latency tolerance | Can accept 50-200ms cross-service overhead | Need <50ms but can accept Container startup | Latency-critical, must stay in Workers |
| Operational appetite | Team comfortable with multi-platform ops | Prefer single platform even with complexity | Prefer simpler architecture overall |
| Feature centrality | Peripheral feature, rarely invoked | Core feature, frequently used | Could work differently without loss |
| Existing solution | Already have this running elsewhere | Starting fresh | Open to alternatives |
The less attractive option is forcing the component to fit through increasingly complex workarounds. If you're building elaborate streaming pipelines to avoid memory limits for conceptually simple operations, consider whether the architecture fits. Either accept the external dependency or redesign the feature.
The lock-in spectrum
Not all Cloudflare usage creates equal platform dependency. Evaluate lock-in across three levels:
Low lock-in: portable by design
R2 is S3-compatible at the API level. Code using AWS SDK against R2 works against S3 with configuration changes. Data exports trivially. Workers running stateless logic with standard JavaScript are similarly portable; platform-specific bindings need replacement, but core logic transfers. KV data exports through straightforward APIs.
Medium lock-in: portable with effort
D1 exports SQL that can recreate its schema and data; it does not export a ready-to-use SQLite database file. The schema and queries still need adaptation for a different SQL engine. More consequentially, moving ten thousand tenant databases requires a plan for provisioning, routing and migration progress, even if every row is portable.
Queues implement standard patterns but with Cloudflare-specific APIs. The pattern exists everywhere, but you'd rewrite integration code for SQS or Azure Service Bus.
Workflows provide durable execution with step-based guarantees. The concept maps to Step Functions or Temporal, but migration requires understanding both systems deeply enough to translate semantics, not just syntax.
High lock-in: architectural commitment
Durable Objects combine an actor model, durable storage, routing and platform-specific concurrency guarantees. Other systems offer actor and durable-execution models, but moving requires mapping these guarantees to the replacement. Exporting the stored records does not reproduce placement, alarms, WebSocket behaviour or output gating.
If Durable Objects handle core coordination in your system, you've made a deep platform commitment. Exit cost isn't data migration; it's architectural redesign.
When to accept each level
Lock-in isn't a problem to avoid but rather a trade-off to make consciously.
Accept high lock-in when the capability enables something competitors can't match. If Durable Objects' coordination primitives let you build real-time features that would be prohibitively complex elsewhere, lock-in is the price of competitive advantage.
Accept medium lock-in when the platform's approach fits your domain better than alternatives. If database-per-tenant with D1 simplifies your multi-tenant architecture, migration complexity on exit is worth operational simplicity now.
Prefer low lock-in for commodity functionality where platform differences don't create advantage. Object storage is object storage; R2's S3 compatibility gives you Cloudflare's pricing without meaningful commitment.
When hyperscalers are better
Sometimes the right answer is AWS, Azure, or GCP; not because Cloudflare is inferior, but because hyperscalers fit better.
Deep ecosystem integration
If your architecture deeply integrates cloud-native services, switching cost may exceed any Cloudflare benefit. A machine learning pipeline using SageMaker for training, S3 for data storage, Lambda for preprocessing, and Step Functions for orchestration has dependencies throughout the AWS ecosystem. Each service connects to others with native integrations, shared IAM, and unified tooling.
The same applies to BigQuery-centred analytics stacks, Cosmos DB applications using specific consistency models, and systems built around Azure AD B2C for enterprise identity. These aren't services you happen to use; they're architectural foundations.
Cloudflare can complement these architectures. Workers at the edge can handle caching, authentication, and global routing while backends remain in the hyperscaler. But replacing core services requires justification beyond "Cloudflare is good."
Regulatory and contractual requirements
Some decisions aren't technical. Government contracts may require FedRAMP-authorised infrastructure. Healthcare systems may require particular data-handling controls and contractual assurances. Financial services may require particular audit capabilities.
Cloudflare publishes compliance reports, certifications and data-protection commitments, but their scope may not satisfy every requirement. Verify compliance requirements before platform selection, especially for regulated industries or government contracts.
Existing use of Cloudflare can help with procurement, but adopting application compute or storage changes the services, purposes and data flows under review. Confirm that the relevant agreements and controls cover those changes. Prior approval for a CDN or security product does not settle every residency or access requirement for the Developer Platform.
For enterprises with strict data residency requirements, the Data Localization Suite provides controls over where TLS keys are stored, where requests are processed, and where metadata persists. Match each control to the required data flow, product compatibility and contractual scope. Request processing controls do not by themselves establish where every bound store, backup or external service keeps data. See Chapter 22 for detailed coverage of DLS capabilities and constraints.
Heavy GPU compute
Workers AI provides inference at the edge, running models Cloudflare hosts. But Cloudflare offers no general-purpose GPU compute for training models, running custom CUDA workloads, or GPU tasks beyond supported inference. If your workload requires GPUs, hyperscaler instances or specialised ML platforms remain your options.
Specific hyperscaler advantages
| If You Need... | Consider... | Because... |
|---|---|---|
| ML training and custom models | AWS SageMaker, GCP Vertex AI | GPU compute, MLOps tooling |
| Massive analytical queries | BigQuery, Redshift, Synapse | Columnar storage, query optimisation at scale |
| Globally consistent relational transactions | Spanner | Distributed transactions with external consistency |
| Writes in multiple regions | Cosmos DB | Configurable consistency and conflict handling; multi-region writes do not support its strong consistency mode |
| Streaming data pipelines | Kinesis, Pub/Sub, Event Hubs | Native integration with analytics ecosystem |
| Lift-and-shift migration | EC2, GCE, Azure VMs | Minimal application changes |
Hybrid architectures
The choice isn't always binary. Hybrid architectures use Cloudflare for what it does best while keeping workloads elsewhere where that makes sense.
Where to draw the boundary
A reasonable default: put authentication, caching, rate limiting, and request routing at the edge. Put business logic that doesn't benefit from edge proximity in your backend of choice.
Latency sensitivity favours the edge. Operations where round-trip time to a centralised backend degrades user experience (authentication, personalisation, A/B testing, content transformation) benefit from edge execution.
Data locality favours keeping compute near data. If an operation requires substantial data from a centralised database, moving compute to the edge means moving data across the network. Sometimes edge caching addresses this; sometimes the operation belongs near the data.
Complexity favours simplicity. If an operation requires elaborate workarounds to fit Workers' constraints, complexity cost may exceed latency benefits.
Common hybrid patterns
E-commerce platform: Workers handle product browsing, user sessions, cart management, and checkout flow. Order processing, inventory management, and fulfilment integration run on existing hyperscaler infrastructure with Hyperdrive connections.
SaaS application: Workers handle authentication, tenant routing, and API gateway functions. Core business logic runs on existing infrastructure, with gradual migration of suitable endpoints to Workers as the team builds expertise.
Content platform: Workers handle content delivery, personalisation, and edge caching. Content ingestion, processing pipelines, and search indexing run where data already lives.
The migration gradient
Full platform adoption isn't required for platform benefit. Start with edge concerns that provide immediate value regardless of backend: caching reduces backend load, authentication moves security enforcement forward, rate limiting protects against abuse. These Workers add value alongside any backend without requiring backend changes.
New features can be built natively on Cloudflare while legacy systems remain elsewhere. Migration happens feature by feature, not big-bang. A hybrid steady state, some workloads on Cloudflare and some on hyperscalers, may be the right long-term architecture, not a transitional phase. When you're ready to migrate specific workloads, Chapter 27 provides detailed playbooks for common scenarios.
Exit planning
The best time to plan your exit is when you have no intention of leaving.
What transfers cleanly
Application logic in JavaScript or TypeScript transfers to any JavaScript runtime. Platform-specific bindings need replacement, but logic is portable. D1 exports SQL dumps that can be imported into SQLite or adapted for another database. R2 exports through S3-compatible tooling. Your data isn't trapped.
Architectural knowledge transfers. Teams that understand edge execution, distributed state management, and horizontal scaling apply that knowledge regardless of platform.
What requires reconstruction
Moving from Durable Objects means mapping its execution guarantees to the replacement actor or state-machine system. Prototype one representative entity, including alarms, reconnects and failure recovery. That reveals more about the migration than a successful export of its stored records.
The coordination patterns enabled by Durable Objects (rate limiting, session management, real-time collaboration) have implementations elsewhere, but they're different implementations. Expect to rebuild rather than migrate.
Reducing exit cost
If exit optionality matters, isolate platform-specific code behind abstractions. A service calling env.DB.prepare() directly throughout is harder to migrate than one where database access flows through an abstraction. This adds development overhead and may not be justified if platform commitment is intentional, but it preserves options.
Document architectural decisions explicitly. Future teams considering migration need to understand not just what the system does but why it was built this way.
Making the decision
Start with the requirements that cannot change: resource needs, transaction boundaries, protocols, data obligations and existing dependencies. Determine which Cloudflare primitives fit those requirements and which components should remain elsewhere.
"One object owns a room's conversation" gives the object a clear job. "Every room update needs locks across several objects" should send you back to the boundary you chose. The extra coordination may be justified, but count it as part of the design rather than hiding it behind the platform's scaling promise.
Then test the operating model. Can the team explain the failure paths, diagnose a slow request and recover state after a bad deployment? Include that work and the eventual cost of moving away in the comparison.
Choose Cloudflare where the resulting design is simpler to build and operate. Keep a hybrid or choose another platform where it better serves the workload. A clear boundary is more useful than forcing every component onto one provider.
What comes next
If the workload fits, the next question is how to move it without losing the ability to recover. Chapter 27 covers coexistence, data reconciliation and cutover: the parts of migration that a successful deployment does not prove.