Skip to main content

Chapter 12: Choosing the Right Storage

Where should I store this data, and what trade-offs am I making?


Parts II and III covered compute: Workers, Durable Objects, Workflows, Queues, Containers, and Realtime for media. Part IV turns to where that compute stores its data.

D1’s 10 GB database limit makes partitioning an early design question. Independent tenant datasets can fit naturally into separate databases; a shared relational dataset may belong in an external database reached through Hyperdrive. Start with the boundaries in your data, then choose the storage.

Why Five Storage Options

Traditional databases try to be everything: relational queries, key-value access, blob storage, real-time subscriptions. This works in a single location but fails at the edge, where "single location" is the opposite of the point. No single storage abstraction performs well for all access patterns globally. Global reads require different architecture than relational queries; large objects need different handling than coordination state.

Cloudflare's answer: specialisation. KV for globally cached reads, D1 for relational queries, R2 for objects and files, Durable Objects for coordination, Hyperdrive for existing external databases. Match the primitive to your access pattern, or pay the tax of fighting the platform.

The five primitives​

Each store makes a different promise about access, freshness and coordination. Start with the promise the application needs, then verify the limits and cost of meeting it.

KV trades freshness for inexpensive cached reads near users. Hot keys can be served locally; cold keys require a fetch through KV’s storage hierarchy. Updates can remain invisible at another location for 60 seconds or longer, depending on cached values. Choose it when that stale read is acceptable.

D1 provides managed SQLite databases with optional read replicas. The primary owns writes; readers choose between reaching that authority and using replicated data through the Sessions API. Its 10 GB per-database limit makes the size of the application's transaction boundary an early fit check.

R2 stores files and blobs through an S3-compatible object API, with objects up to 5 TB and no egress charge. The object API retrieves by key rather than querying relationships; analytical tables in R2 Data Catalog and R2 SQL are a separate choice, covered in Chapter 14.

Durable Objects storage keeps an object's state with its coordination logic. SQLite storage and runtime gates help preserve invariants, while the application must account for interleaving across external awaits. Choose an object when one active owner simplifies decisions or connections for that entity; its capacity remains finite.

Hyperdrive trades native integration for compatibility. It accelerates existing PostgreSQL and MySQL databases without migration: connection pooling at the edge, query caching for repeated queries. This is a valid long-term architecture, not merely a bridge. Many production workloads run permanently with Hyperdrive connecting Workers to external databases, particularly when those databases require PostgreSQL-specific features, exist within established operational practices, or serve as a shared data layer across multiple applications.

Choosing partition boundaries​

D1 and Durable Objects put a practical boundary around each unit of state. A tenant database or a room object works well when most operations stay inside it. Splitting a tightly connected dataset across arbitrary keys creates coordination and query work for the application.

Choose a boundary that contains the transactions you need. Then test its largest plausible size, busiest write path and cross-partition queries. Database-per-tenant is useful when tenants are independent; a shared database remains sensible when shared queries dominate and the workload fits its limits.

Chapter 13 examines the operational cost of a database fleet. Horizontal scaling removes one bottleneck by creating a distribution problem you must be prepared to own.

The decision framework​

Assess access pattern, consistency requirements and coordination needs together. A cheap read path is useful only if its freshness and query behaviour meet the requirement.

Access pattern​

Start with how the data will be read and written.

Read-heavy workloads with rare changes are candidates for KV when cached data may be stale. The ratio of reads to writes helps with cost estimation; the consequence of a stale read decides suitability. A feature flag that changes once a year can still require immediate enforcement.

D1 is a candidate for changing data that needs SQL queries. Reads outside the Sessions API go to the primary. Replicated reads use a session, whose bookmark can carry a caller's progress across requests; the application must preserve and pass it. Chapter 13 explains how to choose the initial read constraint.

Frequent mutations make KV's per-key write limit, operation cost and stale reads more significant. Compare D1 for relational updates and Durable Objects for application-owned coordination. Count the rows, indexes and active duration each design would consume.

Never Store Files in D1

Large objects belong in R2 regardless of access pattern. Files, images, documents, exports; anything measured in megabytes; go to R2. SQLite handles blobs poorly, and you'll hit row size limits.

Ask your team: Which reads can return an older value without violating a user-visible promise? Price KV for those paths.

Consistency requirements​

Separate freshness from availability. KV can answer from an older cached value; a request to authoritative state may need to wait for its owning location. Neither service guarantees that every request succeeds. Decide what stale data can cause before trading freshness for shorter reads.

KV provides eventual consistency. A successful write does not mean every location can read it yet; cached values and cached misses can remain visible for 60 seconds or longer. Routing changes can also expose an older value after a user has already seen a newer one.

Safe for data where brief staleness doesn't cause bugs: configuration, cached API responses, feature flags. Dangerous for data where reads must reflect recent writes: counters, balances, inventory levels, anything where concurrent operations must coordinate.

D1 writes become visible on the primary. Replicas can lag; the Sessions API provides sequential consistency within the session. Carry its bookmark when later requests must continue from that progress, or start at the primary when the decision requires current state.

Durable Object storage is strongly consistent, with input and output gates governing storage operations and outgoing messages. This does not make an arbitrary asynchronous handler one atomic transaction. Chapter 7 explains where external awaits introduce interleaving and where an explicit transaction is needed.

Ask your team: Where do you rely on reading data immediately after writing it? Those paths need D1 or Durable Objects, not KV.

Coordination requirements​

Many workloads need more than storage; they need coordination. "Has this user exceeded their rate limit?" requires reading a counter, comparing to a limit, and incrementing; atomically, with no race condition. "What's in this user's shopping cart?" requires immediate consistency as items change. "Who's currently in this chat room?" requires tracking connections and broadcasting messages.

Keep the Check with the Update

A separate read followed by a write can race in D1. A conditional SQL update can make a reservation atomic: update only while capacity remains, then inspect whether a row changed. Use the database transaction boundary when it contains the invariant; use a Durable Object when an active owner also needs to coordinate connections or application state.

A Durable Object puts the counter and the check in one place. A synchronous check-and-update, or a correctly scoped storage transaction, can preserve the invariant. External awaited I/O can allow another request to run, so single-threaded execution does not make an entire asynchronous handler indivisible.

The trade-off is the capacity of one object. Partition independent limits by user or resource. A strict global counter still needs one authority, or a design that explicitly accepts approximation; adding objects does not preserve a global invariant automatically.

Ask your team: Which decisions would become simpler with one active owner? Compare that design with the conditional updates and transactions your database already provides.

Durable Object storage versus D1​

When building with Durable Objects, should the object store state internally (using private SQLite storage) or externally (in D1)?

Use internal storage when data belongs exclusively to that object and doesn't need cross-entity queries. A rate limiter's counter, a chat room's message history, a user session's state; these are scoped to a single object and never queried across objects. Internal storage keeps coordination and data together, simplifying architecture.

Use D1 when you need cross-entity queries or when data outlives the coordination pattern. "Find all chat rooms with more than 100 messages" or "list all sessions for this user" require D1's query capabilities. If you want historical data queryable after real-time coordination ends, flush state to D1 periodically.

In that hybrid, keep the object authoritative and D1 a projection that can be replayed and reconciled. Track the last exported version and tolerate repeated exports. If ownership must transfer, define the handover explicitly; the age of a record does not establish which copy wins.

Geographic considerations​

Durable Objects have different geographic characteristics than Workers. Workers run in over 300 cities worldwide. Durable Objects require heavier infrastructure for storage and coordination guarantees, so they're deployed to fewer locations.

For users in North America, Europe, and much of Asia, this rarely matters; Durable Object locations are dense enough that latency stays low. For users in parts of Africa, South America, and other regions with sparser infrastructure, Durable Objects may add noticeable latency.

If your users concentrate in well-served regions, Durable Objects work well for latency-sensitive coordination. If you're serving users globally including underserved regions, verify latency characteristics before committing to a Durable Objects-heavy architecture.

D1 and KV have different placement models. KV stores values centrally and caches them where they are read; a cold read may travel farther. D1's primary may be distant, while eligible session reads can use replicas closer to users. Compare the paths your application actually uses.

Common patterns​

Abstract frameworks clarify; concrete examples confirm. These patterns show the decision framework applied to real scenarios.

Configuration and feature flags​

KV suits frequently read configuration whose propagation delay is acceptable. An emergency stop, revoked entitlement or release permission needs an authoritative check when stale execution would cause harm. Decide that distinction before putting every feature flag behind the same cache.

A D1 query can cost less than a KV read, depending on rows scanned and included allowances. KV earns its place through cached access near the caller. Benchmark both paths when configuration latency matters, and keep authoritative entitlements out of a cache that can delay revocation.

User sessions​

Sessions are read-heavy (every authenticated request reads the session) with rare writes on login, logout, and refresh. This suggests KV, and for most applications, KV works well.

The nuance is revocation. A logged-out session cached in KV may remain usable elsewhere until that cached value expires. Short token lifetimes bound exposure but do not provide immediate logout. If revocation must be visible on the next request, consult authoritative state on that request.

Choose the acceptable revocation window explicitly for the operation. Reading a public profile and transferring money need not use the same session-validation path.

User profiles and application data​

Profiles need capabilities only D1 provides: queries to find users by email, filters by status, pagination for lists, updates as users change settings, relations connecting users to posts, orders, and subscriptions. D1 exists to serve this relational pattern.

KV can't query across keys. R2 is for files, not records. Durable Objects are overkill for data that doesn't need coordination. D1 is the relational database for the platform; relational data belongs there.

For global applications where read performance matters, D1 with read replicas provides the best balance: strong consistency for writes, low-latency reads from nearby replicas, and the sessions API ensuring users see their own changes immediately.

File storage​

Files and large binary objects usually belong in R2. Presigned URLs let clients transfer directly; a Worker can still authorise the operation before issuing the URL. Worker-mediated streaming avoids buffering the whole file, but direct transfer also avoids the Worker’s inbound request limits and invocation path.

R2 removes outbound bandwidth from the object-store bill. Savings depend on the existing delivery route, request count and storage class; Chapter 14 separates those components.

Never store files in D1. SQLite handles blobs poorly, row size limits constrain you, and database backups balloon with binary data. Store file metadata in D1 if you need to query it; store files themselves in R2.

Rate limiting and counters​

Counters seem like simple key-value data but require coordination. A rate limiter must atomically check the current count, compare against a limit, and increment; no race condition between steps. KV's eventual consistency creates races: two concurrent requests both read "99", both allow their operations, both write "100", exceeding a limit of 100 at 101.

A Durable Object can keep the check and increment together in synchronous application code backed by its own storage. Avoid yielding to an external service between them; another request can run during that await.

Pattern: one object per entity being limited. One rate limiter per user, one counter per resource. Each object handles all operations on its entity serially.

Real-time features​

Chat rooms, collaborative documents, and multiplayer games need to track connections, broadcast messages, and maintain shared state. This means Durable Objects. The hibernatable WebSockets API holds thousands of connections efficiently while single-threaded execution keeps state consistent as messages arrive and users join or leave.

This isn't primarily a storage decision; it's a compute decision. You choose Durable Objects for the execution model and WebSocket support. Storage serves the coordination.

The "just use Postgres" question​

Technical leaders will ask: why not use Aurora or Cloud SQL via Hyperdrive for everything? PostgreSQL is battle-tested, teams know it, and Hyperdrive makes it accessible from Workers.

This is a legitimate architecture. Many production systems run Workers with Hyperdrive connecting to external PostgreSQL or MySQL databases as their permanent data layer. The question isn't whether this works; it does, and well; but what trade-offs you're making.

Latency considerations. Hyperdrive eliminates connection overhead through pooling and reduces round-trips through query caching, but your database still lives in one region. A D1 database with read replicas distributes reads globally; a PostgreSQL database in us-east-1 serves reads from us-east-1 regardless of where your user sits. For read-heavy workloads with global users, this latency difference matters. For workloads concentrated near your database region, or where database queries aren't in the critical path, it may not matter at all.

Cost trade-offs. External databases have their own pricing: instance hours, storage, I/O operations, potentially egress. D1's pricing scales with actual usage without idle instance costs. For variable workloads, D1's model is often cheaper. For steady workloads with reserved capacity, the economics may favour your existing database.

Operational considerations. An external database means managing upgrades, backups, monitoring, and scaling decisions; but your team may already have established practices for this. D1 eliminates this operational burden but introduces a new system to learn. The question is whether operational simplicity or operational continuity matters more for your team.

Capability differences. PostgreSQL offers features D1 lacks: extensions (PostGIS, pg_trgm, pgvector), stored procedures, advanced data types, and decades of ecosystem tooling. If your workload uses these capabilities, Hyperdrive lets you keep them. If you're using PostgreSQL as "SQL that happens to be PostgreSQL," D1 may serve equally well.

Hyperdrive is the right permanent choice when your database requires PostgreSQL-specific features, when your team has established operational practices around PostgreSQL, when your database serves multiple applications (not just Workers), or when migration risk exceeds migration benefit. It's also the right starting point when you're uncertain; you can always migrate to D1 later if the trade-offs favour it.

D1 is the right choice for new projects without PostgreSQL dependencies, for applications benefiting from global read distribution, and for teams wanting fully managed operations within Cloudflare's ecosystem.

Ask your team: What PostgreSQL features do you actually use that SQLite lacks? The answer guides whether Hyperdrive or D1 fits better; not which is "better" in the abstract.

A concrete architecture​

Consider a project management SaaS application. How would its data distribute across Cloudflare's storage options?

Tenant configuration can use KV when stale values are acceptable. Separate display preferences from plan limits and access controls: their revocation requirements may need authoritative reads. A tenant key organises data but does not enforce who may read or change it.

Project and task data lives in per-tenant D1 databases. Each tenant gets their own database, enabling queries within a tenant while maintaining isolation. The 10 GB limit is generous for most tenants; large enterprises might shard by project or time period. Read replicas provide fast global reads for distributed teams.

File attachments live in R2. Task attachments, project documents, and exports are stored as objects with presigned URLs for direct upload and download. D1 stores metadata for querying; R2 stores the bytes.

Real-time presence (who's viewing this project, where their cursor is) lives in Durable Objects. One object per project maintains WebSocket connections, tracks active users, and broadcasts cursor positions. This state is ephemeral; when everyone disconnects, the object sleeps. For collaboration features like simultaneous document editing, the Durable Object coordinates changes before persisting results to D1.

External integrations might use Hyperdrive. If the application integrates with a customer's existing PostgreSQL database for data synchronisation, Hyperdrive accelerates those queries.

The design needs explicit ownership: D1 owns task records, R2 owns attachment bytes, and the project object owns live collaboration state. Caches and historical projections can lag; the application must know which store to consult when freshness matters.

Cost architecture​

Cloudflare prices storage based on what costs them resources.

KV charges more for writes than reads (roughly 10:1 at the marginal operation rates). Frequent reads of cached values fit that model well. Include writes, misses and freshness requirements when comparing it with another store; the price ratio alone does not describe application cost.

D1 charges per row read and written, not per query. A query scanning 10,000 rows costs ten times more than one scanning 1,000. Index optimisation becomes a cost concern, not just performance. An unindexed query scanning your entire table is expensive twice over: slow and costly.

R2 charges for storage and operations, not egress. This inverts the S3 model. For read-heavy workloads with significant egress, R2 saves dramatically.

Durable Objects charge for requests, compute duration, and storage. Pricing reflects the coordination guarantees.

Worked example​

An application serving 10 million daily requests, each requiring one configuration read and one database query averaging 100 rows.

KV for configuration: 10 million reads at $0.50 per million = $5/day. D1 instead: 10 million queries reading 100 rows each = 1 billion rows at $0.001 per million = $1/day, but with higher latency on every request.

D1's cost appears lower, but KV's sub-10ms reads versus D1's variable latency affects user experience. For configuration data that rarely changes, KV's caching model delivers better performance at acceptable cost.

For database queries, one billion rows daily totals 30 billion in a 30-day month. That is $30 at the marginal read rate, or $5 in overages if the account’s full 25-billion-row Paid allowance is available. Storage, writes and Worker execution remain separate. Compare that total with the external database capacity the application actually needs.

The meta-lesson: storage choices affect both performance and cost. Optimise for actual access patterns, not abstract capabilities.

Combining storage options​

Real applications use multiple storage types. The architectural question: how do they interact?

Configuration in KV, everything else in D1 is the simplest combination. KV holds feature flags, settings, and cached values that rarely change. D1 holds application data. No coordination needed; they serve different purposes.

Metadata in D1, objects in R2 separates queryable attributes from blob storage. Store file references, permissions, and descriptive fields in D1. Store actual files in R2. Query D1 to find files; retrieve files directly from R2 via presigned URLs.

Coordination in Durable Objects, projections in D1 keeps durable live state with the object and copies queryable records into D1. Define when that projection updates and how failed updates are replayed. A D1 outage must not silently lose the only copy of a completed action.

Cache in KV, source in D1 implements read-through caching. Check KV first, fall back to D1 on miss, populate KV on miss. Invalidate on write: update D1, delete or update KV. Accept brief staleness.

The key principle: clear ownership. One store is authoritative; others are caches or derivatives. Ambiguous ownership causes consistency bugs.

Storage mismatch failures​

Choosing the wrong storage produces characteristic failures. Naming them helps diagnose problems.

Consistency collision: using KV for data requiring immediate read-after-write visibility. You update a user's subscription status in KV, redirect them to the premium page, and it reads the old value because propagation hasn't completed. The user sees "not subscribed" seconds after paying. Fix: subscription status needs D1's consistency, not KV's performance.

Coordination gap: splitting a check and update across database round trips. Two callers can both read 99 before either writes 100. Use one conditional SQL statement or an atomic batch when SQL can express the invariant; use a Durable Object when the coordination needs application state or live connections.

Query impedance: storing structured data in KV or R2 then needing to query it. You stored user preferences as KV key-value pairs because reads are fast. Now you need "find all users with dark mode enabled." KV can't query; you're stuck iterating all keys. Fix: D1 for queryable data, with KV as a cache layer if needed.

Blob bloat: storing large objects in D1 or KV instead of R2. A 50 MB file in D1 makes your database unwieldy and hits size limits. Large values in KV add latency and hit the 25 MB limit. Fix: R2 for objects, with references stored in D1 or KV.

Partition avoidance: failing to embrace horizontal scaling. You try to store everything in one D1 database and hit the 10 GB limit. You create one Durable Object for all coordination and hit throughput limits. Fix: accept Cloudflare's horizontal model. Database-per-tenant. Object-per-entity.

Migration paths​

Storage choices aren't permanent. As applications evolve, storage needs change.

KV to D1: when you need queries or stronger consistency. If you're encoding structure into keys (user:123:profile, user:123:settings) and parsing them on read, you want a database. Migrate by reading all keys, transforming into rows, and inserting into D1.

D1 to Durable Objects: when coordination needs an application instance that owns the state, connections or decision sequence. Retry loops alone do not prove D1 is the wrong store. First ask whether the invariant fits in a conditional SQL update or batch; move to an object when the execution model earns its cost.

Hyperdrive to D1: an optional migration path, not an inevitable one. If you started with Hyperdrive and D1's model now fits better (perhaps your workload evolved towards multi-tenant patterns, or you want global read distribution), migrate incrementally by table or feature. But many systems stay on Hyperdrive permanently because the external database serves them well.

D1 to Hyperdrive: when you've discovered D1's horizontal model doesn't fit your evolved requirements. If you need cross-database joins, PostgreSQL extensions, or your data model has grown beyond what horizontal partitioning serves well, Hyperdrive connects you to PostgreSQL or MySQL without leaving Cloudflare's compute layer.

Decision checklist​

When uncertain, work through these questions.

QuestionCandidate
Does an existing database or required engine capability govern the choice?Hyperdrive with that database
Are these files or large binary objects?R2
Must application code own a coordinated entity or live connections?Durable Objects
Do relationships, transactions or queries across records matter?D1, if the dataset fits its partition model
Are keyed reads dominant and stale cached values acceptable?KV

Several answers may apply to different parts of the same application.

D1 is the default because it handles the widest variety of workloads adequately and is the easiest to migrate from. Data in D1 can move to KV for read performance, Durable Objects for coordination, or stay where it is. Starting with D1 keeps options open.

What comes next​

Chapter 13 examines D1's relational model and its partition limits. Chapter 14 follows with R2's object access and egress economics; Chapter 15 covers KV caching and retaining external databases through Hyperdrive. Use the framework here to decide which deserves a closer look.