Chapter 15: KV and Hyperdrive
How do I cache effectively and connect to existing databases?
Not all storage needs relational queries or object semantics. Sometimes you need fast key-value lookups for configuration, cached responses, or session data; sometimes you already have a PostgreSQL or MySQL database and migration isn't practical.
KV caches keyed values near readers. Hyperdrive reuses database connections and can cache query results. Both shorten some data access paths; neither makes an old result current or a distant database local.
Workers KV: a CDN you can write to
Think of KV as a CDN for data you control. Same global caching semantics, same eventual consistency, same sweet spot: content read far more often than it changes.
KV stores values centrally and caches reads through its hierarchy. A hot key can be served near the caller; a cold read takes a longer path. Updates and cached misses propagate eventually, so cached access is valuable only when stale data is acceptable.
The API reflects this simplicity:
// Decode JSON and set the edge-cache lifetime
const flags = await env.CONFIG.get("feature-flags", { type: "json", cacheTtl: 300 });
// Write with expiration
await env.CONFIG.put("session:abc123", JSON.stringify(session), { expirationTtl: 3600 });
The cacheTtl parameter controls how long the edge caches a value before checking for updates; you can set it as low as 30 seconds for applications needing fresher data, though the default of 60 seconds suits most use cases. Lower values mean more frequent checks against the coordination layer, trading read performance for freshness. The expirationTtl controls when the value itself expires and is deleted. Confusing them is a common source of unexpected behaviour.
What eventual consistency actually means
KV offers neither atomic read-modify-write operations nor guaranteed read-your-writes. If a decision depends on a current counter or the last accepted update, choose a store that can enforce that requirement. Waiting a fixed number of seconds is not a consistency protocol.
KV provides eventual consistency without monotonic reads; replicas converge given sufficient time without new writes, but a client might see older data after seeing newer data if requests route to different edge locations. User A in London might read value "2", and their next request to Paris might return an older cached value "1", showing regression not because anything went wrong but because KV makes no ordering guarantees across edge locations.
Do not model propagation as a fixed delay by continent. A cached value, including a cached miss, can remain visible for 60 seconds or longer, and same-location visibility is usually immediate but not guaranteed. There is no global consistency-complete check for a key.
Any pattern requiring two requests to agree on a current value is impossible with KV. Counters fail because two requests might both read "5", both write "6", and you lose an increment. Inventory checks fail because two requests might both see "1 remaining" and both decrement. Rate limiting fails because concurrent requests can't coordinate counts. For coordination, counters, or distributed agreement, use Durable Objects.
KV failure modes
Three failure modes recur frequently enough to deserve explicit names.
Stale read after write confuses teams regularly. You write a value, then immediately read it, and the read might return the old value because it was served from an edge cache that hasn't received the write yet. This isn't a race condition; it's expected. The write went to the coordination layer whilst the read came from edge cache; they don't talk to each other.
If you need to display what you just wrote, return the written value directly rather than re-reading it. If you need another request to see the write, consider whether D1 or Durable Objects better fits your consistency requirements. Adding delays is fragile because propagation timing isn't guaranteed.
Negative caching trap catches teams who check for key existence. When you read a key that doesn't exist, KV caches that non-existence at the edge; if you then create the key, some edges continue returning "not found" until their negative cache expires. There's no way to explicitly clear a negative cache entry, so the design-level workaround is not relying on key existence checks for correctness.
Read skew across locations can violate intuitions about consistency. User A in Sydney might see the new value whilst User B in Singapore sees the old value at the same moment, with each edge location serving its local cache state. For data where this inconsistency causes user-visible problems, consider D1 or Durable Objects.
Write failures and retry semantics
KV write operations can fail, and the failure semantics matter.
A successful write confirms acceptance, not global visibility. On an ambiguous timeout, the application may not know whether the write succeeded. Retrying an assignment of the same value is different from retrying a read-modify-write operation; KV cannot make that second pattern atomic.
Transient failures (network issues, temporary overload) deserve retry with exponential backoff. Permanent failures (value too large, namespace doesn't exist) won't succeed on retry. Transient failures manifest as network errors or 5xx responses; permanent failures return specific error codes.
For data that needs authoritative storage, write to D1 first and treat KV as a derived cache. Record or tolerate failed cache refreshes. Independent writes to two stores do not provide certainty unless the application defines which copy wins and how divergence is repaired.
When KV is the right choice
KV fits frequently read data whose freshness policy tolerates eventual visibility. Cached values and misses can survive beyond sixty seconds; do not select it from a presumed one-minute ceiling.
Configuration and feature flags fit when gradual visibility is acceptable. Keep emergency stops and consequential permissions on an authoritative path when stale execution would cause harm.
Cached API responses work well. External API results that don't change frequently can live in KV. TTL influences cache freshness, but eventual propagation does not supply a strict global staleness deadline.
Session data fits only when delayed revocation is acceptable. Infrequent login and logout do not remove that requirement; decide whether the next authenticated request must consult current session state.
Static content metadata fits when stale attributes are harmless. Treat permissions separately if a revocation must take effect immediately.
When KV is the wrong choice
The inverse pattern (frequent writes, consistency requirements, coordination needs) makes KV actively harmful.
Precise counters need an atomic update. KV cannot provide a read-modify-write transaction. Use a conditional D1 update when the invariant fits one database, or a Durable Object when an active owner also needs to coordinate requests or application state.
User-visible data that changes frequently creates bugs. Shopping carts, user preferences, or profile updates that users expect to see immediately will disappoint when served from stale caches. Users add an item, refresh, don't see it, add it again. Now they have duplicates.
Write-heavy workloads pay a premium. Beyond included allowances, KV writes cost $5 per million operations versus $0.50 per million reads. Compare the actual KV operations with the rows D1 would read and write; a fixed write-percentage threshold cannot decide which is cheaper.
Large values defeat the purpose. KV accepts values up to 25 MB, but large values undermine caching benefits and slow propagation. Use R2 for substantial objects. KV shines with small values typically under 100 KB.
The namespace strategy
Use namespaces to separate concerns that need independent bindings, monitoring or deletion. Prefixes are useful for organising keys inside a namespace. Neither choice changes consistency or creates an independent account-level quota.
A namespace per tenant can simplify offboarding, but it consumes the account’s namespace allowance and increases provisioning work. Choose it for an operational boundary you need, not merely because the tenant count fits under the limit.
Cache invalidation: the hard truth
Cache invalidation is famously hard. With KV, the problem has a specific shape: you cannot reliably invalidate.
When you delete or update a key, the change propagates eventually. During propagation, some edges serve old values, some serve new values, some serve "not found." You cannot force immediate global invalidation. You cannot query whether invalidation has completed. You can only wait.
KV caching strategies must embrace TTL-based expiration rather than explicit invalidation. Set a TTL that represents your staleness tolerance. Design your application so this staleness causes no harm.
Write-through caching (updating the cache when you update the source) provides an illusion of freshness. You update your database and immediately update KV. But the KV update propagates on its own schedule, potentially slower than your database replication. Users in Tokyo might see fresh data before users in London, despite London being closer to your origin database.
Cache-aside keeps the source of truth explicit: on a miss, read the source and populate KV. A database update and a cache refresh can still race, leaving an older value cached. Choose expiry and versioning so that outcome is tolerable; KV offers no immediate global invalidation barrier.
Content-addressed keys avoid overwriting cached content: a hash identifies one immutable value. A separate manifest or pointer chooses the version to serve. That pointer still has a freshness policy, and a newly written content key can still be temporarily absent because negative lookups are cached. Immutability removes stale-value ambiguity; it does not remove propagation.
Old versions also accumulate. Set an expiration or cleanup policy, and retain values while a live manifest or permitted rollback can still reference them.
TTL selection: a framework
TTL selection is a tradeoff between freshness and efficiency. Four questions guide the decision.
Is freshness a preference or a deadline? A lower cache TTL can improve typical freshness, but it does not give KV a strict global propagation bound. If a change must be enforced within a fixed deadline, use a path that supplies that contract.
What's your cache miss cost? Measure the source query and its capacity under cold traffic. An external API miss can consume both latency and quota. Higher miss costs justify longer caching only within the accepted freshness policy.
How quickly must a change take effect? A daily update schedule does not justify a 24-hour cache if that one update must be visible within a minute. Set freshness from the effect of missing the change, then use update frequency to estimate cache efficiency.
What's the cost of serving stale data? Sometimes stale data is merely suboptimal (yesterday's exchange rate). Sometimes stale data breaks functionality (showing a user as logged in after logout). The severity determines your tolerance.
Choose a TTL within the acceptable stale-data policy and measure its effect on hit rate and source load. Session expiry does not settle revocation: a session can become invalid before its original lifetime ends. Where current permission is required, consult its authority.
KV versus regional caches
A Redis-compatible regional cache can provide data structures and atomic operations KV does not. KV provides managed, geographically distributed cached reads with eventual consistency. Compare them from the compute location that will issue the request, and include the freshness and operation semantics the application needs. A same-region microsecond benchmark does not describe a cross-continent path.
Hyperdrive: removing repeated connection setup
A new database connection can require several round trips for TCP, TLS and authentication before the first query runs. Those trips are costly when a Worker and its database are far apart.
Hyperdrive handles setup close to the Worker and reuses connections pooled near the database. A cache miss still travels to that database and back; pooling removes repeated setup round trips, not the query’s network round trip. A cache hit can avoid the origin journey, with the freshness trade-off discussed below.
This distinction matters for chatty code. Ten dependent queries can still pay ten long round trips over a pooled connection. Batch independent queries where possible, and consider placing database-heavy execution near the database rather than assuming pooling makes placement irrelevant.
The developer experience
Hyperdrive presents itself as a connection string. Your code doesn't know it's using Hyperdrive:
import postgres from "postgres";
export default {
async fetch(request: Request, env: Env) {
const sql = postgres(env.HYPERDRIVE.connectionString);
const users = await sql`SELECT * FROM users WHERE active = true`;
return Response.json(users);
}
};
The binding provides a connection string routing through Hyperdrive's infrastructure. Your database driver connects to what appears to be a local endpoint. Hyperdrive handles the geographic complexity invisibly.
The client created in the handler belongs to that request; Hyperdrive's origin pool outlives it. Creating a new client on the next invocation therefore need not repeat the database's remote connection setup.
Existing SQL and supported database drivers can usually be retained. Test transaction and session-dependent behaviour against Hyperdrive’s pooling model, and close per-request driver resources according to the integration guidance. Compatibility is broader than a new connection string, but it is not unlimited.
Measuring the improvement
Benchmark cold connections, pooled cache misses and cache hits separately. Then measure a complete request with the real number of sequential queries. Report the Worker and database locations, cache policy and concurrency.
Pooling helps most when setup dominates. Caching helps when results repeat and may be stale. Placement and query batching help when many origin round trips remain. Those are different levers, and their benefits should not be combined into an unexplained speedup figure.
Transaction integrity
Connection pooling raises an obvious concern about transactions if connections are shared, but Hyperdrive handles this correctly. When you begin a transaction, that connection is dedicated to your request until the transaction commits or rolls back, so your transaction isolation is maintained.
Long-running transactions hold connections. Under high load with long transactions, you might exhaust the pool. Keep transactions short (seconds, not minutes) and you'll stay within capacity.
You can also size the pool to the database rather than accept a fixed default. Hyperdrive lets you set the connection count per configuration, with a floor of 5 and a ceiling determined by your Workers plan. A small origin that cannot tolerate hundreds of inbound connections and a large cluster that can are both accommodated, which matters when several Hyperdrive configurations or other clients share the same database's connection budget.
Query caching: trading freshness for speed
Hyperdrive query caching is enabled by default. Eligible repeated reads can return cached results without reaching the database. Use a separate configuration with caching disabled for paths that require fresh reads.
Hyperdrive’s max_age is capped at 3,600 seconds. Account for any stale-while-revalidate allowance separately when setting the freshness policy.
Not all queries benefit equally:
| Query Pattern | Caching Value | Recommendation |
|---|---|---|
| Immutable or versioned record | High | Enable within the supported cache lifetime |
| Reference data (countries, currencies) | High | Enable, TTL matching update frequency |
| User-specific data queried repeatedly | Medium | Enable with short TTL if staleness acceptable |
| Aggregations for dashboards | Medium | Enable, TTL matching reporting cadence |
| Data with temporal conditions | Low | Disable; conditions change results |
| Write-then-read patterns | Negative | Disable; cache prevents seeing writes |
The fundamental tradeoff mirrors KV in that caching trades freshness for speed; writes through Hyperdrive don't automatically invalidate cached reads, so if you write a row and immediately read it, you might get the cached old value.
For read-heavy workloads with repeating queries and tolerance for staleness, enable caching. For write-heavy workloads or queries requiring current data, disable it. There's no middle ground where caching is "smart enough" to know when to invalidate.
Hyperdrive failure modes
Hyperdrive has characteristic failure modes worth naming.
Origin unreachable is straightforward when the database is down, the network is partitioned, or credentials are invalid; queries fail immediately with connection errors, so implement retry logic for transient failures and surface permanent failures to users.
Pool exhaustion under load manifests subtly when all pooled connections are in use and new queries must wait. You'll see rising latency before you see errors, with a query that normally takes 50ms potentially taking 500ms whilst waiting for a connection.
Monitor both query latency and the published pool metrics: active connections, waiting work and configured capacity. Rising tail latency can come from pool pressure, slow queries or network delay; distinguish those causes before increasing the pool and putting more load on the database.
Stale cache after write is the caching consistency trap when you write a row and then read it; if caching is enabled, the read might return the cached pre-write value, so disable caching for those queries if correctness depends on reading your own writes.
Credential rotation race occurs during database credential updates when you update the database to reject old credentials before Hyperdrive fully adopts new ones, causing connections to fail. The safe rotation sequence is to add new credentials to your database, update Hyperdrive configuration, verify new credentials work, and then remove old credentials.
Origin timeout cascade compounds under load when your database becomes slow and queries time out, with each timeout holding a pooled connection. As the pool fills with timing-out queries, new queries wait and then also timeout. The mitigation is aggressive timeouts and circuit breaking, because if your database is struggling, failing fast protects the system better than waiting hopefully.
Multi-region database topologies
If your database has read replicas in multiple regions, Hyperdrive's interaction with that topology deserves consideration.
Hyperdrive connects to a single database endpoint. It doesn't automatically route to the nearest replica. For read-heavy workloads with existing multi-region replicas, consider multiple Hyperdrive configurations, one per replica region. Route read queries to the geographically appropriate instance; writes go to the primary.
If you're building new infrastructure, consider whether D1 with read replicas better fits your needs. D1 handles global distribution natively; Hyperdrive accelerates access to existing infrastructure but doesn't create geographic distribution.
Supported databases
Hyperdrive supports PostgreSQL and MySQL, including custom certificate and mutual-TLS configuration. Verify the driver and database variant your application uses. A Cloudflare-managed integration with a database vendor can simplify provisioning and billing, but engine behaviour, connection budget and recovery remain part of the database choice.
Database connections require TLS. Hyperdrive refuses unencrypted connections. If your database doesn't support TLS, you cannot use Hyperdrive.
Comparing connection proxies
RDS Proxy pools access to supported RDS databases; PgBouncer provides PostgreSQL pooling where you deploy it. Hyperdrive adds integration for geographically distributed Workers and optional result caching. Choose by the compute and database topology, required pooling semantics and operational ownership. None makes an uncached remote query independent of distance.
Choosing between KV, D1, and Hyperdrive
The first question isn't "how fast?" but rather "how stale?" Once you've answered that, the storage choice often makes itself.
Start with consistency
If you need strong consistency (reads always return the latest write), eliminate KV immediately. Choose between D1 (Cloudflare-native, global with read replicas) and Hyperdrive (your existing database, accelerated).
If eventual consistency is acceptable, KV's speed and simplicity win. Don't over-engineer with D1 what KV handles elegantly.
Consider data shape
Relational data with queries across rows belongs in D1 or Hyperdrive. KV is key-value only; you can't query by attributes, join, or filter. If you need "all users in London" or "orders above 100," KV is structurally incapable.
Simple lookups by known key are KV's strength: session by token, configuration by name, cached response by request hash.
Evaluate write frequency
Compare billable operations, not just the percentage of requests that write. KV charges per key operation; D1 charges per row read and written, including index work. The rows scanned by a query can matter more than the read/write ratio.
Account for existing infrastructure
If you have PostgreSQL or MySQL running elsewhere, Hyperdrive provides immediate value without migration. You keep your existing database, backups, and operational knowledge.
If you're starting fresh, D1 eliminates external dependencies. No database to manage outside Cloudflare, no connection strings to secure, no pool exhaustion to monitor.
Hyperdrive as architecture, not just bridge
The framing of Hyperdrive as a "bridge" undersells its value. Many production systems run permanently with Workers connecting to external databases via Hyperdrive. This isn't a compromise or transition state. It's a valid architectural choice.
Hyperdrive fits as a permanent solution when your database requires PostgreSQL or MySQL capabilities D1 lacks (extensions, stored procedures, specific data types), when your database serves multiple applications beyond Workers, when your team has established operational practices around your existing database, or when migration risk exceeds any benefit D1 would provide. The connection pooling and query caching make external databases perform well from the edge; you're not accepting poor performance as the price of compatibility.
D1 fits better when you're starting fresh without external database dependencies, when you want global read distribution through replicas, when you prefer fully managed operations within Cloudflare's ecosystem, or when your data model aligns with D1's horizontal scaling patterns.
The hybrid approach works well for incremental adoption: new features built on D1, existing features continuing to use PostgreSQL via Hyperdrive. This lets you evaluate D1 with real workloads while maintaining continuity for proven systems. Migration can happen feature by feature, or the hybrid state can persist indefinitely if that's what serves your architecture best.
The decision matrix
| Your Situation | Choose | Because |
|---|---|---|
| Read-heavy, staleness acceptable | KV | Sub-10ms global reads, simple model |
| Need relational queries, new project | D1 | Native integration, no external dependencies |
| Have PostgreSQL/MySQL with established operations | Hyperdrive | Accelerates existing infrastructure permanently |
| Need application-owned coordination | Durable Objects | Keep the state and decision in one object |
| Require PostgreSQL-specific features | Hyperdrive | D1 is SQLite-based; keep your extensions |
| Database serves multiple applications | Hyperdrive | Shared data layer across Workers and other systems |
| Configuration and feature flags | KV when delayed visibility is acceptable | Emergency controls need the authority their deadline requires |
| User profiles with search | D1 | Need relational queries |
| Session tokens | Authoritative store, or KV if delayed revocation is acceptable | Expiry and revocation are separate requirements |
| Write-heavy with consistency needs | D1 or Hyperdrive | KV pricing penalises writes; either relational option works |
| Uncertain about D1 migration | Hyperdrive | Start with existing database, evaluate D1 later |
Operational monitoring
Use service metrics alongside application tracing. Hyperdrive exposes query latency, cache behaviour and connection-pool metrics; correlate them with origin database health. For KV, measure whether a result was usable and sufficiently fresh, not just whether get() succeeded.
Instrument complete requests as well as storage calls. A fast stale result and a fast empty result can both produce a broken user journey. An isolated p99 spike is a lead to investigate, not proof of pool exhaustion.
What comes next
Part V combines these data services with AI inference. Documents, search indexes and conversation state have different access patterns; keeping those boundaries clear makes it easier to evaluate whether an AI feature is both useful and affordable.