Skip to main content

Chapter 20: Cost Modelling and Optimisation

What will this actually cost, and how do we optimise?


Technical leaders don't ask "what does Cloudflare charge?" They ask "will this be cheaper than what we're doing now, and when does that change?" Answering requires understanding the economics behind the prices.

Cloudflare's pricing attaches dollar signs to architectural guidance, providing concrete feedback on design decisions. When something costs significantly more, reconsider the approach. Understanding why reveals when Cloudflare's model wins, when it doesn't, and how to design systems where cost efficiency and correctness align.

The economics behind the pricing​

Pricing turns architectural choices into measurable costs. Start with the units each product bills, then estimate how many units your workload consumes. The meter is a useful guide to optimisation; it is not a complete account of Cloudflare's own costs.

Why CPU time, not wall time​

Workers bill for CPU time consumed, not elapsed time. A Worker waiting 500ms for a database response but using only 2ms of CPU pays for 2ms. This seems generous until you understand the model.

Workers exclude I/O wait from their CPU meter. Fan out independent backend calls when that improves response time, while accounting for the requests and work those backends bill.

A Worker that computes for 10ms and waits two seconds for APIs uses the same billable CPU as one that finishes after 10ms of computation. Lambda's duration meter includes the wait, with its cost also depending on allocated memory. That can favour Workers for orchestration and backend-for-frontend workloads; calculate the saving from both platforms' request charges, allowances and execution rates rather than applying a universal multiplier.

Why KV writes cost ten times what reads cost​

KV reads cost $0.50 per million. Writes cost $5.00 per million. This 10:1 ratio reflects the operational difference between reads and writes in a globally distributed cache.

KV caches reads near users, fetching from higher tiers or its central stores when necessary. A write updates the stored value; other locations may continue serving cached values until they refresh. The read/write price difference helps identify economical access patterns, but it does not mean every write immediately replicates to every edge location.

Let Pricing Guide Architecture

Writing to KV more than reading? Wrong tool. The pricing signals the intended use case. Configuration data, feature flags, cached API responses align with KV's economics. Counters, session updates, real-time state fight the pricing model. Use Durable Objects or D1 instead.

Why D1 charges per row​

D1 costs $0.001 per million rows read and $1.00 per million rows written beyond the Paid plan's monthly allowances of 25 billion reads and 50 million writes. The unit is a row, so count rows scanned as well as rows returned.

Index Your Queries

Beyond the included allowance, scanning a million rows costs $0.001; reading ten costs $0.00000001. The ratio is a hundred thousand to one, but the absolute read charge is small. Index first to protect latency and throughput, then measure whether read overages justify further cost optimisation.

D1 runs on SQLite, implemented as a Durable Object underneath. Rows read are a billing metric, not a count of physical disk operations. Large scans still occupy a single-threaded database and delay other queries even when they generate no additional read charge.

Why egress changes the comparison​

R2 does not charge for egress. For applications delivering large objects or exporting data to other providers, removing that volume-dependent charge can materially change the budget. Compare storage, operations and delivery requirements as well: zero egress is one pricing term, not proof that every transfer-heavy workload is cheaper overall.

Why Durable Objects bill for duration​

Durable Objects have request, duration and storage charges. Duration bills the time an object is active or idle while unable to hibernate, using a 128 MB allocation. Eligible idle objects are not charged for duration while waiting to hibernate; the lifecycle's idle timeout is not a minimum billed tail.

Hibernation removes duration charges during eligible idle periods, while stored data can still incur charges. Open WebSockets using the hibernation APIs need not keep the object billable. Timers, unfinished work and non-hibernating connections can change that behaviour, so model the lifecycle the application actually creates.

For sparse sessions, this can avoid paying continuously for idle coordination. For busy objects, duration may dominate. Measure active time and wake-up frequency instead of estimating cost from the number of object identities alone.

Storage choice as economic decision​

Every storage primitive answers a different question. KV: what's the value for this key? D1: which rows match these conditions? Durable Objects: what happened to this specific entity? R2: what's in this file? Choosing wrong doesn't just cost more; it makes your code fight the abstraction.

Decision framework by access pattern​

If your pattern is...UseBecause
Read-heavy, rarely updatedKV$0.50/M reads, global caching
Write-heavy, needs queriesD1$1.00/M writes, SQL flexibility
Per-entity coordinationDurable ObjectsOne active owner for local decisions
Large binary dataR2Zero egress, S3 compatible
External PostgreSQL existsHyperdriveDon't migrate, accelerate

Cost comparison by operation type​

These are marginal operation rates after included allowances; storage and compute add separate costs.

OperationKVD1
Single read$0.0000005$0.000000001 per row
Single write$0.000005$0.000001 per row
1000 reads$0.0005$0.000001 (if 1 row each)
1000 writes$0.005$0.001 (if 1 row each)

D1 has the lower marginal charge for a single-row read; KV buys a globally distributed cache with different latency and consistency properties. A D1 query reading 1000 rows incurs the same row charge as a thousand single-row queries, although the execution overhead differs. Durable Objects have request costs plus duration and storage; the comparison depends on how long objects stay active.

The pattern: use KV for read-heavy key-value access, D1 for write-heavy or query-complex workloads, and Durable Objects when you need coordination or per-entity state.

When Cloudflare wins on cost​

The answer depends on workload characteristics, not volume alone.

High egress relative to compute​

R2 deserves evaluation when delivery or export is a significant part of the bill. Calculate the avoided transfer charge from the actual destination, tier and contract, then account for storage, operations and the delivery path. There is no universal percentage of egress spend at which the whole architecture becomes cheaper.

Spiky, unpredictable traffic​

Usage billing can reduce the cost of idle capacity, while commitments can lower the unit price of predictable work. Both Cloudflare and hyperscalers offer services with different combinations of these terms. Compare the workload's peaks, minimum charges and capacity limits; do not assume every charge scales proportionally with requests.

Request-heavy, compute-light workloads​

Workers can be economical when a request spends much more time waiting than computing. Apply CPU and request rates to the measured workload, then compare the alternative's duration, memory and request rates. Include the backend work either design performs. The difference between 5ms of CPU and 500ms of elapsed time identifies a promising candidate; it is not itself a price ratio.

When hyperscalers win​

Cloudflare's model isn't universally superior.

Compute-intensive processing with predictable volume benefits from reserved capacity pricing. Heavy computation running hours daily on a predictable schedule favours EC2 reserved instances over per-millisecond billing.

Deep AWS service integration makes migration costly. If your architecture depends heavily on SQS, SNS, Step Functions, DynamoDB Streams, and IAM, the rewrite cost may exceed operational savings over any reasonable horizon.

GPU workloads for ML training have no Cloudflare equivalent. Workers AI provides inference, not training.

Large single databases beyond 10 GB require traditional managed databases. D1's horizontal model works for many applications, but some need monolithic relational stores.

Total cost of ownership​

Comparing platforms requires accounting for expenses that don't appear on obvious line items.

Hidden hyperscaler costs​

Technical leaders comparing Cloudflare to AWS routinely underestimate several categories.

NAT Gateway charges apply when outbound traffic is routed through a managed NAT gateway. Model both its fixed charge and the traffic processed. Do not add it mechanically to every private-network design; connectivity and routing choices determine whether the charge exists.

Cross-AZ data transfer costs $0.01 per GB in each direction. Database in one AZ, compute in another: charges accumulate invisibly.

CloudWatch Logs charges for ingestion and storage. Verbose logging in high-volume applications generates surprising bills.

Provisioned concurrency for cold start mitigation costs whether invocations occur or not, accumulating continuously rather than per-use.

A worked comparison​

Consider a SaaS API with these characteristics:

  • 50 million requests monthly
  • Average response time 200ms, average CPU time 15ms
  • 500 GB in D1 and 500 GB in R2, with 5 TB monthly egress
  • Database averaging 50 rows read per request
  • 70% cache hit rate (35M requests served from cache)

This is an illustrative Cloudflare monthly budget. It assumes the account's included allowances are available, one KV lookup per request, negligible cache-fill writes, and an estimated R2 operation charge. D1 requires sharding because each database is limited to 10 GB. Reprice storage, write traffic and the actual delivery path before using it for procurement.

To compare an AWS design, first specify which requests reach API Gateway and Lambda after caching, the API Gateway type, Lambda memory and architecture, and the database capacity needed for equivalent throughput. Include NAT only if the routing design uses it, and apply the same treatment of included allowances and discounts on both sides.

For example, 50 million 200ms Lambda invocations at 512 MB represent 5 million GB-seconds before allowances. Workers' corresponding CPU estimate is 750 million CPU milliseconds. These are different billing units: apply their rates separately, then include request, database and delivery charges.

Cloudflare Workers + D1 + R2 + KV:

ComponentMonthly Cost
Workers (50M requests, 15ms CPU avg, including $5 subscription)$31.40
D1 rows (750M reads, within included allowance)$0
D1 storage (495 GB beyond included storage)$371.25
KV reads (50M lookups, including cache misses)$20
R2 (500 GB + operations)$11
R2 egress (5 TB)$0
Total~$434

Relational storage is the largest charge in this Cloudflare scenario. Zero egress removes one source of growth, but it does not remove storage, operation or computation charges. The alternative design must meet the same performance and recovery requirements before its total is comparable.

How assumptions change the outcome​

This comparison is sensitive to workload characteristics. Understanding the sensitivities matters more than memorising numbers.

Cache hit rate sensitivity:

A higher hit rate reduces D1 queries but does not remove the KV lookup needed to discover each hit or miss. In this example, D1 reads stay within the included allowance at 50%, 70%, and 90% hit rates. Holding average CPU time constant therefore leaves the estimated bill unchanged. Caching can still improve latency and database throughput; calculate savings from measured CPU reductions and cache-fill writes rather than assuming every avoided query saves money.

Egress sensitivity:

R2's egress line remains zero as transferred bytes increase. Other costs need not remain fixed: additional delivery requests, Worker execution or cache-fill operations may accompany that traffic. In the alternative design, apply its delivery tiers, regions and included transfer allowances. This distinguishes a real egress saving from an assumed flat application bill.

CPU time sensitivity:

Avg CPU TimeCloudflare Workers CostTotal Cloudflare
5ms$21.40$424
15ms$31.40$434
50ms$66.40$469
100ms$116.40$519

Higher CPU per request raises the Workers line predictably while the other assumptions stay fixed. In a real workload, profile whether the same change also affects memory, caching or backend work before extrapolating this table.

Model your specific workload. The comparison that matters is yours, not a hypothetical.

AI inference costs​

AI workloads can dominate bills quickly. The economics differ from traditional compute.

Cost structure​

Text-generation models on Workers AI have model-specific input and output token rates. Parameter count alone does not determine the bill, and not every AI task uses a token meter. Use the rate card for the actual model and task selected in Chapter 17.

Worked example​

A support assistant serving 100,000 queries monthly, averaging 500 input tokens and 300 output tokens each, processes 50 million input and 30 million output tokens. If the chosen model's rates are i and o dollars per million tokens, its inference estimate is 50 × i + 30 × o dollars, before retries and any other services.

Compare candidate models on that same workload and an agreed quality measure. A higher per-token rate may still be economical if it requires shorter prompts, fewer retries or less human review. Conversely, paying for capability the task does not need is avoidable expense.

AI cost optimisation​

Right-size models first. Don't default to the largest available. Test smaller models against your actual queries; quality differences may be negligible for your use case.

Set explicit max_tokens limits. Without limits, models can generate unexpectedly long responses, costing tokens you didn't intend to spend.

Cache repeated queries through AI Gateway. Similar questions? Cached responses eliminate inference costs entirely for duplicates.

Measure prefix-cache reuse for multi-turn interactions. Supported models can reuse retained matching prefixes, and session affinity can improve the chance of reaching the same model instance. Price the observed cached-token usage for the selected model, including cold requests and changed prefixes. A growing conversation does not guarantee that every prior token receives the cached rate.

Truncate context aggressively. Every prompt token costs money. Include only context the model needs. Summarise long documents rather than including full text.

Set gateway spend limits as an operational control. AI Gateway’s spend limits (beta) can block requests against budgets scoped by model, provider or trusted user/team metadata. A configured Dynamic Route can use a cheaper fallback when the primary model’s budget is exhausted. Access user attribution requires the protected custom-domain path and a valid user subject.

These limits use estimated costs and eventually consistent accounting; concurrent requests can overshoot before enforcement catches up. They reduce runaway spending but do not establish a strict financial ceiling. Chapter 23 explains admission controls for shared capacity and retry budgets.

Cost as architectural feedback​

On Cloudflare, high costs usually indicate architectural problems, not just scale. The alignment between cost efficiency and correctness isn't accidental; pricing reflects operational costs, which reflect resource consumption.

Compare usage with its cause​

Track cost per successful request, active tenant, processed job or stored gigabyte, whichever reflects the product's value. Separate changes in traffic from changes in work per request. A rising bill with a stable unit cost is a different problem from a deployment that doubles rows scanned.

Inspect the quantities behind each charge: D1 rows read and written, KV read-to-write volume, Durable Object billable duration, and AI input and output tokens. Compare them with your own baseline. Ratios between product bills are easily distorted by included allowances and say little about whether the architecture is efficient.

Common mistakes by workload type​

SaaS APIs can overload D1 by repeatedly querying and updating per-user state. Check contention and write volume before blaming read charges. Use Durable Objects when per-user coordination is the requirement; keep D1 for relational queries across entities. The access pattern determines the fit.

Media applications should choose the delivery path alongside access control. Direct R2 delivery can avoid a Worker invocation when it meets the requirements; a Worker may be justified for authorisation or response handling. Stream large bodies rather than buffering them, and measure CPU used by processing separately from time spent transferring bytes.

Real-time applications can keep Durable Objects billable through unnecessary timers or polling. With hibernating WebSockets, avoid waking application code for heartbeats the runtime can answer automatically. Decide how stale presence information may be before choosing a cheaper update frequency.

AI applications commonly include excessive context in prompts. Every historical message, every retrieved document, every system instruction: tokens accumulate. Summarise conversation history, limit retrieval results, cache repeated context.

The optimisation priority stack​

When optimisation is warranted, prioritise by impact.

First: D1 query efficiency. Run EXPLAIN QUERY PLAN on expensive or frequent queries. Choose indexes that support their filters, joins and ordering, then measure rows scanned. Each index also adds storage and write work; indexing every filtered column independently is not a substitute for understanding the query.

Second: storage choice alignment. Estimate the cost of the observed read and write pattern on each viable store. Change primitives only when the consistency, query and operational requirements still fit.

Third: Durable Object lifecycle. Identify polling, timers or non-hibernating connections that keep objects billable. Use the hibernation APIs where the application can reconstruct its state when work arrives.

Fourth: AI model and prompt efficiency. Test smaller models. Truncate prompts. Cache responses. Set token limits.

Fifth: cache what you serve repeatedly. Workers Cache turns a repeated expensive response into a cache hit that bills only the request: no CPU, no Durable Object invocation, no D1 query. The trade is that enabling it bills every request at the standard rate, including normally-free static asset and service-binding requests, so it pays for itself on compute-heavy paths and can cost money on asset-heavy ones (Chapter 4 covers the mechanics).

Sixth: Workers CPU optimisation. Profile expensive paths, then use the same payback calculation as any other optimisation. A low-volume application with heavy computation can justify this work before a high-volume, compute-light one.

When to optimise​

Optimisation has costs: engineering time, added complexity, potential bugs. Not all optimisation is worthwhile.

The three-month rule​

Optimisation investment should return within three months. If your D1 costs $200/month and you estimate 50% savings, optimisation yields $100/month, or $300 over three months. If the engineering work costs more than $300 in time, don't optimise yet. Wait until scale makes it worthwhile.

This calculation changes at scale. At $5,000/month, the same 50% optimisation yields $7,500 over three months. That justifies substantial engineering investment.

Moving to the paid plan​

The Workers Free plan enforces its daily request limit; exceeding it does not automatically enrol you in usage billing. Moving to the paid plan introduces the base subscription and monthly included usage, followed by overage charges. Model that plan before traffic requires the move, including the storage and observability services the application uses.

Defensive CPU caps​

The limits.cpu_ms configuration does double duty. Most documentation emphasises raising the limit from the default 30 seconds to accommodate compute-heavy workloads. The defensive use, lowering the limit, receives less attention but matters more for cost control.

For a Workers Paid workload with P99 CPU time of 15ms, a 50ms cpu_ms budget leaves headroom while limiting runaway computation. The runtime allows some occasional overrun flexibility, so this is not an exact termination deadline. It limits CPU per invocation, not total spend across an attack's requests.

wrangler.jsonc: set a 50 ms CPU budget
{
"limits": {
"cpu_ms": 50
}
}

Durable Objects, Workflows and Queue consumers have their own execution limits and configuration contexts. Check those separately. The design question remains: what's a reasonable maximum for legitimate requests, with enough headroom for variance but not enough to cause meaningful cost damage if something goes wrong?

This isn't premature optimisation. It's establishing a sensible boundary before you need it. Setting it wrong costs a few failed requests you'll notice and adjust. Not setting it means discovering problems through your invoice.

Balancing cost and reliability​

Cost optimisation can conflict with reliability investment. Removing redundancy saves money but increases risk. Aggressive caching reduces costs but can serve stale data.

Frame this trade-off explicitly: if your error budget is being exceeded, stop optimising for cost. No savings matter if users are suffering. Well under your error budget? Optimise aggressively. You may be over-investing in reliability users don't perceive.

Monitoring and response​

Cost visibility requires instrumentation specific to Cloudflare's model. The Billable Usage dashboard surfaces daily usage-based costs for pay-as-you-go customers and matches the figures that appear on the monthly invoice, which removes a long-standing gap where the only authoritative number was retrospective. Budget alerts fire when spend crosses thresholds you configure, so anomalies show up the same day rather than at month-end. Treat these as the floor of cost monitoring, not the ceiling: dashboard alerts catch absolute spend, but the per-request and per-product signals below catch architectural drift before it shows up as a bill.

What to track​

Cost per request normalises for traffic variation. Cost per request increasing while traffic grows slower? Something has changed architecturally.

D1 rows per query averaged across your application reveals query efficiency trends. Rising averages suggest degrading query patterns or growing data volumes outpacing index effectiveness.

Durable Object active duration per request reveals whether DOs are sleeping appropriately. Duration growing while requests stay flat means DOs are staying active longer than necessary.

KV write-to-read ratio helps identify changing access patterns. If writes grow, compare their cost and consistency requirements with the viable alternatives before changing stores. A ratio alone does not establish that D1 is a suitable replacement.

Responding to anomalies​

When costs spike unexpectedly, follow this diagnostic sequence.

First, correlate with traffic. Did request volume increase proportionally? Costs doubled and traffic doubled means the system is working correctly.

Second, examine per-request metrics. If cost per request increased, identify what changed. Check recent deployments, new features, logging changes.

Third, check D1 query patterns. A removed index or changed query can cause dramatic cost increases. Review recent schema or code changes affecting database access.

Fourth, verify caching behaviour. Cache hit rate drops increase backend costs. If caching degraded, investigate why: TTL changes, cache invalidation bugs, new uncacheable paths.

Cost spikes are symptoms. Diagnosis identifies the cause; budget adjustment doesn't fix anything.

Modelling for growth​

Technical leaders must project costs at scale for informed platform decisions.

Growth scenario template​

Model at least expected traffic, a higher-volume case and a workload change such as heavier writes or longer AI responses. For each, state requests, CPU per request, cache hit rate, data growth and the operations performed on a miss.

Fixed charges may amortise with growth, but unit cost need not fall. Allowances can be exhausted, larger datasets can increase query work, and a hot tenant can require a different partitioning scheme. Show which assumption causes each change in the projected bill.

Presenting to leadership​

Lead with total cost of ownership, not component costs. The comparison that matters is "what does it cost to run this workload?" and that question has four dimensions: infrastructure cost (the monthly bill), development cost (the engineering time to build and migrate), operational cost (the ongoing effort to deploy, monitor, and respond to incidents), and maintenance cost (keeping dependencies current, applying security patches, managing capacity). Most platform comparisons focus exclusively on the first dimension. Cloudflare's advantage often shows most clearly in the second and third, where the elimination of capacity planning, cold start mitigation, and multi-region deployment translates into engineering weeks that never appear on an invoice.

Acknowledge switching costs honestly. Migration requires engineering investment. A comparison showing $500/month savings is meaningless if migration costs $50,000 in engineering time: that's eight years to break even.

Provide ranges, not point estimates. "This workload will cost $400-600 monthly based on traffic projections" gives leadership a budget range while communicating inherent uncertainty.

What comes next​

A cost model gives you expectations to test. Chapter 21 covers the evidence needed to test them: request traces, state transitions, latency distributions and alerts that distinguish a busy system from a failing one.