Skip to main content

Chapter 14: R2: Object Storage Without Egress Fees

How much can we actually save with R2, and when does it matter?


Object storage bills can be dominated by moving data out. R2 removes that egress charge, but the saving depends on where the data goes: some same-cloud transfers and S3-to-CloudFront transfers are already free. Compare the full delivery path, not just the bucket’s headline rate.

Zero egress isn't a discount; it's a different economic model that enables serving the same large file to millions of users, building data interchange systems without transfer costs, and creating backups you can actually afford to restore. Understanding when these economics matter determines whether R2 fits your architecture.

The economics in practice​

Start with the current bill: storage by class, read and write operations, retrieval fees, outbound transfer and any CDN delivery. Then price the same workload on R2, including Workers or other processing used in the access path. Data volume alone cannot price a million small downloads and a thousand large ones.

For example, 1,000 GB kept for a full month in R2 Standard costs $15 before the included storage allowance. Serving it ten times adds no R2 egress fee, but GET requests still count. At an assumed $0.09 per GB of chargeable transfer elsewhere, 10,000 GB would add $900 before allowances and discounts. That assumption applies only to a route actually billed at that rate.

When the saving changes the architecture​

A public dataset or large download can serve more readers without adding an R2 bandwidth bill. A multi-cloud pipeline can reread its source from R2 without paying R2 egress at every stage. Storage, requests, retrieval from Infrequent Access and downstream compute still cost money.

For an existing system, compare annual net savings with migration work and the ongoing cost of another storage service. A high egress percentage is useful evidence, but it is not a migration threshold: $50 and $50,000 deserve different responses even at the same percentage.

Operations pricing: the cost that replaces egress​

Zero egress doesn't mean zero cost beyond storage. R2 charges for operations in two classes: Class A operations (writes and mutations such as PutObject, CopyObject, ListObjects) cost $4.50 per million requests, while Class B operations (reads such as GetObject, HeadObject) cost $0.36 per million requests. Delete operations are free. For comparison, S3 charges $5.00 per million PUT requests and $0.40 per million GET requests, so R2 is marginally cheaper per operation as well as eliminating egress entirely.

R2's free tier absorbs most light workloads: 10 GB of storage, 1 million Class A operations, and 10 million Class B operations per month at no cost. A small application storing a few gigabytes and serving a few million reads per month may pay nothing at all.

For most architectures, operations costs are negligible compared to storage. But for very high-frequency access patterns, they can become the dominant expense. Consider a CDN-replacement pattern serving 300 million objects per month from R2: after the free tier, that's 290 million billable Class B operations at $0.36 per million, totalling roughly $104 per month. The storage for those objects might cost $15. When your read-to-store ratio reaches hundreds or thousands to one, operations costs dwarf everything else. This is the same pattern that makes S3 request pricing accumulate for small-object workloads, but without the compounding effect of egress on top.

The architectural implication: if you're designing a system that serves the same objects millions of times, Cloudflare's CDN cache sits in front of R2 and absorbs repeat requests. Cache hits don't count as R2 operations. The combination of R2 storage with aggressive edge caching keeps operations costs low even for extremely popular content. Architect for cacheability, and operations pricing rarely matters.

The broader hidden comparison​

Egress isn't the only S3 cost that accumulates silently beyond request pricing. Glacier retrieval carries both retrieval fees and per-gigabyte charges. Cross-region replication incurs transfer costs. R2 sidesteps some of these, though R2's Infrequent Access tier carries its own retrieval fees.

R2 trades S3's feature depth for pricing simplicity: two storage classes rather than S3's extensive tiering hierarchy, no true archival tiers for data accessed once per year, and no Object Lock for compliance-grade immutability. If you need those features, no amount of egress savings makes migration worthwhile.

When to stay with S3​

Zero egress doesn't automatically justify migration. The question isn't "can you save money" but "is the saving worth managing two storage systems?"

Ecosystem integration often outweighs egress savings. S3 connects natively to AWS Backup, AWS Analytics, CloudWatch metrics, IAM policies, and dozens of other services. A 50% egress reduction that requires reimplementing your backup strategy isn't a good trade.

Internal AWS transfers are often free. Data moving between S3 and other AWS services in the same region frequently has zero or discounted egress. If your S3 buckets primarily serve Lambda functions and EC2 instances rather than external users, R2's egress advantage evaporates.

Operational complexity has real costs. Two storage systems mean credentials, monitoring and recovery procedures for both. Estimate that work against the expected saving over the period you expect to run the system.

Compliance requirements may mandate specific providers. Regulatory or contractual obligations sometimes specify particular cloud providers or require features like Object Lock.

True archival storage is cheaper elsewhere. For terabytes of data accessed once per year, S3 Glacier Deep Archive at $0.00099/GB-month beats R2's Infrequent Access at $0.01/GB-month, even accounting for retrieval costs and egress fees. R2's Infrequent Access is for data accessed less frequently, not almost never.

R2 deserves serious consideration when chargeable transfer dominates a growing workload and the required object features are supported. Staying with S3 can be cheaper overall when integration or archive retention dominates. Make the comparison in dollars and operational work, not a fixed egress percentage.

S3 compatibility​

R2 implements the S3 API, not a "similar" API but the actual protocol that existing S3 SDKs speak. Point your AWS SDK at R2's endpoint, provide R2 credentials, and most operations work identically.

Compatibility covers core object operations such as reads, writes, listing, multipart upload and presigned URLs. Check the specific headers and API operations your application uses, including infrequent administrative and recovery paths. A successful GET test does not establish migration compatibility.

S3 Compatibility Gaps

S3 compatibility doesn't mean feature parity. R2 lacks certified WORM compliance (Object Lock), S3 Select, deep archival tiers, and S3 Inventory. Verify your required features are supported before committing to migration.

The gaps are real but specific:

WORM retention requirements. S3 Object Lock provides write-once-read-many retention: compliance mode protects an object version from modification or deletion even by the account’s root user during retention; governance mode and legal holds have different administrative rules. R2 Bucket Locks protect against overwriting and deletion for a configured period, but administrators can remove the rules. If the requirement needs protection from privileged deletion or assessed regulatory retention controls, evaluate S3 Object Lock or Azure immutable Blob Storage against that requirement. R2’s removable retention rules are not an equivalent guarantee.

S3 Select lets you query object contents directly, running SQL against CSV or JSON files without downloading them. If you're not already using it on S3, its absence won't affect you.

S3 Inventory provides automated reports of bucket contents and metadata for audits or compliance. The alternative is custom inventory using listing APIs, which works but requires you to build and maintain it.

Glacier and deep archive tiers provide extremely cheap storage for data accessed rarely. R2's Infrequent Access tier serves a different purpose: reducing costs for data accessed occasionally but not frequently. For petabytes with annual access patterns, S3's archival tiers remain cheaper despite egress fees.

The binding model difference​

Through the S3 API, R2 is external storage with credentials in configuration, network round-trips, and SDK initialisation overhead. Through Workers bindings, R2 is part of your compute environment.

Wiring an R2 bucket as env.STORAGE
{
"r2_buckets": [
{
"binding": "STORAGE",
"bucket_name": "my-bucket"
}
]
}

This configuration makes R2 available as env.STORAGE in your Worker. No credentials in code, no connection management, no SDK initialisation. The binding appears as a typed object with methods for storage operations.

Bindings enable streaming patterns where objects flow directly from R2 through your Worker to the response without loading into memory. A Worker with 128 MB of memory can serve a 5 GB video because the content streams through rather than accumulating. R2 becomes less "storage your code calls" and more "storage woven into your execution environment".

This integration is R2's strongest advantage for Cloudflare-centric architectures. If you're using Workers, R2 binding access is simpler, faster, and more natural than S3 SDK access to any storage provider. If you're accessing R2 purely through the S3 API from non-Cloudflare infrastructure, you're missing the deeper integration that distinguishes R2 from "S3 with different pricing."

Access patterns​

Three patterns serve R2 content. Start with Worker-mediated access; deviate to presigned URLs for large transfers or public buckets for public static content.

Worker-mediated access: the default​

Every request passes through a Worker that fetches from R2 and returns the response. You can authenticate, transform, log, rate-limit, or apply any logic before serving content. The Worker sees every request and controls every response.

The costs are real but usually acceptable. Every request consumes Worker invocations and CPU time, but Workers are cheap and fast. Request body limits apply: Free and Pro zones allow 100 MB, Business allows 200 MB, and Enterprise defaults to 500 MB with self-service configuration up to 5 GB. Larger uploads need a different approach or an agreed limit increase; raising the body limit does not remove connection timeouts.

Worker-mediated access fails when uploads exceed request body limits or when you're serving high-volume static content where per-request compute adds cost without adding value. For everything else, it's the right default because it preserves your options. You can always optimise later by moving specific paths to presigned URLs; you can't easily add authentication to content you've already exposed publicly.

Presigned URLs: the escape hatch for large transfers​

Generate time-limited URLs granting direct access to R2, bypassing Workers entirely. A Worker handles the authorisation decision and URL generation; the actual data transfer happens directly between client and R2.

This pattern exists primarily for large file transfers. Upload presigned URLs let clients send files directly to R2 without streaming through your Worker, avoiding the zone's request body limit and Worker CPU consumption during transfer. R2's own upload limits still apply. Download presigned URLs serve large files without Worker involvement.

The tradeoff is control. A presigned URL authorises access until it expires, so issuing one gives up the Worker's opportunity to reassess each request or transform content on the fly. R2 Data Access Logs provide visibility into successful operations through the S3-compatible API, Workers bindings, and public buckets, including direct access through presigned URLs. Enable them on non-jurisdictional buckets when that operational visibility matters, but allow for asynchronous, best-effort delivery and missing events. They exclude responses with status codes of 400 or above, so they are neither a complete access audit nor a substitute for application authorisation records.

Use presigned URLs when uploads exceed your zone's request body limit or when per-request Worker costs matter (high-volume large file downloads). Keep expiration times short: minutes for uploads, and hours at most for downloads.

User-generated content upload pattern​

For applications accepting user uploads, combine presigned URLs with event-driven processing:

Authorise an upload before issuing a presigned PUT URL. Bind it to an application-generated object key and a short expiry. Treat the uploaded bytes as untrusted even when the declared content type looks valid.

Keep incoming files in a quarantine prefix or bucket. An R2 event can queue validation, scanning or transformation; publish the object only after the required checks succeed. Upload completion and publication are separate states. Queue buffering smooths a spike, but the application still needs quotas and a policy for an accumulating processing backlog.

Local uploads: improving performance for global users​

When clients upload data from a different region than your bucket, object data must travel across potentially significant geographic distance. A user in Singapore uploading to a bucket in Europe experiences this latency on every upload. Local Uploads addresses this by writing object data to storage infrastructure near the client, then asynchronously replicating it to your bucket. The object is immediately accessible and remains strongly consistent throughout the process.

The performance improvement is substantial. Cloudflare reports up to 75% reduction in time to last byte for upload requests when Local Uploads is enabled. For applications where upload responsiveness affects user experience (file sharing, media uploads, collaborative tools), this is worth enabling.

The trade-off is jurisdictional. Local Uploads cannot be used with buckets that have jurisdictional restrictions because it requires temporarily routing data through locations outside the bucket's region. If your compliance requirements mandate that data never touch infrastructure outside a specific jurisdiction, even transiently, you cannot use this feature.

There's also a subtle consistency consideration. When Local Uploads places data near the uploader, reads from near the uploader's region are fast immediately. However, reads from near the bucket's region may experience cross-region latency until replication completes. For most applications this is invisible, but if your workflow requires immediate read-after-write from a location near the bucket rather than near the uploader, design accordingly.

Local Uploads is free to enable and has no additional cost. It's a configuration toggle in bucket settings, not an architectural change. For applications with globally distributed users and upload-sensitive workflows, enable it unless jurisdictional restrictions prevent it.

Public bucket access: for public content​

Configure the entire bucket for public access, making all objects retrievable without authentication. This is the simplest pattern for content that is unconditionally public: static assets, software releases, public documentation.

The tradeoff is absolute, making everything in the bucket public with no per-request authentication.

Use public buckets for static assets served to all users (CSS, JavaScript, marketing images), public downloads (open-source releases, public datasets), and CDN origin scenarios where all content is intentionally public. Avoid using them for user-generated content, even if "usually" public, because the moment you need to make one object private, the architecture requires reworking.

R2 as CDN origin​

For architects coming from AWS, R2 integrates naturally with Cloudflare's CDN. Using S3 as a CloudFront origin requires configuring two services: S3 bucket policies, CloudFront distribution settings, origin access controls, cache behaviours. The configuration spans multiple consoles and mental models.

To use Cloudflare’s cache for a public R2 bucket, connect a custom domain and configure caching for the content types you serve. The r2.dev development endpoint is not a production cached delivery path. For Worker-mediated access, implement caching explicitly through the appropriate cache API or configuration; response headers alone do not make every binding response an edge-cache hit.

S3-to-CloudFront transfer is free, although CloudFront delivery and requests have their own prices. R2 cache misses incur object operations and, for Infrequent Access, retrieval fees. Cacheability matters on both paths; compare origin and viewer costs separately.

Image optimisation pattern​

For image-heavy applications, store only original images in R2. Transformed variants (thumbnails, responsive sizes, format conversions) generate on-demand through Cloudflare Image Resizing and cache at the edge.

Resizing an image through its delivery URL
https://example.com/cdn-cgi/image/width=400,quality=80/images/product.jpg

The pattern eliminates variant storage entirely. No need to generate and store thumbnails during upload. No lifecycle rules managing variant proliferation. No synchronisation between originals and derivatives. Store once at maximum quality; transformations happen at request time.

Image transformations are billed by unique source-and-parameter combinations requested during a calendar month. Repeat requests for the same transformation in that month count once, independent of cache lifetime. Cache reuse still improves delivery latency and reduces origin reads.

Transform Rules enable migration from other image CDNs. If you're using Imgix or Cloudinary URL patterns, Transform Rules can rewrite those URLs to Cloudflare's format without application-level changes. This enables zero-downtime migration where you cut over DNS and existing URLs continue working.

Video transformation pattern​

The Media Transformations binding delegates supported video operations to Cloudflare's media service, including resizing and extracting frames or audio. Streaming avoids buffering the entire input in Worker memory, but the media service still imposes input and output limits. Check those limits against the actual asset; longer or larger processing jobs may need a different media pipeline.

The composability with other bindings is where this becomes architecturally interesting. A single event-driven pipeline can respond to a video upload in R2 by extracting a representative frame with the Media binding, classifying its content with Workers AI, extracting the audio track, transcribing it with Workers AI's speech-to-text, and storing the results back in R2 or D1. No Containers, no external services, no ffmpeg installations to maintain. For lighter video operations (resizing for different devices, generating preview clips, extracting thumbnails) this eliminates what previously required either Containers or an external media processing service. Full video transcoding still needs Containers or dedicated infrastructure, but many practical video workflows never actually require transcoding.

Event-driven processing​

R2 can notify Workers when objects change, enabling reactive architectures without polling. Configure notification rules that filter by action, prefix, and suffix; matching events flow to a Queue that triggers your processing Worker. The integration with Queues (Chapter 9) provides automatic retry semantics and dead-letter handling.

This pattern suits workflows where upload and processing can be decoupled. Image uploads trigger thumbnail generation. Document uploads trigger text extraction. File arrivals trigger validation or virus scanning. The upload completes immediately; processing happens asynchronously. Lifecycle-triggered deletions can also generate notifications for audit logging or cleanup workflows.

The decision comes down to one question: must processing complete before you can confirm the upload succeeded? If yes, use synchronous validation. If processing can happen after upload confirmation, event-driven patterns reduce latency and let you scale processing independently.

Event notifications have characteristics you must design around. Events aren't instantaneous, aren't guaranteed in order, and are at-least-once. If processing order matters, your handler must enforce it.

Event Handlers Must Be Idempotent

R2 event notifications are at-least-once. The same event may be delivered multiple times. Design handlers to process duplicate notifications safely.

Storage classes and lifecycle rules​

R2 Standard costs $0.015 per GB-month. Infrequent Access costs $0.01 per GB-month, plus $0.01 per GB retrieved, higher operation rates and a 30-day minimum. Class A operations cost twice Standard’s rate; Class B operations cost 2.5 times as much.

Ignoring operations and allowances, Infrequent Access saves $0.005 per stored GB-month and charges $0.01 per GB retrieved. The saving disappears when monthly retrieval reaches half the stored volume. Higher request costs and short object lifetimes can favour Standard sooner.

Lifecycle rules automate transitions and handle expiration. Configure rules that move objects from Standard to Infrequent Access after a specified period and delete objects after retention expires. Rules operate at the bucket level with prefix filtering, so different object hierarchies can have different policies.

Lifecycle rules belong to the bucket configuration; a Worker binding only grants code access to that bucket. Review retention rules separately from application deployment.

One constraint: objects in Infrequent Access cannot transition back to Standard via lifecycle rules. If cold data becomes hot, you'll need to copy the object (creating a new one in Standard) rather than transition it in place. If you're uncertain whether data will stay cold, the retrieval fees might offset the storage savings.

When infrequent access makes sense​

Infrequent Access works well for data with predictable low-access patterns: application logs older than a month, user-generated content with long-tail access (old photos, archived documents), backup data that exists for disaster recovery but hopefully never gets restored.

Infrequent Access works poorly for data with unpredictable access spikes or many small objects. The per-operation cost increase (roughly double) can outweigh storage savings when you're storing many small files rather than fewer large ones.

What R2's tiers don't provide​

R2's Infrequent Access is not a Glacier equivalent. S3 Glacier Flexible Retrieval and Deep Archive offer lower storage rates for data that can wait for restoration; compare the selected region’s storage, request and retrieval charges together. The tradeoff is retrieval time: Glacier Flexible Retrieval takes minutes to hours, depending on the retrieval option. Deep Archive Standard retrieval typically completes within 12 hours; Bulk can take 48 hours. R2 Infrequent Access needs no preliminary restore job, but charges a per-GB retrieval fee.

For long-retained data with rare reads, evaluate an archive tier against R2’s immediate-access classes. Include the restore time the business can tolerate and the cost of recovery tests, not just the storage rate.

If your workload includes petabytes of cold data alongside terabytes of active data, a hybrid approach makes sense: R2 for active and occasionally-accessed data, S3 Glacier or GCS Archive for truly cold data where per-GB storage cost dominates.

Metadata and queryability​

R2’s object API exposes metadata and prefix listing, not arbitrary queries over every object’s attributes. R2 SQL queries Iceberg tables; it does not turn a bucket of unrelated uploads into a searchable application catalogue.

For simple access patterns, prefix-based organisation suffices. Store user uploads at users/{userId}/uploads/{filename} and listing by prefix gives you a user's files efficiently.

For complex access patterns, separate metadata from content. Store the file in R2 with a generated key (a UUID avoids conflicts and leaking information). Store metadata in D1: the R2 key, original filename, owner, upload timestamp, file size, content type, and application-specific attributes. Query D1 to find objects; fetch from R2 to retrieve content.

D1 gives you full SQL queryability over file metadata: find all PDFs uploaded last month, find the largest files per user, find documents matching a full-text search. R2 gives you cheap, scalable blob storage with zero egress.

The decision point is query complexity. If your access patterns map cleanly to prefix-based listing, skip the D1 layer. If you need queries that cross dimensions (files by date range, files by type, files matching search terms), add D1 metadata from the start.

What R2 enables​

Zero egress makes certain patterns practical that were cost-prohibitive before.

Iterative processing becomes cheap. Re-read source data at each stage of a multi-step pipeline without accumulating transfer costs. Design pipelines for clarity and correctness, not to minimise reads.

Recovery tests avoid R2 egress fees. Downloading an independent backup from R2 does not add bandwidth charges. Reads, Infrequent Access retrieval, target storage and restore compute still cost money. Compare with the actual restore destination: a same-region S3 recovery does not incur the internet-egress rate.

Data interchange has a simpler transfer bill. Consumers in different clouds can read from R2 without R2 egress fees. The source cloud may still charge to populate the bucket, and consumers still pay for their processing.

Generosity becomes affordable. Free tiers, preview content, and public datasets cost the same as serving paying customers. You can be more generous with trials and public goods because serving them doesn't drain your budget.

What R2 gives up​

R2 has no deep archive class. A cold-storage comparison must include retained volume, retrieval frequency and urgency, minimum storage duration, and restore destination. Very low archive storage rates can dominate when data is rarely retrieved, but that conclusion depends on the full access pattern.

No certified WORM compliance. R2's Bucket Locks provide retention policies that prevent deletion and overwriting, which is useful for basic retention requirements, but compliance scenarios requiring certified immutable storage cannot rely on R2 alone. S3's Object Lock in compliance mode is truly immutable with SEC 17a-4 certification, whereas R2's Bucket Locks can be modified by administrators. Financial services firms with regulatory requirements, healthcare organisations with specific HIPAA interpretations, and legal departments with litigation hold obligations should verify whether Bucket Locks satisfy their compliance needs or whether certified WORM storage remains necessary.

Geographic controls differ. A location hint expresses a placement preference; a jurisdiction restricts where object data is stored and processed. Neither is a choice of arbitrary AWS-style region. Match the documented jurisdiction guarantee to the requirement.

Smaller ecosystem. Fewer third-party tools integrate directly with R2. S3 API compatibility means many work with configuration changes, but the ecosystem remains smaller than S3's.

Comparing with other providers​

This comparison helps you decide whether R2 fits your workload, not declare a universal winner. Each service optimises for different constraints.

Compare providers on the delivery route, storage class, retention controls and integrations the workload actually uses. Region-specific list prices are a starting point; request mix, archive retrieval and commercial discounts determine the bill.

R2 wins for egress-heavy workloads and architectures routing traffic through Cloudflare. S3 wins for feature completeness, ecosystem integration, and deep AWS dependencies. Azure Blob wins when you're committed to Azure's platform. GCS wins for Google Cloud integration.

The choice often isn't exclusive. Many architectures use R2 for egress-heavy content delivery while keeping S3 for workloads requiring archival tiers or certified WORM compliance.

R2 as multi-cloud data hub​

R2 can serve as a data interchange layer for consumers in several clouds. Reads incur no R2 egress fee; populating the bucket may incur transfer charges at the source, and consumers still pay for processing. Put shared data here when those avoided transfers justify operating the additional storage path.

Open table formats make that separation useful. Iceberg is one such format: its metadata tracks a table’s data files and snapshots, which represent table state at particular points in time. R2 Data Catalog manages Iceberg tables that compatible engines can query, allowing storage to outlive a particular query service. Check each engine's catalogue authentication and supported Iceberg features before treating compatibility as established. An open format reduces the cost of changing engines; it does not make their SQL, security policies or performance identical.

Cloudflare Pipelines, in open beta, ingests records through HTTP or Worker bindings, applies SQL transformations and writes to R2, including Iceberg tables through Data Catalog. Logpush can also feed this path; Chapter 21 covers its observability use.

R2 SQL, also in open beta, queries those tables without provisioning query servers. It supports joins, aggregations and window functions, with a 10,000-row result cap. It is read-only, and resource-intensive queries can be rejected or time out. Test representative joins and historical ranges before choosing it for a reporting deadline. An established warehouse may remain the better home for its integrations, access controls or workload management, even when R2 stores the shared source data.

Pipelines ingress is free. On the Workers Paid plan, transformations include 50 GB per month, then cost $0.04 per GB; sinks include 50 GB per month, then cost $0.03 per GB for JSON or $0.06 for Parquet and Iceberg. R2 SQL includes 10 GB scanned per month, then costs $2.50 per TB of compressed data scanned, with a 10 MB minimum per query. Storage, object operations and catalogue charges remain separate. Partitioning and selective reads can reduce scans, but a low storage bill says little about repeated joins over the entire history.

Data Catalog can automate compaction and snapshot expiration. Expiration removes data files that retained snapshots no longer reference; it does not clean files that were never referenced by a snapshot. Configure maintenance and recovery retention together, and keep a separate procedure to remove abandoned files.

From operational state to analytical data​

A database per tenant keeps operational decisions local to each tenant. It also makes a seemingly simple management question awkward: how much did every customer use last month? Running that report across live databases couples its completion to the slowest partition and competes with application traffic. The answer may combine tenants read before and after a correction, without describing any common reporting boundary.

Separate analytical data when the report needs history, joins across partitions or repeated scans that the operational path should not carry. This is a workload decision, not a requirement to introduce a data platform into every application.

QuestionStart withWhat you must own
What is this customer's current usage?Bounded operational SQLIndexes, query budget and freshness
What belongs on this known dashboard?Maintained aggregateUpdate, correction and rebuild rules
How does usage vary across tenants and months?Historical analytical storePublication, deduplication and reconciliation
How does this data join the organisation's existing reporting?Existing warehouse or lakehouseSecure ingestion and a shared definition of each measure

For a few tenants and a daily internal report, a controlled scheduled export may suffice. Record which partitions succeeded, their extraction boundaries and any omissions. An explicit partial report is more useful than an apparently complete total assembled from whichever databases answered.

Begin with a business fact​

Consider a document-processing service that reports completed pages by tenant and month. It keeps job metadata in D1 and documents in R2. Decide what counts before selecting an ingestion service: pages successfully extracted, pages accepted by the customer, or pages attempted? Retries make these different quantities. A useful event might mean “job J completed processing revision 3 with 18 accepted pages”. It should describe an authoritative transition, not an HTTP request that happened to return success.

Give that event a stable identifier, authenticated tenant identity, source partition, source history identity and sequence or revision, schema version, business timestamp, quantity and unit. The identifier remains unchanged through delivery retries and recovery of the same fact. After a source restore, use a new history identity for its sequence and reconcile surviving events across the fork; reused sequence numbers must not disguise distinct business facts. Chapter 23 explains that recovery boundary. Record ingestion time separately so late publication remains observable; choosing the month from arrival time would quietly move work across accounting periods.

For this service, define occurrence as the server's accepted completion time, in UTC. A product that reports work performed offline needs another policy for client timestamps, including clock error and deliberate manipulation. Names such as timestamp and usage conceal decisions that reporting eventually forces someone to make.

Do not infer these business facts from sampled request logs. They can explain traffic and performance, but a retried extraction may generate three successful requests for one accepted result. Keep approximate operational telemetry separate from the ledger used to justify a customer charge.

Publish without losing the mutation​

A Worker that updates D1 and then sends an event can crash between those calls. Reversing the calls allows an event for a mutation that never committed. Neither sequence ties the report to the authoritative change.

Use the outbox pattern from Chapter 9: commit the business mutation and its publication record in the same local transaction. For D1, use an appropriate atomic batch or SQL operation within that database; for a Durable Object, use its local transaction boundary. The outbox lives with the owner of the mutation. Writing it into another database reintroduces the gap it was intended to close. Make event insertion conditional on the intended transition actually occurring: an SQL update affecting zero rows is not necessarily a transaction failure.

A publisher reads pending records, submits them and records progress after successful acceptance. Give it a durable retry path independent of the initiating HTTP request: a scheduled sweep over registered partitions, for example, can find pending work even if a prompt queue notification was lost. Set limits on each sweep and track the oldest unpublished event so a large tenant cannot conceal a stalled one.

A crash after submission but before progress is recorded produces another submission with the same event identity. Pipelines documents exactly-once delivery to R2 within its service boundary. That does not atomically couple your D1 transaction to ingestion or establish that two submissions represent the same business fact. The application still needs deduplication by event identity in its curated reporting path. Preserve the raw input when useful for diagnosis; produce totals from the deduplicated representation.

Do not assume the ingestion transform supplies a durable uniqueness constraint. Choose an engine or materialisation job that can implement the required rule, and test duplicates separated by the longest replay window. If the same identifier arrives with a different payload, quarantine the disagreement. Silently selecting one turns an upstream defect into plausible accounting.

Make lateness and correction visible​

Suppose a job completes just before month-end but its publication is delayed until the next morning. A report grouped by occurrence time should include it in the earlier month. That means yesterday's report can change. Define when a report is provisional, what allows it to close and how later adjustments appear.

For example, the application might publish provisional monthly totals immediately, then close them after every included partition has been reconciled through its agreed cutoff. That cutoff needs durable evidence at each source, such as a sequence boundary recorded after all transactions covered by the reporting period. A quiet stream does not prove completeness. A missing tenant cannot become zero usage just because no events arrived.

Keep corrections explicit. If a completed job was misclassified, publish a correction referring to its original event, with the reason and revised quantity or compensating delta defined by the schema. Decide whether historical reports are restated or adjusted in the next period. Avoid overwriting a raw record while retaining an old aggregate: both may look reasonable while disagreeing.

Reconcile by tenant and period against independently calculated source totals: accepted job count, page count and the last covered source position. Retain these control totals or sufficient revisioned history at the same cutoff as the report, with the same correction policy. Totals queried from mutable current rows later may describe a different state. Investigate gaps and duplicate identities before closing. Comparing two dashboards built from the same incomplete event stream proves only that their arithmetic agrees.

For billing, store the approved statement and the version of its calculation inputs. A mutable analytical query is a useful way to prepare a statement, but it is a poor explanation of why an invoice sent six months ago had that amount.

Backfills are another publisher​

A historical export must coexist with live ingestion without duplicating or losing records. Record a source boundary, backfill through it and continue live publication after it; if the paths deliberately overlap, use the same stable identities so the overlap can be removed. A fleet of tenant databases has a set of boundaries, not one transactionally consistent global snapshot.

Current rows may not contain historical events. If an export sees only a job's latest quantity, it cannot reconstruct every earlier completion and correction. Label that import as a snapshot with its effective time rather than inventing an event history. Reports requiring historical state need records retained for that purpose.

Version the event meaning as carefully as its JSON shape. Adding an optional field is different from changing “pages processed” to “pages accepted”. During a migration, test the new transformation against retained fixtures and publish into a separate table or versioned result. Compare reconciled totals before directing readers to it. Throttle backfills so rebuilding analytics does not consume the capacity needed to accept current work.

Extend the tenant boundary through the report​

Combining tenants' events concentrates authority. The ingestion service should derive tenant identity from a trusted producer, while the reporting service checks which rows and measures its caller may see. A shared table's tenant_id column is useful organisation, but it is not an access policy.

Separate internal analytical credentials from customer-facing queries. Expose constrained reports through an application boundary, or use an engine with suitable enforced access policies. Review exports and cached report files as well as interactive queries; a correctly filtered screen can still link to an unfiltered download.

Retention also crosses the pipeline. A tenant deletion may require removal from raw events, curated tables, snapshots, exports and backups under the application's retention policy. R2 SQL cannot perform those writes. Use an Iceberg-compatible engine to perform selective deletion through the catalogue, then account for retained snapshots and file cleanup; deleting underlying R2 files directly can corrupt the table. Simply hiding the tenant in the dashboard leaves its data stored elsewhere.

The acceptance test for this design is concrete: delay one tenant's publisher, duplicate a batch, introduce a correction and rebuild a month. The report should disclose incompleteness, avoid double counting and explain every difference from the source. R2 makes retaining and sharing the data economical. The publication contract makes it trustworthy.

Migration strategies​

Migrating from S3 to R2 is mechanically straightforward: change endpoints, update credentials, copy data. But the approach matters for cost and disruption.

Sippy: incremental migration​

Sippy implements lazy migration, copying objects on first access rather than bulk-copying upfront. When a request arrives for an object not yet in R2, Sippy fetches it from S3, serves it to the client, and copies it to R2 simultaneously. Future requests serve from R2 directly.

Sippy lets normal reads populate R2, avoiding a separate copy of hot objects purely for migration. Source requests and transfer charges can still apply. If the old delivery path used CloudFront’s free S3 origin transfer, copying to R2 is not zero incremental S3 egress.

Combine Sippy with Super Slurper for optimal efficiency: start with Sippy to migrate actively-accessed objects through normal traffic, then use Super Slurper at migration end to bulk-move remaining cold data. Legacy data that's never accessed simply stays at source until you decide whether to migrate or delete it.

One critical consideration: ETags may not match between S3 and R2 when using Sippy. The system makes autonomous decisions about multipart operations for performance optimisation, which affects ETag calculation. If your systems rely on ETag matching for cache validation or change detection, plan accordingly.

Super slurper: bulk migration​

Super Slurper handles bulk migrations from S3 and Google Cloud Storage. For detailed guidance, see Chapter 27. The decision of whether and when to migrate requires more thought.

Migrate existing data when:

  • Expected transfer savings repay migration and ongoing operating costs
  • You don't need features R2 lacks (Object Lock, deep archival tiers)
  • The workload isn't deeply integrated with S3-specific AWS services
  • You have a natural migration window (major version release, architecture revision)

Start new projects on R2 when:

  • Your architecture routes traffic through Cloudflare
  • You're building content delivery, file sharing, or data interchange systems
  • You want to avoid accumulating AWS dependencies
  • Egress-heavy patterns are likely based on the application's nature

Stay on S3 when:

  • Your S3 integration is deep (many AWS services depending on S3 events, policies, integrations)
  • Compliance requires S3-specific features like Object Lock
  • You need true archival storage for cold data at scale
  • Egress costs are minimal relative to total storage costs
  • The operational overhead of changing isn't justified by the savings

Use both when:

  • Different workloads have different characteristics
  • You want to migrate incrementally and prove the model
  • Some content needs S3 features while other content is purely egress-heavy
  • Active data belongs in R2 while archival data belongs in Glacier

For existing systems, migration is rarely urgent. R2's value comes from avoiding future egress costs, not reclaiming past spending. New projects and new data are natural starting points.

What comes next​

Chapter 15 covers data whose access path matters as much as its storage: KV values cached near readers, and queries sent through Hyperdrive to an existing database. In both cases, the useful question is what freshness and latency the application actually needs.