Skip to main content

Chapter 4: Full-Stack Applications

How do I build and deploy a full-stack application with server-side rendering and static assets?


Traditional web architecture splits concerns across services: a CDN for static assets, application servers for rendering, separate API endpoints, perhaps a caching layer between them. Each boundary adds latency and operational complexity. A request might traverse CDN edge, origin load balancer, application server, and database, with each hop adding milliseconds and each service adding deployment surface area.

Workers can combine static asset delivery, server rendering and API logic behind one deployment. Rendering near the user can avoid a trip to a distant application origin, but data access still follows the database's placement and consistency model. The latency benefit depends on the complete request path, including cache misses and backend calls.

Collapsing these architectural layers, however, creates new decisions you must make thoughtfully. When everything can run at the edge, you must decide what should run there.

The rendering strategy decision​

Before choosing frameworks or writing configuration, decide how your application should render. This decision shapes framework choice, data architecture, caching strategy, and cost model.

The options​

Static generation produces HTML at build time. Every user receives identical content from cache, which delivers response times as fast as physically possible. Suits content that changes infrequently: marketing sites, documentation, blogs, landing pages.

Server-side rendering generates HTML at request time. Content can vary per user, per request, or based on real-time data. Slower than static but still fast at the edge. Suits dynamic content needing SEO or fast first paint: product pages with live pricing, personalised homepages, content changing too frequently for rebuild-on-change.

Single-page applications send a minimal HTML shell; JavaScript renders content client-side. Slowest initial load, but subsequent navigation is instant because the application is already running. Suits highly interactive applications where SEO doesn't matter: dashboards, editors, internal tools, authenticated experiences.

Hybrid approaches combine strategies: static for marketing pages, SSR for product pages, SPA for authenticated dashboards. Most real applications end up here.

The decision framework​

Static until proven dynamic. Start with static generation as your default because it's simplest, fastest, and cheapest. Add complexity only when static fails your requirements.

Move to SSR when you can't pre-generate all content variations, when SEO matters for dynamic content, when you need fast first paint with personalisation, or when content changes too frequently for rebuilds. SSR burns CPU time on every request, proportional to traffic.

Move to SPA when the application is highly interactive, when SEO doesn't matter (authenticated users), when subsequent navigation speed matters more than initial load, or when you're building an application rather than a website. The tradeoff: you shift compute cost to users' devices, affecting mobile users disproportionately.

These choices aren't irreversible, but starting with the right strategy saves migration effort.

The cost lens​

Rendering strategy has economic implications that compound at scale.

Workers Static Assets has no additional storage charge within its file limits, and ordinary asset requests are free. Running a Worker before asset delivery or enabling Workers Cache changes request billing; R2-hosted assets have their own storage and operation charges.

SSR costs CPU time per request. A million pageviews means a million renders, each consuming metered resources. Not prohibitive, but proportional to traffic in a way static generation isn't.

SPAs shift compute to the client. Your infrastructure costs drop, but users pay in battery life and data transfer. A 500 KB JavaScript bundle on every visit adds up for metered connections or constrained devices.

This economic lens clarifies decisions. Content that could be static but uses SSR because it feels sophisticated? You're paying for unneeded complexity. An SPA shipping megabytes of JavaScript to users who just want to read an article? You've externalised costs onto the people you're trying to serve.

Static asset serving​

Workers serve static assets from Cloudflare's edge network, deployed alongside your Worker code on the same global infrastructure.

Configuration fundamentals​

Enable static assets by specifying an asset directory:

Default static assets: file matches serve directly
{
"assets": {
"directory": "./public"
}
}

Files in that directory deploy as static assets. Requests matching file paths serve directly; non-matching requests invoke your Worker. Cloudflare handles content types, compression, and caching automatically.

The run_worker_first decision​

By default, Cloudflare checks for matching static assets before invoking your Worker. If a file exists at the request path, it serves directly. This is faster and cheaper for static-heavy sites.

But authentication checks, request logging, header injection, and A/B testing all require your Worker to run before any response. Enable this with run_worker_first = true and an asset binding:

Run the Worker before serving assets
{
"assets": {
"directory": "./public",
"binding": "ASSETS",
"run_worker_first": true
}
}

Now every request hits your Worker. You call env.ASSETS.fetch(request) when you want to serve a static file; this is explicit rather than implicit.

Use default (asset-first) when: Your site is primarily static and you don't need to inspect static asset requests.

Use run_worker_first when: You must authenticate before serving any content, log all requests, inject headers, or have Worker logic determine static content serving.

Start with the default; add Worker interception when needed.

Asset bundling versus R2​

Workers Static Assets supports 100,000 files per version on Paid plans and 20,000 on Free, with a 25 MiB limit per file. These are separate limits from Worker code size.

Choose bundled assets when they should change atomically with code: application scripts, styles and versioned interface assets. Choose R2 when objects have their own lifecycle: user uploads, large media and CMS-managed content. File limits matter, but update ownership is the more durable reason to separate storage from deployment.

Server-side rendering at the edge​

SSR on Workers can generate HTML close to users, shortening the path to the renderer. The benefit depends on where rendering gets its data.

When edge SSR delivers​

Moving the renderer closer to the user can help when the data it needs is nearby or cached. If it makes several sequential calls to a distant database, moving compute near that database may improve the complete response time more.

Edge SSR Requires Edge Data

Edge SSR only delivers its full latency benefits when the data your pages need is also at the edge. If every render requires a round trip to a centralised database, you've moved the rendering but not eliminated the latency.

D1 read replication can move eligible reads closer to users; session choices govern consistency. KV can serve cached values near readers, subject to eventual consistency. R2 requires an appropriate delivery and caching path. Check which data is actually local on a cache miss, rather than treating the product name as a latency guarantee.

Applications tied to a central database may benefit more from backend-proximate rendering. Compare complete render time, including serial database calls and cache misses, before choosing between user-proximate and origin rendering.

Edge SSR wins when: Users are geographically distributed and data is available at the edge through Cloudflare services.

Origin SSR wins when: Rendering requires large unreplicable datasets, CPU requirements exceed Worker limits, or you're migrating incrementally.

Hybrid works when: You can render some pages at the edge and fall back to origin for complex cases. Simple product pages use edge SSR; complex report generation uses origin.

Raw HTML versus frameworks​

At its simplest, server-side rendering means querying data, rendering it through a template and returning HTML. A small template layer can be sufficient, provided it escapes untrusted values for their output context.

Raw HTML generation suits internal tools, admin interfaces, simple applications with few routes, and teams preferring explicit control because the code is obvious, debugging is straightforward, and dependencies are minimal.

But raw generation doesn't scale to complex applications. Dozens of routes, nested layouts, client-side hydration. Frameworks provide structure that raw concatenation lacks.

The choice is about maintenance cost and team preferences, not fundamental capability. Frameworks trade dependency complexity for structural conventions that guide developers. Raw generation trades those conventions away for transparency and simplicity.

Streaming for perceived performance​

Streaming can send useful HTML before every dependency has completed. It is valuable when a page has a fast, useful shell and slower sections that can arrive independently.

Choose the boundary by what the reader can use, not an arbitrary duration threshold. A partial page that cannot be understood or interacted with offers little benefit. Measure first useful content as well as total completion, and decide how failures in later chunks appear after headers have already been sent.

HTMLRewriter: static performance, dynamic reality​

HTMLRewriter is a streaming HTML parser and transformer built into Workers. It parses HTML properly and allows surgical modifications without buffering entire documents in memory. Unlike full SSR, it transforms existing content rather than generating HTML from scratch.

This hybrid model serves static content but transforms it at the edge. Base HTML caches indefinitely; transformations apply per-request. Static performance, dynamic capability.

Key use cases include: injecting user-specific data into static pages (authentication state, personalisation tokens), A/B testing without maintaining separate static versions, adding analytics scripts to every page, localising content based on request properties.

Use HTMLRewriter when page structure is static but specific values must vary per request. You're injecting dynamic values rather than restructuring the HTML.

Use SSR when page structure varies per request based on data or user properties. Content doesn't exist until request time.

HTMLRewriter avoids constructing a full DOM for the document. It fits small transformations of an existing response; a framework remains easier to maintain when the page needs substantial layout logic or coordinated client behaviour.

Framework integration​

Complex applications benefit from component models, routing abstractions, and build optimisations. Workers integrate with major frameworks through adapters that translate framework conventions to edge execution.

Choosing a framework​

The framework decision follows from rendering strategy and team expertise.

React Router (formerly Remix) offers the best Cloudflare integration: first-class support, excellent documentation, designed for edge deployment. The default choice for new React projects.

Astro excels at content-focused sites with its hybrid static/dynamic model. Pages not needing runtime data pre-render at build time; others render on demand. Natural fit for documentation, marketing sites, or content-heavy applications with some dynamic sections.

SvelteKit and Nuxt have official adapters and work well. If your team already knows Svelte or Vue, that learning curve advantage outweighs any integration smoothness difference.

Next.js works through OpenNext, a community-maintained adapter. More complex integration than purpose-built alternatives. For existing Next.js applications you're migrating, the investment makes sense. For new projects, ask whether ecosystem benefits outweigh integration complexity.

The framework tradeoff​

Frameworks provide structure: routing conventions, component models, build optimisation, ecosystem access. Teams familiar with React Router or Astro can be productive immediately.

But frameworks are abstractions over Cloudflare's platform, and abstractions evolve. React Router's Cloudflare integration is excellent today. If project priorities shift or the integration changes, your code adapts to the framework's timeline, not yours.

For long-lived applications, consider how much framework surface area you're adopting. Thin frameworks cost less to maintain than thick ones that abstract away platform specifics. Choose consciously: know what you're gaining (structure, ecosystem, team familiarity) and trading (coupling to abstractions that may evolve differently than your application).

Binding access across frameworks​

Regardless of framework, Cloudflare bindings (D1 databases, KV stores, R2 buckets, Durable Objects) are available through framework-specific context objects. React Router exposes them in context.cloudflare.env. Astro provides them in Astro.locals.runtime.env. SvelteKit uses platform.env.

Syntax differs; capability is identical. Your framework choice doesn't limit access to Cloudflare services, only how you access them.

The Vite plugin​

Many frameworks use Vite as their build tool. Cloudflare's Vite plugin provides development integration across frameworks: hot module replacement, automatic TypeScript type generation for bindings, local simulation of Cloudflare services.

For frameworks without dedicated Cloudflare adapters, the Vite plugin often provides sufficient integration: local development that approximates production behaviour without framework-specific machinery.

The plugin integrates with @vitejs/plugin-rsc, the official Vite plugin for React Server Components. A childEnvironments option allows multiple environments within a single Worker, enabling RSC patterns where a parent environment imports modules from a child environment to access a separate module graph. For a typical RSC setup, you configure a viteEnvironment named "rsc" with childEnvironments: ["ssr"]. This is the lower-level infrastructure that frameworks like React Router build upon, and it means RSC support on Cloudflare doesn't depend exclusively on Next.js adapters. Teams wanting React Server Components with more control over their framework layer can build directly on these primitives.

The plugin also supports auxiliary Workers: additional Workers defined alongside your main application and callable via service bindings. This enables splitting specific concerns (background processing, shared validation logic, isolated security-sensitive operations) without abandoning the unified build and deploy workflow. Your framework handles the main application; auxiliary Workers handle specialised tasks; everything builds and deploys together.

Moving an existing application​

Framework support does not establish application compatibility. Inventory provider-specific APIs, native dependencies, cache assumptions and data placement before changing the deployment target. Next.js applications use the OpenNext Cloudflare adapter; the Pages adapter is not the route for new Workers deployments.

Test one representative dynamic route with its real dependency pattern. A successful static build says little about database access, image processing or background work. Keep a working comparison deployment until you can explain differences in behaviour and operating cost. Chapter 27 develops this migration decision in detail.

Single-page application patterns​

SPAs present a routing challenge unique to client-side applications: every route should serve the same HTML shell, with JavaScript handling navigation.

SPA routing configuration​

Configure this behaviour in your asset settings:

SPA fallback for unmatched navigation requests
{
"assets": {
"directory": "./dist",
"not_found_handling": "single-page-application"
}
}

SPA navigation requests can fall back to /index.html, letting the client-side router interpret the URL. That routing can serve HTML without invoking your Worker, including when a user navigates directly to an API URL unless you configure the Worker to run first for that path.

API routes alongside SPAs​

Most SPAs need backend endpoints for data operations. Set assets.run_worker_first to ["/api/*"] so API paths reach the Worker for both browser navigation and client fetches. Handle those paths in the Worker and return an API response, including an API 404 for unknown endpoints. Let other navigation paths reach the SPA shell.

Everything ships in one deployment. Your SPA, its assets, and its API version together and route through the same Worker.

Edge authentication for SPAs​

Traditional SPAs have a security gap: the JavaScript bundle downloads before any authentication check. Users receive your application code, then check authentication. For many applications this is fine; the code isn't secret and API endpoints are protected.

But some applications can't tolerate this: proprietary logic in the JavaScript, compliance requirements about code access, or simply the desire not to serve application code to unauthenticated users.

Edge authentication closes this gap. Your Worker validates authentication before serving anything: API responses, static assets, or the SPA shell. Unauthenticated requests redirect to login without receiving protected content.

Set assets.run_worker_first to true when every asset needs authentication. Validate the session before invoking asset serving. This provides an integrated equivalent of the authenticated origin or reverse proxy that could protect an origin-hosted SPA.

Existing Pages applications​

Workers supports static assets, server rendering, APIs, builds and Previews. Use it for new applications. An existing Pages application needs a reason to migrate beyond product consolidation: access to a Workers capability, a simpler application boundary, or a deployment requirement Pages does not meet.

Treat migration as a routing and resource review. File-based Functions routes, redirects, middleware, environment bindings and preview isolation need explicit equivalents. Validate those behaviours before switching traffic; matching build output alone does not establish parity.

Structuring full-stack applications​

With static assets, SSR, and API routes unified in Workers, project structure becomes a design decision rather than a platform constraint.

The default: keep things together​

For most applications, consolidation beats distribution across services. Your Worker handles routing, API logic, and SSR together. Static assets deploy alongside. Tests cover the whole application. Deployment is atomic and coordinated.

Splitting into multiple Workers creates coordination overhead: latency on each service binding call, debugging across service boundaries, deployment order dependencies. Microservice benefits (independent scaling, team autonomy, technology heterogeneity) require scale and organisational complexity that most applications don't have.

When splitting makes sense​

Split a monolithic Worker into multiple Workers when you have specific reasons:

Deployment frequency divergence. One part deploys hourly, another monthly. Coupling means unnecessary deployments or blocked changes.

Resource requirement conflicts. One endpoint needs maximum CPU time while another needs minimal latency. Separate Workers allow independent configuration.

Team boundaries. Genuinely independent teams own different services and coordination cost exceeds the cost of service boundaries.

Security isolation. One service handles particularly sensitive operations and you want defence in depth through separation.

Don't split preemptively or because you might need independent scaling someday. Split only when coupling cost exceeds the coordination cost of service boundaries, not before.

If you're using the Vite plugin with a full-stack framework, auxiliary Workers offer a middle ground: define additional Workers in your Vite configuration that build and deploy alongside your main application. You get service binding performance without managing separate deployment pipelines. This works well for extracting specific concerns while maintaining a unified development workflow.

Caching strategy​

Static assets cache automatically without configuration. Cloudflare handles cache headers and edge distribution transparently. Deployment handles invalidation by generating new asset URLs.

Server-rendered HTML benefits from caching when content doesn't vary per user. The question is staleness tolerance. How long can users see potentially outdated content? Minutes for some pages, seconds for others.

The Cache API allows explicit control: store rendered responses with appropriate TTLs, serve from cache when valid, regenerate when stale. The decision isn't whether to cache but how long and under what invalidation conditions.

API responses often shouldn't cache, but read-heavy endpoints with staleness tolerance benefit from short TTLs. Cache for 60 seconds and most users hit cache rather than compute. The tradeoff is staleness versus cost.

Two architectural constraints shape which caching mechanism you reach for. First, the cf caching options on fetch() (cacheTtl, cacheEverything, cacheTtlByStatus) are silently ignored when fetching from an origin proxied through Cloudflare on a different Cloudflare zone. The request routes to the origin zone's edge, where that zone's own cache settings apply; your Worker cannot override them. This matters for multi-Worker architectures where services sit on separate zones, because the caching behaviour you configured simply does not take effect. For cross-zone fetches, use the Cache API or KV instead.

Second, the Cache API is local to the data centre handling the request; Tiered Cache does not apply to its entries. A miss must regenerate or fetch the response. Cache API objects also have a documented size limit, so streaming a response does not make its cacheability unlimited. Choose the Cache API when that locality is acceptable; use a separate data service when the application needs a durable or globally shared record.

For origin responses that vary by language or format, Cache Rules can use the origin's Vary header to keep representations separate. Normalise negotiation headers when several request values should produce the same representation; preserve exact values when those differences matter. Normalisation can rewrite the preferences sent to the origin, so test the returned language and format as well as the cache hit rate. Bypass caching for unexpected or per-user variation such as cookies. A smaller cache key is useful only if it still identifies the right response.

This is CDN cache configuration, distinct from the Cache API and Workers Cache. It needs an explicit Vary rule, consistent Vary headers from the origin, and a purge when a policy change must invalidate existing variants.

Workers Cache: the cache that follows the Worker​

Workers Cache changes this calculus for whole-response caching. It sits in front of your Worker's fetch handler: when an incoming request matches a cached response, Cloudflare serves it without invoking your Worker at all. Configuration is a per-entrypoint enable flag plus the standard Cache-Control headers your responses already carry; there are no rules, no page settings, and no zone involvement. The cache follows the Worker, not the zone, so it behaves identically on custom domains, on workers.dev, and for Workers invoked purely through service bindings, and zone-level cache configuration has no effect on it.

Three properties make it architecturally interesting rather than merely convenient. The cache key includes the entrypoint and ctx.props, the properties supplied by the calling Worker, so a gateway can dispatch to a cached entrypoint with a tenant identifier in ctx.props and get per-tenant cache partitioning without inventing key schemes. Purging is programmatic and scoped: a Worker calls ctx.cache.purge() with cache tags from inside the entrypoint that owns the entries, so a write handler can invalidate exactly the responses it made stale, immediately. And because service-binding and loopback calls are cacheable, you can wrap an expensive Durable Object aggregation in a cached entrypoint: hits skip the wrapper and the object entirely, while the write path purges by tag.

Authenticate and authorise in an uncached gateway before dispatching to a cached entrypoint. Derive ctx.props from verified identity and include every permission scope that changes the response in the cache key; a tenant key is insufficient when users have different document permissions. Cookies, custom headers and hostnames do not automatically separate entries. Use private or no-store when a protected response cannot safely share a representation. Checks inside a skipped handler cannot protect a cache hit.

The billing consequence deserves attention before you enable it. With caching on, every request to the Worker bills at the standard request rate, including cache hits and requests that are normally free (static asset requests and Worker-to-Worker invocations), while CPU time bills only when the Worker actually runs. That's usually an excellent trade for compute-heavy or Durable-Object-backed responses, and a poor one for Workers that mostly serve static assets. The standard bypass rules apply: only GET and HEAD are cached, responses with Set-Cookie and requests with Authorization bypass automatically, and custom RPC methods never touch the cache. At launch, cached response sizes are capped at the free-plan cache limit on every plan, a restriction Cloudflare says is temporary.

Where does this leave the other mechanisms? Workers Cache for whole HTTP responses you own; the Cache API for per-colo fragment caching inside a request; KV for data that must be globally readable on first access; the cf fetch options for controlling how Cloudflare's CDN caches an origin you fetch from. Four tools, four scopes; Workers Cache is the one to reach for when you own the whole response and want to skip the work of producing it again.

What comes next​

Rendering and caching choices need tests that exercise their failure modes. Chapter 5 explains what local development can reproduce, what requires deployed infrastructure, and how to shorten the path from a failing request to a useful diagnosis.