Chapter 16: The AI Stack on Cloudflare
Can we build AI-powered applications on Cloudflare, and what's the architecture?
Cloudflare's AI stack prioritises operational simplicity and platform integration. Its catalogue includes smaller models and larger reasoning models such as Kimi K2.6, with a 262K context window and multimodal input. Evaluate the hosted options against the task before deciding that an external provider is necessary.
AI Gateway can route to external providers through a common interface and centralise operational controls. Changing a model identifier is easy; changing provider behaviour still requires tests for prompts, tools, response formats, latency and cost.
This chapter covers how Cloudflare's AI infrastructure works, where it outperforms alternatives, and where you should use something else. Chapters 17 through 19 cover implementation.
The core trade-off
Workers AI runs models ranging from efficient open-source options (Llama, Mistral, Qwen) to frontier-class models (Kimi K2.6) on Cloudflare's GPU infrastructure. The catalogue spans classification, summarisation, embeddings, generation and tool use, with capabilities varying by model. External providers expand the choice. Neither hosting location nor model size establishes which one will handle your failure cases well.
Judge model quality by the task’s failure cost. A slightly weaker summary may be acceptable in an internal preview; a misclassified payment instruction may not be. Compare acceptable outputs, latency and cost on the same workload before choosing where inference runs.
If your users can't tell the difference between Llama 3 and GPT-4 for your specific task, you're paying a tax on sophistication you don't need. If they can tell the difference and it matters, use the better model. AI Gateway routes those requests while keeping observability on-platform.
The useful boundary is control. Workers AI suits applications that can use its hosted models and serving behaviour. Training, unsupported models, specialised deployment or particular capacity commitments can require another platform, whether AI is a feature or the whole product.
How "edge AI" actually works
"AI at the edge" suggests models running in every Cloudflare location alongside your Workers, but that's not how it works, and the distinction matters for latency calculations.
Workers AI runs on GPU infrastructure rather than inside the Worker isolate. A binding call still includes routing and inference work. Do not infer the GPU’s location or a same-continent guarantee from the location of the Worker.
Measure time to first token and time to complete the requested output separately. Long outputs can be dominated by generation even when network transit is short. Compare providers from your deployed path; an assumed US endpoint is not a fair model of every external service.
Scaling and capacity
Workers AI doesn't have cold starts in the traditional sense because models are loaded and ready, though inference requests queue when GPU capacity is constrained, and latency increases as requests wait for available GPU cycles.
Cloudflare scales GPU capacity automatically, but GPU scaling is slower than Worker scaling, so traffic spikes might see elevated latency for seconds or minutes whilst capacity adjusts. Cloudflare doesn't publish percentile latency guarantees under load, so test your specific workload patterns before committing to latency-sensitive production use.
Measure latency distributions at expected concurrency and during bursts. A median from a model benchmark is not a tail-latency commitment. For a strict waiting budget, define a fallback such as cached content, another provider or an explicit unavailable response.
The product stack
Cloudflare's AI offerings form a coherent stack. Understanding what each solves prevents both over-engineering and under-building.
Workers AI: serverless inference
Workers AI exists because GPU infrastructure is operationally complex, with provisioning GPUs, managing CUDA drivers, handling model loading, and scaling capacity being problems you avoid by calling Workers AI. Simply configure a binding, invoke a model, and receive a response.
The trade-off is serving control. Workers AI hosts a supported catalogue; it is not arbitrary model hosting or a model-training service. LoRA inference is available in open beta for supported models, using adapters trained elsewhere. For a required proprietary model or unsupported adapter, use another provider, optionally through AI Gateway.
Workers AI handles well embeddings for semantic search, classification and extraction, summarisation of moderate-length documents, conversational responses where speed matters more than sophistication, code generation for common patterns, and increasingly, complex agent workloads requiring long context and tool calling via frontier-class models.
External providers remain the right choice when you need specific proprietary models, fine-tuning capabilities, or have tested Workers AI models on your workload and found them inadequate for your quality requirements.
Vectorize: when and why
An embedding represents content as a numerical vector, allowing similarity search to find related material. Vectorize stores and searches those vectors for semantic search, recommendations and retrieval-augmented generation (RAG), which supplies relevant source material to a model before it answers. Should that vector storage be Vectorize?
| Factor | Choose Vectorize | Choose External Vector DB |
|---|---|---|
| Platform integration | Single bill and native resource bindings | Additional vendor, latency for external calls |
| Scale | Handles millions of vectors adequately | Pinecone, Weaviate proven at massive scale |
| Query capabilities | Similarity search plus metadata filtering | Advanced filtering, hybrid search, more query options |
| Operational model | Managed, minimal configuration | More control, more operational burden |
| Team familiarity | New tooling to learn | Pinecone/Weaviate may be known quantities |
Vectorize constraints
- Maximum 1536 dimensions per vector (sufficient for most embedding models)
- Metadata filtering for scoped queries
- Namespaces for multi-tenant isolation
- 20 million vectors per index on paid plans
These constraints rarely limit typical applications, but verify against your requirements.
Vectorize isn't the most feature-rich vector database available, but it's adequate for most RAG and search applications, integrates cleanly with Workers, and eliminates a vendor relationship. If you need advanced capabilities (hybrid keyword-and-semantic search, complex filtering, proven performance at billions of vectors), evaluate Pinecone or Weaviate.
Vectorize is a building block, not a complete RAG solution. You provide embeddings; Vectorize stores and retrieves them. Chunking strategy, embedding model selection, retrieval logic, and re-ranking remain your responsibility. Chapter 18 covers building complete RAG pipelines.
AI Gateway: optionality as architecture
AI Gateway provides a common inference interface and operational controls for supported models. The Workers AI binding can reach both Cloudflare-hosted and external models; REST interfaces support clients outside Workers. Choose provider credentials or Unified Billing deliberately, and decide which requests need gateway logging, caching and routing.
The immediate value is observability; without AI Gateway, understanding AI costs requires aggregating logs from multiple providers, correlating request patterns, and building analytics pipelines. AI Gateway provides per-request logging, cost tracking, latency analysis, and error rates by provider and model, and you can attach custom metadata (user ID, environment, workflow) to break down spend by attributes the dashboards expose natively: free vs. paid users, individual customers, or specific workflows.
The strategic value is a stable integration boundary. Keep provider choice in configuration, then qualify each candidate against the application’s quality, tool-use and data-handling requirements. A fallback that returns an answer is useful only if it preserves those requirements.
The abstraction is not perfect because providers have different prompt conventions, token limits, response shapes, and capabilities; switching from GPT-4 to Llama still requires prompt adjustments and quality testing. AI Gateway handles the API mechanics so you can focus on model-specific concerns rather than HTTP plumbing.
AI Gateway adds clear value when you use multiple providers, want unified observability, need automatic failover (if a provider returns errors or times out, the gateway retries against another), or when repetitive queries benefit from caching. For agentic workloads there is a subtler benefit: the gateway buffers a streamed response independently of the caller's connection, so an agent that drops mid-stream can reconnect and resume rather than re-running the inference and paying for it twice. The overhead isn't worth it when you have a single provider with no plans to change, every millisecond of latency matters, or when highly variable queries yield negligible cache hit rates.
For most applications building on Cloudflare, AI Gateway is worth the minimal overhead. Even if you use only Workers AI today, routing through AI Gateway preserves future flexibility at negligible cost.
AI Search: managed RAG
AI Search handles RAG end-to-end by accepting data sources (R2 buckets, web URLs), chunking documents, generating embeddings, building an index, and answering questions with retrieved context, so the entire pipeline becomes configuration.
AI Search manages ingestion and retrieval while exposing settings for chunking, embedding models and search behaviour. Test those controls against representative questions before building a custom pipeline. Choose custom RAG when a required ingestion path, access rule or retrieval step cannot be expressed in the managed service.
AI Search Constraints
- Up to 100 instances on the Free plan, 5,000 on paid plans
- Up to 100,000 files per instance on Free; 1 million on paid (500,000 if hybrid search is enabled)
- 4 MB maximum file size
- Limited control over chunking parameters
- Built-in storage and vector index (no separate R2 bucket or Vectorize index needed)
- Hybrid search (vector + BM25 with reciprocal rank fusion) and metadata-based relevance boosting available on the instance
These constraints now suit applications well beyond "RAG as a feature." With the ai_search_namespaces Workers binding, instances can be created at runtime, making per-tenant knowledge bases practical for B2B SaaS.
Start with a retrieval evaluation, not a build-versus-buy slogan. If AI Search’s controls meet the quality, access and freshness requirements, custom ingestion adds work without a demonstrated benefit. If a failed requirement needs a different pipeline, that failure tells you what to build.
Agents SDK: stateful AI applications
The Agents SDK provides a framework for AI agents: applications where models maintain conversation state, use tools, and take actions across multiple interactions. The SDK builds on Durable Objects, inheriting their single-threaded consistency guarantees.
When do you need the Agents SDK versus Durable Objects directly? The SDK provides abstractions for common patterns: conversation history management, tool registration and execution, and structured prompting.
If your "agent" is simple (a chatbot with conversation history), you might not need the SDK, as a Durable Object storing messages and calling Workers AI handles that without additional abstraction. The SDK earns its complexity when you need tool use (the model calls functions you define), multi-step reasoning (the model chains multiple operations), or sophisticated conversation management (branching dialogues, context windowing, memory summarisation).
Chapter 19 covers the Agents SDK in detail. For strategic purposes: it's the right choice for serious agent applications, overkill for simple conversational features.
Data privacy and compliance
For enterprises evaluating Cloudflare's AI stack, data privacy is often decisive.
Review data handling for each inference path. Workers AI’s no-training policy does not establish the retention behaviour of every external provider reached through AI Gateway, and gateway payload logging creates a separate stored copy. Set logging and provider controls to match the data being sent.
AI Gateway logging is configurable at three levels of detail. Full logging stores request and response payloads alongside metadata, useful for debugging and quality analysis. Metadata-only logging (controlled per request via the cf-aig-collect-log-payload: false header) skips payload storage while still recording token counts, model, provider, status code, cost, and duration, giving you usage analytics without persisting sensitive prompts or completions. Disabling logging entirely sacrifices all observability for maximum privacy. For most production workloads handling user data, metadata-only logging strikes the right balance: you retain the cost and performance visibility you need for operations without storing the content that creates compliance exposure.
Compliance certifications follow Cloudflare's broader platform: SOC 2 Type II, ISO 27001, and GDPR compliance, with Workers AI inheriting these as part of the Workers platform.
Data residency considerations for Workers AI are less mature than for other Cloudflare primitives. Inference occurs on GPU clusters whose locations aren't published. For workloads with strict geographic requirements, verify residency guarantees directly with Cloudflare.
Document what data leaves the application, which service processes it and which copies are retained. Verify service-specific commitments rather than treating account-wide certifications as proof that a particular inference route meets a requirement.
When things break
AI systems fail in ways that traditional systems don't. Understanding failure modes shapes resilient architectures.
Workers AI failures
Workers AI can fail due to capacity constraints (GPU queuing), model issues, or infrastructure issues (GPU cluster unavailable), which manifest as elevated latency, error responses, or timeouts.
Detection requires monitoring beyond standard request metrics by tracking AI-specific signals: inference latency percentiles, error rates by model, and timeout rates. Alert when latency exceeds thresholds that affect user experience, not just when requests fail outright, because a 5-second response from a feature that should take 500ms is a failure even if it returns HTTP 200.
Recovery strategies depend on criticality. For features that can fail gracefully (a "related content" widget, a summarisation preview), return cached content, show a placeholder, or hide the feature temporarily. For features where AI is essential (a search interface requiring semantic understanding), implement fallbacks to alternative providers through AI Gateway with clear degradation paths when all providers fail.
Model deprecation
Cloudflare's model catalogue evolves and models deprecate with notice, so your application must handle transitions. Design for model flexibility from the start by abstracting model selection behind configuration, testing against multiple models during development, and monitoring for deprecation announcements.
Quality degradation
Traditional software fails loudly. A broken API returns 500 errors, a crashed process stops responding, a timeout trips an alert. Your operational dashboards catch these failures because they manifest as measurable deviations in error rates, latency, and availability. AI features fail differently. A model update might handle your prompts in subtly worse ways, load patterns might cause inconsistent behaviour, and prompt changes that improve one use case might harm another, all while returning HTTP 200 with well-formed responses. Every operational metric stays green; your users are the first to notice something is wrong.
This distinction matters architecturally because it means standard operational monitoring is necessary but insufficient for AI features. You need a separate quality signal that evaluates what the model produces, not just whether it responds. For classification tasks, track accuracy against a maintained set of labelled samples run on a regular schedule. For generation tasks, implement human review sampling or automated quality scoring against known-good examples. For RAG applications, monitor retrieval confidence scores over time as Chapter 18 describes; a decline is a reason to inspect labelled retrieval outcomes, ingestion health and changes in query mix.
Quality monitoring is harder than availability monitoring, but undetected quality degradation is worse than detected outages. An outage is visible, urgent, and gets fixed. A subtle quality regression can persist for weeks, eroding user trust before anyone realises the system is at fault rather than the users' expectations.
Cost architecture
AI costs can dominate infrastructure spending. Understanding the cost structure enables informed decisions.
Price a workload, not a model label
Use the published input, output and cached-input rates for the exact model. For example, @cf/meta/llama-3.1-8b-instruct-fp8 is priced at $0.152 per million input tokens and $0.287 per million output tokens. At 500 input and 200 output tokens per request, 1,000 requests cost about $0.13 in inference before included allowances. Longer prompts, retries and additional agent turns change that estimate.
Compare candidates on the same successful task. A model with a lower token rate can require more prompt context or more repair. At small volumes integration effort may dominate; at sustained high volumes, compare managed inference with the full staffing and capacity cost of self-hosting. No fixed request count decides that trade.
Optimisation strategies
Right-size model selection provides the largest cost impact by using the smallest model that produces acceptable results. Classification tasks rarely need 70B parameter models; 7B models often suffice at 10x lower cost. Test and measure.
AI Gateway caching reduces costs for repetitive queries, with FAQ-style applications possibly seeing 50%+ cache hit rates and effectively halving AI costs. Enable caching and monitor hit rates, then adjust cache TTL based on how quickly your content changes.
Prompt engineering for efficiency matters at scale, since verbose prompts cost more tokens. Halving input tokens halves the uncached input-token charge at the same model rate; output tokens and additional attempts still contribute to the total. Measure the complete request cost while checking that the shorter prompt preserves quality.
Decision frameworks
Should you use Cloudflare's AI stack?
The decision tree is simpler than it appears:
Do you need a specific proprietary model (GPT-4, Claude, Gemini)? Use those providers directly (potentially through AI Gateway for observability). If you need frontier-level quality but are not tied to a particular model family, evaluate whether Workers AI's frontier-class options like Kimi K2.6 meet your requirements before adding an external dependency.
Is latency critical? Benchmark the complete path at expected load and output length. A regional inference label does not establish a sub-200ms response, and an external provider is not automatically slower.
Already building on Cloudflare? Workers AI integrates naturally with zero additional vendors; if not, integration complexity is similar across providers and you should evaluate based on features and cost.
Value multi-provider flexibility? Route through AI Gateway from the start, as the optionality is worth the minimal latency overhead. If you're certain you'll never switch providers (unlikely), direct integration is marginally simpler.
The hybrid strategy
Most serious applications benefit from a hybrid approach, using Workers AI for high-volume, latency-sensitive, quality-tolerant tasks and frontier models for low-volume, quality-critical tasks.
A single application might use Workers AI embeddings for semantic search (high volume, quality adequate, latency matters), Workers AI classification for content categorisation (high volume, quality adequate), GPT-4 through AI Gateway for customer-facing complex responses (lower volume, quality critical), and Workers AI summarisation for internal tools (quality tolerance higher for internal use).
The key is identifying which tasks are quality-tolerant and which aren't; test rather than assume by running your actual prompts through both Workers AI and frontier models. If users can't tell the difference, Workers AI saves money and latency; if they can tell and it matters, use the better model.
AI Gateway makes hybrid strategies practical with unified logging across providers, consistent authentication patterns, and a single interface for cost tracking; without AI Gateway, hybrid strategies require maintaining multiple provider integrations and aggregating observability manually.
Latency budget planning
Give the feature a waiting budget, then measure where it is spent: fetching context, embedding, retrieval, queueing and generation. For interactive output, distinguish first useful content from completion of the whole answer. Streaming can improve progress feedback without reducing total generation time.
Leave room for the slow tail and decide what happens when the budget expires. A partial answer, a cached answer, a queued job and an unavailable response are different product choices. Test the chosen fallback as part of the feature, not only during an outage.
Build vs buy for AI features
Before building AI features on any infrastructure, ask whether building is the right choice.
Many applications don't need custom AI infrastructure but instead need a chatbot, a search feature, or content classification, and purpose-built products often serve better than primitives. Intercom handles support chat, Algolia handles search with AI features, and various vendors handle content moderation.
Build on Cloudflare's AI stack when your AI feature is differentiated (not commodity), you need control over prompts and behaviour, you want to avoid per-seat SaaS pricing at scale, or the AI feature integrates tightly with your application logic.
Buy a purpose-built product when the feature is commodity (standard support chat, basic search), time-to-market matters more than customisation, the vendor's model quality and training data exceed what you'd achieve, and operational simplicity outweighs control.
Compare the product you would buy with the behaviour you need to own: access control, workflow integration, evaluation and operating cost. If the product meets those needs, custom RAG needs a specific reason to exist. If it does not, the missing requirement defines the scope of the build.
What comes next
Chapter 17 makes the inference choice concrete: evaluate a model on your own tasks, budget its latency and token use, and decide what the application should do when inference fails. Chapters 18 and 19 build retrieval and agent behaviour on that foundation.