Chapter 18: Building RAG Applications
How do I build applications that combine search with generation?
A model asked about your company’s policies, product documentation or yesterday’s support tickets may lack the evidence needed to answer. It can still produce a confident fabrication. Prompt wording alone cannot supply missing facts or guarantee that the model recognises what it does not know.
Retrieval-augmented generation gives the model relevant evidence before it answers. It makes some failures easier to diagnose: was the source absent, missed by retrieval or misused during generation? It does not eliminate hallucination, and a citation is useful only if the cited material supports the claim.
A RAG system that retrieves the wrong chunks fails confidently, with citations. Users trust cited answers more than uncited ones. Bad retrieval with good generation produces confidently wrong answers harder to catch than obvious hallucinations.
Retrieval introduces another budget: how much context to include. Too little can omit the answer; too much can bury it or exceed the model’s context window. Effective RAG retrieves enough evidence to support the response and declines to invent what the sources do not establish.
This chapter covers RAG on Cloudflare: the architectural decisions that determine success, when to use managed versus custom pipelines, and edge-specific patterns that distinguish Cloudflare RAG from traditional cloud implementations.
Index when content changes; retrieve for each question
Loading diagram…
RAG economics
Budget indexing and queries separately. Indexing includes extraction, chunking and embedding whenever content changes. Query cost includes embedding the question, retrieval, optional reranking and generation with the retrieved text. Storage and query charges depend on vector dimensions; generation depends on the tokens actually used.
Measure those components on a representative corpus and question set. Generation may dominate a chat product, while frequent reindexing or retrieval without generation produces a different bill. The claim that inference is always 95% of cost is not a useful planning assumption.
Chunk size connects quality to cost. Smaller chunks produce more vectors and may require more retrieval to restore context. Larger chunks can include unnecessary text in every generation. Choose the combination that answers the evaluation questions with adequate evidence, then compare its cost.
RAG on Cloudflare
Workers bindings simplify the integration of inference, vector search and storage. They do not prove that every stage executes at the user’s nearest location or that every regional failure is isolated. Measure the entire deployed path, including external models and content fetches.
Vectorize’s 20-million-vector limit per Paid index allows substantial corpora. Beyond it, partitioning adds routing and result-merging work. Compare a custom Vectorize pipeline, managed AI Search and an external engine against the same corpus, filters, freshness requirement and query load.
Chunking strategy
Chunking decides which evidence can travel together into retrieval. Choose boundaries that preserve the information needed to answer a question, then test whether the retrieved chunks are complete enough without carrying excessive unrelated text.
Smaller chunks (200–300 tokens) improve retrieval precision. When a query matches a small chunk, you know exactly which content is relevant. But small chunks lack context. A paragraph explaining a concept may need the preceding paragraph to make sense; retrieve the second without the first and you get technically accurate but practically useless content.
Larger chunks (500–1000 tokens) preserve context. Retrieve one and you have enough surrounding material to understand it. But larger chunks dilute relevance. If your query matches one sentence in a 500-token chunk, you're injecting 499 irrelevant tokens into your context window: tokens that cost money and attention.
Overlapping chunks reduce the chance of splitting relevant evidence across a boundary. They do not guarantee that a complete answer survives in one chunk, and they increase indexing work and duplicate context. Start with modest overlap and inspect failed retrievals before increasing it.
Semantic chunking respects document structure. Split on paragraph boundaries, section headers, or logical divisions rather than arbitrary token counts. Authors grouped related information together for a reason. But semantic chunks vary wildly in size, complicating context window management.
Choosing a chunking strategy
| Document type | Recommended approach | Rationale |
|---|---|---|
| API documentation | Semantic chunking on endpoints/sections | Each endpoint is a logical unit; splitting mid-endpoint creates useless fragments |
| Legal contracts | Semantic chunking on clauses, 500–800 tokens | Clauses are self-contained; smaller chunks lose critical context |
| Support tickets | Small chunks, 200–300 tokens, minimal overlap | Conversational content has natural boundaries; precision matters more than context |
| Technical guides | Medium chunks, 400–500 tokens, 15% overlap | Balance between context preservation and retrieval precision |
| Mixed corpus | Document-type routing to appropriate strategy | One size doesn't fit all; invest in classification |
Include chunking in evaluation alongside embedding and ranking choices. A 300–500-token semantic chunk with modest overlap is a starting hypothesis. Inspect failed questions: missing context may justify larger chunks, while irrelevant material may justify smaller or more selective ones.
Vectorize
Vectorize stores embeddings and searches for similar vectors through a Workers binding. Your application supplies the content pipeline, access rules and context assembly.
When Vectorize fits
A Paid Vectorize index supports twenty million vectors: four million documents at five chunks each, for example. Newly upserted vectors typically become queryable in under 30 seconds, but that is a median freshness observation, not a deadline. Pipelines needing immediate visibility must track indexing progress or query the authoritative source.
Vectorize struggles at extreme scale. Exceeding twenty million vectors requires multiple indexes with routing logic: a Worker examining query metadata to select the appropriate index, with cross-index queries requiring fan-out and merge. If you need features Vectorize lacks (built-in hybrid search, complex filtering beyond metadata, custom similarity functions), you'll work around limitations or choose alternatives.
Dimensions and embedding models
Select embeddings against representative query-document pairs before indexing the corpus. Test language coverage, domain vocabulary and the exact identifiers users search for. More dimensions increase storage and query units; they do not guarantee better retrieval.
Match the model's supported dimensions and distance metric to the index. Vectorize supports up to 1,536 dimensions. Do not truncate an arbitrary embedding to fit; use a model-supported reduction and evaluate the resulting quality.
Chapter 17 covers model selection. Keep the model version and preprocessing configuration alongside the index so query embeddings can be reproduced consistently.
Different embedding models produce incompatible vector spaces. A 0.9 similarity score means nothing when comparing vectors from different models. Choose once and stay consistent, or reindex everything when you switch.
Index configuration
Dimensions and distance metric are immutable; changing either requires a new index and full reindex.
Cosine similarity is correct for most embedding models because it measures directional alignment independent of vector magnitude. Use Euclidean distance only if your embedding model documentation recommends it. Dot product suits models trained with that metric, which is rare.
Multi-tenancy with namespaces
Namespaces restrict the set of vectors searched within an index. They are a query boundary, not an independent authorisation system. Derive the namespace from the authenticated tenant and prevent callers from selecting another tenant’s namespace.
// Ingestion: accepted mutations become searchable asynchronously
await env.VECTOR_INDEX.upsert(vectors.map(vector => ({
...vector, namespace: `tenant-${tenantId}`
})));
// Later, at query time, use the authenticated tenant’s namespace
const results = await env.VECTOR_INDEX.query(embedding, {
namespace: `tenant-${tenantId}`,
topK: 5
});
This provides data separation without maintaining separate indexes per tenant. The tradeoff: all tenants share the twenty-million-vector limit. If individual tenants have large corpora, you'll need dedicated indexes per tenant anyway.
Metadata design
Metadata enables filtering and provides context for retrieved results. Design the schema carefully; you'll query against it and use it to reconstruct context.
Store references needed to recover the source and explain a citation: document ID, version, title and location. Keep filtering fields such as tenant, language and product version with the vector. Small chunk text can fit in metadata; larger content can live in R2 or D1 and be fetched after retrieval.
Metadata filtering narrows search scope before similarity ranking. Use it when criteria are known at query time: searching only product documentation, excluding old content, restricting to a specific language. Post-filtering offers more flexibility but wastes retrieval capacity on results you'll discard. Filter early when you can; filter late when you must.
Vectorize limits
| Limit | Value | Architectural implication |
|---|---|---|
| Maximum dimensions | 1536 | Constrains embedding model choice |
| Vectors per index | 20,000,000 paid | Roughly 4M documents at 5 chunks each |
| Metadata per vector | 10 KB | Sufficient for chunk text plus metadata |
| Indexes per account | 100 | Enables multi-index architectures for scale |
Approaching the index limit requires a capacity decision. Multiple indexes add routing and result-merging work and remain subject to account limits. Larger chunks reduce vector count but can dilute retrieval relevance. Compare those trade-offs with another engine using the same corpus and query load.
AI Search: managed RAG
AI Search handles chunking, embedding, indexing, and retrieval automatically. Point it at data sources and query directly. AI Search is the correct choice until it isn't. If standard chunking works for your data, you've saved months of engineering.
Choosing between AI Search and custom RAG
| Requirement | AI Search | Custom RAG |
|---|---|---|
| Time to prototype | Hours | Weeks |
| Chunking control | Configurable chunk size and overlap | Custom algorithms and routing |
| Embedding model choice | Supported models chosen at instance creation | Any compatible model |
| Hybrid search | Built-in vector and keyword search | Combine Vectorize with D1 FTS5 or another keyword index |
| Relevance boosting | Built-in (metadata-based) | Build yourself |
| Corpus size | Up to 1M files per instance (paid; 500K with hybrid) | Limited by Vectorize (20M vectors) |
| Instances per account | 100 free / 5,000 paid | N/A |
| Index management | Automatic | Programmatic |
| Supported formats | PDF, TXT, MD, HTML, CSV, JSON | Anything you can parse |
| Ongoing maintenance | Minimal | Significant |
AI Search instances also expose an MCP endpoint, which means the agents covered in Chapter 19 can query your knowledge base as a tool without any custom integration code. Point an MCP client at the instance's MCP URL and the agent gains search capability over your indexed content immediately. This is particularly valuable during prototyping: you can stand up a knowledge base, enable the MCP endpoint, and have an agent searching your documentation within hours rather than building a custom retrieval pipeline.
Choose AI Search when its supported sources, indexing settings, filters and retrieval controls meet the requirement. Build custom RAG when a measured failure requires a processing or retrieval step the managed service cannot express. Hybrid search or configurable chunk size alone is not a reason to rebuild.
Choose the retrieval and generation boundaries separately. If AI Search retrieves the right evidence but its generation interface cannot express a required model control or validation step, use it for retrieval and call the model yourself. This preserves managed ingestion and indexing while giving the application control over context assembly and generation. You take responsibility for citations, token budgets and generation failures at that boundary. Check the managed interface first: an external model alone does not require this split, because AI Search also supports external providers through AI Gateway.
The decision often becomes clear during prototyping. If AI Search disappoints and you can identify why (chunks splitting wrong, important keywords missing from semantic matches, or filtering requirements exceeding what path rules and five custom metadata fields can express), custom RAG addresses those specific problems. If AI Search works acceptably, custom RAG's engineering effort rarely justifies marginal improvements. Measure before rebuilding.
AI Search constraints
AI Search limits shape what you can build. Free plans allow 100 instances and 100,000 files per instance with 20,000 monthly queries. Paid plans raise these to 5,000 instances, 1 million files per instance (500,000 if hybrid search is enabled), and unlimited queries. The 4 MB maximum file size handles most documents but excludes large media files or data dumps. Instances include built-in storage and a built-in vector index, so you don't provision an R2 bucket and a Vectorize index separately and wire them together; uploading a file via uploadAndPoll() triggers immediate indexing on the instance itself.
One operational lever worth knowing: AI Search caches responses to semantically similar queries to cut latency and repeat-inference cost. That cache has a configurable freshness, defaulting to 48 hours and adjustable from 10 minutes to 6 days, with an on-demand purge. The trade-off is the familiar one: a longer window saves more inference but serves staler answers after your source content changes, so shorten it (or purge) when freshness matters more than cost, and lengthen it for stable corpora where the same questions recur.
Path filtering provides control over what gets indexed from website and R2 data sources. Include and exclude rules using glob patterns let you index documentation while skipping drafts, exclude admin pages from results, or limit indexing to specific language directories. This filtering improves result relevance by keeping irrelevant content out of the index, and enables splitting a single data source across multiple AI Search instances for specialised search experiences.
Custom metadata filtering adds query-time control beyond path-based indexing rules. You can define up to five metadata fields per instance (text, number, or boolean), attach metadata to documents via R2 custom headers or HTML meta tags, and filter search results by those fields at query time. A documentation site might filter by product version and category; a support knowledge base might filter by language and whether content is public-facing. This brings AI Search closer to the filtering capabilities you would otherwise need custom RAG to achieve, and it composes with path filtering rather than replacing it: path filtering controls what enters the index, while metadata filtering controls what returns from queries.
Relevance boosting layers business logic on top of retrieval by nudging rankings based on document metadata: prioritise recent documents, weight content from a specific product line higher, or surface official sources above community contributions. Boosting runs after retrieval but before final ranking, so it changes which results rise to the top without changing what gets retrieved. Pair it with hybrid search when you need both lexical-and-semantic recall and editorial control over ranking.
For multi-tenant SaaS, the ai_search_namespaces Workers binding exposes create(), delete(), list(), and search() at the namespace level. You can spin up a dedicated AI Search instance per customer at runtime without redeploying, which makes the 5,000-instances-per-account ceiling on the paid plan a meaningful number rather than an architectural cap: at that scale, namespace-per-tenant is an entirely reasonable shape for B2B SaaS knowledge bases. Approaching 1 million files per instance, evaluate whether content can split across instances using path filters or whether custom RAG offers more headroom.
Custom RAG architecture
When AI Search doesn't fit, you build custom pipelines. The complexity is real but manageable, and the control enables optimisations impossible with managed services.
The indexing pipeline
A custom indexing pipeline coordinates document processing, chunking, embedding, and storage. The orchestration is straightforward; the decisions hide inside the helper functions. Format-specific text extraction, chunking strategy selection based on document type, and metadata schema design determine quality. The pipeline structure itself is less important.
Chunking strategy should vary by content type. Technical documentation benefits from semantic chunking on section boundaries. Support tickets need smaller chunks with less overlap. A mixed corpus requires classification and routing. One-size-fits-all chunking produces mediocre results across everything.
The retrieval pipeline
Retrieval converts queries to vectors, searches the index, and assembles context for generation. The critical decision isn't the search itself (that's a single API call), but what to do when retrieval confidence is low.
// This index uses cosine similarity: higher scores are closer matches
const results = await env.VECTOR_INDEX.query(queryEmbedding, { topK: 5 });
if (!results.matches.length || results.matches[0].score < minimumCosineSimilarity) {
return "I couldn't find relevant information to answer this question.";
}
This comparison assumes a cosine index; use the score direction defined by the configured metric. A retrieval score is not a probability that an answer is correct. Calibrate thresholds against known relevant and irrelevant query-document pairs. Similarity scores depend on the model and distance metric; 0.8 is not a universal conservative starting point. Track false positives and false negatives separately, and consider the cost of an unsupported answer versus an unnecessary refusal.
Hybrid search
Vector search finds semantically similar content; keyword search finds exact term matches. Production RAG often needs both.
Vector search excels at semantic similarity. A query about "resetting credentials" matches content about "password recovery" even without those exact words. But vector search can miss exact matches that matter. A query for "ERR_SSL_PROTOCOL_ERROR" might retrieve generic SSL troubleshooting rather than the specific documentation for that error code.
If you are using AI Search, hybrid search is a flag rather than a build: enable it on the instance and queries run vector and BM25 in parallel, fuse the results via reciprocal rank fusion (or max fusion), and optionally rerank. This is the simplest path to lexical-plus-semantic retrieval and removes most of the reasons to drop down to custom RAG. The custom-RAG hybrid pattern below remains relevant when you need control over the keyword index, want to combine vector results with non-D1 sources, or have already invested in a custom pipeline.
Keyword search through D1's FTS5 extension finds exact term matches, complementing vector search for precise lookups.
async function hybridSearch(query: string, env: Env) {
const [vectorResults, keywordResults] = await Promise.all([
vectorSearch(query, env),
keywordSearch(query, env) // D1 FTS5 query
]);
return mergeResults(vectorResults, keywordResults);
}
The search and merge functions above are application helpers; the sketch shows their orchestration, not an implementation of FTS5 or result fusion. Merging logic determines hybrid search quality. Simple interleaving alternates results: first vector, first keyword, second vector, and so on. This works when both sources produce comparable quality but can elevate mediocre keyword matches above excellent vector matches.
Rank-based fusion avoids pretending that vector similarity and keyword scores share a scale. Reciprocal rank fusion is one option; tune it against the retrieval evaluation set and deduplicate by document or chunk identity.
If combining raw scores, account for each score’s direction and distribution. SQLite FTS5 ranks better matches with numerically lower BM25 values, so a generic “keep the higher score” rule can reverse the keyword ranking. Preserve source ranks for diagnosis.
RAG failure modes
RAG systems fail in characteristic ways. Understanding these helps you design monitoring and graceful degradation.
Retrieval failures
The most common failure: relevant content exists but isn't retrieved. Symptoms include answers that miss obvious information or generic responses when specific answers exist in the corpus.
Causes vary. Poor chunking splits relevant content across chunks, diluting similarity scores. Embedding model mismatch means your model encodes semantics differently than your queries express them; a model trained on formal text may not embed casual queries effectively. Insufficient indexing means relevant content was never processed, through ingestion failures or gaps in source coverage.
Decide whether retrieval may continue when part of its processing fails. AI Search's return_on_failure option permits partial results by default. For exploratory document search, showing the available evidence may be useful. For an answer that depends on finding an exact error code or a policy exception, missing part of retrieval can change the conclusion. Configure failure behaviour to match that consequence, and test the degraded path alongside normal retrieval quality. The application must distinguish “no supporting evidence found” from “the required search could not be completed”; neither justifies inventing an answer.
Context overflow
Models have finite context windows. Retrieve too many chunks and you exceed the limit, causing truncation or errors. More subtly, too many chunks dilute signal with noise as the model attends to irrelevant content.
Symptoms include answers citing irrelevant sources, responses missing the most relevant information despite it being retrieved, or explicit truncation errors.
Solutions: retrieve fewer chunks through better precision, summarise chunks before injection to preserve signal while reducing tokens, or use models with larger context windows at higher cost. Retrieval precision improvements compound; better retrieval means fewer chunks needed, lower inference costs, and better answer quality simultaneously.
Generation failures
Even with good retrieval, generation can fail. The model might ignore context and hallucinate anyway, particularly when context contradicts training data. It might synthesise incorrectly across sources, combining facts that don't belong together. It might quote accurately but miss the actual answer, fixating on related but non-responsive content.
Check whether each material claim is supported by the retrieved source, not merely whether the answer includes citations. Human sampling, labelled evaluation questions and user feedback catch different failures. A model can cite accurately while drawing an unsupported conclusion.
Staleness
RAG answers based on indexed content. If source documents update but indexes don't, answers become stale. Users trust RAG because it cites sources; stale citations betray that trust worse than honest uncertainty.
For slowly-changing corpora, scheduled batch reindexing suffices: nightly or weekly jobs rebuilding indexes from current sources. Frequently-changing content needs incremental indexing on document update, which adds operational complexity. Event-driven reindexing triggers on document changes and requires your storage to emit change events.
Answer caching requires careful invalidation. TTL-based invalidation accepts bounded staleness; answers may be up to N minutes old. Version-keyed caching includes corpus version in the cache key, invalidating everything when the index updates. Content-hash caching includes hashes of retrieved chunks, invalidating only when specific sources change. Choose based on staleness tolerance and cache hit rate requirements.
Monitoring for failures
Log retrieval results alongside user feedback: query, retrieved chunk IDs, similarity scores, generated answer, and explicit feedback (thumbs up/down, corrections, follow-up questions indicating confusion).
Track top-k similarity scores alongside labelled retrieval outcomes. A falling distribution can indicate changed queries, corpus coverage or model behaviour; it does not diagnose the cause. Check ingestion and embedding compatibility before attributing a decline to corpus-query drift.
Sample real queries for human evaluation, including rare query types and known failure cases. Automated metrics and human review catch different errors. Choose coverage deliberately; a small random sample can miss a low-volume but consequential failure.
Correlate retrieval confidence with user satisfaction where feedback exists. If high-confidence retrievals receive positive feedback and low-confidence retrievals receive negative feedback, your thresholds are well-calibrated. Weak correlation means thresholds need adjustment or retrieval quality needs improvement.
Optimising RAG performance
Measure query embedding, retrieval, content fetching, reranking and generation separately where present. Workers bindings simplify calls; they do not guarantee that embedding and vector search together finish within 50 ms. Optimise the measured critical path.
Model selection directly affects latency and cost. Llama 3.1 8B generates faster and cheaper than 70B. For many RAG applications, the smaller model suffices because retrieved context provides the specificity that larger models achieve through extensive training. The smaller model reads your documentation and answers from it; it doesn't need to know everything, just to read well. Benchmark both before assuming you need the larger model.
Streaming responses improve perceived latency dramatically. Users see content appearing immediately rather than waiting for complete generation. Total time doesn't change, but experience does; a 2-second streaming response feels faster than a 1.5-second blocking response.
Answer caching suits RAG with repeated queries. Documentation search often sees the same questions; "how do I reset my password" appears constantly. Caching eliminates generation latency for cache hits. The tradeoff is invalidation complexity. For stable corpora, aggressive caching makes sense. For rapidly-changing content, shorter TTLs or version-keyed caching prevent stale answers.
When RAG is wrong
RAG finds relevant text in unstructured content. If your answer isn't in unstructured text, RAG is the wrong tool.
For a small corpus, compare direct context with retrieval. If all relevant material fits alongside the prompt and output budget, direct context can be simpler. It still sends those tokens on every request and may make relevance harder for the model; page count alone is not a reliable token budget.
Rapidly changing data strains RAG architectures. If source documents update every minute, indexing lag means perpetually stale answers. Real-time data needs different approaches, such as direct database queries, live API calls, or tool use, rather than pre-computed vector indexes.
Queries not benefiting from semantic search waste RAG's strengths. Exact lookups for order status or account balance need database queries, not vector similarity. Computations need code execution, not text retrieval. Structured data queries need SQL, not embeddings.
High-precision requirements may exceed RAG's capabilities. Legal research, medical diagnosis, and other domains where wrong answers cause serious harm need retrieval quality RAG may not achieve. Vector similarity scores don't map to accuracy guarantees; 0.85 similarity doesn't mean 85% confidence the answer is correct. When precision matters more than coverage, evaluate carefully and consider human review for high-stakes queries.
What comes next
Retrieval can supply evidence for an answer or for an action. Chapter 19 covers agents that select tools and maintain progress, with particular attention to who authorises each action and how to recover when it is only partly completed.