RAG architectures for production
Retrieval-Augmented Generation (RAG) is easy to demo and annoyingly hard to ship. The difference between a prototype and a production system isn’t “better prompts”—it’s architecture: data pipelines, retrieval quality, latency budgets, observability, and guardrails. This post lays out proven RAG patterns that survive real traffic, messy corpora, and changing models.
Start with the contract: what “good” looks like
Before components, define the system contract:
- Answer quality: factuality, completeness, citation correctness, refusal behavior.
- Latency: p95 end-to-end target (e.g., 1.5–3s for chat UX, tighter for API).
- Cost: tokens + embedding + vector DB + reranker + tool calls.
- Freshness: time to index updates (minutes vs hours vs days).
- Security: tenant isolation, permission-aware retrieval, audit logs.
If you can’t write these as measurable SLOs, your RAG system will drift into “vibes-based engineering.”
The canonical production pipeline (and why it’s not enough)
A baseline RAG flow:
- Ingest & chunk documents
- Embed chunks and store in a vector index
- Retrieve top-k candidates for a query
- Generate an answer conditioned on retrieved context
In production, you almost always need three additions:
- Hybrid retrieval (lexical + vector)
- Reranking (cross-encoder or LLM-based)
- Grounding & citations (traceability back to sources)
These are the difference between “it found something relevant” and “it found the right thing reliably.”
Indexing: chunking is a product decision
Chunking strategy determines recall, citation quality, and hallucination rate.
Practical guidance:
- Prefer semantic chunks over fixed tokens. Split by headings/sections, then cap size (e.g., 300–800 tokens) with overlap (50–150 tokens).
- Store structured metadata: doc_id, section path, timestamps, authors, product version, tenant_id, ACLs.
- Use multiple representations when needed:
- A “small chunk” index for pinpoint retrieval
- A “parent” index (or parent pointers) to expand context at generation time
A robust pattern is parent-child retrieval: retrieve small chunks, then expand to the parent section for coherent context. This reduces the “fragment soup” that makes answers brittle.
Retrieval: hybrid or bust
Pure vector search is often mediocre on:
- exact IDs (ticket numbers, transaction hashes)
- code symbols, API names
- rare terms, new product names
Production RAG typically uses hybrid retrieval:
- BM25 / lexical for exact matches and sparse signals
- Dense vectors for semantic similarity
- Optional: metadata filters (tenant, product, locale, time)
Then fuse results via simple weighted scoring or reciprocal rank fusion. The opinionated take: if you’re not doing hybrid, you’re paying an LLM tax to compensate for retrieval misses.
Reranking: where quality actually comes from
Top-k retrieval is noisy. Reranking is the cheapest way to buy accuracy.
Common tiers:
- Fast reranker: cross-encoder model (e.g., MiniLM class) for top 50 → top 8
- LLM reranker: ask a small LLM to choose best passages (more expensive, sometimes better for long queries)
In practice:
- Retrieve k=50–200 candidates
- Rerank down to k=5–12
- Feed k=3–8 into generation (depending on token budget)
Also consider diversity constraints (avoid returning five near-duplicates from the same doc section). Diversity improves coverage and reduces overfitting to one misleading chunk.
Generation: enforce grounding, don’t hope for it
Prompting matters, but production requires mechanical sympathy:
- Citations as a hard requirement: each claim should reference a chunk ID or URL.
- Context window budgeting: reserve tokens for the answer; don’t drown the model in retrieved text.
- Answer style policies: concise by default, expandable on request.
- Refusal when context is insufficient: “I don’t have enough info in the provided sources.”
A strong pattern is two-pass answering:
- Draft answer with citations
- Verify: check that each sentence is supported by at least one cited chunk (LLM-as-judge or heuristic checks)
This costs more, but it’s how you ship RAG into regulated or reputation-sensitive environments.
Data freshness and reindexing: treat docs like code
Documents change. Your index must keep up.
- Incremental ingestion: detect deltas, re-embed only changed chunks.
- Versioned documents: store doc_version and effective dates.
- Backfills: when you change embedding models or chunking, you need a safe reindex path.
A useful discipline: index migrations like database migrations. Run old and new indexes side-by-side, compare retrieval metrics, then cut over.
Multi-tenant and permission-aware RAG
If you’re building for enterprise or even just multiple teams, access control is non-negotiable.
- Enforce ACLs at retrieval time, not post-generation.
- Store ACL metadata per chunk (tenant_id, groups, visibility).
- Prefer filtered retrieval (vector DB filter) rather than retrieving then filtering—otherwise you leak relevance signals and waste latency.
In blockchain/game studios, you’ll also want environment separation (devnet/testnet/mainnet docs) so the model doesn’t mix instructions and operational runbooks.
Evaluation: move from anecdote to metrics
You need continuous evaluation because models, data, and user behavior drift.
What to measure:
- Retrieval metrics: recall@k, MRR, nDCG (using labeled query→doc sets)
- Answer metrics: faithfulness/groundedness, citation accuracy, helpfulness
- Operational metrics: p95 latency by stage, cost per request, cache hit rate
How to build the eval set:
- Start with 50–200 real queries from logs
- Label expected sources (doc IDs) and acceptable answers
- Re-run nightly on new index/model versions
This is where most teams underinvest. Without evals, your “improvement” work is indistinguishable from random changes.
Observability and guardrails in the real world
Production RAG needs tracing, not just logs.
Instrument:
- retrieval query, filters, top-k scores
- reranker scores and selected passages
- final prompt size, completion tokens, latency per stage
- citations emitted
Guardrails that actually help:
- PII/secret filters on retrieved chunks (don’t surface keys in context)
- Prompt injection detection: treat retrieved text as untrusted input; strip/flag instructions like “ignore previous directions”
- Rate limiting and caching: cache embeddings for repeated queries; cache retrieval for common intents
A pragmatic stance: you won’t prevent all jailbreaks, but you can make them observable, less frequent, and less damaging.
Reference architecture (opinionated)
A production-ready setup looks like:
- Ingestion service → chunker → embeddings pipeline
- Vector DB (dense) + search engine (BM25)
- Retrieval service with filters + fusion
- Reranker service
- Generation service with citation formatting + verification
- Evaluation harness + nightly regression
- Tracing/metrics across all stages
Don’t cram everything into a single “RAG function.” Separate services so you can tune, cache, and swap components independently.
Conclusion
Production RAG is systems engineering: hybrid retrieval to avoid misses, reranking to stabilize relevance, grounding with citations to control hallucinations, and eval/observability to prevent silent regressions. If you treat RAG as “prompt engineering plus a vector DB,” you’ll ship something that demos well and fails quietly in the wild. Treat it like a search-and-generation product with SLOs, and it becomes a reliable layer you can build real applications on.