RAG architectures for production
Retrieval-Augmented Generation (RAG) is easy to demo and notoriously easy to ship poorly. In production, “just add a vector DB” turns into latency spikes, hallucinated answers with confident citations, runaway token bills, and brittle pipelines that silently stop indexing. The difference between a prototype and a reliable system is architecture: clear stages, measurable quality, and explicit failure modes.
Below is a production-minded RAG blueprint—what to build, where teams usually cut corners, and how to keep it stable as data and traffic grow.
Start with the contract: what RAG must guarantee
Before choosing embeddings or databases, define a contract:
- Groundedness: answers must be attributable to retrieved sources.
- Coverage: the system should find relevant material when it exists.
- Freshness: updates appear within a defined SLA (e.g., 15 minutes).
- Latency: p95 end-to-end response time (e.g., < 2s).
- Cost: $/request and monthly budget constraints.
If you can’t state these, you can’t evaluate tradeoffs like reranking vs. larger context windows.
Ingestion pipeline: treat it like a data product
Most RAG failures start upstream. Production ingestion is not “load PDFs once.” It’s a continuous pipeline with observability.
Key practices:
- Canonical document model: normalize sources (docs, wikis, tickets, code, chats) into a single schema:
doc_id,source,uri,version,timestamp,access_policy, and content blocks. - Deterministic chunking: chunk boundaries should be stable across runs. Use semantic chunking (headings/sections) when available, then cap by tokens. Avoid arbitrary fixed sizes that split tables or code blocks.
- Metadata richness: store tenant/team, product area, language, doc type, and “authoritative level” (spec vs. discussion). This enables filtering and improves ranking.
- Incremental updates: compute diffs; re-embed only changed chunks. Keep old versions if you need auditability.
- Quality gates: reject empty, duplicate, or low-signal chunks; run PII detection/redaction if required.
Opinionated take: if you don’t have an ingestion queue (Kafka/SQS/PubSub) and a dead-letter path, you don’t have production ingestion—you have a script.
Indexing strategy: one size won’t fit your queries
A robust retrieval stack usually combines multiple indexes:
- Dense vectors (semantic search): embeddings in a vector store.
- Sparse / lexical (BM25): great for exact terms, error codes, identifiers.
- Metadata filters: tenant, ACL, product, time.
In production, hybrid retrieval is a default, not a luxury. Dense-only systems miss exact matches; BM25-only systems miss paraphrases. Use hybrid scoring or retrieve from both and merge.
Also decide where to compute embeddings:
- Offline embeddings for your corpus (batch + incremental).
- Online query embeddings per request (cached when possible).
Track embedding model versions and reindex strategy. A model upgrade without a migration plan is how you end up with inconsistent similarity behavior.
Retrieval pipeline: from “top-k” to “right-k”
Production retrieval is a staged funnel:
- Query understanding: lightweight rewrite/expansion (“error 502” → include “bad gateway”), detect language, and classify intent (FAQ vs. troubleshooting vs. policy).
- Candidate retrieval: hybrid search to get, say, 50–200 candidates.
- Reranking: cross-encoder reranker or LLM-based ranker to select the best 5–15. This is often the biggest quality jump per dollar.
- Context assembly: deduplicate, order by relevance, and enforce diversity (avoid 10 near-identical chunks). Apply token budgeting.
A common anti-pattern: retrieving 3 chunks and hoping the model “fills the gaps.” You want the opposite: retrieve broadly, then aggressively rank down.
Generation: constrain the model, don’t “prompt harder”
Generation should be the final step, not the place you compensate for weak retrieval.
Production patterns:
- Citations as first-class output: require the model to cite chunk IDs/URLs. If it can’t, it should say so.
- Answer templates: for common intents (policy, troubleshooting), use structured outputs (JSON) or consistent sections (“Cause / Fix / References”).
- Refusal and escalation paths: when confidence is low, return “I couldn’t find that in the knowledge base” plus suggested searches or ask a clarifying question.
- Context compression: for large contexts, summarize per document then answer from summaries + key excerpts. This reduces token costs while keeping traceability.
If you need the model to be “creative,” keep that creativity out of factual answers. Grounded systems should be boring in the best way.
Caching: the unsung hero of cost and latency
RAG can be expensive because it repeats work. Add caching at multiple layers:
- Query embedding cache: normalized query → embedding.
- Retrieval cache: query + filters → top results (short TTL).
- Rerank cache: (query, candidate IDs) → ranked list.
- Response cache: for high-repeat queries, store final answers with versioned source references.
Be careful: caching must respect ACLs and tenant boundaries, and should be invalidated on index updates if freshness matters.
Security and access control: ACLs must flow end-to-end
In production, your RAG system is a data exfiltration machine if you get ACLs wrong.
Minimum bar:
- Filter at retrieval time using access metadata (don’t retrieve then filter—leaks can still occur via embeddings or logs).
- Tenant-isolated indexes or strong per-tenant filtering, depending on scale.
- Prompt/data logging policy: redact secrets, avoid storing raw prompts by default, and restrict who can replay traces.
For regulated environments, consider per-tenant encryption keys and audit logs that record which sources were used in each answer.
Evaluation: ship metrics, not vibes
You can’t manage RAG quality without automated evaluation.
Practical evaluation stack:
- Golden set: a curated set of real questions with expected sources (not just expected answers).
- Retrieval metrics: recall@k, MRR, and “source coverage” (did we retrieve the right doc?).
- Generation metrics: groundedness (citation correctness), refusal rate, and answer usefulness (human or LLM-judge with calibration).
- Online monitoring: p50/p95 latency per stage, token usage, cache hit rates, and “no-results” rates.
Run evaluations per release: embedding model changes, chunking logic, reranker updates, and prompt changes should all trigger regression tests.
Deployment and ops: design for failure
A production RAG service should degrade gracefully:
- If the vector DB is slow, fall back to BM25.
- If reranking times out, use initial hybrid scores.
- If retrieval returns nothing, ask clarifying questions or return search links.
Implement tracing across stages (ingest → index → retrieve → generate) with request IDs. The most valuable debugging artifact is a single record showing: query, filters, retrieved chunk IDs, rerank scores, final context, and model output.
A reference architecture (put it together)
A typical production layout:
- Ingestion service → queue → chunk/clean workers → embedding workers
- Indexes: vector store + BM25 store + metadata store
- Online API: query understanding → hybrid retrieval → reranker → context builder → LLM generation
- Shared concerns: ACL enforcement, caching, observability, eval harness, and model/version registry
This isn’t overengineering; it’s what prevents “RAG drift” as your corpus and traffic evolve.
Conclusion
Production RAG is less about the model and more about the system around it: disciplined ingestion, hybrid retrieval, reranking, constrained generation with citations, aggressive caching, real ACLs, and continuous evaluation. If you build those pieces explicitly—and instrument every stage—you get a system that scales in both quality and operations. If you don’t, you’ll ship a demo that slowly turns into an expensive, untrustworthy search box with a personality.