RAG architectures for production
Retrieval-Augmented Generation (RAG) looks deceptively simple: embed docs, retrieve top-k, feed them to an LLM. In production, that “hello world” collapses under latency budgets, messy data, multi-tenancy, and the harsh reality that users ask ambiguous questions. Shipping RAG is less about clever prompts and more about building a dependable information system with an LLM at the edge.
Below are the architecture patterns we’ve seen hold up in real deployments—where accuracy, cost, and reliability matter.
The production RAG mental model
Think in three planes:
- Knowledge plane: ingestion, chunking, embedding, indexing, versioning.
- Retrieval plane: query understanding, hybrid search, reranking, filtering.
- Generation plane: context assembly, prompting, citation, guardrails.
Production quality comes from tight coupling between these planes via evaluation and observability. If you can’t measure retrieval quality independently from generation, you will waste weeks “prompt tuning” around a retrieval problem.
Ingestion and indexing: your future incident queue
Most RAG failures start upstream.
Chunking strategy
- Chunk by semantic boundaries (headings, paragraphs, code blocks), not fixed token counts. Use token caps as a guardrail, not the primary rule.
- Preserve structure metadata: document title, section path, URL, author, timestamp, product/version, access level.
- For code and API docs, consider dual representations: a “natural language” chunk and a “signature/identifier” chunk for exact matching.
Embeddings and index design
- Choose embeddings based on your domain. General-purpose models are fine until you have jargon, symbols, or code-heavy content.
- Version everything:
embedding_model,chunker_version,source_snapshot. Without this, you can’t reproduce results.
Incremental updates
- Don’t rebuild the world daily. Implement incremental ingestion with idempotent doc IDs and tombstones for deletions.
- Maintain a background compaction job to clean old vectors after large churn.
Opinionated take: if your ingestion pipeline doesn’t have unit tests and deterministic outputs, your “AI” system will behave like a flaky ETL job—because it is one.
Retrieval: hybrid is the default, not the exception
Pure vector similarity fails on:
- short queries (“pricing”),
- exact terms (error codes, function names),
- new entities (recent product names).
Use hybrid retrieval
- Combine BM25/keyword + vector search.
- Fuse results with Reciprocal Rank Fusion (RRF) or weighted blending.
Metadata filtering is not optional
- Enforce tenant, permissions, locale, product/version, and time windows at retrieval time.
- Do not “retrieve broadly and filter later.” You will leak data.
Reranking is your accuracy lever
- A cross-encoder reranker over the top 50–200 candidates often yields the biggest quality jump per dollar.
- Rerank on the user query + chunk text, but include key metadata (section title, doc type) in the reranker input.
Query rewriting and routing
- Implement lightweight query rewriting for clarity (“What’s the limit?” → “What is the API rate limit for X?”) using rules or a small model.
- Route queries to different indexes (support tickets vs. docs vs. internal wikis). A single monolithic index is convenient—until it isn’t.
Context assembly: treat tokens as a scarce resource
Once you have candidates, you need to assemble context that the model can reliably use.
Deduplicate and diversify
- Remove near-duplicate chunks (common in scraped docs).
- Prefer coverage: 2–3 distinct sources beat 6 chunks from the same page.
Compression over truncation
- If context is large, summarize retrieved chunks into “fact statements” with citations, then feed the compressed pack to the final answer step.
- This two-step approach lowers token cost and improves coherence, but requires careful attribution.
Citations and provenance
- Attach source IDs to every chunk and require the model to cite them.
- In the UI, show clickable citations. This increases trust and makes debugging feasible.
Generation: constrain the model, don’t beg it
A production RAG prompt should be a contract:
- Use only provided sources.
- If sources conflict, say so.
- If insufficient evidence, ask a clarifying question or decline.
Practical pattern: Answer + Evidence + Next step
- Answer: direct response.
- Evidence: bullet list of cited snippets.
- Next step: what to check or ask if ambiguous.
For higher reliability, use tooling:
- A “cite()” tool the model must call with source IDs.
- A “clarify()” tool when retrieval confidence is low.
Evaluation: separate retrieval quality from model quality
If you only measure “did the final answer look good,” you’ll miss the actual failure mode.
Core offline metrics
- Retrieval: Recall@k, MRR, nDCG (against labeled relevant passages).
- Reranking lift: compare before/after rerank.
- Groundedness: percentage of answer statements supported by citations.
Online metrics
- Deflection rate (does it reduce tickets?).
- User edits/corrections.
- Citation click-through (proxy for trust).
Build a small, curated eval set early (50–200 queries) and grow it continuously. Treat it like a product asset.
Observability and operations: your real moat
Production RAG needs first-class telemetry:
- Query, rewritten query, retrieved doc IDs, scores, filters applied.
- Reranker scores and final selected context.
- Token counts, latency per stage, model/version identifiers.
- User feedback tied to traces.
Set budgets:
- p95 end-to-end latency (e.g., 2–5s consumer; 1–2s internal tooling).
- Max context tokens.
- Max rerank candidates.
Caching matters:
- Cache embeddings for repeated queries.
- Cache retrieval results for popular queries with short TTL.
- Cache final answers only when permissions and freshness allow.
Security and safety: the unglamorous requirements
RAG is a data access layer. Treat it accordingly.
- Enforce document-level ACLs at index and query time.
- Defend against prompt injection in retrieved text: strip or sandbox instructions from sources; tell the model to treat retrieved content as untrusted data.
- Keep a “no retrieval” fallback for sensitive prompts (credentials, private keys, internal-only topics).
Reference architecture (battle-tested)
A solid production layout:
- Ingestion service → parses sources, chunks, extracts metadata, computes embeddings.
- Vector store + keyword index → separate but queryable together.
- Retrieval API → applies ACL filters, hybrid search, reranking, returns top evidence.
- Answer service → assembles context, runs generation, enforces citation format.
- Eval/telemetry pipeline → logs traces, runs offline eval nightly, flags regressions.
This separation lets teams iterate on retrieval without redeploying generation—and vice versa.
Conclusion
RAG in production isn’t “LLM magic.” It’s disciplined information retrieval plus careful generation constraints, wrapped in observability and evaluation. Start with hybrid retrieval, add reranking early, version your knowledge pipeline, and instrument everything. If you do those unsexy things well, your model choice becomes a tuning knob—not a desperate hope.