RAG architectures for production

Retrieval-Augmented Generation (RAG) is no longer a novelty pattern—it’s the default way to make LLMs useful on real business data. But most “RAG demos” collapse in production for predictable reasons: brittle indexing, weak retrieval, no observability, and zero governance. Production RAG is an information system with an LLM attached, not the other way around.

Below is a concrete, opinionated blueprint for building RAG that survives real traffic, messy corpora, and compliance constraints.

What “production RAG” actually means

A production RAG system must reliably answer questions with:

  • High recall (it can find relevant sources)
  • High precision (it doesn’t drown the model in noise)
  • Stable latency and cost (P95 matters more than average)
  • Auditable grounding (you can show why an answer is correct)
  • Evolvable pipelines (content changes daily; models change monthly)

If you can’t measure those five, you don’t have production RAG—you have a prompt.

Canonical architecture: the “RAG spine”

Most robust systems converge on the same spine:

  1. Ingestion: connect to sources (docs, tickets, code, wikis, PDFs)
  2. Normalization: de-duplicate, clean, convert formats, extract metadata
  3. Chunking + enrichment: split content and attach structure (titles, headings, owners, ACLs)
  4. Indexing: build vector + lexical indexes (hybrid is the norm)
  5. Retrieval: candidate generation (top-K) using hybrid search + filters
  6. Reranking: cross-encoder / LLM-based rerank to improve precision
  7. Context assembly: pack evidence into a bounded context window
  8. Generation: produce answer with citations and constraints
  9. Post-processing: safety filters, formatting, redaction, caching
  10. Telemetry + evaluation: log, score, and continuously regress

The critical production insight: retrieval and ranking are your product. The LLM is often interchangeable.

Indexing: treat it like a data product

Production failures often start at ingestion.

Best practices that matter:

  • Metadata-first design: store source, timestamp, author, doc_type, tenant_id, and access control fields. Without this, you can’t filter, audit, or enforce permissions.
  • Deterministic chunk IDs: stable identifiers prevent duplicate chunks across re-indexes and enable cache hits.
  • Incremental indexing: reprocessing everything nightly is expensive and error-prone. Use change detection (hashes, ETags, revision IDs).
  • Chunking strategy by doc type: PDFs, code, and tickets chunk differently. Over-chunking increases recall but kills precision.
  • Store raw + parsed text: you’ll need to re-parse when extractors improve.

Opinionated rule: start with 500–1,000 tokens per chunk, 10–20% overlap, then tune based on retrieval metrics—not vibes.

Retrieval: hybrid by default, filters always

Pure vector search is rarely enough. In production, you need:

  • Hybrid retrieval: combine BM25/keyword with embeddings. Keywords catch exact identifiers (“ERC-4337”, function names, error codes). Embeddings catch paraphrases.
  • Structured filters: enforce tenant boundaries, ACLs, doc freshness, language, product version, chain/network, etc.
  • Query rewriting: transform conversational questions into retrieval-friendly queries (e.g., expand acronyms, add product names, extract entities).

Practical approach:

  • Candidate generation: topK=50–200 from hybrid
  • Apply filters early (before vector scoring if possible)
  • Then rerank down to topN=5–12 for the context window

Reranking: the highest ROI component

If you add one “advanced” component, make it reranking.

  • Cross-encoder rerankers dramatically improve precision because they score query–document pairs jointly.
  • For tight latency budgets, use a small reranker model; for high-stakes answers, allow an “accurate mode” with a stronger reranker.

Key operational tip: reranking is where you can implement business logic (“prefer official docs over community posts”, “prefer latest version”, “prefer internal runbooks”). Encode this in features and weighting, not in prompts.

Context assembly: evidence packing is an art

Even with good retrieval, you can lose grounding by stuffing the window poorly.

  • Diversity > redundancy: prefer 6 distinct sources over 6 near-identical chunks.
  • Section-aware packing: include titles, headings, and breadcrumb paths. This improves model comprehension and citation quality.
  • Quote-based snippets: extract the most relevant spans rather than dumping full chunks.
  • Hard constraints: cap context tokens and enforce a minimum evidence threshold. If you can’t retrieve evidence, say so.

A production pattern that works: a “context compiler” that outputs a structured bundle:

  • sources[] with IDs and URLs
  • snippets[] with exact quoted text and offsets
  • notes (metadata like version and timestamp)

Generation: make it boring, controlled, and cited

The generation prompt should be stable and testable.

  • Citations are non-negotiable: require the model to cite source IDs for each claim.
  • Grounding rules: “Answer only from provided context. If missing, ask clarifying questions or state uncertainty.”
  • Output schemas: use JSON schemas for downstream systems (support bots, dashboards, agents).

If you’re shipping to users, implement a “no evidence, no answer” mode for high-risk domains (finance, medical, security, compliance).

Evaluation: measure retrieval separately from generation

Teams fail by evaluating only the final answer. You need a layered eval stack:

  • Retrieval metrics: recall@K, MRR, nDCG, coverage by doc type, freshness coverage
  • Rerank quality: win-rate against baseline, precision@N
  • Grounding: citation correctness, quote overlap, hallucination rate
  • Task success: user resolution, deflection, time-to-answer

Create a golden set of 200–1,000 queries with labeled relevant documents. Update it monthly. This becomes your regression suite when you change chunking, embedding models, or prompts.

Operations: reliability, cost, and governance

Production RAG needs real ops discipline:

  • Caching: cache embeddings, retrieval results, and final answers for repeated queries (with TTL and invalidation by doc revision).
  • Fallbacks: if vector DB is degraded, fall back to lexical; if reranker fails, skip rerank with reduced quality but stable uptime.
  • Observability: log query, retrieved doc IDs, scores, rerank positions, context tokens, model version, latency breakdown, and user feedback.
  • Security: enforce ACL filtering at retrieval time; never rely on the LLM to “not reveal secrets.”
  • Multi-tenancy: isolate indexes or use strong tenant filters with audited tests.

One more opinionated point: avoid building a single monolithic “RAG service” that does everything. Split into indexing pipeline, retrieval API, and generation API so you can scale and deploy independently.

Conclusion: production RAG is search engineering with LLMs

RAG in production isn’t magic—it’s disciplined information retrieval, ranking, and governance. If you invest in metadata, hybrid retrieval, reranking, and measurable evaluation, you’ll get a system that stays accurate as your content grows and your models change. Treat the LLM as the final renderer of an evidence bundle, and your RAG app will stop behaving like a demo and start acting like infrastructure.