Back to blogAI & Software

RAG Pipelines for Production: Lessons from 30+ Engagements

Chunking, retrieval quality, citation UX, and cost control, what actually matters when RAG leaves the notebook and serves paying users.

Share this article

Production RAG pipelines: PAI Technologies blog

Teams spend weeks on system prompts while retrieval sends irrelevant chunks. Fixing retrieval, and proving it with evals, usually moves accuracy more than another prompt tweak. After 30+ RAG engagements and our own PlaywithDB product work, we see the same failure modes repeat.

RAG is an integration problem: documents, embeddings, search, reranking, LLM, UI citations, logging, and re-embedding when upstream content changes. Natural language is only as good as the data layer underneath.

This article covers chunking, hybrid search, citation UX, and operations, the parts that survive legal review and finance scrutiny when RAG serves real users.

RAG is an integration problem, not a prompt problem

Notebook demos hide broken retrieval behind charismatic answers on familiar questions. Production exposes long-tail queries, stale indexes, and tables split mid-row.

We instrument retrieval separately from generation: hit rate, MRR on golden sets, and “correct chunk in top-k” before the LLM runs. PlaywithDB-style systems add schema awareness and governed SQL: RAG over unstructured docs is one pattern, not the only one.

When a client asks for “ChatGPT on our data,” we translate that into retrieval architecture, access control, and eval criteria, not a single prompt.

Chunking and metadata that survive real documents

Structure-aware document chunking for production RAG
Headings, tables, and version metadata beat fixed 512-token splits.

Fixed 512-token chunks break tables, policies, and API reference docs. We prefer structure-aware splitting: headings, rows, or logical sections, with metadata (source, page, product area, version) stored alongside embeddings.

Store source URI and document version so answers can cite “Policy v3.2, section 4.” Re-embed when upstream docs change; version your index and log which index version answered each query.

For codebases, chunk by file and symbol boundaries. For support KBs, chunk by article. For contracts, preserve clause boundaries, legal teams notice when citations span unrelated clauses.

  • Reject chunks below minimum information density
  • Attach ACL metadata at index time, not query time only
  • Track embedding model version in the index manifest

Hybrid search and reranking

Dense vectors miss exact SKUs, error codes, regulation numbers, and internal IDs. We combine keyword search (BM25 or Postgres full-text) with vector search, then rerank top-k before the LLM sees context.

Rerankers can be cross-encoder models or lighter learned scores, the right choice depends on latency SLO and corpus size. Measure reranker lift on your golden set; do not assume it always helps small corpora.

PlaywithDB deployments often pair structured queries with semantic retrieval over documentation: hybrid patterns apply across SQL and text, not only PDFs.

Citation UX and trust

Citation UI for production RAG applications
Inline sources reduce black-box pushback from compliance and legal teams.

Users trust answers with inline citations, even a sidebar listing sources with links reduces “black box” objections. Compliance teams want traceability from answer to source paragraph.

Show confidence honestly: when retrieval score is weak, say “I could not find a definitive source” instead of hallucinating fluently. Refusal paths are a feature.

Log cited chunk IDs with each response for audit. EU clients routinely ask how long those logs are retained: answer in the architecture doc, not in a ticket during launch week.

Evaluation harnesses that block bad releases

Golden sets need not be huge: 50–200 representative questions with expected sources or answer properties often suffice for regression gates. Add adversarial cases (typos, out-of-scope, PII probes).

Nightly eval in CI: retrieval metrics + end-to-end answer checks. Block deploy when metrics drop. Human review samples weekly for drift the automated set missed.

Separate “model upgrade” evals from “index refresh” evals, conflating them makes root cause analysis slow when accuracy moves.

Cost, latency, and operations

Cache embeddings for static corpora. Use a smaller model for routing and retrieval grading; reserve the large model for final answers. Stream tokens so perceived latency drops.

Monitor p95 latency, cost per query, retrieval hit rate, and thumbs-down rate. Alert when any drift beyond SLO. Gordon-scale EU deployments add region and subprocessors to the ops dashboard, not only latency.

We ship RAG in three-to-four-week slices with handover: index pipeline code, eval scripts, and runbooks. If you are planning production RAG, start with retrieval quality and citations, prompts come after.

Share this article