· 6 min read
RAG You Can Trust: 8 Lessons on Embeddings & Retrieval
Production RAG lessons: chunking with metadata, embedding size, filtered vector search, citations, incremental re-indexing and vector lifecycle.
By Sai Ram Sana
Most RAG tutorials end at "embed your docs, put them in a vector database, ask a question." That gets you a demo. It doesn't get you a system people trust with real decisions.
Throughout this post I'll use one example: an internal company knowledge base. Think product docs, help-center articles, HR and finance policies, and the public website, all searchable by an AI assistant. The lessons apply to almost any RAG system.
1. Retrieval quality is decided before the embedding call
The embedding model matters less than what you feed it.
Chunk with a recursive text splitter that prefers natural boundaries (headings, then paragraphs, then sentences) and uses a modest overlap, so an idea that crosses a boundary isn't lost. Attach metadata to every chunk the moment it's created:
type ChunkMetadata = {
sourceId: string; // the document or page this chunk came from
sourceType: "doc" | "policy" | "article" | "web";
collection: string; // e.g. "billing", "security", "onboarding"
heading?: string; // nearest section heading, great for citations
tags?: string[];
audience?: "internal" | "customer";
updatedAt: string;
url?: string;
};
That metadata is what makes the system feel smart later. Without it, every query searches everything, and every answer is a guess about which document it came from.
One small trick: prepend the document title and section heading to each chunk's text before embedding it. A chunk that just says "This applies after 30 days" is meaningless on its own. "Refund Policy › Enterprise Plans: This applies after 30 days" is findable.
2. Pick an embedding model, then size it on purpose
Some modern embedding models, such as OpenAI's text-embedding-3 family, let you shorten the vector to a fraction of its full size. You trade a little accuracy for big savings in storage, memory and query latency.
Don't guess. Build a small evaluation set of real questions with known correct documents, and measure recall at a few sizes. For many knowledge bases, a mid-size vector is the sweet spot.
The rule I follow: one embedding model and one dimension size per index, forever. Mixing them in one index silently wrecks similarity scores. If you change models, build a new index and backfill it.
3. Choose the vector store by access pattern
There's no single best vector database. There's the right one for your workload:
- Dedicated vector databases (Pinecone, Qdrant, Weaviate) shine at large, fast-growing, heavily filtered similarity search.
- Vector search inside your primary database (for example pgvector in Postgres, or the vector features of Elasticsearch/OpenSearch) keeps vectors next to the source records. You get one query, one backup strategy and fewer sync bugs.
Whichever you choose, plan the failure mode. A simple fallback, such as keyword search or an in-memory similarity pass over a small candidate set, means search degrades gracefully instead of returning an error.
4. Filter first, then search
This is the biggest quality jump in most RAG systems, and the most underrated.
Users rarely want "everything we've ever written." They want "what does the refund policy say for enterprise customers" or "the latest security docs for the mobile app." So make the query a metadata-filtered top-k search:
const results = await index.query({
vector: await embed(question),
topK: 20,
filter: {
collection: { $in: ["billing", "legal"] },
...(user.isEmployee ? {} : { audience: "customer" }),
// Many vector stores only range-filter numbers, so store dates as Unix seconds.
updatedAt: { $gte: Date.parse("2026-01-01") / 1000 },
},
includeMetadata: true,
});
Filtering by collection, tags, audience and date before similarity ranking gives you:
- fewer near-duplicates from unrelated documents
- fewer confidently wrong answers that cite outdated policies
- often lower latency, though that depends on your store
How filtering is applied matters. Some stores apply filters inside the nearest-neighbour search. Others retrieve the top-k first and filter afterwards, which can silently return fewer results than you asked for. Check how yours behaves, and raise topK or switch to pre-filtering if needed.
The audience filter is also a security boundary. Retrieval should never return a chunk the user isn't allowed to read. Enforce permissions in the query, not in the prompt.
5. No citation, no answer
If users can't check an answer, they won't trust the next one.
Every retrieved chunk keeps its sourceId, heading and URL. Instruct the model to answer only from the provided context and to cite the sources it used. Then show those sources as clickable links to the exact document and section.
Two practical rules:
- Return sources even when the answer is "I couldn't find that." An honest miss with visible sources beats a fluent hallucination.
- Cite by name. "From: Refund Policy › Enterprise Plans" is far more convincing than "Source 4".
6. Keep indexes fresh without re-embedding everything
Knowledge bases go stale fast. Re-embedding every document every night is slow, expensive and risky.
Refresh incrementally:
- Diff before you embed. Keep a content hash or last-modified timestamp per document, and only re-process what actually changed.
- Swap, don't gap. Write a document's new chunks first, then remove the old ones, so the document never disappears from search mid-update.
- Batch the embedding pass. Embed all changed chunks in one batched run rather than one call per document.
- Key vectors by source. Storing
sourceIdandurlon every vector makes replacing a single document trivial.
7. Vectors need a lifecycle too
Every deleted or archived document leaves embeddings behind unless you plan for it. Orphaned vectors cost money and, worse, leak into answers: someone retires an old policy and the assistant still quotes it.
Treat vectors like any other data with a lifecycle:
- Delete or archive vectors in the same workflow that deletes or archives the source.
- Run a periodic cleanup job that finds vectors whose source no longer exists.
- If sources can be restored, make their vectors restorable too.
It's unglamorous work, and it's what makes "delete" actually mean delete.
8. When the answer isn't in your data, retrieve from the web carefully
Sometimes the right context isn't in your knowledge base at all, such as a vendor's latest release notes or a regulation that changed last week. Web search is just another retrieval source, with stricter trust rules:
- Search only when it helps. A cheap routing step can decide whether outside context is needed at all, instead of searching on every request.
- Filter results by relevance score and cap them. The model should only see the strongest sources.
- Have the model write its own cited answer. Treat any third-party summary as unverified context, never as truth.
- Treat web content as untrusted data. Escaping template syntax stops a page from breaking your prompt template, but it is not a prompt-injection defence on its own. Clearly delimit untrusted content and tell the model to treat it as data, not instructions. Don't give steps that read web content access to sensitive tools without human approval. Validate the model's output before acting on it.
The checklist
- Chunks carry rich metadata from creation (source, collection, heading, audience, date, URL)
- Titles and headings are prepended to chunk text before embedding
- One embedding model and dimension size per index, chosen by evaluation
- Vector store matches the access pattern, with a graceful fallback
- Metadata filters, including permissions, run before similarity search
- Every answer cites clickable sources, including honest misses
- Incremental refresh: diff, upload before delete, batched embedding
- Vector lifecycle tied to source lifecycle
- Web retrieval is relevance-filtered, cited, and treated as untrusted (delimited, tool-limited, output-validated)
Good RAG isn't one clever prompt. It's a pipeline of small, boring decisions that each remove a reason for users to distrust the answer.
I build RAG systems, agent workflows and multi-LLM platforms. If you're working on retrieval in production, connect with me on LinkedIn. I'm always happy to compare notes.
Building something agentic?
I share build notes and connect with builders on LinkedIn.