How a Private RAG Assistant Answers From Your Own Documents
A private RAG assistant has three separate jobs: retrieve the right passage, attach the passage to the claim, and refuse when retrieval comes back empty. The refusal is an engineering output
TL;DR
- A private RAG assistant has three jobs, and only the first one gets attention: retrieve the passage that answers the question, tie every claim in the answer to the passage it came from, and refuse to answer when retrieval returns nothing relevant.
- Anthropic's Contextual Retrieval work put a number on the retrieval half. Prepend 50-100 tokens of chunk-specific context to each chunk before embedding and indexing, and failed retrievals fall by 49%; add reranking and the reduction reaches 67%. The metric behind it is 1 minus recall@20, and the failure rate moved from 5.7% to 2.9% and then to 1.9%.
- Retrieval quality is not a model property. Dense embeddings miss exact identifiers, BM25 misses paraphrases, and Qdrant's own documentation states the case plainly: a search result can look plausible and still be wrong, while your logs record a successful query.
- Citations are what make the answer checkable. The Claude citations API chunks documents into sentences and returns a pointer to the source text behind each claim, so the link between claim and passage is produced by the API rather than by an instruction in the prompt.
- Rolling it out on a 20-person corpus is a measurement project with a chat interface attached: 30 to 50 labelled questions with known answers, the fusion settings and the reranker cut-off treated as measurable dials, and a re-measure after every corpus change.
What a grounded answer is built from
Four stages, and they fail independently. Ingestion splits the corpus into chunks. Indexing turns each chunk into something searchable, usually a vector plus a lexical representation. Retrieval takes the question and returns a shortlist of passages. Assembly hands those passages to a model and asks for an answer built from them.
An assistant is grounded when the fourth stage can only use what the third stage retrieved, and when the answer carries the evidence for every claim it makes. That definition matters commercially: it is the difference between a system that can be checked by the person who asked the question and a system that produces confident prose at the speed of a chat window.

Take a concrete case. Imagine a Casablanca logistics company with 4,000 supplier contracts, customs declarations and internal procedures written in three languages, and a legal team that will only approve a tool where the documents stay on infrastructure the company controls. The corpus is the asset. The assistant is a reader of that corpus, and everything below is about how well it reads.
How retrieval misses the passage that matters
Standard retrieval has one search path, and it has two predictable blind spots. A dense embedding compresses meaning, so a query for a specific contract reference or a document number can return documents about the right subject and miss the exact string. A sparse, term-weighted index has the opposite failure: it finds the identifier and loses the paraphrase, so a question phrased as "end of agreement" never reaches a document that says "termination".
Both failures are invisible from the outside, because both return results. This is the point Qdrant makes at the start of its hybrid search article: dense retrieval can return a document on the right topic while missing an exact identifier copied into the query, sparse retrieval can miss a relevant document when the query uses terms the corpus never used, and either way the query logs record a success.
The blind spot has a second half that costs money rather than accuracy. Adding a second retriever means a sparse vector per point, a second index and an extra search on every query. Qdrant measured that on one container serving one request at a time, and the extra search raised median query latency by 0.60 to 1.47 milliseconds. Adding a sparse vector to a collection that already holds dense vectors requires a new collection and a full reindex, because the vector configuration is fixed when the collection is created.

Why hybrid retrieval changes the failure rate
Run both retrievers on the same query and merge the two ranked lists. The merge is where most implementations quietly go wrong, because a cosine similarity of 0.7 and a BM25 score of 12.4 do not share a scale, and any fixed weight between them is a number that has to be re-tuned every time the corpus or the query mix changes.
Reciprocal Rank Fusion sidesteps the comparison by using rank position only. A document's fused score is the sum of one over (constant plus its rank) across the lists it appears in. Cormack, Clarke and Buettcher introduced the method in 2009 and it remains the default for good reason; Qdrant reports that across five public datasets, default RRF beat the stronger individual retriever on four. Distribution-Based Score Fusion is the alternative, and it preserves the size of score gaps rather than flattening them.
The contextual half of the problem sits upstream of fusion. A chunk that says "revenue grew by 3% over the previous quarter" is retrievable only if the index knows which company and which quarter the sentence belongs to. Anthropic's answer is to generate that missing context with a model and prepend it to the chunk before embedding and before indexing, which raises the chunk's cost slightly and its retrievability a great deal. The published one-time cost was $1.02 per million document tokens using prompt caching, and the measured effect was a 49% reduction in failed retrievals, rising to 67% with reranking.

Citations turn an answer into something you can check
A correct passage in the context window is a necessary condition. It becomes a useful answer only when the reader can see which passage produced which sentence.
The Claude citations API does that mechanically. Citations are enabled per document on the request. Plain text documents are chunked into sentences so that a citation can point at a single sentence or chain consecutive ones, and the API returns, for each claim, a reference to the text the claim came from. Two consequences follow. Source text returned through the citation field does not count toward output tokens, so the evidence is cheaper than asking the model to reproduce quotes. And because the API parses the citations out of the response, citations are guaranteed to contain valid pointers to the provided documents, rather than whatever the model decided to write between quotation marks.
The design implication is straightforward, and it is the part teams skip: if a sentence cannot be traced to a passage, the answer is not finished. Claim-level traceability is a property of how the answer is assembled, and it survives contact with an auditor in a way that a summary never does.
The refusal path is a retrieval measurement
Refusal gets treated as a safety feature or a prompt instruction. In a document assistant it is arithmetic. If the retrieved passages do not clear the relevance bar, there is nothing to answer from, and the honest output is a statement that the corpus does not cover the question.
Which means the quality of refusal is entirely dependent on the quality of the retrieval metric. Anthropic's measurement is 1 minus recall@20: the share of relevant documents that failed to appear in the top 20 chunks. That number is meaningful precisely because it counts misses, not hits. A team that measures only the answers it likes has no way to know what the assistant was never able to see.

Netics' position after building these systems is that the measurement comes first and the prompt goes last. Build 30 to 50 questions with known answers drawn from the actual corpus, record which passage answers each one, then measure. Refusal behaviour is a threshold decision you make with that table in front of you, and it is the same table that tells you whether contextual chunking or reranking is worth the latency.
Rolling it out on a 20-person corpus
A first implementation on a small corpus is a measurement project with a chat interface attached. Keep the extraction pipeline separate from the assistant, so a re-chunk changes one thing. Keep the evaluation questions in the repository next to the code that answers them. Keep the retrieval parameters visible: the fusion method, the prefetch limits and the reranker's cut-off are the dials that move relevance, and each of them can be measured against the same labelled set.
Where the design pays for itself is the boundary. A company that can say which documents the assistant may read, which passage supports each answer and when the assistant refuses is a company that can put the tool in front of a client-facing team: data-residency constraints in France and Morocco are easier to satisfy when the corpus, the index and the model calls all sit on infrastructure you operate. The Netics page on the private AI assistant for your documents describes the offer this pillar supports.

The uncomfortable trade-off is worth stating. Every improvement to recall adds a stage, and every stage adds latency, storage or reindexing work. The goal is a system whose failure rate you can quote, not one whose answers sound better than last quarter's. That is the difference between a demo that impressed a room and an assistant a legal team will sign off on.
If building that measurement is the part you want help with, we run a free 30-minute audit of the pipeline you already have — book one at neticslabs.com and bring the corpus.
Sources
Source: Introducing Contextual Retrieval — anthropic.com, 2024-09-19 (context prepended to each chunk before embedding and BM25 indexing, 50-100 tokens per chunk, 49% and 67% reductions in failed retrievals, 1 minus recall@20 as the metric, $1.02 per million document tokens with prompt caching; the standard RAG and contextual retrieval preprocessing figures reused with credit). Source: Hybrid Search in Qdrant — qdrant.tech, 2026-08-24 (dense and sparse failure modes, Reciprocal Rank Fusion from Cormack, Clarke and Buettcher 2009, RRF beating the stronger individual retriever on four of five public datasets, the 0.60 to 1.47 ms latency measurement and the full-reindex requirement). Source: Citations — docs.claude.com, retrieved 2026-10-04 (per-document citations, sentence chunking of plain text, cited text excluded from output tokens, valid pointers to the provided documents). Internal linkage: RAG vs. Fine-Tuning for Production AI. More on Netics' AI work at neticslabs.com.
Source: Introducing Contextual Retrieval — anthropic.com, 2024-09-19; Hybrid Search in Qdrant — qdrant.tech, 2026-08-24; Claude citations — docs.claude.com, retrieved 2026-10-04. Figures: official Anthropic diagrams, fetched 2026-10-04.