Back to blogAI and MLOps

Retrieval-augmented generation: build an evidence system, not just an index

Design retrieval-augmented generation around authoritative sources, preserved document context, access enforcement, separate retrieval evaluation, and freshness.

Retrieval-augmented generation can make business documents available to an assistant. It does not automatically make the answer correct, current, or authorized. The system can retrieve an obsolete policy, omit an exception from the next page, or combine guidance intended for different regions into one convincing response.

The central design question for business knowledge systems is which evidence the application can legitimately use to answer a particular user’s question. Retrieval quality, document lifecycle, permissions, and answer behavior all contribute to that decision. A vector index is one component of the system, not the knowledge governance model.

Consider a hypothetical service organization building an assistant for customer-support policies. Its repository includes current guidance, old versions, regional amendments, and internal exception procedures. A support agent asks whether a customer qualifies for a replacement. The answer must use the applicable policy without exposing procedures the agent is not permitted to see.

Define the answerable scope before indexing documents

Specify the questions the system is intended to answer. Explaining a published policy differs from deciding whether a specific customer is eligible under that policy. The latter may require customer state and authorized business rules outside the document repository. A text answer should not silently become an eligibility decision merely because it sounds definitive.

Define what constitutes sufficient evidence. The replacement question may require the policy, region, product category, effective date, and relevant exception. If one is missing, the assistant should ask for context, explain the limit, or route the request appropriately. The system needs a meaningful boundary between an answer and an unresolved question.

Identify the authoritative sources for that scope. Informal notes can help locate issues without serving as final policy. Treat source status as metadata rather than hoping the model will infer authority from polished wording. An obsolete but well-written document can be more semantically similar to a question than the current approved instruction.

Start with a coherent collection whose owners can maintain it. Indexing every document first can increase ambiguity and access exposure faster than it increases useful coverage. A bounded policy area allows the team to evaluate completeness and update behavior before expanding into material with different ownership and approval rules.

Preserve meaning during document preparation

Extract text with attention to structure. Tables, headings, footnotes, and attachments can carry conditions essential to interpretation. A policy sentence detached from the table column specifying its region may become misleading. Test extraction on the documents the organization actually uses, including scanned or unusually formatted files.

Choose chunk boundaries that preserve useful context. Small chunks can improve focused matching but omit exceptions or definitions. Large chunks can carry more context while increasing irrelevant material and processing cost. There is no universally correct chunk size; compare candidate approaches against questions whose answers depend on boundaries within the documents.

Keep document identity, section references, effective dates, and access metadata with each retrievable unit. The system should be able to reconstruct where a passage came from and whether it remains applicable. Metadata added only to the original file is ineffective if the retrieval layer cannot enforce or interpret it at the point of use.

Represent linked context where needed. A regional amendment may override part of a base policy rather than replace the entire document. The assistant needs a reliable way to retrieve or apply that relationship. Simply placing both texts in the same index invites the model to resolve authority through inference when the application should supply an explicit rule.

Enforce permissions before information reaches the model

Authorize retrieval using the authenticated user and the application’s access rules. Restricting the final answer is weaker than preventing unauthorized material from entering the model context. Once sensitive text has been retrieved, relying on the model to suppress it creates an avoidable disclosure path.

Propagate changes in access rights to every relevant layer. Indexed chunks, cached retrieval results, generated responses, and logs can preserve material after its source permissions change. Decide how revocation affects these copies and verify that behavior. A correct document repository does not guarantee that derived stores inherit the same protection.

Control boundaryRequired questionCommon gap
Source collectionIs the material approved for this use?Informal or obsolete documents enter the corpus
RetrievalMay this user receive this passage?Authorization happens only after retrieval
GenerationDoes the answer stay within supported evidence?Plausible inference becomes stated policy
Storage and cachingIs retained content still permitted?Revoked access persists in derived copies

Treat retrieved content as untrusted input, even when it comes from an approved collection. Its contents must not redefine the application’s authority. Documents can contain instructions aimed at changing the assistant’s behavior. Application permissions and allowed actions must remain independent of such content, including material that arrived through an otherwise legitimate source workflow.

Answer with authorized evidence: question scope (establish the task and required context); authorized evidence (retrieve applicable permitted sources); supported answer (cite evidence that establishes the claim); unresolved question (clarify missing context or defer).
When evidence is missing or conflicting, the system should clarify or defer rather than invent an answer. View full-size graphic

Evaluate retrieval separately from answer generation

A wrong answer can begin with missing evidence or with incorrect use of evidence that was available. Separate those failure modes during evaluation. If the current regional policy never reached the context, changing the answer prompt may not solve the problem. If it was retrieved but ignored, retrieval tuning alone is equally insufficient.

Build questions linked to reviewed evidence. Include direct lookups, terminology variations, multi-document conditions, and questions that should remain unanswered. Record which passages are necessary and whether competing passages are inappropriate. This provides a practical way to compare retrieval changes without relying solely on whether a final answer appears fluent.

Review ranking under realistic constraints. Permission filtering and region restrictions can alter the candidate set. A retrieval method that performs well on the entire corpus may behave differently for an individual agent’s authorized subset. Test the actual query path, including metadata filters and any reranking stage used in production.

Do not treat similarity scores as universal confidence in correctness. Their meaning depends on the retrieval method and corpus. A highly similar obsolete policy can still be wrong for the task. Use generative AI evaluation to assess the answer separately, including its handling of unsupported and conflicting information.

Design citations and abstention as useful behavior

A citation should point to evidence that supports the specific claim, not merely to a document containing related words. Evaluate whether the cited section establishes the replacement rule and any qualification in the response. A valid-looking link can create false confidence when the attached passage does not support the conclusion.

Make source context visible without overwhelming the agent. The relevant document title, section, and effective information can help a reviewer verify the answer. Ensure that the citation target is accessible under the same user permissions. Sending an agent to a restricted page after quoting its contents is evidence of an access design failure.

Specify what happens when sources conflict. The system may apply a documented precedence rule or ask for clarification, but it should not invent a merged policy. Surface unresolved conflict to the responsible owner. This behavior also creates feedback about the repository itself, where inconsistent guidance may already be causing manual errors.

Abstention should be actionable. Explain the missing context or route to an appropriate support process when possible. A generic refusal frustrates users and encourages reformulation until the model guesses. The assistant becomes more useful when it can distinguish missing region, absent evidence, and an action outside its authority.

Maintain freshness as an operating capability

Define how approved changes reach the index and how old versions stop influencing answers. Update frequency should reflect the policy’s operating needs. A document marked current in the repository can remain stale in the retrieval store if ingestion failed. Monitor ingestion status and document coverage, not only availability of the assistant endpoint.

Test deletion and replacement explicitly. Adding a new document without retiring the old one can increase contradiction. Removing a source file without deleting its indexed chunks leaves inaccessible history retrievable. Maintain a traceable relationship between source versions and derived entries so lifecycle operations are verifiable.

Decide how effective dates apply. A policy approved today may apply next month, while a historical question may require the version applicable when an event occurred. Current and applicable are not always the same condition. The query and source model should support the required interpretation instead of treating the newest file as universally authoritative.

The principles in data quality and readiness are directly relevant: source meaning and availability determine the capability. Better embeddings cannot repair unclear policy ownership or missing amendments. Some failures require improving the knowledge collection rather than changing the model.

Release a maintainable knowledge product

Assign owners for the corpus, access model, retrieval behavior, and answer evaluation. These responsibilities may belong to different teams, but their handoffs need to be explicit. When the assistant gives obsolete guidance, the organization should know who investigates and who can pause the affected policy area while the issue is resolved.

Measure usefulness through completed support work and correction burden. High query volume can reflect curiosity or repeated failed attempts. Track whether agents can find applicable evidence, resolve the authorized question, and recognize when escalation is needed. Keep those outcomes separate from raw response speed and generated token counts.

Preserve release context for diagnosis. Retrieval configuration, source versions, prompts, and model choices can all affect an answer. Record enough information to investigate a disputed response while respecting data minimization and access requirements. A model identifier alone cannot explain which policy evidence the system used.

For the support organization, the assistant becomes dependable when it retrieves the right authorized policy, preserves its conditions, and handles uncertainty honestly. That capability comes from a maintained evidence system around generation. Indexing documents is the starting mechanism; providing defensible answers is the operating product the organization must design and own.

References and further reading

Research and editorial perspectives relevant to this topic. The project guidance and illustrative examples are Ayterate editorial analysis. Some research may require registration or a subscription.

Next steps

Discuss the implications for your project.

Share your current environment and the decision you need to make. We can help assess the relevant service scope.

Get in touch

Let’s explore this for your business

Tell us how this topic relates to your plans and what you want to achieve. We’ll help you identify the next step.

Fields marked * are required

We’ll use your details to respond to your enquiry. Read our privacy policy.