Enterprise RAG: Sources, Permissions, and Answer Quality
On this page
Retrieval-augmented generation (RAG) combines document search with a language model. An enterprise assistant should find current information, cite its sources and respect the user’s access rights. Plan the document collection, permission checks and evaluation together: retrieving a passage does not establish that the generated answer is correct.
Choose the task before the model
Start with five to ten questions people already ask. For example: “Which version of our supplier policy applies to this region?” or “Where is the approved procedure for this exception?” Record who asks, what source they would trust, and what they should do when no current source answers the question.
A search result may be enough. A generated summary adds value when the user needs information brought together from several permitted documents. A recommendation or automated action is a different scope: it needs its own decision rules and review path.
| Scope | Example output | Required control |
|---|---|---|
| Retrieve | Find the current delivery policy | Current source access and version |
| Recommend | Draft an internal follow-up task | Source evidence and human review |
| Act | Create the approved task in CRM | Current write permission, exact approval and recovery |
The controls accumulate as the scope grows. A reliable answer about policy does not authorise a change to an order. See the AI agents in CRM and ERP guide for an approved task-creation example and its failure paths.
Inventory documents and ownership
List the repositories, file types, document owners, update frequency and access rules. Decide which drafts, archives and duplicates should be excluded. Preserve the source URL or identifier, version and update time through ingestion so users can inspect the original.
Test representative PDFs, spreadsheets and scans before estimating extraction for the full collection. Check table rows, column headings and source references: a value detached from its heading can change the meaning of an answer.
Keep permissions attached to the source
A user must not receive an answer derived from a document they cannot open. Record the relevant permission information during ingestion and enforce it during retrieval, before text is sent to the model. Test changes in role, group membership and source-document permissions, as well as cached answers. Microsoft’s document-level access guidance describes permission-aware retrieval patterns; the exact implementation depends on your identity and document systems.
Treat retrieved text as untrusted input. It may contain instructions that should be quoted as document content, not followed by the assistant. The application should keep the user’s authorised task and the retrieved material separate.
Test revocation and freshness, including cached answers
Consider POLICY-1, an approved document at version 3. An operator can initially read it. The example below changes the permission or source version while leaving the indexed citation at version 3. Cached answers undergo the same checks before reuse.
Scroll the table horizontally to compare all columns.
| Scenario | Required outcome | Computed outcome |
|---|---|---|
| Current source, permitted reader | Source checks passed | Source checks passed |
| Permission revoked after indexing | Access denied | Access denied |
| Indexed version 3; source is now version 4 | Refresh required | Refresh required |
| Cached answer; permission revoked | Access denied | Access denied |
| Cached answer; source version changed | Refresh required | Refresh required |
| Source withdrawn from publication | Source unavailable | Source unavailable |
On an access failure, withhold the affected content. If a version changed, refresh the source and retrieve again before producing an answer. If the source was withdrawn, exclude it. A cached answer combining several documents must pass the checks for every supporting source; otherwise discard it and rebuild from permitted, current material. Do not simply remove a citation while retaining the answer derived from it.
This model uses a supplied snapshot of current permissions and versions; it does not connect to a document system, generate answers or measure retrieval quality. Its checks only help if the application receives current source metadata. In a real system, define the permission-sync delay, failure behaviour and how to recheck before delivering a response. Recheck cache reuse and follow-up conversation context too. Blocking new retrieval does not erase text already delivered to a user.
For the chosen connector and permission feature, verify synchronisation delays, failure behaviour and production support. The Microsoft guidance linked above distinguishes application-managed security filters from native permission features; several native options remain in preview.
Select retrieval by evidence
Keyword search works well for exact identifiers and terminology. Semantic retrieval can help with different wording. A hybrid approach, filters and reranking may improve results, but each adds cost or complexity. Compare them on questions your users actually ask instead of assuming one architecture always wins.
Create a small evaluation set with the expected source documents. Check whether the right document is retrieved before grading the generated answer. Changing only the answer-generation prompt cannot supply missing source evidence.
Keep a separate set of questions out of prompt and retrieval tuning. Fix the document versions, user permissions and retrieval settings for each comparison, and record changes between runs. Otherwise, a better score may reflect easier questions or newer sources rather than a better retrieval method.
For an illustrative check, suppose a question requires two passages and the first five results contain one of them. Passage recall at five is 1/2, or 50%, for that question. This applies the standard definition of recall to a fixed number of retrieved passages; it does not measure answer accuracy. For questions that the permitted corpus cannot answer, assess whether the assistant correctly declines instead of assigning a recall score with no relevant passages. Report counts and results by question type so a strong average cannot hide a failing workflow.
Make answers reviewable
Show the passages or documents used for an answer, their dates and a way to open them. Let the system say that it cannot answer when the available sources do not support a conclusion. Identify when a human must review a result, especially if it affects a customer, payment, health or another consequential workflow.
Review incorrect answers with document owners. Distinguish a missing or misread source from a conclusion the source does not support: each needs a different correction. Microsoft’s RAG guidance covers retrieval, citations, access and operating costs.
Price a bounded pilot before expanding it
A one-team, one-source pilot makes the decision testable. Set evaluation thresholds and a review owner before trying it; critical permission failures block expansion. Compare supported answers, correct refusals, time and cost against the existing search workflow. Separate source ingestion and refresh, search/model calls, hosting, human review and required support. Agree usage limits, a spending cap, alerts and who may approve more use.
Plan for change after launch
Documents are replaced; permissions change; new questions appear. Define who updates the corpus, how quickly a change appears in search, how stale answers are identified, and who reviews recurring failures. Monitor usage without copying unnecessary sensitive content into analytics.
If you are planning a knowledge assistant, share the document sources, user roles and questions it should answer. We can scope the first workflow and compare it with the wider AI solutions available for your business software.
