Retrieval-augmented generation, or RAG, connects a language model to your own documents so it can answer questions with your organization's knowledge and cite its sources. A working demo takes days. A system that employees rely on, and that security and compliance teams accept, takes deliberate engineering in four areas: access control, evaluation, safety, and cost.
This article covers each of them, based on the architecture we recommend for internal knowledge assistants.
The moving parts
A production RAG system is a pipeline, not a prompt. Documents are ingested from source systems, split into chunks, converted into embeddings, and stored in a search index. At question time, the system retrieves the most relevant chunks, passes them to the model with instructions, and returns an answer with citations.
Each stage is a place where quality is won or lost. Poor chunking splits a policy across two fragments so neither makes sense alone. Stale ingestion means the assistant confidently quotes last year's process. Most quality problems blamed on the model are actually retrieval problems.
Access control belongs in retrieval, not in the prompt
The most serious risk in an internal assistant is not a wrong answer. It is a correct answer given to someone who was never allowed to see the source document. Instructions such as "do not reveal confidential information" in the prompt are not a control. They can be ignored or worked around.
Enforce permissions at retrieval time instead. Store each chunk with the access rules of its source document, pass the user's identity from your single sign-on provider into the retrieval call, and filter results before anything reaches the model. If a user cannot open the document in SharePoint, Confluence, or your file system, the assistant should not be able to retrieve it for them.
- Carry document permissions into the index alongside each chunk.
- Propagate the signed-in user's identity to every retrieval request.
- Re-sync permissions on a schedule, since access changes more often than content.
- Log which sources were retrieved for which user, so answers can be audited.
Measure quality before users do
"It looks good" is not an evaluation method. Before launch, build a golden question set: real questions from the people who will use the system, with the correct answer and the document it should come from. Fifty to a few hundred questions is a practical starting point for a single domain.
Evaluate retrieval and generation separately, because they fail in different ways and need different fixes.
- Retrieval: does the correct source appear in the top results? Track recall at k and how much irrelevant context is retrieved.
- Groundedness: is every claim in the answer supported by the retrieved text?
- Answer relevance: does the response actually address the question asked?
- Citation accuracy: do the cited sources contain what the answer says they contain?
Automate these checks and run them on every change to prompts, chunking, models, or the index, the same way you run tests before deploying code. Combine automated scoring with regular human review of a sample of real conversations. Models used as judges are useful, but they need to be checked against human judgment too.
Treat retrieved content as untrusted
Any document in the index can contain text written to manipulate the model, whether by accident or on purpose. This is prompt injection, and it becomes more dangerous when the assistant can take actions such as sending email or updating records.
Keep a clear separation between instructions and retrieved content, give tools the minimum permissions they need, require human confirmation for consequential actions, and monitor for unusual behavior. Assume that some content in a large document collection will eventually be hostile.
Keep cost and latency predictable
Costs in RAG systems scale with tokens: how much retrieved text you send to the model, how long answers are, and how many calls each question triggers. Small design choices compound quickly at scale.
- Retrieve fewer, better chunks rather than padding the context window.
- Use smaller, cheaper models for routing, classification, and query rewriting, and reserve larger models for the final answer.
- Cache answers to frequent questions, with invalidation when source documents change.
- Set token budgets per request and attribute spend to the team or product that generates it.
Running it after launch
A RAG system is never finished. Content changes, users ask new kinds of questions, and model providers update their models. Plan for freshness pipelines that keep the index current, a simple way for users to flag bad answers, and a regular review where flagged answers are added to the golden question set.
That feedback loop is what turns a pilot into a system people keep using.
Checklist
- Permissions enforced at retrieval time, using the user's real identity.
- A golden question set from real users, with expected sources.
- Automated retrieval and answer evaluation on every change.
- Retrieved content treated as untrusted; tools with least privilege.
- Token budgets, caching, and cost attribution per team.
- Freshness pipelines and a user feedback loop.
If you are deciding where AI can create value in your organization, and what it would take to run it safely, our AI Opportunity Workshop ends with prioritized use cases, data-readiness notes, and a recommended pilot with clear success criteria.
Filed under AI