AI ImplementationDefinitionFreshLast reviewed: · 8d ago

    RAG (Retrieval-Augmented Generation)

    /ræɡ/
    RAG is an AI architecture that retrieves relevant passages from an external knowledge source, such as a company's internal documents, and passes them to a large language model so that its answer is grounded in, and can cite, those sources. Knowledge is updated by re-indexing documents, not by retraining the model.
    Also known as: Retrieval-Augmented Generation · Retrieval-augmented generation · RAG pipeline · RAG architecture · Grounded generation · Permission-aware enterprise search

    TL;DR

    Quick Answer
    Cited by AI
    RAG retrieves passages from your own documents, filters them by the user's permissions and lets an LLM answer from them with citations.
    Linus Ingemarsson - Author at Alice Labs
    Written by
    Eric Lundberg - Reviewer at Alice Labs
    Reviewed by
    Published ·Updated
    11 min read

    In context

    Internal policy search

    "'Instead of opening five SharePoint sites, employees ask the assistant which travel policy applies and get an answer that links to the exact section of the current policy document.'"

    Permission-aware retrieval

    "'The RAG index stores each document's access list, so a question from someone outside HR never retrieves chunks from HR-only folders, even if the question matches them perfectly.'"

    Architecture decision

    "'Our content already lives in Microsoft 365 and permissions are clean, so we started with Copilot and only built a custom Azure AI Search pipeline for the engineering archive that Copilot does not reach.'"

    RAG vs fine-tuning

    "'We chose RAG over fine-tuning because our product documentation changes every week and every answer has to point to its source.'"

    Related terms

    RAG vs fine-tuning Vector database Security trimming AI Agent AI knowledge base

    Key points

    • RAG answers questions from your own documents: it retrieves relevant passages first and lets the language model write an answer grounded in them, with citations.
    • In enterprise search the hard part is not the model but retrieval quality and permissions. Every chunk in the index must carry the access rights of its source document.
    • Microsoft 365 Copilot and the Microsoft Search API only return content the signed-in user can already access, so overshared SharePoint sites become overshared AI answers.
    • Restricted SharePoint Search is retiring (no new enablement from 31 July 2026). Microsoft points to Restricted Content Discovery, SharePoint Advanced Management and Purview instead. None of them change permissions.
    • For Swedish companies, GDPR applies as soon as indexed documents contain personal data. IMY's guidance on GDPR and AI is the starting point; EU Data Boundary and region choice cover where data is processed.
    • Measure RAG on three things before rollout: does retrieval find the right passages, is the answer faithful to them, and do the citations point to the right source.
    01 / 10Section

    What is RAG, in one paragraph?

    In short

    RAG is a pattern where an AI system searches a knowledge source, usually internal documents, for relevant passages and gives them to a language model that answers from them and cites them.

    A language model only knows its training data, not your travel policy or last quarter's board decision. RAG adds a search step before the answer: the system retrieves the most relevant passages from your documents and places them in the model's context, so the answer rests on your sources. The term comes from a 2020 paper by Lewis et al. at Facebook AI Research.

    For Swedish companies the most common use is internal knowledge search: employees ask questions in Swedish or English across SharePoint, Teams, Confluence, file shares, intranets and ticket systems, and get a short answer with links to the documents it came from. The same pattern powers customer-facing assistants and AI agents that need company knowledge to act.

    02 / 10Section

    How does RAG work, step by step?

    In short

    RAG ingests documents, chunks them, embeds each chunk, indexes chunks with metadata and permissions, retrieves the best chunks per question and generates an answer citing them.

    Every RAG system follows the same seven steps, whatever product or framework it uses:

    1. Ingest. Connectors read documents from source systems and extract text, together with metadata such as owner, date and the document's access list.
    2. Chunking. Documents are split into passages small enough to be precise but large enough to keep their meaning, usually along headings, each keeping a reference to its parent document.
    3. Embeddings. An embedding model converts each chunk into a vector representing its meaning, so a question and a matching passage end up close even with different wording. For Swedish content, test multilingual models on your own documents.
    4. Index. Vectors, original text, metadata and permissions are stored in a search index or vector database, usually with hybrid search (vectors plus keywords) because product codes and names are poorly served by vectors alone.
    5. Retrieval. The system embeds the question, applies filters (permissions first), runs the search and often re-ranks the top results.
    6. Generation. The best chunks go into the prompt with instructions: answer only from the sources, and say so when they do not answer the question.
    7. Citations. Each claim links to the document it came from, so users can verify and the organization can audit.

    When a document is updated, deleted or its permissions change, the index must follow, or the assistant answers from outdated or no longer permitted content.

    03 / 10Section

    RAG vs fine-tuning vs long context: which should you use?

    In short

    Use RAG for large, changing document sets with citations. Use fine-tuning to change behavior or format, not to add facts. Use long context for a few whole documents at once.

    These three approaches are often presented as alternatives, but they solve different problems. A deeper comparison is in RAG vs fine-tuning.

    Approach What it changes Good for Weak at
    RAG What the model sees per question Large, changing document sets, citations, per-user permissions Questions that need reasoning over an entire corpus at once
    Fine-tuning The model's weights Tone, output format, domain-specific behavior Adding or updating facts, traceability, access control
    Long context How much text is sent in one prompt Analyzing a handful of known documents in full Searching thousands of documents, cost and latency per question

    They combine well: RAG can retrieve whole documents into a long context, and a fine-tuned model can be the generator. But only retrieval can enforce who sees which document, so internal search starts with RAG.

    04 / 10Section

    How does permission-aware retrieval and security trimming work?

    In short

    Security trimming removes every result the user may not open before the model sees it. Copilot and the Graph Search API run as the signed-in user; in a custom index, each chunk must carry its document's access list.

    An enterprise assistant must never answer from a document the asker cannot open. How that is enforced depends on where retrieval happens.

    • Microsoft 365 Copilot and Microsoft Graph. Microsoft states that Copilot only surfaces organizational data to which the user has at least view permission, and the Microsoft Search API in Microsoft Graph runs in the context of the signed-in user with delegated permissions, with results scoped by the access control on each item. Retrieval through Graph is security trimmed by design.
    • The oversharing problem. Trimming enforces the permissions you have. A site open to "Everyone except external users" will be used, which is why Copilot projects often start with a permissions review.
    • Restricted SharePoint Search is being retired. It limited organization-wide search and Copilot to an allow list of up to 100 sites, but was never a security boundary. Since 31 July 2026 new enablement is blocked, and Microsoft points to Restricted Content Discovery, SharePoint Advanced Management and Microsoft Purview. Restricted Content Discovery hides selected sites from organization-wide search and Copilot but, like its predecessor, does not change permissions.
    • Custom RAG on Azure AI Search. The generally available option is a security filter: store user or group IDs on each chunk and filter every query by the caller's identities. Native enforcement of SharePoint ACLs, ADLS Gen2 ACLs and Purview sensitivity labels with the user's Entra token is available in preview as of September 2026. Permission changes only take effect after they are synced to the index.
    • Other platforms. Glean enforces source-system permissions on every result, and Elastic offers document level security (beta) for connectors including SharePoint Online. Always test with users from different departments.

    A Swedish-language walkthrough of the SharePoint case is in our article on AI search in SharePoint without exposing documents users lack access to.

    Linus IngemarssonEric Lundberg
    Alice Labs practitioner team

    Talk to the team behind 100+ AI implementations

    30-minute discovery call with a senior Alice Labs consultant. No slide deck, no sales pitch — just a scoping conversation.

    Book a Discovery Call
    05 / 10Section

    Which RAG architecture fits: Copilot, Azure AI Search, a vector database, Glean or Elastic?

    In short

    Copilot fits content in Microsoft 365 with little engineering. Azure AI Search or a vector database fits when you need control. Glean and Elastic fit knowledge spread across many non-Microsoft systems.

    There is no single best stack. The right choice follows from where your documents live, how much control you need and who will operate the system. Note that the Microsoft 365 app is now called the Microsoft Copilot app, while the licensed product is still Microsoft 365 Copilot.

    Option Fits when Permissions Trade-off
    Microsoft 365 Copilot Most knowledge is in SharePoint, OneDrive, Teams and Outlook Inherited from Microsoft 365 via Microsoft Graph Little control over chunking, ranking and evaluation; per-user licenses
    Azure AI Search (custom RAG) You need your own assistant, sources outside M365 or strict answer rules Security filters (GA); native SharePoint ACLs and Purview labels (preview) Requires engineering, monitoring and an owner
    Own vector database Specific hosting, cost or model requirements; product-embedded search You build it: ACL sync and query filters Most control and most responsibility
    Glean Knowledge spread across many SaaS tools Enforces source-system permissions on results Verify hosting region; customer-hosted option in your own GCP or AWS
    Elastic You already run Elasticsearch or need strong keyword plus vector search Document level security for supported connectors (beta) Cluster operations and relevance tuning in-house

    Many organizations end up with two layers: Copilot for everyday work and a dedicated RAG knowledge base for a domain such as support or engineering documentation.

    06 / 10Section

    What do GDPR and EU data residency mean for RAG in Swedish companies?

    In short

    If indexed documents contain personal data, GDPR applies to the whole pipeline. IMY's guidance on GDPR and AI is the Swedish reference; EU Data Boundary or EU regions govern processing location.

    Internal documents almost always contain personal data. The Swedish Authority for Privacy Protection (IMY) states in its guidance on GDPR and AI that GDPR applies when personal data is processed in developing or using AI. For RAG, that means:

    • Scope the index. Decide which sources are indexed and why. Leaving HR case folders out is often easier than protecting them afterwards.
    • Know every processor. Search service, embedding model and language model may be different providers, each needing a data processing agreement.
    • Check where processing happens. Microsoft states that Copilot is an EU Data Boundary service for EU customers, but also notes that Anthropic models offered in Copilot are currently excluded from the EU Data Boundary. Azure services let you choose EU regions such as Sweden Central. For Glean and other SaaS tools, confirm the hosting region in the contract.
    • Deletion and retention. Erasure must reach chunks and embeddings, and question logs need their own retention rule.
    • Impact assessment. Broad indexing of employee data may call for a DPIA. Involve your data protection officer early.

    This is not legal advice, but the checklist Swedish legal and IT teams typically work through.

    Linus IngemarssonEric Lundberg
    Alice Labs practitioner team

    Planning internal AI search over your documents?

    Alice Labs builds permission-aware AI knowledge bases for Swedish companies, on Microsoft 365 Copilot or a custom RAG stack with processing in the EU.

    See the AI knowledge base service
    07 / 10Section

    How do you evaluate a RAG system?

    In short

    Evaluate retrieval quality, faithfulness and citations, plus permission tests, on a fixed set of real questions from your own organization.

    A RAG system is ready when it passes tests you wrote before building it:

    • Retrieval quality. For each test question, is the passage that contains the answer among the top results? Measure this separately from the answer.
    • Faithfulness. Does every statement in the answer follow from the retrieved passages, and does the assistant say "I can't find this" when they do not?
    • Citations. Does each citation point to a passage that actually supports the claim, in the current version of the document?
    • Permissions. Run the same questions as users with different access and confirm restricted content never appears.

    Build the test set from real employee questions, in Swedish and English if both are used. Frameworks such as RAGAS automate parts of the scoring; human review of a sample is still needed. Re-run it after every change.

    08 / 10Section

    What are the most common RAG mistakes?

    In short

    Indexing before cleaning permissions, duplicate versions, vector-only search, no evaluation and no owner after launch.

    • Indexing before cleaning permissions. RAG makes oversharing visible in seconds.
    • Duplicate and outdated versions. Five copies of a policy in different folders produce contradictory answers. Mark authoritative sources and archive the rest.
    • Vector search only. Article numbers, abbreviations and Swedish compound words often need keyword search as well.
    • Losing structure. Tables, headings and scanned PDFs break when text extraction is naive.
    • No "I don't know". Without an instruction and a test for it, the model fills gaps with plausible text.
    • No owner. Someone must own the sources, the sync and the test set after launch.
    09 / 10Section

    What drives the cost of a RAG system?

    In short

    Sources, document volume and change rate, permission complexity, model usage, licenses or service tier, and ongoing operations.

    Costs vary too much for a single figure, but the drivers are predictable:

    • Sources and connectors. Each system with its own format and permission model adds integration work.
    • Volume and change rate. More documents and frequent updates mean more embedding, indexing and storage.
    • Permission complexity. Unique item-level permissions and nested groups are harder to sync than site-level access.
    • Usage. Every question costs model tokens for the retrieved context and the answer; re-ranking and agentic multi-step retrieval add more.
    • Licenses or service tiers. Per-user licenses for Copilot or Glean versus capacity-based pricing for Azure AI Search or Elastic.
    • Operations. Monitoring, evaluation and content ownership recur and are easy to leave out of the business case.
    Linus IngemarssonEric Lundberg
    Alice Labs practitioner team

    Talk to the team behind 100+ AI implementations

    30-minute discovery call with a senior Alice Labs consultant. No slide deck, no sales pitch — just a scoping conversation.

    Book a Discovery Call
    10 / 10Section

    How does Alice Labs build RAG for internal knowledge search?

    In short

    Questions first, then a source and permission audit, an architecture choice, a build with permission-aware retrieval and citations, and evaluation before launch.

    Alice Labs is a Swedish AI consultancy in Stockholm. When we build an AI knowledge base or internal AI search for Swedish companies, the work follows the same order:

    1. Questions first. Collect real questions and turn them into the test set.
    2. Source and permission audit. Map where answers live, which versions are authoritative and where access is too broad.
    3. Architecture choice. Copilot where it is enough, a custom pipeline on Azure AI Search or another index where control requires it, processed in the EU.
    4. Build with guardrails. Permission filters in every query, citations on every answer, and an explicit "not found" behavior.
    5. Evaluate, launch, hand over. Pass the tests, pilot, then hand over ownership of sources, sync and evaluation.

    The Swedish service page is AI-kunskapsbas. When the assistant should also act, for example create tickets or update records, the next step is an AI agent with RAG as one of its tools.

    About the Authors & Reviewers

    Published ·Updated
    Written by
    Linus Ingemarsson - CEO & Co-Founder, Alice Labs at Alice Labs
    Linus Ingemarsson

    CEO & Co-Founder, Alice Labs

    CEO & Co-Founder at Alice Labs. Author of 7 research reports on AI adoption, governance and labor markets cited across EU, OECD and US benchmarks.

    • 8+ years in AI strategy & implementation
    • Top-5 AI Speaker, Sweden (Mindley 2025)
    • 100+ enterprise AI engagements
    Reviewed by
    Eric Lundberg - Co-Founder, Alice Labs at Alice Labs
    Eric Lundberg

    Co-Founder, Alice Labs

    Co-Founder at Alice Labs. Builds AI automation, agent workflows and integration systems that hold up in real business operations.

    • AI automation & agent systems lead
    • Workflow design across 100+ deployments
    • Specialist in RAG, integrations & APIs
    Published · Updated
    Reviewed for technical accuracy, methodology and source integrity.·All claims trace to public sources cited in-line.

    Frequently Asked Questions

    What is RAG in simple terms?

    RAG lets an AI assistant look things up in your documents before answering. It finds relevant passages, gives them to the language model and returns an answer based on them, with source links.

    How do companies in Sweden use RAG?

    Mostly for internal knowledge search: employees ask in Swedish or English across SharePoint, Teams and document archives and get cited answers. Other uses are support assistants and AI agents.

    What is security trimming in RAG?

    Security trimming removes every result the user may not open before the model sees it. In Microsoft 365, retrieval runs as the signed-in user. In a custom index, each chunk stores its document's access list and queries are filtered by identity.

    Does Microsoft 365 Copilot use RAG?

    In effect, yes. Copilot retrieves content through Microsoft Graph and grounds answers in content the user can access. It inherits Microsoft 365 permissions, so overshared sites are available to it too.

    Can Restricted SharePoint Search protect sensitive sites from Copilot?

    Not for new deployments: new enablement is blocked from 31 July 2026, and it was never a security boundary. Microsoft recommends Restricted Content Discovery, SharePoint Advanced Management, Purview and fixing permissions.

    Does Azure AI Search support SharePoint permissions for RAG?

    Yes. Security filters with user or group IDs per document are generally available. Native SharePoint ACL and Purview label enforcement is in preview as of September 2026.

    Is RAG GDPR compliant?

    It can be, depending on setup: which documents are indexed, legal basis, processing agreements, processing location and deletion of embeddings. IMY's guidance on GDPR and AI is the Swedish starting point.

    What is the difference between RAG and fine-tuning?

    RAG gives the model documents at question time; fine-tuning changes its weights to adjust behavior. For internal search RAG is the default, since knowledge changes, answers need citations and access must follow permissions.

    Do you need a vector database for RAG?

    You need an index with vector search, not necessarily a standalone vector database. Azure AI Search and Elasticsearch combine vector and keyword search, which most enterprise RAG needs.

    How long does it take to build a RAG system?

    A prototype is quick. Production depends on sources, the state of permissions and evaluation strictness. The permission audit and test set usually take longer than the pipeline.

    Previous in AI Implementation

    Why AI Projects Fail: 7 Root Causes & How to Avoid Them

    Next in AI Implementation

    What Is MLOps? Machine Learning Operations Explained

    Further reading

    Related services

    Related reading

    Sources

    1. Lewis et al. — Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (arXiv:2005.11401, 2020)(accessed 2026-09-29)
    2. Microsoft Learn — Data, Privacy, and Security for Microsoft Copilot (permissions, EU Data Boundary, no training on Graph data)(accessed 2026-09-29)
    3. Microsoft Learn — Use the Microsoft Search API to query data (requests run as the signed-in user, results scoped by access control)(accessed 2026-09-29)
    4. Microsoft Learn — Restricted SharePoint Search (retiring, 100-site allow list, not a security boundary)(accessed 2026-09-29)
    5. Microsoft Learn — Restrict discovery of SharePoint sites and content (Restricted Content Discovery)(accessed 2026-09-29)
    6. Microsoft Learn — Document-level access control in Azure AI Search (security filters, SharePoint ACLs, Purview labels)(accessed 2026-09-29)
    7. Elastic — Document level security in Elastic connectors(accessed 2026-09-29)
    8. IMY — Vägledning om GDPR och AI (guidance on GDPR and AI)(accessed 2026-09-29)

    Next scheduled review:

    Linus IngemarssonEric Lundberg
    Alice Labs practitioner team

    Talk to the team behind 100+ AI implementations

    30-minute discovery call with a senior Alice Labs consultant. No slide deck, no sales pitch — just a scoping conversation.

    Book a Discovery Call
    Share

    Get in Touch!

    The lab usually responds within 24 hours.

    Need help with AI?Get in touch