Prompt

How do I build a RAG app that connects to my company documents?

Artificial Intelligence / AI Infrastructure1 observationLast seen Oct 9, 2026

Latest observation

Oct 9, 2026OpenAI APIWeb search: off

Building a RAG app for company documents usually means:

  1. Ingest documents
  2. Chunk them
  3. Create embeddings
  4. Store them in a vector database
  5. Retrieve relevant chunks at query time
  6. Pass them to an LLM to generate an answer with citations

Here’s a practical blueprint.


1) Define the scope and access rules first

Before coding, decide:

  • What sources: PDFs, Google Drive, SharePoint, Confluence, Slack, Notion, web pages, ticketing systems, etc.
  • Who can see what: document-level and row-level permissions matter a lot.
  • Update frequency: hourly, daily, near-real-time.
  • Answer style: short answers, summarized answers, citations, policy compliance, etc.

If your company has sensitive docs, permission filtering is not optional. Your retriever must only return docs the user is allowed to access.


2) Choose a basic architecture

A common stack:

  • Document connectors: pull from Drive/SharePoint/S3/etc.
  • Parsing/OCR: extract text from PDFs, scans, Office docs
  • Chunking pipeline: split documents into searchable pieces
  • Embedding model: convert chunks to vectors
  • Vector database: Pinecone, Weaviate, pgvector, Milvus, OpenSearch, FAISS
  • RAG API: search + prompt + LLM response
  • UI: chat interface, search, citations, feedback

A simple flow:

Company docs -> ingest -> clean text -> chunk -> embed -> vector DB
User question -> embed -> retrieve top-k chunks -> prompt LLM -> answer

3) Ingest documents

You need a job that regularly fetches documents from your sources.

Examples:

  • Google Drive API
  • Microsoft Graph API for SharePoint/OneDrive
  • Confluence REST API
  • S3 bucket sync
  • Web crawler for internal docs site

Store metadata for each doc:

  • document ID
  • title
  • source system
  • URL
  • owner
  • department
  • timestamps
  • ACL / group permissions

This metadata is crucial for filtering and citations.


4) Parse and normalize text

Documents often contain messy content:

  • PDF layout issues
  • tables
  • headers/footers
  • OCR errors
  • duplicated text

Use tools like:

  • unstructured
  • pymupdf
  • pdfplumber
  • python-docx
  • OCR: Tesseract, Azure OCR, AWS Textract, Google Document AI

Clean up:

  • remove boilerplate
  • deduplicate repeated headers/footers
  • preserve section headings
  • keep tables if they matter

5) Chunk the text intelligently

Don’t embed entire documents. Split into chunks.

Typical starting point:

  • Chunk size: 300–800 tokens
  • Overlap: 50–150 tokens

Best practices:

  • chunk by headings/sections when possible
  • keep tables together
  • preserve document title and section path in metadata
  • avoid splitting on arbitrary character boundaries only

Example metadata per chunk:

{
  "doc_id": "hr-policy-2025",
  "chunk_id": "hr-policy-2025-004",
  "title": "Remote Work Policy",
  "section": "Time Tracking",
  "source_url": "https://...",
  "permissions": ["hr", "managers"],
  "updated_at": "2026-01-15"
}

6) Create embeddings

Use an embedding model to convert each chunk into a vector.

Options:

  • OpenAI embeddings
  • Cohere embeddings
  • Voyage embeddings
  • SentenceTransformers / local models

For enterprise doc search, strong general-purpose embeddings are usually enough to start.

Important:

  • Use the same embedding model for indexing and querying
  • Re-embed if the model changes
  • Consider multilingual embeddings if docs are multilingual

7) Store in a vector database

You need a database that supports:

  • vector similarity search
  • metadata filtering
  • updates/deletes
  • scalability

Common choices:

  • Postgres + pgvector: great for simpler deployments
  • Pinecone: managed, easy
  • Weaviate: flexible, hybrid search
  • Milvus: scalable open-source
  • OpenSearch/Elasticsearch: useful if combining keyword + vector search

Recommended pattern:

  • store vectors
  • store chunk text
  • store metadata
  • store ACL info for filtering

8) Implement retrieval

At query time:

  1. Embed user question
  2. Search vector DB for top-k similar chunks
  3. Apply metadata filters:
    • user permissions
    • document type
    • business unit
    • date range
  4. Optionally re-rank results with a cross-encoder or LLM reranker
  5. Send the best chunks to the LLM

You’ll usually get better results by combining:

  • vector search
  • keyword/BM25 search
  • reranking

This is called hybrid retrieval and is often better than vector-only.


9) Build the answer prompt

Your prompt should:

  • instruct the model to use only retrieved context
  • answer concisely
  • cite sources
  • say when it doesn’t know

Example pattern:

You are a helpful assistant for internal company documents.
Answer the question using only the provided context.
If the context is insufficient, say you don’t know.
Cite sources for each claim.

Context:
[1] ...
[2] ...

Question: ...
Answer:

For citations, include chunk IDs or source URLs and make the model output them inline or in a references section.


10) Handle permissions securely

This is one of the biggest enterprise pitfalls.

You need:

  • document ACLs
  • user identity propagation
  • group membership checks
  • metadata filtering in retrieval

A safe pattern:

  • index each chunk with permitted users/groups
  • at query time, filter retrieval by the current user’s access rights
  • never rely only on the LLM to “not mention” restricted info

Do not fetch all docs and filter after generation. Filter before retrieval.


11) Improve quality with reranking and evaluation

RAG quality is often limited by retrieval, not generation.

Add:

  • hybrid retrieval
  • reranking
  • query rewriting
  • multi-hop retrieval for complex questions
  • context compression if chunks are too large

Evaluate with:

  • answer correctness
  • citation accuracy
  • retrieval recall@k
  • latency
  • permission leakage tests

Make a small test set of real employee questions and compare outputs.


12) Add update and deletion handling

Company docs change often.

You need:

  • document versioning
  • re-ingestion on update
  • delete chunks when documents are removed
  • incremental sync if possible

Good practice:

  • use a stable doc_id
  • store content_hash
  • only re-embed if content changed
  • soft-delete old chunks or version them

13) Build the app interface

Typical UI features:

  • chat
  • source citations
  • “open document” links
  • feedback buttons
  • suggested questions
  • document search mode

Also include:

  • confidence/“I’m not sure” behavior
  • logs for retrieval debugging
  • admin tools to monitor failed ingestion jobs

14) Recommended starter stack

If you want something practical and fast:

Simple enterprise-friendly stack

  • Backend: Python + FastAPI
  • Parsing: unstructured + OCR as needed
  • Embeddings: OpenAI / Cohere / local embedding model
  • Vector DB: Postgres + pgvector
  • LLM: GPT-style model or equivalent
  • Orchestration: Celery / cron / Airflow / Prefect
  • Frontend: Next.js / React

If you want a fast prototype

  • LangChain or LlamaIndex for orchestration
  • pgvector or Pinecone for retrieval
  • a simple chat UI
  • document connector for just one source first, like Google Drive

15) A minimal implementation outline

Pseudo-flow:

# ingest.py
docs = fetch_documents_from_drive()
for doc in docs:
    text = parse_document(doc)
    chunks = chunk_text(text)
    for chunk in chunks:
        vector = embed(chunk.text)
        vector_db.upsert(
            id=chunk.id,
            vector=vector,
            metadata={
                "doc_id": doc.id,
                "title": doc.title,
                "url": doc.url,
                "permissions": doc.permissions
            },
            text=chunk.text
        )
# query.py
def answer_question(user প্রশ্ন, user_permissions):
    qvec = embed(user_question)
    chunks = vector_db.search(
        vector=qvec,
        top_k=5,
        filter={"permissions": {"$in": user_permissions}}
    )
    prompt = build_prompt(chunks, user_question)
    return llm.generate(prompt)

16) Common mistakes to avoid

  • No access control
  • Chunks too large or too small
  • Using only vector search
  • No citations
  • Ignoring OCR/layout quality
  • Not evaluating retrieval
  • Not handling document updates/deletions
  • Letting the LLM answer without enough context
  • Using raw docs with sensitive info without governance

17) If you want to start this week

A good 3-step MVP:

  1. Pick one source: e.g. Google Drive or SharePoint
  2. Build ingestion + pgvector
  3. Create a chat UI with citations and permission filtering

Then improve:

  • better chunking
  • hybrid search
  • reranking
  • feedback loop
  • more connectors

If you want, I can also give you:

  1. a reference architecture diagram,
  2. a Python starter project structure, or
  3. a working FastAPI + pgvector example for company docs.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.