Run an LLM Locally to Interact with Your Documents: Complete Ollama & Open WebUI Guide
In this article
- The Dual-Model Local RAG Architecture
- 1. Hardware Sizing & Model Selection by Tier
- 2. Step-by-Step Installation & Core Setup
- Step 1: Install Ollama & Pull Models
- Step 2: Install and Launch Open WebUI
- 3. Configuring Document Chunking & Embedding Engines
- Recommended Chunk Size & Overlap Parameters
- 4. Ingesting Documents & Creating Custom Model Profiles
- Step 1: Create a Knowledge Collection
- Step 2: Bind Knowledge to a Custom Model Profile
- Grounded System Prompt Template
- 5. Querying & Operational Best Practices
- Operational Tips
- Summary & Next Steps
- Related Local LLM Architecture Guides
Interacting with internal documents using cloud-hosted AI APIs creates unavoidable privacy and compliance liabilities. When analyzing proprietary source code, patient health records, corporate financial ledgers, legal contracts, or personal journals, uploading raw text to third-party endpoints risks data leakage, training ingestion, and unpredictable API billing.
The solution is a self-contained, 100% private local Retrieval-Augmented Generation (RAG) stack. By pairing Ollama (for local model orchestration and embeddings) with Open WebUI (for an enterprise-grade ChatGPT-like workspace), you can index, chunk, embed, and query large document collections entirely on your local machine with zero external network calls.
In this step-by-step engineering tutorial, we configure an offline document retrieval pipeline: analyzing dual-model embedding-versus-generation mechanics, tuning token chunking and overlap thresholds across hardware tiers, binding knowledge collections to custom model profiles, and engineering grounded system prompts that eliminate hallucinations.
The Dual-Model Local RAG Architecture
A common misconception when deploying document Q&A systems is assuming a single large language model handles both document indexing and conversation. In practice, a production RAG pipeline requires two distinct models working in tandem:
- The Embedding Model (
nomic-embed-text): Converts parsed document paragraphs into high-dimensional numerical vectors (embeddings). When you ask a question, the embedding model converts your query into a vector and performs fast cosine similarity math across your vector database to extract the most relevant text chunks. - The Generative Chat Model (
llama3,gemma3, orqwen2.5): Receives your prompt along with the retrieved document chunks injected into its context window, synthesizing a coherent, factual response strictly grounded in the source text.
flowchart TD
subgraph Ingestion Pipeline
Docs[Raw Documents\nPDF, Markdown, Text] --> Chunker[Text Splitter\nChunk Size + Overlap]
Chunker --> Embedder["Embedding Model\n(nomic-embed-text)"]
Embedder --> VectorDB[(Local Vector Store\nOpen WebUI DB)]
end
subgraph Query & Generation Pipeline
UserQuery[User Question] --> QueryEmbed["Embed Query\n(nomic-embed-text)"]
QueryEmbed --> Similarity[Cosine Similarity Search]
VectorDB --> Similarity
Similarity --> RetrievedChunks[Top-K Document Chunks]
RetrievedChunks --> PromptAssembly[System Prompt + Context + Query]
UserQuery --> PromptAssembly
PromptAssembly --> Generator["Local Chat LLM\n(Ollama GPU Inference)"]
Generator --> FinalAnswer[Synthesized Grounded Answer]
end
1. Hardware Sizing & Model Selection by Tier
Before pulling weights, match your machine's physical RAM and GPU VRAM with appropriate model scales:
| Hardware Tier | System Specifications | Recommended Chat Model | Embedding Model | Target Performance |
|---|---|---|---|---|
| Tier 1 (Constrained) | ≤ 8 GB RAM, CPU-only or integrated graphics | gemma3:1b / qwen2.5:1.5b / llama3.2:1b |
nomic-embed-text |
15–25 tokens/sec (CPU-friendly, low memory footprint) |
| Tier 2 (Mid-Range) | 16 GB RAM, 4–6 GB VRAM (RTX 3060/4050 or M1/M2 Mac) | gemma3:4b / qwen2.5:7b / mistral:7b |
nomic-embed-text |
35–60 tokens/sec (Balanced context retention) |
| Tier 3 (Comfortable) | ≥ 16 GB RAM, 8–12 GB VRAM (RTX 3080/4070 or 32GB Mac) | llama3.3:8b / qwen2.5:14b |
nomic-embed-text |
60–90 tokens/sec (High semantic accuracy) |
| Tier 4 (Workstation) | ≥ 32 GB RAM, 16–24 GB+ VRAM (RTX 3090/4090 / 64GB+ Mac) | qwen2.5:32b / command-r:35b |
bge-m3 / nomic-embed-text |
Extreme reasoning depth and complex multi-hop synthesis |
2. Step-by-Step Installation & Core Setup
Step 1: Install Ollama & Pull Models
Install Ollama from the official installer or package manager. Once running, open your terminal and pull both the embedding model and your selected generation model:
# 1. Pull the high-performance embedding model
ollama pull nomic-embed-text
# 2. Pull your chosen conversational model (e.g., Gemma 3 4B or Llama 3)
ollama pull gemma3:4b
# 3. Verify local models are registered and available
ollama list
Ollama automatically detects GPU acceleration (CUDA on Windows/Linux or Metal on macOS) and exposes an OpenAI-compatible REST server locally on http://localhost:11434.
Step 2: Install and Launch Open WebUI
Open WebUI is an extensible, self-hosted web interface that includes native document parsing, vector indexing, and memory persistence out of the box.
Install Open WebUI via Python's package manager:
# Recommended: Create a clean virtual environment
python -m venv openwebui_env
source openwebui_env/bin/activate # On Windows: .\openwebui_env\Scripts\activate
# Install Open WebUI
pip install open-webui
# Launch the local service
open-webui serve
Once initialized, navigate to http://localhost:8080/ in your web browser. Create an initial admin account (all account credentials and tokens remain stored in your local SQLite/Chroma database).
3. Configuring Document Chunking & Embedding Engines
To prepare Open WebUI for processing local files, configure its ingestion pipeline through the administrative settings:
- Click your profile avatar in the bottom-left corner and navigate to Admin Panel → Settings → Documents.
- Embedding Model Engine: Select Ollama from the dropdown menu (leave the API Key field empty).
- Embedding Model: Enter
nomic-embed-text. - Retrieval Mode: Under the Retrieval section, enable Full Context Mode or configure top-k retrieval (typically
k=4tok=6).
Recommended Chunk Size & Overlap Parameters
Documents cannot be ingested as massive single blobs; they must be parsed into smaller, overlapping segments before vectorization.
- Chunk Size: The maximum number of tokens in a single segment. Smaller chunks yield faster, more pinpoint retrieval; larger chunks preserve surrounding paragraph nuance.
- Chunk Overlap: The number of shared tokens between consecutive chunks. Overlap prevents critical context from being split across chunk boundaries.
| Document Type / Scenario | Chunk Size (Tokens) | Overlap (%) | Engineering Rationale |
|---|---|---|---|
| Dense Technical PDFs / Legal Contracts | 384 – 512 | 15% – 20% | Keeps complete clauses, technical definitions, and code blocks intact without truncating logical arguments. |
| General Business Documentation / Manuals | 256 – 384 | 15% – 20% | The optimal operational balance between vector search speed and semantic clarity. |
| Short Notes, Meeting Minutes & Support Tickets | 128 – 256 | 10% – 15% | Granular text blocks prevent unrelated notes from polluting the context window. |
| Resource-Constrained Systems (≤ 8 GB RAM) | 128 – 256 | 10% – 15% | Minimizes memory allocation during document parsing and vector search. |
[!IMPORTANT] Re-indexing Rule: Once documents are uploaded and indexed, modifying your chunk size or overlap settings does not retroactively update existing files. You must delete the collection and re-upload documents if you adjust chunking parameters.
4. Ingesting Documents & Creating Custom Model Profiles
Step 1: Create a Knowledge Collection
- In Open WebUI, navigate to Workspace → Knowledge.
- Click Add Knowledge Base, provide a descriptive name (e.g.,
Financial-Reports-2026orEngineering-Docs), and click Create. - Drag and drop your target files (PDF, DOCX, TXT, Markdown, CSV).
- Monitor the processing progress. The background worker parses the text, calculates chunk splits, and calls
nomic-embed-textvia Ollama to store embeddings in the local vector index.
Step 2: Bind Knowledge to a Custom Model Profile
Rather than manually tagging documents with # in every single chat turn, bind your Knowledge Collection directly to a custom model persona:
- Navigate to Workspace → Models and click Create a Model.
- Model Name: Give your persona a descriptive name (e.g.,
DocAnalyst-Gemma). - Base Model: Select your downloaded model (e.g.,
gemma3:4b). - Knowledge: Attach the Knowledge Collection you created in Step 1.
- System Prompt: Define strict grounding boundaries to prevent extrapolation.
Grounded System Prompt Template
You are an expert, meticulous document intelligence assistant analyzing proprietary internal documents.
Rules:
1. Ground all answers strictly in the provided document context.
2. If the answer cannot be determined directly from the retrieved context, explicitly state: "The provided documents do not contain sufficient information to answer this question." Do not speculate or extrapolate.
3. When providing facts, data points, or quotes, cite the specific document title and section where the information originated.
4. Maintain a direct, analytical tone without conversational filler.
By enforcing these constraints in the custom profile, the model reliably resists hallucinations and provides trustworthy citations for enterprise auditing.
5. Querying & Operational Best Practices
Open a new conversation, select your custom model from the top dropdown, and begin querying your documentation:
User: What were the key infrastructure bottlenecks identified in the Q3 retrospective?
Assistant: Based on the "Q3-Infrastructure-Retrospective.md" document (Section 4.2), three primary bottlenecks were identified:
1. High Redis cache evictions during peak batch indexing runs.
2. PCIe bus saturation when offloading unquantized 70B weights.
3. Insufficient KV cache allocation during multi-turn agent loops.
Operational Tips
- 5-Minute Streaming Timeout: By default, web browsers close idle server-sent event (SSE) streams after 5 minutes. If querying a massive collection on a slower CPU causes generation to exceed 5 minutes, the interface may pause visibly. Simply refresh the page and click Continue Response to stream the completed generation from disk.
- Monitoring Extraction Errors: If a specific PDF fails to process during upload, inspect your terminal running
open-webui serve. Complex scanned PDFs without OCR layers require pre-processing with a local tool likepdfplumberorpandocbefore ingestion.
Summary & Next Steps
With Ollama managing local GPU inference and Open WebUI orchestrating vector embeddings via nomic-embed-text, you now possess a sovereign, air-gapped document intelligence platform. No corporate data leaves your machine, no cloud API tokens are billed, and queries execute with minimal latency.
Related Local LLM Architecture Guides
- Self-Hosting Your First LLM: Hardware & Deployment Playbook – Detailed VRAM calculations and cloud GPU cost analysis.
- Run LLMs Locally: 6 Practical Methods – Comparing Ollama, vLLM, LM Studio, Jan, llama.cpp, and llamafile.
- Local RAG Tutorial with Ollama & LangChain – Programmatic Python RAG pipeline construction.
- GPU Cloud Pricing Tool – Sizing hardware for large-scale multi-user RAG clusters.
Related Guides
740M EmbeddingGemma 2 Maps Text, Image, Video and Audio
Google DeepMind's Apache 2.0 EmbeddingGemma 2 maps text, code, images, video and audio into one 768-dimension space, scaling from 270M to 740M parameters.
Run LLMs Locally: 6 Practical Methods (Ollama, LM Studio, vLLM, Jan, llama.cpp, and llamafile)
Master 6 proven frameworks to host and run LLMs locally across Windows, macOS, and Linux with GPU acceleration, from one-click GUIs to high-throughput production servers.
Perplexity's 9B Contextual Embedder Beats voyage-context-4 by 14.4 Points
Perplexity releases pplx-embed-v2-context-9b-preview, an open-weights 9B RAG embedder that scores 45.5% Answer Recall@10 and cuts storage 8× vs float32.