A Retrieval-Augmented Generation system designed around data sovereignty and privacy. Every component of the pipeline — LLM inference, embeddings, reranking — runs locally on your hardware through Ollama and HuggingFace models. No document content, no query, no conversation ever leaves your machine.
This matters when you're working with internal company documents, personal data, legal files, or anything you wouldn't paste into a third-party API. Traditional RAG setups send your private chunks to external LLM providers for embedding and generation. This system doesn't.
The architecture combines hybrid search, multi-query expansion, and cross-encoder reranking to maintain retrieval quality without relying on cloud APIs. Cloud providers (Gemini, OpenRouter) are available as optional fallbacks, but the default configuration is fully local and self-contained.
graph TD
subgraph Ingestion
A[Documents .txt] --> B(Recursive Splitter)
B --> C{Vectorization}
C --> D[Dense Embedding: all-MiniLM]
C --> E[Sparse Embedding: BM25]
D --> F[(Qdrant Vector Store)]
E --> F
end
subgraph Retrieval & Generation
G[User Query] --> H{Conversational Memory}
H --> I[Condense Question]
I --> J[Multi-Query Expansion]
J --> K[Hybrid Search in Qdrant]
K --> L[Candidate Documents]
L --> M[FlashRank Reranker]
M --> N[Top-N Context]
N --> O[Local LLM: Llama 3.2]
O --> P[Final Answer]
end
Most RAG implementations rely on external APIs (OpenAI, Anthropic, Google) for embeddings and generation. That means every document chunk you index and every question you ask gets sent to third-party servers. For many use cases — corporate knowledge bases, HR policies, legal contracts, medical records, financial data — that's a non-starter.
This system keeps everything on-premise:
- LLM inference runs on Ollama (Llama 3.2) — no API calls for generation
- Embeddings are computed locally via HuggingFace (
paraphrase-multilingual-MiniLM-L12-v2) — your documents are never sent to external embedding services - Reranking uses a local ONNX model (FlashRank TinyBERT) — no cloud cross-encoders
- Vector storage can run locally via Docker (Qdrant) — your vectors stay on your infrastructure
The only network traffic is between your machine and your own Qdrant instance.
- Conversational Memory — reformulates follow-up questions into standalone queries using chat history
- Multi-Query Expansion — generates query variations via LLM to improve retrieval recall
- Hybrid Search — combines dense (semantic) and sparse (BM25) retrieval in Qdrant
- FlashRank Reranking — re-orders candidates with a lightweight cross-encoder (TinyBERT)
- RAGAS Evaluation — automated quality scoring with Faithfulness and Context Precision metrics
- Multilingual — uses cross-language embeddings, so you can query in any language regardless of document language
- Dual interface — CLI (
app.py) and web UI (streamlit_app.py)
| Component | Tool |
|---|---|
| Framework | LangChain |
| Vector DB | Qdrant |
| LLM | Ollama (Llama 3.2 gen, Llama 3.1 8B eval) |
| Embeddings | HuggingFace paraphrase-multilingual-MiniLM-L12-v2 + FastEmbed BM25 |
| Reranker | FlashRank |
| Evaluation | RAGAS |
| Package Manager | uv |
- Python >= 3.12
- uv installed
- Ollama running locally (
ollama serve) - A Qdrant instance (cloud or local via Docker)
-
Install dependencies:
uv sync
-
Copy
.env.exampleto.envand fill in your values:QDRANT_URLandQDRANT_API_KEYare required- LLM provider keys are optional if using Ollama locally
-
Pull the local models:
ollama pull llama3.2 # generation ollama pull llama3.1:8b # evaluation judge
-
Verify Ollama is running:
ollama list
uv run python src/ingest.pyuv run streamlit run src/streamlit_app.pyuv run python src/app.pyuv run python src/evaluate.pyThe system uses a "Student-Teacher" approach:
- Student (
llama3.2): generates answers during normal usage - Teacher (
llama3.1:8b): scores the student's responses during evaluation
Metrics:
- Faithfulness: is the answer grounded in the retrieved documents?
- Context Precision: were the retrieved documents actually relevant?
Results are saved as JSON in eval_results/.
| Problem | Solution |
|---|---|
ConnectionRefusedError on Ollama |
Make sure ollama serve is running |
| Qdrant connection fails | Check QDRANT_URL and QDRANT_API_KEY in .env |
| Slow first query | Normal — models are loaded into memory on first use |
| Out of memory | Use smaller models or reduce RETRIEVER_K in rag_pipeline.py |
MIT License — see the LICENSE file for details.