An AI-powered PDF RAG (Retrieval-Augmented Generation) Chatbot built using LangChain, Mistral AI, FAISS, Streamlit, HuggingFace Embeddings, and SQLite.
This project allows users to upload PDF documents and ask questions directly from the uploaded content. The chatbot understands document context, supports follow-up questions, maintains chat history, and provides a smooth conversational experience similar to ChatGPT.
Traditional chatbots answer questions using general knowledge. This project uses the RAG (Retrieval-Augmented Generation) architecture, which means the chatbot first retrieves relevant information from uploaded PDF documents and then generates accurate answers using an LLM (Large Language Model).
This allows the chatbot to:
- answer document-specific questions
- understand context
- support follow-up questions
- maintain conversation memory
- provide more accurate and contextual responses
The application is built with a clean Streamlit UI and stores chat history permanently using SQLite database.
Users can upload their own PDF documents and chat with them.
The chatbot remembers previous questions and supports follow-up conversations.
Users can create and manage multiple chat conversations.
All chats are stored using SQLite database and remain available even after refreshing the browser.
Answers are generated with typing effect similar to ChatGPT.
Instead of technical Python errors, users see understandable messages.
Users can permanently delete specific chat conversations.
Complete chat history can be cleared from settings.
Clean and simple Streamlit-based interface.
Displays Good Morning, Good Afternoon, or Good Evening according to system time.
| Technology | Purpose |
|---|---|
| Python | Backend Programming |
| Streamlit | Web Application UI |
| RAG (Retrieval-Augmented Generation) | AI Retrieval + Response Architecture |
| LangChain | RAG Pipeline and Conversational Chain |
| Mistral AI | Large Language Model (LLM) |
| FAISS | Vector Database for Similarity Search |
| HuggingFace Embeddings | Text Embedding Generation |
| SQLite | Persistent Chat Storage |
| PyPDF | PDF Text Extraction |
| Sentence Transformers | Semantic Embedding Model |
This project follows the Retrieval-Augmented Generation pipeline.
The user uploads a PDF document using the Streamlit interface.
LangChain's PyPDFLoader extracts text content from the uploaded PDF.
Large documents are divided into smaller chunks using:
RecursiveCharacterTextSplitterThis improves retrieval quality and embedding generation.
Example:
- Chunk Size = 500
- Chunk Overlap = 50
Each text chunk is converted into vector embeddings using:
sentence-transformers/all-MiniLM-L6-v2These embeddings capture semantic meaning of text.
FAISS vector database stores all embeddings for fast similarity search.
When user asks a question:
- FAISS searches most relevant chunks
- retrieves related information from document
Retrieved chunks + user question are sent to Mistral AI.
Mistral AI generates a contextual answer using:
- uploaded document data
- previous conversation history
The answer appears word-by-word with typing effect for better user experience.
Each chat conversation is permanently stored inside SQLite database.
This allows:
- reopening old chats
- chat persistence after refresh
- multi-chat support
PDF-RAG-Chatbot/
│
├── app.py
├── main.py
├── database.py
├── requirements.txt
├── README.md
├── .gitignore
│
├── domestic_travel.pdf
├── foreign_travel.pdf
│
├── venv/git clone https://github.com/Suryanandankumar2003/PDF-RAG-Chatbot.gitcd PDF-RAG-Chatbotpython -m venv venvpython3 -m venv venvvenv\Scripts\activatesource venv/bin/activateAfter activation you will see:
(venv)in terminal.
pip install -r requirements.txtThis installs:
- LangChain
- Streamlit
- FAISS
- HuggingFace
- Mistral AI SDK
- PyPDF
- SQLite dependencies
Create a file named:
.envAdd your Mistral API key:
mistral_api_key=YOUR_API_KEYYou can get API key from:
streamlit run app.pyStreamlit automatically opens browser.
Usually:
http://localhost:8501
Upload any PDF document using upload button.
Example:
- What is travel reimbursement policy?
- What are leave rules?
- Explain foreign travel process.
The chatbot remembers previous context.
Example:
User:
What is leave policy?
Follow-Up:
What about sick leave?
SQLite database stores:
- chat names
- user messages
- assistant responses
This enables:
- persistent history
- multiple chats
- reopening old chats
- deleting chats
Database file:
chat_history.dbThe application handles:
- invalid PDFs
- empty answers
- API failures
- unexpected errors
Instead of technical errors, users see understandable messages.
- Source citations
- Authentication system
- Docker support
- Cloud deployment
- Voice assistant
- Chat export
- Persistent vector database
- Dark/Light themes
Suryanandan Kumar
GitHub: https://github.com/Suryanandankumar2003
This project is licensed under the MIT License.



