ContextGPT is a company-context chat platform for asking questions against internal files. It combines a React/Vite frontend, a FastAPI backend, Supabase Auth-style app tables plus private Storage, and a vLLM OpenAI-compatible inference server.
The app lets a user sign in, upload company documents, load them as shared context, choose which loaded files are active for a chat session, and stream responses from a local or remote LLM endpoint.
- Email/password signup and login against the app
userstable with JWT access tokens. - Forgot-password flow with time-limited reset tokens stored in Supabase.
- Personal chat sessions with persisted user and assistant messages.
- Streaming chat responses from FastAPI to the frontend.
- vLLM integration through OpenAI-compatible
/v1/chat/completionsand/v1/modelsendpoints. - Company file upload to private Supabase Storage.
- Text extraction is merged into the Load action for uploaded files.
- Automatic metadata sync for files manually added to the Supabase Storage bucket.
- Session-level context file selection, so each chat can use a specific subset of loaded files.
- Basic monitoring for vLLM reachability, active sessions, loaded context tokens, VRAM-style stats, and chart views.
- Podstack/vLLM server commands for running Qwen2.5-7B-Instruct in standard and 128k FP8 modes.
React frontend (project/)
-> FastAPI backend (backend/app/)
-> Supabase Postgres + private Storage
-> vLLM OpenAI-compatible server
The backend owns authentication, file metadata, Storage operations, text extraction, chat persistence, context assembly, and the streaming bridge to vLLM. The frontend provides the signed-in workspace for chat, file management, settings, and monitoring.
-
Create a Supabase project.
-
Open the Supabase SQL Editor and run
backend/supabase/schema.sql. -
Create a private Storage bucket named
company-filesfrom Storage -> New bucket.The backend also tries to create this bucket automatically with the service-role key, but creating it once in the dashboard makes setup explicit.
-
Copy your Project URL, anon key, and service-role key from Project Settings -> API.
-
Create
backend/.envfrombackend/.env.exampleand fill in:
SUPABASE_URL=https://your-project.supabase.co
SUPABASE_ANON_KEY=your-anon-key
SUPABASE_SERVICE_ROLE_KEY=your-service-role-key
SUPABASE_STORAGE_BUCKET=company-files
JWT_SECRET=use-a-long-random-string
FRONTEND_ORIGIN=http://localhost:5173
PASSWORD_RESET_TOKEN_MINUTES=30
VLLM_BASE_URL=http://localhost:8000/v1
VLLM_API_KEY=EMPTY
VLLM_MODEL=Qwen2.5-7B-InstructPassword reset links are generated by the backend from FRONTEND_ORIGIN. This project does not include an SMTP provider yet, so the forgot-password API currently returns the reset URL to the frontend after creating the hashed reset token.
After installing backend dependencies, create the demo login:
cd backend
python3 -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
PYTHONPATH=. python scripts/create_demo_user.pyDemo credentials:
user@company.com
demo
From backend/:
source .venv/bin/activate
uvicorn app.main:app --reload --host 0.0.0.0 --port 8001The backend expects vLLM to be reachable at VLLM_BASE_URL. If your vLLM server is on another port or machine, only change that environment value.
For example, with:
VLLM_BASE_URL=http://localhost:8080/v1the backend sends chat completions to:
http://localhost:8080/v1/chat/completions
From project/:
npm install
cp .env.example .env
npm run devFrontend env:
VITE_API_BASE_URL=http://localhost:8001/apiOpen:
http://localhost:5173
Uploaded files are stored in the private Supabase Storage bucket configured by SUPABASE_STORAGE_BUCKET.
The company_files.file_path column stores the Storage object path, for example:
uploads/3b5f...-LeavePolicy.pdf
defaults/HR_Leave_Policy.md
The backend still stores extracted_text in Postgres because vLLM needs plain text context for chat requests.
If you manually upload files in the Supabase Storage dashboard, refresh the Files page. The backend automatically creates missing company_files metadata rows for Storage objects it finds. Use Load on the Files page before selecting those files in chat.
The Monitoring page shows the current backend view of the inference service and context usage:
- GPU/model label based on vLLM health.
- Total and used VRAM style counters.
- Active chat session count.
- Loaded context token count.
- Tokens/sec, latency, KV cache, and prefix-cache placeholders for richer telemetry.
- Manual refresh and live auto-refresh controls.
The backend monitoring endpoint reads vLLM's /metrics feed, stores recent samples in system_metrics, and combines those live serve counters with app-side session and loaded-context counts. If your vLLM server is not on the same host/port as VLLM_BASE_URL, set VLLM_METRICS_URL explicitly.
These commands run the Qwen2.5-7B-Instruct model behind a vLLM OpenAI-compatible server. Point VLLM_BASE_URL at the server's /v1 base URL.
python --version
pip --versionpip install -U "huggingface_hub[cli]"
mkdir -p /data/models
hf download Qwen/Qwen2.5-7B-Instruct \
--local-dir /data/models/Qwen2.5-7B-Instruct
du -sh /data/models/Qwen2.5-7B-Instructpip install -U vllm
pip install ninja
echo 'export PATH=$HOME/.local/bin:$PATH' >> ~/.bashrc
source ~/.bashrc
python -m vllm.entrypoints.openai.api_server \
--model /data/models/Qwen2.5-7B-Instruct \
--host 0.0.0.0 \
--port 8080 \
--max-model-len 8192 \
--gpu-memory-utilization 0.90 \
--enable-prefix-cachingUse this backend env value when vLLM is running on the same machine:
VLLM_BASE_URL=http://localhost:8080/v1
VLLM_MODEL=/data/models/Qwen2.5-7B-InstructFrom another SSH terminal:
curl http://localhost:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "/data/models/Qwen2.5-7B-Instruct",
"messages": [
{
"role": "user",
"content": "Tell me about leave policy."
}
],
"temperature": 0.8,
"max_tokens": 50
}'For a longer context window, run vLLM with FP8 quantization and YaRN rope scaling:
cd /data/models
source /data/models/.venv/bin/activate
export CUDA_VISIBLE_DEVICES=0
export VLLM_USE_FLASHINFER_SAMPLER=0
python -m vllm.entrypoints.openai.api_server \
--model /data/models/Qwen2.5-7B-Instruct \
--host 0.0.0.0 \
--port 8080 \
--dtype bfloat16 \
--quantization fp8 \
--kv-cache-dtype auto \
--max-model-len 131072 \
--hf-overrides '{"rope_parameters":{"rope_theta":1000000,"rope_type":"yarn","factor":4.0,"original_max_position_embeddings":32768}}' \
--gpu-memory-utilization 0.96 \
--enable-prefix-caching \
--enable-chunked-prefill \
--max-num-seqs 2 \
--max-num-batched-tokens 8196 \
--trust-remote-codeThis mode is intended for long-context experiments. It is more memory intensive and may need GPU-specific tuning.
Backend:
cd backend
source .venv/bin/activate
uvicorn app.main:app --reload --host 0.0.0.0 --port 8001Frontend:
cd project
npm run dev
npm run build
npm run lint