A clinical reference assistant that answers questions from a curated corpus of medical literature, with every claim traceable to the passage it came from.
Live demo → https://ll-med-seven.vercel.app
Decision support from published sources, not a clinical decision. This is a portfolio project and is not intended for patient care.
Ask a clinical question and get a grounded, cited answer streamed token by
token. Each [n] marker is clickable and opens the exact retrieved passage,
with its document, section, page, publication year, and relevance score.
Questions the corpus cannot answer are refused rather than improvised — a confidently wrong dose is the worst failure this system can produce.
530 documents · 14,400 passages · CDC guidelines, FDA drug labels, open-access reviews
flowchart LR
subgraph Client["Vercel"]
UI[React 19 + TypeScript]
end
subgraph API["Cloud Run"]
F[FastAPI]
end
subgraph Data["Neon Postgres 18"]
V[(pgvector HNSW)]
T[(tsvector GIN)]
C[(conversations)]
end
UI -->|SSE| F
F --> V
F --> T
F --> C
F -->|embed + rerank| Cohere
F -->|generate| Groq
question
→ rewrite into a standalone query (multi-turn follow-ups)
→ dense search (pgvector, cosine, HNSW) top 40
→ lexical search (tsvector, ts_rank_cd) top 40
→ Reciprocal Rank Fusion
→ Cohere rerank-v3.5 top 6
→ relevance floor: below 0.05, refuse without calling the LLM
→ prompt with numbered sources → Groq → SSE stream
Sources are emitted before the first token, so the citation panel renders while the answer is still streaming.
A four-way evaluation over 30 questions — 25 answerable, 5 deliberately outside the corpus — isolating each stage's contribution:
| mode | recall@5 | MRR |
|---|---|---|
| dense only | 1.00 | 0.896 |
| lexical only | 0.80 | 0.657 |
| hybrid (RRF) | 0.88 | 0.880 |
| hybrid + local cross-encoder | 0.88 | 0.833 |
| hybrid + Cohere rerank | 0.96 | 0.920 |
The local cross-encoder scored worse than no reranking at all — it demoted correct documents. The no-API-key fallback is therefore plain RRF ordering, not a reranker that makes results worse.
MRR matters more than recall here because only the top 6 chunks reach the LLM: a correct passage at rank 1 versus rank 4 changes the answer.
On significance: at n=25 a single question is worth 0.04 recall, so dense-only versus hybrid+Cohere is not a statistically separable gap. The robust findings are the large ones — lexical-only is clearly weakest, and local reranking actively hurts.
Rerank scores separate cleanly, which is what makes a threshold defensible:
in-corpus questions 0.44 – 0.89
out-of-corpus questions 0.017 – 0.022
floor 0.05
Below the floor the chain returns an explicit refusal and never calls the LLM.
One database, not three. The conventional build is a vector store, a
separate lexical index, and a third store for chat history. All three live in
Postgres here: pgvector for dense retrieval, tsvector for lexical, ordinary
tables for conversations. Deleting a document is one transactional DELETE
rather than three non-atomic writes that can leave a lexical index pointing at
rows that no longer exist.
Hybrid retrieval, because embeddings blur clinical text. Dense vectors treat drug names differing by one morpheme as near-identical and lose ICD codes and dose units. Lexical search catches exact tokens. RRF fuses them by rank position, so two incomparable score scales never need normalizing.
Lexical search uses OR, not AND. plainto_tsquery ANDs every term, so
"what is the recommended dose of finasteride for hair loss" demanded a passage
containing all five terms and returned 1 result out of 40. Postgres normalizes
and drops stopwords, then & is rewritten to |; ts_rank_cd still ranks by
term coverage and proximity.
Embeddings via API, not locally. Local PyTorch embeddings meant a 3 GB image, 1.5 GB RAM, and a 30–60s cold start — unusable on a free tier. The Cohere API brought it to 431 MB and a 0.59s cold start. For a link someone clicks once, cold start is the product.
Citations are a feature, not decoration. [n] markers are rewritten during
markdown parsing, as a rehype pass over the syntax tree. Rewriting the markdown
first would corrupt link syntax and code spans; rewriting rendered HTML would
require dangerouslySetInnerHTML.
No authentication. A sign-in wall defeats a public demo. Per-visitor
history uses a browser-generated UUID in localStorage; abuse is handled by
rate limiting, which is the actual threat.
Corpus selected by prevalence. Real questions cluster on common topics, so drugs and clinical topics come from curated prevalence lists rather than whatever an API returns first.
| Source | License | Role |
|---|---|---|
| CDC (via Europe PMC) | Public domain | Treatment guidelines, MMWR |
| openFDA | Public domain | Drug labels — dosing, contraindications |
| Europe PMC | Open access | Clinical review articles |
Documents are never committed to this repository. Everything is fetched from public APIs at ingest time, so nothing is redistributed and the corpus is reproducible from a clean database.
Superseded editions are handled explicitly: documents carry an
effective_date, same-title documents keep only the newest, and every citation
displays its year so a stale guideline is visible as stale.
Requires a container runtime and uv.
brew install colima docker docker-compose docker-buildx && colima startcp .env.example .env # then add COHERE_API_KEY and GROQ_API_KEY
docker compose up --buildThe API is on :8000, the UI on :5173.
Ingest a corpus:
cd backend
uv run python scripts/ingest.py migrate
uv run python scripts/ingest.py pmc --limit 44
uv run python scripts/ingest.py openfda --curated
uv run python scripts/ingest.py pmc --curated-topics --topics all --limit 6Verify every configured value by actually exercising it:
uv run python scripts/check_env.pyMeasure retrieval:
uv run python scripts/eval_retrieval.py --k 5 --cohereCohere trial keys are capped at 100k tokens/minute and 1,000 calls/month. For local iteration, use local embeddings instead:
uv sync --group local-models
# .env: EMBEDDING_PROVIDER=local, EMBEDDING_DIM=384, RERANK_PROVIDER=noneLocal and production embeddings have different dimensions (384 vs 1024), so they require separate databases — the schema is built for whichever is configured.
See DEPLOY.md. In short: the container goes to Cloud Run, Postgres is Neon, the static frontend is Vercel, and secrets live in Google Secret Manager. Total cost is $0/month on free tiers.
~10 queries/minute globally. Cohere trial keys allow 10 rerank calls per minute across all users, not per user. The per-IP chat limit is 5/minute so one visitor cannot starve the rest.
Cold starts. Cloud Run scales to zero. The frontend fires a health ping on mount so the backend wakes while the visitor is reading, which usually hides it; a banner covers the case where it does not.
Storage. 241 MB of Neon's ~0.5 GB free tier — room to roughly double the corpus, not triple it.
Evaluation size. 25 answerable questions is too few to separate close configurations. Expanding it is the highest-value next step.
Python · TypeScript · FastAPI · psycopg3 · PostgreSQL/pgvector · LangChain · Cohere · Groq · React · Zustand · Tailwind · Docker · Cloud Run · Neon · Vercel