Skip to content
isiahpwilliamsPublic

About

LLM-powered chat bot designed to answer almost any medical question

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

15 Commits

Folders and files

Repository files navigation

LLMed

A clinical reference assistant that answers questions from a curated corpus of medical literature, with every claim traceable to the passage it came from.

Live demo → https://ll-med-seven.vercel.app

Decision support from published sources, not a clinical decision. This is a portfolio project and is not intended for patient care.


What it does

Ask a clinical question and get a grounded, cited answer streamed token by token. Each [n] marker is clickable and opens the exact retrieved passage, with its document, section, page, publication year, and relevance score.

Questions the corpus cannot answer are refused rather than improvised — a confidently wrong dose is the worst failure this system can produce.

530 documents · 14,400 passages · CDC guidelines, FDA drug labels, open-access reviews

Architecture

flowchart LR
    subgraph Client["Vercel"]
        UI[React 19 + TypeScript]
    end
    subgraph API["Cloud Run"]
        F[FastAPI]
    end
    subgraph Data["Neon Postgres 18"]
        V[(pgvector HNSW)]
        T[(tsvector GIN)]
        C[(conversations)]
    end

    UI -->|SSE| F
    F --> V
    F --> T
    F --> C
    F -->|embed + rerank| Cohere
    F -->|generate| Groq
Loading

Query path

question
  → rewrite into a standalone query (multi-turn follow-ups)
  → dense search  (pgvector, cosine, HNSW)     top 40
  → lexical search (tsvector, ts_rank_cd)      top 40
  → Reciprocal Rank Fusion
  → Cohere rerank-v3.5                         top 6
  → relevance floor: below 0.05, refuse without calling the LLM
  → prompt with numbered sources → Groq → SSE stream

Sources are emitted before the first token, so the citation panel renders while the answer is still streaming.


Measured results

A four-way evaluation over 30 questions — 25 answerable, 5 deliberately outside the corpus — isolating each stage's contribution:

mode recall@5 MRR
dense only 1.00 0.896
lexical only 0.80 0.657
hybrid (RRF) 0.88 0.880
hybrid + local cross-encoder 0.88 0.833
hybrid + Cohere rerank 0.96 0.920

The local cross-encoder scored worse than no reranking at all — it demoted correct documents. The no-API-key fallback is therefore plain RRF ordering, not a reranker that makes results worse.

MRR matters more than recall here because only the top 6 chunks reach the LLM: a correct passage at rank 1 versus rank 4 changes the answer.

On significance: at n=25 a single question is worth 0.04 recall, so dense-only versus hybrid+Cohere is not a statistically separable gap. The robust findings are the large ones — lexical-only is clearly weakest, and local reranking actively hurts.

Grounding

Rerank scores separate cleanly, which is what makes a threshold defensible:

in-corpus questions      0.44 – 0.89
out-of-corpus questions  0.017 – 0.022
floor                    0.05

Below the floor the chain returns an explicit refusal and never calls the LLM.


Design decisions

One database, not three. The conventional build is a vector store, a separate lexical index, and a third store for chat history. All three live in Postgres here: pgvector for dense retrieval, tsvector for lexical, ordinary tables for conversations. Deleting a document is one transactional DELETE rather than three non-atomic writes that can leave a lexical index pointing at rows that no longer exist.

Hybrid retrieval, because embeddings blur clinical text. Dense vectors treat drug names differing by one morpheme as near-identical and lose ICD codes and dose units. Lexical search catches exact tokens. RRF fuses them by rank position, so two incomparable score scales never need normalizing.

Lexical search uses OR, not AND. plainto_tsquery ANDs every term, so "what is the recommended dose of finasteride for hair loss" demanded a passage containing all five terms and returned 1 result out of 40. Postgres normalizes and drops stopwords, then & is rewritten to |; ts_rank_cd still ranks by term coverage and proximity.

Embeddings via API, not locally. Local PyTorch embeddings meant a 3 GB image, 1.5 GB RAM, and a 30–60s cold start — unusable on a free tier. The Cohere API brought it to 431 MB and a 0.59s cold start. For a link someone clicks once, cold start is the product.

Citations are a feature, not decoration. [n] markers are rewritten during markdown parsing, as a rehype pass over the syntax tree. Rewriting the markdown first would corrupt link syntax and code spans; rewriting rendered HTML would require dangerouslySetInnerHTML.

No authentication. A sign-in wall defeats a public demo. Per-visitor history uses a browser-generated UUID in localStorage; abuse is handled by rate limiting, which is the actual threat.

Corpus selected by prevalence. Real questions cluster on common topics, so drugs and clinical topics come from curated prevalence lists rather than whatever an API returns first.


Corpus

Source License Role
CDC (via Europe PMC) Public domain Treatment guidelines, MMWR
openFDA Public domain Drug labels — dosing, contraindications
Europe PMC Open access Clinical review articles

Documents are never committed to this repository. Everything is fetched from public APIs at ingest time, so nothing is redistributed and the corpus is reproducible from a clean database.

Superseded editions are handled explicitly: documents carry an effective_date, same-title documents keep only the newest, and every citation displays its year so a stale guideline is visible as stale.


Running locally

Requires a container runtime and uv.

brew install colima docker docker-compose docker-buildx && colima start
cp .env.example .env      # then add COHERE_API_KEY and GROQ_API_KEY
docker compose up --build

The API is on :8000, the UI on :5173.

Ingest a corpus:

cd backend
uv run python scripts/ingest.py migrate
uv run python scripts/ingest.py pmc --limit 44
uv run python scripts/ingest.py openfda --curated
uv run python scripts/ingest.py pmc --curated-topics --topics all --limit 6

Verify every configured value by actually exercising it:

uv run python scripts/check_env.py

Measure retrieval:

uv run python scripts/eval_retrieval.py --k 5 --cohere

Developing without burning API quota

Cohere trial keys are capped at 100k tokens/minute and 1,000 calls/month. For local iteration, use local embeddings instead:

uv sync --group local-models
# .env: EMBEDDING_PROVIDER=local, EMBEDDING_DIM=384, RERANK_PROVIDER=none

Local and production embeddings have different dimensions (384 vs 1024), so they require separate databases — the schema is built for whichever is configured.


Deployment

See DEPLOY.md. In short: the container goes to Cloud Run, Postgres is Neon, the static frontend is Vercel, and secrets live in Google Secret Manager. Total cost is $0/month on free tiers.


Known limits

~10 queries/minute globally. Cohere trial keys allow 10 rerank calls per minute across all users, not per user. The per-IP chat limit is 5/minute so one visitor cannot starve the rest.

Cold starts. Cloud Run scales to zero. The frontend fires a health ping on mount so the backend wakes while the visitor is reading, which usually hides it; a banner covers the case where it does not.

Storage. 241 MB of Neon's ~0.5 GB free tier — room to roughly double the corpus, not triple it.

Evaluation size. 25 answerable questions is too few to separate close configurations. Expanding it is the highest-value next step.


Stack

Python · TypeScript · FastAPI · psycopg3 · PostgreSQL/pgvector · LangChain · Cohere · Groq · React · Zustand · Tailwind · Docker · Cloud Run · Neon · Vercel

About

LLM-powered chat bot designed to answer almost any medical question

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages