Systematic Agent for Meta-analysis
A multi-agent LLM system for automating systematic reviews and meta-analyses
9 AI agents · Multi-database search · R/metafor statistics · Manuscript generation · Explainability report
Note: The implementation of SAM is maintained in a private repository. This repository is a documentation-only overview of the system's architecture and results.
SAM (Systematic Agent for Meta-analysis) is a multi-agent pipeline that automates the full systematic review workflow, from protocol design to manuscript generation. Give it a clinical research question and it will search databases, screen papers, retrieve full texts, extract data, assess quality, run meta-analysis, and write a manuscript.
Developed as part of the Trabajo de Fin de Grado (TFG) at Universidad Pontificia Comillas (ICAI), Engineering Mathematics (iMAT) program.
- 9 specialized agents working sequentially (A1→A9)
- 4 search databases: PubMed, Europe PMC, Semantic Scholar, Google Scholar
- Multi-source full-text retrieval cascade with graceful degradation to abstract-only screening
- 3 quality assessment tools: RoB 2.0 (RCTs), ROBINS-I (observational), JBI (qualitative)
- R/metafor statistical synthesis: REML, Egger's test, trim-and-fill, Rosenthal's Fail-Safe N, leave-one-out, subgroups, meta-regression, GRADE; no LLM for statistics
- RAG-based manuscript generation with FAISS + MiniLM-L6-v2
- Explainability & traceability report (A9): per-paper Excel trace + explainability report in Spanish
- Model-agnostic: auto-detects chat templates, works with any HuggingFace model
- Per-agent LLM parameters (temperature, max tokens, repetition penalty)
- 991 automated tests (879 passing; the rest skipped or xfail for deferred features / GPU-only paths) covering all 9 agents plus end-to-end pipeline
| Agent | Role | Method |
|---|---|---|
| A1 | Protocol Design | PICO/SPIDER/PECO/PCC framework detection, criteria generation, PubMed query |
| A2 | Literature Search | Multi-database search (PubMed, Europe PMC, Semantic Scholar, Google Scholar) + semantic deduplication |
| A3 | Title/Abstract Screening | Ensemble voting (N votes/paper, configurable majority threshold) |
| A4 | Full-Text Retrieval & Screening | Multi-source retrieval cascade + LLM full-text screening |
| A5 | Data Extraction | 3 focused prompts per paper (study characteristics, intervention, outcomes) |
| A6 | Quality Assessment | RoB 2.0 / ROBINS-I / JBI Critical Appraisal with weighted scores |
| A7 | Statistical Synthesis | R/metafor: random-effects REML, heterogeneity, Egger's test, trim-and-fill, Rosenthal's Fail-Safe N, leave-one-out, subgroups, meta-regression, GRADE, R script export |
| A8 | Manuscript Generation | RAG-based writing, section-specific retrieval, population descriptives, LaTeX/PDF output |
| A9 | Explainability & Traceability | Per-paper Excel trace + explainability report in Spanish, no GPU |
- Single GPU: Only one model loaded at a time, VRAM freed between agents
- Conservative fallback: Parse errors → EXCLUDE (A3/A4) or default quality weights (A6)
- No LLM for statistics: A7 uses R/metafor directly via rpy2
- Model-agnostic: Chat template auto-detection + token-level output slicing (works with Qwen, LLaMA, Mistral, GLM, etc.)
| Category | Stack |
|---|---|
| LLM | Qwen2.5-7B (default), any HuggingFace causal LM, 4/8-bit quantization (BitsAndBytes) |
| Embeddings | MiniLM-L6-v2 (sentence-transformers) |
| RAG | LangChain + FAISS |
| Statistics | R / metafor (via rpy2) |
| Search APIs | Biopython (Entrez), Europe PMC, Semantic Scholar, Google Scholar |
| Full-text | PMC, Unpaywall, OpenAlex, CrossRef, DOAJ, CORE |
| PyMuPDF, pdfplumber, pdfminer, pypdfium2 | |
| UI | Streamlit, React + Vite |
| Output | LaTeX (pdflatex), CSV, Excel, RevMan XML |
- Single-GPU constraint: Only one LLM fits in VRAM at a time (8 GB). Ensemble votes are sequential with the same model. Multi-model ensemble supported but adds latency.
- Full-text access: Many articles are publisher-restricted. The multi-source cascade helps, but some papers still fall back to abstract-only.
- Statistical synthesis: Requires papers to report compatible quantitative outcomes. If extracted data lacks effect sizes/CIs, no meta-analysis is produced.
- LLM accuracy: A 7B model can generate incorrect extractions or quality judgments. Human validation is recommended.
- Google Scholar: Aggressive rate limiting and CAPTCHAs may block access after many requests.
Trabajo de Fin de Grado (TFG), Comillas Pontifical University, ICAI. Presented as a poster at CIPIE 2026 (II Congreso Internacional de Psicología, Innovación Tecnológica y Emprendimiento), Madrid, July 2026.
- Ignacio Queipo de Llano Pérez-Gascón¹ (corresponding author): iqueipo.pg24@gmail.com
- Á. López-López⁵
- S. Lumbreras⁵
- P. Collazo-Castiñeira³
- M. Sánchez-Izquierdo³
- I. Echegoyen³ˑ⁴
- E. C. Garrido-Merchán²ˑ⁵
Affiliations (all at Comillas Pontifical University, Madrid, Spain):
- ICAI School of Engineering
- Quantitative Methods Department
- Psychology Department
- Laboratory of Psychology
- Institute for Research in Technology
- Qwen Team (Alibaba) for Qwen2.5-7B
- Hugging Face for the Transformers ecosystem
- Wolfgang Viechtbauer for R/metafor
- Open-source communities: LangChain, FAISS, Streamlit, Biopython
If you use SAM in your research, please cite it. Until the associated paper is published, cite the software:
Queipo de Llano Pérez-Gascón, I., López-López, Á., Lumbreras, S., Collazo-Castiñeira, P., Sánchez-Izquierdo, M., Echegoyen, I., & Garrido-Merchán, E. C. (2026). SAM: Systematic Agent for Meta-analysis, a multi-agent LLM system for automating systematic reviews and meta-analyses [Software]. https://github.com/iqueipopg/SAM-overview
Version: v0.6.0. 9 agents (A1-A9), 991 automated tests (879 passing; the rest skipped or xfail for deferred features / GPU-only paths).
Copyright (c) 2026 Ignacio Queipo de Llano Perez-Gascon and the SAM authors.