INGENIO (CSIC-Universitat Politècnica de València), Valencia, Spain
This repository contains the scripts and the data processing pipeline used in the study 'Bringing public attention to disease into global health priority-setting'. The research maps diseases from the Global Burden of Disease (GBD) 2023 classification across three dimensions, namely epidemiological burden, public attention and research effort, for the period 2016 to 2023. Burden is measured through disability-adjusted life years (DALYs), public attention through Wikipedia pageviews, and research effort through publications indexed in OpenAlex under major MeSH descriptors. The analysis covers 19 disease groups (GBD level-2) and 138 specific diseases (GBD level-3), and it is complemented by a territorial disaggregation for four linguistic areas, German, Persian, Swahili and Vietnamese. Ternary plots position each disease according to the relative balance between knowledge supply and health demands, making visible where research effort departs from epidemiological burden, from public attention, or from both.
A central concern in global health priority-setting is whether the supply of scientific knowledge aligns with health needs and demands. Alignment is usually assessed by comparing research effort with disease burden, which overlooks whether diseases are socially visible and generate public attention. This study develops an analytical framework that treats public attention and epidemiological burden as complementary dimensions of health demand and examines their alignment with knowledge supply. The three dimensions show limited alignment. Cardiovascular diseases account for the largest share of disease burden, mental disorders attract the largest share of public attention, and neoplasms concentrate the largest share of research effort. Public attention and disease burden are weakly correlated at both the disease group and the specific disease scale, so Wikipedia pageviews capture a distinct dimension of health demand. Ternary plots reveal different forms of misalignment, with some diseases dominated by burden, others by research effort, and others by public attention. Territorial analyses show distinct profiles, with HIV/AIDS and sexually transmitted infections especially prominent in the Swahili-speaking area.
1_gbd_dataset_generation.ipynbto5_2_country_level_3.ipynb. Python notebooks that build the datasets, retrieve Wikipedia pageviews, aggregate indicators along the GBD hierarchy, and run the reliability analysis.6_stat_plots.Rto12_ternary_countries.R. R scripts that produce the descriptive figures, the correlation analysis, the supply-to-demand ratios and the ternary plots.results/. Final datasets in TSV format, ready for analysis and visualisation.data/. Input data. Not included in this repository, see the Data section below.LICENSE. GPL-3.0 licence.
| Script | Purpose |
|---|---|
1_gbd_dataset_generation.ipynb |
Reads the GBD 2023 four-level classification and the manual matching to MeSH descriptors and Wikipedia articles, and generates the disease-to-source lookup tables |
2_wikipedia_views.ipynb |
Retrieves pageviews for every matched article across all language editions through the Wikimedia REST API |
3_1_dataset_generation_lv2.ipynb |
Builds the level-2 dataset, aggregating publications, pageviews and DALYs and counting each publication only once per disease group |
3_2_dataset_generation_lv3.ipynb |
Builds the level-3 dataset, incorporating the level-4 diseases inherited by aggregation |
3_3_dataset_generation_lv2_year.ipynb |
Builds the level-2 dataset broken down by publication year |
4_validation.ipynb |
Compares the level-2 publication assignments against the external dataset by Schmallenbach et al. |
5_1_country_level_2.ipynb |
Builds the level-2 dataset for the selected linguistic areas, filtering DALYs by country, publications by first author affiliation and pageviews by language edition |
5_2_country_level_3.ipynb |
Builds the equivalent level-3 dataset for the selected linguistic areas |
6_stat_plots.R |
Descriptive figures of the distribution and the volume of the three dimensions across disease groups |
7_correlations.R |
Pairwise correlations between dimensions using Kendall's rank correlation coefficient |
8_ratios.R |
Ratios of research effort over disease burden and over public attention at the specific disease scale |
9_ternary.R |
Ternary plots for levels 2 and 3, normalised against all diseases |
10_ternary_context.R |
Ternary plots recalculated within each disease group and compared with the overall pattern |
11_country.R |
Distribution of the three dimensions across the four linguistic areas |
12_ternary_countries.R |
Ternary plots by linguistic area, compared with the global aggregate |
The results/ folder contains the analytical datasets produced by the notebooks. Every file reports publication counts, pageviews and DALYs for the period 2016 to 2023.
results_final_lv2.tsv. The 19 GBD level-2 disease groups.results_final_lv3.tsv. The 158 GBD level-3 diseases, of which 138 have complete data across the three dimensions.results_final_lv2_year.tsv. Level-2 disease groups broken down by year.results_final_lv2_country.tsvandresults_final_lv3_country.tsv. Levels 2 and 3 disaggregated by language edition. Fourteen editions are included, of which German, Persian, Swahili and Vietnamese are analysed in the paper.
The complete disease mapping is openly available on Zenodo at 10.5281/zenodo.21694438. The scripts expect the following files in a data/ folder placed in the root of the repository.
matching_2023.xlsx. The GBD 2023 four-level classification (sheetClassification) and the manual matching of each cause to MeSH descriptors and Wikipedia articles (sheetMapping).level_3_MESH.tsvandlevel_3_Wikipedia.tsv. Lookup tables linking each specific disease to its MeSH descriptors, with qualifiers where applicable, and to its Wikipedia article.gbd_all_dalys_1423.csv,gbd_all_dalys_countries_1623_1.csvandgbd_all_dalys_countries_1623_2.csv. DALY estimates exported from the GBD Results Tool, globally and by country.mesh_papers.csvandmesh_papers_with_q.csv. Publications indexed under each MeSH descriptor, with and without qualifiers, extracted from OpenAlex.openalex_countries.csv. Country of the first author affiliation for each publication.wiki_pageviews_user.tsv. Pageviews by article and language edition, generated by notebook 2.level_2_mesh_validation.tsvanddata_schmallenbach/pmid_cause_FAcountry_year.csv. Files used by the reliability analysis.
Publication data were retrieved from an in-house OpenAlex instance, so the raw extraction is not redistributed here. The corresponding queries are documented in the notebooks and can be reproduced through the OpenAlex API.
Data processing was performed in Python 3.13.5 and visualisation in R 4.5.1 through RStudio.
- Python packages.
pandas,requests,urllib3,openpyxl. - R packages.
dplyr,tidyr,reshape2,readxl,ggplot2,ggtern,GGally,ggrepel,ggnewscale,patchwork,scales,wesanderson,DescTools.
-
Clone this repository.
git clone https://github.com/Wences91/gbd_misalignment.git cd gbd_misalignment -
Download the mapping files from Zenodo and place them in a
data/folder in the root of the repository. -
Run the Python notebooks in numerical order. Notebook 2 queries the Wikimedia REST API and requires a valid contact address in the
USER_AGENTvariable, since Wikimedia blocks generic user agents. -
Run the R scripts in numerical order to reproduce the figures reported in the paper. All scripts read the datasets stored in
results/, so the figures can be reproduced without rebuilding the datasets from scratch.
Arroyo-Machado, W., Rafols, I., & Díaz-Faes, A. A. (2026). Bringing public attention to disease into global health priority-setting.
Wenceslao Arroyo-Machado is supported by the Momentum programme (MMT24-INGENIO-01). Funding for these grants comes from the European Union's Recovery and Resilience Facility-Next Generation, in the framework of the General Invitation of the Spanish Government's public business entity Red.es to participate in talent attraction and retention programmes within Investment 4 of Component 19 of the Recovery, Transformation and Resilience Plan. Adrián A. Díaz-Faes acknowledges support from research projects PID2020-112837RJ-I00, funded by MCIN/AEI/10.13039/501100011033, and MMT24-INGENIO-01.
We thank Alysson Mazoni and the University of Campinas for providing access to the OpenAlex publications in-house database. We are also grateful to previous researchers on this topic, in particular Alfredo Yegros, for making the classifications of their articles fully transparent and accessible.
This repository is distributed under the GPL-3.0 licence.