Independent benchmark analysis / Original EMNLP 2019 reasoning-required protocol
PubMedQA
One abstract. Three labels. A surprisingly important boundary.
PubMedQA asks whether a research abstract supports a yes, no or maybe answer. Its engineering lesson is the separation between supplied evidence and withheld conclusions. We examine the labeled test set, the auxiliary training collections and the evaluator as different parts of a run. Our analysis asks what changes when conclusions enter the prompt, why label balance makes accuracy and macro-F1 answer different questions, and how to preserve that distinction in a reproducible comparison. The historical baseline panel reproduces the authors’ measurements; it is not a new experiment or a current model ranking.
01 / What is being tested?
The task, before the score.
- input
- Research question plus its abstract with the conclusion withheld
- output
- One label: yes, no or maybe
- unit
- One research question / abstract pair
- setting
- Reasoning-required; the conclusion may be auxiliary training supervision, not test input
Data origin. Published PubMed abstracts; expert labels for PQA-L, unlabeled articles for PQA-U, and automatically transformed titles for PQA-A. [1][2][3][4]
- Expert-labeled collection
- 1,000 questions
PQA-L; do not describe the auxiliary pools as expert-labeled tests.
§3.1 [1] - Held-out evaluation
- 500 questions
The other 500 labeled examples support ten-fold development.
§3.1 [1] - Unlabeled collection
- 61.2k questions
PQA-U; paper-rounded count.
Table 1 [1] - Artificial collection
- 211.3k questions
PQA-A; paper-rounded count, title transformations.
Table 1 [1] - Output space
- 3 labels
The official evaluator requires exactly the test PMID keys.
evaluation.py [3]
- 01
Select the labeled test manifest
Keep auxiliary training pools separate.
[2] - 02
Construct the permitted input
Use question and non-conclusion abstract text.
[1] - 03
Produce label predictions
Preserve the PMID-to-label mapping.
[2] - 04
Score every expected identifier
The script checks exact key coverage before scoring.
[3]
Dataset anatomy
Label composition of PQA-L
Derived from 55.2% of the 1,000-question PQA-L collection.
Derived from 33.8% of PQA-L.
Derived from 11.0% of PQA-L.
Full labeled collection, not a separately verified test-only distribution. Counts are arithmetic conversions of Table 1 percentages. Bar lengths use the largest listed count as their reference; they are not percentages of a shared population.
02 / Measurement
Accuracy and macro-F1
Higher is better
03 / Measured evidence
Results, with their conditions attached.
Paper-reported results / selected rows
Original reasoning-required baselines
2019 paper, PQA-L held-out test; original training schedules.
Historical paper-reported values. Training and human annotation procedures differ; no uncertainty is supplied in this selected table.
Source: Tables 4–5 [1]
04 / Our original analysis
What follows from the design?
Conclusions change the task
The paper distinguishes reasoning-required and reasoning-free inputs. [1]
Our inference: an input-field audit belongs beside a score comparison, because accidental conclusion inclusion changes the evidence available.
Dataset size is not test size
Large auxiliary pools accompany a much smaller labeled test. [1]
Our inference: separate training-resource scale from evaluation precision in every dataset summary.
05 / Scope of the evidence
Where this benchmark stops.
Article conclusion is the reference
Agreement with a publication’s conclusion does not establish clinical efficacy or a complete literature synthesis. [1]
Single-abstract scope
The task supplies relevant evidence; it does not test the ability to locate competing studies. [1]
Public material and repeated use
The articles and benchmark are public. Exposure cannot be excluded from a model name alone. [2]
- Availability
- Public repository
- License
- MIT repository license
- Conditions
- Underlying article rights are distinct; this site republishes analysis and metadata, not abstracts.
Evidence trail
Read the originals.
- PubMedQA: A Dataset for Biomedical Research Question Answering ↗
Jin et al. / EMNLP-IJCNLP 2019. Original benchmark construction, split, settings and baseline results.
- PubMedQA official data and evaluation repository ↗
PubMedQA authors. Release organization, split procedure and prediction format.
- PubMedQA evaluation.py ↗
PubMedQA authors. Exact test-key check, accuracy and macro-F1 implementation.
- PubMedQA repository license ↗
PubMedQA authors. MIT repository license; does not independently clear underlying article rights.
Benchmark authors retain authorship of their work. This publication provides independent analysis; published rows are not Arcophos evaluation runs. Editorial method.