Independent benchmark analysis / Original EMNLP 2019 reasoning-required protocol

PubMedQA

One abstract. Three labels. A surprisingly important boundary.

PubMedQA asks whether a research abstract supports a yes, no or maybe answer. Its engineering lesson is the separation between supplied evidence and withheld conclusions. We examine the labeled test set, the auxiliary training collections and the evaluator as different parts of a run. Our analysis asks what changes when conclusions enter the prompt, why label balance makes accuracy and macro-F1 answer different questions, and how to preserve that distinction in a reproducible comparison. The historical baseline panel reproduces the authors’ measurements; it is not a new experiment or a current model ranking.

01 / What is being tested?

The task, before the score.

input
Research question plus its abstract with the conclusion withheld
output
One label: yes, no or maybe
unit
One research question / abstract pair
setting
Reasoning-required; the conclusion may be auxiliary training supervision, not test input

Data origin. Published PubMed abstracts; expert labels for PQA-L, unlabeled articles for PQA-U, and automatically transformed titles for PQA-A. [1][2][3][4]

Expert-labeled collection
1,000 questions

PQA-L; do not describe the auxiliary pools as expert-labeled tests.

§3.1 [1]
Held-out evaluation
500 questions

The other 500 labeled examples support ten-fold development.

§3.1 [1]
Unlabeled collection
61.2k questions

PQA-U; paper-rounded count.

Table 1 [1]
Artificial collection
211.3k questions

PQA-A; paper-rounded count, title transformations.

Table 1 [1]
Output space
3 labels

The official evaluator requires exactly the test PMID keys.

evaluation.py [3]
  1. 01

    Select the labeled test manifest

    Keep auxiliary training pools separate.

    [2]
  2. 02

    Construct the permitted input

    Use question and non-conclusion abstract text.

    [1]
  3. 03

    Produce label predictions

    Preserve the PMID-to-label mapping.

    [2]
  4. 04

    Score every expected identifier

    The script checks exact key coverage before scoring.

    [3]

Dataset anatomy

Label composition of PQA-L

Yes

Derived from 55.2% of the 1,000-question PQA-L collection.

552 labeled questions[1]
No

Derived from 33.8% of PQA-L.

338 labeled questions[1]
Maybe

Derived from 11.0% of PQA-L.

110 labeled questions[1]

Full labeled collection, not a separately verified test-only distribution. Counts are arithmetic conversions of Table 1 percentages. Bar lengths use the largest listed count as their reference; they are not percentages of a shared population.

02 / Measurement

Accuracy and macro-F1

Higher is better

Accuracy weights questions equally; macro-F1 weights the three answer classes equally.

Scoring definition

Accuracy = correct / N; macro-F1 = (F1_yes + F1_no + F1_maybe) / 3

Compare identical held-out identifiers and whether conclusions are available. [3][1]

03 / Measured evidence

Results, with their conditions attached.

Paper-reported results / selected rows

Original reasoning-required baselines

2019 paper, PQA-L held-out test; original training schedules.

Accuracy · %
050100
Reported
Majority baselineAlways selects the majority class.
55.20%
BioBERT, multi-phase + auxiliary supervisionExact Table 5 value, rather than the abstract’s rounded 68.1%.
68.08%
Single human annotatorOne annotator under the reasoning-required setting; not a clinician population estimate.
78.00%

Historical paper-reported values. Training and human annotation procedures differ; no uncertainty is supplied in this selected table.

Source: Tables 4–5 [1]

04 / Our original analysis

What follows from the design?

01

A correct label can hide a missing class

Published evidence

The collection includes an uncommon maybe label. [1][3]

Our interpretation

Our inference: report the confusion matrix before interpreting an accuracy gain as better uncertainty handling.

02

Conclusions change the task

Published evidence

The paper distinguishes reasoning-required and reasoning-free inputs. [1]

Our interpretation

Our inference: an input-field audit belongs beside a score comparison, because accidental conclusion inclusion changes the evidence available.

03

Dataset size is not test size

Published evidence

Large auxiliary pools accompany a much smaller labeled test. [1]

Our interpretation

Our inference: separate training-resource scale from evaluation precision in every dataset summary.

05 / Scope of the evidence

Where this benchmark stops.

Article conclusion is the reference

Agreement with a publication’s conclusion does not establish clinical efficacy or a complete literature synthesis. [1]

Single-abstract scope

The task supplies relevant evidence; it does not test the ability to locate competing studies. [1]

Public material and repeated use

The articles and benchmark are public. Exposure cannot be excluded from a model name alone. [2]

06 / Working with the benchmark

Access & reuse.

Open the author’s resource ↗
Availability
Public repository
License
MIT repository license
Conditions
Underlying article rights are distinct; this site republishes analysis and metadata, not abstracts.
[4][2]

Evidence trail

Read the originals.

  1. PubMedQA: A Dataset for Biomedical Research Question Answering ↗

    Jin et al. / EMNLP-IJCNLP 2019. Original benchmark construction, split, settings and baseline results.

  2. PubMedQA official data and evaluation repository ↗

    PubMedQA authors. Release organization, split procedure and prediction format.

  3. PubMedQA evaluation.py ↗

    PubMedQA authors. Exact test-key check, accuracy and macro-F1 implementation.

  4. PubMedQA repository license ↗

    PubMedQA authors. MIT repository license; does not independently clear underlying article rights.

Benchmark authors retain authorship of their work. This publication provides independent analysis; published rows are not Arcophos evaluation runs. Editorial method.

Explore the assumptions ↗