The short answer
PubMedQA and BioASQ are often grouped under biomedical question answering, but that label conceals different evidence contracts. PubMedQA supplies the relevant abstract; BioASQ 12b separates retrieval and answering phases. Our original comparison uses those contracts to help engineers choose an evaluation for a specific failure mode. It does not add the scores together or propose a new benchmark authored by this site.
Map the product behavior to the benchmark input
Imagine a research assistant whose user supplies one abstract and asks whether it supports a narrow proposition. PubMedQA resembles one part of that interaction because the relevant text is already available. Now imagine a system expected to search a literature collection and construct an answer. BioASQ’s retrieval and phase-dependent answering conditions expose additional parts of that pipeline.
These are task-shape comparisons, not claims of clinical equivalence. Before choosing a score, write down what the product receives and which information it must find. List the expected output format as well. A categorical answer and an explanatory evidence summary require different evaluation criteria even when both concern the same scientific question. Start with this input-output description rather than a popular benchmark name.
Use a coverage map instead of a blended number
Our explorer separates supplied evidence, retrieved evidence, exact answers and prose answers. These are editorial categories tied to published tasks. They make a missing evaluation condition visible without pretending that one task is a substitute for another. A team can mark which rows resemble its intended workflow and then investigate the gaps.
A combined percentage would require an explicit population, weighting rule and interpretation. Simply averaging two published percentages supplies none of those. It can also hide a weak retrieval component behind a strong supplied-context result. Keep a small set of named outcomes and explain what each measures. If an organization later creates a composite for an internal decision, publish its weights and sensitivity to those weights separately.
Keep the denominator attached to the measurement
PubMedQA has large auxiliary collections alongside its much smaller labeled evaluation set. BioASQ changes its training collection across editions and divides testing into challenge batches. Dataset scale can therefore refer to training examples, test questions, retrieved documents or answer variants. Those units are not interchangeable, even when each is informally called an example.
Our reporting recommendation is to pair every count with a unit and a role: labeled test questions, auxiliary training questions, retrieved documents, or generated responses. Record language filtering and exclusions after preprocessing. When an upstream source disagrees with another source, preserve both descriptions and resolve the executable run through its actual manifest. Precision in terminology is more useful than an unsupported exact-looking total.
Investigate disagreement before expanding coverage
If a model performs differently across these benchmarks, first inspect the evidence conditions and answer policies. A gain on supplied abstracts could arise without any improvement in finding sources. A gain in document retrieval could fail to improve answer synthesis. These are possible interpretations to test with paired artifacts, not explanations that can be inferred from aggregate scores alone.
A useful comparison package includes the selected task map, fixed run identities, raw outputs and a short set of error categories. Review a question from each failure category using the source’s actual scoring rules. Then decide whether another benchmark answers a remaining question. This keeps expansion purposeful: each added task should expose a new uncertainty rather than merely increase the length of a leaderboard.
References & further reading
These original sources support the methods discussed here. Our suggested planning steps are editorial guidance, not an endorsement by the source authors.
- PubMedQA: A Dataset for Biomedical Research Question Answering ↗Jin et al. / EMNLP-IJCNLP 2019. Original benchmark construction, split, settings and baseline results.
- PubMedQA evaluation.py ↗PubMedQA authors. Exact test-key check, accuracy and macro-F1 implementation.
- Overview of BioASQ Tasks 12b and Synergy12 in CLEF2024 ↗Nentidis et al. / BioASQ organizers. Edition-specific stages and batch counts; training total differs between text and Table 1.
- Twelfth Challenge Winners ↗BioASQ organizers. Official phase-specific ranking rules; no scores from different answer types are interchangeable.
Published by Arcophos. Educational material, not clinical advice or a claim of regulatory compliance. Read our editorial method.