{"publication":"Med Evals","url":"https://medevals.ai","publisher":"Arcophos","updated":"2026-09-28","provenance":"Independent analytical publication. Benchmark creation and experimental results belong to their cited authors. Reported results are source-version snapshots, not new Arcophos runs or a live leaderboard.","benchmarks":[{"slug":"pubmedqa","name":"PubMedQA","shortName":"PubMedQA","version":"Original EMNLP 2019 reasoning-required protocol","creators":"Qiao Jin, Bhuwan Dhingra, Zhengping Liu, William W. Cohen and Xinghua Lu","paperDate":"2019-11","headline":"One abstract. Three labels. A surprisingly important boundary.","summary":"PubMedQA asks whether a research abstract supports a yes, no or maybe answer. Its engineering lesson is the separation between supplied evidence and withheld conclusions. We examine the labeled test set, the auxiliary training collections and the evaluator as different parts of a run. Our analysis asks what changes when conclusions enter the prompt, why label balance makes accuracy and macro-F1 answer different questions, and how to preserve that distinction in a reproducible comparison. The historical baseline panel reproduces the authors’ measurements; it is not a new experiment or a current model ranking.","task":{"input":"Research question plus its abstract with the conclusion withheld","output":"One label: yes, no or maybe","unit":"One research question / abstract pair","setting":"Reasoning-required; the conclusion may be auxiliary training supervision, not test input"},"dataOrigin":"Published PubMed abstracts; expert labels for PQA-L, unlabeled articles for PQA-U, and automatically transformed titles for PQA-A.","facts":[{"label":"Expert-labeled collection","value":"1,000 questions","detail":"PQA-L; do not describe the auxiliary pools as expert-labeled tests.","sourceIds":["pubmedqa-paper"],"locator":"§3.1"},{"label":"Held-out evaluation","value":"500 questions","detail":"The other 500 labeled examples support ten-fold development.","sourceIds":["pubmedqa-paper"],"locator":"§3.1"},{"label":"Unlabeled collection","value":"61.2k questions","detail":"PQA-U; paper-rounded count.","sourceIds":["pubmedqa-paper"],"locator":"Table 1"},{"label":"Artificial collection","value":"211.3k questions","detail":"PQA-A; paper-rounded count, title transformations.","sourceIds":["pubmedqa-paper"],"locator":"Table 1"},{"label":"Output space","value":"3 labels","detail":"The official evaluator requires exactly the test PMID keys.","sourceIds":["pubmedqa-eval"],"locator":"evaluation.py"}],"metric":{"name":"Accuracy and macro-F1","description":"Accuracy weights questions equally; macro-F1 weights the three answer classes equally.","formula":"Accuracy = correct / N; macro-F1 = (F1_yes + F1_no + F1_maybe) / 3","direction":"Higher is better","comparability":"Compare identical held-out identifiers and whether conclusions are available.","sourceIds":["pubmedqa-eval","pubmedqa-paper"]},"workflow":[{"label":"Select the labeled test manifest","detail":"Keep auxiliary training pools separate.","sourceIds":["pubmedqa-repo"]},{"label":"Construct the permitted input","detail":"Use question and non-conclusion abstract text.","sourceIds":["pubmedqa-paper"]},{"label":"Produce label predictions","detail":"Preserve the PMID-to-label mapping.","sourceIds":["pubmedqa-repo"]},{"label":"Score every expected identifier","detail":"The script checks exact key coverage before scoring.","sourceIds":["pubmedqa-eval"]}],"slices":[{"label":"Yes","value":552,"unit":"labeled questions","detail":"Derived from 55.2% of the 1,000-question PQA-L collection.","sourceIds":["pubmedqa-paper"]},{"label":"No","value":338,"unit":"labeled questions","detail":"Derived from 33.8% of PQA-L.","sourceIds":["pubmedqa-paper"]},{"label":"Maybe","value":110,"unit":"labeled questions","detail":"Derived from 11.0% of PQA-L.","sourceIds":["pubmedqa-paper"]}],"sliceTitle":"Label composition of PQA-L","sliceNote":"Full labeled collection, not a separately verified test-only distribution. Counts are arithmetic conversions of Table 1 percentages.","results":[{"id":"original-accuracy","title":"Original reasoning-required baselines","metric":"Accuracy","unit":"%","lower":0,"upper":100,"scope":"2019 paper, PQA-L held-out test; original training schedules.","sourceIds":["pubmedqa-paper"],"locator":"Tables 4–5","rows":[{"label":"Majority baseline","value":55.2,"display":"55.20%","detail":"Always selects the majority class."},{"label":"BioBERT, multi-phase + auxiliary supervision","value":68.08,"display":"68.08%","detail":"Exact Table 5 value, rather than the abstract’s rounded 68.1%."},{"label":"Single human annotator","value":78,"display":"78.00%","detail":"One annotator under the reasoning-required setting; not a clinician population estimate."}],"note":"Historical paper-reported values. Training and human annotation procedures differ; no uncertainty is supplied in this selected table."}],"analysis":[{"heading":"A correct label can hide a missing class","evidence":"The collection includes an uncommon maybe label.","interpretation":"Our inference: report the confusion matrix before interpreting an accuracy gain as better uncertainty handling.","sourceIds":["pubmedqa-paper","pubmedqa-eval"]},{"heading":"Conclusions change the task","evidence":"The paper distinguishes reasoning-required and reasoning-free inputs.","interpretation":"Our inference: an input-field audit belongs beside a score comparison, because accidental conclusion inclusion changes the evidence available.","sourceIds":["pubmedqa-paper"]},{"heading":"Dataset size is not test size","evidence":"Large auxiliary pools accompany a much smaller labeled test.","interpretation":"Our inference: separate training-resource scale from evaluation precision in every dataset summary.","sourceIds":["pubmedqa-paper"]}],"limitations":[{"title":"Article conclusion is the reference","detail":"Agreement with a publication’s conclusion does not establish clinical efficacy or a complete literature synthesis.","sourceIds":["pubmedqa-paper"]},{"title":"Single-abstract scope","detail":"The task supplies relevant evidence; it does not test the ability to locate competing studies.","sourceIds":["pubmedqa-paper"]},{"title":"Public material and repeated use","detail":"The articles and benchmark are public. Exposure cannot be excluded from a model name alone.","sourceIds":["pubmedqa-repo"]}],"access":{"status":"Public repository","license":"MIT repository license","restrictions":"Underlying article rights are distinct; this site republishes analysis and metadata, not abstracts.","url":"https://github.com/pubmedqa/pubmedqa","sourceIds":["pubmedqa-license","pubmedqa-repo"]},"sourceIds":["pubmedqa-paper","pubmedqa-repo","pubmedqa-eval","pubmedqa-license"]},{"slug":"bioasq-12b","name":"BioASQ 12b","shortName":"BioASQ 12b","version":"2024 challenge, frozen edition","creators":"BioASQ organizing team; Nentidis et al.","paperDate":"2024","headline":"Separate finding evidence from answering with evidence.","summary":"BioASQ 12b makes retrieval an explicit experimental condition. Phase A retrieves articles and snippets; Phase A+ answers from system-found evidence; Phase B answers after gold evidence becomes available. We use that structure to explain why a strong biomedical answer score cannot be interpreted without its evidence condition. Our original comparison map keeps question types, retrieval quality and answer judging separate. It also flags an upstream training-count discrepancy instead of silently choosing a convenient total. This is an independent analysis of a completed challenge edition, not the official BioASQ submission service.","task":{"input":"An English biomedical question, with evidence access determined by phase","output":"Retrieved articles/snippets, exact answers and/or ideal prose answers","unit":"One biomedical question in one challenge batch and phase","setting":"Four independent 2024 batches; compare phase and question type together"},"dataOrigin":"Biomedical experts author questions and reference evidence from biomedical literature.","facts":[{"label":"Test questions","value":"340","detail":"Four batches of 85; derived total.","sourceIds":["bioasq-paper"],"locator":"Table 1"},{"label":"Answer families","value":"4","detail":"Yes/no, factoid, list and summary.","sourceIds":["bioasq-paper"],"locator":"Table 1"},{"label":"Experimental phases","value":"3","detail":"A: retrieval; A+: system-evidence answering; B: gold-evidence answering.","sourceIds":["bioasq-paper"],"locator":"§2.1"},{"label":"Training count caveat","value":"5,046 / 5,049","detail":"Official catalog and paper prose say 5,046; paper Table 1 says 5,049. Preserve this discrepancy.","sourceIds":["bioasq-data","bioasq-paper"],"locator":"Dataset catalog; Table 1"}],"metric":{"name":"Task-specific metrics","description":"Official winner selection separates document MAP, snippet F-measure, exact-answer ranks and manual ideal-answer scores.","formula":"Exact-answer ranking combines ranks for yes/no accuracy, factoid MRR and list mean F-measure","direction":"Higher component scores are better; lower rank is better","comparability":"A rank aggregation is not a universal accuracy percentage; phase and batch must match.","sourceIds":["bioasq-results"]},"workflow":[{"label":"Phase A: retrieve","detail":"Return literature documents and snippets.","sourceIds":["bioasq-paper"]},{"label":"Phase A+: answer with retrieved evidence","detail":"Generate answers before receiving gold evidence.","sourceIds":["bioasq-paper"]},{"label":"Phase B: answer with gold evidence","detail":"Answer with manually selected supporting material.","sourceIds":["bioasq-paper"]},{"label":"Judge by answer family","detail":"Keep exact and ideal-answer evaluations distinct.","sourceIds":["bioasq-results"]}],"slices":[{"label":"Yes/no","value":102,"unit":"test questions","detail":"25 + 26 + 24 + 27 across four batches.","sourceIds":["bioasq-paper"]},{"label":"List","value":80,"unit":"test questions","detail":"21 + 18 + 19 + 22.","sourceIds":["bioasq-paper"]},{"label":"Factoid","value":85,"unit":"test questions","detail":"21 + 19 + 26 + 19.","sourceIds":["bioasq-paper"]},{"label":"Summary","value":73,"unit":"test questions","detail":"18 + 22 + 16 + 17.","sourceIds":["bioasq-paper"]}],"sliceTitle":"The 2024 test question mix","sliceNote":"Computed sums of the published batch rows. These counts describe questions, not retrieval documents or answer variants.","results":[],"analysis":[{"heading":"Evidence availability is an experimental variable","evidence":"A+ precedes gold-evidence release; B follows it.","interpretation":"Our inference: evaluate retrieval and answering separately before attributing a weak answer to the language model alone.","sourceIds":["bioasq-paper"]},{"heading":"A mixed answer set requires mixed score interpretation","evidence":"Official exact-answer selection combines ranks across three answer types.","interpretation":"Our inference: preserve the component scores; a rank is sensitive to which competing systems entered the batch.","sourceIds":["bioasq-results"]},{"heading":"Version labels prevent false precision","evidence":"The official sources disagree about the training total.","interpretation":"Our inference: cite the exact downloaded manifest when reproducing training, and label the discrepancy in descriptive coverage.","sourceIds":["bioasq-data","bioasq-paper"]}],"limitations":[{"title":"Edition-specific design","detail":"Later BioASQ editions change datasets and procedures; the page concerns 12b.","sourceIds":["bioasq-data"]},{"title":"Reference enrichment","detail":"The overview distinguishes preliminary from enriched relevant-item references. Retrieval scores need a reference revision.","sourceIds":["bioasq-paper"]},{"title":"Ideal answers require a different judge","detail":"Manual answer quality scores cannot be treated as exact-match accuracy.","sourceIds":["bioasq-results"]}],"access":{"status":"Registration required for official dataset downloads","license":"Dataset license not verified in the inspected catalog","restrictions":"Consult official dataset terms; article CC BY 4.0 is not proof of a dataset license.","url":"https://participants-area.bioasq.org/datasets/","sourceIds":["bioasq-data"]},"sourceIds":["bioasq-paper","bioasq-data","bioasq-results"]}],"explorer":{"kind":"coverage","title":"Explore the biomedical evidence contract","intro":"Filter actual benchmark conditions by how evidence is supplied and how answers are judged. The annotations are our original comparison, not new performance measurements.","caution":"This is our independent analytical map of published tasks. It does not execute an evaluation, predict a model’s performance or establish clinical benefit.","sourceIds":["pubmedqa-paper","pubmedqa-repo","pubmedqa-eval","pubmedqa-license","bioasq-paper","bioasq-data","bioasq-results"],"rows":[{"label":"PubMedQA: question + abstract","category":"Supplied evidence","input":"Non-conclusion abstract and research question","output":"Yes/no/maybe","metric":"Accuracy; macro-F1","constraint":"Conclusion must remain withheld at test time.","benchmarkSlug":"pubmedqa","sourceIds":["pubmedqa-paper","pubmedqa-eval"]},{"label":"PubMedQA: question + conclusion","category":"Supplied evidence","input":"Question and the abstract conclusion","output":"Yes/no/maybe","metric":"Accuracy; macro-F1","constraint":"Reasoning-free condition, not directly interchangeable with the main benchmark.","benchmarkSlug":"pubmedqa","sourceIds":["pubmedqa-paper"]},{"label":"BioASQ: document retrieval","category":"Retrieved evidence","input":"Biomedical question","output":"Ranked relevant articles","metric":"MAP","constraint":"Retrieval reference enrichment and batch must match.","benchmarkSlug":"bioasq-12b","sourceIds":["bioasq-results","bioasq-paper"]},{"label":"BioASQ: snippet retrieval","category":"Retrieved evidence","input":"Biomedical question","output":"Relevant text snippets","metric":"F-measure","constraint":"A document hit and a snippet hit have different units.","benchmarkSlug":"bioasq-12b","sourceIds":["bioasq-results"]},{"label":"BioASQ: yes/no answers","category":"Exact answers","input":"Question plus phase-permitted evidence","output":"Yes or no","metric":"Accuracy","constraint":"A+ and B differ in evidence availability.","benchmarkSlug":"bioasq-12b","sourceIds":["bioasq-results","bioasq-paper"]},{"label":"BioASQ: factoid answers","category":"Exact answers","input":"Question plus phase-permitted evidence","output":"Ranked exact answers","metric":"MRR","constraint":"Ranking an answer is distinct from producing a summary.","benchmarkSlug":"bioasq-12b","sourceIds":["bioasq-results"]},{"label":"BioASQ: list answers","category":"Exact answers","input":"Question plus phase-permitted evidence","output":"Set of answer entities","metric":"Mean F-measure","constraint":"Coverage and extra entries both matter.","benchmarkSlug":"bioasq-12b","sourceIds":["bioasq-results"]},{"label":"BioASQ: ideal answers","category":"Prose answers","input":"Question plus phase-permitted evidence","output":"Explanatory prose answer","metric":"Mean manual score","constraint":"Manual quality judgment is not exact-answer accuracy.","benchmarkSlug":"bioasq-12b","sourceIds":["bioasq-results"]}],"parameters":[]},"references":[{"id":"pubmedqa-paper","title":"PubMedQA: A Dataset for Biomedical Research Question Answering","organization":"Jin et al. / EMNLP-IJCNLP 2019","url":"https://aclanthology.org/D19-1259.pdf","note":"Original benchmark construction, split, settings and baseline results.","locator":"§3; Tables 1, 4–5","version":"November 2019"},{"id":"pubmedqa-repo","title":"PubMedQA official data and evaluation repository","organization":"PubMedQA authors","url":"https://github.com/pubmedqa/pubmedqa","note":"Release organization, split procedure and prediction format.","locator":"README; data; preprocess","version":"Accessed 2026-09-28"},{"id":"pubmedqa-eval","title":"PubMedQA evaluation.py","organization":"PubMedQA authors","url":"https://github.com/pubmedqa/pubmedqa/blob/master/evaluation.py","note":"Exact test-key check, accuracy and macro-F1 implementation.","locator":"evaluation.py","version":"Accessed 2026-09-28"},{"id":"pubmedqa-license","title":"PubMedQA repository license","organization":"PubMedQA authors","url":"https://github.com/pubmedqa/pubmedqa/blob/master/LICENSE","note":"MIT repository license; does not independently clear underlying article rights.","locator":"LICENSE","version":"2019"},{"id":"bioasq-paper","title":"Overview of BioASQ Tasks 12b and Synergy12 in CLEF2024","organization":"Nentidis et al. / BioASQ organizers","url":"https://www.iit.demokritos.gr/wp-content/uploads/2024/10/paper-01-1.pdf","note":"Edition-specific stages and batch counts; training total differs between text and Table 1.","locator":"§2.1; Table 1","version":"CLEF 2024"},{"id":"bioasq-data","title":"BioASQ task b datasets","organization":"BioASQ organizers","url":"https://participants-area.bioasq.org/datasets/","note":"Official edition catalog lists 5,046 questions for Training 12b and requires registration for downloading.","locator":"Task b dataset table","version":"Accessed 2026-09-28"},{"id":"bioasq-results","title":"Twelfth Challenge Winners","organization":"BioASQ organizers","url":"https://taskb.bioasq.org/participate/twelfth-challenge-winners","note":"Official phase-specific ranking rules; no scores from different answer types are interchangeable.","locator":"Task 12b selection strategies","version":"2024 challenge"}]}