Independent benchmark analysis / 2024 challenge, frozen edition

BioASQ 12b

Separate finding evidence from answering with evidence.

BioASQ 12b makes retrieval an explicit experimental condition. Phase A retrieves articles and snippets; Phase A+ answers from system-found evidence; Phase B answers after gold evidence becomes available. We use that structure to explain why a strong biomedical answer score cannot be interpreted without its evidence condition. Our original comparison map keeps question types, retrieval quality and answer judging separate. It also flags an upstream training-count discrepancy instead of silently choosing a convenient total. This is an independent analysis of a completed challenge edition, not the official BioASQ submission service.

01 / What is being tested?

The task, before the score.

input
An English biomedical question, with evidence access determined by phase
output
Retrieved articles/snippets, exact answers and/or ideal prose answers
unit
One biomedical question in one challenge batch and phase
setting
Four independent 2024 batches; compare phase and question type together

Data origin. Biomedical experts author questions and reference evidence from biomedical literature. [5][6][7]

Test questions
340

Four batches of 85; derived total.

Table 1 [5]
Answer families
4

Yes/no, factoid, list and summary.

Table 1 [5]
Experimental phases
3

A: retrieval; A+: system-evidence answering; B: gold-evidence answering.

§2.1 [5]
Training count caveat
5,046 / 5,049

Official catalog and paper prose say 5,046; paper Table 1 says 5,049. Preserve this discrepancy.

Dataset catalog; Table 1 [6][5]
  1. 01

    Phase A: retrieve

    Return literature documents and snippets.

    [5]
  2. 02

    Phase A+: answer with retrieved evidence

    Generate answers before receiving gold evidence.

    [5]
  3. 03

    Phase B: answer with gold evidence

    Answer with manually selected supporting material.

    [5]
  4. 04

    Judge by answer family

    Keep exact and ideal-answer evaluations distinct.

    [7]

Dataset anatomy

The 2024 test question mix

Yes/no

25 + 26 + 24 + 27 across four batches.

102 test questions[5]
List

21 + 18 + 19 + 22.

80 test questions[5]
Factoid

21 + 19 + 26 + 19.

85 test questions[5]
Summary

18 + 22 + 16 + 17.

73 test questions[5]

Computed sums of the published batch rows. These counts describe questions, not retrieval documents or answer variants. Bar lengths use the largest listed count as their reference; they are not percentages of a shared population.

02 / Measurement

Task-specific metrics

Higher component scores are better; lower rank is better

Official winner selection separates document MAP, snippet F-measure, exact-answer ranks and manual ideal-answer scores.

Scoring definition

Exact-answer ranking combines ranks for yes/no accuracy, factoid MRR and list mean F-measure

A rank aggregation is not a universal accuracy percentage; phase and batch must match. [7]

03 / Measured evidence

Results, with their conditions attached.

This dossier does not reproduce a model ranking. The cited primary sources are the authority for experimental results; the analysis here focuses on task construction, measurement, and interpretation.

04 / Our original analysis

What follows from the design?

01

Evidence availability is an experimental variable

Published evidence

A+ precedes gold-evidence release; B follows it. [5]

Our interpretation

Our inference: evaluate retrieval and answering separately before attributing a weak answer to the language model alone.

02

A mixed answer set requires mixed score interpretation

Published evidence

Official exact-answer selection combines ranks across three answer types. [7]

Our interpretation

Our inference: preserve the component scores; a rank is sensitive to which competing systems entered the batch.

03

Version labels prevent false precision

Published evidence

The official sources disagree about the training total. [6][5]

Our interpretation

Our inference: cite the exact downloaded manifest when reproducing training, and label the discrepancy in descriptive coverage.

05 / Scope of the evidence

Where this benchmark stops.

Edition-specific design

Later BioASQ editions change datasets and procedures; the page concerns 12b. [6]

Reference enrichment

The overview distinguishes preliminary from enriched relevant-item references. Retrieval scores need a reference revision. [5]

Ideal answers require a different judge

Manual answer quality scores cannot be treated as exact-match accuracy. [7]

06 / Working with the benchmark

Access & reuse.

Open the author’s resource ↗
Availability
Registration required for official dataset downloads
License
Dataset license not verified in the inspected catalog
Conditions
Consult official dataset terms; article CC BY 4.0 is not proof of a dataset license.
[6]

Evidence trail

Read the originals.

  1. Overview of BioASQ Tasks 12b and Synergy12 in CLEF2024 ↗

    Nentidis et al. / BioASQ organizers. Edition-specific stages and batch counts; training total differs between text and Table 1.

  2. BioASQ task b datasets ↗

    BioASQ organizers. Official edition catalog lists 5,046 questions for Training 12b and requires registration for downloading.

  3. Twelfth Challenge Winners ↗

    BioASQ organizers. Official phase-specific ranking rules; no scores from different answer types are interchangeable.

Benchmark authors retain authorship of their work. This publication provides independent analysis; published rows are not Arcophos evaluation runs. Editorial method.

Explore the assumptions ↗