Evidence availability is an experimental variable
A+ precedes gold-evidence release; B follows it. [5]
Our inference: evaluate retrieval and answering separately before attributing a weak answer to the language model alone.
Independent benchmark analysis / 2024 challenge, frozen edition
BioASQ 12b makes retrieval an explicit experimental condition. Phase A retrieves articles and snippets; Phase A+ answers from system-found evidence; Phase B answers after gold evidence becomes available. We use that structure to explain why a strong biomedical answer score cannot be interpreted without its evidence condition. Our original comparison map keeps question types, retrieval quality and answer judging separate. It also flags an upstream training-count discrepancy instead of silently choosing a convenient total. This is an independent analysis of a completed challenge edition, not the official BioASQ submission service.
01 / What is being tested?
Data origin. Biomedical experts author questions and reference evidence from biomedical literature. [5][6][7]
Four batches of 85; derived total.
Table 1 [5]Yes/no, factoid, list and summary.
Table 1 [5]A: retrieval; A+: system-evidence answering; B: gold-evidence answering.
§2.1 [5]Return literature documents and snippets.
[5]Generate answers before receiving gold evidence.
[5]Answer with manually selected supporting material.
[5]Keep exact and ideal-answer evaluations distinct.
[7]Dataset anatomy
25 + 26 + 24 + 27 across four batches.
21 + 18 + 19 + 22.
21 + 19 + 26 + 19.
18 + 22 + 16 + 17.
Computed sums of the published batch rows. These counts describe questions, not retrieval documents or answer variants. Bar lengths use the largest listed count as their reference; they are not percentages of a shared population.
02 / Measurement
Higher component scores are better; lower rank is better
Official winner selection separates document MAP, snippet F-measure, exact-answer ranks and manual ideal-answer scores.
Exact-answer ranking combines ranks for yes/no accuracy, factoid MRR and list mean F-measure
A rank aggregation is not a universal accuracy percentage; phase and batch must match. [7]
03 / Measured evidence
04 / Our original analysis
A+ precedes gold-evidence release; B follows it. [5]
Our inference: evaluate retrieval and answering separately before attributing a weak answer to the language model alone.
Official exact-answer selection combines ranks across three answer types. [7]
Our inference: preserve the component scores; a rank is sensitive to which competing systems entered the batch.
05 / Scope of the evidence
Later BioASQ editions change datasets and procedures; the page concerns 12b. [6]
The overview distinguishes preliminary from enriched relevant-item references. Retrieval scores need a reference revision. [5]
Manual answer quality scores cannot be treated as exact-match accuracy. [7]
Evidence trail
Nentidis et al. / BioASQ organizers. Edition-specific stages and batch counts; training total differs between text and Table 1.
BioASQ organizers. Official edition catalog lists 5,046 questions for Training 12b and requires registration for downloading.
BioASQ organizers. Official phase-specific ranking rules; no scores from different answer types are interchangeable.
Benchmark authors retain authorship of their work. This publication provides independent analysis; published rows are not Arcophos evaluation runs. Editorial method.