Benchmark dossiers / independent analysis

The benchmark behind the number.

Inspect what each benchmark asks, how its score is produced, and which conclusions its design can support. These are original analyses by Arcophos of the benchmark authors’ work. Any reproduced measurements are labeled with their published source and version.

Dossier01

Original EMNLP 2019 reasoning-required protocol

PubMedQA ↗

One abstract. Three labels. A surprisingly important boundary.

UnitOne research question / abstract pairMeasureAccuracy and macro-F1
Dossier02

2024 challenge, frozen edition

BioASQ 12b ↗

Separate finding evidence from answering with evidence.

UnitOne biomedical question in one challenge batch and phaseMeasureTask-specific metrics
What we contribute. Source reconciliation, task and metric interpretation, and tools that make assumptions inspectable. We do not claim to have created these benchmarks or run the reported models.