# Med Evals > Independent original analysis of PubMedQA and BioASQ 12b: datasets, scoring, evidence phases, version caveats and reproducible biomedical QA comparisons. We analyze two named biomedical question-answering benchmarks at the level of their inputs, references and scoring. PubMedQA isolates answering from a supplied abstract; BioASQ 12b separates finding literature from answering with it. Explore the published task conditions, inspect selected historical measurements and use our original guides to make an evaluation reproducible. Benchmark authors retain credit for the datasets and experiments. Our contribution is the analysis and tools; we have not run the models displayed here. ## Provenance Independent analysis published by Arcophos. Benchmark creation belongs to the credited authors. Result rows are selected paper-reported measurements with their source versions and evaluation conditions, not new Arcophos runs or a live leaderboard. ## Benchmark dossiers - [PubMedQA](https://medevals.ai/benchmarks/pubmedqa/): One abstract. Three labels. A surprisingly important boundary. Source version: Original EMNLP 2019 reasoning-required protocol. - [BioASQ 12b](https://medevals.ai/benchmarks/bioasq-12b/): Separate finding evidence from answering with evidence. Source version: 2024 challenge, frozen edition. ## Original analyses - [PubMedQA: audit the input before interpreting the score](https://medevals.ai/guides/pubmedqa-input-and-scoring/): An original engineering analysis of conclusion leakage, class balance and the official PubMedQA evaluator. - [BioASQ 12b: why the evidence phase belongs beside every score](https://medevals.ai/guides/bioasq-12b-evidence-phases/): Use the A, A+ and B conditions to distinguish retrieval failures from answer-generation failures. - [PubMedQA and BioASQ measure different parts of a biomedical evidence pipeline](https://medevals.ai/guides/pubmedqa-bioasq-comparison/): Compare abstract interpretation with literature retrieval and answer synthesis without inventing a combined leaderboard. ## Inspect the evidence - [Evidence JSON](https://medevals.ai/evidence.json): Task definitions, dataset facts, scoring rules, source-version results, our interpretations, and reference IDs. - [Sources](https://medevals.ai/sources/): Original papers and repositories with evidence locators. - [Editorial method](https://medevals.ai/methodology/): Source reconciliation and interpretation boundaries. - [About](https://medevals.ai/about/): Ownership and corrections. Analysis updated: 2026-09-28