{"id":"65427075-03ee-4d41-ace2-3d78eb20d3c2","arxiv_id":"2508.20532","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The 2024 BioASQ challenge ran four shared tasks, released two new biomedical datasets, and found that transformer-based systems, especially LLM and RAG pipelines, led performance across question answering and named entity recognition.","lead":"This paper reports the results and organization of the 2024 BioASQ shared tasks: four competitions on biomedical question answering and named entity recognition. It is useful as a public record of what was measured, which teams submitted systems, and which methods led the leaderboards.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The abstract's headline counts (37 teams, >700 submissions) are not reconciled with the task-level tallies in Sections 3.1–3.4, which sum to roughly 330 systems/runs; the paper needs a counting convention and a citation for these figures.","rationale":"The reader's CONDITIONAL verdict is sensible, but the identified weakest assumption—that task-12b results are preliminary—is acknowledged in the paper itself in §4.1, so it is not the most consequential hidden soft spot. A more serious unacknowledged issue is the abstract's quantitative claim about the scale of the challenge. Because the strongest claim is that this was a large shared-task evaluation with 37 teams and more than 700 submissions, the accuracy of those numbers is fundamental to the paper's purpose. The body gives task-level counts that sum to roughly 330 systems or runs; even if a multiplier such as per-batch counting is intended, no definition or source is provided. This is a missing-support problem, not a disagreement with consensus, and it is directly checkable. The task-12b preliminary caveat affects only the secondary claim about state-of-the-art improvement, whereas the count issue affects the primary identity of the paper as an accurate record. I therefore keep the reader's CONDITIONAL verdict rather than moving it: independent verification of the official logs could fully resolve the concern, but as written the abstract and the body are not reconciled. If the official logs confirm the abstract, the condition is satisfied and the verdict can move toward acceptance; if not, the abstract should be revised.","tokens_in":20364,"tokens_out":7529,"duration_ms":69729,"concrete_test":"Recompute the totals from Sections 3.1–3.4 using the official platform logs: the BioASQ participants-area URLs given in §4.1 for tasks 12b and Synergy, CodaLab competition 16464 for BioNNE, and the MultiCardioNER overview paper [31]. Define distinct submission exactly as the organizers' evaluation system does; if the official number of unique submissions is at least 700, the discrepancy is resolved and no correction is needed. If it is approximately 330 or any other number below 700, the abstract's headline figures are wrong and must be replaced with the verified totals and a definition. Also verify the unique-team count across the four tasks from registration and submission logs; if it is 37, say so; if it is 39 or another number, correct the abstract.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central factual claim in the abstract is that 37 teams participated with more than 700 distinct submissions. The only task-level numbers given are: §3.1 reports 89 different systems for task 12b, §3.2 reports 16 distinct systems for Synergy 12, §3.3 reports 70 runs for MultiCardioNER, and §3.4 reports 155 runs for BioNNE. These four figures sum to 330, not >700. One can invent a reconciliation—for example, counting every task-12b system in each of four batches, or every Synergy system in each of four rounds—but neither the term distinct submissions nor such a multiplier is defined anywhere in the paper, and no external citation is given for the abstract's numbers. The 37-team figure is likewise not derivable from the text: the task-level counts are 26, 4, 7 submitted, and 5 submitted; after the three known overlaps between task 12b and Synergy, the most direct sum gives 39 unique teams, and reaching 37 requires unknown overlaps. Because the paper is first and foremost a record of a large shared-task evaluation, these headline quantities are load-bearing. Unlike the preliminary task-12b results, which are explicitly caveated in §4.1, this discrepancy is unacknowledged. If official challenge logs define distinct submission differently, the paper should state that definition and show the per-task breakdown; otherwise the abstract overstates the scale of the event.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper is the official overview of the twelfth BioASQ challenge run at CLEF 2024. It describes four shared tasks: task 12b (English biomedical question answering with phases A, A+, and B), Synergy 12 (iterative question answering for open developing problems over four rounds), MultiCardioNER (disease and drug named-entity recognition in cardiology clinical texts in Spanish, English, and Italian), and BioNNE (nested named-entity recognition in English and Russian). For each task it gives corpus statistics, participant counts, system descriptions, and evaluation results, with full results made available online. The abstract's headline quantitative claim is that 37 teams submitted more than 700 distinct submissions across the four tasks, and the conclusions frame the results as continued advancement of the state of the art.","tokens_in":20651,"tokens_out":8164,"duration_ms":71528,"significance":"The paper's value is as an archival record of a large shared-task evaluation. It is not a methods paper; its main contributions are the task definitions, dataset releases, participant and system inventories, and aggregated results, with URLs to data repositories and full result pages. The writing is generally clear, and the per-task descriptions are internally consistent. I credit the authors for explicitly labeling the task 12b results as preliminary in Section 4.1 and for acknowledging in Section 4.3 that the observed benefit of the CardioCCC data may reflect data volume rather than domain adaptation. If the participation-count discrepancy is resolved and the preliminary-result caveat is carried into the abstract and conclusions, the paper would serve as a reliable reference for the community.","major_comments":[{"comment":"The abstract states that 37 competing teams participated with more than 700 distinct submissions, but the task-level tallies in Sections 3.1-3.4 (89 systems for task 12b, 16 systems for Synergy 12, 70 runs for MultiCardioNER, and 155 runs for BioNNE) sum to 330, and the text does not define how a \"distinct submission\" is counted. If submissions are counted multiplicatively, for example per phase, round, or language, that convention needs to be stated explicitly and a per-task breakdown should be provided; otherwise the headline figures are not reproducible from the paper. The 37-team figure is likewise not derivable from the text: the per-task counts are 26 + 4 + 7 + 5 = 42, and with only the overlapping teams named in the paper one arrives at 39 rather than 37. Please add a summary table with the counting convention for both teams and submissions, and cite the official challenge logs for these figures.","section":"Abstract; Sections 3.1-3.4"},{"comment":"The task 12b results are explicitly preliminary: Section 4.1 says that final results depend on the ongoing manual assessment of system responses and the enrichment of the ground truth. Despite this, the abstract claims \"continuous advancement of the state-of-the-art\" and Section 5 asserts that top-performing systems \"were able to improve over the state-of-the-art performance from previous years.\" These claims should be explicitly qualified as based on preliminary results, or deferred until the final task 12b scores are available, since the manual assessment could change the reported rankings and conclusions.","section":"Section 4.1; Section 5; Abstract"}],"minor_comments":[{"comment":"The sentence \"The full 12b results are available online\" in the Synergy results subsection should read \"The full Synergy 12 results are available online,\" since the preceding paragraph concerns the Synergy task.","section":"Section 4.2"},{"comment":"The conclusion says BioASQ has been pushing the research frontier \"for eleven years now,\" but the paper describes the twelfth edition of the challenge; this should be corrected to twelve years or rephrased as \"since 2013.\"","section":"Section 5"},{"comment":"In the evaluation metric formula, F1rel_c is used without a clear definition; the text says it is the macro F1-score averaged across all relevance classes, but it would be clearer to state that F1rel_c is the per-class relevance F1 and that the formula is the macro-average over the eight entity classes.","section":"Section 4.4"},{"comment":"References [16] and [35] contain stray commas in the author lists; please check and normalize the reference formatting.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The participation-count discrepancy is the main substantive issue. If the authors can provide the official counting convention and correct the abstract, I would support acceptance after revision; the rest of the paper is a solid and useful overview."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nYou should know this paper is a competent yearly record, not a research contribution. It documents BioASQ 2024: four tasks, including two genuinely new shared tasks (MultiCardioNER for multilingual clinical NER in cardiology, BioNNE for nested NER in English/Russian), a new phase A+ in task 12b, and new datasets (CardioCCC, the BioNNE test set, and the multilingual DrugTEMIST release). The task descriptions are clear, the summary tables are internally consistent, and the authors are explicit that the task 12b results are preliminary pending manual assessment and ground-truth enrichment (Section 4.1). They also point to companion overview papers for full results. For readers who want a snapshot of what biomedical NLP systems could do in 2024, this is a useful reference.\n\nThe soft spot is the abstract's headline arithmetic. It claims 37 competing teams and 'more than 700 distinct submissions in total.' The body gives per-task counts: 89 systems for task 12b, 16 for Synergy 12, 70 runs for MultiCardioNER, 155 runs for BioNNE—summing to 330. The team counts are 26, 4, 7, and 5; even after the three known overlaps between 12b and Synergy, the most direct sum gives 39 unique teams, not 37. You can invent a reconciliation—counting each batch or round as a separate submission, say—but the paper never defines the counting unit, and the abstract says 'distinct submissions,' which is the same word the body uses for 'systems' in 12b. This is a blemish on the paper's primary purpose as an accurate record of scale. It is not a fatal flaw: the per-task numbers are in the text, and the discrepancy is visible to anyone who checks. But it should be fixed before publication, either by matching the abstract to the body or by stating the convention.\n\nOther soft spots are minor. Only top-6 results are shown for several tasks, so the claim that top systems improved over previous years rests partly on external pages and omitted runs. That is acceptable practice for an overview, but it means the paper's headline conclusions are not fully self-contained. The heavy self-citation is normal for a series overview and not a problem here.\n\nWho is this for? Organizers, participants, and anyone using the released datasets or tracking biomedical NLP evaluation. It deserves a serious referee—shared-task overviews are archival documents that need checking—and the reviewer's main job should be to make the counts consistent.\n\nMy recommendation: accept subject to a minor revision that reconciles the abstract's numbers with the task-level tallies.","headline":"A solid, honest overview of the 2024 BioASQ edition with two real additions (phase A+ and two new tasks), undercut by an abstract whose participation totals don't add up against the tables in the body.","tokens_in":21234,"tokens_out":3260,"would_cite":true,"duration_ms":28399,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The 2024 BioASQ challenge ran four shared tasks with 37 teams and over 700 submissions, and the top systems—built almost entirely on LLMs and transformers—continued the series' upward trend in biomedical question answering and clinical…","keywords":["biomedical question answering","semantic indexing","shared task evaluation","named entity recognition","large language models","Retrieval Augmented Generation","multilingual clinical NLP","BioASQ 2024"],"falsifier":"A re-evaluation of task 12b after the ground-truth enrichment is complete would settle the point: if the final yes/no macro-F1 scores drop materially or the system ranking changes, the claim of continued state-of-the-art progress would not survive. Likewise, an ablation training MultiCardioNER models on the same total number of documents but with the cardiology-specific corpus replaced by equal-sized mixed-specialty text would show whether the recall gain is due to domain adaptation or just to more data.","tokens_in":20203,"feed_emoji":"🧬","tokens_out":7150,"duration_ms":64883,"temperature":0.7,"pith_summary":"This paper is the official overview of the twelfth BioASQ challenge, a shared-task evaluation series for biomedical semantic indexing and question answering. It reports that four tasks ran in 2024: the established task b and Synergy, plus two new tasks, MultiCardioNER (multilingual detection of diseases and drugs in cardiology case reports) and BIONNE (nested named-entity recognition in Russian and English biomedical abstracts), with 37 participating teams and more than 700 total submissions. The paper's central claim is that the top systems, almost all built on large language models and transformer architectures, kept raising the state of the art, with particularly strong performance on yes/no questions and with the two new tasks generating reusable benchmark datasets. A sympathetic reader would care because the results define the current capability baseline for biomedical question answering and clinical named-entity recognition, including multilingual and domain-adaptation settings.","feed_headline":"BioASQ 2024: LLM systems lead four biomedical shared tasks","feed_subtitle":"New benchmarks test multilingual clinical drug detection and Russian-English nested biomedical entity recognition.","key_machinery":"The central machinery is the shared-task benchmark design: each task combines a public dataset, a scoring protocol, and expert assessment, so that systems are compared on identical inputs. For task 12b the measures are MAP for documents, character-overlap F-measure for snippets, F1/MRR/macro-F1 for exact answers, and manual expert scores for ideal answers; phase A+ isolates answer generation from retrieval. For MultiCardioNER, the key object is the CardioCCC corpus (508 cardiology clinical case reports, 250 held out for testing), used alongside DisTEMIST and DrugTEMIST to test domain adaptation. For BioNNE, the key object is the nested-entity dataset built from NEREL-BIO, with eight biomedical entity types and a macro-F1 metric averaged over classes. These datasets and measures are what let the paper claim meaningful comparisons.","core_discovery":"The paper claims that the 2024 edition of the challenge demonstrates continued improvement in biomedical question-answering systems, especially on yes/no questions, where top systems approached or reached perfect macro-F1 on some test batches, while factoid and list questions remain harder and more variable. It claims that the new phase A+ shows systems can generate competitive exact and ideal answers without being given manually selected relevant documents, and that providing such documents in phase B still improves answer quality. For the new MultiCardioNER task, the paper argues that incorporating cardiology-specific training data (the CardioCCC corpus) is the decisive factor for disease detection, since systems trained only on mixed-specialty clinical text achieve high precision but lower recall on cardiology-specific entities; drug detection performance is higher overall and fairly comparable across languages, with Italian somewhat lower because fewer clinical pretrained models exist. For BIONNE, the paper claims that a bi-encoder contrastive model fine-tuned on the supplied data clearly outperforms a zero-shot LLM-based extractor, indicating that specialized training data is necessary for nested biomedical named-entity recognition. The paper also reports that the Synergy iteration process enabled experts to reach answers for about 78% of the open questions, with about 51% receiving at least one ideal answer judged to be of ground-truth quality.","pith_inferences":["Because the task 12b scores are explicitly preliminary, the ongoing ground-truth enrichment may shift the reported rankings; treating the specific numbers as authoritative would overread the paper.","The MultiCardioNER result suggests an ablative explanation the paper leaves open: the benefit of CardioCCC might come from simply having more training instances rather than from domain adaptation, and a controlled data-volume experiment would separate the two.","The BioNNE result invites a follow-up test the paper does not run: fine-tuning an LLM on the same nested-entity training data would tell whether the gap is due to model architecture or to the absence of supervised signal."],"forward_implications":["Top 12b systems reached perfect or near-perfect macro-F1 on yes/no questions in some batches, while factoid and list questions still showed inconsistent performance.","Phase A+ results showed that systems can produce competitive answers without manually selected relevant documents, but releasing such documents in phase B still improved answer quality.","In MultiCardioNER, systems that incorporated the cardiology-specific CardioCCC corpus clearly outperformed those using only mixed-specialty clinical data, implying that domain-specific training data drives recall on specialty entities.","In BioNNE, the fine-tuned bi-encoder model scored 0.7044 F1 on the bilingual test set, far above the 0.3479 of a zero-shot LLM-based system, implying that nested biomedical named-entity recognition needs supervised training data.","The Synergy task reached answer-ready status for about 78% of the 73 open questions, and roughly 51% received at least one ideal answer judged ground-truth quality."],"supporting_citations":[{"why":"It establishes the BioASQ challenge frame of large-scale biomedical semantic indexing and question answering that this edition extends.","marker":"[50]"},{"why":"It provides the task-specific details and results for tasks 12b and Synergy12 that this overview summarizes.","marker":"[37]"},{"why":"It defines the MultiCardioNER task, its datasets, and its results, including the observed effect of the CardioCCC corpus.","marker":"[31]"},{"why":"It defines the BioNNE task and reports the dataset construction and participating-system results for nested biomedical NER.","marker":"[12]"},{"why":"It supplies the manually curated BioASQ-QA corpus used as the training set for task 12b.","marker":"[24]"},{"why":"It is the ancestor DisTEMIST task whose annotation guidelines and corpus underpin the MultiCardioNER disease-detection subtrack.","marker":"[35]"},{"why":"It provides the NEREL-BIO dataset from which the BioNNE training and validation data were derived.","marker":"[34]"},{"why":"It defines the manual scoring protocol for ideal answers and the exact-answer measures used across the QA tasks.","marker":"[6]"}],"fun_headline_variants":["BioASQ 2024: 37 teams, 700+ runs, yes/no QA near perfect","BioASQ 2024: New tasks target cardiology NER and nested entities","BioASQ 2024: Phase A+ answers without curated documents","BioASQ 2024: Factoid questions still tough for top systems","BioASQ 2024: Specialized data beats zero-shot LLMs for NER"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper's conclusions rely on treating the preliminary task 12b scores as valid evidence of system quality even though the ground truth is still being actively enriched by manual assessment.","fun_headline_variants_meta":{"raw":{"variants":["BioASQ 2024: 37 teams, 700+ runs, yes/no QA near perfect","BioASQ 2024: New tasks target cardiology NER and nested entities","BioASQ 2024: Phase A+ answers without curated documents","BioASQ 2024: Factoid questions still tough for top systems","BioASQ 2024: Specialized data beats zero-shot LLMs for NER"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000271,"raw_usage":{"total_tokens":1642,"prompt_tokens":970,"completion_tokens":672,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":586,"completion_tokens_details":{"reasoning_tokens":561}},"tokens_in":586,"tokens_out":672,"duration_ms":5621,"temperature":1.0,"reasoning_tokens":561,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T16:42:21.705539+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A re-evaluation of task 12b after the ground-truth enrichment is complete would settle the point: if the final yes/no macro-F1 scores drop materially or the system ranking changes, the claim of continued state-of-the-art progress would not survive. Likewise, an ablation training MultiCardioNER models on the same total number of documents but with the cardiology-specific corpus replaced by equal-sized mixed-specialty text would show whether the recall gain is due to domain adaptation or just to more data.","supporting_citations":[{"cited_title":"BMC Bioinformatics 16, 138 (2015)","cited_arxiv_id":null,"evidence_quote":"It establishes the BioASQ challenge frame of large-scale biomedical semantic indexing and question answering that this edition extends."},{"cited_title":"In: Faggioli, G., Ferro, N., Galuˇ sˇ c´ akov´ a, P., Garc´ ıa Seco de Herrera, A","cited_arxiv_id":null,"evidence_quote":"It defines the MultiCardioNER task, its datasets, and its results, including the observed effect of the CardioCCC corpus."},{"cited_title":"In: CLEF Working Notes (2024)","cited_arxiv_id":null,"evidence_quote":"It defines the BioNNE task and reports the dataset construction and participating-system results for nested biomedical NER."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It is the ancestor DisTEMIST task whose annotation guidelines and corpus underpin the MultiCardioNER disease-detection subtrack."},{"cited_title":"Project deliverable D4.1, UPMC (2013)","cited_arxiv_id":null,"evidence_quote":"It defines the manual scoring protocol for ideal answers and the exact-answer measures used across the QA tasks."}],"review_version":2}