{"id":"5584edb1-383f-48a5-ab1a-053421389afb","arxiv_id":"2508.20554","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The 2025 BioASQ challenge overview reports participation and preliminary results for six biomedical NLP shared tasks, including four new benchmarks.","lead":"This paper reports the setup and results of the BioASQ 2025 challenge, which ran six biomedical language tasks with 83 teams and over 1,000 submissions. It introduces four new shared tasks, including multilingual clinical summarization and Greek clinical coding, and summarizes how well participating systems performed.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Task 13b results are explicitly preliminary and the paper's headline benchmark claim rests on them; re-verification against final results is the decisive check.","rationale":"The reader's weakest assumption is that reported performances are accurate and final; the paper itself says they are preliminary for task 13b. This is exactly the load-bearing issue. The challenge overview's usefulness as a benchmark depends on the final results, and the paper currently presents conclusions drawn from provisional scores. I do not see a stronger concern: the participation figures are internally consistent across the abstract, Section 3, and the task sections; the four new tasks report final evaluation numbers with tables; and the dataset description inconsistencies (e.g., MultiClinSum 1,280/534/200/200 versus Table 2 rows that sum differently, and the 'six languages' claim that does not match the MultilingualClinicalSummarization task description which mentions English, Spanish, French, Portuguese, Russian, Greek) are present but do not undercut the central participation/benchmark claim. The condition is precisely the one the reader identified: finalize and validate the 13b results before treating them as benchmark results, and fix the descriptive inconsistencies. Hence CONDITIONAL remains the right verdict, and agreement is 'agree' rather than 'partial' because the reader's weakest assumption is the same one I would flag.","tokens_in":21882,"tokens_out":1705,"duration_ms":15164,"concrete_test":"Check the official BioASQ 13b results pages after ground-truth enrichment and manual assessment are complete. Recompute the top-system metrics for Phase A, A+, and B, and the per-question-type comparisons in Figure 1, using the final ground-truth files. If any top-system ranking changes, or if the gap to 12b changes by more than a few points in MRR or F1, then the claim that 13b systems achieved scores comparable or higher to 12b needs revision. Also verify that the final Synergy 13 results use the same per-round answer-ready sets, or re-run the comparison on a fixed question set if a subset was finalized.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim that the BioASQ 2025 datasets and results provide a valid benchmark, and that systems achieved competitive performance, rests heavily on Section 4.1. The paper states the task 13b results are preliminary, pending manual assessment and ground-truth enrichment with additional documents, snippets, answer elements, and synonyms. The reported task 13b Phase A and exact-answer scores, including Figure 1's claim that top systems matched or exceeded 12b, are therefore provisional. If enrichment adds accepted answers or synonyms, MRR and F1 scores for factoid and list questions will shift, and rankings may move. This is not a mere plotting issue: the paper's own conclusion about 'high scores, especially for yes/no answer generation' and 'less consistent' factoid/list performance in Phase A+ is directly drawn from this preliminary data. The same caveat applies to Synergy 13 results, which rely on expert feedback and answer-ready question subsets that were assessed during the rounds; Table 8's declining Top MAP and snippet F1 across rounds could reflect changing question sets or expert assessment timing. A second-order concern is that Figure 1 and Tables 7-8 present top scores without significance testing, error bars, or interval estimates for batch/round effects, so even final scores would need a paired or bootstrap analysis to support claims of improvement over 12b. The concern is not that the organizers are wrong, but that the paper publishes headline results before the official evaluation is final; this is internally acknowledged rather than hidden.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents the thirteenth edition of the BioASQ challenge, held at CLEF 2025, covering six shared tasks: the established biomedical QA task 13b and Synergy 13, and four new tasks (MultiClinSum, BioNNE-L, ELCardioCC, GutBrainIE). For each task, the authors describe the dataset, evaluation measures, participant systems, and reported results, and they conclude that participation increased and that several systems achieved competitive performance across tasks and languages.","tokens_in":22256,"tokens_out":5028,"duration_ms":43747,"significance":"If the reported results hold, the paper provides a useful record of the BioASQ 2025 benchmark resources and of the state of the art in six biomedical NLP tasks spanning six languages, three document types, and specialized domains such as cardiology and gut-brain interaction. The new tasks (MultiClinSum, BioNNE-L, ELCardioCC, GutBrainIE) introduce publicly available datasets and baselines that are likely to be reused by the community. The paper is candid in Section 4.1 that the 13b results are preliminary, which is appropriate for a challenge overview, and the system results come from independent competing teams measured against externally grounded resources such as UMLS, PubMed, and expert annotations. The main weaknesses are internal inconsistencies in the reported dataset statistics and the presentation of preliminary 13b results as a firm conclusion in the abstract and conclusions.","major_comments":[{"comment":"The abstract states that \"several participating systems achieved competitive performance\" and the conclusions draw on the 13b scores, but Section 4.1 explicitly states that the 13b results are preliminary, pending manual assessment and ground-truth enrichment that may add documents, snippets, answer elements, and synonyms. Because the benchmark-validity claim rests on these results, the abstract and the conclusions should either be updated to reflect the final results or explicitly state that the 13b scores, including Figure 1, are preliminary and that rankings and scores may change after enrichment. As written, the paper overstates the firmness of its headline results.","section":"Abstract and Section 4.1"},{"comment":"The MultiClinSum dataset statistics are internally inconsistent. The text says the gold standard comprises 1,280 English, 534 Spanish, 200 Portuguese, and 200 French pairs, and that translation yields 1,976 pairs per language, but Table 2 reports 988, 988, 1061, and 1034 pairs for the gold-standard sub-tracks and 28.902 for the large-scale sub-tracks. These numbers do not match either the stated originals or the stated post-translation totals, and the table also duplicates the sub-track label \"MultiClinSum-ls-es\" in the EN row. The dataset is a core contribution of the paper, so the counts need to be reconciled and corrected.","section":"Section 2.3, Table 2"},{"comment":"The comparison between 13b and 12b in Figure 1 and the batch-level observations in Table 7 are presented without error bars, confidence intervals, or significance testing. Since the scores are single top-system values on small batches (85 questions each), the claims that \"the top systems achieved scores comparable or higher to those of 12b\" and that the last batch is more challenging are not statistically supported. The authors should either add interval estimates or hedge the wording to reflect that these are point estimates from a single run of the evaluation.","section":"Section 4.1, Figure 1, Tables 7-8"}],"minor_comments":[{"comment":"The MultiClinSum conclusion mentions \"English and Italian text,\" but Italian is not among the task languages (English, Spanish, French, Portuguese); this appears to be an error and should be corrected.","section":"Section 5"},{"comment":"The large-scale row for English is labeled \"MultiClinSum-ls-es EN\"; the sub-track label should likely be \"MultiClinSum-ls-en,\" and the duplicate \"ls-es\" label should be fixed. In addition, the count \"28.902\" should use a thousands separator (28,902) for readability.","section":"Table 2"},{"comment":"Footnote 12 contains a URL with a space (\"BioNNE-L Shared Task\"); it should be percent-encoded or replaced with a stable link.","section":"Section 2.4"},{"comment":"There is a typo in the task name \"GrutBrainIE\" (should be \"GutBrainIE\").","section":"Section 5"},{"comment":"The phrase \"registered 17 teams submitting runs\" is ambiguous; it likely means 17 teams registered and submitted runs. Please rephrase for clarity.","section":"Section 3.6"}],"recommendation":"major_revision","confidential_remarks":"The paper is a challenge overview rather than a novel-methodology paper, which is appropriate for the venue. The main revisions needed are factual: reconciling the MultiClinSum numbers and aligning the abstract/conclusions with the explicitly preliminary 13b results. The Table 2 inconsistency is the kind of error that could mislead readers who use the dataset, so it should be fixed before acceptance. The self-evaluation aspect is standard for challenge overviews and does not concern me."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The short version: this year's BioASQ overview is a functional reference document, and the four new tasks are a genuine contribution to biomedical NLP benchmarking. The one thing to remember before citing it: the Task 13b numbers in Section 4.1 are explicitly preliminary, and the abstract's confident \"competitive performance\" line does not carry that caveat. That is not a hidden flaw; the paper says so. But it means anyone using those scores should wait for the final results or treat them as provisional.\n\nWhat is actually new: four shared tasks—MultiClinSum (multilingual clinical summarization), BioNNE-L (nested entity linking in RU/EN), ELCardioCC (Greek ICD-10 coding), GutBrainIE (gut-brain information extraction). The datasets and baselines for these are new, and the paper does a decent job of laying out the statistics and the evaluation setting. The participation numbers (83 teams, 1000+ runs) are useful, and the overview points to separate task-specific papers for deeper detail, which is the right division of labor. Credit is due for being transparent about the BERTScore limitation in non-English MultiClinSum subtracks and about the ongoing ground-truth enrichment for 13b.\n\nSoft spots, in order. First, Table 2 does not match the text: the prose says 1,280/534/200/200 native pairs for EN/ES/PT/FR, but the table lists 988/988/1034/1061 for the gold-standard tracks. Someone misread a spreadsheet. Second, the conclusions paragraph attributes cardiology and Italian subtracks to MultiClinSum, which is simply wrong—cardiology is ELCardioCC, and Italian is not in this year's tasks. That kind of copy-paste errors makes me less willing to trust the prose without checking the tables. Third, the abstract and conclusions describe 13b performance as 'competitive' without the preliminary qualifier; the body is honest, the summary overstates. A careful revision should fix all three. Minor: no significance tests or intervals on the top-system scores, so the Figure 1 comparisons with 12b are indicative, not conclusive. Also, the Synergy round-over-round numbers are hard to interpret because the answer-ready question set changes each round.\n\nThe core contribution—new benchmarks that the community can reuse—holds up. The inconsistencies are editorial, not methodological, and the preliminary-results concern is explicitly acknowledged. This paper is for people who want a map of the 2025 BioASQ landscape and citations for the new datasets. It deserves a serious peer review; I would send it out with a request to fix the data-reporting errors and to align the abstract with the preliminary caveat.","headline":"BioASQ 2025 overview is a genuinely useful benchmark paper; the 13b scores are preliminary and the MultiClinSum table counts don't match the text, so treat it as a solid draft, not a final reference.","tokens_in":22778,"tokens_out":3962,"would_cite":true,"duration_ms":34685,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The thirteenth BioASQ challenge benchmarks biomedical AI across six tasks and six languages.","keywords":["biomedical question answering","semantic indexing","shared tasks","clinical summarization","entity linking","information extraction","multilingual NLP","benchmark datasets"],"falsifier":"Re-run the final, post-enrichment evaluation of task 13b once the ground truth is completed, or compare the preliminary leaderboard with the published final one; if the added synonyms and answer elements move top systems' MRR or macro-F1 by more than a few points, or if any top run fails to reproduce on the released test set, the conclusion that 2025 systems are competitive with 2024 would need revision.","tokens_in":21714,"feed_emoji":"🧬","tokens_out":4816,"duration_ms":41579,"temperature":0.7,"pith_summary":"This paper reports on the thirteenth BioASQ challenge, a set of six shared tasks that test how well computer systems can find, summarize, and structure biomedical information. The challenge asks whether current natural-language systems can answer expert biomedical questions, index the scientific literature, summarize clinical case reports in four languages, link nested medical terms in English and Russian, assign cardiology ICD-10 codes to Greek discharge letters, and extract entities and relations from gut-brain research abstracts. The paper's claim is that the 2025 edition produced valid benchmark datasets and that 83 teams with over 1000 submissions showed performance at or above the level of previous years, so the state of the art continues to advance. The value of the claim, if correct, is that these datasets become reference points for measuring biomedical NLP progress and for showing where the largest difficulties lie.","feed_headline":"Biomedical AI meets six languages and six new benchmarks","feed_subtitle":"Four new multilingual tasks join the long-running QA benchmark; results show where LLMs help and where they stall.","key_machinery":"The carrying object is the challenge itself: a collection of six shared tasks, each with a labelled dataset, an evaluation measure, and a baseline system. The evaluation stack combines automatic metrics (MAP and F1 for retrieval, MRR for factoid answers, macro-F1 for yes/no, BERTScore and ROUGE-LSum for summarization, Accuracy@k and MRR for entity linking, micro-F1 for coding and information extraction) with manual expert assessment for the open-ended ideal answers of tasks 13b and Synergy 13. These measures convert each task into comparable leaderboards, and the baselines provide the bar systems must beat.","core_discovery":"In the thirteenth edition, BioASQ ran two established tasks (13b and Synergy 13) and four new ones (MultiClinSum, BioNNE-L, ELCardioCC, GutBrainIE), producing gold-standard datasets in six languages. According to the reported results, the best systems matched or exceeded the previous edition on task 13b, with notably good yes/no answering even when no relevant documents were supplied, while the new tasks showed that multilingual clinical summarization works best in English and that nested entity linking is best handled by domain-specific biomedical BERT-style retrieval rather than general LLMs. The paper also reports that in the gut-brain information extraction task, named-entity recognition reached a top micro-F1 of 0.84, whereas the hardest subtask, mention-based relation extraction, reached only 0.46, so the paper's conclusion that several participating systems achieved competitive performance holds task by task.","pith_inferences":["The near-parity of Phase A+ and Phase B yes/no scores suggests retrieval quality may matter least for binary questions; a clean ablation would be to feed a top system random documents and measure the drop.","Because MultiClinSum rows were machine-translated across languages, the non-English leaderboards measure translation-plus-summarization jointly; a future edition could score only native-language texts to separate the two.","The Greek cardiology task's effective mBERT baseline hints that a simple fine-tuned multilingual encoder is a difficult target for clinical coding in low-resource languages, which other languages could adopt as a cheap baseline.","The BioNNE-L dictionary sizes, about 1.8 million English UMLS concepts versus 92 thousand Russian ones, make the bilingual track as much a test of vocabulary coverage as of linking, so cross-lingual gains could be analysed by stratifying on whether the English concept exists in Russian."],"forward_implications":["The six released datasets become public benchmarks, so future systems can be compared on the same questions, discharge letters, and abstracts.","The yes/no result suggests that for binary questions, capable LLMs may not need retrieved golden documents, which could change how QA pipelines are built.","Domain-specific biomedical encoders, not general LLMs, carried the entity-linking task, indicating that pre-training on UMLS concepts still matters.","The 0.84 versus 0.46 gap in GutBrainIE shows where the next bottleneck is: locating and typing the exact mention pairs in a relation, not finding entities.","Multilingual summarization results were best in English, so language-specific data remains a limiting factor even within one clinical domain."],"supporting_citations":[{"why":"defines the six-task structure of the 2025 edition and frames the challenge.","marker":"[50]"},{"why":"reports the task 13b and Synergy 13 datasets, participation, and preliminary results.","marker":"[55]"},{"why":"supplies the MultiClinSum corpus and its evaluation.","marker":"[66]"},{"why":"describes the BioNNE-L nested entity linking data and dictionary.","marker":"[67]"},{"why":"defines the ELCardioCC Greek cardiology coding task and its baselines.","marker":"[18]"},{"why":"provides the GutBrainIE dataset and the NER and relation extraction baselines.","marker":"[48]"},{"why":"introduces the BioASQ framework that this edition extends.","marker":"[77]"},{"why":"provides the training corpus used for task 13b system development.","marker":"[34]"},{"why":"is the previous-edition overview against which 13b progress is compared.","marker":"[52]"}],"fun_headline_variants":["BioASQ 2025: Six tasks, multilingual, mixed results","BioASQ 13th: New multilingual tasks, LLMs still struggle on relations","BioASQ 2025: Best yes/no answers, weakest relation extraction","Four new BioASQ tasks span six languages, performance varies","BioASQ 2025: English summarization leads, BERT edges out LLMs for entities"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The reported scores are accurate and final; the paper itself states that task 13b results are preliminary, pending manual assessment and enrichment of the ground truth with additional synonyms and answer elements, so the rankings and the conclusion about competitive performance could change.","fun_headline_variants_meta":{"raw":{"variants":["BioASQ 2025: Six tasks, multilingual, mixed results","BioASQ 13th: New multilingual tasks, LLMs still struggle on relations","BioASQ 2025: Best yes/no answers, weakest relation extraction","Four new BioASQ tasks span six languages, performance varies","BioASQ 2025: English summarization leads, BERT edges out LLMs for entities"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000299,"raw_usage":{"total_tokens":1711,"prompt_tokens":911,"completion_tokens":800,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":527,"completion_tokens_details":{"reasoning_tokens":696}},"tokens_in":527,"tokens_out":800,"duration_ms":6234,"temperature":1.0,"reasoning_tokens":696,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T16:42:14.272756+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the final, post-enrichment evaluation of task 13b once the ground truth is completed, or compare the preliminary leaderboard with the published final one; if the added synonyms and answer elements move top systems' MRR or macro-F1 by more than a few points, or if any top run fails to reproduce on the released test set, the conclusion that 2025 systems are competitive with 2024 would need revision.","supporting_citations":[{"cited_title":"In: Arampatzis, A., Kanoulas, E., Tsikrika, T., Vrochidis, S., Giachanou, 24 A","cited_arxiv_id":null,"evidence_quote":"is the previous-edition overview against which 13b progress is compared."},{"cited_title":"In: Faggioli, G., Ferro, N., Rosso, P., Spina, D","cited_arxiv_id":null,"evidence_quote":"provides the GutBrainIE dataset and the NER and relation extraction baselines."},{"cited_title":"In: Findings of the Association for Computational Linguistics: EMNLP 26 A","cited_arxiv_id":null,"evidence_quote":"introduces the BioASQ framework that this edition extends."},{"cited_title":"Scientific Data 10(1), 170 (2023)","cited_arxiv_id":null,"evidence_quote":"provides the training corpus used for task 13b system development."},{"cited_title":"(eds.) CLEF 2025 Working Notes (2025)","cited_arxiv_id":null,"evidence_quote":"defines the six-task structure of the 2025 edition and frames the challenge."},{"cited_title":"In: Faggioli, G., Ferro, N., Galuˇ sˇ c´ akov´ a, P., Garc´ ıa Seco de Herrera, A","cited_arxiv_id":null,"evidence_quote":"reports the task 13b and Synergy 13 datasets, participation, and preliminary results."},{"cited_title":"In: Faggioli, G., Ferro, N., Rosso, P., Spina, D","cited_arxiv_id":null,"evidence_quote":"supplies the MultiClinSum corpus and its evaluation."},{"cited_title":"In: Faggioli, G., Ferro, N., Rosso, P., Spina, D","cited_arxiv_id":null,"evidence_quote":"describes the BioNNE-L nested entity linking data and dictionary."},{"cited_title":"In: Faggioli, G., Ferro, N., Rosso, P., Spina, D","cited_arxiv_id":null,"evidence_quote":"defines the ELCardioCC Greek cardiology coding task and its baselines."}],"review_version":2}