Pith. sign in

REVIEW 3 major objections 5 minor 1 cited by

Overview of BioASQ 2025: The Thirteenth BioASQ Challenge on Large-Scale Biomedical Semantic Indexing and Question Answering

T0 review · 3 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read The thirteenth BioASQ challenge benchmarks biomedical AI across six tasks and six languages.

desk verdict BioASQ 2025 overview is a genuinely useful benchmark paper; the 13b scores are preliminary and the MultiClinSum table counts don't match the text, so treat it as a solid draft, not a final reference. read the letter →

arxiv 2508.20554 v1 pith:QZC736TG submitted 2025-08-28 cs.CL cs.AIcs.IR

classification cs.CLcs.AIcs.IR
keywords biomedicalquestionansweringsemanticindexingsharedtasksclinicalsummarizationentitylinkinginformationextractionmultilingualNLPbenchmarkdatasets
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper reports on the thirteenth BioASQ challenge, a set of six shared tasks that test how well computer systems can find, summarize, and structure biomedical information. The challenge asks whether current natural-language systems can answer expert biomedical questions, index the scientific literature, summarize clinical case reports in four languages, link nested medical terms in English and Russian, assign cardiology ICD-10 codes to Greek discharge letters, and extract entities and relations from gut-brain research abstracts. The paper's claim is that the 2025 edition produced valid benchmark datasets and that 83 teams with over 1000 submissions showed performance at or above the level of previous years, so the state of the art continues to advance. The value of the claim, if correct, is that these datasets become reference points for measuring biomedical NLP progress and for showing where the largest difficulties lie.

What carries the argument

The carrying object is the challenge itself: a collection of six shared tasks, each with a labelled dataset, an evaluation measure, and a baseline system. The evaluation stack combines automatic metrics (MAP and F1 for retrieval, MRR for factoid answers, macro-F1 for yes/no, BERTScore and ROUGE-LSum for summarization, Accuracy@k and MRR for entity linking, micro-F1 for coding and information extraction) with manual expert assessment for the open-ended ideal answers of tasks 13b and Synergy 13. These measures convert each task into comparable leaderboards, and the baselines provide the bar systems must beat.

What would settle it

Re-run the final, post-enrichment evaluation of task 13b once the ground truth is completed, or compare the preliminary leaderboard with the published final one; if the added synonyms and answer elements move top systems' MRR or macro-F1 by more than a few points, or if any top run fails to reproduce on the released test set, the conclusion that 2025 systems are competitive with 2024 would need revision.

Watch

Extended reading notes

Core claim

In the thirteenth edition, BioASQ ran two established tasks (13b and Synergy 13) and four new ones (MultiClinSum, BioNNE-L, ELCardioCC, GutBrainIE), producing gold-standard datasets in six languages. According to the reported results, the best systems matched or exceeded the previous edition on task 13b, with notably good yes/no answering even when no relevant documents were supplied, while the new tasks showed that multilingual clinical summarization works best in English and that nested entity linking is best handled by domain-specific biomedical BERT-style retrieval rather than general LLMs. The paper also reports that in the gut-brain information extraction task, named-entity recognition reached a top micro-F1 of 0.84, whereas the hardest subtask, mention-based relation extraction, reached only 0.46, so the paper's conclusion that several participating systems achieved competitive performance holds task by task.

Load-bearing premise

The reported scores are accurate and final; the paper itself states that task 13b results are preliminary, pending manual assessment and enrichment of the ground truth with additional synonyms and answer elements, so the rankings and the conclusion about competitive performance could change.

Editorial extensions

If this is right

  • The six released datasets become public benchmarks, so future systems can be compared on the same questions, discharge letters, and abstracts.
  • The yes/no result suggests that for binary questions, capable LLMs may not need retrieved golden documents, which could change how QA pipelines are built.
  • Domain-specific biomedical encoders, not general LLMs, carried the entity-linking task, indicating that pre-training on UMLS concepts still matters.
  • The 0.84 versus 0.46 gap in GutBrainIE shows where the next bottleneck is: locating and typing the exact mention pairs in a relation, not finding entities.
  • Multilingual summarization results were best in English, so language-specific data remains a limiting factor even within one clinical domain.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The near-parity of Phase A+ and Phase B yes/no scores suggests retrieval quality may matter least for binary questions; a clean ablation would be to feed a top system random documents and measure the drop.
  • Because MultiClinSum rows were machine-translated across languages, the non-English leaderboards measure translation-plus-summarization jointly; a future edition could score only native-language texts to separate the two.
  • The Greek cardiology task's effective mBERT baseline hints that a simple fine-tuned multilingual encoder is a difficult target for clinical coding in low-resource languages, which other languages could adopt as a cheap baseline.
  • The BioNNE-L dictionary sizes, about 1.8 million English UMLS concepts versus 92 thousand Russian ones, make the bilingual track as much a test of vocabulary coverage as of linking, so cross-lingual gains could be analysed by stratifying on whether the English concept exists in Russian.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. This paper presents the thirteenth edition of the BioASQ challenge, held at CLEF 2025, covering six shared tasks: the established biomedical QA task 13b and Synergy 13, and four new tasks (MultiClinSum, BioNNE-L, ELCardioCC, GutBrainIE). For each task, the authors describe the dataset, evaluation measures, participant systems, and reported results, and they conclude that participation increased and that several systems achieved competitive performance across tasks and languages.

Significance. If the reported results hold, the paper provides a useful record of the BioASQ 2025 benchmark resources and of the state of the art in six biomedical NLP tasks spanning six languages, three document types, and specialized domains such as cardiology and gut-brain interaction. The new tasks (MultiClinSum, BioNNE-L, ELCardioCC, GutBrainIE) introduce publicly available datasets and baselines that are likely to be reused by the community. The paper is candid in Section 4.1 that the 13b results are preliminary, which is appropriate for a challenge overview, and the system results come from independent competing teams measured against externally grounded resources such as UMLS, PubMed, and expert annotations. The main weaknesses are internal inconsistencies in the reported dataset statistics and the presentation of preliminary 13b results as a firm conclusion in the abstract and conclusions.

major comments (3)
  1. [Abstract and Section 4.1] The abstract states that "several participating systems achieved competitive performance" and the conclusions draw on the 13b scores, but Section 4.1 explicitly states that the 13b results are preliminary, pending manual assessment and ground-truth enrichment that may add documents, snippets, answer elements, and synonyms. Because the benchmark-validity claim rests on these results, the abstract and the conclusions should either be updated to reflect the final results or explicitly state that the 13b scores, including Figure 1, are preliminary and that rankings and scores may change after enrichment. As written, the paper overstates the firmness of its headline results.
  2. [Section 2.3, Table 2] The MultiClinSum dataset statistics are internally inconsistent. The text says the gold standard comprises 1,280 English, 534 Spanish, 200 Portuguese, and 200 French pairs, and that translation yields 1,976 pairs per language, but Table 2 reports 988, 988, 1061, and 1034 pairs for the gold-standard sub-tracks and 28.902 for the large-scale sub-tracks. These numbers do not match either the stated originals or the stated post-translation totals, and the table also duplicates the sub-track label "MultiClinSum-ls-es" in the EN row. The dataset is a core contribution of the paper, so the counts need to be reconciled and corrected.
  3. [Section 4.1, Figure 1, Tables 7-8] The comparison between 13b and 12b in Figure 1 and the batch-level observations in Table 7 are presented without error bars, confidence intervals, or significance testing. Since the scores are single top-system values on small batches (85 questions each), the claims that "the top systems achieved scores comparable or higher to those of 12b" and that the last batch is more challenging are not statistically supported. The authors should either add interval estimates or hedge the wording to reflect that these are point estimates from a single run of the evaluation.
minor comments (5)
  1. [Section 5] The MultiClinSum conclusion mentions "English and Italian text," but Italian is not among the task languages (English, Spanish, French, Portuguese); this appears to be an error and should be corrected.
  2. [Table 2] The large-scale row for English is labeled "MultiClinSum-ls-es EN"; the sub-track label should likely be "MultiClinSum-ls-en," and the duplicate "ls-es" label should be fixed. In addition, the count "28.902" should use a thousands separator (28,902) for readability.
  3. [Section 2.4] Footnote 12 contains a URL with a space ("BioNNE-L Shared Task"); it should be percent-encoded or replaced with a stable link.
  4. [Section 5] There is a typo in the task name "GrutBrainIE" (should be "GutBrainIE").
  5. [Section 3.6] The phrase "registered 17 teams submitting runs" is ambiguous; it likely means 17 teams registered and submitted runs. Please rephrase for clarity.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the paper reports externally grounded measurements of systems from 83 independent teams, and no prediction reduces to a fitted input or to a self-citation.

full rationale

The paper contains no derivation chain of the kind that can be circular. Its central content is a shared-task overview: 83 teams submitted more than 1000 runs, and each task's evaluation is grounded in external resources rather than in the paper's own outputs. Task 13b uses expert-created gold documents, snippets, and answers; BioNNE-L uses UMLS dictionaries and NEREL-BIO annotations; ELCardioCC uses expert-annotated Greek discharge letters with ICD-10 codes; GutBrainIE uses expert-annotated PubMed abstracts; and MultiClinSum compares generated summaries against summaries written by the original case-report authors. The many self-citations, such as [50], [54], [55], [66], [67], [18], and [48], are contextual references to prior editions and to task overviews; they are not load-bearing for the reported system performances. The paper explicitly flags in Section 4.1 that the Task 13b results are preliminary, pending manual assessment and ground-truth enrichment, which is an empirical validity caveat about possibly shifting rankings, not a circularity step: the final scores will still come from independent measurements against enriched gold data. No fitted parameter is renamed as a prediction, no uniqueness theorem from the authors' prior work is invoked to force a choice, and no known empirical result is repackaged under new coordinates. The overview is therefore self-consistent and non-circular.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The paper's conclusions rest on the validity of expert-annotated gold data, standard evaluation measures, and the public resources the tasks are built on. It introduces no fitted parameters and no new theoretical entities.

assumptions (3)
  • domain assumption Expert-annotated gold data are reliable ground truth for evaluation
    Task descriptions assume expert annotations and manual assessment are the reference standard for measuring system performance (Sections 2.3-2.6 and Section 4).
  • standard math Standard evaluation measures are valid proxies for system quality
    The paper computes leaderboards using MAP, F1, MRR, BERTScore, and ROUGE without questioning their validity (Section 4).
  • domain assumption Publicly available resources are comprehensive and correctly used
    BioNNE-L and MultiClinSum rely on UMLS dictionaries and PMC-Patients as source data (Sections 2.3-2.4).

how reviews work

0 comments
Cite this review

Pith. "Pith review of Overview of BioASQ 2025: The Thirteenth BioASQ Challenge on Large-Scale Biomedical Semantic Indexing and Question Answering." pith.science (2026). https://pith.science/paper/QZC736TG

@misc{pith2026250820554,
  author       = {Pith},
  title        = {Pith review of: Overview of BioASQ 2025: The Thirteenth BioASQ Challenge on Large-Scale Biomedical Semantic Indexing and Question Answering},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/QZC736TG}},
  note         = {Machine review of arXiv:2508.20554}
}
read the original abstract

This is an overview of the thirteenth edition of the BioASQ challenge in the context of the Conference and Labs of the Evaluation Forum (CLEF) 2025. BioASQ is a series of international challenges promoting advances in large-scale biomedical semantic indexing and question answering. This year, BioASQ consisted of new editions of the two established tasks, b and Synergy, and four new tasks: a) Task MultiClinSum on multilingual clinical summarization. b) Task BioNNE-L on nested named entity linking in Russian and English. c) Task ELCardioCC on clinical coding in cardiology. d) Task GutBrainIE on gut-brain interplay information extraction. In this edition of BioASQ, 83 competing teams participated with more than 1000 distinct submissions in total for the six different shared tasks of the challenge. Similar to previous editions, several participating systems achieved competitive performance, indicating the continuous advancement of the state-of-the-art in the field.

Figures

Figures reproduced from arXiv: 2508.20554 by the authors.

Figure 1
Figure 1. The scores of the top systems in exact answer generation, for Phase A+ (dashed lines) and B (solid lines), across the test sets of task 13b and task 12b [60]. miss some synonyms or alternative terms submitted by the participants. Such synonyms will be detected during the enrichment process and will be considered for the final results [PITH_FULL_IMAGE:figures/full_fig_p012_1.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. From Voting to Agent Collaboration: Answer-Type-Aware LLM Pipelines for BioASQ 14b

    cs.CL 2026-07 conditional novelty 3.0 of 10

    A question-type-specific LLM ensemble and multi-agent pipeline achieved competitive results on BioASQ 14b Task B, including first place in the factoid subtask of Batch 4.

Reference graph

Works this paper leans on

89 extracted references · 68 canonical work pages · cited by 1 Pith paper

  1. [1]

    In: Faggioli, G., Ferro, N., Rosso, P., Spina, D

    Andersen, L.R., Gardshodn, M.I., Dolmer, M.H., Rodriguez, J.M., Dell’Aglio, D.: Trusting Gut Instincts: Transformer-Based Extraction of Structured Data from Gut-Brain Axis Publications. In: Faggioli, G., Ferro, N., Rosso, P., Spina, D. (eds.) CLEF 2025 Working Notes (2025)

  2. [2]

    In: Faggioli, G., Ferro, N., Rosso, P., Spina, D

    Angulo, J., Yeste, V.: AQAMS and AQAMS2: Multi Agent Systems for Biomedical Question Answering . In: Faggioli, G., Ferro, N., Rosso, P., Spina, D. (eds.) CLEF 2025 Working Notes (2025)

  3. [3]

    Integrative Medicine: A Clinician’s Journal 17(4), 28 (2018)

    Appleton, J.: The gut-brain axis: influence of microbiota on mood and mental health. Integrative Medicine: A Clinician’s Journal 17(4), 28 (2018)

  4. [4]

    In: Faggioli, G., Ferro, N., Rosso, P., Spina, D

    Ateia, S., Kruschwitz, U.: Can Language Models Critique Themselves? Investi- gating Self-Feedback for Retrieval Augmented Generation at BioASQ 2025 . In: Faggioli, G., Ferro, N., Rosso, P., Spina, D. (eds.) CLEF 2025 Working Notes (2025)

  5. [5]

    In: EMNLP (2019)

    Beltagy, I., Lo, K., Cohan, A.: SciBERT: Pretrained Language Model for Scientific Text. In: EMNLP (2019)

  6. [6]

    In: Faggioli, G., Ferro, N., Rosso, P., Spina, D

    Bing-Chen, C., Han, J.C., Hung, H.C., Tsai, R.T.H.: NCU-IISR: Biomedical Ques- tion Answering via Gemini and GPT APIs in the BioASQ 13b Phase B Challenge . In: Faggioli, G., Ferro, N., Rosso, P., Spina, D. (eds.) CLEF 2025 Working Notes (2025)

  7. [7]

    Nucleic acids research 32(suppl 1), D267–D270 (2004)

    Bodenreider, O.: The unified medical language system (UMLS): integrating biomedical terminology. Nucleic acids research 32(suppl 1), D267–D270 (2004)

  8. [8]

    Bogdanov, S., Constantin, A., Bernard, T., Crabb´ e, B., Bernard, E.: NuNER: Entity Recognition Encoder Pre-training via LLM-Annotated Data (2024)

Show all 89 references
  1. [9]

    In: Faggioli, G., Ferro, N., Rosso, P., Spina, D

    Borazio, F., Croce, D., Basili, R.: UniTor at BioASQ 2025: Modular Biomedical QA with Synthetic Snippets and Multiple Task Answer Generation . In: Faggioli, G., Ferro, N., Rosso, P., Spina, D. (eds.) CLEF 2025 Working Notes (2025)

  2. [10]

    In: Faggioli, G., Ferro, N., Rosso, P., Spina, D

    Burlova, A.: Navigating Partial UMLS Terminology: GAT Embeddings and Confi- dence Analysis for Multilingual Concept Linking. In: Faggioli, G., Ferro, N., Rosso, P., Spina, D. (eds.) CLEF 2025 Working Notes (2025)

  3. [11]

    Annals of gastroenterology: quarterly publication of the Hellenic Society of Gastroenterology 28(2), 203 (2015)

    Carabotti, M., Scirocco, A., Maselli, M.A., Severi, C.: The gut-brain axis: interac- tions between enteric microbiota, central and enteric nervous systems. Annals of gastroenterology: quarterly publication of the Hellenic Society of Gastroenterology 28(2), 203 (2015)

  4. [12]

    In: International Conference on Learning Representations (2020)

    Clark, K., Luong, M.T., Le, Q.V., Manning, C.D.: ELECTRA: Pre-training Text Encoders as Discriminators Rather Than Generators. In: International Conference on Learning Representations (2020)

  5. [13]

    In: Faggioli, G., Ferro, N., Rosso, P., Spina, D

    Concei¸ c˜ ao, S.I.R., Lopes, P.R.C., Couto, F.M.: lasigeBioTM at BioASQ25 Task GutBrainIE - Lean Large language models with syntactic features. In: Faggioli, G., Ferro, N., Rosso, P., Spina, D. (eds.) CLEF 2025 Working Notes (2025)

  6. [14]

    The Lancet Neurology19(2), 179–194 (2020) Overview of BioASQ 2025 21

    Cryan, J.F., O’Riordan, K.J., Sandhu, K., Peterson, V., Dinan, T.G.: The gut microbiome in neurological disorders. The Lancet Neurology19(2), 179–194 (2020) Overview of BioASQ 2025 21

  7. [15]

    D. Lain, A., Lee, C., Doneva, S.E., Rodr´ ıguez-Cubillos, M.J., Castagnari, E., Simp- son, T.I., , Posma, J.M.: Multilingual and Nested Biomedical Named Entity Nor- malisation via Candidate Retrieval and Lightweight Large Language Model Dis- ambiguation. In: Faggioli, G., Ferr...

  8. [16]

    In: Faggioli, G., Ferro, N., Rosso, P., Spina, D

    Datseris, A., Kuzmanov, M., Nikolova-Koleva, I., Taskov, D., Boytcheva, S.: Graph- wise @ CLEF-2025 GutBrainIE: Towards Automated Discovery of Gut-Brain Inter- actions: Deep Learning for NER and Relation Extraction from PubMed Abstracts. In: Faggioli, G., Ferro, N., Rosso, P.,...

  9. [17]

    In: Burstein, J., Doran, C., Solorio, T

    Devlin, J., Chang, M.W., Lee, K., Toutanova, K.: BERT: Pre-training of deep bidirectional transformers for language understanding. In: Burstein, J., Doran, C., Solorio, T. (eds.) Proceedings of the 2019 Conference of the North American Chapter of the Association for Computatio...

  10. [18]

    In: Faggioli, G., Ferro, N., Rosso, P., Spina, D

    Dimitriadis, D., Patsiou, V., Stoikopoulou, E., Toumpas, A., Kipouros, A., Pa- padopoulos, D., Bekiaridou, A., Barmpagiannos, K., Vasilopoulou, A., Barmpa- giannos, A., Samaras, A., Giannakoulas, G., Tsoumakas, G.: Overview of ElCar- dioCC Task on Clinical Coding in Cardiology...

  11. [19]

    In: Faggioli, G., Ferro, N., Rosso, P., Spina, D

    Due˜ nas Romero, S., Ure˜ na-L´ opez, L.A., Mart´ ınez-C´ amara, E.: SINAI at CLEF 2025: A Multi-Stage RAG Pipeline for Biomedical Semantic Question Answering . In: Faggioli, G., Ferro, N., Rosso, P., Spina, D. (eds.) CLEF 2025 Working Notes (2025)

  12. [20]

    In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)

    Feng, F., Yang, Y., Cer, D., Arivazhagan, N., Wang, W.: Language-agnostic BERT sentence embedding. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). pp. 878–891. ACL, Dublin, Ireland (May 2022). https://doi.org...

  13. [21]

    In: Faggioli, G., Ferro, N., Rosso, P., Spina, D

    Galat, D., Molla-Aliod, D.: LLM Ensemble for RAG: Role of context length in zero-shot Question Answering for BioASQ Challenge . In: Faggioli, G., Ferro, N., Rosso, P., Spina, D. (eds.) CLEF 2025 Working Notes (2025)

  14. [22]

    Evaluation of advance hierarchical classification techniques for scientific literature, patents and clinical trials

    Gasco, L., Nentidis, A., Krithara, A., Estrada-Zavala, D., Toshiyuki Murasaki, R., Primo-Pe˜ na, E., Bojo-Canales, C., Paliouras, G., Krallinger, M.: Overview of BioASQ 2021-MESINESP track. Evaluation of advance hierarchical classification techniques for scientific literature,...

  15. [23]

    Pharmacology & therapeutics 158, 52–62 (2016)

    Ghaisas, S., Maher, J., Kanthasamy, A.: Gut microbiome in health and disease: Linking the microbiome–gut–brain axis and environmental factors in the patho- genesis of systemic and neurodegenerative diseases. Pharmacology & therapeutics 158, 52–62 (2016)

  16. [24]

    In: Faggioli, G., Ferro, N., Rosso, P., Spina, D

    Grazhdanski, G.: Group relative policy optimization for spanish clinical case report summarization. In: Faggioli, G., Ferro, N., Rosso, P., Spina, D. (eds.) CLEF 2025 Working Notes (2025)

  17. [25]

    ACM Transactions on Computing for Healthcare (HEALTH) 3(1), 1–23 (2021)

    Gu, Y., Tinn, R., Cheng, H., Lucas, M., Usuyama, N., Liu, X., Naumann, T., Gao, J., Poon, H.: Domain-specific language model pretraining for biomedical natural language processing. ACM Transactions on Computing for Healthcare (HEALTH) 3(1), 1–23 (2021)

  18. [26]

    In: Faggioli, G., Ferro, N., Rosso, P., Spina, D

    Gupta, H.P., Banerjee, R.: LLMs for Biomedical NER. In: Faggioli, G., Ferro, N., Rosso, P., Spina, D. (eds.) CLEF 2025 Working Notes (2025) 22 A. Nentidis et al

  19. [27]

    In: Faggioli, G., Ferro, N., Rosso, P., Spina, D

    Han, J., Liu, Y.: GutUZH at CLEF2025 BioASQ Task 6: a method of SOTA performance with the best results at GutBrainIE NER subtask 1. In: Faggioli, G., Ferro, N., Rosso, P., Spina, D. (eds.) CLEF 2025 Working Notes (2025)

  20. [28]

    In: Faggioli, G., Ferro, N., Rosso, P., Spina, D

    Huang, B.: Clinical entity recognition and linking in greek discharge letters using multilingual-llm-based multi-stage system. In: Faggioli, G., Ferro, N., Rosso, P., Spina, D. (eds.) CLEF 2025 Working Notes (2025)

  21. [29]

    In: Findings of the Association for Computational Linguistics: EMNLP 2021

    Huguet Cabot, P.L., Navigli, R.: REBEL: Relation extraction by end-to-end lan- guage generation. In: Findings of the Association for Computational Linguistics: EMNLP 2021. pp. 2370–2381. ACL, Punta Cana, Dominican Republic (Nov 2021), https://aclanthology.org/2021.findings-emnlp.204

  22. [30]

    In: Fag- gioli, G., Ferro, N., Rosso, P., Spina, D

    Jonker, R.A.A., Almeida, T., Almeida, J., Matos, S.: BIT.UA at BioASQ 13B: Revisiting Evaluation, DPRF-Enhanced Retrieval and Fine-Tuned LLMs. In: Fag- gioli, G., Ferro, N., Rosso, P., Spina, D. (eds.) CLEF 2025 Working Notes (2025)

  23. [31]

    In: Faggioli, G., Ferro, N., Rosso, P., Spina, D

    Kantz, B., Waldert, P., Lengauer, S., Schreck, T.: Constrained Linked Entity AN- notation using RAG (CLEANR). In: Faggioli, G., Ferro, N., Rosso, P., Spina, D. (eds.) CLEF 2025 Working Notes (2025)

  24. [32]

    In: Faggioli, G., Ferro, N., Rosso, P., Spina, D

    Keinan, R., Cohen, A.D.N., Tsarfaty, R.: From Named Entities to Relations: End- to-End Biomedical Information Extraction. In: Faggioli, G., Ferro, N., Rosso, P., Spina, D. (eds.) CLEF 2025 Working Notes (2025)

  25. [33]

    In: Faggioli, G., Ferro, N., Rosso, P., Spina, D

    Kim, H., Lee, H., Cho, Y., Park, J., Park, J., Park, S., Chok, Y.T., Baek, S., Lee, D., Kang, J.: Prompting Matters: Snippet-Aware Strategies for Biomedical QA with LLMs in BioASQ 13b . In: Faggioli, G., Ferro, N., Rosso, P., Spina, D. (eds.) CLEF 2025 Working Notes (2025)

  26. [34]

    Scientific Data 10(1), 170 (2023)

    Krithara, A., Nentidis, A., Bougiatiotis, K., Paliouras, G.: BioASQ-QA: A manu- ally curated corpus for Biomedical Question Answering. Scientific Data 10(1), 170 (2023)

  27. [35]

    In: Proceedings of the Fourth BioASQ workshop (2016), https://www.aclweb.org/anthology/W16-3101.pdf

    Krithara, A., Nentidis, A., Paliouras, G., Kakadiaris, I.: Results of the 4th edition of BioASQ Challenge. In: Proceedings of the Fourth BioASQ workshop (2016), https://www.aclweb.org/anthology/W16-3101.pdf

  28. [36]

    In: Advances in Information Retrieval: ECIR 2021, Virtual Event, March 28–April 1, 2021, Proceedings, Part II 43

    Krithara, A., Nentidis, A., Paliouras, G., Krallinger, M., Miranda, A.: BioASQ at CLEF2021: large-scale biomedical semantic indexing and question answering. In: Advances in Information Retrieval: ECIR 2021, Virtual Event, March 28–April 1, 2021, Proceedings, Part II 43. pp. 62...

  29. [37]

    Journal of the American Medical Informatics Association p

    Krithara, A., Nentidis, A., Vandorou, E., Katsimpras, G., Almirantis, Y., Ar- nal, M., Bunevicius, A., Farre-Maduell, E., Kassiss, M., Konstantakos, V., Matis- Mitchell, S., Polychronopoulos, D., Rodriguez-Pascual, J., Samaras, E.G., Samio- taki, M., Sanoudou, D., Vozi, A., Pa...

  30. [38]

    In: Faggioli, G., Ferro, N., Rosso, P., Spina, D

    Lee, C., Doneva, S., Rodriguez-Cubillos, M., Castagnari, E., Lain, A., Posma, J., Simpson, T.I.: Understanding Gut-Brain Interplay in Scientific Literature: A Hybrid Approach from Classification to Generative LLM Reasoning. In: Faggioli, G., Ferro, N., Rosso, P., Spina, D. (ed...

  31. [39]

    Bioinformatics 36(4), 1234–1240 (09 2019)

    Lee, J., Yoon, W., Kim, S., Kim, D., Kim, S., So, C.H., Kang, J.: BioBERT: a pre-trained biomedical language representation model for biomedical text mining. Bioinformatics 36(4), 1234–1240 (09 2019). https://doi.org/10.1093/bioinformatics/btz682

  32. [40]

    In: Faggioli, G., Ferro, N., Rosso, P., Spina, D

    Li, C., Zheng, X., Liu, S.: BIBERT on Biomedical Nested Named Entity Linking at BioASQ 2025. In: Faggioli, G., Ferro, N., Rosso, P., Spina, D. (eds.) CLEF 2025 Working Notes (2025) Overview of BioASQ 2025 23

  33. [41]

    In: Pro- ceedings of the ACL workshop ‘Text Summarization Branches Out’

    Lin, C.Y.: ROUGE: A package for automatic evaluation of summaries. In: Pro- ceedings of the ACL workshop ‘Text Summarization Branches Out’. pp. 74–81. Barcelona, Spain (2004)

  34. [42]

    In: Proceedings of the 2021 Confer- ence of the North American Chapter of the Association for Computational Lin- guistics: Human Language Technologies

    Liu, F., Shareghi, E., Meng, Z., Basaldella, M., Collier, N.: Self-alignment pre- training for biomedical entity representations. In: Proceedings of the 2021 Confer- ence of the North American Chapter of the Association for Computational Lin- guistics: Human Language Technolog...

  35. [43]

    Liu, F., Vuli´ c, I., Korhonen, A., Collier, N.: Learning domain-specialised rep- resentations for cross-lingual biomedical entity linking. In: Proceedings of the 59th Annual Meeting of the Association for Computational Linguis- tics and the 11th International Joint Conference...

  36. [44]

    In: Faggioli, G., Ferro, N., Rosso, P., Spina, D

    Liu, Y.: LYX DMIIP FDU at BioASQ 2025: Utilizing BERT embeddings for biomedical text mining. In: Faggioli, G., Ferro, N., Rosso, P., Spina, D. (eds.) CLEF 2025 Working Notes (2025)

  37. [45]

    Bioinformatics (04 2023)

    Loukachevitch, N., Manandhar, S., Baral, E., Rozhkov, I., Braslavski, P., Ivanov, V., Batura, T., Tutubalina, E.: NEREL-BIO: A Dataset of Biomedi- cal Abstracts Annotated with Nested Named Entities. Bioinformatics (04 2023). https://doi.org/10.1093/bioinformatics/btad161, btad161

  38. [46]

    In: Pro- ceedings of the 2024 Joint International Conference on Computational Linguis- tics, Language Resources and Evaluation (LREC-COLING 2024)

    Loukachevitch, N., Sakhovskiy, A., Tutubalina, E.: Biomedical concept normal- ization over nested entities with partial UMLS terminology in Russian. In: Pro- ceedings of the 2024 Joint International Conference on Computational Linguis- tics, Language Resources and Evaluation (...

  39. [47]

    Malakasiotis, P., Pavlopoulos, I., Androutsopoulos, I., Nentidis, A.: Evaluation measures for task b. Tech. rep., Tech. rep. BioASQ (2022), http://participants- area.bioasq.org/Tasks/b/eval meas 2022

  40. [48]

    In: Faggioli, G., Ferro, N., Rosso, P., Spina, D

    Martinelli, M., Silvello, G., Bonato, V., Di Nunzio, G.M., Ferro, N., Irrera, O., Marchesin, S., Menotti, L., Vezzani, F.: Overview of GutBrainIE@CLEF 2025: Gut-Brain Interplay Information Extraction. In: Faggioli, G., Ferro, N., Rosso, P., Spina, D. (eds.) CLEF 2025 Working N...

  41. [49]

    Mehta, R.: Enhancing Biomedical Named Entity Recognition using GLiNER- BioMed with Targeted Dictionary-Based Post-processing for BioASQ 2025 task

  42. [50]

    (eds.) CLEF 2025 Working Notes (2025)

    In: Faggioli, G., Ferro, N., Rosso, P., Spina, D. (eds.) CLEF 2025 Working Notes (2025)

  43. [51]

    In: Advances in Information Retrieval

    Nentidis, A., Katsimpras, G., Krithara, A., Krallinger, M., Ortega, M.R., Loukachevitch, N., Sakhovskiy, A., Tutubalina, E., Tsoumakas, G., Giannakoulas, G., Bekiaridou, A., Samaras, A., Di Nunzio, G.M., Ferro, N., Marchesin, S., Menotti, L., Silvello, G., Paliouras, G.: BioAS...

  44. [52]

    In: Arampatzis, A., Kanoulas, E., Tsikrika, T., Vrochidis, S., Giachanou, 24 A

    Nentidis, A., Katsimpras, G., Krithara, A., Lima L´ opez, S., Farr´ e-Maduell, E., Gasco, L., Krallinger, M., Paliouras, G.: Overview of BioASQ 2023: The Eleventh BioASQ Challenge on Large-Scale Biomedical Semantic Indexing and Question An- swering. In: Arampatzis, A., Kanoula...

  45. [53]

    In: Experimental IR Meets Mul- tilinguality, Multimodality, and Interaction

    Nentidis, A., Katsimpras, G., Krithara, A., Lima-L´ opez, S., Farr´ e-Maduell, E., Krallinger, M., Loukachevitch, N., Davydova, V., Tutubalina, E., Paliouras, G.: Overview of BioASQ 2024: The twelfth BioASQ challenge on Large-Scale Biomed- ical Semantic Indexing and Question A...

  46. [54]

    In: CEUR Workshop Proceedings (2023)

    Nentidis, A., Katsimpras, G., Krithara, A., Paliouras, G.: Overview of BioASQ Tasks 11b and Synergy11 in CLEF2023. In: CEUR Workshop Proceedings (2023)

  47. [55]

    In: Faggioli, G., Ferro, N., Galuˇ sˇ c´ akov´ a, P., Garc´ ıa Seco de Herrera, A

    Nentidis, A., Katsimpras, G., Krithara, A., Paliouras, G.: Overview of BioASQ Tasks 12b and Synergy12 in CLEF2024. In: Faggioli, G., Ferro, N., Galuˇ sˇ c´ akov´ a, P., Garc´ ıa Seco de Herrera, A. (eds.) CLEF Working Notes (2024)

  48. [56]

    In: Faggioli, G., Ferro, N., Rosso, P., Spina, D

    Nentidis, A., Katsimpras, G., Krithara, A., Paliouras, G.: Overview of BioASQ Tasks 13b and Synergy13 in CLEF2025. In: Faggioli, G., Ferro, N., Rosso, P., Spina, D. (eds.) CLEF 2025 Working Notes (2025)

  49. [57]

    In: International Conference of the Cross-Language Evaluation Forum for European Languages

    Nentidis, A., Katsimpras, G., Vandorou, E., Krithara, A., Gasco, L., Krallinger, M., Paliouras, G.: Overview of BioASQ 2021: The Ninth BioASQ Challenge on Large- Scale Biomedical Semantic Indexing and Question Answering. In: International Conference of the Cross-Language Evalu...

  50. [58]

    In: Experimental IR Meets Multilinguality, Multimodality, and Inter- action

    Nentidis, A., Katsimpras, G., Vandorou, E., Krithara, A., Miranda-Escalada, A., Gasco, L., Krallinger, M., Paliouras, G.: Overview of BioASQ 2022: The Tenth BioASQ Challenge on Large-Scale Biomedical Semantic Indexing and Question Answering. In: Experimental IR Meets Multiling...

  51. [59]

    In: Proceedings of the 9th BioASQ Workshop A challenge on large-scale biomedical semantic indexing and question answering

    Nentidis, A., Katsimpras, G., Vandorou, E., Krithara, A., Paliouras, G.: Overview of BioASQ Tasks 9a, 9b and Synergy in CLEF2021. In: Proceedings of the 9th BioASQ Workshop A challenge on large-scale biomedical semantic indexing and question answering. CEUR Workshop Proceeding...

  52. [60]

    In: CEUR Workshop Proceedings

    Nentidis, A., Katsimpras, G., Vandorou, E., Krithara, A., Paliouras, G.: Overview of BioASQ Tasks 10a, 10b and Synergy10 in CLEF2022. In: CEUR Workshop Proceedings. vol. 3180, pp. 171–178 (2022)

  53. [61]

    In: ECIR2024

    Nentidis, A., Krithara, A., Paliouras, G., Krallinger, M., Sanchez, L.G., Lima, S., Farre, E., Loukachevitch, N., Davydova, V., Tutubalina, E.: BioASQ at CLEF2024: The Twelfth Edition of the Large-Scale Biomedical Semantic Indexing and Ques- tion Answering Challenge. In: ECIR2...

  54. [62]

    ArXiv abs/1807.03748 (2018), https://api.semanticscholar.org/CorpusID:49670925

    van den Oord, A., Li, Y., Vinyals, O.: Representation learning with contrastive predictive coding. ArXiv abs/1807.03748 (2018), https://api.semanticscholar.org/CorpusID:49670925

  55. [63]

    In: Faggioli, G., Ferro, N., Rosso, P., Spina, D

    Pamio, L., Di Nunzio, G.M.: BioASQ task GutBrainIE 2025 Task 6: Comparing CRF vs BERT Models for Named Entity Recognition and Relation Extraction. In: Faggioli, G., Ferro, N., Rosso, P., Spina, D. (eds.) CLEF 2025 Working Notes (2025)

  56. [64]

    In: Faggioli, G., Ferro, N., Rosso, P., Spina, D

    Panou, D., Dimopoulos, A., Koubarakis, M., Reczko, M.: Harnessing Collective Intelligence of LLMs for Robust Biomedical QA: A Multi-Model Approach. In: Faggioli, G., Ferro, N., Rosso, P., Spina, D. (eds.) CLEF 2025 Working Notes (2025) Overview of BioASQ 2025 25

  57. [65]

    In: Faggioli, G., Ferro, N., Rosso, P., Spina, D

    Pe˜ na Gnecco, D., Serrano, J., Puertas, E., Mart´ ınez-Santos, J.C.: Hybrid Re- ranking for Biomedical Entity Linking using SapBERT Embeddings: A High- Performance System for BioNNE-L 2025-1. In: Faggioli, G., Ferro, N., Rosso, P., Spina, D. (eds.) CLEF 2025 Working Notes (2025)

  58. [66]

    In: Faggioli, G., Ferro, N., Rosso, P., Spina, D

    Piron, S., Di Nunzio, G.M.: Named Entity Recognition with GLiNER and Relation Extraction with LLMs. In: Faggioli, G., Ferro, N., Rosso, P., Spina, D. (eds.) CLEF 2025 Working Notes (2025)

  59. [67]

    In: Faggioli, G., Ferro, N., Rosso, P., Spina, D

    Rodr´ ıguez-Ortega, M., Rodr´ ıguez-Lopez, E., Lima-L´ opez, S., Escolano, C., Melero, M., Pratesi, L., Vigil-Gimenez, L., Fernandez, L., Farr´ e-Maduell, E., Krallinger, M.: Overview of MultiClinSum task at BioASQ 2025: evaluation of clinical case summarization strategies for...

  60. [68]

    In: Faggioli, G., Ferro, N., Rosso, P., Spina, D

    Sakhovskiy, A., Loukachevitch, N., Tutubalina, E.: Overview of the BioASQ BioNNE-L Task on Biomedical Nested Entity Linking in CLEF 2025. In: Faggioli, G., Ferro, N., Rosso, P., Spina, D. (eds.) CLEF 2025 Working Notes (2025)

  61. [69]

    In: Experimental IR Meets Multilin- guality, Multimodality, and Interaction

    Sakhovskiy, A., Semenova, N., Kadurin, A., Tutubalina, E.: Graph-enriched biomedical entity representation transformer. In: Experimental IR Meets Multilin- guality, Multimodality, and Interaction. pp. 109–120. Springer Nature Switzerland, Cham (2023)

  62. [70]

    In: Findings of the Association for Computational Linguistics: NAACL 2024

    Sakhovskiy, A., Semenova, N., Kadurin, A., Tutubalina, E.: Biomedical entity representation with graph-augmented multi-objective transformer. In: Findings of the Association for Computational Linguistics: NAACL 2024. pp. 4626–4643. ACL, Mexico City, Mexico (Jun 2024). https://...

  63. [71]

    Journal of the American Society for Information Science41(4), 288–297 (jun 1990)

    Salton, G., Buckley, C.: Improving retrieval performance by relevance feedback. Journal of the American Society for Information Science41(4), 288–297 (jun 1990). https://doi.org/10.1002/(SICI)1097-4571(199006)41:4¡288::AID-ASI8¿3.0.CO;2-H

  64. [72]

    In: Faggioli, G., Ferro, N., Rosso, P., Spina, D

    Schneider, E.T.R., Schneider, F.H., Paraiso, E.C., Britto Jr, A.S., Cruz, R.M.O.: MedGemma-Sum-Pt: A Lightweight Model for Portuguese Clinical Summarization. In: Faggioli, G., Ferro, N., Rosso, P., Spina, D. (eds.) CLEF 2025 Working Notes (2025)

  65. [73]

    In: Faggioli, G., Ferro, N., Rosso, P., Spina, D

    Stachura, D., Konieczna, J., Nowak, A.: Are Smaller Open-Weight LLMs Closing the Gap to Proprietary Models for Biomedical Question Answering? . In: Faggioli, G., Ferro, N., Rosso, P., Spina, D. (eds.) CLEF 2025 Working Notes (2025)

  66. [74]

    In: Proceedings of the 58th Annual Meeting of the Asso- ciation for Computational Linguistics

    Sung, M., Jeon, H., Lee, J., Kang, J.: Biomedical entity representations with syn- onym marginalization. In: Proceedings of the 58th Annual Meeting of the Asso- ciation for Computational Linguistics. pp. 3641–3650. ACL, Online (Jul 2020). https://doi.org/10.18653/v1/2020.acl-main.335

  67. [75]

    In: Faggioli, G., Ferro, N., Rosso, P., Spina, D

    Tang, J., Yang, H., Xiong, K., Li, H., Quaresma, P., Yu, H., Zhang, W., Song, M., Jiang, Y.: Applying DeepSeek to BioASQ Task 13B: Using Supervised Fine- Tuning and Few-Shot Learning . In: Faggioli, G., Ferro, N., Rosso, P., Spina, D. (eds.) CLEF 2025 Working Notes (2025)

  68. [76]

    In: Faggioli, G., Ferro, N., Rosso, P., Spina, D

    Taylor, S., Dil, C., Shah, A., Jannat, Oldham, C., Upadhyay, A., Varughese, J., Yazbeck, N., McInnes, B.T.: NLP@VCU at BioASQ2025: Information Extraction on the GutBrainIE dataset. In: Faggioli, G., Ferro, N., Rosso, P., Spina, D. (eds.) CLEF 2025 Working Notes (2025)

  69. [77]

    In: Findings of the Association for Computational Linguistics: EMNLP 26 A

    Tedeschi, S., Maiorca, V., Campolungo, N., Cecconi, F., Navigli, R.: WikiNEu- Ral: Combined neural and knowledge-based silver data creation for multilingual NER. In: Findings of the Association for Computational Linguistics: EMNLP 26 A. Nentidis et al

  70. [78]

    In: Faggioli, G., Ferro, N., Rosso, P., Spina, D

    Vachharajani, P.: Multilingual embedding and prompt-driven approaches for named entity recognition, entity linking, and clinical code prediction in greek dis- charge summaries. In: Faggioli, G., Ferro, N., Rosso, P., Spina, D. (eds.) CLEF 2025 Working Notes (2025)

  71. [79]

    BMC Bioinformatics 16, 138 (2015)

    Tsatsaronis, G., Balikas, G., Malakasiotis, P., Partalas, I., Zschunke, M., Alvers, M.R., Weissenborn, D., Krithara, A., Petridis, S., Polychronopoulos, D., Almiran- tis, Y., Pavlopoulos, J., Baskiotis, N., Gallinari, P., Artieres, T., Ngonga, A., Heino, N., Gaussier, E., Barr...

  72. [80]

    In: Faggioli, G., Ferro, N., Rosso, P., Spina, D

    Velichkov, B., Datseris, A., Vassileva, S., Boytcheva, S.: Enigma @ ElCardioCC: Bridging NER and ICD-10 Entity Linking - A Hybrid Method for Greek Clinical Narratives. In: Faggioli, G., Ferro, N., Rosso, P., Spina, D. (eds.) CLEF 2025 Working Notes (2025)

  73. [81]

    In: Fag- gioli, G., Ferro, N., Rosso, P., Spina, D

    Vachharajani, P.: pjmathematician at MultiClinSUM 2025: A Novel Automated Prompt Optimization Framework for Multilingual Clinical Summarization. In: Fag- gioli, G., Ferro, N., Rosso, P., Spina, D. (eds.) CLEF 2025 Working Notes (2025)

  74. [82]

    In: Kakadiaris, I.A., Paliouras, G., Krithara, A

    Yang, Z., Zhou, Y., Nyberg, E.: Learning to answer biomedical questions: OAQA at BioASQ 4B. In: Kakadiaris, I.A., Paliouras, G., Krithara, A. (eds.) Proceedings of the Fourth BioASQ workshop. pp. 23–37. ACL, Berlin, Germany (Aug 2016). https://doi.org/10.18653/v1/W16-3104

  75. [83]

    In: Faggioli, G., Ferro, N., Rosso, P., Spina, D

    Verma, S., Jiang, F., Xue, X.: Beyond Retrieval: Ensembling Cross-Encoders and GPT Rerankers with LLMs for Biomedical QA . In: Faggioli, G., Ferro, N., Rosso, P., Spina, D. (eds.) CLEF 2025 Working Notes (2025)

  76. [84]

    In: Duh, K., Gomez, H., Bethard, S

    Zaratiana, U., Tomeh, N., Holat, P., Charnois, T.: GLiNER: Generalist model for named entity recognition using bidirectional transformer. In: Duh, K., Gomez, H., Bethard, S. (eds.) Proceedings of the 2024 Conference of the North American Chapter of the Association for Computat...

  77. [85]

    In: Association for Computational Linguistics (ACL) (2022)

    Yasunaga, M., Leskovec, J., Liang, P.: LinkBERT: Pretraining Language Models with Document Links. In: Association for Computational Linguistics (ACL) (2022)

  78. [86]

    In: Proceedings of the AAAI Conference on Artificial Intelligence (2021)

    Zhou, W., Huang, K., Ma, T., Huang, J.: Document-level relation extraction with adaptive thresholding and localized context pooling. In: Proceedings of the AAAI Conference on Artificial Intelligence (2021)

  79. [87]

    In: International Conference on Learning Rep- resentations (ICLR) (2020), https://arxiv.org/abs/1904.09675

    Zhang, T., Kishore, V., Wu, F., Weinberger, K.Q., Artzi, Y.: BERTScore: Evalu- ating Text Generation with BERT. In: International Conference on Learning Rep- resentations (ICLR) (2020), https://arxiv.org/abs/1904.09675

  80. [89]

    In: Li, S., Sun, M., Liu, Y., Wu, H., Liu, K., Che, W., He, S., Rao, G

    Zhuang, L., Wayne, L., Ya, S., Jun, Z.: A robustly optimized BERT pre-training approach with post-training. In: Li, S., Sun, M., Liu, Y., Wu, H., Liu, K., Che, W., He, S., Rao, G. (eds.) Proceedings of the 20th Chinese National Conference on Computational Linguistics. pp. 1218...

  81. [2021]

    2521–2533

    pp. 2521–2533. ACL, Punta Cana, Dominican Republic (Nov 2021). https://doi.org/10.18653/v1/2021.findings-emnlp.215

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.