REVIEW 3 major objections 5 minor 1 cited by
Overview of BioASQ 2025: The Thirteenth BioASQ Challenge on Large-Scale Biomedical Semantic Indexing and Question Answering
T0 review · 3 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read The thirteenth BioASQ challenge benchmarks biomedical AI across six tasks and six languages.
desk verdict BioASQ 2025 overview is a genuinely useful benchmark paper; the 13b scores are preliminary and the MultiClinSum table counts don't match the text, so treat it as a solid draft, not a final reference. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying object is the challenge itself: a collection of six shared tasks, each with a labelled dataset, an evaluation measure, and a baseline system. The evaluation stack combines automatic metrics (MAP and F1 for retrieval, MRR for factoid answers, macro-F1 for yes/no, BERTScore and ROUGE-LSum for summarization, Accuracy@k and MRR for entity linking, micro-F1 for coding and information extraction) with manual expert assessment for the open-ended ideal answers of tasks 13b and Synergy 13. These measures convert each task into comparable leaderboards, and the baselines provide the bar systems must beat.
What would settle it
Re-run the final, post-enrichment evaluation of task 13b once the ground truth is completed, or compare the preliminary leaderboard with the published final one; if the added synonyms and answer elements move top systems' MRR or macro-F1 by more than a few points, or if any top run fails to reproduce on the released test set, the conclusion that 2025 systems are competitive with 2024 would need revision.
Extended reading notes
Core claim
In the thirteenth edition, BioASQ ran two established tasks (13b and Synergy 13) and four new ones (MultiClinSum, BioNNE-L, ELCardioCC, GutBrainIE), producing gold-standard datasets in six languages. According to the reported results, the best systems matched or exceeded the previous edition on task 13b, with notably good yes/no answering even when no relevant documents were supplied, while the new tasks showed that multilingual clinical summarization works best in English and that nested entity linking is best handled by domain-specific biomedical BERT-style retrieval rather than general LLMs. The paper also reports that in the gut-brain information extraction task, named-entity recognition reached a top micro-F1 of 0.84, whereas the hardest subtask, mention-based relation extraction, reached only 0.46, so the paper's conclusion that several participating systems achieved competitive performance holds task by task.
Load-bearing premise
The reported scores are accurate and final; the paper itself states that task 13b results are preliminary, pending manual assessment and enrichment of the ground truth with additional synonyms and answer elements, so the rankings and the conclusion about competitive performance could change.
Editorial extensions
If this is right
- The six released datasets become public benchmarks, so future systems can be compared on the same questions, discharge letters, and abstracts.
- The yes/no result suggests that for binary questions, capable LLMs may not need retrieved golden documents, which could change how QA pipelines are built.
- Domain-specific biomedical encoders, not general LLMs, carried the entity-linking task, indicating that pre-training on UMLS concepts still matters.
- The 0.84 versus 0.46 gap in GutBrainIE shows where the next bottleneck is: locating and typing the exact mention pairs in a relation, not finding entities.
- Multilingual summarization results were best in English, so language-specific data remains a limiting factor even within one clinical domain.
Reading between the lines
- The near-parity of Phase A+ and Phase B yes/no scores suggests retrieval quality may matter least for binary questions; a clean ablation would be to feed a top system random documents and measure the drop.
- Because MultiClinSum rows were machine-translated across languages, the non-English leaderboards measure translation-plus-summarization jointly; a future edition could score only native-language texts to separate the two.
- The Greek cardiology task's effective mBERT baseline hints that a simple fine-tuned multilingual encoder is a difficult target for clinical coding in low-resource languages, which other languages could adopt as a cheap baseline.
- The BioNNE-L dictionary sizes, about 1.8 million English UMLS concepts versus 92 thousand Russian ones, make the bilingual track as much a test of vocabulary coverage as of linking, so cross-lingual gains could be analysed by stratifying on whether the English concept exists in Russian.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper presents the thirteenth edition of the BioASQ challenge, held at CLEF 2025, covering six shared tasks: the established biomedical QA task 13b and Synergy 13, and four new tasks (MultiClinSum, BioNNE-L, ELCardioCC, GutBrainIE). For each task, the authors describe the dataset, evaluation measures, participant systems, and reported results, and they conclude that participation increased and that several systems achieved competitive performance across tasks and languages.
Significance. If the reported results hold, the paper provides a useful record of the BioASQ 2025 benchmark resources and of the state of the art in six biomedical NLP tasks spanning six languages, three document types, and specialized domains such as cardiology and gut-brain interaction. The new tasks (MultiClinSum, BioNNE-L, ELCardioCC, GutBrainIE) introduce publicly available datasets and baselines that are likely to be reused by the community. The paper is candid in Section 4.1 that the 13b results are preliminary, which is appropriate for a challenge overview, and the system results come from independent competing teams measured against externally grounded resources such as UMLS, PubMed, and expert annotations. The main weaknesses are internal inconsistencies in the reported dataset statistics and the presentation of preliminary 13b results as a firm conclusion in the abstract and conclusions.
major comments (3)
- [Abstract and Section 4.1] The abstract states that "several participating systems achieved competitive performance" and the conclusions draw on the 13b scores, but Section 4.1 explicitly states that the 13b results are preliminary, pending manual assessment and ground-truth enrichment that may add documents, snippets, answer elements, and synonyms. Because the benchmark-validity claim rests on these results, the abstract and the conclusions should either be updated to reflect the final results or explicitly state that the 13b scores, including Figure 1, are preliminary and that rankings and scores may change after enrichment. As written, the paper overstates the firmness of its headline results.
- [Section 2.3, Table 2] The MultiClinSum dataset statistics are internally inconsistent. The text says the gold standard comprises 1,280 English, 534 Spanish, 200 Portuguese, and 200 French pairs, and that translation yields 1,976 pairs per language, but Table 2 reports 988, 988, 1061, and 1034 pairs for the gold-standard sub-tracks and 28.902 for the large-scale sub-tracks. These numbers do not match either the stated originals or the stated post-translation totals, and the table also duplicates the sub-track label "MultiClinSum-ls-es" in the EN row. The dataset is a core contribution of the paper, so the counts need to be reconciled and corrected.
- [Section 4.1, Figure 1, Tables 7-8] The comparison between 13b and 12b in Figure 1 and the batch-level observations in Table 7 are presented without error bars, confidence intervals, or significance testing. Since the scores are single top-system values on small batches (85 questions each), the claims that "the top systems achieved scores comparable or higher to those of 12b" and that the last batch is more challenging are not statistically supported. The authors should either add interval estimates or hedge the wording to reflect that these are point estimates from a single run of the evaluation.
minor comments (5)
- [Section 5] The MultiClinSum conclusion mentions "English and Italian text," but Italian is not among the task languages (English, Spanish, French, Portuguese); this appears to be an error and should be corrected.
- [Table 2] The large-scale row for English is labeled "MultiClinSum-ls-es EN"; the sub-track label should likely be "MultiClinSum-ls-en," and the duplicate "ls-es" label should be fixed. In addition, the count "28.902" should use a thousands separator (28,902) for readability.
- [Section 2.4] Footnote 12 contains a URL with a space ("BioNNE-L Shared Task"); it should be percent-encoded or replaced with a stable link.
- [Section 5] There is a typo in the task name "GrutBrainIE" (should be "GutBrainIE").
- [Section 3.6] The phrase "registered 17 teams submitting runs" is ambiguous; it likely means 17 teams registered and submitted runs. Please rephrase for clarity.
Circularity Check
No significant circularity: the paper reports externally grounded measurements of systems from 83 independent teams, and no prediction reduces to a fitted input or to a self-citation.
full rationale
The paper contains no derivation chain of the kind that can be circular. Its central content is a shared-task overview: 83 teams submitted more than 1000 runs, and each task's evaluation is grounded in external resources rather than in the paper's own outputs. Task 13b uses expert-created gold documents, snippets, and answers; BioNNE-L uses UMLS dictionaries and NEREL-BIO annotations; ELCardioCC uses expert-annotated Greek discharge letters with ICD-10 codes; GutBrainIE uses expert-annotated PubMed abstracts; and MultiClinSum compares generated summaries against summaries written by the original case-report authors. The many self-citations, such as [50], [54], [55], [66], [67], [18], and [48], are contextual references to prior editions and to task overviews; they are not load-bearing for the reported system performances. The paper explicitly flags in Section 4.1 that the Task 13b results are preliminary, pending manual assessment and ground-truth enrichment, which is an empirical validity caveat about possibly shifting rankings, not a circularity step: the final scores will still come from independent measurements against enriched gold data. No fitted parameter is renamed as a prediction, no uniqueness theorem from the authors' prior work is invoked to force a choice, and no known empirical result is repackaged under new coordinates. The overview is therefore self-consistent and non-circular.
Assumptions & free parameters
assumptions (3)
- domain assumption Expert-annotated gold data are reliable ground truth for evaluation
- standard math Standard evaluation measures are valid proxies for system quality
- domain assumption Publicly available resources are comprehensive and correctly used
Cite this review
Pith. "Pith review of Overview of BioASQ 2025: The Thirteenth BioASQ Challenge on Large-Scale Biomedical Semantic Indexing and Question Answering." pith.science (2026). https://pith.science/paper/QZC736TG
@misc{pith2026250820554,
author = {Pith},
title = {Pith review of: Overview of BioASQ 2025: The Thirteenth BioASQ Challenge on Large-Scale Biomedical Semantic Indexing and Question Answering},
year = {2026},
howpublished = {\url{https://pith.science/paper/QZC736TG}},
note = {Machine review of arXiv:2508.20554}
}
read the original abstract
This is an overview of the thirteenth edition of the BioASQ challenge in the context of the Conference and Labs of the Evaluation Forum (CLEF) 2025. BioASQ is a series of international challenges promoting advances in large-scale biomedical semantic indexing and question answering. This year, BioASQ consisted of new editions of the two established tasks, b and Synergy, and four new tasks: a) Task MultiClinSum on multilingual clinical summarization. b) Task BioNNE-L on nested named entity linking in Russian and English. c) Task ELCardioCC on clinical coding in cardiology. d) Task GutBrainIE on gut-brain interplay information extraction. In this edition of BioASQ, 83 competing teams participated with more than 1000 distinct submissions in total for the six different shared tasks of the challenge. Similar to previous editions, several participating systems achieved competitive performance, indicating the continuous advancement of the state-of-the-art in the field.
Figures
Forward citations
Cited by 1 Pith paper
-
From Voting to Agent Collaboration: Answer-Type-Aware LLM Pipelines for BioASQ 14b
A question-type-specific LLM ensemble and multi-agent pipeline achieved competitive results on BioASQ 14b Task B, including first place in the factoid subtask of Batch 4.
Reference graph
Works this paper leans on
-
[1]
In: Faggioli, G., Ferro, N., Rosso, P., Spina, D
Andersen, L.R., Gardshodn, M.I., Dolmer, M.H., Rodriguez, J.M., Dell’Aglio, D.: Trusting Gut Instincts: Transformer-Based Extraction of Structured Data from Gut-Brain Axis Publications. In: Faggioli, G., Ferro, N., Rosso, P., Spina, D. (eds.) CLEF 2025 Working Notes (2025)
2025
-
[2]
In: Faggioli, G., Ferro, N., Rosso, P., Spina, D
Angulo, J., Yeste, V.: AQAMS and AQAMS2: Multi Agent Systems for Biomedical Question Answering . In: Faggioli, G., Ferro, N., Rosso, P., Spina, D. (eds.) CLEF 2025 Working Notes (2025)
2025
-
[3]
Integrative Medicine: A Clinician’s Journal 17(4), 28 (2018)
Appleton, J.: The gut-brain axis: influence of microbiota on mood and mental health. Integrative Medicine: A Clinician’s Journal 17(4), 28 (2018)
2018
-
[4]
In: Faggioli, G., Ferro, N., Rosso, P., Spina, D
Ateia, S., Kruschwitz, U.: Can Language Models Critique Themselves? Investi- gating Self-Feedback for Retrieval Augmented Generation at BioASQ 2025 . In: Faggioli, G., Ferro, N., Rosso, P., Spina, D. (eds.) CLEF 2025 Working Notes (2025)
2025
-
[5]
In: EMNLP (2019)
Beltagy, I., Lo, K., Cohan, A.: SciBERT: Pretrained Language Model for Scientific Text. In: EMNLP (2019)
2019
-
[6]
In: Faggioli, G., Ferro, N., Rosso, P., Spina, D
Bing-Chen, C., Han, J.C., Hung, H.C., Tsai, R.T.H.: NCU-IISR: Biomedical Ques- tion Answering via Gemini and GPT APIs in the BioASQ 13b Phase B Challenge . In: Faggioli, G., Ferro, N., Rosso, P., Spina, D. (eds.) CLEF 2025 Working Notes (2025)
2025
-
[7]
Nucleic acids research 32(suppl 1), D267–D270 (2004)
Bodenreider, O.: The unified medical language system (UMLS): integrating biomedical terminology. Nucleic acids research 32(suppl 1), D267–D270 (2004)
2004
-
[8]
Bogdanov, S., Constantin, A., Bernard, T., Crabb´ e, B., Bernard, E.: NuNER: Entity Recognition Encoder Pre-training via LLM-Annotated Data (2024)
2024
Show all 89 references
-
[9]
In: Faggioli, G., Ferro, N., Rosso, P., Spina, D
Borazio, F., Croce, D., Basili, R.: UniTor at BioASQ 2025: Modular Biomedical QA with Synthetic Snippets and Multiple Task Answer Generation . In: Faggioli, G., Ferro, N., Rosso, P., Spina, D. (eds.) CLEF 2025 Working Notes (2025)
2025
-
[10]
In: Faggioli, G., Ferro, N., Rosso, P., Spina, D
Burlova, A.: Navigating Partial UMLS Terminology: GAT Embeddings and Confi- dence Analysis for Multilingual Concept Linking. In: Faggioli, G., Ferro, N., Rosso, P., Spina, D. (eds.) CLEF 2025 Working Notes (2025)
2025
-
[11]
Annals of gastroenterology: quarterly publication of the Hellenic Society of Gastroenterology 28(2), 203 (2015)
Carabotti, M., Scirocco, A., Maselli, M.A., Severi, C.: The gut-brain axis: interac- tions between enteric microbiota, central and enteric nervous systems. Annals of gastroenterology: quarterly publication of the Hellenic Society of Gastroenterology 28(2), 203 (2015)
2015
-
[12]
In: International Conference on Learning Representations (2020)
Clark, K., Luong, M.T., Le, Q.V., Manning, C.D.: ELECTRA: Pre-training Text Encoders as Discriminators Rather Than Generators. In: International Conference on Learning Representations (2020)
2020
-
[13]
In: Faggioli, G., Ferro, N., Rosso, P., Spina, D
Concei¸ c˜ ao, S.I.R., Lopes, P.R.C., Couto, F.M.: lasigeBioTM at BioASQ25 Task GutBrainIE - Lean Large language models with syntactic features. In: Faggioli, G., Ferro, N., Rosso, P., Spina, D. (eds.) CLEF 2025 Working Notes (2025)
2025
-
[14]
The Lancet Neurology19(2), 179–194 (2020) Overview of BioASQ 2025 21
Cryan, J.F., O’Riordan, K.J., Sandhu, K., Peterson, V., Dinan, T.G.: The gut microbiome in neurological disorders. The Lancet Neurology19(2), 179–194 (2020) Overview of BioASQ 2025 21
2020
-
[15]
D. Lain, A., Lee, C., Doneva, S.E., Rodr´ ıguez-Cubillos, M.J., Castagnari, E., Simp- son, T.I., , Posma, J.M.: Multilingual and Nested Biomedical Named Entity Nor- malisation via Candidate Retrieval and Lightweight Large Language Model Dis- ambiguation. In: Faggioli, G., Ferr...
2025
-
[16]
In: Faggioli, G., Ferro, N., Rosso, P., Spina, D
Datseris, A., Kuzmanov, M., Nikolova-Koleva, I., Taskov, D., Boytcheva, S.: Graph- wise @ CLEF-2025 GutBrainIE: Towards Automated Discovery of Gut-Brain Inter- actions: Deep Learning for NER and Relation Extraction from PubMed Abstracts. In: Faggioli, G., Ferro, N., Rosso, P.,...
2025
-
[17]
In: Burstein, J., Doran, C., Solorio, T
Devlin, J., Chang, M.W., Lee, K., Toutanova, K.: BERT: Pre-training of deep bidirectional transformers for language understanding. In: Burstein, J., Doran, C., Solorio, T. (eds.) Proceedings of the 2019 Conference of the North American Chapter of the Association for Computatio...
2019 doi
-
[18]
In: Faggioli, G., Ferro, N., Rosso, P., Spina, D
Dimitriadis, D., Patsiou, V., Stoikopoulou, E., Toumpas, A., Kipouros, A., Pa- padopoulos, D., Bekiaridou, A., Barmpagiannos, K., Vasilopoulou, A., Barmpa- giannos, A., Samaras, A., Giannakoulas, G., Tsoumakas, G.: Overview of ElCar- dioCC Task on Clinical Coding in Cardiology...
2025
-
[19]
In: Faggioli, G., Ferro, N., Rosso, P., Spina, D
Due˜ nas Romero, S., Ure˜ na-L´ opez, L.A., Mart´ ınez-C´ amara, E.: SINAI at CLEF 2025: A Multi-Stage RAG Pipeline for Biomedical Semantic Question Answering . In: Faggioli, G., Ferro, N., Rosso, P., Spina, D. (eds.) CLEF 2025 Working Notes (2025)
2025
-
[20]
In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Feng, F., Yang, Y., Cer, D., Arivazhagan, N., Wang, W.: Language-agnostic BERT sentence embedding. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). pp. 878–891. ACL, Dublin, Ireland (May 2022). https://doi.org...
2022 doi
-
[21]
In: Faggioli, G., Ferro, N., Rosso, P., Spina, D
Galat, D., Molla-Aliod, D.: LLM Ensemble for RAG: Role of context length in zero-shot Question Answering for BioASQ Challenge . In: Faggioli, G., Ferro, N., Rosso, P., Spina, D. (eds.) CLEF 2025 Working Notes (2025)
2025
-
[22]
Evaluation of advance hierarchical classification techniques for scientific literature, patents and clinical trials
Gasco, L., Nentidis, A., Krithara, A., Estrada-Zavala, D., Toshiyuki Murasaki, R., Primo-Pe˜ na, E., Bojo-Canales, C., Paliouras, G., Krallinger, M.: Overview of BioASQ 2021-MESINESP track. Evaluation of advance hierarchical classification techniques for scientific literature,...
2021
-
[23]
Pharmacology & therapeutics 158, 52–62 (2016)
Ghaisas, S., Maher, J., Kanthasamy, A.: Gut microbiome in health and disease: Linking the microbiome–gut–brain axis and environmental factors in the patho- genesis of systemic and neurodegenerative diseases. Pharmacology & therapeutics 158, 52–62 (2016)
2016
-
[24]
In: Faggioli, G., Ferro, N., Rosso, P., Spina, D
Grazhdanski, G.: Group relative policy optimization for spanish clinical case report summarization. In: Faggioli, G., Ferro, N., Rosso, P., Spina, D. (eds.) CLEF 2025 Working Notes (2025)
2025
-
[25]
ACM Transactions on Computing for Healthcare (HEALTH) 3(1), 1–23 (2021)
Gu, Y., Tinn, R., Cheng, H., Lucas, M., Usuyama, N., Liu, X., Naumann, T., Gao, J., Poon, H.: Domain-specific language model pretraining for biomedical natural language processing. ACM Transactions on Computing for Healthcare (HEALTH) 3(1), 1–23 (2021)
2021
-
[26]
In: Faggioli, G., Ferro, N., Rosso, P., Spina, D
Gupta, H.P., Banerjee, R.: LLMs for Biomedical NER. In: Faggioli, G., Ferro, N., Rosso, P., Spina, D. (eds.) CLEF 2025 Working Notes (2025) 22 A. Nentidis et al
2025
-
[27]
In: Faggioli, G., Ferro, N., Rosso, P., Spina, D
Han, J., Liu, Y.: GutUZH at CLEF2025 BioASQ Task 6: a method of SOTA performance with the best results at GutBrainIE NER subtask 1. In: Faggioli, G., Ferro, N., Rosso, P., Spina, D. (eds.) CLEF 2025 Working Notes (2025)
2025
-
[28]
In: Faggioli, G., Ferro, N., Rosso, P., Spina, D
Huang, B.: Clinical entity recognition and linking in greek discharge letters using multilingual-llm-based multi-stage system. In: Faggioli, G., Ferro, N., Rosso, P., Spina, D. (eds.) CLEF 2025 Working Notes (2025)
2025
-
[29]
In: Findings of the Association for Computational Linguistics: EMNLP 2021
Huguet Cabot, P.L., Navigli, R.: REBEL: Relation extraction by end-to-end lan- guage generation. In: Findings of the Association for Computational Linguistics: EMNLP 2021. pp. 2370–2381. ACL, Punta Cana, Dominican Republic (Nov 2021), https://aclanthology.org/2021.findings-emnlp.204
2021
-
[30]
In: Fag- gioli, G., Ferro, N., Rosso, P., Spina, D
Jonker, R.A.A., Almeida, T., Almeida, J., Matos, S.: BIT.UA at BioASQ 13B: Revisiting Evaluation, DPRF-Enhanced Retrieval and Fine-Tuned LLMs. In: Fag- gioli, G., Ferro, N., Rosso, P., Spina, D. (eds.) CLEF 2025 Working Notes (2025)
2025
-
[31]
In: Faggioli, G., Ferro, N., Rosso, P., Spina, D
Kantz, B., Waldert, P., Lengauer, S., Schreck, T.: Constrained Linked Entity AN- notation using RAG (CLEANR). In: Faggioli, G., Ferro, N., Rosso, P., Spina, D. (eds.) CLEF 2025 Working Notes (2025)
2025
-
[32]
In: Faggioli, G., Ferro, N., Rosso, P., Spina, D
Keinan, R., Cohen, A.D.N., Tsarfaty, R.: From Named Entities to Relations: End- to-End Biomedical Information Extraction. In: Faggioli, G., Ferro, N., Rosso, P., Spina, D. (eds.) CLEF 2025 Working Notes (2025)
2025
-
[33]
In: Faggioli, G., Ferro, N., Rosso, P., Spina, D
Kim, H., Lee, H., Cho, Y., Park, J., Park, J., Park, S., Chok, Y.T., Baek, S., Lee, D., Kang, J.: Prompting Matters: Snippet-Aware Strategies for Biomedical QA with LLMs in BioASQ 13b . In: Faggioli, G., Ferro, N., Rosso, P., Spina, D. (eds.) CLEF 2025 Working Notes (2025)
2025
-
[34]
Scientific Data 10(1), 170 (2023)
Krithara, A., Nentidis, A., Bougiatiotis, K., Paliouras, G.: BioASQ-QA: A manu- ally curated corpus for Biomedical Question Answering. Scientific Data 10(1), 170 (2023)
2023
-
[35]
In: Proceedings of the Fourth BioASQ workshop (2016), https://www.aclweb.org/anthology/W16-3101.pdf
Krithara, A., Nentidis, A., Paliouras, G., Kakadiaris, I.: Results of the 4th edition of BioASQ Challenge. In: Proceedings of the Fourth BioASQ workshop (2016), https://www.aclweb.org/anthology/W16-3101.pdf
2016
-
[36]
In: Advances in Information Retrieval: ECIR 2021, Virtual Event, March 28–April 1, 2021, Proceedings, Part II 43
Krithara, A., Nentidis, A., Paliouras, G., Krallinger, M., Miranda, A.: BioASQ at CLEF2021: large-scale biomedical semantic indexing and question answering. In: Advances in Information Retrieval: ECIR 2021, Virtual Event, March 28–April 1, 2021, Proceedings, Part II 43. pp. 62...
2021
-
[37]
Journal of the American Medical Informatics Association p
Krithara, A., Nentidis, A., Vandorou, E., Katsimpras, G., Almirantis, Y., Ar- nal, M., Bunevicius, A., Farre-Maduell, E., Kassiss, M., Konstantakos, V., Matis- Mitchell, S., Polychronopoulos, D., Rodriguez-Pascual, J., Samaras, E.G., Samio- taki, M., Sanoudou, D., Vozi, A., Pa...
2024
-
[38]
In: Faggioli, G., Ferro, N., Rosso, P., Spina, D
Lee, C., Doneva, S., Rodriguez-Cubillos, M., Castagnari, E., Lain, A., Posma, J., Simpson, T.I.: Understanding Gut-Brain Interplay in Scientific Literature: A Hybrid Approach from Classification to Generative LLM Reasoning. In: Faggioli, G., Ferro, N., Rosso, P., Spina, D. (ed...
2025
-
[39]
Bioinformatics 36(4), 1234–1240 (09 2019)
Lee, J., Yoon, W., Kim, S., Kim, D., Kim, S., So, C.H., Kang, J.: BioBERT: a pre-trained biomedical language representation model for biomedical text mining. Bioinformatics 36(4), 1234–1240 (09 2019). https://doi.org/10.1093/bioinformatics/btz682
2019 doi
-
[40]
In: Faggioli, G., Ferro, N., Rosso, P., Spina, D
Li, C., Zheng, X., Liu, S.: BIBERT on Biomedical Nested Named Entity Linking at BioASQ 2025. In: Faggioli, G., Ferro, N., Rosso, P., Spina, D. (eds.) CLEF 2025 Working Notes (2025) Overview of BioASQ 2025 23
2025
-
[41]
In: Pro- ceedings of the ACL workshop ‘Text Summarization Branches Out’
Lin, C.Y.: ROUGE: A package for automatic evaluation of summaries. In: Pro- ceedings of the ACL workshop ‘Text Summarization Branches Out’. pp. 74–81. Barcelona, Spain (2004)
2004
-
[42]
In: Proceedings of the 2021 Confer- ence of the North American Chapter of the Association for Computational Lin- guistics: Human Language Technologies
Liu, F., Shareghi, E., Meng, Z., Basaldella, M., Collier, N.: Self-alignment pre- training for biomedical entity representations. In: Proceedings of the 2021 Confer- ence of the North American Chapter of the Association for Computational Lin- guistics: Human Language Technolog...
2021 doi
-
[43]
Liu, F., Vuli´ c, I., Korhonen, A., Collier, N.: Learning domain-specialised rep- resentations for cross-lingual biomedical entity linking. In: Proceedings of the 59th Annual Meeting of the Association for Computational Linguis- tics and the 11th International Joint Conference...
2021 doi
-
[44]
In: Faggioli, G., Ferro, N., Rosso, P., Spina, D
Liu, Y.: LYX DMIIP FDU at BioASQ 2025: Utilizing BERT embeddings for biomedical text mining. In: Faggioli, G., Ferro, N., Rosso, P., Spina, D. (eds.) CLEF 2025 Working Notes (2025)
2025
-
[45]
Bioinformatics (04 2023)
Loukachevitch, N., Manandhar, S., Baral, E., Rozhkov, I., Braslavski, P., Ivanov, V., Batura, T., Tutubalina, E.: NEREL-BIO: A Dataset of Biomedi- cal Abstracts Annotated with Nested Named Entities. Bioinformatics (04 2023). https://doi.org/10.1093/bioinformatics/btad161, btad161
2023 doi
-
[46]
In: Pro- ceedings of the 2024 Joint International Conference on Computational Linguis- tics, Language Resources and Evaluation (LREC-COLING 2024)
Loukachevitch, N., Sakhovskiy, A., Tutubalina, E.: Biomedical concept normal- ization over nested entities with partial UMLS terminology in Russian. In: Pro- ceedings of the 2024 Joint International Conference on Computational Linguis- tics, Language Resources and Evaluation (...
2024
-
[47]
Malakasiotis, P., Pavlopoulos, I., Androutsopoulos, I., Nentidis, A.: Evaluation measures for task b. Tech. rep., Tech. rep. BioASQ (2022), http://participants- area.bioasq.org/Tasks/b/eval meas 2022
2022
-
[48]
In: Faggioli, G., Ferro, N., Rosso, P., Spina, D
Martinelli, M., Silvello, G., Bonato, V., Di Nunzio, G.M., Ferro, N., Irrera, O., Marchesin, S., Menotti, L., Vezzani, F.: Overview of GutBrainIE@CLEF 2025: Gut-Brain Interplay Information Extraction. In: Faggioli, G., Ferro, N., Rosso, P., Spina, D. (eds.) CLEF 2025 Working N...
2025
-
[49]
Mehta, R.: Enhancing Biomedical Named Entity Recognition using GLiNER- BioMed with Targeted Dictionary-Based Post-processing for BioASQ 2025 task
2025
-
[50]
(eds.) CLEF 2025 Working Notes (2025)
In: Faggioli, G., Ferro, N., Rosso, P., Spina, D. (eds.) CLEF 2025 Working Notes (2025)
2025
-
[51]
In: Advances in Information Retrieval
Nentidis, A., Katsimpras, G., Krithara, A., Krallinger, M., Ortega, M.R., Loukachevitch, N., Sakhovskiy, A., Tutubalina, E., Tsoumakas, G., Giannakoulas, G., Bekiaridou, A., Samaras, A., Di Nunzio, G.M., Ferro, N., Marchesin, S., Menotti, L., Silvello, G., Paliouras, G.: BioAS...
2025
-
[52]
In: Arampatzis, A., Kanoulas, E., Tsikrika, T., Vrochidis, S., Giachanou, 24 A
Nentidis, A., Katsimpras, G., Krithara, A., Lima L´ opez, S., Farr´ e-Maduell, E., Gasco, L., Krallinger, M., Paliouras, G.: Overview of BioASQ 2023: The Eleventh BioASQ Challenge on Large-Scale Biomedical Semantic Indexing and Question An- swering. In: Arampatzis, A., Kanoula...
2023
-
[53]
In: Experimental IR Meets Mul- tilinguality, Multimodality, and Interaction
Nentidis, A., Katsimpras, G., Krithara, A., Lima-L´ opez, S., Farr´ e-Maduell, E., Krallinger, M., Loukachevitch, N., Davydova, V., Tutubalina, E., Paliouras, G.: Overview of BioASQ 2024: The twelfth BioASQ challenge on Large-Scale Biomed- ical Semantic Indexing and Question A...
2024
-
[54]
In: CEUR Workshop Proceedings (2023)
Nentidis, A., Katsimpras, G., Krithara, A., Paliouras, G.: Overview of BioASQ Tasks 11b and Synergy11 in CLEF2023. In: CEUR Workshop Proceedings (2023)
2023
-
[55]
In: Faggioli, G., Ferro, N., Galuˇ sˇ c´ akov´ a, P., Garc´ ıa Seco de Herrera, A
Nentidis, A., Katsimpras, G., Krithara, A., Paliouras, G.: Overview of BioASQ Tasks 12b and Synergy12 in CLEF2024. In: Faggioli, G., Ferro, N., Galuˇ sˇ c´ akov´ a, P., Garc´ ıa Seco de Herrera, A. (eds.) CLEF Working Notes (2024)
2024
-
[56]
In: Faggioli, G., Ferro, N., Rosso, P., Spina, D
Nentidis, A., Katsimpras, G., Krithara, A., Paliouras, G.: Overview of BioASQ Tasks 13b and Synergy13 in CLEF2025. In: Faggioli, G., Ferro, N., Rosso, P., Spina, D. (eds.) CLEF 2025 Working Notes (2025)
2025
-
[57]
In: International Conference of the Cross-Language Evaluation Forum for European Languages
Nentidis, A., Katsimpras, G., Vandorou, E., Krithara, A., Gasco, L., Krallinger, M., Paliouras, G.: Overview of BioASQ 2021: The Ninth BioASQ Challenge on Large- Scale Biomedical Semantic Indexing and Question Answering. In: International Conference of the Cross-Language Evalu...
2021
-
[58]
In: Experimental IR Meets Multilinguality, Multimodality, and Inter- action
Nentidis, A., Katsimpras, G., Vandorou, E., Krithara, A., Miranda-Escalada, A., Gasco, L., Krallinger, M., Paliouras, G.: Overview of BioASQ 2022: The Tenth BioASQ Challenge on Large-Scale Biomedical Semantic Indexing and Question Answering. In: Experimental IR Meets Multiling...
2022 doi
-
[59]
In: Proceedings of the 9th BioASQ Workshop A challenge on large-scale biomedical semantic indexing and question answering
Nentidis, A., Katsimpras, G., Vandorou, E., Krithara, A., Paliouras, G.: Overview of BioASQ Tasks 9a, 9b and Synergy in CLEF2021. In: Proceedings of the 9th BioASQ Workshop A challenge on large-scale biomedical semantic indexing and question answering. CEUR Workshop Proceeding...
2021
-
[60]
In: CEUR Workshop Proceedings
Nentidis, A., Katsimpras, G., Vandorou, E., Krithara, A., Paliouras, G.: Overview of BioASQ Tasks 10a, 10b and Synergy10 in CLEF2022. In: CEUR Workshop Proceedings. vol. 3180, pp. 171–178 (2022)
2022
-
[61]
In: ECIR2024
Nentidis, A., Krithara, A., Paliouras, G., Krallinger, M., Sanchez, L.G., Lima, S., Farre, E., Loukachevitch, N., Davydova, V., Tutubalina, E.: BioASQ at CLEF2024: The Twelfth Edition of the Large-Scale Biomedical Semantic Indexing and Ques- tion Answering Challenge. In: ECIR2...
2024
-
[62]
ArXiv abs/1807.03748 (2018), https://api.semanticscholar.org/CorpusID:49670925
van den Oord, A., Li, Y., Vinyals, O.: Representation learning with contrastive predictive coding. ArXiv abs/1807.03748 (2018), https://api.semanticscholar.org/CorpusID:49670925
2018 arXiv
-
[63]
In: Faggioli, G., Ferro, N., Rosso, P., Spina, D
Pamio, L., Di Nunzio, G.M.: BioASQ task GutBrainIE 2025 Task 6: Comparing CRF vs BERT Models for Named Entity Recognition and Relation Extraction. In: Faggioli, G., Ferro, N., Rosso, P., Spina, D. (eds.) CLEF 2025 Working Notes (2025)
2025
-
[64]
In: Faggioli, G., Ferro, N., Rosso, P., Spina, D
Panou, D., Dimopoulos, A., Koubarakis, M., Reczko, M.: Harnessing Collective Intelligence of LLMs for Robust Biomedical QA: A Multi-Model Approach. In: Faggioli, G., Ferro, N., Rosso, P., Spina, D. (eds.) CLEF 2025 Working Notes (2025) Overview of BioASQ 2025 25
2025
-
[65]
In: Faggioli, G., Ferro, N., Rosso, P., Spina, D
Pe˜ na Gnecco, D., Serrano, J., Puertas, E., Mart´ ınez-Santos, J.C.: Hybrid Re- ranking for Biomedical Entity Linking using SapBERT Embeddings: A High- Performance System for BioNNE-L 2025-1. In: Faggioli, G., Ferro, N., Rosso, P., Spina, D. (eds.) CLEF 2025 Working Notes (2025)
2025
-
[66]
In: Faggioli, G., Ferro, N., Rosso, P., Spina, D
Piron, S., Di Nunzio, G.M.: Named Entity Recognition with GLiNER and Relation Extraction with LLMs. In: Faggioli, G., Ferro, N., Rosso, P., Spina, D. (eds.) CLEF 2025 Working Notes (2025)
2025
-
[67]
In: Faggioli, G., Ferro, N., Rosso, P., Spina, D
Rodr´ ıguez-Ortega, M., Rodr´ ıguez-Lopez, E., Lima-L´ opez, S., Escolano, C., Melero, M., Pratesi, L., Vigil-Gimenez, L., Fernandez, L., Farr´ e-Maduell, E., Krallinger, M.: Overview of MultiClinSum task at BioASQ 2025: evaluation of clinical case summarization strategies for...
2025
-
[68]
In: Faggioli, G., Ferro, N., Rosso, P., Spina, D
Sakhovskiy, A., Loukachevitch, N., Tutubalina, E.: Overview of the BioASQ BioNNE-L Task on Biomedical Nested Entity Linking in CLEF 2025. In: Faggioli, G., Ferro, N., Rosso, P., Spina, D. (eds.) CLEF 2025 Working Notes (2025)
2025
-
[69]
In: Experimental IR Meets Multilin- guality, Multimodality, and Interaction
Sakhovskiy, A., Semenova, N., Kadurin, A., Tutubalina, E.: Graph-enriched biomedical entity representation transformer. In: Experimental IR Meets Multilin- guality, Multimodality, and Interaction. pp. 109–120. Springer Nature Switzerland, Cham (2023)
2023
-
[70]
In: Findings of the Association for Computational Linguistics: NAACL 2024
Sakhovskiy, A., Semenova, N., Kadurin, A., Tutubalina, E.: Biomedical entity representation with graph-augmented multi-objective transformer. In: Findings of the Association for Computational Linguistics: NAACL 2024. pp. 4626–4643. ACL, Mexico City, Mexico (Jun 2024). https://...
2024 doi
-
[71]
Journal of the American Society for Information Science41(4), 288–297 (jun 1990)
Salton, G., Buckley, C.: Improving retrieval performance by relevance feedback. Journal of the American Society for Information Science41(4), 288–297 (jun 1990). https://doi.org/10.1002/(SICI)1097-4571(199006)41:4¡288::AID-ASI8¿3.0.CO;2-H
1990 doi
-
[72]
In: Faggioli, G., Ferro, N., Rosso, P., Spina, D
Schneider, E.T.R., Schneider, F.H., Paraiso, E.C., Britto Jr, A.S., Cruz, R.M.O.: MedGemma-Sum-Pt: A Lightweight Model for Portuguese Clinical Summarization. In: Faggioli, G., Ferro, N., Rosso, P., Spina, D. (eds.) CLEF 2025 Working Notes (2025)
2025
-
[73]
In: Faggioli, G., Ferro, N., Rosso, P., Spina, D
Stachura, D., Konieczna, J., Nowak, A.: Are Smaller Open-Weight LLMs Closing the Gap to Proprietary Models for Biomedical Question Answering? . In: Faggioli, G., Ferro, N., Rosso, P., Spina, D. (eds.) CLEF 2025 Working Notes (2025)
2025
-
[74]
In: Proceedings of the 58th Annual Meeting of the Asso- ciation for Computational Linguistics
Sung, M., Jeon, H., Lee, J., Kang, J.: Biomedical entity representations with syn- onym marginalization. In: Proceedings of the 58th Annual Meeting of the Asso- ciation for Computational Linguistics. pp. 3641–3650. ACL, Online (Jul 2020). https://doi.org/10.18653/v1/2020.acl-main.335
2020 doi
-
[75]
In: Faggioli, G., Ferro, N., Rosso, P., Spina, D
Tang, J., Yang, H., Xiong, K., Li, H., Quaresma, P., Yu, H., Zhang, W., Song, M., Jiang, Y.: Applying DeepSeek to BioASQ Task 13B: Using Supervised Fine- Tuning and Few-Shot Learning . In: Faggioli, G., Ferro, N., Rosso, P., Spina, D. (eds.) CLEF 2025 Working Notes (2025)
2025
-
[76]
In: Faggioli, G., Ferro, N., Rosso, P., Spina, D
Taylor, S., Dil, C., Shah, A., Jannat, Oldham, C., Upadhyay, A., Varughese, J., Yazbeck, N., McInnes, B.T.: NLP@VCU at BioASQ2025: Information Extraction on the GutBrainIE dataset. In: Faggioli, G., Ferro, N., Rosso, P., Spina, D. (eds.) CLEF 2025 Working Notes (2025)
2025
-
[77]
In: Findings of the Association for Computational Linguistics: EMNLP 26 A
Tedeschi, S., Maiorca, V., Campolungo, N., Cecconi, F., Navigli, R.: WikiNEu- Ral: Combined neural and knowledge-based silver data creation for multilingual NER. In: Findings of the Association for Computational Linguistics: EMNLP 26 A. Nentidis et al
-
[78]
In: Faggioli, G., Ferro, N., Rosso, P., Spina, D
Vachharajani, P.: Multilingual embedding and prompt-driven approaches for named entity recognition, entity linking, and clinical code prediction in greek dis- charge summaries. In: Faggioli, G., Ferro, N., Rosso, P., Spina, D. (eds.) CLEF 2025 Working Notes (2025)
2025
-
[79]
BMC Bioinformatics 16, 138 (2015)
Tsatsaronis, G., Balikas, G., Malakasiotis, P., Partalas, I., Zschunke, M., Alvers, M.R., Weissenborn, D., Krithara, A., Petridis, S., Polychronopoulos, D., Almiran- tis, Y., Pavlopoulos, J., Baskiotis, N., Gallinari, P., Artieres, T., Ngonga, A., Heino, N., Gaussier, E., Barr...
2015
-
[80]
In: Faggioli, G., Ferro, N., Rosso, P., Spina, D
Velichkov, B., Datseris, A., Vassileva, S., Boytcheva, S.: Enigma @ ElCardioCC: Bridging NER and ICD-10 Entity Linking - A Hybrid Method for Greek Clinical Narratives. In: Faggioli, G., Ferro, N., Rosso, P., Spina, D. (eds.) CLEF 2025 Working Notes (2025)
2025
-
[81]
In: Fag- gioli, G., Ferro, N., Rosso, P., Spina, D
Vachharajani, P.: pjmathematician at MultiClinSUM 2025: A Novel Automated Prompt Optimization Framework for Multilingual Clinical Summarization. In: Fag- gioli, G., Ferro, N., Rosso, P., Spina, D. (eds.) CLEF 2025 Working Notes (2025)
2025
-
[82]
In: Kakadiaris, I.A., Paliouras, G., Krithara, A
Yang, Z., Zhou, Y., Nyberg, E.: Learning to answer biomedical questions: OAQA at BioASQ 4B. In: Kakadiaris, I.A., Paliouras, G., Krithara, A. (eds.) Proceedings of the Fourth BioASQ workshop. pp. 23–37. ACL, Berlin, Germany (Aug 2016). https://doi.org/10.18653/v1/W16-3104
2016 doi
-
[83]
In: Faggioli, G., Ferro, N., Rosso, P., Spina, D
Verma, S., Jiang, F., Xue, X.: Beyond Retrieval: Ensembling Cross-Encoders and GPT Rerankers with LLMs for Biomedical QA . In: Faggioli, G., Ferro, N., Rosso, P., Spina, D. (eds.) CLEF 2025 Working Notes (2025)
2025
-
[84]
In: Duh, K., Gomez, H., Bethard, S
Zaratiana, U., Tomeh, N., Holat, P., Charnois, T.: GLiNER: Generalist model for named entity recognition using bidirectional transformer. In: Duh, K., Gomez, H., Bethard, S. (eds.) Proceedings of the 2024 Conference of the North American Chapter of the Association for Computat...
2024 doi
-
[85]
In: Association for Computational Linguistics (ACL) (2022)
Yasunaga, M., Leskovec, J., Liang, P.: LinkBERT: Pretraining Language Models with Document Links. In: Association for Computational Linguistics (ACL) (2022)
2022
-
[86]
In: Proceedings of the AAAI Conference on Artificial Intelligence (2021)
Zhou, W., Huang, K., Ma, T., Huang, J.: Document-level relation extraction with adaptive thresholding and localized context pooling. In: Proceedings of the AAAI Conference on Artificial Intelligence (2021)
2021
-
[87]
In: International Conference on Learning Rep- resentations (ICLR) (2020), https://arxiv.org/abs/1904.09675
Zhang, T., Kishore, V., Wu, F., Weinberger, K.Q., Artzi, Y.: BERTScore: Evalu- ating Text Generation with BERT. In: International Conference on Learning Rep- resentations (ICLR) (2020), https://arxiv.org/abs/1904.09675
2020 arXiv
-
[89]
In: Li, S., Sun, M., Liu, Y., Wu, H., Liu, K., Che, W., He, S., Rao, G
Zhuang, L., Wayne, L., Ya, S., Jun, Z.: A robustly optimized BERT pre-training approach with post-training. In: Li, S., Sun, M., Liu, Y., Wu, H., Liu, K., Che, W., He, S., Rao, G. (eds.) Proceedings of the 20th Chinese National Conference on Computational Linguistics. pp. 1218...
2021
-
[2021]
2521–2533
pp. 2521–2533. ACL, Punta Cana, Dominican Republic (Nov 2021). https://doi.org/10.18653/v1/2021.findings-emnlp.215
2021 doi
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.