Pith. sign in

REVIEW 2 major objections 4 minor 56 references

Overview of BioASQ 2024: The twelfth BioASQ challenge on Large-Scale Biomedical Semantic Indexing and Question Answering

T0 review · 2 major / 4 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read The 2024 BioASQ challenge ran four shared tasks with 37 teams and over 700 submissions, and the top systems—built almost entirely on LLMs and transformers—continued the series' upward trend in biomedical question answering and clinical…

desk verdict A solid, honest overview of the 2024 BioASQ edition with two real additions (phase A+ and two new tasks), undercut by an abstract whose participation totals don't add up against the tables in the body. read the letter →

arxiv 2508.20532 v1 pith:V6VV3VEB submitted 2025-08-28 cs.CL cs.AIcs.IR

classification cs.CLcs.AIcs.IR
keywords biomedicalquestionansweringsemanticindexingsharedtaskevaluationnamedentityrecognitionlargelanguagemodelsRetrievalAugmentedGenerationmultilingualclinicalNLPBioASQ2024
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper is the official overview of the twelfth BioASQ challenge, a shared-task evaluation series for biomedical semantic indexing and question answering. It reports that four tasks ran in 2024: the established task b and Synergy, plus two new tasks, MultiCardioNER (multilingual detection of diseases and drugs in cardiology case reports) and BIONNE (nested named-entity recognition in Russian and English biomedical abstracts), with 37 participating teams and more than 700 total submissions. The paper's central claim is that the top systems, almost all built on large language models and transformer architectures, kept raising the state of the art, with particularly strong performance on yes/no questions and with the two new tasks generating reusable benchmark datasets. A sympathetic reader would care because the results define the current capability baseline for biomedical question answering and clinical named-entity recognition, including multilingual and domain-adaptation settings.

What carries the argument

The central machinery is the shared-task benchmark design: each task combines a public dataset, a scoring protocol, and expert assessment, so that systems are compared on identical inputs. For task 12b the measures are MAP for documents, character-overlap F-measure for snippets, F1/MRR/macro-F1 for exact answers, and manual expert scores for ideal answers; phase A+ isolates answer generation from retrieval. For MultiCardioNER, the key object is the CardioCCC corpus (508 cardiology clinical case reports, 250 held out for testing), used alongside DisTEMIST and DrugTEMIST to test domain adaptation. For BioNNE, the key object is the nested-entity dataset built from NEREL-BIO, with eight biomedical entity types and a macro-F1 metric averaged over classes. These datasets and measures are what let the paper claim meaningful comparisons.

What would settle it

A re-evaluation of task 12b after the ground-truth enrichment is complete would settle the point: if the final yes/no macro-F1 scores drop materially or the system ranking changes, the claim of continued state-of-the-art progress would not survive. Likewise, an ablation training MultiCardioNER models on the same total number of documents but with the cardiology-specific corpus replaced by equal-sized mixed-specialty text would show whether the recall gain is due to domain adaptation or just to more data.

Watch

Extended reading notes

Core claim

The paper claims that the 2024 edition of the challenge demonstrates continued improvement in biomedical question-answering systems, especially on yes/no questions, where top systems approached or reached perfect macro-F1 on some test batches, while factoid and list questions remain harder and more variable. It claims that the new phase A+ shows systems can generate competitive exact and ideal answers without being given manually selected relevant documents, and that providing such documents in phase B still improves answer quality. For the new MultiCardioNER task, the paper argues that incorporating cardiology-specific training data (the CardioCCC corpus) is the decisive factor for disease detection, since systems trained only on mixed-specialty clinical text achieve high precision but lower recall on cardiology-specific entities; drug detection performance is higher overall and fairly comparable across languages, with Italian somewhat lower because fewer clinical pretrained models exist. For BIONNE, the paper claims that a bi-encoder contrastive model fine-tuned on the supplied data clearly outperforms a zero-shot LLM-based extractor, indicating that specialized training data is necessary for nested biomedical named-entity recognition. The paper also reports that the Synergy iteration process enabled experts to reach answers for about 78% of the open questions, with about 51% receiving at least one ideal answer judged to be of ground-truth quality.

Load-bearing premise

The paper's conclusions rely on treating the preliminary task 12b scores as valid evidence of system quality even though the ground truth is still being actively enriched by manual assessment.

Editorial extensions

If this is right

  • Top 12b systems reached perfect or near-perfect macro-F1 on yes/no questions in some batches, while factoid and list questions still showed inconsistent performance.
  • Phase A+ results showed that systems can produce competitive answers without manually selected relevant documents, but releasing such documents in phase B still improved answer quality.
  • In MultiCardioNER, systems that incorporated the cardiology-specific CardioCCC corpus clearly outperformed those using only mixed-specialty clinical data, implying that domain-specific training data drives recall on specialty entities.
  • In BioNNE, the fine-tuned bi-encoder model scored 0.7044 F1 on the bilingual test set, far above the 0.3479 of a zero-shot LLM-based system, implying that nested biomedical named-entity recognition needs supervised training data.
  • The Synergy task reached answer-ready status for about 78% of the 73 open questions, and roughly 51% received at least one ideal answer judged ground-truth quality.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Because the task 12b scores are explicitly preliminary, the ongoing ground-truth enrichment may shift the reported rankings; treating the specific numbers as authoritative would overread the paper.
  • The MultiCardioNER result suggests an ablative explanation the paper leaves open: the benefit of CardioCCC might come from simply having more training instances rather than from domain adaptation, and a controlled data-volume experiment would separate the two.
  • The BioNNE result invites a follow-up test the paper does not run: fine-tuning an LLM on the same nested-entity training data would tell whether the gap is due to model architecture or to the absence of supervised signal.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 4 minor

Summary. This paper is the official overview of the twelfth BioASQ challenge run at CLEF 2024. It describes four shared tasks: task 12b (English biomedical question answering with phases A, A+, and B), Synergy 12 (iterative question answering for open developing problems over four rounds), MultiCardioNER (disease and drug named-entity recognition in cardiology clinical texts in Spanish, English, and Italian), and BioNNE (nested named-entity recognition in English and Russian). For each task it gives corpus statistics, participant counts, system descriptions, and evaluation results, with full results made available online. The abstract's headline quantitative claim is that 37 teams submitted more than 700 distinct submissions across the four tasks, and the conclusions frame the results as continued advancement of the state of the art.

Significance. The paper's value is as an archival record of a large shared-task evaluation. It is not a methods paper; its main contributions are the task definitions, dataset releases, participant and system inventories, and aggregated results, with URLs to data repositories and full result pages. The writing is generally clear, and the per-task descriptions are internally consistent. I credit the authors for explicitly labeling the task 12b results as preliminary in Section 4.1 and for acknowledging in Section 4.3 that the observed benefit of the CardioCCC data may reflect data volume rather than domain adaptation. If the participation-count discrepancy is resolved and the preliminary-result caveat is carried into the abstract and conclusions, the paper would serve as a reliable reference for the community.

major comments (2)
  1. [Abstract; Sections 3.1-3.4] The abstract states that 37 competing teams participated with more than 700 distinct submissions, but the task-level tallies in Sections 3.1-3.4 (89 systems for task 12b, 16 systems for Synergy 12, 70 runs for MultiCardioNER, and 155 runs for BioNNE) sum to 330, and the text does not define how a "distinct submission" is counted. If submissions are counted multiplicatively, for example per phase, round, or language, that convention needs to be stated explicitly and a per-task breakdown should be provided; otherwise the headline figures are not reproducible from the paper. The 37-team figure is likewise not derivable from the text: the per-task counts are 26 + 4 + 7 + 5 = 42, and with only the overlapping teams named in the paper one arrives at 39 rather than 37. Please add a summary table with the counting convention for both teams and submissions, and cite the official challenge logs for these figures.
  2. [Section 4.1; Section 5; Abstract] The task 12b results are explicitly preliminary: Section 4.1 says that final results depend on the ongoing manual assessment of system responses and the enrichment of the ground truth. Despite this, the abstract claims "continuous advancement of the state-of-the-art" and Section 5 asserts that top-performing systems "were able to improve over the state-of-the-art performance from previous years." These claims should be explicitly qualified as based on preliminary results, or deferred until the final task 12b scores are available, since the manual assessment could change the reported rankings and conclusions.
minor comments (4)
  1. [Section 4.2] The sentence "The full 12b results are available online" in the Synergy results subsection should read "The full Synergy 12 results are available online," since the preceding paragraph concerns the Synergy task.
  2. [Section 5] The conclusion says BioASQ has been pushing the research frontier "for eleven years now," but the paper describes the twelfth edition of the challenge; this should be corrected to twelve years or rephrased as "since 2013."
  3. [Section 4.4] In the evaluation metric formula, F1rel_c is used without a clear definition; the text says it is the macro F1-score averaged across all relevance classes, but it would be clearer to state that F1rel_c is the per-class relevance F1 and that the formula is the macro-average over the eight entity classes.
  4. [References] References [16] and [35] contain stray commas in the author lists; please check and normalize the reference formatting.

Circularity Check

0 steps flagged · score 0.0 of 10

No circular derivation: the paper is a shared-task overview whose reported scores come from independent participating teams and expert assessment, and its self-citations only document evaluation conventions and dataset provenance.

full rationale

This paper does not perform a derivation chain; it records the organization and outcomes of the BioASQ 2024 shared tasks. The central reported quantities — team participation, system descriptions, and evaluation scores — originate from independent participating teams (26, 4, 7, and 5 teams across tasks 12b, Synergy 12, MultiCardioNER, and BioNNE respectively) and from manual assessment by BioASQ experts. The paper's heavy citation of previous BioASQ publications is used to state evaluation measures (e.g., modified MAP since BioASQ8, character-overlap F-measure since BioASQ9, manual scores for ideal answers), dataset construction choices, and baseline systems such as OAQA; these are conventions and provenance statements, not fitted parameters that are later relabeled as predictions. No equation in the paper takes an input and returns the same quantity under a new name, and no result is justified solely by a self-citation that itself assumes the target result. Section 4.1 explicitly marks task 12b results as preliminary pending manual ground-truth enrichment, which weakens confidence in those rankings but is not circularity. The abstract's headline figures (37 teams, over 700 submissions) are not straightforwardly reconciled with the task-level tallies summing to about 330 runs/systems, and the counting convention is undefined; however, that is an internal-consistency and completeness issue, not a case of a claimed prediction reducing to its input by construction. Since the paper's substantive claims are externally produced and independently assessable, the appropriate finding is no significant circularity.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The paper introduces no free parameters, and no invented entities. Its main reliance is on domain assumptions about the validity of its evaluation setup, including the manual expert assessments that were still in progress for task 12b.

assumptions (3)
  • domain assumption The evaluation measures (MAP, F-measure, MRR, macro F1) are appropriate for the task comparisons.
    Section 4 uses these metrics for ranking without a comparative justification of metric choice, although they are inherited from prior BioASQ editions.
  • domain assumption Manual expert assessment of ideal and exact answers reflects answer quality accurately.
    Sections 2 and 4 rely on BioASQ expert manual scores for ideal answers and for ground-truth enrichment, which is inherently subjective and is described as still in progress.
  • domain assumption The test datasets are reliable, unbiased samples of the underlying information needs.
    The paper assumes the question sets and annotation guidelines, e.g., DisTEMIST and DrugTEMIST guidelines, produce valid benchmarks, citing them without independent validation.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Overview of BioASQ 2024: The twelfth BioASQ challenge on Large-Scale Biomedical Semantic Indexing and Question Answering." pith.science (2026). https://pith.science/paper/V6VV3VEB

@misc{pith2026250820532,
  author       = {Pith},
  title        = {Pith review of: Overview of BioASQ 2024: The twelfth BioASQ challenge on Large-Scale Biomedical Semantic Indexing and Question Answering},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/V6VV3VEB}},
  note         = {Machine review of arXiv:2508.20532}
}
read the original abstract

This is an overview of the twelfth edition of the BioASQ challenge in the context of the Conference and Labs of the Evaluation Forum (CLEF) 2024. BioASQ is a series of international challenges promoting advances in large-scale biomedical semantic indexing and question answering. This year, BioASQ consisted of new editions of the two established tasks b and Synergy, and two new tasks: a) MultiCardioNER on the adaptation of clinical entity detection to the cardiology domain in a multilingual setting, and b) BIONNE on nested NER in Russian and English. In this edition of BioASQ, 37 competing teams participated with more than 700 distinct submissions in total for the four different shared tasks of the challenge. Similarly to previous editions, most of the participating systems achieved competitive performance, suggesting the continuous advancement of the state-of-the-art in the field.

Figures

Figures reproduced from arXiv: 2508.20532 by the authors.

Figure 1
Figure 1. The preliminary results for task 12b, reveal that the participating [PITH_FULL_IMAGE:figures/full_fig_p014_1.png] view at source ↗
Figure 1
Figure 1. The evaluation scores of the best-performing systems in task B, Phase B, for exact answers, across the twelve years of BioASQ. Since BioASQ6, accuracy (Acc) was replaced by macro F1 as the official measure for Yes/No questions. The black dots indicate an additional batch with questions from new experts [38]. presented in [PITH_FULL_IMAGE:figures/full_fig_p015_1.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

56 extracted references · 41 canonical work pages

  1. [1]

    In: Faggioli, G., Ferro, N., Galuˇ sˇ c´ akov´ a, P., Garc´ ıa Seco de Herrera, A

    Aksenova, A., Datseris, A., Vassileva, S., Boytcheva, S.: Transformer-Based Dis- ease and Drug Named Entity Recognition in Multilingual Clinical Texts: Multi- CardioNER challenge. In: Faggioli, G., Ferro, N., Galuˇ sˇ c´ akov´ a, P., Garc´ ıa Seco de Herrera, A. (eds.) CLEF Working Notes (2024)

  2. [2]

    In: Faggioli, G., Ferro, N., Galuˇ sˇ c´ akov´ a, P., Garc´ ıa Seco de Herrera, A

    Almeida, T., Jonker, R., Reis, J., Almeida, J., Matos, S.: From Retrieval to An- swer Generation: Insights from BioASQ 12 Task B. In: Faggioli, G., Ferro, N., Galuˇ sˇ c´ akov´ a, P., Garc´ ıa Seco de Herrera, A. (eds.) CLEF Working Notes (2024)

  3. [3]

    In: Faggioli, G., Ferro, N., Galuˇ sˇ c´ akov´ a, P., Garc´ ıa Seco de Herrera, A

    Anaya, C., Fernandes, M., Couto, F.: LLM fine-tuning with biomedical open-source data. In: Faggioli, G., Ferro, N., Galuˇ sˇ c´ akov´ a, P., Garc´ ıa Seco de Herrera, A. (eds.) CLEF Working Notes (2024)

  4. [4]

    In: Faggioli, G., Ferro, N., Galuˇ sˇ c´ akov´ a, P., Garc´ ıa Seco de Herrera, A

    Ateia, S., Kruschwitz, U.: Can Open-Source LLMs Compete with Commercial Models? Exploring the Few-Shot Performance of Current GPT Models in Biomed- ical Tasks. In: Faggioli, G., Ferro, N., Galuˇ sˇ c´ akov´ a, P., Garc´ ıa Seco de Herrera, A. (eds.) CLEF Working Notes (2024) 22 A. Nentidis et al

  5. [5]

    Available from World Wide Web: http://alias-i

    Baldwin, B., Carpenter, B.: Lingpipe. Available from World Wide Web: http://alias-i. com/lingpipe (2003)

  6. [6]

    Project deliverable D4.1, UPMC (2013)

    Balikas, G., Partalas, I., Kosmopoulos, A., Petridis, S., Malakasiotis, P., Pavlopou- los, I., Androutsopoulos, I., Baskiotis, N., Gaussier, E., Artieres, T., Gallinari, P.: Evaluation framework specifications. Project deliverable D4.1, UPMC (2013)

  7. [7]

    In: EMNLP (2019)

    Beltagy, I., Lo, K., Cohan, A.: Scibert: Pretrained language model for scientific text. In: EMNLP (2019)

  8. [8]

    Journal of Biomedical Informatics 144, 104431 (2023)

    Buonocore, T.M., Crema, C., Redolfi, A., Bellazzi, R., Parimbelli, E.: Localizing in-domain adaptation of transformer-based biomedical language models. Journal of Biomedical Informatics 144, 104431 (2023)

Show all 56 references
  1. [9]

    In: Faggioli, G., Ferro, N., Galuˇ sˇ c´ akov´ a, P., Garc´ ıa Seco de Herrera, A

    Chih, B.C., Han, J.C., Tzong-Han Tsai, R.: NCU-IISR: Enhancing Biomedical Question Answering with GPT-4 and Retrieval Augmented Generation in BioASQ 12b Phase B. In: Faggioli, G., Ferro, N., Galuˇ sˇ c´ akov´ a, P., Garc´ ıa Seco de Herrera, A. (eds.) CLEF Working Notes (2024)

  2. [10]

    CoRR abs/1911.02116 (2019), http://arxiv.org/abs/1911.02116

    Conneau, A., Khandelwal, K., Goyal, N., Chaudhary, V., Wenzek, G., Guzm´ an, F., Grave, E., Ott, M., Zettlemoyer, L., Stoyanov, V.: Unsupervised cross-lingual representation learning at scale. CoRR abs/1911.02116 (2019), http://arxiv.org/abs/1911.02116

  3. [11]

    In: Faggioli, G., Ferro, N., Galuˇ sˇ c´ akov´ a, P., Garc´ ıa Seco de Herrera, A

    Danu, M.D., Marica, V.G., Suciu, C., Itu, L.M., Farri, O.: Multilingual Clinical NER for Diseases and Medications Recognition in Cardiology Texts using BERT Embeddings. In: Faggioli, G., Ferro, N., Galuˇ sˇ c´ akov´ a, P., Garc´ ıa Seco de Herrera, A. (eds.) CLEF Working Notes (2024)

  4. [12]

    In: CLEF Working Notes (2024)

    Davydova, V., Loukachevitch, N., Tutubalina, E.: Overview of BioNNE Task on Biomedical Nested Named Entity Recognition at BioASQ 2024. In: CLEF Working Notes (2024)

  5. [13]

    CoRR abs/1810.04805 (2018), http://arxiv.org/abs/1810.04805

    Devlin, J., Chang, M., Lee, K., Toutanova, K.: BERT: pre-training of deep bidirec- tional transformers for language understanding. CoRR abs/1810.04805 (2018), http://arxiv.org/abs/1810.04805

  6. [14]

    In: Faggioli, G., Ferro, N., Galuˇ sˇ c´ akov´ a, P., Garc´ ıa Seco de Herrera, A

    Galat, D., Moshkin, S.: Refining Zero-short Approaches for Biomedical Question Answering. In: Faggioli, G., Ferro, N., Galuˇ sˇ c´ akov´ a, P., Garc´ ıa Seco de Herrera, A. (eds.) CLEF Working Notes (2024)

  7. [15]

    In: Faggioli, G., Ferro, N., Galuˇ sˇ c´ akov´ a, P., Garc´ ıa Seco de Herrera, A

    Gao, Y., Zong, L., Li, Y.: Enhancing Biomedical Question Answering with Parameter-Efficient Fine-Tuning and Hierarchical Retrieval Augmented Genera- tion. In: Faggioli, G., Ferro, N., Galuˇ sˇ c´ akov´ a, P., Garc´ ıa Seco de Herrera, A. (eds.) CLEF Working Notes (2024)

  8. [16]

    Evaluation of advance hierarchical classification techniques for scientific literature, patents and clinical trials

    Gasco, L., Nentidis, A., Krithara, A., Estrada-Zavala, D., , Murasaki, R.T., Primo- Pe˜ na, E., Bojo-Canales, C., Paliouras, G., Krallinger, M.: Overview of BioASQ 2021-MESINESP track. Evaluation of advance hierarchical classification techniques for scientific literature, pate...

  9. [17]

    In: Faggioli, G., Ferro, N., Galuˇ sˇ c´ akov´ a, P., Garc´ ıa Seco de Herrera, A

    Gon¸ calves, R., Lam´ urias, A.: Team NOVA LINCS @ BIOASQ12 MultiCardioNER Track: Entity Recognition with Additional Entity Types. In: Faggioli, G., Ferro, N., Galuˇ sˇ c´ akov´ a, P., Garc´ ıa Seco de Herrera, A. (eds.) CLEF Working Notes (2024)

  10. [18]

    ACM Transactions on Computing for Healthcare (HEALTH) 3(1), 1–23 (2021)

    Gu, Y., Tinn, R., Cheng, H., Lucas, M., Usuyama, N., Liu, X., Naumann, T., Gao, J., Poon, H.: Domain-specific language model pretraining for biomedical natural language processing. ACM Transactions on Computing for Healthcare (HEALTH) 3(1), 1–23 (2021)

  11. [19]

    He, P., Gao, J., Chen, W.: Debertav3: Improving deberta using electra-style pre- training with gradient-disentangled embedding sharing (2021)

  12. [20]

    In: International Conference on Learning Representations (2021), https://openreview.net/forum?id=XPZIaotutsD Overview of BioASQ 2024 23

    He, P., Liu, X., Gao, J., Chen, W.: Deberta: Decoding-enhanced bert with disentan- gled attention. In: International Conference on Learning Representations (2021), https://openreview.net/forum?id=XPZIaotutsD Overview of BioASQ 2024 23

  13. [21]

    In: Faggioli, G., Ferro, N., Galuˇ sˇ c´ akov´ a, P., Garc´ ıa Seco de Herrera, A

    Huang, B.W.: Generative Large Language Models Augmented Hybrid Re- trieval System for Biomedical Question Answering. In: Faggioli, G., Ferro, N., Galuˇ sˇ c´ akov´ a, P., Garc´ ıa Seco de Herrera, A. (eds.) CLEF Working Notes (2024)

  14. [22]

    Jiang, A.Q., Sablayrolles, A., Roux, A., Mensch, A., Savary, B., Bamford, C., Chaplot, D.S., de las Casas, D., Hanna, E.B., Bressand, F., Lengyel, G., Bour, G., Lample, G., Lavaud, L.R., Saulnier, L., Lachaux, M.A., Stock, P., Subramanian, S., Yang, S., Antoniak, S., Scao, T.L...

  15. [23]

    In: Faggioli, G., Ferro, N., Galuˇ sˇ c´ akov´ a, P., Garc´ ıa Seco de Herrera, A

    Jonker, R., Almeida, T., Matos, S.: BIT.UA at MultiCardioNER: Adapting a Multi-head CRF for Cardiology. In: Faggioli, G., Ferro, N., Galuˇ sˇ c´ akov´ a, P., Garc´ ıa Seco de Herrera, A. (eds.) CLEF Working Notes (2024)

  16. [24]

    Scientific Data 10(1), 170 (2023)

    Krithara, A., Nentidis, A., Bougiatiotis, K., Paliouras, G.: BioASQ-QA: A manu- ally curated corpus for Biomedical Question Answering. Scientific Data 10(1), 170 (2023)

  17. [25]

    In: Advances in Information Retrieval: ECIR 2021, Virtual Event, March 28–April 1, 2021, Proceedings, Part II 43

    Krithara, A., Nentidis, A., Paliouras, G., Krallinger, M., Miranda, A.: BioASQ at CLEF2021: large-scale biomedical semantic indexing and question answering. In: Advances in Information Retrieval: ECIR 2021, Virtual Event, March 28–April 1, 2021, Proceedings, Part II 43. pp. 62...

  18. [26]

    In: Faggioli, G., Ferro, N., Galuˇ sˇ c´ akov´ a, P., Garc´ ıa Seco de Herrera, A

    Lee, C., Simpson, T.I., Posma, J.M., Lain, A.D.: Comparative Analyses of Multi- lingual Drug Entity Recognition Systems for Clinical Case Reports In Cardiology. In: Faggioli, G., Ferro, N., Galuˇ sˇ c´ akov´ a, P., Garc´ ıa Seco de Herrera, A. (eds.) CLEF Working Notes (2024)

  19. [27]

    Database J

    Li, J., Sun, Y., Johnson, R.J., Sciaky, D., Wei, C., Leaman, R., Davis, A.P., Mattingly, C.J., Wiegers, T.C., Lu, Z.: Biocreative V CDR task cor- pus: a resource for chemical disease relation extraction. Database J. Biol. Databases Curation 2016 (2016). https://doi.org/10.1093...

  20. [28]

    Procesamiento del Lenguaje Natural 71 (2023)

    Lima-L´ opez, S., Farr´ e-Maduell, E., Briv´ a-Escalada, V., Gasc´ o, L., Krallinger, M.: MEDDOPLACE Shared Task overview: recognition, normalization and classifi- cation of locations and patient movement in clinical texts. Procesamiento del Lenguaje Natural 71 (2023)

  21. [29]

    In: Proceedings of the BioCreative VIII Challenge and Workshop: Curation and Evaluation in the era of Generative Models (2023)

    Lima-L´ opez, S., Farr´ e-Maduell, E., Gasco-S´ anchez, L., Rodr´ ıguez-Miret, J., Krallinger, M.: Overview of SympTEMIST at BioCreative VIII: Corpus, Guide- lines and Evaluation of Systems for the Detection and Normalization of Symptoms, Signs and Findings from Text. In: Proc...

  22. [30]

    In: Working Notes of CLEF 2023 (2023)

    Lima-L´ opez, S., Farr´ e-Maduell, E., Gasc´ o, L., Nentidis, A., Krithara, A., Katsim- pras, G., Paliouras, G., Krallinger, M.: Overview of medprocner task on medical procedure detection and entity linking at bioasq 2023. In: Working Notes of CLEF 2023 (2023)

  23. [31]

    In: Faggioli, G., Ferro, N., Galuˇ sˇ c´ akov´ a, P., Garc´ ıa Seco de Herrera, A

    Lima-L´ opez, S., Farr´ e-Maduell, E., Rodr´ ıguez-Miret, J., Rodr´ ıguez-Ortega, M., Lilli, L., Lenkowicz, J., Ceroni, G., Kossoff, J., Shah, A., Nentidis, A., Krithara, A., Katsimpras, G., Paliouras, G., Krallinger, M.: Overview of MultiCardioNER task at BioASQ 2024 on Medic...

  24. [32]

    arXiv preprint arXiv:1907.11692 (2019) 24 A

    Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., Stoyanov, V.: Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692 (2019) 24 A. Nentidis et al

  25. [33]

    Language Resources and Evaluation pp

    Loukachevitch, N., Artemova, E., Batura, T., Braslavski, P., Ivanov, V., Manand- har, S., Pugachev, A., Rozhkov, I., Shelmanov, A., Tutubalina, E., et al.: Nerel: a russian information extraction dataset with rich annotation for nested entities, relations, and wikidata entity ...

  26. [34]

    Bioin- formatics (04 2023)

    Loukachevitch, N., Manandhar, S., Baral, E., Rozhkov, I., Braslavski, P., Ivanov, V., Batura, T., Tutubalina, E.: NEREL-BIO: A Dataset of Biomedical Abstracts Annotated with Nested Named Entities. Bioin- formatics (04 2023). https://doi.org/10.1093/bioinformatics/btad161, http...

  27. [35]

    Miranda-Escalada, A., Gasc´ o, L., Lima-L´ opez, S., Farr´ e-Maduell, E., Estrada, D., Nentidis, A., Krithara, A., Katsimpras, G., Paliouras, G., , Krallinger, M.: Overview of DISTEMIST at BioASQ: Automatic detection and normalization of diseases from clinical texts: results, ...

  28. [36]

    In: Experimental IR Meets Multilinguality, Multimodality, and Interac- tion

    Nentidis, A., Katsimpras, G., Krithara, A., Lima-L´ opez, S., Farr´ e-Maduell, E., Gasco, L., Krallinger, M., Paliouras, G.: Overview of BioASQ 2023: The eleventh BioASQ challenge on Large-Scale Biomedical Semantic Indexing and Question An- swering. In: Experimental IR Meets M...

  29. [37]

    In: Faggioli, G., Ferro, N., Galuˇ sˇ c´ akov´ a, P., Garc´ ıa Seco de Herrera, A

    Nentidis, A., Katsimpras, G., Krithara, A., Paliouras, G.: Overview of BioASQ Tasks 12b and Synergy12 in CLEF2024. In: Faggioli, G., Ferro, N., Galuˇ sˇ c´ akov´ a, P., Garc´ ıa Seco de Herrera, A. (eds.) CLEF Working Notes (2024)

  30. [38]

    In: International Conference of the Cross-Language Evaluation Forum for European Languages

    Nentidis, A., Katsimpras, G., Vandorou, E., Krithara, A., Gasco, L., Krallinger, M., Paliouras, G.: Overview of BioASQ 2021: The Ninth BioASQ Challenge on Large- Scale Biomedical Semantic Indexing and Question Answering. In: International Conference of the Cross-Language Evalu...

  31. [39]

    In: Experimental IR Meets Multilinguality, Multimodality, and Inter- action (2022)

    Nentidis, A., Katsimpras, G., Vandorou, E., Krithara, A., Miranda-Escalada, A., Gasco, L., Krallinger, M., Paliouras, G.: Overview of BioASQ 2022: The Tenth BioASQ Challenge on Large-Scale Biomedical Semantic Indexing and Question Answering. In: Experimental IR Meets Multiling...

  32. [40]

    In: International Conference of the Cross-Language Evaluation Forum for European Languages

    Nentidis, A., Krithara, A., Bougiatiotis, K., Krallinger, M., Rodriguez-Penagos, C., Villegas, M., Paliouras, G.: Overview of BioASQ 2020: The Eighth BioASQ Chal- lenge on Large-Scale Biomedical Semantic Indexing and Question Answering. In: International Conference of the Cros...

  33. [41]

    In: Advances in Information Retrieval: ECIR 2023, Dublin, Ireland, April 2–6, 2023, Proceedings, Part III

    Nentidis, A., Krithara, A., Paliouras, G., Farre-Maduell, E., Lima-Lopez, S., Krallinger, M.: BioASQ at CLEF2023: The Eleventh Edition of the Large-Scale Biomedical Semantic Indexing and Question Answering Challenge. In: Advances in Information Retrieval: ECIR 2023, Dublin, Ir...

  34. [42]

    In: Advances in Information Retrieval: ECIR 2022, Stavanger, Norway, April 10–14, 2022, Proceedings, Part II

    Nentidis, A., Krithara, A., Paliouras, G., Gasco, L., Krallinger, M.: BioASQ at CLEF2022: The Tenth Edition of the Large-scale Biomedical Semantic Indexing and Question Answering Challenge. In: Advances in Information Retrieval: ECIR 2022, Stavanger, Norway, April 10–14, 2022,...

  35. [43]

    In: ECIR2024

    Nentidis, A., Krithara, A., Paliouras, G., Krallinger, M., Sanchez, L.G., Lima, S., Farre, E., Loukachevitch, N., Davydova, V., Tutubalina, E.: BioASQ at CLEF2024: The Twelfth Edition of the Large-Scale Biomedical Semantic Indexing and Ques- tion Answering Challenge. In: ECIR2...

  36. [44]

    In: Faggioli, G., Ferro, N., Galuˇ sˇ c´ akov´ a, P., Garc´ ıa Seco de Herrera, A

    Panou, D., Dimopoulos, A., Reczko, M.: Farming Open LLMs for Biomedical Ques- tion Answering. In: Faggioli, G., Ferro, N., Galuˇ sˇ c´ akov´ a, P., Garc´ ıa Seco de Herrera, A. (eds.) CLEF Working Notes (2024)

  37. [45]

    In: CLEF Working Notes (2024)

    Rehana, H., Bansal, B., Bengisu C ¸ am, N., Zheng, J., He, Y., ¨Ozg¨ ur, A., Hur, J.: Nested Named Entity Recognition using Multilayer BERT-based Model. In: CLEF Working Notes (2024)

  38. [46]

    In: Faggioli, G., Ferro, N., Galuˇ sˇ c´ akov´ a, P., Garc´ ıa Seco de Herrera, A

    Reimer, J.H., Bondarenko, A., Hagen, M., Viehweger, A.: MiBi at BioASQ 2024: Retrieval-Augmented Generation for Answering Biomedical Questions. In: Faggioli, G., Ferro, N., Galuˇ sˇ c´ akov´ a, P., Garc´ ıa Seco de Herrera, A. (eds.) CLEF Working Notes (2024)

  39. [47]

    In: Faggioli, G., Ferro, N., Galuˇ sˇ c´ akov´ a, P., Garc´ ıa Seco de Herrera, A

    Romano, A., Riccio, G., Postiglione, M., Moscato, V.: Identifying Cardiological Disorders in Spanish via Data Augmentation and Fine-Tuned Language Models. In: Faggioli, G., Ferro, N., Galuˇ sˇ c´ akov´ a, P., Garc´ ıa Seco de Herrera, A. (eds.) CLEF Working Notes (2024)

  40. [48]

    Pattern Recognition and Image Analysis 33(2), 122–131 (2023)

    Rozhkov, I., Loukachevitch, N.: Prompts in few-shot named entity recognition. Pattern Recognition and Image Analysis 33(2), 122–131 (2023)

  41. [49]

    In: Faggioli, G., Ferro, N., Galuˇ sˇ c´ akov´ a, P., Garc´ ıa Seco de Herrera, A

    Styll, P., Campillos-Llanos, L., Kusa, W., Hanbury, A.: Cross-Linguistic Disease and Drug Detection in Cardiology Clinical Texts: Methods and Outcomes. In: Faggioli, G., Ferro, N., Galuˇ sˇ c´ akov´ a, P., Garc´ ıa Seco de Herrera, A. (eds.) CLEF Working Notes (2024)

  42. [50]

    BMC Bioinformatics 16, 138 (2015)

    Tsatsaronis, G., Balikas, G., Malakasiotis, P., Partalas, I., Zschunke, M., Alvers, M.R., Weissenborn, D., Krithara, A., Petridis, S., Polychronopoulos, D., Almiran- tis, Y., Pavlopoulos, J., Baskiotis, N., Gallinari, P., Artieres, T., Ngonga, A., Heino, N., Gaussier, E., Barr...

  43. [51]

    ACL 2016 p

    Yang, Z., Zhou, Y., Eric, N.: Learning to answer biomedical questions: Oaqa at bioasq 4b. ACL 2016 p. 23 (2016)

  44. [52]

    In: Association for Computational Linguistics (ACL) (2022)

    Yasunaga, M., Leskovec, J., Liang, P.: Linkbert: Pretraining language models with document links. In: Association for Computational Linguistics (ACL) (2022)

  45. [53]

    In: The Eleventh International Conference on Learning Representations (2022)

    Zhang, S., Cheng, H., Gao, J., Poon, H.: Optimizing bi-encoder for named entity recognition via contrastive learning. In: The Eleventh International Conference on Learning Representations (2022)

  46. [54]

    In: CLEF Working Notes (2024)

    Zhou, W.: Biomedical Nested NER with Large Language Model and UMLS Heuris- tics. In: CLEF Working Notes (2024)

  47. [55]

    In: Faggioli, G., Ferro, N., Galuˇ sˇ c´ akov´ a, P., Garc´ ıa Seco de Herrera, A

    Zhou, W., Ngo, T.H.: Using Pretrained Large Language Model with Prompt Engi- neering to Answer Biomedical Questions. In: Faggioli, G., Ferro, N., Galuˇ sˇ c´ akov´ a, P., Garc´ ıa Seco de Herrera, A. (eds.) CLEF Working Notes (2024)

  48. [56]

    In: Faggioli, G., Ferro, N., Galuˇ sˇ c´ akov´ a, P., Garc´ ıa Seco de Herrera, A

    S ¸erbet¸ ci, O., Wang, X.D., Leser, U.: HU-WBI at BioASQ12B Phase A: Exploring Rank Fusion of Dense Retrievers for Biomedical Question Answering. In: Faggioli, G., Ferro, N., Galuˇ sˇ c´ akov´ a, P., Garc´ ıa Seco de Herrera, A. (eds.) CLEF Working Notes (2024)

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.