MMed-Bench-IR is a new heterogeneous benchmark spanning 6 languages and three non-overlapping tasks that exposes severe cross-lingual drops in biomedical retrieval performance.
Apollo: Lightweight multilingual medical llms towards democratizing medical ai to 6b people
4 Pith papers cite this work. Polarity classification is still indexing.
representative citing papers
EHRBench uses an EHR-LLM-KB pipeline to automatically create 960,067 reliable QA items spanning diagnosis, treatment, and prognosis for large-scale LLM evaluation in clinical decision making.
HuatuoGPT-o1 achieves superior medical complex reasoning by using a verifier to curate reasoning trajectories for fine-tuning and then applying RL with verifier-based rewards.
citing papers explorer
-
MMed-Bench-IR: A Heterogeneous Benchmark for Multilingual Medical Information Retrieval
MMed-Bench-IR is a new heterogeneous benchmark spanning 6 languages and three non-overlapping tasks that exposes severe cross-lingual drops in biomedical retrieval performance.
-
EHRBench: An Automated and Reliable EHR-based Benchmark for Clinical Decision Making with LLMs
EHRBench uses an EHR-LLM-KB pipeline to automatically create 960,067 reliable QA items spanning diagnosis, treatment, and prognosis for large-scale LLM evaluation in clinical decision making.
-
HuatuoGPT-o1, Towards Medical Complex Reasoning with LLMs
HuatuoGPT-o1 achieves superior medical complex reasoning by using a verifier to curate reasoning trajectories for fine-tuning and then applying RL with verifier-based rewards.
- Language corpora for the Dutch medical domain