A new Polish-English medical exam benchmark shows GPT-4o answering at or above average human level, with smaller and medical-specific models lagging and persistent cross-lingual gaps.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Polish-English medical knowledge transfer: A new benchmark and results
A new Polish-English medical exam benchmark shows GPT-4o answering at or above average human level, with smaller and medical-specific models lagging and persistent cross-lingual gaps.