Pith. sign in

REVIEW 3 major objections 5 minor 1 cited by

Polish-English medical knowledge transfer: A new benchmark and results

T0 review · 3 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read A benchmark built from Polish medical licensing exams shows GPT-4o performing at roughly the level of an average medical student, while most other models answer English versions of the same questions better than Polish ones.

desk verdict Useful Polish medical benchmark and professional PL-EN parallel data, but the near-human GPT-4o results rest on an acknowledged and unmitigated contamination risk. read the letter →

arxiv 2412.00559 v2 pith:NKXFVS66 submitted 2024-11-30 cs.CL cs.AI

classification cs.CLcs.AI
keywords PolishmedicalexamsLEKLDEKPESquestionansweringcross-lingualevaluationlargelanguagemodelsbenchmark
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper assembles a benchmark from Polish medical licensing and specialization exams (LEK, LDEK, and PES) that were published by the Medical Examination Center and the Chief Medical Chamber, producing over 22,000 usable multiple-choice questions plus a parallel subset in which the English versions are professional translations made by the examination center itself. It uses this benchmark to ask how well large language models answer Polish medical questions, how their scores compare with those of human examinees, and how much medical knowledge transfers from English to Polish. On the benchmark, GPT-4o scores near the average human on LEK and LDEK, passes 68 of 72 PES exams, and outperforms the median human on a majority of PES specializations, while almost every other tested model scores higher on English than on Polish versions of identical questions. The authors conclude that general-purpose models outperform medical-specific models, that the Polish-English gap narrows as model quality improves, and that performance remains too uneven across medical specialties for unsupervised clinical use.

What carries the argument

The load-bearing object is the parallel Polish-English benchmark. Its engine is the examination center's professionally translated English versions of LEK and LDEK questions, which make the Polish and English forms semantically equivalent and thereby let a score difference between languages be read as a language-transfer effect rather than translation noise. The three exam types give the benchmark a difficulty gradient: LEK and LDEK are final licensing exams with a high share of questions from a public bank, while PES is a harder specialization exam whose questions are not public. The same questions are paired with anonymized human score distributions, which is what allows model accuracy to be converted into human-percentile comparisons and per-specialty pass/fail judgments.

What would settle it

Have the Medical Examination Center or an independent panel write new LEK, LDEK, and PES-style questions that have never been published, give the same models the same prompt, and compare accuracy on the private set with accuracy on the public benchmark; if GPT-4o's near-human scores drop substantially on the private set while human scores do not, the public benchmark's headline numbers overestimate model capability due to memorization.

Watch

Extended reading notes

Core claim

The central claim is that a structured benchmark built from publicly available LEK, LDEK, and PES exam questions can support valid measurement of Polish medical question answering and of cross-lingual medical knowledge transfer, because the English portion is a human-expert translation of the Polish portion produced by the examination center. On that benchmark, the paper reports GPT-4o answering 89.4% of LEK questions correctly, 75.35% of PES questions correctly, and scoring within one standard deviation of the average human on LEK and LDEK, while Meta-Llama-3.1-70B-Instruct is the best open model. It reports that most models score higher on English versions of the same questions, that the gap narrows as overall performance improves (about 13 percentage points for Llama-3.1-8B on LEK versus less than 2 points for Llama-3.1-70B), and that GPT-4o scores slightly higher in Polish than in English. The paper also claims that medical-specific models fine-tuned on English data do not beat general-purpose models on these Polish exams, which it attributes to the language mismatch of the fine-tuning data.

Load-bearing premise

The whole measurement stands or falls on the assumption that a model's score reflects its medical knowledge rather than its memory of questions it already saw during training, because every question was publicly available before several evaluated models' training cutoffs and the paper does not filter for overlap.

Editorial extensions

If this is right

  • If the benchmark measures capability as intended, GPT-4o is the only tested model that consistently performs at the level of an average medical student, and the only one that passes nearly every PES specialization.
  • Most models answering identical questions better in English means that English-centric training leaves a residual Polish medical knowledge deficit; for smaller models, evaluating in English would overstate their Polish clinical competence.
  • The narrowing of the Polish-English gap with model scale implies that larger models transfer medical knowledge across languages more effectively, so cross-lingual capability should be reported as a function of model size.
  • General-purpose models beating medical-specific models implies that English-only medical fine-tuning does not translate into an advantage on Polish exams; specialized models need per-language evidence before deployment.
  • Per-specialty differences, with dental specialties hardest and laboratory diagnostics easiest, imply that LLM deployment in Polish medicine should be specialty-aware rather than uniform.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A private held-out set of newly written, never-published exam questions would separate genuine medical reasoning from memorization; if GPT-4o's near-human scores drop sharply on that set, the public benchmark's headline numbers should be read as upper bounds.
  • The benchmark's Polish-English parallel structure could be extended by having the exam center translate a PES subset; this would test whether the cross-lingual gap grows with specialization difficulty, which the current data cannot answer.
  • The question-level agreement pattern suggests a practical clinician-facing filter: asking a model the same question in both languages and flagging answers that disagree could highlight low-confidence responses for human review, a use the paper only hypothesizes.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper introduces a new benchmark for Polish medical question answering, built from public LEK, LDEK, and PES licensing and specialization exams, including a Polish-English parallel subset with professionally produced English translations. The authors evaluate sixteen LLMs, reporting per-exam accuracy and pass rates (Tables 2 and 3), cross-lingual comparisons on matched Polish/English question subsets (Tables 4 and 5), and comparisons with human exam performance (Tables 6 and 7). The headline claims are that GPT-4o achieves near-human performance on these exams, that most models perform better in English than in Polish, and that general-purpose multilingual models outperform medical-specific models.

Significance. If the benchmark measures what it claims, it is a useful resource for non-English medical QA evaluation and cross-lingual knowledge-transfer studies. The dataset draws on an authoritative examination source, includes professionally translated parallel questions rather than machine-translated ones, covers a long time span (2008-2024), and is publicly released. The appendices document data acquisition and preprocessing in unusual detail, and the inclusion of human results is a valuable feature. The main significance hinges on whether the public, pre-cutoff exam questions can support capability claims rather than memorization-based scores; this is the central validity question the paper does not resolve.

major comments (3)
  1. [Limitations (final section); Tables 2, 3, and 7] The contamination risk is acknowledged but not mitigated, and it is load-bearing for the paper's central claim. All exam questions are publicly available on the CEM and NIL websites and predate the training cutoffs of several evaluated models, including gpt-4o-2024-08-06 and Llama-3.1-70B-Instruct. The Limitations paragraph states that 'there is a potential risk for these exams being included in the training datasets of evaluated LLMs' and then compares the situation to MMLU, but that analogy concedes rather than resolves the problem. The statement that 'the training dataset is not provided' is also not a mitigation, because the questions themselves were already public. As a result, the near-human GPT-4o results in Table 7, the pass counts in Table 3, and the accuracy ranking in Table 2 are not interpretable as measures of medical reasoning ability unless the authors provide a concrete contamination analysis. I would ask for at least one of the following: a question-level overlap analysis with known training corpora, an evaluation restricted to questions published after the model training cutoffs, an answer-only memorization probe, or a clear reframing of the results as an upper-bound benchmark in which memorization is an explicitly uncontrolled factor.
  2. [Section 4, Tables 2-5; Section 5, Tables 4-5] All accuracy results are single-run point estimates with no error bars, confidence intervals, or significance tests. For exams with roughly 200 questions, a difference of two to three percentage points can easily lie within binomial sampling noise, yet the paper makes comparative claims such as 'GPT-4o is the best performing model overall' and 'the performance gap between languages narrows' without any uncertainty quantification. The authors should report repeated runs with the decoding parameters (temperature, top-p, number of runs), or at minimum binomial confidence intervals for each reported accuracy. Without this, the ranking of closely clustered models in Tables 2 and 4 and the cross-lingual gap analysis cannot be evaluated.
  3. [Section 6, Table 7] The human comparison may not be apples-to-apples. The dataset description states that questions containing images and invalidated questions were removed, and Table 1 reports 'invalidated questions' for every sub-dataset. However, Table 7 compares raw model scores (e.g., 184 out of 200 for GPT-4o) with published human averages and standard deviations from the CEM webpage. It is not stated whether the human statistics were recomputed on the same filtered question subset or whether the model scores were rescaled to account for excluded questions. In addition, the text says the comparison covers '977 LEK and 984 LDEK questions' from four sessions of each exam, but each LEK and LDEK session has only 200 questions, so four sessions yield at most 800 questions; this number needs clarification. The authors should specify the identical question set used for both humans and models, and recompute human statistics on that set if necessary.
minor comments (5)
  1. [Abstract; Table 1] The abstract states that the dataset 'comprises over 24,000 exam questions,' but Table 1 sums to 22,604 valid questions plus 436 invalidated questions, i.e., 23,040 total. Section 1 says 'over 22,000 questions.' The numbers should be reconciled, and the abstract's figure should be corrected.
  2. [Section 3, Table 1] The text says 'For the PES dataset, we collected a total of 180,712 questions,' but Table 1 reports 8,532 valid PES questions. If the 180,712 figure is the raw scraped count before filtering and selection of the most recent exam per specialty, this should be stated explicitly and connected to the analysis set.
  3. [Appendix F, Tables 10 and 11] In Table 10, the last column header 'Incorrect PL, Incorrect EN' appears to be a typo for 'Incorrect PL, Correct EN' based on the category definitions in the same appendix. Table 11 uses the correct 'Correct EN, Incorrect PL' label for the corresponding column.
  4. [Section 4, model list and Table 2] The model naming is inconsistent: the text lists 'GPT-4-o' while Table 2 uses 'gpt-4o-2024-08-06' and 'gpt-4o-mini-2024-07-18.' Standardizing model names across the text and tables would improve reproducibility.
  5. [Limitations section] The sentence 'To prevent models from being trained on the benchmark data, the training dataset is not provided' is confusing because the benchmark itself is released on Hugging Face. Clarify what is withheld, and note that public availability of the original exam questions makes this statement ineffective as a contamination safeguard.

Circularity Check

0 steps flagged · score 0.0 of 10

No circular derivation: the benchmark scores are external measurements against official exam keys; the acknowledged training-data contamination risk is a validity concern, not circularity.

full rationale

The paper contains no derivation chain in which an output is defined in terms of its own input. The benchmark is constructed from publicly available official Polish medical exam questions (LEK, LDEK, PES) with answers provided by the Medical Examination Center and the Supreme Medical Chamber, and the English portions are professional translations supplied by the examination center. LLM performance is then measured by prompting models on these fixed questions and scoring against official answer keys. There are no fitted parameters, no model-derived labels, and no prediction that is equivalent to a construction choice. The comparison with human examinees uses published anonymized human results, again an external reference. The only self-citation, Pokrywka et al. (2024), is used as related-work context for prior PES observations and is not load-bearing for the benchmark's validity or for any of the reported results. The paper explicitly acknowledges in the Limitations that the public exam questions may have appeared in LLM training corpora: 'There is a potential risk for these exams being included in the training datasets of evaluated LLMs.' That is a legitimate threat to the interpretability of the accuracy numbers as measures of reasoning rather than memorization, but it is not a circularity of the kind where a claimed prediction reduces to an input by construction. The evaluation is self-contained against external data and official answer keys, so the appropriate finding is no significant circularity.

Assumptions & free parameters 0 free parameters · 5 assumptions · 0 invented entities

The benchmark relies on official exam sources, preprocessing choices, and several comparability assumptions. All are disclosed, but contamination, translation equivalence, and human-score comparability are not quantitatively validated.

assumptions (5)
  • domain assumption Official exam answer keys from CEM and NIL are correct and current.
    Benchmark accuracy is measured against the published correct answers. If a key is wrong or outdated, reported model scores are affected. The paper discusses invalidated and annulled questions in Appendix E.
  • domain assumption Publicly available exam questions were not memorized by the evaluated models.
    The paper's performance claims assume no training-data contamination. The Limitations section acknowledges this risk explicitly but provides no test or mitigation.
  • domain assumption The English translations are semantically equivalent and comparable in difficulty to the Polish originals.
    Section 5 relies on exam-center translations being faithful and equally difficult. The paper asserts equivalence but does not verify it with back-translation or human difficulty ratings.
  • domain assumption LEK and LDEK human exam scores are approximately normally distributed.
    Section 6 interprets average plus or minus standard deviation as a typical student range, which assumes normality of the human score distribution.
  • domain assumption Aggregating human PES results across selected sessions is comparable to the single most recent LLM exam.
    Appendix G joins human results from 12 sessions with one LLM exam per specialty. Specialty-name normalization and small populations make this comparison approximate.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Polish-English medical knowledge transfer: A new benchmark and results." pith.science (2026). https://pith.science/paper/NKXFVS66

@misc{pith2026241200559,
  author       = {Pith},
  title        = {Pith review of: Polish-English medical knowledge transfer: A new benchmark and results},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/NKXFVS66}},
  note         = {Machine review of arXiv:2412.00559}
}
read the original abstract

Large Language Models (LLMs) have demonstrated significant potential in handling specialized tasks, including medical problem-solving. However, most studies predominantly focus on English-language contexts. This study introduces a novel benchmark dataset based on Polish medical licensing and specialization exams (LEK, LDEK, PES) taken by medical doctor candidates and practicing doctors pursuing specialization. The dataset was web-scraped from publicly available resources provided by the Medical Examination Center and the Chief Medical Chamber. It comprises over 24,000 exam questions, including a subset of parallel Polish-English corpora, where the English portion was professionally translated by the examination center for foreign candidates. By creating a structured benchmark from these existing exam questions, we systematically evaluate state-of-the-art LLMs, including general-purpose, domain-specific, and Polish-specific models, and compare their performance against human medical students. Our analysis reveals that while models like GPT-4o achieve near-human performance, significant challenges persist in cross-lingual translation and domain-specific understanding. These findings underscore disparities in model performance across languages and medical specialties, highlighting the limitations and ethical considerations of deploying LLMs in clinical practice.

Figures

Figures reproduced from arXiv: 2412.00559 by the authors.

Figure 1
Figure 1. Models performance on different specialties on PES exams (part 1/2). Dotted lines indicate the passing [PITH_FULL_IMAGE:figures/full_fig_p014_1.png] view at source ↗
Figure 2
Figure 2. Models performance on different specialties on PES exams (part 2/2). Dotted lines indicate the passing [PITH_FULL_IMAGE:figures/full_fig_p015_2.png] view at source ↗
Figure 3
Figure 3. Quiz interface on the Medical Examination [PITH_FULL_IMAGE:figures/full_fig_p016_3.png] view at source ↗
Figures from the paper (6 more)
Figure 5
Figure 5. Figure 5: Example of missing data caused by an absent [PITH_FULL_IMAGE:figures/full_fig_p017_5.png]
Figure 4
Figure 4. Figure 4: Data acquisition and processing workflow [PITH_FULL_IMAGE:figures/full_fig_p017_4.png]
Figure 7
Figure 7. Figure 7: Answer options presented horizontally, verti [PITH_FULL_IMAGE:figures/full_fig_p017_7.png]
Figure 8
Figure 8. Figure 8: Students performance compared to top-performing LLMs on different specialties on PES exam (part 1/3). [PITH_FULL_IMAGE:figures/full_fig_p020_8.png]
Figure 9
Figure 9. Figure 9: Students performance compared to top-performing LLMs on different specialties on PES exam (part 2/3). [PITH_FULL_IMAGE:figures/full_fig_p021_9.png]
Figure 10
Figure 10. Figure 10: Students performance compared to top-performing LLMs on different specialties on PES exam (part [PITH_FULL_IMAGE:figures/full_fig_p022_10.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LLMzSz{\L}: a comprehensive LLM benchmark for Polish

    cs.CL 2025-01 conditional novelty 6.0 of 10

    LLMzSzŁ is a new benchmark of almost 19,000 Polish national exam questions with evaluations of 38 language models and comparisons to human results.

Reference graph

Works this paper leans on

46 extracted references · 18 canonical work pages · cited by 1 Pith paper

  1. [1]

    Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774

  2. [2]

    Rohan Anil, Andrew M Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, et al. 2023. Palm 2 technical report. arXiv preprint arXiv:2305.10403

  3. [3]

    Do Large Language Models have Shared Weaknesses in Medical Question Answering?

    Andrew M. Bean, Karolina Korgul, Felix Krones, Robert McCraith, and Adam Mahdi. 2024. http://arxiv.org/abs/2310.07225 Do large language models have shared weaknesses in medical question answering?

  4. [4]

    Martin Juan Jos \'e Bucher and Marco Martini. 2024. Fine-tuned'small'llms (still) significantly outperform zero-shot generative ai models in text classification. arXiv preprint arXiv:2406.08660

  5. [5]

    Yupeng Chang, Xu Wang, Jindong Wang, Yuan Wu, Linyi Yang, Kaijie Zhu, Hao Chen, Xiaoyuan Yi, Cunxiang Wang, Yidong Wang, et al. 2024. A survey on evaluation of large language models. ACM Transactions on Intelligent Systems and Technology, 15(3):1--45

  6. [6]

    Michelle Clark and Sharon Bailey. 2024. Chatbots in health care: Connecting patients to information. Canadian Journal of Health Technologies, 4(1)

  7. [7]

    Jan Clusmann, Fiona R Kolbinger, Hannah Sophie Muti, Zunamys I Carrero, Jan-Niklas Eckardt, Narmin Ghaffari Laleh, Chiara Maria Lavinia L \"o ffler, Sophie-Caroline Schwarzkopf, Michaela Unger, Gregory P Veldhuizen, et al. 2023. The future landscape of large language models in medicine. Communications medicine, 3(1):141

  8. [8]

    Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. 2024. The llama 3 herd of models. arXiv preprint arXiv:2407.21783

Show all 46 references
  1. [9]

    Clémentine Fourrier, Nathan Habib, Alina Lozovskaya, Konrad Szafer, and Thomas Wolf. 2024. Open llm leaderboard v2. https://huggingface.co/spaces/open-llm-leaderboard/open_llm_leaderboard

  2. [10]

    Gemini, Rohan Anil, Sebastian Borgeaud, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, Katie Millican, et al. 2023. Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805

  3. [11]

    Stefan Harrer. 2023. Attention is not all you need: the complicated case of ethically using large language models in healthcare and medicine. EBioMedicine, 90

  4. [12]

    Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2020. Measuring massive multitask language understanding. arXiv preprint arXiv:2009.03300

  5. [13]

    Niclas Hertzberg and Anna Lokrantz. 2024. Medqa-swe-a clinical question & answer dataset for swedish. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), pages 11178--11186

  6. [14]

    Albert Qiaochu Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de Las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, L'elio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, T...

  7. [15]

    Di Jin, Eileen Pan, Nassim Oufattole, Wei-Hung Weng, Hanyi Fang, and Peter Szolovits. 2021. What disease does this patient have? a large-scale open domain question answering dataset from medical exams. Applied Sciences, 11(14):6421

  8. [16]

    Qiao Jin, Bhuwan Dhingra, Zhengping Liu, William W Cohen, and Xinghua Lu. 2019. Pubmedqa: A dataset for biomedical research question answering. arXiv preprint arXiv:1909.06146

  9. [17]

    Yiqiao Jin, Mohit Chandra, Gaurav Verma, Yibo Hu, Munmun De Choudhury, and Srijan Kumar. 2024. https://doi.org/10.1145/3589334.3645643 Better to ask in english: Cross-lingual evaluation of large language models for healthcare queries . In Proceedings of the ACM Web Conference ...

  10. [18]

    johnsnowlabs. 2024. Jsl-medllama-3-8b-v2.0. https://huggingface.co/johnsnowlabs/JSL-MedLlama-3-8B-v2.0. Accessed: 2024-11-02

  11. [19]

    Mert Karabacak and Konstantinos Margetis. 2023. Embracing large language models for medical applications: opportunities and challenges. Cureus, 15(5)

  12. [20]

    Jungo Kasai, Yuhei Kasai, Keisuke Sakaguchi, Yutaro Yamada, and Dragomir Radev. 2023. http://arxiv.org/abs/2303.18027 Evaluating gpt-4 and chatgpt on japanese medical licensing examinations

  13. [21]

    Markus Kipp. 2024. https://doi.org/10.3390/info15090543 From gpt-3.5 to gpt-4.o: A leap in ai’s medical exam performance . Information, 15(9)

  14. [22]

    Jakub Kufel, Iga Paszkiewicz, Micha Biel \'o wka, Wiktoria Bartnikowska, Micha Janik, Magdalena Stencel, ukasz Czogalik, Katarzyna Gruszczy \'n ska, and Sylwia Mielcarska. 2023. Will chatgpt pass the polish specialty exam in radiology and diagnostic imaging? insights into stre...

  15. [23]

    Yanis Labrak, Adrien Bazoge, Emmanuel Morin, Pierre-Antoine Gourraud, Mickael Rouvier, and Richard Dufour. 2024. http://arxiv.org/abs/2402.10373 Biomistral: A collection of open-source pretrained large language models for medical domains

  16. [24]

    Peter Lee, Sebastien Bubeck, and Joseph Petro. 2023. https://doi.org/10.1056/NEJMsr2214184 Benefits, limits, and risks of gpt-4 as an ai chatbot for medicine . New England Journal of Medicine, 388(13):1233--1239

  17. [25]

    Hanzhou Li, John T Moon, Saptarshi Purkayastha, Leo Anthony Celi, Hari Trivedi, and Judy W Gichoya. 2023. Ethics of large language models in medicine and medical research. The Lancet Digital Health, 5(6):e333--e335

  18. [26]

    Jialin Liu, Changyu Wang, and Siru Liu. 2023. https://doi.org/10.2196/48568 Utility of chatgpt in clinical practice . J Med Internet Res, 25:e48568

  19. [27]

    Junling Liu, Peilin Zhou, Yining Hua, Dading Chong, Zhongyu Tian, Andrew Liu, Helin Wang, Chenyu You, Zhenhua Guo, Lei Zhu, et al. 2024. Benchmarking large language models on cmexam-a comprehensive chinese medical exam dataset. Advances in Neural Information Processing Systems, 36

  20. [28]

    Shervin Minaee, Tomas Mikolov, Narjes Nikzad, Meysam Chenaghlu, Richard Socher, Xavier Amatriain, and Jianfeng Gao. 2024. Large language models: A survey. arXiv preprint arXiv:2402.06196

  21. [29]

    Zabir Al Nazi and Wei Peng. 2024. Large language models in healthcare and medical domain: A review. In Informatics, volume 11, page 57. MDPI

  22. [30]

    Harsha Nori, Yin Tat Lee, Sheng Zhang, Dean Carignan, Richard Edgar, Nicolo Fusi, Nicholas King, Jonathan Larson, Yuanzhi Li, Weishung Liu, et al. 2023. Can generalist foundation models outcompete special-purpose tuning? case study in medicine. arXiv preprint arXiv:2311.16452

  23. [31]

    Krzysztof Ociepa, Łukasz Flis, Krzysztof Wróbel, Adrian Gwoździej, and SpeakLeash Team and Cyfronet Team . 2024. https://huggingface.co/speakleash/Bielik-7B-v0.1 Introducing bielik-7b-v0.1: Polish language model . Accessed: 2024-11-02

  24. [32]

    OpenMeditron. 2024. Meditron3-70b. https://huggingface.co/OpenMeditron/Meditron3-70B. Accessed: 2024-11-02

  25. [33]

    Ankit Pal, Logesh Kumar Umapathi, and Malaikannan Sankarasubbu. 2022. Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question answering. In Conference on health, inference, and learning, pages 248--260. PMLR

  26. [34]

    Ye-Jean Park, Abhinav Pillai, Jiawen Deng, Eddie Guo, Mehul Gupta, Mike Paget, and Christopher Naugler. 2024. Assessing the research landscape and clinical utility of large language models: A scoping review. BMC Medical Informatics and Decision Making, 24(1):72

  27. [35]

    Jakub Pokrywka, Jeremi Kaczmarek, and Edward Gorzelańczyk. 2024. http://arxiv.org/abs/2405.01589 Gpt-4 passes most of the 297 written polish board certification examinations

  28. [36]

    Maciej Roso , Jakub S G a sior, Jonasz aba, Kacper Korzeniewski, and Marcel M y \'n czak. 2023. Evaluation of the performance of gpt-3.5 and gpt-4 on the polish medical final examination. Scientific Reports, 13(1):20512

  29. [37]

    Karan Singhal, Tao Tu, Juraj Gottweis, Rory Sayres, Ellery Wulczyn, Mohamed Amin, Le Hou, Kevin Clark, Stephen R Pfohl, Heather Cole-Lewis, et al. 2025. Toward expert-level medical question answering with large language models. Nature Medicine, pages 1--8

  30. [38]

    S Suwa a, P Szulc, A Dudek, A Bia czyk, K Koperska, and R Junik. 2023. Chatgpt fails the internal medicine state specialization exam in poland: artificial intelligence still has much to learn. Pol Arch Intern Med, 133(11):16608

  31. [39]

    Qwen Team. 2024. https://qwenlm.github.io/blog/qwen2.5/ Qwen2.5: A party of foundation models

  32. [40]

    Ehsan Ullah, Anil Parwani, Mirza Mansoor Baig, and Rajendra Singh. 2024. Challenges and barriers of using large language models (llm) such as chatgpt for diagnostic medicine with a focus on digital pathology--a recent scoping review. Diagnostic pathology, 19(1):43

  33. [41]

    Simona Wojcik, Anna Rulkiewicz, Piotr Pruszczyk, Wojciech Lisik, Marcin Poboży, Iwona Pilchowska, and Justyna Domienik-Karłowicz. 2023. https://doi.org/10.20944/preprints202309.1100.v1 Beyond human understanding: Benchmarking language models for polish cariology expertise

  34. [42]

    T Wolf. 2019. Huggingface's transformers: State-of-the-art natural language processing. arXiv preprint arXiv:1910.03771

  35. [43]

    R. Yang, T. F. Tan, W. Lu, A. J. Thirunavukarasu, D. S. W. Ting, and N. Liu. 2023. https://doi.org/10.1002/hcs2.61 Large language models in health care: Development, applications, and challenges . Health Care Science, 2(4):255--263

  36. [44]

    T. Zhou, D. Salman, and A. H. McGregor. 2024. https://doi.org/10.1186/s12891-024-07468-0 Recent clinical practice guidelines for the management of low back pain: a global comparison . BMC Musculoskeletal Disorders, 25(1):344

  37. [45]

    URL: " 'urlintro :=

    ENTRY address author booktitle chapter edition editor howpublished institution journal key month note number organization pages publisher school series title type volume year eprint doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before...

  38. [46]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.