Pith. sign in

REVIEW 3 major objections 6 minor 1 cited by

PersianMedQA: Evaluating Large Language Models on a Persian-English Bilingual Medical Question Answering Benchmark

T0 review · 3 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read PersianMedQA, a 20,785-question expert-validated exam benchmark, shows GPT-4.1 at 83.09% in Persian, Persian-tuned models near chance, and 3–10% of items answerable only in Persian because translation erases local context.

desk verdict A genuinely useful Persian medical QA benchmark, but the contamination check is too thin to trust the headline accuracies yet. read the letter →

arxiv 2506.00250 v4 pith:HIMRQZA2 submitted 2025-05-30 cs.CL cs.ITmath.IT

classification cs.CLcs.ITmath.IT
keywords Persianmedicalquestionansweringbilingualbenchmarklargelanguagemodelevaluationlow-resourceNLPexaminationdatasettranslationlossculturalcontextdatacontamination
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper introduces PersianMedQA, a 20,785-question multiple-choice benchmark built from 14 years of Iranian national medical residency exams, with expert-validated answers, 23 specialties, and English translations checked by a board-certified internist. Its purpose is to measure whether large language models can handle real medical questions in Persian, and to separate language ability from medical knowledge. The authors report that closed-weight general models lead, with GPT-4.1 reaching 83.09% in Persian and 80.71% in English, while Persian-specialized models perform near or below chance. They also report that translating questions into English hurts performance on 3–10% of items whose correct answers depend on Iranian protocols and cultural cues, so English-only evaluation misses part of the clinical signal. If the benchmark holds, it gives the field a way to test culturally grounded medical reasoning rather than English-centric translate-then-test shortcuts.

What carries the argument

The central object is the PersianMedQA test set itself: 5,236 stratified multiple-choice items drawn from 20,785 expert-validated questions, each paired with an official answer key, a medical-specialty label, a clinical or non-clinical label, and a Gemini-2.5-Flash English translation validated by a board-certified internist. The benchmark's analytic machinery is the bilingual comparison—same model, same prompt, Persian text versus English translation—together with an answer-only partial-input protocol that isolates guessing from reasoning. These comparisons let the authors attribute performance differences to language, domain adaptation, and answer-choice artifacts rather than to dataset noise.

What would settle it

Collect a fresh set of Persian medical exam questions written after the evaluated models' training cutoffs, run the paper's zero-shot protocol, and compare the new accuracy with the reported 83.09%; a large drop would indicate that the benchmark numbers depend on memorized items.

Watch

Extended reading notes

Core claim

The paper's central claim is that PersianMedQA is a valid instrument for evaluating LLMs on Persian and English medical questions, and that its measured results characterize current models. On the 5,236-question test set, GPT-4.1 achieves 83.09% in Persian and 80.71% in English; the best open-weight model, DeepSeek-Chat-V3, scores 68.05% in Persian and 73.30% in English, while Persian-only models drop to the 24–35% range, often because of instruction-following failures. The benchmark also documents a translation asymmetry: 3–10% of items are answered correctly only in Persian, with manual analysis attributing this to regional vaccination schedules, antibiotic protocols, and Persian medical terms that machine translation alters or loses.

Load-bearing premise

The load-bearing premise is that the PersianMedQA questions have not appeared in the training data of the evaluated models; if some have, the reported accuracies and rankings would be inflated by memorization.

Editorial extensions

If this is right

  • Closed-weight general-purpose models, not Persian-tuned or medical-tuned models, currently give the best Persian medical question answering, so practical deployments for Persian speakers should start with general-purpose models.
  • Because a consistent 3–10% of items are answered correctly only in the original Persian, translate-then-test evaluation understates models' local medical competence and hides culturally specific failure modes.
  • English is not a ceiling for non-English medical QA: several top models score higher on the Persian text than on its English translation, and that gap marks the clinical context that translation removes.
  • Model scale alone does not fix domain or language gaps: several larger specialized models score below smaller general-purpose models on the same questions.
  • Answer-only evaluation in fields like medical ethics shows that multiple-choice accuracy can be driven by answer-choice artifacts, so benchmark scores should not be read as proof of genuine clinical or ethical reasoning.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The authors frame their reported scores as conservative lower bounds because API cost, rate limits, and licensing constrained the evaluation and precluded fine-tuning; an open-weight fine-tuning study on the released training split would test whether open models can close the gap to GPT-4.1.
  • The contamination check reported in Section 3.4 examines exact matches on a random sample of under 50 questions; a full-corpus near-duplicate audit against training corpora would make the benchmark's validity claim stronger.
  • The bilingual comparison design transfers to other low-resource languages: publish native and translated test sets, report per-item language asymmetry, and annotate which items depend on regional protocols.
  • The answer-choice artifact result suggests adding a partial-input protocol to future medical multiple-choice benchmarks so that memorized patterns can be separated from domain reasoning.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper introduces PersianMedQA, a bilingual (Persian-English) multiple-choice medical QA benchmark of 20,785 questions collected from 14 years of Iranian national residency and pre-residency exams, spanning 23 specialties. The authors evaluate 41 LLMs in zero-shot and chain-of-thought settings, reporting that closed-weight models such as GPT-4.1 achieve 83.09% (Persian) and 80.71% (English), while Persian-tuned models underperform substantially. The paper further claims that 3-10% of questions are correctly answerable only in Persian because machine translation loses culturally and clinically specific context, and it includes an answer-only control experiment showing that models can exploit answer-choice artifacts. The dataset and a bilingual dictionary are released publicly.

Significance. If the dataset is clean and the evaluation is robust, PersianMedQA would fill a clear gap in multilingual, low-resource medical QA evaluation. The scale (20,785 items), the use of official national exam questions with institutional answer verification, the inclusion of a bilingual medical dictionary, the breadth of model coverage (41 models), and the answer-only control experiment are all strengths. The paper also makes a useful empirical point about translation-induced loss of culturally grounded clinical context. However, the credibility of the headline accuracy figures and model rankings rests on the contamination check in §3.4, which is currently too weak to support the stated 'no leakage' conclusion; the lack of statistical significance testing further tempers the ranking claims. These issues are fixable, and the underlying resource remains valuable.

major comments (3)
  1. [§3.4] The data contamination analysis is not sufficient to support the claim that PersianMedQA is free of training-data leakage. The exact-match search was run on a 'randomly sampled subset of under 50 questions' out of 20,785 total items; a sample of this size cannot reliably detect contamination at rates that would affect the reported accuracies. For example, if 2% of the 5,236 test questions were memorized in training (≈105 items), a 50-item sample has only about a 64% chance of containing at least one such item, and exact-match search would miss paraphrased or translated variants even if one were sampled. The temporal analysis in the same section is not decisive for GPT-4.1 and Gemini-2.5-Flash, whose training cutoffs are not documented, and it cannot distinguish batch-scraped data from legitimate post-cutoff exposure. Moreover, the Ethics statement (Section 7) states that the exam papers have been 'publicly released after each examination cycle and have been widely circulated... for over a decade,' which directly weakens the 'Secure Sourcing' paragraph in §3.4. I recommend a full-dataset contamination audit (e.g., n-gram overlap against public corpora, exact and near-match search on all test items), a power analysis for any sampled check, or validation on a held-out exam that post-dates model training cutoffs; at minimum, the sample size and detection limits must be reported, and the 'no leakage' conclusion should be reworded as conditional.
  2. [§4, Table 4] The paper reports accuracy differences between models without any significance testing, yet the central claims include model rankings and the existence of a 3-10% subset of questions answerable only in Persian. Points such as GPT-4.1 (83.09%) versus Gemini-2.5-Flash-Preview (82.37%) on the Persian test set, or the +0.033 ensemble gain over the best open-weight model, could easily fall within sampling variability. Because the same 5,236 questions are used for all models, paired comparisons are appropriate; the authors should report confidence intervals or apply McNemar's test (or an equivalent per-question paired test) to the differences that drive the abstract and conclusion. Without this, the ranking statements and the '3-10% only-in-Persian' percentages should be treated as preliminary rather than established.
  3. [§3.5, §9.2] The English evaluation set and the translation-impact analysis rest on expert validation by a single board-certified medical specialist, and the paper acknowledges in §9.2 that no inter-annotator agreement could be computed. This is load-bearing for the English accuracy figures and for the claim that translation produces a loss of culturally specific clinical context. A single expert's judgment, however well qualified, cannot by itself establish that the translations preserve medical meaning or that the 3-10% Persian-only items are genuinely due to cultural loss rather than translation errors or annotator subjectivity. I recommend a second independent medical reviewer on a stratified sample of, say, 200-300 translated questions, with agreement rates reported, or at least a detailed, pre-specified rubric for what counts as 'preservation of clinical meaning' and 'loss of cultural context.'
minor comments (6)
  1. [Throughout] There are many typographical and formatting inconsistencies, including 'official' and 'sufficient' (ligature artifacts), inconsistent use of 'Sanjeshp' versus 'Sanjesh', and inconsistent rendering of 'CoT' as 'Co T'. A careful proofreading pass is needed.
  2. [Table 4] The column header 'Fa (%)' should be expanded or defined in the table caption; 'Fa' is not obviously Persian to all readers, and other parts of the paper use 'Persian' and 'English'.
  3. [§3.4, §7] The 'Secure Sourcing' claim in §3.4 and the Ethics statement's statement that the exam papers are 'widely circulated' and available in commercial study materials appear contradictory. Please clarify whether the dataset is scrape-resistant and report any copyright or terms-of-use considerations for the released dataset.
  4. [§4.1] The statement that results on the 5,236-question test set are 'consistent' with full-dataset runs is not quantified. Please provide the correlation or maximum deviation observed for the models used in that validation.
  5. [Figure 5] The interpretation of the 2020-2021 performance dip as 'increased exam difficulty during the COVID-19 pandemic' is speculative. The paper should either provide supporting evidence (e.g., pass-rate statistics or expert testimony) or label this as a hypothesis.
  6. [Appendix references] The text refers to 'Appendix 15' and 'Appendix 16', but the appendices are numbered inconsistently in the source; please ensure all appendix cross-references match the final numbering.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the benchmark is built from external official exams with fixed answer keys, and the reported accuracies are direct measurements rather than derived predictions.

full rationale

PersianMedQA is constructed from externally sourced Iranian national residency examination questions with official answer keys, then cleaned, annotated, and split. The central claims are the reported model accuracies on the fixed 5,236-question test set, and these are direct measurements against those fixed answers. There is no fitted parameter that is later renamed as a prediction: the few-shot experiments use training-split examples only as retrieved context and report no gains, and the translation-impact analysis is an empirical comparison of model correctness across Persian and English versions of the same items. The answer-only control in Section 4.5 is an independent falsification test rather than a circular justification. The paper contains no load-bearing self-citation chain, no uniqueness theorem imported from the authors' prior work, and no ansatz smuggled in via citation. The weak contamination check in Section 3.4 (a sample of under 50 questions out of 20,785), the single-expert translation validation, and the Ethics statement's acknowledgment that exam papers have been publicly circulated for over a decade are genuine validity threats, but they concern data contamination and inter-annotator reliability, not circular reasoning. Even if contamination were present, the accuracy numbers would be inflated measurements, not conclusions that reduce by construction to the paper's own inputs. Therefore no circular step is identifiable, and the appropriate circularity score is 0.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

No parameters were fitted; the benchmark's credibility rests on institutional answer verification, freedom from training-data contamination, and translation fidelity. These are assumed rather than established with strong independent evidence, which is why the axiom ledger lists four domain assumptions and the correctness risk is medium.

assumptions (4)
  • domain assumption Official Iranian medical exam answer keys are correct for all 20,785 questions.
    The benchmark ground truth relies on Sanjesh's institutional three-level verification and the authors' manual filtering, but no independent re-grading of the full dataset is reported.
  • domain assumption The dataset has not appeared in LLM training corpora.
    Leakage would inflate reported accuracies; the paper's evidence is an exact-match search on fewer than 50 questions and a temporal consistency argument, which are weak supports for the full 20,785-question set.
  • domain assumption Gemini-2.5-Flash English translations preserve clinical meaning and cultural context.
    The English evaluation and the translation-impact analysis depend on translation fidelity; validation was performed by a single board-certified specialist without a quantitative agreement measure.
  • domain assumption Accuracy on this multiple-choice test set is a valid measure of medical reasoning ability.
    The paper interprets accuracy as medical capability, though its own answer-only experiment shows that models can exploit answer-choice patterns, so the benchmark partially reflects artifact recognition rather than pure reasoning.

how reviews work

0 comments
Cite this review

Pith. "Pith review of PersianMedQA: Evaluating Large Language Models on a Persian-English Bilingual Medical Question Answering Benchmark." pith.science (2026). https://pith.science/paper/HIMRQZA2

@misc{pith2026250600250,
  author       = {Pith},
  title        = {Pith review of: PersianMedQA: Evaluating Large Language Models on a Persian-English Bilingual Medical Question Answering Benchmark},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/HIMRQZA2}},
  note         = {Machine review of arXiv:2506.00250}
}
read the original abstract

Large Language Models (LLMs) have achieved remarkable performance on a wide range of Natural Language Processing (NLP) benchmarks, often surpassing human-level accuracy. However, their reliability in high-stakes domains such as medicine, particularly in low-resource languages, remains underexplored. In this work, we introduce PersianMedQA, a large-scale dataset of 20,785 expert-validated multiple-choice Persian medical questions from 14 years of Iranian national medical exams, spanning 23 medical specialties and designed to evaluate LLMs in both Persian and English. We benchmark 41 state-of-the-art models, including general-purpose, Persian, and medical LLMs, in zero-shot and chain-of-thought (CoT) settings. Our results show that closed-weight general models (e.g., GPT-4.1) consistently outperform all other categories, achieving 83.09% accuracy in Persian and 80.7% in English, while Persian LLMs such as Dorna underperform significantly (e.g., 34.9% in Persian), often struggling with both instruction-following and domain reasoning. We also analyze the impact of translation, showing that while English performance is generally higher, 3-10% of questions can only be answered correctly in Persian due to cultural and clinical contextual cues that are lost in translation. Finally, we demonstrate that model size alone is insufficient for robust performance without strong domain or language adaptation. PersianMedQA provides a foundation for evaluating bilingual and culturally grounded medical reasoning in LLMs. The dataset, along with a bilingual medical dictionary, is available: https://huggingface.co/datasets/MohammadJRanjbar/PersianMedQA .

Figures

Figures reproduced from arXiv: 2506.00250 by the authors.

Figure 1
Figure 1. A translated medical question example from the dataset. ply translating questions is inadequate, as such pipelines can strip away critical terminology, subtle cultural cues, and localized standards of care, po￾tentially leading to life-threatening consequences in clinical practice (Mehandru et al., 2022). Medical practice is inherently shaped by con￾arXiv:2506.00250v4 [cs.CL] 26 May 2026 [PITH_FULL_IMAGE:figures/fu… view at source ↗
Figure 2
Figure 2. Overview of the PersianMedQA dataset construction process, including data collection, clean￾ing, annotation, and partitioning steps. textual factors, including sociocultural, socioeco￾nomic, regional, and healthcare system variables that extend beyond language translation (Klein￾man, 1978; Betancourt et al., 2003). Clinical decision-making protocols and symptom interpre￾tation vary significantly across healthcare sy… view at source ↗
Figure 3
Figure 3. Distribution of medical fields in the dataset. advances include DORNA (Team, 2024), a large￾scale Persian language model. In the medical do￾main, SINA-BERT (Taghizadeh et al., 2021) repre￾sents an early attempt at Persian medical NLP, uti￾lizing pre-training on large-scale medical corpora including both formal and informal medical texts from diverse online resources. Furthermore, ex￾isting Persian medical NLP effort… view at source ↗
Figures from the paper (10 more)
Figure 4
Figure 4. Figure 4: Overall accuracy of models on Persian and English test sets. 3.4. Data Contamination and Evaluation Integrity To ensure the reliability of our medical evaluation, we implemented multiple safeguards against data contamination and memorization artifacts: Secure Sourcing:…
Figure 5
Figure 5. Figure 5: LLM performance across exam years (2011-2024) [PITH_FULL_IMAGE:figures/full_fig_p005_5.png]
Figure 6
Figure 6. Figure 6: The zero-shot prompt used for evalua￾tion [PITH_FULL_IMAGE:figures/full_fig_p006_6.png]
Figure 7
Figure 7. Figure 7: Heatmap showing the accuracy of each model across all medical specialties in the Persian￾MedQA dataset. Each cell represents the accuracy for a particular model-field pair. The full model list is available in Appendix 15. cal practice. To quantify this effect, we trans…
Figure 8
Figure 8. Figure 8: Example demonstrating regional vaccination protocol differences affecting tetanus post￾exposure prophylaxis decisions. Few-shot learning: For every test question, we drew the in-context examples exclusively from the PersianMedQA training split (up to k = 5 per query). …
Figure 9
Figure 9. Figure 9: Telegram interface for expert subject classification of ambiguous questions. 13.2. CoT Reasoning Interface To analyze the reasoning behind model outputs, we designed an interface that presented the ex￾pert with a curated 200-question subset. For each question, the expe…
Figure 10
Figure 10. Figure 10: Telegram interface for expert annota￾tion of reasoning categories and explanations. 14. Persian Medical Dictionary [PITH_FULL_IMAGE:figures/full_fig_p015_10.png]
Figure 11
Figure 11. Figure 11: Category 1: correct in both languages. Category 2: Correct Only After Translation These questions benefit from the model’s stronger English medical training, particularly in specialized terminology. Category 2: Specialized Anatomical Pathology English: Which finger’s …
Figure 12
Figure 12. Figure 12: Category 2: correct only after translation. Category 3: Correct Only in Persian These questions involve Iran-specific medical practices or clinical contexts that are altered or lost during translation. Category 3: Iran-Specific Clinical Protocols Example A: Regional A…
Figure 13
Figure 13. Figure 13: Category 3: correct only in Persian due to Iran-specific protocols [PITH_FULL_IMAGE:figures/full_fig_p016_13.png]

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Biomedical Machine Translation for Low-Resource Arabic-Script Languages via Cross-Lingual Transfer and LoRA Adapter Merging

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Zero-data merging of Arabic and Persian biomedical LoRA adapters comes within 1.4–3.5 CHrF++ points of supervised adaptation for Dari; Pashto and Sorani Kurdish stay below usable quality.

Reference graph

Works this paper leans on

24 extracted references · 24 canonical work pages · cited by 1 Pith paper

  1. [1]

    Expert Evaluation: The model follows a Western protocol; how- ever, local clinical practice requires urgent biopsy due to high mor- tality risk

    Contextual Mismatch Question: What is the next step in an immunocompromised pa- tient with nasal congestion and suspected invasive fungal sinusi- tis? Correct Answer: Endoscopy and biopsy Model Answer: Imaging (MRI) is needed before biopsy. Expert Evaluation: The model follows a Western protocol; how- ever, local clinical practice requires urgent biopsy d...

  2. [2]

    Expert Evaluation: The model selected a technically true but con- textually incorrect answer; expert notes ambiguity in phrasing and clinical intent

    Ambiguity in Options Question: What is the most common malignant neoplasm of the liver? Correct Answer: Hepatocellular carcinoma (HCC) Model Answer: Metastasis is more common overall, so we choose that. Expert Evaluation: The model selected a technically true but con- textually incorrect answer; expert notes ambiguity in phrasing and clinical intent

  3. [3]

    Expert Evaluation: The patient’s immunosuppression requires a different clinical approach, which the model failed to identify

    Reasoning Failure Question: What is the correct order of action in a 25-year-old with lymphoma and meningitis signs but no neurologic deficits? Correct Answer: Blood culture → Lumbar puncture → Empiric antibiotics Model Answer: CT scan should be done first due to immunosup- pression. Expert Evaluation: The patient’s immunosuppression requires a different ...

  4. [4]

    ACM T ransactions on Computing for Healthcare, 3(1):1–23

    Domain-specific language model pretrain- ing for biomedical natural language processing . ACM T ransactions on Computing for Healthcare, 3(1):1–23. Niclas Hertzberg and Anna Lokrantz. 2024. MedQA-SWE - a clinical question & answer dataset for Swedish . In Proceedings of the 2024 Joint International Conference on Com- putational Linguistics, Language Resou...

  5. [5]

    Explain why each incorrect option is wrong, and the chosen one is correct

  6. [6]

    CoT": "your step-by-step reasoning

    Explicitly state which option (1, 2, 3, or 4) is your final answer. Response format (JSON): { "CoT": "your step-by-step reasoning", "Final_Answer": 1 | 2 | 3 | 4, "Reasoning": "concise justification" } Be methodical, precise, and thorough in your analysis. Your ex- pertise as {english_specialty} is critical for answering these specialized questions correctly

  7. [8]

    Medical Specialist Background The medical specialist involved in this study is a board-certified internal medicine physician. She graduated from the University of Tehran with a de- gree in general medicine and completed her spe- cialty training in internal medicine at Shahid Be- heshti University of Medical Sciences. She has 5 years of clinical practice a...

  8. [9]

    Data Verification and Quality Assurance 9.1. Answer Verification Process As described in Section 3, all questions under- went a rigorous three-level verification process by the National Center for Medical Education Assess- ment (Sanjesh): (1) Initial expert committee re- view by board-certified medical professionals, (2) Public comment period where medica...

Show all 24 references
  1. [10]

    For each example, we highlight the clinical context, the correct answer, the model’s response, and a summary of the ex- pert’s evaluation

    Examples of CoT Error Patterns This section presents representative error patterns identified in model-generated Co T outputs, as an- notated by our clinical expert. For each example, we highlight the clinical context, the correct answer, the model’s response, and a summary of...

  2. [11]

    Few-shot Prompt You are a medical expert tasked with answering multiple-choice medical questions

    Few-shot Evaluation Prompt In-context examples are drawn from the Persian- MedQA training split using LaBSE cosine similar- ity, TF-IDF , and random selection (up to k = 5). Few-shot Prompt You are a medical expert tasked with answering multiple-choice medical questions. In-co...

  3. [12]

    Expert Evaluation: Model lacks pharmacologic mechanism knowledge and defaults to common treatments

    Knowledge Gap Question: Which drug works via motilin receptor stimulation for gastroparesis? Correct Answer: Erythromycin Model Answer: Metoclopramide is commonly used for gastro- paresis, so we chose that. Expert Evaluation: Model lacks pharmacologic mechanism knowledge and d...

  4. [13]

    User Interfaces To facilitate expert interaction throughout vari- ous phases of our study, we developed multiple Telegram-based interfaces to streamline collabora- tion with our medical specialist. 13.1. Subject Annotation Interface We created a Telegram annotation bot to supp...

  5. [14]

    For each question, please:

    CoT Reasoning Prompt CoT Prompt You are a medical expert taking a medical board examination. For each question, please:

  6. [15]

    Read and understand the question carefully

  7. [16]

    Analyze the options (1–4) systematically

  8. [17]

    Apply your medical knowledge step by step

  9. [18]

    Show your chain-of-thought (Co T) reasoning clearly

  10. [22]

    T able 3: Distribution of extracted Persian medical terms

    Persian Medical Dictionary T able 3 summarizes the number of unique med- ical terms extracted per category in the bilingual Persian medical dictionary released alongside the dataset. T able 3: Distribution of extracted Persian medical terms. Category Unique Terms Medical Devic...

  11. [23]

    Models are sorted by average performance

    Overall Performance Comparison T able4 shows the zero-shot accuracy of all 41 eval- uated models on the original Persian questions, the translated English questions, and their aver- age. Models are sorted by average performance. The five lowest-performing models struggled sig-...

  12. [24]

    Repre- sentative examples for each category are provided below

    Cross-Linguistic Performance Analysis Our cross-linguistic analysis revealed three dis- tinct performance patterns across models. Repre- sentative examples for each category are provided below. Category 1: Correct in Both Languages These questions involve standardized clinical...

  13. [2020]

    Josepha Campinha-Bacote

    Language models are few-shot learners . Josepha Campinha-Bacote. 2002. The process of cultural competence in the delivery of healthcare services: a model of care . Journal of T ranscul- tural Nursing, 13(3):181–201. Yu Guan Cao, Fei Liu, Pippa Simpson, Lamont Antieau, Andrew B...

  14. [2021]

    Neural Pro- cessing Letters, 53(6):3831–3847

    Parsbert: Transformer-based model for persian language understanding . Neural Pro- cessing Letters, 53(6):3831–3847. Fangxiaoyu Feng, Yinfei Yang, Daniel Cer, Naveen Arivazhagan, and Wei Wang. 2020. Language- agnostic bert sentence embedding . Yu Gu, Robert Tinn, Hao Cheng, Mi...

  15. [2023]

    Laurence J

    Better to ask in english: Cross-lingual evaluation of large language models for health- care queries. Laurence J. Kirmayer. 2001. Cultural variations in the clinical presentation of depression and anx- iety: implications for diagnosis and treatment. The Journal of Clinical Psy...

  16. [2024]

    Artificial Intelligence in Medicine , 155:102938

    Medexpqa: Multilingual benchmarking of large language models for medical question answering. Artificial Intelligence in Medicine , 155:102938. Sofia J. Athenikos and Hyoil Han. 2010. Biomed- ical question answering: A survey . Computer Methods and Programs in Biomedicine, 99(1...

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.