{"id":"aaebb73f-f7d9-4638-88fc-06698d9c35f2","arxiv_id":"2507.19885","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"On an authentic Brazilian Portuguese residency exam, top general-purpose AI models scored near the human average on text-only questions but dropped when questions included medical images.","lead":"AI models took a real Brazilian medical residency exam in Portuguese without any training on it. The best ones matched the average human applicant on text questions but struggled more when the questions included images.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The zero-shot framing depends on an unverifiable and date-questionable claim that no evaluated model saw the exam questions; a cutoff-date check is needed before the human-comparison result can be interpreted as zero-shot.","rationale":"The reader identified the zero-shot status as the weakest assumption, and my stress-test agrees: the no-exposure assertion is not independently verifiable and is particularly suspicious for the newer proprietary models that are the headline performers. This concern is more fundamental than the internal numeric inconsistencies, because if contamination occurred, the accuracy comparisons themselves no longer measure what the title promises. However, this is a point that can be partially addressed with documentation and does not necessarily invalidate the underlying observation that some models perform well on Portuguese medical text. The current conditional verdict remains appropriate: the paper should either provide evidence that the models' training cutoffs predate the exam's public release, or revise the claims to acknowledge the unverifiable nature of the zero-shot assumption. I do not recommend rejection because the dataset and scripts are publicly available, enabling this check to be performed, and the directional findings about multimodal difficulty are plausible even under a weaker interpretation.","tokens_in":14091,"tokens_out":12791,"duration_ms":131542,"concrete_test":"Obtain the official Fuvest publication date for the HCFMUSP 2023/2024 exam questions and the documented training-data cutoffs for Claude-3.5-Sonnet, Claude-3-Opus, and GPT-4 Turbo from provider model cards. If any model's cutoff is later than the publication date, the Methods claim is contradicted for that model and the zero-shot interpretation fails. If all cutoffs predate the publication date, the contamination concern is substantially weakened, though not fully eliminated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim, including the title, is specifically about zero-shot performance. Section 2 states that 'None of the models were exposed to the questions of this test during their training phase,' and the Introduction asserts the models were trained before the exam questions were released. For proprietary models such as Claude-3.5-Sonnet and GPT-4 Turbo, this cannot be independently audited, and the timeline is doubtful: the exam is the HCFMUSP 2023/2024 residency test, while Claude-3.5-Sonnet was released in mid-2024 with a training window that may cover the period when the questions became public. If any headline model encountered the questions during training, the accuracy figures would reflect memorization or retrieval rather than zero-shot reasoning, and the comparison with human candidates would no longer support the stated conclusion. The open-weight models (LLaMA-3, Mixtral) also cannot be fully verified without inspecting their training corpora. The paper provides no per-model training-data evidence, model version identifiers, or cutoff dates; it only asserts non-exposure. This is load-bearing because the entire interpretation of the results as a measure of unseen-task generalization rests on it.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript reports a zero-shot evaluation of ten generative AI models (six text-only LLMs and four multimodal Claude models) on the Brazilian Portuguese HCFMUSP medical residency entrance exam. The authors use 117 valid multiple-choice questions, 74 of which are text-only, run five shuffled trials per model, and compare model accuracy with the distribution of human candidate scores. They also conduct a secondary experiment in which three physicians evaluate the explanations produced by Claude-3-Opus for concordance with the chosen answer and potential patient harm, using Gwet's AC1 agreement coefficient. The central claim is that some proprietary models, especially Claude-3.5-Sonnet and Claude-3-Opus, reach accuracy comparable to typical human applicants on text questions, while performance drops on image-based questions, and that this reveals language and modality gaps for non-English medical AI.","tokens_in":14290,"tokens_out":3634,"duration_ms":36677,"significance":"If the core claims hold, this is a useful contribution to multilingual medical LLM evaluation. The study uses a real, high-stakes, publicly disclosed exam rather than a translated or synthetic benchmark, includes both text-only and multimodal conditions under identical prompting, applies repeated-measures statistics over five trials, and provides code, data, and model explanations in a public repository. The expert-evaluation component of Claude-3-Opus explanations is a valuable addition, as is the explicit comparison with the human candidate score distribution. The external and non-circular nature of the benchmark is a strength. However, the central zero-shot interpretation depends on an unverifiable training-data assumption, and several quantitative inconsistencies across the text and tables must be resolved before the main comparisons can be accepted.","major_comments":[{"comment":"The zero-shot framing is load-bearing but rests on the assertion in §2 that 'None of the models were exposed to the questions of this test during their training phase.' For proprietary models like Claude-3.5-Sonnet, GPT-4 Turbo, and Command R+, this cannot be independently audited, and no model version identifiers, API access dates, or training-cutoff dates are provided. The HCFMUSP exam questions were publicly disclosed, and the timeline relative to each model's training window is not established. The authors should either provide concrete cutoff-date and release-date evidence, reframe the study as a closed-book or post-publication assessment, or add a clear limitation paragraph explaining that the zero-shot characterization is conditional on an unverifiable assumption.","section":"§2 Methods and Title/Abstract"},{"comment":"Accuracy values are inconsistent between the narrative and the tables. For example, §3.1 reports Claude-3-sonnet at 72.97% and GPT-4 Turbo at 66.22%, while §3.3 reports 75% and 68.3% for the same models; §3.2 reports Claude-3-haiku at 44.44% and Claude-3-opus at 63.59%, while §3.3 reports 61.1% and 63.7% and refers to a nonexistent 'Claude-3-instant.' Table 1 also lists LLaMA-3-70b at 58.11%, which §3.1 does not mention in its summary of moderate performers. These discrepancies directly affect the headline claims and must be reconciled with a single verified set of numbers across text, tables, and figures.","section":"§3.1, §3.2, §3.3, Tables 1 and 2"},{"comment":"The comparison between Experiment 1 model accuracy and the human candidate distribution is not apples-to-apples. Experiment 1 uses only the 74 text-only questions, but the human candidate scores are for the full 117-question exam, which includes image-based questions. If humans perform systematically better or worse on image questions, the statement that top models are 'comparable to human candidates' in the text-only condition could be misleading. The authors should either restrict the human comparison to Experiment 2, compute a human text-only score if the underlying data allow it, or explicitly discuss the direction and magnitude of this mismatch as a limitation.","section":"§3.3 and Figure 5"},{"comment":"The claim that several models 'achieved accuracy levels comparable to human candidates' is based solely on visual overlap between model point estimates and the smoothed human score density in Figure 5. No statistical test quantifies this comparison, and the human density is estimated from aggregated candidate scores without uncertainty or per-question type information. An equivalence test or a formal interval-based comparison would provide a much stronger basis for the central claim; otherwise, the language should be softened to 'within the observed human score range.'","section":"§3.3 and Statistical Analysis"}],"minor_comments":[{"comment":"The phrase 'Brazilian spoken portuguese' should be 'Brazilian Portuguese'; the manuscript also contains repeated typos such as 'LLhama' instead of 'LLaMA.'","section":"Abstract, §1"},{"comment":"The 'Similarity' column entries are concatenated with model names (for example, 'aLLhama-3-8b' and 'gClaude-3.5-sonnet'), making the tables difficult to read. Use a separate compact-letter column or line breaks.","section":"Tables 1 and 2"},{"comment":"In the provided manuscript, Figures 1-5 are represented only by captions; the actual images are missing. Ensure the figures are included in the submission, as the panels are referenced extensively in the results.","section":"Figures"},{"comment":"Reference [26] is cited as the 'Haus Lin package' but the package is called 'hausekeep'; also, the Wikipedia citations for language statistics could be replaced with more authoritative sources, although this is not central to the findings.","section":"References and Methods"},{"comment":"The sentence 'mean processing time per question increases were observed with the addiction of questions containing images' should read 'addition of questions containing images.'","section":"§3.2"},{"comment":"The phrase 'this in congruent with the fact' should be 'this is incongruent with the fact'; also, clarify whether the exam was administered in 2023 or 2024, since §2 says the exam 'applied in 2024' but earlier refers to the 'HCFMUSP 2023' exam.","section":"§4 Discussion"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a reasonable empirical study, but the zero-shot claim is the central selling point and it is currently supported only by an unverifiable assertion. The numeric inconsistencies between the narrative and the tables make it hard to trust the current version without a careful audit. The author team appears to have access to the raw data and scripts, so correcting these issues is feasible. If the data availability statement is accurate, the reproducibility aspect is commendable. I would not reject, but the revision must address the contamination framing and the human-comparison mismatch."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear [colleague], you should know two things about this paper. First, it's a genuinely useful benchmark: ten models (six text-only, four multimodal) evaluated zero-shot on an authentic Brazilian Portuguese medical residency exam, benchmarked against human candidates, with clinician-rated explanation analysis and all data and scripts on Dataverse. That combination is new. Second, it has a load-bearing credibility problem: the zero-shot claim is unverifiable and possibly wrong, and some of the numbers in the text don't match the tables.\n\nThe design is actually reasonable. Five trials, shuffled order, repeated-measures mixed-effects models with Holm-adjusted post-hocs, and the human performance density comparison is a nice touch. The clinician safety/concordance analysis of Claude-3-Opus explanations is thoughtful, and using Gwet's AC1 for observer agreement is appropriate. The core directional finding—top models land near the human median on text, and all models drop on image questions—is probably robust, and it's the kind of evidence the field needs for non-English medical AI evaluation.\n\nWhere it gets soft: the zero-shot framing. The methods assert 'None of the models were exposed to the questions of this test during their training phase.' For proprietary models like Claude-3.5-Sonnet, released around the same time as this 2024 exam, that is an assumption, not a fact, and the paper gives no model version numbers, no knowledge-cutoff dates, and no leakage checks. If any headline model saw the questions, the comparison with human candidates isn't zero-shot any more. This needs to be fixed before the central interpretation is taken seriously.\n\nSecond, the manuscript has internal numeric inconsistencies. In Section 3.3 the reported best accuracies (Claude-3-sonnet 75%, Claude-3-opus 70.1%, Claude-3-haiku 69.1%, GPT-4 Turbo 68.3%) don't match the tables (72.97%, 70.54%, 61.35%, 66.22%). And it mentions 'Claude-3-instant,' which is not a model in this study. These errors are easy to fix, but they undermine trust in the results until done.\n\nAlso, the abstract highlights Claude-3.5-Sonnet and Opus as matching human levels, which is selective—Claude-3-Sonnet outscored both on text-only. Minor, but it colors the narrative.\n\nBottom line: this paper deserves a serious referee, not a desk reject. The benchmark is useful, the data are open, and the language/modality question is worth asking. But the authors need to substantiate or soften the zero-shot claim, align the reported numbers with the tables, and clean up the model names. If a good reviewer pushes on those, it could come out a solid contribution. I'd bring it to a reading group only after those fixes.","headline":"A useful, well-designed Brazilian Portuguese medical exam benchmark with a load-bearing zero-shot claim that needs verification and numeric inconsistencies that need cleanup.","tokens_in":14849,"tokens_out":4323,"would_cite":true,"duration_ms":37283,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"General-purpose AI models reach applicant-level scores on a Brazilian Portuguese medical residency exam, with image questions as the remaining bottleneck.","keywords":["generative AI","large language models","multimodal large language models","medical residency exam","Brazilian Portuguese","zero-shot evaluation","image interpretation","medical education"],"falsifier":"A contamination check would settle it: if any tested model can reproduce verbatim or near-verbatim text of any of the 117 exam questions when probed, or if its answer pattern matches the official key on items that appear in public training corpora, the zero-shot interpretation fails. The cleanest test is to rerun the same protocol on a newly written, never-published Portuguese residency exam and see whether the accuracy gap to human candidates persists.","tokens_in":13929,"feed_emoji":"🩺","tokens_out":10530,"duration_ms":95805,"temperature":0.7,"pith_summary":"Using 117 real questions from the 2024 HCFMUSP medical residency entrance exam, this study asks whether general-purpose language models can answer high-stakes clinical questions written in Brazilian Portuguese without prior exposure to those items. Six text-only and four multimodal (image-capable) models were prompted once in a standardized zero-shot protocol, and their accuracy was positioned against the actual distribution of human applicant scores. The paper's central finding is that the strongest models, Claude-3.5-Sonnet and Claude-3-Opus, perform within the main range of human candidates, while smaller and open-weights models fall well below. Accuracy drops systematically when questions include radiological or non-radiological images, and a physician review of one model's explanations shows that wrong answers usually come with unsound, sometimes unsafe, reasoning. The result matters because it tests AI on a real, linguistically authentic non-English medical exam rather than a translated benchmark, and it locates the remaining weakness in multimodal understanding rather than in Portuguese itself.","feed_headline":"Top AI models match applicants on a Brazilian medical exam","feed_subtitle":"Best scores land near the median human candidate; image questions remain the main weakness","key_machinery":"The central object is the HCFMUSP residency exam itself: a publicly released, standardized 117-question multiple-choice test written in authentic Brazilian Portuguese, split into 74 text-only and 43 image-based items, with a known distribution of real applicant scores. The evaluation machinery is a zero-shot prompting protocol in which each model receives the same standardized instruction, answers five shuffled trials per question set, and has its answers extracted by regular expressions; accuracy is compared with the human score density, and image questions are further stratified into radiological and non-radiological categories. For explanation quality, one model's answers are reviewed by three physicians under concordance and safety criteria, with agreement quantified by Gwet's AC1 coefficient. This design lets the authors separate language effects from image-understanding effects without building a new benchmark or translating an existing one.","core_discovery":"On the paper's own terms, the discovery is that a zero-shot commercial model can already perform like a typical medical residency applicant in Brazilian Portuguese, and that the main performance gap is modality-specific. On 74 text-only questions, Claude-3-Sonnet reached 72.97% accuracy, Claude-3-Opus 70.54%, Claude-3.5-Sonnet 70.27%, and GPT-4 Turbo 66.22%, placing the best models at or above the 65–70% peak of the human applicant distribution. On all 117 questions including images, Claude-3.5-Sonnet sustained 69.57%, while Claude-3-Opus fell to 63.59%, Claude-3-Sonnet to 54.70%, and Claude-3-Haiku to 44.44%, with radiological images producing the lowest scores. The authors further report that a three-physician review of Claude-3-Opus explanations judged roughly 94% of correct answers as coherently explained, while most incorrect answers were accompanied by flawed rationales that reviewers often deemed unsafe. These results are used to argue that Portuguese itself is not the binding constraint for top models, and that improved multimodal reasoning and language-specific fine-tuning are the needed next steps.","pith_inferences":["The authors leave implicit that a controlled Portuguese-versus-English translation study of the same 117 questions would sharply test the language-disparity reading; the current data support it only indirectly.","Because the exam and candidate scores are released publicly by law, the same protocol could be rerun yearly as a rolling benchmark, with each new exam serving as genuinely unseen data.","The pattern that experts disagreed most about the model's wrong answers suggests some of those items are genuinely ambiguous, so part of the accuracy gap may reflect question quality rather than model deficiency.","A testable prediction of the paper's image-bottleneck reading is that fine-tuning on Portuguese radiology and clinical image-text pairs should substantially raise image-question accuracy; if it does not, the limitation is architectural rather than data-driven."],"forward_implications":["Top-tier general-purpose models can serve as passable first-pass answerers on Portuguese clinical text, at roughly the level of the median residency applicant.","Image-based clinical questions, especially radiology, are the current reliability ceiling; any deployment in specialties where imaging is central should treat model output as draft rather than decision.","Because wrong answers almost always came with flawed explanations, an answer-only accuracy score overstates clinical usefulness; explanation safety needs separate evaluation.","The strongest models' comparable performance in Portuguese, Spanish, and English suggests language alone is not the dominant barrier, so investing in multimodal training data may pay off more than further English-to-Portuguese translation.","Annual public release of such exams creates a repeatable, contamination-aware evaluation loop: each new exam can serve as a fresh held-out test for the next generation of models."],"supporting_citations":[{"why":"Supplies the 117-question HCFMUSP exam and the real applicant score distribution that model accuracy is measured against.","marker":"[21]"},{"why":"Provides the Spanish MIR results used to interpret the Portuguese scores and the sharp accuracy drop on image questions.","marker":"[20]"},{"why":"Grounds the choice of leading models by their prior clinical-knowledge benchmark performance.","marker":"[8]"},{"why":"Provides the USMLE ChatGPT baseline that GPT-4 Turbo's Portuguese score is compared with.","marker":"[22]"},{"why":"Defines the standardized multiple-choice prompting protocol used for zero-shot answer extraction.","marker":"[23]"},{"why":"Supplies the prompt-engineering template for multiple-choice medical questions followed across all models.","marker":"[24]"},{"why":"Another standardized USMLE-style prompting method adopted to keep trials consistent and reproducible.","marker":"[25]"},{"why":"Supplies the Gwet AC1 inter-rater statistic used to measure physician agreement on explanation concordance and safety.","marker":"[30]"},{"why":"Documents that humans can benefit from images in exam questions, sharpening the contrast with the models' image-related accuracy drop.","marker":"[31]"}],"fun_headline_variants":["AI matches Brazilian med applicants on text, stumbles on images","Top AI rivals med exam applicants, images trip them up","AI equals med applicants on text, images stump it"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The comparison is only meaningful as a zero-shot test if none of the tested models encountered these exam questions during training; that claim is asserted but not independently verifiable for proprietary models.","fun_headline_variants_meta":{"raw":{"variants":["AI matches Brazilian med applicants on text, stumbles on images","Top AI rivals med exam applicants, images trip them up","AI equals med applicants on text, images stump it"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001664,"raw_usage":{"total_tokens":6700,"prompt_tokens":1138,"completion_tokens":5562,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":754,"completion_tokens_details":{"reasoning_tokens":5509}},"tokens_in":754,"tokens_out":5562,"duration_ms":36066,"temperature":1.0,"reasoning_tokens":5509,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T17:51:25.931057+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A contamination check would settle it: if any tested model can reproduce verbatim or near-verbatim text of any of the 117 exam questions when probed, or if its answer pattern matches the official key on items that appear in public training corpora, the zero-shot interpretation fails. The cleanest test is to rerun the same protocol on a newly written, never-published Portuguese residency exam and see whether the accuracy gap to human candidates persists.","supporting_citations":[{"cited_title":"Evaluating the Efficacy of ChatGPT in Navigating the Spanish Medical Residency Entrance Examination (MIR): Promising Horizons for AI in Clinical Medicine","cited_arxiv_id":null,"evidence_quote":"Supplies the 117-question HCFMUSP exam and the real applicant score distribution that model accuracy is measured against."},{"cited_title":"Residˆ encia M´ edica 2024 – FUVEST divulga quest˜ oes de prova objetiva de concurso para residˆ encia m´ edica da FMUSP – Fuvest; 2024","cited_arxiv_id":null,"evidence_quote":"Provides the USMLE ChatGPT baseline that GPT-4 Turbo's Portuguese score is compared with."},{"cited_title":"ChatGPT-4 Performance on USMLE Step 1 Style Questions and Its Implications for Medical Education: A Comparative Study Across Systems and Disciplines","cited_arxiv_id":null,"evidence_quote":"Another standardized USMLE-style prompting method adopted to keep trials consistent and reproducible."},{"cited_title":"Better to be in agreement than in bad company: A critical analysis of many kappa-like tests","cited_arxiv_id":null,"evidence_quote":"Supplies the Gwet AC1 inter-rater statistic used to measure physician agreement on explanation concordance and safety."}],"review_version":1}