{"id":"14fd520a-5f55-4f9a-bbe6-ea4b68e23f5a","arxiv_id":"2506.21578","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"On a new Portuguese healthcare benchmark, leading language models score high overall but drop sharply in several specialties, such as neurosurgery and social work.","lead":"Researchers compiled 5,632 questions from Brazil's national medical and allied health exams into a new benchmark called HealthQA-BR and tested 21 large language models on it. Top models average about 86% accuracy overall, but scores fall to about 60% in neurosurgery and 68% in social work, suggesting aggregate scores hide important gaps.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Specialty-level accuracy claims and the systemic claim rest on tiny question counts and an unaudited metadata basis; a one-revision check on counts and per-model spread would decide whether the central narrative survives.","rationale":"The reader's weakest assumption identifies the same core issue — small sample sizes and unreliable metadata undermine the quantitative specialty-level percentages and hence the spikiness narrative. I agree with that diagnosis. However, the reader places the center of gravity on the quantitative precision issue (no confidence intervals, small n, repeated 68.4%). I think the additional load-bearing vulnerability is the paper's universal 'systemic across all models' claim, which is asserted with no supporting per-model specialty table in the submission; the paper only shows GPT 4.1 figures and gives qualitative assurance about DeepSeek R1 and GPT-4o. In a stress-test on the central claim, the universal claim is the strongest sentence in the abstract and conclusion, and it is the least supported by presented evidence. Therefore I partially agree with the reader: the quantitative uncertainty is serious, but the universal claim's missing evidence is equally load-bearing and is distinct from a confidence-interval fix. Those two issues are independently addressable: adding confidence intervals fixes the small-n precision problem; publishing a per-model specialty matrix tests the universal claim. Both are required by the CONDITIONAL verdict. I also note that the paper's own Section 2.2 audit of 82 questions (1.5%) is an internal limitation statement that the stress-test rules require flagging; the authors themselves restrict the metadata validation to that sample, so the burden is on them to show the specialty denominators are correct. The paper deserves credit for the novelty of the system-wide benchmark and the overall model ranking plausibility, but the article cannot currently support the quantified gap magnitudes or the universality claim without the additional evidence and uncertainty quantification. UNCHANGED would keep the reader's CONDITIONAL; I agree with CONDITIONAL rather than ACCEPT because the requested revisions are material to the central claim, not cosmetic. I do not see grounds for REJECT: the dataset construction description is detailed, the source exams are public, the overall accuracy hierarchy is internally consistent, and all issues are testable by release of artifacts and recomputation.","tokens_in":9108,"tokens_out":2680,"duration_ms":25102,"concrete_test":"Run one focused re-derivation: request the authors to release the per-model specialty-level accuracy matrix for all 21+ models together with per-specialty counts, and recompute the key numbers from the raw answer logs. Specifically (a) verify the 68.4% values for Orthopedics and Social Work are not a copy error by recomputing from per-question predictions; (b) compute 95% Wilson confidence intervals for Neurosurgery (n=25), Dermatology (n=14), Anesthesiology (n=10), and Rheumatology (n=8) and check whether the claimed 60.0% gap versus Ophthalmology 98.7% survives; (c) for at least DeepSeek R1 and GPT-4o, produce the specialty-level accuracy and a dispersion measure (e.g., min-max or coefficient of variation) to test whether the spiky pattern is systemic rather than GPT 4.1-specific; (d) cross-check the specialty tags on the full set of small categories (all categories with n<50)…","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim is that a single accuracy score is dangerously misleading because accuracy is highly uneven across specialties, e.g., GPT 4.1 drops to 60.0% in Neurosurgery and 68.4% in Social Work, and this spikiness is systemic across all models. The load-bearing empirical support is the specialty-level accuracy table (Table 2/Fig 2) and the allied-health table (Table 3/Fig 3). But Table 2 lists only 25 Neurosurgery and 14 Dermatology questions, and Table 3 lists 75 Psychology and 79 Social Work questions; the paper reports two-decimal percentages without confidence intervals. For n=25, a single answer change shifts accuracy by 4%; for n=14, by ~7%. The 60.0% Neurosurgery figure is ±~20% at 95% CI and is statistically indistinguishable from many other specialties. The 68.4% Orthopedics (n=79) and 68.4% Social Work (n=79) identical values are suspicious of a copy/paste error. Second, the paper's universal claim ('spiky profile is systemic across all models') is asserted in Section 4.2 but only GPT 4.1 is shown in detail. The claim that other models also show variance is stated qualitatively with no table, no variance measure, and no per-model specialty breakdown; Table 4 only reports overall accuracy and the three exam-level sub-scores. Third, the metadata tags that define the specialty counts are the infrastructure for all specialty claims, yet the QA audit validated metadata on only 82 of 5,632 questions, and no inter-annotator agreement or error-rate breakdown is reported; a small tagging error rate in small categories directly changes the denominator and the headline accuracy. Finally, no confidence intervals for the overall scores, no calibration of the reported values, and no data/code link are provided, so the exact numerical claims (e.g., 60.0%, 68.4%) cannot be independently verified.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces HealthQA-BR, a benchmark of 5,632 multiple-choice questions drawn from Brazilian national licensing and residency exams (Revalida and Enare), covering medicine, nursing, dentistry, psychology, social work, and other allied health fields. The authors report zero-shot accuracy for over 20 LLMs, finding that top models such as GPT 4.1 reach 86.6% overall accuracy but show strong variability across specialties, with scores as low as 60.0% in Neurosurgery and 68.4% in Social Work. They argue that aggregate scores create a dangerous illusion of competence and that specialty- and profession-level reporting is necessary for safe deployment of medical AI.","tokens_in":9333,"tokens_out":2436,"duration_ms":29024,"significance":"If its specialty-level claims are supported, HealthQA-BR addresses a genuine gap in LLM evaluation by providing a Portuguese-language, interprofessional healthcare benchmark. The dataset construction from official high-stakes exams, the effort to deduplicate questions, and the public release of the benchmark and evaluation suite are concrete contributions. The aggregate model rankings in Table 4 are based on large question counts and appear internally consistent, so the overall performance hierarchy is likely robust. However, the central narrative of a systematically 'spiky' knowledge profile depends on specialty-level accuracy estimates that are currently not statistically supported, and the claimed universality across models is not directly demonstrated.","major_comments":[{"comment":"The specialty-level accuracy estimates for GPT 4.1 are presented to one decimal place without confidence intervals, despite very small category sizes: Neurosurgery has 25 questions, Dermatology 14, Anesthesiology 10, and Rheumatology 8. For n=25, a one-answer change moves accuracy by 4 percentage points; for n=8, by 12.5 points. The headline figure of 60.0% in Neurosurgery therefore has a wide confidence interval and is statistically indistinguishable from many other specialties. The paper should report exact hit counts and confidence intervals (or credible intervals) for every specialty, and should either aggregate low-sample categories further or explicitly mark them as unreliable.","section":"Table 2 and Figure 2"},{"comment":"The claim that the 'spiky' profile is 'a systemic issue observed across all models' is asserted but not supported by any per-model specialty data. Only GPT 4.1 is analyzed in detail; Table 4 reports only overall and exam-level scores, not specialty breakdowns. To support the universal claim, the authors should provide specialty-level accuracy for all evaluated models, or at least a variance measure per model, and demonstrate that the unevenness is not driven by small-sample noise. Without this, the paper's central generalization remains an unsupported extrapolation from a single model.","section":"Section 4.2 and Abstract"},{"comment":"The metadata tags are the infrastructure for every specialty-level result, yet the final quality assurance audit validated only 82 of 5,632 questions (~1.5%), and no per-tag error rate or inter-annotator agreement is reported. Because a single mislabeled question in a small specialty can change the reported accuracy by several percentage points, the authors should either expand the audit substantially, provide a detailed breakdown of audit results by specialty, or run the analysis under realistic assumptions about metadata error rates to show the conclusions are unchanged.","section":"Section 2.2, item 6"}],"minor_comments":[{"comment":"The identical reported value of 68.4% for Orthopedics and Traumatology (n=79) and Social Work (n=79) is suspicious; please verify the underlying counts and report the raw hit counts (e.g., 54/79) so readers can check the arithmetic.","section":"Table 2 vs Table 3"},{"comment":"The 'Most Common Answer' baseline is stated as 21.89% for a five-option MCQ. This value needs a derivation or citation, since random uniform guessing would be 20%.","section":"Section 3.3"},{"comment":"The figures would be more informative if they included sample sizes or error bars, and the captions should state that the percentages are point estimates without confidence intervals.","section":"Figures 2 and 3"},{"comment":"The statement that findings are 'strongly corroborated' by AfriMed-QA goes beyond what is demonstrated in the paper; the comparison is qualitative and no systematic analysis of variability is provided. Please soften or support this claim.","section":"Section 5"},{"comment":"The abstract says 'over 20' models while Section 3.1 says 'over 21'; please make the count consistent throughout, and list the exact number of models in the main text.","section":"Abstract and Section 3.1"}],"recommendation":"major_revision","confidential_remarks":"The benchmark itself and the aggregate comparison are valuable, and the paper is within scope for a computational linguistics or medical AI venue. The main concern is that the paper's headline claims about specialty-level deficiencies and their universality are not yet statistically supported. These issues are fixable with additional analysis and reporting, which is why I recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"HealthQA-BR is a real contribution: the first large Portuguese-language benchmark covering not just medicine but nursing, psychology, social work and other allied professions, built from 5,632 official licensing and residency questions. That fills a genuine gap, and the aggregate evaluation of 21 models is useful. The overall rankings look plausible, and the authors deserve credit for framing the work around interprofessional care rather than another physician-only MCQ set.\n\nThe central warning is also right: an 86.6% overall score can hide 60% in Neurosurgery. The qualitative point that knowledge is uneven is well taken. But the quantitative case is thinner than the paper suggests. Several specialty categories are tiny: 25 Neurosurgery, 14 Dermatology, 10 Anesthesiology, 8 Rheumatology. A one-answer change moves those percentages by 4 to 12 points, and the paper gives no confidence intervals. So the specific gap between 98.7% and 60.0% is real in direction but not in magnitude.\n\nThe bigger problem is the \"systemic across all models\" claim in Section 4.2. The paper only shows GPT 4.1's specialty breakdown in detail. There is no per-model table or variance measure for the other 20 models, so the universal claim is an assertion, not a result. The metadata audit on 82 of 5,632 questions is also too thin to support fine-grained categories; a few tagging errors in a category of 8 or 14 questions directly changes the headline numbers.\n\nTwo smaller things: the identical 68.4% for Orthopedics and Social Work looks like a copy-paste error, and the paper promises public release but provides no link or repository. Training-data contamination from public exams is also not discussed, which matters for a benchmark of public questions.\n\nWho is this for? People building or evaluating medical LLMs, especially for Portuguese or allied health contexts, will get value from the resource and the discussion. The conclusion—report granular accuracy, not a single score—holds up even if the exact percentages don't.\n\nRecommendation: accept for peer review, but with revisions. The dataset and question are important enough to warrant referee time. The authors need to release the data, add confidence intervals, show per-model specialty spread, fix the duplicated number, and tighten the language from \"systemic\" to \"observed in the models we examined.\" As it stands, the paper is a good idea with fragile supporting numbers.","headline":"HealthQA-BR is a genuinely useful new benchmark and the aggregate-score warning is right, but the specialty-level numbers are too fragile to support the current claims.","tokens_in":9984,"tokens_out":2603,"would_cite":true,"duration_ms":30185,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that aggregate accuracy scores on healthcare benchmarks hide a \"spiky\" knowledge profile, where a top model like GPT 4.1 can score 86.6% overall yet only 60.0% in neurosurgery and 68.4% in social work.","keywords":["large language models","healthcare benchmark","medical AI evaluation","Portuguese-language NLP","Brazilian licensing exams","allied health professions","specialty-level accuracy","zero-shot evaluation"],"falsifier":"Compute exact binomial confidence intervals for each specialty-level accuracy from the reported question counts; if the intervals for top and bottom specialties overlap, the spiky-profile claim loses much of its evidentiary support. An independent re-tagging of the full dataset's specialty labels would likewise test whether the observed gaps are artifacts of metadata noise.","tokens_in":8838,"feed_emoji":"🩺","tokens_out":4879,"duration_ms":50939,"temperature":0.7,"pith_summary":"The paper argues that a single top-line accuracy score creates a dangerous illusion of competence in medical AI evaluation. To test this, it introduces HealthQA-BR, a 5,632-question benchmark built from Brazil's national licensing and residency exams that spans medicine, nursing, dentistry, psychology, social work, and other allied health professions. Zero-shot evaluation of more than 20 large language models shows that top models meet or exceed overall passing thresholds, but specialty-level accuracy varies wildly: GPT 4.1 drops from 98.7% in ophthalmology to 60.0% in neurosurgery and 68.4% in social work. The paper claims this uneven, \"spiky\" profile is systemic across all evaluated models, so reporting only aggregate scores is insufficient for safety validation.","feed_headline":"86.6% accuracy masks 60% neurosurgery score","feed_subtitle":"A new 5,632-question Brazilian benchmark shows why medical AI needs specialty-level reporting, not just overall scores.","key_machinery":"The central instrument is HealthQA-BR, a 5,632-question multiple-choice dataset drawn from Brazilian licensing and residency examinations and tagged with more than 30 professional and subspecialty categories. The mechanism that carries the argument is the fine-grained metadata tagging, which lets the authors compute accuracy per specialty and profession rather than only overall; the \"spiky profile\" claim is exactly the variance among those per-category scores.","core_discovery":"The central discovery is that high overall performance masks a spiky, specialty-dependent knowledge profile: the same model that nears perfection in one area can barely reach the passing threshold in another, and this pattern is not unique to one model but appears across the entire cohort. The paper presents this as direct evidence that single-score evaluation of medical AI is inadequate and that granular, specialty- and profession-level auditing is a prerequisite for assessing safety and reliability.","pith_inferences":["If the spiky-profile finding generalizes, a natural extension is that safety validation for medical LLMs should require a per-specialty coverage map rather than a single pass/fail threshold; the paper implies this but does not formalize an evaluation standard.","Because specialty-level samples are small and no confidence intervals are reported, the headline gaps should be treated as provisional until replicated on larger or independently tagged question sets.","A direct test of the proposed training-data bias would be to compare model performance on public-health and primary-care items against specialized items while controlling for question difficulty or wording."],"forward_implications":["Even top-performing models cannot be certified as clinically safe on the basis of overall accuracy; specialty-level validation is argued to be a fundamental prerequisite.","Disciplines integral to public health, such as social work and collective medicine, show notably lower scores, suggesting a bias in training data toward specialized medicine at the expense of community-focused care.","Public release of the benchmark and evaluation suite lets other research groups audit additional models and run similar system-wide evaluations for Portuguese-speaking healthcare.","Developers can use the specialty-level breakdown to target remediation, such as fine-tuning or retrieval-augmented generation on the specific areas where a model is weakest."],"supporting_citations":[{"why":"Supplies the prior context of physician-centric licensing-exam evaluation that the paper argues creates an illusion of competence.","marker":"[1]"},{"why":"Provides the cross-country benchmark whose reported specialty variability the paper cites as corroboration that uneven knowledge is systemic.","marker":"[2]"},{"why":"Grounds the framing of healthcare as an interprofessional team activity that requires system-wide evaluation.","marker":"[3]"},{"why":"Official source of one of the three examination components used to build the HealthQA-BR dataset.","marker":"[8]"},{"why":"Official source providing both medical and allied health residency examination questions, including the multiprofessional component.","marker":"[9]"},{"why":"Documents the model used for the headline granular analysis of specialty-level accuracy.","marker":"[10]"}],"fun_headline_variants":["Healthcare AI scores range from 98.7% to 60% by specialty","New Brazil health benchmark: AI top scores hide 60% lows","System-wide health AI audit: high overall scores, spiky gaps","AI medical knowledge: 98.7% in eyes, 60% in neurosurgery"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"Specialty-level accuracy estimates are treated as reliable even though several categories contain very few questions, so the claimed gaps may partly reflect sampling noise.","fun_headline_variants_meta":{"raw":{"variants":["Healthcare AI scores range from 98.7% to 60% by specialty","New Brazil health benchmark: AI top scores hide 60% lows","System-wide health AI audit: high overall scores, spiky gaps","AI medical knowledge: 98.7% in eyes, 60% in neurosurgery"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001018,"raw_usage":{"total_tokens":4265,"prompt_tokens":885,"completion_tokens":3380,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":501,"completion_tokens_details":{"reasoning_tokens":3297}},"tokens_in":501,"tokens_out":3380,"duration_ms":26145,"temperature":1.0,"reasoning_tokens":3297,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T00:37:08.164734+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compute exact binomial confidence intervals for each specialty-level accuracy from the reported question counts; if the intervals for top and bottom specialties overlap, the spiky-profile claim loses much of its evidentiary support. An independent re-tagging of the full dataset's specialty labels would likewise test whether the observed gaps are artifacts of metadata noise.","supporting_citations":[{"cited_title":"How large language models perform on the United States Medical Licensing Examination: a systematic review","cited_arxiv_id":null,"evidence_quote":"Supplies the prior context of physician-centric licensing-exam evaluation that the paper argues creates an illusion of competence."},{"cited_title":"Interprofessional Collaborative Practice","cited_arxiv_id":null,"evidence_quote":"Grounds the framing of healthcare as an interprofessional team activity that requires system-wide evaluation."},{"cited_title":"Revalida","cited_arxiv_id":null,"evidence_quote":"Official source of one of the three examination components used to build the HealthQA-BR dataset."},{"cited_title":"Exame Nacional de Residência (Enare)","cited_arxiv_id":null,"evidence_quote":"Official source providing both medical and allied health residency examination questions, including the multiprofessional component."},{"cited_title":"Introducing GPT-4.1 in the API [Internet]","cited_arxiv_id":null,"evidence_quote":"Documents the model used for the headline granular analysis of specialty-level accuracy."}],"review_version":1}