{"id":"3310a891-6901-4cbe-aafa-90ef7342e95d","arxiv_id":"2502.01243","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"OphthBench is a new 591-question Chinese ophthalmology benchmark on which 39 LLMs score around 70% (after normalization), showing a clear gap between current models and clinical readiness.","lead":"The authors introduce OphthBench, a set of 591 Chinese-language ophthalmology questions spanning education, triage, diagnosis, treatment, and prognosis, and they score 39 large language models on it. The results suggest that today's LLMs, especially non-Chinese ones, still fall well short of what would be needed for real clinical use in Chinese eye care.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central 'gap' claim is uncalibrated: §3.3 Eq. 1 normalizes every task's best LLM to 90, so the leaderboard ranks models relative to each other and never measures performance against a clinical threshold or human expert baseline.","rationale":"I read the paper as attempting to show that contemporary LLMs are not yet clinically useful in Chinese ophthalmology, and the load-bearing evidence is the empirical evaluation. The reader's weakest assumption focuses on the unvalidated LLM-as-judge protocol for OEQ scoring, which is a real threat to the reliability of individual task scores. However, I find a more fundamental issue: even if the judge were perfectly aligned with human experts, the paper's reported scores are normalized relative to the best observed model. Eq. 1 makes the top model a 90 on every task and discards absolute performance information. Therefore, the leaderboard cannot support the paper's central quantitative claim about a 'substantial gap with practical requirements' unless an external anchor is supplied. The paper hints at an absolute 70% performance rate in Fig. 4, but the main tables are normalized and no human or chance baseline is reported, so the reader cannot verify the magnitude of the gap. This concern does not attack the benchmark's construction or the value of collecting 591 questions; it attacks the inferential step from raw or normalized scores to a clinical-utility conclusion. The proposed test—comparing a blinded expert sample and random baselines against unnormalized model scores—would directly calibrate the claim. Because the concern is addressable by adding baselines and reporting raw scores, I keep the reader's CONDITIONAL verdict rather than escalating to REJECT, and I partially agree with the reader: the OEQ judge issue is valid, but the normalization/calibration issue is the more load-bearing obstacle to the central claim.","tokens_in":16795,"tokens_out":7692,"duration_ms":76344,"concrete_test":"Recruit at least three board-certified Chinese ophthalmologists, none of whom authored the benchmark, to independently answer a stratified random sample of at least 100 questions (including open-ended questions) under the same conditions as the models, and compute their raw accuracy, F1, and quality scores plus inter-rater agreement. Then recompute the OphthBench leaderboard without the Eq. 1 normalization and compare model scores against (a) random-chance baselines and (b) the expert distribution. If the best models fall within the expert range or meet a prespecified clinical threshold, the 'substantial gap' claim is unsupported; if they remain significantly below experts, the claim stands.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The headline conclusion is that current LLMs show a substantial gap between capabilities and practical requirements. The evidence offered in §4 is the leaderboard, but the protocol in §3.3 (Eq. 1) explicitly renormalizes each task by dividing by the best model's raw score and multiplying by 90. Consequently, the top model on each task is always assigned 90; every other score is a ratio to that observed maximum. These numbers contain no information about whether the best model is 50%, 80%, or 95% correct, so they cannot calibrate 'practical requirements.' The paper also never reports human expert performance on the same 591 questions or a random-chance baseline, and no clinical acceptance threshold is defined. A model could rank first with a normalized score near 90 while still being far below expert-level accuracy, or the whole field could be near expert level and the normalized scores would look similar. The phrase 'scoring rate of approximately 70%' in §4 (Fig. 4) gestures at an absolute figure, but the caption and text do not state whether this is raw accuracy, and the main tables are normalized. Without a prespecified reference standard, the central 'substantial gap' conclusion is not derivable from the reported results.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces OphthBench, a benchmark of 591 questions for evaluating LLMs in Chinese ophthalmology, organized into five clinical workflow scenarios (Education, Triage, Diagnosis, Treatment, Prognosis) and nine tasks, with single-choice, multiple-choice, and open-ended questions. The authors evaluate 39 LLMs under two prompting regimes (common and advanced) and report leaderboard scores. The central claim is that the results reveal a substantial gap between current model capabilities and practical clinical requirements in Chinese ophthalmology, with secondary findings that Chinese models outperform non-Chinese models, medical-specific LLMs do not surpass general LLMs, and prompt engineering substantially improves performance.","tokens_in":17071,"tokens_out":5381,"duration_ms":47894,"significance":"If properly validated, OphthBench would be a useful resource: it is one of the few benchmarks explicitly aligned with a clinical workflow in a non-English medical context, it combines three question formats, and the evaluation includes both common and advanced prompting protocols. The authors also include an external MMLU reference column, report a t-test for one subgroup comparison, and describe their model deployment and output-constraint methods, which aids reproducibility. However, the paper's headline conclusion of a \"substantial gap\" between model capabilities and practical requirements is not currently supported by the reported metrics, because the normalization in Eq. (1) removes absolute score information and no human expert, random-chance, or clinical-threshold baseline is provided. The benchmark construction and the multi-model evaluation are credible contributions, but the measurement validity needs to be demonstrated before the practical-utility conclusion can be accepted.","major_comments":[{"comment":"The normalization in Eq. (1) divides every task score by the best observed model score and multiplies by 90, so the reported leaderboard values are relative to the best model on each task; they contain no information about absolute correctness. The abstract and Section 4 conclude that there is a \"substantial gap between current model capabilities and practical requirements,\" but no clinical threshold, human expert baseline, or random-chance baseline is defined anywhere in the manuscript. The \"performance rate of approximately 70%\" mentioned in Section 4 is not tied to the tables, and the text does not state whether it is raw accuracy; if it refers to normalized scores it is an artifact of the 90-point ceiling. Please report raw accuracy/F1 per task and scenario, add human expert performance on the same 591 questions and a random-chance baseline, and provide confidence intervals, so that the conclusion about a gap to practical requirements can actually be tested.","section":"Section 3.3, Eq. (1); Section 4"},{"comment":"For OEQs, reference answers were drafted by GPT-4o and then reviewed by ophthalmologists, and responses are scored by CompassJudger-1-7B comparing them to those references on seven dimensions. The manuscript provides no calibration evidence for this judge on this benchmark: no human-scored subset, no inter-rater agreement, and no correlation between judge scores and expert opinion. Because GPT-4o was also used to refine the advanced OEQ prompts, the evaluation loop is self-referential: an LLM helped create the gold responses and an LLM judges the answers. If the judge's dimension preferences diverge from expert judgment, the Triage, Diagnosis, and Treatment scenario scores would be systematically biased. Please add a human-validation sub-study on a stratified random sample of responses and report agreement statistics (e.g., Cohen's kappa or correlation) before using the judge scores as the basis for task rankings.","section":"Sections 3.2 and 3.3, OEQ protocol"},{"comment":"The harmonic-mean aggregation across tasks is not justified and mixes relative per-task normalized scores into the \"All scenarios\" column; moreover, no confidence intervals or significance tests are reported for leaderboard differences, so adjacent ranks (e.g., ranks 1-3 in Table 3) may not be statistically distinguishable. The t-test for Chinese vs. non-Chinese 6B-9B models gives P=0.096 and P=0.085, which is at best marginal, yet the text states that Chinese models \"consistently outperformed\" non-Chinese models. Provide bootstrap confidence intervals around the aggregated scores and report the t-test details (e.g., whether it is paired, the sample size, and the effect size); otherwise, these claims are not supported.","section":"Section 4, Tables 3 and 4, t-test"},{"comment":"The stated limitations in Section 5 mention incomplete model coverage and limited dataset size, but do not acknowledge the three most consequential threats to the paper's central claim: the absence of a human expert baseline, the lack of calibration for the LLM-as-judge, and the fact that the normalization in Eq. (1) removes absolute score information. Since the conclusion of a \"substantial gap\" depends on these measurement properties, the limitations section should be extended accordingly.","section":"Section 5, Limitations"}],"minor_comments":[{"comment":"The text refers to \"HuatuoGPT2-o1-7B\" as being based on Qwen2.5-7B, but Tables 2-4 list \"HuatuoGPT-o1-7B\"; the naming is inconsistent and should be corrected.","section":"Section 4, paragraph on medical LLMs"},{"comment":"The sentence ending with \"These advancements\" is cut off in the preprint and should be completed.","section":"Section 4, final paragraph of the first subsection"},{"comment":"The statistics panel in Figure 2 is not legible in the preprint; please provide a vector graphic with labeled axes and clearly readable text.","section":"Figure 2"},{"comment":"The MMLU column is not part of OphthBench; clarify in the caption or text whether these are external reference scores and how they were obtained, and state why they are included in an OphthBench leaderboard.","section":"Tables 3 and 4, MMLU column"},{"comment":"The abstract says the benchmark was developed by \"three experienced Chinese ophthalmologists,\" while Section 3.2 describes \"three junior ophthalmologists\" plus a senior reviewer; please make the count and seniority consistent.","section":"Section 3.2 vs. Abstract"},{"comment":"Figure 5 reports a maximum improvement of 175.0% and a decline of 62.2%, but the text does not explain which tasks or models produced these extremes; please clarify.","section":"Figure 5"}],"recommendation":"major_revision","confidential_remarks":"The paper presents a potentially useful benchmark and a broad model evaluation, but the headline 'substantial gap' claim is not yet supported because of the normalization and missing baselines. These issues are fixable within the manuscript's scope: the authors already have the raw scores and have access to ophthalmologists, so adding human and random-chance baselines and validating the judge would materially strengthen the paper. I do not see a fundamental flaw in the benchmark construction itself, but the current presentation overstates what the normalized scores can show. No conflict of interest."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"OphthBench is the kind of paper that's worth reading even when its headline conclusion doesn't hold. The contribution is the benchmark itself: a Chinese ophthalmology evaluation set organized around a clinical workflow—education, triage, diagnosis, treatment, prognosis—with 591 questions across 9 task types, including both multi-choice and open-ended items. That's genuinely new; I'm not aware of another public Chinese ophthalmology benchmark with this scope, and the workflow taxonomy is a useful organizing principle for the area.\n\nThe evaluation effort is real: 39 LLMs, two prompt settings, careful attention to formatting. The comparative ranking is probably the most useful part of the paper. But the central claim that 'current model capabilities lag practical requirements' is not derivable from the results as reported. Equation (1) normalizes each task by dividing by the best model's raw score and multiplying by 90, so the top model always gets exactly 90. That means the leaderboard is purely relative—it can tell you which model is best, but it cannot tell you whether any model is close to clinical competence. Without a human expert baseline or a pre-specified acceptance threshold, the 'gap' language is unjustified. This is the main weakness and it's a load-bearing one for the paper's narrative.\n\nThe open-ended evaluation has a secondary but real problem. Reference answers were drafted by GPT-4o and reviewed by three ophthalmologists, which is a reasonable start, but the scoring is done by CompassJudger-1-7B with no human-judge calibration or inter-rater agreement statistics. If that judge's preferences diverge from an expert panel's, the OEQ scenario scores (triage, diagnosis, treatment) would be systematically off.\n\nMinor issues: no random-chance baseline, no confidence intervals, and the dataset/code aren't released, which makes the benchmark hard to adopt. The 'approximately 70%' in Figure 4 isn't clearly identified as raw accuracy.\n\nNone of this is fatal to the benchmark itself. The taxonomy, question construction, and model coverage are solid enough to deserve referee time. I'd recommend the editors treat this as a candidate for peer review with major revision: re-analyze without the normalization, add a human expert comparison, validate or at least sanity-check the judge, and release the data.","headline":"Useful Chinese ophthalmology benchmark, but the headline 'substantial gap' claim is an artifact of the per-task normalization to 90, not a measured fact.","tokens_in":17570,"tokens_out":4034,"would_cite":true,"duration_ms":33189,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A 591-question benchmark built around the Chinese eye-care workflow shows that today's large language models still fall short of clinical readiness.","keywords":["large language models","ophthalmology","Chinese medical benchmark","clinical workflow","LLM-as-a-judge","medical question answering","prompt engineering"],"falsifier":"Select a random sample of open-ended responses from the triage, diagnosis, and treatment tasks, have three independent ophthalmologists score them without seeing the judge's scores, and measure agreement; if agreement is low, the benchmark's open-ended scenario scores are not trustworthy.","tokens_in":16626,"feed_emoji":"👁️","tokens_out":5257,"duration_ms":46295,"temperature":0.7,"pith_summary":"OphthBench is a benchmark for deciding whether large language models can be useful in Chinese ophthalmology, built by splitting a typical clinical workflow into education, triage, diagnosis, treatment, and prognosis. It contains 591 questions across nine tasks, mixing single-choice, multiple-choice, and open-ended formats, and it evaluates 39 popular LLMs under two prompting regimes. The paper's central finding is that even the best models reach only the mid-80s on a normalized scale, with most scoring near 70 percent, leaving a substantial gap between current capabilities and practical clinical requirements. A sympathetic reader would take this as evidence that the field needs both workflow-aligned evaluation and further model development before deployment.","feed_headline":"39 LLMs fall short on a 591-question Chinese eye-care benchmark","feed_subtitle":"Workflow-aligned test spanning education, triage, diagnosis, treatment, and prognosis exposes a large gap to real practice.","key_machinery":"The central object is OphthBench itself: a five-scenario, nine-task question set that maps onto a typical ophthalmic clinical workflow and includes single-choice, multiple-choice, and open-ended questions. The load-bearing scoring machinery is a normalized, rule-calibrated multi-dimensional LLM-as-a-judge protocol, in which the CompassJudger-1-7B model scores open-ended answers against ophthalmologist-reviewed reference responses on seven dimensions, task scores are normalized by the best model's score, and scenario totals are computed with harmonic means. Two prompting regimes, a common prompt and an advanced prompt generated with the CO-STAR framework, are used to reduce prompt-sensitivity bias.","core_discovery":"OphthBench's central claim is that LLM performance in Chinese ophthalmology should be measured against a workflow-shaped yardstick rather than isolated exam questions, and that when such a yardstick is built, 39 popular models show a substantial gap between their scores and practical requirements. The paper constructs this yardstick from 591 questions organized into five scenarios and nine tasks, with answers reviewed by three experienced ophthalmologists, and evaluates models under both common and optimized prompts. The strongest observed models reach the mid-80s on the normalized scale while many models sit near or below 70 percent, which the paper interprets as evidence that current LLMs are not yet ready for real-world ophthalmic use.","pith_inferences":["Our inference: the same five-scenario template could be reused for other Chinese clinical specialties, producing comparable cross-specialty readiness scores.","Our inference: because the benchmark contains only text, adding fundus photography and OCT interpretation would probably lower diagnostic scores further and better reflect real eye-care work.","Our inference: the judge-model scores could be converted into a training signal by collecting a small human-rated sample and using those preferences to fine-tune either the judge or the clinical model, something the paper does not attempt."],"forward_implications":["If OphthBench reflects real clinical utility, current LLMs, including the best commercial systems, are not yet dependable enough for unsupervised use in Chinese eye care because the observed ceiling sits near the mid-80s rather than near 100.","Medical-specific LLMs trained in Chinese did not outperform general-purpose models on this benchmark, suggesting that domain post-training alone does not translate into ophthalmic competence.","Prompt engineering can raise most models' scores substantially, so part of the capability gap is accessible through better instruction design rather than only through better models.","The prognosis scenario proved easiest and the education scenario hardest, giving developers a task-level map of where to focus future work.","The benchmark's workflow structure provides a concrete way to compare models on triage, diagnosis, and treatment skills rather than on general medical knowledge alone."],"supporting_citations":[{"why":"Defines the tertiary eye-center clinical workflow that grounds the five scenarios.","marker":"[46]"},{"why":"Supplies CompassJudger-1-7B, the judge model used to score open-ended responses.","marker":"[42]"},{"why":"Contributes the rule-calibrated multi-dimensional point-wise LLM-as-a-judge evaluation method.","marker":"[48]"},{"why":"Provides the CO-STAR prompt framework used to generate advanced prompts for open-ended questions.","marker":"[47]"},{"why":"Shows prior work on a localized Chinese medical benchmark that OphthBench extends to ophthalmology.","marker":"[40]"},{"why":"Serves as the Chinese medical-exam benchmark baseline against which OphthBench's workflow design is distinguished.","marker":"[32]"}],"fun_headline_variants":["OphthBench: 39 LLMs fail 591-question eye-care test","Eye-care AI benchmark: 39 models underperform in clinic","591 questions, 5 workflows: LLMs not ready for eye care","Chinese eye-care benchmark shows LLMs gap to practice","OphthBench: LLMs score below practical need in Chinese eye care"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the automated judge's scores on open-ended answers agree with what expert ophthalmologists would consider clinically correct, but the paper reports no human validation or inter-rater agreement to demonstrate this.","fun_headline_variants_meta":{"raw":{"variants":["OphthBench: 39 LLMs fail 591-question eye-care test","Eye-care AI benchmark: 39 models underperform in clinic","591 questions, 5 workflows: LLMs not ready for eye care","Chinese eye-care benchmark shows LLMs gap to practice","OphthBench: LLMs score below practical need in Chinese eye care"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00033,"raw_usage":{"total_tokens":1824,"prompt_tokens":916,"completion_tokens":908,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":532,"completion_tokens_details":{"reasoning_tokens":814}},"tokens_in":532,"tokens_out":908,"duration_ms":613606,"temperature":1.0,"reasoning_tokens":814,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-09T15:55:09.566905+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Select a random sample of open-ended responses from the triage, diagnosis, and treatment tasks, have three independent ophthalmologists score them without seeing the judge's scores, and measure agreement; if agreement is low, the benchmark's open-ended scenario scores are not trustworthy.","supporting_citations":[{"cited_title":"Large language models and their impact in ophthalmology","cited_arxiv_id":null,"evidence_quote":"Defines the tertiary eye-center clinical workflow that grounds the five scenarios."},{"cited_title":"Alignbench: Benchmarking chinese alignment of large language models","cited_arxiv_id":null,"evidence_quote":"Contributes the rule-calibrated multi-dimensional point-wise LLM-as-a-judge evaluation method."},{"cited_title":"How i won singapore’s gpt-4 prompt engineering competition","cited_arxiv_id":null,"evidence_quote":"Provides the CO-STAR prompt framework used to generate advanced prompts for open-ended questions."},{"cited_title":"Benchmarking large language models on cmexam-a comprehensive chinese medical exam dataset","cited_arxiv_id":null,"evidence_quote":"Serves as the Chinese medical-exam benchmark baseline against which OphthBench's workflow design is distinguished."}],"review_version":1}