{"id":"7282fad3-ef45-43d6-a738-e89d7b91dbd3","arxiv_id":"2507.15717","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"BELO is a new ophthalmology benchmark of 900 expert-checked multiple-choice questions with reasoning, used to evaluate six LLMs on accuracy and explanation quality.","lead":"Researchers created BELO, a set of 900 eye-care questions with expert-written answer explanations, to test how well AI language models know ophthalmology and explain their answers. They used it to rank six AI models and found that top reasoning models like OpenAI o1 answer accurately, but all models give weak written explanations.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The claimed reasoning assessment rests on five text-generation metrics whose validity for clinical reasoning is never demonstrated; the small expert evaluation is not compared with them and excludes wrong-answer cases, so Table 2's reasoning ranking is unsubstantiated.","rationale":"I read the paper as a benchmark-construction paper whose primary novelty is the expert-curated ophthalmology MCQ set; the accuracy comparisons are defensible. The load-bearing step for the broader claim is the inference from \"scores on five text-generation metrics\" to \"clinical reasoning quality.\" That inference has no supporting validity evidence in the paper. The human evaluation is not used to validate the automatic metrics; it is reported as a separate demonstration, and its design (only correct-answer items, only three models) makes it unable to support the metric ranking. I therefore agree with the reader's weakest_assumption, and my concern is essentially the same one. I do not think the paper should be rejected: the dataset resource and accuracy results stand, but the reasoning-evaluation claim is conditional. A focused correlation analysis on already-collected human ratings would settle it.","tokens_in":15706,"tokens_out":4308,"duration_ms":48058,"concrete_test":"Re-score the 50 human-evaluated outputs (and, ideally, 50 additional outputs sampled without restricting to correct answers) with the five text-generation metrics exactly as in Table 2, and compute per-metric Spearman correlations with the two ophthalmologists' mean \"accuracy\" and \"completeness\" scores. If no metric exceeds a pre-specified correlation threshold (e.g., rho >= 0.3) with a lower 95% CI above zero, the reasoning-quality interpretation of the weighted normalized score is falsified for this dataset.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim that BELO evaluates clinical reasoning depends entirely on the composite \"weighted normalized score\" in Table 2, which averages five reference-based text-generation metrics (ROUGE-L, METEOR, BERTScore, BARTScore, AlignScore) after min-max normalization. No evidence is provided that any of these metrics track clinical reasoning quality in ophthalmology. ROUGE-L and METEOR are n-gram/lexical overlap measures; BERTScore is semantic similarity; BARTScore is generation likelihood; AlignScore measures factual consistency between model output and reference. The only human validation performed (Table S2, Supplementary Table 5) scored 50 outputs from three models on accuracy/completeness/readability, but (i) the sample was restricted to questions where all three models chose the correct answer, removing the cases where reasoning errors are most likely to appear; (ii) the three qualitative models exclude o1, o3-mini, and DeepSeek-R1, which are exactly the models ranked top by the automatic metrics; and (iii) the paper never reports a correlation or agreement between the five metrics and the ophthalmologists' ratings. Consequently, the metric-based ranking (o1 > o3-mini > GPT-4o > DeepSeek-R1 > Llama > Gemini) and the conclusion that \"performance on text-generation metrics ... indicates room for improvement in clinical reasoning\" are not supported. This is a validity gap, not merely a wording issue: the benchmark's advertised reasoning dimension is currently measured by an unvalidated proxy. The accuracy dimension of BELO remains credible; the reasoning dimension does not.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces BELO, a benchmark of 900 ophthalmology multiple-choice questions curated from five existing medical QA datasets, with multiple rounds of expert checking by ophthalmologists and expert-written reasoning references. To demonstrate the benchmark's utility, the authors evaluate six LLMs on accuracy, macro-F1, and five text-generation metrics, and additionally report a small qualitative evaluation by two ophthalmologists on 50 outputs from three models. The central claim is that BELO is a robust, clinically relevant benchmark for assessing both ophthalmological knowledge and clinical reasoning.","tokens_in":15930,"tokens_out":3788,"duration_ms":40194,"significance":"If the reasoning-assessment claim were supported, BELO would be a valuable community resource: it is externally grounded, expert-curated, hold-out in design, and accompanied by a public leaderboard. The curation pipeline is clearly described, the expert checking is a genuine strength, and the hold-out decision is methodologically sound. However, the reasoning dimension rests on an unvalidated composite of text-generation metrics, and the human validation is too small, too biased, and not linked to those metrics. Since the paper's advertised novelty over prior ophthalmology benchmarks is precisely the evaluation of reasoning, this validity gap is load-bearing. The accuracy and macro-F1 results are informative, but the reasoning claims need either validation or reframing before the central conclusion is warranted.","major_comments":[{"comment":"The composite 'weighted normalized score' is defined as the equal-weighted average of five min-max normalized text-generation metrics (ROUGE-L, METEOR, BERTScore, BARTScore, AlignScore). These are reference-based lexical and semantic similarity measures; the paper provides no evidence that any of them tracks clinical reasoning quality in ophthalmology. The Abstract and Discussion use this composite to conclude that low scores indicate 'room for improvement in clinical reasoning,' so the reasoning claim is entirely dependent on an unvalidated proxy. Please either validate these metrics against expert judgments (e.g., correlation or agreement with ophthalmologist ratings) or explicitly reframe the reasoning claims as claims about textual similarity to expert-written references.","section":"Methods, 'Demonstrative Quantitative Analysis'; Table 2"},{"comment":"The qualitative evaluation is restricted to 50 questions on which GPT-4o, Llama-3-8B, and Gemini 1.5 Pro all selected the correct answer, and it excludes o1, o3-mini, and DeepSeek-R1—the models that rank highest on the automatic metrics. No correlation or agreement is reported between the ophthalmologists' Likert ratings and any of the five text-generation metrics. Consequently, this human evaluation cannot validate the metric-based reasoning ranking in Table 2. Please report a suitable validation analysis, or remove the implication that the qualitative evaluation supports the quantitative reasoning metric.","section":"Methods, 'Demonstrative Qualitative Evaluation'; Supplementary Table 5"},{"comment":"The macro-F1 subset is inconsistently specified. The Methods state that macro-F1 was computed using 'only questions with four options, which were questions from BCSC and MedMCQA,' but the Figure 3 legend and Table 2 report 872 questions and explicitly include MedQA. Since BCSC (260) plus MedMCQA (572) totals 832, while adding MedQA (40) gives 872, this discrepancy materially affects reproducibility. Please correct the Methods or the figure/table and state the actual subset used.","section":"Methods, 'Demonstrative Quantitative Analysis'; Figure 3 legend; Table 2"}],"minor_comments":[{"comment":"The abstract reports o1's macro-F1 as '0.78, 95% CI: 0.869–0.910,' which is inconsistent with Table 2's value of 0.890 and the same CI; the point estimate appears to be a typo.","section":"Abstract"},{"comment":"The Accuracy row does not report a 95% confidence interval for Llama-3-8B, although intervals are given in the Results text and for all other models in the table.","section":"Table 2"},{"comment":"The description of the manual check lists 'two optometrists (SY, WTL)' and then 'six research staff (SS, XA, TWSL, SY, WTL, YC),' with SY and WTL appearing in both lists; the personnel counts should be clarified.","section":"Methods, 'QA Quality Check'"},{"comment":"The terms 'comprehensiveness' (Abstract) and 'completeness' (Methods, Results, Supplementary Table S5) are used interchangeably for the same construct; this should be harmonized.","section":"Throughout"},{"comment":"The score is called a 'weighted normalized score' even though the five normalized metrics are averaged with equal weights; 'equal-weight normalized score' would be less misleading.","section":"Table 2, note"}],"recommendation":"major_revision","confidential_remarks":"The benchmark curation is a solid contribution, and the accuracy comparison is publishable. However, the paper's advertised reasoning dimension needs to be made credible before acceptance: either provide a proper validation of the text-generation metrics against expert reasoning judgments or soften the reasoning claims. The current qualitative evaluation is too narrow to bridge the gap. I would require the authors to add the missing validation or reframing as a condition of publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nRead BELO. The genuinely new thing is the dataset: 900 ophthalmology MCQs from five sources, filtered and re-annotated through multiple rounds of expert checking by 13 ophthalmologists, with explanations for every answer. That is a real contribution. Prior ophthalmology benchmarks are accuracy-only or smaller; this one adds a reasoning reference standard and intends to be a hold-out evaluation set. The curation pipeline is described well enough to reproduce the sampling, and oversampling the small datasets is sensible. The accuracy benchmarking of six LLMs on 900 items is credible as a demonstration.\n\nThe soft spot is the reasoning evaluation. The \"weighted normalized score\" in Table 2 averages five text-generation metrics (ROUGE-L, METEOR, BERTScore, BARTScore, AlignScore) after min-max normalization. The paper never shows that any of these correlate with clinical reasoning quality. The human evaluation is separate — 50 questions where all three models got the right answer, covering GPT-4o, Llama, and Gemini only — and there is no reported agreement or correlation between those metrics and the ophthalmologists' ratings. So the claim that o1 ranks first in reasoning is not supported. The authors call the evaluation \"demonstrative,\" which softens this, but the abstract and conclusion still present reasoning assessment as a core feature. That needs either a validation study or a more cautious claim.\n\nOther issues are minor. The abstract's macro-F1 for o1 is misreported (0.78 in the abstract, 0.890 in Table 2). The BARTScore normalization direction is easy to misread. No code or processed data are released, which is understandable given the hold-out design but still limits external audit; at least the leaderboard is public.\n\nWho is this for? Anyone building or evaluating ophthalmology LLMs. The dataset itself is worth having even if the reasoning metric is fixed. I'd send it to review; the referee should push on the reasoning-metric validation and the accuracy of the reported numbers.\n\nRecommendation: accept with major revisions, not desk reject.","headline":"A genuinely useful expert-curated ophthalmology QA resource, but the reasoning score in Table 2 rests on unvalidated text-generation metrics; the accuracy results are solid.","tokens_in":16663,"tokens_out":2099,"would_cite":true,"duration_ms":20660,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that a 900-question, expert-refined benchmark, BELO, can serve as a standardized test of both accuracy and reasoning for ophthalmology LLMs, and that on its first run the o1 model leads while all models' explanations lag.","keywords":["ophthalmology benchmark","large language models","clinical reasoning evaluation","multiple-choice questions","text-generation metrics","expert-curated dataset","hold-out evaluation","LLM leaderboard"],"falsifier":"Have two or more ophthalmologists blindly rate, on a 5-point scale, a random sample (say 100) of model explanations from BELO for reasoning quality, then compute the rank correlation between the mean expert ratings and the five text-generation metric scores; a near-zero or negative correlation would refute the claim that the benchmark measures clinical reasoning.","tokens_in":15469,"feed_emoji":"👁️","tokens_out":9006,"duration_ms":85122,"temperature":0.7,"pith_summary":"BELO (BEnchmarking LLMs for Ophthalmology) is a new evaluation benchmark built from 900 multiple-choice questions filtered from five medical QA datasets and refined through multiple rounds of checking by 13 ophthalmologists. The authors' central claim is that BELO supplies a standardized, held-out test that measures not only whether a language model picks the correct answer but also the quality of the reasoning it gives, by comparing model explanations against expert-written reference explanations. To show the benchmark works, they ran six current LLMs through it: the o1 model achieved the highest accuracy (0.882) and macro-F1 (0.890), while explanation-quality scores from five text-generation metrics were uniformly modest, with the best aggregate reasoning score at 0.804 on a 0-1 scale. If this claim holds, the field gains a common ruler for comparing ophthalmology LLMs and a concrete target—improving explanation quality—that accuracy alone had hidden.","feed_headline":"900 expert-checked questions put six LLMs to the ophthalmology test","feed_subtitle":"The best model answers 88% correctly, yet every model's explanations score low on reasoning metrics.","key_machinery":"The central object is the BELO dataset itself: 900 MCQ items drawn from BCSC, BioASQ, MedMCQA, MedQA, and PubMedQA, extracted by keyword matching plus a fine-tuned PubMedBERT classifier, then graded, amended, and adjudicated by 13 ophthalmologists so that every question carries a reference explanation. The evaluation machinery has two parts: accuracy and macro-F1 measure whether the chosen answer is right, while five text-generation metrics (ROUGE-L, METEOR, BERTScore, BARTScore, AlignScore) measure how closely each model's explanation matches the expert reference, collapsed into a weighted normalized score. The dataset is kept as a hold-out, evaluation-only set with a public leaderboard, so future models face the same questions without prior exposure.","core_discovery":"The paper's discovery is a validation set rather than a new model: 900 expert-checked ophthalmology MCQs, each paired with a reference explanation, that can rank LLMs on two dimensions at once. On the demonstration run, answer accuracy and macro-F1 clearly separated the models (o1 first at 0.882/0.890, Gemini 1.5 Pro last at 0.596/0.639), but the five explanation metrics (ROUGE-L, METEOR, BERTScore, BARTScore, AlignScore) all produced low absolute scores across the board once rescaled, with o1 again first on the weighted normalized score (0.804) and Gemini 1.5 Pro last (0.037). The authors interpret this gap as evidence that current models can often identify the right answer yet provide explanations that fall short of expert-level reasoning.","pith_inferences":["The five text-generation metrics are never validated against expert ratings in the paper; a direct calibration study comparing metric scores with ophthalmologists' judgments on the same outputs would show whether the reasoning ranking is measuring reasoning or just lexical similarity.","The low absolute text-generation scores could partly reflect reference explanations being longer and more detailed than model outputs, so a length-controlled or rubric-based evaluation might reorder the models relative to the reported ranking.","Because the dataset is evaluation-only and held out, model developers cannot build a development set from it; without a companion training or validation split, BELO cannot be used to drive iterative improvement, only to assess finished models.","The extraction pipeline's keyword method had 0.718 sensitivity on MedMCQA, meaning nearly a third of true ophthalmology questions were missed; a future version could recover those by applying the higher-sensitivity classifier-only approach followed by expert screening."],"forward_implications":["Any future LLM in ophthalmology can now be compared against six already-published scores on the same 900 questions, because BELO is a fixed, held-out benchmark.","The near-universally low explanation-metric scores indicate that reasoning quality, not answer selection, is the current bottleneck in ophthalmology LLMs.","The benchmark's composition—572 exam-style questions from MedMCQA and 260 from BCSC out of 900—means current results mostly measure knowledge recall rather than real-world clinical management.","Because the dataset is withheld from public release, models cannot be legitimately fine-tuned on it, preserving the validity of future head-to-head comparisons.","The planned extensions to visual question answering and clinical scenario tasks will broaden the benchmark beyond text-only MCQs."],"supporting_citations":[{"why":"Supplies the prior method for grading reasoning quality and the set of text-generation metrics that BELO adapts to ophthalmology.","marker":"12"},{"why":"MedQA is one of the five source datasets, contributing 40 USMLE-style four-option items without original reference reasoning.","marker":"25"},{"why":"MedMCQA contributes 572 of the 900 items and provides the subject-level labels used to train and evaluate the extraction methods.","marker":"34"},{"why":"PubMedQA contributes 18 expert-reasoned research items and shows how BELO blends exam and literature-derived questions.","marker":"35"},{"why":"PubMedBERT is the model fine-tuned to classify ophthalmology questions, the higher-sensitivity arm of the extraction pipeline.","marker":"36"},{"why":"ROUGE-L is one of the five metrics used to score model explanations against reference reasoning.","marker":"39"},{"why":"BERTScore is one of the five metrics, measuring semantic similarity via contextual embeddings.","marker":"40"},{"why":"BARTScore is one of the five metrics, evaluating the generated reasoning as a text-generation task.","marker":"41"},{"why":"AlignScore is one of the five metrics, scoring factual consistency between model output and reference.","marker":"42"},{"why":"METEOR is the fifth metric, combining precision, recall, stemming, and synonym matching for the explanation comparison.","marker":"47"}],"fun_headline_variants":["900 expert questions: LLMs get answers right, reasoning wrong","Benchmark: 900 eye QAs show LLMs struggle to reason like experts","Ophthalmology benchmark: 6 LLMs, 900 questions, one big reasoning gap","New eye benchmark ranks LLMs on both accuracy and reasoning","Expert-checked benchmark: LLMs ace answers, flunk explanations"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the five text-generation metrics, when comparing a model's explanation to the expert reference, genuinely measure the quality of clinical reasoning; if those metrics track style or wording instead of reasoning, BELO's claim to assess reasoning is not supported.","fun_headline_variants_meta":{"raw":{"variants":["900 expert questions: LLMs get answers right, reasoning wrong","Benchmark: 900 eye QAs show LLMs struggle to reason like experts","Ophthalmology benchmark: 6 LLMs, 900 questions, one big reasoning gap","New eye benchmark ranks LLMs on both accuracy and reasoning","Expert-checked benchmark: LLMs ace answers, flunk explanations"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000334,"raw_usage":{"total_tokens":1910,"prompt_tokens":1058,"completion_tokens":852,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":674,"completion_tokens_details":{"reasoning_tokens":770}},"tokens_in":674,"tokens_out":852,"duration_ms":8046,"temperature":1.0,"reasoning_tokens":770,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T15:24:44.024446+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have two or more ophthalmologists blindly rate, on a 5-point scale, a random sample (say 100) of model explanations from BELO for reasoning quality, then compute the rank correlation between the mean expert ratings and the five text-generation metric scores; a near-zero or negative correlation would refute the claim that the benchmark measures clinical reasoning.","supporting_citations":[{"cited_title":"Can OpenAI o1’s Enhanced Reasoning Capabilities Extend to Ophthalmology? A Benchmark Study Across Large Language Models and Text Generation Metrics [Internet]","cited_arxiv_id":null,"evidence_quote":"Supplies the prior method for grading reasoning quality and the set of text-generation metrics that BELO adapts to ophthalmology."},{"cited_title":"Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question answering","cited_arxiv_id":null,"evidence_quote":"MedMCQA contributes 572 of the 900 items and provides the subject-level labels used to train and evaluate the extraction methods."},{"cited_title":"PubMedQA: A Dataset for Biomedical Research Question Answering","cited_arxiv_id":null,"evidence_quote":"PubMedQA contributes 18 expert-reasoned research items and shows how BELO blends exam and literature-derived questions."},{"cited_title":"Domain-Specific Language Model Pretraining for Biomedical Natural Language Processing","cited_arxiv_id":null,"evidence_quote":"PubMedBERT is the model fine-tuned to classify ophthalmology questions, the higher-sensitivity arm of the extraction pipeline."},{"cited_title":"ROUGE: A Package for Automatic Evaluation of Summaries [Internet]","cited_arxiv_id":null,"evidence_quote":"ROUGE-L is one of the five metrics used to score model explanations against reference reasoning."},{"cited_title":"Meteor: an automatic metric for MT evaluation with high levels of correlation with human judgments","cited_arxiv_id":null,"evidence_quote":"METEOR is the fifth metric, combining precision, recall, stemming, and synonym matching for the explanation comparison."}],"review_version":1}