{"id":"8d472beb-7193-4329-9548-cad2bbc46eda","arxiv_id":"2506.00991","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"GOBench measures how well multimodal AI models generate and understand geometric optics, finding that even top models make frequent physical errors.","lead":"This paper introduces GOBench, a benchmark that tests whether image-generating and image-understanding AI models respect basic rules of light, such as shadows, reflections, and refraction. The results show that today's best models still fail often, with the top understanding model scoring about 37% on the optical judgment tasks.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The headline 37.35% rests on an unreported expert-agreement level; without inter-rater reliability, low SPS may reflect expert noise rather than an optical understanding deficit.","rationale":"The reader's weakest_assumption identifies exactly the load-bearing issue: the six-expert ground truth is treated as valid and reliable without any inter-rater agreement or variance data. My stress-test sharpens this into a quantitative mechanism: SPS compares each MLLM to the averaged expert score with a fixed tolerance, so expert disagreement directly deflates SPS. Since the paper provides no evidence that experts agree within the 0.5 tolerance, the reported 37.35% cannot be separated from noise. The abstract's use of 'accuracy' for SPS is a further interpretive error, but the deeper issue is that the metric's denominator—a noisy subjective mean—is unvalidated. Other concerns (manual filtering, missing code) are real but secondary; they affect generalizability rather than the core validity of the score. A concrete reliability analysis would settle whether the concern lands: if expert agreement is high (e.g., alpha > 0.7 and most pairwise differences below 0.5), the SPS interpretation would be substantially supported; if not, the central quantitative claim weakens. The reader's CONDITIONAL verdict remains appropriate, so I recommend no change.","tokens_in":12991,"tokens_out":4444,"duration_ms":51515,"concrete_test":"Compute inter-rater reliability on the six experts' Optical Authenticity scores (e.g., Krippendorff's alpha or Fleiss' kappa), then recompute each MLLM's SPS against each individual expert instead of the pooled mean. If the median pairwise expert difference on a 0–5 scale exceeds delta_max = 0.5, or if the spread of per-expert SPS values for the top model is comparable to the reported 37.35%, the low SPS is not diagnostic of an optical understanding deficit. Also report the sensitivity of Table 3 to delta_max (e.g., delta_max = 0.25 and 1.0) to show whether the ranking is robust to the tolerance choice.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative claim—that Gemini-2.5Pro attains only 37.35% on optical understanding—depends on treating the six-expert average as ground truth and interpreting the Scaled Proximity Score (Eq. 1) as a meaningful performance measure. Section 2.2 states that six experts score each image, but no inter-rater reliability or variance is reported. The SPS is computed against the averaged expert score with delta_max = 0.5. If the six experts disagree on an image by more than 0.5 points on the 0–5 Authenticity scale, then even an MLLM that perfectly matches one expert's judgment will receive SPS = 0 against the mean. Low SPS would then reflect expert disagreement, not optical misunderstanding. This is not a hypothetical artifact: the paper gives no evidence that expert scores cluster within 0.5, and the 'Cannot be determined' option (0.5 points) is explicitly designed for ambiguous cases, which likely inflates cross-expert variance. Additionally, the abstract and Section 3.3 call this SPS 'accuracy,' but SPS is a tolerance-weighted proximity to a subjective mean, not the percentage of correctly answered optical questions. Without expert-agreement data, the 37.35% figure cannot be interpreted as evidence that MLLMs fail at optical understanding; it may only show that MLLM ratings differ from the average of six humans who may themselves differ substantially.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces GOBench, a benchmark for evaluating multi-modal large language models (MLLMs) on two tasks: generating images that obey geometric optics principles and understanding optical phenomena in images. The authors construct GOBench-Gen-1k, a dataset of 1,000 images generated by three MLLMs (GPT-4o-Image, Seedream 3.0, Imagen 3) from 340 curated prompts across three optical categories (direct light, reflection, refraction). A panel of six human experts scores each image on Optical Authenticity, Aesthetic Quality, and Instruction Fidelity, and the averaged expert scores serve as ground truth. Eleven MLLMs are then prompted to score the same images on the same dimensions, and their outputs are compared against the human ground truth using a Scaled Proximity Score (SPS) with a hand-set tolerance delta_max = 0.5. The main reported findings are that GPT-4o-Image achieves the highest generation Authenticity score (4.09/5) but still makes optical errors, and that the best understanding model, Gemini-2.5Pro, attains only 37.35% SPS on Optical Authenticity, leading the authors to conclude that current MLLMs face significant challenges in both optical generation and understanding.","tokens_in":13350,"tokens_out":2555,"duration_ms":26709,"significance":"If the measurement is trustworthy, GOBench would be a useful first systematic benchmark for geometric optics in MLLMs, an underexplored but visually important domain. The paper has concrete strengths: the scenario taxonomy is sensible, the dataset and code are promised to be public, the evaluation protocol is described in enough detail to be partially reproducible, and the qualitative conclusion that MLLMs struggle with fine-grained optical phenomena is plausible and consistent with related perception benchmarks. However, the quantitative headline results are fragile because they rest entirely on the reliability of a six-expert subjective ground truth and on the chosen SPS tolerance; neither is validated or characterized. The abstract's characterization of SPS as 'accuracy' overstates what the metric measures. The significance of the numerical rankings in Table 3 therefore cannot currently be assessed.","major_comments":[{"comment":"The central quantitative claim, that Gemini-2.5Pro attains only 37.35% on optical understanding, depends on treating the six-expert average as ground truth, but the paper reports no inter-rater reliability statistic (e.g., Fleiss' kappa, ICC, or per-image expert variance). Since the SPS gives zero credit whenever an MLLM's score differs from the averaged expert score by more than delta_max = 0.5, large expert disagreement directly deflates all model scores. If the six experts themselves disagree by more than 0.5 points on a meaningful fraction of images, then low SPS values may reflect expert noise rather than an optical understanding deficit. Please report per-image expert score distributions and an inter-rater reliability measure, and show how the conclusions change under alternative ground-truth aggregation rules (e.g., median or per-expert agreement).","section":"Section 2.2 and Section 3.3"},{"comment":"The Scaled Proximity Score is not accuracy, but the abstract and Section 3.3 describe the 37.35% figure as 'accuracy.' SPS is a tolerance-weighted proximity to a subjective mean, and its scale depends on the arbitrarily chosen delta_max = 0.5. A model that exactly matches one expert but differs from the averaged ground truth by more than 0.5 receives zero for that case, so the numerical value is not interpretable as percent correct. Please report standard per-question accuracy on the Yes/No/Cannot-determine answers as a separate metric, provide a sensitivity analysis over delta_max, and state the chance-level baseline for SPS given the three-option response format.","section":"Equation (1), Abstract, Section 3.3"},{"comment":"The ground truth for physical optical correctness is entirely based on the judgments of six experts viewing MLLM-generated images that the authors manually filtered. There is no external validation against physically grounded references, such as ray-traced renders of the same prompts or known-correct photographs. Without such validation, the benchmark risks measuring 'agreement with six experts' rather than 'agreement with optical physics.' At minimum, please describe the manual filtering criteria quantitatively, report how many images were removed per model, and validate a random subset of expert judgments against an external physical ground truth or against a second independent panel.","section":"Section 2.1 and Section 2.2"},{"comment":"The scoring of 'Cannot be determined' as 0.5 points is arbitrary and may systematically affect both the human ground truth and the MLLM comparison. For instance, if experts use this option for genuinely ambiguous images, the averaged ground-truth score will be pulled toward 0.5, while MLLMs that commit to Yes or No will be penalized under SPS. Please report the frequency of 'Cannot be determined' responses for both experts and each MLLM, and analyze whether the SPS ranking is robust to excluding such cases or to treating them as a separate response category.","section":"Section 2.2, 'Cannot be determined' and Section 3.3"}],"minor_comments":[{"comment":"The manuscript contains several typographical and grammatical errors, e.g., 'We curates high-quality prompts' (abstract), 'This involve 340 unique scenarios' (Section 1), and 'generated by by GPT-4o-Image' (Section 2.2). A careful proofreading pass is needed.","section":"Throughout"},{"comment":"The generative model name is written inconsistently as 'SeeDream3.0' and 'Seedream 3.0.' Please use one consistent name.","section":"Section 2.1 and Table 2"},{"comment":"Some citations do not match their referents: reference [24] is cited for Grok-2-Vision but appears to be a paper on AI adoption, and reference [17] is cited for GLM4V but points to CogVLM2. Please verify all model-specific citations.","section":"References"},{"comment":"The differences between adjacent models in the Authenticity column are small (e.g., 31.53% vs. 31.93%), and no confidence intervals or significance tests are reported. Please add error bars or statistical comparisons before claiming a strict ranking.","section":"Table 3"},{"comment":"The statement that an inter-dimension correlation of 0.70 indicates the dimensions are 'largely independent' is overstated; a correlation of 0.7 is substantial and suggests Aesthetic Quality and Instruction Fidelity share considerable variance. Please soften the interpretation or provide a statistical test for independence.","section":"Section 3.1"}],"recommendation":"major_revision","confidential_remarks":"The qualitative conclusion that MLLMs struggle with geometric optics is plausible, but the quantitative backbone of the paper—the SPS rankings and the 37.35% headline—needs substantially more validation before the benchmark can be considered reliable. The missing inter-rater reliability and the unvalidated ground truth are the main concerns. The paper would also benefit from a discussion of how the benchmark relates to existing perception-focused MLLM benchmarks (e.g., BLINK), which are cited but not compared in detail."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my take on GOBench. The genuinely new thing is the benchmark itself: 1k images across direct light, reflection, and refraction, with a structured taxonomy and prompts designed to probe specific optical failures. That's a real gap, and the dataset is likely to be reused. The qualitative conclusion—state-of-the-art MLLMs produce physically wrong reflections, shadows, and refractions and struggle to judge them correctly—is plausible and matches what I'd expect.\n\nWhat it does well: the task design is sensible, the examples in Figures 3 and 5 are revealing, and the authors are upfront about the subjective nature of the ground truth. The correlation analysis between dimensions (Table 1) is a small but responsible touch.\n\nThe soft spots are real, and they mostly concern the quantitative claim. The abstract and Section 3.3 call 37.35% \"accuracy,\" but the SPS is a tolerance-weighted proximity to the averaged expert score, not a percentage of correct answers. That's a mislabel that will mislead readers. More seriously, there is no inter-rater reliability reported for the six experts. The stress-test point is fair: if experts disagree by more than 0.5 on a 0–5 scale, a model that matches one expert's judgment gets SPS=0 against the mean. Given there is an explicit \"Cannot be determined\" option worth 0.5 points, expert variance is likely non-trivial. Without agreement data, the 37.35% could partly reflect expert noise. The delta_max=0.5 is also hand-chosen with no sensitivity analysis. The manual filtering of generated images before scoring is described too vaguely to be reproduced.\n\nNone of this topples the core qualitative finding. Even if the SPS rankings shift, the overall picture—models well below 50% alignment with human judgment—remains. The authors should add inter-rater agreement (e.g., Fleiss' kappa per dimension), report SPS variance or confidence intervals, clarify the metric's name, and give operational filtering criteria plus a commit hash.\n\nThis paper deserves a serious referee. It's a contribution to multimodal evaluation and physical reasoning, with real data and a reusable benchmark, but the revisions matter. I'd read the revised version carefully, and I'd probably cite it as the first GOBench reference once the metric issues are fixed.","headline":"A useful first benchmark for geometric optics in MLLMs; the qualitative finding holds, but the 37.35% headline is an SPS proximity score, not accuracy, and expert-agreement data are missing.","tokens_in":13825,"tokens_out":3441,"would_cite":true,"duration_ms":31857,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A new benchmark finds that state-of-the-art multimodal models score barely above chance at judging shadows, reflections, and refraction, and still err when generating them.","keywords":["geometric optics","multimodal large language models","optical authenticity","image generation evaluation","visual understanding","shadow rendering","reflection and refraction","Scaled Proximity Score"],"falsifier":"Re-run the expert panel with a fresh set of six physicists on the same 1k images and compute inter-rater agreement; independently ray-trace the 340 prompts to produce physical ground truth. If expert scores disagree with each other or with the simulated optics, the SPS numbers and the claim that MLLMs fail at optical understanding are not established.","tokens_in":12764,"feed_emoji":"🔍","tokens_out":5062,"duration_ms":46896,"temperature":0.7,"pith_summary":"GOBench is the first benchmark to test multimodal large language models (MLLMs) on geometric optics, both as image generators and as image understanders. The paper builds GOBench-Gen-1k, a dataset of 1,000 images generated by GPT-4o-Image, Seedream 3.0, and Imagen 3 across 340 scenarios covering direct light, reflection, and refraction, and has six experts score each image for Optical Authenticity, Aesthetic Quality, and Instruction Fidelity. On the generation side, the paper reports that even the top model, GPT-4o-Image, does not perfectly complete all generation tasks and shows flaws that violate optical principles. On the understanding side, eleven MLLMs are asked the same questions as the experts, and the best model, Gemini-2.5Pro, reaches only 37.35% Scaled Proximity Score on Optical Authenticity, which the paper reads as evidence that current models lack robust physical understanding of optics.","feed_headline":"Best MLLMs score just 37% on optical understanding","feed_subtitle":"GOBench benchmark shows GPT-4o and Gemini still misjudge shadows, reflections, and refraction.","key_machinery":"The load-bearing object is the GOBench-Gen-1k dataset plus its evaluation protocol. The dataset is built from 340 curated prompts (108 direct-light, 121 reflection, 111 refraction), each with a tailored five-question Optical Authenticity rubric (Yes/No/Cannot be determined), and each scenario is rendered by GPT-4o-Image, Seedream 3.0, and Imagen 3. The metric that carries the understanding claim is the Scaled Proximity Score (SPS), which linearly credits an MLLM's rating for being within a maximum permitted difference $\\delta_{\\max} = 0.5$ of the averaged expert score per case, then averages across cases.","core_discovery":"The paper's central claim is that state-of-the-art MLLMs can generate visually appealing images yet fail to respect geometric optics, and that they cannot evaluate optical phenomena the way trained human experts do. The evidence is a benchmark in which 1k images from three generative models are scored by six human experts using five per-scenario Optical Authenticity questions based on rectilinear propagation, reflection, and Snell's law, and in which 11 MLLMs are given the same questions as a test of understanding. The reported numbers—GPT-4o-Image's authenticity score of 4.09/5, Gemini-2.5Pro's 37.35% SPS on authenticity against expert ground truth—are intended to quantify the gap between current capability and physical correctness.","pith_inferences":["If the expert gold standard were replaced by a ray-traced physical simulation of each prompt, the SPS numbers could be recomputed as a direct measure of physical accuracy, which might shift model rankings if expert judgment disagrees with the renderer.","Because the benchmark's questions are binary (Yes/No/Cannot determine), the 37% SPS may underestimate models' latent optical knowledge; a free-response or part-scored rubric might reveal partial understanding.","The same scenario design could be extended to video frames, where dynamic light behavior (e.g., moving shadows, caustics) would test whether MLLMs understand optics as a process rather than a static pictorial feature.","A testable extension is to fine-tune an open MLLM on GOBench-Gen-1k with the expert labels as supervision; if scores jump after fine-tuning, the optics gap is at least partly learnable from visual data alone."],"forward_implications":["Until models are trained with explicit physical supervision, optical generation errors such as double shadows, impossible shadow angles, and discontinuous refracted beams will persist in high-fidelity content creation.","The 37% SPS ceiling implies current MLLMs are not reliable judges of their own or others' optical realism, so automated filtering of AI-generated images for physical plausibility cannot yet be trusted.","The low correlation between Aesthetics and Authenticity (PLCC 0.2213) implies that looking good and physically right are separate axes, so improving visual appeal will not by itself fix physical realism.","Benchmark scores on GOBench can serve as a target metric for future physically grounded training objectives, giving a concrete number to optimize beyond human preference ratings."],"supporting_citations":[{"why":"Supplies the geometric optics laws (rectilinear propagation, reflection, Snell's law) that the Optical Authenticity questions encode.","marker":"[23]"},{"why":"Provides the worked optical phenomena that ground the scenario taxonomy of direct light, reflection, and refraction.","marker":"[29]"},{"why":"Prior result that MLLMs can see but not perceive, which the paper extends to geometric optics.","marker":"[12]"},{"why":"Documents the inconsistent lighting and optical errors in diffusion-based generation that motivate the generation evaluation.","marker":"[44]"},{"why":"Earlier evidence of LLM/MLLM limitations in optical tasks, extended by GOBench's understanding test.","marker":"[47]"}],"fun_headline_variants":["MLLMs fail geometric optics, scoring just 37%","AI sees images but not the physics behind them","GOBench: Top AI models get optics wrong","GPT-4o and Gemini falter on optical physics","New benchmark exposes AI's weak spot in optics"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole benchmark depends on the assumption that six experts' answers to the authors' Optical Authenticity questions are a valid and reliable ground truth for physical correctness, yet no inter-rater reliability measure is reported and the questions were not validated against an external physical ground truth such as ray-traced renders.","fun_headline_variants_meta":{"raw":{"variants":["MLLMs fail geometric optics, scoring just 37%","AI sees images but not the physics behind them","GOBench: Top AI models get optics wrong","GPT-4o and Gemini falter on optical physics","New benchmark exposes AI's weak spot in optics"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000214,"raw_usage":{"total_tokens":1420,"prompt_tokens":938,"completion_tokens":482,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":554,"completion_tokens_details":{"reasoning_tokens":406}},"tokens_in":554,"tokens_out":482,"duration_ms":4784,"temperature":1.0,"reasoning_tokens":406,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T11:52:25.050460+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the expert panel with a fresh set of six physicists on the same 1k images and compute inter-rater agreement; independently ray-trace the 340 prompts to produce physical ground truth. If expert scores disagree with each other or with the simulated optics, the SPS numbers and the claim that MLLMs fail at optical understanding are not established.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the geometric optics laws (rectilinear propagation, reflection, Snell's law) that the Optical Authenticity questions encode."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the worked optical phenomena that ground the scenario taxonomy of direct light, reflection, and refraction."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Prior result that MLLMs can see but not perceive, which the paper extends to geometric optics."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Documents the inconsistent lighting and optical errors in diffusion-based generation that motivate the generation evaluation."},{"cited_title":"LLM-Optic: Unveiling the Capabilities of Large Language Models for Universal Visual Grounding","cited_arxiv_id":"2405.17104","evidence_quote":"Earlier evidence of LLM/MLLM limitations in optical tasks, extended by GOBench's understanding test."}],"review_version":1}