{"id":"bd574066-bdbc-4cb6-8a16-483cfc8fb646","arxiv_id":"2508.04059","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"A new benchmark shows multimodal AI models consistently underperform humans at perceiving occluded objects, and the gap is not closed by model scaling or extended reasoning.","lead":"The paper introduces O-Bench, a benchmark of 1,365 synthetic images and 4,588 questions that tests how well multimodal AI models understand occluded objects. It finds that 22 current models all fall far short of human performance on these occlusion tasks, and that bigger models or more reasoning time do not close the gap.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Benchmark construct validity is the key risk: synthetic composites may let models exploit compositing artifacts, so the human-model gap may not reflect genuine occlusion perception.","rationale":"The reader's weakest assumption is benchmark construct validity, specifically that the synthetic layered scenes measure occlusion perception rather than compositing artifacts. I agree with this identification; it is the most load-bearing concern because the paper's headline contribution is the existence and persistence of a human-model gap, and every secondary finding (failure patterns, scale-insensitivity) is built on that gap. If the benchmark is solvable by low-level cues or if the models are merely sensitive to synthetic seams, the central claim is unsubstantiated. The abstract alone offers no evidence for semantic coherence or annotation reliability, and the full text is corrupted in this review, so the concern remains live. The proposed concrete test—re-rendering with Poisson blending and checking score stability—directly tests whether low-level artifacts are a confound. If the scores are unchanged, the concern is largely retired; if they change, the benchmark needs revalidation. The reader's UNVERDICTED verdict remains appropriate because the full method is unavailable and the central assumption is untested. My read does not move the verdict; it reinforces the need for the authors to provide the method and such a control before the claim can be assessed.","tokens_in":13777,"tokens_out":3864,"duration_ms":47213,"concrete_test":"Re-render a random subset of O-Bench images using Poisson blending to eliminate compositing seams, with matched illumination and shadowing so the inserted occluder appears physically coherent. Re-run the same 22 MLLMs on these re-rendered images. If average accuracy changes by more than 5% absolute, or if the human-model gap narrows substantially, the original scores are partly explained by compositing artifacts rather than occlusion perception. Additionally, have a separate set of human annotators label which images appear semantically coherent; if model errors concentrate on 'incoherent' images while human accuracy does not, that further confirms the construct-validity risk.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—a structural occlusion-perception deficit that scaling and thinking cannot bridge—depends entirely on O-Bench measuring occlusion reasoning rather than sensitivity to synthetic compositing artifacts. The abstract's 'layered synthesis' of SA-1B masks likely creates scenes where correct answers can be inferred from low-level cues: mask boundary discontinuities, lighting/seam inconsistencies, or atypical object arrangements. The 'semi-automatic workflow' is asserted to yield reliable ground truth, but no ambiguity metrics (e.g., inter-annotator agreement) or semantic-coherence validation are reported in the abstract. Because the full method is unreadable in this review, this concern cannot be dismissed. If models fail on O-Bench not because they lack occlusion perception but because they are confused by cut-and-paste artifacts, then the reported human-model gap and the 'cannot be bridged by scaling' conclusion are misattributed. This is the most load-bearing assumption: every downstream claim (the gap, its scale-insensitivity, the three failure patterns) collapses if the benchmark does not actually isolate occlusion perception.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces O-Bench, a visual question answering benchmark for occlusion perception. The benchmark is constructed from SA-1B images using a layered synthesis approach, yielding 1,365 images and 4,588 QA pairs across five tasks. The authors evaluate 22 multimodal large language models against a human baseline and report a significant performance gap that, they claim, is not sufficiently bridged by model scaling or thinking processes. They also identify three failure patterns: an overly conservative bias, fragile gestalt prediction, and difficulty with quantitative tasks. The benchmark is planned for public release.","tokens_in":13889,"tokens_out":3017,"duration_ms":39910,"significance":"If the central claim holds, O-Bench would be a valuable community resource and the finding that the human-model gap is insensitive to scale and reasoning effort is a substantive, falsifiable statement about current MLLMs. The paper's strengths include a concrete benchmark construction, a broad model suite, a human baseline, and an attempt to characterize failure modes. The claim is externally measurable and therefore testable rather than circular. The main risk is construct validity: whether O-Bench measures occlusion reasoning rather than sensitivity to artifacts of synthetic composition. The paper currently provides insufficient evidence on this point, and the strong scaling/thinking-invariance conclusion depends on this evidence.","major_comments":[{"comment":"The central claim that MLLMs have a structural deficit in occlusion perception assumes that O-Bench isolates occlusion reasoning. The abstract mentions a 'layered synthesis approach' and a 'reliable, semi-automatic workflow,' but no evidence is provided that models cannot solve the task using low-level compositing artifacts (e.g., seams, lighting mismatches, or atypical object arrangements). The manuscript should report a control analysis: for example, compare performance on composite images against the same images with the occluder removed, or ablate the synthesis pipeline, and demonstrate that errors track occlusion complexity rather than artifact presence. Without this, the reported human-model gap could misattribute failure to occlusion perception.","section":"Abstract; Benchmark Construction section"},{"comment":"The abstract states there is a 'significant performance gap' between MLLMs and humans, but no quantitative results, confidence intervals, or inter-annotator agreement are reported in the summary. The paper needs to provide the human baseline size, per-question agreement (e.g., Fleiss' kappa), and statistical testing (e.g., paired bootstrap or Wilcoxon) to establish that the gap is not within annotation noise or due to ambiguous QA pairs. This is load-bearing because the headline conclusion depends on the gap being real and attributable to occlusion perception.","section":"Abstract; Experiments section"},{"comment":"The claim that the gap 'cannot be sufficiently bridged by model scaling or thinking process' is a strong negative result. The manuscript should specify the exact model sizes, the definition of 'thinking process' (e.g., chain-of-thought prompting, self-consistency, or inference-time compute), and the criterion for 'sufficiently.' A trend plot across at least three distinct model families with monotonically increasing scale, and an apples-to-apples comparison with and without reasoning prompts, is needed. The current abstract-level description does not rule out that a different prompting strategy or larger model would close the gap.","section":"Abstract; Model scaling and thinking analysis"},{"comment":"Because the benchmark is built from SA-1B, which may appear in MLLM training corpora, the paper should address potential contamination. The authors should report near-duplicate image search against common training sets, or demonstrate that models do not perform anomalously well on any subset, and describe any test-set access restrictions. If templated QA generation is used, a leakage check for the exact question templates is also necessary. This concern is central to interpreting the absolute performance numbers.","section":"Evaluation protocol; Benchmark Construction"}],"minor_comments":[{"comment":"The terms 'overly conservative bias' and 'fragile gestalt prediction' are used without operational definitions. Please provide concrete examples and quantitative criteria for each failure pattern.","section":"Abstract"},{"comment":"Clarify the distribution of the 4,588 QA pairs across the five tasks and whether the human baseline and model evaluations are balanced per task. Reporting per-task accuracy with confidence intervals would help readers assess whether the overall gap is driven by one task.","section":"Benchmark tasks"},{"comment":"The manuscript text supplied for review is heavily garbled in places, which prevents independent verification of the methods and tables. Please ensure a readable version is provided with the submission.","section":"Full text"},{"comment":"The limitations section should explicitly acknowledge the synthetic-to-real transfer risk and state whether real-world occlusion images were used for validation. This is closely tied to the construct-validity concern and should be addressed in the paper body, not only a future-work sentence.","section":"Limitations"}],"recommendation":"major_revision","confidential_remarks":"The core idea is timely and the benchmark could be a useful contribution, but the provided text is partially unreadable and the key methodological safeguards are not visible. I recommend major revision: the authors should supply the missing statistical and control analyses, clarify the construct-validity evidence, and make the full methods readable. If the missing analyses were already present in the original PDF, then the rating might improve, but I cannot verify from the supplied copy."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things you should know about this one. The abstract describes O-Bench, an occlusion-perception VQA benchmark built from SA-1B with layered synthesis, 1,365 images, 4,588 QA pairs, five tasks, 22 MLLMs, and a human baseline. The headline result is that models trail humans and scaling/thinking time doesn't close the gap. That's a genuinely useful negative result for anyone picking models for spatial tasks. The second thing: I couldn't read the paper. The full-text extraction is mojibake. So my take is on the abstract and the reader's notes, not the methods.\n\nWhat's good: the idea of a dedicated occlusion benchmark with a human baseline is solid. Layered synthesis from SA-1B is a plausible way to make controlled occlusion scenarios, and 22 models is a broad sweep. The three failure patterns (conservative bias, fragile gestalt, quantitative struggle) are the kind of qualitative analysis that makes a benchmark usable.\n\nWhere it gets shaky: the central claim is that the gap reflects occlusion perception, not sensitivity to synthetic compositing artifacts. That's a construct-validity burden. The abstract says the scenes are 'semantically coherent' and the workflow 'reliable, semi-automatic,' but there are no numbers to back that up—no inter-annotator agreement, no leakage checks, no error bars on the human baseline. The 'cannot be bridged by scaling' conclusion is strong and needs to be robust across model families and prompt variations. And 'first' is unverifiable here because the related work is unreadable.\n\nIs this fatal? Not on its own. The risk is real, but it's the kind of thing a good referee can probe. The paper deserves a serious referee. I'd send it to review. The benchmark is useful and the negative result is worth having on record. But the review should require evidence that the benchmark isolates occlusion reasoning: artifact analysis, human agreement, and ideally control images with the same construction but no occlusion.\n\nFor your purposes: if you work on MLLM evaluation, this is worth citing once it's out, and worth reading carefully once a readable version exists. Not the last word, but a reasonable tool.","headline":"O-Bench is a plausible and useful occlusion-perception benchmark with a strong human baseline, but the full text is unreadable in this copy and the central negative result rests on construct validity we cannot check.","tokens_in":14483,"tokens_out":2029,"would_cite":true,"duration_ms":24309,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper introduces O-Bench, a visual question answering benchmark showing that twenty-two multimodal large language models all underperform humans at occlusion perception, and that the gap persists with scaling and longer reasoning.","keywords":["occlusion perception","multimodal large language models","O-Bench benchmark","visual question answering","layered image synthesis","SA-1B","human baseline comparison","failure mode analysis"],"falsifier":"On a sample of O-Bench items, replace the synthetic occluder with a real photograph of an occluding object matched for lighting and edge statistics, keeping the question unchanged. If model accuracy rises substantially on the same items, the reported gap is partly an artifact of low-level compositing differences rather than occlusion perception.","tokens_in":13570,"feed_emoji":"👁","tokens_out":4186,"duration_ms":49297,"temperature":0.7,"pith_summary":"This paper introduces O-Bench, a visual question answering benchmark built to test whether multimodal large language models can perceive and reason about partially hidden objects. Using 1,365 images composed from SA-1B segment masks, the authors created 4,588 question-answer pairs across five occlusion-specific tasks, and scored 22 MLLMs against a human baseline. The central finding is that every tested model falls short of humans, and the shortfall does not shrink when models are scaled up or given more time to reason. The paper argues this indicates a structural limitation in how current models represent occluded scenes, and identifies three recurring failure modes: an overly conservative bias, fragile gestalt completion, and poor performance on quantitative occlusion questions. If the benchmark is valid, it gives the field a targeted evaluation tool and a concrete agenda for improving occlusion perception.","feed_headline":"22 MLLMs fail occlusion tests humans pass","feed_subtitle":"A 4,588-question benchmark finds the gap holds even with bigger models and longer reasoning.","key_machinery":"The central object is the layered synthesis pipeline: the benchmark builds occlusion scenes by taking object segment masks from SA-1B and compositing them into semantically coherent arrangements, so that each image has known ground-truth occlusion structure. This provides controllable, minimally ambiguous occlusion scenarios whose correct answers require genuine occlusion reasoning. The five task types and the semi-automatic annotation workflow turn these composites into a VQA benchmark with 4,588 question-answer pairs.","core_discovery":"The paper's central claim is that occlusion perception—the ability to infer the presence, shape, and layout of objects that are partially hidden—is a distinct visual capability in which current multimodal large language models are systematically deficient. The authors report that across 22 representative models, performance on O-Bench is substantially below the human baseline, and that this gap persists across model scales and with extended reasoning or 'thinking process.' They attribute this to a fundamental limitation in models' visual reasoning rather than to insufficient compute, and they characterize the failures as an overly conservative bias, fragile gestalt prediction (incomplete or","pith_inferences":["The paper leaves open whether the gap is caused by the visual encoder's lack of amodal representations; a natural next test is whether models trained with explicit amodal segmentation or 3D depth supervision close the gap on O-Bench.","One way to stress-test the benchmark's construct validity would be to compare model performance on questions about the same scene with and without the occluder: if accuracy drops only when occlusion is present, the task isolates occlusion reasoning.","The conservative bias pattern suggests models may default to 'cannot tell' when objects are partially hidden; calibration studies could quantify this and possibly lead to better answer strategies.","O-Bench could be extended to dynamic occlusion (video) or to active perception where a model chooses a new viewpoint, testing whether the limitation is in static inference or in requesting more information."],"forward_implications":["Current multimodal large language models, including large and reasoning-augmented models, cannot match human accuracy on occlusion perception tasks.","The gap is not a compute or scale effect, so simply training larger models or adding more inference-time thinking will not close it.","Three failure modes—overly conservative bias, fragile gestalt prediction, and weak quantitative judgments—can serve as targeted targets for model improvement.","O-Bench provides a reusable, public evaluation tool for future work on occlusion-aware visual reasoning."],"supporting_citations":[],"fun_headline_variants":["O-Bench: MLLMs lag humans on occlusion perception","Occlusion perception gap: 22 MLLMs vs humans on O-Bench","Scaling doesn't close occlusion gap in MLLMs","MLLMs fail hidden-object tests humans ace","O-Bench: 4,588 Q&A pairs expose MLLM occlusion blind spots"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The benchmark's claim rests on assuming its composited images really require occlusion reasoning, not just guessing from seams, lighting mismatches, or question wording.","fun_headline_variants_meta":{"raw":{"variants":["O-Bench: MLLMs lag humans on occlusion perception","Occlusion perception gap: 22 MLLMs vs humans on O-Bench","Scaling doesn't close occlusion gap in MLLMs","MLLMs fail hidden-object tests humans ace","O-Bench: 4,588 Q&A pairs expose MLLM occlusion blind spots"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000922,"raw_usage":{"total_tokens":3779,"prompt_tokens":722,"completion_tokens":3057,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":466,"completion_tokens_details":{"reasoning_tokens":2964}},"tokens_in":466,"tokens_out":3057,"duration_ms":24568,"temperature":1.0,"reasoning_tokens":2964,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T00:54:55.768748+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"On a sample of O-Bench items, replace the synthetic occluder with a real photograph of an occluding object matched for lighting and edge statistics, keeping the question unchanged. If model accuracy rises substantially on the same items, the reported gap is partly an artifact of low-level compositing differences rather than occlusion perception.","supporting_citations":[],"review_version":1}