{"id":"99ea4b4c-853e-4af8-b5f6-65af0e82ea7a","arxiv_id":"2602.15892","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Most vision-language models fail Level-2 visual perspective taking: they report the camera's view rather than the 180°-rotated string, even though they often recognize that another agent sees differently.","lead":"This paper introduces FlipSet, a benchmark that asks vision-language models what a character string looks like when viewed by a monkey on the opposite side of a card. Across 103 models, most answer with the camera-view string, exposing a systematic egocentric bias and an apparent gap between knowing another agent sees differently and being able to compute what they see.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The stated scene geometry (upright opaque card, monkey facing the blank back) invalidates the ground-truth labels; every accuracy result depends on the monkey seeing the rotated string, which the described setup does not allow.","rationale":"The reader's weakest assumption—that the benchmark's ground truth presupposes the monkey sees the rotated string—is exactly the most load-bearing concern. I independently derived the same inconsistency from the Methods section: an upright, opaque A4 card with printing on the camera-facing side presents a blank back to the monkey. Because the paper does not say the card is transparent or double-sided, or that it lies flat on a table, the ground truth '18' has no basis in the described scene. All headline numbers (91.3% below chance, 75.88% egocentric errors, ToM=90.4%, MR=26.1%, L2=10.3%, compositional deficit) are computed against these labels; if the labels are wrong, the entire empirical edifice collapses. I do not think this warrants outright rejection because the fix is straightforward: release the stimuli or correct the geometry description. The conditional verdict is appropriate. The product-rule baseline is a separate weakness, but it is secondary and also addressable. I agree with the reader that the raw observation—VLMs often answer from the camera's viewpoint—may survive a corrected setup, but it cannot be credited on the current text.","tokens_in":12807,"tokens_out":6220,"duration_ms":57050,"concrete_test":"Release a random sample of at least 10 original FlipSet images (or a precise 3D render of the scene). Independently inspect the card's orientation and opacity. If the card is opaque and upright with printing only on the camera-facing side, the monkey's view is blank and all ground-truth labels are invalid; re-evaluate the entire benchmark with corrected labels or a changed physical setup. If the card is transparent or lying flat, request a corrected Methods description and verify that the rotation is a 180° in-plane rotation; then re-run the main statistics on a subset to confirm they are unchanged.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The benchmark's correctness labels and all derived statistics depend on the assumption that the plush monkey sees the 180°-rotated string (e.g., '18' for '81'). The Methods section states the card is 'placed upright on a wooden floor' and that the monkey 'sits on the opposite side, facing the card's back.' A white A4 card is opaque; its back is blank. Under the described geometry, the monkey sees no characters, so none of the four options is correct. For the labels to hold, the card would have to be transparent, printed on both sides, or lying flat with the monkey across the table (the six/nine paradigm in Zhao et al. 2016); none of this is specified. This is not a wording nit: the main result (91.3% below chance, 75.88% egocentric) and the control tasks (ToM, MR, L2 VPT) all compare model outputs to these labels. If the monkey sees a blank back, a model answering '81' is reporting the only visible content, not necessarily exhibiting egocentric bias; and the ToM item, which asks whether the monkey sees a different string than in the image, would have ground truth 'No' rather than 'Yes,' so the reported 90.4% ToM accuracy would be uninterpretable. The paper neither releases the stimulus images nor provides a diagram of the actual layout, so a reader cannot resolve this ambiguity. This is the single most load-bearing concern because it threatens the validity of every quantitative claim in the paper. A secondary issue is the asserted product-rule baseline (L2 = ToM × MR), which lacks independent justification, but it is moot if the labels are wrong.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces FlipSet, a benchmark for Level-2 visual perspective taking (L2 VPT) in vision-language models. Each item shows a card with a character string and a plush monkey on the opposite side; the model must choose what the monkey sees. The authors report that 91.3% of 103 VLMs perform below 25% chance, with roughly 75.88% of errors being egocentric (outputting the camera viewpoint). Control experiments on 24 models separate theory of mind (ToM), mental rotation (MR), and L2 VPT, reporting high ToM accuracy, near-chance MR, and catastrophically low L2 VPT. They further claim a compositional deficit: L2 VPT is below the product ToM × MR. The paper frames these results as evidence that VLMs lack mechanisms for integrating social awareness with spatial transformation.","tokens_in":13121,"tokens_out":8953,"duration_ms":81097,"significance":"If the benchmark and its labels are valid, this is a valuable large-scale diagnostic result with clear implications for model architecture: it would demonstrate a systematic and specific failure in situated social reasoning. The scale (103 models), the error-type taxonomy, and the separation of ToM and MR are strong points. However, the entire quantitative edifice rests on the stimulus geometry producing the ground-truth 180° rotations, and the compositional-deficit conclusion relies on an unvalidated product-rule baseline. Both are load-bearing and currently unsupported.","major_comments":[{"comment":"The described stimulus geometry is inconsistent with the ground-truth labels. The text says the card is 'placed upright on a wooden floor' and the monkey sits 'on the opposite side, facing the card's back.' An upright opaque card shows the monkey a blank back; a transparent card would show a mirror reflection (e.g., 'd' would appear 'b', not 'p'). The 180° rotations in Table 1 (d→p, nod→pou) correspond to the six/nine paradigm where the card lies flat between the two agents, not to the described vertical-card layout. Since all accuracy, egocentric-error, and ToM/MR/L2 statistics are computed against these labels, the benchmark's validity is at stake. The paper releases neither stimuli images nor a precise diagram. Please correct the geometry, provide the actual layout, or re-run the evaluation under the intended setup.","section":"Methods (Main Experiments)"},{"comment":"The compositional-deficit claim rests on the asserted product rule L2 = ToM × MR. This rule is not derived and is not a neutral baseline: the L2 task is four-way multiple choice while the ToM task is binary, and L2 adds the demand of recognizing which option corresponds to the transformed string. Any such added task demand will automatically produce L2 < ToM × MR even without an integration deficit. For example, with ToM = 1.0 and MR = 0.505, the product is 0.505, but a 0.339 L2 score is not evidence of a binding failure unless the product rule is justified or replaced with a matched-task baseline. Please provide a formal task model or empirical calibration; otherwise the 'deficit' is an artifact of task design.","section":"Control Results / Figure 3"}],"minor_comments":[{"comment":"The ToM task is described as 'recognizing that another agent's view differs', but the question is actually a Level-1 visibility judgment ('Is the monkey seeing a different string...?'). This conflation should be acknowledged more explicitly, as it may overstate what the high ToM accuracy measures.","section":"Control Experiment"},{"comment":"The claim that MR performance is 'above chance' (mean 26.1% vs 25%) is not statistically supported. With only 28 items per model per task, the standard error for a single model is about 8 percentage points; a t-test or confidence interval across the 24 models is needed before concluding MR is above chance.","section":"Results / Figure 3"},{"comment":"The selection criterion for the 24 models used in the control experiments is not described. Since these models drive the compositional-deficit analysis, a clear sampling procedure or a justification for the subset is needed to rule out selection bias.","section":"Model Evaluations"},{"comment":"Figure A1 shows fluctuations in confusable errors of up to 12 percentage points across answer layouts. The claim that 'answer position has limited influence' is supported only by visual inspection; a statistical test (e.g., ANOVA or chi-square) would strengthen the claim. Also, the main text contains typos such as 'In this respectm' in the Introduction.","section":"Appendix A1"}],"recommendation":"major_revision","confidential_remarks":"The geometry issue is severe: if the actual stimuli images do not match the intended 180°-rotation layout, the paper is not salvageable without re-running the experiments. I recommend the editors require the authors to submit the actual stimulus images or a precise experimental diagram before any further consideration, and to re-analyze the data if the layout differs from the text."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The core observation here is probably right: when VLMs are asked what another agent sees, they overwhelmingly repeat the camera's view. That is a useful, likely-robust empirical finding. FlipSet's 2D-rotation design and the five-way error taxonomy (correct/egocentric/confusable/random/fail) are good ideas, and evaluating 103 models with ToM/MR/L2 control tasks is a substantial step. The 75.88% egocentric error rate is vivid and plausible.\n\nThe problem is that the benchmark as described does not support the labels. The Methods say the card is \"placed upright on a wooden floor\" and the monkey \"sits on the opposite side, facing the card's back.\" A white, upright, opaque A4 card has a blank back. Under that geometry, the monkey sees no characters, so none of the four options is correct; the ToM item's ground truth is \"No\" rather than \"Yes\"; and every accuracy and error-type statistic derived from these labels collapses. The classic six/nine paradigm they cite works because the card lies flat on a table and the other viewer sees the same printed characters inverted 180 degrees. The paper never says the card is horizontal, double-sided, or transparent, and no stimulus images are released. This is not a minor wording issue; it is load-bearing.\n\nA second soft spot is the \"compositional deficit\" conclusion. The claim that L2 VPT should equal ToM × MR is asserted, not justified. Any additional task cost—prompt comprehension, response formatting, the extra step of binding a social judgment to a spatial transformation—automatically produces a \"deficit.\" The MR–L2 correlation (r=0.75) is real evidence that mental rotation contributes, but the binding-integrative interpretation is stronger than the product rule supports.\n\nAlso, no dataset, code, or evaluation logs are released, so the 103-model numbers cannot be independently checked.\n\nI would still engage with the work. The egocentric-bias signal is likely real, and the dissociative control design is clever. But the current manuscript is not acceptable as-is. It needs a clear correction/diagram of the actual scene geometry, public stimuli and code, and a defended or replaced product baseline. If the intended geometry is a flat card, that changes the task description substantially; if the card really is upright and opaque, then the results are unmeasurable from the description.\n\nSend it to peer review—seriously—but expect major revision before it is publishable. I might bring it to reading group as a case study in how a one-sentence geometry detail can sink a large empirical study.","headline":"Promising diagnostic design and a likely-real egocentric-bias finding, but the stated physical setup would keep the monkey from seeing any characters—invalidating the labels as described.","tokens_in":13698,"tokens_out":4763,"would_cite":false,"duration_ms":49419,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"FlipSet, a 28-item benchmark that asks vision-language models what a monkey sees on the back of a rotated card, finds that 91.3% of 103 models score below the 25% chance level and that 75.9% of errors merely copy the camera's viewpoint.","keywords":["visual perspective taking","egocentric bias","vision-language models","mental rotation","theory of mind","benchmark","Level-2 VPT","compositional deficit"],"falsifier":"Run the same FlipSet battery with the character string printed on both sides of the card, so the monkey genuinely sees the rotated characters. If model accuracy jumps from near 10% to at or above chance, the original opaque-card geometry—not a compositional deficit—explains the failure; if accuracy stays near chance, the integration deficit is genuine. A complementary check is to ask human raters what the monkey sees in the original images; if they say the back is blank, the ground-truth labels are invalid.","tokens_in":12639,"feed_emoji":"🐵","tokens_out":6618,"duration_ms":61680,"temperature":0.7,"pith_summary":"The paper tries to establish that current vision-language models systematically fail at Level-2 visual perspective taking—inferring how a symbol appears from another agent's viewpoint—and that the failure is not simply missing component skills. It introduces FlipSet, a 28-item benchmark that asks a model what a plush monkey sees on the back of a card bearing a 2D character string, requiring a 180-degree mental rotation. Across 103 models, 91.3% score below the 25% chance level, and 75.88% of all errors reproduce the camera's own view, a pattern the paper calls egocentric bias. Control experiments on 24 models find high theory-of-mind accuracy (90.4%), near-chance isolated mental rotation (26.1%), and catastrophic full-task performance (10.3%), with L2 performance below the product of the component scores for 22 of 24 models. A sympathetic reader would care because the result suggests that today's multimodal systems can recognize that others see differently and can rotate shapes in isolation, yet cannot bind the two operations together—a specific architectural gap rather than a generic weakness.","feed_headline":"Vision-language models flunk a monkey's-eye-view test","feed_subtitle":"A 103-model benchmark: 91% score below chance, and 76% of errors just copy the camera's view.","key_machinery":"FlipSet's central device is a controlled multiple-choice item: an upright white card showing a 2D string such as '81', with a plush monkey on the opposite side facing the card's back. The question 'What does the monkey see on the card?' requires mentally rotating the string 180° (to '18'). Every item's four options are designed to diagnose the failure mode: correct perspective-transformed answer, egocentric camera-view answer, a contour-confusable distractor, and an unrelated random distractor, with 12 counterbalanced layouts per item to remove position bias. A companion control set reuses the same images under three prompts—theory-of-mind visibility judgment, pure mental rotation, and full","core_discovery":"On the paper's own terms, FlipSet shows that 91.3% of 103 vision-language models perform below the 25% random baseline on a task that only requires reading a short character string and rotating it 180 degrees from a monkey's viewpoint. Mean accuracy is 8.96%, median 5.36%, and egocentric responses account for 75.88% of all answers. In the control battery, the same models average 90.4% on theory-of-mind recognition (does the monkey see a different string?), 26.1% on isolated mental rotation (what does the string become under 180 degrees?), and 10.3% on the full L2 task. The authors interpret the gap between the product of the component scores and the observed L2 accuracy—a deficit present in","pith_inferences":["Beyond the paper: the ground-truth labels assume the monkey sees the rotated string, but the described setup—an upright, opaque card with characters printed on the camera-facing side—would show the monkey a blank back. Re-rendering the stimuli with the string visible on both sides (or a transparent card) would test whether the reported egocentric bias is partly a visual-geometry artifact.","Beyond the paper: the product-rule baseline treats theory of mind and mental rotation as independent and free to combine. Because the full L2 task adds prompt comprehension and coordination demands, any real task cost will automatically read as a 'compositional deficit'; adding a two-step control task would calibrate this baseline.","Beyond the paper: the benchmark only uses 180-degree rotations. Extending FlipSet to 90- and 270-degree rotations would test whether model error scales with rotation angle the way human response time does, linking the result to classical mental-rotation findings.","Beyond the paper: the egocentric bias may partly reflect training statistics, since front-view text is far more common than rotated text. Fine-tuning on multi-view or egocentric-to-allocentric data is a testable intervention that, if it reduces egocentric errors, would support a data-driven rather than architectural explanation."],"forward_implications":["Chain-of-thought prompting does not fix the egocentric bias and often amplifies it, implying the limitation is not a lack of verbal reasoning steps.","Models with near-perfect theory-of-mind scores and above-chance mental rotation still fail the integrated task, so the two component skills do not automatically compose in current architectures.","Mental rotation accuracy correlates strongly with L2 accuracy (r = 0.746) while theory of mind does not (r = 0.010), pointing to spatial transformation as the bottleneck skill.","FlipSet's 28-item, zero-shot protocol with counterbalanced answer positions offers a reusable diagnostic for tracking perspective-taking progress in future vision-language models."],"fun_headline_variants":["91% of vision-language models flunk a simple viewpoint swap","VLMs can't see from a monkey's eyes — egocentric bias rules","Theory of mind yes, but mental rotation? VLMs crash when combined","FlipSet: VLMs below chance on 180-degree perspective shift"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The claims rest on two assumptions: that the monkey actually sees the rotated character string, and that expected L2 performance equals theory-of-mind accuracy multiplied by mental-rotation accuracy; if either is wrong, the egocentric-bias and compositional-deficit conclusions lose their footing.","fun_headline_variants_meta":{"raw":{"variants":["91% of vision-language models flunk a simple viewpoint swap","VLMs can't see from a monkey's eyes — egocentric bias rules","Theory of mind yes, but mental rotation? VLMs crash when combined","FlipSet: VLMs below chance on 180-degree perspective shift"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000281,"raw_usage":{"total_tokens":1483,"prompt_tokens":710,"completion_tokens":773,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":454,"completion_tokens_details":{"reasoning_tokens":696}},"tokens_in":454,"tokens_out":773,"duration_ms":7562,"temperature":1.0,"reasoning_tokens":696,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T02:54:09.119552+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same FlipSet battery with the character string printed on both sides of the card, so the monkey genuinely sees the rotated characters. If model accuracy jumps from near 10% to at or above chance, the original opaque-card geometry—not a compositional deficit—explains the failure; if accuracy stays near chance, the integration deficit is genuine. A complementary check is to ask human raters what the monkey sees in the original images; if they say the back is blank, the ground-truth labels are invalid.","supporting_citations":[],"review_version":1}