{"id":"b500eb65-0b6d-484f-867b-2961f7ff9892","arxiv_id":"2605.09883","paper_version":2,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Reformulating 53 visual reasoning tasks in polar coordinates causes frontier MLLMs to drop from 70-83% to 31-39% accuracy while preserving logical equivalence, revealing a Cartesian shortcut in current benchmarks.","lead":"Multimodal large language models exploit Cartesian grid layouts in visual benchmarks by using textual coordinates for reasoning. Converting tasks to polar coordinates causes their performance to plummet from 70-83% to 31-39%, exposing reliance on shortcuts rather than genuine visual understanding.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"Polar reformulation may introduce new perceptual or transformation difficulties not controlled for in the 53 tasks.","rationale":"The reader's weakest_assumption directly identifies the equivalence of the reformulation as the load-bearing point. The full-text placeholder does not alter this; without explicit controls (human baselines, transformation ablations, or per-task equivalence proofs), the central empirical contrast remains conditional on an unverified invariance claim. No stronger internal inconsistency appears from the given claims.","tokens_in":1679,"tokens_out":313,"duration_ms":14440,"concrete_test":"Collect human accuracy on a matched subset of 10 Cartesian vs. Polar task pairs; if humans show >15% drop on Polar versions while models show 40%+, the equivalence assumption is undermined and the headline degradation cannot be isolated to the Cartesian shortcut.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The strongest claim requires that Polaris-Bench versions preserve identical logical constraints and task semantics. Converting visual reasoning problems (e.g., counting, spatial relations, path finding) to polar coordinates necessarily changes how distances, angles, and alignments are represented and perceived in the image plane. Even if textual descriptions are adjusted for equivalence, the visual input itself may impose additional demands on coordinate conversion, curvature handling, or discretization that Cartesian grids avoid. The abstract states preservation but provides no quantitative check (e.g., human accuracy parity or ablation on transformation complexity) that would confirm the performance drop is attributable solely to removal of the orthogonal shortcut rather than these side effects.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper claims that MLLMs exploit a pervasive 'Cartesian Shortcut' in visual reasoning benchmarks that rely on orthogonal grid layouts, allowing models to discretize inputs into explicit textual coordinates and bypass robust visual understanding. To expose this, the authors introduce Polaris-Bench, which reformulates 53 tasks into polar coordinate space (with paired Cartesian references) while asserting preservation of logical constraints and task semantics; evaluation of 14 frontier MLLMs shows accuracy collapsing from 70-83% on Cartesian versions to 31-39% on polar equivalents, with diminished reasoning gains, indicating a lack of topology-invariant visual reasoning.","tokens_in":1819,"tokens_out":468,"duration_ms":17736,"significance":"If the equivalence of logical constraints holds, the result would be significant for the field: it would demonstrate that high benchmark scores on canonical visual reasoning tasks do not reflect genuine visual or topological understanding, but rather exploitation of grid-based priors. The scale of the evaluation (14 models, 53 tasks) and the introduction of a new benchmark with explicit Cartesian-polar pairs would provide a concrete, falsifiable test of shortcut reliance, potentially influencing future benchmark design and training objectives.","major_comments":[{"comment":"Abstract: the central claim that performance collapse is attributable to removal of the orthogonal shortcut rests on the assertion that polar reformulations 'preserve consistent logical constraints and task semantics'; however, no quantitative controls (human accuracy parity, ablation on transformation complexity, or discretization effects) are referenced to rule out new perceptual or representational difficulties introduced by polar coordinates.","section":"Abstract"},{"comment":"The abstract reports clear performance drops across 14 models but supplies no implementation details, verification of logical equivalence between paired tasks, error bars, or controls for task difficulty; this information is load-bearing for interpreting whether the 31-39% polar scores reflect the claimed deficiency rather than uncontrolled side effects of the coordinate transformation.","section":"Abstract"}],"minor_comments":[{"comment":"The invented term 'Cartesian Shortcut' is used without an early formal definition or pseudocode for how the shortcut is operationalized in the 53 tasks.","section":null}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the detailed and constructive feedback. We address the two major comments on the abstract below, clarifying the basis for our claims while committing to revisions that add requested controls and details.","responses":[{"response":"Logical equivalence is maintained by direct, rule-preserving reformulation: each polar task is constructed from its Cartesian pair by remapping only the coordinate representation while keeping identical logical constraints, objectives, and solution paths (detailed with examples in Section 3). We agree that explicit quantitative validation would further rule out transformation artifacts. In revision we will add (i) human accuracy parity results on a 10-task subset and (ii) an ablation examining discretization granularity and transformation complexity.","revision_made":"yes","referee_comment":"[Abstract] Abstract: the central claim that performance collapse is attributable to removal of the orthogonal shortcut rests on the assertion that polar reformulations 'preserve consistent logical constraints and task semantics'; however, no quantitative controls (human accuracy parity, ablation on transformation complexity, or discretization effects) are referenced to rule out new perceptual or representational difficulties introduced by polar coordinates."},{"response":"Space limits prevent the abstract from containing these elements; they appear in the full manuscript (model implementation and evaluation protocol in Section 4; paired-task equivalence verification and difficulty matching via Cartesian-polar design in Section 3). We will incorporate error bars on all reported accuracies and add an explicit discussion of potential side effects in the revised abstract and main text.","revision_made":"partial","referee_comment":"[Abstract] The abstract reports clear performance drops across 14 models but supplies no implementation details, verification of logical equivalence between paired tasks, error bars, or controls for task difficulty; this information is load-bearing for interpreting whether the 31-39% polar scores reflect the claimed deficiency rather than uncontrolled side effects of the coordinate transformation."}],"tokens_in":1394,"tokens_out":405,"duration_ms":24340,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The paper's key observation is that 14 MLLMs drop from 70-83% accuracy on standard visual reasoning tasks to 31-39% when the same 53 tasks are rewritten in polar coordinates. The authors tie this to models exploiting the orthogonal grid structure that Cartesian layouts make easy to turn into text coordinates.\n\nPolaris-Bench itself is the concrete new piece. Creating paired Cartesian and polar versions of the tasks gives a direct way to measure how much performance depends on that grid property. The evaluation is broad enough to show the pattern holds across frontier models and that reasoning gains seen on Cartesian versions largely vanish on the polar ones.\n\nThe soft spot is the claim that the polar versions preserve identical logical constraints and task semantics. Shifting to polar coordinates changes how distances, angles, and alignments are rendered in the image. That shift can add demands around curvature, discretization, or visual parsing that Cartesian grids avoid. The abstract states the versions are equivalent but gives no human accuracy parity check, no ablation on transformation complexity, and no other control that would isolate the shortcut as the only cause. Without those, part of the drop could come from the reformulation itself.\n\nThis work is for people who build or evaluate multimodal models and care about whether benchmarks actually test visual reasoning. Anyone running robustness tests or designing new visual tasks would get direct use from the paired dataset.\n\nIt deserves peer review. The empirical pattern is new and the question matters for how we measure MLLM capabilities, even if the current version needs added controls to pin down the cause.","headline":"The performance collapse on polar reformulations is the main result, but the paper still needs to show those versions are truly equivalent in difficulty.","tokens_in":2303,"tokens_out":389,"would_cite":false,"duration_ms":20800,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Models achieve high visual reasoning scores by discretizing grid images into text but collapse when the same tasks use polar coordinates.","keywords":["visual reasoning","multimodal large language models","Cartesian shortcut","polar coordinates","benchmarks","spatial topology","coordinate invariance"],"falsifier":"A controlled test in which models reach similar accuracy on the polar and Cartesian versions even when image-to-text coordinate conversion is blocked would falsify the shortcut claim.","tokens_in":2597,"feed_emoji":"🌀","tokens_out":608,"duration_ms":22897,"temperature":0.7,"pith_summary":"The paper argues that current multimodal models exploit an orthogonal grid prior in benchmarks, converting layouts into explicit textual coordinates for deductive reasoning instead of performing visual analysis. To test this, the authors create paired Cartesian and polar versions of 53 tasks that keep identical logic and semantics. Frontier models that reach 70-83 percent accuracy on the Cartesian versions fall to 31-39 percent on the polar versions, and any reasoning gains seen on grids largely disappear. The result indicates that these models lack visual reasoning that remains stable across different spatial topologies.","feed_headline":"Vision models drop from 80% to 35% when grids turn polar","feed_subtitle":"Reformulating 53 tasks in polar space shows models rely on text discretization of orthogonal layouts rather than visual reasoning.","key_machinery":"Polaris-Bench, a collection of 53 visual reasoning tasks reformulated in polar coordinate space together with their Cartesian counterparts, engineered to eliminate the orthogonal grid structure that models exploit.","core_discovery":"The Cartesian Shortcut allows models to bypass genuine visual processing by turning grid-based images into discrete text coordinates; when tasks are instead expressed in polar space while preserving every logical constraint, performance of leading MLLMs drops sharply and stays low, showing that their reasoning is not invariant to coordinate topology.","pith_inferences":["The same shortcut could appear in other non-grid representations such as spherical or cylindrical coordinates used in 3D or panoramic data.","Applications that naturally produce radial or angular observations, including certain robotics or medical imaging settings, may reveal comparable weaknesses.","Augmenting training data with polar-transformed versions of existing tasks could encourage more invariant internal representations."],"forward_implications":["High scores on standard Cartesian benchmarks do not demonstrate robust visual understanding.","Reasoning improvements measured on grid-based tasks largely fail to transfer when the spatial representation changes.","Current models cannot reliably solve the same logical problems once the underlying topology is altered.","Development of topology-invariant visual reasoning is required before benchmark saturation can be treated as genuine progress."],"fun_headline_variants":["Text shortcuts explain high Cartesian vision scores","MLLMs lack topology invariant visual reasoning","Cartesian grids turn vision into text deduction","Polar reformulation removes grid based advantages"],"cache_read_input_tokens":64,"weakest_assumption_plain":"Converting the original tasks into polar coordinates keeps the logical constraints and task meaning exactly the same and adds no extra visual or reasoning difficulty.","fun_headline_variants_meta":{"raw":{"variants":["Text shortcuts explain high Cartesian vision scores","MLLMs lack topology invariant visual reasoning","Cartesian grids turn vision into text deduction","Polar reformulation removes grid based advantages"]},"model":"grok-4.3","cost_usd":0.009801,"raw_usage":{"total_tokens":4341,"prompt_tokens":626,"num_sources_used":0,"completion_tokens":50,"cost_in_usd_ticks":98012000,"prompt_tokens_details":{"text_tokens":626,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":3665,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":626,"tokens_out":50,"duration_ms":28353,"temperature":1.0,"reasoning_tokens":3665,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-30T22:55:46.332649+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A controlled test in which models reach similar accuracy on the polar and Cartesian versions even when image-to-text coordinate conversion is blocked would falsify the shortcut claim.","supporting_citations":[],"review_version":2}