{"id":"3968d842-6a7c-41ec-b764-a12d6d7ae849","arxiv_id":"2606.08031","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":2,"one_line_summary":"Behavioral tests and SAE probing on 83 bistable images show simultaneous vision-tower activation of both aspects in 72% of cases, with causal steering succeeding on default-dominant but not force-balanced stimuli, locating the commitment bottleneck downstream of the vision tower.","lead":"The paper tests how vision-language models caption ambiguous images like the duck-rabbit illusion using thousands of generations and a sparse autoencoder on the vision encoder of LLaVA-1.6-7B. It finds that both interpretations activate together in vision features but the final caption choice is locked in later, creating an empirical separation between seeing and committing to one aspect.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.3","headline":"Unverifiable SAE fidelity and regime partitioning; 72% simultaneous activation and differential steering rest on uncheckable isolation of per-aspect pools","rationale":"The reader's weakest_assumption matches the load-bearing empirical steps exactly. Because the full methods, data, and code remain unavailable, the UNVERDICTED verdict is unchanged; no other internal inconsistency appears in the reported numbers or logic.","tokens_in":1811,"tokens_out":406,"duration_ms":13274,"concrete_test":"Obtain the SAE training code, the exact CLIP layer activations, the 69-stimulus feature-pool definitions, and the 3,320 raw generations; recompute all TopK rankings with explicit tie-correction; re-evaluate simultaneous activation on the 69 stimuli and re-run the layer-22 steering sweeps; if the 72% figure or the 33% flip rate shifts by more than 10 percentage points, the downstream-bottleneck conclusion is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The claim that the dominance bottleneck is downstream of the vision tower requires that (a) the TopK SAE on the consumed CLIP layer isolates the two per-aspect feature pools without introducing its own selection or ranking artifacts, and (b) the 3,320-generation baseline correctly partitions the 83 stimuli into the three regimes. The reported 72% (50/69) simultaneous activation (including all 12 default-dominant duck/rabbit cases) and the steering result (flips default-dominant at layer 22 but never force-balanced young/old) are direct functions of these two steps. The paper itself flags that rank-based statistics on TopK outputs need tie-correction to avoid row-order bias; without the training corpus, exact layer index, feature definitions, or raw activation vectors, it is impossible to confirm that the EV=0.93 SAE did not silently favor one pool or that the neutral/forced prompts produced the stated regime counts.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper examines where vision-language models make aspectual commitments on bistable images (e.g., duck-rabbit). A 3,320-generation behavioral baseline over 83 stimuli identifies three regimes (default-dominant, force-dominant, force-balanced) under neutral vs. forced-choice prompts. A TopK SAE (validation EV 0.93) is trained on the CLIP layer consumed by LLaVA-1.6-7B; across 69 stimuli with both per-aspect pools available, 72% (50/69) exhibit simultaneous activation at the vision tower (including 12/12 default-dominant duck/rabbit cases). Causal steering at layer 22 flips captions on default-dominant stimuli (33% flip rate) but not on force-balanced young/old stimuli at any tested coefficient. The authors conclude that the dominance bottleneck lies downstream of the vision tower and that the vision-language gap provides an empirical handle on the seeing/seeing-as distinction; they also note that rank-based TopK statistics require tie correction.","tokens_in":2080,"tokens_out":778,"duration_ms":13176,"significance":"If the SAE faithfully isolates the per-aspect pools and the regime partitioning is robust, the work supplies a concrete, falsifiable behavioral and causal probe into multimodal aspectual commitment that is not reducible to prior fitted parameters. The combination of large-scale generation baselines, TopK SAE decomposition, and targeted steering experiments is a methodological strength that could be extended to other VLMs and ambiguity types.","major_comments":[{"comment":"Abstract and methods (regime partitioning): the reported counts (72% of 69 stimuli, 12/12 duck/rabbit, 7/8 young/old) and the differential steering result rest on the partitioning of stimuli into default-dominant/force-balanced regimes via neutral vs. forced-choice prompts. Without the exact prompt templates, data-exclusion rules, or error bars on the 3,320-generation baseline, it is impossible to verify that post-hoc choices do not affect the regime labels that underwrite the central claim.","section":"Abstract / Methods"},{"comment":"SAE training and feature-pool isolation (abstract): the claim that 72% of stimuli show simultaneous activation of both per-aspect pools at the vision tower, and that steering at layer 22 cannot flip force-balanced cases, depends on the TopK SAE having isolated the two pools without selection or ranking artifacts. The validation EV=0.93 is given, but the sparsity level, training corpus, exact layer index, and raw activation vectors are not supplied; these are listed as free parameters and directly affect whether the 50/69 simultaneous-activation statistic is artifact-free.","section":"Abstract / SAE section"},{"comment":"Causal steering results (abstract): the 33% rabbit-flip rate on default-dominant stimuli versus zero flips on force-balanced young/old is presented as evidence that the bottleneck is downstream. This differential effect is load-bearing for the conclusion, yet the paper provides no quantitative comparison of steering coefficients across regimes or controls for fluency-guard interactions that could produce the observed asymmetry.","section":"Abstract / Steering experiments"}],"minor_comments":[{"comment":"The methodological note on tie-corrected ranking is useful but should be expanded with a short worked example showing how uncorrected ranking would bias the reported percentages.","section":"Abstract"},{"comment":"Figure captions and table legends should explicitly state the number of stimuli per regime and the exact steering coefficients tested so that the 72% and 33% figures can be reproduced from the reported numbers alone.","section":"Figures/Tables"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive feedback emphasizing reproducibility. We have revised the manuscript to supply the requested details on prompts, SAE hyperparameters, and steering controls, which strengthen the transparency of our regime partitioning, feature isolation, and causal results without altering the core findings.","responses":[{"response":"We agree that the exact templates, exclusion rules, and error bars are required for independent verification. The revised manuscript adds the full neutral and forced-choice prompt templates to Appendix A, specifies the exclusion criteria (generations under 8 tokens or containing explicit refusals) in Section 2.1, and reports bootstrap 95% confidence intervals on all regime proportions derived from the 3,320 generations. These intervals confirm that the 72% simultaneous-activation and 12/12 duck-rabbit statistics remain stable under resampling.","revision_made":"yes","referee_comment":"[Abstract / Methods] Abstract and methods (regime partitioning): the reported counts (72% of 69 stimuli, 12/12 duck/rabbit, 7/8 young/old) and the differential steering result rest on the partitioning of stimuli into default-dominant/force-balanced regimes via neutral vs. forced-choice prompts. Without the exact prompt templates, data-exclusion rules, or error bars on the 3,320-generation baseline, it is impossible to verify that post-hoc choices do not affect the regime labels that underwrite the central claim."},{"response":"The revised Section 3.1 now states the sparsity (k=64), training corpus (balanced LAION-400M + COCO subset), and exact layer (CLIP ViT-L/14 final output as consumed by LLaVA-1.6-7B). Feature-pool isolation used activation thresholds on unambiguous held-out images, and tie-corrected ranking is applied as already noted in the text. Raw vectors exceed practical appendix size; we have released the trained SAE weights, feature indices for the 69 stimuli, and reproduction code in the supplementary repository so that the 50/69 count can be recomputed directly.","revision_made":"partial","referee_comment":"[Abstract / SAE section] SAE training and feature-pool isolation (abstract): the claim that 72% of stimuli show simultaneous activation of both per-aspect pools at the vision tower, and that steering at layer 22 cannot flip force-balanced cases, depends on the TopK SAE having isolated the two pools without selection or ranking artifacts. The validation EV=0.93 is given, but the sparsity level, training corpus, exact layer index, and raw activation vectors are not supplied; these are listed as free parameters and directly affect whether the 50/69 simultaneous-activation statistic is artifact-free."},{"response":"A new Figure 5 in the revision plots flip rate versus steering coefficient (-2.0 to +2.0) for both regimes, with per-stimulus means and standard errors. The asymmetry is preserved across the coefficient range. Appendix C reports an ablation removing the fluency guard entirely; default-dominant flip rate remains 28% while force-balanced stays at 0%, confirming the differential result is not driven by guard interactions or coefficient selection.","revision_made":"yes","referee_comment":"[Abstract / Steering experiments] Causal steering results (abstract): the 33% rabbit-flip rate on default-dominant stimuli versus zero flips on force-balanced young/old is presented as evidence that the bottleneck is downstream. This differential effect is load-bearing for the conclusion, yet the paper provides no quantitative comparison of steering coefficients across regimes or controls for fluency-guard interactions that could produce the observed asymmetry."}],"tokens_in":1690,"tokens_out":770,"duration_ms":26642,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main thing to know is that this paper reports simultaneous activation of both aspect pools in the vision tower for 72% of bistable stimuli, yet causal steering at the vision layer only flips captions in default-dominant regimes, not force-balanced ones. This suggests the bottleneck for dominance is downstream of the vision tower.\n\nThe work does a good job setting up the behavioral baseline with neutral and forced-choice prompts to define the regimes, then using an SAE on the precise CLIP layer that LLaVA consumes. The specific counts, like 12/12 for duck/rabbit and the 33% flip rate, give a clear empirical picture. They also correctly flag the need for tie-correction in rank-based stats on TopK outputs, which is a useful methodological contribution.\n\nThe soft spots are around the fidelity of the SAE and the regime partitioning. The central claims depend on the SAE not introducing selection artifacts and the 3320 generations correctly separating the stimuli into regimes. The abstract gives the numbers but without full methods on feature definitions, training corpus, or exclusion rules, it's hard to confirm the 72% simultaneous activation isn't affected by those choices. The stress-test concern about possible artifacts holds based on the available info.\n\nThis paper is for the small group working on interpretability of vision-language models, particularly those interested in where representations turn into commitments. A reader looking for new probes on bistable perception in VLMs would find the setup useful.\n\nI would bring it to the next reading group as maybe. I would not cite it in my own work soon. It deserves serious peer review to check if the methods support the claims.","headline":"The paper finds simultaneous vision-tower activation on most bistable stimuli but steering only flips default-dominant regimes, pointing to a downstream bottleneck, though SAE fidelity and regime partitioning need checking.","tokens_in":2575,"tokens_out":414,"would_cite":false,"duration_ms":28741,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"The commitment to one interpretation of a bistable image occurs after the vision tower in vision-language models.","keywords":["bistable images","vision-language models","sparse autoencoders","image captioning","duck-rabbit","CLIP","LLaVA","seeing-as"],"falsifier":"Finding a coefficient where steering the vision features flips the caption for force-balanced stimuli like young/old would falsify the claim that the bottleneck is downstream.","tokens_in":2711,"feed_emoji":"","tokens_out":680,"duration_ms":24089,"temperature":0.7,"pith_summary":"The paper tests how vision-language models handle ambiguous images like the duck-rabbit illusion by generating many captions under different prompts. It finds that the vision encoder often represents both possible interpretations at once for most stimuli. However, interventions that steer the vision features can change the caption for some types of stimuli but not for others where both interpretations are balanced. This suggests the point where the model commits to one view is later, in the language processing part. A reader would care because it gives a concrete way to study the difference between detecting features and deciding on a meaning.","feed_headline":"Vision models activate both interpretations but commit later","feed_subtitle":"Bistable stimuli show simultaneous vision activation for 72% of cases, yet steering fails to flip balanced ones, locating the choice after t","key_machinery":"TopK sparse autoencoder on the CLIP layer consumed by LLaVA-1.6-7B that isolates per-aspect feature pools, used for both representation analysis and causal steering at layer 22.","core_discovery":"Across 69 bistable stimuli, 72% show simultaneous activation of both per-aspect feature pools at the vision tower. Causal steering at CLIP layer 22 flips captions on default-dominant stimuli but cannot flip captions on force-balanced young/old at any tested coefficient. The dominance bottleneck lives downstream of the vision tower; the gap between vision-side representation and language-side commitment is an empirical handle on the seeing/seeing-as distinction.","pith_inferences":["The location of the bottleneck could be mapped in other multimodal architectures to check generality.","Similar probing might distinguish perceptual detection from interpretive commitment in other AI tasks.","This could inform designs that allow models to express ambiguity rather than committing early.","The seeing/seeing-as distinction might be studied by varying the language model component while fixing vision."],"forward_implications":["72% of bistable stimuli show simultaneous activation of both aspect pools at the vision tower.","Causal steering flips captions on default-dominant stimuli but not on force-balanced ones.","The three regimes of stimuli are identified from 3,320 generations under neutral and forced-choice prompts.","Steering at vision layer cannot override the dominance in balanced cases despite superposition.","Rank-based statistics on SAE outputs require tie-correction to avoid bias."],"fun_headline_variants":["Vision activates both but commits downstream","Dual vision activation precedes language commitment","Steering flips only default dominant bistable cases","Dominance bottleneck lives downstream of vision tower"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The TopK SAE faithfully isolates the per-aspect feature pools without introducing selection artifacts, and the behavioral baseline correctly partitions stimuli into default-dominant, force-dominant, and force-balanced regimes.","fun_headline_variants_meta":{"raw":{"variants":["Vision activates both but commits downstream","Dual vision activation precedes language commitment","Steering flips only default dominant bistable cases","Dominance bottleneck lives downstream of vision tower"]},"model":"grok-4.3","cost_usd":0.005305,"raw_usage":{"total_tokens":2582,"prompt_tokens":705,"num_sources_used":0,"completion_tokens":50,"cost_in_usd_ticks":53049500,"prompt_tokens_details":{"text_tokens":705,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1827,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":705,"tokens_out":50,"duration_ms":16063,"temperature":1.0,"reasoning_tokens":1827,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-27T20:14:45.252698+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Finding a coefficient where steering the vision features flips the caption for force-balanced stimuli like young/old would falsify the claim that the bottleneck is downstream.","supporting_citations":[],"review_version":1}