REVIEW 2 major objections 69 references
Interleaved multimodal reasoning works better when models generate only the visual changes between steps, not full images each time.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-10 07:30 UTC pith:OAEQ2Y47
load-bearing objection Clean update-vs-full-image result is real; the 8.4% in-domain headline mixes paradigm with exclusive StructCoT data. the 2 major comments →
DeltaV: Thinking with Visual State Updates in Unified Large Multimodal Models
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The bottleneck in interleaved multimodal reasoning is not pixel-level generation itself but treating every intermediate visual state as independent full-image generation. Modeling reasoning as a base state plus sparse visual updates, with token budgets allocated by temporal similarity and diminishing reconstruction gain, concentrates supervision on reasoning-critical transitions, cuts redundant visual tokens, and improves reasoning.
What carries the argument
Visual state updates with a TSIM Router: each step emits variable-length update tokens conditioned on history; temporal similarity scores group change magnitudes, fitted token–fidelity curves stop allocation when marginal reconstruction gain falls below a threshold, and a learned end token transfers that length policy to inference.
Load-bearing premise
The offline map from how similar successive images look to how many update tokens they need, fitted on one validation set, remains a good general rule for when an update is finished—even at test time when only a learned stop token is available and the true future image is not.
What would settle it
A controlled same-architecture, same-data comparison in which full-image intermediate states match or beat TSIM-routed updates on both reconstruction fidelity and multi-step reasoning accuracy under equal average visual-token budgets would overturn the central efficiency and supervision claim.
If this is right
- Interleaved multimodal systems can treat visual continuity as a first-class resource instead of paying full-image cost at every step.
- Training signal for multi-step visual reasoning can be densified on state transitions rather than repeated low-level reconstruction.
- StructCoT’s scale and seven reasoning-structure categories become a standard training resource for update-centric multimodal models.
- Adaptive token budgets for visual intermediates become a practical alternative to fixed high-resolution image tokens in unified models.
Where Pith is reading between the lines
- Hybrid pipelines that emit compact updates for small changes and call external tools for deterministic edits (crop, rotate, enhance) follow naturally from the paper’s efficiency argument.
- The same continuity-plus-delta idea could apply to video, robotics, or GUI agents where successive frames share most pixels.
- Tasks that need fine local detail may still require higher-capacity perception branches; update tokens alone may under-serve pure recognition.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. DeltaV reformulates interleaved multimodal reasoning in ULMMs as progressive visual state updates rather than full-image generation at each step. Conditioned on historical visual tokens, the model predicts compact update tokens ΔZ_t; a TSIM Router (Eqs. 3–9) allocates token budgets from offline-fitted token–SSIM curves by stopping when marginal reconstruction gain falls below τ. The authors also release StructCoT (1.05M samples, 44 domains). Controlled ablations on Zebra-CoT show reconstruction parity at ~64 vs 144 tokens (−55.6% new visual tokens; Fig. 8) and +3.3% reasoning over full-image modeling (Tab. 1). With StructCoT and large-scale data, DeltaV-2B reports +8.4% over larger open models on in-domain tests (Tab. 4) and +5.9% over Qwen3-VL-2B on external benchmarks (Tab. 5).
Significance. If the controlled update-vs-full-image results hold, the paper offers a concrete, efficiency-oriented alternative to the dominant full-image intermediate-state paradigm in ULMMs: denser supervision on sparse transitions, fewer visual tokens, and better temporal consistency (Figs. 10–11). StructCoT is a substantial data contribution for the community. Strengths include matched-architecture ablations (Tabs. 1, 3; Fig. 8), component studies (Tabs. 6–7; Fig. 9), and promised release of code, models, and data. The work is systems-empirical rather than theoretical; its main value is a clear modeling shift plus evidence that reconstruction-aware token routing can improve reasoning under a fixed architecture.
major comments (2)
- Abstract and §5.5.1 / Tab. 4 package the +8.4% in-domain gain as if it primarily validates the visual-update paradigm, but the manuscript states that baselines are not fine-tuned on StructCoT. That confounds exclusive access to a 1.05M, 44-domain interleaved corpus with the architectural claim. The load-bearing evidence for the paradigm is the matched Zebra-CoT comparison (+3.3%, Tab. 1; −55.6% tokens, Fig. 8). Please restate headline claims so that data exclusivity is not attributed to visual updates, and add a same-data baseline (e.g., full-image or text-only trained on StructCoT) for Tab. 4-style numbers.
- §3.2–3.3 and §5.2: TSIM budgets are calibrated offline on Zebra-CoT validation (DINOv2 features, SSIM curves, fixed τ; Eq. 9) and transferred at inference only via a learned <|vision_end|> token, without online TSIM. Fig. 9(b) and App. A.1 support reconstruction transfer and threshold robustness, but the paper still treats reconstruction fidelity as a proxy for reasoning-useful token allocation (axiom underlying the Router). A short analysis of cases where high reconstruction quality does not improve reasoning (or where the learned stop token under/over-allocates relative to oracle TSIM) would strengthen the central transfer claim.
Circularity Check
No significant circularity: empirical systems paper with offline calibration as design choice, not a derivation that forces claimed gains by construction.
full rationale
DeltaV is an empirical multimodal systems paper. Its load-bearing claims are controlled ablations (visual-update vs full-image under matched architecture/data; Tab. 1, Fig. 8) and external benchmark numbers, not first-principles predictions. The TSIM Router (Sec. 3.2, Eq. 5–9; Sec. 5.2) is an offline calibration that maps temporal similarity intervals to token budgets via fitted SSIM–token curves and a marginal-gain threshold τ; those budgets then supervise training and a learned <|vision_end|> stop token. That is standard hyperparameter/teacher-policy design: reconstruction fidelity at the chosen operating point is intentionally controlled, but the reported reasoning gains (+3.3% over full-image) and external benchmark improvements are measured on held-out tasks whose objectives are not the SSIM curves used for calibration. Fixed-budget visual updates already beat full-image reconstruction and reasoning under the same token count, so the paradigm’s benefit is not forced solely by the slope rule. There is no self-definitional loop (X defined as Y then “derived”), no uniqueness theorem imported from the authors, no ansatz smuggled via self-citation, and no renaming of a known closed-form result. Self-citations, if any, are ordinary related-work context and not load-bearing for the central claims. Score 0 is therefore appropriate.
Axiom & Free-Parameter Ledger
free parameters (5)
- TSIM marginal-gain threshold τ
- Temporal decay α in TSIM
- Number of TSIM intervals M
- Discrete token budget set K and base K0=144
- Loss weights λ_rec, λ_p, λ_adv, λ_sem, λ_vis, β, λ_ent
axioms (4)
- domain assumption Consecutive intermediate visual states in interleaved reasoning are highly correlated; reasoning-critical information is concentrated in sparse changes.
- ad hoc to paper Reconstruction fidelity (SSIM/rFID/PSNR under TSIM-Tok) is a valid proxy for allocating tokens that later help reasoning.
- domain assumption A frozen foundation visual backbone (SigLIP2) plus discrete update slots can represent task-relevant visual transitions for AR language modeling.
- ad hoc to paper Offline DINOv2-based temporal similarity is a stable measure of visual change magnitude for budget assignment.
invented entities (3)
-
TSIM (temporal similarity) Router
no independent evidence
-
Visual update tokens ΔZ_t / TSIM-Tok
no independent evidence
-
StructCoT dataset
no independent evidence
read the original abstract
Current Unified Large Multimodal Models (ULMMs) support interleaved multimodal reasoning through textual reasoning and intermediate visual states, but typically generate each visual state as a full image. This full-image generation paradigm introduces substantial visual-token redundancy and dilutes supervision on sparse yet reasoning-critical state transitions. We propose DeltaV, a ULMM that replaces full-image generation with visual updates. Conditioned on historical visual states, DeltaV incrementally predicts compact update tokens that capture the visual changes across reasoning steps, avoiding repeated modeling of unchanged content. To align the token budget of each update with the magnitude of visual change, DeltaV introduces a temporal similarity (TSIM) Router, which stops allocating tokens once the marginal reconstruction gain falls below a threshold. To support more diverse and generalizable reasoning, we further construct StructCoT, a large-scale interleaved multimodal reasoning dataset with 1.05M samples spanning 44 task domains. Experiments show that the visual-update paradigm reduces newly generated visual tokens by 55.6\% on average without compromising reconstruction fidelity, and improves multimodal reasoning by 3.3\% over full-image generation. Trained with StructCoT and large-scale multimodal data, DeltaV-2B further outperforms substantially larger open-source models by 8.4\% on in-domain multimodal reasoning evaluations and surpasses the comparable-scale Qwen3-VL-2B by 5.9\% on external multimodal reasoning and understanding benchmarks. Code, models, and StructCoT will be released at https://github.com/Pengjie-W/DeltaV.
Reference graph
Works this paper leans on
-
[1]
Wang, Weiyun and Gao, Zhangwei and Gu, Lixin and Pu, Hengjun and Cui, Long and Wei, Xingguang and Liu, Zhaoyang and Jing, Linglin and Ye, Shenglong and Shao, Jie and others , journal=
-
[2]
Wu, Zhiyu and Chen, Xiaokang and Pan, Zizheng and Liu, Xingchao and Liu, Wen and Dai, Damai and Gao, Huazuo and Ma, Yiyang and Wu, Chengyue and Wang, Bingxuan and others , journal=
-
[3]
Bai, Shuai and Cai, Yuxuan and Chen, Ruizhe and Chen, Keqin and Chen, Xionghui and Cheng, Zesen and Deng, Lianghao and Ding, Wei and Gao, Chang and Ge, Chunjiang and others , journal=
-
[4]
Yunzhuo Hao and Jiawei Gu and Huichen Will Wang and Linjie Li and Zhengyuan Yang and Lijuan Wang and Yu Cheng , booktitle=. Can
-
[5]
Dongzhi Jiang and Renrui Zhang and Ziyu Guo and Yanwei Li and Yu Qi and Xinyan Chen and Liuhui Wang and Jianhan Jin and Claire Guo and Shen Yan and Bo Zhang and Chaoyou Fu and Peng Gao and Hongsheng Li , booktitle=
-
[6]
Image-of-Thought Prompting for Visual Reasoning Refinement in Multimodal Large Language Models
Image-of-thought prompting for visual reasoning refinement in multimodal large language models , author=. arXiv preprint arXiv:2405.13872 , year=
work page internal anchor Pith review Pith/arXiv arXiv
-
[7]
Zheng, Ziwei and Yang, Michael and Hong, Jack and Zhao, Chenxiao and Xu, Guohai and Yang, Le and Shen, Chao and Yu, Xing , journal=
-
[8]
Hu, Yushi and Shi, Weijia and Fu, Xingyu and Roth, Dan and Ostendorf, Mari and Zettlemoyer, Luke and Smith, Noah A and Krishna, Ranjay , booktitle=
-
[9]
Chen, Xiaokang and Wu, Zhiyu and Liu, Xingchao and Pan, Zizheng and Liu, Wen and Xie, Zhenda and Yu, Xingkai and Ruan, Chong , journal=
-
[10]
Emerging Properties in Unified Multimodal Pretraining
Emerging properties in unified multimodal pretraining , author=. arXiv preprint arXiv:2505.14683 , year=
work page internal anchor Pith review Pith/arXiv arXiv
-
[11]
Li, Ang and Wang, Charles and Fu, Deqing and Yue, Kaiyu and Cai, Zikui and Zhu, Wang Bill and Liu, Ollie and Guo, Peng and Neiswanger, Willie and Huang, Furong and others , journal=
-
[12]
Gu, Jiawei and Hao, Yunzhuo and Wang, Huichen Will and Li, Linjie and Shieh, Michael Qizhe and Choi, Yejin and Krishna, Ranjay and Cheng, Yu , journal=
-
[13]
Proceedings of the International Conference on Machine Learning , year=
Imagine While Reasoning in Space: Multimodal Visualization-of-Thought , author=. Proceedings of the International Conference on Machine Learning , year=
-
[14]
Advances in Neural Information Processing Systems , year=
Vision Foundation Models as Effective Visual Tokenizers for Autoregressive Generation , author=. Advances in Neural Information Processing Systems , year=
-
[15]
Du, Sinan and Guo, Jiahao and Li, Bo and Cui, Shuhao and Xu, Zhengzhuo and Luo, Yifu and Wei, Yongxian and Gai, Kun and Wang, Xinggang and Wu, Kai and others , journal=
-
[16]
He, Xin and Wei, Longhui and Ouyang, Jianbo and Xie, Lingxi and Tian, Qi , journal=
-
[17]
Emu3.5: Native Multimodal Models are World Learners
Emu3.5: Native multimodal models are world learners , author=. arXiv preprint arXiv:2510.26583 , year=
work page internal anchor Pith review Pith/arXiv arXiv
-
[18]
and Krishna, Ranjay , booktitle=
Bigverdi, Mahtab and Luo, Zelun and Hsieh, Cheng-Yu and Shen, Ethan and Chen, Dongping and Shapiro, Linda G. and Krishna, Ranjay , booktitle=. Perception Tokens Enhance Visual Reasoning in Multimodal Language Models , year=
-
[19]
Introducing Visual Perception Token into Multimodal Large Language Model
Introducing visual perception token into multimodal large language model , author=. arXiv preprint arXiv:2502.17425 , year=
work page internal anchor Pith review Pith/arXiv arXiv
-
[20]
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=
Visual programming: Compositional visual reasoning without training , author=. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=
-
[21]
Pixel Reasoner: Incentivizing Pixel-Space Reasoning with Curiosity-Driven Reinforcement Learning
Pixel reasoner: Incentivizing pixel-space reasoning with curiosity-driven reinforcement learning , author=. arXiv preprint arXiv:2505.15966 , year=
work page internal anchor Pith review Pith/arXiv arXiv
-
[22]
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=
Interleaved-modal chain-of-thought , author=. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=
-
[23]
Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing , pages=
Whiteboard-of-thought: Thinking step-by-step across modalities , author=. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing , pages=
work page 2024
-
[24]
Reinforced Visual Perception with Tools
Reinforced visual perception with tools , author=. arXiv preprint arXiv:2509.01656 , year=
work page internal anchor Pith review Pith/arXiv arXiv
- [25]
-
[26]
Chen, Zhangquan and Zhao, Ruihui and Luo, Chuwei and Sun, Mingze and Yu, Xinlei and Kang, Yangyang and Huang, Ruqi , booktitle=
-
[27]
Zhao, Qingqing and Lu, Yao and Kim, Moo Jin and Fu, Zipeng and Zhang, Zhuoyang and Wu, Yecheng and Li, Zhaoshuo and Ma, Qianli and Han, Song and Finn, Chelsea and others , booktitle=
-
[28]
Shi, Weikang and Yu, Aldrich and Fang, Rongyao and Ren, Houxing and Wang, Ke and Zhou, Aojun and Tian, Changyao and Fu, Xinyu and Hu, Yuxuan and Lu, Zimu and others , journal=
-
[29]
Tong, Shengbang and Fan, David and Li, Jiachen and Xiong, Yunyang and Chen, Xinlei and Sinha, Koustuv and Rabbat, Michael and LeCun, Yann and Xie, Saining and Liu, Zhuang , booktitle=
-
[30]
Lin, Bin and Li, Zongjian and Cheng, Xinhua and Niu, Yuwei and Ye, Yang and He, Xianyi and Yuan, Shenghai and Yu, Wangbo and Wang, Shaodong and Ge, Yunyang and others , journal=
-
[31]
Wu, Chenyuan and Zheng, Pengfei and Yan, Ruiran and Xiao, Shitao and Luo, Xin and Wang, Yueze and Li, Wanli and Jiang, Xiyan and Liu, Yexin and Zhou, Junjie and others , journal=
-
[32]
Fang, Rongyao and Duan, Chengqi and Wang, Kun and Li, Hao and Huang, Linjiang and Tian, Hao and Zeng, Xingyu and Zhao, Rui and Dai, Jifeng and Li, Hongsheng and others , booktitle=
-
[33]
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=
Infinity: Scaling bitwise autoregressive modeling for high-resolution image synthesis , author=. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=
-
[34]
Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation
Mogao: An omni foundation model for interleaved multi-modal generation , author=. arXiv preprint arXiv:2505.05472 , year=
work page internal anchor Pith review Pith/arXiv arXiv
-
[35]
Yang, Ling and Tian, Ye and Li, Bowen and Zhang, Xinchen and Shen, Ke and Tong, Yunhai and Wang, Mengdi , booktitle=
-
[36]
Li, Shufan and Gu, Jiuxiang and Liu, Kangning and Lin, Zhe and Wei, Zijun and Grover, Aditya and Kuen, Jason , journal=
-
[37]
Xin, Yi and Qin, Qi and Luo, Siqi and Zhu, Kaiwen and Yan, Juncheng and Tai, Yan and Lei, Jiayi and Cao, Yuewen and Wang, Keqi and Wang, Yibin and others , journal=. Lumina-
-
[38]
arXiv preprint arXiv:2511.23469 , year=
Visual Generation Tuning , author=. arXiv preprint arXiv:2511.23469 , year=
-
[39]
Liu, Yuan and Duan, Haodong and Zhang, Yuanhan and Li, Bo and Zhang, Songyang and Zhao, Wangbo and Yuan, Yike and Wang, Jiaqi and He, Conghui and Liu, Ziwei and others , booktitle=. 2024 , organization=
work page 2024
-
[40]
Advances in Neural Information Processing Systems , volume=
Are we on the right way for evaluating large vision-language models? , author=. Advances in Neural Information Processing Systems , volume=
-
[41]
Zou, Kai and Huang, Ziqi and Dong, Yuhao and Tian, Shulin and Zheng, Dian and Liu, Hongbo and He, Jingwen and Liu, Bin and Qiao, Yu and Liu, Ziwei , journal=
-
[42]
Beyond language modeling: An exploration of multimodal pretraining, March 2026
Beyond language modeling: An exploration of multimodal pretraining , author=. arXiv preprint arXiv:2603.03276 , year=
-
[43]
Zhang, Haichao and Li, Yijiang and He, Shwai and Nagarajan, Tushar and Chen, Mingfei and Lu, Jianglin and Li, Ang and Fu, Yun , journal=
-
[44]
Advances in Neural Information Processing Systems , volume=
Vision Foundation Models as Effective Visual Tokenizers for Autoregressive Generation , author=. Advances in Neural Information Processing Systems , volume=
-
[45]
Qwen2.5-VL Technical Report , author=. arXiv preprint arXiv:2502.13923 , year=
work page internal anchor Pith review Pith/arXiv arXiv
-
[46]
Chern, Ethan and Su, Jiadi and Ma, Yan and Liu, Pengfei , journal=
-
[47]
Chameleon: Mixed-Modal Early-Fusion Foundation Models
Chameleon: Mixed-modal early-fusion foundation models , author=. arXiv preprint arXiv:2405.09818 , year=
work page internal anchor Pith review Pith/arXiv arXiv
-
[48]
Hurst, Aaron and Lerer, Adam and Goucher, Adam P and Perelman, Adam and Ramesh, Aditya and Clark, Aidan and Ostrow, AJ and Welihinda, Akila and Hayes, Alan and Radford, Alec and others , journal=
-
[49]
Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities , author=. arXiv preprint arXiv:2507.06261 , year=
work page internal anchor Pith review Pith/arXiv arXiv
-
[50]
arXiv preprint arXiv:2505.02567 , year=
Unified multimodal understanding and generation models: Advances, challenges, and opportunities , author=. arXiv preprint arXiv:2505.02567 , year=
-
[51]
DINOv2: Learning Robust Visual Features without Supervision
Oquab, Maxime and Darcet, Timoth. arXiv preprint arXiv:2304.07193 , year=
work page internal anchor Pith review Pith/arXiv arXiv
-
[52]
Singh, Aaditya and Fry, Adam and Perelman, Adam and Tart, Adam and Ganesh, Adi and El-Kishky, Ahmed and McLaughlin, Aidan and Low, Aiden and Ostrow, AJ and Ananthram, Akhila and others , journal=
-
[53]
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=
Monkey: Image resolution and text label are important things for large multi-modal models , author=. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=
-
[54]
arXiv preprint arXiv:2602.08524 , year=
GeoFocus: Blending Efficient Global-to-Local Perception for Multimodal Geometry Problem-Solving , author=. arXiv preprint arXiv:2602.08524 , year=
-
[55]
International Journal of Computer Vision , volume=
Liquid: Language models are scalable and unified multi-modal generators , author=. International Journal of Computer Vision , volume=. 2026 , publisher=
work page 2026
-
[56]
Synthetic Similarity Search in Automotive Production
Synthetic Similarity Search in Automotive Production , author=. arXiv preprint arXiv:2505.07256 , year=
work page internal anchor Pith review Pith/arXiv arXiv
-
[57]
Advances in Neural Information Processing Systems , volume=
A tale of two features: Stable diffusion complements dino for zero-shot semantic correspondence , author=. Advances in Neural Information Processing Systems , volume=
-
[58]
Vision Foundation Models as Generalist Tokenizers for Image Generation
Vision Foundation Models as Generalist Tokenizers for Image Generation , author=. arXiv preprint arXiv:2605.18390 , year=
work page internal anchor Pith review Pith/arXiv arXiv
-
[59]
Du, Sinan and Guo, Jiahao and Li, Bo and Cui, Shuhao and Xu, Zhengzhuo and Luo, Yifu and Wei, Yongxian and Gai, Kun and Wang, Xinggang and Wu, Kai and others , booktitle=
-
[60]
Tschannen, Michael and Gritsenko, Alexey and Wang, Xiao and Naeem, Muhammad Ferjad and Alabdulmohsin, Ibrahim and Parthasarathy, Nikhil and Evans, Talfan and Beyer, Lucas and Xia, Ye and Mustafa, Basil and others , journal=
-
[61]
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=
Machine mental imagery: Empower multimodal reasoning with latent visual tokens , author=. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=
-
[62]
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=
Monet: Reasoning in Latent Visual Space Beyond Image and Language , author=. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=
-
[63]
Update to GPT-5 System Card: GPT-5.2 , year =
-
[64]
Chen, Qiguang and Qin, Libo and Zhang, Jin and Chen, Zhi and Xu, Xiao and Che, Wanxiang , booktitle=
-
[65]
Lu, Pan and Bansal, Hritik and Xia, Tony and Liu, Jiacheng and Li, Chunyuan and Hajishirzi, Hannaneh and Cheng, Hao and Chang, Kai-Wei and Galley, Michel and Gao, Jianfeng , booktitle=
-
[66]
Xu, Weiye and Wang, Jiahao and Wang, Weiyun and Chen, Zhe and Zhou, Wengang and Yang, Aijun and Lu, Lewei and Li, Houqiang and Wang, Xiaohua and Zhu, Xizhou and Wang, Wenhai and Dai, Jifeng and Zhu, Jinguo , booktitle=
-
[67]
Fu, Chaoyou and Chen, Peixian and Shen, Yunhang and Qin, Yulei and Zhang, Mengdan and Lin, Xu and Yang, Jinrui and Zheng, Xiawu and Li, Ke and Sun, Xing and Wu, Yunsheng and Ji, Rongrong and Shan, Caifeng and He, Ran , booktitle=
-
[68]
Wu, Penghao and Xie, Saining , booktitle=
-
[69]
Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal
Tong, Shengbang and Liu, Zhuang and Zhai, Yuexiang and Ma, Yi and LeCun, Yann and Xie, Saining , booktitle=. Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.