REVIEW 2 major objections 4 minor 60 references
RADIO1D: Elastic Representations for Condensed Vision Modeling
T0 review · 2 major / 4 minor · reviewed 2026-07-12 · grok-4.5
Pith's one-line read VLMs do not need fixed 2D patch grids; variable-length 1D tokens can match or beat them while letting you dial accuracy against cost.
desk verdict Solid systems paper: elastic continuous 1D tokens from multi-teacher distillation give a real VLM accuracy-latency Pareto front and better composition retrieval, with the hierarchical-ordering story only partially isolated. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
RADIO1D: an encoder-decoder student that maps an image to an elastic 1D token sequence via multi-teacher distillation (SigLIP2, DINOv3, SAM3) and nested dropout (triangular length prior, early tokens kept preferentially), with Patch Merging internalized at block 24; the decoder is used only at training time to restore a 2D grid for teacher alignment.
What would settle it
Train an otherwise identical continuous 1D auto-encoder without nested dropout (or with a uniform length prior) and measure whether the single-token and low-rate ADE20K mIoU and VLM scores collapse relative to the nested-dropout RADIO1D checkpoint.
Extended reading notes
Core claim
Fixed 2D patch grids are not required for strong VLM performance. A hierarchical 1D sequence produced by multi-teacher distillation and nested dropout can summarize an image so effectively that a single token already yields non-trivial scene understanding, and adjustable token counts give continuous accuracy-efficiency trade-offs that match or exceed fixed-grid baselines on multimodal benchmarks.
Load-bearing premise
The paper treats the hierarchical ordering created by nested dropout plus the chosen multi-teacher mix as the main reason for strong low-token performance, without fully isolating that mechanism from the continuous embedding space or the rest of the auto-encoder design.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper argues that VLMs do not require fixed 2D patch grids. Analysis of SigLIP2, DINOv3 and VLM-finetuned C-RADIOv4 shows that image–text training produces abstract, low-spatial-coherence features and a few specialized global tokens. RADIO1D is then introduced: a multi-teacher (SigLIP2, DINOv3, SAM3) distilled encoder that maps an image to a continuous, variable-length 1D sequence via nested dropout (triangular length prior) and an internal Patch-Merging stage (κ=24, ρ=2). A training-only decoder reconstructs 2D features for teacher alignment. Experiments demonstrate hierarchical summarization (single-token ADE20K mIoU 40.23, strong ImageNet k-NN on early tokens), superior rate–distortion curves, a new composition-aware retrieval metric (Comp@K) on which RADIO1D outperforms baselines, and a continuous accuracy–latency Pareto front on ten VLM benchmarks when the same checkpoint is sliced to 1–256 tokens and paired with a fixed 9B Nemotron LLM.
Significance. If the empirical results hold, the work supplies a practical, elastic vision backbone that lets practitioners trade accuracy for latency at inference time without retraining, while matching or exceeding fixed-token SigLIP2 and C-RADIOv4 baselines. The composition-aware retrieval metric and the controlled Nemotron-VL ablation (identical LLM, data and schedule) are concrete, reusable contributions. The analysis of specialized global tokens and the loss of spatial coherence under VLM fine-tuning is also of independent interest to the foundation-model community. Model checkpoints are released under a permissive license, increasing the work’s immediate utility.
major comments (2)
- Section 4.1 and Figure 5 ablate the length prior and down-scaling position, yet never isolate nested dropout itself from the rest of the auto-encoder + multi-teacher recipe. Because hierarchical ordering is repeatedly invoked as the mechanism behind the low-rate regime (Figures 8–9, single-token mIoU 40.23), a controlled ablation that freezes the architecture and teachers while removing or randomizing the nested-dropout schedule would strengthen the causal claim; without it the systems-level Pareto front remains solid but the mechanistic story is only correlational.
- Table 1 reports single-run VLM numbers with no error bars or multi-seed statistics. Given that the central claim is a continuous accuracy–latency trade-off that “matches or exceeds” fixed-256-token baselines, modest run-to-run variance could reorder the Pareto ranking at intermediate token counts (e.g., 128 vs. 192). At least three independent SFT seeds for the key RADIO1D points would make the comparison statistically robust.
minor comments (4)
- Figure 1 caption and Table 1: TTFT is measured with a fixed 32-image / 128-token context; a short note on how the measurement changes with variable tile counts would help readers extrapolate to other deployment settings.
- Section 3.3: the claim “Minimizing the objective ensures that I(T_i;Y)>I(T_j;Y) ∀i<j” is stated without a formal derivation; a short information-theoretic argument or a pointer to the nested-dropout literature would clarify the guarantee.
- Appendix L (CKA regularization) is interesting but ultimately discarded; a one-sentence summary in the main text of why the regularizer was abandoned would prevent readers from wondering whether the final model still contains residual 2-D structure.
- Typographical: “A verage” in Table 1 header; “L VBench” should be “LVBench”; “pre-trained C-RADIOv4-H and C-RADIOv4-H (fine-tuned in a VLM)” in Figure 6 caption is slightly redundant.
Circularity Check
No significant circularity: empirical VLM and retrieval results rest on external benchmarks and standard multi-teacher distillation; self-citations supply only the base 2D RADIO recipe.
-
self citation load bearing
[Section 3.1 Agglomerative Training; also D. Initialization]
"We adopt the training recipe of C-RADIOv4 [26] to produce RADIO1D. ... We initialize RADIO1D from a pre-trained standard 2D Vision Transformer checkpoint (e.g., C-RADIOv4) using a dedicated conversion procedure."
The base agglomerative multi-teacher recipe and weight-initialization procedure are taken from the authors’ own prior C-RADIO papers. This is ordinary engineering reuse, not a load-bearing uniqueness claim or a definition that forces the new 1D elastic numbers; the variable-length results and VLM tables remain independent measurements.
full rationale
RADIO1D is an empirical systems paper. Its load-bearing claims (variable-token VLM accuracy–latency Pareto front on ten public benchmarks, single-token ADE20K mIoU, composition-aware retrieval on MS-COCO/nuImages) are measured against held-out external data with a controlled Nemotron-VL setup that freezes the LLM, data, and schedule. Nested dropout and the triangular length prior are training regularizers whose hierarchical effect is verified post-hoc by prefix/suffix curves and rate-distortion plots, not definitional identities that force the reported numbers. Multi-teacher losses are ordinary MSE/cosine distillation. The sole self-citations (C-RADIOv4 training recipe, Net2WiderNet-style initialization) supply the starting 2D backbone and are not used to prove uniqueness or to derive the 1D elastic results; those results are new measurements. No fitted parameter is renamed a prediction, no uniqueness theorem is imported, and no known result is merely re-labeled. Score 1 reflects only the minor, non-load-bearing self-citation of the authors’ prior RADIO line.
Assumptions & free parameters
free parameters (4)
- triangular PDF for token length ℓ =
p(x)=2-2x on [0,1]
- downscaling block position κ =
24
- embedding expansion factor ρ =
2
- teacher loss weights λ_t
assumptions (3)
- domain assumption Nested dropout on a 1D sequence induces a strict hierarchical ordering of mutual information I(T_i;Y) > I(T_j;Y) for i<j
- domain assumption Multi-teacher continuous distillation from DINOv3, SigLIP2 and SAM3 yields a student whose early tokens are globally informative
- ad hoc to paper The bipartite gIoU-based composition score correctly quantifies scene layout similarity for retrieval
invented entities (2)
-
RADIO1D elastic 1D continuous token sequence
independent evidence
-
Composition-aware retrieval metric (Comp@K)
Cite this review
Pith. "Pith review of RADIO1D: Elastic Representations for Condensed Vision Modeling." pith.science (2026). https://pith.science/paper/5PAEBXSG
@misc{pith2026260703624,
author = {Pith},
title = {Pith review of: RADIO1D: Elastic Representations for Condensed Vision Modeling},
year = {2026},
howpublished = {\url{https://pith.science/paper/5PAEBXSG}},
note = {Machine review of arXiv:2607.03624}
}
read the original abstract
This paper challenges the assumption that vision-language models (VLMs) require fixed patch-based 2D vision features. Analyzing fine-tuned vision encoders, we find that representations become increasingly abstract and less spatially coherent during VLM training. Notably, models trained with image-text alignment (such as SigLIP2) develop a small number of specialized tokens that effectively summarize global image content. Building on this, we introduce RADIO1D, which compresses images into a compact, variable-length 1D token sequence using multi-teacher knowledge distillation and an autoencoder design. The resulting representations exhibit strong hierarchical summarization, enabling accurate scene understanding - even with a single token - and support improved composition-aware image retrieval. In VLMs, RADIO1D provides flexible accuracy-efficiency tradeoffs through adjustable token counts, delivering competitive performance on diverse multimodal benchmarks with lower computational overhead and better accuracy.
Reference graph
Works this paper leans on
-
[1]
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agar- wal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning transferable visual models from natural language supervision. In Marina Meila and Tong Zhang, editors,Proceedings of the 38th International Conference on Machin...
2021
-
[2]
BLIP-2: Bootstrapping language- image pre-training with frozen image encoders and large language models
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. BLIP-2: Bootstrapping language- image pre-training with frozen image encoders and large language models. In Andreas Krause, Emma Brunskill, Kyunghyun Cho, Barbara En- gelhardt, Sivan Sabato, and Jonathan Scarlett, editors,Proceedings of the 40th International Conference on Machine Learning, volume 202...
2023
-
[3]
Language is not all you need: Aligning perception with language models
Shaohan Huang, Li Dong, Wenhui Wang, Yaru Hao, Saksham Singhal, Shuming Ma, Tengchao Lv, Lei Cui, Owais Khan Mohammed, Barun Patra, Qiang Liu, Kriti Aggarwal, Zewen Chi, Johan Bjorck, Vishrav Chaudhary, Subhojit Som, Xia Song, and Furu Wei. Language is not all you need: Aligning perception with language models. InAdvances in Neural Information Processing ...
2023
-
[4]
Visual instruction tuning
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning. InAd- vances in Neural Information Processing Systems, volume 36, 2023
2023
-
[5]
Sigmoid loss for language image pre-training
Xiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, and Lucas Beyer. Sigmoid loss for language image pre-training. InProceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 11975–11986, October 2023
2023
-
[6]
Michael Tschannen, Alexey Gritsenko, Xiao Wang, Muhammad Ferjad Naeem, Ibrahim Alab- dulmohsin, Nikhil Parthasarathy, Talfan Evans, Lucas Beyer, Ye Xia, Basil Mustafa, Olivier Hé- naff, Jeremiah Harmsen, Andreas Steiner, and Xiaohua Zhai. Siglip 2: Multilingual vision- language encoders with improved semantic under- standing, localization, and dense featu...
arXiv 2025
-
[7]
Paligemma: A versatile 3b vlm for transfer, 2024
Lucas Beyer, Andreas Steiner, André Su- sano Pinto, Alexander Kolesnikov, Xiao Wang, Daniel Salz, Maxim Neumann, Ibrahim Al- abdulmohsin, Michael Tschannen, Emanuele Bugliarello, Thomas Unterthiner, Daniel Keysers, Skanda Koppula, Fangyu Liu, Adam Grycner, Alexey Gritsenko, Neil Houlsby, Manoj Kumar, Keran Rong, Julian Eisenschlos, Rishabh Kabra, Matthias...
arXiv 2024
-
[8]
What matters when building vision-language models?, 2024
Hugo Laurençon, Léo Tronchon, Matthieu Cord, and Victor Sanh. What matters when building vision-language models?, 2024. URLhttps:// arxiv.org/abs/2405.02246
arXiv 2024
Show all 60 references
-
[9]
Cambrian-1: A fully open, vision-centric exploration of multimodal llms
Shengbang Tong, Ellis Brown, Penghao Wu, Sanghyun Woo, Manoj Middepogu, Sai Charitha Akula, Jihan Yang, Shusheng Yang, Adithya Iyer, Xichen Pan, Austin Wang, Rob Fergus, Yann Le- Cun, and Saining Xie. Cambrian-1: A fully open, vision-centric exploration of multimodal llms. In ...
2024
-
[10]
Nvila: Efficient frontier visual language models, 2024
Zhijian Liu, Ligeng Zhu, Baifeng Shi, Zhuoyang Zhang, Yuming Lou, Shang Yang, Haocheng Xi, Shiyi Cao, Yuxian Gu, Dacheng Li, Xiuyu Li, Yunhao Fang, Yukang Chen, Cheng-Yu Hsieh, De-An Huang, An-Chieh Cheng, Vishwesh Nath, Jinyi Hu, Sifei Liu, Ranjay Krishna, Daguang Xu, Xiaolon...
2024 arXiv
-
[11]
Qwen3-vl technical report, 2025
Shuai Bai, Yuxuan Cai, Ruizhe Chen, Keqin Chen, Xionghui Chen, Zesen Cheng, Lianghao Deng, Wei Ding, Chang Gao, Chunjiang Ge, Wenbin Ge, Zhifang Guo, Qidong Huang, Jie Huang, Fei Huang, Binyuan Hui, Shutong Jiang, Zhaohai Li, Mingsheng Li, Mei Li, Kaixin Li, Zicheng Lin, Junya...
2025 arXiv
-
[12]
Radiov2.5: Improved baselines for agglomerative vision foun- dation models
Greg Heinrich, Mike Ranzinger, Hongxu Danny Yin, Yao Lu, Jan Kautz, Andrew Tao, Bryan Catanzaro, and Pavlo Molchanov. Radiov2.5: Improved baselines for agglomerative vision foun- dation models. In2025 IEEE/CVF Confer- ence on Computer Vision and Pattern Recog- nition (CVPR), p...
2025 doi
-
[13]
Amala Sanjay Deshmukh, Kateryna Chu- machenko, Tuomas Rintamaki, Matthieu Le, Tyler Poon, Danial Mohseni-Taheri, Ilia Kar- manov, Guilin Liu, Jarno Seppänen, Guo Chen, Karan Sapra, Zhi-Wei Yu, Adi Renduchin- tala, Charles Wang, Peter Jin, Arushi Goel, Mike Ranzinger, Lukas Voe...
2025
-
[14]
Oriane Siméoni, Huy V. Vo, Maximilian Seitzer, Federico Baldassarre, Maxime Oquab, Cijo Jose, Vasil Khalidov, Marc Szafraniec, Seungeun Yi, Michaël Ramamonjisoa, Francisco Massa, Daniel Haziza, Luca Wehrstedt, Jianyuan Wang, Tim- othée Darcet, Théo Moutakanni, Leonel Sen- tana...
2025 arXiv
-
[15]
Sam 3: Segment anything with 11 RADIO1D: Elastic Representations for Condensed Vision Modeling concepts, 2025
Nicolas Carion, Laura Gustafson, Yuan-Ting Hu, and et al. Sam 3: Segment anything with 11 RADIO1D: Elastic Representations for Condensed Vision Modeling concepts, 2025. URL https://arxiv.org/abs/ 2511.16719
2025 arXiv
-
[16]
Eagle: Exploring the design space for multimodal LLMs with mixture of encoders
Min Shi, Fuxiao Liu, Shihao Wang, Shijia Liao, Subhashree Radhakrishnan, Yilin Zhao, De-An Huang, Hongxu Yin, Karan Sapra, Yaser Yacoob, Humphrey Shi, Bryan Catanzaro, Andrew Tao, Jan Kautz, Zhiding Yu, and Guilin Liu. Eagle: Exploring the design space for multimodal LLMs with...
2025
-
[17]
VILA-u: a unified foundation model integrating visual understanding and generation
Yecheng Wu, Zhuoyang Zhang, Junyu Chen, Hao- tian Tang, Dacheng Li, Yunhao Fang, Ligeng Zhu, Enze Xie, Hongxu Yin, Li Yi, Song Han, and Yao Lu. VILA-u: a unified foundation model integrating visual understanding and generation. InThe Thirteenth International Conference on Lear...
2025
-
[18]
Qwen-VL: A versatile vision-language model for under- standing, localization, text reading, and beyond,
Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou. Qwen-VL: A versatile vision-language model for under- standing, localization, text reading, and beyond,
-
[19]
URL https://openreview.net/forum? id=qrGjFJVl3m
-
[20]
Llava- uhd v2: an mllm integrating high-resolution se- mantic pyramid via hierarchical window trans- former, 2025
Yipeng Zhang, Yifan Liu, Zonghao Guo, Yi- dan Zhang, Xuesong Yang, Xiaoying Zhang, Chi Chen, Jun Song, Bo Zheng, Yuan Yao, Zhiyuan Liu, Tat-Seng Chua, and Maosong Sun. Llava- uhd v2: an mllm integrating high-resolution se- mantic pyramid via hierarchical window trans- former, ...
2025 arXiv
-
[21]
Scene parsing through ADE20K dataset
Bolei Zhou, Hang Zhao, Xavier Puig, Sanja Fidler, Adela Barriuso, and Antonio Torralba. Scene parsing through ADE20K dataset. InPro- ceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 5122–5130, 2017. doi: 10.1109/CVPR.2017.544. URL https:...
2017 doi
-
[22]
Schwing, Alexander Kirillov, and Rohit Girdhar
Bowen Cheng, Ishan Misra, Alexander G. Schwing, Alexander Kirillov, and Rohit Girdhar. Masked-attention mask transformer for univer- sal image segmentation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 1290–1299, 2022
2022
-
[23]
Donghao Zhang, Yimin Chen, Kauê T. N. Duarte, Taha Aslan, Mohamed AlShamrani, Brij Karmur, Yan Wan, Shengcai Chen, Bo Hu, Bi- joy K. Menon, and Wu Qiu. Benchmarking dinov3 for multi-task stroke analysis on non- contrast ct, 2025. URL https://arxiv.org/ abs/2509.23132
2025
-
[24]
DINOv3-driven se- mantic segmentation for landslide mapping in mountainous regions.Sensors, 26(2), 2026
Zhiyi Dou, Edore Akpokodje, Yuelin He, Yuxin Liu, Zixuan Ni, Chang’an Xu, Muhammad Aslam, and Meng Tang. DINOv3-driven se- mantic segmentation for landslide mapping in mountainous regions.Sensors, 26(2), 2026. doi: 10.3390/s26020406. URL https://doi.org/10. 3390/s26020406
2026 doi
-
[25]
Similarity of neural network representations revisited
Simon Kornblith, Mohammad Norouzi, Honglak Lee, and Geoffrey Hinton. Similarity of neural network representations revisited. In Kamalika Chaudhuri and Ruslan Salakhutdi- nov, editors,Proceedings of the 36th Interna- tional Conference on Machine Learning, vol- ume 97 ofProceedi...
-
[26]
URL https://proceedings.mlr.press/ v97/kornblith19a.html
-
[27]
Lawrence Zit- nick, and Piotr Dollár
Tsung-Yi Lin, Michael Maire, Serge Belongie, Lubomir Bourdev, Ross Girshick, James Hays, Pietro Perona, Deva Ramanan, C. Lawrence Zit- nick, and Piotr Dollár. Microsoft coco: Common objects in context, 2015
2015
-
[28]
C-radiov4 (tech report), 2026
Mike Ranzinger, Greg Heinrich, Collin McCarthy, Jan Kautz, Andrew Tao, Bryan Catanzaro, and Pavlo Molchanov. C-radiov4 (tech report), 2026. URLhttps://arxiv.org/abs/2601.17237
2026
-
[29]
Pereira, and William Bialek
Naftali Tishby, Fernando C. Pereira, and William Bialek. The information bottle- neck method. InProceedings of the 37th Annual Allerton Conference on Communica- tion, Control, and Computing, pages 368– 377, 1999. URL https://www.cs.huji.ac.il/ ~tishby/papers/IB-Allerton.pdf
1999
-
[30]
Modeling by shortest data description.Automatica, 14(5):465–471, 1978
Jorma Rissanen. Modeling by shortest data description.Automatica, 14(5):465–471, 1978. doi: 10.1016/0005-1098(78)90005-5. URLhttps: //doi.org/10.1016/0005-1098(78)90005-5
1978 doi
-
[31]
AM-RADIO: Agglomerative Vision Foundation Model Reduce All Domains Into One
Mike Ranzinger, Greg Heinrich, Jan Kautz, and Pavlo Molchanov. AM-RADIO: Agglomerative Vision Foundation Model Reduce All Domains Into One . In2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 12490–12500, Los Alamitos, CA, USA, June 2024. IEEE ...
2024 doi
-
[32]
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexan- der Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition ...
2021
-
[33]
Flextok: Resam- pling images into 1d token sequences of flexible length
Roman Bachmann, Jesse Allardice, David Mizrahi, Enrico Fini, Oğuzhan Fatih Kar, Elmira Amirloo, Alaaeldin El-Nouby, Amir Za- mir, and Afshin Dehghan. Flextok: Resam- pling images into 1d token sequences of flexible length. InForty-second International Confer- ence on Machine L...
2025
-
[34]
Swin transformer: Hierarchical vision transformer using shifted windows
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yix- uan Wei, Zheng Zhang, Stephen Lin, and Bain- ing Guo. Swin transformer: Hierarchical vision transformer using shifted windows. InProceed- ings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 10012–10022, October 2021
2021
-
[35]
Efros, Jenia Jitsev, Yair Carmon, Lud- wig Schmidt, and Vaishaal Shankar
Samir Yitzhak Gadre, Gabriel Ilharco, Alex Fang, Jonathan Hayase, Georgios Smyrnis, Thao Nguyen, Ryan Marten, Mitchell Wortsman, Dhruba Ghosh, Jieyu Zhang, Ekin Dogus Cubuk, Alexei A. Efros, Jenia Jitsev, Yair Carmon, Lud- wig Schmidt, and Vaishaal Shankar. Datacomp: In search...
2023 arXiv
-
[36]
Getting vit in shape: Scaling laws for compute-optimal model design
Ibrahim Alabdulmohsin, Xiaohua Zhai, Alexan- der Kolesnikov, and Lucas Beyer. Getting vit in shape: Scaling laws for compute-optimal model design. InAdvances in Neural Informa- tion Processing Systems, volume 36, 2023. URL https://arxiv.org/abs/2305.13035
2023 arXiv
-
[37]
Cover and Joy A
Thomas M. Cover and Joy A. Thomas. Rate distortion theory. InElements of Informa- tion Theory, chapter 10, pages 301–340. Wiley- Interscience, Hoboken, NJ, 2nd edition, 2006. doi: 10.1002/047174882X.ch10
2006 doi
-
[38]
McAfee, Laya Sleiman, Leon Derczynski, Luis Vega, Maer Rodrigues de Melo, Makesh Nar- simhan Sreedhar, Marcin Chochowski, Mark Cai, Markus Kliegl, Marta M
Nvidia Aarti Basant, Abhijit Khairnar, Ab- hijit Paithankar, Abhinav Khattar, Adi Ren- duchintala, Adi Renduchintala, Aditya Malte, Akhiad Bercovich, Akshay Hazare, Alejandra Rico, Aleksander Ficek, Alex Kondratenko, Alex Shaposhnikov, Ali Taghibakhshi, Amelia Barton, Ameya Ma...
2025 arXiv
-
[39]
Towards vqa models that can read
Amanpreet Singh, Vivek Natarajan, Meet Shah, YuJiang, XinleiChen, DhruvBatra, DeviParikh, and Marcus Rohrbach. Towards vqa models that can read. InThe IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2019
2019
-
[40]
Minesh Mathew, Dimosthenis Karatzas, and C.V. Jawahar. Docvqa: Adatasetforvqaondocument images, 2020. URL https://arxiv.org/abs/ 2007.00398
2020 arXiv
-
[41]
Minesh Mathew, Viraj Bagal, Rubén Tito, Di- mosthenis Karatzas, Ernest Valveny, and C.V. Jawahar. Infographicvqa. InProceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (W ACV), pages 1697–1706, January 2022
2022
-
[42]
OCRBench: On the hidden mys- tery of OCR in large multimodal models.Sci- ence China Information Sciences, 67(12):220102, dec 2024
Yuliang Liu, Zhang Li, Mingxin Huang, Biao Yang, Wenwen Yu, Chunyuan Li, Xu-Cheng Yin, Cheng-Lin Liu, Lianwen Jin, and Xi- ang Bai. OCRBench: On the hidden mys- tery of OCR in large multimodal models.Sci- ence China Information Sciences, 67(12):220102, dec 2024. ISSN 1869-1919...
2024
-
[43]
Ocrbench v2: An improved bench- mark for evaluating large multimodal models on visual text localization and reasoning, 2025
Ling Fu, Zhebin Kuang, Jiajun Song, Mingxin Huang, Biao Yang, Yuzhe Li, Linghao Zhu, Qidi Luo, Xinyu Wang, Hao Lu, Zhang Li, Guozhi Tang, Bin Shan, Chunhui Lin, Qi Liu, Binghong Wu, Hao Feng, Hao Liu, Can Huang, Jingqun Tang, Wei Chen, Lianwen Jin, Yuliang Liu, and Xiang Bai. ...
2025 arXiv
-
[44]
A diagram is worth a dozen images
Aniruddha Kembhavi, Mike Salvato, Eric Kolve, Minjoon Seo, Hannaneh Hajishirzi, and Ali Farhadi. A diagram is worth a dozen images. InEuropean Conference on Computer Vision (ECCV), pages 235–251. Springer, 2016
2016
-
[45]
ChartQA: A benchmark for question answering about charts with visual and logical reasoning
Ahmed Masry, Xuan Long Do, Jia Qing Tan, Shafiq Joty, and Enamul Hoque. ChartQA: A benchmark for question answering about charts with visual and logical reasoning. InFindings of the Association for Computational Linguis- tics: ACL 2022, pages 2263–2279, Dublin, Ire- land, May ...
2022 doi
-
[46]
findings-acl.177
URL https://aclanthology.org/2022. findings-acl.177
2022
-
[47]
Mmmu: A massive multi-discipline multi- modal understanding and reasoning benchmark for expert agi
Xiang Yue, Yuansheng Ni, Kai Zhang, Tianyu Zheng, Ruoqi Liu, Ge Zhang, Samuel Stevens, Dongfu Jiang, Weiming Ren, Yuxuan Sun, Cong Wei, BotaoYu, RuibinYuan, RenliangSun, Ming Yin, Boyuan Zheng, Zhenzhu Yang, Yibo Liu, Wenhao Huang, Huan Sun, Yu Su, and Wenhu Chen. Mmmu: A mass...
2024
-
[48]
Seed-bench: Benchmarking multimodal large language models
Bohao Li, Yuying Ge, Yixiao Ge, Guangzhi Wang, Rui Wang, Ruimao Zhang, and Ying Shan. Seed-bench: Benchmarking multimodal large language models. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR),pages13299–13308, June 2024
2024
-
[49]
Longvideobench: A benchmark for long- context interleaved video-language understand- ing, 2024
Haoning Wu, Dongxu Li, Bei Chen, and Junnan Li. Longvideobench: A benchmark for long- context interleaved video-language understand- ing, 2024. URL https://arxiv.org/abs/2407. 15754
2024
-
[50]
Token merging: Your ViT but faster
Daniel Bolya, Cheng-Yang Fu, Xiaoliang Dai, Peizhao Zhang, Christoph Feichtenhofer, and Judy Hoffman. Token merging: Your ViT but faster. InThe Eleventh International Conference on Learning Representations, 2023. URL https: //openreview.net/forum?id=aZ8qbRkUql
2023
-
[51]
Beyond attention or similarity: Maximizing conditional diversity for token prun- ing in mllms.arXiv preprint arXiv:2506.10967, 2025
Qizhe Zhang, Mengzhen Liu, Lichen Li, Ming Lu, Yuan Zhang, Junwen Pan, Qi She, and Shang- hang Zhang. Beyond attention or similarity: Maximizing conditional diversity for token prun- ing in mllms.arXiv preprint arXiv:2506.10967, 2025. 14 RADIO1D: Elastic Representations for Co...
2025 arXiv
-
[52]
Generalized intersection over union
Hamid Rezatofighi, Nathan Tsoi, JunYoung Gwak, Amir Sadeghian, Ian Reid, and Silvio Savarese. Generalized intersection over union. June 2019
2019
-
[53]
Harold W. Kuhn. The hungarian method for the assignment problem.Naval Re- search Logistics Quarterly, 2:83–97, 1955. doi: 10.1002/nav.3800020109. URL https://onlinelibrary.wiley.com/doi/ abs/10.1002/nav.3800020109
1955 doi
-
[54]
Lang, Sourabh Vora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Giancarlo Baldan, and Oscar Beijbom
Holger Caesar, Varun Bankiti, Alex H. Lang, Sourabh Vora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Giancarlo Baldan, and Oscar Beijbom. nuscenes: A multimodal dataset for autonomous driving.arXiv preprint arXiv:1903.11027, 2019
1903 arXiv
-
[55]
Vision transform- ers need registers
Timothée Darcet, Maxime Oquab, Julien Mairal, and Piotr Bojanowski. Vision transform- ers need registers. InThe Twelfth Interna- tional Conference on Learning Representations,
-
[56]
URL https://openreview.net/forum? id=2dnO3LLiJ1
-
[57]
An image is worth 32 tokens for reconstruction and generation
Qihang Yu, Mark Weber, Xueqing Deng, Xiaohui Shen, Daniel Cremers, and Liang-Chieh Chen. An image is worth 32 tokens for reconstruction and generation. InThe Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024. URLhttps://openreview.net/ forum?id=tOXoQPRzPL
2024
-
[58]
Net2net: Accelerating learning via knowl- edge transfer, 2016
Tianqi Chen, Ian Goodfellow, and Jonathon Shlens. Net2net: Accelerating learning via knowl- edge transfer, 2016. URLhttps://arxiv.org/ abs/1511.05641. Presented at ICLR 2016
2016 arXiv
-
[59]
Claude E. Shannon. Coding theorems for a dis- crete source with a fidelity criterion.IRE Na- tional Convention Record, 7(4):142–163, 1959
1959
-
[60]
512min_T2
Zhe Chen, Jiannan Wu, Wenhai Wang, and et al. Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks, 2023. URL https://arxiv.org/abs/ 2312.14238. 15 RADIO1D: Elastic Representations for Condensed Vision Modeling A. RADIO1D Architecture ...
2023 arXiv
Reviewed July 12, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.