REVIEW 3 major objections 4 minor 84 references
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding?
T0 review · 3 major / 4 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read HRScene, a new benchmark spanning 25 real-world and two synthetic high-resolution image tasks, finds that current vision-language models average about 50% accuracy and show a U-shaped 'lost-in-the-middle' pattern in Manhattan distance…
desk verdict A useful real-world high-resolution VLM benchmark, but the diagnostic 'Manhattan lost-in-the-middle' finding is underdetermined until 2D positional-embedding artifacts are ruled out. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the pair of synthetic diagnostic constructions built on the needle-in-a-haystack idea. WhiteBackground places a small natural image with its question as the needle in an $N \times N$ grid of identical white cells and quantifies Regional Divergence as the difference between the best-performing region and the mean over regions. ComplexGrid places the needle among visually similar distractor images and asks the model to output the needle's row and column, with accuracy plotted against Manhattan distance $\lvert x-1\rvert + \lvert y-1\rvert$ from the upper-left cell at row 1, column 1. The real-world arm supplies the accuracy baseline and the scale range: 25 curated tasks, partly re-annotated by graduate annotators, spanning eight categories with images from 1,024 × 1,024 up to 35,503 × 26,627 pixels.
What would settle it
A control experiment that embeds the same needle images into a continuous real high-resolution scene instead of a white or similar-image grid, then measures accuracy against the ground-truth pixel position of the target, would settle whether the Manhattan-distance U-shape is a general property of high-resolution perception or an artifact of grid composition; the paper's supplementary ablation rules out only the linear patch-order alternative.
Extended reading notes
Core claim
The central discovery is that current vision-language models do not yet understand high-resolution images in a spatially uniform way. On HRScene's real-world multiple-choice tasks, the average accuracy of all 28 evaluated models is 49.68%, and the best model reaches only about 62%. The two diagnostic datasets locate the failure: in WhiteBackground, a single natural-image needle is pasted onto white $N \times N$ grids of increasing size, and accuracy drops while the gap between the best and average region grows; in ComplexGrid, the needle sits among visually similar distractors and models must report its row and column, and accuracy follows a U-shape in Manhattan distance to the top-left corner rather than in linear patch order. The paper argues that this is a distinct spatial version of lost-in-the-middle for high-resolution images, not the linear-context phenomenon, and that the position-dependent drop is robust across model families and sizes.
Load-bearing premise
The diagnostic tests assume that a small natural image pasted onto a white or similar-image grid faithfully represents how a model uses regions of a genuinely high-resolution real scene, and that the resulting position-dependent accuracy reflects HRI utilization rather than an artifact of the collage layout.
Editorial extensions
If this is right
- If HRScene is taken as the measurement, no current VLM is close to reliable high-resolution understanding: average accuracy near 50% versus roughly 65% for the human annotators.
- Native-resolution input handling is a visible lever: the best-performing model is the only one above 60% overall and is especially strong on very large images, indicating that resolution-adaptive architectures matter as much as parameter count.
- Regional divergence means that where the answer sits in a high-resolution image materially changes a model's measured ability, so position-uncontrolled benchmarks can mislead.
- The U-shaped Manhattan-distance curve gives a concrete target: reducing the middle-region drop should improve accuracy on grid-like and multi-subimage inputs.
- Model-size scaling improves accuracy only logarithmically, so simply making models bigger is unlikely to close the high-resolution gap.
Reading between the lines
- The paper does not test whether the two diagnostic patterns come from the vision encoder's positional bias rather than from the language model's spatial reasoning; a control that moves a real object within one continuous high-resolution scene would separate 'regional utilization failure' from 'collage layout artifact.'
- If the U-shape reflects distance from the first visual patch, reordering how an image is tiled (for example, center-out or serpentine scanning) could be a cheap, training-free fix for middle-region accuracy; the paper explores no reordering.
- Exact-match scoring on multiple-choice options may hide partial knowledge; a likelihood-based or open-ended scoring variant could change model rankings and give a finer signal for small improvements.
- Because the synthetic diagnostics reuse a single source of needle images, porting the same construction to radiology, remote sensing, and document domains would show whether regional divergence is content-dependent or a general property of current VLMs.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper introduces HRScene, a benchmark for high-resolution image (HRI) understanding that combines 25 real-world datasets (7,068 samples, resolutions from 1K to roughly 35,000×26,000 pixels) with two synthetic NIAH-style diagnostic datasets. The authors evaluate 28 VLMs (6 proprietary, 22 open-weights) and report that current models reach only about 50% average accuracy on real-world tasks, with the best model (Qwen2-VL 72B) scoring 61.85% versus a non-expert human baseline of 64.72%. On the synthetic diagnostics, the paper observes Regional Divergence (accuracy varies across grid positions) and a 'lost-in-the-middle-Manhattan' U-shape in which accuracy declines with Manhattan distance from the top-left corner of the grid.
Significance. The real-world part of HRScene is a substantial and useful contribution: it covers an unusually wide range of scene types and resolutions, includes 2,008 re-annotated or scratch-annotated samples, and provides a consistent evaluation of 28 models. The headline finding that all current VLMs, including Gemini 2.0 Flash and GPT-4o, perform around 50% on these tasks is credible and likely to be widely cited. The diagnostic claims, however, are only as strong as the assumption that per-position performance differences in the synthetic grids reflect content-dependent region utilization; the paper currently does not rule out vision-encoder positional or tiling artifacts. If the authors can close that gap with additional ablations, the diagnostic contribution would be a valuable tool for the field.
major comments (3)
- [§4.3 and Supplementary §8] The lost-in-the-middle-Manhattan claim is not adequately separated from vision-encoder spatial priors. The supplementary ablation only rules out linear patch-order distance; it does not control for 2D positional-embedding geometry, row/column main effects, distance to image border, tiling boundaries, or the effect of reversing the visual token sequence. Because the needle is a low-resolution VQAv2 image re-encoded inside a synthetic grid by each model's vision transformer, per-position accuracy differences could arise from the encoder's positional geometry rather than from content-dependent attention to image regions. Since the paper explicitly claims this phenomenon is novel and distinct from text NIAH, this confound is load-bearing for the diagnostic contribution. Please add ablations such as random permutation of grid cells, reversed or shuffled visual token order, or regression on row, column, and distance-to-border to establish that the U-shape reflects region utilization rather than encoder artifacts.
- [§4.3, Table 4] Regional Divergence is defined as the difference between the best region's accuracy and the mean accuracy across regions. With roughly 500 samples split across up to 100 grid regions on the 10×10 setting, per-region estimates are based on about five samples each, so the reported divergence values (e.g., Gemini-2.0-Flash at 39.85%) may be dominated by sampling noise. The paper should report confidence intervals, bootstrap estimates, or significance tests to establish that the divergence is not an artifact of small per-region sample sizes.
- [§3.2, §3.4, Table 1] The benchmark reuses datasets that are standard VLM training corpora (VQAv2, InfographicVQA, DocStruct4M, NovaChart) without a contamination analysis. Since HRScene's purpose is to measure current capability gaps, the paper should test for overlap with common training data or explicitly discuss how contamination would affect the headline numbers. This concern is especially relevant for the WhiteBackground diagnostic, which is constructed directly from VQAv2 and inherits any question-answer memorization present in the evaluated models.
minor comments (4)
- [Table 3] The human performance for Medical is 23.81%, only 1.5 points above random; the annotators are graduate students rather than domain experts, so the 'human' row should be labeled as non-expert human performance or accompanied by a caveat for expert-dependent categories.
- [§4.1 and throughout] There are recurring typos in model names and metrics: 'exact math' should be 'exact match', 'Calude' should be 'Claude', and 'molMo' appears for 'Molmo'.
- [Supplementary §8] The sentence 'no significant patter can be observed' contains a typo ('patter' for 'pattern'), and the figure caption should state the number of models and samples used for the patch-order ablation, which is currently missing.
- [Figure 2] The x-axis label '1k 2k 3k 4k 8k 5×10' is unclear; '5×10' appears to denote 5×10^6 pixels but this is not defined in the caption.
Circularity Check
No circularity found; HRScene is an empirical benchmark whose diagnostics are measured statistics, not derived predictions.
full rationale
HRScene is an evaluation benchmark rather than a derivation: every headline result (roughly 50% average real-world accuracy, Region↓ values, and the Manhattan-distance U-shape) is a directly measured statistic over model outputs on fixed inputs. No fitted parameter is later renamed as a prediction. Regional Divergence is explicitly defined in Section 4.3 as 'the difference between the highest performance region and the mean performance of every region' and then computed from observed per-position accuracies; the name is a label for the measured spread, not a quantity derived from that definition. The lost-in-the-middle claim is supported by the empirical curves in Figure 4 and by the supplementary Section 8 patch-order ablation, which is an experimental control rather than a citation or an assumed premise. The possible confound that 2D positional embeddings or tiling artifacts cause the U-shape is a construct-validity concern about what the diagnostic measures, not circularity: the phenomenon is observed, not assumed into existence through the defining equations. The paper does not rely on a self-citation chain to license its main conclusions; references to NIAH, MM-NIAH, and Visual Haystack are prior work by other groups used for positioning, not proof of HRScene's own results. I therefore find no step that reduces a claimed prediction or first-principles claim to its own inputs.
Assumptions & free parameters
assumptions (2)
- domain assumption The sampled subsets of existing datasets remain valid evaluation tasks after reannotation, with unambiguous options.
- domain assumption The synthetic NIAH tests (white grid, complex grid) isolate regional utilization rather than confounds such as needle scale or positional priors.
Cite this review
Pith. "Pith review of HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding?." pith.science (2026). https://pith.science/paper/AXMK6UAA
@misc{pith2026250418406,
author = {Pith},
title = {Pith review of: HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding?},
year = {2026},
howpublished = {\url{https://pith.science/paper/AXMK6UAA}},
note = {Machine review of arXiv:2504.18406}
}
abstract
High-resolution image (HRI) understanding aims to process images with a large number of pixels, such as pathological images and agricultural aerial images, both of which can exceed 1 million pixels. Vision Large Language Models (VLMs) can allegedly handle HRIs, however, there is a lack of a comprehensive benchmark for VLMs to evaluate HRI understanding. To address this gap, we introduce HRScene, a novel unified benchmark for HRI understanding with rich scenes. HRScene incorporates 25 real-world datasets and 2 synthetic diagnostic datasets with resolutions ranging from 1,024 $\times$ 1,024 to 35,503 $\times$ 26,627. HRScene is collected and re-annotated by 10 graduate-level annotators, covering 25 scenarios, ranging from microscopic to radiology images, street views, long-range pictures, and telescope images. It includes HRIs of real-world objects, scanned documents, and composite multi-image. The two diagnostic evaluation datasets are synthesized by combining the target image with the gold answer and distracting images in different orders, assessing how well models utilize regions in HRI. We conduct extensive experiments involving 28 VLMs, including Gemini 2.0 Flash and GPT-4o. Experiments on HRScene show that current VLMs achieve an average accuracy of around 50% on real-world tasks, revealing significant gaps in HRI understanding. Results on synthetic datasets reveal that VLMs struggle to effectively utilize HRI regions, showing significant Regional Divergence and lost-in-middle, shedding light on future research.
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[1]
12 Kucherlapati Raju 13, Genome data analysis: Baylor College of Medicine Creighton Chad J
Brigham & Women’s Hospital & Harvard Medical School Chin Lynda 9 11 Park Peter J. 12 Kucherlapati Raju 13, Genome data analysis: Baylor College of Medicine Creighton Chad J. 22 23 Donehower Lawrence A. 22 23 24 25, Institute for Systems Biology Reynolds Sheila 31 Kreisberg Richard B. 31 Bernard Brady 31 Bressler Ryan 31 Erkkila Timo 32 Lin Jake 31 Thorsso...
2012
-
[2]
Phi-3 technical report: A highly capable language model locally on your phone, 2024
Marah Abdin, Jyoti Aneja, Hany Awadalla, Ahmed Awadal- lah, Ammar Ahmad Awan, Nguyen Bach, Amit Bahree, Arash Bakhtiari, Jianmin Bao, Harkirat Behl, et al. Phi-3 technical report: A highly capable language model locally on your phone, 2024. URL https://arxiv. org/abs/2404.14219,
arXiv 2024
-
[3]
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ah- mad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. Gpt-4 technical report. arXiv preprint arXiv:2303.08774 ,
-
[4]
The claude 3 model family: Opus, sonnet, haiku
AI Anthropic. The claude 3 model family: Opus, sonnet, haiku. Claude-3 Model Card, 2024. 1, 2, 6
2024
-
[5]
Bach: Grand challenge on breast cancer histology im- ages
Guilherme Aresta, Teresa Ara ´ujo, Scotty Kwok, Sai Saketh Chennamsetty, Mohammed Safwan, Varghese Alex, Bahram Marami, Marcel Prastawa, Monica Chan, Michael Donovan, et al. Bach: Grand challenge on breast cancer histology im- ages. Medical image analysis, 56:122–139, 2019. 3
2019
-
[6]
Jauregui, and Juan Andr ´es Cardoso
Darwin Alexis Arrechea-Castillo, Paula Espitia-Buitrago, Ronald David Arboleda, Luis Miguel Hernandez, Rosa N. Jauregui, and Juan Andr ´es Cardoso. High-resolution image dataset for the automatic classification of phenological stage and identification of racemes in urochloa spp. hybrids. Data in Brief, 57:110928, 2024. 4
2024
-
[7]
Ef- ficient high-resolution deep learning: A survey
Arian Bakhtiarnia, Qi Zhang, and Alexandros Iosifidis. Ef- ficient high-resolution deep learning: A survey. ACM Com- puting Surveys, 56(7):1–35, 2024. 1, 3
work page 2024
-
[8]
Effi- cient high-resolution deep learning: A survey.ACM Comput
Arian Bakhtiarnia, Qi Zhang, and Alexandros Iosifidis. Effi- cient high-resolution deep learning: A survey.ACM Comput. Surv., 2024. 2, 8
work page 2024
Show all 84 references
-
[9]
Janus- pro: Unified multimodal understanding and generation with data and model scaling
Xiaokang Chen, Zhiyu Wu, Xingchao Liu, Zizheng Pan, Wen Liu, Zhenda Xie, Xingkai Yu, and Chong Ruan. Janus- pro: Unified multimodal understanding and generation with data and model scaling. arXiv preprint arXiv:2501.17811 ,
-
[10]
Expanding performance boundaries of open-source multimodal models with model, data, and test- time scaling
Zhe Chen, Weiyun Wang, Yue Cao, Yangzhou Liu, Zhang- wei Gao, Erfei Cui, Jinguo Zhu, Shenglong Ye, Hao Tian, Zhaoyang Liu, et al. Expanding performance boundaries of open-source multimodal models with model, data, and test- time scaling. arXiv preprint arXiv:2412.05271, 2024. 6, 2
2024 arXiv
-
[11]
How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites
Zhe Chen, Weiyun Wang, Hao Tian, Shenglong Ye, Zhang- wei Gao, Erfei Cui, Wenwen Tong, Kongzhi Hu, Jiapeng Luo, Zheng Ma, et al. How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites. arXiv preprint arXiv:2404.16821, 2024
2024 arXiv
-
[12]
Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks
Zhe Chen, Jiannan Wu, Wenhai Wang, Weijie Su, Guo Chen, Sen Xing, Muyan Zhong, Qinglong Zhang, Xizhou Zhu, Lewei Lu, et al. Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks. In Pro- ceedings of the IEEE/CVF Conference on Computer...
2024
-
[13]
Functional map of the world
Gordon Christie, Neil Fendley, James Wilson, and Ryan Mukherjee. Functional map of the world. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 6172–6180, 2018. 3
2018
-
[14]
The cityscapes dataset for semantic urban scene understanding
Marius Cordts, Mohamed Omran, Sebastian Ramos, Timo Rehfeld, Markus Enzweiler, Rodrigo Benenson, Uwe Franke, Stefan Roth, and Bernt Schiele. The cityscapes dataset for semantic urban scene understanding. In Proceed- ings of the IEEE conference on computer vision and pattern re...
2016
-
[15]
Molmo 9 and pixmo: Open weights and open data for state-of-the-art multimodal models
Matt Deitke, Christopher Clark, Sangho Lee, Rohun Tri- pathi, Yue Yang, Jae Sung Park, Mohammadreza Salehi, Niklas Muennighoff, Kyle Lo, Luca Soldaini, et al. Molmo 9 and pixmo: Open weights and open data for state-of-the-art multimodal models. arXiv preprint arXiv:2409.17146,...
2024 arXiv
-
[16]
LungHist700: A dataset of histological images for deep learning in pulmonary pathology
Jorge Diosdado, Pere Gilabert, Santi Segu ´ı, and Henar Bor- rego. LungHist700: A dataset of histological images for deep learning in pulmonary pathology. Scientific Data, 11 (1):1088, 2024. 1
2024
-
[17]
LungHist700: A dataset of histological images for deep learning in pulmonary pathology
Jorge Diosdado, Pere Gilabert, Santi Segu ´ı, and Henar Bor- rego. LungHist700: A dataset of histological images for deep learning in pulmonary pathology. Scientific Data, 11 (1):1088, 2024. 4
2024
-
[18]
Internlm-xcomposer2-4khd: A pioneer- ing large vision-language model handling resolutions from 336 pixels to 4k HD
Xiaoyi Dong, Pan Zhang, Yuhang Zang, Yuhang Cao, Bin Wang, Linke Ouyang, Songyang Zhang, Haodong Duan, Wenwei Zhang, Yining Li, Hang Yan, Yang Gao, Zhe Chen, Xinyue Zhang, Wei Li, Jingwen Li, Wenhai Wang, Kai Chen, Conghui He, Xingcheng Zhang, Jifeng Dai, Yu Qiao, Dahua Lin, a...
2024
-
[19]
The llama 3 herd of models
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Ab- hishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. The llama 3 herd of models. arXiv preprint arXiv:2407.21783 ,
-
[20]
Describing differences in image sets with natural language
Lisa Dunlap, Yuhui Zhang, Xiaohan Wang, Ruiqi Zhong, Trevor Darrell, Jacob Steinhardt, Joseph E Gonzalez, and Serena Yeung-Levy. Describing differences in image sets with natural language. InProceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, p...
2024
-
[21]
Floorplancad: A large-scale cad draw- ing dataset for panoptic symbol spotting
Zhiwen Fan, Lingjie Zhu, Honghua Li, Xiaohao Chen, Siyu Zhu, and Ping Tan. Floorplancad: A large-scale cad draw- ing dataset for panoptic symbol spotting. In Proceedings of the IEEE/CVF international conference on computer vision, pages 10128–10137, 2021. 4
2021
-
[22]
Mini-internvl: A flexible-transfer pocket multimodal model with 5% parameters and 90% perfor- mance
Zhangwei Gao, Zhe Chen, Erfei Cui, Yiming Ren, Weiyun Wang, Jinguo Zhu, Hao Tian, Shenglong Ye, Junjun He, Xizhou Zhu, et al. Mini-internvl: A flexible-transfer pocket multimodal model with 5% parameters and 90% perfor- mance. arXiv preprint arXiv:2410.16261, 2024. 6, 2
-
[23]
Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing
Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Ba- tra, and Devi Parikh. Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing. In Proceedings of the IEEE conference on computer vision and pattern recognition , pages 6904–691...
2017
-
[24]
Llava-uhd: An LMM perceiving any aspect ratio and high- resolution images
Zonghao Guo, Ruyi Xu, Yuan Yao, Junbo Cui, Zanlin Ni, Chunjiang Ge, Tat-Seng Chua, Zhiyuan Liu, and Gao Huang. Llava-uhd: An LMM perceiving any aspect ratio and high- resolution images. In European Conference on Computer Vision (ECCV), 2024. 3
2024
-
[25]
Cogagent: A visual lan- guage model for GUI agents
Wenyi Hong, Weihan Wang, Qingsong Lv, Jiazheng Xu, Wenmeng Yu, Junhui Ji, Yan Wang, Zihan Wang, Yuxiao Dong, Ming Ding, and Jie Tang. Cogagent: A visual lan- guage model for GUI agents. In IEEE/CVF Computer Vision and Pattern Recognition Conference (CVPR), 2024. 2
2024
-
[26]
mplug-docowl 1.5: Unified structure learning for ocr-free document understanding
Anwen Hu, Haiyang Xu, Jiabo Ye, Ming Yan, Liang Zhang, Bo Zhang, Ji Zhang, Qin Jin, Fei Huang, and Jingren Zhou. mplug-docowl 1.5: Unified structure learning for ocr-free document understanding. In Conference on Empirical Meth- ods in Natural Language Processing (EMNLP) Findings ,
-
[27]
mplug-docowl2: High-resolution compressing for ocr- free multi-page document understanding
Anwen Hu, Haiyang Xu, Liang Zhang, Jiabo Ye, Ming Yan, Ji Zhang, Qin Jin, Fei Huang, and Jingren Zhou. mplug-docowl2: High-resolution compressing for ocr- free multi-page document understanding. arXiv preprint arXiv:2409.03420, 2024. 1, 4
2024 arXiv
-
[28]
Novachart: A large- scale dataset towards chart understanding and generation of multimodal large language models
Linmei Hu, Duokang Wang, Yiming Pan, Jifan Yu, Yingxia Shao, Chong Feng, and Liqiang Nie. Novachart: A large- scale dataset towards chart understanding and generation of multimodal large language models. In Proceedings of the 32nd ACM International Conference on Multimedia , p...
2024
-
[29]
High reso- lution image quality database
Huang Huang, Qiang Wan, and Jari Korhonen. High reso- lution image quality database. In ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 3105–3109. IEEE, 2024. 4
2024
-
[30]
Mini-monkey: Alleviating the semantic saw- tooth effect for lightweight MLLMs via complementary im- age pyramid
Mingxin Huang, Yuliang Liu, Dingkang Liang, Lianwen Jin, and Xiang Bai. Mini-monkey: Alleviating the semantic saw- tooth effect for lightweight MLLMs via complementary im- age pyramid. In International Conference on Learning Rep- resentations (ICLR), 2025. 3
2025
-
[31]
The apolloscape dataset for autonomous driving
Xinyu Huang, Xinjing Cheng, Qichuan Geng, Binbin Cao, Dingfu Zhou, Peng Wang, Yuanqing Lin, and Ruigang Yang. The apolloscape dataset for autonomous driving. InProceed- ings of the IEEE conference on computer vision and pattern recognition workshops, pages 954–960, 2018. 3
2018
-
[32]
Gpt-4o system card
Aaron Hurst, Adam Lerer, Adam P Goucher, Adam Perel- man, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Weli- hinda, Alan Hayes, Alec Radford, et al. Gpt-4o system card. arXiv preprint arXiv:2410.21276, 2024. 2, 6
2024 arXiv
-
[33]
Multi-source multi-scale counting in extremely dense crowd images
Haroon Idrees, Imran Saleemi, Cody Seibert, and Mubarak Shah. Multi-source multi-scale counting in extremely dense crowd images. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2547–2554,
-
[34]
Composition loss for counting, density map estima- tion and localization in dense crowds
Haroon Idrees, Muhmmad Tayyab, Kishan Athrey, Dong Zhang, Somaya Al-Maadeed, Nasir Rajpoot, and Mubarak Shah. Composition loss for counting, density map estima- tion and localization in dense crowds. In Proceedings of the European conference on computer vision (ECCV), pages 53...
2018
-
[35]
Cosmoclip: Generaliz- ing large vision-language models for astronomical imaging
Raza Imam, Mohammed Talha Alam, Umaima Rahman, Mohsen Guizani, and Fakhri Karray. Cosmoclip: Generaliz- ing large vision-language models for astronomical imaging. arXiv preprint arXiv:2407.07315, 2024. 1
2024 arXiv
-
[36]
Mmad: The first-ever comprehensive benchmark for mul- timodal large language models in industrial anomaly detec- tion
Xi Jiang, Jian Li, Hanqiu Deng, Yong Liu, Bin-Bin Gao, Yifeng Zhou, Jialin Li, Chengjie Wang, and Feng Zheng. Mmad: The first-ever comprehensive benchmark for mul- timodal large language models in industrial anomaly detec- tion. arXiv preprint arXiv:2410.09453, 2024. 4
-
[37]
A diagram is 10 worth a dozen images
Aniruddha Kembhavi, Mike Salvato, Eric Kolve, Minjoon Seo, Hannaneh Hajishirzi, and Ali Farhadi. A diagram is 10 worth a dozen images. In Computer Vision–ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11–14, 2016, Proceedings, Part IV 14 , pages 235–
2016
-
[38]
Berg, Wan-Yen Lo, Piotr Doll ´ar, and Ross B
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chlo´e Rolland, Laura Gustafson, Tete Xiao, Spencer White- head, Alexander C. Berg, Wan-Yen Lo, Piotr Doll ´ar, and Ross B. Girshick. Segment anything. In International Con- ference on Computer Vision (ICCV), 2023. 3
2023
-
[39]
J. Lau, S. Gayen, A. Ben Abacha, and D. Demner-Fushman. A dataset of clinically generated visual questions and an- swers about radiology images. Scientific Data , 5:180251,
-
[40]
A dataset of clinically generated visual questions and answers about radiology images
Jason J Lau, Soumya Gayen, Asma Ben Abacha, and Dina Demner-Fushman. A dataset of clinically generated visual questions and answers about radiology images. Scientific data, 5(1):1–10, 2018. 1
2018
-
[41]
Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation, 2022
Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation, 2022. 2
2022
-
[42]
Hrvqa: A visual question answering benchmark for high-resolution aerial images
Kun Li, George V osselman, and Michael Ying Yang. Hrvqa: A visual question answering benchmark for high-resolution aerial images. ISPRS Journal of Photogrammetry and Re- mote Sensing, 214:65–81, 2024. 4
2024
-
[43]
To- wards streaming perception
Mengtian Li, Yu-Xiong Wang, and Deva Ramanan. To- wards streaming perception. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23– 28, 2020, Proceedings, Part II 16, pages 473–488. Springer,
2020
-
[44]
Mon- key: Image resolution and text label are important things for large multi-modal models
Zhang Li, Biao Yang, Qiang Liu, Zhiyin Ma, Shuo Zhang, Jingxu Yang, Yabo Sun, Yuliang Liu, and Xiang Bai. Mon- key: Image resolution and text label are important things for large multi-modal models. In IEEE/CVF Computer Vision and Pattern Recognition Conference (CVPR), 2024. 3
2024
-
[45]
The artbench dataset: Benchmarking generative models with art- works
Peiyuan Liao, Xiuyu Li, Xihui Liu, and Kurt Keutzer. The artbench dataset: Benchmarking generative models with art- works. arXiv preprint arXiv:2206.11404, 2022. 4
2022 arXiv
-
[46]
Visual instruction tuning
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning. Advances in neural information processing systems, 36:34892–34916, 2023. 6, 3
2023
-
[47]
Visual instruction tuning
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning. In Conference on Neural Informa- tion Processing Systems (NeurIPS), 2023. 2
2023
-
[48]
Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang
Nelson F. Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang. Lost in the middle: How language models use long contexts. Trans- actions of the Association for Computational Linguistics, 12: 157–173, 2024. 1, 3, 5, 7, 2
2024
-
[49]
A convnet for the 2020s
Zhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feicht- enhofer, Trevor Darrell, and Saining Xie. A convnet for the 2020s. In IEEE/CVF Computer Vision and Pattern Recogni- tion Conference (CVPR), 2022. 2
2022
-
[50]
Deepseek-vl: towards real-world vision- language understanding
Haoyu Lu, Wen Liu, Bo Zhang, Bingxuan Wang, Kai Dong, Bo Liu, Jingxiang Sun, Tongzheng Ren, Zhuoshu Li, Hao Yang, et al. Deepseek-vl: towards real-world vision- language understanding. arXiv preprint arXiv:2403.05525,
-
[51]
Feast your eyes: Mixture- of-resolution adaptation for multimodal large language mod- els
Gen Luo, Yiyi Zhou, Yuxin Zhang, Xiawu Zheng, Xi- aoshuai Sun, and Rongrong Ji. Feast your eyes: Mixture- of-resolution adaptation for multimodal large language mod- els. In International Conference on Learning Representa- tions (ICLR), 2025. 2, 6
2025
-
[52]
Infographicvqa
Minesh Mathew, Viraj Bagal, Rub `en Tito, Dimosthenis Karatzas, Ernest Valveny, and CV Jawahar. Infographicvqa. In Proceedings of the IEEE/CVF Winter Conference on Ap- plications of Computer Vision , pages 1697–1706, 2022. 1, 4
2022
-
[53]
MM1: methods, analysis and insights from multimodal LLM pre-training
Brandon McKinzie, Zhe Gan, Jean-Philippe Fauconnier, Sam Dodge, Bowen Zhang, Philipp Dufter, Dhruti Shah, Xi- anzhi Du, Futang Peng, Anton Belyi, Haotian Zhang, Karan- jeet Singh, Doug Kang, Hongyu H `e, Max Schwarzer, Tom Gunter, Xiang Kong, Aonan Zhang, Jianyu Wang, Chong Wa...
2024
-
[54]
The mame dataset: on the relevance of high resolution and variable shape image properties.Applied Intelligence, 52(10):11703–11724, 2022
Ferran Par ´es, Anna Arias-Duart, Dario Garcia-Gasulla, Gema Campo-Franc ´es, Nina Viladrich, Eduard Ayguad ´e, and Jes ´us Labarta. The mame dataset: on the relevance of high resolution and variable shape image properties.Applied Intelligence, 52(10):11703–11724, 2022. 4
2022
-
[55]
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning transferable visual models from natural language supervision. In International Con...
2021
-
[56]
The synthia dataset: A large collection of synthetic images for semantic segmentation of urban scenes
German Ros, Laura Sellart, Joanna Materzynska, David Vazquez, and Antonio M Lopez. The synthia dataset: A large collection of synthetic images for semantic segmentation of urban scenes. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 3234–3243,
-
[57]
Jhu-crowd++: Large-scale crowd counting dataset and a benchmark method
Vishwanath A Sindagi, Rajeev Yasarla, and Vishal M Pa- tel. Jhu-crowd++: Large-scale crowd counting dataset and a benchmark method. IEEE transactions on pattern analysis and machine intelligence, 44(5):2594–2609, 2020. 3
2020
-
[58]
Milebench: Benchmark- ing mllms in long context
Dingjie Song, Shunian Chen, Guiming Hardy Chen, Fei Yu, Xiang Wan, and Benyou Wang. Milebench: Benchmark- ing mllms in long context. arXiv preprint arXiv:2404.18532,
-
[59]
Gemini: a family of highly capable multimodal models
Gemini Team, Rohan Anil, Sebastian Borgeaud, Jean- Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, Katie Millican, et al. Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805, 2023. 1, 2
2023 arXiv
-
[60]
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of con- text
Gemini Team, Petko Georgiev, Ving Ian Lei, Ryan Burnell, Libin Bai, Anmol Gulati, Garrett Tanzer, Damien Vincent, Zhufeng Pan, Shibo Wang, et al. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of con- text. arXiv preprint arXiv:2403.05530, 2024. 2, 6, 3
2024 arXiv
-
[61]
Predicting breast tumor proliferation from whole-slide im- ages: the tupac16 challenge
Mitko Veta, Yujing J Heng, Nikolas Stathonikos, Babak Eht- eshami Bejnordi, Francisco Beca, Thomas Wollmann, Karl 11 Rohr, Manan A Shah, Dayong Wang, Mikael Rousson, et al. Predicting breast tumor proliferation from whole-slide im- ages: the tupac16 challenge. Medical image an...
2019
-
[62]
Muirbench: A comprehensive bench- mark for robust multi-image understanding
Fei Wang, Xingyu Fu, James Y Huang, Zekun Li, Qin Liu, Xiaogeng Liu, Mingyu Derek Ma, Nan Xu, Wenxuan Zhou, Kai Zhang, et al. Muirbench: A comprehensive bench- mark for robust multi-image understanding. arXiv preprint arXiv:2406.09411, 2024. 4
2024 arXiv
-
[64]
Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution
Peng Wang, Shuai Bai, Sinan Tan, Shijie Wang, Zhihao Fan, Jinze Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, et al. Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution. arXiv preprint arXiv:2409.12191, 2024. 1
2024 arXiv
-
[65]
Divide, conquer and combine: A training-free framework for high-resolution im- age perception in multimodal large language models
Wenbin Wang, Liang Ding, Minyan Zeng, Xiabin Zhou, Li Shen, Yong Luo, and Dacheng Tao. Divide, conquer and combine: A training-free framework for high-resolution im- age perception in multimodal large language models. arXiv preprint, 2024. 1, 2
2024
-
[66]
Divide, conquer and combine: A training-free framework for high-resolution im- age perception in multimodal large language models
Wenbin Wang, Liang Ding, Minyan Zeng, Xiabin Zhou, Li Shen, Yong Luo, and Dacheng Tao. Divide, conquer and combine: A training-free framework for high-resolution im- age perception in multimodal large language models. arXiv preprint arXiv:2408.15556, 2024. 4
2024 arXiv
-
[67]
Needle in a multimodal haystack
Weiyun Wang, Shuibo Zhang, Yiming Ren, Yuchen Duan, Tiantong Li, Shuo Liu, Mengkang Hu, Zhe Chen, Kaipeng Zhang, Lewei Lu, et al. Needle in a multimodal haystack. Advances in Neural Information Processing Systems , 37: 20540–20565, 2025. 1, 2, 3
2025
-
[68]
Panda: A gigapixel- level human-centric video dataset
Xueyang Wang, Xiya Zhang, Yinheng Zhu, Yuchen Guo, Xiaoyun Yuan, Liuyu Xiang, Zerun Wang, Guiguang Ding, David Brady, Qionghai Dai, et al. Panda: A gigapixel- level human-centric video dataset. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognitio...
2020
-
[69]
Weiser, P.L
E.L. Weiser, P.L. Flint, D.K. Marks, B.S. Shults, H.M. Wil- son, S.J. Thompson, and J.B. Fischer. Aerial photo imagery from fall waterfowl surveys, izembek lagoon, alaska, 2017-
2017
-
[70]
V?: Guided visual search as a core mechanism in multimodal llms
Penghao Wu and Saining Xie. V?: Guided visual search as a core mechanism in multimodal llms. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13084–13094, 2024. 4
2024
-
[71]
Visual haystacks: A vision-centric needle-in-a- haystack benchmark
Tsung-Han Wu, Giscard Biamby, Jerome Quenum, Ritwik Gupta, Joseph E Gonzalez, Trevor Darrell, and David M Chan. Visual haystacks: A vision-centric needle-in-a- haystack benchmark. arXiv preprint arXiv:2407.13766 ,
-
[73]
Deepseek-vl2: Mixture-of- experts vision-language models for advanced multimodal understanding
Zhiyu Wu, Xiaokang Chen, Zizheng Pan, Xingchao Liu, Wen Liu, Damai Dai, Huazuo Gao, Yiyang Ma, Chengyue Wu, Bingxuan Wang, et al. Deepseek-vl2: Mixture-of- experts vision-language models for advanced multimodal understanding. arXiv preprint arXiv:2412.10302, 2024. 2
2024 arXiv
-
[74]
Deep learning for detecting corona virus disease 2019 (covid-19) on high-resolution computed tomography: a pilot study
Shuyi Yang, Longquan Jiang, Zhuoqun Cao, Liya Wang, Jiawang Cao, Rui Feng, Zhiyong Zhang, Xiangyang Xue, Yuxin Shi, and Fei Shan. Deep learning for detecting corona virus disease 2019 (covid-19) on high-resolution computed tomography: a pilot study. Annals of translational med...
2019
-
[75]
Bdd100k: A diverse driving dataset for heterogeneous multitask learning
Fisher Yu, Haofeng Chen, Xin Wang, Wenqi Xian, Yingying Chen, Fangchen Liu, Vashisht Madhavan, and Trevor Dar- rell. Bdd100k: A diverse driving dataset for heterogeneous multitask learning. In Proceedings of the IEEE/CVF con- ference on computer vision and pattern recognition ...
2020
-
[76]
Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for ex- pert agi
Xiang Yue, Yuansheng Ni, Kai Zhang, Tianyu Zheng, Ruoqi Liu, Ge Zhang, Samuel Stevens, Dongfu Jiang, Weiming Ren, Yuxuan Sun, et al. Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for ex- pert agi. In Proceedings of the IEEE/CVF Conference on...
2024
-
[77]
Single-image crowd counting via multi-column convolutional neural network
Yingying Zhang, Desen Zhou, Siqin Chen, Shenghua Gao, and Yi Ma. Single-image crowd counting via multi-column convolutional neural network. In 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 589–597, 2016. 3
2016
-
[78]
Video instruction tuning with synthetic data
Yuanhan Zhang, Jinming Wu, Wei Li, Bo Li, Zejun Ma, Zi- wei Liu, and Chunyuan Li. Video instruction tuning with synthetic data. arXiv preprint arXiv:2410.02713, 2024. 6, 3
2024 arXiv
-
[79]
Mme-realworld: Could your multimodal llm challenge high-resolution real-world scenarios that are difficult for humans? arXiv preprint arXiv:2408.13257, 2024
Yi-Fan Zhang, Huanyu Zhang, Haochen Tian, Chaoyou Fu, Shuangqing Zhang, Junfei Wu, Feng Li, Kun Wang, Qing- song Wen, Zhang Zhang, et al. Mme-realworld: Could your multimodal llm challenge high-resolution real-world scenarios that are difficult for humans? arXiv preprint arXiv...
2024 arXiv
-
[81]
Monitoring Extracted from MME-Realworld, this dataset features images taken from public safety cameras in diverse scenarios
Dataset Details Autonomous Driving We extract samples from the MME- Realworld dataset to evaluate a model’s embodied intelli- gence, focusing on perception tasks such as distant object perception, attribute recognition, and counting, as well as reasoning tasks including intent...
-
[82]
Ablation on Lost-in-the-Middle To test whether the observed U-shape is a trivial extension of existing work [48], we further evaluate the model with flattened distance, the metric used in the original lost-in- the-middle that measures the linear distance of the starting token ...
-
[83]
Phi- 3.5 [2] is a lightweight model designed for efficient lan- guage understanding and generation
Model Details We include a total of 28 models in our experiment. Phi- 3.5 [2] is a lightweight model designed for efficient lan- guage understanding and generation. We include Phi 3.5 vision instruct [2] for experiments. DeepSeek Janus Pro 7B [9] is a model that integrates mul...
-
[84]
The scores are the average per- formance of all samples in val, test, testmini splits
Performance Details on Real-world Datasets Table 8 and Table 9 display the performance of all VLMs on every real-world dataset. The scores are the average per- formance of all samples in val, test, testmini splits. 3 Table 8. Performance of all VLMs on every real-world dataset...
-
[85]
question n Give an answer with this format: <ans>ANSWER</ans>, no redundant words. For example: <ans>A</ans>
Prompts and Metrics For ComplexGrid dataset, our prompt is “The image is composed of multiple sub-images. The left upper corner is row 1 column 1. We also add the row and column numbers under each image. You need to identify the sub-image that best suits the caption: {caption}...
-
[86]
We compress the images to display them in the paper
Examples Table 10 to 32 show examples from HRScene real-world datasets. We compress the images to display them in the paper. Table 10. Example from HRScene – ArtBench The painting in the picture belongs to which of the following categories? (A) Surrealism (B) Expressionism (C)...
-
[2019]
Geological Survey data release, 2022
U.S. Geological Survey data release, 2022. 4
2022
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.