Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T10:22:52.691595Z
Paper Citation Record · LEDGER
As of 19 August 2026, this Paper Citation Record lists 84 of 84 outbound references and 0 inbound Pith citation observations for arXiv:2504.18406.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T10:22:52.691595Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
84 of 84 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation f9b1899d-2417-4785-8d5f-28b113856ef8 · outbound
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? 12 Kucherlapati Raju 13, Genome data analysis: Baylor College of Medicine Creighton Chad J
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d0c885f-e035-4a56-9bb4-5492d0030c35 · outbound
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1175f303-e6a7-4395-97ef-3c9cd122ee9e · outbound
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? GPT-4 Technical Report
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cbc5567b-d8f4-429e-9f34-44a9c47784e2 · outbound
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? The claude 3 model family: Opus, sonnet, haiku
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15c957c5-8cd7-4cd8-bd1a-b94bde5faa64 · outbound
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Bach: Grand challenge on breast cancer histology im- ages
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd92b8e1-1e15-4332-b442-9ff0bc14226d · outbound
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Jauregui, and Juan Andr ´es Cardoso
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3a4f4f7-944b-474c-995c-f7149645fcae · outbound
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Ef- ficient high-resolution deep learning: A survey
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 19bb4625-5def-4fbf-9d05-ce0b6ba78736 · outbound
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Effi- cient high-resolution deep learning: A survey.ACM Comput
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation e944eaa2-510a-49db-9955-3c21580f39a6 · outbound
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 888b5071-0d94-441c-bd07-3b0a4ffc64c3 · outbound
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6070e0a-7b27-4f7e-a792-7a5bf964a236 · outbound
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e97359c0-4da3-485b-8341-aa52329fbaa5 · outbound
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 16647775-f89e-440e-b7ef-ec1659f534dd · outbound
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Functional map of the world
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation c6c8c29c-eb5e-4f8d-89e7-bfb21486adfd · outbound
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? The cityscapes dataset for semantic urban scene understanding
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 57a333d4-255e-4b52-86fb-ddd5aea2bf8f · outbound
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e34ee8a0-b7f5-476e-b3ed-e84af2704903 · outbound
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? LungHist700: A dataset of histological images for deep learning in pulmonary pathology
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation b299ab0e-d64d-4bbd-87ba-8cadd9c49f01 · outbound
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? LungHist700: A dataset of histological images for deep learning in pulmonary pathology
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ab042208-0d94-4f79-8ecc-e81ad17c7e34 · outbound
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Internlm-xcomposer2-4khd: A pioneer- ing large vision-language model handling resolutions from 336 pixels to 4k HD
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation a263f2d1-568b-4dbf-b65f-3d6ffda9c112 · outbound
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? The Llama 3 Herd of Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3543c1bb-b91e-41ad-a5a0-6f8843f862db · outbound
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Describing differences in image sets with natural language
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation f65c327e-f379-4aad-89a6-efc6604d2669 · outbound
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Floorplancad: A large-scale cad draw- ing dataset for panoptic symbol spotting
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation b33ecc19-b636-4868-807b-f9f156c0d2a3 · outbound
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67467734-d1fc-46cc-8cea-14e87c0e9ea4 · outbound
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 5dae7f81-ff1f-428d-8965-120282bb8756 · outbound
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Llava-uhd: An LMM perceiving any aspect ratio and high- resolution images
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation a7cb86e0-1b89-49ee-8103-6ee52159a44f · outbound
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Cogagent: A visual lan- guage model for GUI agents
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation e71e3d52-d4ac-49ac-bb8c-d59361ed0628 · outbound
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? mplug-docowl 1.5: Unified structure learning for ocr-free document understanding
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation c82a385b-598a-4c8f-87ef-8d3fcbdbcdda · outbound
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 694fe1a0-1e68-4e9d-a306-5ab338b52b13 · outbound
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Novachart: A large- scale dataset towards chart understanding and generation of multimodal large language models
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 320bb115-4ebc-421b-a50b-ed91ba62d796 · outbound
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? High reso- lution image quality database
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation dbd2eaae-922f-4a31-b77a-d201cdb248a9 · outbound
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Mini-monkey: Alleviating the semantic saw- tooth effect for lightweight MLLMs via complementary im- age pyramid
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation a3c99a96-fffb-4241-a0cb-e6cceaf6ca5c · outbound
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? The apolloscape dataset for autonomous driving
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation f5021136-94e3-4344-99ca-29f1255b301d · outbound
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? GPT-4o System Card
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47e467f5-93b5-46f2-8aa4-3c00497d8f1d · outbound
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Multi-source multi-scale counting in extremely dense crowd images
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 7771275f-05a5-40d4-b808-9a8c39b60bd7 · outbound
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Composition loss for counting, density map estima- tion and localization in dense crowds
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation bbdc62ea-4066-464f-9852-4d5b711ee143 · outbound
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? CosmoCLIP: Generalizing Large Vision-Language Models for Astronomical Imaging
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af02aa70-aabd-4317-b138-8e2890980087 · outbound
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? MMAD: A Comprehensive Benchmark for Multimodal Large Language Models in Industrial Anomaly Detection
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f97d40ff-5c0e-4917-8a81-08ae1b554f3b · outbound
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? A diagram is 10 worth a dozen images
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ef01dd84-2979-419b-b91b-f2006978501c · outbound
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Berg, Wan-Yen Lo, Piotr Doll ´ar, and Ross B
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 652cbbda-5d7a-4ae4-b779-259e86ac5003 · outbound
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Unresolved cited work
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ce38ad38-cf11-496e-b34a-23bafbf7087c · outbound
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? A dataset of clinically generated visual questions and answers about radiology images
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation dd87e033-17dd-40d5-a7ac-b5ca145bb28e · outbound
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation, 2022
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 090ce644-3b98-469d-9077-51d19d9bb127 · outbound
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Hrvqa: A visual question answering benchmark for high-resolution aerial images
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 5ea0ebcb-a443-4f56-bce3-67b7cd2542a7 · outbound
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? To- wards streaming perception
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation e02da5d8-5072-4c2c-b1f7-49c0d86fa976 · outbound
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Mon- key: Image resolution and text label are important things for large multi-modal models
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 8834e4b0-29c1-49bb-ac17-2cf49e132b87 · outbound
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? The ArtBench Dataset: Benchmarking Generative Models with Artworks
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 702d130a-c5fd-47a4-bb23-c923f4826bf1 · outbound
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Visual instruction tuning
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation b662ac6b-21e7-47e4-91a6-3b735978b878 · outbound
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Visual instruction tuning
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 5c0f2fe2-da3d-4828-b7db-7357e9005fd2 · outbound
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation f5d8e1cc-4b44-4e31-8afd-14b7e3d78322 · outbound
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? A convnet for the 2020s
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation f3b305db-fadf-4584-b15e-fa91c6b88242 · outbound
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? DeepSeek-VL: Towards Real-World Vision-Language Understanding
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 997a93e9-ecbe-44a5-9a1d-bf4f0c1d7d36 · outbound
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Feast your eyes: Mixture- of-resolution adaptation for multimodal large language mod- els
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 432380a8-118f-4daf-b63d-c7788b99631b · outbound
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Infographicvqa
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 304adae2-21fe-4aed-8c1e-cadeca871c33 · outbound
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? MM1: methods, analysis and insights from multimodal LLM pre-training
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 36d208a9-8be8-4387-b4e5-9c87de2704f5 · outbound
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? The mame dataset: on the relevance of high resolution and variable shape image properties.Applied Intelligence, 52(10):11703–11724, 2022
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 6d458e48-6f4a-4836-86b6-9d15bbb56e64 · outbound
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Learning transferable visual models from natural language supervision
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 418f9234-1f9a-48e7-b5ec-48bed3326402 · outbound
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? The synthia dataset: A large collection of synthetic images for semantic segmentation of urban scenes
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6dbd83c3-87e4-46ae-bc16-6a5df342a570 · outbound
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Jhu-crowd++: Large-scale crowd counting dataset and a benchmark method
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 2effe67c-8df2-4ba9-be1d-f9082195663a · outbound
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? MileBench: Benchmarking MLLMs in Long Context
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2249752f-28cf-48b5-afd6-ee52df310d5b · outbound
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Gemini: A Family of Highly Capable Multimodal Models
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0073b7ca-a1b8-42e1-bd00-913a855c079c · outbound
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8d37346-5381-4628-8b22-f894d82d57e9 · outbound
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Predicting breast tumor proliferation from whole-slide im- ages: the tupac16 challenge
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 3581a251-a555-4e0f-81f8-bd8810a09e88 · outbound
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8038171f-0e4d-4270-88fd-fe9b7f7436c3 · outbound
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99e3a955-7b77-43ed-a0d4-4de8baafce46 · outbound
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Divide, conquer and combine: A training-free framework for high-resolution im- age perception in multimodal large language models
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 99dd5e4f-374c-4d32-bc37-8bb84f7c2773 · outbound
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Divide, Conquer and Combine: A Training-Free Framework for High-Resolution Image Perception in Multimodal Large Language Models
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4cab3b4-c321-46b0-9e4b-28dae3ee7bd5 · outbound
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Needle in a multimodal haystack
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation eef52f33-274d-440e-b097-3312e6487ebe · outbound
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Panda: A gigapixel- level human-centric video dataset
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 57526d2d-fb65-4948-b8c8-6d0d0a0fe354 · outbound
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Weiser, P.L
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 2c891846-3d79-4421-aae7-cdf56950113c · outbound
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? V?: Guided visual search as a core mechanism in multimodal llms
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation d38f6978-372a-4e46-914f-b76d9b486cff · outbound
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Visual Haystacks: A Vision-Centric Needle-In-A-Haystack Benchmark
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88fce6b9-61f5-46f8-b8c4-77343d86eccb · outbound
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8fdb897-c59f-4dad-b020-0499b7480244 · outbound
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Deep learning for detecting corona virus disease 2019 (covid-19) on high-resolution computed tomography: a pilot study
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 319de3f4-0f78-4693-ad79-70d0235c154c · outbound
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Bdd100k: A diverse driving dataset for heterogeneous multitask learning
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation a3770264-13b9-43f9-a15f-822ebf6caee3 · outbound
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for ex- pert agi
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6e95b13-d20a-4042-9b3e-21ee90b4f30c · outbound
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Single-image crowd counting via multi-column convolutional neural network
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 4ff3db0e-5e4a-45ea-8d17-2c8a00b98a54 · outbound
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce0ae181-b1cf-4695-9ff0-773db659f6a4 · outbound
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ad8a1ae-f2b4-4ff2-accb-87d5da2029f5 · outbound
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Monitoring Extracted from MME-Realworld, this dataset features images taken from public safety cameras in diverse scenarios
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation f10ccd4a-4de5-4ce5-a358-e5474c129936 · outbound
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Unresolved cited work
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation c18d8af0-22b6-4773-8268-b177c9265cec · outbound
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Phi- 3.5 [2] is a lightweight model designed for efficient lan- guage understanding and generation
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 4439af5e-c13a-4957-b8fe-e865bc6755f0 · outbound
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? The scores are the average per- formance of all samples in val, test, testmini splits
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ea35b337-831a-47a2-a900-74d0eeac204d · outbound
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? question n Give an answer with this format: <ans>ANSWER</ans>, no redundant words. For example: <ans>A</ans>
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 4b9f70eb-5177-4308-8a37-d78e5fd7e01a · outbound
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? We compress the images to display them in the paper
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation a0c9f485-4bef-49bc-b1b0-d659d904e87f · outbound
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding? Geological Survey data release, 2022
Reference 2019
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
No inbound Pith citation observations are available.