Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:40:57.963146Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 74 of 74 outbound references and 8 inbound Pith citation observations for arXiv:2506.01663.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:40:57.963146Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-02T13:24:17.538850Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T15:09:55.338873Z
74 of 74 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 45c38f8e-d3ce-48bd-ba4b-8707c222c4be · outbound
Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Phi-3 technical report: A highly capable language model locally on your phone
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 21388ee8-55fa-40a8-a688-794109a643c0 · outbound
Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Qwen technical report
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 76b437c0-1477-4cef-99da-25e08e82863f · outbound
Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Qwen2.5-VL Technical Report
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3692d56d-5932-4216-8d21-bc4b56c42684 · outbound
Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement DualFocus: Integrating Macro and Micro Perspectives in Multi-modal Large Language Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 334684bc-b081-41e7-ac79-568dc72da4c9 · outbound
Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4887c136-36cd-4ac4-abbc-6bf9f3025d34 · outbound
Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Sharegpt4v: Improving large multi-modal models with better captions
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 665676c7-b8e3-4351-8709-731161ae88a9 · outbound
Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32bca65c-6b13-4103-9b7f-e8af1db770fa · outbound
Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3d43ec70-be5b-486a-8192-5cd6012f69af · outbound
Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4d04e08-e5a8-4602-bb1a-6b7f10711d09 · outbound
Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Chain-of-verification reduces hallucination in large language models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f61066cc-98dd-4a66-8559-fed13358bda7 · outbound
Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Internlm-xcomposer2: Mastering free-form text-image composition and comprehension in vision-language large model
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0013e321-7f3e-43e4-b1af-7a9c19f08130 · outbound
Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Unresolved cited work
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7294c1b8-4b96-4bbf-b770-25084a9a1065 · outbound
Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Active vision: The psychology of looking and seeing
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7586c170-250b-4a29-ba4a-6c91de48c5ee · outbound
Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Small language model can self-correct
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d8daadd2-8b5b-4aae-a3e5-1c645ef670c0 · outbound
Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Eye movements in natural behavior
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c30b6367-8779-4ae2-a245-70b18d5a1b9f · outbound
Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Human gaze control during real-world scene perception.Trends in cognitive sciences, 7(11):498–504, 2003
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 33795074-5261-4f2d-a544-74a39440397d · outbound
Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement CogVLM2: Visual Language Models for Image and Video Understanding
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25624c85-eccc-45a0-a88a-5a25ead4f616 · outbound
Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3b73bc6-8def-4a2e-a70e-bd96ec54c53e · outbound
Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6bd05da0-5767-45c1-a096-ec6091e15086 · outbound
Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Ma- tryoshka query transformer for large vision-language models
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5d96b365-9739-4cf7-8d02-28482aa219f5 · outbound
Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Aggregate-and-adapt natural language prompts for downstream generalization of clip
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9b0380ac-9b30-4c9e-8040-fa3c11b07d07 · outbound
Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Large language models cannot self-correct reasoning yet
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21a94415-319b-4bb7-a817-e02e381b4212 · outbound
Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement GPT-4o System Card
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0392638f-e796-4657-9068-5aea7dfe2cac · outbound
Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Maven: An effective multi-granularity hybrid visual encoding framework for multimodal large language model
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 512ac1f3-7669-4d76-a8f4-3eb2c4078c43 · outbound
Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement When can LLMs actually correct their own mistakes? a critical survey of self-correction of LLMs
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 72cb6998-059e-4966-8f19-ba19b769d6a5 · outbound
Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Language models can solve computer tasks
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 287cf32b-7167-4fbb-94c5-96fe62a9704e · outbound
Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Training language models to self-correct via reinforcement learning
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8b5917de-feb5-493f-a3f4-011f9e3d746a · outbound
Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement What matters when building vision-language models? Advances in Neural Information Processing Systems, 37:87874–87907, 2024
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 686006e4-f2dd-40fa-b86d-8ae837c87679 · outbound
Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Llava-next: Stronger llms supercharge multimodal capabilities in the wild
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4634c676-34a9-4239-8107-53d9cc6d8430 · outbound
Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Otter: A Multi-Modal Model with In-Context Instruction Tuning
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca30a7ad-6ad2-4fd2-92f0-9978ef0cbcd1 · outbound
Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement DyFo: A Training-Free Dynamic Focus Visual Search for Enhancing LMMs in Fine-Grained Visual Understanding
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a47c34b-7eec-4c16-82b3-8439b44816ff · outbound
Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Cumo: Scaling multimodal llm with co-upcycled mixture-of-experts
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b92f1c4f-9973-4e2e-97ba-4a759a90d28b · outbound
Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Hrvqa: A visual question answering benchmark for high-resolution aerial images
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f22effdc-6b0a-420f-ab2e-5fbf73395dd8 · outbound
Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement When hindsight is not 20/20: Testing limits on reflective thinking in large language models
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e62d103a-67ea-41fc-bcea-d92ef42eb996 · outbound
Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Mini-gemini: Mining the potential of multi-modality vision language models
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8c015068-9151-440b-b307-0fba713fcd06 · outbound
Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2d1b6cf-dd28-42b1-ad6a-2b21150e6076 · outbound
Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Vila: On pre-training for visual language models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91980365-2df3-49ee-a963-527143961fc3 · outbound
Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Visual anchors are strong information aggregators for multimodal large language model
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 93631198-2d5c-4fcf-beb6-697d249952ac · outbound
Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Improved baselines with visual instruction tuning
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f3c7515-4b67-40c4-9bf5-5d175824de0e · outbound
Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Llava-next: improved reasoning, ocr, and world knowledge (2024)
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 00e4b129-d5d7-472c-aba8-8778e2700f68 · outbound
Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation afe2df36-233b-4b4b-8ef6-4474a7ff8dfb · outbound
Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94f934a8-1323-4fa7-93f5-d281f039bb05 · outbound
Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement DeepSeek-VL: Towards Real-World Vision-Language Understanding
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a24f198-bdbf-43fe-b694-f2d496b58f50 · outbound
Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Task-to-instance prompt learning for vision-language models at test time
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c86508c6-2314-49f8-aa28-9a37f1f5abea · outbound
Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement TextCoT: Zoom In for Enhanced Multimodal Text-Rich Image Understanding
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2867ea04-cbad-4032-8c42-81fbca559218 · outbound
Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Feast your eyes: Mixture-of-resolution adaptation for multimodal large language models
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation abb0e6ed-2781-46e2-b6a9-a5564a20848d · outbound
Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Self-refine: Iterative refinement with self-feedback
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a1d21a7-d9c7-4166-b330-aaf67753a2ba · outbound
Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Gpt-4v(ision) system card
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 65f80211-1a44-4583-b83f-d1d8bb65a67b · outbound
Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Automatically correcting large language models: Surveying the landscape of diverse automated correction strategies
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4455c1e7-ce09-4d99-a68d-c76918055fbc · outbound
Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Orienting of attention
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5b6a1295-b337-4ebf-b08c-5bc9c626abce · outbound
Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Zoomer: Adaptive image focus optimization for black-box mllm
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ae2e871-eccd-4eeb-95d6-0cdc5fddfa96 · outbound
Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Self-Reflection in LLM Agents: Effects on Problem-Solving Performance
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74a5a5f2-1872-4cd6-9c98-002717c1c428 · outbound
Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Visual cot: Advancing multi-modal language models with a comprehen- sive dataset and benchmark for chain-of-thought reasoning
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4440d2ba-e57c-45c3-94bf-08cbf63fdc6d · outbound
Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement ZoomEye: Enhancing Multimodal LLMs with Human-Like Zooming Capabilities through Tree-Based Image Exploration
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02828600-11bf-4813-bdbd-0c7d797acf12 · outbound
Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Reflexion: Language agents with verbal reinforcement learning.Advances in Neural Information Processing Systems, 36:8634–8652, 2023
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66162ec5-8ac7-4456-aadc-a8731e1eb92b · outbound
Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Gemini: A Family of Highly Capable Multimodal Models
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0fbc42b6-179f-4d12-882c-11916fca7544 · outbound
Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Cambrian-1: A fully open, vision-centric exploration of multimodal llms
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7416fb6-828b-4b77-929d-edacb27efd05 · outbound
Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement LLMs cannot find reasoning errors, but can correct them given the error location
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 71067d3b-6c02-484e-9bcd-25fc548b3be8 · outbound
Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Lever- aging visual tokens for extended text contexts in multi-modal learning
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d9328b18-c20f-4bdd-9f52-cda4009c3eda · outbound
Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement The dark days are overcast: Iron-bearing clouds on HD 209458 b and WASP-43 b can explain low dayside albedos
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df6afcfc-3157-495e-9f19-70f11b2885d3 · outbound
Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e386970a-a070-4514-8af6-c3a44be88688 · outbound
Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Cogvlm: Visual expert for pretrained language models
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4ac4896-052e-41bf-9282-a694a9887d5a · outbound
Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Divide, conquer and combine: A training-free framework for high-resolution image perception in multimodal large language models
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 01d0dbda-89ac-45ec-a669-c6cc4003d3a5 · outbound
Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Visionllm v2: An end-to-end generalist multimodal large language model for hundreds of vision-language tasks
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6b5bb8f-f7ee-4c4e-a119-32592b238f36 · outbound
Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement V*: Guided visual search as a core mechanism in multimodal llms
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5220909d-b610-42d8-b132-db6725d818af · outbound
Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Large language models can self-correct with key condition verification
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c6cd0429-912b-4186-b656-91bc683d6c52 · outbound
Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Graph- based unsupervised disentangled representation learning via multimodal large language models
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 208476fa-b365-4ac9-acdf-1f97ce4381b6 · outbound
Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement mPLUG-2: A modularized multi-modal foundation model across text, image and video
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 71d2c1b3-0784-455e-a21a-b735f55ceba6 · outbound
Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Yi: Open foundation models by 01
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8f38b95d-d9c2-471f-b671-55de3805bb75 · outbound
Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Learning to Check: Unleashing Potentials for Self-Correction in Large Language Models
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4db93dbb-ef84-4fd0-b971-c203f35a4d55 · outbound
Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6dccc8c-113e-4639-aa31-360a77cd65aa · outbound
Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Beyond llava-hd: Diving into high-resolution large multimodal models
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f0c9d9eb-6e6f-4cae-b789-a64b09108131 · outbound
Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement Wings: Learning multimodal llms without text-only forgetting
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 24a627c5-91d6-4e12-86c9-66262efea867 · outbound
Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9868b5db-7fe2-451f-b3ad-670da9d701f7 · inbound
Perceptual Flow Network for Visually Grounded Reasoning Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 21821802-1710-4230-af71-ed63661cdd3e · inbound
Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Policy Self-Distillation Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e1e983bc-ebd2-4237-8dfa-6156a944dcb8 · inbound
Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Policy Self-Distillation Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 73bd0da4-65a6-45c9-8aea-a993e3567ffe · inbound
Mags-RL: Wearing Multimodal LLMs a Magnifying Glass via Agentic Reinforcement Learning For Complex Scene Reasoning Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ee861032-4467-4f75-a999-7ab0da331cb7 · inbound
DeepLatent: Think with Images via Parallel Latent Visual Reasoning Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8c912615-7240-44c2-bdf9-0a1a1cb58bb8 · inbound
From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement
Reference 138
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6cdee9f2-199e-463c-a468-e0352d01435a · inbound
VisReflect: Latent Visual Reflection for Fine-Grained Perception in Long Visual Context Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ba875588-6208-4359-8884-f766a96d17ba · inbound
Perceive-to-Reason: Decoupling Perception and Reasoning for Fine-Grained Visual Reasoning Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.