Pith. sign in

Paper Citation Record · LEDGER

LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

As of 17 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 43 inbound Pith citation observations for arXiv:2403.11703.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2403.11703 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 43 of 43 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T04:13:03.169164Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T02:36:26.526436Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 7dc71625-7470-4bd5-9191-21555ca9f18b · inbound

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites cites this paper.

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 126

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T20:58:59.266253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-12T20:58:58.849040Z digest=sha256:184e4a2b54ab9baacde50b69f13288f52de5da9a1dd2fd6cc103821bd7658319

Observation 4160050a-63f9-4da2-a747-14b2a3401ff5 · inbound

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output cites this paper.

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 158

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T10:46:28.791681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-17T10:46:28.447347Z digest=sha256:5e34351053bb6ba561df6608fc5e5b588d6b9ad9842c943f66bb66063e716902

Observation 7eb48ea3-aab4-46d7-89a3-843fd6901bcf · inbound

MiniCPM-V: A GPT-4V Level MLLM on Your Phone cites this paper.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 107

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T21:07:32.061335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:939edce60ed8526b458d3c5359842623548ed232b71a5f768826460ad10d3d24

Observation 2a176405-3573-42a4-9762-43da5b515c7e · inbound

MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans? cites this paper.

MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans? LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 64

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T07:59:32.722067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T07:59:32.638758Z digest=sha256:c1f89113aa92095d9ef84ddddf3d326496ff86dd985dbd98216bd8985cb5efcb

Observation bff25321-8cc8-4f57-ac0e-fab80c0b6773 · inbound

PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction cites this paper.

PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T12:12:14.716016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-15T12:12:14.613620Z digest=sha256:a195fe4ee0d1568e6d2e9336384815c424de224a2eb384d7649ad2827691b109

Observation 4c9380e0-8f64-479a-9a71-12683b17e959 · inbound

Advancing Fine-Grained Visual Understanding with Multi-Scale Alignment in Multi-Modal Models cites this paper.

Advancing Fine-Grained Visual Understanding with Multi-Scale Alignment in Multi-Modal Models LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T20:27:20.055623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:27:20.055623Z digest=sha256:8e73b273498a0ed77e332612b079f3d66a2c71695fd811fb5128606b046a6fa7

Observation 6aee29cb-5595-4836-be2c-d468dbc7c14c · inbound

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning cites this paper.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.873349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.873349Z digest=sha256:4c9a2e22d1bb6d3fe4f6ef3bc029da9c06cf3d4ab133de70c39f38a360ce5960

Observation b5ccce93-dd71-4f15-af6e-bde51b409499 · inbound

LMM-driven Semantic Image-Text Coding for Ultra Low-bitrate Learned Image Compression cites this paper.

LMM-driven Semantic Image-Text Coding for Ultra Low-bitrate Learned Image Compression LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T17:00:24.637561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:00:24.637561Z digest=sha256:f82a075a0cb4889a836b8afbb2e3fd56affe3faad74917973ea7c79a67fc5822

Observation adf3e0c0-7b68-4965-945b-5221e8622735 · inbound

FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression cites this paper.

FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T15:28:37.486326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:28:37.486326Z digest=sha256:d603fa89d2f758c4f27837860f21c0ebf1713ccb897782b0df628866e9156b50

Observation 243f9db7-0ee0-452f-875f-e68730c06957 · inbound

Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models cites this paper.

Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:06.921149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:06.921149Z digest=sha256:fa4bb424bfe5ff01d6ff9d47bbf44ff7eb0c0e769aacdb40a894d8bb1eb8892c

Observation e2fd5e40-6ffa-47b7-a945-2e14b621c775 · inbound

GMAI-VL & GMAI-VL-5.5M: A Large Vision-Language Model and A Comprehensive Multimodal Dataset Towards General Medical AI cites this paper.

GMAI-VL & GMAI-VL-5.5M: A Large Vision-Language Model and A Comprehensive Multimodal Dataset Towards General Medical AI LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-12T15:16:27.913904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:16:27.913904Z digest=sha256:e27e98b18a842e6f4eeabe83f62d57920dbeb68f39f4f5fa2dc864f0a36fe13b

Observation ef32f889-edb0-484f-86f7-91a1467a4adb · inbound

Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy cites this paper.

Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T14:21:34.872391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:21:34.872391Z digest=sha256:18f06fdc608266eb5f1dc0c6c0af46d3503c23b2a09670cda91ed3166b67a1f2

Observation 33b7389b-5513-4a1d-9481-c34d3bcb9e98 · inbound

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding cites this paper.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.866411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.866411Z digest=sha256:e9b49331868a2a573d20cb9dbb9d20bad450debc4125c7c36bea7725383b0906

Observation 5647cc14-a07e-43cd-b4f1-8ea8d72e7d65 · inbound

PerLA: Perceptive 3D Language Assistant cites this paper.

PerLA: Perceptive 3D Language Assistant LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-12T05:52:39.162003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:52:39.162003Z digest=sha256:88c4c280a894c1092fb8ee1760ebb974d5eca45c6e8dd326eafe299fecaaf123

Observation d9beaaad-9d20-4101-9fe7-7fbb08b79fcd · inbound

Accelerating Multimodal Large Language Models by Searching Optimal Vision Token Reduction cites this paper.

Accelerating Multimodal Large Language Models by Searching Optimal Vision Token Reduction LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T05:21:10.945578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:21:10.945578Z digest=sha256:d8d6f8e4c9883c7645fa237f92da3a3c26a588326323cc0ad2723da23aa1daed

Observation 54a415b3-819d-41fb-977b-ce08a5b70f4b · inbound

Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs cites this paper.

Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-12T00:57:32.547350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:57:32.547350Z digest=sha256:ce9bc012afde466f02bba3928098c6de62574841d4f17fba4e49158f234a3654

Observation a0472031-8831-4bed-a82c-bec1764fb727 · inbound

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining cites this paper.

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T16:10:29.030128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:10:29.030128Z digest=sha256:11e052271429866c59a34b1d87e06bca45b5868aa63ef8841891ea5ece14023a

Observation dbb3d671-2012-410a-884e-ad8830cd6e8c · inbound

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding cites this paper.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 120

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:08.371099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:08.371099Z digest=sha256:e178ea0eda0bf1cc79ea7e18126f7aab844ec3edb3944bae1c0eaf87be661a97

Observation bd986c29-9567-499f-ab33-9bc948658b71 · inbound

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation cites this paper.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.309295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.309295Z digest=sha256:041323f3f1fe1846843da2791f2c61897ea6e89403c19a3d8e7dd8b802f57190

Observation ca19e749-1942-4d8e-87fb-4d3f901b9cce · inbound

MBQ: Modality-Balanced Quantization for Large Vision-Language Models cites this paper.

MBQ: Modality-Balanced Quantization for Large Vision-Language Models LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T00:21:52.300528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:21:52.300528Z digest=sha256:6f6ba8c5b6e0f2d30d6b4fcc80f4de506418c543179857facf1a805cf9bb3cfe

Observation 6bfcd838-d87c-49f2-bbd7-cf1f88539ebb · inbound

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends cites this paper.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:28.060042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:28.060042Z digest=sha256:3d3d6878138315c4cf871135aff8b013d5f1a96d84272bb4bd19d1a5693ae28a

Observation 97733769-5075-41b0-9ffd-ae8eb34423ac · inbound

Mitigating Hallucinations on Object Attributes using Multiview Images and Negative Instructions cites this paper.

Mitigating Hallucinations on Object Attributes using Multiview Images and Negative Instructions LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T19:28:28.962345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:28:28.962345Z digest=sha256:e0b6fd1c4f7435374fcc19f764eb5b84bd4360df42a74b57bcf7e3c70197e143

Observation 641d1f68-7aca-4ac3-b248-8dcc02586bb5 · inbound

Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models cites this paper.

Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 216

Resolution
unresolved
no resolver link, observed 2026-08-10T18:04:34.961445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:04:34.961445Z digest=sha256:a26b9e737af0c2418f5d697605f1e2d01d1d01bda45ed1f5c1cd96e10a8840c4

Observation c73ca218-7d4b-46d1-83c3-fb6010d0fa22 · inbound

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models cites this paper.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.962910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.962910Z digest=sha256:968c90ee110302ff8d29a934309c772f7368e1482b75ba8d2836cde745550dad

Observation 5b74ac38-e830-468f-a941-ae1a69344cba · inbound

Task-Oriented Semantic Communication in Large Multimodal Models-based Vehicle Networks cites this paper.

Task-Oriented Semantic Communication in Large Multimodal Models-based Vehicle Networks LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T01:01:03.854216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:01:03.854216Z digest=sha256:cdad26b32ce449c99de3d2898493aec4f0219cba534573b0bd003136440469de

Observation 4207f9a5-e042-4533-96ab-33b5cfcdf685 · inbound

RESAnything: Attribute Prompting for Arbitrary Referring Segmentation cites this paper.

RESAnything: Attribute Prompting for Arbitrary Referring Segmentation LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-16T04:13:03.169164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:13:03.169164Z digest=sha256:4ff5ca52b8211cf85f41e8011ae998f03249cb742e1081303d77f9b84fe91251

Observation 2a7fd5c2-e049-4bdb-8a1d-5184c95a36b1 · inbound

Adaptive Chain-of-Focus Reasoning via Dynamic Visual Search and Zooming for Efficient VLMs cites this paper.

Adaptive Chain-of-Focus Reasoning via Dynamic Visual Search and Zooming for Efficient VLMs LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:35:13.220175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-17T05:35:13.118221Z digest=sha256:78f9f681c769216a0a527a406801cdff79b1095e5cdb7cdb491ab89dba47c623

Observation a948cb9a-ed4a-49d7-a9c0-9c532f81ba55 · inbound

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs cites this paper.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:08.624340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:08.624340Z digest=sha256:4d8fe4699a57d82544b7ccc038e993e0ed01a615cdb3c404d0a23ceb43f2163b

Observation d30a2810-ca8f-4506-841d-cab57a499d09 · inbound

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation cites this paper.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:09.966497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:09.966497Z digest=sha256:54c1cbe160a546400a696cf41e397b71cec9dd5206d88882a3954e5908f3cd72

Observation 04f360b7-f5ba-4a69-b6cf-8a1077010caa · inbound

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs cites this paper.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:45.114851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:45.114851Z digest=sha256:bebbafd1b09ed21dd3906d23e29ee13893a8155995cf5479a394df74d267b97c

Observation 8a51f921-683d-4c95-aaa1-d22e69acf980 · inbound

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance cites this paper.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.946036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.946036Z digest=sha256:4d11468b15dce6a4f38241c192d25625159654c0866041b5325edf43bf1949ac

Observation c856420f-0375-4530-bb92-9893575ac663 · inbound

ChartM$^3$: Benchmarking Chart Editing with Multimodal Instructions cites this paper.

ChartM$^3$: Benchmarking Chart Editing with Multimodal Instructions LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T18:00:46.494815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:00:46.494815Z digest=sha256:24a394212c3c4686a0021b10b2992491f772ea2c9f598f28b06b538520065516

Observation 1e2f5763-d0fa-42bd-93e6-b8142f783d72 · inbound

VisReason: A Large-Scale Dataset for Visual Chain-of-Thought Reasoning cites this paper.

VisReason: A Large-Scale Dataset for Visual Chain-of-Thought Reasoning LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-03T20:57:33.650037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:57:33.650037Z digest=sha256:95ae9d47b180ab7f3a7d35168954ca364e5343fab8644bc59dda1490ac99498c

Observation ee24e842-e62e-47ed-ac14-e8fe6817bf62 · inbound

Visual Funnel: Resolving Contextual Blindness in Multimodal Large Language Models cites this paper.

Visual Funnel: Resolving Contextual Blindness in Multimodal Large Language Models LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-16T23:48:41.926907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T23:47:08.562575Z digest=sha256:1b570e5cc337b650ab57d03d14ca9edeada9d65cddce9444a7d5d425c9ebac83

Observation 1877bc03-d0cb-490a-927b-beb65ecab6fa · inbound

E-VLA: Event-Augmented Vision-Language-Action Model for Dark and Blurred Scenes cites this paper.

E-VLA: Event-Augmented Vision-Language-Action Model for Dark and Blurred Scenes LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-13T09:40:35.631188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T09:40:35.631188Z digest=sha256:9c38a5bf0e5487c7309f3c583c98d4f3c13b98cd9df6d9ebf14c6fe239fc5ca9

Observation 653deced-ed06-4e0f-b05e-96b9ec49eb7f · inbound

Less Detail, Better Answers: Degradation-Driven Prompting for VQA cites this paper.

Less Detail, Better Answers: Degradation-Driven Prompting for VQA LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:05:47.957190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T20:17:01.867903Z digest=sha256:0d5e121aca481116794da1dac0ed0cba08f1b50f4d2df9d9cb688a3b187a7c50

Observation 7d1d40ee-1341-44cb-abad-283d687a1dc1 · inbound

How Many Visual Tokens Do Multimodal Language Models Need? Scaling Visual Token Pruning with F^3A cites this paper.

How Many Visual Tokens Do Multimodal Language Models Need? Scaling Visual Token Pruning with F^3A LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:53:49.169616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-20T22:52:28.333209Z digest=sha256:e6565ea9dd9e2c47730950a791b43ff4f902c082ff2dbe0fbfe05524a744bc5e

Observation 9de7d28e-7dd8-4052-8e29-2b5b9566c68d · inbound

Toward Native Multimodal Modeling: A Roadmap cites this paper.

Toward Native Multimodal Modeling: A Roadmap LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 209

Resolution
verified exact
arxiv_id, observed 2026-06-29T23:04:01.689687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-29T22:58:38.610609Z digest=sha256:3e01e3a7bf53f1652e1eb151355d204759e0da18a82d6c570f0d43b1d00fadcd

Observation 08aafe95-02e6-4d43-93a5-4bcb90d34c15 · inbound

Self-Prophetic Decoding to Unlock Visual Search in LVLMs cites this paper.

Self-Prophetic Decoding to Unlock Visual Search in LVLMs LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T13:03:26.895231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-29T12:53:40.783281Z digest=sha256:a7e7aaa97d838547ea513452a8603ccf3fb83a1b3017d503ca49338360b02ea6

Observation e7bdc8b3-0d3d-42e5-bd3d-270ddb991127 · inbound

When Attention Collapses: Stage-Aware Visual Token Pruning from Structure to Semantics cites this paper.

When Attention Collapses: Stage-Aware Visual Token Pruning from Structure to Semantics LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T02:36:26.528104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-28T10:51:38.455604Z digest=sha256:28903b85f1b6d77e3aa86f4a23f9f4da97bc9f7063f6de33ba74fee645499fc9

Observation d0d2ec27-054d-4647-b706-b20d7448e68f · inbound

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception cites this paper.

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 228

Resolution
unresolved
no resolver link, observed 2026-07-12T04:17:40.198357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T04:17:40.198357Z digest=sha256:7dd68234047cd8acf49ebdf284c32632ba44df08c6a0a1c93d1e77d2ae8ace7a

Observation c955fe02-0de6-4385-83ac-b075f505bd63 · inbound

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning cites this paper.

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-31T15:16:36.341288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T15:16:36.341288Z digest=sha256:44705bbdbd7ee9178678568e7ac671686da221f5a3dbd8787e1a2602e6e2e53c

Observation ba25bc3d-d543-41ea-ba6b-fe2f83f6904b · inbound

VAD: Attributing Visual Evidence for Target Reconstruction in Multimodal On-Policy Distillation cites this paper.

VAD: Attributing Visual Evidence for Target Reconstruction in Multimodal On-Policy Distillation LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 93

Resolution
unresolved
no resolver link, observed 2026-07-31T02:56:58.293552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T02:56:58.293552Z digest=sha256:9e9c0a68adf63e7527856d3205d16f0546febe1733f61f757c00e5380af5012c