Pith. sign in

Paper Citation Record · LEDGER

Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 17 inbound Pith citation observations for arXiv:2406.08487.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.08487 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T12:54:56.387608Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T19:52:01.847188Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a9e8b27d-9201-4c1e-92c0-a0ba5b52c41d · inbound

MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans? cites this paper.

MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans? Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:59:32.756615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T07:59:32.638758Z digest=sha256:76ce84d70b887438c70934583e53267a7f26a1992e00ae777365f4ac173338e1

Observation 050d4782-a562-4c0b-ad4b-9dfb6841be5f · inbound

Large Language Model-Brained GUI Agents: A Survey cites this paper.

Large Language Model-Brained GUI Agents: A Survey Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models

Reference 235

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:08:27.814263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T11:08:27.472508Z digest=sha256:3ad3abb601edafa07ee89a1597c446d0239ea6a64dea8e23f6a2233c079f53ce

Observation 2a1add6b-2778-4f07-9329-22dcd1aed3ee · inbound

VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction cites this paper.

VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:08:19.635675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T21:08:19.570050Z digest=sha256:c21c7a574889ffe7c24a0de08c77f561c4e60464b4aaf71189eb69900a47fc32

Observation 8c8b5f18-f5f1-4841-8f7e-5927283508ab · inbound

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering cites this paper.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.387608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.387608Z digest=sha256:931a21c22c17c9967d39c8bf9ef23c74821108ceeae663bb7bdef19d7bd57b72

Observation f8a1bd3f-8f1a-4e44-ac29-90882517e885 · inbound

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment cites this paper.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:50.516471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:50.516471Z digest=sha256:5d177e6a3b49a7b10f97064fb46ba0cce36068fe61be83973b2ca7ae030e9392

Observation 55c7d61c-77b3-48d8-a6b7-88c4aa8ead28 · inbound

FLARE: Fully Integration of Vision-Language Representations for Deep Cross-Modal Understanding cites this paper.

FLARE: Fully Integration of Vision-Language Representations for Deep Cross-Modal Understanding Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-05-22T19:52:01.848900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T19:49:00.961388Z digest=sha256:50c5c88cdabc080d1cb12baf2d38d409a8c908691ced9fdd3d3a429baa034222

Observation e9b2a7ad-a09d-464c-adfa-027d2a0f65bb · inbound

Beyond Hard and Soft: Hybrid Context Compression for Balancing Local and Global Information Retention cites this paper.

Beyond Hard and Soft: Hybrid Context Compression for Balancing Local and Global Information Retention Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T15:16:01.697010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:16:01.697010Z digest=sha256:bdb6f2396b9e3be3ace40f0a6a2ac73e8fd058ac7518a8fc5536d5128260d6de

Observation 80f2e1f6-1317-4010-b328-cee0148f0e27 · inbound

AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs cites this paper.

AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-07T13:40:25.686235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:40:25.686235Z digest=sha256:b967829a63fabeea3d4a44862ea6f4ad4517dbc55c9f655ce6b3f428f1674886

Observation 309da78b-8d1e-4153-836e-aa64ee4ed65e · inbound

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models cites this paper.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:34.329669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:34.329669Z digest=sha256:5c0259cd8d8ab018a710d884ba76d4a6322d5f137e1c381e3c5fc8e3e337a87b

Observation 8e3c2543-df55-4b6f-86a9-bb839450d42e · inbound

GenRecal: Generation after Recalibration from Large to Small Vision-Language Models cites this paper.

GenRecal: Generation after Recalibration from Large to Small Vision-Language Models Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models

Reference 123

Resolution
unresolved
no resolver link, observed 2026-08-06T23:57:28.778358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:57:28.778358Z digest=sha256:74b74c6e0ef27bf1fea1972087e9cf8f6b7b193961b0cf1cf3981df2ac3885ff

Observation 858439a1-fa26-45ca-af18-97a616067348 · inbound

High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning cites this paper.

High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-19T06:12:07.101562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T06:10:57.219445Z digest=sha256:71733be85bebce100ea9ffd67d08598e9d0cc41e851834e918fb80da309d3031

Observation 2f9c8825-b548-405f-a355-b60ea1a8aa1c · inbound

Mitigating Information Loss under High Pruning Rates for Efficient Large Vision Language Models cites this paper.

Mitigating Information Loss under High Pruning Rates for Efficient Large Vision Language Models Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T05:51:16.113242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:51:16.113242Z digest=sha256:7289d3136c49ae197ff72be1df8a829a06dfb1a2ebbee5831c3d12964159aa5d

Observation 3bfbd8f8-02de-442e-a511-6d69283bce5f · inbound

CogVLA: Cognition-Aligned Vision-Language-Action Model via Instruction-Driven Routing & Sparsification cites this paper.

CogVLA: Cognition-Aligned Vision-Language-Action Model via Instruction-Driven Routing & Sparsification Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-05T14:42:32.838564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:42:32.838564Z digest=sha256:3027cedbb5171533f482ef363cc63e3dc024a02fc4e0575aa6fdcdfd4d6c208a

Observation 54d06397-c0d5-4151-a839-fb563950427b · inbound

Mitigating Coordinate Prediction Bias from Positional Encoding Failures cites this paper.

Mitigating Coordinate Prediction Bias from Positional Encoding Failures Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:12:23.524070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T05:11:30.734633Z digest=sha256:d189ccf39d5daeac8af0249209571a7d7557416cde1b166af35a362eb53c54ca

Observation 2825c939-e432-4838-a4b1-6456fd772dab · inbound

HART: High-Resolution Annotation-Free Reasoning Technique through a Closed-loop Framework cites this paper.

HART: High-Resolution Annotation-Free Reasoning Technique through a Closed-loop Framework Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-02T20:19:17.878795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:19:17.878795Z digest=sha256:8e668e157cea4c6c84cf8b17cd24b7821d18134657a0402923e8d6313422d43d

Observation 2517de30-451d-4e1f-ad06-5ddc7ee99579 · inbound

ConsisVLA-4D: Advancing Spatiotemporal Consistency in Efficient 3D-Perception and 4D-Reasoning for Robotic Manipulation cites this paper.

ConsisVLA-4D: Advancing Spatiotemporal Consistency in Efficient 3D-Perception and 4D-Reasoning for Robotic Manipulation Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models

Reference 87

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:06:06.159651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T16:40:19.057979Z digest=sha256:7a2ca14c49f3b5169c4e31b62ea53987aaf9918180dd08b8529d7b676f722a19

Observation aa753e71-a650-4be4-807e-36b50e213599 · inbound

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception cites this paper.

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models

Reference 199

Resolution
unresolved
no resolver link, observed 2026-07-12T04:17:40.198357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T04:17:40.198357Z digest=sha256:1903654fed0027ea6e060001850c284fc3691192dc3c5bd388d3ebdf7e23722b