Pith. sign in

Paper Citation Record · LEDGER

Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models

As of 19 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 34 inbound Pith citation observations for arXiv:2406.08487.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.08487 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 34 of 34 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 34 of 34 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:28:36.781880Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T19:52:01.847188Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a9e8b27d-9201-4c1e-92c0-a0ba5b52c41d · inbound

MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans? cites this paper.

MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans? Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:59:32.756615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-16T07:59:32.638758Z digest=sha256:2bf0e23eb8cb424a1bd2ded4369a6d65050830898f828108f748fa1090945abc

Observation c85c23ed-b540-42fa-b5ae-eb93a40efafd · inbound

FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression cites this paper.

FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T15:28:37.702245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:28:37.702245Z digest=sha256:694f771c3e94490ceac23981eafe39586b681213ec3dc3faf3b6f5b3bca124b5

Observation eec66a7f-c702-4e86-bd74-b72d55c96ba4 · inbound

SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis cites this paper.

SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:04.636677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:04.636677Z digest=sha256:b0eaf7d8c9a79605a0a0a5aaba00360a4edb0138ed78ec96ff52b4c503fcb486

Observation 050d4782-a562-4c0b-ad4b-9dfb6841be5f · inbound

Large Language Model-Brained GUI Agents: A Survey cites this paper.

Large Language Model-Brained GUI Agents: A Survey Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models

Reference 235

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:08:27.814263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T11:08:27.472508Z digest=sha256:01b4f59dd68ed67ed83c4754f1a08ae0e925ea732f4be07b5be8fb87ece7a933

Observation 6e61ed9f-6618-42e0-922a-6d32942ed904 · inbound

Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing cites this paper.

Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T10:18:49.027719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:18:49.027719Z digest=sha256:c21efa95fc33be7795897089306d489413d68f99d254cb48209bd97f9b748577

Observation fd389ee4-b4b8-4a57-9470-38c4779e0df1 · inbound

Accelerating Multimodal Large Language Models by Searching Optimal Vision Token Reduction cites this paper.

Accelerating Multimodal Large Language Models by Searching Optimal Vision Token Reduction Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T05:21:10.973615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:21:10.973615Z digest=sha256:a219fe106cee57b3d08df90731258d5dd93189bc932dd5eeb464ec866afc7606

Observation 2293816a-e9c4-45b7-9785-2041aec0a79d · inbound

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion cites this paper.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:52.470510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:52.470510Z digest=sha256:3100a114daf70be99872f387661bbae96a7c01b95ddd537fc3ad391fe1442bce

Observation ee01eec7-4e0c-4fbb-bccc-f7315f330684 · inbound

p-MoD: Building Mixture-of-Depths MLLMs via Progressive Ratio Decay cites this paper.

p-MoD: Building Mixture-of-Depths MLLMs via Progressive Ratio Decay Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T21:31:44.624154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:31:44.624154Z digest=sha256:2babe416ef2179ab71050f41ae8b9e7dad43855909f3a9042ef7db0db222e6ff

Observation b17128a6-cf98-48ea-bcac-7682aca48906 · inbound

LLaVA-Zip: Adaptive Visual Token Compression with Intrinsic Image Information cites this paper.

LLaVA-Zip: Adaptive Visual Token Compression with Intrinsic Image Information Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T17:38:56.747030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:38:56.747030Z digest=sha256:334a61f544c9420eebeec9dfcc1ecab4641a9f8c37374fba8728cca151eb6195

Observation d3f4f408-d529-4635-8f5f-6201c447c5eb · inbound

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer cites this paper.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models

Reference 119

Resolution
unresolved
no resolver link, observed 2026-08-11T12:47:00.028820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:47:00.028820Z digest=sha256:b5b82fc30b6c05c78b4ce27b09af1e95356bc2a285ae19735d5027113600eb1c

Observation df1d01a3-4cf9-421d-9494-3812d1d04ec7 · inbound

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders cites this paper.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:13.683823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:13.683823Z digest=sha256:7098770c4d9e539e9c746c9d9abdc988e879ccbe68969a2cdc2352e791e43466

Observation 2a1add6b-2778-4f07-9329-22dcd1aed3ee · inbound

VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction cites this paper.

VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:08:19.635675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-17T21:08:19.570050Z digest=sha256:4202525c268abd687f62f317e93b2fecf9b7fc9e97c36ae1921497058a0a8939

Observation b9ef01a3-a29c-4a5e-9094-8278d42682ff · inbound

CL3DOR: Contrastive Learning for 3D Large Multimodal Models via Odds Ratio on High-Resolution Point Clouds cites this paper.

CL3DOR: Contrastive Learning for 3D Large Multimodal Models via Odds Ratio on High-Resolution Point Clouds Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-10T21:48:27.570775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:48:27.570775Z digest=sha256:c75f02814b59daa3fba06f9e1f565ab9cb2ea9c5d7566ae913e464db1d085533

Observation b2894911-4be6-4cfa-b667-da2a3ece02c6 · inbound

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs cites this paper.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models

Reference 110

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:16.392386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:31:16.392386Z digest=sha256:6e039c81d4ddb49ebcb43ebf48fb5f56bed5887d91990f2c39d0e331ee48608a

Observation 06e0b997-010a-4be1-bb5a-16b4b227a6d5 · inbound

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding cites this paper.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:35.628119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:35.628119Z digest=sha256:253cdb60c44286f414166c44d8050c8ee99eae2d92065cdbd951e2b45e970a62

Observation fe5d7b0f-5675-4012-8c79-da07e7ebffe9 · inbound

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler cites this paper.

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T14:18:18.042072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:18:18.042072Z digest=sha256:2d790bc04976d566ed556d46f9ff1c6add42ff5b2ef016f928e993d002831bf6

Observation b382da80-0c94-48de-b22d-f2e8fd4ba5eb · inbound

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers cites this paper.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.919283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.919283Z digest=sha256:37083169fbf4c72c8be45129f8843f2e690907a35e3d6a55feca50c0670e7923

Observation 8c8b5f18-f5f1-4841-8f7e-5927283508ab · inbound

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering cites this paper.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.387608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.387608Z digest=sha256:a3cd3472e5a1df66130eb3d56b621c46abd57f330d63b6b5e9c011574bc834b0

Observation f8a1bd3f-8f1a-4e44-ac29-90882517e885 · inbound

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment cites this paper.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:50.516471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:50.516471Z digest=sha256:405d4be07f8246d1f18c96712ee852c1155f17371c84e4ca7795a3cbd839241e

Observation 55c7d61c-77b3-48d8-a6b7-88c4aa8ead28 · inbound

FLARE: Fully Integration of Vision-Language Representations for Deep Cross-Modal Understanding cites this paper.

FLARE: Fully Integration of Vision-Language Representations for Deep Cross-Modal Understanding Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-05-22T19:52:01.848900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T19:49:00.961388Z digest=sha256:023bfe923a72cfeefdd40436c85f3d105b15445ad799cfa523584ac1ebd99d36

Observation fc0e204c-eed9-4d35-9456-4e775fb30aa4 · inbound

ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding cites this paper.

ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-16T10:28:36.781880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:28:36.781880Z digest=sha256:a68480ce73e7c28eaa4c1e8384993380a8d94a1991d02946af7343b7343d2fa7

Observation b1c3bbed-f522-4f09-a95f-6f74449b07b3 · inbound

RESAnything: Attribute Prompting for Arbitrary Referring Segmentation cites this paper.

RESAnything: Attribute Prompting for Arbitrary Referring Segmentation Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-16T04:13:03.230614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:13:03.230614Z digest=sha256:658fb3d6227eee99ca311544c2cb7d993390da6075b6c1b35a906ba75a5e2a73

Observation e9b2a7ad-a09d-464c-adfa-027d2a0f65bb · inbound

Beyond Hard and Soft: Hybrid Context Compression for Balancing Local and Global Information Retention cites this paper.

Beyond Hard and Soft: Hybrid Context Compression for Balancing Local and Global Information Retention Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T15:16:01.697010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:16:01.697010Z digest=sha256:d3d782414c24d61b12bc5923d085152ceff2a03191dd0af421cecb56beedf043

Observation 80f2e1f6-1317-4010-b328-cee0148f0e27 · inbound

AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs cites this paper.

AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-07T13:40:25.686235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:40:25.686235Z digest=sha256:79a12ed9550c2cd5255bb6f4f845ae27b704da4b856721348dd53ace85fd2f91

Observation 309da78b-8d1e-4153-836e-aa64ee4ed65e · inbound

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models cites this paper.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:34.329669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:34.329669Z digest=sha256:5f081030457ffadd9f4ff16a20cc8f5a4b00aea3b0807971e27773346dd1a4ec

Observation 8e3c2543-df55-4b6f-86a9-bb839450d42e · inbound

GenRecal: Generation after Recalibration from Large to Small Vision-Language Models cites this paper.

GenRecal: Generation after Recalibration from Large to Small Vision-Language Models Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models

Reference 123

Resolution
unresolved
no resolver link, observed 2026-08-06T23:57:28.778358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:57:28.778358Z digest=sha256:7d66438dde39b8b5c0a457032ae94e9313c49be82d0ec76b1b3cd88e0a9fb24d

Observation 858439a1-fa26-45ca-af18-97a616067348 · inbound

High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning cites this paper.

High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-19T06:12:07.101562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T06:10:57.219445Z digest=sha256:a3c0fe7a3403dbc4f909fc30397da2513e2e1df18921c152bf37ecd1929e814b

Observation 2f9c8825-b548-405f-a355-b60ea1a8aa1c · inbound

Mitigating Information Loss under High Pruning Rates for Efficient Large Vision Language Models cites this paper.

Mitigating Information Loss under High Pruning Rates for Efficient Large Vision Language Models Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T05:51:16.113242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:51:16.113242Z digest=sha256:8d83f11a3e8ad5d2feedc4c14455cb5441fba0b0625406656e831abd5f5b6be1

Observation 3bfbd8f8-02de-442e-a511-6d69283bce5f · inbound

CogVLA: Cognition-Aligned Vision-Language-Action Model via Instruction-Driven Routing & Sparsification cites this paper.

CogVLA: Cognition-Aligned Vision-Language-Action Model via Instruction-Driven Routing & Sparsification Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-05T14:42:32.838564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:42:32.838564Z digest=sha256:ba6ea9ccef39c0c1303a21642eea3c1b9e8ea5a780d4e9da0719c0c9aa42c59c

Observation 54d06397-c0d5-4151-a839-fb563950427b · inbound

Mitigating Coordinate Prediction Bias from Positional Encoding Failures cites this paper.

Mitigating Coordinate Prediction Bias from Positional Encoding Failures Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:12:23.524070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T05:11:30.734633Z digest=sha256:9493d2f03afd3f62182840497b15c381c356a26db7727af208ab9ab4f23baba2

Observation 2825c939-e432-4838-a4b1-6456fd772dab · inbound

HART: High-Resolution Annotation-Free Reasoning Technique through a Closed-loop Framework cites this paper.

HART: High-Resolution Annotation-Free Reasoning Technique through a Closed-loop Framework Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-02T20:19:17.878795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:19:17.878795Z digest=sha256:86ea1d29cad7bd1e97246aa604f3b59ef26363020952458d465a9a52b5d5b93b

Observation 2517de30-451d-4e1f-ad06-5ddc7ee99579 · inbound

ConsisVLA-4D: Advancing Spatiotemporal Consistency in Efficient 3D-Perception and 4D-Reasoning for Robotic Manipulation cites this paper.

ConsisVLA-4D: Advancing Spatiotemporal Consistency in Efficient 3D-Perception and 4D-Reasoning for Robotic Manipulation Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models

Reference 87

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:06:06.159651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-08T16:40:19.057979Z digest=sha256:fd9e9752a60aa3425a872de48b5e8e9860ae2790f3229522a05845ed4ed1d95d

Observation aa753e71-a650-4be4-807e-36b50e213599 · inbound

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception cites this paper.

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models

Reference 199

Resolution
unresolved
no resolver link, observed 2026-07-12T04:17:40.198357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T04:17:40.198357Z digest=sha256:886bb27b10352aa0321197f70290c237047da7fe772cb48fffdd1d028f1835ba

Observation 0a18acc9-376b-4343-b92c-c026baf70eed · inbound

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin cites this paper.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.350680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.350680Z digest=sha256:eecd770549828008a1ab5d42a8050a40a1fb5cd569b991b6ccca7f096cd6caa9