Pith. sign in

Paper Citation Record · LEDGER

Efficient Large Multi-modal Models via Visual Context Compression

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2406.20092.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.20092 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:11:32.736571Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T10:48:03.161959Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 46f9e84c-7eac-4789-9d20-70dc9040a77d · inbound

PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction cites this paper.

PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction Efficient Large Multi-modal Models via Visual Context Compression

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:12:14.743952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:12:14.613620Z digest=sha256:e43177210ae38987a5e6ced13dabca0337caa8dada33b3ae52a4a32674a63615

Observation dd31a8fc-f16e-4dcd-bc52-ab91dbd3186a · inbound

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling cites this paper.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Efficient Large Multi-modal Models via Visual Context Compression

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:02:43.682863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:ca3d2aabe5e6086436b5ee53f82cd69601a06783cd4ecc8b508c8338c6bfba1f

Observation 5c4ded37-74fb-480f-a64a-039396ba00c7 · inbound

Growing a Multi-head Twig via Distillation and Reinforcement Learning to Accelerate Large Vision-Language Models cites this paper.

Growing a Multi-head Twig via Distillation and Reinforcement Learning to Accelerate Large Vision-Language Models Efficient Large Multi-modal Models via Visual Context Compression

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-23T00:02:17.873890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T23:58:57.819555Z digest=sha256:5770cb2eb6ee6147373ae01276371671aab71c7e143ef9b5b6e051b112da07b3

Observation 2167d069-857a-41eb-9072-65fe975ff2e5 · inbound

Efficient Multi-modal Long Context Learning for Training-free Adaptation cites this paper.

Efficient Multi-modal Long Context Learning for Training-free Adaptation Efficient Large Multi-modal Models via Visual Context Compression

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:32.736571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:11:32.736571Z digest=sha256:ed286998f077a6db6cef1497d1216530513b251455bc51c2c75d47366d5fc56b

Observation 28c77754-5715-4cfa-a80a-afb162c9921b · inbound

GreedyPrune: Retenting Critical Visual Token Set for Large Vision Language Models cites this paper.

GreedyPrune: Retenting Critical Visual Token Set for Large Vision Language Models Efficient Large Multi-modal Models via Visual Context Compression

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T00:42:53.603924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:42:53.603924Z digest=sha256:e46ed4b371bf4ce50d5deaa996490271052543d39171a3278c66db6d77b6f342

Observation 0e65442a-8ef8-44e2-86c1-fa56cf9150ee · inbound

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding cites this paper.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Efficient Large Multi-modal Models via Visual Context Compression

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:58.561451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:58.561451Z digest=sha256:dcbd2938278cc8d92baa648919416bdbf14fe101d812707ef5585e25c312faa4

Observation 9366842e-7f78-4ba8-a8a1-f0bdb62c1076 · inbound

Structured Prompting and Multi-Agent Knowledge Distillation for Traffic Video Interpretation and Risk Inference cites this paper.

Structured Prompting and Multi-Agent Knowledge Distillation for Traffic Video Interpretation and Risk Inference Efficient Large Multi-modal Models via Visual Context Compression

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T19:06:39.208373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T19:06:39.208373Z digest=sha256:0d8a15bcd4464c53c1f40755d225df83a764fb39416764c038f754fdcd321e22

Observation 4ca98475-15be-4e2d-b40d-3420c76b12b1 · inbound

Why and When Visual Token Pruning Fails? A Study on Relevant Visual Information Shift in MLLMs Decoding cites this paper.

Why and When Visual Token Pruning Fails? A Study on Relevant Visual Information Shift in MLLMs Decoding Efficient Large Multi-modal Models via Visual Context Compression

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:31:01.569823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T14:50:37.022338Z digest=sha256:22f6eb96095489176bd38ae936bd975e42068c85d82e71ef4608f9088593c9e0

Observation 97a43270-b6fe-47e1-b35c-e6153636e3bc · inbound

Task-Aware Structured Memory for Dynamic Multi-modal In-Context Learning cites this paper.

Task-Aware Structured Memory for Dynamic Multi-modal In-Context Learning Efficient Large Multi-modal Models via Visual Context Compression

Reference 210

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T09:07:48.419690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T10:28:11.440915Z digest=sha256:64860b2ebaeccc649938555d4c18519bd370e29a9e16b20e5cdf2e5ef90dc466

Observation b00b2214-67db-454d-bb4b-8352925e2055 · inbound

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning cites this paper.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Efficient Large Multi-modal Models via Visual Context Compression

Reference 295

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T10:48:03.163310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:f9ef280840e5527270167447205fbca43dcfef9caa82a188b9267026be162fed