Pith. sign in

Paper Citation Record · LEDGER

VoCo-LLaMA: Towards Vision Compression with Large Language Models

As of 16 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 19 inbound Pith citation observations for arXiv:2406.12275.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.12275 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 19 of 19 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:55:03.218092Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-16T12:55:40.367899Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 49af9acd-90f2-4ab2-8674-a7dcf84265aa · inbound

Efficient Multi-modal Large Language Models via Visual Token Grouping cites this paper.

Efficient Multi-modal Large Language Models via Visual Token Grouping VoCo-LLaMA: Towards Vision Compression with Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T12:25:00.386323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:25:00.386323Z digest=sha256:27eb320b41dfb7242f1d545d8b611bd6a3eccf5237a3bb9625c2db87e74cb000

Observation 809edccb-1c91-4a68-be7a-c59c4b8baf20 · inbound

Accelerating Multimodal Large Language Models by Searching Optimal Vision Token Reduction cites this paper.

Accelerating Multimodal Large Language Models by Searching Optimal Vision Token Reduction VoCo-LLaMA: Towards Vision Compression with Large Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T05:21:10.954733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:21:10.954733Z digest=sha256:fbc67b41f58c10256cb966984de162f12a5b77bef78c00fcdcdb185b330ad1b4

Observation f352bcc1-23e7-4d6a-ade1-4a0ad0759621 · inbound

Dynamic-LLaVA: Efficient Multimodal Large Language Models via Dynamic Vision-language Context Sparsification cites this paper.

Dynamic-LLaVA: Efficient Multimodal Large Language Models via Dynamic Vision-language Context Sparsification VoCo-LLaMA: Towards Vision Compression with Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T05:01:00.359007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:01:00.359007Z digest=sha256:4ceae7e215ac0c591c8216e037bae2f70e101284ed184d7e6e8c853ccdb729d3

Observation 9307404b-c882-4586-aeef-3fab82364977 · inbound

Ponder & Press: Advancing Visual GUI Agent towards General Computer Control cites this paper.

Ponder & Press: Advancing Visual GUI Agent towards General Computer Control VoCo-LLaMA: Towards Vision Compression with Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:29.638163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:29.638163Z digest=sha256:38486888270bb24a05e971e60100fbd6b9aad7c99e6c6c899aee01fa8253e177

Observation 94e33742-a456-418b-8ce9-4b18cac4efd5 · inbound

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs cites this paper.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs VoCo-LLaMA: Towards Vision Compression with Large Language Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-11T22:35:24.370458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:35:24.370458Z digest=sha256:145f8fdb180e521c51fb11152ee406af9f97a791b021637ac4478ef2fada931a

Observation 6bcebf37-bfbd-47d8-be5d-00105a08efd5 · inbound

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs cites this paper.

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs VoCo-LLaMA: Towards Vision Compression with Large Language Models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-11T20:23:18.713337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:23:18.713337Z digest=sha256:88312c43e28d726652faedc5d66a48ef9cc198817ac9818c20d95c643bb2af8e

Observation bbb1f662-c167-462f-b560-654daa488e16 · inbound

ST$^3$: Accelerating Multimodal Large Language Model by Spatial-Temporal Visual Token Trimming cites this paper.

ST$^3$: Accelerating Multimodal Large Language Model by Spatial-Temporal Visual Token Trimming VoCo-LLaMA: Towards Vision Compression with Large Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T23:39:18.195610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:39:18.195610Z digest=sha256:ff0004bcebc7139ca2d1bdcdf9e72e27439d2317f6ce0c6656403d27089cd918

Observation 08ff8a72-db80-4fa8-8245-9c42b0332b51 · inbound

DyMU: Dynamic Merging and Virtual Unmerging for Efficient VLMs cites this paper.

DyMU: Dynamic Merging and Virtual Unmerging for Efficient VLMs VoCo-LLaMA: Towards Vision Compression with Large Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T10:55:03.218092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:55:03.218092Z digest=sha256:c57030d00af938c45202263a4e6a0b4091c52b5b33e511c32dc28c1a2b2b8259

Observation 0d1d4fab-a1c0-4627-895a-05d81901127d · inbound

UniVG-R1: Reasoning Guided Universal Visual Grounding with Reinforcement Learning cites this paper.

UniVG-R1: Reasoning Guided Universal Visual Grounding with Reinforcement Learning VoCo-LLaMA: Towards Vision Compression with Large Language Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:56.264830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:56.264830Z digest=sha256:bd19162b1ea208673b7be952af3d5ea48fe418ea19123b86a2383975f92c5de8

Observation 03a1e32b-abe3-466e-a411-0abe0a4039db · inbound

Towards General Continuous Memory for Vision-Language Models cites this paper.

Towards General Continuous Memory for Vision-Language Models VoCo-LLaMA: Towards Vision Compression with Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:46:44.748456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:46:44.748456Z digest=sha256:beb196e1301e9c8839233b8e1edefdba3aba79073009bf28c62834cb1fabbf99

Observation 949cbecf-4b31-44f1-933a-4ae41c6f8d5f · inbound

VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning cites this paper.

VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning VoCo-LLaMA: Towards Vision Compression with Large Language Models

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:55:40.369819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-16T12:55:40.245908Z digest=sha256:42f55cf094e936ee72a1288bd0b0a60f55ee9285adc4c149507a7403fa504a6d

Observation 085c94df-aa16-4d8f-926b-7d6a5fec40e2 · inbound

CoreMatching: A Co-adaptive Sparse Inference Framework with Token and Neuron Pruning for Comprehensive Acceleration of Vision-Language Models cites this paper.

CoreMatching: A Co-adaptive Sparse Inference Framework with Token and Neuron Pruning for Comprehensive Acceleration of Vision-Language Models VoCo-LLaMA: Towards Vision Compression with Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:26:17.253571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:26:17.253571Z digest=sha256:32f3c92c787ce570626b9236f9635141846df2ff1fe7868561b3828134bcfcba

Observation fa0c8107-91a4-4810-9bac-fc901fbcce76 · inbound

GEM: Empowering LLM for both Embedding Generation and Language Understanding cites this paper.

GEM: Empowering LLM for both Embedding Generation and Language Understanding VoCo-LLaMA: Towards Vision Compression with Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:50.940626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:50:50.940626Z digest=sha256:ce6d845e4d6569e566badc544661a714ce55f792755959c2e6ecea04b7e28e2e

Observation 68ed76ec-112b-4f43-92a6-a1c7b91af129 · inbound

Task-Aware KV Compression For Cost-Effective Long Video Understanding cites this paper.

Task-Aware KV Compression For Cost-Effective Long Video Understanding VoCo-LLaMA: Towards Vision Compression with Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:39.267392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:39.267392Z digest=sha256:baf1d32e6036c78b5e580c47b67a3610fb8c14957c1a7783e8782efc76491a4f

Observation a52528c2-d11a-43dd-a239-7af0352ca0dd · inbound

LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs cites this paper.

LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs VoCo-LLaMA: Towards Vision Compression with Large Language Models

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-06T22:24:35.303741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:24:35.303741Z digest=sha256:27f1b12eab6cf6de09661adc399d87ecfd246dd8212d79dd0801fd2281971613

Observation 451ae934-c65c-42c8-9c5d-bad3c4c82881 · inbound

PUMA: Layer-Pruned Language Model for Efficient Unified Multimodal Retrieval with Modality-Adaptive Learning cites this paper.

PUMA: Layer-Pruned Language Model for Efficient Unified Multimodal Retrieval with Modality-Adaptive Learning VoCo-LLaMA: Towards Vision Compression with Large Language Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T18:39:07.070965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:39:07.070965Z digest=sha256:e3f5b00e0b3ca96cacea679e80ce5677111a4d987fb7631a6632c8948aae4789

Observation 84c13542-3f1b-49c4-b010-e4b1694cf24c · inbound

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding cites this paper.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding VoCo-LLaMA: Towards Vision Compression with Large Language Models

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:03.529576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:03.529576Z digest=sha256:81d1a4f8c1e4e0471785fd08236836eb536b2e39a983b5a2c6cddce564f55cc0

Observation 876f6bfa-4589-495d-878e-6a2f0db7cf72 · inbound

VisReason: A Large-Scale Dataset for Visual Chain-of-Thought Reasoning cites this paper.

VisReason: A Large-Scale Dataset for Visual Chain-of-Thought Reasoning VoCo-LLaMA: Towards Vision Compression with Large Language Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-03T20:57:34.402360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:57:34.402360Z digest=sha256:a4a2b387f74e84615a3fee63730a34793b5ee8aa9b348698ed51c513a1312e75

Observation d3e064d2-1d8b-4cfb-8768-b50960fa318b · inbound

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception cites this paper.

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception VoCo-LLaMA: Towards Vision Compression with Large Language Models

Reference 201

Resolution
unresolved
no resolver link, observed 2026-07-12T04:17:40.198357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T04:17:40.198357Z digest=sha256:44eb2f818ad8849b7996705cf7023b261c578110fc2cf091647d201f78d58474