Pith. sign in

Paper Citation Record · LEDGER

LLaVA-Zip: Adaptive Visual Token Compression with Intrinsic Image Information

As of 15 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2412.08771.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.08771 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T17:38:56.881537Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact0
  • verified fuzzy5
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3da70c1f-c91b-4a78-ac6f-1045cbf86f16 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

LLaVA-Zip: Adaptive Visual Token Compression with Intrinsic Image Information LLaVA-OneVision: Easy Visual Task Transfer

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T17:38:56.716772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:38:56.716772Z digest=sha256:41c6a9f809bb8e83826e7a8ff3d0c8b609af39b325c39875f1840cd2622b4958

Observation a876d2af-f907-4e5f-9bd7-fb52c28479a1 · outbound

This paper cites Improved baselines with visual instruction tuning.

LLaVA-Zip: Adaptive Visual Token Compression with Intrinsic Image Information Improved baselines with visual instruction tuning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T17:38:56.724332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:38:56.724332Z digest=sha256:008d0d3ccfdb34d71e518ae37268d56db8ea75beac1a13ec5871f2300c2dd16a

Observation b2f2f848-a4e1-4498-ae03-80ec1bd017a6 · outbound

This paper cites Visual instruction tuning.

LLaVA-Zip: Adaptive Visual Token Compression with Intrinsic Image Information Visual instruction tuning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T17:38:56.730122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:38:56.730122Z digest=sha256:32ae2369d5c2473e54e6cd4b493090909783c6276f1fd6c359586c269fc740f8

Observation af58c3fd-2e36-494f-9fd3-029e451b498c · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

LLaVA-Zip: Adaptive Visual Token Compression with Intrinsic Image Information Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:38:57.311285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:38:56.735300Z digest=sha256:7d739110eb8dec9434a12858d57bec4a4317cc9f4735ff39098cb6640a7e7d67

Observation 20390628-8090-428a-907a-0c1a15a05ccd · outbound

This paper cites Less is More: A Simple yet Effective Token Reduction Method for Efficient Multi-modal LLMs.

LLaVA-Zip: Adaptive Visual Token Compression with Intrinsic Image Information Less is More: A Simple yet Effective Token Reduction Method for Efficient Multi-modal LLMs

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T17:38:56.742024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:38:56.742024Z digest=sha256:de4ad451bf8c13848520b45bbda616835eaace04dc856ea2960a9709856a3e9c

Observation b17128a6-cf98-48ea-bcac-7682aca48906 · outbound

This paper cites Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models.

LLaVA-Zip: Adaptive Visual Token Compression with Intrinsic Image Information Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T17:38:56.747030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:38:56.747030Z digest=sha256:23c9d3935b06fd42b966e223cf832879e18fb41fa5436bc72bf66b876a37056a

Observation 082004d0-bfd8-4a5a-a00f-5ae86d088b6f · outbound

This paper cites HiRED: Attention-Guided Token Dropping for Efficient Inference of High-Resolution Vision-Language Models.

LLaVA-Zip: Adaptive Visual Token Compression with Intrinsic Image Information HiRED: Attention-Guided Token Dropping for Efficient Inference of High-Resolution Vision-Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T17:38:56.755180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:38:56.755180Z digest=sha256:20709a6754436a729d0e162c1ea7eeb857b12d5321d334e5f3745a3ec8a5f9dd

Observation eefffb81-b279-4120-ada8-0c530d5cf11d · outbound

This paper cites Matryoshka Multimodal Models.

LLaVA-Zip: Adaptive Visual Token Compression with Intrinsic Image Information Matryoshka Multimodal Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T17:38:56.762019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:38:56.762019Z digest=sha256:b942bede0be2730f85ec86c298aa2b863c0b98e6b42479006bd0e887bdfea7f3

Observation 6b22601a-7576-4731-bf58-c7d046f3fc20 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

LLaVA-Zip: Adaptive Visual Token Compression with Intrinsic Image Information Flamingo: a visual language model for few-shot learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T17:38:56.767350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:38:56.767350Z digest=sha256:402daf646398b325b99b92a56d8ff2c346e90134420643499bb5f53f4f2efc08

Observation 81522011-0c70-4c50-b4e6-f46a7adcdf3e · outbound

This paper cites FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression.

LLaVA-Zip: Adaptive Visual Token Compression with Intrinsic Image Information FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T17:38:56.772419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:38:56.772419Z digest=sha256:3adeae63324b6c82e74a579d529bd65f4352ab5125eeb1a45b8d12ba3b6c20e9

Observation 66b5eed6-4f90-499c-9ac5-849f2171c11c · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

LLaVA-Zip: Adaptive Visual Token Compression with Intrinsic Image Information Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T17:38:56.776843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:38:56.776843Z digest=sha256:9e8e436edbf771af7b5643fd97c644f5c6904e625c539047ce6ece48291df45c

Observation 51d5bba8-1fb1-493c-b520-d423a3a07391 · outbound

This paper cites Learning transferable visual models from natural language supervision.

LLaVA-Zip: Adaptive Visual Token Compression with Intrinsic Image Information Learning transferable visual models from natural language supervision

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T17:38:56.781775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:38:56.781775Z digest=sha256:30f2a552d321814258f23c71e381d7663443bf20e083f8c51821e47fa2989039

Observation 886b59ed-d335-4c77-9271-7b21ec3dfb4a · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.

LLaVA-Zip: Adaptive Visual Token Compression with Intrinsic Image Information Judging llm-as-a-judge with mt-bench and chatbot arena

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T17:38:56.787378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:38:56.787378Z digest=sha256:0f4e32b0b0f9ea606366a56cde08ffec3d8a146dbd17c5747928a966ab79ba55

Observation 0462a536-8991-45ab-a93b-062bd723a3a9 · outbound

This paper cites Lmms-eval: Accelerating the development of large multimoal models, March 2024.

LLaVA-Zip: Adaptive Visual Token Compression with Intrinsic Image Information Lmms-eval: Accelerating the development of large multimoal models, March 2024

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:38:57.263040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:38:56.792098Z digest=sha256:7b771bc4dd00be97280fa52dd847f54afaf74c3ec6e64b38da2cf928fa031741

Observation ff118de8-15c3-40fb-b0f8-34609c1c92f9 · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and com- positional question answering.

LLaVA-Zip: Adaptive Visual Token Compression with Intrinsic Image Information Gqa: A new dataset for real-world visual reasoning and com- positional question answering

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:38:57.245227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:38:56.796485Z digest=sha256:a760d1361ddc629aeec2e27190177e71eb58dd83a8db446e22c22a89198f49e2

Observation 1574031e-206e-434c-86c5-f0d001fa7d8b · outbound

This paper cites Towards vqa models that can read.

LLaVA-Zip: Adaptive Visual Token Compression with Intrinsic Image Information Towards vqa models that can read

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T17:38:56.806224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:38:56.806224Z digest=sha256:fa7d94814fee3133e3edec387f7593e367f4732d3d5ff45f1b16e545be25eea4

Observation 522ece48-3d03-42b5-834c-122ee1b67971 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answering.

LLaVA-Zip: Adaptive Visual Token Compression with Intrinsic Image Information Making the v in vqa matter: Elevating the role of image understanding in visual question answering

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T17:38:56.811310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:38:56.811310Z digest=sha256:bc65544034119448cbb4692107a5bc3caf40774685b7938d350e63b75bc75e9a

Observation 21ed474f-098d-4ec1-aa90-6b8977507d4d · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

LLaVA-Zip: Adaptive Visual Token Compression with Intrinsic Image Information Evaluating Object Hallucination in Large Vision-Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T17:38:56.816229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:38:56.816229Z digest=sha256:79e5f76e9c31f5f31c4ffc8468051fe720987de1c80192f48a71f8bc9dcb3919

Observation e9f6b0d1-199e-4e39-9103-f02365561dc9 · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

LLaVA-Zip: Adaptive Visual Token Compression with Intrinsic Image Information MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T17:38:56.821977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:38:56.821977Z digest=sha256:a014a74276c76f34190cdd180e55880a3d10d929f15ae3c6fad0ce1974c50ed4

Observation 7136a7d0-242b-432d-bf2d-ed56bf59d1e0 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

LLaVA-Zip: Adaptive Visual Token Compression with Intrinsic Image Information MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T17:38:56.826692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:38:56.826692Z digest=sha256:2ee3e6fac26a3950886f7bd048f5e6cce7deb6f497c8099c07633ba3a707d0b1

Observation 9ba0145d-d6af-4293-9300-b83a1639c68b · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

LLaVA-Zip: Adaptive Visual Token Compression with Intrinsic Image Information SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T17:38:56.832530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:38:56.832530Z digest=sha256:6b5c59acb4070ca305f2bb8478abe512d0dd309900538774bb67f066cfa2a2d6

Observation deac9b71-a821-4eb7-8923-97a186334a82 · outbound

This paper cites Microsoft coco: Common objects in context.

LLaVA-Zip: Adaptive Visual Token Compression with Intrinsic Image Information Microsoft coco: Common objects in context

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T17:38:56.837330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:38:56.837330Z digest=sha256:ccdaceefdbc28471b3f8debacc66a3477fdd2b632182bcc7cb6e9dd6ccaf5c62

Observation 63e0699d-71d1-4c86-a7f4-0503c5dbf140 · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

LLaVA-Zip: Adaptive Visual Token Compression with Intrinsic Image Information Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T17:38:56.844734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:38:56.844734Z digest=sha256:c109279b4c965b8412038704c0762ada17da9d3264eae3cc66f022ac5c03cb8d

Observation 32b3dd32-1c17-48d5-b54e-60ce82f730ef · outbound

This paper cites Ray: A distributed framework for emerging {AI} applications.

LLaVA-Zip: Adaptive Visual Token Compression with Intrinsic Image Information Ray: A distributed framework for emerging {AI} applications

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:38:57.201146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:38:56.851157Z digest=sha256:8324da03e07e2cd02efbae18fbfaa6da06bc2bc23530542d5fbe499637d0141c

Observation d516df20-98d7-4ee1-b139-f49f5193a639 · outbound

This paper cites Ffcv: Accelerating training by removing data bottlenecks.

LLaVA-Zip: Adaptive Visual Token Compression with Intrinsic Image Information Ffcv: Accelerating training by removing data bottlenecks

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:38:57.182475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:38:56.861083Z digest=sha256:5ae789e5ac0f3f82e670378c4dc118d2f93cdd61b2996f7471d5eaa2e666c281

Observation 6e4c7e81-8d2f-4afb-a905-1232853e6952 · outbound

This paper cites The Llama 3 Herd of Models.

LLaVA-Zip: Adaptive Visual Token Compression with Intrinsic Image Information The Llama 3 Herd of Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T17:38:56.867338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:38:56.867338Z digest=sha256:5684ab256ce59314371b056e834db8da6870fd1ba829b0a28b082f6b3e7f329a

Observation 885f5b81-3a33-4fd7-9e6d-36c80160fe4d · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

LLaVA-Zip: Adaptive Visual Token Compression with Intrinsic Image Information Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T17:38:56.881537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:38:56.881537Z digest=sha256:d144a0aff44028b50a84b2c464fde52df39c7215435a28da1276038076fa1a82

Pith citing papers

No inbound Pith citation observations are available.