Pith. sign in

Paper Citation Record · LEDGER

LLaVA-Zip: Adaptive Visual Token Compression with Intrinsic Image Information

As of 15 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2412.08771.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.08771 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T17:38:56.881537Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact0
  • verified fuzzy5
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3da70c1f-c91b-4a78-ac6f-1045cbf86f16 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

LLaVA-Zip: Adaptive Visual Token Compression with Intrinsic Image Information LLaVA-OneVision: Easy Visual Task Transfer

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T17:38:56.716772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:38:56.716772Z digest=sha256:188dcf246409955fe37a7e4f1903ef4c3335e2fb55d4c8ba6b46d141ce45e989

Observation a876d2af-f907-4e5f-9bd7-fb52c28479a1 · outbound

This paper cites Improved baselines with visual instruction tuning.

LLaVA-Zip: Adaptive Visual Token Compression with Intrinsic Image Information Improved baselines with visual instruction tuning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T17:38:56.724332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:38:56.724332Z digest=sha256:a2dc98eb088f0e698ebc00dbdfe80139d69dfdb544494a2c9b54a299fa5cfe30

Observation b2f2f848-a4e1-4498-ae03-80ec1bd017a6 · outbound

This paper cites Visual instruction tuning.

LLaVA-Zip: Adaptive Visual Token Compression with Intrinsic Image Information Visual instruction tuning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T17:38:56.730122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:38:56.730122Z digest=sha256:f0b0117f18db6389411d6a7a8c3a622ee725891d6e7d788a96420a8dc5e43e78

Observation af58c3fd-2e36-494f-9fd3-029e451b498c · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

LLaVA-Zip: Adaptive Visual Token Compression with Intrinsic Image Information Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:38:57.311285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:38:56.735300Z digest=sha256:67934722a3e3144b01730c81ea1fefc789592e4b69c492072adf4c397e24d1e8

Observation 20390628-8090-428a-907a-0c1a15a05ccd · outbound

This paper cites Less is More: A Simple yet Effective Token Reduction Method for Efficient Multi-modal LLMs.

LLaVA-Zip: Adaptive Visual Token Compression with Intrinsic Image Information Less is More: A Simple yet Effective Token Reduction Method for Efficient Multi-modal LLMs

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T17:38:56.742024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:38:56.742024Z digest=sha256:73a881e25bef14eaa4e8e36a7480986a1c71a6775fa0c94a1d8c2b6699318b28

Observation b17128a6-cf98-48ea-bcac-7682aca48906 · outbound

This paper cites Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models.

LLaVA-Zip: Adaptive Visual Token Compression with Intrinsic Image Information Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T17:38:56.747030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:38:56.747030Z digest=sha256:e20eb18df7e96cda1c8708f02c05c3ed4e24765391d47bfe860419d8b2303667

Observation 082004d0-bfd8-4a5a-a00f-5ae86d088b6f · outbound

This paper cites HiRED: Attention-Guided Token Dropping for Efficient Inference of High-Resolution Vision-Language Models.

LLaVA-Zip: Adaptive Visual Token Compression with Intrinsic Image Information HiRED: Attention-Guided Token Dropping for Efficient Inference of High-Resolution Vision-Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T17:38:56.755180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:38:56.755180Z digest=sha256:8ffa7605e26dc69950bd43118587a0b71f13e39776d79041cd7f0225a7505948

Observation eefffb81-b279-4120-ada8-0c530d5cf11d · outbound

This paper cites Matryoshka Multimodal Models.

LLaVA-Zip: Adaptive Visual Token Compression with Intrinsic Image Information Matryoshka Multimodal Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T17:38:56.762019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:38:56.762019Z digest=sha256:1e5dd83eeaa2019fcdb535ff0c3f0fa9feaf8c0c8bbbc2b6d83dbf7d18927ad3

Observation 6b22601a-7576-4731-bf58-c7d046f3fc20 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

LLaVA-Zip: Adaptive Visual Token Compression with Intrinsic Image Information Flamingo: a visual language model for few-shot learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T17:38:56.767350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:38:56.767350Z digest=sha256:6b20cca8e0d71d6718affce96fb42692471f6034a6f5368f75338763f9784f1b

Observation 81522011-0c70-4c50-b4e6-f46a7adcdf3e · outbound

This paper cites FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression.

LLaVA-Zip: Adaptive Visual Token Compression with Intrinsic Image Information FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T17:38:56.772419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:38:56.772419Z digest=sha256:8f4dc7ad3d14cde719fe5176a3ef61956614f20cefcc964fbed34aeaad55e311

Observation 66b5eed6-4f90-499c-9ac5-849f2171c11c · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

LLaVA-Zip: Adaptive Visual Token Compression with Intrinsic Image Information Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T17:38:56.776843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:38:56.776843Z digest=sha256:fd27281d62687cd263ea77cd26ee6f3d83b3b2c26cc93b10024ff5d24bb376ab

Observation 51d5bba8-1fb1-493c-b520-d423a3a07391 · outbound

This paper cites Learning transferable visual models from natural language supervision.

LLaVA-Zip: Adaptive Visual Token Compression with Intrinsic Image Information Learning transferable visual models from natural language supervision

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T17:38:56.781775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:38:56.781775Z digest=sha256:722d2375847008e0f39fbf4beb51f3e654ee95c0bbe808c40a58aa20929d0636

Observation 886b59ed-d335-4c77-9271-7b21ec3dfb4a · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.

LLaVA-Zip: Adaptive Visual Token Compression with Intrinsic Image Information Judging llm-as-a-judge with mt-bench and chatbot arena

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T17:38:56.787378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:38:56.787378Z digest=sha256:4c59f8746d2b0b4c47cd2a306e9ed48490dbf9c0f78ccf856b35ae639b0c8f1f

Observation 0462a536-8991-45ab-a93b-062bd723a3a9 · outbound

This paper cites Lmms-eval: Accelerating the development of large multimoal models, March 2024.

LLaVA-Zip: Adaptive Visual Token Compression with Intrinsic Image Information Lmms-eval: Accelerating the development of large multimoal models, March 2024

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:38:57.263040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:38:56.792098Z digest=sha256:c52859d64a6dcbcff40a87cde605282101dc344447682af584e657cf46985948

Observation ff118de8-15c3-40fb-b0f8-34609c1c92f9 · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and com- positional question answering.

LLaVA-Zip: Adaptive Visual Token Compression with Intrinsic Image Information Gqa: A new dataset for real-world visual reasoning and com- positional question answering

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:38:57.245227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:38:56.796485Z digest=sha256:8d00eaa7f8a85fe931f1eab443a7fa633a0f210efc5ab84f8909e164d275b3a4

Observation 1574031e-206e-434c-86c5-f0d001fa7d8b · outbound

This paper cites Towards vqa models that can read.

LLaVA-Zip: Adaptive Visual Token Compression with Intrinsic Image Information Towards vqa models that can read

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T17:38:56.806224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:38:56.806224Z digest=sha256:4f08cd0e68bd4dd3708034e7352c88f1a43c6ab2c7cdadd6c8fcdfa377d6eb9b

Observation 522ece48-3d03-42b5-834c-122ee1b67971 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answering.

LLaVA-Zip: Adaptive Visual Token Compression with Intrinsic Image Information Making the v in vqa matter: Elevating the role of image understanding in visual question answering

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T17:38:56.811310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:38:56.811310Z digest=sha256:2004f688d78190e853ad3fac16b62feb02bbab5d72a56df3cdab4cb03e96857f

Observation 21ed474f-098d-4ec1-aa90-6b8977507d4d · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

LLaVA-Zip: Adaptive Visual Token Compression with Intrinsic Image Information Evaluating Object Hallucination in Large Vision-Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T17:38:56.816229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:38:56.816229Z digest=sha256:f2e1974bbcdfe4a27275fac5a3757a716cfd36192a432e159ce495ef636a59be

Observation e9f6b0d1-199e-4e39-9103-f02365561dc9 · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

LLaVA-Zip: Adaptive Visual Token Compression with Intrinsic Image Information MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T17:38:56.821977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:38:56.821977Z digest=sha256:6971935daf33c48d9e2ebe2d4ca92f0c760a1b6cfff70ddd23d0882161c5482d

Observation 7136a7d0-242b-432d-bf2d-ed56bf59d1e0 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

LLaVA-Zip: Adaptive Visual Token Compression with Intrinsic Image Information MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T17:38:56.826692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:38:56.826692Z digest=sha256:0b75ff4b7a7b09e8b5514438231eb6e273f1affc7db923bda8702b0f05ca8aed

Observation 9ba0145d-d6af-4293-9300-b83a1639c68b · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

LLaVA-Zip: Adaptive Visual Token Compression with Intrinsic Image Information SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T17:38:56.832530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:38:56.832530Z digest=sha256:5692570c843947b141fc5bd2a2b4d639217f9ac888b044a77b9911a7e7d6413a

Observation deac9b71-a821-4eb7-8923-97a186334a82 · outbound

This paper cites Microsoft coco: Common objects in context.

LLaVA-Zip: Adaptive Visual Token Compression with Intrinsic Image Information Microsoft coco: Common objects in context

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T17:38:56.837330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:38:56.837330Z digest=sha256:ed33d28b53d318b6d7914bdb9a793051ecf1a6d309f503df023603f819e8a910

Observation 63e0699d-71d1-4c86-a7f4-0503c5dbf140 · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

LLaVA-Zip: Adaptive Visual Token Compression with Intrinsic Image Information Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T17:38:56.844734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:38:56.844734Z digest=sha256:cdc743c2291588a50ce4aa2c032ad2cef477346effe56df43394c51a846016e4

Observation 32b3dd32-1c17-48d5-b54e-60ce82f730ef · outbound

This paper cites Ray: A distributed framework for emerging {AI} applications.

LLaVA-Zip: Adaptive Visual Token Compression with Intrinsic Image Information Ray: A distributed framework for emerging {AI} applications

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:38:57.201146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:38:56.851157Z digest=sha256:f08c1c2fdc5c47e1772e0001766d46dd47ee3862d7a26a3d9e28853a7da5c49e

Observation d516df20-98d7-4ee1-b139-f49f5193a639 · outbound

This paper cites Ffcv: Accelerating training by removing data bottlenecks.

LLaVA-Zip: Adaptive Visual Token Compression with Intrinsic Image Information Ffcv: Accelerating training by removing data bottlenecks

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:38:57.182475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:38:56.861083Z digest=sha256:91ea450485c932c19be39aa006deaa059869a68dc7cba596faf9e6d481a5afe8

Observation 6e4c7e81-8d2f-4afb-a905-1232853e6952 · outbound

This paper cites The Llama 3 Herd of Models.

LLaVA-Zip: Adaptive Visual Token Compression with Intrinsic Image Information The Llama 3 Herd of Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T17:38:56.867338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:38:56.867338Z digest=sha256:18dc48b77b4c3366fdd3d03c29d4359ed28c462fd1ac1be16ed619378fd9ec29

Observation 885f5b81-3a33-4fd7-9e6d-36c80160fe4d · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

LLaVA-Zip: Adaptive Visual Token Compression with Intrinsic Image Information Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T17:38:56.881537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:38:56.881537Z digest=sha256:32110f1e6f25ccd40208901c10ca3335c530a4e3e0ca4ed574b0c9914d88d26a

Pith citing papers

No inbound Pith citation observations are available.