Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T03:42:16.353153Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 0 inbound Pith citation observations for arXiv:2607.16326.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T03:42:16.353153Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
30 of 30 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 22e10980-24c3-4109-b9ec-fe6d9a67ee10 · outbound
CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference A survey of state of the art large vision language models: Benchmark evaluations and challenges,
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2e7303d-8d63-4d0b-a2ec-b2d66a03a22c · outbound
CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models,
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ffc841a0-501a-4e7a-8a12-df7467060223 · outbound
CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference [cls] attention is all you need for training-free visual token pruning: Make vlm inference faster,
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb80ec59-6fda-4a39-a281-9eb459c21c3f · outbound
CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference Stop Looking for Important Tokens in Multimodal Language Models: Duplication Matters More
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fcb18766-f096-4f16-9ae0-344f2ddf65fd · outbound
CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference Visionzip: Longer is better but not necessary in vision language models,
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3f840ac-ea02-43e2-9cf5-8fcc758c8011 · outbound
CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference Divprune: Diversity- based visual token pruning for large multimodal models,
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2747e3b-6ba8-4f41-84f1-43088a6a4d92 · outbound
CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference Sparsevlm: Visual token sparsification for efficient vision-language model inference,
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation deb8fc81-c2ea-47af-a648-ca7099b13bd2 · outbound
CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference Which experimental design is better suited for vqa tasks?: Eye tracking study on cognitive load, performance, and gaze allocations,
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf119c7b-0c1e-4101-8c49-904514bb2413 · outbound
CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference Making the invisible visible: Verbal but not visual cues enhance visual detection,
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84348696-6677-4790-919f-ea30c9956231 · outbound
CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference Beyond text-visual attention: Exploiting visual cues for effective token pruning in vlms,
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation afcfed36-0ac4-44ac-8e3d-7e56bab9427b · outbound
CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a87a12c8-fe66-4e61-b471-5a5b9fb3f6f8 · outbound
CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference Improved baselines with visual instruction tuning,
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f349ea5-b27c-4a40-aabb-1980a500fb4f · outbound
CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference Llava-next: Improved reasoning, ocr, and world knowledge,
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b78eb07-ce38-47e4-906f-82d0d86ed7e6 · outbound
CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference A-vl: Adaptive attention for large vision-language models,
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43bb17fe-eb23-4fad-ba9e-8ac6e71a801b · outbound
CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference Llava-prumerge: Adaptive token reduction for efficient large multimodal models,
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6c1fd54-da0e-4dc4-85b6-278d6826381c · outbound
CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference Hired: Attention-guided token dropping for efficient inference of high-resolution vision-language models,
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e8570fb-3b5f-477d-a0d0-4f7da959d9b4 · outbound
CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e70aee6f-94b2-4f0e-893a-450614b14839 · outbound
CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference Token merging: Your vit but faster,
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9be3d9bc-25e7-4233-9703-6ef794af667e · outbound
CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference Less is more: A simple yet effective token reduction method for efficient multi-modal llms,
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2cac17e1-2c7a-41fb-9ea6-fe26463ff5d8 · outbound
CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb6895fa-9c7e-47b1-948a-5c12c1dfccec · outbound
CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference en core web sm: spacy english small model,
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa5edcd7-3e17-4763-95cc-a265e3619c42 · outbound
CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference Making the v in vqa matter: Elevating the role of image understanding in visual question answering,
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ee3588a-2ed4-48ec-a572-b6008311a5bb · outbound
CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference Vizwiz grand challenge: Answering visual questions from blind people,
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b291d16-1ac2-4445-9add-45ac222e2926 · outbound
CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference Gqa: A new dataset for real-world visual reasoning and compositional question answering,
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e23e9de6-19a2-423b-9be8-701099206de4 · outbound
CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference Learn to explain: Multimodal reasoning via thought chains for science question answering,
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ac73fdb-ead5-4f4c-a603-1df1de942469 · outbound
CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference Towards vqa models that can read,
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0dc2b26-e124-494d-8b8a-f0c40c015cc9 · outbound
CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference Evaluating object hallucination in large vision-language models,
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf4a3d85-1401-493d-9782-227dffb6d8ee · outbound
CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ccc400f-956c-4576-9587-b505246e1706 · outbound
CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference Mmbench: Is your multi-modal model an all-around player?
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6462f743-c873-4b9f-8e30-c2f19f1637fd · outbound
CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference Mm-vet: evaluating large multimodal models for integrated capabilities,
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.