Pith. sign in

Paper Citation Record · LEDGER

CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference

As of 9 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 0 inbound Pith citation observations for arXiv:2607.16326.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.16326 v1

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T03:42:16.353153Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

30 of 30 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved30
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 22e10980-24c3-4109-b9ec-fe6d9a67ee10 · outbound

This paper cites A survey of state of the art large vision language models: Benchmark evaluations and challenges,.

CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference A survey of state of the art large vision language models: Benchmark evaluations and challenges,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T03:42:13.962870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:42:13.962870Z digest=sha256:9dcb025266ed92f3eed6344e0f011ab0455e66d21e884f1f142b0bf3e903f72f

Observation a2e7303d-8d63-4d0b-a2ec-b2d66a03a22c · outbound

This paper cites An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models,.

CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T03:42:14.029230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:42:14.029230Z digest=sha256:eff33346206750bedcda3c6ff076b8551c60e826f0bbc63272fd28b18fdc267c

Observation ffc841a0-501a-4e7a-8a12-df7467060223 · outbound

This paper cites [cls] attention is all you need for training-free visual token pruning: Make vlm inference faster,.

CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference [cls] attention is all you need for training-free visual token pruning: Make vlm inference faster,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T03:42:14.088093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:42:14.088093Z digest=sha256:da921c4e71eb3c900f7a183c0151a2a115ed7f9c122fa87c8d7e8e1244834c0d

Observation cb80ec59-6fda-4a39-a281-9eb459c21c3f · outbound

This paper cites Stop Looking for Important Tokens in Multimodal Language Models: Duplication Matters More.

CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference Stop Looking for Important Tokens in Multimodal Language Models: Duplication Matters More

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T03:42:14.167493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:42:14.167493Z digest=sha256:8ef7a445bfbc9d36c04447b4641b30058ba7f366419b97081c293e581779b61c

Observation fcb18766-f096-4f16-9ae0-344f2ddf65fd · outbound

This paper cites Visionzip: Longer is better but not necessary in vision language models,.

CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference Visionzip: Longer is better but not necessary in vision language models,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T03:42:14.225585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:42:14.225585Z digest=sha256:64c34e331563e11f59511c21bc09ae680045cca18545c89ca3f162c75f5fd88b

Observation a3f840ac-ea02-43e2-9cf5-8fcc758c8011 · outbound

This paper cites Divprune: Diversity- based visual token pruning for large multimodal models,.

CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference Divprune: Diversity- based visual token pruning for large multimodal models,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T03:42:14.268413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:42:14.268413Z digest=sha256:bc2be19a292da2986e9f0e66458be0eac040840263139eaaba95c30f7324411d

Observation b2747e3b-6ba8-4f41-84f1-43088a6a4d92 · outbound

This paper cites Sparsevlm: Visual token sparsification for efficient vision-language model inference,.

CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference Sparsevlm: Visual token sparsification for efficient vision-language model inference,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T03:42:14.328001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:42:14.328001Z digest=sha256:a3457f736e057216974af45efb0a42fcfd3fbc726f94f68b2c82b0e212258d18

Observation deb8fc81-c2ea-47af-a648-ca7099b13bd2 · outbound

This paper cites Which experimental design is better suited for vqa tasks?: Eye tracking study on cognitive load, performance, and gaze allocations,.

CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference Which experimental design is better suited for vqa tasks?: Eye tracking study on cognitive load, performance, and gaze allocations,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T03:42:14.387140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:42:14.387140Z digest=sha256:1be393efa2f68700ba65004249a3a78526e2e36c79e3e558d0aeabad1e7cb92d

Observation cf119c7b-0c1e-4101-8c49-904514bb2413 · outbound

This paper cites Making the invisible visible: Verbal but not visual cues enhance visual detection,.

CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference Making the invisible visible: Verbal but not visual cues enhance visual detection,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T03:42:14.451408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:42:14.451408Z digest=sha256:c1586067f89dd4808fad7fb9c05bb3968e9905000d3fbb83910257dfff0ad30a

Observation 84348696-6677-4790-919f-ea30c9956231 · outbound

This paper cites Beyond text-visual attention: Exploiting visual cues for effective token pruning in vlms,.

CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference Beyond text-visual attention: Exploiting visual cues for effective token pruning in vlms,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T03:42:14.510468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:42:14.510468Z digest=sha256:e8894ecbe71b0cb45dfa248ecaab3c3d707b03925b9485e5229329cdcb37d5c3

Observation afcfed36-0ac4-44ac-8e3d-7e56bab9427b · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T03:42:14.568395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:42:14.568395Z digest=sha256:55019bf066c3a477069e02beb2b896022a3fda36bac6bb5a9e67f5b60874c267

Observation a87a12c8-fe66-4e61-b471-5a5b9fb3f6f8 · outbound

This paper cites Improved baselines with visual instruction tuning,.

CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference Improved baselines with visual instruction tuning,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T03:42:14.614845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:42:14.614845Z digest=sha256:8efcdfe82d5ed768649e516a2c184e49c2882ef1ba96dcb44323b9ab45a145c4

Observation 0f349ea5-b27c-4a40-aabb-1980a500fb4f · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge,.

CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference Llava-next: Improved reasoning, ocr, and world knowledge,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T03:42:14.672964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:42:14.672964Z digest=sha256:e180368656c244e1606fd1bf87cef97a7aada7e237b23f4bd52818e34fd1f055

Observation 4b78eb07-ce38-47e4-906f-82d0d86ed7e6 · outbound

This paper cites A-vl: Adaptive attention for large vision-language models,.

CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference A-vl: Adaptive attention for large vision-language models,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T03:42:14.747206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:42:14.747206Z digest=sha256:d05fee9b6211d18052eb6884f1a39bd2759bbf3eb4678759b9405fb4124fca50

Observation 43bb17fe-eb23-4fad-ba9e-8ac6e71a801b · outbound

This paper cites Llava-prumerge: Adaptive token reduction for efficient large multimodal models,.

CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference Llava-prumerge: Adaptive token reduction for efficient large multimodal models,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T03:42:14.810397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:42:14.810397Z digest=sha256:5e6aa85960e34e34195c8cd4d33cf488339de4a0b84a3eb2c6288c1e94532506

Observation a6c1fd54-da0e-4dc4-85b6-278d6826381c · outbound

This paper cites Hired: Attention-guided token dropping for efficient inference of high-resolution vision-language models,.

CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference Hired: Attention-guided token dropping for efficient inference of high-resolution vision-language models,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T03:42:14.869453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:42:14.869453Z digest=sha256:ef43e525689937882f859c544e1365d8ddd254c25114b7eea340999efc90fce7

Observation 7e8570fb-3b5f-477d-a0d0-4f7da959d9b4 · outbound

This paper cites PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction.

CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T03:42:14.917470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:42:14.917470Z digest=sha256:2d9d9f2d9c3e0800dd977d4a71beab448accc385b5c59f378b355955932cf161

Observation e70aee6f-94b2-4f0e-893a-450614b14839 · outbound

This paper cites Token merging: Your vit but faster,.

CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference Token merging: Your vit but faster,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T03:42:14.968369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:42:14.968369Z digest=sha256:39b0a7b1bd9e012c53927a0bce172c5b8282dcd6ce7546f62f3d3252259f06cb

Observation 9be3d9bc-25e7-4233-9703-6ef794af667e · outbound

This paper cites Less is more: A simple yet effective token reduction method for efficient multi-modal llms,.

CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference Less is more: A simple yet effective token reduction method for efficient multi-modal llms,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T03:42:14.973153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:42:14.973153Z digest=sha256:92b185558514724c6fc198a9087959faac10d22e3a1559064de2d4bed290a836

Observation 2cac17e1-2c7a-41fb-9ea6-fe26463ff5d8 · outbound

This paper cites Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model.

CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T03:42:14.989028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:42:14.989028Z digest=sha256:9ae3a3449cd0c90a054f56da06f253cebd777041488fa60efb5311fc1b35960d

Observation eb6895fa-9c7e-47b1-948a-5c12c1dfccec · outbound

This paper cites en core web sm: spacy english small model,.

CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference en core web sm: spacy english small model,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T03:42:15.064940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:42:15.064940Z digest=sha256:4ddd200c6a9a929701e0d7694f26ce49395225c8283b0182e296773604cd13a0

Observation aa5edcd7-3e17-4763-95cc-a265e3619c42 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answering,.

CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference Making the v in vqa matter: Elevating the role of image understanding in visual question answering,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T03:42:15.224141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:42:15.224141Z digest=sha256:24271ead53681f026895ee4be1e2f1296b03a3b63e5563052b164f6f3c2eff84

Observation 6ee3588a-2ed4-48ec-a572-b6008311a5bb · outbound

This paper cites Vizwiz grand challenge: Answering visual questions from blind people,.

CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference Vizwiz grand challenge: Answering visual questions from blind people,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T03:42:15.385887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:42:15.385887Z digest=sha256:26e680e33f52fda8470be43ac8875850c8a981819ec7c44210a625042646c03b

Observation 0b291d16-1ac2-4445-9add-45ac222e2926 · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering,.

CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference Gqa: A new dataset for real-world visual reasoning and compositional question answering,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T03:42:15.538454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:42:15.538454Z digest=sha256:73f3e3b1aeca5ff2109ac84994d32a49c4f781c66e1e5bebc0bab45327d61945

Observation e23e9de6-19a2-423b-9be8-701099206de4 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering,.

CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference Learn to explain: Multimodal reasoning via thought chains for science question answering,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T03:42:15.652552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:42:15.652552Z digest=sha256:986db579f135da12cebf3d546f99e072d6855d065b40161f21447fe6883bbea9

Observation 6ac73fdb-ead5-4f4c-a603-1df1de942469 · outbound

This paper cites Towards vqa models that can read,.

CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference Towards vqa models that can read,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T03:42:15.811440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:42:15.811440Z digest=sha256:3bc40c89d8f2b989bea46e8fc7ad00b36c4c0b45e66e3c73d343e0ea22c5d534

Observation a0dc2b26-e124-494d-8b8a-f0c40c015cc9 · outbound

This paper cites Evaluating object hallucination in large vision-language models,.

CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference Evaluating object hallucination in large vision-language models,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T03:42:15.972178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:42:15.972178Z digest=sha256:2d779bc56b3b0e276061fa562c696dfe0b8597785dd9abb43d100ed239cf4564

Observation cf4a3d85-1401-493d-9782-227dffb6d8ee · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T03:42:16.140370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:42:16.140370Z digest=sha256:8768a6940e46aaca0345f7b3cce8f14541eb0925daf694df465563d90ac15f58

Observation 9ccc400f-956c-4576-9587-b505246e1706 · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player?.

CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference Mmbench: Is your multi-modal model an all-around player?

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T03:42:16.307600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:42:16.307600Z digest=sha256:69d3126bd2bb027ba85acc511a604ef55370921012e4d6ff5297d85e6b947931

Observation 6462f743-c873-4b9f-8e30-c2f19f1637fd · outbound

This paper cites Mm-vet: evaluating large multimodal models for integrated capabilities,.

CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference Mm-vet: evaluating large multimodal models for integrated capabilities,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T03:42:16.353153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:42:16.353153Z digest=sha256:1aea315136047caa2808c33fc972a5f909cebecc863841e1bd9ae6a17ee8ae97

Pith citing papers

No inbound Pith citation observations are available.