Pith. sign in

Paper Citation Record · LEDGER

When Do Fewer Visual Tokens Accelerate Multimodal Inference? A Break-Even Study Across Decision Locations and Hardware

As of 8 August 2026, this Paper Citation Record lists 12 of 12 outbound references and 0 inbound Pith citation observations for arXiv:2608.03649.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.03649 v1

Coverage vector

measured 12 of 12 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T15:19:02.848690Z

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

12 of 12 outbound references displayed

  • verified exact2
  • verified fuzzy1
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a9ac022f-09d5-4619-ada0-72b387edb7e8 · outbound

This paper cites Qwen2.5-VL Technical Report.

When Do Fewer Visual Tokens Accelerate Multimodal Inference? A Break-Even Study Across Decision Locations and Hardware Qwen2.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:02.785738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:02.785738Z digest=sha256:cce4ed4946b13f445e50c2621b522b1fa9e3d7a672f46c8c6d80160c8d74a7ec

Observation 12a72b7e-cfe8-4a6d-a7f5-4f4b45bf6057 · outbound

This paper cites Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering.

When Do Fewer Visual Tokens Accelerate Multimodal Inference? A Break-Even Study Across Decision Locations and Hardware Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:02.790931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:02.790931Z digest=sha256:1ae68f8827fc57df253faea25fbd622509f3d765fb8af3a744f926b3c072589c

Observation c58f92f9-2a10-4044-b4dd-ae0044801909 · outbound

This paper cites CARES : Context-aware resolution selector for VLM s.

When Do Fewer Visual Tokens Accelerate Multimodal Inference? A Break-Even Study Across Decision Locations and Hardware CARES : Context-aware resolution selector for VLM s

Reference 3

Resolution
verified exact
doi, observed 2026-08-05T15:19:03.166692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-05T15:19:02.799318Z digest=sha256:a220e04bdaa7e83f6e7571cb7b2e72ff7154bf70e9981ad1e2f1001f3afbd4a5

Observation a842a616-3f4c-4e81-9abb-82e1759158ba · outbound

This paper cites TokenPacker: Efficient Visual Projector for Multimodal LLM.

When Do Fewer Visual Tokens Accelerate Multimodal Inference? A Break-Even Study Across Decision Locations and Hardware TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:02.805044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:02.805044Z digest=sha256:6954b559e678144abd5e4283d4ddb844b8c47ef405ced7b61a8688d3010af87b

Observation 8b1fc09e-7b6b-41d9-9f83-aa2e0de2ad58 · outbound

This paper cites ResAdapt : Adaptive resolution for efficient multimodal reasoning.

When Do Fewer Visual Tokens Accelerate Multimodal Inference? A Break-Even Study Across Decision Locations and Hardware ResAdapt : Adaptive resolution for efficient multimodal reasoning

Reference 5

Resolution
verified exact
doi, observed 2026-08-05T15:19:03.130483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-05T15:19:02.810877Z digest=sha256:0170b0867350fbdb1b3b870b86a50ffacd880adb0997f0c5d976eacec3f54d14

Observation cbff600b-51cf-4dc4-b64e-65e8b4e056c8 · outbound

This paper cites AdaptVision : Efficient vision-language models via adaptive visual acquisition.

When Do Fewer Visual Tokens Accelerate Multimodal Inference? A Break-Even Study Across Decision Locations and Hardware AdaptVision : Efficient vision-language models via adaptive visual acquisition

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:19:03.342589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-05T15:19:02.815910Z digest=sha256:df13544287076b2cda8182b0bec3dd45322c49799629df6ab6e6c01164b4de59

Observation 4b7b91bc-e4f1-478f-b1b8-e9ffb00a8c82 · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

When Do Fewer Visual Tokens Accelerate Multimodal Inference? A Break-Even Study Across Decision Locations and Hardware ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:02.822062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:02.822062Z digest=sha256:decc3e465b971eeb210a33e56f28a8cb94dce3e806324824a1f776eeee23fbf8

Observation 477b4d13-068d-49db-ad08-8707a305e08d · outbound

This paper cites LLaVA-PruMerge : Adaptive token reduction for efficient large multimodal models.

When Do Fewer Visual Tokens Accelerate Multimodal Inference? A Break-Even Study Across Decision Locations and Hardware LLaVA-PruMerge : Adaptive token reduction for efficient large multimodal models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:02.827885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:02.827885Z digest=sha256:1200885cb8668d05f354e58239badf1052ec1403b2a5305da71c523fd7f0889c

Observation 40a67815-76bc-4e7b-9c45-59bb231065ae · outbound

This paper cites Towards VQA models that can read.

When Do Fewer Visual Tokens Accelerate Multimodal Inference? A Break-Even Study Across Decision Locations and Hardware Towards VQA models that can read

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:02.833221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:02.833221Z digest=sha256:2f88b1030704d4301a54d304369974e898827d7934a12bb4556af8381d6ab9c4

Observation 03fcfec7-111b-4c18-8d98-5ee76158c127 · outbound

This paper cites Fit and Prune: Fast and Training-free Visual Token Pruning for Multi-modal Large Language Models.

When Do Fewer Visual Tokens Accelerate Multimodal Inference? A Break-Even Study Across Decision Locations and Hardware Fit and Prune: Fast and Training-free Visual Token Pruning for Multi-modal Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:02.837652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:02.837652Z digest=sha256:e01caec3bc556bfbbf32e48c673fbaf3654d225e2bbc877a73564fca9fcd6fb2

Observation b8e56e5f-2d4b-4c0a-ba66-b94aabab9a82 · outbound

This paper cites Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs.

When Do Fewer Visual Tokens Accelerate Multimodal Inference? A Break-Even Study Across Decision Locations and Hardware Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:02.843428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:02.843428Z digest=sha256:100598642d305c0c7674cfb4f525299d1cf0d231d3f6664757649c808517a7a8

Observation 4a8e9594-1b6f-4d93-94d8-584de078c6d5 · outbound

This paper cites SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference.

When Do Fewer Visual Tokens Accelerate Multimodal Inference? A Break-Even Study Across Decision Locations and Hardware SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:02.848690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:02.848690Z digest=sha256:10739b20e088f9dad2421350b42ce97d863aea4c8e495789bb49e96584449aea

Pith citing papers

No inbound Pith citation observations are available.