Pith. sign in

Paper Citation Record · LEDGER

VisionPulse: Dynamic Visual Sparsity for Efficient Multimodal Reasoning

As of 12 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 0 inbound Pith citation observations for arXiv:2605.31457.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.31457 v1

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-28T23:13:15.593994Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

24 of 24 outbound references displayed

  • verified exact17
  • verified fuzzy0
  • unresolved3
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch4

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fd4d505e-01e1-42d6-975f-b038b984bff1 · outbound

This paper cites Qwen3-VL Technical Report.

VisionPulse: Dynamic Visual Sparsity for Efficient Multimodal Reasoning Qwen3-VL Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-06-29T00:12:50.578005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-28T23:13:15.593994Z digest=sha256:821a055f131d1e06bc30de9a0002a5b61c67af1facb640fff57a8f30ef42115a

Observation 997dde5f-5dcc-44bc-9f54-065a0fbd3774 · outbound

This paper cites Think Clearly: Improving Reasoning via Redundant Token Pruning.

VisionPulse: Dynamic Visual Sparsity for Efficient Multimodal Reasoning Think Clearly: Improving Reasoning via Redundant Token Pruning

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T00:12:50.580956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-28T23:13:15.593994Z digest=sha256:cb925aaefb15b1e3a50cb0ed68c50ed4aa29622ca7daa5cdbb889e27a2c99ec7

Observation 31b1ba56-24a3-49a8-92b8-a5e6d4485e5e · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

VisionPulse: Dynamic Visual Sparsity for Efficient Multimodal Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-06-29T00:12:50.586292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-28T23:13:15.593994Z digest=sha256:06a009e0fc8c9fa81e81163bdc123071dd26a9be5205b556bccfeda1939605ba

Observation 5c8640b3-9191-4a69-8019-a3f804756657 · outbound

This paper cites OpenAI o1 System Card.

VisionPulse: Dynamic Visual Sparsity for Efficient Multimodal Reasoning OpenAI o1 System Card

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-06-29T00:12:50.586463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-28T23:13:15.593994Z digest=sha256:b2938e24bd2008364d3a4ae79575701292bd7f3ddfa7eee08f636b8fd5ee86d6

Observation 494d7288-de93-4018-834a-47fb4884390c · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

VisionPulse: Dynamic Visual Sparsity for Efficient Multimodal Reasoning LLaVA-OneVision: Easy Visual Task Transfer

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-06-29T00:12:50.584225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-28T23:13:15.593994Z digest=sha256:905269c5a676ae4bda4af8225db399a6640d3d5649ea0466db88cf223cc7b70c

Observation c000258d-a7ff-47ae-9c8c-a15336907d71 · outbound

This paper cites Not All Patches are What You Need: Expediting Vision Transformers via Token Reorganizations.

VisionPulse: Dynamic Visual Sparsity for Efficient Multimodal Reasoning Not All Patches are What You Need: Expediting Vision Transformers via Token Reorganizations

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-06-29T00:12:50.564283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-28T23:13:15.593994Z digest=sha256:1e5be9a52a0b8ff628b6780cdc040c9c2faaac89c97802f45ff6df4e34211d7c

Observation 0cacf445-a2be-4995-a822-07bc6d941093 · outbound

This paper cites HiPrune: Hierarchical Attention for Efficient Token Pruning in Vision-Language Models.

VisionPulse: Dynamic Visual Sparsity for Efficient Multimodal Reasoning HiPrune: Hierarchical Attention for Efficient Token Pruning in Vision-Language Models

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-06-29T00:12:50.588896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-28T23:13:15.593994Z digest=sha256:ca26bfb4a5f3ecda109c28bf0b60af504580b3940add7cd5cd919e83a933e6ae

Observation d5c5d5ae-8a3d-4f3c-b403-ef3e649f7a75 · outbound

This paper cites L., Tan, J.

VisionPulse: Dynamic Visual Sparsity for Efficient Multimodal Reasoning L., Tan, J

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-28T23:13:15.593994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:13:15.593994Z digest=sha256:3625be7a27f2757b781f07c6eaac08e055786246ac665348bf428c2d46fac663

Observation f6f59feb-9fda-4e9d-bea5-f7a375758e55 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

VisionPulse: Dynamic Visual Sparsity for Efficient Multimodal Reasoning Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-06-29T00:12:50.541483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-28T23:13:15.593994Z digest=sha256:d21e7c4bf081af20ce25452d680e913f695f36aa7651ee11b002434fdf96beaa

Observation ecfb1240-c0a0-4fb5-a808-9049115ef076 · outbound

This paper cites Kimi-VL Technical Report.

VisionPulse: Dynamic Visual Sparsity for Efficient Multimodal Reasoning Kimi-VL Technical Report

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-06-29T00:12:50.569596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-28T23:13:15.593994Z digest=sha256:a7ea348454dd184eccec6f24cc440cc52dcfb42ab1498a52e53538761bf31159

Observation 1c2908e8-a975-435b-a5e8-182033b85bd2 · outbound

This paper cites LOOK-M: Look-Once Optimization in KV Cache for Efficient Multimodal Long-Context Inference.

VisionPulse: Dynamic Visual Sparsity for Efficient Multimodal Reasoning LOOK-M: Look-Once Optimization in KV Cache for Efficient Multimodal Long-Context Inference

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-06-29T00:12:50.575478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-28T23:13:15.593994Z digest=sha256:752bd1fc5833a449001cfd151dff67358bb165f2da7621525bfd1d21324212e2

Observation 392264aa-5eb5-4da6-a67e-93caf0c54a51 · outbound

This paper cites ONLY: One-Layer Intervention Sufficiently Mitigates Hallucinations in Large Vision-Language Models.

VisionPulse: Dynamic Visual Sparsity for Efficient Multimodal Reasoning ONLY: One-Layer Intervention Sufficiently Mitigates Hallucinations in Large Vision-Language Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-06-29T00:12:50.561358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-28T23:13:15.593994Z digest=sha256:600213b968f6260de07acb878695c800d259393997194d1434707cd8a82e8ab4

Observation 78efca44-1cf1-48dc-9efd-9bf72f00011b · outbound

This paper cites VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning.

VisionPulse: Dynamic Visual Sparsity for Efficient Multimodal Reasoning VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-06-29T00:12:50.573182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-28T23:13:15.593994Z digest=sha256:6fe2acd571260b607332e5581ef0c23a33ee4ab23c3181eff8ed1bd65fd8a368

Observation a664e121-c2e0-42cc-9158-cf2bd93c1ac6 · outbound

This paper cites Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey.

VisionPulse: Dynamic Visual Sparsity for Efficient Multimodal Reasoning Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-06-29T00:12:50.581682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-28T23:13:15.593994Z digest=sha256:1107e71bcfdba9c831acb8561d9c4dc8287989fd16226a4ea32ab792702c8271

Observation 6fd908c0-611d-4160-adf7-a299b677d698 · outbound

This paper cites LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding.

VisionPulse: Dynamic Visual Sparsity for Efficient Multimodal Reasoning LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-06-29T00:12:50.596604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-28T23:13:15.593994Z digest=sha256:4fb83f603a35403839a7e7088efa031a6a96137a6e2e3dcd98cda48c6a607bfe

Observation b9586db2-5c1a-4ebe-ade8-410680ee40f5 · outbound

This paper cites Tokenskip: Controllable chain-of-thought compression in llms.arXiv preprint arXiv:2502.12067.

VisionPulse: Dynamic Visual Sparsity for Efficient Multimodal Reasoning Tokenskip: Controllable chain-of-thought compression in llms.arXiv preprint arXiv:2502.12067

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-06-29T00:12:50.578033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-28T23:13:15.593994Z digest=sha256:bdd5e6fffcffb888c7118c9b0cd7d87cefd0f548cef413f457b7c42bf7f2894b

Observation 2ac26f35-e0db-4748-85dd-67f48dec80ee · outbound

This paper cites arXiv preprint arXiv:2504.18458 , year=.

VisionPulse: Dynamic Visual Sparsity for Efficient Multimodal Reasoning arXiv preprint arXiv:2504.18458 , year=

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T00:12:50.567176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-28T23:13:15.593994Z digest=sha256:2cb655de19532fbd3a175babf2f6e20647bae12b4034ab0f1b4ada00a38d7529

Observation 85b4cc93-84c7-4410-802e-2f00d4df2a9e · outbound

This paper cites PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction.

VisionPulse: Dynamic Visual Sparsity for Efficient Multimodal Reasoning PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-06-29T00:12:50.569823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-28T23:13:15.593994Z digest=sha256:6b3ffda99c2e9bc59d74fde0f7306d09f1e5da6ff34067d14fe0e5a0f2bf96e5

Observation 141616eb-2574-40b2-8685-dae74dd5d590 · outbound

This paper cites Look-Back: Implicit Visual Re-focusing in MLLM Reasoning.

VisionPulse: Dynamic Visual Sparsity for Efficient Multimodal Reasoning Look-Back: Implicit Visual Re-focusing in MLLM Reasoning

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T00:12:50.591396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-28T23:13:15.593994Z digest=sha256:3df3eb921496733995deeacd8c4d1d0bbd04abdd7d3571ee346eb57c8b6d6354

Observation d9a5c939-091c-4390-9be3-727c6c91a4dd · outbound

This paper cites Demystifying Long Chain-of-Thought Reasoning in LLMs.

VisionPulse: Dynamic Visual Sparsity for Efficient Multimodal Reasoning Demystifying Long Chain-of-Thought Reasoning in LLMs

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-06-29T00:12:50.561905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-28T23:13:15.593994Z digest=sha256:9f18cee7b1e7d3f466fca8a9a61ea2dd58950ac122d526a170d2f9c490543e18

Observation e4e7d9a4-4ab9-404c-aa76-6d5a1505cf3c · outbound

This paper cites Unified Visual Transformer Compression.

VisionPulse: Dynamic Visual Sparsity for Efficient Multimodal Reasoning Unified Visual Transformer Compression

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-06-29T00:12:50.553383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-28T23:13:15.593994Z digest=sha256:38dff5a55d75fba17b70f099b60d1f818a569dae3b31faafe6514a76afa1f93b

Observation 4c923cbf-ef0c-43e1-9256-0d4e6636c162 · outbound

This paper cites AdaptThink: Reasoning Models Can Learn When to Think.

VisionPulse: Dynamic Visual Sparsity for Efficient Multimodal Reasoning AdaptThink: Reasoning Models Can Learn When to Think

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T00:12:50.594042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-28T23:13:15.593994Z digest=sha256:724ae1a7d387f1e6b3a6c1f2b8b2fce5122959cce40937cde6b6420702fa6eee

Observation e2b5de62-db43-45a2-97aa-08f36abe5aaa · outbound

This paper cites Additional Analysis A.1.

VisionPulse: Dynamic Visual Sparsity for Efficient Multimodal Reasoning Additional Analysis A.1

Reference 23

Resolution
unresolved
no resolver link, observed 2026-06-28T23:13:15.593994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:13:15.593994Z digest=sha256:c3b5b4522e77b33ac208dc8f9331a29441496a489b99fbe91f199a035376ca60

Observation c041c2ca-250b-4483-a790-3a80bbab479c · outbound

This paper cites This suggests that VisionPulse is orthogonal to static visual token pruning methods such as FastV and can be combined with them for additional efficiency improvements.

VisionPulse: Dynamic Visual Sparsity for Efficient Multimodal Reasoning This suggests that VisionPulse is orthogonal to static visual token pruning methods such as FastV and can be combined with them for additional efficiency improvements

Reference 24

Resolution
unresolved
no resolver link, observed 2026-06-28T23:13:15.593994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T23:13:15.593994Z digest=sha256:f594e6147e47a2031908c9b134f049caeb140e5a6cd8785bae3f4586a3f687a6

Pith citing papers

No inbound Pith citation observations are available.