Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 23 inbound Pith citation observations for arXiv:2502.17425.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T17:46:28.298491Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-10T07:36:57.997924Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation d7787977-b193-42c0-a166-9520d8a30dd0 · inbound
VeriThinker: Learning to Verify Makes Reasoning Model Efficient Introducing Visual Perception Token into Multimodal Large Language Model
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59e799c9-745e-4f80-8775-79913f385c70 · inbound
VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Introducing Visual Perception Token into Multimodal Large Language Model
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15fad10e-f36d-4452-91f7-f6b04cb1e2f2 · inbound
Qwen Look Again: Guiding Vision-Language Reasoning Models to Re-attention Visual Information Introducing Visual Perception Token into Multimodal Large Language Model
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7c3e071-3a60-42c5-97e2-4e2182473919 · inbound
MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought Reasoning Introducing Visual Perception Token into Multimodal Large Language Model
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c9b82e6-839c-49f1-9544-ad7b2b8ada81 · inbound
Enhancing Spatial Reasoning through Visual and Textual Thinking Introducing Visual Perception Token into Multimodal Large Language Model
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a86ab97-3722-41b1-92e4-34065fbedba8 · inbound
Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Introducing Visual Perception Token into Multimodal Large Language Model
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38802939-0095-4799-931c-53674b9679dc · inbound
MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs Introducing Visual Perception Token into Multimodal Large Language Model
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 5d91ee8c-713d-4a67-8c63-db2c2e9241c1 · inbound
MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs Introducing Visual Perception Token into Multimodal Large Language Model
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation f19d2e7b-d5e3-4403-a4dc-f0baa4491e32 · inbound
MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs Introducing Visual Perception Token into Multimodal Large Language Model
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c78b2ec3-ebf7-47d3-9400-cb82fd17fa16 · inbound
Token Warping Helps MLLMs Look from Nearby Viewpoints Introducing Visual Perception Token into Multimodal Large Language Model
Reference 115
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 7795feb2-0425-4bc2-b194-dc9cc6212f70 · inbound
Visual Enhanced Depth Scaling for Multimodal Latent Reasoning Introducing Visual Perception Token into Multimodal Large Language Model
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation b25ca432-1a0f-4544-822d-9ab5f7a9720e · inbound
Visual Enhanced Depth Scaling for Multimodal Latent Reasoning Introducing Visual Perception Token into Multimodal Large Language Model
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 929cf2b2-0a0f-4923-83d1-dae6d1b70c35 · inbound
Visual Enhanced Depth Scaling for Multimodal Latent Reasoning Introducing Visual Perception Token into Multimodal Large Language Model
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation ad7551bf-0755-407d-bbea-79925e110763 · inbound
Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs Introducing Visual Perception Token into Multimodal Large Language Model
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 59d66061-9aeb-4188-bce1-40be84d536aa · inbound
From Failure to Feedback: Group Revision Unlocks Hard Cases in Object-Level Grounding Introducing Visual Perception Token into Multimodal Large Language Model
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation ec8d8ba5-57cd-40f9-85d4-db1af218863e · inbound
ROVER: Routing Object-Centric Visual Evidence for Grounded Multi-Image Reasoning Introducing Visual Perception Token into Multimodal Large Language Model
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 14c371c7-2ec0-4565-88da-64bd82c79528 · inbound
MathVis-Fine: Aligning Visual Supervision with Necessity via Progressive Dependency-Guided Training for Multimodal Mathematical Reasoning Introducing Visual Perception Token into Multimodal Large Language Model
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 65361e8c-8d12-408b-8624-dcf4fb2cfda2 · inbound
From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models Introducing Visual Perception Token into Multimodal Large Language Model
Reference 139
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation e8758001-cf6e-4400-8108-1c661c2ce96b · inbound
Latent Noise Mask for Reducing Visual Redundancy in Multimodal Large Language Models Introducing Visual Perception Token into Multimodal Large Language Model
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 63e626de-969a-4252-b038-bb09a77b5df3 · inbound
DeltaV: Thinking with Visual State Updates in Unified Large Multimodal Models Introducing Visual Perception Token into Multimodal Large Language Model
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation c3b6a2db-9f62-4656-9d02-1af7581f551f · inbound
Visual Access Boundaries in Vision-Language Model Reasoning Introducing Visual Perception Token into Multimodal Large Language Model
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82654470-ac0b-4dcf-aea7-a6cf86912c7e · inbound
Contextualized Evaluation of Vision Language Models through Dynamic, Multi-turn Interactions Introducing Visual Perception Token into Multimodal Large Language Model
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db43555a-3cc1-4100-86fd-a1e9da023427 · inbound
Mitigating Visual Degradation in MLLMs via Spatial-Spectral Visual Anchor Learning Introducing Visual Perception Token into Multimodal Large Language Model
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.