Pith. sign in

Paper Citation Record · LEDGER

Introducing Visual Perception Token into Multimodal Large Language Model

As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 23 inbound Pith citation observations for arXiv:2502.17425.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.17425 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 23 of 23 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:46:28.298491Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T07:36:57.997924Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d7787977-b193-42c0-a166-9520d8a30dd0 · inbound

VeriThinker: Learning to Verify Makes Reasoning Model Efficient cites this paper.

VeriThinker: Learning to Verify Makes Reasoning Model Efficient Introducing Visual Perception Token into Multimodal Large Language Model

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:15.423664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:15.423664Z digest=sha256:b8e69242ccc91c32dccd03b2718292ac5b8f28b2c10e31f60fed7ff878a085cf

Observation 59e799c9-745e-4f80-8775-79913f385c70 · inbound

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning cites this paper.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Introducing Visual Perception Token into Multimodal Large Language Model

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:54.345474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:54.345474Z digest=sha256:25549d4f9ef220ba6e0877838e11a07eee90813919638bc299151459fc80f361

Observation 15fad10e-f36d-4452-91f7-f6b04cb1e2f2 · inbound

Qwen Look Again: Guiding Vision-Language Reasoning Models to Re-attention Visual Information cites this paper.

Qwen Look Again: Guiding Vision-Language Reasoning Models to Re-attention Visual Information Introducing Visual Perception Token into Multimodal Large Language Model

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T12:47:01.250685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:47:01.250685Z digest=sha256:28b3692aa46c8490ab54065279fdeacf0d0576c774441988b01be316085c7e2a

Observation d7c3e071-3a60-42c5-97e2-4e2182473919 · inbound

MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought Reasoning cites this paper.

MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought Reasoning Introducing Visual Perception Token into Multimodal Large Language Model

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:54.026686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:54.026686Z digest=sha256:7d854fa2f11fbec3e63f0ce3509016a292529a525b9d197e129134c33a63fe17

Observation 4c9b82e6-839c-49f1-9544-ad7b2b8ada81 · inbound

Enhancing Spatial Reasoning through Visual and Textual Thinking cites this paper.

Enhancing Spatial Reasoning through Visual and Textual Thinking Introducing Visual Perception Token into Multimodal Large Language Model

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T17:46:28.298491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:46:28.298491Z digest=sha256:ca32416551a5793acc0b463366389c2ff94605a7fca1d3c2082a255f31f5182a

Observation 5a86ab97-3722-41b1-92e4-34065fbedba8 · inbound

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding cites this paper.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Introducing Visual Perception Token into Multimodal Large Language Model

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:38.749150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:38.749150Z digest=sha256:3b604065f112c6361d66b4de39241d4d2fb20953cc4720ded06192a6c0982918

Observation 38802939-0095-4799-931c-53674b9679dc · inbound

MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs cites this paper.

MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs Introducing Visual Perception Token into Multimodal Large Language Model

Reference 95

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:16:34.442685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-15T20:12:46.385646Z digest=sha256:bd455702c21346fd653d2975b357229ccf4efaebe83c2285103ceb146ca7cb40

Observation 5d91ee8c-713d-4a67-8c63-db2c2e9241c1 · inbound

MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs cites this paper.

MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs Introducing Visual Perception Token into Multimodal Large Language Model

Reference 95

Resolution
verified exact
arxiv_id, observed 2026-05-22T10:31:25.388070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-22T10:30:06.829915Z digest=sha256:991bff172a4ad965456f7766a377e13f2d574bcef2dc990077193ebe6668b414

Observation f19d2e7b-d5e3-4403-a4dc-f0baa4491e32 · inbound

MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs cites this paper.

MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs Introducing Visual Perception Token into Multimodal Large Language Model

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-02T22:00:12.709967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:00:12.709967Z digest=sha256:22d582b4f149f4e833ad8b338d90f2c59349fe277cf68a0bb8297e1f76c875cb

Observation c78b2ec3-ebf7-47d3-9400-cb82fd17fa16 · inbound

Token Warping Helps MLLMs Look from Nearby Viewpoints cites this paper.

Token Warping Helps MLLMs Look from Nearby Viewpoints Introducing Visual Perception Token into Multimodal Large Language Model

Reference 115

Resolution
verified exact
arxiv_id, observed 2026-05-13T21:08:17.430134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-13T21:07:55.062113Z digest=sha256:44bfbc05303474a51157aa6459fa2dc42f6a6c4e023f71f74f3a923c30ada73b

Observation 7795feb2-0425-4bc2-b194-dc9cc6212f70 · inbound

Visual Enhanced Depth Scaling for Multimodal Latent Reasoning cites this paper.

Visual Enhanced Depth Scaling for Multimodal Latent Reasoning Introducing Visual Perception Token into Multimodal Large Language Model

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:11:04.430716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T16:46:36.010169Z digest=sha256:5893571632fdb6fce6c2b7fc1d614384f29bc142585326a3a54830e28a432ce2

Observation b25ca432-1a0f-4544-822d-9ab5f7a9720e · inbound

Visual Enhanced Depth Scaling for Multimodal Latent Reasoning cites this paper.

Visual Enhanced Depth Scaling for Multimodal Latent Reasoning Introducing Visual Perception Token into Multimodal Large Language Model

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:06:24.110121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-12T04:38:06.774877Z digest=sha256:fb570e1db4a9a9c2eec0acd7e712a68ac6bcf758c170a6145c7ee400aa263f32

Observation 929cf2b2-0a0f-4923-83d1-dae6d1b70c35 · inbound

Visual Enhanced Depth Scaling for Multimodal Latent Reasoning cites this paper.

Visual Enhanced Depth Scaling for Multimodal Latent Reasoning Introducing Visual Perception Token into Multimodal Large Language Model

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:27:29.237405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-13T07:26:59.917150Z digest=sha256:5deb19efac0c4ba6b2882a591c3d9c7d754a6c65deb7bcde37a3747b329e9733

Observation ad7551bf-0755-407d-bbea-79925e110763 · inbound

Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs cites this paper.

Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs Introducing Visual Perception Token into Multimodal Large Language Model

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:26:01.904902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T15:29:25.650175Z digest=sha256:5515534a19bbfc2565fe2ed392fa062b2f7d73582ab86c701d086db99e5a89b0

Observation 59d66061-9aeb-4188-bce1-40be84d536aa · inbound

From Failure to Feedback: Group Revision Unlocks Hard Cases in Object-Level Grounding cites this paper.

From Failure to Feedback: Group Revision Unlocks Hard Cases in Object-Level Grounding Introducing Visual Perception Token into Multimodal Large Language Model

Reference 92

Resolution
verified exact
arxiv_id, observed 2026-05-20T18:43:38.894957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T18:39:11.904941Z digest=sha256:da66b849acea67b0a8e79793047b667c3cd379cb9b705711f2077e1adbf37a41

Observation ec8d8ba5-57cd-40f9-85d4-db1af218863e · inbound

ROVER: Routing Object-Centric Visual Evidence for Grounded Multi-Image Reasoning cites this paper.

ROVER: Routing Object-Centric Visual Evidence for Grounded Multi-Image Reasoning Introducing Visual Perception Token into Multimodal Large Language Model

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:43:28.670041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-29T13:41:44.049230Z digest=sha256:d66e657104a22cbe91862bcba1b653791ccddc7359eeb0b13bc6acb7c9d6ac37

Observation 14c371c7-2ec0-4565-88da-64bd82c79528 · inbound

MathVis-Fine: Aligning Visual Supervision with Necessity via Progressive Dependency-Guided Training for Multimodal Mathematical Reasoning cites this paper.

MathVis-Fine: Aligning Visual Supervision with Necessity via Progressive Dependency-Guided Training for Multimodal Mathematical Reasoning Introducing Visual Perception Token into Multimodal Large Language Model

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T20:18:57.676800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-27T01:23:40.564561Z digest=sha256:1ec8d13c6089c91941157c0ff53c0272d916316ca8e25e2ea317bac54cde691f

Observation 65361e8c-8d12-408b-8624-dcf4fb2cfda2 · inbound

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models cites this paper.

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models Introducing Visual Perception Token into Multimodal Large Language Model

Reference 139

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:09:55.102881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-26T01:50:54.242508Z digest=sha256:ee66c7ac022df007a589ef0d71fc1bfacd5d2be098141e71227a498e8f20fcbb

Observation e8758001-cf6e-4400-8108-1c661c2ce96b · inbound

Latent Noise Mask for Reducing Visual Redundancy in Multimodal Large Language Models cites this paper.

Latent Noise Mask for Reducing Visual Redundancy in Multimodal Large Language Models Introducing Visual Perception Token into Multimodal Large Language Model

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-06-30T06:44:19.799258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T06:35:16.868238Z digest=sha256:c5dfb3a58612b403d0899349d838989af318ea764b9b6e5e2dbbfedee227fb41

Observation 63e626de-969a-4252-b038-bb09a77b5df3 · inbound

DeltaV: Thinking with Visual State Updates in Unified Large Multimodal Models cites this paper.

DeltaV: Thinking with Visual State Updates in Unified Large Multimodal Models Introducing Visual Perception Token into Multimodal Large Language Model

Reference 19

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T07:36:57.999269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-10T07:31:26.225257Z digest=sha256:b312bbb5a011c1bc96a5d1bca39322a326862b9a8db7d12aab9940e2800f058a

Observation c3b6a2db-9f62-4656-9d02-1af7581f551f · inbound

Visual Access Boundaries in Vision-Language Model Reasoning cites this paper.

Visual Access Boundaries in Vision-Language Model Reasoning Introducing Visual Perception Token into Multimodal Large Language Model

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T06:24:29.245780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:24:29.245780Z digest=sha256:d43e0a9d02c3098f10216d23c77e407ea43d5edc553dfd0438e804acd1219bdf

Observation 82654470-ac0b-4dcf-aea7-a6cf86912c7e · inbound

Contextualized Evaluation of Vision Language Models through Dynamic, Multi-turn Interactions cites this paper.

Contextualized Evaluation of Vision Language Models through Dynamic, Multi-turn Interactions Introducing Visual Perception Token into Multimodal Large Language Model

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T01:57:57.496660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:57:57.496660Z digest=sha256:c52022ca2fb9a5093075d06cece1e371036316e6be229de16f010c972503565b

Observation db43555a-3cc1-4100-86fd-a1e9da023427 · inbound

Mitigating Visual Degradation in MLLMs via Spatial-Spectral Visual Anchor Learning cites this paper.

Mitigating Visual Degradation in MLLMs via Spatial-Spectral Visual Anchor Learning Introducing Visual Perception Token into Multimodal Large Language Model

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T15:10:13.231969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:10:13.231969Z digest=sha256:c756c4c690549acd7b61117145944e085cc29424f611ea1cfc08ab6e1e589f69