Pith. sign in

Paper Citation Record · LEDGER

Introducing Visual Perception Token into Multimodal Large Language Model

As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 23 inbound Pith citation observations for arXiv:2502.17425.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.17425 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 23 of 23 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:46:28.298491Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T07:36:57.997924Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d7787977-b193-42c0-a166-9520d8a30dd0 · inbound

VeriThinker: Learning to Verify Makes Reasoning Model Efficient cites this paper.

VeriThinker: Learning to Verify Makes Reasoning Model Efficient Introducing Visual Perception Token into Multimodal Large Language Model

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:15.423664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:15.423664Z digest=sha256:b8e69242ccc91c32dccd03b2718292ac5b8f28b2c10e31f60fed7ff878a085cf

Observation 59e799c9-745e-4f80-8775-79913f385c70 · inbound

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning cites this paper.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Introducing Visual Perception Token into Multimodal Large Language Model

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:54.345474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:54.345474Z digest=sha256:25549d4f9ef220ba6e0877838e11a07eee90813919638bc299151459fc80f361

Observation 15fad10e-f36d-4452-91f7-f6b04cb1e2f2 · inbound

Qwen Look Again: Guiding Vision-Language Reasoning Models to Re-attention Visual Information cites this paper.

Qwen Look Again: Guiding Vision-Language Reasoning Models to Re-attention Visual Information Introducing Visual Perception Token into Multimodal Large Language Model

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T12:47:01.250685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:47:01.250685Z digest=sha256:28b3692aa46c8490ab54065279fdeacf0d0576c774441988b01be316085c7e2a

Observation d7c3e071-3a60-42c5-97e2-4e2182473919 · inbound

MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought Reasoning cites this paper.

MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought Reasoning Introducing Visual Perception Token into Multimodal Large Language Model

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:54.026686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:54.026686Z digest=sha256:7d854fa2f11fbec3e63f0ce3509016a292529a525b9d197e129134c33a63fe17

Observation 4c9b82e6-839c-49f1-9544-ad7b2b8ada81 · inbound

Enhancing Spatial Reasoning through Visual and Textual Thinking cites this paper.

Enhancing Spatial Reasoning through Visual and Textual Thinking Introducing Visual Perception Token into Multimodal Large Language Model

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T17:46:28.298491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:46:28.298491Z digest=sha256:ca32416551a5793acc0b463366389c2ff94605a7fca1d3c2082a255f31f5182a

Observation 5a86ab97-3722-41b1-92e4-34065fbedba8 · inbound

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding cites this paper.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Introducing Visual Perception Token into Multimodal Large Language Model

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:38.749150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:38.749150Z digest=sha256:3b604065f112c6361d66b4de39241d4d2fb20953cc4720ded06192a6c0982918

Observation 38802939-0095-4799-931c-53674b9679dc · inbound

MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs cites this paper.

MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs Introducing Visual Perception Token into Multimodal Large Language Model

Reference 95

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:16:34.442685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T20:12:46.385646Z digest=sha256:e8f93173569f60a6144e8a4918f0ab1888cd185dfa617ccf0cb4b8b8c3d73018

Observation 5d91ee8c-713d-4a67-8c63-db2c2e9241c1 · inbound

MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs cites this paper.

MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs Introducing Visual Perception Token into Multimodal Large Language Model

Reference 95

Resolution
verified exact
arxiv_id, observed 2026-05-22T10:31:25.388070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T10:30:06.829915Z digest=sha256:1586a3a4d7015eef676aac64fe4910d04c5755e602f048af24f9940da0d42cd2

Observation f19d2e7b-d5e3-4403-a4dc-f0baa4491e32 · inbound

MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs cites this paper.

MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs Introducing Visual Perception Token into Multimodal Large Language Model

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-02T22:00:12.709967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:00:12.709967Z digest=sha256:22d582b4f149f4e833ad8b338d90f2c59349fe277cf68a0bb8297e1f76c875cb

Observation c78b2ec3-ebf7-47d3-9400-cb82fd17fa16 · inbound

Token Warping Helps MLLMs Look from Nearby Viewpoints cites this paper.

Token Warping Helps MLLMs Look from Nearby Viewpoints Introducing Visual Perception Token into Multimodal Large Language Model

Reference 115

Resolution
verified exact
arxiv_id, observed 2026-05-13T21:08:17.430134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-13T21:07:55.062113Z digest=sha256:05a8eedda4aa7774e6e1187c42f87321a81f4e538e4730a2f177e0fb14995882

Observation 7795feb2-0425-4bc2-b194-dc9cc6212f70 · inbound

Visual Enhanced Depth Scaling for Multimodal Latent Reasoning cites this paper.

Visual Enhanced Depth Scaling for Multimodal Latent Reasoning Introducing Visual Perception Token into Multimodal Large Language Model

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:11:04.430716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T16:46:36.010169Z digest=sha256:9624c30744db0846d7879577c82430f8cedf82816ee9ee7405197ff059b6b97e

Observation b25ca432-1a0f-4544-822d-9ab5f7a9720e · inbound

Visual Enhanced Depth Scaling for Multimodal Latent Reasoning cites this paper.

Visual Enhanced Depth Scaling for Multimodal Latent Reasoning Introducing Visual Perception Token into Multimodal Large Language Model

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:06:24.110121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-12T04:38:06.774877Z digest=sha256:ba4861f043751497e9272c4aea45e1ffac414962f8a594cc98618d461fd94141

Observation 929cf2b2-0a0f-4923-83d1-dae6d1b70c35 · inbound

Visual Enhanced Depth Scaling for Multimodal Latent Reasoning cites this paper.

Visual Enhanced Depth Scaling for Multimodal Latent Reasoning Introducing Visual Perception Token into Multimodal Large Language Model

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:27:29.237405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-13T07:26:59.917150Z digest=sha256:d3c43fbb089f746e8b0745791b3bc8369c5b0f39d49f2848234703cb879eb015

Observation ad7551bf-0755-407d-bbea-79925e110763 · inbound

Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs cites this paper.

Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs Introducing Visual Perception Token into Multimodal Large Language Model

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:26:01.904902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T15:29:25.650175Z digest=sha256:d454f20b22d2849bfab8f01fff9013e7f12c67ca1d3626173b05ef5304dde14f

Observation 59d66061-9aeb-4188-bce1-40be84d536aa · inbound

From Failure to Feedback: Group Revision Unlocks Hard Cases in Object-Level Grounding cites this paper.

From Failure to Feedback: Group Revision Unlocks Hard Cases in Object-Level Grounding Introducing Visual Perception Token into Multimodal Large Language Model

Reference 92

Resolution
verified exact
arxiv_id, observed 2026-05-20T18:43:38.894957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T18:39:11.904941Z digest=sha256:cb4c72d12609b3e90e0e9983044e60a72b163078eaebefee89a261c74e876345

Observation ec8d8ba5-57cd-40f9-85d4-db1af218863e · inbound

ROVER: Routing Object-Centric Visual Evidence for Grounded Multi-Image Reasoning cites this paper.

ROVER: Routing Object-Centric Visual Evidence for Grounded Multi-Image Reasoning Introducing Visual Perception Token into Multimodal Large Language Model

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:43:28.670041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-29T13:41:44.049230Z digest=sha256:cc22215203302ebe99ae48890202d2a788499674029c71b15568ab1a6b91fa91

Observation 14c371c7-2ec0-4565-88da-64bd82c79528 · inbound

MathVis-Fine: Aligning Visual Supervision with Necessity via Progressive Dependency-Guided Training for Multimodal Mathematical Reasoning cites this paper.

MathVis-Fine: Aligning Visual Supervision with Necessity via Progressive Dependency-Guided Training for Multimodal Mathematical Reasoning Introducing Visual Perception Token into Multimodal Large Language Model

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T20:18:57.676800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-27T01:23:40.564561Z digest=sha256:ce4ff98b44a4cea171674308b9d2e500777ac0565a3684fefeb93eadda789690

Observation 65361e8c-8d12-408b-8624-dcf4fb2cfda2 · inbound

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models cites this paper.

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models Introducing Visual Perception Token into Multimodal Large Language Model

Reference 139

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:09:55.102881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-26T01:50:54.242508Z digest=sha256:2647d907d2e0ceba40a2ed88ddf12e93c274172e456af88e9ecfdcb647618cea

Observation e8758001-cf6e-4400-8108-1c661c2ce96b · inbound

Latent Noise Mask for Reducing Visual Redundancy in Multimodal Large Language Models cites this paper.

Latent Noise Mask for Reducing Visual Redundancy in Multimodal Large Language Models Introducing Visual Perception Token into Multimodal Large Language Model

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-06-30T06:44:19.799258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T06:35:16.868238Z digest=sha256:44cba86a8c57a0cafbc22cdca86e5444a7ad4460aa63b7829cc8f22dbe6d38d4

Observation 63e626de-969a-4252-b038-bb09a77b5df3 · inbound

DeltaV: Thinking with Visual State Updates in Unified Large Multimodal Models cites this paper.

DeltaV: Thinking with Visual State Updates in Unified Large Multimodal Models Introducing Visual Perception Token into Multimodal Large Language Model

Reference 19

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T07:36:57.999269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-07-10T07:31:26.225257Z digest=sha256:eefd59c12a4633f5acab29837617ac8a8f35587260280d49769f7f17e7a0e59b

Observation c3b6a2db-9f62-4656-9d02-1af7581f551f · inbound

Visual Access Boundaries in Vision-Language Model Reasoning cites this paper.

Visual Access Boundaries in Vision-Language Model Reasoning Introducing Visual Perception Token into Multimodal Large Language Model

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T06:24:29.245780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:24:29.245780Z digest=sha256:d43e0a9d02c3098f10216d23c77e407ea43d5edc553dfd0438e804acd1219bdf

Observation 82654470-ac0b-4dcf-aea7-a6cf86912c7e · inbound

Contextualized Evaluation of Vision Language Models through Dynamic, Multi-turn Interactions cites this paper.

Contextualized Evaluation of Vision Language Models through Dynamic, Multi-turn Interactions Introducing Visual Perception Token into Multimodal Large Language Model

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T01:57:57.496660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:57:57.496660Z digest=sha256:c52022ca2fb9a5093075d06cece1e371036316e6be229de16f010c972503565b

Observation db43555a-3cc1-4100-86fd-a1e9da023427 · inbound

Mitigating Visual Degradation in MLLMs via Spatial-Spectral Visual Anchor Learning cites this paper.

Mitigating Visual Degradation in MLLMs via Spatial-Spectral Visual Anchor Learning Introducing Visual Perception Token into Multimodal Large Language Model

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T15:10:13.231969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:10:13.231969Z digest=sha256:c756c4c690549acd7b61117145944e085cc29424f611ea1cfc08ab6e1e589f69