Pith. sign in

Paper Citation Record · LEDGER

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models

As of 6 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 23 inbound Pith citation observations for arXiv:2505.04921.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.04921 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 23 of 23 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T16:31:57.845882Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T02:44:28.049569Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation bffba20d-c726-44ca-9288-01e43290b833 · inbound

Why Relational Graphs Will Save the Next Generation of Vision Foundation Models? cites this paper.

Why Relational Graphs Will Save the Next Generation of Vision Foundation Models? Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T16:31:57.845882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:31:57.845882Z digest=sha256:b7c7c9ef1013e09719c3fc7c8c453673b59ee9927ba488c318033399404e53ec

Observation e45b5ec3-a23d-49ae-98a7-e04075498a79 · inbound

Toward Edge General Intelligence with Agentic AI and Agentification: Concepts, Technologies, and Future Directions cites this paper.

Toward Edge General Intelligence with Agentic AI and Agentification: Concepts, Technologies, and Future Directions Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models

Reference 133

Resolution
unresolved
no resolver link, observed 2026-08-05T16:20:59.685324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:20:59.685324Z digest=sha256:274ce33335c34f701286e240f092c86f087f481af22bc0c95575b28a332d6b0c

Observation 194346aa-9f27-42e6-b08f-ae432e5de4b5 · inbound

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models cites this paper.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:01.493981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:01.493981Z digest=sha256:bba1f9fd48200f53ba2781400f0dde8390faf127eafca4b43cb5a0567757c565

Observation d65d6905-acca-4367-9f24-339f66b6618b · inbound

AI Reasoning for Wireless Communications and Networking: A Survey and Perspectives cites this paper.

AI Reasoning for Wireless Communications and Networking: A Survey and Perspectives Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T19:34:25.492417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:34:25.492417Z digest=sha256:12cb9b5329f991bb3e714d9c0b6d74f5fc13fe238aab91a2405b13fab5bb62bd

Observation 616ab416-5f6c-42a8-bd47-e9155f0ad6a5 · inbound

Mitigating Visual Context Degradation in Large Multimodal Models: A Training-Free Decoupled Agentic Framework cites this paper.

Mitigating Visual Context Degradation in Large Multimodal Models: A Training-Free Decoupled Agentic Framework Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-18T12:32:36.361359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T12:31:25.257879Z digest=sha256:934fb84c583847c1383b41808eaae0c82795a5d09a589054d49607181ecbae50

Observation 36a7b0d3-5897-4037-a822-8ca9115891bd · inbound

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents cites this paper.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:11:30.087068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T03:09:31.161760Z digest=sha256:a0699210e172ab1979e39c469e786defedfecb6868a0f5c1430c841a254d18ea

Observation ec5d3588-2cf9-4939-b169-9d4cc35722ef · inbound

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents cites this paper.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:45.298957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:45.298957Z digest=sha256:18f39905b6d3c51ec3a55d7c017af89de9535972c6aa0baa8671ccc3eecd2a93

Observation 6cd204ca-3d25-4713-9dd6-08f368f410ad · inbound

LookWise: Knowing When and Where to Look for Fine-Grained Visual Reasoning in Multimodal Large Language Models cites this paper.

LookWise: Knowing When and Where to Look for Fine-Grained Visual Reasoning in Multimodal Large Language Models Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T20:30:56.951607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:30:56.951607Z digest=sha256:57f2859dedd6a6869955252232d9ea02faed02402ac692f56ccea02bfe5b47ac

Observation 6d1c70cb-d8a9-40b9-b5b8-ec7df293a230 · inbound

OmniJigsaw: Enhancing Omni-Modal Reasoning via Modality-Orchestrated Reordering cites this paper.

OmniJigsaw: Enhancing Omni-Modal Reasoning via Modality-Orchestrated Reordering Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T06:11:01.552354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T17:45:51.528645Z digest=sha256:93249282ca722d4f75178618875cb462975e3deb63998ee7b1c9d94f6c1bb1b3

Observation 438d5b62-a28e-4181-bc8a-a59a304a203a · inbound

Decoupled Similarity for Task-Aware Token Pruning in Large Vision-Language Models cites this paper.

Decoupled Similarity for Task-Aware Token Pruning in Large Vision-Language Models Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:56:05.744774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:15:04.261855Z digest=sha256:4ee49e69a11747c3c08192cee02572c55e9b2079ffedf8787edf36c5d00f3a5a

Observation 65ac6982-1253-44bc-8b74-f1915ca8ca32 · inbound

Beyond Binary Contrast: Modeling Continuous Skeleton Action Spaces with Transitional Anchors cites this paper.

Beyond Binary Contrast: Modeling Continuous Skeleton Action Spaces with Transitional Anchors Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:51:03.578653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T04:32:04.547197Z digest=sha256:d23a55ec30a6cdcf62c3bdd0a4b98d71b87cb2b7b6e37f07f70c2cd9ea18c4b9

Observation c0c8f5af-aa88-4182-a8e6-315ab58e9420 · inbound

ViSRA: A Video-based Spatial Reasoning Agent for Multi-modal Large Language Models cites this paper.

ViSRA: A Video-based Spatial Reasoning Agent for Multi-modal Large Language Models Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:46:36.421527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T04:00:23.681682Z digest=sha256:f1417df7d212a93c672f0e0d581d6682a9dd245b65909e816d2cb79486f10732

Observation 5fc18d1e-fa3c-4bc0-9b50-8b7785c4f707 · inbound

NASiC: 3D NAND-based CAM-Selected Multibit CIM Architecture for Efficient On-Device Mixture-of-Experts LLM Inference cites this paper.

NASiC: 3D NAND-based CAM-Selected Multibit CIM Architecture for Efficient On-Device Mixture-of-Experts LLM Inference Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-25T03:00:16.019653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-25T02:57:04.813106Z digest=sha256:89c4abcba2bc358b857ad47d2b8ce469cded9fc1c704f8a3e6d69f56acd961e5

Observation f2fe28ed-01c0-4b75-87bb-6b6533eebe8f · inbound

PEDESTRIANQA: A Benchmark for Vision-Language Models on Pedestrian Intention and Trajectory Prediction cites this paper.

PEDESTRIANQA: A Benchmark for Vision-Language Models on Pedestrian Intention and Trajectory Prediction Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:34:40.936877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T13:25:04.053283Z digest=sha256:83e2c58c8c6621a598d64a9813758da8a30b55e1c5aee20cc4584ff147582a9b

Observation 86f6ae4c-1b6e-4e9b-82de-283950f3170c · inbound

AnE: Pushing the Reasoning Frontier of Multimodal LLMs via Anchor Evolution cites this paper.

AnE: Pushing the Reasoning Frontier of Multimodal LLMs via Anchor Evolution Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-30T00:14:04.662917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T22:56:39.504430Z digest=sha256:d6ad41c1f4fe8482bbf605af36b46f94dcf709d4f81346611ef6ee97a1321cdb

Observation 50254a08-a76b-4ad7-94c6-df15f8d46f46 · inbound

Mags-RL: Wearing Multimodal LLMs a Magnifying Glass via Agentic Reinforcement Learning For Complex Scene Reasoning cites this paper.

Mags-RL: Wearing Multimodal LLMs a Magnifying Glass via Agentic Reinforcement Learning For Complex Scene Reasoning Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T13:43:29.018395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T13:38:01.819121Z digest=sha256:420f0eb66f142d2a56a0c6f46ec3f95846cfc5b62153fbb4cc2cf364d297e12b

Observation 3db0dc68-3a88-4d5a-b518-32cb56fcfb3b · inbound

Watch, Remember, Reason: Human-View Video Understanding with MLLMs cites this paper.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:37:14.037520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:f3a97ab063e9683dac9bae0a6bc196aaceab9cf6dbc3c9d8c69567d2905a3b09

Observation f137f6d5-c43a-4677-9922-2d3efa2a795b · inbound

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs cites this paper.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models

Reference 62

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T23:06:21.347654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:97df803c31777ae26079827f5c2b44455037cb58d4fc1b15e087c0a9105f89a5

Observation 815cbf74-c26e-40c7-9108-d33440545e1c · inbound

Empowering GUI Agents via Autonomous Experience Exploration and Hindsight Experience Utilization for Task Planning cites this paper.

Empowering GUI Agents via Autonomous Experience Exploration and Hindsight Experience Utilization for Task Planning Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T14:29:53.387832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-26T03:51:51.827622Z digest=sha256:9b3c715f55b3a87b3f434c00cbb2de6af9fd379e9b327f542f7da990eb8fec65

Observation 9729d942-8013-4e75-8fe9-66e95fc67465 · inbound

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs cites this paper.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T02:44:28.050948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:da11ae7c5993713ed96381a6e1e1c2c12db472e95981f60566b7f4da0d0b6430

Observation 39982139-c36d-4678-a8ab-a2e0f5789ddd · inbound

Mixture of Cognitive Experts in Large Vision-Language Models cites this paper.

Mixture of Cognitive Experts in Large Vision-Language Models Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-14T09:13:07.507164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T09:13:07.507164Z digest=sha256:c7f1c1cbc413d0d83750d2e590a3890df25f581422c3d44973b3fb753d095f80

Observation d79251d8-84ac-45b3-bd19-4b1fb35c38a3 · inbound

Knowing the Self, Understanding the World: A Dual-Cognition Benchmark for UAV Spatio-temporal Reasoning with MLLMs cites this paper.

Knowing the Self, Understanding the World: A Dual-Cognition Benchmark for UAV Spatio-temporal Reasoning with MLLMs Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T21:08:19.564996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:08:19.564996Z digest=sha256:b296f5973f13bba69f7770881a1c57a89d38843917f9e67c9f99dda535457418

Observation 6f897abd-8868-4151-b05f-2845cbbfeee4 · inbound

DC-Leap: Training-Free Acceleration of dLLMs via Draft-Guided Contiguous Leaping Decoding cites this paper.

DC-Leap: Training-Free Acceleration of dLLMs via Draft-Guided Contiguous Leaping Decoding Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-02T13:42:25.191504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:42:25.191504Z digest=sha256:72f90c367becfdd3d641964c3b43783128e2552dbf1af0a8e12bb7fa1b3e207d