Pith. sign in

Paper Citation Record · LEDGER

Thinking with Generated Images

As of 3 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 27 inbound Pith citation observations for arXiv:2505.22525.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.22525 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-03T06:30:56.289259+00:00

measured 27 of 27 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T17:06:54.616428Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T19:20:06.640927Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 5355db36-3228-47ca-a384-b6e9f417446f · inbound

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models cites this paper.

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models Thinking with Generated Images

Reference 130

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:40:42.013852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T08:40:40.910461Z digest=sha256:f726fbfaac8578376ea57f51ec8ec6b7a4d6d0b5d9a03b460f87df06c1f1d838

Observation 4dd8abe7-10b0-41ed-88cd-7c1ed1d406fc · inbound

IRG-MotionLLM: Interleaving Motion Generation, Assessment and Refinement for Text-to-Motion Generation cites this paper.

IRG-MotionLLM: Interleaving Motion Generation, Assessment and Refinement for Text-to-Motion Generation Thinking with Generated Images

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T17:06:54.616428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:06:54.616428Z digest=sha256:3d61bcc80c01380abe7a2c0a7506eb1ccc99eed7d1b940e15e20994459b35357

Observation a264c351-f033-42d4-bb37-a1cd73388f81 · inbound

Mull-Tokens: Modality-Agnostic Latent Thinking cites this paper.

Mull-Tokens: Modality-Agnostic Latent Thinking Thinking with Generated Images

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:58:38.739619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-16T22:57:04.802871Z digest=sha256:ca5087ecf9e070cd57e8562df9ddc0a295227297f1f9a37e873edb0f099fc59b

Observation 99abcb62-8025-46b4-be1f-713c25f01b33 · inbound

UniDrive-WM: Unified Understanding, Planning and Generation World Model for Autonomous Driving cites this paper.

UniDrive-WM: Unified Understanding, Planning and Generation World Model for Autonomous Driving Thinking with Generated Images

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T12:06:26.284752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:06:26.284752Z digest=sha256:02e4cc2f7cc532af4b0f6ee369536103fa916161b1500a8b86f2e11da2ef509f

Observation d4b40090-f755-4108-9e19-4339e0d548f3 · inbound

Omni-R1: Towards the Unified Generative Paradigm for Multimodal Reasoning cites this paper.

Omni-R1: Towards the Unified Generative Paradigm for Multimodal Reasoning Thinking with Generated Images

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-16T14:37:59.927761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-16T14:37:05.402850Z digest=sha256:f5f07ec55e40c365cf63ab410458f12644d1240496c6bdc5e3a3b20b71ab210c

Observation d3cbe0b3-fd42-4ba0-b7d5-525d590253ea · inbound

MentisOculi: Revealing the Limits of Reasoning with Mental Imagery cites this paper.

MentisOculi: Revealing the Limits of Reasoning with Mental Imagery Thinking with Generated Images

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T05:23:30.774736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:23:30.774736Z digest=sha256:bd2dec922e2afbf2c930fc239b18b7cf35085d3df73fd8f6333759f72c963b04

Observation 8aea9ef2-3fea-40fc-8cf6-4dde4ab511b6 · inbound

Thinking with Drafting: Optical Decompression via Logical Reconstruction cites this paper.

Thinking with Drafting: Optical Decompression via Logical Reconstruction Thinking with Generated Images

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-16T03:30:32.939197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-16T03:28:32.264246Z digest=sha256:f921b81098bc17ccac328293efb78492ce7818868525664d38f2b749b91c4f32

Observation 65f490ae-9825-4a71-8343-4a97fdfa1f22 · inbound

MMRad-22K: A Structured Multimodal Evidence Dataset for Chest X-ray Report Generation cites this paper.

MMRad-22K: A Structured Multimodal Evidence Dataset for Chest X-ray Report Generation Thinking with Generated Images

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T23:43:29.982407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:43:29.982407Z digest=sha256:fa67d82381ccfb9d1d49cbb11c7761cb33f3b9e709620209fc5deb0755129359

Observation 2a9425a5-e1ce-406e-95de-221bb169e5ae · inbound

Demystifying Video Reasoning cites this paper.

Demystifying Video Reasoning Thinking with Generated Images

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-13T23:27:11.006580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:27:11.006580Z digest=sha256:fc2863acbfc0703bda0d6d8854b25a9a5059a9f08539ed6095b1e6d0c6fd3a2d

Observation b2339c8b-b782-4231-8d29-031878628030 · inbound

Demystifying Video Reasoning cites this paper.

Demystifying Video Reasoning Thinking with Generated Images

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T02:33:53.570433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:33:53.570433Z digest=sha256:d85080069f7f44381c95e255f9c2d65f1fc7701e1951112efce962beaa1a9066

Observation 4285b2c7-93d6-458b-9d60-4cacfd868741 · inbound

Visual Enhanced Depth Scaling for Multimodal Latent Reasoning cites this paper.

Visual Enhanced Depth Scaling for Multimodal Latent Reasoning Thinking with Generated Images

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:11:04.491939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T16:46:36.010169Z digest=sha256:6d384651fea5c41377e6ee482c8bdc916ee2a293f8fe1fd9aaf8fbe31aadad1d

Observation a33719c0-ca8d-4e16-a933-39054753fef4 · inbound

Visual Enhanced Depth Scaling for Multimodal Latent Reasoning cites this paper.

Visual Enhanced Depth Scaling for Multimodal Latent Reasoning Thinking with Generated Images

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:06:24.076055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T04:38:06.774877Z digest=sha256:409ebcc4d85adc240c157fe70c864520a1eb61446af4e8336524870ef1ae09e3

Observation f44a1e8d-573d-41e0-a096-7a13d3e0ddee · inbound

Visual Enhanced Depth Scaling for Multimodal Latent Reasoning cites this paper.

Visual Enhanced Depth Scaling for Multimodal Latent Reasoning Thinking with Generated Images

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:27:29.194424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T07:26:59.917150Z digest=sha256:58b9daa6f829aaad48758935af8e63bba35ae7e620e07ac53b02f6e132e9da6d

Observation b031ab72-d0c4-4ddd-93a2-57750730bffb · inbound

Steering Visual Generation in Unified Multimodal Models with Understanding Supervision cites this paper.

Steering Visual Generation in Unified Multimodal Models with Understanding Supervision Thinking with Generated Images

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:41:10.130970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T14:48:22.805268Z digest=sha256:e78ad19d96f8bf94c0f8f93cf79e299db8eb05784d56ef36a20c54723bc9bcc5

Observation 30e42e67-9dd5-4aab-a589-02f1a8f94387 · inbound

Self-Consistent Latent Reasoning: Long Latent Sequence Reasoning for Vision-Language Model cites this paper.

Self-Consistent Latent Reasoning: Long Latent Sequence Reasoning for Vision-Language Model Thinking with Generated Images

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:17:29.232593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T07:14:48.918959Z digest=sha256:637bccf4fc5d7b0bc921b78242e047654541bb8b23d44681250317ab86c40fc4

Observation 98405ad6-7447-4c9e-985f-8f224292daeb · inbound

Self-Consistent Latent Reasoning: Long Latent Sequence Reasoning for Vision-Language Model cites this paper.

Self-Consistent Latent Reasoning: Long Latent Sequence Reasoning for Vision-Language Model Thinking with Generated Images

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:48:00.657129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-14T21:47:50.595481Z digest=sha256:3a060f5f37ec515becc65fa6ec1bdd695bcb538e0434672e0152468aec1ca1d8

Observation 6d518a9c-3248-42b5-b003-177e1c6a0e78 · inbound

DeepLatent: Think with Images via Parallel Latent Visual Reasoning cites this paper.

DeepLatent: Think with Images via Parallel Latent Visual Reasoning Thinking with Generated Images

Reference 65

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T19:52:36.120842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-06-28T18:44:39.545911Z digest=sha256:1cc2e3ac2ebbbe186efebc92fb89d5274d31c802198e7cd1c601028decd60ae5

Observation 04d0cf82-0ed9-45c2-b3e6-816187139f14 · inbound

Robust-U1: Can MLLMs Self-Recover Corrupted Visual Content for Robust Understanding? cites this paper.

Robust-U1: Can MLLMs Self-Recover Corrupted Visual Content for Robust Understanding? Thinking with Generated Images

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T20:47:23.363872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-27T20:06:11.170282Z digest=sha256:ac76c0dfbf369d80cbe1d9e3cdc807da2d322e598a8641fc7bacaf6262d932c3

Observation 0c75e93e-1f34-4947-adf5-bd51bc909cb1 · inbound

Bridging Modal Isolation in Interleaved Thinking: Supervising Modality Transitions via Stepwise Reinforcement cites this paper.

Bridging Modal Isolation in Interleaved Thinking: Supervising Modality Transitions via Stepwise Reinforcement Thinking with Generated Images

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-03T13:38:19.358184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-27T07:43:09.538833Z digest=sha256:f9362827bdfbca9c3eb3e1de1f000567c66578f55fc4d0cd99a30a6fe16a40ec

Observation 3be24b96-e035-4eb6-bd48-00e41960e062 · inbound

Bridging Modal Isolation in Interleaved Thinking: Supervising Modality Transitions via Stepwise Reinforcement cites this paper.

Bridging Modal Isolation in Interleaved Thinking: Supervising Modality Transitions via Stepwise Reinforcement Thinking with Generated Images

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-12T14:14:01.413340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T14:14:01.413340Z digest=sha256:f8622e550275b53fe5ce21b6ddfeea48a1903e8b78a9112dc2f2afcfce963c0e

Observation 04c29cee-4cc0-4fdf-8713-d82534076e2c · inbound

Universal Image Restoration via Internalized Chain-of-Thought Reasoning cites this paper.

Universal Image Restoration via Internalized Chain-of-Thought Reasoning Thinking with Generated Images

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:08:56.575921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-27T01:34:34.269899Z digest=sha256:ad9c5f3e656c3179487f5d5d1d7af2f4dcb5b7d682f5feb6e86e88931e308909

Observation de32478d-e81c-4b36-ab19-46b755620299 · inbound

V-Zero: Answer-Label-Free On-Policy Distillation with Contrastive Evidence Gating for Fine-Grained Visual Reasoning cites this paper.

V-Zero: Answer-Label-Free On-Policy Distillation with Contrastive Evidence Gating for Fine-Grained Visual Reasoning Thinking with Generated Images

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-07-04T19:20:06.642493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-06-25T21:23:10.051805Z digest=sha256:054c193efa5ddea4b7a205539e4bc0bddeb34956483b88d0a11f28ee1cd2f810

Observation b72142c9-a4ca-4733-accd-8e3ee32e6da2 · inbound

Einstein World Models cites this paper.

Einstein World Models Thinking with Generated Images

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T13:59:51.812525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-06-26T04:47:59.575489Z digest=sha256:34c7c69bca048ac9f99918934923cbdbd74771fff56a0c1a31caa72822781295

Observation 874dc02b-fec6-4ff3-98fe-8e055c61edb0 · inbound

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception cites this paper.

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception Thinking with Generated Images

Reference 286

Resolution
unresolved
no resolver link, observed 2026-07-12T04:17:40.198357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T04:17:40.198357Z digest=sha256:a3b681451f515bb9ffd3e52740e8b03ae664247b20476c591f3d2746521dc4da

Observation f7c47963-7d2c-47f7-ad85-3b8d8bc5a317 · inbound

Bridging Interleaved Multi-Modal Reasoning as a Unified Decision Process cites this paper.

Bridging Interleaved Multi-Modal Reasoning as a Unified Decision Process Thinking with Generated Images

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-12T00:14:19.104498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:14:19.104498Z digest=sha256:d4892fb2a82c4e1ae4e0585774341cb63240c2752c8f54f39a54232af055c897

Observation 705ca317-21e5-4db6-ad7f-2798b656a2ad · inbound

OPLD: On-Policy Latent Distillation for Multimodal Reasoning cites this paper.

OPLD: On-Policy Latent Distillation for Multimodal Reasoning Thinking with Generated Images

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-31T16:13:17.436811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T16:13:17.436811Z digest=sha256:dcbdbad196b3e30d62066cd5503d4eed592f596e7b3914185985d623fbab1cff

Observation 83c505a3-8516-404f-bb1d-243a002acfdb · inbound

FaithEyes: Towards Faithful Tool Use via Multi-Agent Process-Image Verification cites this paper.

FaithEyes: Towards Faithful Tool Use via Multi-Agent Process-Image Verification Thinking with Generated Images

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-07-31T14:08:11.168680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T14:08:11.168680Z digest=sha256:c2d336cfab806699ed1ab70ace5d67a71d9530d1bc18f79010ee88db2019970d