Pith. sign in

Paper Citation Record · LEDGER

4DThinker: Thinking with 4D Imagery for Dynamic Spatial Understanding

As of 10 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 0 inbound Pith citation observations for arXiv:2605.05997.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.05997 v2

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-25T06:15:33.062980Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

28 of 28 outbound references displayed

  • verified exact24
  • verified fuzzy4
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 01600f00-44bf-4106-8d3a-81f21576e130 · outbound

This paper cites Qwen2.5-VL Technical Report.

4DThinker: Thinking with 4D Imagery for Dynamic Spatial Understanding Qwen2.5-VL Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-25T06:16:39.927531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T06:15:33.062980Z digest=sha256:4e11523d60c20854c40b49c2ddde49ee5ca68b6a21693303dc23d80f4309ca4f

Observation fa29e92c-af28-48a0-a4ed-1caaaf27aa2f · outbound

This paper cites SpatialBot: Precise Spatial Understanding with Vision Language Models.

4DThinker: Thinking with 4D Imagery for Dynamic Spatial Understanding SpatialBot: Precise Spatial Understanding with Vision Language Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-25T06:16:39.992535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T06:15:33.062980Z digest=sha256:4896b230c21154670454e367e66d1af6363ab57bae9631c7f38bab08c3f5950f

Observation 6d090e5e-d90c-48cc-9833-b202f07ebd64 · outbound

This paper cites SAM 3: Segment Anything with Concepts.

4DThinker: Thinking with 4D Imagery for Dynamic Spatial Understanding SAM 3: Segment Anything with Concepts

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-25T06:16:39.977799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T06:15:33.062980Z digest=sha256:ea5f3c9462afb5bc66093d9c3317879e19aadfa4c38788e1d4d75a75dd52f935

Observation f6ef8b51-9c5f-4b38-8360-fbe703d6b05f · outbound

This paper cites Adagar: Adaptive gabor representation for dynamic scene reconstruction.arXiv preprint arXiv:2601.00796.

4DThinker: Thinking with 4D Imagery for Dynamic Spatial Understanding Adagar: Adaptive gabor representation for dynamic scene reconstruction.arXiv preprint arXiv:2601.00796

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-25T06:16:40.007580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T06:15:33.062980Z digest=sha256:1b79f8ed205382abc07d4d7849deff0ce9e10af36f3af7e6b9f4ff40b94a4cc7

Observation c756657c-3ea6-4f43-8b77-ba01de29f7fb · outbound

This paper cites Think with 3d: Geometric imagination grounded spatial reasoning from limited views.arXiv preprint arXiv:2510.18632, 2025a.

4DThinker: Thinking with 4D Imagery for Dynamic Spatial Understanding Think with 3d: Geometric imagination grounded spatial reasoning from limited views.arXiv preprint arXiv:2510.18632, 2025a

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-25T06:16:40.021069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T06:15:33.062980Z digest=sha256:931903d834784f968f22f9b893fc9c1d5fced16da84999ce2303d9382c137630

Observation c529e4e0-952b-44b6-81fd-b692fe75d8db · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

4DThinker: Thinking with 4D Imagery for Dynamic Spatial Understanding Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-25T06:16:39.982717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T06:15:33.062980Z digest=sha256:823d95cccdbdee1ea7f5c9470fbe831683dab8371b3a7187f8a12f15c95821b7

Observation 6a5834d6-bde9-4e5b-ac64-08429524b9a8 · outbound

This paper cites From Explicit CoT to Implicit CoT: Learning to Internalize CoT Step by Step.

4DThinker: Thinking with 4D Imagery for Dynamic Spatial Understanding From Explicit CoT to Implicit CoT: Learning to Internalize CoT Step by Step

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-25T06:16:39.960788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T06:15:33.062980Z digest=sha256:2606416701318f55a3d24c257cb0806776fb1d2c221fd2ede682a3a3031a20b3

Observation 7b1db54e-df36-495e-8276-412ce2a58024 · outbound

This paper cites VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction.

4DThinker: Thinking with 4D Imagery for Dynamic Spatial Understanding VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-25T06:16:39.987125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T06:15:33.062980Z digest=sha256:4d077c76ec5acdf53a526e970683ac6089562ee6393213c962884ec3fd3575fa

Observation 14890bca-94bc-42b8-a51e-7e46bef1126e · outbound

This paper cites Think before you speak: Training Language Models With Pause Tokens.

4DThinker: Thinking with 4D Imagery for Dynamic Spatial Understanding Think before you speak: Training Language Models With Pause Tokens

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-25T06:16:39.972605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T06:15:33.062980Z digest=sha256:ed7bdac41b338070be30aba6c1cb613dfcf082b9cfdc9efd109af43969f7f81d

Observation 916f0389-7738-4d75-ab05-74bed9fe690a · outbound

This paper cites Training Large Language Models to Reason in a Continuous Latent Space.

4DThinker: Thinking with 4D Imagery for Dynamic Spatial Understanding Training Large Language Models to Reason in a Continuous Latent Space

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-25T06:16:40.011759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T06:15:33.062980Z digest=sha256:057ef3125715dafef52478e38f81a9dedeafb596ad932ef7e6b55bb034705ce0

Observation 72cf7429-17e7-4569-be8f-5dd04b1b5c38 · outbound

This paper cites Thinking in dynamics: How multimodal large language models perceive, track, and reason dynamics in physical 4d world.arXiv preprint arXiv:2603.12746.

4DThinker: Thinking with 4D Imagery for Dynamic Spatial Understanding Thinking in dynamics: How multimodal large language models perceive, track, and reason dynamics in physical 4d world.arXiv preprint arXiv:2603.12746

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-25T06:16:39.926537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T06:15:33.062980Z digest=sha256:c391f6ac2e449c1ea1d993f28bccf7fd450338c1827459ce3f84af646a3bd7db

Observation 831f9b8b-eb6e-4250-a1a8-54bb0d40f8bd · outbound

This paper cites Latent Visual Reasoning.

4DThinker: Thinking with 4D Imagery for Dynamic Spatial Understanding Latent Visual Reasoning

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-25T06:16:39.937928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T06:15:33.062980Z digest=sha256:9fe9b1c8c84d577c6967ac80b5762678fdb06010633a58bd5cec5cba669da490

Observation 5ce32704-f9b8-4113-829b-a30e07d5e669 · outbound

This paper cites Llava-st: A multimodal large language model for fine-grained spatial-temporal understanding.

4DThinker: Thinking with 4D Imagery for Dynamic Spatial Understanding Llava-st: A multimodal large language model for fine-grained spatial-temporal understanding

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-25T06:16:39.997846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T06:15:33.062980Z digest=sha256:1a07771a5367c8d37b9048077743fa6f3ee5d769167fee18a3c816aa82e43f63

Observation 735994be-fcfe-4f8e-b488-ed43f427ee0b · outbound

This paper cites SpatialCoT: Advancing Spatial Reasoning through Coordinate Alignment and Chain-of-Thought for Embodied Task Planning.

4DThinker: Thinking with 4D Imagery for Dynamic Spatial Understanding SpatialCoT: Advancing Spatial Reasoning through Coordinate Alignment and Chain-of-Thought for Embodied Task Planning

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-25T06:16:39.949775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T06:15:33.062980Z digest=sha256:d9b365d345795126f865ee5a9cb3893a62019255f6b597ee0e03c4436775a4ac

Observation d16c1eea-dfdc-4961-82e6-089d8489944c · outbound

This paper cites SpaceR: Reinforcing MLLMs in Video Spatial Reasoning.

4DThinker: Thinking with 4D Imagery for Dynamic Spatial Understanding SpaceR: Reinforcing MLLMs in Video Spatial Reasoning

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-25T06:16:39.932213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T06:15:33.062980Z digest=sha256:adb69bd89065f47a34875d80be1b6d5b8c3608c789c571585675806f8f5ce09d

Observation db52918a-cef0-4ba8-abff-7003f23b0660 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

4DThinker: Thinking with 4D Imagery for Dynamic Spatial Understanding DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-25T06:16:39.916356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T06:15:33.062980Z digest=sha256:6dd97d494bdf3c251d8b3a428a4465da33dc07092a1bccf6886d756b8943669a

Observation 3a11d632-0aa1-4573-a9ac-afdd4507ac5b · outbound

This paper cites OpenAI GPT-5 System Card.

4DThinker: Thinking with 4D Imagery for Dynamic Spatial Understanding OpenAI GPT-5 System Card

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-25T06:16:39.904303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T06:15:33.062980Z digest=sha256:3a882cdc834be4448f96a24a34fd763ba2620a010ba814e52e7e0d1a42140a1d

Observation 8b4b4f60-76da-4052-ba68-522cf93c2ba1 · outbound

This paper cites Spatialvid: A large-scale video dataset with spatial annotations.arXiv preprint arXiv:2509.09676, 2025a.

4DThinker: Thinking with 4D Imagery for Dynamic Spatial Understanding Spatialvid: A large-scale video dataset with spatial annotations.arXiv preprint arXiv:2509.09676, 2025a

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-25T06:16:39.966490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T06:15:33.062980Z digest=sha256:f37cd0808cbd193ccf26ec79e96165368f2ec87fac7361204bd0babf29377225

Observation 3fbb694d-5e16-4783-8678-d7fec5ab29e9 · outbound

This paper cites Learning how to use tools, not just when: Pattern-aware tool-integrated reasoning.MATH-AI @ NeurIPS 2025.

4DThinker: Thinking with 4D Imagery for Dynamic Spatial Understanding Learning how to use tools, not just when: Pattern-aware tool-integrated reasoning.MATH-AI @ NeurIPS 2025

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:16:40.103623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T06:15:33.062980Z digest=sha256:8db31ffee5b15a532400bf5099a1966c7f9a563394b6f179da766617d6f1c234

Observation eac84f28-44e0-44c6-aa02-2fd5f6c66137 · outbound

This paper cites Thinking in space: How multimodal large language models see, remember, and recall spaces.

4DThinker: Thinking with 4D Imagery for Dynamic Spatial Understanding Thinking in space: How multimodal large language models see, remember, and recall spaces

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-25T06:16:39.955239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T06:15:33.062980Z digest=sha256:7834186faf0a3b0aa87c3dcf46ce9f42f07d73e2c9c18067b1e1d283e61694e1

Observation b669276a-569e-4ce3-96fb-3873649ca946 · outbound

This paper cites Mllm-4d: Towards visual-based spatial-temporal intelligence.arXiv preprint arXiv:2603.00515.

4DThinker: Thinking with 4D Imagery for Dynamic Spatial Understanding Mllm-4d: Towards visual-based spatial-temporal intelligence.arXiv preprint arXiv:2603.00515

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-25T06:16:39.922421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T06:15:33.062980Z digest=sha256:c9e3c723a30c1f7ab81092ff9c6175da9ef71fc4ea95e8c90dde18af7a49de70

Observation 2edae53f-eb7f-45f3-9d17-4835d1aef7c2 · outbound

This paper cites The Latent Space: Foundation, Evolution, Mechanism, Ability, and Outlook.

4DThinker: Thinking with 4D Imagery for Dynamic Spatial Understanding The Latent Space: Foundation, Evolution, Mechanism, Ability, and Outlook

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-06-08T02:03:55.088792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T06:15:33.062980Z digest=sha256:5000cb4f69868befd6ae537187975cbfbac80b08bd30eb37a6c15a9c2b7cf6d6

Observation 360b37fa-c8b9-4188-9374-b4298a8fafbd · outbound

This paper cites Dsi-bench: A benchmark for dynamic spatial intelligence.arXiv preprint arXiv:2510.18873.

4DThinker: Thinking with 4D Imagery for Dynamic Spatial Understanding Dsi-bench: A benchmark for dynamic spatial intelligence.arXiv preprint arXiv:2510.18873

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-25T06:16:39.943395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T06:15:33.062980Z digest=sha256:6d0b302b8ad22cd75d1ab8ff791308cfe8e857c2d2f5bd328df5d4fed75b3336

Observation a980be5e-9add-4cf2-a111-b69e9f7eadf1 · outbound

This paper cites Learning from videos for 3d world: Enhancing mllms with 3d vision geometry priors.arXiv preprint arXiv:2505.24625.

4DThinker: Thinking with 4D Imagery for Dynamic Spatial Understanding Learning from videos for 3d world: Enhancing mllms with 3d vision geometry priors.arXiv preprint arXiv:2505.24625

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-25T06:16:40.016165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T06:15:33.062980Z digest=sha256:679f8817f869b105df0902f21d95a3a819f4d2db69cd749a0fe0923d4112df95

Observation 427c47a6-02ff-4d9c-a46f-e5317ee75d61 · outbound

This paper cites Learning to reason in 4d: Dynamic spatial understanding for vision language models.arXiv preprint arXiv:2512.20557, 2025a.

4DThinker: Thinking with 4D Imagery for Dynamic Spatial Understanding Learning to reason in 4d: Dynamic spatial understanding for vision language models.arXiv preprint arXiv:2512.20557, 2025a

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-25T06:16:39.910870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T06:15:33.062980Z digest=sha256:f4e01e47a29cd410d807ea0b1fad076cafe37445c7f5cbd610514bb09eea2bf5

Observation 5ea82e2a-6c93-4ba6-b2de-0f3849733e92 · outbound

This paper cites the red car on the left lane.

4DThinker: Thinking with 4D Imagery for Dynamic Spatial Understanding the red car on the left lane

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:16:40.109865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T06:15:33.062980Z digest=sha256:7d3b5b288f3299c33899847cce92c85ebab25bb0e61cf814c874099a355005bf

Observation 002578bd-e508-412e-801f-b12db3442e9e · outbound

This paper cites mental imagery.

4DThinker: Thinking with 4D Imagery for Dynamic Spatial Understanding mental imagery

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:16:40.100566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T06:15:33.062980Z digest=sha256:2f51054e36dc50f046eda3fffd3971597ae388a1365a98966c6608703a75c03a

Observation 6e44524e-8e1c-48e1-b7a6-b3ac243ac190 · outbound

This paper cites Absolute.

4DThinker: Thinking with 4D Imagery for Dynamic Spatial Understanding Absolute

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:16:40.106521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T06:15:33.062980Z digest=sha256:a643d23532fbeb9c91b205e4186695d030087b7ef413922ec69d2ad04eb517b0

Pith citing papers

No inbound Pith citation observations are available.