Pith. sign in

Paper Citation Record · LEDGER

Coarse Correspondences Boost Spatial-Temporal Reasoning in Multimodal Language Model

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2408.00754.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2408.00754 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:42:52.468598Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T09:27:43.983079Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 9f2284d1-ed7c-4516-aa29-aaf7658178a5 · inbound

Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces cites this paper.

Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces Coarse Correspondences Boost Spatial-Temporal Reasoning in Multimodal Language Model

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:27:43.990119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T09:27:43.919941Z digest=sha256:f8f4e62be8dbd18cb1ab8779171901d7d06714c7416a1fe45deadb07a1d5dd1c

Observation b48a3527-234a-4dac-87a2-444f5473da37 · inbound

UniVG-R1: Reasoning Guided Universal Visual Grounding with Reinforcement Learning cites this paper.

UniVG-R1: Reasoning Guided Universal Visual Grounding with Reinforcement Learning Coarse Correspondences Boost Spatial-Temporal Reasoning in Multimodal Language Model

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:52.468598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:52.468598Z digest=sha256:df7ae273a8665a98dec2e341bed0324ee726f39aafe5b0e2ee07963343f54114

Observation 68399f19-443a-4216-9391-869add710de0 · inbound

Out of Sight, Not Out of Context? Egocentric Spatial Reasoning in VLMs Across Disjoint Frames cites this paper.

Out of Sight, Not Out of Context? Egocentric Spatial Reasoning in VLMs Across Disjoint Frames Coarse Correspondences Boost Spatial-Temporal Reasoning in Multimodal Language Model

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:18.039306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:18.039306Z digest=sha256:d7d736debf1f73147fc1cfb27fcdf9613b743766881dddc962e37a7b3277c402

Observation 06ca9977-97ca-4c6e-a229-b1cf9a2edc0a · inbound

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces cites this paper.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Coarse Correspondences Boost Spatial-Temporal Reasoning in Multimodal Language Model

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:12.740878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:12.740878Z digest=sha256:1dee6506bdfcc15a3524a1c235a42326ea1ff308e54491ffd058278bf70e6af4

Observation 1adc13ec-88b0-4da0-9873-fc780c13da75 · inbound

AD^2-Bench: A Hierarchical CoT Benchmark for MLLM in Autonomous Driving under Adverse Conditions cites this paper.

AD^2-Bench: A Hierarchical CoT Benchmark for MLLM in Autonomous Driving under Adverse Conditions Coarse Correspondences Boost Spatial-Temporal Reasoning in Multimodal Language Model

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T04:49:28.115899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:49:28.115899Z digest=sha256:9b78c37f58c7baade202c9936353427dfddbc13853b29a438baa3b93fe2f7bbf

Observation ed1433d7-667b-469a-9390-b35f87b0fbe0 · inbound

Ego-R1: Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning cites this paper.

Ego-R1: Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning Coarse Correspondences Boost Spatial-Temporal Reasoning in Multimodal Language Model

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:32.084749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:34:32.084749Z digest=sha256:c62ef3db38a6a0b4ed8f98effa59cea7ab023d6d8cd11ed91fc530c4f1b08e8f

Observation 31c2b4e6-2685-4110-bcb6-df4e1146e7c6 · inbound

High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning cites this paper.

High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning Coarse Correspondences Boost Spatial-Temporal Reasoning in Multimodal Language Model

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-19T06:12:07.045708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T06:10:57.219445Z digest=sha256:5b69d5843573bd46ee5939907b7155c4e9622e64c61ffcefe683050c332d96a6

Observation 1c0cf0b0-9d36-4abf-b355-42620e1141c5 · inbound

From Correspondence to Actions: Human-Like Multi-Image Spatial Reasoning in Multi-modal Large Language Models cites this paper.

From Correspondence to Actions: Human-Like Multi-Image Spatial Reasoning in Multi-modal Large Language Models Coarse Correspondences Boost Spatial-Temporal Reasoning in Multimodal Language Model

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T03:16:50.003291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:16:50.003291Z digest=sha256:dbc5345220b507b278bf8aaa4aaed5d7c9776f10c07149c1aa8a21a3b2d06646