Pith. sign in

Paper Citation Record · LEDGER

ST-Think: How Multimodal Large Language Models Reason About 4D Worlds from Ego-Centric Videos

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2503.12542.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.12542 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:14:29.112728Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T17:27:15.721443Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation dde903a5-3660-47ac-b6f7-df37b6899a28 · inbound

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence cites this paper.

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence ST-Think: How Multimodal Large Language Models Reason About 4D Worlds from Ego-Centric Videos

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:34:36.943066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T08:34:36.824053Z digest=sha256:105d5d94b0e16bef214fba1f9225487a4007e7441859d994bed150320b50232f

Observation 6654a363-dc28-444d-82c0-a699d1dcfae2 · inbound

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence cites this paper.

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence ST-Think: How Multimodal Large Language Models Reason About 4D Worlds from Ego-Centric Videos

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-22T01:00:51.454558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T00:59:13.826054Z digest=sha256:3bd7f224affbd922d634a67a6f47f65b1365533cc764e5ca5b3d3e280e6a7a0a

Observation 7825ea91-ba7e-42fb-a506-0c2806d66248 · inbound

EgoVLM: Policy Optimization for Egocentric Video Understanding cites this paper.

EgoVLM: Policy Optimization for Egocentric Video Understanding ST-Think: How Multimodal Large Language Models Reason About 4D Worlds from Ego-Centric Videos

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:29.112728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:14:29.112728Z digest=sha256:bb06edd1c905bc307f8b60cdf96547f4fd94f1205a599a3afab69b44e5abd998

Observation 81f72875-eaac-4ac7-9633-7c3f7746bcc0 · inbound

SpatialBench: Benchmarking Multimodal Large Language Models for Spatial Cognition cites this paper.

SpatialBench: Benchmarking Multimodal Large Language Models for Spatial Cognition ST-Think: How Multimodal Large Language Models Reason About 4D Worlds from Ego-Centric Videos

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-17T04:59:04.287792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T04:54:59.903644Z digest=sha256:46493929bbb539e9bf3d58736f04e6a2bd8ef22d5cddae93c52c795be8d6795c

Observation 53e5f12f-4bd8-436b-ac18-5a7cf1b0c975 · inbound

CT-1: Vision-Language-Camera Models Transfer Spatial Reasoning Knowledge to Camera-Controllable Video Generation cites this paper.

CT-1: Vision-Language-Camera Models Transfer Spatial Reasoning Knowledge to Camera-Controllable Video Generation ST-Think: How Multimodal Large Language Models Reason About 4D Worlds from Ego-Centric Videos

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:56:00.403306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T16:53:54.166457Z digest=sha256:f8e40fa2b0e81e38c08a3de5ecd30d234e8f8c8c1ade57727ec2adab72f070fe

Observation b2c6297c-2616-41dd-aa8e-d6ede03348d7 · inbound

ViSRA: A Video-based Spatial Reasoning Agent for Multi-modal Large Language Models cites this paper.

ViSRA: A Video-based Spatial Reasoning Agent for Multi-modal Large Language Models ST-Think: How Multimodal Large Language Models Reason About 4D Worlds from Ego-Centric Videos

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:46:35.883391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T04:00:23.681682Z digest=sha256:9d99d3956a77c31af4c3cc1874bccc1064066b81fc9dd51632e96cabd6086359

Observation 2cb0bb14-a96f-4199-a15b-2cfaab223378 · inbound

EgoProx: Evaluating MLLMs on Egocentric 3D Proximity Reasoning Across a Cognitive Hierarchy cites this paper.

EgoProx: Evaluating MLLMs on Egocentric 3D Proximity Reasoning Across a Cognitive Hierarchy ST-Think: How Multimodal Large Language Models Reason About 4D Worlds from Ego-Centric Videos

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:34:40.492377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T13:28:43.538541Z digest=sha256:8649efe754a1f6738be180faa3200e7f02e47998d9b64277dbfa91be5322a304

Observation c2728fcc-64a1-4f5c-b8ad-b65e11ea6ebc · inbound

Watch, Remember, Reason: Human-View Video Understanding with MLLMs cites this paper.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs ST-Think: How Multimodal Large Language Models Reason About 4D Worlds from Ego-Centric Videos

Reference 238

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:27:15.723158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:a2b9b1a44d4136b69a0f05a3f27640ea0ce61d1eb44de44429e823a3ff8fc2a8