Pith. sign in

Paper Citation Record · LEDGER

Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2412.00493.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.00493 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:46:13.494993Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T01:00:51.358037Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 31acac97-89dd-432e-855f-aef3ea07a7bd · inbound

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts cites this paper.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:13.494993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:13.494993Z digest=sha256:8601015d1f7c280470529b3d14428fe27da95047d20c2593cf131ee060ea7ac1

Observation 0c60e752-4b9b-46f1-8d89-b0217515b7c5 · inbound

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence cites this paper.

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:34:36.872348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T08:34:36.824053Z digest=sha256:5f49a54f827c36674ae2662ddca2e4624c9e3e19fb8e5ecadd3364b973a2ce7b

Observation a4c06c78-5058-4645-811b-0778b0d90f73 · inbound

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence cites this paper.

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-22T01:00:51.361125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T00:59:13.826054Z digest=sha256:edc050efa461a4a39004c1e92db1cc951e828129d3f484c71b8cbc924f4ddb76

Observation 3eabc801-a4d0-4b7b-ae40-b2bc915861bf · inbound

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing cites this paper.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.802988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.802988Z digest=sha256:a2690301622c14a57be3180577a0d4e9d4066b802b7e4a6190a32fc85a5456cf

Observation 2d0cd3b5-f8a4-4fb0-95d8-c5b13fa3842e · inbound

Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture cites this paper.

Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T11:40:24.704510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:40:24.704510Z digest=sha256:126cc9fb2ca4b7f03d5d0e05b470f27e11b3adde1e1806f6f72fd17b45fdf5e9

Observation 9f03306a-48ba-4fb2-813d-0922a73fea03 · inbound

MAG-3D: Multi-Agent Grounded Reasoning for 3D Understanding cites this paper.

MAG-3D: Multi-Agent Grounded Reasoning for 3D Understanding Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:51:19.088377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T17:25:31.097385Z digest=sha256:34f7aa5f0fef22dea5680784b2f0542903b32b8232d806c807e08490a0d97114

Observation 6f6f27bb-e7d2-4e8a-8ae7-ca3dc3495693 · inbound

A Progressive Training Strategy for Vision-Language Models to Counteract Spatio-Temporal Hallucinations in Embodied Reasoning cites this paper.

A Progressive Training Strategy for Vision-Language Models to Counteract Spatio-Temporal Hallucinations in Embodied Reasoning Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T09:41:01.809714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:53:52.775348Z digest=sha256:628886d8574a57d2155094375a502ab66b620752bf8013d9a824c7907e7110c0

Observation 40688cb6-5a6f-412d-afde-7fe87b94e9ca · inbound

Seeing Once is Enough? Online Geometry-Aware Token Pruning for 3D Question Answering cites this paper.

Seeing Once is Enough? Online Geometry-Aware Token Pruning for 3D Question Answering Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-11T21:50:49.625343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T21:50:49.625343Z digest=sha256:9370194ed0883fd6a2cad0ce26e61e80e5666cbc29ff15b85c1b4a2da23973f9