Pith. sign in

Paper Citation Record · LEDGER

Language-Image Models with 3D Understanding

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2405.03685.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.03685 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T12:06:26.385648Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T19:05:00.618076Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 8f3fc286-a2ef-424e-988d-e4f0a0befd3e · inbound

EMMA: End-to-End Multimodal Model for Autonomous Driving cites this paper.

EMMA: End-to-End Multimodal Model for Autonomous Driving Language-Image Models with 3D Understanding

Reference 197

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T05:08:54.485083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-15T05:08:54.368109Z digest=sha256:ee5ba574fa5b104192453181c6de5b996c4d480be10f8052d6ad1d1d6b265150

Observation a046a27a-9795-44b2-90a5-db4f197a2d93 · inbound

Abstract 3D Perception for Spatial Intelligence in Vision-Language Models cites this paper.

Abstract 3D Perception for Spatial Intelligence in Vision-Language Models Language-Image Models with 3D Understanding

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-17T22:45:24.425330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T22:43:16.761970Z digest=sha256:989c1752a4302ddd399e112fd0091ab7bbf6120e975fdbff764c26c0a6d8f0d4

Observation 4497d9cb-7957-4e6f-855c-acacef9c1b15 · inbound

UniDrive-WM: Unified Understanding, Planning and Generation World Model for Autonomous Driving cites this paper.

UniDrive-WM: Unified Understanding, Planning and Generation World Model for Autonomous Driving Language-Image Models with 3D Understanding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T12:06:26.385648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:06:26.385648Z digest=sha256:e7f639428f3a9136b844efdfcf975ef445290eae9637c5a3dad33fc8313eaeee

Observation a19c185b-4b09-4a03-a3ac-67ea31f57fbe · inbound

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies cites this paper.

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies Language-Image Models with 3D Understanding

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:05:10.021577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T13:04:30.544504Z digest=sha256:4469170dbc420acb0b7145c8201370b4eae0620d2af497c1402cf26c91e5a5e0

Observation a825d5c6-361f-4032-84ec-2cb6bbbc3e42 · inbound

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies cites this paper.

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies Language-Image Models with 3D Understanding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T21:37:10.906144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:37:10.906144Z digest=sha256:466d9ebc0c69eea87118ad8dae74ca8602bb0ee653cdd6157abc9d972665e09e

Observation cfcca018-d144-4c9b-a1d9-90444d627d1c · inbound

Boxer: Robust Lifting of Open-World 2D Bounding Boxes to 3D cites this paper.

Boxer: Robust Lifting of Open-World 2D Bounding Boxes to 3D Language-Image Models with 3D Understanding

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:20:49.836772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T19:14:24.020697Z digest=sha256:cbd6006db11eafc6928ad018f15204493904a5b2e1bac4bd2b7e2199ff035d5b

Observation 25ca5235-8fe2-4e90-b1e9-8032acddda0f · inbound

GeoWorld-VLM: Geometry from World Models for Vision-Language Models cites this paper.

GeoWorld-VLM: Geometry from World Models for Vision-Language Models Language-Image Models with 3D Understanding

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-20T17:58:49.597274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T17:57:00.909897Z digest=sha256:d2f3d315766d11c0997a5248ab51d54bfe77db6ec4393e582d944b729ded97cf

Observation d7732fa2-b776-4ef5-9ed9-568181bcdb28 · inbound

GeoWorld-VLM: Geometry from World Models for Vision-Language Models cites this paper.

GeoWorld-VLM: Geometry from World Models for Vision-Language Models Language-Image Models with 3D Understanding

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-06-30T19:05:00.620307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T19:02:05.125937Z digest=sha256:d44782b86571006a0444bf65cdd6494c3d4afe173f00aa009e9a3dbb93ac1377

Observation fedeb9dc-1b3c-471d-885b-bfb7843046fe · inbound

GReFEM: Multimodal LLMs as Zero-Shot Semantic Assistants for Physics-Guided 3D Mesh Refinement cites this paper.

GReFEM: Multimodal LLMs as Zero-Shot Semantic Assistants for Physics-Guided 3D Mesh Refinement Language-Image Models with 3D Understanding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-13T06:42:45.558324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T06:42:45.558324Z digest=sha256:8cddda243f72f1636eaec7c1ce9122aeccf3622f3d2af1093391be8ef8e6f8e9

Observation 988ffc2f-282c-434c-ab9e-b9f5c0cf5df4 · inbound

RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation cites this paper.

RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation Language-Image Models with 3D Understanding

Reference 198

Resolution
unresolved
no resolver link, observed 2026-08-01T14:39:52.305768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:39:52.305768Z digest=sha256:ea6f1e40e2eaa127654f0c872df6fe58e59de8f2dc6400ca1457abfee24c033f