Pith. sign in

Paper Citation Record · LEDGER

Language-Image Models with 3D Understanding

As of 24 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2405.03685.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.03685 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:53:09.797343Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T19:05:00.618076Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 8f3fc286-a2ef-424e-988d-e4f0a0befd3e · inbound

EMMA: End-to-End Multimodal Model for Autonomous Driving cites this paper.

EMMA: End-to-End Multimodal Model for Autonomous Driving Language-Image Models with 3D Understanding

Reference 197

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T05:08:54.485083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-15T05:08:54.368109Z digest=sha256:a211451eee7787a89b5d1676388798a4cb720dae5d16662e5f509a0e3904a7e2

Observation add2d72e-3b04-461f-bfcc-1d9e03910ff2 · inbound

Do large language vision models understand 3D shapes? cites this paper.

Do large language vision models understand 3D shapes? Language-Image Models with 3D Understanding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T15:32:59.771499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:32:59.771499Z digest=sha256:c11451ecbac4c0f10a1732590b99ee819c3818ea640da037c2658c737bb54996

Observation 35ee3faa-f307-44e3-9b6c-d28be498676b · inbound

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving cites this paper.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Language-Image Models with 3D Understanding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.797343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.797343Z digest=sha256:81f2427b184a62bb5b72dbd0ce3301b329f4392becf153f71439ec8719dd47f9

Observation a046a27a-9795-44b2-90a5-db4f197a2d93 · inbound

Abstract 3D Perception for Spatial Intelligence in Vision-Language Models cites this paper.

Abstract 3D Perception for Spatial Intelligence in Vision-Language Models Language-Image Models with 3D Understanding

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-17T22:45:24.425330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T22:43:16.761970Z digest=sha256:b8deac62c8baad717f2f3eac70a241102f3c3677a9c097860ac1a5afef5f226e

Observation 4497d9cb-7957-4e6f-855c-acacef9c1b15 · inbound

UniDrive-WM: Unified Understanding, Planning and Generation World Model for Autonomous Driving cites this paper.

UniDrive-WM: Unified Understanding, Planning and Generation World Model for Autonomous Driving Language-Image Models with 3D Understanding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T12:06:26.385648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:06:26.385648Z digest=sha256:a0f5806f3fb76f8d69bf41fbfbf7f543b80382cd6badb144c85c2a0c18904344

Observation a19c185b-4b09-4a03-a3ac-67ea31f57fbe · inbound

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies cites this paper.

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies Language-Image Models with 3D Understanding

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:05:10.021577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T13:04:30.544504Z digest=sha256:a7ca1ae8cbf3f0a4aa33df675478b3278421ae82f6bfb45cb37745bee11d7ebc

Observation a825d5c6-361f-4032-84ec-2cb6bbbc3e42 · inbound

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies cites this paper.

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies Language-Image Models with 3D Understanding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T21:37:10.906144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:37:10.906144Z digest=sha256:97779430d9e9b5c41d26d01f0cca7f076aa3002c6213561e6a7b47522684d84f

Observation cfcca018-d144-4c9b-a1d9-90444d627d1c · inbound

Boxer: Robust Lifting of Open-World 2D Bounding Boxes to 3D cites this paper.

Boxer: Robust Lifting of Open-World 2D Bounding Boxes to 3D Language-Image Models with 3D Understanding

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:20:49.836772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T19:14:24.020697Z digest=sha256:d6900fcab6852c345d71945acb276283d15fbc3eb1cbfd14c29dd3d7369a2c84

Observation 25ca5235-8fe2-4e90-b1e9-8032acddda0f · inbound

GeoWorld-VLM: Geometry from World Models for Vision-Language Models cites this paper.

GeoWorld-VLM: Geometry from World Models for Vision-Language Models Language-Image Models with 3D Understanding

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-20T17:58:49.597274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-20T17:57:00.909897Z digest=sha256:241fc33cb0918589c175457a515f8df12d26bde2e52276dac4ded9f7ead7cd03

Observation d7732fa2-b776-4ef5-9ed9-568181bcdb28 · inbound

GeoWorld-VLM: Geometry from World Models for Vision-Language Models cites this paper.

GeoWorld-VLM: Geometry from World Models for Vision-Language Models Language-Image Models with 3D Understanding

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-06-30T19:05:00.620307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-30T19:02:05.125937Z digest=sha256:83aad6fc71429ec39fc90d979cc0012d08c467efa716104f7a142be835c9a109

Observation fedeb9dc-1b3c-471d-885b-bfb7843046fe · inbound

GReFEM: Multimodal LLMs as Zero-Shot Semantic Assistants for Physics-Guided 3D Mesh Refinement cites this paper.

GReFEM: Multimodal LLMs as Zero-Shot Semantic Assistants for Physics-Guided 3D Mesh Refinement Language-Image Models with 3D Understanding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-13T06:42:45.558324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T06:42:45.558324Z digest=sha256:21b4dd05e84000936071e010ef9404f6971ab72c8a2b0ddfa09cd0520fb858a7

Observation 988ffc2f-282c-434c-ab9e-b9f5c0cf5df4 · inbound

RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation cites this paper.

RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation Language-Image Models with 3D Understanding

Reference 198

Resolution
unresolved
no resolver link, observed 2026-08-01T14:39:52.305768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:39:52.305768Z digest=sha256:a5ba5afa3d48275220ba605652ad1f1b5861b5e02640230a25a52df4116266a2