Pith. sign in

Paper Citation Record · LEDGER

LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 53 inbound Pith citation observations for arXiv:2409.18125.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2409.18125 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 53 of 53 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:46:14.423259Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T15:39:56.857028Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ea7272f7-bad2-4ed9-978f-d9fab1682bfe · inbound

Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces cites this paper.

Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 107

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T09:27:44.102329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T09:27:43.919941Z digest=sha256:1e7cf9e565220f582ca72ea8c0fb9ff54441790eafe5edb6956aa93878866e79

Observation fbe72696-13d7-474c-bec0-7d43145b806e · inbound

SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model cites this paper.

SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:12:19.853033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T06:12:19.643111Z digest=sha256:c91e28672bdccba61953464fc262959994794eec103c6705c831debbdbb51ba3

Observation b26fd254-490e-46e6-b794-35981a26791a · inbound

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts cites this paper.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:14.423259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:14.423259Z digest=sha256:7d79eef498a1807a3b8a090c9688aa523223c427eedaf30169d8846af92909dd

Observation 6ac13126-0ca5-4894-912f-2fb74e07e39c · inbound

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence cites this paper.

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:34:36.868541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T08:34:36.824053Z digest=sha256:c80a5b899cf9d7b1900373e52354a81d85d78923e040d15f4dfa66a4e3393c0c

Observation 506780a9-b225-41c3-82f0-71e1a9ab68f3 · inbound

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence cites this paper.

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-22T01:00:51.268241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T00:59:13.826054Z digest=sha256:21f731f31a473a4c338e062f7bfc7e2d73f641fbf4e802392dba08d521138070

Observation e9eba971-990b-4037-98c2-afd389d58a14 · inbound

Out of Sight, Not Out of Context? Egocentric Spatial Reasoning in VLMs Across Disjoint Frames cites this paper.

Out of Sight, Not Out of Context? Egocentric Spatial Reasoning in VLMs Across Disjoint Frames LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:20.069930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:20.069930Z digest=sha256:6e135fe03da9ea61ab4ae987d302eeea5cdf1230dd25fe80ad51bd87c0ac779c

Observation 7046edbd-9db3-4a06-8615-8c6489c47695 · inbound

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces cites this paper.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 111

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:16.730134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:16.730134Z digest=sha256:25f449d824eaf05209cd3c33850de473a59246318926352ffd94acc1cfd06fbc

Observation 264df855-651c-4abf-a942-66b432d95dc4 · inbound

Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs cites this paper.

Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T10:25:48.356054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:25:48.356054Z digest=sha256:812ab26cbc21100006ba3b3141b289da9eb8ee0f2a78d9bc569350c4b92bc1b5

Observation 7ab239f0-022e-42d6-bc5d-4a37fbd8672b · inbound

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing cites this paper.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.588064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.588064Z digest=sha256:0b61a3260ae0a6d9c2a0ba646518278b758cd018868ad473da66c46caaa8eafa

Observation 50cb826f-7ff1-4bd9-9c9f-50f230d11047 · inbound

Pts3D-LLM: Studying the Impact of Token Structure for 3D Scene Understanding With Large Language Models cites this paper.

Pts3D-LLM: Studying the Impact of Token Structure for 3D Scene Understanding With Large Language Models LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T10:20:37.480142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:20:37.480142Z digest=sha256:143f50472b77d405c0f425feb0b1102eb69287e110fe7aeac7133e5aa334f77b

Observation c8499945-4b1d-4d4d-89e3-eaac8adcd815 · inbound

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making cites this paper.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:27.599419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:27.599419Z digest=sha256:7f3c39b2238255ea0f34b94bfe87733b7604e52bd3afa95af2470bbe89f4486b

Observation baa76daf-d0cc-4046-867a-c8c4e18cb39f · inbound

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering cites this paper.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:12.924934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:12.924934Z digest=sha256:b72568f312f099bf72d43ba2c2ecf1f680d4df0c78203217aca9db2285ce51b5

Observation b4048b41-fcc1-4dee-a24f-dc6722516e0a · inbound

3D-R1: Enhancing Reasoning in 3D VLMs for Unified Scene Understanding cites this paper.

3D-R1: Enhancing Reasoning in 3D VLMs for Unified Scene Understanding LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T10:49:47.146212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T10:49:47.146212Z digest=sha256:336b38113b2630f55bcc53055aeab80a511ba76855e320d9195a1cfb837f0975

Observation 3cb81618-21a3-405b-8dac-28d034d2a138 · inbound

PySeizure: A single machine learning classifier framework to detect seizures in diverse datasets cites this paper.

PySeizure: A single machine learning classifier framework to detect seizures in diverse datasets LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T22:18:05.150726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:18:05.150726Z digest=sha256:78a082a77372efc7e2d8a6fc1457bb94363aef1ae8da4b2679d1f29bc06e33ae

Observation 8fb6f571-f63e-4466-985d-9a52c58cb6d2 · inbound

GeoVLA: Empowering 3D Representations in Vision-Language-Action Models cites this paper.

GeoVLA: Empowering 3D Representations in Vision-Language-Action Models LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T21:16:50.649376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:16:50.649376Z digest=sha256:7c0eeb759a4fb6307f06c9a84d1c1dcaddb5a14ac5443c7d1673f4ee08f5de6a

Observation d947ec6d-b7b5-4312-a67d-8132fdde69d8 · inbound

GAE: Unleashing Physical Potential of VLM with Generalizable Action Expert cites this paper.

GAE: Unleashing Physical Potential of VLM with Generalizable Action Expert LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-04T11:38:13.467839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:38:13.467839Z digest=sha256:7befddaae664f2aec9f339a3ac76427ea839b1ce9421d2329d47d2a1d6d51851

Observation 8b515b32-dab1-44d5-ace8-129bf2380628 · inbound

POMA-3D: The Point Map Way to 3D Scene Understanding cites this paper.

POMA-3D: The Point Map Way to 3D Scene Understanding LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:30:11.559924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T20:27:27.347592Z digest=sha256:2d40273e7502285e33fed8f126c3a5be8a4e9ba50c385afd2a32064d66663598

Observation 9ece6334-1468-428b-868d-cd6c9d1295ac · inbound

Boosting Reasoning in Large Multimodal Models via Activation Replay cites this paper.

Boosting Reasoning in Large Multimodal Models via Activation Replay LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:09:03.922650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T05:05:48.682057Z digest=sha256:15288c3e818079b67c34f4abb246176a30e7bebb32572c8d4a8bce04917def85

Observation 7ea85dfe-8b0f-466b-b79d-b8a9c07fb2f3 · inbound

Vision-Language Memory for Spatial Reasoning cites this paper.

Vision-Language Memory for Spatial Reasoning LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-03T20:15:37.516145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:15:37.516145Z digest=sha256:7e363039275d07c68ab0fb155dfce3294851325e0420c27750d58d1acd384dac

Observation 6e8a4b7d-4e8b-426a-888d-96b9dceb4bfa · inbound

SpatialBench: Benchmarking Multimodal Large Language Models for Spatial Cognition cites this paper.

SpatialBench: Benchmarking Multimodal Large Language Models for Spatial Cognition LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-17T04:59:04.240779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T04:54:59.903644Z digest=sha256:10916fda92a344dbf5e1ca36129b1badd57ccced368d65eb59263f6e74f1c105

Observation fef0e861-6800-4855-80ee-a03f7a27a63b · inbound

SpaceDrive: Infusing Spatial Awareness into VLM-based Autonomous Driving cites this paper.

SpaceDrive: Infusing Spatial Awareness into VLM-based Autonomous Driving LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-05-22T12:21:31.269741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T12:17:54.325055Z digest=sha256:661f6270a9af12475dd2278e95376661c5f5ec576a9a9d5ced44072a36344519

Observation 0c2d3562-4638-4175-90ac-aafe8eccde47 · inbound

SpaceDrive: Infusing Spatial Awareness into VLM-based Autonomous Driving cites this paper.

SpaceDrive: Infusing Spatial Awareness into VLM-based Autonomous Driving LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-03T17:07:45.189496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:07:45.189496Z digest=sha256:2bc5090187b23ed0a808127e127c55df0a1c138a622dea07b465df8c59c3b0d3

Observation 725628d7-5b50-4f3c-9e4c-88091d41cc87 · inbound

OpenGround: Planning-based Online Perception for Open-World 3D Visual Grounding cites this paper.

OpenGround: Planning-based Online Perception for Open-World 3D Visual Grounding LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-03T13:46:30.457026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:46:30.457026Z digest=sha256:5e608587cf76a872fffe5a8f67e13b15d701a82b79348b7f73cf0f975811a92f

Observation bf521ba6-1604-4758-b59a-e5446d40d410 · inbound

Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models cites this paper.

Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:26:26.916148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T18:25:21.621268Z digest=sha256:2d31fe00be6a4eb3f73759fcf7ddad7ff8e63a69b456220d58832d58ba15e33d

Observation 46529768-f4c3-4547-991f-32b41b4863c2 · inbound

Boosting MLLM Spatial Reasoning with Geometrically Referenced 3D Scene Representations cites this paper.

Boosting MLLM Spatial Reasoning with Geometrically Referenced 3D Scene Representations LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:35:56.050224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T14:31:03.909336Z digest=sha256:a14e026ef0f2f9b3eff530bea00850c8109464e435c84e56e729873258934c0b

Observation 78606104-cf6a-4120-8cd3-97f27074cf27 · inbound

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models cites this paper.

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.995051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.995051Z digest=sha256:09f2377d8b81c772f3b29bf0bc019ece66d05d7f56863c0410597930bc93855f

Observation 57ad9ea2-ad06-4ac9-a736-0307cecf070c · inbound

VIEW2SPACE: Studying Multi-View Visual Reasoning from Sparse Observations cites this paper.

VIEW2SPACE: Studying Multi-View Visual Reasoning from Sparse Observations LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 53

Resolution
unresolved
no resolver link, observed 2026-07-13T23:42:44.515158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:42:44.515158Z digest=sha256:0753de2bcfc7637a3842a3b04260a8b448ed07393910421128bd67cc3eb0e0b3

Observation 829037dc-1aad-4ab7-8b1b-df4face8d24e · inbound

Feeling the Space: Egomotion-Aware Video Representation for Efficient and Accurate 3D Scene Understanding cites this paper.

Feeling the Space: Egomotion-Aware Video Representation for Efficient and Accurate 3D Scene Understanding LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 69

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T09:30:22.395942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T09:30:03.668178Z digest=sha256:2c963c9115769d6a080ce29dc15f992f9a0b4628aa5565aec3239d343311b4a8

Observation c438ce0f-c7a5-40bd-9c12-975277b5176a · inbound

Unleashing Spatial Reasoning in Multimodal Large Language Models via Textual Representation Guided Reasoning cites this paper.

Unleashing Spatial Reasoning in Multimodal Large Language Models via Textual Representation Guided Reasoning LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-15T00:09:35.116047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T00:08:22.362342Z digest=sha256:c5600c7ebedf7a1c3755dc690b1ffbad328ea4747c7e8d662e2980d990ff8183

Observation 0691fced-5885-4c15-adcb-4003ed9e02a6 · inbound

3D-IDE: 3D Implicit Depth Emergent cites this paper.

3D-IDE: 3D Implicit Depth Emergent LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:38:11.453485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T22:34:04.833557Z digest=sha256:3989232a4d355ed44390eadcc2619c398239d54c5912ad8ad56620cbe494d4e9

Observation 3a769354-f6c4-494b-89d9-0e9f76277df1 · inbound

EgoMind: Activating Spatial Cognition through Linguistic Reasoning in MLLMs cites this paper.

EgoMind: Activating Spatial Cognition through Linguistic Reasoning in MLLMs LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:43:22.828897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T22:41:09.840792Z digest=sha256:f17ef4969f4998eedf056d183ffe404298d99cdbde7c80cf6c8b26b6a7b48955

Observation e8389b68-bbd3-450c-a0e2-ab92616c3637 · inbound

EgoMind: Activating Spatial Cognition through Linguistic Reasoning in MLLMs cites this paper.

EgoMind: Activating Spatial Cognition through Linguistic Reasoning in MLLMs LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 61

Resolution
unresolved
no resolver link, observed 2026-07-13T14:39:55.177552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T14:39:55.177552Z digest=sha256:1cb1eb2a647ca9ff707ee820310e39562f784bee0dd8c0e17eb63937032290d9

Observation 2b917ffe-6cf8-440d-804d-e651014702dc · inbound

UniMesh: Unifying 3D Mesh Understanding and Generation cites this paper.

UniMesh: Unifying 3D Mesh Understanding and Generation LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 76

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T06:56:47.338138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T06:55:42.679323Z digest=sha256:5d486fe22140315ade6aeaa1b15bb07618ebb837472af6e787f9e1eacac9330e

Observation a819ce1a-ddab-48f3-8388-30975d06eab5 · inbound

Flame3D: Zero-shot Compositional Reasoning of 3D Scenes with Agentic Language Models cites this paper.

Flame3D: Zero-shot Compositional Reasoning of 3D Scenes with Agentic Language Models LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 70

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T02:11:15.854529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T02:09:09.684786Z digest=sha256:c00cc52d329b3a36847443308a4be94053c018546f7a29e42bf053cdf8bb5842

Observation 05ddfdf3-00fb-458a-aa00-d13e6dd877d0 · inbound

Distilling 3D Spatial Reasoning into a Lightweight Vision-Language Model with CoT cites this paper.

Distilling 3D Spatial Reasoning into a Lightweight Vision-Language Model with CoT LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:31:25.837388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T02:42:09.922300Z digest=sha256:6828ab550dd692f671485be40eb53d4d2f3a2aed85a6e9ffc66d29fad794f49d

Observation b2e9ba17-1ab5-4936-a7cf-ef96753ab33d · inbound

Unlocking Dense Metric Depth Estimation in VLMs cites this paper.

Unlocking Dense Metric Depth Estimation in VLMs LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 79

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T19:23:40.971811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T19:20:04.468206Z digest=sha256:f45b9b5eac53920fd459d8fa95cd438cb2f8c3bf9a61aa0592828f38a47114ab

Observation c7c1b832-f5bc-48bb-9d9c-d284f8a40baf · inbound

Unlocking Dense Metric Depth Estimation in VLMs cites this paper.

Unlocking Dense Metric Depth Estimation in VLMs LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 86

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T07:59:51.117080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T07:54:52.926995Z digest=sha256:5debce4d68f3f4017beed385711d8a6cae00a060dda834536eb4f2e7f47db585

Observation c568eb3c-3fee-4e39-bcc1-a6c77c8e4a7e · inbound

CaMo: Camera Motion Grounded Evaluation and Training for Vision-Language Models cites this paper.

CaMo: Camera Motion Grounded Evaluation and Training for Vision-Language Models LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 55

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T05:28:04.610018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-20T05:27:30.938311Z digest=sha256:d1a33f0c31cd6b7268659927ca688695b308801ea7014902fba9613d2008539c

Observation d8ccdd82-7f95-41e5-8e78-e71a7b07fca7 · inbound

Understanding the Impact of Geometric Foundation Models on Vision-Language-Action Models cites this paper.

Understanding the Impact of Geometric Foundation Models on Vision-Language-Action Models LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 54

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T13:44:40.748826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T13:40:59.330788Z digest=sha256:2521d097a3d1c8f1cb0475ed083e3df58c4f222de8b4f39224eaf30351a742bd

Observation 4350b7bb-d452-4ce5-aed8-9627f153003b · inbound

Perceive-then-Plan: Layout-as-Policy for Monocular 3D Scene Layout Estimation cites this paper.

Perceive-then-Plan: Layout-as-Policy for Monocular 3D Scene Layout Estimation LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-06-29T23:14:01.279614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T23:11:59.282627Z digest=sha256:83103bdf0336b8e3e1343b14f997d2413d03c3f898b3fb59ad91700a6f03d2a7

Observation 19ace1ec-6239-45a8-9243-d82e6f3fa22b · inbound

Geometry-Aware Representation Denoising for Robust Multi-view 3D Reconstruction cites this paper.

Geometry-Aware Representation Denoising for Robust Multi-view 3D Reconstruction LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-06-29T23:14:01.824595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T23:08:52.333329Z digest=sha256:6d311906cf69f3e354f44272e6d9c801226c55b5992bf2bf50bb30a3795dab48

Observation c55ddd02-4122-41fc-81b9-b0161d46f9b0 · inbound

Beyond 3D VQAs: Injecting 3D Spatial Priors into Vision-Language Models for Enhanced Geometric Reasoning cites this paper.

Beyond 3D VQAs: Injecting 3D Spatial Priors into Vision-Language Models for Enhanced Geometric Reasoning LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-06-29T07:53:13.622650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T07:47:52.739735Z digest=sha256:db26b6f0b1c8fb3b025287e385b057198b844307e7e52abc9cce6c518564849a

Observation a27de77b-f796-415c-8585-73b8247538b8 · inbound

Thinking with Imagination: Agentic Visual Spatial Reasoning with World Simulators cites this paper.

Thinking with Imagination: Agentic Visual Spatial Reasoning with World Simulators LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:06:56.162004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T02:25:30.998989Z digest=sha256:0b828f92d2d34c79d65f6047abf0473508cfb66066719bf8751a2db892b44bd5

Observation 75ee56c0-637b-4033-af2f-4716bda53048 · inbound

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks cites this paper.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 91

Resolution
malformed identifier
arxiv_id, observed 2026-07-03T01:27:30.682419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:4ad64dbf5273a6460db0fd9f236eaa2662a4fd9ef349d385247e3e3074d839d2

Observation 7eb1e127-ffc6-40f8-ab3c-a070e54e1fa4 · inbound

Reinforcing Dual-Path Reasoning in Spatial Vision Language Models cites this paper.

Reinforcing Dual-Path Reasoning in Spatial Vision Language Models LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 86

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T20:08:55.586365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T01:42:30.005911Z digest=sha256:1cdaccee16c668393c0f08d3a5d2f7286527cdf75ca09bedee30c1ec3510825f

Observation 2d49d67e-8080-4b9b-8b8c-83bc07d1cde2 · inbound

Occ-VLM: Occupancy Grounded Vision Language Model for Indoor Scene Understanding cites this paper.

Occ-VLM: Occupancy Grounded Vision Language Model for Indoor Scene Understanding LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 59

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T02:59:25.806993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T18:40:20.588652Z digest=sha256:f4691a4085a8c51c41f57f2ba020c6832d9218ef2e079d10350c80f4a67c3227

Observation 63d561c5-8a01-4c23-84f7-bc3d0e40f601 · inbound

3D-PLOT-LLM: Part-Level Object Tokens for 3D Large Language Models cites this paper.

3D-PLOT-LLM: Part-Level Object Tokens for 3D Large Language Models LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T03:09:29.331044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-26T18:25:55.011125Z digest=sha256:df37dd904f984f58f0df955a4bfcc40a0ee9f8f2f85f373135f80d10f548cff8

Observation 0fcf1d7a-fc6c-4154-9227-0f0858810e3d · inbound

CVSBench: A Comprehensive Benchmark for Cross-view Spatial Reasoning and Dreaming cites this paper.

CVSBench: A Comprehensive Benchmark for Cross-view Spatial Reasoning and Dreaming LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 70

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T08:49:42.526702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T10:54:49.774490Z digest=sha256:a38c2dbfb1016419d2f488977fd06dc9f535185d08a66b965f4a223874facb5c

Observation 5b1597f1-6ccf-4f79-951a-0010d5918c3a · inbound

ObsGraph: Hierarchical Observation Representation for Embodied Reasoning and Exploration cites this paper.

ObsGraph: Hierarchical Observation Representation for Embodied Reasoning and Exploration LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T15:39:56.858556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T01:30:05.294887Z digest=sha256:703c3cf2acc9310a43593821a6cb65a009635f525efef89919f14074b1ee2d72

Observation c41b482a-6788-499c-a549-4641c73976b2 · inbound

Self in Space: Benchmarking Self-Awareness and Spatial Cognition in UAV Embodied Intelligence cites this paper.

Self in Space: Benchmarking Self-Awareness and Spatial Cognition in UAV Embodied Intelligence LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-02T06:33:31.366247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:33:31.366247Z digest=sha256:f1b8a24abf70ff9a39fc89daf9f09c045e61a4827e499540f168cc68a640851c

Observation d4fbc204-770b-4f7d-8a80-f4f86b9f6505 · inbound

Towards Enhancing 3D Spatial Reasoning in Medical Multimodal Large Language Models cites this paper.

Towards Enhancing 3D Spatial Reasoning in Medical Multimodal Large Language Models LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T03:31:46.403528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:31:46.403528Z digest=sha256:87a3663c463793624685aff07335af6ff800110505f379594df9524964d8466c

Observation 18e083b6-e336-4298-b764-a396607929ae · inbound

IGGT4D: Streaming 4D Instance-Grounded Geometry Transformer cites this paper.

IGGT4D: Streaming 4D Instance-Grounded Geometry Transformer LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T13:07:03.267487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:07:03.267487Z digest=sha256:119dcca8a447531c434089043835999d7ed90768d5e97f0e796bb078e423a7ce

Observation 8fd70c46-7875-4f27-8e60-647ad55a4dc6 · inbound

ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? cites this paper.

ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-01T09:09:24.877243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T09:09:24.877243Z digest=sha256:1db79290c33abbab33874818b73f5afbf46ceaeb406beded04ff2a0ca1a8d9f3