Pith. sign in

hub

Learning from videos for 3d world: Enhancing mllms with 3d vision geometry priors.arXiv preprint arXiv:2505.24625

28 Pith papers cite this work. Polarity classification is still indexing.

28 Pith papers citing it

hub tools

citation-role summary

background 3 dataset 1

citation-polarity summary

fields

cs.CV 27 cs.RO 1

years

2026 26 2025 2

representative citing papers

OmniSpace: Efficient Geometry Awareness for Autonomous Vehicles MLLMs

cs.CV · 2026-06-21 · unverdicted · novelty 6.0

OmniSpace is a plug-and-play method that improves spatial reasoning in MLLMs for AV by injecting camera pose, using epipolar attention across views, and distilling 3D geometric knowledge to overcome weak cross-view correspondence and depth estimation.

Unlocking Dense Metric Depth Estimation in VLMs

cs.CV · 2026-05-15 · unverdicted · novelty 6.0 · 2 refs

DepthVLM converts a standard VLM into a dense metric depth predictor by attaching a lightweight head and training under unified vision-text supervision, outperforming prior VLMs and some pure vision models on a new indoor-outdoor benchmark.

citing papers explorer

Showing 28 of 28 citing papers.