Pith. sign in

hub Canonical reference

Bridgevla: Input-output alignment for efficient 3d manipulation learning with vision-language models

Canonical reference. 90% of citing Pith papers cite this work as background.

24 Pith papers citing it
Background 90% of classified citations

hub tools

citation-role summary

background 10

citation-polarity summary

fields

cs.RO 21 cs.CV 3

years

2026 21 2025 3

polarities

background 9 unclear 1

representative citing papers

THOM: Generating Physically Plausible Hand-Object Meshes From Text

cs.CV · 2026-04-03 · unverdicted · novelty 7.0

THOM is a training-free two-stage framework that generates physically plausible hand-object 3D meshes directly from text by combining text-guided Gaussians with contact-aware physics optimization and VLM refinement.

GeoAlign: Beyond Semantics with State-Guided Spatial Alignment in VLA Models

cs.RO · 2026-06-02 · unverdicted · novelty 5.0

GeoAlign post-trains an RGB geometry branch on robot RGB-D data to produce GEP features that are queried by proprioceptive state to generate phase-dependent geometry tokens, yielding 99.0% on LIBERO, 85.3% on SimplerEnv-Fractal, and 78.8% on real ALOHA tasks.

R3D: Revisiting 3D Policy Learning

cs.CV · 2026-04-16 · unverdicted · novelty 5.0

A transformer 3D encoder plus diffusion decoder architecture, with 3D-specific augmentations, outperforms prior 3D policy methods on manipulation benchmarks by improving training stability.

citing papers explorer

Showing 24 of 24 citing papers.