Pith. sign in

REVIEW 2 cited by

Spatiality-guided Transformer for 3D Dense Captioning on Point Clouds

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2204.10688 v1 pith:XBL6S4L4 submitted 2022-04-22 cs.CV

Spatiality-guided Transformer for 3D Dense Captioning on Point Clouds

classification cs.CV
keywords captioningdenseobjectobjectsscenespacap3dcloudspoint
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Dense captioning in 3D point clouds is an emerging vision-and-language task involving object-level 3D scene understanding. Apart from coarse semantic class prediction and bounding box regression as in traditional 3D object detection, 3D dense captioning aims at producing a further and finer instance-level label of natural language description on visual appearance and spatial relations for each scene object of interest. To detect and describe objects in a scene, following the spirit of neural machine translation, we propose a transformer-based encoder-decoder architecture, namely SpaCap3D, to transform objects into descriptions, where we especially investigate the relative spatiality of objects in 3D scenes and design a spatiality-guided encoder via a token-to-token spatial relation learning objective and an object-centric decoder for precise and spatiality-enhanced object caption generation. Evaluated on two benchmark datasets, ScanRefer and ReferIt3D, our proposed SpaCap3D outperforms the baseline method Scan2Cap by 4.94% and 9.61% in CIDEr@0.5IoU, respectively. Our project page with source code and supplementary files is available at https://SpaCap3D.github.io/ .

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet

    cs.CV 2026-07 conditional novelty 6.0

    PVCap combines instance-mixing data augmentation with pseudo-labels and a voxel-based captioning network to achieve new state-of-the-art on 3D dense captioning benchmarks ScanRefer and Nr3D.

  2. Curvature-Aware Captioning:Leveraging Geodesic Attention for 3D Scene Understanding

    cs.CV 2026-05 unverdicted novelty 6.0

    A new framework combines self-attention on the Oblique manifold with bidirectional geodesic cross-attention on the Lorentz hyperboloid to improve both localization accuracy and descriptive coherence in 3D dense captioning.