REVIEW 2 cited by
Multi-view 3D Reconstruction with Transformer
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
abstract
Deep CNN-based methods have so far achieved the state of the art results in multi-view 3D object reconstruction. Despite the considerable progress, the two core modules of these methods - multi-view feature extraction and fusion, are usually investigated separately, and the object relations in different views are rarely explored. In this paper, inspired by the recent great success in self-attention-based Transformer models, we reformulate the multi-view 3D reconstruction as a sequence-to-sequence prediction problem and propose a new framework named 3D Volume Transformer (VolT) for such a task. Unlike previous CNN-based methods using a separate design, we unify the feature extraction and view fusion in a single Transformer network. A natural advantage of our design lies in the exploration of view-to-view relationships using self-attention among multiple unordered inputs. On ShapeNet - a large-scale 3D reconstruction benchmark dataset, our method achieves a new state-of-the-art accuracy in multi-view reconstruction with fewer parameters ($70\%$ less) than other CNN-based methods. Experimental results also suggest the strong scaling capability of our method. Our code will be made publicly available.
Forward citations
Cited by 2 Pith papers
-
Multistream Network for LiDAR and Camera-based 3D Object Detection in Outdoor Scenes
MuStD, a multistream LiDAR-camera fusion network with a novel UV-Polar block, reports competitive 3D detection on the KITTI benchmark.
-
Refine3DNet: Scaling Precision in 3D Object Reconstruction from Multi-View RGB Images using Attention
A hybrid CNN-transformer 3D reconstruction method claims state-of-the-art IoU on ShapeNet, yet the architectural description, IoU equation, and baseline tables are internally inconsistent.
Discussion (0). Continue with ORCID to comment.