Pith. sign in

REVIEW 1 cited by

MVT: Multi-view Vision Transformer for 3D Object Recognition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2110.13083 v1 pith:S44R2WWA submitted 2021-10-25 cs.CV

classification cs.CV
keywords recognitionobjecttransformermulti-viewviewsvisionachievedcommunications
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Inspired by the great success achieved by CNN in image recognition, view-based methods applied CNNs to model the projected views for 3D object understanding and achieved excellent performance. Nevertheless, multi-view CNN models cannot model the communications between patches from different views, limiting its effectiveness in 3D object recognition. Inspired by the recent success gained by vision Transformer in image recognition, we propose a Multi-view Vision Transformer (MVT) for 3D object recognition. Since each patch feature in a Transformer block has a global reception field, it naturally achieves communications between patches from different views. Meanwhile, it takes much less inductive bias compared with its CNN counterparts. Considering both effectiveness and efficiency, we develop a global-local structure for our MVT. Our experiments on two public benchmarks, ModelNet40 and ModelNet10, demonstrate the competitive performance of our MVT.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. NeuroVoxel-LM: Language-Aligned 3D Perception via Dynamic Voxelization and Meta-Embedding

    cs.CV 2025-07 conditional novelty 4.0 of 10

    NeuroVoxel-LM combines dynamic multi-resolution voxelization with attention-based pooling of NeRF weights, reporting faster 3D feature extraction and modestly better NeRF captioning than fixed-resolution and max-pooli...

Pith tools