REVIEW 15 cited by
Rethinking Network Design and Local Geometry in Point Cloud: A Simple Residual MLP Framework
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Rethinking Network Design and Local Geometry in Point Cloud: A Simple Residual MLP Framework
read the original abstract
Point cloud analysis is challenging due to irregularity and unordered data structure. To capture the 3D geometries, prior works mainly rely on exploring sophisticated local geometric extractors using convolution, graph, or attention mechanisms. These methods, however, incur unfavorable latency during inference, and the performance saturates over the past few years. In this paper, we present a novel perspective on this task. We notice that detailed local geometrical information probably is not the key to point cloud analysis -- we introduce a pure residual MLP network, called PointMLP, which integrates no sophisticated local geometrical extractors but still performs very competitively. Equipped with a proposed lightweight geometric affine module, PointMLP delivers the new state-of-the-art on multiple datasets. On the real-world ScanObjectNN dataset, our method even surpasses the prior best method by 3.3% accuracy. We emphasize that PointMLP achieves this strong performance without any sophisticated operations, hence leading to a superior inference speed. Compared to most recent CurveNet, PointMLP trains 2x faster, tests 7x faster, and is more accurate on ModelNet40 benchmark. We hope our PointMLP may help the community towards a better understanding of point cloud analysis. The code is available at https://github.com/ma-xu/pointMLP-pytorch.
Forward citations
Cited by 15 Pith papers
-
On the Generation and Mitigation of Harmful Geometry in Image-to-3D Models
Image-to-3D models successfully generate harmful geometries in most cases with under 0.3% caught by commercial filters; existing safeguards are weak but a stacked defense cuts harmful outputs to under 1% at 11% false-...
-
RelFlexformer: Efficient Attention 3D-Transformers for Integrable Relative Positional Encodings
RelFlexformers enable flexible integrable 3D RPE in attention via NU-FFT, generalizing prior methods to heterogeneous token positions with O(L log L) complexity.
-
DAP: Doppler-aware Point Network for Heterogeneous mmWave Action Recognition
Introduces the first heterogeneous multi-source mmWave point cloud HAR dataset and DAP-Net architecture with Doppler reparameterization and text alignment for cross-source robustness.
-
Detecting Dental Landmarks from Intraoral 3D Scans: the 3DTeethLand challenge
A public benchmark dataset and competition results for 3D dental landmark detection from intraoral scans, with the top team reaching 0.91 rank score using a stratified transformer and DBSCAN.
-
LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention
LLaMA-Adapter turns frozen LLaMA 7B into a capable instruction follower using only 1.2M new parameters and zero-init attention, matching Alpaca while extending to image-conditioned reasoning on ScienceQA and COCO.
-
DAP: Doppler-aware Point Network for Heterogeneous mmWave Action Recognition
Introduces the first heterogeneous multi-source mmWave point cloud HAR dataset and DAP-Net, which uses Doppler patterns for source-invariant action recognition and outperforms prior methods.
-
Beyond Defenses: Manifold-Aligned Regularization for Intrinsic 3D Point Cloud Robustness
MAPR aligns latent and intrinsic geometries in 3D point cloud models via regularization on curvature and diffusion features plus consistency loss, yielding +20% average robustness gains on ModelNet40 without adversari...
-
Beyond Defenses: Manifold-Aligned Regularization for Intrinsic 3D Point Cloud Robustness
MAPR improves adversarial robustness in 3D point cloud networks by aligning latent predictions with intrinsic manifold geometry via curvature/diffusion features and a consistency loss.
-
Delaunay Canopy: Building Wireframe Reconstruction from Airborne LiDAR Point Clouds via Delaunay Graph
Delaunay Canopy uses Delaunay graphs as a geometric prior with region-wise curvature scoring to reconstruct accurate building wireframes from sparse and noisy airborne LiDAR point clouds.
-
CLAMP: Contrastive Learning for 3D Multi-View Action-Conditioned Robotic Manipulation Pretraining
CLAMP pretrains 3D multi-view encoders with contrastive learning on point clouds and actions, then initializes diffusion policies for more sample-efficient fine-tuning on robotic tasks.
-
$\text{VG}^2$GT: Voxel-Gaussian Splatting Visual Geometry Grounded Transformer
VG²GT regresses Gaussian primitive parameters from multi-scale voxel features of a frozen VFM and uses stochastic solid volume rendering for depth supervision to produce geometrically accurate reconstructions that out...
-
Full-field prediction for engineering-scale three-dimensional aircraft with multigrid-hierarchical learning
MHLF combines multigrid geometry representation with hierarchical learning to predict full flow fields for engineering-scale 3D aircraft, accelerating CFD convergence 3-8x across subsonic to supersonic regimes without...
-
A Camera-Cooperative ISAC Framework for Multimodal Non-Cooperative UAVs Sensing
The CC-ISAC framework aligns camera visuals with radio echoes via cross-attention and fuses multimodal data to reduce beam steering overhead by 71% and tracking overhead by 1.69-11.15% on the DeepSense 6G dataset whil...
-
Heterogeneous and Adept Snapshot Distillation for 3D Semantic Segmentation
HAS-KD combines information-oriented heterogeneous distillation from multi-modal models with adept snapshot distillation from training checkpoints to reach SOTA 3D semantic segmentation on ScanNetV2 and S3DIS without ...
-
Scene Reconstruction as Mapping Priors for 3D Detection
Automatically constructed mapping priors from sensor aggregation are integrated via the MPA3D framework to achieve state-of-the-art 3D detection results on the Waymo Open Dataset.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.