Pith. sign in

REVIEW 3 major objections 6 minor 107 references

Improving the Transferability of 3D Point Cloud Attack via Spectral-aware Admix and Optimization Designs

T0 review · 3 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read Mixing point clouds in spectral space makes adversarial attacks transfer across 3D classifiers.

desk verdict Promising transfer-attack idea with large reported gains, but the spectral Admix is under-specified: it only works if all mixed clouds share one eigenbasis, which the paper never states. read the letter →

arxiv 2412.12626 v1 pith:KY3WXTKD submitted 2024-12-17 cs.CV cs.CR

classification cs.CVcs.CR
keywords 3Dpointcloudattackadversarialtransferabilityblack-boxGraphFourierTransformspectraldomainAdmixclassificationModelNet40
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper targets a practical weakness of 3D adversarial attacks: examples that fool one point-cloud classifier usually do not fool a second one, so black-box attacks have low success. It proposes SAAO, a transfer-based attack that mixes the point cloud being attacked with other-category point clouds in the graph spectral domain instead of in coordinates, then optimizes the adversarial spectral feature along selected mixing paths. The paper reports that on ModelNet40 this raises black-box transfer attack success by roughly 8-22 percentage points over prior methods across PointNet, PointNet++, PointConv, and DGCNN, with lower or comparable perturbation distances and unchanged 100% white-box success. If that is right, adversarial examples for 3D perception can be made to work against models the attacker never saw.

What carries the argument

The central object is the Graph Fourier Transform of a point cloud: the projection $\tilde P = Q^T P$ of the coordinate matrix onto the eigenbasis $Q$ of the graph Laplacian of a K-NN graph. This projection converts an unordered point set into an ordered spectral vector, which is what makes linear mixing of two point clouds meaningful. The argument is carried by two mixing weights: a learnable diagonal positive matrix $M$ that approximates a Mahalanobis distance to guide the adversarial sample toward class boundaries, and a fixed spectral mask $M_s$ that keeps the top-32 low-frequency components close to the original to preserve geometry and imperceptibility. A path-selection step ranks candidate mixing point clouds by the cosine similarity between the adversarial gradient and each candidate's averaged gradient. Inverse GFT returns the final adversarial sample to coordinate space, so the entire optimization happens in spectral space while the delivered perturbation is spatial.

What would settle it

One concrete check: run the attack on a fixed set of point-cloud pairs twice, first mixing all clouds in the eigenbasis of the target point cloud (as the equations imply), then transforming each cloud by its own graph Laplacian eigenbasis before mixing and reporting what happens. If the shared-basis run does not beat the independent-basis run, or if the independent-basis run breaks down, then the spectral alignment the method silently depends on is not responsible for the reported gains.

Watch

Extended reading notes

Core claim

The paper's central claim is that Admix-style input mixing, transferred from 2D images to point clouds, becomes effective only when the mixing is done in the graph spectral domain. For a target point cloud $P$, the method builds a K-NN graph, takes the graph Laplacian eigenvector matrix $Q$ as the Graph Fourier Transform basis, and represents each point cloud as spectral features $\tilde P = Q^T P$. It then mixes the adversarial spectral feature with those of point clouds from other categories, using a learnable diagonal weight matrix $M$ (initialized from inverse variance as a stable stand-in for a Mahalanobis distance) and a fixed spectral mask $M_s$ that preserves the low-frequency shape components. A warm-up phase computes cosine similarity between the adversarial gradient and the averaged gradient for each candidate mix, selects the best augmentation paths, and the main optimization runs along those paths. The adversarial spectral feature is mapped back to a point cloud with $P^{\mathrm{adv}} = Q \tilde P^{\mathrm{adv}}$. The paper claims that on ModelNet40 this method achieves the best transfer attack success among the methods compared, with the lowest or near-lowest perturbation distances, while keeping white-box success at 100%.

Load-bearing premise

The method assumes that every point cloud used for mixing can be projected onto the graph Laplacian eigenbasis of the original target point cloud, so that their spectral features line up componentwise; the paper never states this shared-basis condition.

Editorial extensions

If this is right

  • Transfer-based black-box attacks on point clouds become more effective: SAAO reports black-box success-rate gains of roughly 8-22 percentage points over the compared methods, depending on the target classifier.
  • Adversarial point clouds generated this way have lower or comparable Hausdorff, Chamfer, and MSE distances than the baselines, so the attack is harder to spot by shape distortion.
  • White-box attack success stays at 100%, so the transferability gains do not come at the cost of direct attack strength.
  • The method keeps a large advantage over several defenses (SRS, SOR, DUP-Net) but its transfer success is cut by about 35-50 points under the IF-Defense variants, which shows the strongest defenses can still blunt it.
  • The recipe transfers across four different point-cloud classifiers in both directions, meaning the surrogate and victim models do not need to share architecture.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper's equations implicitly assume all mixed point clouds are projected onto the eigenbasis $Q$ of the original target point cloud; if each cloud used its own graph Laplacian eigenbasis, the spectral additions and distances would live in different coordinate systems and the method would be ill-defined. That shared-basis premise is never stated and should be checked in the implementation.
  • Because the fixed spectral mask relies on energy concentrating in the top-32 components, the cutoff is likely dataset-dependent; one testable extension is to make the cutoff adaptive per object or per class.
  • The spectral-mixing idea could be carried to other unordered geometric representations, such as LiDAR sweeps or graph-based 3D scene data, wherever a graph Laplacian can be constructed.
  • Reported numbers already show the strongest defense (IF-Defense) sharply reduces transfer success, so a natural follow-up is to combine SAAO with optimization designed specifically to resist shape-restoration defenses.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. This paper proposes SAAO, a transfer-based black-box attack against 3D point cloud classifiers. Instead of mixing point clouds by coordinate-level addition, the method transforms point clouds into the graph spectral domain via the Graph Fourier Transform, performs a weighted Admix update on spectral features, selects augmentation paths by gradient cosine similarity, and returns to the data domain by inverse GFT. The experimental section reports lower Chamfer/Hausdorff/L2 perturbation distances than several baselines and substantially higher transfer attack success rates than FGSM, 3D-Adv, GeoA, AdvPC, and SS-Attack on ModelNet40, with additional results under six defenses.

Significance. If confirmed, the reported transfer improvements are substantial: for example, from a PointNet surrogate to a DGCNN target, transfer ASR rises from 30.0% for SS-Attack to 64.3% for Ours-F, and from a PointNet++ surrogate to PointConv the rate rises from 46.1% to 71.6%. The idea of realizing Admix in the graph spectral domain is a reasonable and potentially useful adaptation to unordered point sets, and the evaluation covers four surrogate/target models and six defenses. The paper does not release code or report exact evaluation sizes, and the central spectral-mixing operation is not fully specified, so the current manuscript is more a promising empirical proposal than an independently verifiable result; the contribution is empirical rather than theoretical, with several design choices inherited from prior spectral-domain attack papers.

major comments (3)
  1. [§3.3, Eqs. (2)-(4), Algorithm 1 line 7] The central Admix operation is well-defined only if all spectral features are coefficients in one common orthonormal basis. Section 3.3 defines the GFT basis Q from the graph Laplacian of the point cloud P being attacked, giving φ(P)=Q^T P. Equations (3) and (5) and Algorithm 1 line 7 then add or mix φ of the current adversarial cloud with φ of other randomly selected point clouds, but the paper never states that those other clouds are projected onto the same Q. Since point clouds are unordered and the selected clouds are different objects, a shared Q alone is not enough: applying Q^T to a different cloud P' requires a node correspondence between the rows of P' and the graph nodes of P, which is not provided. Without a common basis and a row correspondence, the mixed vector in Eq. (3) is not the spectrum of any point cloud in a fixed basis, and the IGFT in Eq. (4) has no well-defined geometric meaning. The paper also does not state whether Q is recomputed during optimization, although Eq. (4) requires a fixed original Q. Please state the shared-basis and correspondence assumption explicitly, or reformulate the update so that it is well-posed; this is load-bearing because all transfer results depend on it.
  2. [§3.5 and §3.7, Eq. (8) vs. Algorithm 1 line 7] The mixing formulas in the paper are mutually inconsistent. Eq. (8) forms γ_i M_s P̂_adv + γ_i η_i(I-M_s)P̂', while Algorithm 1 line 7 forms β_i M_s P̂_adv + (1-β_i)(I-M_s) f(P,P_j;M) eP_j, and Eq. (3) has neither M_s nor f. The parameter β_i in the algorithm is initialized from a lower/upper bound but is not related to γ_i and η_i in the equations, and f(P,P_j;M) is a scalar that does not appear in Eq. (8). Since Algorithm 1 is presented as the implemented optimization, the results in Tables 1-3 cannot be traced to a single unambiguous update rule. Please unify Eq. (3), Eq. (8), and Algorithm 1, clarify the roles of f, M, and M_s, and state which formula produced the experimental numbers.
  3. [§4.1-§4.2] The evaluation protocol is underspecified in ways that directly affect the headline transferability numbers. Section 4.1 says the authors 'randomly select a number of instances' from the ModelNet40 testing set, but it never gives the number of test examples, the selection criterion beyond 'well classified', or whether the same examples were used for all attacks. Section 4.2 gives only a total of 500 iterations and does not report the values of k and k' used in Algorithm 2, the number of candidate paths n', the number of path-selection steps, or the batch size b. Without these values, the comparisons in Tables 2 and 3 cannot be reproduced, and the absence of variance or confidence information is hard to interpret. Please report the exact protocol, including sample count, random seed or sample indices, and all path-selection and optimization hyperparameters.
minor comments (6)
  1. [§1 and §3.2] There are several typographical errors, including 'Pitcure 1' in §3.2, 'Specficially' in §3.4, 'diagnoal' and 'caculate' in §3.3-§3.4, and 'spectral-awared' in the title of Algorithm 1; these should be corrected.
  2. [§3.3-§3.7] The symbols P̂ and eP are used inconsistently for spectral-domain quantities; the paper should define the spectral representation once and use it consistently in Eqs. (2)-(8) and Algorithm 1.
  3. [§3.7, Algorithm 1] Algorithm 1 outputs P_adv_k, but line 12 updates only the spectral feature; the final IGFT mapping from Eq. (4) is missing from the pseudocode and should be added.
  4. [§3.4, Eq. (6)] The initialization M_0 = Diag(1/(Var+ε)) does not define which variable the variance is computed over (coordinates, spectral channels, or batch elements); this should be specified.
  5. [§4.3, Table 3] The caption 'on PointNet model' is ambiguous because the table has four model columns; it should state explicitly that PointNet is the surrogate model and the columns are the target models under each defense.
  6. [§4.3] The sentence 'one failure in DGCNN also implies that it is a harder classification model to attack' is unclear; if DGCNN transfer results are one of the cases, the claim should be stated more precisely and tied to the data.

Circularity Check

0 steps flagged · score 2.0 of 10

No significant circularity: the transferability claims are empirical and out-of-sample; the main self-citation is a minor design support, not a load-bearing derivation.

full rationale

The paper's central claims are empirical rather than derived, so the circularity tests largely do not apply. The headline transferability results in Table 2 are measured on held-out classifiers (PointNet, PointNet++, PointConv, DGCNN) against standard baselines, and no fitted parameter is relabeled as a prediction. The optimization procedure in Eq. (3), Eq. (8), and Algorithm 1 is a constructive method, not a derivation that assumes its own conclusion. The only notable self-citation is in Section 3.5, where the low/high-frequency weight split is justified by the authors' own prior spectral attack paper [15] via the claim that 'the lower components of the spectral feature represents the shape and the higher components represents the details' and that energy concentrates on the top-32 dimensions. This is a minor self-citation supporting a design heuristic, but it is not circular: [15] is an independent empirical observation about spectral energy concentration, not a statement of SAAO's own transferability, and the paper's reported attack success rates are computed from actual classifier outputs rather than from that citation. The skeptic's concern that Admix mixes spectral coefficients expressed in different eigenbases is a genuine correctness-and-reproducibility gap, because Eq. (3), Eq. (5), and Algorithm 1 require a shared basis Q, but that gap is an ill-posedness issue rather than a circular reduction of a claimed result to its input. Overall, no predicted quantity reduces by construction to an input or to a self-citation chain, so the circularity score is low.

Assumptions & free parameters 16 free parameters · 6 assumptions · 0 invented entities

The method introduces many hand-set hyperparameters and relies on several domain assumptions about graph spectral representations. The implicit basis-sharing assumption is the most load-bearing undocumented premise.

free parameters (16)
  • alpha_low = 0.9
    Spectral mask weight for low-frequency components in Eq. (7), chosen by hand to preserve geometry; not derived.
  • alpha_high = 0.25
    Spectral mask weight for high-frequency components in Eq. (7), chosen by hand to limit high-frequency noise.
  • top_k_spectral = 32
    Dimension cutoff for the low-frequency mask; justified by an unstated claim that spectral energy concentrates in top-32 dimensions, from prior work without evidence in this paper.
  • lambda_1 = 0.5
    Weight for MSE loss in Eq. (12), set without sensitivity analysis.
  • lambda_2 = 20
    Weight for Chamfer distance in Eq. (12).
  • lambda_3 = 50
    Weight for Hausdorff distance in Eq. (12).
  • admix_lower_bound = 0.1
    Lower bound b_l for Admix interpolation in Algorithm 1.
  • admix_upper_bound = 0.9
    Upper bound b_u for Admix interpolation in Algorithm 1.
  • admix_m = 20
    Number of scaled copies in Admix, Eq. (3).
  • admix_n = 9
    Number of mixing point clouds, Section 4.2.
  • knn_k = 10
    Number of neighbors for K-NN graph construction, Section 4.2.
  • learning_rate = 0.01
    Adam learning rate for perturbation and M updates.
  • momentum = 0.9
    Adam momentum.
  • iterations = 500
    Total optimization iterations.
  • epsilon = small positive
    Added to variance in Eq. (6) to avoid division by zero.
  • M0_init = Diag(1/(Var+epsilon))
    Initial diagonal weight matrix for Mahalanobis-like distance, learned during attack; initialization is hand-chosen.
assumptions (6)
  • domain assumption The graph Laplacian eigen-decomposition provides an ordered, geometry-aware representation of point clouds suitable for linear mixing.
    Invoked in Section 3.3 to justify spectral-domain Admix; relies on graph signal processing literature [20,54] and the authors' prior spectral attack works.
  • domain assumption Spectral energy of point clouds concentrates in the top-32 low-frequency components.
    Used in Section 3.5 to set the mask M_s; stated without evidence in this paper, presumably from the authors' prior work [15].
  • ad hoc to paper The GFT basis Q of the original point cloud is used to compute spectral features of all mixing point clouds.
    Implicit in Eq. (3) and Eq. (5); without a shared basis, component-wise spectral mixing and distance are ill-defined. Never stated explicitly.
  • domain assumption Cosine similarity between the adversarial gradient and the mixed-sample gradient is a valid criterion for selecting augmentation paths.
    Adopted from GRA [101] in Section 3.6; no comparison to alternative selection metrics.
  • domain assumption Surrogate and target classifiers share similar decision boundaries, enabling transferability.
    Standard assumption in transfer attacks, stated in Section 3.1.
  • standard math The adversarial loss and distance losses are differentiable with respect to the spectral perturbation and weight matrix.
    Required for gradient descent updates in Algorithm 1; holds for the chosen losses assuming classifier differentiability.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Improving the Transferability of 3D Point Cloud Attack via Spectral-aware Admix and Optimization Designs." pith.science (2026). https://pith.science/paper/KY3WXTKD

@misc{pith2026241212626,
  author       = {Pith},
  title        = {Pith review of: Improving the Transferability of 3D Point Cloud Attack via Spectral-aware Admix and Optimization Designs},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/KY3WXTKD}},
  note         = {Machine review of arXiv:2412.12626}
}
read the original abstract

Deep learning models for point clouds have shown to be vulnerable to adversarial attacks, which have received increasing attention in various safety-critical applications such as autonomous driving, robotics, and surveillance. Existing 3D attackers generally design various attack strategies in the white-box setting, requiring the prior knowledge of 3D model details. However, real-world 3D applications are in the black-box setting, where we can only acquire the outputs of the target classifier. Although few recent works try to explore the black-box attack, they still achieve limited attack success rates (ASR). To alleviate this issue, this paper focuses on attacking the 3D models in a transfer-based black-box setting, where we first carefully design adversarial examples in a white-box surrogate model and then transfer them to attack other black-box victim models. Specifically, we propose a novel Spectral-aware Admix with Augmented Optimization method (SAAO) to improve the adversarial transferability. In particular, since traditional Admix strategy are deployed in the 2D domain that adds pixel-wise images for perturbing, we can not directly follow it to merge point clouds in coordinate domain as it will destroy the geometric shapes. Therefore, we design spectral-aware fusion that performs Graph Fourier Transform (GFT) to get spectral features of the point clouds and add them in the spectral domain. Afterward, we run a few steps with spectral-aware weighted Admix to select better optimization paths as well as to adjust corresponding learning weights. At last, we run more steps to generate adversarial spectral feature along the optimization path and perform Inverse-GFT on the adversarial spectral feature to obtain the adversarial example in the data domain. Experiments show that our SAAO achieves better transferability compared to existing 3D attack methods.

Figures

Figures reproduced from arXiv: 2412.12626 by the authors.

Figure 1
Figure 1. Transferbility is the ability of a attack method us [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Our Attack baseline using weighted Admix and path selection. [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

107 extracted references · 55 canonical work pages

  1. [1]

    Matan Atzmon, Haggai Maron, and Yaron Lipman. 2018. Point convolutional neural networks by extension operators. arXiv preprint arXiv:1803.10091 (2018)

  2. [2]

    Xiaowen Cai, Yunbo Tao, Daizong Liu, Pan Zhou, Xiaoye Qu, Jianfeng Dong, Keke Tang, and Lichao Sun. 2024. Frequency-Aware GAN for Imperceptible Transfer Attack on 3D Point Clouds. InProceedings of the 32nd ACM International Conference on Multimedia. 6162–6171

  3. [3]

    Nicholas Carlini and David Wagner. 2017. Towards evaluating the robustness of neural networks. In 2017 IEEE Symposium on Security and Privacy (SP) . 39–57

  4. [4]

    Xiaozhi Chen, Huimin Ma, Ji Wan, Bo Li, and Tian Xia. 2017. Multi-view 3d object detection network for autonomous driving. In Proceedings of the IEEE conference on Computer Vision and Pattern Recognition (CVPR) . 1907–1915

  5. [5]

    Ricardo L De Queiroz and Philip A Chou. 2017. Transform coding for point clouds using a Gaussian process model. IEEE Transactions on Image Processing 26, 7 (2017), 3507–3517

  6. [6]

    Yinpeng Dong, Fangzhou Liao, Tianyu Pang, Hang Su, Jun Zhu, Xiaolin Hu, and Jianguo Li. 2018. Boosting adversarial attacks with momentum. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) . 9185– 9193

  7. [7]

    Yueqi Duan, Yu Zheng, Jiwen Lu, Jie Zhou, and Qi Tian. 2019. Structural relational reasoning of point clouds. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) . 949–958

  8. [8]

    Xiang Fang, Arvind Easwaran, Blaise Genest, and Ponnuthurai Nagaratnam Suganthan. 2024. Your Data Is Not Perfect: Towards Cross-Domain Out-of- Distribution Detection in Class-Imbalanced Data. Expert Systems With Applica- tions (2024)

Show all 107 references
  1. [9]

    Xiang Fang, Daizong Liu, Pan Zhou, and Guoshun Nan. 2023. You can ground earlier than see: An effective and efficient pipeline for temporal sentence ground- ing in compressed videos. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 2448–2460

  2. [10]

    Hao Feng, Qi Liu, Hao Liu, Wengang Zhou, Houqiang Li, and Can Huang. 2023. Docpedia: Unleashing the power of large multimodal model in the frequency domain for versatile document understanding. arXiv preprint arXiv:2311.11810 (2023)

  3. [11]

    Hao Feng, Zijian Wang, Jingqun Tang, Jinghui Lu, Wengang Zhou, Houqiang Li, and Can Huang. 2023. Unidoc: A universal large multimodal model for simultaneous text detection, recognition, spotting and understanding. arXiv preprint arXiv:2308.11592 (2023)

  4. [12]

    Xiang Gao, Wei Hu, and Guo-Jun Qi. 2020. GraphTER: Unsupervised learning of graph transformation equivariant representations via auto-encoding node-wise transformations. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 7163–7172

  5. [13]

    Ian J Goodfellow, Jonathon Shlens, and Christian Szegedy. 2014. Explaining and harnessing adversarial examples. arXiv preprint arXiv:1412.6572 (2014)

  6. [14]

    Abdullah Hamdi, Sara Rojas, Ali Thabet, and Bernard Ghanem. 2020. Advpc: Transferable adversarial perturbations on 3d point clouds. In European Confer- ence on Computer Vision (ECCV) . 241–257

  7. [15]

    Qianjiang Hu, Daizong Liu, and Wei Hu. 2022. Exploring the Devil in Graph Spectral Domain for 3D Point Cloud Attacks. arXiv:2202.07261 [cs.CV]

  8. [16]

    Wei Hu, Jiahao Pang, Xianming Liu, Dong Tian, Chia-Wen Lin, and Anthony Vetro. 2021. Graph Signal Processing for Geometric Data and Beyond: Theory and Applications. IEEE Transactions on Multimedia (TMM) (2021)

  9. [17]

    Qidong Huang, Xiaoyi Dong, Dongdong Chen, Hang Zhou, Weiming Zhang, and Nenghai Yu. 2022. Shape-invariant 3D Adversarial Point Clouds. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 15335– 15344

  10. [18]

    Wencan Huang, Daizong Liu, and Wei Hu. 2023. Dense object grounding in 3d scenes. In Proceedings of the 31st ACM International Conference on Multimedia . 5017–5026. Improving the Transferability of 3D Point Cloud Attack via Spectral-aware Admix and Optimization Designs xxxx, 2024,

  11. [19]

    Wencan Huang, Daizong Liu, and Wei Hu. 2024. Advancing 3d object grounding beyond a single 3d scene. InProceedings of the 32nd ACM International Conference on Multimedia. 7995–8004

  12. [20]

    Kipf and Max Welling

    Thomas N. Kipf and Max Welling. 2017. Semi-Supervised Classification with Graph Convolutional Networks. arXiv:1609.02907 [cs.LG]

  13. [21]

    Alexey Kurakin, Ian Goodfellow, and Samy Bengio. 2016. Adversarial machine learning at scale. arXiv preprint arXiv:1611.01236 (2016)

  14. [22]

    Kibok Lee, Zhuoyuan Chen, Xinchen Yan, Raquel Urtasun, and Ersin Yumer

  15. [23]

    Xinke Li, Zhiru Chen, Yue Zhao, Zekun Tong, Yabang Zhao, Andrew Lim, and Joey Tianyi Zhou. 2021. PointBA: Towards Backdoor Attacks in 3D Point Cloud. In Proceedings of the IEEE International Conference on Computer Vision (ICCV)

  16. [24]

    Daizong Liu, Xiang Fang, Xiaoye Qu, Jianfeng Dong, He Yan, Yang Yang, Pan Zhou, and Yu Cheng. 2024. Unsupervised Domain Adaptative Temporal Sentence Localization with Mutual Information Maximization. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 38. 3567–3575

  17. [25]

    Daizong Liu, Xiang Fang, Pan Zhou, Xing Di, Weining Lu, and Yu Cheng. 2023. Hypotheses tree building for one-shot temporal sentence localization. In Pro- ceedings of the AAAI Conference on Artificial Intelligence , Vol. 37. 1640–1648

  18. [26]

    Daizong Liu and Wei Hu. 2022. Imperceptible transfer attack and defense on 3d point cloud classification. IEEE Transactions on Pattern Analysis and Machine Intelligence (2022)

  19. [27]

    Daizong Liu and Wei Hu. 2024. Explicitly Perceiving and Preserving the Local Geometric Structures for 3D Point Cloud Attack. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 38. 3576–3584

  20. [28]

    Daizong Liu, Wei Hu, and Xin Li. 2023. Point cloud attacks in graph spectral domain: When 3d geometry meets graph signal processing. IEEE Transactions on Pattern Analysis and Machine Intelligence (2023)

  21. [29]

    Daizong Liu, Wei Hu, and Xin Li. 2023. Robust geometry-dependent attack for 3D point clouds. IEEE Transactions on Multimedia (2023)

  22. [30]

    Daizong Liu, Yang Liu, Wencan Huang, and Wei Hu. 2024. A Survey on Text- guided 3D Visual Grounding: Elements, Recent Advances, and Future Directions. arXiv preprint arXiv:2406.05785 (2024)

  23. [31]

    Daizong Liu, Xi Ouyang, Shuangjie Xu, Pan Zhou, Kun He, and Shiping Wen

  24. [32]

    Daizong Liu, Xiaoye Qu, Jianfeng Dong, Guoshun Nan, Pan Zhou, Zichuan Xu, Lixing Chen, He Yan, and Yu Cheng. 2023. Filling the Information Gap between Video and Query for Language-Driven Moment Retrieval. In Proceedings of the 31st ACM International Conference on Multimedia . ...

  25. [33]

    Neurocomputing 413 (2020), 145–157

    SAANet: Siamese action-units attention network for improving dynamic facial expression recognition. Neurocomputing 413 (2020), 145–157

  26. [34]

    Daizong Liu, Yunbo Tao, Pan Zhou, and Wei Hu. 2024. Hard-Label Black-Box Attacks on 3D Point Clouds. arXiv preprint arXiv:2412.00404 (2024)

  27. [35]

    Daizong Liu, Xiaoye Qu, Jianfeng Dong, Pan Zhou, Zichuan Xu, Haozhao Wang, Xing Di, Weining Lu, and Yu Cheng. 2024. Transform-Equivariant Consistency Learning for Temporal Sentence Grounding. ACM Transactions on Multimedia Computing, Communications and Applications 20, 4 (2024), 1–19

  28. [36]

    Daizong Liu, Mingyu Yang, Xiaoye Qu, Pan Zhou, Yu Cheng, and Wei Hu. 2024. A survey of attacks on large vision-language models: Resources, advances, and future trends. arXiv preprint arXiv:2407.07403 (2024)

  29. [37]

    Daizong Liu, Shuangjie Xu, Xiao-Yang Liu, Zichuan Xu, Wei Wei, and Pan Zhou

  30. [38]

    Daniel Liu, Ronald Yu, and Hao Su. 2019. Extending adversarial attacks and defenses to deep 3d point cloud classifiers. In 2019 IEEE International Conference on Image Processing (ICIP) . 2279–2283

  31. [39]

    Daizong Liu, Pan Zhou, Zichuan Xu, Haozhao Wang, and Ruixuan Li. 2022. Few-shot temporal sentence grounding via memory-guided semantic learning. IEEE Transactions on Circuits and Systems for Video Technology 33, 5 (2022), 2491–2505

  32. [40]

    Daizong Liu, Mingyu Yang, Xiaoye Qu, Pan Zhou, Xiang Fang, Keke Tang, Yao Wan, and Lichao Sun. 2024. Pandora’s Box: Towards Building Universal Attackers against Real-World Large Vision-Language Models. In The Thirty- eighth Annual Conference on Neural Information Processing Systems

  33. [41]

    Yongcheng Liu, Bin Fan, Gaofeng Meng, Jiwen Lu, Shiming Xiang, and Chun- hong Pan. 2019. Densepoint: Learning densely contextual representation for efficient point cloud processing. In Proceedings of the IEEE International Confer- ence on Computer Vision (ICCV) . 5239–5248

  34. [42]

    Yongcheng Liu, Bin Fan, Shiming Xiang, and Chunhong Pan. 2019. Relation- shape convolutional neural network for point cloud analysis. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) . 8895– 8904

  35. [43]

    Daizong Liu, Jiahao Zhu, Xiang Fang, Zeyu Xiong, Huan Wang, Renfu Li, and Pan Zhou. 2023. Conditional Video Diffusion Network for Fine-grained Temporal Sentence Grounding. IEEE Transactions on Multimedia (2023)

  36. [44]

    Yang Liu, Daizong Liu, and Wei Hu. 2024. Joint Top-Down and Bottom-Up Frameworks for 3D Visual Grounding. arXiv preprint arXiv:2410.15615 (2024)

  37. [45]

    Yuliang Liu, Jiaxin Zhang, Dezhi Peng, Mingxin Huang, Xinyu Wang, Jingqun Tang, Can Huang, Dahua Lin, Chunhua Shen, Xiang Bai, et al. 2023. Spts v2: single-point scene text spotting. IEEE Transactions on Pattern Analysis and Machine Intelligence (2023)

  38. [46]

    Yang Liu, Daizong Liu, Zongming Guo, and Wei Hu. 2024. Cross-task knowledge transfer for semi-supervised joint 3d grounding and captioning. In Proceedings of the 32nd ACM International Conference on Multimedia . 3818–3827

  39. [47]

    Chao Ma, Yulan Guo, Jungang Yang, and Wei An. 2018. Learning multi-view representation with LSTM for 3-D shape recognition and retrieval. IEEE Trans- actions on Multimedia (TMM) 21, 5 (2018), 1169–1182

  40. [48]

    Chengcheng Ma, Weiliang Meng, Baoyuan Wu, Shibiao Xu, and Xiaopeng Zhang. 2020. Efficient joint gradient based attack against sor defense for 3d point cloud classification. InProceedings of the 28th ACM International Conference on Multimedia. 1819–1827

  41. [49]

    Yuyang Long, Qilong Zhang, Boheng Zeng, Lianli Gao, Xianglong Liu, Jian Zhang, and Jingkuan Song. 2022. Frequency Domain Model Augmentation for Adversarial Attack. arXiv:2207.05382 [cs.CV]

  42. [50]

    Prasanta Chandra Mahalanobis. 1936. On the generalized distance in statistics. https://api.semanticscholar.org/CorpusID:117765088

  43. [51]

    Charles R Qi, Hao Su, Kaichun Mo, and Leonidas J Guibas. 2017. Pointnet: Deep learning on point sets for 3d classification and segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 652–660

  44. [52]

    Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu. 2017. Towards deep learning models resistant to adversarial attacks. arXiv preprint arXiv:1706.06083 (2017)

  45. [53]

    Shi Qiu, Saeed Anwar, and Nick Barnes. 2021. Geometric back-projection network for point cloud classification. IEEE Transactions on Multimedia (TMM) 24 (2021), 1943–1955

  46. [54]

    Aliaksei Sandryhaila and José M. F. Moura. 2013. Discrete signal processing on graphs: Graph fourier transform. In 2013 IEEE International Conference on Acoustics, Speech and Signal Processing . 6167–6170. https://doi.org/10.1109/ ICASSP.2013.6638850

  47. [55]

    Charles R Qi, Li Yi, Hao Su, and Leonidas J Guibas. 2017. Pointnet++: Deep hierarchical feature learning on point sets in a metric space. Advances in Neural Information Processing Systems (NIPS) (2017)

  48. [56]

    Yiru Shen, Chen Feng, Yaoqing Yang, and Dong Tian. 2018. Mining point cloud local structures by kernel correlation and graph pooling. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) . 4548–4557

  49. [57]

    Martin Simonovsky and Nikos Komodakis. 2017. Dynamic edge-conditioned filters in convolutional neural networks on graphs. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) . 3693–3702

  50. [58]

    Yiting Shao, Zhaobin Zhang, Zhu Li, Kui Fan, and Ge Li. 2017. Attribute com- pression of 3D point clouds using Laplacian sparsity optimized graph transform. In 2017 IEEE Visual Communications and Image Processing (VCIP) . IEEE, 1–4

  51. [59]

    Hang Su, Subhransu Maji, Evangelos Kalogerakis, and Erik Learned-Miller

  52. [60]

    Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow, and Rob Fergus. 2013. Intriguing properties of neural networks. arXiv preprint arXiv:1312.6199 (2013)

  53. [61]

    Satya P Singh, Lipo Wang, Sukrit Gupta, Haveesh Goli, Parasuraman Padman- abhan, and Balázs Gulyás. 2020. 3D deep learning on medical images: a review. Sensors 20, 18 (2020), 5097

  54. [62]

    Jingqun Tang, Chunhui Lin, Zhen Zhao, Shu Wei, Binghong Wu, Qi Liu, Hao Feng, Yang Li, Siqi Wang, Lei Liao, et al . 2024. TextSquare: Scaling up Text- Centric Visual Instruction Tuning. arXiv preprint arXiv:2404.12803 (2024)

  55. [63]

    Jingqun Tang, Wenming Qian, Luchuan Song, Xiena Dong, Lan Li, and Xiang Bai

  56. [64]

    Jingqun Tang, Su Qiao, Benlei Cui, Yuhang Ma, Sheng Zhang, and Dimitrios Kanoulas. 2022. You can even annotate text with voice: Transcription-only- supervised text spotting. In Proceedings of the 30th ACM International Conference on Multimedia. 4154–4163

  57. [65]

    Jingqun Tang, Weidong Du, Bin Wang, Wenyang Zhou, Shuqi Mei, Tao Xue, Xing Xu, and Hai Zhang. 2023. Character recognition competition for street view shop signs. National Science Review 10, 6 (2023), nwad141

  58. [66]

    Yunbo Tao, Daizong Liu, Pan Zhou, Yulai Xie, Wei Du, and Wei Hu. 2023. 3DHacker: Spectrum-based decision boundary generation for hard-label 3D point cloud attack. In Proceedings of the IEEE/CVF International Conference on Computer Vision. 14340–14350

  59. [67]

    Gusi Te, Wei Hu, Amin Zheng, and Zongming Guo. 2018. Rgcnn: Regular- ized graph cnn for point cloud segmentation. In Proceedings of the 26th ACM xxxx, 2024, Shiyu Hu, Daizong Liu, Wei Hu international conference on Multimedia . 746–754

  60. [68]

    Hugues Thomas, Charles R Qi, Jean-Emmanuel Deschaud, Beatriz Marcotegui, François Goulette, and Leonidas J Guibas. 2019. Kpconv: Flexible and deformable convolution for point clouds. In Proceedings of the IEEE International Conference on Computer Vision (ICCV) . 6411–6420

  61. [69]

    Tzungyu Tsai, Kaichen Yang, Tsung-Yi Ho, and Yier Jin. 2020. Robust adversarial objects against deep learning models. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 34. 954–962

  62. [70]

    Jingqun Tang, Wenqing Zhang, Hongye Liu, MingKun Yang, Bo Jiang, Guang- long Hu, and Xiang Bai. 2022. Few could be better than all: Feature sampling and grouping for scene text detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 4563–4572

  63. [71]

    An-Lan Wang, Bin Shan, Wei Shi, Kun-Yu Lin, Xiang Fei, Guozhi Tang, Lei Liao, Jingqun Tang, Can Huang, and Wei-Shi Zheng. 2024. ParGo: Bridging Vision-Language with Partial and Global Views. arXiv:2408.12928 [cs.CV] https://arxiv.org/abs/2408.12928

  64. [72]

    Xiaosen Wang, Xuanran He, Jingdong Wang, and Kun He. 2021. Admix: En- hancing the Transferability of Adversarial Attacks. arXiv:2102.00436 [cs.CV]

  65. [73]

    Yue Wang, Yongbin Sun, Ziwei Liu, Sanjay E Sarma, Michael M Bronstein, and Justin M Solomon. 2019. Dynamic graph cnn for learning on point clouds. Acm Transactions On Graphics (TOG) 38, 5 (2019), 1–12

  66. [74]

    Yuxin Wen, Jiehong Lin, Ke Chen, CL Philip Chen, and Kui Jia. 2020. Geometry- aware generation of adversarial point clouds. IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI) (2020)

  67. [75]

    Chun-Chen Tu, Paishun Ting, Pin-Yu Chen, Sijia Liu, Huan Zhang, Jinfeng Yi, Cho-Jui Hsieh, and Shin-Ming Cheng. 2019. Autozoom: Autoencoder-based zeroth order optimization method for attacking black-box neural networks. In Proceedings of the AAAI Conference on Artificial Intel...

  68. [76]

    Matthew Wicker and Marta Kwiatkowska. 2019. Robustness of 3d deep learning in an adversarial setting. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) . 11767–11775

  69. [77]

    Wenxuan Wu, Zhongang Qi, and Li Fuxin. 2020. PointConv: Deep Convolutional Networks on 3D Point Clouds. arXiv:1811.07246 [cs.CV]

  70. [78]

    Ziyi Wu, Yueqi Duan, He Wang, Qingnan Fan, and Leonidas J Guibas. 2020. If-defense: 3d adversarial point cloud defense via implicit function based restora- tion. arXiv preprint arXiv:2010.05272 (2020)

  71. [79]

    Zhirong Wu, Shuran Song, Aditya Khosla, Fisher Yu, Linguang Zhang, Xiaoou Tang, and Jianxiong Xiao. 2015. 3d shapenets: A deep representation for vol- umetric shapes. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 1912–1920

  72. [80]

    Tingyu Weng, Jun Xiao, Feilong Yan, and Haiyong Jiang. 2022. Context-Aware 3D Point Cloud Semantic Segmentation With Plane Guidance.IEEE Transactions on Multimedia (TMM) (2022)

  73. [81]

    Zhen Xiang, David J Miller, Siheng Chen, Xi Li, and George Kesidis. 2021. A Backdoor Attack against 3D Point Cloud Classifiers. In Proceedings of the IEEE International Conference on Computer Vision (ICCV)

  74. [82]

    Qiangeng Xu, Xudong Sun, Cho-Ying Wu, Panqu Wang, and Ulrich Neumann

  75. [83]

    Bo Yang, Jianan Wang, Ronald Clark, Qingyong Hu, Sen Wang, Andrew Markham, and Niki Trigoni. 2019. Learning object bounding boxes for 3d instance segmentation on point clouds. arXiv preprint arXiv:1906.01140 (2019)

  76. [85]

    Chong Xiang, Charles R Qi, and Bo Li. 2019. Generating 3d adversarial point clouds. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 9136–9144

  77. [86]

    Kaichen Yang, Xuan-Yi Lin, Yixin Sun, Tsung-Yi Ho, and Yier Jin. 2021. 3D- Adv: Black-Box Adversarial Attacks against Deep Learning Models through 3D Sensors. In 2021 58th ACM/IEEE Design Automation Conference (DAC) . 547–552. https://doi.org/10.1109/DAC18074.2021.9586275

  78. [87]

    Mingyu Yang, Daizong Liu, Keke Tang, Pan Zhou, Lixing Chen, and Junyang Chen. 2025. Hiding Imperceptible Noise in Curvature-Aware Patches for 3D Point Cloud Attack. In European Conference on Computer Vision . Springer, 431– 448

  79. [88]

    In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)

    Grid-gcn for fast and scalable point cloud learning. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) . 5661–5670

  80. [89]

    Manzil Zaheer, Satwik Kottur, Siamak Ravanbakhsh, Barnabas Poczos, Russ R Salakhutdinov, and Alexander J Smola. 2017. Deep Sets. Advances in Neural Information Processing Systems (NIPS) 30 (2017)

  81. [90]

    Jinlai Zhang, Yinpeng Dong, Jun Zhu, Jihong Zhu, Minchi Kuang, and Xiaming Yuan. 2024. Improving transferability of 3D adversarial attacks with scale and shear transformations. Information Sciences 662 (2024), 120245

  82. [91]

    Jiancheng Yang, Qiang Zhang, Bingbing Ni, Linguo Li, Jinxian Liu, Mengdie Zhou, and Qi Tian. 2019. Modeling point clouds with self-attention and gumbel subset sampling. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 3323–3332

  83. [92]

    Qiang Zhang, Jiancheng Yang, Rongyao Fang, Bingbing Ni, Jinxian Liu, and Qi Tian. 2019. Adversarial attack and defense on point sets. arXiv preprint arXiv:1902.10899 (2019)

  84. [93]

    Yu Zhang, Gongbo Liang, Tawfiq Salem, and Nathan Jacobs. 2019. Defense- pointnet: Protecting pointnet against adversarial attacks. In 2019 IEEE Interna- tional Conference on Big Data (Big Data) . 5654–5660

  85. [94]

    Tan Yu, Jingjing Meng, and Junsong Yuan. 2018. Multi-view harmonized bilinear network for 3d object recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) . 186–194

  86. [95]

    Zhen Zhao, Jingqun Tang, Chunhui Lin, Binghong Wu, Can Huang, Hao Liu, Xin Tan, Zhizhong Zhang, and Yuan Xie. 2024. Multi-modal In-Context Learning Makes an Ego-evolving Scene Text Recognizer. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition ...

  87. [96]

    Zhen Zhao, Jingqun Tang, Binghong Wu, Chunhui Lin, Shu Wei, Hao Liu, Xin Tan, Zhizhong Zhang, Can Huang, and Yuan Xie. 2024. Harmonizing Visual Text Comprehension and Generation. arXiv preprint arXiv:2407.16364 (2024)

  88. [97]

    Jianping Zhang, Jen tse Huang, Wenxuan Wang, Yichen Li, Weibin Wu, Xiaosen Wang, Yuxin Su, and Michael R. Lyu. 2023. Improving the Transferability of Adversarial Samples by Path-Augmented Method. arXiv:2303.15735 [cs.CV]

  89. [98]

    Hang Zhou, Dongdong Chen, Jing Liao, Kejiang Chen, Xiaoyi Dong, Kunlin Liu, Weiming Zhang, Gang Hua, and Nenghai Yu. 2020. Lg-gan: Label guided adversarial network for flexible targeted attack of point cloud based deep net- works. In Proceedings of the IEEE Conference on Compu...

  90. [99]

    Hang Zhou, Kejiang Chen, Weiming Zhang, Han Fang, Wenbo Zhou, and Neng- hai Yu. 2019. Dup-net: Denoiser and upsampler network for 3d adversarial point clouds defense. In Proceedings of the IEEE International Conference on Computer Vision (ICCV). 1961–1970

  91. [100]

    Yue Zhao, Yuwei Wu, Caihua Chen, and Andrew Lim. 2020. On isometry robustness of deep 3d point cloud models under adversarial attacks. In Proceed- ings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) . 1201–1210

  92. [101]

    Hegui Zhu, Yuchen Ren, Xiaoyan Sui, Lianping Yang, and Wuming Jiang. 2023. Boosting Adversarial Transferability via Gradient Relevance Attack. In 2023 IEEE/CVF International Conference on Computer Vision (ICCV) . 4718–4727. https: //doi.org/10.1109/ICCV51070.2023.00437

  93. [102]

    Jiahao Zhu, Daizong Liu, Pan Zhou, Xing Di, Yu Cheng, Song Yang, Wenzheng Xu, Zichuan Xu, Yao Wan, Lichao Sun, et al. 2023. Rethinking the video sam- pling and reasoning strategies for temporal sentence grounding. arXiv preprint arXiv:2301.00514 (2023)

  94. [103]

    Tianhang Zheng, Changyou Chen, Junsong Yuan, Bo Li, and Kui Ren. 2019. Pointcloud saliency maps. In Proceedings of the IEEE International Conference on Computer Vision (ICCV). 1598–1606

  95. [106]

    Hanqi Zhu, Jiajun Deng, Yu Zhang, Jianmin Ji, Qiuyu Mao, Houqiang Li, and Yanyong Zhang. 2022. Vpfnet: Improving 3d object detection with virtual point based lidar and stereo data fusion. IEEE Transactions on Multimedia (TMM) (2022)

  96. [2015]

    In Proceedings of the IEEE International Conference on Computer Vision (ICCV)

    Multi-view convolutional neural networks for 3d shape recognition. In Proceedings of the IEEE International Conference on Computer Vision (ICCV) . 945–953

  97. [2020]

    arXiv preprint arXiv:2005.11626 (2020)

    ShapeAdv: Generating Shape-Aware Adversarial 3D Point Clouds. arXiv preprint arXiv:2005.11626 (2020)

  98. [2021]

    In Proceedings of the AAAI Conference on Artificial Intelligence, Vol

    Spatiotemporal graph neural network based mask reconstruction for video object segmentation. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 35. 2100–2108

  99. [2022]

    In European Conference on Computer Vision

    Optimal boxes: boosting end-to-end scene text recognition by adjusting annotated bounding boxes via reinforcement learning. In European Conference on Computer Vision. Springer, 233–248

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.