Pith. sign in

REVIEW 3 major objections 71 references

Multiple-Noise-Resilient Nonadiabatic Geometric Quantum Control of Solid-State Spins in Diamond

T0 review · 3 major / 0 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read A multiple-noise-resilient nonadiabatic geometric gate for diamond NV-center spins is reported to keep single-qubit gates stable under Rabi-scale detuning fluctuations, extend electron-spin coherence to 690 ± 30 µs, and reach a quantum-proc

desk verdict The abstract describes a plausible NV geometric gate result, but the supplied full text is a different audio-detection paper—nothing to review. read the letter →

arxiv 2508.12221 v1 pith:HKBWJQZG submitted 2025-08-17 quant-ph

classification quant-ph
keywords nonadiabaticgeometricquantumgatediamondNVcentercontrolrobustnessdetuningfluctuationprocesstomographycoherencetimesolid-statespinqubitphase
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper claims that a carefully designed nonadiabatic geometric quantum gate (MNR-NGQG) makes single-qubit control of diamond NV-center spins robust to simultaneous detuning, amplitude, and phase fluctuations, the main sources of control error. The evidence would show that gate performance barely degrades even when detuning fluctuations are as large as the maximum Rabi frequency, that electron-spin coherence time reaches 690 µs (3.5 times that of a standard dynamical gate in the same setup), and that quantum process tomography puts single-qubit gate fidelity at 0.9992(1). If correct, the gate offers an experimentally feasible way to achieve high-fidelity, robust quantum control in NV centers without demanding extremely stable hardware. A sympathetic reader would care because robustness to realistic noise is a key obstacle for scalable diamond-based quantum information processing.

What carries the argument

The central object is the MNR-NGQG, a nonadiabatic geometric quantum gate designed for NV centers. A geometric gate encodes the logic operation in the geometric (Berry) phase accumulated along a cyclic evolution of the qubit state, so the gate's action is determined by the path in parameter space rather than by the detailed pulse shape; this is what makes it resilient to certain classes of control noise. The added 'multiple-noise-resilient' construction is engineered to cancel or suppress errors from simultaneous detuning, amplitude, and phase fluctuations of the driving field.

What would settle it

Run the same MNR-NGQG and the dynamical comparison gate under calibrated detuning fluctuations at exactly the maximum Rabi frequency, but measure gate fidelity with randomized benchmarking instead of process tomography; if the geometric gate's fidelity drops significantly below 0.999 while the dynamical gate's drops even more, the central claim would be falsified. Alternatively, if the coherence time of the electron spin under the geometric gate is not 3.5 times the dynamical one when measured with the same pulse sequence and environment, the 3.5x claim fails.

Watch

Extended reading notes

Core claim

The core claim is that a multiple-noise-resilient nonadiabatic geometric quantum gate outperforms the conventional dynamical gate on both robustness and coherence in the same diamond NV-center device. The design keeps the qubit's evolution geometric so that it acquires phase without relying on adiabaticity, making it faster than adiabatic schemes while remaining protected against control errors. The presented results are: gate performance almost unchanged when detuning fluctuation range is comparable to the maximum Rabi frequency; electron-spin coherence time of 690 ± 30 µs, 3.5 times longer than the naive dynamical counterpart; and single-qubit gate fidelity of 0.9992(1) measured by quantum

Load-bearing premise

The load-bearing premise is that the reported 0.9992(1) process-tomography fidelity is a gate fidelity corrected for state preparation and measurement errors, and that the tested error channel really is simultaneous detuning, amplitude, and phase fluctuations with detuning dominating; the abstract alone provides no experimental parameters to verify either.

Editorial extensions

If this is right

  • If the claims hold, geometric single-qubit gates in diamond NV centers can replace dynamical gates in high-fidelity control, because they offer the same or better fidelity plus longer coherence.
  • Detuning fluctuations up to the maximum Rabi frequency can be tolerated, which relaxes hardware requirements for frequency stabilization and magnetic-field control.
  • The 690-µs electron-spin coherence time suggests that geometric gates could be used for longer-lived storage or for multi-qubit operations before decoherence sets in.
  • The experiment-friendly design implies the scheme could be adopted in other solid-state spin qubits or other platforms where nonadiabatic geometric control is feasible.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The reported robustness against detuning fluctuations at the Rabi scale suggests the same design principle might suppress noise in other qubit platforms, such as trapped ions or superconducting circuits, where detuning and amplitude errors are common.
  • Since only process-tomography fidelity is reported, a randomized-benchmarking measurement would be a stronger test; if the two disagree, the 0.9992(1) number may not reflect the gate's operational fidelity under SPAM-free conditions.
  • The supplied full text is an unrelated audio-detection manuscript, so the experimental claims here are supported only by the abstract; a reader should rely on the original published version for the derivation and experimental parameters.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 0 minor

Summary. The manuscript claims an experimental demonstration of a multiple-noise-resilient nonadiabatic geometric quantum gate (MNR-NGQG) on nitrogen-vacancy (NV) centers in diamond. The abstract reports: (i) single-qubit gate robustness when the detuning fluctuation range is comparable to the maximum Rabi frequency; (ii) electron-spin coherence time of 690 ± 30 µs, stated as 3.5 times that of a 'naive dynamical counterpart'; and (iii) single-qubit gate fidelity of 0.9992(1) from quantum process tomography. The supplied full text, however, is a completely different paper, 'Exploring Self-Supervised Audio Models for Generalized Anomalous Sound Detection' (arXiv:2508.12230). That text contains no Hamiltonian, no geometric-phase derivation, no pulse sequence, no NV-center data, no coherence measurement, and no QPT procedure. The manuscript therefore consists of an abstract with no supporting body.

Significance. If substantiated, the reported combination—geometric single-qubit gates with fidelities near 0.9992, Rabi-scale detuning robustness, and a 3.5× coherence-time advantage over a dynamical comparison—would be a valuable experimental result for NV-center quantum control. However, the absence of any supporting text makes the claims impossible to assess. The reader cannot check the gate construction, the noise model, the calibration of the QPT, or the equality of conditions between geometric and dynamical gates. No derivations, data, or code are supplied. Because the manuscript as submitted does not contain its own evidence, its significance is currently unverifiable rather than established.

major comments (3)
  1. [Full text (entire body)] The entire body text is an unrelated paper on anomalous sound detection (arXiv:2508.12230). None of the abstract's claims about MNR-NGQG, NV centers, coherence time, or QPT appear in the body. This is not a presentation defect; it removes the evidentiary basis for every load-bearing claim. The manuscript as supplied cannot be reviewed scientifically.
  2. [Abstract, QPT fidelity claim] The claim that 'the fidelity of single-qubit gates reaches 0.9992(1), as characterized by quantum process tomography' is unsupported. There is no description of the gates implemented, the QPT pulse sequence, the readout calibration, SPAM error mitigation, or the fitting procedure. Without these, the fidelity number cannot be evaluated.
  3. [Abstract, robustness and coherence claims] The claims of robustness to detuning fluctuations comparable to the maximum Rabi frequency and coherence time 690±30 µs (3.5× the naive dynamical counterpart) require the pulse construction, the noise model, the comparison gate definition, and the experimental conditions. None of these are present. In particular, the 'naive dynamical counterpart' is never defined, so the ratio 3.5 cannot be checked.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity found: the abstract reports empirical results, and the supplied full text is an unrelated audio-detection paper, so there is no derivation chain to audit and no exhibited reduction of a prediction to an input.

full rationale

The claim under review is an experimental report: a multiple-noise-resilient nonadiabatic geometric gate implemented on diamond NV centers, with robustness to detuning fluctuations, coherence time 690 ± 30 µs, a 3.5x improvement over a dynamical gate, and QPT fidelity 0.9992(1). None of these numbers is presented as derived from a fit, and the abstract contains no equation in which an output is defined in terms of the target quantity. The supplied full text is not the NV-center paper at all: it is 'Exploring Self-Supervised Audio Models for Generalized Anomalous Sound Detection' (arXiv:2508.12230), containing no Hamiltonian, no geometric-phase construction, no pulse parameters, no noise model, and no quantum process tomography. This body/abstract mismatch means the derivation chain cannot be audited, but an unverifiable claim is not the same as a circular claim. The circularity patterns listed—self-definition, fitted input renamed as prediction, load-bearing self-citation, imported uniqueness, ansatz smuggled via citation, renaming a known result—all require exhibiting a specific reduction in the paper's own text. No such reduction can be exhibited here. Any concern about whether the SPAM calibration, noise model, or choice of dynamical comparison makes the headline numbers favorable is a correctness/experimental-fairness question, not a circularity question, and speculating about hidden circularity in the missing body would violate the instruction not to manufacture circularity. Therefore the honest finding is no significant circularity, score 0.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The abstract alone reveals load-bearing premises: the two-level model of the NV spin with the stated error channels; the mathematical validity of the MNR-NGQG construction; and the reliability of quantum process tomography as a fidelity estimator. None could be audited because the body text is an unrelated paper. No free parameters are visible at the abstract level (Rabi frequency, detuning amplitudes, and durations are not stated), and no invented physical entities are introduced; MNR-NGQG is a control protocol, not a postulated entity.

assumptions (3)
  • domain assumption A diamond NV-center electron spin can be modeled as a driven two-level system whose dominant errors are control-pulse amplitude noise, phase noise, and detuning fluctuations
    Abstract: 'control pulses inevitably introduce multiple errors, leading to decoherence.' The specific error model underpinning the claimed multiple-noise resilience is not stated in the accessible text.
  • domain assumption The MNR-NGQG construction is mathematically valid: the prescribed pulse sequence generates the target rotation as a geometric phase along a closed loop, with the claimed suppression of the stated noise channels
    This is the theoretical core of the robustness and coherence claims. No equations or derivation appear in the supplied manuscript body, which is an unrelated paper.
  • domain assumption Quantum process tomography of the electron spin yields an unbiased estimate of gate fidelity, with state preparation and measurement errors properly accounted for
    Abstract: 'fidelity of single-qubit gates reaches 0.9992(1), as characterized by quantum process tomography.' The reconstruction procedure and SPAM calibration are absent.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Multiple-Noise-Resilient Nonadiabatic Geometric Quantum Control of Solid-State Spins in Diamond." pith.science (2026). https://pith.science/paper/HKBWJQZG

@misc{pith2026250812221,
  author       = {Pith},
  title        = {Pith review of: Multiple-Noise-Resilient Nonadiabatic Geometric Quantum Control of Solid-State Spins in Diamond},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/HKBWJQZG}},
  note         = {Machine review of arXiv:2508.12221}
}
abstract

Reliable and robust control lies at the core of implementing quantum information processing with diamond nitrogen-vacancy (NV) centers. However, control pulses inevitably introduce multiple errors, leading to decoherence and hindering scalable applications. Here, we experimentally report an experiment-friendly multiple-noise-resilient nonadiabatic geometric quantum gate~(MNR-NGQG) that can significantly improve conventional dynamical gate in both robustness and coherence. Notably, even when the detuning fluctuation range is comparable to the maximum Rabi frequency, the single-qubit gate performance of the MNR-NGQG remains almost unchanged. Besides, the coherence time of the electron spin is significantly extended to 690 $\pm$ 30 $ \mu$s, 3.5 times that of the naive dynamical counterpart. As a result, the fidelity of single-qubit gates reaches 0.9992(1), as characterized by quantum process tomography. With its experimentally feasible design and relaxed hardware requirements, our work offers a solid paradigm for achieving high-fidelity quantum control in NV center system, paving the way for practical applications in quantum information science.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

71 extracted references · 59 canonical work pages

  1. [1]

    Exploring large scale pre-trained models for robust machine anomalous sound detection,

    B. Han, Z. Lv, A. Jiang, W. Huang, Z. Chen, Y . Deng, J. Ding, C. Lu, W.- Q. Zhang, P. Fan, J. Liu, and Y . Qian, “Exploring large scale pre-trained models for robust machine anomalous sound detection,” in ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2024, pp. 1–5

  2. [2]

    Anopatch: Towards better consistency in machine anomalous sound detection,

    A. Jiang, B. Han, Z. Lv, Y . Deng, W.-Q. Zhang, X. Chen, Y . Qian, J. Liu, and P. Fan, “Anopatch: Towards better consistency in machine anomalous sound detection,” in Interspeech 2024, 2024, pp. 107–111

  3. [3]

    Description and Discussion on DCASE2020 Challenge Task2: Unsupervised Anomalous Sound Detection for Machine Condition Monitoring

    Y . Koizumi, Y . Kawaguchi, K. Imoto, T. Nakamura, Y . Nikaido, R. Tan- abe, H. Purohit, K. Suefusa, T. Endo, M. Yasuda et al. , “Description and discussion on dcase2020 challenge task2: Unsupervised anomalous sound detection for machine condition monitoring,” arXiv preprint arXiv:2006.05822, 2020. JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2015 14

  4. [4]

    Description and Discussion on DCASE 2021 Challenge Task 2: Unsupervised Anomalous Sound Detection for Machine Condition Monitoring under Domain Shifted Conditions

    Y . Kawaguchi, K. Imoto, Y . Koizumi, N. Harada, D. Niizumi, K. Dohi, R. Tanabe, H. Purohit, and T. Endo, “Description and discussion on dcase 2021 challenge task 2: Unsupervised anomalous sound detection for machine condition monitoring under domain shifted conditions,” arXiv preprint arXiv:2106.04492 , 2021

  5. [5]

    Description and Discussion on DCASE 2022 Challenge Task 2: Unsupervised Anomalous Sound Detection for Machine Condition Monitoring Applying Domain Generalization Techniques

    K. Dohi, K. Imoto, N. Harada, D. Niizumi, Y . Koizumi, T. Nishida, H. Purohit, T. Endo, M. Yamamoto, and Y . Kawaguchi, “Description and discussion on dcase 2022 challenge task 2: Unsupervised anomalous sound detection for machine condition monitoring applying domain generalization techniques,” arXiv preprint arXiv:2206.05876 , 2022

  6. [6]

    Description and Discussion on DCASE 2023 Challenge Task 2: First-Shot Unsupervised Anomalous Sound Detection for Machine Condition Monitoring

    K. Dohi, K. Imoto, N. Harada, D. Niizumi, Y . Koizumi, T. Nishida, H. Purohit, R. Tanabe, T. Endo, and Y . Kawaguchi, “Description and discussion on dcase 2023 challenge task 2: First-shot unsupervised anomalous sound detection for machine condition monitoring,” arXiv preprint arXiv:2305.07828, 2023

  7. [7]

    Description and Discussion on DCASE 2024 Challenge Task 2: First-Shot Unsupervised Anomalous Sound Detection for Machine Condition Monitoring

    T. Nishida, N. Harada, D. Niizumi, D. Albertini, R. Sannino, S. Pradolini, F. Augusti, K. Imoto, K. Dohi, H. Purohit et al., “Descrip- tion and discussion on dcase 2024 challenge task 2: First-shot unsu- pervised anomalous sound detection for machine condition monitoring,” arXiv preprint arXiv:2406.07250 , 2024

  8. [8]

    A unifying review of deep and shallow anomaly detection,

    L. Ruff, J. R. Kauffmann, R. A. Vandermeulen, G. Montavon, W. Samek, M. Kloft, T. G. Dietterich, and K.-R. M ¨uller, “A unifying review of deep and shallow anomaly detection,” Proceedings of the IEEE , vol. 109, no. 5, pp. 756–795, 2021

Show all 71 references
  1. [9]

    Anomalous sound detection based on interpolation deep neural network,

    K. Suefusa, T. Nishida, H. Purohit, R. Tanabe, T. Endo, and Y . Kawaguchi, “Anomalous sound detection based on interpolation deep neural network,” in ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2020, pp. 271–275

  2. [10]

    Flow- based self-supervised density estimation for anomalous sound detection,

    K. Dohi, T. Endo, H. Purohit, R. Tanabe, and Y . Kawaguchi, “Flow- based self-supervised density estimation for anomalous sound detection,” in ICASSP 2021-2021 Proc. IEEE Int. Conf. Acoust., Speech, Signal Process. (ICASSP). IEEE, 2021, pp. 336–340

  3. [11]

    Unsupervised anomalous sound detection using self- supervised classification and group masked autoencoder for density estimation,

    R. Giri, S. V . Tenneti, K. Helwani, F. Cheng, U. Isik, and A. Kr- ishnaswamy, “Unsupervised anomalous sound detection using self- supervised classification and group masked autoencoder for density estimation,” DCASE2020 Challenge, Tech. Rep., July 2020

  4. [12]

    Unsupervised anomaly detection and localization of machine audio: A gan-based approach,

    A. Jiang, W.-Q. Zhang, Y . Deng, P. Fan, and J. Liu, “Unsupervised anomaly detection and localization of machine audio: A gan-based approach,” in ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2023, pp. 1–5

  5. [13]

    Self- supervised representation learning for unsupervised anomalous sound detection under domain shift,

    H. Chen, Y . Song, L.-R. Dai, I. McLoughlin, and L. Liu, “Self- supervised representation learning for unsupervised anomalous sound detection under domain shift,” in ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2022...

  6. [14]

    Bert: Pre-training of deep bidirectional transformers for language understanding,

    J. Devlin, “Bert: Pre-training of deep bidirectional transformers for language understanding,” arXiv preprint arXiv:1810.04805 , 2018

  7. [15]

    Superb: Speech processing universal performance benchmark,

    S.-w. Yang, P.-H. Chi, Y .-S. Chuang, C.-I. J. Lai, K. Lakhotia, Y . Y . Lin, A. T. Liu, J. Shi, X. Chang, G.-T. Lin et al., “Superb: Speech processing universal performance benchmark,” arXiv preprint arXiv:2105.01051 , 2021

  8. [16]

    BEATs: Audio pre-training with acoustic tokenizers,

    S. Chen, Y . Wu, C. Wang, S. Liu, D. Tompkins, Z. Chen, W. Che, X. Yu, and F. Wei, “BEATs: Audio pre-training with acoustic tokenizers,” in Proceedings of the 40th International Conference on Machine Learning , ser. Proceedings of Machine Learning Research, vol. 202. PMLR, 23–...

  9. [17]

    Large-scale self-supervised speech representation learning for automatic speaker verification,

    Z. Chen, S. Chen, Y . Wu, Y . Qian, C. Wang, S. Liu, Y . Qian, and M. Zeng, “Large-scale self-supervised speech representation learning for automatic speaker verification,” in Proc. IEEE ICASSP . IEEE, 2022, pp. 6147–6151

  10. [18]

    Attention is all you need,

    A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in Proc. NIPS, 2017, pp. 5998–6008

  11. [19]

    Sub-cluster adacos: Learning representations for anomalous sound detection,

    K. Wilkinghoff, “Sub-cluster adacos: Learning representations for anomalous sound detection,” in 2021 International Joint Conference on Neural Networks (IJCNN) . IEEE, 2021, pp. 1–8

  12. [20]

    Diffusion augmentation sub-center modeling for unsupervised anomalous sound detection with partially attribute-unavailable conditions,

    J. Yin, Y . Gao, W. Zhang, T. Wang, and M. Zhang, “Diffusion augmentation sub-center modeling for unsupervised anomalous sound detection with partially attribute-unavailable conditions,” in ICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Processi...

  13. [21]

    Adaptive prototype learning for anomalous sound detection with partially known attributes,

    A. Jiang, X. Zheng, B. Han, Y . Qiu, P. Fan, W.-Q. Zhang, C. Lu, and J. Liu, “Adaptive prototype learning for anomalous sound detection with partially known attributes,” in ICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEE...

  14. [22]

    Disentangling hierarchical features for anomalous sound detection under domain shift,

    J. Guan, J. Tian, Q. Zhu, F. Xiao, H. Zhang, and X. Liu, “Disentangling hierarchical features for anomalous sound detection under domain shift,” in ICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2025, pp. 1–5

  15. [23]

    Joint generative-contrastive representation learning for anomalous sound detection,

    X.-M. Zeng, Y . Song, Z. Zhuo, Y . Zhou, Y .-H. Li, H. Xue, L.-R. Dai, and I. McLoughlin, “Joint generative-contrastive representation learning for anomalous sound detection,” in ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)...

  16. [24]

    Sw-wavenet: learning represen- tation from spectrogram and wavegram using wavenet for anomalous sound detection,

    H. Chen, L. Ran, X. Sun, and C. Cai, “Sw-wavenet: learning represen- tation from spectrogram and wavegram using wavenet for anomalous sound detection,” in ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2023, pp. 1–5

  17. [25]

    Anomalous sound de- tection using self-attention-based frequency pattern analysis of machine sounds,

    H. Zhang, J. Guan, Q. Zhu, F. Xiao, and Y . Liu, “Anomalous sound de- tection using self-attention-based frequency pattern analysis of machine sounds,” in INTERSPEECH 2023, 2023, pp. 336–340

  18. [26]

    Efficient algorithms for mining outliers from large data sets,

    S. Ramaswamy, R. Rastogi, and K. Shim, “Efficient algorithms for mining outliers from large data sets,” in Proc. 2000 ACM SIGMOD Int. Conf. Manag. Data , 2000, pp. 427–438

  19. [27]

    Lof: identifying density-based local outliers,

    M. M. Breunig, H.-P. Kriegel, R. T. Ng, and J. Sander, “Lof: identifying density-based local outliers,” in Proceedings of the 2000 ACM SIGMOD international conference on Management of data , 2000, pp. 93–104

  20. [28]

    Deep autoencoding gaussian mixture model for unsupervised anomaly detection,

    B. Zong, Q. Song, M. R. Min, W. Cheng, C. Lumezanu, D. Cho, and H. Chen, “Deep autoencoding gaussian mixture model for unsupervised anomaly detection,” in International conference on learning representa- tions, 2018

  21. [29]

    wav2vec 2.0: A framework for self-supervised learning of speech representations,

    A. Baevski, Y . Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A framework for self-supervised learning of speech representations,” Advances in neural information processing systems, vol. 33, pp. 12 449– 12 460, 2020

  22. [30]

    Hubert: Self-supervised speech representation learning by masked prediction of hidden units,

    W.-N. Hsu, B. Bolte, Y .-H. H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed, “Hubert: Self-supervised speech representation learning by masked prediction of hidden units,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 29, pp. 3451–3460, 2021

  23. [31]

    Unispeech: Unified speech representation learning with labeled and unlabeled data,

    C. Wang, Y . Wu, Y . Qian, K. Kumatani, S. Liu, F. Wei, M. Zeng, and X. Huang, “Unispeech: Unified speech representation learning with labeled and unlabeled data,” in International Conference on Machine Learning. PMLR, 2021, pp. 10 937–10 947

  24. [32]

    Wavlm: Large-scale self-supervised pre- training for full stack speech processing,

    S. Chen, C. Wang, Z. Chen, Y . Wu, S. Liu, Z. Chen, J. Li, N. Kanda, T. Yoshioka, X. Xiao et al. , “Wavlm: Large-scale self-supervised pre- training for full stack speech processing,” IEEE Journal of Selected Topics in Signal Processing , vol. 16, no. 6, pp. 1505–1518, 2022

  25. [33]

    AST: Audio Spectrogram Trans- former,

    Y . Gong, Y .-A. Chung, and J. Glass, “AST: Audio Spectrogram Trans- former,” in Proc. Interspeech 2021 , 2021, pp. 571–575

  26. [34]

    An image is worth 16x16 words: Transformers for image recognition at scale,

    A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly et al., “An image is worth 16x16 words: Transformers for image recognition at scale,” arXiv preprint arXiv:2010.11929 , 2020

  27. [35]

    Imagebind: One embedding space to bind them all,

    R. Girdhar, A. El-Nouby, Z. Liu, M. Singh, K. V . Alwala, A. Joulin, and I. Misra, “Imagebind: One embedding space to bind them all,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 15 180–15 190

  28. [36]

    Self-supervised audio teacher-student transformer for both clip-level and frame-level tasks,

    X. Li, N. Shao, and X. Li, “Self-supervised audio teacher-student transformer for both clip-level and frame-level tasks,” IEEE/ACM Trans- actions on Audio, Speech, and Language Processing , 2024

  29. [37]

    Emotion recognition from speech using wav2vec 2.0 embeddings,

    L. Pepino, P. Riera, and L. Ferrer, “Emotion recognition from speech using wav2vec 2.0 embeddings,” arXiv preprint arXiv:2104.03502 , 2021

  30. [38]

    Sparsely shared lora on whisper for child speech recognition,

    W. Liu, Y . Qin, Z. Peng, and T. Lee, “Sparsely shared lora on whisper for child speech recognition,” in ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2024, pp. 11 751–11 755

  31. [39]

    Prefix-tuning: Optimizing continuous prompts for generation,

    X. L. Li and P. Liang, “Prefix-tuning: Optimizing continuous prompts for generation,” arXiv preprint arXiv:2101.00190 , 2021

  32. [40]

    Efficient adapter transfer of self- supervised speech models for automatic speech recognition,

    B. Thomas, S. Kessler, and S. Karout, “Efficient adapter transfer of self- supervised speech models for automatic speech recognition,” in ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2022, pp. 7102–7106

  33. [41]

    Librispeech: an asr corpus based on public domain audio books,

    V . Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: an asr corpus based on public domain audio books,” in2015 IEEE international conference on acoustics, speech and signal processing (ICASSP). IEEE, 2015, pp. 5206–5210. JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. ...

  34. [42]

    Audio set: An ontology and human- labeled dataset for audio events,

    J. F. Gemmeke, D. P. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter, “Audio set: An ontology and human- labeled dataset for audio events,” in 2017 IEEE international conference on acoustics, speech and signal processing (ICASSP) . IEEE, 2017,...

  35. [43]

    Attentive statistics pooling for deep speaker embedding,

    K. Okabe, T. Koshinaka, and K. Shinoda, “Attentive statistics pooling for deep speaker embedding,” arXiv preprint arXiv:1803.10963 , 2018

  36. [44]

    Lora: Low-rank adaptation of large language models,

    E. J. Hu, Y . Shen, P. Wallis, Z. Allen-Zhu, Y . Li, S. Wang, L. Wang, and W. Chen, “Lora: Low-rank adaptation of large language models,” arXiv preprint arXiv:2106.09685 , 2021

  37. [45]

    Robust anomaly sound detection framework for machine condition monitoring,

    Y . Zeng, H. Liu, L. Xu, Y . Zhou, and L. Gan, “Robust anomaly sound detection framework for machine condition monitoring,” DCASE2022 Challenge, Tech. Rep., July 2022

  38. [46]

    Arcface: Additive angular margin loss for deep face recognition,

    J. Deng, J. Guo, N. Xue, and S. Zafeiriou, “Arcface: Additive angular margin loss for deep face recognition,” in Proc. CVPR, 2019, pp. 4690– 4699

  39. [47]

    Large margin softmax loss for speaker verification,

    Y . Liu, L. He, and J. Liu, “Large margin softmax loss for speaker verification,” in Proc. ISCA Interspeech , G. Kubin and Z. Kacic, Eds., 2019, pp. 2873–2877

  40. [48]

    Why do angular margin losses work well for semi-supervised anomalous sound detection?

    K. Wilkinghoff and F. Kurth, “Why do angular margin losses work well for semi-supervised anomalous sound detection?” IEEE/ACM Transac- tions on Audio, Speech, and Language Processing , 2023

  41. [49]

    Toyadmos: A dataset of miniature-machine operating sounds for anomalous sound detection,

    Y . Koizumi, S. Saito, H. Uematsu, N. Harada, and K. Imoto, “Toyadmos: A dataset of miniature-machine operating sounds for anomalous sound detection,” in 2019 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA) . IEEE, 2019, pp. 313–317

  42. [50]

    Toyadmos2: Another dataset of miniature-machine operating sounds for anomalous sound detection under domain shift conditions,

    N. Harada, D. Niizumi, D. Takeuchi, Y . Ohishi, M. Yasuda, and S. Saito, “Toyadmos2: Another dataset of miniature-machine operating sounds for anomalous sound detection under domain shift conditions,” arXiv preprint arXiv:2106.02369, 2021

  43. [51]

    Mimii dataset: Sound dataset for malfunction- ing industrial machine investigation and inspection,

    H. Purohit, R. Tanabe, K. Ichige, T. Endo, Y . Nikaido, K. Suefusa, and Y . Kawaguchi, “Mimii dataset: Sound dataset for malfunction- ing industrial machine investigation and inspection,” arXiv preprint arXiv:1909.09347, 2019

  44. [52]

    Mimii dg: Sound dataset for mal- functioning industrial machine investigation and inspection for domain generalization task,

    K. Dohi, T. Nishida, H. Purohit, R. Tanabe, T. Endo, M. Yamamoto, Y . Nikaido, and Y . Kawaguchi, “Mimii dg: Sound dataset for mal- functioning industrial machine investigation and inspection for domain generalization task,” arXiv preprint arXiv:2205.13879 , 2022

  45. [53]

    Imad-ds: A dataset for industrial multi-sensor anomaly detection under domain shift conditions,

    D. Albertini, F. Augusti, K. Esmer, A. Bernardini, and R. Sannino, “Imad-ds: A dataset for industrial multi-sensor anomaly detection under domain shift conditions,” in Proceedings of the Detection and Classi- fication of Acoustic Scenes and Events 2024 Workshop (DCASE2024) , T...

  46. [54]

    Mobilefacenets: Efficient cnns for accurate real-time face verification on mobile devices,

    S. Chen, Y . Liu, X. Gao, and Z. Han, “Mobilefacenets: Efficient cnns for accurate real-time face verification on mobile devices,” in Biometric Recognition: 13th Chinese Conference, CCBR 2018, Urumqi, China, August 11-12, 2018, Proceedings 13 . Springer, 2018, pp. 428–438

  47. [55]

    Ensemble of complemen- tary anomaly detectors under domain shifted conditions,

    J. Lopez, G. Stemmer, and P. Lopez-Meyer, “Ensemble of complemen- tary anomaly detectors under domain shifted conditions,” DCASE2021 Challenge, Tech. Rep., July 2021

  48. [56]

    First- shot anomaly sound detection for machine condition monitoring: A do- main generalization baseline,

    N. Harada, D. Niizumi, Y . Ohishi, D. Takeuchi, and M. Yasuda, “First- shot anomaly sound detection for machine condition monitoring: A do- main generalization baseline,” in 2023 31st European Signal Processing Conference (EUSIPCO). IEEE, 2023, pp. 191–195

  49. [57]

    Specaugment: A simple data augmentation method for automatic speech recognition,

    D. S. Park, W. Chan, Y . Zhang, C. Chiu, B. Zoph, E. D. Cubuk, and Q. V . Le, “Specaugment: A simple data augmentation method for automatic speech recognition,” in Proc. ISCA Interspeech , 2019, pp. 2613–2617

  50. [58]

    Anomalous sound detection using cnn-based features by self supervised learning,

    K. Morita, T. Yano, and K. Tran, “Anomalous sound detection using cnn-based features by self supervised learning,” Tech. Rep., Challenge on Detection and Classification of Acoustic Scenes and Events (DCASE Challenge), 2021

  51. [59]

    Ced: Consistent ensemble distillation for audio tagging,

    H. Dinkel, Y . Wang, Z. Yan, J. Zhang, and Y . Wang, “Ced: Consistent ensemble distillation for audio tagging,” in ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2024, pp. 291–295

  52. [60]

    Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,

    Y . Wu, K. Chen, T. Zhang, Y . Hui, T. Berg-Kirkpatrick, and S. Dubnov, “Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,” in ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (I...

  53. [61]

    Anomalous sound detection using spectral-temporal information fusion,

    Y . Liu, J. Guan, Q. Zhu, and W. Wang, “Anomalous sound detection using spectral-temporal information fusion,” in ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2022, pp. 816–820

  54. [62]

    An effective anomalous sound detection method based on represen- tation learning with simulated anomalies,

    H. Chen, Y . Song, Z. Zhuo, Y . Zhou, Y .-H. Li, H. Xue, and I. McLough- lin, “An effective anomalous sound detection method based on represen- tation learning with simulated anomalies,” in ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processi...

  55. [63]

    Time- weighted frequency domain audio representation with gmm estimator for anomalous sound detection,

    J. Guan, Y . Liu, Q. Zhu, T. Zheng, J. Han, and W. Wang, “Time- weighted frequency domain audio representation with gmm estimator for anomalous sound detection,” in ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2023, pp. 1–5

  56. [64]

    Anomalous sound detection based on self-supervised learning,

    J. Jie, “Anomalous sound detection based on self-supervised learning,” DCASE2023 Challenge, Tech. Rep., June 2023

  57. [65]

    Self-supervised learning for anomalous sound detec- tion,

    K. Wilkinghoff, “Self-supervised learning for anomalous sound detec- tion,” arXiv preprint arXiv:2312.09578 , 2023

  58. [66]

    Stream-based active learning for anomalous sound detection in machine condition monitoring,

    T. V . Ho, K. Dohi, and Y . Kawaguchi, “Stream-based active learning for anomalous sound detection in machine condition monitoring,” in Interspeech 2024, 2024, pp. 102–106

  59. [67]

    A dual-path frame- work with frequency-and-time excited network for anomalous sound detection,

    Y . Zhang, J. Liu, Y . Tian, H. Liu, and M. Li, “A dual-path frame- work with frequency-and-time excited network for anomalous sound detection,” in ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2024, pp. 1266–1270

  60. [68]

    Aithu system for first-shot unsupervised anomalous sound detection,

    Z. Lv, A. Jiang, B. Han, Y . Liang, Y . Qian, X. Chen, J. Liu, and P. Fan, “Aithu system for first-shot unsupervised anomalous sound detection,” DCASE2024 Challenge, Tech. Rep., June 2024

  61. [69]

    Thuee system for first-shot unsupervised anomalous sound detection,

    A. Jiang, X. Zheng, Y . Qiu, W. Zhang, B. Chen, P. Fan, W.-Q. Zhang, C. Lu, and J. Liu, “Thuee system for first-shot unsupervised anomalous sound detection,” DCASE2024 Challenge, Tech. Rep., June 2024

  62. [70]

    Enhanced unsupervised anomalous sound detection using conditional autoencoder for machine condition monitoring,

    R. Zhao, K. Ren, and L. Zou, “Enhanced unsupervised anomalous sound detection using conditional autoencoder for machine condition monitoring,” DCASE2024 Challenge, Tech. Rep., June 2024

  63. [71]

    Smote: synthetic minority over-sampling technique,

    N. V . Chawla, K. W. Bowyer, L. O. Hall, and W. P. Kegelmeyer, “Smote: synthetic minority over-sampling technique,” Journal of artificial intel- ligence research, vol. 16, pp. 321–357, 2002

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.