Pith. sign in

REVIEW 3 cited by

End-to-End Audio Strikes Back: Boosting Augmentations Towards An Efficient Audio Classification Network

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2204.11479 v5 pith:Q5LGLTII submitted 2022-04-25 cs.SD cs.CVeess.AS

classification cs.SDcs.CVeess.AS
keywords audioaugmentationsclassificationefficientend-to-endarchitectureslargenetwork
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

While efficient architectures and a plethora of augmentations for end-to-end image classification tasks have been suggested and heavily investigated, state-of-the-art techniques for audio classifications still rely on numerous representations of the audio signal together with large architectures, fine-tuned from large datasets. By utilizing the inherited lightweight nature of audio and novel audio augmentations, we were able to present an efficient end-to-end network with strong generalization ability. Experiments on a variety of sound classification sets demonstrate the effectiveness and robustness of our approach, by achieving state-of-the-art results in various settings. Public code is available at: \href{https://github.com/Alibaba-MIIL/AudioClassfication}{this http url}

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SPBA: Utilizing Speech Large Language Model for Backdoor Attacks on Speech Classification Models

    cs.SD 2025-06 conditional novelty 6.0 of 10

    A speech backdoor attack uses SLLM-generated timbre and emotion triggers with MGDA-balanced training to implant multiple effective backdoors in speech classifiers.

  2. OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A shared-backbone transformer with pairwise modality training reports top results across 25 datasets spanning 12 modalities.

  3. Adaptive Loss Balancing for Multi-Task Bioacoustic Classification of Bird Species and Call Types

    cs.SD 2026-07 conditional novelty 4.5 of 10

    On WiWa, factorised multi-task bird species and call-type classification with adaptive loss balancing improves call-type recognition most consistently, while preferred weighting and adaptation depth depend on backbone...

Pith tools