Pith. sign in

REVIEW 4 cited by

TorchAudio: Building Blocks for Audio and Speech Processing

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2110.15018 v2 pith:XPDN7PX3 submitted 2021-10-28 eess.AS cs.SD

classification eess.AScs.SD
keywords torchaudioaudioblocksbuildingspeechapplicationsavailablebenchmarks
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

This document describes version 0.10 of TorchAudio: building blocks for machine learning applications in the audio and speech processing domain. The objective of TorchAudio is to accelerate the development and deployment of machine learning applications for researchers and engineers by providing off-the-shelf building blocks. The building blocks are designed to be GPU-compatible, automatically differentiable, and production-ready. TorchAudio can be easily installed from Python Package Index repository and the source code is publicly available under a BSD-2-Clause License (as of September 2021) at https://github.com/pytorch/audio. In this document, we provide an overview of the design principles, functionalities, and benchmarks of TorchAudio. We also benchmark our implementation of several audio and speech operations and models. We verify through the benchmarks that our implementations of various operations and models are valid and perform similarly to other publicly available implementations.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. StemFX: Learning Mixing Style Representations via Autoregressive FX Chain Prediction on Source-Separated Stems

    cs.SD 2026-07 conditional novelty 6.0 of 10

    StemFX predicts tokenized per-stem audio-effect chains with a jointly-trained Transformer encoder-decoder, beating contrastive and prior FX-encoding methods on effect-chain retrieval and real-mix style transfer.

  2. Representation Purification for End-to-End Speech Translation

    cs.CL 2024-12 conditional novelty 6.0 of 10

    SRPSE removes content-agnostic components from speech representations using orthogonal projection, supervision from speaker and SNR labels, and consistency and mutual information losses, improving speech translation quality.

  3. Data-independent Beamforming for End-to-end Multichannel Multi-speaker ASR

    cs.SD 2025-09 conditional novelty 5.0 of 10

    A data-independent, sector-based beamformer used as a training-free front-end improves word error rate by up to 11% relative on the AMI meeting corpus for a multichannel MFCCA-based multi-speaker ASR system.

  4. Multimodal Sentiment Analysis based on Video and Audio Inputs

    cs.SD 2024-12 reject novelty 3.0 of 10

    Combining probabilities from fine-tuned wav2vec2 and ViViT models gives modest emotion-recognition results, but the evaluation design does not support a reliable conclusion.

Pith tools