Pith. sign in

REVIEW 4 cited by

Cross-domain Neural Pitch and Periodicity Estimation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2301.12258 v3 pith:7GDNARYK submitted 2023-01-28 eess.AS cs.SD

classification eess.AScs.SD
keywords pitchneuralmusicestimatorsnetworkspennperiodicityspeech
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Pitch is a foundational aspect of our perception of audio signals. Pitch contours are commonly used to analyze speech and music signals and as input features for many audio tasks, including music transcription, singing voice synthesis, and prosody editing. In this paper, we describe a set of techniques for improving the accuracy of widely-used neural pitch and periodicity estimators to achieve state-of-the-art performance on both speech and music. We also introduce a novel entropy-based method for extracting periodicity and per-frame voiced-unvoiced classifications from statistical inference-based pitch estimators (e.g., neural networks), and show how to train a neural pitch estimator to simultaneously handle both speech and music data (i.e., cross-domain estimation) without performance degradation. Our estimator implementations run 11.2x faster than real-time on a Intel i9-9820X 10-core 3.30 GHz CPU$\unicode{x2014}$approaching the speed of state-of-the-art DSP-based pitch estimators$\unicode{x2014}$or 408x faster than real-time on a NVIDIA GeForce RTX 3090 GPU. We release all of our code and models as Pitch-Estimating Neural Networks (penn), an open-source, pip-installable Python module for training, evaluating, and performing inference with pitch- and periodicity-estimating neural networks. The code for penn is available at https://github.com/interactiveaudiolab/penn.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. The Prosody of Emojis

    cs.CL 2025-08 conditional novelty 6.0 of 10

    Speakers systematically modify prosody when reading sentences with different emojis, and listeners can recover the intended emoji from prosody alone, with larger semantic differences yielding larger prosodic shifts.

  2. Improving Neural Pitch Estimation with SWIPE Kernels

    cs.SD 2025-07 conditional novelty 6.0 of 10

    SWIPE scores outperform a self-supervised neural pitch estimator (PESTO) on common benchmarks and, as an audio frontend, let a 647-parameter network match or beat the larger model.

  3. RepeaTTS: Towards Feature Discovery through Repeated Fine-Tuning

    cs.CL 2025-07 conditional novelty 6.0 of 10

    PCA of repeated synthesized utterances with fixed inputs can reveal controllable prosodic features that can be enrolled as new prompts via fine-tuning.

  4. ProMode: A Speech Prosody Model Conditioned on Acoustic and Textual Inputs

    eess.AS 2025-08 unverdicted novelty 5.0 of 10

    ProMode learns a prosody embedding from partially masked audio and text, improving F0 and energy prediction over baseline style encoders.

Pith tools