Pith. sign in

REVIEW 2 cited by

Acoustic-to-articulatory Speech Inversion with Multi-task Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2205.13755 v1 pith:RLNPR3KW submitted 2022-05-27 eess.AS

classification eess.AS
keywords speechacousticinversionlearningacoustic-to-articulatoryarticulatorymodelmulti-task
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Multi-task learning (MTL) frameworks have proven to be effective in diverse speech related tasks like automatic speech recognition (ASR) and speech emotion recognition. This paper proposes a MTL framework to perform acoustic-to-articulatory speech inversion by simultaneously learning an acoustic to phoneme mapping as a shared task. We use the Haskins Production Rate Comparison (HPRC) database which has both the electromagnetic articulography (EMA) data and the corresponding phonetic transcriptions. Performance of the system was measured by computing the correlation between estimated and actual tract variables (TVs) from the acoustic to articulatory speech inversion task. The proposed MTL based Bidirectional Gated Recurrent Neural Network (RNN) model learns to map the input acoustic features to nine TVs while outperforming the baseline model trained to perform only acoustic to articulatory inversion.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Enhancing Acoustic-to-Articulatory Speech Inversion by Incorporating Nasality

    eess.AS 2025-06 conditional novelty 6.0 of 10

    A multi-task speech inversion model jointly estimates oral tract variables, velar port constriction via nasalance, and source features, achieving modest accuracy gains over independent models.

  2. Acoustic to Articulatory Speech Inversion for Children with Velopharyngeal Insufficiency

    eess.AS 2025-09 conditional novelty 4.0 of 10

    An audio-only speech inversion system, fine-tuned on children with velopharyngeal insufficiency, estimates nasalance with improved correlation over a prior adult baseline.

Pith tools