Pith. sign in

REVIEW

Multimodal Segmentation for Vocal Tract Modeling

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.15754 v1 pith:RLYPOOI6 submitted 2024-06-22 cs.CV cs.CLcs.LGcs.SDeess.AS

classification cs.CVcs.CLcs.LGcs.SDeess.AS
keywords vocaltractmodelingrt-mrisegmentationarticulatorsdatasetinternal
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Accurate modeling of the vocal tract is necessary to construct articulatory representations for interpretable speech processing and linguistics. However, vocal tract modeling is challenging because many internal articulators are occluded from external motion capture technologies. Real-time magnetic resonance imaging (RT-MRI) allows measuring precise movements of internal articulators during speech, but annotated datasets of MRI are limited in size due to time-consuming and computationally expensive labeling methods. We first present a deep labeling strategy for the RT-MRI video using a vision-only segmentation approach. We then introduce a multimodal algorithm using audio to improve segmentation of vocal articulators. Together, we set a new benchmark for vocal tract modeling in MRI video segmentation and use this to release labels for a 75-speaker RT-MRI dataset, increasing the amount of labeled public RT-MRI data of the vocal tract by over a factor of 9. The code and dataset labels can be found at \url{rishiraij.github.io/multimodal-mri-avatar/}.

Discussion (0). Continue with ORCID to comment.

Pith tools