Pith. sign in

REVIEW 1 cited by

Voice Conversion for Lombard Speaking Style with Implicit and Explicit Acoustic Feature Conditioning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2507.09310 v1 pith:DDQR5ERQ submitted 2025-07-12 cs.SD cs.CLeess.AS

Voice Conversion for Lombard Speaking Style with Implicit and Explicit Acoustic Feature Conditioning

classification cs.SD cs.CLeess.AS
keywords lombardspeakingstyleacousticspeakerconditioningconversionexplicit
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Text-to-Speech (TTS) systems in Lombard speaking style can improve the overall intelligibility of speech, useful for hearing loss and noisy conditions. However, training those models requires a large amount of data and the Lombard effect is challenging to record due to speaker and noise variability and tiring recording conditions. Voice conversion (VC) has been shown to be a useful augmentation technique to train TTS systems in the absence of recorded data from the target speaker in the target speaking style. In this paper, we are concerned with Lombard speaking style transfer. Our goal is to convert speaker identity while preserving the acoustic attributes that define the Lombard speaking style. We compare voice conversion models with implicit and explicit acoustic feature conditioning. We observe that our proposed implicit conditioning strategy achieves an intelligibility gain comparable to the model conditioned on explicit acoustic features, while also preserving speaker similarity.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Synthesizing the Lombard Effect: Multi-Level Control of Speech Clarity and Vocal Effort in TTS

    cs.SD 2026-06 unverdicted novelty 5.0

    A flow-matching TTS model with pseudo-label training enables continuous disentangled control of vocal effort, articulation, and word-level emphasis to replicate Lombard effect intelligibility gains in noise.