Imitation Learning for Elder-Facing Speech Synthesis

· 2026 · cs.SD · arXiv 2606.21053

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it

open full Pith review browse 1 citing papers arXiv PDF

abstract

Recent advances in text-to-speech (TTS) synthesis have achieved highly natural and expressive speech generation. However, these systems are designed for general adults and overlook older adults' speech comprehension needs due to age-related sensory and cognitive decline. Prior work involves older adults by collecting preference feedback to tune model parameters. However, obtaining sufficient preference data is costly and difficult, as older adults quickly become fatigued during collection. In this paper, we propose a novel imitation learning (IL) framework to learn TTS models from expert demonstrations. We further improve Group Relative Policy Optimization (GRPO) with two-stage on-policy reward learning (OPRL) to mitigate reward hacking under limited supervision from expert demonstration. Experimental results show that GRPO w/ OPRL outperforms GRPO and supervised baselines in objective and subjective metrics. Audio samples are available at https://dongru1.github.io/demo/im-efss

representative citing papers

Imitation Learning for Elder-Facing Speech Synthesis

cs.SD · 2026-06-19 · unverdicted · novelty 5.0

An imitation learning approach with two-stage on-policy reward learning enhances TTS for elderly listeners and outperforms standard GRPO and supervised baselines.

citing papers explorer

Showing 1 of 1 citing paper.

Imitation Learning for Elder-Facing Speech Synthesis cs.SD · 2026-06-19 · unverdicted · none · ref 1 · internal anchor
An imitation learning approach with two-stage on-policy reward learning enhances TTS for elderly listeners and outperforms standard GRPO and supervised baselines.

Imitation Learning for Elder-Facing Speech Synthesis

fields

years

verdicts

representative citing papers

citing papers explorer