Pith. sign in

REVIEW 1 cited by

MultiActor-Audiobook: Zero-Shot Audiobook Generation with Faces and Voices of Multiple Speakers

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2505.13082 v1 pith:K2VAJBOD submitted 2025-05-19 cs.SD cs.AIeess.AS

classification cs.SDcs.AIeess.AS
keywords multiactor-audiobookgenerationprosodyspeakeraudiobookaudiobooksconsistentexpressive
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce MultiActor-Audiobook, a zero-shot approach for generating audiobooks that automatically produces consistent, expressive, and speaker-appropriate prosody, including intonation and emotion. Previous audiobook systems have several limitations: they require users to manually configure the speaker's prosody, read each sentence with a monotonic tone compared to voice actors, or rely on costly training. However, our MultiActor-Audiobook addresses these issues by introducing two novel processes: (1) MSP (**Multimodal Speaker Persona Generation**) and (2) LSI (**LLM-based Script Instruction Generation**). With these two processes, MultiActor-Audiobook can generate more emotionally expressive audiobooks with a consistent speaker prosody without additional training. We compare our system with commercial products, through human and MLLM evaluations, achieving competitive results. Furthermore, we demonstrate the effectiveness of MSP and LSI through ablation studies.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Evaluating the Expressive Appropriateness of Speech in Rich Contexts

    eess.AS 2026-05 unverdicted novelty 7.0 of 10

    CEAEval is a context-aware evaluation system for speech expressive appropriateness, supported by a new Mandarin dataset with multi-dimensional human annotations and a model that outperforms prior systems.

Pith tools