Pith. sign in

REVIEW 2 cited by

GLOBE: A High-quality English Corpus with Global Accents for Zero-shot Speaker Adaptive Text-to-Speech

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.14875 v1 pith:XQDQ66AU submitted 2024-06-21 cs.SD eess.AS

classification cs.SDeess.AS
keywords globecorpusspeakeraccentsadaptiveenglishspeakersspeech
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper introduces GLOBE, a high-quality English corpus with worldwide accents, specifically designed to address the limitations of current zero-shot speaker adaptive Text-to-Speech (TTS) systems that exhibit poor generalizability in adapting to speakers with accents. Compared to commonly used English corpora, such as LibriTTS and VCTK, GLOBE is unique in its inclusion of utterances from 23,519 speakers and covers 164 accents worldwide, along with detailed metadata for these speakers. Compared to its original corpus, i.e., Common Voice, GLOBE significantly improves the quality of the speech data through rigorous filtering and enhancement processes, while also populating all missing speaker metadata. The final curated GLOBE corpus includes 535 hours of speech data at a 24 kHz sampling rate. Our benchmark results indicate that the speaker adaptive TTS model trained on the GLOBE corpus can synthesize speech with better speaker similarity and comparable naturalness than that trained on other popular corpora. We will release GLOBE publicly after acceptance. The GLOBE dataset is available at https://globecorpus.github.io/.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Omni-Diffusion: Unified Multimodal Understanding and Generation with Masked Discrete Diffusion

    cs.CV 2026-03 conditional novelty 6.0 of 10

    A unified mask-based discrete diffusion model jointly models text, speech, and image tokens and matches or exceeds several any-to-any and specialist multimodal baselines.

  2. MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model

    eess.AS 2025-01 conditional novelty 6.0 of 10

    A small 70M-parameter codec-based TTS model, combining hierarchical decoding and several stabilization tricks, matches or beats far larger models on expressive reference cloning, especially in speaker similarity.

Pith tools