Pith. sign in

REVIEW 2 cited by

SwissDial: Parallel Multidialectal Corpus of Spoken Swiss German

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2103.11401 v1 pith:GAEPJP34 submitted 2021-03-21 cs.CL

classification cs.CL
keywords germanswisscorpusdialectsannotatedparallelspokenstandard
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Swiss German is a dialect continuum whose natively acquired dialects significantly differ from the formal variety of the language. These dialects are mostly used for verbal communication and do not have standard orthography. This has led to a lack of annotated datasets, rendering the use of many NLP methods infeasible. In this paper, we introduce the first annotated parallel corpus of spoken Swiss German across 8 major dialects, plus a Standard German reference. Our goal has been to create and to make available a basic dataset for employing data-driven NLP applications in Swiss German. We present our data collection procedure in detail and validate the quality of our corpus by conducting experiments with the recent neural models for speech synthesis.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Advancing STT for Low-Resource Real-World Speech

    cs.CL 2025-06 conditional novelty 6.0 of 10

    Fine-tuning Whisper on a new 303-hour corpus of spontaneous Swiss German broadcast speech improves WER from 21% to 17% for large-v3 and outperforms models trained on sentence-level corpora.

  2. Voice Adaptation for Swiss German

    cs.CL 2025-05 conditional novelty 5.0 of 10

    Fine-tuning XTTS-v2 on about 5,000 hours of weakly labeled Swiss podcast audio produces a voice adaptation model that renders Standard German text in seven Swiss German dialect regions with near-reference quality in h...

Pith tools