Pith. sign in

Paper Citation Record · LEDGER

Voice Adaptation for Swiss German

As of 7 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 1 inbound Pith citation observation for arXiv:2505.22054.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.22054 v1

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:23:12.222004Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:23:04.116674Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T13:23:12.457477Z

Reference resolution

28 of 28 outbound references displayed

  • verified exact0
  • verified fuzzy17
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9360fa95-2080-4783-9ccf-5cf0dfdfd366 · outbound

This paper cites It is now possible to clone a voice across languages with less than a minute of audio required [3, 4].

Voice Adaptation for Swiss German It is now possible to clone a voice across languages with less than a minute of audio required [3, 4]

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:23:16.935149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:23:04.038461Z digest=sha256:7772a4c6cfa518f354c6b6626515e732212ba0929408b3fe3baa057095432559

Observation 73889d1b-4869-4a5e-805c-de04eb51344b · outbound

This paper cites Voice Adaptation for Swiss German.

Voice Adaptation for Swiss German Voice Adaptation for Swiss German

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T13:23:12.560265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:23:04.116674Z digest=sha256:f4f7bc2fa99670aed92da004561151bb45b90147f4bae8fd8a7881f14b65f5d5

Observation c3ddeacb-2b21-408f-8d4f-af3d3e24ce66 · outbound

This paper cites For our first model, we fine- tuned XTTS-v2 using the SRG and STT4SG-350 data mix.

Voice Adaptation for Swiss German For our first model, we fine- tuned XTTS-v2 using the SRG and STT4SG-350 data mix

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:23:16.674564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:23:04.252180Z digest=sha256:d1286eec699ee6331cc1e770c5603d3cb6a73bc3aaabc0cd62aa4269fe315366

Observation 8cef58c1-db32-4a48-b1e9-610963262ba3 · outbound

This paper cites The learning rate was changed to 6e-5 from the original 5e-5 due to internal tests and listening to the generated audio files.

Voice Adaptation for Swiss German The learning rate was changed to 6e-5 from the original 5e-5 due to internal tests and listening to the generated audio files

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:23:16.458662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:23:04.357760Z digest=sha256:e09fea2e8069df12b200dc75aa008740ac0c7205cffb13caaf6ddb76062b202e

Observation 6f335fa5-5f71-4d6a-899f-8b6aa9af0001 · outbound

This paper cites The training and test sets were pre- partitioned by the dataset authors to ensure speaker indepen- dence, such that no speaker or sample appears in both splits.

Voice Adaptation for Swiss German The training and test sets were pre- partitioned by the dataset authors to ensure speaker indepen- dence, such that no speaker or sample appears in both splits

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:23:16.199244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:23:04.486066Z digest=sha256:d4d98dff56065761879ab52f206602305d912a9f1dbc7dada6420428135feb2f

Observation 0d4607ae-b7db-4300-820e-fe1f53b6300b · outbound

This paper cites We showed that translating Standard German text to Swiss German dialect speech is feasible and yields satisfactory results.

Voice Adaptation for Swiss German We showed that translating Standard German text to Swiss German dialect speech is feasible and yields satisfactory results

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:23:16.007737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:23:04.666971Z digest=sha256:5b208613556a1035d5248e4a226703325b8f87e4f231f2430fe5e27a40bf974e

Observation 07c2fa73-8827-4748-8982-68848deb21f1 · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

Voice Adaptation for Swiss German Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:23:04.803829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:23:04.803829Z digest=sha256:1eb248b519ee785528e3ef8ed89c7bbccf1b3127fec027d8c902e1a8c81f8093

Observation 78f0abd3-8987-4043-b332-5830629f06d2 · outbound

This paper cites VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers.

Voice Adaptation for Swiss German VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:23:05.343214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:23:05.343214Z digest=sha256:0cde56e042dc88fa2a5a4ba5632f750105729e118e252534c54895cc548c5346

Observation f148466c-20d9-4340-a9ed-b9aedfd7a3f6 · outbound

This paper cites Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling.

Voice Adaptation for Swiss German Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:23:07.320866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:23:07.320866Z digest=sha256:ca63363b36679446d34ecbcc798380d8e5dafcefff3bccfd0dccfb192521e00a

Observation c72a5b82-fef3-45bd-bdd5-26bfe6611af9 · outbound

This paper cites XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model,.

Voice Adaptation for Swiss German XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:23:15.733994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:23:07.455851Z digest=sha256:562d45fedf902927ae2fed53e1b9971f7f37b872c49afb7ddfaa221ce48c6685

Observation b9c8fbaa-7db7-4946-8806-96a193e5e12e · outbound

This paper cites Atten- tion is all you need,.

Voice Adaptation for Swiss German Atten- tion is all you need,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:23:15.526781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:23:07.623190Z digest=sha256:417a00917d699b48c0370afff051bcf217605d4422ad7e129f792236f2f76c3b

Observation eee9640c-9dc1-4413-b0db-56dd9709b897 · outbound

This paper cites Libriheavy: A 50,000 hours asr corpus with punc- tuation casing and context,.

Voice Adaptation for Swiss German Libriheavy: A 50,000 hours asr corpus with punc- tuation casing and context,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:23:15.233252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:23:08.110373Z digest=sha256:4b48bd1d5368d44dd0ce4409db92b7b95993cf2a9040d18e91fec9b205b7fa09

Observation 2ff7b0fe-c7bf-4b3f-a4ad-2c442e8cb43e · outbound

This paper cites Wenetspeech4tts: A 12,800-hour mandarin tts corpus for large speech generation model bench- mark,.

Voice Adaptation for Swiss German Wenetspeech4tts: A 12,800-hour mandarin tts corpus for large speech generation model bench- mark,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T13:23:10.048074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:23:10.048074Z digest=sha256:4fbd99fd939f4864b9acc321b36964474aab5d91b74e14931b4855ef4bb44a27

Observation 4f9c09f7-e28a-4e74-951e-cec81e9bae97 · outbound

This paper cites Autoprep: An automatic preprocessing framework for in-the-wild speech data,.

Voice Adaptation for Swiss German Autoprep: An automatic preprocessing framework for in-the-wild speech data,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:23:15.060135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:23:10.189848Z digest=sha256:e830b297b84fc619ae7b2f24948f213380fa9a70cf8b171c87f7e298336b0287

Observation 9610d4f2-2459-4f65-bc20-0db5b7d752f4 · outbound

This paper cites SDS-200: A Swiss German speech to Standard German text corpus,.

Voice Adaptation for Swiss German SDS-200: A Swiss German speech to Standard German text corpus,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:23:14.853461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:23:10.299302Z digest=sha256:3df03bee3ab888bdc4328c6a4cd2b91eb78e9c7eebbd57770418849cceeaed42

Observation ccab4ab0-f4ef-4840-b927-e429cad7b4ef · outbound

This paper cites STT4SG-350: A speech corpus for all Swiss German dialect regions,.

Voice Adaptation for Swiss German STT4SG-350: A speech corpus for all Swiss German dialect regions,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:23:14.674152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:23:10.426587Z digest=sha256:778c2415080e3960d8656bd6ed790c1f89cb7ba5b8b086fa84a37e9d434418dc

Observation c96d4746-dcf6-4885-9db6-e185de5cc42b · outbound

This paper cites Fine-tuning Whisper on Low-Resource Languages for Real-World Applications.

Voice Adaptation for Swiss German Fine-tuning Whisper on Low-Resource Languages for Real-World Applications

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:23:10.577473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:23:10.577473Z digest=sha256:8bee0b8edab91f99df412a39249994a21d5a6450464ae26a6ca24953b81a1f1d

Observation d36cf610-c113-49cd-ae69-a0f96157642f · outbound

This paper cites SwissDial: Parallel Multidialectal Corpus of Spoken Swiss German.

Voice Adaptation for Swiss German SwissDial: Parallel Multidialectal Corpus of Spoken Swiss German

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:23:10.730467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:23:10.730467Z digest=sha256:9f5630532495b433e9d8d2683eef762998a6c158165af799b18b23b64abbcbdc

Observation 07d8835e-427e-4c65-a01e-8b43d7627762 · outbound

This paper cites Natural tts synthesis by condi- tioning wavenet on mel spectrogram predictions,.

Voice Adaptation for Swiss German Natural tts synthesis by condi- tioning wavenet on mel spectrogram predictions,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:23:14.426205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:23:10.900109Z digest=sha256:b099febfd58c137967d5d49858306c7448a21920d96f6e5ca7f91f636934fd8a

Observation 6c2527e8-c2fa-417a-b316-56e2321ca07d · outbound

This paper cites pyannote.audio 2.1 speaker diarization pipeline: prin- ciple, benchmark, and recipe,.

Voice Adaptation for Swiss German pyannote.audio 2.1 speaker diarization pipeline: prin- ciple, benchmark, and recipe,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:23:14.127369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:23:11.089755Z digest=sha256:1c53c7b3511c3934e8470e6b508c441053d7db93a3f36a071e56bca069eb1ae6

Observation 45c10899-69c6-4eb3-9ffc-da8cd68cb9ed · outbound

This paper cites Powerset multi-class cross entropy loss for neural speaker diarization,.

Voice Adaptation for Swiss German Powerset multi-class cross entropy loss for neural speaker diarization,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:23:13.763019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:23:11.230304Z digest=sha256:26c9c1be83527b6dc30d01cc74a5901aa6310af25a4a46ce2588767ce90d4f54

Observation 383a0702-687b-4fad-a8b0-68d07dd8345f · outbound

This paper cites ELAN (Version 6.8) [Computer software],.

Voice Adaptation for Swiss German ELAN (Version 6.8) [Computer software],

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:23:13.405616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:23:11.428780Z digest=sha256:a9f7712cd48cc513ee0c748e8e3009d7f52bcd6d49e87743df14784af223844d

Observation 3cacc949-1138-4e5d-bfb0-bf129a7e2be8 · outbound

This paper cites Robust speech recognition via large-scale weak su- pervision,.

Voice Adaptation for Swiss German Robust speech recognition via large-scale weak su- pervision,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T13:23:11.567094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:23:11.567094Z digest=sha256:33be3bc86a35ee7894755ca35af361527a89d91a4411c37cc2e3f5cbb4bce190

Observation fa6ec5c3-4192-4638-a18f-9b3c9a5926cf · outbound

This paper cites Automatische erkennung schweizerdeutscher dialekte anhand von audiodaten via phonem- transkriptionen,.

Voice Adaptation for Swiss German Automatische erkennung schweizerdeutscher dialekte anhand von audiodaten via phonem- transkriptionen,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:23:13.150766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:23:11.681298Z digest=sha256:6a48fcb8b3fe0161dcf1f876a158a5d3f827ecdac978f4be3e099f9336a11cab

Observation bd712c5f-effc-41a8-9975-fac9d42d2d4b · outbound

This paper cites Simple and Effective Zero-shot Cross-lingual Phoneme Recognition.

Voice Adaptation for Swiss German Simple and Effective Zero-shot Cross-lingual Phoneme Recognition

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T13:23:11.843471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:23:11.843471Z digest=sha256:57bc97a82ac6b3f7d078058b06921c1bb73fdf43620106d1017aa5ad500585b9

Observation 2125b9b7-a794-4ec2-9470-b4b64e7caa41 · outbound

This paper cites Common voice: A massively-multilingual speech corpus,.

Voice Adaptation for Swiss German Common voice: A massively-multilingual speech corpus,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T13:23:11.955633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:23:11.955633Z digest=sha256:d7acb2c9488d90b0d905d0240330b421cfdf5494d2323d3f42a446545f425cd2

Observation 7b48d7f8-d108-4065-92a3-d65c95915eeb · outbound

This paper cites Ecapa2: A hybrid neural net- work architecture and training strategy for robust speaker embed- dings,.

Voice Adaptation for Swiss German Ecapa2: A hybrid neural net- work architecture and training strategy for robust speaker embed- dings,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T13:23:12.102534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:23:12.102534Z digest=sha256:279406019980193f7e4f0a92be07942e5d7883276547c4644143c4cabf05ad7a

Observation e8a41c36-d42a-4dff-8cf8-28f5895874fa · outbound

This paper cites Dialect transfer for Swiss German speech translation,.

Voice Adaptation for Swiss German Dialect transfer for Swiss German speech translation,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:23:12.860467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:23:12.222004Z digest=sha256:bfee119b18d9fdb88ca0d35c17c436958faa8767e01769786216a27970daaa41

Pith citing papers

Observation 73889d1b-4869-4a5e-805c-de04eb51344b · inbound

Voice Adaptation for Swiss German cites this paper.

Voice Adaptation for Swiss German Voice Adaptation for Swiss German

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T13:23:12.560265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:23:04.116674Z digest=sha256:f4f7bc2fa99670aed92da004561151bb45b90147f4bae8fd8a7881f14b65f5d5