Pith. sign in

Paper Citation Record · LEDGER

Voice Adaptation for Swiss German

As of 12 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 1 inbound Pith citation observation for arXiv:2505.22054.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.22054 v1

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:23:12.222004Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:23:04.116674Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T13:23:12.457477Z

Reference resolution

28 of 28 outbound references displayed

  • verified exact0
  • verified fuzzy17
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9360fa95-2080-4783-9ccf-5cf0dfdfd366 · outbound

This paper cites It is now possible to clone a voice across languages with less than a minute of audio required [3, 4].

Voice Adaptation for Swiss German It is now possible to clone a voice across languages with less than a minute of audio required [3, 4]

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:23:16.935149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T13:23:04.038461Z digest=sha256:05f73de2a00e25a6e12eba48cae8f6cb55c7db2eaf4a2e80c5e184258d936e9d

Observation 73889d1b-4869-4a5e-805c-de04eb51344b · outbound

This paper cites Voice Adaptation for Swiss German.

Voice Adaptation for Swiss German Voice Adaptation for Swiss German

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T13:23:12.560265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T13:23:04.116674Z digest=sha256:183d56a22c07e7e62fa02145256bef030b2101f7c64ef5956ee07c20580ec486

Observation c3ddeacb-2b21-408f-8d4f-af3d3e24ce66 · outbound

This paper cites For our first model, we fine- tuned XTTS-v2 using the SRG and STT4SG-350 data mix.

Voice Adaptation for Swiss German For our first model, we fine- tuned XTTS-v2 using the SRG and STT4SG-350 data mix

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:23:16.674564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T13:23:04.252180Z digest=sha256:fb8d5264ba82aeb145924ed6cb315409f593dac8373944d978f7cb7fba145540

Observation 8cef58c1-db32-4a48-b1e9-610963262ba3 · outbound

This paper cites The learning rate was changed to 6e-5 from the original 5e-5 due to internal tests and listening to the generated audio files.

Voice Adaptation for Swiss German The learning rate was changed to 6e-5 from the original 5e-5 due to internal tests and listening to the generated audio files

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:23:16.458662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T13:23:04.357760Z digest=sha256:3675181d9a8f1777ff94d74945ab9272e883cec369815218ab3617b1af4af113

Observation 6f335fa5-5f71-4d6a-899f-8b6aa9af0001 · outbound

This paper cites The training and test sets were pre- partitioned by the dataset authors to ensure speaker indepen- dence, such that no speaker or sample appears in both splits.

Voice Adaptation for Swiss German The training and test sets were pre- partitioned by the dataset authors to ensure speaker indepen- dence, such that no speaker or sample appears in both splits

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:23:16.199244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T13:23:04.486066Z digest=sha256:c33f967a71c098956b464f0ab304bd5d45716422ce1a67d5e13591b3434eb2fc

Observation 0d4607ae-b7db-4300-820e-fe1f53b6300b · outbound

This paper cites We showed that translating Standard German text to Swiss German dialect speech is feasible and yields satisfactory results.

Voice Adaptation for Swiss German We showed that translating Standard German text to Swiss German dialect speech is feasible and yields satisfactory results

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:23:16.007737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T13:23:04.666971Z digest=sha256:7e67428e4ccd8a09e4a01d1b011b650cc3bca5b7a9673428653dec0489abafb9

Observation 07c2fa73-8827-4748-8982-68848deb21f1 · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

Voice Adaptation for Swiss German Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:23:04.803829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:23:04.803829Z digest=sha256:91318f3d17cdaa62d5bdcd023179422e68de2c75a34181ea556cf885b96cfbac

Observation 78f0abd3-8987-4043-b332-5830629f06d2 · outbound

This paper cites VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers.

Voice Adaptation for Swiss German VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:23:05.343214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:23:05.343214Z digest=sha256:1358a1fc011af2fd05ffe252da5d2c8cf6e99933a7fcdda33b6644db06765709

Observation f148466c-20d9-4340-a9ed-b9aedfd7a3f6 · outbound

This paper cites Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling.

Voice Adaptation for Swiss German Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:23:07.320866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:23:07.320866Z digest=sha256:f84c4b0454c34fe15fecfb921cf02468bc89e3ee5283e4d3f4d9c96419b6a0bc

Observation c72a5b82-fef3-45bd-bdd5-26bfe6611af9 · outbound

This paper cites XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model,.

Voice Adaptation for Swiss German XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:23:15.733994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T13:23:07.455851Z digest=sha256:4bcc500b9e538b6583503242926d451da985cb2da62758eaad334e091f23157d

Observation b9c8fbaa-7db7-4946-8806-96a193e5e12e · outbound

This paper cites Atten- tion is all you need,.

Voice Adaptation for Swiss German Atten- tion is all you need,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:23:15.526781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T13:23:07.623190Z digest=sha256:12a98501ad8116cdd62b3d849a126c0646765a1234dc53f6982db1807263def3

Observation eee9640c-9dc1-4413-b0db-56dd9709b897 · outbound

This paper cites Libriheavy: A 50,000 hours asr corpus with punc- tuation casing and context,.

Voice Adaptation for Swiss German Libriheavy: A 50,000 hours asr corpus with punc- tuation casing and context,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:23:15.233252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T13:23:08.110373Z digest=sha256:3ce043e42618e79b4a97db2ad18e1e349611a4d490b6388ea7d5fbfcfe578b80

Observation 2ff7b0fe-c7bf-4b3f-a4ad-2c442e8cb43e · outbound

This paper cites Wenetspeech4tts: A 12,800-hour mandarin tts corpus for large speech generation model bench- mark,.

Voice Adaptation for Swiss German Wenetspeech4tts: A 12,800-hour mandarin tts corpus for large speech generation model bench- mark,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T13:23:10.048074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:23:10.048074Z digest=sha256:cd43536c71e6344e2b66a2a0cc2be326b9f072fcaf1cd41ec9b0d05e606bd272

Observation 4f9c09f7-e28a-4e74-951e-cec81e9bae97 · outbound

This paper cites Autoprep: An automatic preprocessing framework for in-the-wild speech data,.

Voice Adaptation for Swiss German Autoprep: An automatic preprocessing framework for in-the-wild speech data,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:23:15.060135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T13:23:10.189848Z digest=sha256:3b99e42c4d1ab807c3b96ddc8085babc9914594d06a75ca61be8de1852e443ed

Observation 9610d4f2-2459-4f65-bc20-0db5b7d752f4 · outbound

This paper cites SDS-200: A Swiss German speech to Standard German text corpus,.

Voice Adaptation for Swiss German SDS-200: A Swiss German speech to Standard German text corpus,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:23:14.853461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T13:23:10.299302Z digest=sha256:6b3b64c9f67d70db4db30d0646c1278c62b09c1ab5f62f347a49e04743d21be3

Observation ccab4ab0-f4ef-4840-b927-e429cad7b4ef · outbound

This paper cites STT4SG-350: A speech corpus for all Swiss German dialect regions,.

Voice Adaptation for Swiss German STT4SG-350: A speech corpus for all Swiss German dialect regions,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:23:14.674152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T13:23:10.426587Z digest=sha256:02fbbf68a18a2b9de375e0f87afce1ea6899a1909f24825d7c119d024cf62541

Observation c96d4746-dcf6-4885-9db6-e185de5cc42b · outbound

This paper cites Fine-tuning Whisper on Low-Resource Languages for Real-World Applications.

Voice Adaptation for Swiss German Fine-tuning Whisper on Low-Resource Languages for Real-World Applications

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:23:10.577473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:23:10.577473Z digest=sha256:e44c89ac8c578f7f6b72b8056a578c99f75905e326db1e8b1e1a5e759a122f6d

Observation d36cf610-c113-49cd-ae69-a0f96157642f · outbound

This paper cites SwissDial: Parallel Multidialectal Corpus of Spoken Swiss German.

Voice Adaptation for Swiss German SwissDial: Parallel Multidialectal Corpus of Spoken Swiss German

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:23:10.730467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:23:10.730467Z digest=sha256:d7be228fabb61319638aae71ea407c380060fb396ef8964bf83ea46eea5c285b

Observation 07d8835e-427e-4c65-a01e-8b43d7627762 · outbound

This paper cites Natural tts synthesis by condi- tioning wavenet on mel spectrogram predictions,.

Voice Adaptation for Swiss German Natural tts synthesis by condi- tioning wavenet on mel spectrogram predictions,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:23:14.426205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T13:23:10.900109Z digest=sha256:b5b5f06ff301fb96d3360646442ed975c68b9e3abbc37e0a4df4947028db7c7d

Observation 6c2527e8-c2fa-417a-b316-56e2321ca07d · outbound

This paper cites pyannote.audio 2.1 speaker diarization pipeline: prin- ciple, benchmark, and recipe,.

Voice Adaptation for Swiss German pyannote.audio 2.1 speaker diarization pipeline: prin- ciple, benchmark, and recipe,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:23:14.127369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T13:23:11.089755Z digest=sha256:f7155ab0db87da62ab5ec6639eab465c7d09c4616a870c77ebba938b6c5ddf07

Observation 45c10899-69c6-4eb3-9ffc-da8cd68cb9ed · outbound

This paper cites Powerset multi-class cross entropy loss for neural speaker diarization,.

Voice Adaptation for Swiss German Powerset multi-class cross entropy loss for neural speaker diarization,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:23:13.763019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T13:23:11.230304Z digest=sha256:f6f44b04b8f8d18e1d0c540b4fbfe81a5853f789c91c0e9443909931a1979297

Observation 383a0702-687b-4fad-a8b0-68d07dd8345f · outbound

This paper cites ELAN (Version 6.8) [Computer software],.

Voice Adaptation for Swiss German ELAN (Version 6.8) [Computer software],

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:23:13.405616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T13:23:11.428780Z digest=sha256:df7dd945eaf4cb526278fa48b55e1b2aa8a7de2aca5e526e49846cf2a8beaf4d

Observation 3cacc949-1138-4e5d-bfb0-bf129a7e2be8 · outbound

This paper cites Robust speech recognition via large-scale weak su- pervision,.

Voice Adaptation for Swiss German Robust speech recognition via large-scale weak su- pervision,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T13:23:11.567094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:23:11.567094Z digest=sha256:708f623779d3f0d6c7fd41126988d1e93df1903e22b9d5644a8e7f6c08a00a6d

Observation fa6ec5c3-4192-4638-a18f-9b3c9a5926cf · outbound

This paper cites Automatische erkennung schweizerdeutscher dialekte anhand von audiodaten via phonem- transkriptionen,.

Voice Adaptation for Swiss German Automatische erkennung schweizerdeutscher dialekte anhand von audiodaten via phonem- transkriptionen,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:23:13.150766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T13:23:11.681298Z digest=sha256:7a14a59c8685c09ecc84ea44238993789b698451d24f964971d6d1853c2fe433

Observation bd712c5f-effc-41a8-9975-fac9d42d2d4b · outbound

This paper cites Simple and Effective Zero-shot Cross-lingual Phoneme Recognition.

Voice Adaptation for Swiss German Simple and Effective Zero-shot Cross-lingual Phoneme Recognition

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T13:23:11.843471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:23:11.843471Z digest=sha256:99209a49b30dfc5b19d950baad34c4586814ea293c5e75f2ade349c1ff1a8035

Observation 2125b9b7-a794-4ec2-9470-b4b64e7caa41 · outbound

This paper cites Common voice: A massively-multilingual speech corpus,.

Voice Adaptation for Swiss German Common voice: A massively-multilingual speech corpus,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T13:23:11.955633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:23:11.955633Z digest=sha256:754d8e2f6a024f39ea8ebacf1c7d92ca655945231fe23e524dc4e188ae47ff0c

Observation 7b48d7f8-d108-4065-92a3-d65c95915eeb · outbound

This paper cites Ecapa2: A hybrid neural net- work architecture and training strategy for robust speaker embed- dings,.

Voice Adaptation for Swiss German Ecapa2: A hybrid neural net- work architecture and training strategy for robust speaker embed- dings,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T13:23:12.102534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:23:12.102534Z digest=sha256:5c6a34686d9902aa3a9f657f35f1951c42787d51abbfbee9147fbb205190b641

Observation e8a41c36-d42a-4dff-8cf8-28f5895874fa · outbound

This paper cites Dialect transfer for Swiss German speech translation,.

Voice Adaptation for Swiss German Dialect transfer for Swiss German speech translation,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:23:12.860467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T13:23:12.222004Z digest=sha256:a5b58175bfffe20a8ebddeab40b3b886f80ccd440c99e094aa01b4e2aa14e2f3

Pith citing papers

Observation 73889d1b-4869-4a5e-805c-de04eb51344b · inbound

Voice Adaptation for Swiss German cites this paper.

Voice Adaptation for Swiss German Voice Adaptation for Swiss German

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T13:23:12.560265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T13:23:04.116674Z digest=sha256:183d56a22c07e7e62fa02145256bef030b2101f7c64ef5956ee07c20580ec486