Pith. sign in

Paper Citation Record · LEDGER

Voice Adaptation for Swiss German

As of 19 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 1 inbound Pith citation observation for arXiv:2505.22054.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.22054 v1

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:23:12.222004Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:23:04.116674Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T13:23:12.457477Z

Reference resolution

28 of 28 outbound references displayed

  • verified exact0
  • verified fuzzy17
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9360fa95-2080-4783-9ccf-5cf0dfdfd366 · outbound

This paper cites It is now possible to clone a voice across languages with less than a minute of audio required [3, 4].

Voice Adaptation for Swiss German It is now possible to clone a voice across languages with less than a minute of audio required [3, 4]

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:23:16.935149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:23:04.038461Z digest=sha256:212df0f9b6dff58cb8f02b1efc4eafbdb36adffabad4dbff48d68ce5b0cfb252

Observation 73889d1b-4869-4a5e-805c-de04eb51344b · outbound

This paper cites Voice Adaptation for Swiss German.

Voice Adaptation for Swiss German Voice Adaptation for Swiss German

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T13:23:12.560265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:23:04.116674Z digest=sha256:4678b039cdada35304198cf34c10016390a27a7bb89a38f3e1a249f2ecb01d17

Observation c3ddeacb-2b21-408f-8d4f-af3d3e24ce66 · outbound

This paper cites For our first model, we fine- tuned XTTS-v2 using the SRG and STT4SG-350 data mix.

Voice Adaptation for Swiss German For our first model, we fine- tuned XTTS-v2 using the SRG and STT4SG-350 data mix

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:23:16.674564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:23:04.252180Z digest=sha256:8bc137a0abcd4ff2bde06fc1dcab7f37e06bf26f3d99b70665e477afc92f11fa

Observation 8cef58c1-db32-4a48-b1e9-610963262ba3 · outbound

This paper cites The learning rate was changed to 6e-5 from the original 5e-5 due to internal tests and listening to the generated audio files.

Voice Adaptation for Swiss German The learning rate was changed to 6e-5 from the original 5e-5 due to internal tests and listening to the generated audio files

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:23:16.458662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:23:04.357760Z digest=sha256:1c5371c4d982e5ffc90ec3b9a63a8a8ae9dab758e78b6f3d689926c9737272a1

Observation 6f335fa5-5f71-4d6a-899f-8b6aa9af0001 · outbound

This paper cites The training and test sets were pre- partitioned by the dataset authors to ensure speaker indepen- dence, such that no speaker or sample appears in both splits.

Voice Adaptation for Swiss German The training and test sets were pre- partitioned by the dataset authors to ensure speaker indepen- dence, such that no speaker or sample appears in both splits

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:23:16.199244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:23:04.486066Z digest=sha256:ebe1107034e09b3aac21c403d9e35555648e381195c7232e3dbf315ad63511c4

Observation 0d4607ae-b7db-4300-820e-fe1f53b6300b · outbound

This paper cites We showed that translating Standard German text to Swiss German dialect speech is feasible and yields satisfactory results.

Voice Adaptation for Swiss German We showed that translating Standard German text to Swiss German dialect speech is feasible and yields satisfactory results

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:23:16.007737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:23:04.666971Z digest=sha256:db9014ccb1ce1875e18da502efc4b4b3d7c2a31dd37e447cfb93f509d5804ed7

Observation 07c2fa73-8827-4748-8982-68848deb21f1 · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

Voice Adaptation for Swiss German Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:23:04.803829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:23:04.803829Z digest=sha256:a83c9a67842f1c094396b0a643232681d1970d694543a38494eaf2ebe5d797ba

Observation 78f0abd3-8987-4043-b332-5830629f06d2 · outbound

This paper cites VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers.

Voice Adaptation for Swiss German VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:23:05.343214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:23:05.343214Z digest=sha256:78f95c4a6c4ead847a23c92ddbb3eee8daff0c439e2a66270588d1b6881b9cfc

Observation f148466c-20d9-4340-a9ed-b9aedfd7a3f6 · outbound

This paper cites Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling.

Voice Adaptation for Swiss German Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:23:07.320866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:23:07.320866Z digest=sha256:9c70cb0d4c85a6f627e58dd55d33ce5a66a05b48935e70b938313b97ebb85217

Observation c72a5b82-fef3-45bd-bdd5-26bfe6611af9 · outbound

This paper cites XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model,.

Voice Adaptation for Swiss German XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:23:15.733994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:23:07.455851Z digest=sha256:949e02a42f6b3396b4521976040df7dfef026c19d3dfa5bd3b9ead4e37c446be

Observation b9c8fbaa-7db7-4946-8806-96a193e5e12e · outbound

This paper cites Atten- tion is all you need,.

Voice Adaptation for Swiss German Atten- tion is all you need,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:23:15.526781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:23:07.623190Z digest=sha256:e98d5d6f2c66c3f39d3d03444186f1fe276b824fec2ade27c85fffa1ad9319cf

Observation eee9640c-9dc1-4413-b0db-56dd9709b897 · outbound

This paper cites Libriheavy: A 50,000 hours asr corpus with punc- tuation casing and context,.

Voice Adaptation for Swiss German Libriheavy: A 50,000 hours asr corpus with punc- tuation casing and context,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:23:15.233252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:23:08.110373Z digest=sha256:e30024a3e66f8aaefd9ecbf16486b8e654e2a58168840f96e47acb4c78ddcfe0

Observation 2ff7b0fe-c7bf-4b3f-a4ad-2c442e8cb43e · outbound

This paper cites Wenetspeech4tts: A 12,800-hour mandarin tts corpus for large speech generation model bench- mark,.

Voice Adaptation for Swiss German Wenetspeech4tts: A 12,800-hour mandarin tts corpus for large speech generation model bench- mark,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T13:23:10.048074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:23:10.048074Z digest=sha256:d5b410eb08e5bfa27ade0ace80cec08170955e5e7a7f75ca8288cd8d5c3df734

Observation 4f9c09f7-e28a-4e74-951e-cec81e9bae97 · outbound

This paper cites Autoprep: An automatic preprocessing framework for in-the-wild speech data,.

Voice Adaptation for Swiss German Autoprep: An automatic preprocessing framework for in-the-wild speech data,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:23:15.060135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:23:10.189848Z digest=sha256:b05a2defe047831cd735dddc983c0ca31353377cb2e6d6e7bd13adbf7b42d8c9

Observation 9610d4f2-2459-4f65-bc20-0db5b7d752f4 · outbound

This paper cites SDS-200: A Swiss German speech to Standard German text corpus,.

Voice Adaptation for Swiss German SDS-200: A Swiss German speech to Standard German text corpus,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:23:14.853461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:23:10.299302Z digest=sha256:3d354c54ab08af470dd34c1e6851a684eb32cd0e0ecde984c2b6ecd919421e6c

Observation ccab4ab0-f4ef-4840-b927-e429cad7b4ef · outbound

This paper cites STT4SG-350: A speech corpus for all Swiss German dialect regions,.

Voice Adaptation for Swiss German STT4SG-350: A speech corpus for all Swiss German dialect regions,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:23:14.674152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:23:10.426587Z digest=sha256:2d6cf3d158ed382ac6b5d802a0e5cdde88618a559146bc4c706999e74089dccb

Observation c96d4746-dcf6-4885-9db6-e185de5cc42b · outbound

This paper cites Fine-tuning Whisper on Low-Resource Languages for Real-World Applications.

Voice Adaptation for Swiss German Fine-tuning Whisper on Low-Resource Languages for Real-World Applications

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:23:10.577473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:23:10.577473Z digest=sha256:e186855eef1b013e5aa0eda8e7af0a64477825d7a25eefa2e6005cddf1f4e45d

Observation d36cf610-c113-49cd-ae69-a0f96157642f · outbound

This paper cites SwissDial: Parallel Multidialectal Corpus of Spoken Swiss German.

Voice Adaptation for Swiss German SwissDial: Parallel Multidialectal Corpus of Spoken Swiss German

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:23:10.730467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:23:10.730467Z digest=sha256:b379bea3a93ab84dd8ebc7e4c76cbea103d837d8dd6564e38588bf275a1eb68e

Observation 07d8835e-427e-4c65-a01e-8b43d7627762 · outbound

This paper cites Natural tts synthesis by condi- tioning wavenet on mel spectrogram predictions,.

Voice Adaptation for Swiss German Natural tts synthesis by condi- tioning wavenet on mel spectrogram predictions,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:23:14.426205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:23:10.900109Z digest=sha256:af50778718804fe3190456c2d100691de8558e0a1849cacecb5d9bb6b7aa58fb

Observation 6c2527e8-c2fa-417a-b316-56e2321ca07d · outbound

This paper cites pyannote.audio 2.1 speaker diarization pipeline: prin- ciple, benchmark, and recipe,.

Voice Adaptation for Swiss German pyannote.audio 2.1 speaker diarization pipeline: prin- ciple, benchmark, and recipe,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:23:14.127369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:23:11.089755Z digest=sha256:6ae22635389816152a55ff5ff8f5d36884cc33f3168cacb9cd606038698ef64d

Observation 45c10899-69c6-4eb3-9ffc-da8cd68cb9ed · outbound

This paper cites Powerset multi-class cross entropy loss for neural speaker diarization,.

Voice Adaptation for Swiss German Powerset multi-class cross entropy loss for neural speaker diarization,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:23:13.763019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:23:11.230304Z digest=sha256:c28185444bf811b3a7c382d8366059bad709abf7353612e6e75d1ff87d0a68db

Observation 383a0702-687b-4fad-a8b0-68d07dd8345f · outbound

This paper cites ELAN (Version 6.8) [Computer software],.

Voice Adaptation for Swiss German ELAN (Version 6.8) [Computer software],

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:23:13.405616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:23:11.428780Z digest=sha256:f916fa4a752e49b826a1e8ed56509a9ac5c695ddd264eeffc6126eb0832eb170

Observation 3cacc949-1138-4e5d-bfb0-bf129a7e2be8 · outbound

This paper cites Robust speech recognition via large-scale weak su- pervision,.

Voice Adaptation for Swiss German Robust speech recognition via large-scale weak su- pervision,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T13:23:11.567094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:23:11.567094Z digest=sha256:0268254e23fdc6b4f0ba26405fdba332a2f614bd92441fba715145be86e0517a

Observation fa6ec5c3-4192-4638-a18f-9b3c9a5926cf · outbound

This paper cites Automatische erkennung schweizerdeutscher dialekte anhand von audiodaten via phonem- transkriptionen,.

Voice Adaptation for Swiss German Automatische erkennung schweizerdeutscher dialekte anhand von audiodaten via phonem- transkriptionen,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:23:13.150766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:23:11.681298Z digest=sha256:4058eeb011a20b87e6448ab13fff710a8a75bc620096b57351044017dc3f6076

Observation bd712c5f-effc-41a8-9975-fac9d42d2d4b · outbound

This paper cites Simple and Effective Zero-shot Cross-lingual Phoneme Recognition.

Voice Adaptation for Swiss German Simple and Effective Zero-shot Cross-lingual Phoneme Recognition

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T13:23:11.843471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:23:11.843471Z digest=sha256:c4f9a3a7db5486cdae28ebdcd35878487e6f5046813ffddd0d419044fc030dbd

Observation 2125b9b7-a794-4ec2-9470-b4b64e7caa41 · outbound

This paper cites Common voice: A massively-multilingual speech corpus,.

Voice Adaptation for Swiss German Common voice: A massively-multilingual speech corpus,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T13:23:11.955633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:23:11.955633Z digest=sha256:e0e08264c4f0bdd0dbcf89bdebe13bf602c5ac04ba36ec591e6485a764e6d75d

Observation 7b48d7f8-d108-4065-92a3-d65c95915eeb · outbound

This paper cites Ecapa2: A hybrid neural net- work architecture and training strategy for robust speaker embed- dings,.

Voice Adaptation for Swiss German Ecapa2: A hybrid neural net- work architecture and training strategy for robust speaker embed- dings,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T13:23:12.102534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:23:12.102534Z digest=sha256:701665058edfb820dcd5fba83cc07b377c4ad507c76742a5f550ef8be0cbd2ab

Observation e8a41c36-d42a-4dff-8cf8-28f5895874fa · outbound

This paper cites Dialect transfer for Swiss German speech translation,.

Voice Adaptation for Swiss German Dialect transfer for Swiss German speech translation,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:23:12.860467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:23:12.222004Z digest=sha256:c3e863e2edfe18b68e6ba746cc5a437ecce76771d34fa69fe7ce59beabcae08d

Pith citing papers

Observation 73889d1b-4869-4a5e-805c-de04eb51344b · inbound

Voice Adaptation for Swiss German cites this paper.

Voice Adaptation for Swiss German Voice Adaptation for Swiss German

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T13:23:12.560265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T13:23:04.116674Z digest=sha256:4678b039cdada35304198cf34c10016390a27a7bb89a38f3e1a249f2ecb01d17