Pith. sign in

Paper Citation Record · LEDGER

Designed Vocalizations Dataset: Sound-Designed Human and Animal Voices for Non-human Voice Conversion

As of 22 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 1 inbound Pith citation observation for arXiv:2607.20951.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.20951 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T08:57:41.015800Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T08:57:38.957612Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

29 of 29 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation db0ef954-6acb-451c-be74-158df5ec607b · outbound

This paper cites Designed Vocalizations Dataset: Sound-Designed Human and Animal Voices for Non-human Voice Conversion.

Designed Vocalizations Dataset: Sound-Designed Human and Animal Voices for Non-human Voice Conversion Designed Vocalizations Dataset: Sound-Designed Human and Animal Voices for Non-human Voice Conversion

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T08:57:38.957612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:57:38.957612Z digest=sha256:0a45a4ff8113c2f450db2ebc91617b5aa5e204fbc1054c5d58e228fe240f49f0

Observation 15f88e8d-1b25-4ea3-97d5-f4f3af235d39 · outbound

This paper cites an unresolved cited work.

Designed Vocalizations Dataset: Sound-Designed Human and Animal Voices for Non-human Voice Conversion Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T08:57:39.036247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:57:39.036247Z digest=sha256:19e002c635de3f1dd5d775596574e12b7cdda1ff3bf879e9ba4876459612b2d6

Observation 051019da-42f9-46a2-bddb-b71140df6a91 · outbound

This paper cites The experiments assessed model performance on the four seen/unseen combinations of sourcex s and reference xr =x (p) using the test set, following training on the train- ing set.

Designed Vocalizations Dataset: Sound-Designed Human and Animal Voices for Non-human Voice Conversion The experiments assessed model performance on the four seen/unseen combinations of sourcex s and reference xr =x (p) using the test set, following training on the train- ing set

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T08:57:39.120104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:57:39.120104Z digest=sha256:b26fe512eca3bb185619db8e704c96cf319c26823f89bcdb3519770a2a62f4ec

Observation f59ee706-1af8-4933-b177-1f22d4d9501f · outbound

This paper cites The dataset in- cludes raw audio and preset-specific designed variants gener- ated through sound-design operators, covering diverse linguistic and non-linguistic sources.

Designed Vocalizations Dataset: Sound-Designed Human and Animal Voices for Non-human Voice Conversion The dataset in- cludes raw audio and preset-specific designed variants gener- ated through sound-design operators, covering diverse linguistic and non-linguistic sources

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T08:57:39.206188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:57:39.206188Z digest=sha256:3839add3d1d07571c563b9d2297af933d478adb9fc61df5b283985a4f8efe08a

Observation 593894c0-3926-4da9-9b04-1e1d95d07344 · outbound

This paper cites RS-2025- 25441313, Professional AI Talent Development Program for Multimodal AI Agents, Contribution: 50%).

Designed Vocalizations Dataset: Sound-Designed Human and Animal Voices for Non-human Voice Conversion RS-2025- 25441313, Professional AI Talent Development Program for Multimodal AI Agents, Contribution: 50%)

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T08:57:39.300620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:57:39.300620Z digest=sha256:3ccac19fc0163e4d78249a5367e049a87c940db983975c4bcf833db30dc49a5f

Observation bf667cdd-f6c7-4d0a-9a1f-e3d438921604 · outbound

This paper cites The scientific content, analysis, and conclusions were developed by the authors.

Designed Vocalizations Dataset: Sound-Designed Human and Animal Voices for Non-human Voice Conversion The scientific content, analysis, and conclusions were developed by the authors

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T08:57:39.414624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:57:39.414624Z digest=sha256:6cb04143da53d502b87c276114ec4607c36aad409a69f56abf07ddd3abd284d3

Observation 1ceec710-28fe-4844-8cea-c06ab3379490 · outbound

This paper cites Speak like a dog: Human to non-human creature voice conversion,.

Designed Vocalizations Dataset: Sound-Designed Human and Animal Voices for Non-human Voice Conversion Speak like a dog: Human to non-human creature voice conversion,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T08:57:39.500161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:57:39.500161Z digest=sha256:76519a529ac17c5034b1955f88dc6680f1ef524e236d11859bc02674a50d7118

Observation 6f2899f6-01ed-416d-aced-8c7f40098eef · outbound

This paper cites Ai-assisted human-pet artistic musical co-creation for wellness therapy,.

Designed Vocalizations Dataset: Sound-Designed Human and Animal Voices for Non-human Voice Conversion Ai-assisted human-pet artistic musical co-creation for wellness therapy,

Reference 8

Resolution
verified exact
doi, observed 2026-08-01T08:58:18.734855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-01T08:57:39.612881Z digest=sha256:dbc196a6ce2b75fb8ec2dbde0dd9e3616184b8fbc4ab9cae14b0417d39a1e23f

Observation 1144d8a0-7adc-41d8-ac45-9783f7f3d21e · outbound

This paper cites When humans growl and birds speak: High-fidelity voice conversion from human to animal and designed sounds,.

Designed Vocalizations Dataset: Sound-Designed Human and Animal Voices for Non-human Voice Conversion When humans growl and birds speak: High-fidelity voice conversion from human to animal and designed sounds,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T08:57:39.694495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:57:39.694495Z digest=sha256:0994d5a53fbf5d98f7e456b158808eb49475cbec6e812098105fa80cdde07643

Observation 37fa11f5-8eee-4cec-be53-b6045c1b30d0 · outbound

This paper cites Cartoonsing: Unifying human and nonhuman timbres in singing generation,.

Designed Vocalizations Dataset: Sound-Designed Human and Animal Voices for Non-human Voice Conversion Cartoonsing: Unifying human and nonhuman timbres in singing generation,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T08:57:39.775443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:57:39.775443Z digest=sha256:c2438e5bd6a429c970c0b10c5363e8e8a260f6f9db33a7ca021918c46d455015

Observation 793b03e8-ea3c-4056-8974-2bb816976678 · outbound

This paper cites Cstr vctk corpus: English multi-speaker corpus for cstr voice cloning toolkit,.

Designed Vocalizations Dataset: Sound-Designed Human and Animal Voices for Non-human Voice Conversion Cstr vctk corpus: English multi-speaker corpus for cstr voice cloning toolkit,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T08:57:39.838048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:57:39.838048Z digest=sha256:6748dbff9634a53430a5f2aac53539e8eae75f7779391d0ce87a8d9f50bcbd37

Observation cbb2ffe4-b744-4ad8-a5a3-a5177518c525 · outbound

This paper cites LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech.

Designed Vocalizations Dataset: Sound-Designed Human and Animal Voices for Non-human Voice Conversion LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T08:57:39.910949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:57:39.910949Z digest=sha256:ff6118f57398e7ee7ab60f48dd08b48fde7f13f488b690c4c1e79f06a3a172ec

Observation 60afadfa-bd4f-4d70-b88f-0ed7d2ae4e28 · outbound

This paper cites Libri-light: A benchmark for asr with limited or no su- pervision,.

Designed Vocalizations Dataset: Sound-Designed Human and Animal Voices for Non-human Voice Conversion Libri-light: A benchmark for asr with limited or no su- pervision,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T08:57:39.984092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:57:39.984092Z digest=sha256:b445216363057d704e5cd91e235943397eadecf95ac93046c176e20a1fb07b64

Observation 0c4ec660-1768-463d-8887-34e4cb624f23 · outbound

This paper cites Hi-Fi Multi-Speaker English TTS Dataset.

Designed Vocalizations Dataset: Sound-Designed Human and Animal Voices for Non-human Voice Conversion Hi-Fi Multi-Speaker English TTS Dataset

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T08:57:40.034274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:57:40.034274Z digest=sha256:6d268cb0311ea8b39536b2a1438ff6b6f5f3678ab26f657782eddafb396f9ba1

Observation 80d95c99-fb0b-480a-a397-11d413eb16be · outbound

This paper cites Natural tts synthesis by conditioning wavenet on mel spectrogram pre- dictions,.

Designed Vocalizations Dataset: Sound-Designed Human and Animal Voices for Non-human Voice Conversion Natural tts synthesis by conditioning wavenet on mel spectrogram pre- dictions,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T08:57:40.144985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:57:40.144985Z digest=sha256:cb1be9ad33247cf70d3a93b5e7ec1409ec43eea050f2a0133d9f74462069f8f8

Observation d0c683b4-48b1-4d44-8145-b7b42e8bd38a · outbound

This paper cites FastSpeech 2: Fast and High-Quality End-to-End Text to Speech.

Designed Vocalizations Dataset: Sound-Designed Human and Animal Voices for Non-human Voice Conversion FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T08:57:40.230200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:57:40.230200Z digest=sha256:d3f46507e8067a947467feb55caca82721eeb9244933819db9710547d0a740ba

Observation 20d865b3-da13-42d9-9b48-84dc2c9eb5f8 · outbound

This paper cites Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,.

Designed Vocalizations Dataset: Sound-Designed Human and Animal Voices for Non-human Voice Conversion Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T08:57:40.280776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:57:40.280776Z digest=sha256:7795f53c126d0967248788c1d6239094c5b4c282ab8958b962088e4614febac0

Observation f9a22e12-cc4b-4fce-b4a7-5872891e0f9a · outbound

This paper cites Yourtts: Towards zero-shot multi-speaker tts and zero-shot voice conversion for everyone,.

Designed Vocalizations Dataset: Sound-Designed Human and Animal Voices for Non-human Voice Conversion Yourtts: Towards zero-shot multi-speaker tts and zero-shot voice conversion for everyone,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T08:57:40.344794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:57:40.344794Z digest=sha256:8f3c2683671fcb74cc9eb210972b99010946bd19e249275dc928a5ed5557c2f7

Observation 5c022861-d44f-439c-9d20-6c63466a45e7 · outbound

This paper cites Freevc: Towards high-quality text-free one-shot voice conversion,.

Designed Vocalizations Dataset: Sound-Designed Human and Animal Voices for Non-human Voice Conversion Freevc: Towards high-quality text-free one-shot voice conversion,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T08:57:40.401540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:57:40.401540Z digest=sha256:88aacb2e104c30f8258fc981a0489b09c906a432b4218b087d7fd420ede62b35

Observation a057cb02-5773-45ea-9a64-40046e056c16 · outbound

This paper cites Hiervst: Hierarchical adap- tive zero-shot voice style transfer,.

Designed Vocalizations Dataset: Sound-Designed Human and Animal Voices for Non-human Voice Conversion Hiervst: Hierarchical adap- tive zero-shot voice style transfer,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T08:57:40.503856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:57:40.503856Z digest=sha256:290d67261cdfcc37e4d6fcb9f1e68ae0034c0a9655ffbaf1518556a3db5a0d83

Observation fdae1ef9-568e-4abf-90f2-275fe77d8098 · outbound

This paper cites Diff-hiervc: Diffusion-based hierar- chical voice conversion with robust pitch generation and masked prior for zero-shot speaker adaptation,.

Designed Vocalizations Dataset: Sound-Designed Human and Animal Voices for Non-human Voice Conversion Diff-hiervc: Diffusion-based hierar- chical voice conversion with robust pitch generation and masked prior for zero-shot speaker adaptation,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T08:57:40.555395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:57:40.555395Z digest=sha256:9f4e6e889849e5630b0d898ce339c1751eb50b00a6dffcd300922e8c754f321e

Observation c6a4aa27-03e1-4842-a9ea-0ce169265804 · outbound

This paper cites Dddm-vc: Decoupled denoising dif- fusion models with disentangled representation and prior mixup for verified robust voice conversion,.

Designed Vocalizations Dataset: Sound-Designed Human and Animal Voices for Non-human Voice Conversion Dddm-vc: Decoupled denoising dif- fusion models with disentangled representation and prior mixup for verified robust voice conversion,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T08:57:40.623523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:57:40.623523Z digest=sha256:e0354ee4a8c3cb200e78fbca7196884f2ec8e26c1d84a5f2aedb0c2041943105

Observation 93a88d64-2b22-4b1b-b43a-123f6ca52a72 · outbound

This paper cites Dehumaniser 2,.

Designed Vocalizations Dataset: Sound-Designed Human and Animal Voices for Non-human Voice Conversion Dehumaniser 2,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T08:57:40.663277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:57:40.663277Z digest=sha256:0544f5e5c508870640df43f781e55fe2c860691406769a8ee772aac1e6bc6438

Observation d3e89e7e-e71e-4a87-82f1-0b82b29f7774 · outbound

This paper cites Cubase 12 pro,.

Designed Vocalizations Dataset: Sound-Designed Human and Animal Voices for Non-human Voice Conversion Cubase 12 pro,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T08:57:40.726035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:57:40.726035Z digest=sha256:6ef268e75d182f9df9b2174c999fcac643c5d8363466a50ba247702f37559f9b

Observation 456411ab-7685-4a32-94ae-9c8bf1fd4f03 · outbound

This paper cites Freesound datasets: a platform for the creation of open audio datasets,.

Designed Vocalizations Dataset: Sound-Designed Human and Animal Voices for Non-human Voice Conversion Freesound datasets: a platform for the creation of open audio datasets,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T08:57:40.783472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:57:40.783472Z digest=sha256:9d97c331d5d0af2b4a4ed4ebe37eaf175c0dcbbe32929c03db2e3d98ec66ada0

Observation d2fd2a91-f5a8-40ae-8901-be636268fbd9 · outbound

This paper cites High-fidelity audio compression with improved rvqgan,.

Designed Vocalizations Dataset: Sound-Designed Human and Animal Voices for Non-human Voice Conversion High-fidelity audio compression with improved rvqgan,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T08:57:40.848500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:57:40.848500Z digest=sha256:bdfe6698acc7273a4f5843adfb53d1f2ffef15fb5bc08e09b94b4b4225df7d4c

Observation 7dcfb2ee-e89a-494d-bece-fd29b75b586f · outbound

This paper cites BEATs: Audio pre-training with acoustic tokenizers,.

Designed Vocalizations Dataset: Sound-Designed Human and Animal Voices for Non-human Voice Conversion BEATs: Audio pre-training with acoustic tokenizers,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T08:57:40.903670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:57:40.903670Z digest=sha256:cd92bf74ae00efd3b65e9e9876b0e967fdbe08f5023d740f65bd730ea5052a28

Observation 8a3178cb-68dc-4196-bb47-a2b9f0436af3 · outbound

This paper cites SALMONN: Towards generic hearing abilities for large language models,.

Designed Vocalizations Dataset: Sound-Designed Human and Animal Voices for Non-human Voice Conversion SALMONN: Towards generic hearing abilities for large language models,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T08:57:40.958939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:57:40.958939Z digest=sha256:ba49d247e6c365f872691fd061ac2c6dddc83480989c90577e7f41feef7cdec7

Observation b08caead-7770-44e2-bb24-01fdd9c45c2c · outbound

This paper cites Robust speech recognition via large-scale weak su- pervision,.

Designed Vocalizations Dataset: Sound-Designed Human and Animal Voices for Non-human Voice Conversion Robust speech recognition via large-scale weak su- pervision,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T08:57:41.015800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:57:41.015800Z digest=sha256:f515c4c545af07cafdcd41e9afc854ccdecf0549ee4e9630f25f16130d02683a

Pith citing papers

Observation db0ef954-6acb-451c-be74-158df5ec607b · inbound

Designed Vocalizations Dataset: Sound-Designed Human and Animal Voices for Non-human Voice Conversion cites this paper.

Designed Vocalizations Dataset: Sound-Designed Human and Animal Voices for Non-human Voice Conversion Designed Vocalizations Dataset: Sound-Designed Human and Animal Voices for Non-human Voice Conversion

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T08:57:38.957612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:57:38.957612Z digest=sha256:0a45a4ff8113c2f450db2ebc91617b5aa5e204fbc1054c5d58e228fe240f49f0