Pith. sign in

Paper Citation Record · LEDGER

When Humans Growl and Birds Speak: High-Fidelity Voice Conversion from Human to Animal and Designed Sounds

As of 7 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 1 inbound Pith citation observation for arXiv:2505.24336.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.24336 v1

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:30:40.211506Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:30:37.536147Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T12:30:40.292035Z

Reference resolution

36 of 36 outbound references displayed

  • verified exact0
  • verified fuzzy32
  • unresolved2
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ce1e0469-2921-4255-99b6-dbf3dd596347 · outbound

This paper cites When Humans Growl and Birds Speak: High-Fidelity Voice Conversion from Human to Animal and Designed Sounds.

When Humans Growl and Birds Speak: High-Fidelity Voice Conversion from Human to Animal and Designed Sounds When Humans Growl and Birds Speak: High-Fidelity Voice Conversion from Human to Animal and Designed Sounds

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T12:30:40.356114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:30:37.536147Z digest=sha256:1fcaf3f4dc94242a71004ed9fb59da7408c55b8f28d6cfae2297d7405ac276f0

Observation ce5c8d29-6113-4bdf-ba07-9b0dde979e60 · outbound

This paper cites designed voices.

When Humans Growl and Birds Speak: High-Fidelity Voice Conversion from Human to Animal and Designed Sounds designed voices

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:30:46.525352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:30:37.627362Z digest=sha256:278d9e9c7f83f5195d677b411a5fa42356dcd7ec38bbdca0a32537119c3b8014

Observation 2ed239d4-430e-48be-b6e8-06d96807855c · outbound

This paper cites Figure 2 presents an overview of the workflow of the system.

When Humans Growl and Birds Speak: High-Fidelity Voice Conversion from Human to Animal and Designed Sounds Figure 2 presents an overview of the workflow of the system

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:30:46.353849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:30:37.720404Z digest=sha256:298c5132171dd2cbcc4a367b935a7ba411cd823b703f6635e0d75bd315c828a1

Observation 3ea5963b-3398-47b5-8a84-626c63b2b273 · outbound

This paper cites w/o PP” indicates the proposed model using conventional (speech-focused) preprocessing method instead of our proposed pipeline, “w/ SEED.

When Humans Growl and Birds Speak: High-Fidelity Voice Conversion from Human to Animal and Designed Sounds w/o PP” indicates the proposed model using conventional (speech-focused) preprocessing method instead of our proposed pipeline, “w/ SEED

Reference 4

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T12:30:46.208169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:30:37.814312Z digest=sha256:bcb1df25e610e9570fb8ef7b1cb436df537c89e68ff4aad1b2239181a71947a0

Observation 65575883-45c6-4fdb-bde1-57649d26e2c2 · outbound

This paper cites It accu- rately captured transient-rich signals and wide-frequency–range of non-human sounds.

When Humans Growl and Birds Speak: High-Fidelity Voice Conversion from Human to Animal and Designed Sounds It accu- rately captured transient-rich signals and wide-frequency–range of non-human sounds

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:30:46.033796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:30:37.934981Z digest=sha256:b1c6ab569f19930c571dae5238819f235c8a7cac34ffeb1567ba04c28cdea984

Observation 34b340fd-8d36-4971-ba1f-6580f5437c6f · outbound

This paper cites an unresolved cited work.

When Humans Growl and Birds Speak: High-Fidelity Voice Conversion from Human to Animal and Designed Sounds Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:30:45.875286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:30:38.010806Z digest=sha256:9f1ca94f4e529c5468a9cebb6a2ce4db1a32c3e6b1272452d832c8ff3282a4bb

Observation afd7a08d-8a6f-4a2b-8c3d-ab7056e6f3e2 · outbound

This paper cites Neural anal- ysis and synthesis: Reconstructing speech from self-supervised representations,.

When Humans Growl and Birds Speak: High-Fidelity Voice Conversion from Human to Animal and Designed Sounds Neural anal- ysis and synthesis: Reconstructing speech from self-supervised representations,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:30:45.766768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:30:38.074547Z digest=sha256:c4375bd19b18b943526f0efde33526601c1be6fcec29b9307e1d07975f2c5c35

Observation 8c33a3df-cd0c-4308-a0d6-a0a40642fb4c · outbound

This paper cites Dddm-vc: Decoupled denoising dif- fusion models with disentangled representation and prior mixup for verified robust voice conversion,.

When Humans Growl and Birds Speak: High-Fidelity Voice Conversion from Human to Animal and Designed Sounds Dddm-vc: Decoupled denoising dif- fusion models with disentangled representation and prior mixup for verified robust voice conversion,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:30:45.625448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:30:38.150161Z digest=sha256:5cabdb5388cde567fb8946269d8bcf4c0b2ae6f8a3c1a62129205b101780b165

Observation 616ff8cc-6499-4ff3-9fb0-56b2e7d1dca3 · outbound

This paper cites Diff-hiervc: Diffusion-based hierar- chical voice conversion with robust pitch generation and masked prior for zero-shot speaker adaptation,.

When Humans Growl and Birds Speak: High-Fidelity Voice Conversion from Human to Animal and Designed Sounds Diff-hiervc: Diffusion-based hierar- chical voice conversion with robust pitch generation and masked prior for zero-shot speaker adaptation,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:30:45.477131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:30:38.207028Z digest=sha256:9c2935d095a8c8fc28be9199f6de37b8bad8d2638b8abcac545cf79ed4813c85

Observation 12ead880-7eaa-4d05-8fee-8d91af3271b7 · outbound

This paper cites Hiervst: Hierarchical adap- tive zero-shot voice style transfer,.

When Humans Growl and Birds Speak: High-Fidelity Voice Conversion from Human to Animal and Designed Sounds Hiervst: Hierarchical adap- tive zero-shot voice style transfer,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:30:45.320632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:30:38.246277Z digest=sha256:efba318bdaa24c374ad28c09412dca1e9a698a89c15f500060e89b56749d2417

Observation 8c3bc694-73bc-4dbf-88aa-a93f29d78fe0 · outbound

This paper cites FreeVC: Towards high-quality text- free one-shot voice conversion,.

When Humans Growl and Birds Speak: High-Fidelity Voice Conversion from Human to Animal and Designed Sounds FreeVC: Towards high-quality text- free one-shot voice conversion,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:30:45.167592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:30:38.294734Z digest=sha256:3a51f961fbbb20678b9114255689a8d64b76584c7455fa00e73536a3237c4759

Observation 9c860285-8e4d-4fb4-8e35-b57f923964e7 · outbound

This paper cites Speak like a dog: Human to non-human creature voice conversion,.

When Humans Growl and Birds Speak: High-Fidelity Voice Conversion from Human to Animal and Designed Sounds Speak like a dog: Human to non-human creature voice conversion,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:30:45.060722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:30:38.371345Z digest=sha256:fe0b831f5053c5a9e46794711ade906656a1a54a0cac3da45c0c01269c46dc30

Observation bb14cb09-195f-4854-9453-ebb30ce6e28c · outbound

This paper cites High-fidelity audio compression with improved rvqgan,.

When Humans Growl and Birds Speak: High-Fidelity Voice Conversion from Human to Animal and Designed Sounds High-fidelity audio compression with improved rvqgan,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:30:44.928302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:30:38.464196Z digest=sha256:6a179d9a7c9dd40ca5221fb2d3619c381a599a41a50c49b1b48eeec0485989e0

Observation 2cce21f5-3ca7-4eae-a73b-202bc91d8348 · outbound

This paper cites DDSP-SFX: Acoustically-guided sound ef- fects generation with differentiable digital signal processing,.

When Humans Growl and Birds Speak: High-Fidelity Voice Conversion from Human to Animal and Designed Sounds DDSP-SFX: Acoustically-guided sound ef- fects generation with differentiable digital signal processing,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:30:44.826311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:30:38.556266Z digest=sha256:a4be6c4b1e195611cdbfadbc88ff316d34cc47d374afa8dcd7b2915581070ec3

Observation da9aadb5-e5db-4be9-abbb-e49a6bd62450 · outbound

This paper cites T-FOLEY: A controllable waveform-domain diffusion model for temporal-event-guided fo- ley sound synthesis,.

When Humans Growl and Birds Speak: High-Fidelity Voice Conversion from Human to Animal and Designed Sounds T-FOLEY: A controllable waveform-domain diffusion model for temporal-event-guided fo- ley sound synthesis,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:30:44.679612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:30:38.655167Z digest=sha256:6dcaa95e178e2861898f37ffe8d22f505e14a75b8712ef1b5bd73a6ed017c2b4

Observation 2d893310-cb1c-4b26-92d0-16dc827ec744 · outbound

This paper cites Dehumaniser2 - creature & monster sound design.

When Humans Growl and Birds Speak: High-Fidelity Voice Conversion from Human to Animal and Designed Sounds Dehumaniser2 - creature & monster sound design

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:30:44.431287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:30:38.713623Z digest=sha256:bf245055545d3866fc9b01dd5bf3bcbea226b4f5ac53423286101d2e7b10db11

Observation 4f2c9c77-7bd6-4d98-8a89-e2e6cbfc3759 · outbound

This paper cites Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,.

When Humans Growl and Birds Speak: High-Fidelity Voice Conversion from Human to Animal and Designed Sounds Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:30:44.162873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:30:38.791781Z digest=sha256:97d1b26a3df7fa2abb3c7cfbbf4dbe11159cdf0bb9a2ca40e0f29567588ee4d5

Observation 3d8050ef-7d75-4ee6-828b-0903422fb25f · outbound

This paper cites Wave-tacotron: Spectrogram-free end-to-end text- to-speech synthesis,.

When Humans Growl and Birds Speak: High-Fidelity Voice Conversion from Human to Animal and Designed Sounds Wave-tacotron: Spectrogram-free end-to-end text- to-speech synthesis,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:30:43.858160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:30:38.886114Z digest=sha256:4993f876b9e860f721702985146931a24508a852a3bc0d33db7b81aec2a0260b

Observation 3a0db536-09c8-4875-899e-52043bb6d076 · outbound

This paper cites A systematic explo- ration of joint-training for singing voice synthesis,.

When Humans Growl and Birds Speak: High-Fidelity Voice Conversion from Human to Animal and Designed Sounds A systematic explo- ration of joint-training for singing voice synthesis,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:30:43.614823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:30:38.972288Z digest=sha256:584927023dde42f19831155661b770a8101afad9b8b8dfb493a7962225ed5e19

Observation a3578694-ab93-4a68-a18f-bfcc9d1b6fbf · outbound

This paper cites Audioldm: Text-to-audio generation with latent diffusion models,.

When Humans Growl and Birds Speak: High-Fidelity Voice Conversion from Human to Animal and Designed Sounds Audioldm: Text-to-audio generation with latent diffusion models,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:30:43.423797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:30:39.069329Z digest=sha256:562ee2b812bbcf988cc136c977ded95ffc4f1291b8eb2215de3ebf1747264253

Observation 47f788c6-406e-49f5-aa01-c01ad0e19694 · outbound

This paper cites V AE with a vampprior,.

When Humans Growl and Birds Speak: High-Fidelity Voice Conversion from Human to Animal and Designed Sounds V AE with a vampprior,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:30:43.227352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:30:39.202650Z digest=sha256:b565036e9f883d88560bacde0ef594ee0ae04ce2999f4c88c435e727a8b9aa17

Observation f3a8ba71-ace4-4c55-bd7e-560ecb0ed47a · outbound

This paper cites Flowtron: an autoregressive flow-based generative network for text-to- speech synthesis,.

When Humans Growl and Birds Speak: High-Fidelity Voice Conversion from Human to Animal and Designed Sounds Flowtron: an autoregressive flow-based generative network for text-to- speech synthesis,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:30:43.013472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:30:39.250419Z digest=sha256:addcf7449f5a03f03ab660da36bd1ed20cbc3e31b7207e8a83655dae260d0cfc

Observation 897c58e1-6066-48dd-bdf2-776706484c74 · outbound

This paper cites Meta-stylespeech: Multi-speaker adaptive text-to-speech generation,.

When Humans Growl and Birds Speak: High-Fidelity Voice Conversion from Human to Animal and Designed Sounds Meta-stylespeech: Multi-speaker adaptive text-to-speech generation,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:30:42.725905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:30:39.326861Z digest=sha256:7804affe829c7d521338c383d274f0bafdec66e6f9ba304823b306685538d637

Observation 3d08d1db-b1fc-490e-b86e-eea02b4d5870 · outbound

This paper cites XLS-R: Self-supervised cross-lingual speech representation learning at scale,.

When Humans Growl and Birds Speak: High-Fidelity Voice Conversion from Human to Animal and Designed Sounds XLS-R: Self-supervised cross-lingual speech representation learning at scale,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:30:42.507870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:30:39.388254Z digest=sha256:aee1f777593469d7b70d051ff85474e92a373ad8d8c8b23aa371052f80387ded

Observation 8e9ed2a6-93eb-42e2-99f7-502a94a8842c · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech representa- tions,.

When Humans Growl and Birds Speak: High-Fidelity Voice Conversion from Human to Animal and Designed Sounds wav2vec 2.0: A framework for self-supervised learning of speech representa- tions,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:30:42.282016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:30:39.440296Z digest=sha256:04795ed03a4ac5a9f857039da627ea521eb15a8fc590c76ada33a21edb08dab8

Observation 3997f8f2-f4a4-458b-822e-191ae80f84b4 · outbound

This paper cites Praat: doing phonetics by computer (Computer program),.

When Humans Growl and Birds Speak: High-Fidelity Voice Conversion from Human to Animal and Designed Sounds Praat: doing phonetics by computer (Computer program),

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:30:42.085126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:30:39.515800Z digest=sha256:4996012d00c69408c800d9491c3c74ae0b31815edc679f65d45bb773fcd6b2f8

Observation 730b4df6-32ed-4e68-b513-296b8b8c32b0 · outbound

This paper cites Yet another algorithm for pitch tracking,.

When Humans Growl and Birds Speak: High-Fidelity Voice Conversion from Human to Animal and Designed Sounds Yet another algorithm for pitch tracking,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:30:39.581396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:30:39.581396Z digest=sha256:ee7d67d1a8441e24e9eb302a7052421539ed9756db8904e0c8b94eebac47ce70

Observation d9a875a8-8f6c-43ba-9fff-cdb6ca32b708 · outbound

This paper cites Crepe: A convolu- tional representation for pitch estimation,.

When Humans Growl and Birds Speak: High-Fidelity Voice Conversion from Human to Animal and Designed Sounds Crepe: A convolu- tional representation for pitch estimation,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:30:41.883587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:30:39.683835Z digest=sha256:c1c4bf30b0d0f4c2d8df53eef15baa80a2435736934e4a49c83a8fb808912a04

Observation e408e754-2ed8-4c3b-b63d-c526218b226c · outbound

This paper cites SPICE: Self-supervised pitch estimation,.

When Humans Growl and Birds Speak: High-Fidelity Voice Conversion from Human to Animal and Designed Sounds SPICE: Self-supervised pitch estimation,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:30:41.665000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:30:39.740688Z digest=sha256:fb044775fe3172e264961c8ee2c0c2b630d24fa85ffbf5bb0afcd60455cd51bc

Observation c8d0484b-7abb-4003-a8ba-58bd678d2958 · outbound

This paper cites PESTO: Pitch estimation with self-supervised transposition-equivariant objec- tive,.

When Humans Growl and Birds Speak: High-Fidelity Voice Conversion from Human to Animal and Designed Sounds PESTO: Pitch estimation with self-supervised transposition-equivariant objec- tive,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:30:41.417805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:30:39.796020Z digest=sha256:70179b9b12a32fdafb5d0ff0593e955338c33d262837e85831661f64d11bd61f

Observation e0481ced-f44c-401f-bd49-53ffd9c136ea · outbound

This paper cites Autoencoding beyond pixels using a learned similarity metric,.

When Humans Growl and Birds Speak: High-Fidelity Voice Conversion from Human to Animal and Designed Sounds Autoencoding beyond pixels using a learned similarity metric,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:30:41.141141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:30:39.852908Z digest=sha256:090183b8b13534ecd65d87a27017e8a1a93ac8212da8e1cc99f182f35b14f8f7

Observation 70e3a205-8925-49ed-9247-09f0d98c6cf3 · outbound

This paper cites Least squares generative adversarial networks,.

When Humans Growl and Birds Speak: High-Fidelity Voice Conversion from Human to Animal and Designed Sounds Least squares generative adversarial networks,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:30:40.942889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:30:39.952430Z digest=sha256:a7db45748621318a4e592576034d21c15c045ae0b6121a29eea821c81d3aec62

Observation dae955db-4b70-4757-9a85-fbd62b8fafdc · outbound

This paper cites Learning discourse-level di- versity for neural dialog models using conditional variational au- toencoders,.

When Humans Growl and Birds Speak: High-Fidelity Voice Conversion from Human to Animal and Designed Sounds Learning discourse-level di- versity for neural dialog models using conditional variational au- toencoders,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:30:40.875316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:30:40.022568Z digest=sha256:d57ad0025b7be720c52bbcde8f0eea64d1bd523bfdeee38197c55921fb46cede

Observation a32ec00f-3a89-454b-aebd-ceb903e25868 · outbound

This paper cites Pro sound effects - core 6 library.

When Humans Growl and Birds Speak: High-Fidelity Voice Conversion from Human to Animal and Designed Sounds Pro sound effects - core 6 library

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:30:40.754692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:30:40.077343Z digest=sha256:d3d2b030df892a7c908f5162f55fadae596b099d5de47a0a5f8bab07fc2130d3

Observation 2ffab1d5-8140-435e-9450-43d62567711c · outbound

This paper cites Robust speech recognition via large-scale weak su- pervision,.

When Humans Growl and Birds Speak: High-Fidelity Voice Conversion from Human to Animal and Designed Sounds Robust speech recognition via large-scale weak su- pervision,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:30:40.611219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:30:40.153968Z digest=sha256:6ee7f8687cd89fedc08d84880a7c77bc402d8d909c61247d5f3f724ec506d231

Observation bcabf7a6-6a07-4780-a4e3-01db46a49f3f · outbound

This paper cites HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis,.

When Humans Growl and Birds Speak: High-Fidelity Voice Conversion from Human to Animal and Designed Sounds HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:30:40.497679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:30:40.211506Z digest=sha256:28c3dbbb38d07abd582a452e7889d3e4e6bcfa287a8ad9e01a0fd150f5da1f76

Pith citing papers

Observation ce1e0469-2921-4255-99b6-dbf3dd596347 · inbound

When Humans Growl and Birds Speak: High-Fidelity Voice Conversion from Human to Animal and Designed Sounds cites this paper.

When Humans Growl and Birds Speak: High-Fidelity Voice Conversion from Human to Animal and Designed Sounds When Humans Growl and Birds Speak: High-Fidelity Voice Conversion from Human to Animal and Designed Sounds

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T12:30:40.356114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:30:37.536147Z digest=sha256:1fcaf3f4dc94242a71004ed9fb59da7408c55b8f28d6cfae2297d7405ac276f0