Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 34 inbound Pith citation observations for arXiv:2411.09943.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:58:57.726231Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T08:29:41.883949Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation e4e2bbb6-2e69-4f3a-8511-830be7742da8 · inbound
Kimi-Audio Technical Report Zero-shot Voice Conversion with Diffusion Transformers
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0b9c0f93-9ca5-4759-9b6e-400d1b5d827d · inbound
EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion Zero-shot Voice Conversion with Diffusion Transformers
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4926be1-b5a6-4c69-9098-2f81b5e375b4 · inbound
IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Zero-shot Voice Conversion with Diffusion Transformers
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f9872d1-24ce-41d6-8e04-0265e059ac42 · inbound
De-AntiFake: Rethinking the Protective Perturbations Against Voice Cloning Attacks Zero-shot Voice Conversion with Diffusion Transformers
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6fcb33aa-9703-4594-a6f5-c5c54c7c1d50 · inbound
The Man Behind the Sound: Demystifying Audio Private Attribute Profiling via Multimodal Large Language Model Agents Zero-shot Voice Conversion with Diffusion Transformers
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb5a0ba8-41c0-4959-a554-c5bd674b3d0c · inbound
SpeechFake: A Large-Scale Multilingual Speech Deepfake Dataset Incorporating Cutting-Edge Generation Methods Zero-shot Voice Conversion with Diffusion Transformers
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b728c6e2-5022-4993-a1fa-bde3f00b2e67 · inbound
REF-VC: Robust, Expressive and Fast Zero-Shot Voice Conversion with Diffusion Transformers Zero-shot Voice Conversion with Diffusion Transformers
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac23beba-29e1-4989-9b7f-6b0eebea3e0b · inbound
Semantic-Aware Ship Detection with Vision-Language Integration Zero-shot Voice Conversion with Diffusion Transformers
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c260918-7ea5-4d2c-8c26-e641f28fb380 · inbound
Entropy-based Coarse and Compressed Semantic Speech Representation Learning Zero-shot Voice Conversion with Diffusion Transformers
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0efb3374-c69e-4d13-9310-fc9df2c2a120 · inbound
Intelligent Agents with Emotional Intelligence: Current Trends, Challenges, and Future Prospects Zero-shot Voice Conversion with Diffusion Transformers
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a6cdb647-77bb-46ec-b961-3b52975bdb6c · inbound
QASA: Quality-Aware Semantic Augmentation for Robust Multimodal Sentiment Analysis Zero-shot Voice Conversion with Diffusion Transformers
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c308a249-ce11-45e9-bd90-e015265ee38c · inbound
Universal Speech Content Factorization Zero-shot Voice Conversion with Diffusion Transformers
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c8828e4-e183-4b8f-8777-4669fb22fe49 · inbound
AT-ADD: All-Type Audio Deepfake Detection Challenge Evaluation Plan Zero-shot Voice Conversion with Diffusion Transformers
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c463817e-4ee0-4514-900f-962ebbd3c4e3 · inbound
X-VC: Zero-shot Streaming Voice Conversion in Codec Space Zero-shot Voice Conversion with Diffusion Transformers
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 48646aca-6227-401e-acaa-d2dcad05656b · inbound
From Seeing it to Experiencing it: Interactive Evaluation of Intersectional Voice Bias in Human-AI Speech Interaction Zero-shot Voice Conversion with Diffusion Transformers
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a5ef4e07-2da1-43d8-94f8-0df1aa567665 · inbound
How Far Are Video Models from True Multimodal Reasoning? Zero-shot Voice Conversion with Diffusion Transformers
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2ff6df34-4f54-4a2a-b6f1-6ff99c44e7b2 · inbound
RTCFake: Speech Deepfake Detection in Real-Time Communication Zero-shot Voice Conversion with Diffusion Transformers
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 854b3826-4e6a-4e10-8493-9af9f355a1ac · inbound
Poly-SVC: Polyphony-Aware Singing Voice Conversion with Harmonic Modeling Zero-shot Voice Conversion with Diffusion Transformers
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0ea1d867-762f-42d1-b092-b9539be5185b · inbound
SpeechEditBench: A Bilingual Multi-Attribute Benchmark for Instruction-Guided Speech Editing Zero-shot Voice Conversion with Diffusion Transformers
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation dfd36285-41fe-4150-8876-fa50210360b2 · inbound
From A to B to A: Palindromic Zero-Shot Voice Conversion with Non-Parallel Data Zero-shot Voice Conversion with Diffusion Transformers
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4c810e7a-48ce-47fc-9a6f-d83682abe6fe · inbound
From A to B to A: Palindromic Zero-Shot Voice Conversion with Non-Parallel Data Zero-shot Voice Conversion with Diffusion Transformers
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4ef1ae2b-b912-4f11-afe5-0354b55e4378 · inbound
MeanVC 2: Robust Low-Latency Streaming Zero-Shot Voice Conversion Zero-shot Voice Conversion with Diffusion Transformers
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c2d6f366-a499-4e1a-b1db-2a481ae472c7 · inbound
Vibrato Expression Control for Singing Voice Conversion with Improving Independent Control Zero-shot Voice Conversion with Diffusion Transformers
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e19300d7-d149-422d-b35a-0fc20c528157 · inbound
Zero-VC: Zero-Lookahead Streaming Voice Conversion via Speaker Anonymization Zero-shot Voice Conversion with Diffusion Transformers
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 99764790-3ccd-4382-9982-029011139817 · inbound
Speaker Identity in Non-Verbal Vocalizations: Conditional Distillation and Mixture of Experts Approach Zero-shot Voice Conversion with Diffusion Transformers
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 53d830a9-414e-4786-b458-eedce1f71606 · inbound
ProsoCodec: Prosody-Oriented Speech Codec for Voice Conversion Zero-shot Voice Conversion with Diffusion Transformers
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation bb720f14-37a2-42bd-b2fb-03139d65bf02 · inbound
AugCodec: A Low-Bitrate Disentangled Neural Speech Codec via Data Augmentation Zero-shot Voice Conversion with Diffusion Transformers
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5a422359-a41d-4708-96e6-b045c0e9070f · inbound
TRACE: Temporal Relationship-Aware Conversational Entrainment Detection in Dyadic Speech Zero-shot Voice Conversion with Diffusion Transformers
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bed75580-beb7-4272-b054-acaa075688fb · inbound
Enhancing Flow Matching with A Unified Guidance Framework for Efficient and Robust Speech Synthesis Zero-shot Voice Conversion with Diffusion Transformers
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 93f6fa2d-b120-4162-88d0-753d170d0aab · inbound
Beyond Words: Towards Effective Modeling of Non-Verbal Vocalizations in ASR Zero-shot Voice Conversion with Diffusion Transformers
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 298657dc-1185-40b7-9a76-4246648b1dcc · inbound
GRAFT: Grafted Reference Audio for Fine-grained Pronunciation in Zero-shot Text-to-Speech Zero-shot Voice Conversion with Diffusion Transformers
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f34df43-1869-47d7-b6a6-8d82c20fac19 · inbound
VoxENES 2026: Benchmarking Generalization of Speech Spoofing Detectors Against LLM-Era TTS and Voice Conversion Zero-shot Voice Conversion with Diffusion Transformers
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2287fb6-a75e-473e-9877-3627d630a0ef · inbound
A Geometry-Limited Identification Floor and Its Consequences for Voice-Clone Attribution in Professional Voice Actors Zero-shot Voice Conversion with Diffusion Transformers
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c19194d9-67a2-41f6-87b9-cd07aaaeab6b · inbound
A Geometry-Limited Identification Floor and Its Consequences for Voice-Clone Attribution in Professional Voice Actors Zero-shot Voice Conversion with Diffusion Transformers
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.