Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T14:21:56.940453Z
Paper Citation Record · LEDGER
As of 14 August 2026, this Paper Citation Record lists 100 of 150 outbound references and 4 inbound Pith citation observations for arXiv:2501.15442.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T14:21:56.940453Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-08T13:26:19.327889Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-01T11:45:47.119925Z
100 of 150 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 9fb8894c-94b5-4523-a1a4-14a7563ad28b · outbound
Overview of the Amphion Toolkit (v0.2) GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9233c72d-7318-445b-bbe7-98b22bc823a2 · outbound
Overview of the Amphion Toolkit (v0.2) Powerset Multi-class Cross Entropy Loss for Neural Speaker Diarization
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e954cec-5b6b-4a73-95fd-468bc96c2cb1 · outbound
Overview of the Amphion Toolkit (v0.2) Seed-TTS: A Family of High-Quality Versatile Speech Generation Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 032b69dd-de93-4a3c-b326-18a55af66f2f · outbound
Overview of the Amphion Toolkit (v0.2) Sd-eval: A benchmark dataset for spoken dialogue understand- ing beyond words
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f20de216-c0fa-4ae2-9415-b9ade29e7593 · outbound
Overview of the Amphion Toolkit (v0.2) Common Voice: A Massively-Multilingual Speech Corpus
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1136f15b-1eb4-43a5-9430-4b4e1ecbcb12 · outbound
Overview of the Amphion Toolkit (v0.2) Wav2vec 2.0: A framework for self-supervised learning of speech representations
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 688d8dcb-6d35-44b8-8bb4-1aedb81c7778 · outbound
Overview of the Amphion Toolkit (v0.2) METEOR: An automatic metric for MT evaluation with improved correlation with human judgments
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5661bae0-aca0-45ef-952a-39ebd08de0fa · outbound
Overview of the Amphion Toolkit (v0.2) SoundStorm: Efficient Parallel Audio Generation
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5e20aa3-9ce1-40da-9fdf-3916d0d7fb11 · outbound
Overview of the Amphion Toolkit (v0.2) AudioLM: A Language Modeling Approach to Audio Generation
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26be6c64-0534-4b4b-a87e-3cb69bf44c24 · outbound
Overview of the Amphion Toolkit (v0.2) Data augmentation and loss normalization for deep noise suppression
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce52798e-64eb-4523-8620-78f139ec9098 · outbound
Overview of the Amphion Toolkit (v0.2) pyannote.audio 2.1 speaker diarization pipeline: principle, benchmark, and recipe
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1be097d8-abe8-47d7-a462-1244fe0e9379 · outbound
Overview of the Amphion Toolkit (v0.2) Maskgit: Masked generative image transformer
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8abaa3f6-38f5-4bd5-bcd8-ea540c1fb82d · outbound
Overview of the Amphion Toolkit (v0.2) Gigaspeech: An evolving, multi-domain ASR corpus with 10, 000 hours of transcribed audio
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a63c61f-543a-49c8-a523-d3a49e22b287 · outbound
Overview of the Amphion Toolkit (v0.2) Wavlm: Large-scale self-supervised pre-training for full stack speech processing
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25d91c24-3010-4728-88c7-82ed1d181b0c · outbound
Overview of the Amphion Toolkit (v0.2) Qwen2-Audio Technical Report
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d4fddf2-f978-4b66-9acb-64778e5b95f1 · outbound
Overview of the Amphion Toolkit (v0.2) Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e3ebb6d-94a5-48f9-bbc1-35b26dbb6386 · outbound
Overview of the Amphion Toolkit (v0.2) Simple and controllable music generation
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4909080-fe5a-4c88-aa4c-a781a86325a4 · outbound
Overview of the Amphion Toolkit (v0.2) Librimix: An open-source dataset for generalizable speech separation, 2020
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18a403e8-054d-4017-9905-cca3b1f37f8c · outbound
Overview of the Amphion Toolkit (v0.2) Overview of the 2023 icassp sp clarity challenge: Speech enhancement for hearing aids
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe5c9afc-8d95-457f-b74f-a82f72f7f2cb · outbound
Overview of the Amphion Toolkit (v0.2) Moshi: a speech-text foundation model for real-time dialogue
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65ccbcd5-bc1d-4004-a135-99fa31108b3b · outbound
Overview of the Amphion Toolkit (v0.2) Real time speech enhancement in the waveform domain
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f3c0e4f-b8e6-4c5d-ada2-41b889e16743 · outbound
Overview of the Amphion Toolkit (v0.2) The ncte transcripts: A dataset of elementary math classroom transcripts
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7269c64f-6c19-4c72-9841-efe3467e1a05 · outbound
Overview of the Amphion Toolkit (v0.2) Pengi: An audio language model for audio tasks
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2b59e0b-f70f-4619-b884-a011376618fd · outbound
Overview of the Amphion Toolkit (v0.2) BERT: pre-training of deep bidirectional transformers for language understanding
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e92f1155-03a1-4184-a47e-7bf04eac8021 · outbound
Overview of the Amphion Toolkit (v0.2) Exploring speech enhancement with generative adversarial networks for robust speech recognition
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7cbc66a-d73e-4557-a2a3-08c99735048a · outbound
Overview of the Amphion Toolkit (v0.2) Cosyvoice: A scalable multilingual zero-shot text-to-speech synthesizer based on supervised semantic tokens, 2024
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a814e7ce-67c7-4743-bf9c-e206f38bb230 · outbound
Overview of the Amphion Toolkit (v0.2) High Fidelity Neural Audio Compression
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf10d203-f963-45d5-a399-8fdafed36048 · outbound
Overview of the Amphion Toolkit (v0.2) Introduction to audio data, 2025
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e87a630-419f-452a-9a12-fb6acc2db29f · outbound
Overview of the Amphion Toolkit (v0.2) Super-scotus: A multi- sourced dataset for the supreme court of the us
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b0625f8-1a5c-4afe-88c1-573aebba36b8 · outbound
Overview of the Amphion Toolkit (v0.2) LLaMA-Omni: Seamless Speech Interaction with Large Language Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7d37e2e-d5cb-4911-88fd-5825de27fe83 · outbound
Overview of the Amphion Toolkit (v0.2) Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61454f8a-6c84-4c56-9390-336464a314b1 · outbound
Overview of the Amphion Toolkit (v0.2) Liu, Leonid Karlinsky, and James R
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7f833b2-cd6c-4116-a82e-8d36dfa621e4 · outbound
Overview of the Amphion Toolkit (v0.2) Multi-scale sub-band constant-q transform discriminator for high-fidelity vocoder
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31794911-d5a6-49c4-9034-5a5b781d7844 · outbound
Overview of the Amphion Toolkit (v0.2) Conformer: Convolution- augmented transformer for speech recognition
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 405f8fce-1278-4214-b2b5-4009a4c5dfde · outbound
Overview of the Amphion Toolkit (v0.2) FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c48b55f6-9957-4b36-b49d-8c15668fad55 · outbound
Overview of the Amphion Toolkit (v0.2) Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3570b13a-c15e-46ec-af69-4eae64aaa75c · outbound
Overview of the Amphion Toolkit (v0.2) Noro: Noise-Robust One-shot Voice Conversion with Hidden Speaker Representation Learning
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd7ea5f8-283f-4c19-9d8b-b92432394cba · outbound
Overview of the Amphion Toolkit (v0.2) Gaussian Error Linear Units (GELUs)
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d08adf07-3e93-4824-802d-47f53c0c4102 · outbound
Overview of the Amphion Toolkit (v0.2) Denoising diffusion probabilistic models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97633959-b0ba-49be-a899-40fa4a94d72e · outbound
Overview of the Amphion Toolkit (v0.2) Hubert: Self-supervised speech representation learning by masked prediction of hidden units
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46b82778-5d37-4368-9455-c69442a9b454 · outbound
Overview of the Amphion Toolkit (v0.2) WavLLM: Towards Robust and Adaptive Speech Large Language Model
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5dc29f4e-4bf9-4c3a-9ca8-87c654f2be6b · outbound
Overview of the Amphion Toolkit (v0.2) The singing voice conversion challenge 2023
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7bd97b3b-1cbb-42f1-be7f-75c68711ab92 · outbound
Overview of the Amphion Toolkit (v0.2) Debatts: Zero-shot debating text-to-speech synthesis, 2024
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23025484-369e-46e7-adcf-e8480f20a8a6 · outbound
Overview of the Amphion Toolkit (v0.2) Repcodec: A speech representation codec for speech tokenization, 2024
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4517ac67-2e6b-4774-bab2-2db7892ac8a5 · outbound
Overview of the Amphion Toolkit (v0.2) The lj speech dataset
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7438ec5-f919-4fca-b9e2-c60e45301ba5 · outbound
Overview of the Amphion Toolkit (v0.2) Naturalspeech 3: Zero-shot speech synthesis with factorized codec and diffusion models
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 815bf137-36a6-411e-9f2d-bb98746a11de · outbound
Overview of the Amphion Toolkit (v0.2) AASIST: Audio Anti-Spoofing Using Integrated Spectro- Temporal Graph Attention Networks
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be519a90-6c3f-4f22-9f83-28eafb974ad8 · outbound
Overview of the Amphion Toolkit (v0.2) Libri-light: A benchmark for ASR with limited or no supervision
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d92074a6-22a3-4d9a-9706-27ef687f3323 · outbound
Overview of the Amphion Toolkit (v0.2) Librivox: Free public domain audiobooks
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 586b9c88-8913-45f8-9926-b97032ee208d · outbound
Overview of the Amphion Toolkit (v0.2) Fréchet audio distance: A reference-free metric for evaluating music enhancement algorithms
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5050ad88-c8e4-492b-a097-9bbdc02eaeff · outbound
Overview of the Amphion Toolkit (v0.2) Hifi-gan: generative adversarial networks for efficient and high fidelity speech synthesis
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c6468c7-728e-40b7-bb9a-3b787f908e8c · outbound
Overview of the Amphion Toolkit (v0.2) Audio Flamingo: A Novel Audio Language Model with Few-Shot Learning and Dialogue Abilities
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d2025e7-fcab-4402-a6a6-d24f779c06d3 · outbound
Overview of the Amphion Toolkit (v0.2) AudioGen: Textually guided audio generation
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 698abf5a-afbd-43c5-9f7b-940df9565bd3 · outbound
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae1fe7df-c22c-4a55-bb19-11d2ebc267b4 · outbound
Overview of the Amphion Toolkit (v0.2) V oiceBox: Text-guided Multilingual Universal Speech Generation at Scale.Advances in Neural Information Processing Systems, 36, 2024
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c80c56d8-2980-4eb1-9bb4-9f80f190ab9f · outbound
Overview of the Amphion Toolkit (v0.2) Textless speech-to-speech translation on real data
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 112cb62f-e905-4a89-9c1e-53f76967c6c1 · outbound
Overview of the Amphion Toolkit (v0.2) Bigvgan: A universal neural vocoder with large-scale training
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c73151f-6e9f-4e2d-b653-dddd8d5697c6 · outbound
Overview of the Amphion Toolkit (v0.2) HierSpeech++: Bridging the Gap between Semantic and Acoustic Representation of Speech by Hierarchical Variational Inference for Zero-shot Speech Synthesis
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd943c17-751e-4947-9d43-29e526936860 · outbound
Overview of the Amphion Toolkit (v0.2) Storm: A diffusion- based stochastic regeneration model for speech enhancement and dereverberation
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38a957ad-aa67-4582-aca1-9154e7447fb9 · outbound
Overview of the Amphion Toolkit (v0.2) Improved masked image generation with token-critic
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 553ae97f-bfde-4c59-b35b-235c2a236f57 · outbound
Overview of the Amphion Toolkit (v0.2) Espnet- se: End-to-end speech enhancement and separation toolkit designed for asr integration
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 106a6a93-dd3d-4b49-9ea0-2849d55d6a87 · outbound
Overview of the Amphion Toolkit (v0.2) Investigating neural audio codecs for speech language model-based speech generation
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 928c8276-4e12-4f2d-953c-4abfb09d1e62 · outbound
Overview of the Amphion Toolkit (v0.2) MaskSR: Masked Language Model for Full-band Speech Restoration
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 629dac48-25a1-4db7-895a-0165dad5fdb5 · outbound
Overview of the Amphion Toolkit (v0.2) ROUGE: A package for automatic evaluation of summaries
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d7fe1ae-b39e-425d-b731-a4159e6ff9dc · outbound
Overview of the Amphion Toolkit (v0.2) Unresolved cited work
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5614432e-54a3-47cb-b24b-806c370c10a0 · outbound
Overview of the Amphion Toolkit (v0.2) Audiosr: Versatile audio super-resolution at scale
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad2e3afc-6fa0-46f3-af86-0d2235bcf956 · outbound
Overview of the Amphion Toolkit (v0.2) Mandic, Wenwu Wang, and Mark D
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 751382a5-04aa-4965-81d4-8ccb9b9d786e · outbound
Overview of the Amphion Toolkit (v0.2) VoiceFixer: A Unified Framework for High-Fidelity Speech Restoration
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9daf869-0e4f-421d-84da-6f08dd825df4 · outbound
Overview of the Amphion Toolkit (v0.2) AudioLDM 2: Learning holistic audio generation with self-supervised pretraining
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0daa479a-36a8-4d05-9548-85956eaa945e · outbound
Overview of the Amphion Toolkit (v0.2) Spmis: An investigation of synthetic spoken misinformation detection
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a15d40d-4942-49c7-927e-2320b348bbf8 · outbound
Overview of the Amphion Toolkit (v0.2) Unresolved cited work
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c69d4e2-7363-4e51-9465-699839311db4 · outbound
Overview of the Amphion Toolkit (v0.2) SingVisio: Visual Analytics of Diffusion Model for Singing V oice Conversion
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62909fd7-b90a-46aa-a123-24b808a9daba · outbound
Overview of the Amphion Toolkit (v0.2) emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d771f84d-4c8b-48ae-87d7-8b2e412e3f38 · outbound
Overview of the Amphion Toolkit (v0.2) emotion2vec: Self-supervised pre-training for speech emotion representation
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e083cee-0276-452e-8084-e17077f69460 · outbound
Overview of the Amphion Toolkit (v0.2) WenetSpeech4TTS: A 12,800-hour Mandarin TTS Corpus for Large Speech Generation Model Benchmark
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f78c517f-beb7-4566-8cf3-daece1325725 · outbound
Overview of the Amphion Toolkit (v0.2) Good debt or bad debt: Detecting semantic orientations in economic texts
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a5faf88-a32e-412b-b364-f63a82f54698 · outbound
Overview of the Amphion Toolkit (v0.2) Metrics for polyphonic sound event detection
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 772a9847-6895-4a12-a802-c361164763e7 · outbound
Overview of the Amphion Toolkit (v0.2) NISQA: A Deep CNN-Self-Attention Model for Multidimensional Speech Quality Prediction with Crowdsourced Datasets
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be6239ff-6433-45ef-88b9-8b13d5589a94 · outbound
Overview of the Amphion Toolkit (v0.2) An overview of voice conversion systems
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1062cb0d-7448-46ff-98f5-c496848d655b · outbound
Overview of the Amphion Toolkit (v0.2) Hansard speeches 1979-2021: Version 3.1.0, May 2021
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3475df66-9183-4e4e-92c4-78fcb57b39e8 · outbound
Overview of the Amphion Toolkit (v0.2) Gpt-4o, 2024
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 364f2c8b-2a90-4991-aadc-2312150043dc · outbound
Overview of the Amphion Toolkit (v0.2) Bleu: a method for automatic evaluation of machine translation
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8df1c6a-9f23-425f-a70f-6d89282ab138 · outbound
Overview of the Amphion Toolkit (v0.2) Scalable diffusion models with transformers
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ee8aea3-97e5-414e-a0a0-ee80d83cfa38 · outbound
Overview of the Amphion Toolkit (v0.2) V oicecraft: Zero-shot speech editing and text-to-speech in the wild.ACL, 2024
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fab353a4-e01e-4ab0-be87-7f19a792c114 · outbound
Overview of the Amphion Toolkit (v0.2) Powerset multi-class cross entropy loss for neural speaker diarization
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e68ccfaa-670d-4791-95c7-cf8f8cafe421 · outbound
Overview of the Amphion Toolkit (v0.2) MLS: A large-scale multilingual dataset for speech research
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6fc6744-af09-4801-b219-6af0621faa9b · outbound
Overview of the Amphion Toolkit (v0.2) Autovc: Zero-shot voice style transfer with only autoencoder loss
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 658cd23e-3afd-4df8-aec7-15e1ab0dad33 · outbound
Overview of the Amphion Toolkit (v0.2) OpenVoice: Versatile Instant Voice Cloning
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8830185-5307-4992-b3b7-562759b70695 · outbound
Overview of the Amphion Toolkit (v0.2) Robust speech recognition via large-scale weak supervision
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 306cfa2c-6713-4be6-af9b-f5eada8dbcd2 · outbound
Overview of the Amphion Toolkit (v0.2) The interspeech 2020 deep noise suppression challenge: Datasets, subjective testing framework, and challenge results
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a81be4b7-0060-4613-9148-1aa974f936cf · outbound
Overview of the Amphion Toolkit (v0.2) DNSMOS P.835: A Non-intrusive Perceptual Objective Speech Quality Metric to Evaluate Noise Suppressors
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1e98c69-2bcf-45be-85a6-892097c982be · outbound
Overview of the Amphion Toolkit (v0.2) Speech enhancement and dereverberation with diffusion-based generative models
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 24bb73cf-5e5c-452b-8e8d-ea820233e06f · outbound
Overview of the Amphion Toolkit (v0.2) High-resolution image synthesis with latent diffusion models
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1cd1feb1-b882-4d8c-8e0e-605c54e75617 · outbound
Overview of the Amphion Toolkit (v0.2) SpeechBERTScore: Reference-Aware Automatic Evaluation of Speech Generation Leveraging NLP Evaluation Metrics
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f167429b-ce25-434e-951c-1977be006263 · outbound
Overview of the Amphion Toolkit (v0.2) Evaluating unsupervised text classification: Zero-shot and similarity-based approaches
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 6b004366-4aec-4027-a007-7715c7c60ad7 · outbound
Overview of the Amphion Toolkit (v0.2) NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e9e726e-7e2a-42f3-98dd-581de78ff745 · outbound
Overview of the Amphion Toolkit (v0.2) An overview of voice conversion and its challenges: From statistical modeling to deep learning
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 4ae5d636-4c8e-49bc-a738-07d449a545d6 · outbound
Overview of the Amphion Toolkit (v0.2) Unresolved cited work
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 6b9fd21b-7514-4948-8813-2015322df657 · outbound
Overview of the Amphion Toolkit (v0.2) Score-Based Generative Modeling through Stochastic Differential Equations
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44d61262-95fb-4c55-8672-d87041da05c4 · outbound
Overview of the Amphion Toolkit (v0.2) Roformer: Enhanced transformer with rotary position embedding
Reference 100
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4d9660f-dee0-4fd3-836c-94376ab66820 · inbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Overview of the Amphion Toolkit (v0.2)
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8598f5e2-524d-4196-afb4-8750e7a734b9 · inbound
Zero-Shot Text-to-Speech for Vietnamese Overview of the Amphion Toolkit (v0.2)
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57e53a4d-8694-414f-81ca-a15a61e36d7c · inbound
FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model Overview of the Amphion Toolkit (v0.2)
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 81f71f1e-67c0-4bba-a587-c6a478490971 · inbound
TRACE-EVC: Text-Guided Relative Affective Control for Zero-Shot Emotional Voice Conversion Overview of the Amphion Toolkit (v0.2)
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.