Pith. sign in

Paper Citation Record · LEDGER

Overview of the Amphion Toolkit (v0.2)

As of 14 August 2026, this Paper Citation Record lists 100 of 150 outbound references and 4 inbound Pith citation observations for arXiv:2501.15442.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.15442 v2

Coverage vector

measured 100 of 150 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T14:21:56.940453Z

measured 104 of 104 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T13:26:19.327889Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T11:45:47.119925Z

Reference resolution

100 of 150 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved97
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9fb8894c-94b5-4523-a1a4-14a7563ad28b · outbound

This paper cites GPT-4 Technical Report.

Overview of the Amphion Toolkit (v0.2) GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.664753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.664753Z digest=sha256:2ac1f38ce998f02e248fb23c2964b4870923dec2482cc0edface623fd364bdce

Observation 9233c72d-7318-445b-bbe7-98b22bc823a2 · outbound

This paper cites Powerset Multi-class Cross Entropy Loss for Neural Speaker Diarization.

Overview of the Amphion Toolkit (v0.2) Powerset Multi-class Cross Entropy Loss for Neural Speaker Diarization

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.669353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.669353Z digest=sha256:457cf09ca1e4b68323beeb917586a33ea656968049487be373b34f5ea0b000a3

Observation 6e954cec-5b6b-4a73-95fd-468bc96c2cb1 · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

Overview of the Amphion Toolkit (v0.2) Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.672064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.672064Z digest=sha256:1192bd122a9c57bf25cb8a49cbc25cad5c6bee14588e72952e1fe48a7ddb3749

Observation 032b69dd-de93-4a3c-b326-18a55af66f2f · outbound

This paper cites Sd-eval: A benchmark dataset for spoken dialogue understand- ing beyond words.

Overview of the Amphion Toolkit (v0.2) Sd-eval: A benchmark dataset for spoken dialogue understand- ing beyond words

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.676180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.676180Z digest=sha256:df602b7698eaa5d82ce048b1e91751f8980792167323d2c80264a3a1694ff93c

Observation f20de216-c0fa-4ae2-9415-b9ade29e7593 · outbound

This paper cites Common Voice: A Massively-Multilingual Speech Corpus.

Overview of the Amphion Toolkit (v0.2) Common Voice: A Massively-Multilingual Speech Corpus

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.679058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.679058Z digest=sha256:78be0877a744c784ba7d1116c94e2840832511342a1e1c9faae14eaa6202d0f9

Observation 1136f15b-1eb4-43a5-9430-4b4e1ecbcb12 · outbound

This paper cites Wav2vec 2.0: A framework for self-supervised learning of speech representations.

Overview of the Amphion Toolkit (v0.2) Wav2vec 2.0: A framework for self-supervised learning of speech representations

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.682705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.682705Z digest=sha256:d7ac5ddb03fc86c7e7edf5306683d564edb389e326a4e3b53442e2d4e0537ce3

Observation 688d8dcb-6d35-44b8-8bb4-1aedb81c7778 · outbound

This paper cites METEOR: An automatic metric for MT evaluation with improved correlation with human judgments.

Overview of the Amphion Toolkit (v0.2) METEOR: An automatic metric for MT evaluation with improved correlation with human judgments

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.686284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.686284Z digest=sha256:6ee9a94313b9c90c85b92989fd69d94c5f755a8a5e707e99861a3c4ee694c035

Observation 5661bae0-aca0-45ef-952a-39ebd08de0fa · outbound

This paper cites SoundStorm: Efficient Parallel Audio Generation.

Overview of the Amphion Toolkit (v0.2) SoundStorm: Efficient Parallel Audio Generation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.688812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.688812Z digest=sha256:9dac7fcf9456ce987eb74e3c918b3d42be8790a6109c60ff1a2dc6240ba1e8a2

Observation f5e20aa3-9ce1-40da-9fdf-3916d0d7fb11 · outbound

This paper cites AudioLM: A Language Modeling Approach to Audio Generation.

Overview of the Amphion Toolkit (v0.2) AudioLM: A Language Modeling Approach to Audio Generation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.691517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.691517Z digest=sha256:89e4e53f7bedfeff7e247141cef95d3885ec4baf8fba9e4385969febbe03b90d

Observation 26be6c64-0534-4b4b-a87e-3cb69bf44c24 · outbound

This paper cites Data augmentation and loss normalization for deep noise suppression.

Overview of the Amphion Toolkit (v0.2) Data augmentation and loss normalization for deep noise suppression

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.694127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.694127Z digest=sha256:faf56c0ae6a60e9ac661f3a56875bb2d61e4ca291417e462ef55936a36c23b3d

Observation ce52798e-64eb-4523-8620-78f139ec9098 · outbound

This paper cites pyannote.audio 2.1 speaker diarization pipeline: principle, benchmark, and recipe.

Overview of the Amphion Toolkit (v0.2) pyannote.audio 2.1 speaker diarization pipeline: principle, benchmark, and recipe

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.697092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.697092Z digest=sha256:0028657b61ed99671f058205d3dec68f3fc0b8066f4b69b84e94841503795378

Observation 1be097d8-abe8-47d7-a462-1244fe0e9379 · outbound

This paper cites Maskgit: Masked generative image transformer.

Overview of the Amphion Toolkit (v0.2) Maskgit: Masked generative image transformer

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.699996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.699996Z digest=sha256:3d8051c771ad3c5d0e47687ff55890b32e587c335a2ce64083ff3f9841538069

Observation 8abaa3f6-38f5-4bd5-bcd8-ea540c1fb82d · outbound

This paper cites Gigaspeech: An evolving, multi-domain ASR corpus with 10, 000 hours of transcribed audio.

Overview of the Amphion Toolkit (v0.2) Gigaspeech: An evolving, multi-domain ASR corpus with 10, 000 hours of transcribed audio

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.702701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.702701Z digest=sha256:4f3a49ca233e5912d38207a6dbdc1cd8716b3a757b491e53ed350698dfd649a4

Observation 4a63c61f-543a-49c8-a523-d3a49e22b287 · outbound

This paper cites Wavlm: Large-scale self-supervised pre-training for full stack speech processing.

Overview of the Amphion Toolkit (v0.2) Wavlm: Large-scale self-supervised pre-training for full stack speech processing

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.705348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.705348Z digest=sha256:cac0fc115b0df52b4f04f5d78733d4c709d4889ddc36ef5134cf71c517856883

Observation 25d91c24-3010-4728-88c7-82ed1d181b0c · outbound

This paper cites Qwen2-Audio Technical Report.

Overview of the Amphion Toolkit (v0.2) Qwen2-Audio Technical Report

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.708233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.708233Z digest=sha256:b8975c0b7fe268b4709e6a90201a2767a76b0faa3525a0d1c4f19f35449f0150

Observation 0d4fddf2-f978-4b66-9acb-64778e5b95f1 · outbound

This paper cites Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models.

Overview of the Amphion Toolkit (v0.2) Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.711506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.711506Z digest=sha256:20165129d6f2909e3f8a1b1167e9e5670a4e294e55aec63cfec284943b237f7b

Observation 0e3ebb6d-94a5-48f9-bbc1-35b26dbb6386 · outbound

This paper cites Simple and controllable music generation.

Overview of the Amphion Toolkit (v0.2) Simple and controllable music generation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.714682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.714682Z digest=sha256:848dc9af21bc711dbb0fa8d3c051ee5b3a5707f5f8bfdcdd4d700e5b6f5fc25f

Observation e4909080-fe5a-4c88-aa4c-a781a86325a4 · outbound

This paper cites Librimix: An open-source dataset for generalizable speech separation, 2020.

Overview of the Amphion Toolkit (v0.2) Librimix: An open-source dataset for generalizable speech separation, 2020

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.717556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.717556Z digest=sha256:0cd6bad7daf2dc498074c74ba87367f274dacf0bcdaaa86202b74ee8c631dc8a

Observation 18a403e8-054d-4017-9905-cca3b1f37f8c · outbound

This paper cites Overview of the 2023 icassp sp clarity challenge: Speech enhancement for hearing aids.

Overview of the Amphion Toolkit (v0.2) Overview of the 2023 icassp sp clarity challenge: Speech enhancement for hearing aids

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.720088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.720088Z digest=sha256:63f7e588534b5e38848dc4213a3d03b8d8a13f6fd40a15ae8f78758e866827b2

Observation fe5c9afc-8d95-457f-b74f-a82f72f7f2cb · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

Overview of the Amphion Toolkit (v0.2) Moshi: a speech-text foundation model for real-time dialogue

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.722997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.722997Z digest=sha256:287594f8fee10df0e9d6644708acf22d0f20868ecebafdb6de279006294d840b

Observation 65ccbcd5-bc1d-4004-a135-99fa31108b3b · outbound

This paper cites Real time speech enhancement in the waveform domain.

Overview of the Amphion Toolkit (v0.2) Real time speech enhancement in the waveform domain

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.726307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.726307Z digest=sha256:fc7f76dd07d3063f714e35c6e666434b7a289db85375bbfa0803df1941c8560a

Observation 0f3c0e4f-b8e6-4c5d-ada2-41b889e16743 · outbound

This paper cites The ncte transcripts: A dataset of elementary math classroom transcripts.

Overview of the Amphion Toolkit (v0.2) The ncte transcripts: A dataset of elementary math classroom transcripts

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.728929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.728929Z digest=sha256:9c4d09ed0f7cf3146a8deaf70d35372ce8ed0b713ff10951293e1a360a1f0a15

Observation 7269c64f-6c19-4c72-9841-efe3467e1a05 · outbound

This paper cites Pengi: An audio language model for audio tasks.

Overview of the Amphion Toolkit (v0.2) Pengi: An audio language model for audio tasks

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.731481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.731481Z digest=sha256:4e3d1f033f1b2899efd41eceed2b74fe791beab808e1652d60ad66c144ee9445

Observation f2b59e0b-f70f-4619-b884-a011376618fd · outbound

This paper cites BERT: pre-training of deep bidirectional transformers for language understanding.

Overview of the Amphion Toolkit (v0.2) BERT: pre-training of deep bidirectional transformers for language understanding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.734023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.734023Z digest=sha256:dc788d036d0a95f671e572f2d7337cb2bbc10d16e3b42fc73177002f0775a356

Observation e92f1155-03a1-4184-a47e-7bf04eac8021 · outbound

This paper cites Exploring speech enhancement with generative adversarial networks for robust speech recognition.

Overview of the Amphion Toolkit (v0.2) Exploring speech enhancement with generative adversarial networks for robust speech recognition

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.736555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.736555Z digest=sha256:d6f009572c077bfe516a533a4c252994eb92bd6ac0572bde194855c651b91e0f

Observation d7cbc66a-d73e-4557-a2a3-08c99735048a · outbound

This paper cites Cosyvoice: A scalable multilingual zero-shot text-to-speech synthesizer based on supervised semantic tokens, 2024.

Overview of the Amphion Toolkit (v0.2) Cosyvoice: A scalable multilingual zero-shot text-to-speech synthesizer based on supervised semantic tokens, 2024

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.739122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.739122Z digest=sha256:0c3a80777733266be79a5c759d010582a3f7bd65901bbc6dee5eeb61e39b83b3

Observation a814e7ce-67c7-4743-bf9c-e206f38bb230 · outbound

This paper cites High Fidelity Neural Audio Compression.

Overview of the Amphion Toolkit (v0.2) High Fidelity Neural Audio Compression

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.741411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.741411Z digest=sha256:70a6a132247d6ef5e5b9676189d4c9f3f3fa35ba76fa08d52c07be15bba9e9b8

Observation cf10d203-f963-45d5-a399-8fdafed36048 · outbound

This paper cites Introduction to audio data, 2025.

Overview of the Amphion Toolkit (v0.2) Introduction to audio data, 2025

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.743913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.743913Z digest=sha256:89ca8aaba0d104bd7f20389429d9fc86fe0dc513954dbc6b2ce5783d9d4a7f87

Observation 8e87a630-419f-452a-9a12-fb6acc2db29f · outbound

This paper cites Super-scotus: A multi- sourced dataset for the supreme court of the us.

Overview of the Amphion Toolkit (v0.2) Super-scotus: A multi- sourced dataset for the supreme court of the us

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.746599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.746599Z digest=sha256:540180aa93b4febddc9b215cead25c5a58c72bc66a346c08aac0efc699317d0d

Observation 2b0625f8-1a5c-4afe-88c1-573aebba36b8 · outbound

This paper cites LLaMA-Omni: Seamless Speech Interaction with Large Language Models.

Overview of the Amphion Toolkit (v0.2) LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.749133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.749133Z digest=sha256:b1221522cf4d37804ea09c112e594ebdc33297dfeed69e0c1f418d473035b918

Observation f7d37e2e-d5cb-4911-88fd-5825de27fe83 · outbound

This paper cites Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model.

Overview of the Amphion Toolkit (v0.2) Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.752505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.752505Z digest=sha256:93b766e27fa06306d51d86e02f0ae18cb7d8a54db651345d32bf32eeebd8e249

Observation 61454f8a-6c84-4c56-9390-336464a314b1 · outbound

This paper cites Liu, Leonid Karlinsky, and James R.

Overview of the Amphion Toolkit (v0.2) Liu, Leonid Karlinsky, and James R

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.755993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.755993Z digest=sha256:0d69b82dde7dd06d51a949e3462f1fe61ec3f5e0e622012ed18881abd26364b0

Observation f7f833b2-cd6c-4116-a82e-8d36dfa621e4 · outbound

This paper cites Multi-scale sub-band constant-q transform discriminator for high-fidelity vocoder.

Overview of the Amphion Toolkit (v0.2) Multi-scale sub-band constant-q transform discriminator for high-fidelity vocoder

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.758548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.758548Z digest=sha256:2e89bb2f90ffb69a503b52c61fc5d8aa15de09a7bb2d0147725e6ceab2ffecee

Observation 31794911-d5a6-49c4-9034-5a5b781d7844 · outbound

This paper cites Conformer: Convolution- augmented transformer for speech recognition.

Overview of the Amphion Toolkit (v0.2) Conformer: Convolution- augmented transformer for speech recognition

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.761260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.761260Z digest=sha256:66589b4453eebcecc057d357edc9dbe670445fec3868993b409498c6f974122c

Observation 405f8fce-1278-4214-b2b5-4009a4c5dfde · outbound

This paper cites FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications.

Overview of the Amphion Toolkit (v0.2) FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.764205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.764205Z digest=sha256:42638e9f34c89e5a2dcfb63d075d56f9d926f1cf7e2f2e64c285cf35cffc6afa

Observation c48b55f6-9957-4b36-b49d-8c15668fad55 · outbound

This paper cites Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation.

Overview of the Amphion Toolkit (v0.2) Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.767466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.767466Z digest=sha256:eb5542ba89b99d0cfe75396c1e7f3fbe01318370f8e36e52eb3b6d55fb57e9f5

Observation 3570b13a-c15e-46ec-af69-4eae64aaa75c · outbound

This paper cites Noro: Noise-Robust One-shot Voice Conversion with Hidden Speaker Representation Learning.

Overview of the Amphion Toolkit (v0.2) Noro: Noise-Robust One-shot Voice Conversion with Hidden Speaker Representation Learning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.770180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.770180Z digest=sha256:245ddc662bc4f53f614241a75f72ea4a075903f9dcd79785e4e31fe3a0db8fdd

Observation bd7ea5f8-283f-4c19-9d8b-b92432394cba · outbound

This paper cites Gaussian Error Linear Units (GELUs).

Overview of the Amphion Toolkit (v0.2) Gaussian Error Linear Units (GELUs)

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.773121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.773121Z digest=sha256:0d053b9a9db1dc8732476ffafc1a6129df0ab0b5028d5656c77a4d4e16c5b568

Observation d08adf07-3e93-4824-802d-47f53c0c4102 · outbound

This paper cites Denoising diffusion probabilistic models.

Overview of the Amphion Toolkit (v0.2) Denoising diffusion probabilistic models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.776184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.776184Z digest=sha256:9940214189ce6e22a1159e6579a85b579d5d42093c892521213b540a7dac2f0b

Observation 97633959-b0ba-49be-a899-40fa4a94d72e · outbound

This paper cites Hubert: Self-supervised speech representation learning by masked prediction of hidden units.

Overview of the Amphion Toolkit (v0.2) Hubert: Self-supervised speech representation learning by masked prediction of hidden units

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.778941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.778941Z digest=sha256:368f68d9b82c5b88500c6c46bf456bb448705b1657ef0c792be21c26a8e93ba9

Observation 46b82778-5d37-4368-9455-c69442a9b454 · outbound

This paper cites WavLLM: Towards Robust and Adaptive Speech Large Language Model.

Overview of the Amphion Toolkit (v0.2) WavLLM: Towards Robust and Adaptive Speech Large Language Model

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.781684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.781684Z digest=sha256:3caa8bb08a434b866c2232576dcd658ef7b90851f7822d655f2184d93559fa64

Observation 5dc29f4e-4bf9-4c3a-9ca8-87c654f2be6b · outbound

This paper cites The singing voice conversion challenge 2023.

Overview of the Amphion Toolkit (v0.2) The singing voice conversion challenge 2023

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.784989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.784989Z digest=sha256:b4ca0592e5012583aad4199c74073cec5d00f6c44266487caf6fe10dd3ef06dc

Observation 7bd97b3b-1cbb-42f1-be7f-75c68711ab92 · outbound

This paper cites Debatts: Zero-shot debating text-to-speech synthesis, 2024.

Overview of the Amphion Toolkit (v0.2) Debatts: Zero-shot debating text-to-speech synthesis, 2024

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.787173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.787173Z digest=sha256:4895b97759073408a25355dea7d8c64b5682f6c8b51d970a474dd646f1672059

Observation 23025484-369e-46e7-adcf-e8480f20a8a6 · outbound

This paper cites Repcodec: A speech representation codec for speech tokenization, 2024.

Overview of the Amphion Toolkit (v0.2) Repcodec: A speech representation codec for speech tokenization, 2024

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.789486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.789486Z digest=sha256:7c314c1a89efde06d14e2d4872d22b1eb3b37717f25fa4fdeef8a8fbec1ba261

Observation 4517ac67-2e6b-4774-bab2-2db7892ac8a5 · outbound

This paper cites The lj speech dataset.

Overview of the Amphion Toolkit (v0.2) The lj speech dataset

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.791767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.791767Z digest=sha256:d606734bb8ba929aad1cc0a6ab55e2347ecba2f4638d3e1d3e637f5354368bd7

Observation e7438ec5-f919-4fca-b9e2-c60e45301ba5 · outbound

This paper cites Naturalspeech 3: Zero-shot speech synthesis with factorized codec and diffusion models.

Overview of the Amphion Toolkit (v0.2) Naturalspeech 3: Zero-shot speech synthesis with factorized codec and diffusion models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.794848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.794848Z digest=sha256:a071681381773acfdd89bdd4d06abcd6ad8e564ef2f0ae9f7e53028a84280d0c

Observation 815bf137-36a6-411e-9f2d-bb98746a11de · outbound

This paper cites AASIST: Audio Anti-Spoofing Using Integrated Spectro- Temporal Graph Attention Networks.

Overview of the Amphion Toolkit (v0.2) AASIST: Audio Anti-Spoofing Using Integrated Spectro- Temporal Graph Attention Networks

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.798397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.798397Z digest=sha256:664ba257719aa5eac2ebdeed6e2b23ed7a84a58168e25c2170be2b449329f2df

Observation be519a90-6c3f-4f22-9f83-28eafb974ad8 · outbound

This paper cites Libri-light: A benchmark for ASR with limited or no supervision.

Overview of the Amphion Toolkit (v0.2) Libri-light: A benchmark for ASR with limited or no supervision

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.800834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.800834Z digest=sha256:0d0ce931766acd20e2721974abad042c2d1ba6c8e55757f5401ac7ded71eac4d

Observation d92074a6-22a3-4d9a-9706-27ef687f3323 · outbound

This paper cites Librivox: Free public domain audiobooks.

Overview of the Amphion Toolkit (v0.2) Librivox: Free public domain audiobooks

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.804065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.804065Z digest=sha256:d74288c70ecf9e05b56d75a3d6053f57b1e1c02e94abc82e419302670bcaad4b

Observation 586b9c88-8913-45f8-9926-b97032ee208d · outbound

This paper cites Fréchet audio distance: A reference-free metric for evaluating music enhancement algorithms.

Overview of the Amphion Toolkit (v0.2) Fréchet audio distance: A reference-free metric for evaluating music enhancement algorithms

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.807463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.807463Z digest=sha256:aefad0da52a4b289bee8b4cc64353043833b8b5fa83d400635d8586cd8016717

Observation 5050ad88-c8e4-492b-a097-9bbdc02eaeff · outbound

This paper cites Hifi-gan: generative adversarial networks for efficient and high fidelity speech synthesis.

Overview of the Amphion Toolkit (v0.2) Hifi-gan: generative adversarial networks for efficient and high fidelity speech synthesis

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.810249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.810249Z digest=sha256:f76ddcee5121ac19df6c860ea156ad62d31d671695437d3195a36135c3b253b1

Observation 9c6468c7-728e-40b7-bb9a-3b787f908e8c · outbound

This paper cites Audio Flamingo: A Novel Audio Language Model with Few-Shot Learning and Dialogue Abilities.

Overview of the Amphion Toolkit (v0.2) Audio Flamingo: A Novel Audio Language Model with Few-Shot Learning and Dialogue Abilities

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.812438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.812438Z digest=sha256:04036f9c323b55f3a0631b0002a369322aee83bc55a3ea24bceb7d6f36c01f69

Observation 2d2025e7-fcab-4402-a6a6-d24f779c06d3 · outbound

This paper cites AudioGen: Textually guided audio generation.

Overview of the Amphion Toolkit (v0.2) AudioGen: Textually guided audio generation

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.815009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.815009Z digest=sha256:3b0feacb15f928eb5edb2e7ec6fb58ab97a24a943fb041e03ef50bc7c1ff862a

Observation 698abf5a-afbd-43c5-9f7b-940df9565bd3 · outbound

This paper cites Kumar, P.

Overview of the Amphion Toolkit (v0.2) Kumar, P

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.817661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.817661Z digest=sha256:f3bde8998c76cf83f6b4886ec9a92d77de9647317502a4ec90f21b8b58691b31

Observation ae1fe7df-c22c-4a55-bb19-11d2ebc267b4 · outbound

This paper cites V oiceBox: Text-guided Multilingual Universal Speech Generation at Scale.Advances in Neural Information Processing Systems, 36, 2024.

Overview of the Amphion Toolkit (v0.2) V oiceBox: Text-guided Multilingual Universal Speech Generation at Scale.Advances in Neural Information Processing Systems, 36, 2024

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.819883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.819883Z digest=sha256:eaf05bc26eeaeb561d364a0870b89763b40e5dede1384ed11d4ada95b235b325

Observation c80c56d8-2980-4eb1-9bb4-9f80f190ab9f · outbound

This paper cites Textless speech-to-speech translation on real data.

Overview of the Amphion Toolkit (v0.2) Textless speech-to-speech translation on real data

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.822149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.822149Z digest=sha256:4d6559468caa823120d2f6bbef656e59daa8031beb065d714d7d1e58dadaef47

Observation 112cb62f-e905-4a89-9c1e-53f76967c6c1 · outbound

This paper cites Bigvgan: A universal neural vocoder with large-scale training.

Overview of the Amphion Toolkit (v0.2) Bigvgan: A universal neural vocoder with large-scale training

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.824471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.824471Z digest=sha256:30cbfaecdb3babd57c2a18b56ac063c5e60fd41853218620f4dd4237f5dc712e

Observation 4c73151f-6e9f-4e2d-b653-dddd8d5697c6 · outbound

This paper cites HierSpeech++: Bridging the Gap between Semantic and Acoustic Representation of Speech by Hierarchical Variational Inference for Zero-shot Speech Synthesis.

Overview of the Amphion Toolkit (v0.2) HierSpeech++: Bridging the Gap between Semantic and Acoustic Representation of Speech by Hierarchical Variational Inference for Zero-shot Speech Synthesis

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.827008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.827008Z digest=sha256:0a51a44f21ff48a280cf5584dec7aeda06d1a5d8fd9d7d4855aec44fb8e9074f

Observation dd943c17-751e-4947-9d43-29e526936860 · outbound

This paper cites Storm: A diffusion- based stochastic regeneration model for speech enhancement and dereverberation.

Overview of the Amphion Toolkit (v0.2) Storm: A diffusion- based stochastic regeneration model for speech enhancement and dereverberation

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.829455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.829455Z digest=sha256:b81917a7e1a8f856862f78ac13fdc1eb403c9dcfc997b48ee566d9838d303bef

Observation 38a957ad-aa67-4582-aca1-9154e7447fb9 · outbound

This paper cites Improved masked image generation with token-critic.

Overview of the Amphion Toolkit (v0.2) Improved masked image generation with token-critic

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.832581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.832581Z digest=sha256:afce1599d299a560d3ebfb310460ddc9c552cbf170cf4216d77dad0db43c368f

Observation 553ae97f-bfde-4c59-b35b-235c2a236f57 · outbound

This paper cites Espnet- se: End-to-end speech enhancement and separation toolkit designed for asr integration.

Overview of the Amphion Toolkit (v0.2) Espnet- se: End-to-end speech enhancement and separation toolkit designed for asr integration

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.835894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.835894Z digest=sha256:eb448d3e1d9e77140ccaa0401abf88b472d6256fdf36451a39e7347539d3f3e1

Observation 106a6a93-dd3d-4b49-9ea0-2849d55d6a87 · outbound

This paper cites Investigating neural audio codecs for speech language model-based speech generation.

Overview of the Amphion Toolkit (v0.2) Investigating neural audio codecs for speech language model-based speech generation

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.838266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.838266Z digest=sha256:ecde0e6ddaddbfb3b13bde6d545b2bf79fe0caecf140478886d68c1b46a6448f

Observation 928c8276-4e12-4f2d-953c-4abfb09d1e62 · outbound

This paper cites MaskSR: Masked Language Model for Full-band Speech Restoration.

Overview of the Amphion Toolkit (v0.2) MaskSR: Masked Language Model for Full-band Speech Restoration

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.840664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.840664Z digest=sha256:926dfdc9857760293750a1395e2d99a70ff9066fc2f3abfc95e5895d76471225

Observation 629dac48-25a1-4db7-895a-0165dad5fdb5 · outbound

This paper cites ROUGE: A package for automatic evaluation of summaries.

Overview of the Amphion Toolkit (v0.2) ROUGE: A package for automatic evaluation of summaries

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.843675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.843675Z digest=sha256:904c3deb5c58db52a2b65c641bc314138ea184731a0dd5ff19f0871f7e5b5f9b

Observation 8d7fe1ae-b39e-425d-b731-a4159e6ff9dc · outbound

This paper cites an unresolved cited work.

Overview of the Amphion Toolkit (v0.2) Unresolved cited work

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.846090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.846090Z digest=sha256:f8c8933ae5b9321d2cfd9c384b5e746bdb0126d13bce7e014a372b5af30000a1

Observation 5614432e-54a3-47cb-b24b-806c370c10a0 · outbound

This paper cites Audiosr: Versatile audio super-resolution at scale.

Overview of the Amphion Toolkit (v0.2) Audiosr: Versatile audio super-resolution at scale

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.848535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.848535Z digest=sha256:e8042f899ed0ac284473e0209b7ffc59baa972ab7beed7a89aea61c046f49529

Observation ad2e3afc-6fa0-46f3-af86-0d2235bcf956 · outbound

This paper cites Mandic, Wenwu Wang, and Mark D.

Overview of the Amphion Toolkit (v0.2) Mandic, Wenwu Wang, and Mark D

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.850991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.850991Z digest=sha256:49dce936f49858cc3e53aa3af84532cf1377f9861a9d62e3ac0018249c0e3fff

Observation 751382a5-04aa-4965-81d4-8ccb9b9d786e · outbound

This paper cites VoiceFixer: A Unified Framework for High-Fidelity Speech Restoration.

Overview of the Amphion Toolkit (v0.2) VoiceFixer: A Unified Framework for High-Fidelity Speech Restoration

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.853600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.853600Z digest=sha256:7e5045d216a8dbda3734b3abdaa0cd9a12d9039bf6a8e9b17a9e2dc4d3a64d5d

Observation e9daf869-0e4f-421d-84da-6f08dd825df4 · outbound

This paper cites AudioLDM 2: Learning holistic audio generation with self-supervised pretraining.

Overview of the Amphion Toolkit (v0.2) AudioLDM 2: Learning holistic audio generation with self-supervised pretraining

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.856216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.856216Z digest=sha256:1c629093686abe397c6bbc30c0be79b0aff032ff10b24c340af0258bf49ebcfa

Observation 0daa479a-36a8-4d05-9548-85956eaa945e · outbound

This paper cites Spmis: An investigation of synthetic spoken misinformation detection.

Overview of the Amphion Toolkit (v0.2) Spmis: An investigation of synthetic spoken misinformation detection

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.858980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.858980Z digest=sha256:f00ee3e3fc7ffa0994b5dba81ec867f283e30910030c0b190ff7407d5a94c510

Observation 7a15d40d-4942-49c7-927e-2320b348bbf8 · outbound

This paper cites an unresolved cited work.

Overview of the Amphion Toolkit (v0.2) Unresolved cited work

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.861389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.861389Z digest=sha256:659849e4bc7c8dec003b67df80005110fd9a2a847d8786a9aa5e8f0fae942307

Observation 8c69d4e2-7363-4e51-9465-699839311db4 · outbound

This paper cites SingVisio: Visual Analytics of Diffusion Model for Singing V oice Conversion.

Overview of the Amphion Toolkit (v0.2) SingVisio: Visual Analytics of Diffusion Model for Singing V oice Conversion

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.863539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.863539Z digest=sha256:26e6523f1c74db04352987f287184285516c2df0f653c70f0b0ecccafde2d379

Observation 62909fd7-b90a-46aa-a123-24b808a9daba · outbound

This paper cites emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation.

Overview of the Amphion Toolkit (v0.2) emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.866566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.866566Z digest=sha256:4225e062f57cbbd2d1ed82e2264bb9284255dec51875a0bab21d5b2300c195f4

Observation d771f84d-4c8b-48ae-87d7-8b2e412e3f38 · outbound

This paper cites emotion2vec: Self-supervised pre-training for speech emotion representation.

Overview of the Amphion Toolkit (v0.2) emotion2vec: Self-supervised pre-training for speech emotion representation

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.869648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.869648Z digest=sha256:89f946852eb2b502d3e5d550c194864f876cbb776c0931c0aa3c5d5566cdc697

Observation 4e083cee-0276-452e-8084-e17077f69460 · outbound

This paper cites WenetSpeech4TTS: A 12,800-hour Mandarin TTS Corpus for Large Speech Generation Model Benchmark.

Overview of the Amphion Toolkit (v0.2) WenetSpeech4TTS: A 12,800-hour Mandarin TTS Corpus for Large Speech Generation Model Benchmark

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.871990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.871990Z digest=sha256:ad9ece5306283799a73e605b66e1c905f1dc00cf65ffa6e17ef0df818e73dc07

Observation f78c517f-beb7-4566-8cf3-daece1325725 · outbound

This paper cites Good debt or bad debt: Detecting semantic orientations in economic texts.

Overview of the Amphion Toolkit (v0.2) Good debt or bad debt: Detecting semantic orientations in economic texts

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.874228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.874228Z digest=sha256:3a6ec75dba105c7a717704e6f942eefa0d4c072b80cd9f1bad7423233b033392

Observation 2a5faf88-a32e-412b-b364-f63a82f54698 · outbound

This paper cites Metrics for polyphonic sound event detection.

Overview of the Amphion Toolkit (v0.2) Metrics for polyphonic sound event detection

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.877007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.877007Z digest=sha256:8a1c07909f42fd7048ae47e3c9967e6294629abc1bade9ad95c8062da2a452cd

Observation 772a9847-6895-4a12-a802-c361164763e7 · outbound

This paper cites NISQA: A Deep CNN-Self-Attention Model for Multidimensional Speech Quality Prediction with Crowdsourced Datasets.

Overview of the Amphion Toolkit (v0.2) NISQA: A Deep CNN-Self-Attention Model for Multidimensional Speech Quality Prediction with Crowdsourced Datasets

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.879183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.879183Z digest=sha256:b78848673c353bbbc133819312e82c6c96e163f83319bcb4a1473a9f55ebdc82

Observation be6239ff-6433-45ef-88b9-8b13d5589a94 · outbound

This paper cites An overview of voice conversion systems.

Overview of the Amphion Toolkit (v0.2) An overview of voice conversion systems

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.882423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.882423Z digest=sha256:51f1e2c4afde43bf91cda06d4f224b83e8b957c6295e516af6741751b74a7ce3

Observation 1062cb0d-7448-46ff-98f5-c496848d655b · outbound

This paper cites Hansard speeches 1979-2021: Version 3.1.0, May 2021.

Overview of the Amphion Toolkit (v0.2) Hansard speeches 1979-2021: Version 3.1.0, May 2021

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.886037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.886037Z digest=sha256:0aa88dd18a59c191a2b777a5213c843b5c622e2ab69614714984e78836a6eacb

Observation 3475df66-9183-4e4e-92c4-78fcb57b39e8 · outbound

This paper cites Gpt-4o, 2024.

Overview of the Amphion Toolkit (v0.2) Gpt-4o, 2024

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.889158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.889158Z digest=sha256:854aa300526af151330d2f79b894943040a90d9188c77d85983038bab736c44e

Observation 364f2c8b-2a90-4991-aadc-2312150043dc · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation.

Overview of the Amphion Toolkit (v0.2) Bleu: a method for automatic evaluation of machine translation

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.891618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.891618Z digest=sha256:43b80d45c5db51e0d93135300764cdbff1cd4909984e5fa896579e3a470e4580

Observation a8df1c6a-9f23-425f-a70f-6d89282ab138 · outbound

This paper cites Scalable diffusion models with transformers.

Overview of the Amphion Toolkit (v0.2) Scalable diffusion models with transformers

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.893787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.893787Z digest=sha256:8b086e75a1dab9e87af216fa34357726e9f126a6c3589b19d8e2ec37986eb698

Observation 2ee8aea3-97e5-414e-a0a0-ee80d83cfa38 · outbound

This paper cites V oicecraft: Zero-shot speech editing and text-to-speech in the wild.ACL, 2024.

Overview of the Amphion Toolkit (v0.2) V oicecraft: Zero-shot speech editing and text-to-speech in the wild.ACL, 2024

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.896152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.896152Z digest=sha256:f3561383c87dbb2de55e951a81c71f5dbd3ff012a6f611f7b2f01ee47f264747

Observation fab353a4-e01e-4ab0-be87-7f19a792c114 · outbound

This paper cites Powerset multi-class cross entropy loss for neural speaker diarization.

Overview of the Amphion Toolkit (v0.2) Powerset multi-class cross entropy loss for neural speaker diarization

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.899225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.899225Z digest=sha256:c2b8e048b61f602f4bd494336e684ea70c6f813a19958fabd1783b5136e2dcf7

Observation e68ccfaa-670d-4791-95c7-cf8f8cafe421 · outbound

This paper cites MLS: A large-scale multilingual dataset for speech research.

Overview of the Amphion Toolkit (v0.2) MLS: A large-scale multilingual dataset for speech research

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.901830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.901830Z digest=sha256:46de33622b61ad1b90311f411dff83054f0b3d1106c0e43019d7b32b4698784c

Observation b6fc6744-af09-4801-b219-6af0621faa9b · outbound

This paper cites Autovc: Zero-shot voice style transfer with only autoencoder loss.

Overview of the Amphion Toolkit (v0.2) Autovc: Zero-shot voice style transfer with only autoencoder loss

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.905232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.905232Z digest=sha256:a030f484e68dd91aca9fb44873e851f71588592961effb2d7ce780a347727aa5

Observation 658cd23e-3afd-4df8-aec7-15e1ab0dad33 · outbound

This paper cites OpenVoice: Versatile Instant Voice Cloning.

Overview of the Amphion Toolkit (v0.2) OpenVoice: Versatile Instant Voice Cloning

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.907998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.907998Z digest=sha256:880853e8f706087d80b151161e2b26f2c72102ac272e35cc52f1d02263c4f3d4

Observation d8830185-5307-4992-b3b7-562759b70695 · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

Overview of the Amphion Toolkit (v0.2) Robust speech recognition via large-scale weak supervision

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.910543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.910543Z digest=sha256:8f9baa6a0e2a80df6c655346fe9778f61c1cfbf08a2f128ba4e60e29d2f1ec30

Observation 306cfa2c-6713-4be6-af9b-f5eada8dbcd2 · outbound

This paper cites The interspeech 2020 deep noise suppression challenge: Datasets, subjective testing framework, and challenge results.

Overview of the Amphion Toolkit (v0.2) The interspeech 2020 deep noise suppression challenge: Datasets, subjective testing framework, and challenge results

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.913433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.913433Z digest=sha256:ebdaaeef0ad9ac0ec1aa23350e1122a96ccb2cb64cf86e714ad369ceb96c0cc7

Observation a81be4b7-0060-4613-9148-1aa974f936cf · outbound

This paper cites DNSMOS P.835: A Non-intrusive Perceptual Objective Speech Quality Metric to Evaluate Noise Suppressors.

Overview of the Amphion Toolkit (v0.2) DNSMOS P.835: A Non-intrusive Perceptual Objective Speech Quality Metric to Evaluate Noise Suppressors

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.916029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.916029Z digest=sha256:7a42942059b47249ebedb41f5f590462e4c8c5c665065db2634e030c23af0790

Observation b1e98c69-2bcf-45be-85a6-892097c982be · outbound

This paper cites Speech enhancement and dereverberation with diffusion-based generative models.

Overview of the Amphion Toolkit (v0.2) Speech enhancement and dereverberation with diffusion-based generative models

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:21:57.729870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T14:21:56.919107Z digest=sha256:386de29fce6c5b6559e53231516604b8ad2c5f3e1ad87e584a4e62031d88b91b

Observation 24bb73cf-5e5c-452b-8e8d-ea820233e06f · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Overview of the Amphion Toolkit (v0.2) High-resolution image synthesis with latent diffusion models

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.921691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.921691Z digest=sha256:88fdef4c567830bf77251bd5b01cb446e92affb374078da5f7619766edf521df

Observation 1cd1feb1-b882-4d8c-8e0e-605c54e75617 · outbound

This paper cites SpeechBERTScore: Reference-Aware Automatic Evaluation of Speech Generation Leveraging NLP Evaluation Metrics.

Overview of the Amphion Toolkit (v0.2) SpeechBERTScore: Reference-Aware Automatic Evaluation of Speech Generation Leveraging NLP Evaluation Metrics

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.924127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.924127Z digest=sha256:937abcf4f6f5896759a74775756f32ea96821b2d23c512538c6897a271617c2b

Observation f167429b-ce25-434e-951c-1977be006263 · outbound

This paper cites Evaluating unsupervised text classification: Zero-shot and similarity-based approaches.

Overview of the Amphion Toolkit (v0.2) Evaluating unsupervised text classification: Zero-shot and similarity-based approaches

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:21:57.717454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T14:21:56.926732Z digest=sha256:1e64af953061a647ca57631b85a8a9441372f40314b59681e7bd55b4bcae1c02

Observation 6b004366-4aec-4027-a007-7715c7c60ad7 · outbound

This paper cites NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers.

Overview of the Amphion Toolkit (v0.2) NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.928999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.928999Z digest=sha256:5c1eb5c842ec626cd1242e8b35a5f8e47dc2e42f482c04fb4bd56009a89ddb66

Observation 3e9e726e-7e2a-42f3-98dd-581de78ff745 · outbound

This paper cites An overview of voice conversion and its challenges: From statistical modeling to deep learning.

Overview of the Amphion Toolkit (v0.2) An overview of voice conversion and its challenges: From statistical modeling to deep learning

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:21:57.709301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T14:21:56.932067Z digest=sha256:233c7a8741bc9ebb1b6a55dcc55f6314b74aabeb2b5208c08cd862939d0ae908

Observation 4ae5d636-4c8e-49bc-a738-07d449a545d6 · outbound

This paper cites an unresolved cited work.

Overview of the Amphion Toolkit (v0.2) Unresolved cited work

Reference 98

Resolution
unresolved
raw_fallback, observed 2026-08-10T14:21:57.701060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T14:21:56.935434Z digest=sha256:1830c77851a6789e779d945be5e41b97ab0146389701f9721c00a60b653a13a2

Observation 6b9fd21b-7514-4948-8813-2015322df657 · outbound

This paper cites Score-Based Generative Modeling through Stochastic Differential Equations.

Overview of the Amphion Toolkit (v0.2) Score-Based Generative Modeling through Stochastic Differential Equations

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.937957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.937957Z digest=sha256:70edeb2d67ad730e9b878ac34c015d5ce43a6713cdba47e3eb880d1cc46921bf

Observation 44d61262-95fb-4c55-8672-d87041da05c4 · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding.

Overview of the Amphion Toolkit (v0.2) Roformer: Enhanced transformer with rotary position embedding

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.940453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.940453Z digest=sha256:bd129c5e36a6af98aea4119c2c7aa27133d4b5951ae0ac2c0b2ab27b4c0d2486

Pith citing papers

Observation d4d9660f-dee0-4fd3-836c-94376ab66820 · inbound

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement cites this paper.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Overview of the Amphion Toolkit (v0.2)

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.327889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.327889Z digest=sha256:ccb2a83707cd08365a739f2bf9c70a6ba40f2c22e117263d0ef8a8c827e9e1dd

Observation 8598f5e2-524d-4196-afb4-8750e7a734b9 · inbound

Zero-Shot Text-to-Speech for Vietnamese cites this paper.

Zero-Shot Text-to-Speech for Vietnamese Overview of the Amphion Toolkit (v0.2)

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:50:04.356814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:50:04.356814Z digest=sha256:3d4a307ded49c88f43c1e41f45b58a36a1038dd343801c160a81b264d15ff5d0

Observation 57e53a4d-8694-414f-81ca-a15a61e36d7c · inbound

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model cites this paper.

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model Overview of the Amphion Toolkit (v0.2)

Reference 85

Resolution
verified exact
arxiv_id, observed 2026-07-01T11:45:47.122493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-07-01T03:50:26.873406Z digest=sha256:3ccc6b2808a43cee9615ff1ed3d107264c78243c12127b464baf2f45ad2ccf93

Observation 81f71f1e-67c0-4bba-a587-c6a478490971 · inbound

TRACE-EVC: Text-Guided Relative Affective Control for Zero-Shot Emotional Voice Conversion cites this paper.

TRACE-EVC: Text-Guided Relative Affective Control for Zero-Shot Emotional Voice Conversion Overview of the Amphion Toolkit (v0.2)

Reference 59

Resolution
unresolved
no resolver link, observed 2026-07-12T00:50:15.291565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:50:15.291565Z digest=sha256:32a06127ba5ebde77d7a07162f6fe56b330651cb953063611dc89c8454414fd5