Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-11T22:58:27.449735Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 78 of 78 outbound references and 0 inbound Pith citation observations for arXiv:2607.03928.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-11T22:58:27.449735Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
78 of 78 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 074ad78a-7a56-4abe-91f1-dbf1495cf856 · outbound
TokAN: Accent Normalization Using Self-Supervised Speech Tokens Foreign accent conversion by synthesizing speech from phonetic posteriorgrams
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7919503b-bd8b-487f-953a-a67ba4037d07 · outbound
TokAN: Accent Normalization Using Self-Supervised Speech Tokens Foreign accent conver- sion in computer assisted pronunciation training,
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87fa0dd0-52ea-4def-8aae-ec80ba227cbd · outbound
TokAN: Accent Normalization Using Self-Supervised Speech Tokens Subband based voice conversion
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a00cacc-46d3-4e67-9884-dfaab3152a17 · outbound
TokAN: Accent Normalization Using Self-Supervised Speech Tokens Personalized, cross- lingual tts using phonetic posteriorgrams
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ff375ac-bd1e-44c5-a018-7fac4dea7d8c · outbound
TokAN: Accent Normalization Using Self-Supervised Speech Tokens Accent conversion using phonetic posteriorgrams,
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09ce1fa3-b9b5-4336-a7e8-785bd0bfd7ca · outbound
TokAN: Accent Normalization Using Self-Supervised Speech Tokens Improving Accent Conversion with Reference Encoder and End-To-End Text-To-Speech
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 186ed197-6104-4806-888e-f583890f4234 · outbound
TokAN: Accent Normalization Using Self-Supervised Speech Tokens Accentron: Foreign accent conversion to arbitrary non-native speakers using zero-shot learning,
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27be2d3d-f1ab-403c-ba20-41f524bb2ee7 · outbound
TokAN: Accent Normalization Using Self-Supervised Speech Tokens Vevo: Control- lable zero-shot voice imitation with self-supervised disentanglement,
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 158b673c-3712-45c1-b99e-e78ab7813fe4 · outbound
TokAN: Accent Normalization Using Self-Supervised Speech Tokens Converting foreign accent speech without a reference,
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39ad60c4-67d2-41a1-8b68-cb44d6b6fab9 · outbound
TokAN: Accent Normalization Using Self-Supervised Speech Tokens Accent conversion using pre-trained model and synthesized data from voice conversion
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2581cdb3-85e8-4934-907d-aea0a1cb3af3 · outbound
TokAN: Accent Normalization Using Self-Supervised Speech Tokens Zero-shot foreign accent conversion without a native reference,
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6c7199d-8c8e-4a52-8c5d-ba40a1706685 · outbound
TokAN: Accent Normalization Using Self-Supervised Speech Tokens End-to-end accent conversion without using native utterances,
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 933e74ec-4179-4a57-8892-74978113acdf · outbound
TokAN: Accent Normalization Using Self-Supervised Speech Tokens V oice- preserving zero-shot multiple accent conversion,
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c98be191-48ef-4fe7-8c4d-3dce81dcf0ec · outbound
TokAN: Accent Normalization Using Self-Supervised Speech Tokens Tts-guided training for accent conversion without parallel data,
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87e22484-a3a4-4004-b399-7756b1bb8cff · outbound
TokAN: Accent Normalization Using Self-Supervised Speech Tokens Transfer the linguistic representations from tts to accent conversion with non-parallel data,
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dfaa7b54-f1a8-4438-9685-f5155d3daf8a · outbound
TokAN: Accent Normalization Using Self-Supervised Speech Tokens Diffusion-based method with tts guidance for foreign accent conver- sion,
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28f74490-93aa-418c-9178-e4c65b0f85cd · outbound
TokAN: Accent Normalization Using Self-Supervised Speech Tokens Improving pronunciation and accent conversion through knowledge distillation and synthetic ground-truth from native tts,
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19ceb256-865f-4c54-8c07-f1ceb2768d2b · outbound
TokAN: Accent Normalization Using Self-Supervised Speech Tokens Convert and speak: Zero-shot accent conversion with minimum supervision,
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50b6b147-6632-455d-8a68-36e690440c64 · outbound
TokAN: Accent Normalization Using Self-Supervised Speech Tokens Hubert: Self-supervised speech representation learning by masked prediction of hidden units,
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f464a9a0-f335-4bcc-a05d-83b4dcc10b39 · outbound
TokAN: Accent Normalization Using Self-Supervised Speech Tokens Wavlm: Large-scale self-supervised pre- training for full stack speech processing,
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 941b0aaf-c9aa-4cf2-9c48-2ce0aacbaa30 · outbound
TokAN: Accent Normalization Using Self-Supervised Speech Tokens Self-supervised speech representations are more phonetic than semantic,
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2addf8d6-64c3-4f91-aced-2fa35659832a · outbound
TokAN: Accent Normalization Using Self-Supervised Speech Tokens High fidelity neural audio compression,
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ccd3e735-b6e6-4367-b0bf-3942c3c1e25a · outbound
TokAN: Accent Normalization Using Self-Supervised Speech Tokens CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b43622a8-11bb-4ef6-9c43-4c1ca758b30f · outbound
TokAN: Accent Normalization Using Self-Supervised Speech Tokens Accent conversion using discrete units with parallel data synthesized from controllable accented tts,
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef111922-0620-4947-a95c-79cfc308c3ba · outbound
TokAN: Accent Normalization Using Self-Supervised Speech Tokens Accent normalization using self-supervised discrete tokens with non-parallel data,
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4f1d709-e018-4144-bcaa-63ef10663601 · outbound
TokAN: Accent Normalization Using Self-Supervised Speech Tokens DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2a57529-46bc-42dd-9d10-7f4bcc210f15 · outbound
TokAN: Accent Normalization Using Self-Supervised Speech Tokens L2-ARCTIC: A Non-native English Speech Corpus,
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d306acef-da72-496e-b36d-bcf7321ed14b · outbound
TokAN: Accent Normalization Using Self-Supervised Speech Tokens Cosyaccent: Duration-controllable accent normalization using source-synthesis train- ing data,
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c62e4b77-b73d-445f-9f5b-3261687116d0 · outbound
TokAN: Accent Normalization Using Self-Supervised Speech Tokens Evaluating methods for ground-truth-free foreign accent conversion,
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4939373a-50a8-4888-865b-d4ba112ebb01 · outbound
TokAN: Accent Normalization Using Self-Supervised Speech Tokens Fac- facodec: Controllable zero-shot foreign accent conversion with factor- ized speech codec,
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17700096-1bce-4046-a9ec-70f1e49687b8 · outbound
TokAN: Accent Normalization Using Self-Supervised Speech Tokens Any-to-one sequence-to- sequence voice conversion using self-supervised discrete speech rep- resentations,
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 116c96e2-d9ab-429c-80f6-0ed848db004f · outbound
TokAN: Accent Normalization Using Self-Supervised Speech Tokens Speak, read and prompt: High-fidelity text-to-speech with minimal supervision,
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad085ab3-9f44-4f3a-b303-a0df36c0f75a · outbound
TokAN: Accent Normalization Using Self-Supervised Speech Tokens On generative spoken language modeling from raw audio,
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c2d2ce4-5ad8-40d6-bfb2-03e0c6f4f867 · outbound
TokAN: Accent Normalization Using Self-Supervised Speech Tokens Direct speech-to- speech translation with discrete units,
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a130e201-e309-4fdf-bb17-135aa45b6047 · outbound
TokAN: Accent Normalization Using Self-Supervised Speech Tokens LLaMA-Omni: Seamless Speech Interaction with Large Language Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b157ec2f-ba50-442f-9077-63590cf4212a · outbound
TokAN: Accent Normalization Using Self-Supervised Speech Tokens W2v-bert: Combining contrastive learning and masked language modeling for self-supervised speech pre-training,
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c18cd026-6178-414d-9745-84a69103b0f4 · outbound
TokAN: Accent Normalization Using Self-Supervised Speech Tokens CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b97e921-b519-4960-ba94-85616d24910a · outbound
TokAN: Accent Normalization Using Self-Supervised Speech Tokens Flow matching for generative modeling,
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2dba284c-1f85-4d9a-8ec0-341ccadc4a07 · outbound
TokAN: Accent Normalization Using Self-Supervised Speech Tokens Matcha-tts: A fast tts architecture with conditional flow matching,
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74bd6d07-d215-49ee-9028-9d8e3e9dbbb4 · outbound
TokAN: Accent Normalization Using Self-Supervised Speech Tokens V oicebox: Text- guided multilingual universal speech generation at scale,
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f76e9c6-3349-4df9-a5cc-8cb5600df6de · outbound
TokAN: Accent Normalization Using Self-Supervised Speech Tokens NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 583ca04e-88aa-4b6f-9be3-b3e2b5e19865 · outbound
TokAN: Accent Normalization Using Self-Supervised Speech Tokens Total- duration-aware duration modeling for text-to-speech systems,
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99ffa8e6-3877-481e-91b9-036198a4a82a · outbound
TokAN: Accent Normalization Using Self-Supervised Speech Tokens Training language models to follow instructions with human feedback,
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3e70b09-04ea-4c7b-9e7c-4cbdad581494 · outbound
TokAN: Accent Normalization Using Self-Supervised Speech Tokens Group relative policy optimization for speech recognition,
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93f51340-2c8f-42bf-8208-adefc7efdcee · outbound
TokAN: Accent Normalization Using Self-Supervised Speech Tokens Reinforcement Learning for Emotional Text-to-Speech Synthesis with Improved Emotion Discriminability,
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2fc2671-6920-4029-b009-6d54cb33a8cc · outbound
TokAN: Accent Normalization Using Self-Supervised Speech Tokens Dmospeech 2: Reinforcement learning for duration prediction in metric-optimized speech synthesis,
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 28fbc2ab-04d0-4da4-b1bc-1611e12afa4c · outbound
TokAN: Accent Normalization Using Self-Supervised Speech Tokens Proximal Policy Optimization Algorithms
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b985e1da-a11c-4068-a91a-4b070c8403dd · outbound
TokAN: Accent Normalization Using Self-Supervised Speech Tokens Attention is all you need,
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5654a8b-a98e-4725-b2dd-098ac5cf8eb5 · outbound
TokAN: Accent Normalization Using Self-Supervised Speech Tokens Roformer: En- hanced transformer with rotary position embedding,
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6297c80-a820-4125-aa38-519ef75fe7d3 · outbound
TokAN: Accent Normalization Using Self-Supervised Speech Tokens HiFTNet: A Fast High-Quality Neural Vocoder with Harmonic-plus-Noise Filter and Inverse Short Time Fourier Transform
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16ac4dc7-c07d-4922-8a32-09f8236e7b3c · outbound
TokAN: Accent Normalization Using Self-Supervised Speech Tokens Neural discrete representation learning,
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6998574-cf37-4df8-a500-b0ab11a7d03a · outbound
TokAN: Accent Normalization Using Self-Supervised Speech Tokens Scalable diffusion models with transformers,
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6ebb85a-a923-4026-a14b-2e5b1f76af65 · outbound
TokAN: Accent Normalization Using Self-Supervised Speech Tokens Film: Visual reasoning with a general conditioning layer,
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bab2d095-1b49-44ef-ac8a-a78fd6a8f5d7 · outbound
TokAN: Accent Normalization Using Self-Supervised Speech Tokens Classifier-free diffusion guidance,
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb3be712-cb5b-4d6a-afc5-fddfc33fa0ed · outbound
TokAN: Accent Normalization Using Self-Supervised Speech Tokens wav2vec 2.0: A framework for self-supervised learning of speech representations,
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32774286-2cb2-44de-8eef-48d367269dc1 · outbound
TokAN: Accent Normalization Using Self-Supervised Speech Tokens LibriTTS-R: A Restored Multi-Speaker Text-to-Speech Corpus,
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4bcf9eea-ddaf-4dfa-a760-9b904c458b7c · outbound
TokAN: Accent Normalization Using Self-Supervised Speech Tokens Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation,
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1c01422-40c3-4ab3-adf6-ba59ec250f5f · outbound
TokAN: Accent Normalization Using Self-Supervised Speech Tokens Connection- ist temporal classification: labelling unsegmented sequence data with recurrent neural networks,
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a67accf-03a9-4bbf-9365-e31f888595c5 · outbound
TokAN: Accent Normalization Using Self-Supervised Speech Tokens BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 512e4cfc-3c80-44a4-81b1-c2d973cca182 · outbound
TokAN: Accent Normalization Using Self-Supervised Speech Tokens DAPO: An open-source LLM reinforcement learning system at scale,
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4605241b-b83f-443f-872f-53e3b60ead3a · outbound
TokAN: Accent Normalization Using Self-Supervised Speech Tokens Robust speech recognition via large-scale weak super- vision,
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d89b01b5-6631-4332-9f0f-f279de943cf2 · outbound
TokAN: Accent Normalization Using Self-Supervised Speech Tokens Com- monAccent: Exploring Large Acoustic Pretrained Models for Accent Classification Based on Common V oice,
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b351a1d1-eaa1-4634-b124-1d8acc53f432 · outbound
TokAN: Accent Normalization Using Self-Supervised Speech Tokens Common voice: A massively-multilingual speech corpus,
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 155cfd64-b754-45ce-9901-3333797b27d2 · outbound
TokAN: Accent Normalization Using Self-Supervised Speech Tokens GLOBE: A High-quality English Corpus with Global Accents for Zero-shot Speaker Adaptive Text-to-Speech,
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8becd556-ce31-46db-b81e-ffeb64d487d0 · outbound
TokAN: Accent Normalization Using Self-Supervised Speech Tokens The cmu arctic speech databases,
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24d521dc-5fa8-4f3c-9c56-b08285c93d1c · outbound
TokAN: Accent Normalization Using Self-Supervised Speech Tokens Fastspeech 2: Fast and high-quality end-to-end text to speech,
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4938e5fb-8b08-4012-9cda-adc7dce563fa · outbound
TokAN: Accent Normalization Using Self-Supervised Speech Tokens Bfa: Real- time multilingual text-to-speech forced alignment,
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1b14742-0654-4f34-a310-4df802d2160a · outbound
TokAN: Accent Normalization Using Self-Supervised Speech Tokens Unresolved cited work
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0930dcb2-32cd-434c-ae88-16be0304ab48 · outbound
TokAN: Accent Normalization Using Self-Supervised Speech Tokens A comparison of best-worst scaling and rating scale for timbre characterisation,
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 640d0203-7ff2-4800-ae54-912bb46967e0 · outbound
TokAN: Accent Normalization Using Self-Supervised Speech Tokens The t05 system for the VoiceMOS Challenge 2024: Transfer learning from deep image classifier to naturalness MOS prediction of high-quality synthetic speech,
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99ba0f15-0e79-4f3f-931e-6fc892b363ad · outbound
TokAN: Accent Normalization Using Self-Supervised Speech Tokens ECAPA-TDNN: Emphasized channel attention, propagation and aggregation in TDNN based speaker verification,
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 583cbe99-ad68-4af9-9d60-9313843e8350 · outbound
TokAN: Accent Normalization Using Self-Supervised Speech Tokens High-fidelity neural phonetic posteriorgrams,
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9d3a208-51a2-460c-8dca-ef50aed6d4b7 · outbound
TokAN: Accent Normalization Using Self-Supervised Speech Tokens Exploring ssl discrete tokens for multilingual asr,
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 356384d3-f1a0-4f4a-92bd-721a0862fef9 · outbound
TokAN: Accent Normalization Using Self-Supervised Speech Tokens Towards universal speech discrete tokens: A case study for asr and tts,
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fffe06c9-7956-48d9-8ad6-c2c6048d923d · outbound
TokAN: Accent Normalization Using Self-Supervised Speech Tokens Exploring speech recognition, translation, and understanding with discrete speech units: A comparative study,
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fecae4a7-6d80-4247-9c7d-9c020be60fa3 · outbound
TokAN: Accent Normalization Using Self-Supervised Speech Tokens Codecmos-accent: A mos benchmark of resynthesized and tts speech from neural codecs across english accents,
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d73e881-3b89-4247-b2a3-6582851bac86 · outbound
TokAN: Accent Normalization Using Self-Supervised Speech Tokens Montreal forced aligner: Trainable text-speech alignment using kaldi
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32566a22-3d21-4afe-9932-a6abfbbc8e1f · outbound
TokAN: Accent Normalization Using Self-Supervised Speech Tokens Duanmu,The Phonology of Standard Chinese, 2nd ed
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.