Pith. sign in

Paper Citation Record · LEDGER

TokAN: Accent Normalization Using Self-Supervised Speech Tokens

As of 7 August 2026, this Paper Citation Record lists 78 of 78 outbound references and 0 inbound Pith citation observations for arXiv:2607.03928.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.03928 v1

Coverage vector

measured 78 of 78 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-11T22:58:27.449735Z

measured 78 of 78 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

78 of 78 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved77
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 074ad78a-7a56-4abe-91f1-dbf1495cf856 · outbound

This paper cites Foreign accent conversion by synthesizing speech from phonetic posteriorgrams.

TokAN: Accent Normalization Using Self-Supervised Speech Tokens Foreign accent conversion by synthesizing speech from phonetic posteriorgrams

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-11T22:58:27.449735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:58:27.449735Z digest=sha256:3e8a574bde50cf08c1dccf5693e9e20cdc64b1dc88ec5eadbdfbd93346c5cada

Observation 7919503b-bd8b-487f-953a-a67ba4037d07 · outbound

This paper cites Foreign accent conver- sion in computer assisted pronunciation training,.

TokAN: Accent Normalization Using Self-Supervised Speech Tokens Foreign accent conver- sion in computer assisted pronunciation training,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-11T22:58:27.449735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:58:27.449735Z digest=sha256:ebc68b04b4109e02d8d61f6b955f4392a8b702f0b37f8353e26d3d769dda8b26

Observation 87fa0dd0-52ea-4def-8aae-ec80ba227cbd · outbound

This paper cites Subband based voice conversion.

TokAN: Accent Normalization Using Self-Supervised Speech Tokens Subband based voice conversion

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-11T22:58:27.449735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:58:27.449735Z digest=sha256:d7ecc38767deca1b554bc2da45808266c72c43dee88121e908b2d7a924a52ff7

Observation 4a00cacc-46d3-4e67-9884-dfaab3152a17 · outbound

This paper cites Personalized, cross- lingual tts using phonetic posteriorgrams.

TokAN: Accent Normalization Using Self-Supervised Speech Tokens Personalized, cross- lingual tts using phonetic posteriorgrams

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-11T22:58:27.449735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:58:27.449735Z digest=sha256:b833b5c77ba42aea6cab3e3f6bf97d7ad051b0578c2c7d9d34c7bd2c73bc2606

Observation 6ff375ac-bd1e-44c5-a018-7fac4dea7d8c · outbound

This paper cites Accent conversion using phonetic posteriorgrams,.

TokAN: Accent Normalization Using Self-Supervised Speech Tokens Accent conversion using phonetic posteriorgrams,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-11T22:58:27.449735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:58:27.449735Z digest=sha256:e859df586858e0ae4777ef91c8b435ea9c52003881e7da51da34531089f28494

Observation 09ce1fa3-b9b5-4336-a7e8-785bd0bfd7ca · outbound

This paper cites Improving Accent Conversion with Reference Encoder and End-To-End Text-To-Speech.

TokAN: Accent Normalization Using Self-Supervised Speech Tokens Improving Accent Conversion with Reference Encoder and End-To-End Text-To-Speech

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-11T22:58:27.449735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:58:27.449735Z digest=sha256:d1c98c7f18c5c4609952974580cb326e6090a28b4a50a0efa243bd20494fef76

Observation 186ed197-6104-4806-888e-f583890f4234 · outbound

This paper cites Accentron: Foreign accent conversion to arbitrary non-native speakers using zero-shot learning,.

TokAN: Accent Normalization Using Self-Supervised Speech Tokens Accentron: Foreign accent conversion to arbitrary non-native speakers using zero-shot learning,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-11T22:58:27.449735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:58:27.449735Z digest=sha256:9e9625c02370ac5a33a42239a17dbe96f912a93c19707a50addfdcab6d157377

Observation 27be2d3d-f1ab-403c-ba20-41f524bb2ee7 · outbound

This paper cites Vevo: Control- lable zero-shot voice imitation with self-supervised disentanglement,.

TokAN: Accent Normalization Using Self-Supervised Speech Tokens Vevo: Control- lable zero-shot voice imitation with self-supervised disentanglement,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-11T22:58:27.449735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:58:27.449735Z digest=sha256:b30cc103e4c1267b723c9f47413cef070246a410a9757756b5adde89df3b0250

Observation 158b673c-3712-45c1-b99e-e78ab7813fe4 · outbound

This paper cites Converting foreign accent speech without a reference,.

TokAN: Accent Normalization Using Self-Supervised Speech Tokens Converting foreign accent speech without a reference,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-11T22:58:27.449735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:58:27.449735Z digest=sha256:ed68487a9c42011de3a7af5c056a5ff0c75fed95acd612fbbf1ef868172efa96

Observation 39ad60c4-67d2-41a1-8b68-cb44d6b6fab9 · outbound

This paper cites Accent conversion using pre-trained model and synthesized data from voice conversion.

TokAN: Accent Normalization Using Self-Supervised Speech Tokens Accent conversion using pre-trained model and synthesized data from voice conversion

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-11T22:58:27.449735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:58:27.449735Z digest=sha256:a2fa6fe4a6ad117ae04982ac7b406bb515e294cf3a0cca95ff896173ed7fe30c

Observation 2581cdb3-85e8-4934-907d-aea0a1cb3af3 · outbound

This paper cites Zero-shot foreign accent conversion without a native reference,.

TokAN: Accent Normalization Using Self-Supervised Speech Tokens Zero-shot foreign accent conversion without a native reference,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-11T22:58:27.449735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:58:27.449735Z digest=sha256:b8b1facc69d938a0cf63faa16641c84eb5f5c1a4b5e830f88a066f6e66ed4e5c

Observation c6c7199d-8c8e-4a52-8c5d-ba40a1706685 · outbound

This paper cites End-to-end accent conversion without using native utterances,.

TokAN: Accent Normalization Using Self-Supervised Speech Tokens End-to-end accent conversion without using native utterances,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-11T22:58:27.449735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:58:27.449735Z digest=sha256:25b0623a33034ca8d9023c3e7c4040bf7c7bbcbf088fb9279ada0b2fc9bbc809

Observation 933e74ec-4179-4a57-8892-74978113acdf · outbound

This paper cites V oice- preserving zero-shot multiple accent conversion,.

TokAN: Accent Normalization Using Self-Supervised Speech Tokens V oice- preserving zero-shot multiple accent conversion,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-11T22:58:27.449735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:58:27.449735Z digest=sha256:9026c924c8f49913441248edcb89473616be82cf5722288ba539320b351120e8

Observation c98be191-48ef-4fe7-8c4d-3dce81dcf0ec · outbound

This paper cites Tts-guided training for accent conversion without parallel data,.

TokAN: Accent Normalization Using Self-Supervised Speech Tokens Tts-guided training for accent conversion without parallel data,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-11T22:58:27.449735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:58:27.449735Z digest=sha256:305f819471a18be28d8daf6397374263f780a7aae1b8932701ddda1f6d0fe572

Observation 87e22484-a3a4-4004-b399-7756b1bb8cff · outbound

This paper cites Transfer the linguistic representations from tts to accent conversion with non-parallel data,.

TokAN: Accent Normalization Using Self-Supervised Speech Tokens Transfer the linguistic representations from tts to accent conversion with non-parallel data,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-11T22:58:27.449735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:58:27.449735Z digest=sha256:f66c4f0a86469797466939af21be3af421927b59fb12645ab0b43a71ea485e4c

Observation dfaa7b54-f1a8-4438-9685-f5155d3daf8a · outbound

This paper cites Diffusion-based method with tts guidance for foreign accent conver- sion,.

TokAN: Accent Normalization Using Self-Supervised Speech Tokens Diffusion-based method with tts guidance for foreign accent conver- sion,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-11T22:58:27.449735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:58:27.449735Z digest=sha256:304d2dd155aac7ae2246070a8ab117cb8bd977c7d1445177876671a159095057

Observation 28f74490-93aa-418c-9178-e4c65b0f85cd · outbound

This paper cites Improving pronunciation and accent conversion through knowledge distillation and synthetic ground-truth from native tts,.

TokAN: Accent Normalization Using Self-Supervised Speech Tokens Improving pronunciation and accent conversion through knowledge distillation and synthetic ground-truth from native tts,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-11T22:58:27.449735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:58:27.449735Z digest=sha256:44c20f513433a642723884beeea3440ddd1ce7576930d906fc8e0bd2ad31985f

Observation 19ceb256-865f-4c54-8c07-f1ceb2768d2b · outbound

This paper cites Convert and speak: Zero-shot accent conversion with minimum supervision,.

TokAN: Accent Normalization Using Self-Supervised Speech Tokens Convert and speak: Zero-shot accent conversion with minimum supervision,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-11T22:58:27.449735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:58:27.449735Z digest=sha256:61e03e77d8468f44e4def9c5dfa195594187adfb1c2be0539b728618a9af91fd

Observation 50b6b147-6632-455d-8a68-36e690440c64 · outbound

This paper cites Hubert: Self-supervised speech representation learning by masked prediction of hidden units,.

TokAN: Accent Normalization Using Self-Supervised Speech Tokens Hubert: Self-supervised speech representation learning by masked prediction of hidden units,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-11T22:58:27.449735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:58:27.449735Z digest=sha256:2a3c80692f94bbe1da18b6a89e5f994de42a0ea5969a2e4438235305476b05a3

Observation f464a9a0-f335-4bcc-a05d-83b4dcc10b39 · outbound

This paper cites Wavlm: Large-scale self-supervised pre- training for full stack speech processing,.

TokAN: Accent Normalization Using Self-Supervised Speech Tokens Wavlm: Large-scale self-supervised pre- training for full stack speech processing,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-11T22:58:27.449735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:58:27.449735Z digest=sha256:8e496978873e117da0de3eed3606eecde0b1d8391705577c511cf69eec368cc5

Observation 941b0aaf-c9aa-4cf2-9c48-2ce0aacbaa30 · outbound

This paper cites Self-supervised speech representations are more phonetic than semantic,.

TokAN: Accent Normalization Using Self-Supervised Speech Tokens Self-supervised speech representations are more phonetic than semantic,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-11T22:58:27.449735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:58:27.449735Z digest=sha256:40aeab84c757b47de826b2e4626508f93ef85fdb4ed8d9afc6b07019dad55b6d

Observation 2addf8d6-64c3-4f91-aced-2fa35659832a · outbound

This paper cites High fidelity neural audio compression,.

TokAN: Accent Normalization Using Self-Supervised Speech Tokens High fidelity neural audio compression,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-11T22:58:27.449735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:58:27.449735Z digest=sha256:f10b9ecf620cacf456fefb09ca4a5e2872eb72832bb4f634a5d9a5a45439a377

Observation ccd3e735-b6e6-4367-b0bf-3942c3c1e25a · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

TokAN: Accent Normalization Using Self-Supervised Speech Tokens CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-11T22:58:27.449735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:58:27.449735Z digest=sha256:f5350ac2a89815a5c3e6ab2088157b229f1f565b027fd42c8275092d74779a35

Observation b43622a8-11bb-4ef6-9c43-4c1ca758b30f · outbound

This paper cites Accent conversion using discrete units with parallel data synthesized from controllable accented tts,.

TokAN: Accent Normalization Using Self-Supervised Speech Tokens Accent conversion using discrete units with parallel data synthesized from controllable accented tts,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-11T22:58:27.449735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:58:27.449735Z digest=sha256:568f792c7957fdfc7c298effa88feed9742dac87b401bce69a906180c19bc26f

Observation ef111922-0620-4947-a95c-79cfc308c3ba · outbound

This paper cites Accent normalization using self-supervised discrete tokens with non-parallel data,.

TokAN: Accent Normalization Using Self-Supervised Speech Tokens Accent normalization using self-supervised discrete tokens with non-parallel data,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-11T22:58:27.449735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:58:27.449735Z digest=sha256:8e688b943c1e316a03ae97f11a8e3933aed6968539713cebedebf9b8cd3e622f

Observation e4f1d709-e018-4144-bcaa-63ef10663601 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

TokAN: Accent Normalization Using Self-Supervised Speech Tokens DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-11T22:58:27.449735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:58:27.449735Z digest=sha256:779fe6b1a8a1f6284e49b3841fa299fd4b2f3aa5f44759d53e33a5cfc662ea4b

Observation f2a57529-46bc-42dd-9d10-7f4bcc210f15 · outbound

This paper cites L2-ARCTIC: A Non-native English Speech Corpus,.

TokAN: Accent Normalization Using Self-Supervised Speech Tokens L2-ARCTIC: A Non-native English Speech Corpus,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-11T22:58:27.449735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:58:27.449735Z digest=sha256:32a784bb20f1baeb491c5d098ac1643a3d38d74f4be8d9fdd0d714480d8e07dd

Observation d306acef-da72-496e-b36d-bcf7321ed14b · outbound

This paper cites Cosyaccent: Duration-controllable accent normalization using source-synthesis train- ing data,.

TokAN: Accent Normalization Using Self-Supervised Speech Tokens Cosyaccent: Duration-controllable accent normalization using source-synthesis train- ing data,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-11T22:58:27.449735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:58:27.449735Z digest=sha256:09457af6994fec86fbe72d7287162fcb95ac1d4313f4fc4ba1efa49a26be239b

Observation c62e4b77-b73d-445f-9f5b-3261687116d0 · outbound

This paper cites Evaluating methods for ground-truth-free foreign accent conversion,.

TokAN: Accent Normalization Using Self-Supervised Speech Tokens Evaluating methods for ground-truth-free foreign accent conversion,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-11T22:58:27.449735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:58:27.449735Z digest=sha256:40a7dcb8747349ce1e37af1d575b4ab13f8903509338f2f8e60a9562890d1ca8

Observation 4939373a-50a8-4888-865b-d4ba112ebb01 · outbound

This paper cites Fac- facodec: Controllable zero-shot foreign accent conversion with factor- ized speech codec,.

TokAN: Accent Normalization Using Self-Supervised Speech Tokens Fac- facodec: Controllable zero-shot foreign accent conversion with factor- ized speech codec,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-11T22:58:27.449735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:58:27.449735Z digest=sha256:59834e24e4e35189eae29da9869ba475535ecf401da8468e03ac7bd3861b622f

Observation 17700096-1bce-4046-a9ec-70f1e49687b8 · outbound

This paper cites Any-to-one sequence-to- sequence voice conversion using self-supervised discrete speech rep- resentations,.

TokAN: Accent Normalization Using Self-Supervised Speech Tokens Any-to-one sequence-to- sequence voice conversion using self-supervised discrete speech rep- resentations,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-11T22:58:27.449735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:58:27.449735Z digest=sha256:43c15cd7ad34ea8389a50ae47d93ab0a7ecce67c5796a4c384422e44e4714540

Observation 116c96e2-d9ab-429c-80f6-0ed848db004f · outbound

This paper cites Speak, read and prompt: High-fidelity text-to-speech with minimal supervision,.

TokAN: Accent Normalization Using Self-Supervised Speech Tokens Speak, read and prompt: High-fidelity text-to-speech with minimal supervision,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-11T22:58:27.449735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:58:27.449735Z digest=sha256:e2fa61ce22f7f8e83e227102ce37069a74af493454e2745e16262ef2678acbd7

Observation ad085ab3-9f44-4f3a-b303-a0df36c0f75a · outbound

This paper cites On generative spoken language modeling from raw audio,.

TokAN: Accent Normalization Using Self-Supervised Speech Tokens On generative spoken language modeling from raw audio,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-11T22:58:27.449735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:58:27.449735Z digest=sha256:1e8c36ca1991877378a13f51977ea3de38b8cc95787e54a3dbd44b40f46ec68d

Observation 8c2d2ce4-5ad8-40d6-bfb2-03e0c6f4f867 · outbound

This paper cites Direct speech-to- speech translation with discrete units,.

TokAN: Accent Normalization Using Self-Supervised Speech Tokens Direct speech-to- speech translation with discrete units,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-11T22:58:27.449735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:58:27.449735Z digest=sha256:d181a8544a72d2422d0de2fcd404591713b33c02a6568ce161d0f3b37341061d

Observation a130e201-e309-4fdf-bb17-135aa45b6047 · outbound

This paper cites LLaMA-Omni: Seamless Speech Interaction with Large Language Models.

TokAN: Accent Normalization Using Self-Supervised Speech Tokens LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-11T22:58:27.449735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:58:27.449735Z digest=sha256:2b9e60f6f8de0a09daebd747360a5f3e30b11d3ec3a787d6fa3b1728b42557a8

Observation b157ec2f-ba50-442f-9077-63590cf4212a · outbound

This paper cites W2v-bert: Combining contrastive learning and masked language modeling for self-supervised speech pre-training,.

TokAN: Accent Normalization Using Self-Supervised Speech Tokens W2v-bert: Combining contrastive learning and masked language modeling for self-supervised speech pre-training,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-11T22:58:27.449735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:58:27.449735Z digest=sha256:99b1c5c907422e9fa7bda090fa543933afb1008637a56edb97dd6447a8b1c988

Observation c18cd026-6178-414d-9745-84a69103b0f4 · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

TokAN: Accent Normalization Using Self-Supervised Speech Tokens CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-11T22:58:27.449735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:58:27.449735Z digest=sha256:c9a9c5bc5ea34bbaf1f5f8fb124b5d65138e58cd7fc8c465d0bad8fbe2f0bd74

Observation 9b97e921-b519-4960-ba94-85616d24910a · outbound

This paper cites Flow matching for generative modeling,.

TokAN: Accent Normalization Using Self-Supervised Speech Tokens Flow matching for generative modeling,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-11T22:58:27.449735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:58:27.449735Z digest=sha256:b14d2ac4cb32fc96fd494b38514f6655e58d09c55c4788be24459386e991bb29

Observation 2dba284c-1f85-4d9a-8ec0-341ccadc4a07 · outbound

This paper cites Matcha-tts: A fast tts architecture with conditional flow matching,.

TokAN: Accent Normalization Using Self-Supervised Speech Tokens Matcha-tts: A fast tts architecture with conditional flow matching,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-11T22:58:27.449735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:58:27.449735Z digest=sha256:c81bb84a53b6d6a65ba7eb5775a3346c9c08402a1dd159d606e92676f8cd6fb4

Observation 74bd6d07-d215-49ee-9028-9d8e3e9dbbb4 · outbound

This paper cites V oicebox: Text- guided multilingual universal speech generation at scale,.

TokAN: Accent Normalization Using Self-Supervised Speech Tokens V oicebox: Text- guided multilingual universal speech generation at scale,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-11T22:58:27.449735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:58:27.449735Z digest=sha256:69d6a962b2c046ec0f34ae41fee76cedd81916f97a4a1472e6938d08111009dc

Observation 4f76e9c6-3349-4df9-a5cc-8cb5600df6de · outbound

This paper cites NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models.

TokAN: Accent Normalization Using Self-Supervised Speech Tokens NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-11T22:58:27.449735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:58:27.449735Z digest=sha256:9f0d204d63572cc4f919f11ef1d1e3672747a8494ec5ef6c1c37b81af8302a2b

Observation 583ca04e-88aa-4b6f-9be3-b3e2b5e19865 · outbound

This paper cites Total- duration-aware duration modeling for text-to-speech systems,.

TokAN: Accent Normalization Using Self-Supervised Speech Tokens Total- duration-aware duration modeling for text-to-speech systems,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-11T22:58:27.449735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:58:27.449735Z digest=sha256:338abf3a7abbf2a62430daf40d9fc1db0af335851d9a2ac3946f69ad58be487f

Observation 99ffa8e6-3877-481e-91b9-036198a4a82a · outbound

This paper cites Training language models to follow instructions with human feedback,.

TokAN: Accent Normalization Using Self-Supervised Speech Tokens Training language models to follow instructions with human feedback,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-11T22:58:27.449735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:58:27.449735Z digest=sha256:af2c37bb1896e24af9f6b0fc5f430063c43b8f2e6756a06e01382227ca644784

Observation d3e70b09-04ea-4c7b-9e7c-4cbdad581494 · outbound

This paper cites Group relative policy optimization for speech recognition,.

TokAN: Accent Normalization Using Self-Supervised Speech Tokens Group relative policy optimization for speech recognition,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-11T22:58:27.449735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:58:27.449735Z digest=sha256:51262d9ca971076d6d8d19c8b136c9e4c9fe97abe58ee63272b685f9db10be71

Observation 93f51340-2c8f-42bf-8208-adefc7efdcee · outbound

This paper cites Reinforcement Learning for Emotional Text-to-Speech Synthesis with Improved Emotion Discriminability,.

TokAN: Accent Normalization Using Self-Supervised Speech Tokens Reinforcement Learning for Emotional Text-to-Speech Synthesis with Improved Emotion Discriminability,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-07-11T22:58:27.449735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:58:27.449735Z digest=sha256:26db21b92d8243ee39b78174d2a309892381dfb65d9b5a0f285e7a64deb77513

Observation a2fc2671-6920-4029-b009-6d54cb33a8cc · outbound

This paper cites Dmospeech 2: Reinforcement learning for duration prediction in metric-optimized speech synthesis,.

TokAN: Accent Normalization Using Self-Supervised Speech Tokens Dmospeech 2: Reinforcement learning for duration prediction in metric-optimized speech synthesis,

Reference 46

Resolution
verified exact
doi, observed 2026-07-11T23:08:21.667101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-11T22:58:27.449735Z digest=sha256:68c90d6fff46c2a3b2a9bd5c08a4549ce136fb60cd9bb5d837237812d8819b2d

Observation 28fbc2ab-04d0-4da4-b1bc-1611e12afa4c · outbound

This paper cites Proximal Policy Optimization Algorithms.

TokAN: Accent Normalization Using Self-Supervised Speech Tokens Proximal Policy Optimization Algorithms

Reference 47

Resolution
unresolved
no resolver link, observed 2026-07-11T22:58:27.449735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:58:27.449735Z digest=sha256:af44a0de6e458897dc636b7d292b4cf36f703f0d7e0606e01a74a193f317c7aa

Observation b985e1da-a11c-4068-a91a-4b070c8403dd · outbound

This paper cites Attention is all you need,.

TokAN: Accent Normalization Using Self-Supervised Speech Tokens Attention is all you need,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-07-11T22:58:27.449735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:58:27.449735Z digest=sha256:8dd78a809e5b307606e666a29db64c2fa60d3268622df9a351d2de462d006fcb

Observation a5654a8b-a98e-4725-b2dd-098ac5cf8eb5 · outbound

This paper cites Roformer: En- hanced transformer with rotary position embedding,.

TokAN: Accent Normalization Using Self-Supervised Speech Tokens Roformer: En- hanced transformer with rotary position embedding,

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-11T22:58:27.449735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:58:27.449735Z digest=sha256:d1c021a4719a9e1a369d2a0bc0bc8bfe162fa58deadb9d751e94f07ef7a00d5e

Observation a6297c80-a820-4125-aa38-519ef75fe7d3 · outbound

This paper cites HiFTNet: A Fast High-Quality Neural Vocoder with Harmonic-plus-Noise Filter and Inverse Short Time Fourier Transform.

TokAN: Accent Normalization Using Self-Supervised Speech Tokens HiFTNet: A Fast High-Quality Neural Vocoder with Harmonic-plus-Noise Filter and Inverse Short Time Fourier Transform

Reference 50

Resolution
unresolved
no resolver link, observed 2026-07-11T22:58:27.449735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:58:27.449735Z digest=sha256:d4a52f838423637e5e2c2fa1dfa41d042de1a666e79b5a69684e5b08618b1368

Observation 16ac4dc7-c07d-4922-8a32-09f8236e7b3c · outbound

This paper cites Neural discrete representation learning,.

TokAN: Accent Normalization Using Self-Supervised Speech Tokens Neural discrete representation learning,

Reference 51

Resolution
unresolved
no resolver link, observed 2026-07-11T22:58:27.449735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:58:27.449735Z digest=sha256:b2b4039b52f65bede2bd19081d14e532db6344590640e0aa1daaf55ae93ca379

Observation e6998574-cf37-4df8-a500-b0ab11a7d03a · outbound

This paper cites Scalable diffusion models with transformers,.

TokAN: Accent Normalization Using Self-Supervised Speech Tokens Scalable diffusion models with transformers,

Reference 52

Resolution
unresolved
no resolver link, observed 2026-07-11T22:58:27.449735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:58:27.449735Z digest=sha256:91987566b9de91066ffa71f58e9c7c6cadd5524040673ec7945ec8c3f9244e7d

Observation f6ebb85a-a923-4026-a14b-2e5b1f76af65 · outbound

This paper cites Film: Visual reasoning with a general conditioning layer,.

TokAN: Accent Normalization Using Self-Supervised Speech Tokens Film: Visual reasoning with a general conditioning layer,

Reference 53

Resolution
unresolved
no resolver link, observed 2026-07-11T22:58:27.449735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:58:27.449735Z digest=sha256:76c23d07ed2eaf2ea2a8b86c2e0428cde6b4bc4fcb96093dbbd24476ee37c77e

Observation bab2d095-1b49-44ef-ac8a-a78fd6a8f5d7 · outbound

This paper cites Classifier-free diffusion guidance,.

TokAN: Accent Normalization Using Self-Supervised Speech Tokens Classifier-free diffusion guidance,

Reference 54

Resolution
unresolved
no resolver link, observed 2026-07-11T22:58:27.449735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:58:27.449735Z digest=sha256:da33c7dd9a9eee1ab1d0c0fcee8a2092f9b79510582676223876c1a355ecfa5d

Observation eb3be712-cb5b-4d6a-afc5-fddfc33fa0ed · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech representations,.

TokAN: Accent Normalization Using Self-Supervised Speech Tokens wav2vec 2.0: A framework for self-supervised learning of speech representations,

Reference 55

Resolution
unresolved
no resolver link, observed 2026-07-11T22:58:27.449735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:58:27.449735Z digest=sha256:18b9e4f6317943bd45eecdf01b5af4e82377024160b15920eea5d81de11de386

Observation 32774286-2cb2-44de-8eef-48d367269dc1 · outbound

This paper cites LibriTTS-R: A Restored Multi-Speaker Text-to-Speech Corpus,.

TokAN: Accent Normalization Using Self-Supervised Speech Tokens LibriTTS-R: A Restored Multi-Speaker Text-to-Speech Corpus,

Reference 56

Resolution
unresolved
no resolver link, observed 2026-07-11T22:58:27.449735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:58:27.449735Z digest=sha256:a383adc293ad2d85f827c58dbd50836aae9ff03e3c2ad2f56e451e5dcb8e6ea4

Observation 4bcf9eea-ddaf-4dfa-a760-9b904c458b7c · outbound

This paper cites Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation,.

TokAN: Accent Normalization Using Self-Supervised Speech Tokens Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation,

Reference 57

Resolution
unresolved
no resolver link, observed 2026-07-11T22:58:27.449735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:58:27.449735Z digest=sha256:3958bede9f4ba92d74a3fa6be79dd9b3104da74bdb9212e100e502865c63aeab

Observation e1c01422-40c3-4ab3-adf6-ba59ec250f5f · outbound

This paper cites Connection- ist temporal classification: labelling unsegmented sequence data with recurrent neural networks,.

TokAN: Accent Normalization Using Self-Supervised Speech Tokens Connection- ist temporal classification: labelling unsegmented sequence data with recurrent neural networks,

Reference 58

Resolution
unresolved
no resolver link, observed 2026-07-11T22:58:27.449735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:58:27.449735Z digest=sha256:8a52732c0da5071ca93011baaa60c3fb8fd01588c3be427417d4e618af54018f

Observation 2a67accf-03a9-4bbf-9365-e31f888595c5 · outbound

This paper cites BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension.

TokAN: Accent Normalization Using Self-Supervised Speech Tokens BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension

Reference 59

Resolution
unresolved
no resolver link, observed 2026-07-11T22:58:27.449735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:58:27.449735Z digest=sha256:363c24a04739e8c2699ef16edc0fb5b8b15e2c89da5b3a91288782e087b9de1c

Observation 512e4cfc-3c80-44a4-81b1-c2d973cca182 · outbound

This paper cites DAPO: An open-source LLM reinforcement learning system at scale,.

TokAN: Accent Normalization Using Self-Supervised Speech Tokens DAPO: An open-source LLM reinforcement learning system at scale,

Reference 60

Resolution
unresolved
no resolver link, observed 2026-07-11T22:58:27.449735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:58:27.449735Z digest=sha256:2e57c492442cd96d2911e4ee7be14c3ce454b79c7fbbd0ff1856f9b701c0b336

Observation 4605241b-b83f-443f-872f-53e3b60ead3a · outbound

This paper cites Robust speech recognition via large-scale weak super- vision,.

TokAN: Accent Normalization Using Self-Supervised Speech Tokens Robust speech recognition via large-scale weak super- vision,

Reference 61

Resolution
unresolved
no resolver link, observed 2026-07-11T22:58:27.449735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:58:27.449735Z digest=sha256:c2185eb2bf4e7b30ae15666f259240a40adf63412bc96a6c83372bb6fbb13a65

Observation d89b01b5-6631-4332-9f0f-f279de943cf2 · outbound

This paper cites Com- monAccent: Exploring Large Acoustic Pretrained Models for Accent Classification Based on Common V oice,.

TokAN: Accent Normalization Using Self-Supervised Speech Tokens Com- monAccent: Exploring Large Acoustic Pretrained Models for Accent Classification Based on Common V oice,

Reference 62

Resolution
unresolved
no resolver link, observed 2026-07-11T22:58:27.449735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:58:27.449735Z digest=sha256:623b42dee2ba8666001b749b7c61849fe855b77025abaadf954bc54d6cd5f9b3

Observation b351a1d1-eaa1-4634-b124-1d8acc53f432 · outbound

This paper cites Common voice: A massively-multilingual speech corpus,.

TokAN: Accent Normalization Using Self-Supervised Speech Tokens Common voice: A massively-multilingual speech corpus,

Reference 63

Resolution
unresolved
no resolver link, observed 2026-07-11T22:58:27.449735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:58:27.449735Z digest=sha256:65c711952ac9cb32f3a61ee8d3d392fd72c7a415a9377a1a6b550ba56bc93cb2

Observation 155cfd64-b754-45ce-9901-3333797b27d2 · outbound

This paper cites GLOBE: A High-quality English Corpus with Global Accents for Zero-shot Speaker Adaptive Text-to-Speech,.

TokAN: Accent Normalization Using Self-Supervised Speech Tokens GLOBE: A High-quality English Corpus with Global Accents for Zero-shot Speaker Adaptive Text-to-Speech,

Reference 64

Resolution
unresolved
no resolver link, observed 2026-07-11T22:58:27.449735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:58:27.449735Z digest=sha256:e057be7ac34fa88101502241947c489aa61e26985e5af98cd2d8cb243ea278d5

Observation 8becd556-ce31-46db-b81e-ffeb64d487d0 · outbound

This paper cites The cmu arctic speech databases,.

TokAN: Accent Normalization Using Self-Supervised Speech Tokens The cmu arctic speech databases,

Reference 65

Resolution
unresolved
no resolver link, observed 2026-07-11T22:58:27.449735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:58:27.449735Z digest=sha256:e9e9ff139c0aca2f89fcdba4a39c9daecd19fe6cb22d60d0f9305440b0b0b450

Observation 24d521dc-5fa8-4f3c-9c56-b08285c93d1c · outbound

This paper cites Fastspeech 2: Fast and high-quality end-to-end text to speech,.

TokAN: Accent Normalization Using Self-Supervised Speech Tokens Fastspeech 2: Fast and high-quality end-to-end text to speech,

Reference 66

Resolution
unresolved
no resolver link, observed 2026-07-11T22:58:27.449735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:58:27.449735Z digest=sha256:f588631299bd2dbccaa57203da4663606f2be76a7fff09ef02213a7dac4d4372

Observation 4938e5fb-8b08-4012-9cda-adc7dce563fa · outbound

This paper cites Bfa: Real- time multilingual text-to-speech forced alignment,.

TokAN: Accent Normalization Using Self-Supervised Speech Tokens Bfa: Real- time multilingual text-to-speech forced alignment,

Reference 67

Resolution
unresolved
no resolver link, observed 2026-07-11T22:58:27.449735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:58:27.449735Z digest=sha256:c4e85e8abcdfda15fec8d1d47f2b453a2eb89a1d6b7ea0eb87ebaf5b65f50ff6

Observation f1b14742-0654-4f34-a310-4df802d2160a · outbound

This paper cites an unresolved cited work.

TokAN: Accent Normalization Using Self-Supervised Speech Tokens Unresolved cited work

Reference 68

Resolution
unresolved
no resolver link, observed 2026-07-11T22:58:27.449735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:58:27.449735Z digest=sha256:6519ca7db0a656e484cb0a846aa264a4972616ed896f8e77419ddfaed82b279d

Observation 0930dcb2-32cd-434c-ae88-16be0304ab48 · outbound

This paper cites A comparison of best-worst scaling and rating scale for timbre characterisation,.

TokAN: Accent Normalization Using Self-Supervised Speech Tokens A comparison of best-worst scaling and rating scale for timbre characterisation,

Reference 69

Resolution
unresolved
no resolver link, observed 2026-07-11T22:58:27.449735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:58:27.449735Z digest=sha256:9337e10c978dc188267ae94e756187e810001d2aa5a7e30cf8dedaa17d62a289

Observation 640d0203-7ff2-4800-ae54-912bb46967e0 · outbound

This paper cites The t05 system for the VoiceMOS Challenge 2024: Transfer learning from deep image classifier to naturalness MOS prediction of high-quality synthetic speech,.

TokAN: Accent Normalization Using Self-Supervised Speech Tokens The t05 system for the VoiceMOS Challenge 2024: Transfer learning from deep image classifier to naturalness MOS prediction of high-quality synthetic speech,

Reference 70

Resolution
unresolved
no resolver link, observed 2026-07-11T22:58:27.449735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:58:27.449735Z digest=sha256:e98aae421a07cb3912dfafa8f0807c86353cd7c153aa44661b3ce08573624bff

Observation 99ba0f15-0e79-4f3f-931e-6fc892b363ad · outbound

This paper cites ECAPA-TDNN: Emphasized channel attention, propagation and aggregation in TDNN based speaker verification,.

TokAN: Accent Normalization Using Self-Supervised Speech Tokens ECAPA-TDNN: Emphasized channel attention, propagation and aggregation in TDNN based speaker verification,

Reference 71

Resolution
unresolved
no resolver link, observed 2026-07-11T22:58:27.449735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:58:27.449735Z digest=sha256:b12e0d7f4e7f3d38ea8c0f336f013e2ba05e823a799448923a1590029026a90c

Observation 583cbe99-ad68-4af9-9d60-9313843e8350 · outbound

This paper cites High-fidelity neural phonetic posteriorgrams,.

TokAN: Accent Normalization Using Self-Supervised Speech Tokens High-fidelity neural phonetic posteriorgrams,

Reference 72

Resolution
unresolved
no resolver link, observed 2026-07-11T22:58:27.449735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:58:27.449735Z digest=sha256:86eea3cede029c1ea626165cb1a4c5219a292107945505d177279f9a453a38f2

Observation e9d3a208-51a2-460c-8dca-ef50aed6d4b7 · outbound

This paper cites Exploring ssl discrete tokens for multilingual asr,.

TokAN: Accent Normalization Using Self-Supervised Speech Tokens Exploring ssl discrete tokens for multilingual asr,

Reference 73

Resolution
unresolved
no resolver link, observed 2026-07-11T22:58:27.449735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:58:27.449735Z digest=sha256:9844ddbec98dac7ecafd984a4e14a0fe33400d45193e1010d5f47d8f2a459624

Observation 356384d3-f1a0-4f4a-92bd-721a0862fef9 · outbound

This paper cites Towards universal speech discrete tokens: A case study for asr and tts,.

TokAN: Accent Normalization Using Self-Supervised Speech Tokens Towards universal speech discrete tokens: A case study for asr and tts,

Reference 74

Resolution
unresolved
no resolver link, observed 2026-07-11T22:58:27.449735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:58:27.449735Z digest=sha256:0fb30fd49c888674f2a54df57b448d36f875ecdefa964c37f703a4060363efcc

Observation fffe06c9-7956-48d9-8ad6-c2c6048d923d · outbound

This paper cites Exploring speech recognition, translation, and understanding with discrete speech units: A comparative study,.

TokAN: Accent Normalization Using Self-Supervised Speech Tokens Exploring speech recognition, translation, and understanding with discrete speech units: A comparative study,

Reference 75

Resolution
unresolved
no resolver link, observed 2026-07-11T22:58:27.449735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:58:27.449735Z digest=sha256:f2fbaf734ad026c1de94999dc1858b884911c62fe8ce95b746282479fb314f6f

Observation fecae4a7-6d80-4247-9c7d-9c020be60fa3 · outbound

This paper cites Codecmos-accent: A mos benchmark of resynthesized and tts speech from neural codecs across english accents,.

TokAN: Accent Normalization Using Self-Supervised Speech Tokens Codecmos-accent: A mos benchmark of resynthesized and tts speech from neural codecs across english accents,

Reference 76

Resolution
unresolved
no resolver link, observed 2026-07-11T22:58:27.449735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:58:27.449735Z digest=sha256:22884cdff2e7040d5d42670606a8ab4319d72b1c7d5d60c66c9d655bde78fad5

Observation 5d73e881-3b89-4247-b2a3-6582851bac86 · outbound

This paper cites Montreal forced aligner: Trainable text-speech alignment using kaldi.

TokAN: Accent Normalization Using Self-Supervised Speech Tokens Montreal forced aligner: Trainable text-speech alignment using kaldi

Reference 77

Resolution
unresolved
no resolver link, observed 2026-07-11T22:58:27.449735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:58:27.449735Z digest=sha256:ea2512aa9f9df99f78cfa7cc62aba986369a9af7d2d3683359e2a2fbb0f809e9

Observation 32566a22-3d21-4afe-9932-a6abfbbc8e1f · outbound

This paper cites Duanmu,The Phonology of Standard Chinese, 2nd ed.

TokAN: Accent Normalization Using Self-Supervised Speech Tokens Duanmu,The Phonology of Standard Chinese, 2nd ed

Reference 78

Resolution
unresolved
no resolver link, observed 2026-07-11T22:58:27.449735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:58:27.449735Z digest=sha256:d6e41db0740618cbab3a94131bd7a9b2e2c2a4ef5f7966d7130dbbf1478d3aa9

Pith citing papers

No inbound Pith citation observations are available.