Pith. sign in

Paper Citation Record · LEDGER

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning

As of 22 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 1 inbound Pith citation observation for arXiv:2506.04527.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.04527 v1

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:45:15.786786Z

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:45:15.390168Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T10:45:15.943522Z

Reference resolution

49 of 49 outbound references displayed

  • verified exact0
  • verified fuzzy45
  • unresolved3
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8d183f01-308a-4a4d-b3a1-f50a73b0244b · outbound

This paper cites For training high-quality and diverse-styled TTS models, a large amount of text-speech paired data is re- quired [4, 5].

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning For training high-quality and diverse-styled TTS models, a large amount of text-speech paired data is re- quired [4, 5]

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.968584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:45:15.376258Z digest=sha256:13f0e8cb5781090ee6aad43f131a8e1c52478cecf7a928f9679711b8e5c4d09e

Observation 3828ef2d-8103-4827-9723-0e4ff294770b · outbound

This paper cites Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T10:45:15.953103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:45:15.390168Z digest=sha256:77c4d26f908e7cf5ba2c99abdbedc329f809a9d7a2724a825e7495a1f1db32a5

Observation d83c57ff-e9c2-4051-a2b8-5ffde5a2fcff · outbound

This paper cites <blank>.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning <blank>

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.947649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:45:15.397152Z digest=sha256:6b3126e179ac300828a9e473638385299222c34f8bfc313d8c8a91b5ace6d434

Observation 283541ee-6f26-4a93-adb5-f166d39e8b93 · outbound

This paper cites Evaluation of proposed annotation model 4.1.1.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Evaluation of proposed annotation model 4.1.1

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.928154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:45:15.408097Z digest=sha256:b21ab6ab0967ee476e83401b6b6448206adc25fd378f35a462d36d2990269994

Observation dd5fb3c8-6404-4130-b3ef-a50ee30c9db0 · outbound

This paper cites Specifically, we utilized the basic5000 subset along with its manually annotated TTS labels1.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Specifically, we utilized the basic5000 subset along with its manually annotated TTS labels1

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.903592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:45:15.416347Z digest=sha256:6d610dd011af35dccc2157612a159a28563c40c2da1b8e4de0a49d6a2ed7cdf9

Observation 88bb5455-04e4-464a-b6ce-7b7f26498546 · outbound

This paper cites ”, (2) Accent change from low to high “[.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning ”, (2) Accent change from low to high “[

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.884422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:45:15.422376Z digest=sha256:5c2cb9250c38028e312fc08fcde0c11f128637ade22ec80e832da83719efdfcb

Observation 8a7ac17f-3c7a-49fd-b1c8-b4a1f7ba2937 · outbound

This paper cites Applying this method to downstream tasks be- yond textual accent estimation is a challenge for future work.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Applying this method to downstream tasks be- yond textual accent estimation is a challenge for future work

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.865438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:45:15.428341Z digest=sha256:c68f8d9781498754e68971d94625047d0ba9b8fecacd70dff609adbb4608b4b0

Observation 471f62fb-b9d2-4ecb-8445-85cff0353220 · outbound

This paper cites Tacotron: Towards end-to-end speech synthesis,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Tacotron: Towards end-to-end speech synthesis,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.843284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:45:15.436595Z digest=sha256:1d5c41e979e617bbda618296ab8239b976425716a6b8298e22b5665f79b698de

Observation 0d18b20c-8ab5-4230-87dd-921714906ef3 · outbound

This paper cites Fastspeech 2: Fast and high-quality end-to-end text to speech,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Fastspeech 2: Fast and high-quality end-to-end text to speech,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.820773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:45:15.444523Z digest=sha256:0911f615bfe3d2dad425815c34b9daa721027dab41fde1851e576915d3b25bbd

Observation 76c0dd0c-e7cd-4f32-8924-f033a51a921c · outbound

This paper cites Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.800603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:45:15.455133Z digest=sha256:9540c55163986a056b497416777ce9c6c2690261a89444641994fa5edea4ba5b

Observation 30c4032d-9850-4d5d-b62e-ae3728234cc1 · outbound

This paper cites Naturalspeech 2: Latent diffusion models are natural and zero-shot speech and singing synthesiz- ers,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Naturalspeech 2: Latent diffusion models are natural and zero-shot speech and singing synthesiz- ers,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.780616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:45:15.461525Z digest=sha256:188a44240c5b3b44366f8f5e9d05be1d22631257a92af82786768a8e369c7fb6

Observation b78bc22a-fcd7-4a2c-84c2-66b3954ae4d2 · outbound

This paper cites Mega-TTS: Zero-Shot Text-to-Speech at Scale with Intrinsic Inductive Bias.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Mega-TTS: Zero-Shot Text-to-Speech at Scale with Intrinsic Inductive Bias

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T10:45:15.469169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:45:15.469169Z digest=sha256:9d88613f9c16afa4b539fc4ac645e148752c647e7f8f8dfaef0b8668ae68692a

Observation f942b0ff-5a44-45d4-8b9d-f6c1fa025cb7 · outbound

This paper cites V oicebox: Text-guided multilin- gual universal speech generation at scale,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning V oicebox: Text-guided multilin- gual universal speech generation at scale,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.760835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:45:15.479055Z digest=sha256:2a04cb5a3c08c14f1a0429086a0c5c656236d4b2b28ab4e9faa0d1934247a6f3

Observation bc11e927-d6f0-4d57-8504-d2201b043660 · outbound

This paper cites Listening while speaking: Speech chain by deep learning,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Listening while speaking: Speech chain by deep learning,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.741415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:45:15.515750Z digest=sha256:4c6968f26325bc09d11b3f2cd52c5273243e1a36fa7644683fe53bd5a8728050

Observation 868b8ddd-ff18-4177-8dfd-146abb81dbe7 · outbound

This paper cites Almost unsupervised text to speech and automatic speech recognition,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Almost unsupervised text to speech and automatic speech recognition,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.720787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:45:15.522695Z digest=sha256:08f77210f21e7fad82c7baaab47c56962f423b6d431dc6eeddd429c4fc60555b

Observation f8a5af9e-90d8-4ad9-a796-5392a21170be · outbound

This paper cites Prosodic features con- trol by symbols as input of sequence-to-sequence acoustic mod- eling for neural TTS,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Prosodic features con- trol by symbols as input of sequence-to-sequence acoustic mod- eling for neural TTS,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.701489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:45:15.530970Z digest=sha256:a679df1f7ce20537855f7d308437dedbb6d1bfeef8be13fbe596f3172c9a6f4d

Observation ef53aded-23e9-4299-9d45-52066e46cc17 · outbound

This paper cites Investigation of enhanced tacotron text-to-speech synthesis systems with self- attention for pitch accent language,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Investigation of enhanced tacotron text-to-speech synthesis systems with self- attention for pitch accent language,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.678475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:45:15.536055Z digest=sha256:51d876be3e21758ad09054f03e83b13a75a4f3735b62d6cd6f830d870b2879ea

Observation c08a13df-5cfb-450e-9d4b-c0447b39af09 · outbound

This paper cites A unified sequence-to-sequence front-end model for mandarin text-to-speech synthesis,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning A unified sequence-to-sequence front-end model for mandarin text-to-speech synthesis,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.646962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:45:15.543348Z digest=sha256:e853c8bc7743382b5847560df2023b93316d742a4e088b80486feae89733b203

Observation 9435527f-a9d1-4d16-9010-aa869e1c7848 · outbound

This paper cites A unified accent esti- mation method based on multi-task learning for Japanese text-to- speech,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning A unified accent esti- mation method based on multi-task learning for Japanese text-to- speech,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.619564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:45:15.549588Z digest=sha256:b3c60c895ff3214624230b4b06f6f5e465e684d94d89b710561b9b07b77a57e8

Observation a50ed950-a929-4c7e-93a5-db966a4ba857 · outbound

This paper cites Polyphone disambigua- tion and accent prediction using pre-trained language models in Japanese TTS front-end,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Polyphone disambigua- tion and accent prediction using pre-trained language models in Japanese TTS front-end,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.599984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:45:15.555619Z digest=sha256:47406282757df3c98d51b275bd2662555ab37449ad03038e0072129a4672c985

Observation f3392d7b-db93-4477-98bf-158eef3a6723 · outbound

This paper cites Enhancing Japanese text-to-speech ac- curacy with a novel combination Transformer-BERT-based G2P: Integrating pronunciation dictionaries and accent sandhi,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Enhancing Japanese text-to-speech ac- curacy with a novel combination Transformer-BERT-based G2P: Integrating pronunciation dictionaries and accent sandhi,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.570769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:45:15.565465Z digest=sha256:ab10626dbf17e811410c6a61445f42e8938713173f1694e3db01ac617d11d64c

Observation f552c249-4b3b-47c8-8c9d-4d3c2c4c0196 · outbound

This paper cites Audio- conditioned phonemic and prosodic annotation for building text- to-speech models from unlabeled speech data,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Audio- conditioned phonemic and prosodic annotation for building text- to-speech models from unlabeled speech data,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.544361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:45:15.572435Z digest=sha256:47e2782c430b5fb13a4917d53f07424957fbdb617aac5fb6a03930a8532df73a

Observation d1ed3e51-589b-422a-9cc4-5904314ad472 · outbound

This paper cites Robust speech recognition via large-scale weak supervision,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Robust speech recognition via large-scale weak supervision,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.517621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:45:15.577753Z digest=sha256:79bc2b02b19e464d1df0c9342a65b7e4671296204246f8ef9a974cd6253cec47

Observation c724ebc8-0b5b-4c22-b3c5-84e02602b416 · outbound

This paper cites Reazonspeech: A free and massive corpus for Japanese asr,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Reazonspeech: A free and massive corpus for Japanese asr,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.493668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:45:15.583381Z digest=sha256:b20465b410339a85871fbc0a73c39f2dbfa2aef6e142a9d2ca5e7ff490823abc

Observation 6e54d56c-00de-446f-88a5-9df6dee64ba7 · outbound

This paper cites YODAS: Youtube-oriented dataset for audio and speech,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning YODAS: Youtube-oriented dataset for audio and speech,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.471599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:45:15.591574Z digest=sha256:5b8934efdad8696b0f28a82f3a71efda0f1fbc00417424997f996eaf81214d40

Observation 3f41bf25-c650-4f98-9b8a-8920bdf1e30a · outbound

This paper cites OWSM-CTC: An open encoder-only speech foundation model for speech recog- nition, translation, and language identification,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning OWSM-CTC: An open encoder-only speech foundation model for speech recog- nition, translation, and language identification,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.440054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:45:15.598323Z digest=sha256:24c7dc327c52978b80be1bd54e0938792a8931a836314581d04df124d0830897

Observation 19463acd-44b6-48c9-86c0-122a72c50bce · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T10:45:15.603784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:45:15.603784Z digest=sha256:b2b7af89b22507a02e1907ecbd303ab16020e0923df3a312f74fe8351f891867

Observation 5d8ac709-023a-4726-bbaa-09be72f8e728 · outbound

This paper cites DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T10:45:15.616398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:45:15.616398Z digest=sha256:6d67abefdfafd316319afd3f317da9c256628c7ddd2dcdc8e97affa866dfc75b

Observation 17825b3e-71ae-4e4a-a71a-3f083d27a2c4 · outbound

This paper cites PnG BERT: Augmented BERT on phonemes and graphemes for neural TTS,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning PnG BERT: Augmented BERT on phonemes and graphemes for neural TTS,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.411517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:45:15.626660Z digest=sha256:b809f0a932dace89eee10bbc465f3df7f6ca7a6331689109b4f285003900d6ae

Observation 4fb61dce-a47f-449a-bdf0-d4cc80b894ab · outbound

This paper cites Phoneme-level BERT for enhanced prosody of text-to-speech with grapheme pre- dictions,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Phoneme-level BERT for enhanced prosody of text-to-speech with grapheme pre- dictions,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.385283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:45:15.633625Z digest=sha256:e8c0e7cf7a0e2d3d1f9a2a12a205bd7cb436a11bc45a858b68d1ca22d2b4f900

Observation 2ce861f4-1eea-4496-9484-b587b5dc1434 · outbound

This paper cites Miipher: A robust speech restoration model integrating self-supervised speech and text rep- resentations,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Miipher: A robust speech restoration model integrating self-supervised speech and text rep- resentations,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.366480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:45:15.648375Z digest=sha256:13dcaf806d21ebab24ff8a54efe9a233952623910f44d5aa924480ab4013e351

Observation 31d5dd8b-a730-4d18-9463-76b826057f19 · outbound

This paper cites Japanese text-to-speech syn- thesis system: Open JTalk,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Japanese text-to-speech syn- thesis system: Open JTalk,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.346905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:45:15.656518Z digest=sha256:1179bc5b8074d5e6bba6c0abb4c2279aa48f696e2baef1d5a9154c65c44ca89a

Observation 732fba98-843d-4fa3-a9c3-04c6d9d80334 · outbound

This paper cites Release of pre-trained mod- els for the Japanese language,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Release of pre-trained mod- els for the Japanese language,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.327855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:45:15.665499Z digest=sha256:795e6706ac04df7d540efe07d824b69e3f11e2a7e992ca496333090badea57b7

Observation 583a47e6-ec4b-4d60-9037-fd4384f561c0 · outbound

This paper cites Joint ctc-attention based end- to-end speech recognition using multi-task learning,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Joint ctc-attention based end- to-end speech recognition using multi-task learning,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.304977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:45:15.671693Z digest=sha256:0af00bb6bc948908d19d601458b9d5139da4f371ba08ec6d9320ef14db6c2c32

Observation 43022c1b-e346-4bc8-8cde-d01027fbcd08 · outbound

This paper cites Relaxing the conditional indepen- dence assumption of ctc-based asr by conditioning on interme- diate predictions,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Relaxing the conditional indepen- dence assumption of ctc-based asr by conditioning on interme- diate predictions,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.285924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:45:15.678701Z digest=sha256:237c568b22fba4547ede3408baf0b8f49157cc4f1e1c403f53e29107958641c4

Observation 47071f34-7d6e-493e-900b-270466e0f06d · outbound

This paper cites BERT meets CTC: New formulation of end-to-end speech recognition with pre-trained masked language model,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning BERT meets CTC: New formulation of end-to-end speech recognition with pre-trained masked language model,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.256179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:45:15.686812Z digest=sha256:c72a56c5b64a9793bbef53efb93289a1a1383e873007e0df9cc1c6f942b99c3a

Observation fc735343-fa35-4e9a-a7d3-262905b81062 · outbound

This paper cites Improving speech recognition error prediction for modern and off-the-shelf speech recognizers,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Improving speech recognition error prediction for modern and off-the-shelf speech recognizers,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.236603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:45:15.692902Z digest=sha256:d7b8383be9a572b9ba84864b64d3a7da60defa9243a7f2fb91aafca7470ab967

Observation b26f4c93-ca02-47c9-bb40-7bb3d732ce36 · outbound

This paper cites The theory of dynamic programming,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning The theory of dynamic programming,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.216220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:45:15.698387Z digest=sha256:8857d2f7b67d13dbce8e4b103937f6fd92d4b4d7cbf89d4ad21dba3b478399b2

Observation cbda27c5-89e2-45a0-8e39-e82525343f51 · outbound

This paper cites JSUT and JVS: Free Japanese voice corpora for accelerating speech synthesis re- search,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning JSUT and JVS: Free Japanese voice corpora for accelerating speech synthesis re- search,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.195859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:45:15.704647Z digest=sha256:1fe1008e2894ed09e2f0581cc8f9944a3dc61d34f0237004c7c0442ddab57da9

Observation 2059ae9e-9677-44c5-8b28-8d2c822ac4c3 · outbound

This paper cites Applying conditional random fields to Japanese morphological analysis,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Applying conditional random fields to Japanese morphological analysis,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.176497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:45:15.711298Z digest=sha256:a838e611b925d0068b108a96ff2be75d98ec06482e9ab333d313a3bdec0d031b

Observation 5822a7ae-f875-4f4f-8bca-30cb0505e99b · outbound

This paper cites A proper approach to Japanese morphological analysis: Dictionary, model, and eval- uation.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning A proper approach to Japanese morphological analysis: Dictionary, model, and eval- uation

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.151737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:45:15.717972Z digest=sha256:f2cf04602012e1b1bd5bc9b2d88c6bd45a6beaab3e45b3beb651043085a2df8e

Observation eaa20038-d109-494e-84f1-59dd387d048a · outbound

This paper cites Period VITS: Varia- tional inference with explicit pitch modeling for end-to-end emo- tional speech synthesis,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Period VITS: Varia- tional inference with explicit pitch modeling for end-to-end emo- tional speech synthesis,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.129554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:45:15.723592Z digest=sha256:73b20c79d7d744c276a5969c97f8a60f311e6ccbb725611a1d200b6d0f8db2f0

Observation dd86049e-a605-499f-986c-042b9a825cad · outbound

This paper cites DEMAND: A col- lection of multi-channel recordings of acoustic noise in diverse environments,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning DEMAND: A col- lection of multi-channel recordings of acoustic noise in diverse environments,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.111311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:45:15.731270Z digest=sha256:a2dc2f81c31013df8dc5760f5704f47661f4f2c3f8d39a7881544c6e865d368b

Observation 1a35e2c5-68bb-42df-8bc1-73b3a4f86434 · outbound

This paper cites The ACE challenge—corpus description and performance evaluation,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning The ACE challenge—corpus description and performance evaluation,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.088377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:45:15.740563Z digest=sha256:683973bcb4362d74e6864352f0e424dc9a638f8138818fc772a4ee0cf5466c86

Observation 296c0a6e-4eca-4d0f-9891-b8969d28c888 · outbound

This paper cites End-to-end ASR to jointly predict transcriptions and linguistic annotations,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning End-to-end ASR to jointly predict transcriptions and linguistic annotations,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.069677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:45:15.750147Z digest=sha256:549ba816505a65f144356df52b5a6f8c71c7974c223f9aac291a7dcd4a706283

Observation 2aaf2e9b-4a5f-415d-b22c-84b3264f98ac · outbound

This paper cites Building competitive direct acoustics-to-word models for english conversa- tional speech recognition,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Building competitive direct acoustics-to-word models for english conversa- tional speech recognition,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.051902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:45:15.760108Z digest=sha256:212ab91f002d950911051134941ec42c4868ebe99cccd9e7e74b2feb0bc86dd2

Observation f0633d88-a512-411c-8e16-5b87fe41b9e0 · outbound

This paper cites Joint speech recognition and speaker diarization via sequence transduction,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Joint speech recognition and speaker diarization via sequence transduction,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.028606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:45:15.769922Z digest=sha256:cabfa342b876b1d5106ce74093a689d03cae3d354744a1461732b4a9a25e3b37

Observation 1f62b35d-1388-466c-87de-fbdaad28620a · outbound

This paper cites Uncon- strained many-to-many alignment for automatic pronunciation an- notation,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Uncon- strained many-to-many alignment for automatic pronunciation an- notation,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.004563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:45:15.777135Z digest=sha256:dd41d8bd4f07ec483675596ef16f2dec1d8fdf901270cf772ed822a4d8c57b31

Observation 800a57a2-ed89-4474-9bb9-d7ba9a01ac0e · outbound

This paper cites Evaluation of many-to-many alignment algorithm by auto- matic pronunciation annotation using web text mining,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Evaluation of many-to-many alignment algorithm by auto- matic pronunciation annotation using web text mining,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:15.982204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:45:15.786786Z digest=sha256:e946b1bcc371f5a4310a79eb23812e7db6922fa1942f4847cc047dd3d78f2c73

Pith citing papers

Observation 3828ef2d-8103-4827-9723-0e4ff294770b · inbound

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning cites this paper.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T10:45:15.953103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:45:15.390168Z digest=sha256:77c4d26f908e7cf5ba2c99abdbedc329f809a9d7a2724a825e7495a1f1db32a5