Pith. sign in

Paper Citation Record · LEDGER

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning

As of 7 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 1 inbound Pith citation observation for arXiv:2506.04527.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.04527 v1

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:45:15.786786Z

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:45:15.390168Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T10:45:15.943522Z

Reference resolution

49 of 49 outbound references displayed

  • verified exact0
  • verified fuzzy45
  • unresolved3
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8d183f01-308a-4a4d-b3a1-f50a73b0244b · outbound

This paper cites For training high-quality and diverse-styled TTS models, a large amount of text-speech paired data is re- quired [4, 5].

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning For training high-quality and diverse-styled TTS models, a large amount of text-speech paired data is re- quired [4, 5]

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.968584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:45:15.376258Z digest=sha256:d7a6f218da7f4a6f20c7a1150781fe86b75909e4c1e8f8f4a648ef961fabaf66

Observation 3828ef2d-8103-4827-9723-0e4ff294770b · outbound

This paper cites Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T10:45:15.953103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:45:15.390168Z digest=sha256:907e9e85cbd8ebf4bc3496c1efa114f651ae95103867c361a9853fe93ea8354a

Observation d83c57ff-e9c2-4051-a2b8-5ffde5a2fcff · outbound

This paper cites <blank>.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning <blank>

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.947649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:45:15.397152Z digest=sha256:c872cd1df3cff7186d1a8cc46f469075df8e03d816b17d842d301834580a7812

Observation 283541ee-6f26-4a93-adb5-f166d39e8b93 · outbound

This paper cites Evaluation of proposed annotation model 4.1.1.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Evaluation of proposed annotation model 4.1.1

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.928154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:45:15.408097Z digest=sha256:84256de2eafbd155324e2940735d0476816ce223f51a98e34aca4c7d9b00806f

Observation dd5fb3c8-6404-4130-b3ef-a50ee30c9db0 · outbound

This paper cites Specifically, we utilized the basic5000 subset along with its manually annotated TTS labels1.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Specifically, we utilized the basic5000 subset along with its manually annotated TTS labels1

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.903592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:45:15.416347Z digest=sha256:68013bc82a2e5289eb93d553d6584c54f7c59e80d724fadf5cdf8de1c5ea48bf

Observation 88bb5455-04e4-464a-b6ce-7b7f26498546 · outbound

This paper cites ”, (2) Accent change from low to high “[.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning ”, (2) Accent change from low to high “[

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.884422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:45:15.422376Z digest=sha256:00da2cd512f110075c1f45d1880cb24ef9e9aad5eaa2fbab99f0401579858545

Observation 8a7ac17f-3c7a-49fd-b1c8-b4a1f7ba2937 · outbound

This paper cites Applying this method to downstream tasks be- yond textual accent estimation is a challenge for future work.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Applying this method to downstream tasks be- yond textual accent estimation is a challenge for future work

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.865438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:45:15.428341Z digest=sha256:da1a2f6e783eff8e24fcde535318b0a9df57ee632addaeca0adc13e4e33bb2ef

Observation 471f62fb-b9d2-4ecb-8445-85cff0353220 · outbound

This paper cites Tacotron: Towards end-to-end speech synthesis,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Tacotron: Towards end-to-end speech synthesis,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.843284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:45:15.436595Z digest=sha256:ab1f0b61468ac08625d89e90813aee0a901c2282f796fa13101bfd0acbc4af5b

Observation 0d18b20c-8ab5-4230-87dd-921714906ef3 · outbound

This paper cites Fastspeech 2: Fast and high-quality end-to-end text to speech,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Fastspeech 2: Fast and high-quality end-to-end text to speech,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.820773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:45:15.444523Z digest=sha256:4e4126dd1bc8b8eb8fc73840b0fc7f9c3d931eae4e05a72a2bfea3f9bf63cc80

Observation 76c0dd0c-e7cd-4f32-8924-f033a51a921c · outbound

This paper cites Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.800603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:45:15.455133Z digest=sha256:8e5bdb2aeb82d1abba20c775279c6462c53dd1ea99917b5626cd7dceff964cf5

Observation 30c4032d-9850-4d5d-b62e-ae3728234cc1 · outbound

This paper cites Naturalspeech 2: Latent diffusion models are natural and zero-shot speech and singing synthesiz- ers,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Naturalspeech 2: Latent diffusion models are natural and zero-shot speech and singing synthesiz- ers,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.780616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:45:15.461525Z digest=sha256:ef21f3ba1460fc569409ec9764f64b175f4581e2ec0d79097227c3e18a120804

Observation b78bc22a-fcd7-4a2c-84c2-66b3954ae4d2 · outbound

This paper cites Mega-TTS: Zero-Shot Text-to-Speech at Scale with Intrinsic Inductive Bias.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Mega-TTS: Zero-Shot Text-to-Speech at Scale with Intrinsic Inductive Bias

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T10:45:15.469169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:45:15.469169Z digest=sha256:bb04e457163f180b71463001767547f4fc83ccd21da70611857448143d2fac73

Observation f942b0ff-5a44-45d4-8b9d-f6c1fa025cb7 · outbound

This paper cites V oicebox: Text-guided multilin- gual universal speech generation at scale,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning V oicebox: Text-guided multilin- gual universal speech generation at scale,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.760835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:45:15.479055Z digest=sha256:e72535ded0e6f6d8aab486c573b676f1bb1a2587c1c40d7c7bda00b5e35eee46

Observation bc11e927-d6f0-4d57-8504-d2201b043660 · outbound

This paper cites Listening while speaking: Speech chain by deep learning,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Listening while speaking: Speech chain by deep learning,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.741415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:45:15.515750Z digest=sha256:c27c095a5b075c242a762e6ec9820f2412100bdf17f745c87440b5d219638f0f

Observation 868b8ddd-ff18-4177-8dfd-146abb81dbe7 · outbound

This paper cites Almost unsupervised text to speech and automatic speech recognition,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Almost unsupervised text to speech and automatic speech recognition,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.720787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:45:15.522695Z digest=sha256:726816ad95d975f4ab099b063ee6aecacc2b522daded6ff00113d8a61d14240a

Observation f8a5af9e-90d8-4ad9-a796-5392a21170be · outbound

This paper cites Prosodic features con- trol by symbols as input of sequence-to-sequence acoustic mod- eling for neural TTS,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Prosodic features con- trol by symbols as input of sequence-to-sequence acoustic mod- eling for neural TTS,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.701489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:45:15.530970Z digest=sha256:bca1cf295777be0ad966d5a74a0c24780119da8d598dc014126ad3b066b5a1a0

Observation ef53aded-23e9-4299-9d45-52066e46cc17 · outbound

This paper cites Investigation of enhanced tacotron text-to-speech synthesis systems with self- attention for pitch accent language,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Investigation of enhanced tacotron text-to-speech synthesis systems with self- attention for pitch accent language,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.678475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:45:15.536055Z digest=sha256:d81ecb5b58b9e44c7895e0a59ee9c985144de5506ad0b97d5112a392e74b5058

Observation c08a13df-5cfb-450e-9d4b-c0447b39af09 · outbound

This paper cites A unified sequence-to-sequence front-end model for mandarin text-to-speech synthesis,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning A unified sequence-to-sequence front-end model for mandarin text-to-speech synthesis,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.646962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:45:15.543348Z digest=sha256:13ebcd63e5cad546ea1f3a424b5d37c19d3a1e35f2ce2b520765ba9a5eb30738

Observation 9435527f-a9d1-4d16-9010-aa869e1c7848 · outbound

This paper cites A unified accent esti- mation method based on multi-task learning for Japanese text-to- speech,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning A unified accent esti- mation method based on multi-task learning for Japanese text-to- speech,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.619564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:45:15.549588Z digest=sha256:7fb983bfd20a7f401492110de922850cf4fcaf2568666833a87395d2b9cf2d21

Observation a50ed950-a929-4c7e-93a5-db966a4ba857 · outbound

This paper cites Polyphone disambigua- tion and accent prediction using pre-trained language models in Japanese TTS front-end,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Polyphone disambigua- tion and accent prediction using pre-trained language models in Japanese TTS front-end,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.599984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:45:15.555619Z digest=sha256:024e991ee57d3a3b9e0522886010569efe7a9eff2adc407d159739d3485c8820

Observation f3392d7b-db93-4477-98bf-158eef3a6723 · outbound

This paper cites Enhancing Japanese text-to-speech ac- curacy with a novel combination Transformer-BERT-based G2P: Integrating pronunciation dictionaries and accent sandhi,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Enhancing Japanese text-to-speech ac- curacy with a novel combination Transformer-BERT-based G2P: Integrating pronunciation dictionaries and accent sandhi,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.570769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:45:15.565465Z digest=sha256:635d31ad7fa09c5bb580ead24771c50a4479f1ddeda2699d5c23507cc55764f6

Observation f552c249-4b3b-47c8-8c9d-4d3c2c4c0196 · outbound

This paper cites Audio- conditioned phonemic and prosodic annotation for building text- to-speech models from unlabeled speech data,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Audio- conditioned phonemic and prosodic annotation for building text- to-speech models from unlabeled speech data,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.544361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:45:15.572435Z digest=sha256:24f51fa4750f14cdc8908e482a0dda35c7a4fea6d616d778d6781983e0467e38

Observation d1ed3e51-589b-422a-9cc4-5904314ad472 · outbound

This paper cites Robust speech recognition via large-scale weak supervision,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Robust speech recognition via large-scale weak supervision,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.517621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:45:15.577753Z digest=sha256:63ed1b6be78515d3d96a7b37652e19eb1a6b94a5c1797c6b9fbf2df2b75d49de

Observation c724ebc8-0b5b-4c22-b3c5-84e02602b416 · outbound

This paper cites Reazonspeech: A free and massive corpus for Japanese asr,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Reazonspeech: A free and massive corpus for Japanese asr,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.493668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:45:15.583381Z digest=sha256:ae328b3e4c9af6814db5a4d2feef067f85831cac8144f5286834d1df3e4654dc

Observation 6e54d56c-00de-446f-88a5-9df6dee64ba7 · outbound

This paper cites YODAS: Youtube-oriented dataset for audio and speech,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning YODAS: Youtube-oriented dataset for audio and speech,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.471599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:45:15.591574Z digest=sha256:fcc0e04cebc8bbdddb0bb2c622e041774052a792224ed4dca52c494dde35e329

Observation 3f41bf25-c650-4f98-9b8a-8920bdf1e30a · outbound

This paper cites OWSM-CTC: An open encoder-only speech foundation model for speech recog- nition, translation, and language identification,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning OWSM-CTC: An open encoder-only speech foundation model for speech recog- nition, translation, and language identification,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.440054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:45:15.598323Z digest=sha256:84bdf42e5bff1221118fce940c18f7eef009fb95d4b0112224fadf8242a95010

Observation 19463acd-44b6-48c9-86c0-122a72c50bce · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T10:45:15.603784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:45:15.603784Z digest=sha256:e4ec22edaf32eabfbf8d450cd130b70534b6c4904b69fa040b1d1cb5d5cd84d1

Observation 5d8ac709-023a-4726-bbaa-09be72f8e728 · outbound

This paper cites DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T10:45:15.616398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:45:15.616398Z digest=sha256:45d19a5950956cb91621bf7e7020990f4926df63e5bfcfd9c240e122e8055ab9

Observation 17825b3e-71ae-4e4a-a71a-3f083d27a2c4 · outbound

This paper cites PnG BERT: Augmented BERT on phonemes and graphemes for neural TTS,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning PnG BERT: Augmented BERT on phonemes and graphemes for neural TTS,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.411517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:45:15.626660Z digest=sha256:998f8884aa8b8001e507640f3b8da86e75d4e576c021b3d8c28a57cb9797e310

Observation 4fb61dce-a47f-449a-bdf0-d4cc80b894ab · outbound

This paper cites Phoneme-level BERT for enhanced prosody of text-to-speech with grapheme pre- dictions,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Phoneme-level BERT for enhanced prosody of text-to-speech with grapheme pre- dictions,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.385283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:45:15.633625Z digest=sha256:4737e2d274f6d80d126ecbbcc70b67b7100206adcb9a6c6652538f62fd135c94

Observation 2ce861f4-1eea-4496-9484-b587b5dc1434 · outbound

This paper cites Miipher: A robust speech restoration model integrating self-supervised speech and text rep- resentations,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Miipher: A robust speech restoration model integrating self-supervised speech and text rep- resentations,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.366480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:45:15.648375Z digest=sha256:78f764d3d47c33ce4f11223fcad5002ae7800b29219506fac2719bf33fd8b017

Observation 31d5dd8b-a730-4d18-9463-76b826057f19 · outbound

This paper cites Japanese text-to-speech syn- thesis system: Open JTalk,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Japanese text-to-speech syn- thesis system: Open JTalk,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.346905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:45:15.656518Z digest=sha256:182d4d5d622ec618e54cd4bf17d7bdc58aece8d821e070a739f5d257dbdcbded

Observation 732fba98-843d-4fa3-a9c3-04c6d9d80334 · outbound

This paper cites Release of pre-trained mod- els for the Japanese language,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Release of pre-trained mod- els for the Japanese language,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.327855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:45:15.665499Z digest=sha256:447e13ad1eb38655ef19f20f01ceefb6a727b2d34d2d52c8ad984c9134035c83

Observation 583a47e6-ec4b-4d60-9037-fd4384f561c0 · outbound

This paper cites Joint ctc-attention based end- to-end speech recognition using multi-task learning,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Joint ctc-attention based end- to-end speech recognition using multi-task learning,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.304977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:45:15.671693Z digest=sha256:1d8c475c597a0e2019d53f8969788adfd8ffa00d655ab3698c16db32cfc111c0

Observation 43022c1b-e346-4bc8-8cde-d01027fbcd08 · outbound

This paper cites Relaxing the conditional indepen- dence assumption of ctc-based asr by conditioning on interme- diate predictions,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Relaxing the conditional indepen- dence assumption of ctc-based asr by conditioning on interme- diate predictions,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.285924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:45:15.678701Z digest=sha256:93085057dca4636a2a1fae1b531a4f3b94aeceb724858affa022a2e5a366850f

Observation 47071f34-7d6e-493e-900b-270466e0f06d · outbound

This paper cites BERT meets CTC: New formulation of end-to-end speech recognition with pre-trained masked language model,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning BERT meets CTC: New formulation of end-to-end speech recognition with pre-trained masked language model,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.256179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:45:15.686812Z digest=sha256:87f340c6b11033dad63961860b25cb4ec335374171dd05e1d9a14979d9656ec3

Observation fc735343-fa35-4e9a-a7d3-262905b81062 · outbound

This paper cites Improving speech recognition error prediction for modern and off-the-shelf speech recognizers,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Improving speech recognition error prediction for modern and off-the-shelf speech recognizers,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.236603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:45:15.692902Z digest=sha256:400d06a832b3bbfc0aafdf6668b3a4163acf70d9c4722f154bf41c318f4e7884

Observation b26f4c93-ca02-47c9-bb40-7bb3d732ce36 · outbound

This paper cites The theory of dynamic programming,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning The theory of dynamic programming,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.216220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:45:15.698387Z digest=sha256:3cbf737cd09ed04bb506ed857f61b5f5c517386f7bd16d6f49eae7da6db19aa5

Observation cbda27c5-89e2-45a0-8e39-e82525343f51 · outbound

This paper cites JSUT and JVS: Free Japanese voice corpora for accelerating speech synthesis re- search,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning JSUT and JVS: Free Japanese voice corpora for accelerating speech synthesis re- search,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.195859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:45:15.704647Z digest=sha256:ace52e953b339b7fee197654eef2e3a6eb9ad23b0fd109ea457c3b26df7874ea

Observation 2059ae9e-9677-44c5-8b28-8d2c822ac4c3 · outbound

This paper cites Applying conditional random fields to Japanese morphological analysis,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Applying conditional random fields to Japanese morphological analysis,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.176497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:45:15.711298Z digest=sha256:b91bbfa3e7c02d0aa5546f350e55816fa782043a223cd0f7859a76012780f71c

Observation 5822a7ae-f875-4f4f-8bca-30cb0505e99b · outbound

This paper cites A proper approach to Japanese morphological analysis: Dictionary, model, and eval- uation.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning A proper approach to Japanese morphological analysis: Dictionary, model, and eval- uation

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.151737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:45:15.717972Z digest=sha256:db888bf7d91eb3868284436f76edfad319811ecc30f492bf3ea33dd365d2b639

Observation eaa20038-d109-494e-84f1-59dd387d048a · outbound

This paper cites Period VITS: Varia- tional inference with explicit pitch modeling for end-to-end emo- tional speech synthesis,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Period VITS: Varia- tional inference with explicit pitch modeling for end-to-end emo- tional speech synthesis,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.129554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:45:15.723592Z digest=sha256:e5da078916a2e7603a430af45a06236dbe1335da861b5f357682a7c31eb29447

Observation dd86049e-a605-499f-986c-042b9a825cad · outbound

This paper cites DEMAND: A col- lection of multi-channel recordings of acoustic noise in diverse environments,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning DEMAND: A col- lection of multi-channel recordings of acoustic noise in diverse environments,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.111311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:45:15.731270Z digest=sha256:ea037118c6d5e429bbfe654b76229cfc62cc27e0597a7d2d56fc7d832822f936

Observation 1a35e2c5-68bb-42df-8bc1-73b3a4f86434 · outbound

This paper cites The ACE challenge—corpus description and performance evaluation,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning The ACE challenge—corpus description and performance evaluation,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.088377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:45:15.740563Z digest=sha256:2419a113763ebe55bcdb9ee27f034363f8c7e317553396b570f3b0379773153f

Observation 296c0a6e-4eca-4d0f-9891-b8969d28c888 · outbound

This paper cites End-to-end ASR to jointly predict transcriptions and linguistic annotations,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning End-to-end ASR to jointly predict transcriptions and linguistic annotations,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.069677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:45:15.750147Z digest=sha256:9ce745c0413029a2fb1f5270e8c7d787c6d8a94d224054ebbc4035d22adb3693

Observation 2aaf2e9b-4a5f-415d-b22c-84b3264f98ac · outbound

This paper cites Building competitive direct acoustics-to-word models for english conversa- tional speech recognition,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Building competitive direct acoustics-to-word models for english conversa- tional speech recognition,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.051902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:45:15.760108Z digest=sha256:63c4ecdc693885ae68eee42a0ee1529c34e041ceaacefb8bdcc9fd9de66f550c

Observation f0633d88-a512-411c-8e16-5b87fe41b9e0 · outbound

This paper cites Joint speech recognition and speaker diarization via sequence transduction,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Joint speech recognition and speaker diarization via sequence transduction,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.028606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:45:15.769922Z digest=sha256:006e1ae5fac249be3dc40340ee3192a77335faaa1b4065f1775470b731d2ca4a

Observation 1f62b35d-1388-466c-87de-fbdaad28620a · outbound

This paper cites Uncon- strained many-to-many alignment for automatic pronunciation an- notation,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Uncon- strained many-to-many alignment for automatic pronunciation an- notation,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:16.004563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:45:15.777135Z digest=sha256:99f6380b62adcb5c3f45a2edc278c546f71571f25ec57c417d90018d7ac11c79

Observation 800a57a2-ed89-4474-9bb9-d7ba9a01ac0e · outbound

This paper cites Evaluation of many-to-many alignment algorithm by auto- matic pronunciation annotation using web text mining,.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Evaluation of many-to-many alignment algorithm by auto- matic pronunciation annotation using web text mining,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:45:15.982204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:45:15.786786Z digest=sha256:d0db4cceaf2fbc367f44dfa45b9641b874103297b4cabc345c8ee736b295f414

Pith citing papers

Observation 3828ef2d-8103-4827-9723-0e4ff294770b · inbound

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning cites this paper.

Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T10:45:15.953103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:45:15.390168Z digest=sha256:907e9e85cbd8ebf4bc3496c1efa114f651ae95103867c361a9853fe93ea8354a