Pith. sign in

Paper Citation Record · LEDGER

Prosody Labeling with Phoneme-BERT and Speech Foundation Models

As of 19 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 1 inbound Pith citation observation for arXiv:2507.03912.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.03912 v1

Coverage vector

measured 53 of 53 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:03:18.031484Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:03:14.249592Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T20:03:18.073334Z

Reference resolution

53 of 53 outbound references displayed

  • verified exact1
  • verified fuzzy43
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a8e9006e-c63d-4fb7-8a0a-0c41862ebe45 · outbound

This paper cites Prosody Labeling with Phoneme-BERT and Speech Foundation Models.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models Prosody Labeling with Phoneme-BERT and Speech Foundation Models

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T20:03:18.076461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T20:03:14.249592Z digest=sha256:cee0d5047022b1e9d06638395652a86d9c480cbb4fb9ae73800b142fd1ea4922

Observation 7a37b1f9-63a0-4b0b-b6d1-b84466d39149 · outbound

This paper cites [8] have also conducted automatic prosody an- notation for constructing a speech synthesis database, similar to the objective of our study.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models [8] have also conducted automatic prosody an- notation for constructing a speech synthesis database, similar to the objective of our study

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:03:18.570098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T20:03:14.334997Z digest=sha256:aee383f6fcedb2eb34f404e11a7fcae8e0950f137451c37eb77b2d2476dad69a

Observation 486d9ad2-4af9-4aea-bb4d-14ec30cd5b01 · outbound

This paper cites ashita wa hare desuka.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models ashita wa hare desuka

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:03:18.562280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T20:03:14.448986Z digest=sha256:783bee959baa461cdbfc36c09565a20cc9209df791f4a8f10be0dc6db3a20054

Observation 25f7c816-b74d-4ccd-bd15-7b8514a2d4d3 · outbound

This paper cites Figure 3 shows the outline of the proposed model.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models Figure 3 shows the outline of the proposed model

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:03:18.554087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T20:03:14.576338Z digest=sha256:790a70dc2d7ea664ad6e6c431f517d973bd64711b306ac0a94c6c305e5db6c29

Observation 7f5ba96d-e444-4a68-8274-f3d160d541e9 · outbound

This paper cites *”: Other symbol that was neither a high-low transition nor a phrase boundary. • “[.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models *”: Other symbol that was neither a high-low transition nor a phrase boundary. • “[

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:03:18.546322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T20:03:14.718140Z digest=sha256:f2de654acb047786e15b87db68c7fb7559e765fb3176c72a33c655542c979790

Observation d85e9b1d-ba7e-45fe-98fd-6753788a252e · outbound

This paper cites an unresolved cited work.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:03:18.530571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T20:03:14.941364Z digest=sha256:dec3d54b9608575093fd4e698424260c7fa419efd28d4309c0456e8fae43e75a

Observation ad3aa484-253a-4abe-8ef5-5b8e17d39938 · outbound

This paper cites Automatic prosody annotation with pre-trained text- speech model,.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models Automatic prosody annotation with pre-trained text- speech model,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:03:18.474594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T20:03:15.460095Z digest=sha256:8ff593284c1d8027301c490b8b34e7f9f29f5d67ef9d5119b7b25352b3100f84

Observation 24699a4f-1edb-44b9-aef4-f4cd2e9745ab · outbound

This paper cites Mora-level prosody prediction for text-to-speech us- ing japanese BERT without accentual labels,.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models Mora-level prosody prediction for text-to-speech us- ing japanese BERT without accentual labels,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:03:18.522972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T20:03:14.995512Z digest=sha256:7481dc51c42177cff74d9e02c7a60877e253071d4713b26303dc870e96f4f8c8

Observation 91978ff0-5f1a-43d0-8fac-dce5a42e6dec · outbound

This paper cites Semi-supervised prosody model- ing using deep Gaussian process latent variable model,.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models Semi-supervised prosody model- ing using deep Gaussian process latent variable model,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:03:18.515166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T20:03:15.059899Z digest=sha256:a41fa186758e74095e060a40d804e62205e4d9f7ce19cf5649d7220dfdfbf322

Observation 2800bd16-57a5-4d68-b532-08a393ce229a · outbound

This paper cites Ac- cent modeling of low-resourced dialect in pitch accent language using variational autoencoder,.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models Ac- cent modeling of low-resourced dialect in pitch accent language using variational autoencoder,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:03:18.506457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T20:03:15.121236Z digest=sha256:9005dbb90adab22ecefd65fbe82d5a5a945a78967457b8befd3d1d352e3c4a43

Observation 17c90688-3bf3-4167-b89c-4e05771cc5d0 · outbound

This paper cites Cross-dialect text- to-speech in pitch-accent language incorporating multi-dialect phoneme-level BERT,.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models Cross-dialect text- to-speech in pitch-accent language incorporating multi-dialect phoneme-level BERT,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:03:18.498411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T20:03:15.224190Z digest=sha256:5edc76c2cb999ed9ca0504a41b7628aed32a7375c9119070e88dcbddb9b0b6c6

Observation 7752be24-91c1-4332-92de-fe68656049ef · outbound

This paper cites Multi-modal automatic prosody anno- tation with contrastive pretraining of speech-silence and word- punctuation,.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models Multi-modal automatic prosody anno- tation with contrastive pretraining of speech-silence and word- punctuation,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:03:18.490401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T20:03:15.277753Z digest=sha256:9131bfd688b215950b9292983ce327e93624c0237412c77c32fa46822e2f4dd2

Observation 9f4f8240-b7e3-4f27-b812-aa1f6842bcc1 · outbound

This paper cites Language-independent prosody-enhanced speech representations for multilingual speech synthesis,.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models Language-independent prosody-enhanced speech representations for multilingual speech synthesis,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:03:18.482741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T20:03:15.376794Z digest=sha256:eff7339fd5a1f8062883b0ab95a75a9b363179941ba1073e7e557a18d1a60f15

Observation 8377d84f-481f-43c6-b405-5126743c9a7e · outbound

This paper cites BERT: Pre- training of deep bidirectional transformers for language under- standing,.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models BERT: Pre- training of deep bidirectional transformers for language under- standing,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:03:18.419024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T20:03:15.971177Z digest=sha256:984b952388720958de722e535541b99d5e973e2c840cbfbff3a7a77f8059406b

Observation 501b9959-6d09-43de-bc6f-d6a5715126f7 · outbound

This paper cites Audio- conditioned phonemic and prosodic annotation for building text- to-speech models from unlabeled speech data,.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models Audio- conditioned phonemic and prosodic annotation for building text- to-speech models from unlabeled speech data,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:03:18.465831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T20:03:15.540553Z digest=sha256:97099b3b65453b5d2cb618a9dc1b739eea282a4d38bd9dfa3b5fb7322b42e62c

Observation 02fafbbd-40ca-4ff9-8c6b-9d4d91b2a8c8 · outbound

This paper cites Low-resourced phonetic and prosodic feature estimation with self-supervised-learning-based acoustic modeling,.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models Low-resourced phonetic and prosodic feature estimation with self-supervised-learning-based acoustic modeling,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:03:18.457670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T20:03:15.611761Z digest=sha256:b02dfdc2df23ec1984b3b1d9462494a73b799cb26923e59073fcc11aa5812670

Observation 3a4335c3-5762-4f48-90b7-61b6610ca6c1 · outbound

This paper cites Cross-lingual speech- based ToBI label generation using bidirectional LSTM,.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models Cross-lingual speech- based ToBI label generation using bidirectional LSTM,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:03:18.449384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T20:03:15.676428Z digest=sha256:7958afb342c2ebd4ef809d5f33ca32228ec3d5d801ac87e40f80cf5883b1f69c

Observation 994affb5-0eb8-48ed-8edd-474a1541b990 · outbound

This paper cites A unified accent esti- mation method based on multi-task learning for Japanese text-to- speech,.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models A unified accent esti- mation method based on multi-task learning for Japanese text-to- speech,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:03:18.440586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T20:03:15.742927Z digest=sha256:903a03147471c5d9e8d78d6018aee909cbf12cf917a4ef76ec0466d4aaa81c30

Observation 73a9539d-6bdd-4c62-bbfc-1831c6547496 · outbound

This paper cites Wav2ToBI: a new approach to automatic ToBI transcription,.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models Wav2ToBI: a new approach to automatic ToBI transcription,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:03:18.432635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T20:03:15.809588Z digest=sha256:bb854c8ea2aee45d2ec808755abfd6e96ae4d9ccb968eb4cde450522d0d5ff79

Observation 7d985065-12a4-469d-ae23-b642b0469279 · outbound

This paper cites Conformer: Convolution-augmented transformer for speech recognition,.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models Conformer: Convolution-augmented transformer for speech recognition,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T20:03:15.886329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:03:15.886329Z digest=sha256:5e5213ee437385591ff1af29a8142d226d26dcb18a3d815346626c9034893270

Observation 75ce9da2-3cf3-4a12-8a6a-6f6eaa5606a5 · outbound

This paper cites Investigation of enhanced Tacotron text-to-speech synthesis systems with self- attention for pitch accent language,.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models Investigation of enhanced Tacotron text-to-speech synthesis systems with self- attention for pitch accent language,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:03:18.361910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T20:03:16.479507Z digest=sha256:0b28ef26f058fd0a9fd93ab4b2f9fddbec39911c2908a04cb8055a7e5944ead3

Observation 5b927204-f5ad-474e-aaba-5cad605fe1e3 · outbound

This paper cites Robust speech recognition via large-scale weak su- pervision,.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models Robust speech recognition via large-scale weak su- pervision,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:03:18.411280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T20:03:16.049594Z digest=sha256:7a37f823584a49f5f0c8ea42bb043fe58707d2d248f03daf4249cf3f339e5811

Observation 01d1492d-4e95-403e-88d1-d12a7e03f4e1 · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech representa- tions,.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models wav2vec 2.0: A framework for self-supervised learning of speech representa- tions,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:03:18.403453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T20:03:16.124677Z digest=sha256:0ea8c426c43a3e697dcb4591c03bd2506474978a062cad6694e830e3224dd6ff

Observation e28cb59e-5e18-41cf-9318-cede3a3f7cf9 · outbound

This paper cites Pre-trained text embeddings for enhanced text-to- speech synthesis,.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models Pre-trained text embeddings for enhanced text-to- speech synthesis,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:03:18.395326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T20:03:16.187592Z digest=sha256:d5144b4a1604b3dc6182a1af7d78d38c41cccc9d5f90964e058df84cb197574d

Observation 91b6ee43-c70f-4403-9818-071a4513fae6 · outbound

This paper cites Improving prosody with linguistic and BERT derived features in multi-speaker based Mandarin Chinese neural TTS,.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models Improving prosody with linguistic and BERT derived features in multi-speaker based Mandarin Chinese neural TTS,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:03:18.387703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T20:03:16.256372Z digest=sha256:9f94d9288ba753dc1ead1b7404481ca1e74ad09492ca10f2d7ab98d5fc7322c0

Observation 0596f701-8685-4c3f-93ef-90ef31690aa2 · outbound

This paper cites Improving the prosody of RNN-based English text-to-speech synthesis by incorporating a BERT model,.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models Improving the prosody of RNN-based English text-to-speech synthesis by incorporating a BERT model,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:03:18.379448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T20:03:16.342536Z digest=sha256:52aca3e53c1d0a94dffdd335d6a308bf8ea454b72753ec907db774fba2e902e0

Observation e1b73488-3522-47eb-b05f-08f9e55a9403 · outbound

This paper cites Im- proving prosody modelling with cross-utterance BERT embed- dings for end-to-end speech synthesis,.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models Im- proving prosody modelling with cross-utterance BERT embed- dings for end-to-end speech synthesis,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:03:18.371358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T20:03:16.410019Z digest=sha256:cbc926c0275360fc4594015ae03b09eeca1aa09ab8c9dd5a24e9ca39dc9fc936

Observation cc6a775b-2392-47c7-b5e8-6c6e12926128 · outbound

This paper cites WavLM: Large-scale self-supervised pre-training for full stack speech processing,.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models WavLM: Large-scale self-supervised pre-training for full stack speech processing,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:03:18.315366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T20:03:17.032913Z digest=sha256:3196c25e02aa2b2986dc684b7eda5881a23c9f7153f16995f40e0cce0cd674a8

Observation 57503311-5675-44f1-bb7e-bd5a1c06b197 · outbound

This paper cites PE- Wav2vec: A prosody-enhanced speech model for self-supervised prosody learning in TTS,.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models PE- Wav2vec: A prosody-enhanced speech model for self-supervised prosody learning in TTS,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:03:18.353407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T20:03:16.565219Z digest=sha256:b69dff67a96e6cf2f811bf836ef4271e3e49b232f974b713c2e6251d352777db

Observation f97fc26c-1ad4-4f1c-8e71-e2eda8f71880 · outbound

This paper cites PnG BERT: Aug- mented BERT on phonemes and graphemes for neural TTS,.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models PnG BERT: Aug- mented BERT on phonemes and graphemes for neural TTS,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:03:18.344923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T20:03:16.622279Z digest=sha256:59abbeaebe516e31d388e993489b12acc9a78d3914a7c524b2ce5f7cf0c39dcf

Observation b8b7e871-807a-4ee9-9e2b-a67c10d1f833 · outbound

This paper cites Phoneme-Level BERT for Enhanced Prosody of Text-to-Speech with Grapheme Predictions.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models Phoneme-Level BERT for Enhanced Prosody of Text-to-Speech with Grapheme Predictions

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:03:18.063898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T20:03:16.688013Z digest=sha256:2f0f84e31b3d9d71004edd4b90588ae21dd2f922b8bbe48882926f5f5fc41e3e

Observation acfc20ea-a1a6-4a01-8151-3d210e10a9e6 · outbound

This paper cites Detection of prosodic bound- aries in speech using wav2vec 2.0,.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models Detection of prosodic bound- aries in speech using wav2vec 2.0,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:03:18.336829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T20:03:16.765347Z digest=sha256:d3824952131cda25427517be9108db74168e651c756b097eafeaa00b472cbda1

Observation 0f49b92c-8d03-47be-89cf-189aa850975d · outbound

This paper cites Corpus of Spontaneous Japanese: Its design and evaluation,.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models Corpus of Spontaneous Japanese: Its design and evaluation,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:03:18.328914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T20:03:16.852515Z digest=sha256:9b42922b90b1f83c26c16c4ad544dcefdaea10507e1504cf60cd33967bbf7059

Observation 060b2b47-c026-4190-bffb-129380289285 · outbound

This paper cites HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T20:03:16.935390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:03:16.935390Z digest=sha256:b3652d809c33ccd9610ed61326f9fa41479afb2e56671dd92297cf64898a87a9

Observation a4060926-a9e1-4dc4-9ad1-c32572eba2a7 · outbound

This paper cites V AE-based phoneme alignment using gradient an- nealing and SSL acoustic features,.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models V AE-based phoneme alignment using gradient an- nealing and SSL acoustic features,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:03:18.261266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T20:03:17.608199Z digest=sha256:acbf848f2155c17a8f504dd4b047d432660efa31bc816d66202c745e644b8671

Observation c7b819f2-d590-4756-a8f6-d33de3c60e10 · outbound

This paper cites Self-supervised speech representation learning: A review,.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models Self-supervised speech representation learning: A review,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:03:18.307988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T20:03:17.101847Z digest=sha256:756fc06f34e27744b3f6b94dced6e599d168a9aec071f674db5dad53efd5f383

Observation abc9ed54-b6e4-446c-8c81-fef129ee3414 · outbound

This paper cites StyleTTS 2: Towards human-level text-to-speech through style diffusion and adversarial training with large speech language models,.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models StyleTTS 2: Towards human-level text-to-speech through style diffusion and adversarial training with large speech language models,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:03:18.300322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T20:03:17.116416Z digest=sha256:604d9bc9bc34c8f476e0254602b97aa8e3a3de13c0d477f1cdbb4e28ed4dda09

Observation 19f23418-8170-4b75-a43d-390fc79a55a4 · outbound

This paper cites Emobox: Multilingual multi-corpus speech emotion recognition toolkit and benchmark,.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models Emobox: Multilingual multi-corpus speech emotion recognition toolkit and benchmark,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:03:18.292565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T20:03:17.161090Z digest=sha256:4c1b492fdd57e1eac1d5ffcdd98c0b4a56e16c489d7e2262b96856c6dacf45a0

Observation 2e81d025-b644-4923-bc15-00202f2e2917 · outbound

This paper cites Non-intrusive speech intelligibility pre- diction for hearing-impaired users using intermediate ASR fea- tures and human memory models,.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models Non-intrusive speech intelligibility pre- diction for hearing-impaired users using intermediate ASR fea- tures and human memory models,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:03:18.284318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T20:03:17.249134Z digest=sha256:75688b0cae6fedec6d5062b8c4a9101a0a8bdf2bb4a8bdd754c39471e04f36cb

Observation 0b479e7d-188f-4172-bba5-2fae76a05cfe · outbound

This paper cites Montreal forced aligner: Trainable text-speech align- ment using kaldi.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models Montreal forced aligner: Trainable text-speech align- ment using kaldi

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:03:18.276771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T20:03:17.361914Z digest=sha256:e08286ec847a4f18cf7c822c7b817e77da7568a7dc0ca167e580628732f0a795

Observation 36e6563f-7f5b-4834-8ede-aff580a51a1c · outbound

This paper cites One TTS alignment to rule them all,.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models One TTS alignment to rule them all,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:03:18.269135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T20:03:17.465887Z digest=sha256:c7eb94383977c0df68a693e1f5ec65dcdc99f95d12677f1a7064cbe31f1da101

Observation 6fe82536-100d-48da-b81a-b35f370b1b9a · outbound

This paper cites X-JToBI: an extended J-ToBI for spontaneous speech,.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models X-JToBI: an extended J-ToBI for spontaneous speech,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:03:18.207838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T20:03:18.020758Z digest=sha256:f83cd8aab171eecf4ab20ca34e4946b2f7d67eb4d8c13947184b17d41b7203ba

Observation 1e7d9279-efec-43b0-a8ed-86e58662a4a5 · outbound

This paper cites Stress in Thai,.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models Stress in Thai,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:03:18.253843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T20:03:17.750478Z digest=sha256:50079963959d9d6a64e6bec62da2f71aeb3a927c2b30a25d2444860060aa565e

Observation 30199321-7b5d-41a0-9bdd-cedac8f77223 · outbound

This paper cites ee shumi ga ongaku nandesu keredomo.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models ee shumi ga ongaku nandesu keredomo

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:03:18.538756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T20:03:14.805676Z digest=sha256:7e8b37140954359b358d9d269875cdbf105cec7c3f8a3fca562483a9b499ca5b

Observation 4a37a47a-7458-4da5-87f0-7374f43f504d · outbound

This paper cites an unresolved cited work.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:03:18.246286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T20:03:17.846726Z digest=sha256:3fe12ce75303bc4c1faf4bf80e2b8cb9af5643fbf359ed838a5ea7405a4eea97

Observation ee997eda-f97a-411d-add7-3492446b1997 · outbound

This paper cites an unresolved cited work.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:03:18.238744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T20:03:17.852370Z digest=sha256:f78f67189ffd8d218d5893ce1939cff00e3a2e9af8dbd206adb129b1c8e09d2f

Observation 5e13959f-06c3-468c-90a2-24e9a228116f · outbound

This paper cites an unresolved cited work.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:03:18.231361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T20:03:17.978850Z digest=sha256:7f015a0749544691cb3fc51d2b196320f58953c0040d19bee3b03e78c3d19c06

Observation 1043fdb1-3323-4995-85ab-f70300f9fcd1 · outbound

This paper cites Adam: A method for stochastic opti- mization,.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models Adam: A method for stochastic opti- mization,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:03:18.223716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T20:03:18.015268Z digest=sha256:59fce67fd553cba6c581d53e566349f90799c035e3b94b8d10384f3abb8d279e

Observation 03c44fd4-d47f-4fac-88aa-b3e519a329b4 · outbound

This paper cites Prosodic features con- trol by symbols as input of sequence-to-sequence acoustic mod- eling for neural TTS,.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models Prosodic features con- trol by symbols as input of sequence-to-sequence acoustic mod- eling for neural TTS,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:03:18.216152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T20:03:18.018239Z digest=sha256:3240793e124c9deac660a27ca6629d6e2b7fb393640a044814bb3ba2725fd010

Observation ed323aca-ccbe-48df-8dd5-3fdc755c5c3f · outbound

This paper cites ESPnet: End-to-end speech processing toolkit,.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models ESPnet: End-to-end speech processing toolkit,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:03:18.199524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T20:03:18.023597Z digest=sha256:d99a597e8beed432c6932e4b648e21262e3f9cebef609b6ad347e4ba29050891

Observation 7a77b45e-0bef-43fa-a63d-dde25d8ad60f · outbound

This paper cites an unresolved cited work.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:03:18.191205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T20:03:18.026215Z digest=sha256:6cfbc4c5d4f04a7c6fc17df645a196268f67e5698e9dc7a7f1921c2cc914f4ee

Observation 4a17447c-9180-42b4-a351-1ad51b06dae9 · outbound

This paper cites an unresolved cited work.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:03:18.092629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T20:03:18.028687Z digest=sha256:62b9feb55f0926a06d97c3f1a3d5f59260df380d86aefa4991b04ea154ed77e9

Observation 566d1044-1005-4142-a880-81b694d682dd · outbound

This paper cites WORLD: A vocoder- based high-quality speech synthesis system for real-time appli- cations,.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models WORLD: A vocoder- based high-quality speech synthesis system for real-time appli- cations,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:03:18.084938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T20:03:18.031484Z digest=sha256:e31c284281dad67a06359c45a35348264bc0405c4f6826994994bc1bf8824f4a

Pith citing papers

Observation a8e9006e-c63d-4fb7-8a0a-0c41862ebe45 · inbound

Prosody Labeling with Phoneme-BERT and Speech Foundation Models cites this paper.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models Prosody Labeling with Phoneme-BERT and Speech Foundation Models

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T20:03:18.076461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T20:03:14.249592Z digest=sha256:cee0d5047022b1e9d06638395652a86d9c480cbb4fb9ae73800b142fd1ea4922