Pith. sign in

Paper Citation Record · LEDGER

Prosody Labeling with Phoneme-BERT and Speech Foundation Models

As of 16 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 1 inbound Pith citation observation for arXiv:2507.03912.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.03912 v1

Coverage vector

measured 53 of 53 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:03:18.031484Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:03:14.249592Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T20:03:18.073334Z

Reference resolution

53 of 53 outbound references displayed

  • verified exact1
  • verified fuzzy43
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a8e9006e-c63d-4fb7-8a0a-0c41862ebe45 · outbound

This paper cites Prosody Labeling with Phoneme-BERT and Speech Foundation Models.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models Prosody Labeling with Phoneme-BERT and Speech Foundation Models

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T20:03:18.076461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T20:03:14.249592Z digest=sha256:21136e38ae0b96e7f29867254e3e8b9070155dcc8adea1d09311232c94d2aa0f

Observation 7a37b1f9-63a0-4b0b-b6d1-b84466d39149 · outbound

This paper cites [8] have also conducted automatic prosody an- notation for constructing a speech synthesis database, similar to the objective of our study.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models [8] have also conducted automatic prosody an- notation for constructing a speech synthesis database, similar to the objective of our study

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:03:18.570098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T20:03:14.334997Z digest=sha256:fec3616b1ec97886e8ce64366aa8437dd4c352134317a329521cca776c0eedba

Observation 486d9ad2-4af9-4aea-bb4d-14ec30cd5b01 · outbound

This paper cites ashita wa hare desuka.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models ashita wa hare desuka

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:03:18.562280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T20:03:14.448986Z digest=sha256:2797344442580d4a09c716b50f72b1a70fedc634887a821b4e5455eeea0a61e8

Observation 25f7c816-b74d-4ccd-bd15-7b8514a2d4d3 · outbound

This paper cites Figure 3 shows the outline of the proposed model.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models Figure 3 shows the outline of the proposed model

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:03:18.554087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T20:03:14.576338Z digest=sha256:04f67b4d8294a518753a761d38c4b439748ee5a1ea6719a2df5458a0ab71c523

Observation 7f5ba96d-e444-4a68-8274-f3d160d541e9 · outbound

This paper cites *”: Other symbol that was neither a high-low transition nor a phrase boundary. • “[.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models *”: Other symbol that was neither a high-low transition nor a phrase boundary. • “[

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:03:18.546322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T20:03:14.718140Z digest=sha256:09f76a12a6da128b79c27cd131fb718649a7387bfe6397f5e76794954f6fcc5d

Observation d85e9b1d-ba7e-45fe-98fd-6753788a252e · outbound

This paper cites an unresolved cited work.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:03:18.530571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T20:03:14.941364Z digest=sha256:4f4183eeed28b097692d2ac14d2d9016a4401ea3a43b19c8bc530939cbd273be

Observation ad3aa484-253a-4abe-8ef5-5b8e17d39938 · outbound

This paper cites Automatic prosody annotation with pre-trained text- speech model,.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models Automatic prosody annotation with pre-trained text- speech model,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:03:18.474594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T20:03:15.460095Z digest=sha256:4a764dfe4cd8641f6fc26f0888244ea00f34f994b5c94a1c60d5fbdf6952b9f6

Observation 24699a4f-1edb-44b9-aef4-f4cd2e9745ab · outbound

This paper cites Mora-level prosody prediction for text-to-speech us- ing japanese BERT without accentual labels,.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models Mora-level prosody prediction for text-to-speech us- ing japanese BERT without accentual labels,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:03:18.522972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T20:03:14.995512Z digest=sha256:5814eef22c34a30e77f123d0791c5fa87af6291f1b415cb4e351dca8648036fa

Observation 91978ff0-5f1a-43d0-8fac-dce5a42e6dec · outbound

This paper cites Semi-supervised prosody model- ing using deep Gaussian process latent variable model,.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models Semi-supervised prosody model- ing using deep Gaussian process latent variable model,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:03:18.515166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T20:03:15.059899Z digest=sha256:e96fd1ac7a94896125ac9f374c0730a245a8ea2fd19aae54d0822f79b2adeb56

Observation 2800bd16-57a5-4d68-b532-08a393ce229a · outbound

This paper cites Ac- cent modeling of low-resourced dialect in pitch accent language using variational autoencoder,.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models Ac- cent modeling of low-resourced dialect in pitch accent language using variational autoencoder,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:03:18.506457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T20:03:15.121236Z digest=sha256:fe441d4bce54f093c41cbaab4b0d1ca5f9becdf23fc5d08e0601e69aa5983e72

Observation 17c90688-3bf3-4167-b89c-4e05771cc5d0 · outbound

This paper cites Cross-dialect text- to-speech in pitch-accent language incorporating multi-dialect phoneme-level BERT,.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models Cross-dialect text- to-speech in pitch-accent language incorporating multi-dialect phoneme-level BERT,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:03:18.498411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T20:03:15.224190Z digest=sha256:b0bbc1e126331db1a83580780aeef2b6bc424a66dd42053a2f3a779f0f340864

Observation 7752be24-91c1-4332-92de-fe68656049ef · outbound

This paper cites Multi-modal automatic prosody anno- tation with contrastive pretraining of speech-silence and word- punctuation,.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models Multi-modal automatic prosody anno- tation with contrastive pretraining of speech-silence and word- punctuation,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:03:18.490401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T20:03:15.277753Z digest=sha256:f0b952ff20fcbbabd032556d3a29c0300dd468da3fb4ab2a25cfb04a37c1ee6b

Observation 9f4f8240-b7e3-4f27-b812-aa1f6842bcc1 · outbound

This paper cites Language-independent prosody-enhanced speech representations for multilingual speech synthesis,.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models Language-independent prosody-enhanced speech representations for multilingual speech synthesis,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:03:18.482741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T20:03:15.376794Z digest=sha256:6596b1758728cca24138dc3bc001a77a35619ad8cd2748cf2b659103bc526765

Observation 8377d84f-481f-43c6-b405-5126743c9a7e · outbound

This paper cites BERT: Pre- training of deep bidirectional transformers for language under- standing,.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models BERT: Pre- training of deep bidirectional transformers for language under- standing,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:03:18.419024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T20:03:15.971177Z digest=sha256:a0d189ea950fe6323f344a76607bd1036a316f0d5a0178d9e447ec86379285a1

Observation 501b9959-6d09-43de-bc6f-d6a5715126f7 · outbound

This paper cites Audio- conditioned phonemic and prosodic annotation for building text- to-speech models from unlabeled speech data,.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models Audio- conditioned phonemic and prosodic annotation for building text- to-speech models from unlabeled speech data,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:03:18.465831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T20:03:15.540553Z digest=sha256:571a61858ca02e752bcc2a1a099302382c1bfc18aff0984c5d45479966e49dd1

Observation 02fafbbd-40ca-4ff9-8c6b-9d4d91b2a8c8 · outbound

This paper cites Low-resourced phonetic and prosodic feature estimation with self-supervised-learning-based acoustic modeling,.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models Low-resourced phonetic and prosodic feature estimation with self-supervised-learning-based acoustic modeling,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:03:18.457670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T20:03:15.611761Z digest=sha256:ecde0a2d386009deb2713bd24e527068e6ec16cfd0d61010ddef5564c220af25

Observation 3a4335c3-5762-4f48-90b7-61b6610ca6c1 · outbound

This paper cites Cross-lingual speech- based ToBI label generation using bidirectional LSTM,.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models Cross-lingual speech- based ToBI label generation using bidirectional LSTM,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:03:18.449384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T20:03:15.676428Z digest=sha256:948fc38cf6d2c9ac6bfe2c3aa75b54b991b1f6bb80d797be9082821b56c93063

Observation 994affb5-0eb8-48ed-8edd-474a1541b990 · outbound

This paper cites A unified accent esti- mation method based on multi-task learning for Japanese text-to- speech,.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models A unified accent esti- mation method based on multi-task learning for Japanese text-to- speech,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:03:18.440586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T20:03:15.742927Z digest=sha256:8734ca2e75ff79e0ec7f977f27277e431253a15677324501b07ad48abe0783a6

Observation 73a9539d-6bdd-4c62-bbfc-1831c6547496 · outbound

This paper cites Wav2ToBI: a new approach to automatic ToBI transcription,.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models Wav2ToBI: a new approach to automatic ToBI transcription,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:03:18.432635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T20:03:15.809588Z digest=sha256:ee4af513ef1242d53a5041f6bd8ec76260bf45420f82cab5c249bf1a2ea73866

Observation 7d985065-12a4-469d-ae23-b642b0469279 · outbound

This paper cites Conformer: Convolution-augmented transformer for speech recognition,.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models Conformer: Convolution-augmented transformer for speech recognition,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T20:03:15.886329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:03:15.886329Z digest=sha256:fa3fbe3992146cfbc67ed671927964afb48fb44f3f2ef29c11aca7eb101a502b

Observation 75ce9da2-3cf3-4a12-8a6a-6f6eaa5606a5 · outbound

This paper cites Investigation of enhanced Tacotron text-to-speech synthesis systems with self- attention for pitch accent language,.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models Investigation of enhanced Tacotron text-to-speech synthesis systems with self- attention for pitch accent language,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:03:18.361910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T20:03:16.479507Z digest=sha256:39b5a86fcb8daf26a28b29e49c10b770dc929871ae7bd12063e2203ef2596865

Observation 5b927204-f5ad-474e-aaba-5cad605fe1e3 · outbound

This paper cites Robust speech recognition via large-scale weak su- pervision,.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models Robust speech recognition via large-scale weak su- pervision,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:03:18.411280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T20:03:16.049594Z digest=sha256:400fb3853e028092354c86d66b5547ffd0dcea140968701a7507d3d99fc40f2b

Observation 01d1492d-4e95-403e-88d1-d12a7e03f4e1 · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech representa- tions,.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models wav2vec 2.0: A framework for self-supervised learning of speech representa- tions,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:03:18.403453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T20:03:16.124677Z digest=sha256:cd61c5127c1afe8608928e7a242ac19aec589a5d5dbddf3469c7e19f57614f9b

Observation e28cb59e-5e18-41cf-9318-cede3a3f7cf9 · outbound

This paper cites Pre-trained text embeddings for enhanced text-to- speech synthesis,.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models Pre-trained text embeddings for enhanced text-to- speech synthesis,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:03:18.395326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T20:03:16.187592Z digest=sha256:dd9d6c174ebf5fdbeaebcdb2a5dd1250e1c5a8492fc9133771735bbd4c8a2e98

Observation 91b6ee43-c70f-4403-9818-071a4513fae6 · outbound

This paper cites Improving prosody with linguistic and BERT derived features in multi-speaker based Mandarin Chinese neural TTS,.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models Improving prosody with linguistic and BERT derived features in multi-speaker based Mandarin Chinese neural TTS,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:03:18.387703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T20:03:16.256372Z digest=sha256:889a1646b467d140b445947e15adeecb233b9668a380e6197922fbb75adee674

Observation 0596f701-8685-4c3f-93ef-90ef31690aa2 · outbound

This paper cites Improving the prosody of RNN-based English text-to-speech synthesis by incorporating a BERT model,.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models Improving the prosody of RNN-based English text-to-speech synthesis by incorporating a BERT model,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:03:18.379448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T20:03:16.342536Z digest=sha256:ba59e0568f4ad43e8252c54b293480373d235ef79b781e77ec604a72c9ecfac0

Observation e1b73488-3522-47eb-b05f-08f9e55a9403 · outbound

This paper cites Im- proving prosody modelling with cross-utterance BERT embed- dings for end-to-end speech synthesis,.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models Im- proving prosody modelling with cross-utterance BERT embed- dings for end-to-end speech synthesis,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:03:18.371358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T20:03:16.410019Z digest=sha256:9a294b3c8763680aa9e91f7e8991eb1b77f0ea9b8a60ed52d0df1fe2169eb0e7

Observation cc6a775b-2392-47c7-b5e8-6c6e12926128 · outbound

This paper cites WavLM: Large-scale self-supervised pre-training for full stack speech processing,.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models WavLM: Large-scale self-supervised pre-training for full stack speech processing,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:03:18.315366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T20:03:17.032913Z digest=sha256:d5bf844e49641e2e101626d6909c0634aafd9f7105ed33620e9d1e5ff9d93c2a

Observation 57503311-5675-44f1-bb7e-bd5a1c06b197 · outbound

This paper cites PE- Wav2vec: A prosody-enhanced speech model for self-supervised prosody learning in TTS,.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models PE- Wav2vec: A prosody-enhanced speech model for self-supervised prosody learning in TTS,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:03:18.353407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T20:03:16.565219Z digest=sha256:01b9ee3f6c63175292ac2fde5e00d85ba87225aa51cf775259682b8512e0218a

Observation f97fc26c-1ad4-4f1c-8e71-e2eda8f71880 · outbound

This paper cites PnG BERT: Aug- mented BERT on phonemes and graphemes for neural TTS,.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models PnG BERT: Aug- mented BERT on phonemes and graphemes for neural TTS,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:03:18.344923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T20:03:16.622279Z digest=sha256:56978749bba35f3adc933ccf9e4cd6684f60384b5ee75e1cbffa4180d22d7c8a

Observation b8b7e871-807a-4ee9-9e2b-a67c10d1f833 · outbound

This paper cites Phoneme-Level BERT for Enhanced Prosody of Text-to-Speech with Grapheme Predictions.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models Phoneme-Level BERT for Enhanced Prosody of Text-to-Speech with Grapheme Predictions

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:03:18.063898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T20:03:16.688013Z digest=sha256:eec27d8c6ced99058885df929a9bedafb4f53b7f58ada2aa2414f895d2beda8d

Observation acfc20ea-a1a6-4a01-8151-3d210e10a9e6 · outbound

This paper cites Detection of prosodic bound- aries in speech using wav2vec 2.0,.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models Detection of prosodic bound- aries in speech using wav2vec 2.0,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:03:18.336829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T20:03:16.765347Z digest=sha256:4b245ccc926b535c760eb93a4d604ca627f443f07259b7d7e6e0af9b5fc61a8c

Observation 0f49b92c-8d03-47be-89cf-189aa850975d · outbound

This paper cites Corpus of Spontaneous Japanese: Its design and evaluation,.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models Corpus of Spontaneous Japanese: Its design and evaluation,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:03:18.328914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T20:03:16.852515Z digest=sha256:e2dbd85dbdd961a3de6cb8689341cdf29c97df45c6addb4d389be07ea0b7531c

Observation 060b2b47-c026-4190-bffb-129380289285 · outbound

This paper cites HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T20:03:16.935390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:03:16.935390Z digest=sha256:92703017ecf84ebf1dc341d9d63da42671b3a08d46b662812ce201ec1cfce345

Observation a4060926-a9e1-4dc4-9ad1-c32572eba2a7 · outbound

This paper cites V AE-based phoneme alignment using gradient an- nealing and SSL acoustic features,.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models V AE-based phoneme alignment using gradient an- nealing and SSL acoustic features,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:03:18.261266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T20:03:17.608199Z digest=sha256:39615e3a4a3900b3c4c42234d2f0e2a7f2809877c387cbe6adf89ca20f47370e

Observation c7b819f2-d590-4756-a8f6-d33de3c60e10 · outbound

This paper cites Self-supervised speech representation learning: A review,.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models Self-supervised speech representation learning: A review,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:03:18.307988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T20:03:17.101847Z digest=sha256:b218a88df2c761ce2b02c5f971208ba142eb50ed43ce4c72640214eb2ebfcc24

Observation abc9ed54-b6e4-446c-8c81-fef129ee3414 · outbound

This paper cites StyleTTS 2: Towards human-level text-to-speech through style diffusion and adversarial training with large speech language models,.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models StyleTTS 2: Towards human-level text-to-speech through style diffusion and adversarial training with large speech language models,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:03:18.300322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T20:03:17.116416Z digest=sha256:21b996c20b27154ba9fdd780747b646bb2d07ded50d8e65379e90598084e0353

Observation 19f23418-8170-4b75-a43d-390fc79a55a4 · outbound

This paper cites Emobox: Multilingual multi-corpus speech emotion recognition toolkit and benchmark,.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models Emobox: Multilingual multi-corpus speech emotion recognition toolkit and benchmark,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:03:18.292565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T20:03:17.161090Z digest=sha256:be49674ed222b5d88e571dac558934f6c908f482c6d90ba01caf566cf6b037fe

Observation 2e81d025-b644-4923-bc15-00202f2e2917 · outbound

This paper cites Non-intrusive speech intelligibility pre- diction for hearing-impaired users using intermediate ASR fea- tures and human memory models,.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models Non-intrusive speech intelligibility pre- diction for hearing-impaired users using intermediate ASR fea- tures and human memory models,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:03:18.284318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T20:03:17.249134Z digest=sha256:0c687077904799390e00f9cbc44f52fad8e797f4b0caad8e425f30a600e9ee3d

Observation 0b479e7d-188f-4172-bba5-2fae76a05cfe · outbound

This paper cites Montreal forced aligner: Trainable text-speech align- ment using kaldi.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models Montreal forced aligner: Trainable text-speech align- ment using kaldi

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:03:18.276771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T20:03:17.361914Z digest=sha256:7f01bb1f4a703d5a04cecf94fa2a0bbcdcd614dafd3e07ceddaedd5b3643497b

Observation 36e6563f-7f5b-4834-8ede-aff580a51a1c · outbound

This paper cites One TTS alignment to rule them all,.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models One TTS alignment to rule them all,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:03:18.269135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T20:03:17.465887Z digest=sha256:bac71cc9d88bea34cdc30162180220ee7e0c17f425389dd390bdd2b293602ab2

Observation 6fe82536-100d-48da-b81a-b35f370b1b9a · outbound

This paper cites X-JToBI: an extended J-ToBI for spontaneous speech,.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models X-JToBI: an extended J-ToBI for spontaneous speech,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:03:18.207838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T20:03:18.020758Z digest=sha256:206a44825878f2af7d62651806cad8ad7c61c502613bb0a6c9312d9efb17582c

Observation 1e7d9279-efec-43b0-a8ed-86e58662a4a5 · outbound

This paper cites Stress in Thai,.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models Stress in Thai,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:03:18.253843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T20:03:17.750478Z digest=sha256:987be79e0beafa8134f00048ac2a4a4b107273290aa67bdcd4ac14072a934108

Observation 30199321-7b5d-41a0-9bdd-cedac8f77223 · outbound

This paper cites ee shumi ga ongaku nandesu keredomo.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models ee shumi ga ongaku nandesu keredomo

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:03:18.538756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T20:03:14.805676Z digest=sha256:dee1b6c5a1afb6fca67af3ca9dfddbbf49d44565d33ae1bd6b9d3ee3e335b7ee

Observation 4a37a47a-7458-4da5-87f0-7374f43f504d · outbound

This paper cites an unresolved cited work.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:03:18.246286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T20:03:17.846726Z digest=sha256:abf5037723d6ecd1bacaeb37764c75d72e8ce010eb7ddc8d7cf12273a1db095c

Observation ee997eda-f97a-411d-add7-3492446b1997 · outbound

This paper cites an unresolved cited work.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:03:18.238744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T20:03:17.852370Z digest=sha256:940d538544aeede7d554f4df4bcf23684d4bc5a0d627cab17b926217cc879a1c

Observation 5e13959f-06c3-468c-90a2-24e9a228116f · outbound

This paper cites an unresolved cited work.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:03:18.231361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T20:03:17.978850Z digest=sha256:77a309c3971adf0df25441ec2508bec3f478e205370a5aa4d780cdc28914c0fe

Observation 1043fdb1-3323-4995-85ab-f70300f9fcd1 · outbound

This paper cites Adam: A method for stochastic opti- mization,.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models Adam: A method for stochastic opti- mization,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:03:18.223716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T20:03:18.015268Z digest=sha256:2d77a58d827b92f9f511e41bfef859bed83d36c2138a70480bbc236d5e3953fb

Observation 03c44fd4-d47f-4fac-88aa-b3e519a329b4 · outbound

This paper cites Prosodic features con- trol by symbols as input of sequence-to-sequence acoustic mod- eling for neural TTS,.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models Prosodic features con- trol by symbols as input of sequence-to-sequence acoustic mod- eling for neural TTS,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:03:18.216152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T20:03:18.018239Z digest=sha256:3239cc6687ed8172a1850d4e75d0069f8c8f7d0e6c0e99a7935267975c580b6e

Observation ed323aca-ccbe-48df-8dd5-3fdc755c5c3f · outbound

This paper cites ESPnet: End-to-end speech processing toolkit,.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models ESPnet: End-to-end speech processing toolkit,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:03:18.199524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T20:03:18.023597Z digest=sha256:74d0f7f081ff2e1ee2a444c369efdccdac00884ca9afeb2efa6dea3072eac9e1

Observation 7a77b45e-0bef-43fa-a63d-dde25d8ad60f · outbound

This paper cites an unresolved cited work.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:03:18.191205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T20:03:18.026215Z digest=sha256:6a31313a6e65305292181e45aad8db345a759d5a5fd3b81e97224ee629d708ef

Observation 4a17447c-9180-42b4-a351-1ad51b06dae9 · outbound

This paper cites an unresolved cited work.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:03:18.092629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T20:03:18.028687Z digest=sha256:d72e654dc0b29c6edde87d432c1d9912a532d5fc93f1c6cb5da6abb97353a7a0

Observation 566d1044-1005-4142-a880-81b694d682dd · outbound

This paper cites WORLD: A vocoder- based high-quality speech synthesis system for real-time appli- cations,.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models WORLD: A vocoder- based high-quality speech synthesis system for real-time appli- cations,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:03:18.084938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T20:03:18.031484Z digest=sha256:8f6f52219a7b755949121b690c3307d4c360e3538948118d6198b3a0ecc960ff

Pith citing papers

Observation a8e9006e-c63d-4fb7-8a0a-0c41862ebe45 · inbound

Prosody Labeling with Phoneme-BERT and Speech Foundation Models cites this paper.

Prosody Labeling with Phoneme-BERT and Speech Foundation Models Prosody Labeling with Phoneme-BERT and Speech Foundation Models

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T20:03:18.076461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T20:03:14.249592Z digest=sha256:21136e38ae0b96e7f29867254e3e8b9070155dcc8adea1d09311232c94d2aa0f