Pith. sign in

Paper Citation Record · LEDGER

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset

As of 8 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 1 inbound Pith citation observation for arXiv:2505.20341.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.20341 v1

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:28:48.254853Z

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:28:44.020833Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T14:28:48.777510Z

Reference resolution

37 of 37 outbound references displayed

  • verified exact2
  • verified fuzzy23
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6f217865-cea5-42ed-87f4-44b0d1568529 · outbound

This paper cites Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:28:48.852397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:28:44.020833Z digest=sha256:11d28a7acd7dc4e7abe96826604d0391bf48310e353b4b3621a33d0c76a9098c

Observation 8eee5942-3a49-489d-a7e9-4e21042c05b8 · outbound

This paper cites sadness” as an ex- ample): “Now, I will give you a sentence. Please modify only one or two words to change the emotion to sadness. Please output only one modified sentence.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset sadness” as an ex- ample): “Now, I will give you a sentence. Please modify only one or two words to change the emotion to sadness. Please output only one modified sentence

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:28:54.751761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:28:44.161073Z digest=sha256:76485869217c0fc3293c3f918fe31b89e752182f7b3b19fb087b81e71a86ccc5

Observation 43ed3c79-9812-44be-b80f-5831f9cd1cd9 · outbound

This paper cites The first two modules require pre- training, and then the third module is trained end-to-end.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset The first two modules require pre- training, and then the third module is trained end-to-end

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:28:54.248425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:28:44.341967Z digest=sha256:ead8a20b208d9488dc95cc35cdfeadce672d7070420c8ee5e02c868132b77bd0

Observation e2ce15a0-d0e7-44e1-ac88-97f32f3890ae · outbound

This paper cites Experimental Setup We evaluate EmoCorrector on the ECD-TSE.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset Experimental Setup We evaluate EmoCorrector on the ECD-TSE

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:28:54.011124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:28:44.448642Z digest=sha256:1735f0c904ed116c91a9749524783de0eaccbb2fd0f2fa27e5c768c8d01e008c

Observation 135f474c-6533-4f97-aec9-9636001b2759 · outbound

This paper cites Experimental results demonstrate that the proposed framework improves emotional consistency in the edited speech.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset Experimental results demonstrate that the proposed framework improves emotional consistency in the edited speech

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:28:53.764375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:28:44.578272Z digest=sha256:9325277e4f2f8c85b2cd8861ca43b70dd8e052fd4f12fb1b98b22b0e412cc571

Observation 7ada997d-43c7-48da-b9e4-be65e03beab1 · outbound

This paper cites 62206136), the General Program (No.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset 62206136), the General Program (No

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:28:53.558362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:28:44.687416Z digest=sha256:661896a951753511d209939c845fdb4d955afb90dc149d26953428679e8b1f74

Observation 2b4da371-5ea5-4768-b7db-d6dab5a649ce · outbound

This paper cites Emotional voice con- version: Theory, databases and esd,.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset Emotional voice con- version: Theory, databases and esd,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:28:45.477922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:28:45.477922Z digest=sha256:9743127ae044a4ea585e39483dbec8b3bdae3b5370c5150068f75d5abfacb939

Observation 5bf0d3c4-5ff9-412c-b02f-b716b75d951d · outbound

This paper cites Fluentspeech: Stutter-oriented automatic speech editing with context-aware diffusion models,.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset Fluentspeech: Stutter-oriented automatic speech editing with context-aware diffusion models,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:28:53.302583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:28:44.859707Z digest=sha256:2eb3bd7725b7b491d0a819acef8e29fd15e13183d12bafe47d6042ae67e5298a

Observation a957226c-a261-4f04-90b0-3968a953c34a · outbound

This paper cites A3t: Alignment-aware acoustic and text pretraining for speech synthe- sis and editing,.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset A3t: Alignment-aware acoustic and text pretraining for speech synthe- sis and editing,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:28:53.069985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:28:44.980567Z digest=sha256:761fdb318c43fc0c3a463a2d92afd07ede93e0b1e352312022f3ec342e6deb04

Observation 0b8cb2be-9168-4c7a-87d8-d7fd5af52ee8 · outbound

This paper cites FluentEditor: Text-based Speech Editing by Considering Acoustic and Prosody Consistency.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset FluentEditor: Text-based Speech Editing by Considering Acoustic and Prosody Consistency

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:28:45.077048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:28:45.077048Z digest=sha256:ea9f436afdd4e4c401e77c3e5064f4db4b83b74350e194308fdb86bd8827a9a0

Observation 0a4f1427-3dea-4a21-bdb3-d8564011123e · outbound

This paper cites Speechx: Neural codec language model as a versatile speech transformer,.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset Speechx: Neural codec language model as a versatile speech transformer,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:28:52.829602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:28:45.176163Z digest=sha256:be2efff112bde9cfdd9dba6e42b99ba96d1e6c4b5f2e1671f616ab3332838e29

Observation 96ad5d9e-5159-48cf-a12f-b918656bc5c2 · outbound

This paper cites V oice- craft: Zero-shot speech editing and text-to-speech in the wild,.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset V oice- craft: Zero-shot speech editing and text-to-speech in the wild,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:28:52.574513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:28:45.291603Z digest=sha256:dfe3b42370fc9c1633bb33e8b7a12d3ede5bb2e72718e4ee9f291d317f94f03a

Observation 673ba17a-ef8f-4be3-af29-a216d3511379 · outbound

This paper cites Tackling modality heterogeneity with multi-view calibration net- work for multimodal sentiment detection,.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset Tackling modality heterogeneity with multi-view calibration net- work for multimodal sentiment detection,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:28:52.338641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:28:45.392572Z digest=sha256:c4970a93bf54210a01164dadb90c85adf3795f3c79b6dc5c059d6e599ade0d7d

Observation 31dd990b-f13a-4ad6-8b46-7f45de78aaaa · outbound

This paper cites MSP-Podcast SER Challenge 2024: L'antenne du Ventoux Multimodal Self-Supervised Learning for Speech Emotion Recognition.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset MSP-Podcast SER Challenge 2024: L'antenne du Ventoux Multimodal Self-Supervised Learning for Speech Emotion Recognition

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:28:46.273638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:28:46.273638Z digest=sha256:cd2be686a7e08254ffe43939855cd24ae082a62579c01001d4ce3484c8bb00bb

Observation 4c16ec55-e528-412b-b2bc-b0758292b0cf · outbound

This paper cites Decoupling speaker-independent emotions for voice conversion via source-filter networks,.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset Decoupling speaker-independent emotions for voice conversion via source-filter networks,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:28:52.135108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:28:45.573461Z digest=sha256:325b5b9f3ae70dc6438a7eeacff737517f1241b29c9394fcf051c68ece90be50

Observation 9a8e4392-99f9-488b-8d6f-500b10c82f14 · outbound

This paper cites Retrieval- augmented generation for knowledge-intensive nlp tasks,.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset Retrieval- augmented generation for knowledge-intensive nlp tasks,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:28:51.903067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:28:45.691045Z digest=sha256:e15e98402320f61d28ac3ddaf17fe778011f1c273747869730688aa772685534

Observation e2d3ddee-1e00-47c9-a852-ce3304452638 · outbound

This paper cites CALM: Contrastive Cross-modal Speaking Style Modeling for Expressive Text-to-Speech Synthesis.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset CALM: Contrastive Cross-modal Speaking Style Modeling for Expressive Text-to-Speech Synthesis

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:28:48.621297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:28:45.816756Z digest=sha256:f7869de5874d257a51583a677edfaf568b72ae46f91b55b776dbce1c697a14c1

Observation 0e07d195-a889-4926-be63-00af7372900b · outbound

This paper cites Cross-speaker emotion disentangling and transfer for end-to-end speech synthe- sis,.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset Cross-speaker emotion disentangling and transfer for end-to-end speech synthe- sis,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:28:51.659972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:28:45.921557Z digest=sha256:7a681caab5313d422eaa78af5e3d5a8f11b121df91acd2158dd8a975f50777f0

Observation 7ce033de-03a7-4c40-ad50-22df9f5efbfa · outbound

This paper cites Ultimately, the Azure system synthesizes speech for 5 speakers, CosyV oice2 synthe- sizes speech for another 5 speakers, and F5-TTS synthesizes speech for an additional 2 speakers.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset Ultimately, the Azure system synthesizes speech for 5 speakers, CosyV oice2 synthe- sizes speech for another 5 speakers, and F5-TTS synthesizes speech for an additional 2 speakers

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:28:54.460940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:28:44.265368Z digest=sha256:ae9e3bf932c1415932b0ab1758e71a64a790de9ae766d783fa5ffca9f76c488a

Observation 2642dd13-49f6-4254-a0bf-594957080dcc · outbound

This paper cites Disentangling correlated speaker and noise for speech synthesis via data augmentation and adversarial factoriza- tion,.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset Disentangling correlated speaker and noise for speech synthesis via data augmentation and adversarial factoriza- tion,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:28:51.328227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:28:46.042428Z digest=sha256:540490a181ba6a3b6ea9d2f7e7d211849e09a089197cafd62258084894f14f75

Observation 10093d43-8a68-498b-bdda-59b7196ec64b · outbound

This paper cites Seen and unseen emo- tional style transfer for voice conversion with a new emotional speech dataset,.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset Seen and unseen emo- tional style transfer for voice conversion with a new emotional speech dataset,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:28:51.073704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:28:46.173639Z digest=sha256:7c3824e38a355c86946cce6cd19be98dfb763be394fd068a2bc59b47a3fb8bd3

Observation d76e586a-0303-4c6d-9aeb-3a27757aee85 · outbound

This paper cites Iemocap: Interactive emotional dyadic motion capture database,.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset Iemocap: Interactive emotional dyadic motion capture database,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:28:46.386970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:28:46.386970Z digest=sha256:2cae857df427cd34df08f25503fe283e3fc283138134a0187f889d83b1d28d4d

Observation 68eb2da2-3e00-45b8-b83c-4b10dff8cee3 · outbound

This paper cites Azure speech studio,.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset Azure speech studio,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:28:50.841506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:28:46.491925Z digest=sha256:ecd74789aa50f3e662a514d34255e75761ae471ce3bcc8c098a407ce5da98b6d

Observation 25852870-80ad-436a-ad82-6a8eac9a38a9 · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:28:46.620609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:28:46.620609Z digest=sha256:1bad9979ffa00a9d82c01eba83a941a990a2249f3647374f16b72455a1269dc0

Observation c41608d8-a453-4986-893b-393c8db9e335 · outbound

This paper cites F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:28:46.725703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:28:46.725703Z digest=sha256:b5403f11af38999a011087a71ad006b28bf7a755f57aa257e948e52bf2cc7a0c

Observation 4013cdb8-27d9-47ae-ae3c-7eec84735a0c · outbound

This paper cites Mead: A large-scale audio-visual dataset for emotional talking-face generation,.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset Mead: A large-scale audio-visual dataset for emotional talking-face generation,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:28:50.578718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:28:46.821401Z digest=sha256:4c2dd77f81b87ccb5e1f2e4939cb041b422434a445c6f16afc05eb63e370ab25

Observation 778845df-7422-4fb4-91fb-f946b3164c71 · outbound

This paper cites Clap learning audio concepts from natural language supervision,.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset Clap learning audio concepts from natural language supervision,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:28:46.937080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:28:46.937080Z digest=sha256:6f055f68e4ee42fd04df3f9c2d258834ea735ea129bd61fcdaba646485ccfab6

Observation 1d26394d-f402-4627-92fc-1e5aaf6f69b8 · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:28:47.071239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:28:47.071239Z digest=sha256:cb16529a7943b31345390d8188f0fa3c9768fcd6e93996f88c1bc453510c8d1c

Observation 6594f6f7-d31f-4e70-84cd-8276353c998b · outbound

This paper cites emotion2vec: Self-supervised pre-training for speech emotion representation.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset emotion2vec: Self-supervised pre-training for speech emotion representation

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:28:50.256556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:28:47.187645Z digest=sha256:c943547b75f9346d301c248bab8a875676bc7d26b148e44301dfa904dc87b511

Observation e872d0b9-17ac-46a6-bac7-d619c8bda871 · outbound

This paper cites Generspeech: Towards style transfer for generalizable out-of-domain text-to- speech,.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset Generspeech: Towards style transfer for generalizable out-of-domain text-to- speech,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:28:50.029173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:28:47.306630Z digest=sha256:4ba431777afa4f551e2d2819df7b86e22354a002f26c4186d0c680431231bfbe

Observation b207e2eb-9bbc-49bf-b31f-a04100c052dc · outbound

This paper cites Montreal forced aligner: Trainable text-speech align- ment using kaldi.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset Montreal forced aligner: Trainable text-speech align- ment using kaldi

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:28:47.452788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:28:47.452788Z digest=sha256:3c94583264e5c3216ea8ba2aaf8f94447df0a9cd3059d2b297e63a93977b1d45

Observation 9c9d3d00-fd43-4ea4-a043-9bed53fb2a6a · outbound

This paper cites Style tokens: Un- supervised style modeling, control and transfer in end-to-end speech synthesis,.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset Style tokens: Un- supervised style modeling, control and transfer in end-to-end speech synthesis,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:28:49.785963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:28:47.581107Z digest=sha256:5c65ee823e3c259567a6269a2c04119369174a696824da86303e47efb97d39e4

Observation 70bae382-b6d5-4cdf-819f-89643bb7203d · outbound

This paper cites Hifi-gan: Generative adversarial net- works for efficient and high fidelity speech synthesis,.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset Hifi-gan: Generative adversarial net- works for efficient and high fidelity speech synthesis,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:28:47.720959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:28:47.720959Z digest=sha256:783207e25bbd15c5f3e0f42967cf1191c0350277711f525163e4de05194250d2

Observation 0f2aecc1-bdad-4bd6-96b2-0e5b86de5ee5 · outbound

This paper cites Qwen2-Audio Technical Report.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset Qwen2-Audio Technical Report

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:28:47.826489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:28:47.826489Z digest=sha256:80a9b225682be34b99d7fb9f1813dfd3971cefa4cf55c8b4d8607251d62b1989

Observation 443f867f-75c4-4789-91af-1d32735db711 · outbound

This paper cites Editspeech: A text based speech editing system using partial in- ference and bidirectional fusion,.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset Editspeech: A text based speech editing system using partial in- ference and bidirectional fusion,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:28:49.387912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:28:47.969837Z digest=sha256:372c9fa9df443092d5c5a433c8f93a71d681187cdadd1bf63476cacd0c1ab3fd

Observation 5b8b15e4-92e7-44f9-be3b-f24a4c2ec7f6 · outbound

This paper cites Transforming Spectrum and Prosody for Emotional Voice Conversion with Non-Parallel Training Data.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset Transforming Spectrum and Prosody for Emotional Voice Conversion with Non-Parallel Training Data

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:28:48.097024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:28:48.097024Z digest=sha256:dbec6a2784a3db956b11d5f92bcc0bea1e1f94c49ee695ebb7a0ffcd610ce84d

Observation 8900faa3-d864-46b1-b66e-89201e0ed880 · outbound

This paper cites Converting anyone’s emotion: Towards speaker-independent emotional voice conver- sion,.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset Converting anyone’s emotion: Towards speaker-independent emotional voice conver- sion,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:28:49.073961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:28:48.254853Z digest=sha256:34dca3a72d6882b8797c229e8d6ebf8738321de443f2b82e3567934cb4175d64

Pith citing papers

Observation 6f217865-cea5-42ed-87f4-44b0d1568529 · inbound

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset cites this paper.

Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:28:48.852397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:28:44.020833Z digest=sha256:11d28a7acd7dc4e7abe96826604d0391bf48310e353b4b3621a33d0c76a9098c