Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T21:29:43.405193Z
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 0 inbound Pith citation observations for arXiv:2501.04904.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T21:29:43.405193Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
42 of 42 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 455ed74b-dcef-479c-9e4e-35ffe4ac91b9 · outbound
JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis Natural TTS Synthesis by Conditioning Wavenet on Mel Spectrogram Predictions,
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5775c4de-2f6f-4eb9-989d-47f62652191c · outbound
JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis FastSpeech 2: Fast and High-Quality End-to-End Text to Speech,
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e7e85684-bfa0-4440-a5d5-4eb2e7cadf91 · outbound
JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis HierSpeech: Bridging the Gap between Text and Speech by Hierarchical Variational Inference using Self-supervised Representations for Speech Synthesis,
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 90b748e9-fa3a-4628-835d-8b599e50947f · outbound
JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis Matcha-TTS: A fast TTS Architecture with Conditional Flow Matching,
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e20bbb16-47d6-40d3-ba20-4e328af25522 · outbound
JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis V oice- Flow: Efficient Text-To-Speech with Rectified Flow Matching,
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 32510b16-9547-40a3-8b5f-d1a683c714f9 · outbound
JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis Conversational End-to-End TTS for V oice Agents,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ba724b39-767d-444f-9a78-f4f13c82626b · outbound
JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis Enhancing Speaking Styles in Conversational Text-to- Speech Synthesis with Graph-Based Multi-Modal Context Modeling,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ee3e98f9-7070-44e0-931e-9431589ab592 · outbound
JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis M2-CTTS: End-to-End Multi-Scale Multi-Modal Conversational Text-to-Speech Synthesis,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 742097bf-bc69-4f58-8f98-1d2a851855ec · outbound
JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis Concss: Contrastive-based Context Comprehension for Dialogue-Appropriate Prosody in Conversational Speech Synthesis,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5ce4b5e3-cffa-4fc9-9454-d64d312ae447 · outbound
JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis Considering Temporal Connection between Turns for Conversational Speech Synthesis,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation fb2c1b8d-15ba-40ce-bab2-c8de5594d452 · outbound
JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis A new recurrent neural-network architecture for visual pattern recognition,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 3a40749b-a6f4-4121-b242-d43c29bd8288 · outbound
JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis Classification of drowsiness levels based on a deep spatio-temporal convolutional bidirectional LSTM network using electroencephalogra- phy signals,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation beffdf87-2a96-4cb6-8721-f96253610690 · outbound
JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis Towards an EEG-based intuitive BCI communication system using imagined speech and visual imagery,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 164e2906-c334-4023-94c1-e87a15e31658 · outbound
JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis A multi-view cnn with novel variance layer for motor imagery brain computer interface,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation fcafa22c-1ec9-4a12-a8e9-a0e61124b666 · outbound
JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis An adaptive deep reinforcement learning framework enables curling robots with human-like performance in real-world conditions,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9c98865d-edeb-494d-960a-ebc03960a0c9 · outbound
JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis PeriodWave: Multi-Period Flow Matching for High-Fidelity Waveform Generation
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ab269ed-6f19-44d1-88c0-899b507cf3f8 · outbound
JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis Emoq-tts: Emotion intensity quantization for fine-grained controllable emotional text-to-speech,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6996cda2-e285-40fb-830d-fb887bd96225 · outbound
JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis Diffprosody: Diffusion-based latent prosody generation for expressive speech synthe- sis with prosody conditional adversarial training,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 984f69b8-44a3-4d20-8b4a-8b5c772f9de0 · outbound
JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis EmoSphere-TTS: Emotional Style and Intensity Modeling via Spherical Emotion Vector for Controllable Emotional Text-to-Speech,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 341a8fea-0dac-4930-8223-911bc5be0baf · outbound
JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis DurFlex-EVC: Duration-Flexible Emotional Voice Conversion Leveraging Discrete Representations without Text Alignment
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94e8086e-f874-41c0-a17f-f6512c75925b · outbound
JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis Emotion rendering for conversational speech synthesis with heterogeneous graph- based context modeling,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 7e155f12-d29e-4db2-ac40-b730b9212873 · outbound
JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4b4056f-d47d-4147-a27a-19bb8f70547a · outbound
JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis BLSP-KD: Bootstrapping Language-Speech Pre-training via Knowledge Distillation
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e31169c-973e-422b-8cda-896d7e88e3e7 · outbound
JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis Whisper-AT: Noise-Robust Automatic Speech Recognizers are Also Strong General Audio Event Taggers,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 72786add-657c-47f6-a067-d0c711000053 · outbound
JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis BLIP-2: Boot- strapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d6b51203-c677-478f-bc65-b3d7ff262897 · outbound
JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis Robust Speech Recognition via Large-Scale Weak Supervision,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2b3da1c3-3ed5-40f3-b68a-f60507af2742 · outbound
JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis Joint Audio and Speech Understanding,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4677b17d-40d3-4301-8b47-3c7cf683dc18 · outbound
JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis HiFi-GAN: Gen- erative Adversarial Networks for Efficient and High Fidelity Speech Synthesis,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 147b7bcf-18f5-4439-b418-170a3605770e · outbound
JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis DailyTalk: Spoken Dialogue Dataset for Conversational Text-to-Speech,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f0ae0780-a0db-4eb4-a9f3-39cd9a8dff17 · outbound
JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis DailyDialog: A Manually Labelled Multi-turn Dialogue Dataset,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 52c67a9b-2bec-4a87-9803-ea1867d58737 · outbound
JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis CREMA-D: Crowd-Sourced Emotional Multimodal Actors Dataset,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 826908ef-0d70-4fe4-9dcb-0992498ea9d1 · outbound
JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis The Emotional Voices Database: Towards Controlling the Emotion Dimension in Voice Generation Systems
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 641c9827-6c87-4460-9c1b-7c46aef6d3cc · outbound
JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis IEMOCAP: Interactive emotional dyadic motion capture database,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 7171c184-92f1-43ca-8ef0-bf1fc0f118e5 · outbound
JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis MEAD: A Large-scale Audio-visual Dataset for Emotional Talking-face Generation,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6c25328d-a94d-4d36-a1fe-da76096e92e1 · outbound
JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis Toronto emotional speech set (tess)-younger talker happy,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2c0d3525-bab5-4ed6-b275-f2e7825b8c83 · outbound
JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis Vicuna: An Open-Source Chatbot Impressing GPT-4 with 90%* ChatGPT Quality,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4def9cc2-dc2c-448f-ae21-4702a0626c30 · outbound
JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis LLaMA: Open and Efficient Foundation Language Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7648c2f8-f3da-4d9d-ae90-fab86882c09a · outbound
JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis Decoupled Weight Decay Regular- ization,
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 67b43322-bc93-41fb-b1cf-9886ae39087c · outbound
JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation,
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5f18c2de-8774-4623-9571-2f397f39f659 · outbound
JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations,
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9c39e78d-c66c-41b2-9c47-f32b48e68740 · outbound
JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis Sequence-to-Sequence Acoustic Modeling for V oice Conversion,
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4df5316b-c29f-4c04-9415-635dde855aaa · outbound
JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis LoRA: Low-Rank Adaptation of Large Language Models,
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
No inbound Pith citation observations are available.