Pith. sign in

Paper Citation Record · LEDGER

JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis

As of 12 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 0 inbound Pith citation observations for arXiv:2501.04904.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.04904 v1

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T21:29:43.405193Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

42 of 42 outbound references displayed

  • verified exact0
  • verified fuzzy36
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 455ed74b-dcef-479c-9e4e-35ffe4ac91b9 · outbound

This paper cites Natural TTS Synthesis by Conditioning Wavenet on Mel Spectrogram Predictions,.

JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis Natural TTS Synthesis by Conditioning Wavenet on Mel Spectrogram Predictions,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:29:43.878455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T21:29:43.250452Z digest=sha256:5101898df18a7042071d2cd65648e0fffaa6050ac63f78bddb0498a989899e99

Observation 5775c4de-2f6f-4eb9-989d-47f62652191c · outbound

This paper cites FastSpeech 2: Fast and High-Quality End-to-End Text to Speech,.

JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis FastSpeech 2: Fast and High-Quality End-to-End Text to Speech,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:29:43.866233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T21:29:43.254928Z digest=sha256:5b75adf67ac698fb62532201bcc12eb02fe0052bc3a8cb9fa4149a8232e7e53a

Observation e7e85684-bfa0-4440-a5d5-4eb2e7cadf91 · outbound

This paper cites HierSpeech: Bridging the Gap between Text and Speech by Hierarchical Variational Inference using Self-supervised Representations for Speech Synthesis,.

JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis HierSpeech: Bridging the Gap between Text and Speech by Hierarchical Variational Inference using Self-supervised Representations for Speech Synthesis,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:29:43.853674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T21:29:43.258540Z digest=sha256:3dcd59b815516ec092eca12bb586f544d1c0949bab7b0dd2293294db7bdad7f1

Observation 90b748e9-fa3a-4628-835d-8b599e50947f · outbound

This paper cites Matcha-TTS: A fast TTS Architecture with Conditional Flow Matching,.

JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis Matcha-TTS: A fast TTS Architecture with Conditional Flow Matching,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:29:43.841512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T21:29:43.262268Z digest=sha256:d3c0932561f3049d8023a6cdd604a935e3d6e09f23eaf87d12884ebf346cfba8

Observation e20bbb16-47d6-40d3-ba20-4e328af25522 · outbound

This paper cites V oice- Flow: Efficient Text-To-Speech with Rectified Flow Matching,.

JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis V oice- Flow: Efficient Text-To-Speech with Rectified Flow Matching,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:29:43.829346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T21:29:43.266357Z digest=sha256:af2bbcaaf93748cf2be72474734eb0e1a7e0afe7a30f54931cb32fa892b91203

Observation 32510b16-9547-40a3-8b5f-d1a683c714f9 · outbound

This paper cites Conversational End-to-End TTS for V oice Agents,.

JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis Conversational End-to-End TTS for V oice Agents,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:29:43.817053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T21:29:43.269897Z digest=sha256:10de73c0f63a2ef1ceeef3e38f53ca56f83bf43beabeecb038e8693eebb55741

Observation ba724b39-767d-444f-9a78-f4f13c82626b · outbound

This paper cites Enhancing Speaking Styles in Conversational Text-to- Speech Synthesis with Graph-Based Multi-Modal Context Modeling,.

JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis Enhancing Speaking Styles in Conversational Text-to- Speech Synthesis with Graph-Based Multi-Modal Context Modeling,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:29:43.804447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T21:29:43.273627Z digest=sha256:6d7f06a629a23d743920f9cd0a389bbccdd78d93260cfd313fe8520910f5e418

Observation ee3e98f9-7070-44e0-931e-9431589ab592 · outbound

This paper cites M2-CTTS: End-to-End Multi-Scale Multi-Modal Conversational Text-to-Speech Synthesis,.

JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis M2-CTTS: End-to-End Multi-Scale Multi-Modal Conversational Text-to-Speech Synthesis,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:29:43.791811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T21:29:43.276720Z digest=sha256:0e0a79124e97a5cb2284f2fd2fffa1abec2ecc757c4125cfee1400b4aaf76f4a

Observation 742097bf-bc69-4f58-8f98-1d2a851855ec · outbound

This paper cites Concss: Contrastive-based Context Comprehension for Dialogue-Appropriate Prosody in Conversational Speech Synthesis,.

JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis Concss: Contrastive-based Context Comprehension for Dialogue-Appropriate Prosody in Conversational Speech Synthesis,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:29:43.780401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T21:29:43.280053Z digest=sha256:3d8df32a12f0911fcfb887529d62ca79965238ab29dc180c74b4c02d0efb0082

Observation 5ce4b5e3-cffa-4fc9-9454-d64d312ae447 · outbound

This paper cites Considering Temporal Connection between Turns for Conversational Speech Synthesis,.

JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis Considering Temporal Connection between Turns for Conversational Speech Synthesis,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:29:43.771439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T21:29:43.283619Z digest=sha256:fc1afc99c9b57376b2e66b91b967034c6ee9a0d7e02120a6c839a637b4e37010

Observation fb2c1b8d-15ba-40ce-bab2-c8de5594d452 · outbound

This paper cites A new recurrent neural-network architecture for visual pattern recognition,.

JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis A new recurrent neural-network architecture for visual pattern recognition,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:29:43.762313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T21:29:43.287141Z digest=sha256:b1855493b3cb9561b9c94a71b49f0e38dc88eefa0b80862c8398ad7bbc8b1491

Observation 3a40749b-a6f4-4121-b242-d43c29bd8288 · outbound

This paper cites Classification of drowsiness levels based on a deep spatio-temporal convolutional bidirectional LSTM network using electroencephalogra- phy signals,.

JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis Classification of drowsiness levels based on a deep spatio-temporal convolutional bidirectional LSTM network using electroencephalogra- phy signals,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:29:43.751693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T21:29:43.290406Z digest=sha256:378d13606e8afdd871563d3d42b524ad03cf07767e0c4cb3d11f61eb60729d8f

Observation beffdf87-2a96-4cb6-8721-f96253610690 · outbound

This paper cites Towards an EEG-based intuitive BCI communication system using imagined speech and visual imagery,.

JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis Towards an EEG-based intuitive BCI communication system using imagined speech and visual imagery,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:29:43.741084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T21:29:43.293873Z digest=sha256:c85cdd25ab3f2054b14a8de3f35c0bf40acc9ab2cc167aabd194479f3e9009da

Observation 164e2906-c334-4023-94c1-e87a15e31658 · outbound

This paper cites A multi-view cnn with novel variance layer for motor imagery brain computer interface,.

JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis A multi-view cnn with novel variance layer for motor imagery brain computer interface,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:29:43.729368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T21:29:43.297415Z digest=sha256:3398a7523a02986a249c056491f26ee30d5592cf8317441d1f4756769f773406

Observation fcafa22c-1ec9-4a12-a8e9-a0e61124b666 · outbound

This paper cites An adaptive deep reinforcement learning framework enables curling robots with human-like performance in real-world conditions,.

JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis An adaptive deep reinforcement learning framework enables curling robots with human-like performance in real-world conditions,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:29:43.717871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T21:29:43.300931Z digest=sha256:9a3e4853fd8c4a4105c728d47bf0fc28492b29fa54c0f88e7f6b803e535e86ff

Observation 9c98865d-edeb-494d-960a-ebc03960a0c9 · outbound

This paper cites PeriodWave: Multi-Period Flow Matching for High-Fidelity Waveform Generation.

JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis PeriodWave: Multi-Period Flow Matching for High-Fidelity Waveform Generation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T21:29:43.304427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:29:43.304427Z digest=sha256:a5dad36df13bd7fdc1c65684c10656f1ada599bedf0cf7d5cc79295457ecb4d2

Observation 6ab269ed-6f19-44d1-88c0-899b507cf3f8 · outbound

This paper cites Emoq-tts: Emotion intensity quantization for fine-grained controllable emotional text-to-speech,.

JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis Emoq-tts: Emotion intensity quantization for fine-grained controllable emotional text-to-speech,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:29:43.707472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T21:29:43.308697Z digest=sha256:274f92e3f838196036c6c119284690e24bce36c9e848e0b6b10aad8d298d63cf

Observation 6996cda2-e285-40fb-830d-fb887bd96225 · outbound

This paper cites Diffprosody: Diffusion-based latent prosody generation for expressive speech synthe- sis with prosody conditional adversarial training,.

JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis Diffprosody: Diffusion-based latent prosody generation for expressive speech synthe- sis with prosody conditional adversarial training,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:29:43.696860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T21:29:43.312310Z digest=sha256:cf381f137a24a5e77cc074c533dedc4a1fd6eba0af65d841eafd0564507184ef

Observation 984f69b8-44a3-4d20-8b4a-8b5c772f9de0 · outbound

This paper cites EmoSphere-TTS: Emotional Style and Intensity Modeling via Spherical Emotion Vector for Controllable Emotional Text-to-Speech,.

JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis EmoSphere-TTS: Emotional Style and Intensity Modeling via Spherical Emotion Vector for Controllable Emotional Text-to-Speech,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:29:43.686283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T21:29:43.316220Z digest=sha256:769122d58e128f1b231d966f5a9b5d40cbaa38db86ebfc94b385e9f2e6c9abcd

Observation 341a8fea-0dac-4930-8223-911bc5be0baf · outbound

This paper cites DurFlex-EVC: Duration-Flexible Emotional Voice Conversion Leveraging Discrete Representations without Text Alignment.

JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis DurFlex-EVC: Duration-Flexible Emotional Voice Conversion Leveraging Discrete Representations without Text Alignment

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T21:29:43.320033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:29:43.320033Z digest=sha256:397c777a9c63146cefdf2924a675b6dd2ffab5c698d103726786f6e74a98526d

Observation 94e8086e-f874-41c0-a17f-f6512c75925b · outbound

This paper cites Emotion rendering for conversational speech synthesis with heterogeneous graph- based context modeling,.

JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis Emotion rendering for conversational speech synthesis with heterogeneous graph- based context modeling,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:29:43.675754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T21:29:43.324595Z digest=sha256:0aa2a1069be868d5dea62a6f797368ceb2995044d3aa8c7a8aa64c926796bd5a

Observation 7e155f12-d29e-4db2-ac40-b730b9212873 · outbound

This paper cites InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model.

JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T21:29:43.328351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:29:43.328351Z digest=sha256:df55b07c5f0bea3c857f11bb8e960a14e774120835fcc5cb672c77c732386b80

Observation d4b4056f-d47d-4147-a27a-19bb8f70547a · outbound

This paper cites BLSP-KD: Bootstrapping Language-Speech Pre-training via Knowledge Distillation.

JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis BLSP-KD: Bootstrapping Language-Speech Pre-training via Knowledge Distillation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T21:29:43.332430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:29:43.332430Z digest=sha256:49dac1c85e097b6e61118045f6c125ac9b12d3d0fd112687dea3af1db8a1ac09

Observation 6e31169c-973e-422b-8cda-896d7e88e3e7 · outbound

This paper cites Whisper-AT: Noise-Robust Automatic Speech Recognizers are Also Strong General Audio Event Taggers,.

JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis Whisper-AT: Noise-Robust Automatic Speech Recognizers are Also Strong General Audio Event Taggers,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:29:43.665088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T21:29:43.336305Z digest=sha256:3ddabda52539f5e3a0754de9dcf9b5ddb33ae4b52d1be9af9b48e2672dfe3612

Observation 72786add-657c-47f6-a067-d0c711000053 · outbound

This paper cites BLIP-2: Boot- strapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models,.

JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis BLIP-2: Boot- strapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:29:43.654723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T21:29:43.340004Z digest=sha256:8b1093bf28e0cd0b499668d5684033ea431d00b6231f1b5ac74caa0a48a488be

Observation d6b51203-c677-478f-bc65-b3d7ff262897 · outbound

This paper cites Robust Speech Recognition via Large-Scale Weak Supervision,.

JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis Robust Speech Recognition via Large-Scale Weak Supervision,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:29:43.646248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T21:29:43.343722Z digest=sha256:33aa5046cd2ad8b77a21548913e62f8f7217965919711be926b6309d0de2d9c2

Observation 2b3da1c3-3ed5-40f3-b68a-f60507af2742 · outbound

This paper cites Joint Audio and Speech Understanding,.

JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis Joint Audio and Speech Understanding,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:29:43.637554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T21:29:43.347055Z digest=sha256:bd4227b91199994689d97b933e0a21043f9802cbf1b95dc85ab5b74849ae566b

Observation 4677b17d-40d3-4301-8b47-3c7cf683dc18 · outbound

This paper cites HiFi-GAN: Gen- erative Adversarial Networks for Efficient and High Fidelity Speech Synthesis,.

JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis HiFi-GAN: Gen- erative Adversarial Networks for Efficient and High Fidelity Speech Synthesis,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:29:43.627949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T21:29:43.350771Z digest=sha256:714db11f722d50789a06fcd4da37ce94ca8aaf159ea412fbcc98c05ef290b06b

Observation 147b7bcf-18f5-4439-b418-170a3605770e · outbound

This paper cites DailyTalk: Spoken Dialogue Dataset for Conversational Text-to-Speech,.

JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis DailyTalk: Spoken Dialogue Dataset for Conversational Text-to-Speech,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:29:43.617661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T21:29:43.354260Z digest=sha256:f58f1df0303a0bad7863b25f86bb38128fc8dcc060e193318cf3f53ea10735a2

Observation f0ae0780-a0db-4eb4-a9f3-39cd9a8dff17 · outbound

This paper cites DailyDialog: A Manually Labelled Multi-turn Dialogue Dataset,.

JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis DailyDialog: A Manually Labelled Multi-turn Dialogue Dataset,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:29:43.607229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T21:29:43.357664Z digest=sha256:fb07801d2282eec49981c2478497fe1ed2040a77d3886318f04a6c4ac4e9ef3e

Observation 52c67a9b-2bec-4a87-9803-ea1867d58737 · outbound

This paper cites CREMA-D: Crowd-Sourced Emotional Multimodal Actors Dataset,.

JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis CREMA-D: Crowd-Sourced Emotional Multimodal Actors Dataset,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:29:43.596339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T21:29:43.361632Z digest=sha256:03a10b324e5354459e6f1af9755a6d6a20694ea8e76b95d1bb86e5751fe8a126

Observation 826908ef-0d70-4fe4-9dcb-0992498ea9d1 · outbound

This paper cites The Emotional Voices Database: Towards Controlling the Emotion Dimension in Voice Generation Systems.

JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis The Emotional Voices Database: Towards Controlling the Emotion Dimension in Voice Generation Systems

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T21:29:43.365529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:29:43.365529Z digest=sha256:5ad407d814eb01358b06bd7282575a79b061155a0be8e814f9b599e3a8bad676

Observation 641c9827-6c87-4460-9c1b-7c46aef6d3cc · outbound

This paper cites IEMOCAP: Interactive emotional dyadic motion capture database,.

JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis IEMOCAP: Interactive emotional dyadic motion capture database,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:29:43.585460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T21:29:43.369873Z digest=sha256:ac5e3dd778d26d58d38084dd43917c2316289f943b561e3ae236be3c667009ca

Observation 7171c184-92f1-43ca-8ef0-bf1fc0f118e5 · outbound

This paper cites MEAD: A Large-scale Audio-visual Dataset for Emotional Talking-face Generation,.

JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis MEAD: A Large-scale Audio-visual Dataset for Emotional Talking-face Generation,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:29:43.574845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T21:29:43.373630Z digest=sha256:701e6d0f4254caa56fc09b5d8cb6f155d12cfa50fa936b4e3eba943b1327d932

Observation 6c25328d-a94d-4d36-a1fe-da76096e92e1 · outbound

This paper cites Toronto emotional speech set (tess)-younger talker happy,.

JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis Toronto emotional speech set (tess)-younger talker happy,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:29:43.564236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T21:29:43.377953Z digest=sha256:3b6c720c181f0d951fe6e28136e1b485bdbe7c95f4294f6c12d554cf6ae42199

Observation 2c0d3525-bab5-4ed6-b275-f2e7825b8c83 · outbound

This paper cites Vicuna: An Open-Source Chatbot Impressing GPT-4 with 90%* ChatGPT Quality,.

JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis Vicuna: An Open-Source Chatbot Impressing GPT-4 with 90%* ChatGPT Quality,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:29:43.552689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T21:29:43.381638Z digest=sha256:f7bd32b74e8b3e052ee094d4d9a959b9f8e1266862b6a9277e53e1823fff25d3

Observation 4def9cc2-dc2c-448f-ae21-4702a0626c30 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis LLaMA: Open and Efficient Foundation Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T21:29:43.385693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:29:43.385693Z digest=sha256:41fdd9f7dc4237c89623e13327bdd325bb825267e8998fd7f01fd174e428c0c5

Observation 7648c2f8-f3da-4d9d-ae90-fab86882c09a · outbound

This paper cites Decoupled Weight Decay Regular- ization,.

JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis Decoupled Weight Decay Regular- ization,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:29:43.542313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T21:29:43.389703Z digest=sha256:6bab99b5f173eb01e1e17587bb1a32318f560fece93210dcff9fddea7bbf2a87

Observation 67b43322-bc93-41fb-b1cf-9886ae39087c · outbound

This paper cites emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation,.

JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:29:43.531861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T21:29:43.393629Z digest=sha256:2c574e019999e40c8e4101ef44a1b4eb55c6aaee84f4ab0b55e552331f4307f4

Observation 5f18c2de-8774-4623-9571-2f397f39f659 · outbound

This paper cites wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations,.

JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:29:43.521601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T21:29:43.397315Z digest=sha256:549b5243bb57671ce8050a528e457a86bc3b91652416559d001313e6c271d426

Observation 9c39e78d-c66c-41b2-9c47-f32b48e68740 · outbound

This paper cites Sequence-to-Sequence Acoustic Modeling for V oice Conversion,.

JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis Sequence-to-Sequence Acoustic Modeling for V oice Conversion,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:29:43.511085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T21:29:43.401154Z digest=sha256:a5aab246a8e1e804259237242f03eaebed67fbe8b69b00c86562a14a6cd6ddea

Observation 4df5316b-c29f-4c04-9415-635dde855aaa · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models,.

JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis LoRA: Low-Rank Adaptation of Large Language Models,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:29:43.499432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T21:29:43.405193Z digest=sha256:ad33cc9798c30d4335b90caa6b46a9fecc716ebca358f2ddc882c12daefa4dc6

Pith citing papers

No inbound Pith citation observations are available.