Pith. sign in

Paper Citation Record · LEDGER

MELT: Towards Automated Multimodal Emotion Data Annotation by Leveraging LLM Embedded Knowledge

As of 21 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 1 inbound Pith citation observation for arXiv:2505.24493.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.24493 v1

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:27:30.556529Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:27:26.864720Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T12:27:31.120281Z

Reference resolution

39 of 39 outbound references displayed

  • verified exact0
  • verified fuzzy20
  • unresolved17
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 402ac5de-66b2-4199-bf35-69f4371c4609 · outbound

This paper cites MELT: Towards Automated Multimodal Emotion Data Annotation by Leveraging LLM Embedded Knowledge.

MELT: Towards Automated Multimodal Emotion Data Annotation by Leveraging LLM Embedded Knowledge MELT: Towards Automated Multimodal Emotion Data Annotation by Leveraging LLM Embedded Knowledge

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T12:27:31.225846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:27:26.864720Z digest=sha256:b689bf18f48fa97c1e5450a55bb6c8c0112c521dbb4bc1a0d66befe5549c1cec

Observation 0a6528c0-cd32-4f50-9cfd-fd596d54520c · outbound

This paper cites 1st Customer.

MELT: Towards Automated Multimodal Emotion Data Annotation by Leveraging LLM Embedded Knowledge 1st Customer

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:27:34.805142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:27:26.942234Z digest=sha256:2d44d46773bddca6f27749bbdcea25768d61208e25f6a903be27dd97efb5fffb

Observation 633642c8-a18f-4dbf-bec1-c711c7d9406c · outbound

This paper cites Emo Prediction.

MELT: Towards Automated Multimodal Emotion Data Annotation by Leveraging LLM Embedded Knowledge Emo Prediction

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:27:34.608199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:27:27.087019Z digest=sha256:525be8e517cd37a850abfa5ab8197b5eb52e5b78b6960ed17dcd2e150cf4be2c

Observation 26ad3144-fb9d-4890-a481-7c9e6c2b4a38 · outbound

This paper cites Subjective Experiment To assess the emotion annotation quality, we invited 20 par- ticipants, comprising 11 males and 9 females to conduct a Mean Opinion Score (MOS) experiment.

MELT: Towards Automated Multimodal Emotion Data Annotation by Leveraging LLM Embedded Knowledge Subjective Experiment To assess the emotion annotation quality, we invited 20 par- ticipants, comprising 11 males and 9 females to conduct a Mean Opinion Score (MOS) experiment

Reference 4

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T12:27:31.044694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:27:27.228099Z digest=sha256:ef7152180154423418237e6a623bb2ddeb33ebdb173dfc14657ab96496651b33

Observation 77a5abfc-b9ba-4998-a0a1-536b1eb404c4 · outbound

This paper cites Performance The overall MOS result is shown in Fig.

MELT: Towards Automated Multimodal Emotion Data Annotation by Leveraging LLM Embedded Knowledge Performance The overall MOS result is shown in Fig

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:27:34.513190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:27:27.361044Z digest=sha256:f97e4236feefad12bbd0dd1ebf3bd556a9762bdab75a9b24eb9e01de26dfcd03

Observation 2870d4d2-c87c-4e41-b056-05e6767a9f00 · outbound

This paper cites To achieve this, we developed a prompting strategy incorporating cross-validation and CoT rea- soning to ensure consistent and accurate annotations.

MELT: Towards Automated Multimodal Emotion Data Annotation by Leveraging LLM Embedded Knowledge To achieve this, we developed a prompting strategy incorporating cross-validation and CoT rea- soning to ensure consistent and accurate annotations

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:27:34.373866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:27:27.428756Z digest=sha256:596e2487525f5418137fa66660de11a009cc9001ee0f877752961b2e5343d02f

Observation 40bc5abf-dbd0-4a94-9992-2aad67ef0218 · outbound

This paper cites Schuller is also with the Munich Data Science Insti- tute and the Konrad Zuse School of Excellence in Reliable AI, both in Munich, Germany.

MELT: Towards Automated Multimodal Emotion Data Annotation by Leveraging LLM Embedded Knowledge Schuller is also with the Munich Data Science Insti- tute and the Konrad Zuse School of Excellence in Reliable AI, both in Munich, Germany

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:27:34.233138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:27:27.498174Z digest=sha256:2222d1e6118eba6f07c55c585c6b0343a180a011f88bc25d3c1e7c0a81322ce4

Observation 67387979-9e87-4696-b7c9-e968f89b0446 · outbound

This paper cites Be- yond deep learning: Charting the next frontiers of affective com- puting,.

MELT: Towards Automated Multimodal Emotion Data Annotation by Leveraging LLM Embedded Knowledge Be- yond deep learning: Charting the next frontiers of affective com- puting,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:27:34.043674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:27:27.560638Z digest=sha256:ac382ae6c237a9aaeaa209116d27eaea739c2b5b02e76648a03dc7df18d46661

Observation 3f2e4355-2c90-4e2f-b94c-a6a4d105c234 · outbound

This paper cites En- hancing emotional text-to-speech controllability with natural lan- guage guidance through contrastive learning and diffusion mod- els,.

MELT: Towards Automated Multimodal Emotion Data Annotation by Leveraging LLM Embedded Knowledge En- hancing emotional text-to-speech controllability with natural lan- guage guidance through contrastive learning and diffusion mod- els,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:27:33.903281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:27:27.677962Z digest=sha256:9a6079b5f21fe0882af1024e8017b8119cc93fdf01b9ef5fca2c9e86ad15a62c

Observation c5e9f569-d767-4100-bfc5-57040ed1b31d · outbound

This paper cites Emotion recognition in context,.

MELT: Towards Automated Multimodal Emotion Data Annotation by Leveraging LLM Embedded Knowledge Emotion recognition in context,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:27:33.725487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:27:27.799259Z digest=sha256:acd1603ad3900567e568fc47e2abeb2e59555d6714d5c11dc0f10db2e7ab7550

Observation feafe0e5-f059-4a73-84ab-a7cc6783b03b · outbound

This paper cites Contextual Emotion Recognition using Large Vision Language Models.

MELT: Towards Automated Multimodal Emotion Data Annotation by Leveraging LLM Embedded Knowledge Contextual Emotion Recognition using Large Vision Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:27:27.873897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:27:27.873897Z digest=sha256:64c41bc662d509dc87c64b65f73361fc3f17ded4242919b89b79340febab0ae0

Observation a8eff7a5-9bcd-4917-a2c1-47ab01bb0441 · outbound

This paper cites The human in emotion recognition on social media: Attitudes, outcomes, risks,.

MELT: Towards Automated Multimodal Emotion Data Annotation by Leveraging LLM Embedded Knowledge The human in emotion recognition on social media: Attitudes, outcomes, risks,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:27:33.548809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:27:27.960259Z digest=sha256:4f58ce19f4907ad3a129c598bbfa96be86a05f36d261781d1f3f52259d65a74d

Observation e6423b5b-6bd8-4dc8-96b2-940c3cc04980 · outbound

This paper cites Language models are unsupervised multitask learners,.

MELT: Towards Automated Multimodal Emotion Data Annotation by Leveraging LLM Embedded Knowledge Language models are unsupervised multitask learners,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:27:28.088535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:27:28.088535Z digest=sha256:f5105c487f85f952e1ef39af0b1328887b6789859b2a91d6869e16c5eaa9760e

Observation c4c1fadf-d6eb-4ee0-b898-87f0b163e5a7 · outbound

This paper cites Language Models are Few-Shot Learners.

MELT: Towards Automated Multimodal Emotion Data Annotation by Leveraging LLM Embedded Knowledge Language Models are Few-Shot Learners

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:27:28.164916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:27:28.164916Z digest=sha256:232dec8acc3e0df782d1ee1008a3221f420a7c16077fd13bec8fd47992963d6d

Observation 507770af-6dac-4c0d-b692-82ef45a2283f · outbound

This paper cites GPT-4o System Card.

MELT: Towards Automated Multimodal Emotion Data Annotation by Leveraging LLM Embedded Knowledge GPT-4o System Card

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:27:28.249407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:27:28.249407Z digest=sha256:22d596d11030da1173b623a6f9a18b161671c9c919b5971b20e5bd8e24584b6f

Observation 24e35159-b0a7-4c28-9c50-f89fddd93f71 · outbound

This paper cites Large Language Models for Data Annotation and Synthesis: A Survey.

MELT: Towards Automated Multimodal Emotion Data Annotation by Leveraging LLM Embedded Knowledge Large Language Models for Data Annotation and Synthesis: A Survey

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:27:28.359993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:27:28.359993Z digest=sha256:1af83bde266e3afbade49796a09f55230b0bf83a3234ee95cc284da918312059

Observation 5b00d479-fa07-4bc1-8d12-dd421f3e78cf · outbound

This paper cites Chatgpt vs. human annotators: A comprehen- sive analysis of chatgpt for text annotation,.

MELT: Towards Automated Multimodal Emotion Data Annotation by Leveraging LLM Embedded Knowledge Chatgpt vs. human annotators: A comprehen- sive analysis of chatgpt for text annotation,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:27:33.412896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:27:28.424390Z digest=sha256:2fc17007155aaf5c688b29e0115b2ef331a299b4510d7d082eeaa61ba2ecb97e

Observation 2b63ef89-f5eb-45c1-9b2a-d360e614254f · outbound

This paper cites Chatgpt outperforms crowd workers for text-annotation tasks,.

MELT: Towards Automated Multimodal Emotion Data Annotation by Leveraging LLM Embedded Knowledge Chatgpt outperforms crowd workers for text-annotation tasks,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:27:33.259232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:27:28.492592Z digest=sha256:5aa4514e60db1e8fa08dd1619c9fa40c13df481e17d2ef9e780689db89fff478

Observation bc585383-0a91-4d08-9cc7-15eb89bb6556 · outbound

This paper cites Wavcaps: A chatgpt-assisted weakly- labelled audio captioning dataset for audio-language multimodal research,.

MELT: Towards Automated Multimodal Emotion Data Annotation by Leveraging LLM Embedded Knowledge Wavcaps: A chatgpt-assisted weakly- labelled audio captioning dataset for audio-language multimodal research,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:27:33.108845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:27:28.583808Z digest=sha256:5d414a51bf04581a3987a176ec74f9c60032e57b30db22bb703519ef874d58f9

Observation a3d040bf-46f1-4e6d-bd62-7a2bc5540e12 · outbound

This paper cites Pengi: An audio language model for audio tasks,.

MELT: Towards Automated Multimodal Emotion Data Annotation by Leveraging LLM Embedded Knowledge Pengi: An audio language model for audio tasks,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:27:28.671061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:27:28.671061Z digest=sha256:a615a25b1b598db4af859d7f531b777cd319b74a533295111a3cc7181bbca047

Observation bb199048-d0c0-410e-a418-3b9cbe7bbb88 · outbound

This paper cites Secap: Speech emotion captioning with large language model,.

MELT: Towards Automated Multimodal Emotion Data Annotation by Leveraging LLM Embedded Knowledge Secap: Speech emotion captioning with large language model,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:27:32.890406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:27:28.759546Z digest=sha256:989397bb5cf38a4a89b32f7488489e6eddb5bc284c793440f1b0cc3ebed898d2

Observation 829657dd-f225-475a-a7ac-269940d2f011 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

MELT: Towards Automated Multimodal Emotion Data Annotation by Leveraging LLM Embedded Knowledge LLaMA: Open and Efficient Foundation Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:27:28.833637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:27:28.833637Z digest=sha256:3e86752ce2aa2a6ded474edcc3b437fb097da035404004786ea97fd169b690c8

Observation cb9d5785-08b4-4876-b801-17776bf9b817 · outbound

This paper cites Meld: A multimodal multi-party dataset for emo- tion recognition in conversations,.

MELT: Towards Automated Multimodal Emotion Data Annotation by Leveraging LLM Embedded Knowledge Meld: A multimodal multi-party dataset for emo- tion recognition in conversations,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:27:32.684728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:27:28.903608Z digest=sha256:55cb5df634909f8ef738f8c2aa870fab86e0a5cf4a5091ce247a677d3b952aeb

Observation e304351d-024f-49df-a304-ea7058bb1faa · outbound

This paper cites On the time course of vocal emotion recognition,.

MELT: Towards Automated Multimodal Emotion Data Annotation by Leveraging LLM Embedded Knowledge On the time course of vocal emotion recognition,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:27:32.534749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:27:29.042826Z digest=sha256:919999fb9b40b2cfd032576037b6102815135b3aeccd8ad428b870e0a80b2dd2

Observation 81739cf6-dc75-4887-9ef3-f7b436edad95 · outbound

This paper cites Applying tdnn architectures for an- alyzing duration dependencies on speech emotion recognition.

MELT: Towards Automated Multimodal Emotion Data Annotation by Leveraging LLM Embedded Knowledge Applying tdnn architectures for an- alyzing duration dependencies on speech emotion recognition

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:27:32.234416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:27:29.117137Z digest=sha256:e72f2e346351a573c83a02edeb8f1ed4a64c0cfff738fe5388a88ef2596742ec

Observation dd452cb9-5977-497a-9524-dcdbe58347ed · outbound

This paper cites A wide evaluation of chatgpt on affective computing tasks,.

MELT: Towards Automated Multimodal Emotion Data Annotation by Leveraging LLM Embedded Knowledge A wide evaluation of chatgpt on affective computing tasks,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:27:32.062542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:27:29.189552Z digest=sha256:c5016c4e64b640c6c255485e3b06bcf21ed7bb63286feb3a8dc8c73b4a4ed669

Observation 8a311e23-6e8a-4094-aeec-2a6c6432be05 · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech repre- sentations,.

MELT: Towards Automated Multimodal Emotion Data Annotation by Leveraging LLM Embedded Knowledge wav2vec 2.0: A framework for self-supervised learning of speech repre- sentations,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:27:29.295963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:27:29.295963Z digest=sha256:44623bd050ebb5ad1e559501a7efd297ea75665adbb66e1478430f795118ba62

Observation 64cdf607-ff54-4605-a33a-0040025a012e · outbound

This paper cites Dawn of the trans- former era in speech emotion recognition: closing the valence gap,.

MELT: Towards Automated Multimodal Emotion Data Annotation by Leveraging LLM Embedded Knowledge Dawn of the trans- former era in speech emotion recognition: closing the valence gap,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:27:29.431282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:27:29.431282Z digest=sha256:d9147deda7ebef7bc609055b5e5b0f26c4d21bbbdcf97c7f7ea7975d96ef8d69

Observation d0c74c29-d9fe-4bd8-abe6-0afdd43d52d7 · outbound

This paper cites Hubert: Self-supervised speech represen- tation learning by masked prediction of hidden units,.

MELT: Towards Automated Multimodal Emotion Data Annotation by Leveraging LLM Embedded Knowledge Hubert: Self-supervised speech represen- tation learning by masked prediction of hidden units,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:27:29.512545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:27:29.512545Z digest=sha256:ee3b229ed9c7c53f5370192d9efa14745802f2ff854e2ee52bb7ee2148bea711

Observation 7dd3754d-f5d6-4243-b857-115f071b1fed · outbound

This paper cites Wavlm: Large-scale self- supervised pre-training for full stack speech processing,.

MELT: Towards Automated Multimodal Emotion Data Annotation by Leveraging LLM Embedded Knowledge Wavlm: Large-scale self- supervised pre-training for full stack speech processing,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:27:29.656517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:27:29.656517Z digest=sha256:7c5fc34b113b9a2f97e24a80f494102f9f5c103184dee7261649fc1f8ee3092a

Observation 58958196-2d1a-4589-9851-37059db372d5 · outbound

This paper cites Iemocap: Interactive emotional dyadic motion capture database,.

MELT: Towards Automated Multimodal Emotion Data Annotation by Leveraging LLM Embedded Knowledge Iemocap: Interactive emotional dyadic motion capture database,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:27:29.800246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:27:29.800246Z digest=sha256:3aeecedd63c234f6a7979648f482e4cf2f49c5fc341dd28802c48cf0d6b17e55

Observation 77085dc2-383e-455c-b349-9bcb8a4a2a87 · outbound

This paper cites Toronto emotional speech set (tess).

MELT: Towards Automated Multimodal Emotion Data Annotation by Leveraging LLM Embedded Knowledge Toronto emotional speech set (tess)

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:27:29.942112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:27:29.942112Z digest=sha256:936d075a0083d14b1d22f8c334ce1c86417e9e24a6968bf55550ea03a6988bad

Observation a5a16bbb-10cf-4525-9873-a9418b6712ab · outbound

This paper cites The ryerson audio-visual database of emotional speech and song (ravdess): A dynamic, multimodal set of facial and vocal expressions in north american english,.

MELT: Towards Automated Multimodal Emotion Data Annotation by Leveraging LLM Embedded Knowledge The ryerson audio-visual database of emotional speech and song (ravdess): A dynamic, multimodal set of facial and vocal expressions in north american english,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:27:30.046700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:27:30.046700Z digest=sha256:52a5a131553bce0f8652e9a61b850593bb0f720282ba3ea7d9443b583d6fe9ea

Observation 66577d96-c517-4924-aeb0-82a1de589378 · outbound

This paper cites Crema-d: Crowd-sourced emotional multimodal actors dataset,.

MELT: Towards Automated Multimodal Emotion Data Annotation by Leveraging LLM Embedded Knowledge Crema-d: Crowd-sourced emotional multimodal actors dataset,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:27:30.136880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:27:30.136880Z digest=sha256:135ccbded65888d6b312a210c5ed485e26b33e1e8dcda0f16a55c49c418f5c40

Observation c35b71e7-64cb-4d7c-9fe7-cb23a372826d · outbound

This paper cites EMO-SUPERB: An In-depth Look at Speech Emotion Recognition.

MELT: Towards Automated Multimodal Emotion Data Annotation by Leveraging LLM Embedded Knowledge EMO-SUPERB: An In-depth Look at Speech Emotion Recognition

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T12:27:30.198470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:27:30.198470Z digest=sha256:729891a7e10764c9b5fc08578461607b7db3aa17f9cb49ca13277f047ce0816b

Observation bd2260ac-bae1-4883-9ce6-16fcdbcefecc · outbound

This paper cites Paraclap– towards a general language-audio model for computational par- alinguistic tasks,.

MELT: Towards Automated Multimodal Emotion Data Annotation by Leveraging LLM Embedded Knowledge Paraclap– towards a general language-audio model for computational par- alinguistic tasks,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:27:31.857628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:27:30.294929Z digest=sha256:bff06bef54332b4d1cb32016181ef1b665abd03499f1ada94c140e5b8a953b0d

Observation a0867d04-8d09-4afa-88d6-ba04896c76f8 · outbound

This paper cites The geneva minimalistic acoustic parameter set (gemaps) for voice research and affective computing,.

MELT: Towards Automated Multimodal Emotion Data Annotation by Leveraging LLM Embedded Knowledge The geneva minimalistic acoustic parameter set (gemaps) for voice research and affective computing,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T12:27:30.409565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:27:30.409565Z digest=sha256:9717f0980e977cc4c6a2f4656e4a81614e10748b1dd7d5b858427e5f0508a77d

Observation 75f7473e-3a93-4c5f-b694-1ce0eeedf5b5 · outbound

This paper cites openSMILE: the Munich versatile and fast open-source audio feature extractor,.

MELT: Towards Automated Multimodal Emotion Data Annotation by Leveraging LLM Embedded Knowledge openSMILE: the Munich versatile and fast open-source audio feature extractor,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:27:31.684716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:27:30.481144Z digest=sha256:908f85814ef988d02933accfc7a35b406df68072c899c76523c0ffb3558ea7c1

Observation d99fd321-fc6d-4b52-aa62-e1f19f317c61 · outbound

This paper cites How do we describe other people from voices and faces?.

MELT: Towards Automated Multimodal Emotion Data Annotation by Leveraging LLM Embedded Knowledge How do we describe other people from voices and faces?

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:27:31.467360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:27:30.556529Z digest=sha256:aaf093b1f46a5f1ecb3a48e68ef1a2e7d4c5bf56ec558f4e7642048327bb1d3a

Pith citing papers

Observation 402ac5de-66b2-4199-bf35-69f4371c4609 · inbound

MELT: Towards Automated Multimodal Emotion Data Annotation by Leveraging LLM Embedded Knowledge cites this paper.

MELT: Towards Automated Multimodal Emotion Data Annotation by Leveraging LLM Embedded Knowledge MELT: Towards Automated Multimodal Emotion Data Annotation by Leveraging LLM Embedded Knowledge

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T12:27:31.225846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:27:26.864720Z digest=sha256:b689bf18f48fa97c1e5450a55bb6c8c0112c521dbb4bc1a0d66befe5549c1cec