Pith. sign in

Paper Citation Record · LEDGER

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues

As of 12 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 0 inbound Pith citation observations for arXiv:2412.17292.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.17292 v1

Coverage vector

measured 64 of 64 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T05:41:33.390250Z

measured 64 of 64 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

64 of 64 outbound references displayed

  • verified exact0
  • verified fuzzy25
  • unresolved39
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 10a95dde-38d3-4839-87f4-e1ca93295a73 · outbound

This paper cites Empathy Through Multimodality in Conversational Interfaces.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues Empathy Through Multimodality in Conversational Interfaces

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T05:41:32.836243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:41:32.836243Z digest=sha256:847080079d612b7caa48386bf3f0c5967bbaef604bf28dfffca05af7e0e74f59

Observation 5caa6b3b-b711-4899-95df-ed6e1491726d · outbound

This paper cites Rahmani, and Ramesh Jain.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues Rahmani, and Ramesh Jain

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:41:34.464912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T05:41:32.842235Z digest=sha256:cd4d94c52f15ea836d963d4c23eb35de0ec086678251cf38a0a1aea5686d9952

Observation 85c12ff7-8a16-40da-8665-2e0a3715138a · outbound

This paper cites GPT-4 Technical Report.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues GPT-4 Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T05:41:32.847576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:41:32.847576Z digest=sha256:66badd59b1e494016786c07501b9f5eb0d68016acc9afd3a274e0d4b40badb05

Observation 97c2493c-aef2-490d-89d0-d52181a4266e · outbound

This paper cites Facechat: An emotion-aware face-to-face dialogue framework.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues Facechat: An emotion-aware face-to-face dialogue framework

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:41:34.447857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T05:41:33.069472Z digest=sha256:f4621324e267f0962f8a42cd293e55938b27919eb79a6ad1b080bec3b5c55b06

Observation bbe1a122-f4fc-4dbc-b32f-0f1a8bd7e0ac · outbound

This paper cites PaLM 2 Technical Report.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues PaLM 2 Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T05:41:33.074736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:41:33.074736Z digest=sha256:d43c5515e1f0a766944ab6b8ba81aac10c3ba0008b99a823473e9d78daed3e8b

Observation 91b55cb7-7ec2-419a-9915-154f9b0191c5 · outbound

This paper cites Common Voice: A Massively-Multilingual Speech Corpus.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues Common Voice: A Massively-Multilingual Speech Corpus

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T05:41:33.080449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:41:33.080449Z digest=sha256:7206ed2f5c209fa60883691bc9897a08c32faec1fe531b4a0f314356641b33da

Observation 1165a748-2539-49b5-94f8-188866968abd · outbound

This paper cites Non-verbal communication in human social interaction.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues Non-verbal communication in human social interaction

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:41:34.430997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T05:41:33.086252Z digest=sha256:a077c43cf4d4c312559edfabd99269042c4ec9dfd75fb675011ad6472dcbd2cf

Observation bd8c445e-b357-4cb2-986f-38a17e003b6b · outbound

This paper cites Qwen Technical Report.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues Qwen Technical Report

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T05:41:33.091119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:41:33.091119Z digest=sha256:dc64023c706e1307dc5f99555af44aae327b769a5e088b244ce5fc4733aec6f1

Observation 5e29f143-3003-4d08-976d-b6ee62fba3f2 · outbound

This paper cites Meteor: An automatic metric for mt evaluation with improved correlation with hu- man judgments.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues Meteor: An automatic metric for mt evaluation with improved correlation with hu- man judgments

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:41:34.414928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T05:41:33.097389Z digest=sha256:09ac71b81e020473a744ebbe58f37efd7857e8f0a881995e526e8867f720b5c8

Observation e1fcdb86-810a-40de-8830-5dbe02ddfaca · outbound

This paper cites A neural probabilistic language model.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues A neural probabilistic language model

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:41:34.397506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T05:41:33.101838Z digest=sha256:399c25b0c5a891a30602359e40659ef6fc5b2ad658b879f5b1a6800a51b63735

Observation 9a2599c3-1b19-4067-bd61-18105ac3e172 · outbound

This paper cites Lan- guage models are few-shot learners.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues Lan- guage models are few-shot learners

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T05:41:33.107326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:41:33.107326Z digest=sha256:c6ad89c92bfbf230ca0e22add41f8291618b0f57e1a24c9c9441f44ad60443f0

Observation aad6f30a-782c-4647-a182-b26077c14612 · outbound

This paper cites How far are we from solving the 2d & 3d face alignment problem? (and a dataset of 230,000 3d facial landmarks).

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues How far are we from solving the 2d & 3d face alignment problem? (and a dataset of 230,000 3d facial landmarks)

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T05:41:33.112033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:41:33.112033Z digest=sha256:c1af374b4fdbfe798f21061c45e2571306ea12d5ac04760c9326a4c68ee6bfd6

Observation fa8681dc-d0c3-4886-8af0-bf68a0f6dc01 · outbound

This paper cites Crema-d: Crowd-sourced emotional multimodal actors dataset.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues Crema-d: Crowd-sourced emotional multimodal actors dataset

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:41:34.357690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T05:41:33.117351Z digest=sha256:ea9132824ebbbe0d21fbb358631dbc8458aac3165630041ae3da940dcb4464fb

Observation 7cf50eda-3079-4e43-8608-02e13d757dec · outbound

This paper cites End-to- end object detection with transformers.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues End-to- end object detection with transformers

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T05:41:33.122541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:41:33.122541Z digest=sha256:6d3ffd54eec74af85ad326e2452c60af5524b79c519d1a518077a0f844eb7cf5

Observation 24beb63f-628e-4a50-ad59-36a4a5a335d0 · outbound

This paper cites GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T05:41:33.127633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:41:33.127633Z digest=sha256:de39f487fbdcfa654acf8ae0a8bb57771cb4c39073dc817bd9d884af14eb6804

Observation a5f0462c-2ba7-41fb-b957-f79247347c4f · outbound

This paper cites Palm: Scaling language modeling with pathways.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues Palm: Scaling language modeling with pathways

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T05:41:33.133408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:41:33.133408Z digest=sha256:49897ab2be5b1f5c0d4bac024e0263d21ca75859b97a4b274beb493b6e7a68a1

Observation bb900e1d-4d39-45d1-b8c2-4cc399c98be3 · outbound

This paper cites Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T05:41:33.138391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:41:33.138391Z digest=sha256:e53b9283775cdd6f6b7347c7a207b5cdb04a853b966290f0d744c36ba534a0a7

Observation 053d322a-ba49-424b-86ff-8dbfe45ddf01 · outbound

This paper cites Towards multimodal emotional support con- versation systems, 2024.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues Towards multimodal emotional support con- versation systems, 2024

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:41:34.320642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T05:41:33.143449Z digest=sha256:86d19f7d10c92fd4bde0e724cd186d17933c819a66f859f941d9dec9212fbcd3

Observation a1aef6e3-e7d0-4f97-8552-90ed59f3653d · outbound

This paper cites GoEmotions: A Dataset of Fine-Grained Emotions.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues GoEmotions: A Dataset of Fine-Grained Emotions

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T05:41:33.148396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:41:33.148396Z digest=sha256:ef3f76d373a162ca7f8ef8add7000cf8e6d88b621a206c9a7b3b622513a47d95

Observation 1b9dc277-c2cf-49ef-8775-d2a1dbd601ac · outbound

This paper cites Retinaface: Single-shot multi- level face localisation in the wild.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues Retinaface: Single-shot multi- level face localisation in the wild

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:41:34.304532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T05:41:33.153793Z digest=sha256:ebf7abe3a376bca53ad6d1793808591b0928178014e61e8adfd8df708f76dc49

Observation a801b367-aceb-47b3-9db8-eb8a763541c1 · outbound

This paper cites EmpathyEar: An Open-source Avatar Multimodal Empathetic Chatbot.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues EmpathyEar: An Open-source Avatar Multimodal Empathetic Chatbot

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T05:41:33.159002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:41:33.159002Z digest=sha256:311d90415ee34b6c47311dec4e941eb239bd4b876d7510322e185aaab7fb7e1b

Observation df1fcccf-4332-4586-9d3c-7e4eceaa3ab8 · outbound

This paper cites Imagebind: One embedding space to bind them all,.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues Imagebind: One embedding space to bind them all,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:41:34.287083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T05:41:33.164662Z digest=sha256:4a66450f907a5c0448f0be0f022aa4b785136637311af2facc7ccb0d620676cb

Observation c9633150-aece-4a77-9ed0-9059406c1eb7 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues LoRA: Low-Rank Adaptation of Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T05:41:33.171100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:41:33.171100Z digest=sha256:4f9f1e992a90f030ff21ae74a6752c57915a9bf017ca6e8b942521364e063773

Observation f4f11a56-8622-43b2-9f1f-fe98f6afb096 · outbound

This paper cites AudioGPT: Understanding and Generating Speech, Music, Sound, and Talking Head.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues AudioGPT: Understanding and Generating Speech, Music, Sound, and Talking Head

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T05:41:33.176818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:41:33.176818Z digest=sha256:3bd89e258943ebc66dfec8ac7b780c806c4b883806b91c25778606d81fa8593e

Observation 32dce23d-ac17-4de5-8b3b-99ab082b7ad5 · outbound

This paper cites Perceiver: General perception with iterative attention.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues Perceiver: General perception with iterative attention

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T05:41:33.182934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:41:33.182934Z digest=sha256:2c602cf41dc05498f2bd50ea1e60abb085c83e2701ce17628013bc886f597ed6

Observation e4b3b818-95d7-4f05-ab8b-ab48d29d055f · outbound

This paper cites CoLLaVO: Crayon Large Language and Vision mOdel.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues CoLLaVO: Crayon Large Language and Vision mOdel

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T05:41:33.187905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:41:33.187905Z digest=sha256:0b05191a5c927a5bd9136386a0af51c0d52e2efc766e563743c582647ded0ac9

Observation 49604cd8-84cc-4946-8f70-1c133ee6d42a · outbound

This paper cites A Diversity-Promoting Objective Function for Neural Conversation Models.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues A Diversity-Promoting Objective Function for Neural Conversation Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T05:41:33.193401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:41:33.193401Z digest=sha256:a94cefb4cabd8b85344fcb79c72477ecbfe42b8392718b978b3b75c93404b6f0

Observation f49ec87b-b44e-4ce6-98d7-e511d6024f4e · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T05:41:33.198727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:41:33.198727Z digest=sha256:01ac3a998aae5c5b16b2962181d1e1eee9e6371814f10450867e3e8a91759300

Observation ef1bc8fc-ff9a-49cc-9b51-e64d8ca4cce7 · outbound

This paper cites Rouge: A package for automatic evaluation of summaries.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues Rouge: A package for automatic evaluation of summaries

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:41:34.248264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T05:41:33.203694Z digest=sha256:aa737ea30d029dd58690916d6aa0b5b4a20b867b0e62a73bcc62492c27c803f7

Observation 8fef244b-6d9e-4e28-ab4a-ecd1e1e3a98a · outbound

This paper cites Paralinguistics-enhanced large language modeling of spoken dialogue.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues Paralinguistics-enhanced large language modeling of spoken dialogue

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:41:34.231184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T05:41:33.209055Z digest=sha256:d1303afb945255c8bc61f35e10a0e2b13317deb259895e719183a91d429e66c2

Observation 141e1eaa-952c-45d7-b268-6b88002b2ca9 · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T05:41:33.213822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:41:33.213822Z digest=sha256:bb345dee09df73a12593c807cb36cbc72cf3c882cb8f5552a09b7b34b975cc54

Observation a010c5e9-99e9-4651-aef4-227a4274d111 · outbound

This paper cites Visual instruction tuning.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues Visual instruction tuning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T05:41:33.218859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:41:33.218859Z digest=sha256:29949c60c5bfeffda8743fb6972bccddb73507ffc814b961b050de8ad25d2aca

Observation f538aa0b-225b-40f7-ad69-b948551e7fcb · outbound

This paper cites The ryer- son audio-visual database of emotional speech and song (ravdess): A dynamic, multimodal set of facial and vocal expressions in north american english.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues The ryer- son audio-visual database of emotional speech and song (ravdess): A dynamic, multimodal set of facial and vocal expressions in north american english

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:41:34.192672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T05:41:33.224336Z digest=sha256:a29c9b6369bfa79c3cabae2ea9a20953a129cd5fd9da35aab010888852b79c27

Observation 78d20b3b-01ce-4009-b82c-e1629321318e · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T05:41:33.229353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:41:33.229353Z digest=sha256:fc7c04a6a38b5c373b9973d8cd0f92019e8d11ebd3bd36b836763e40a1edf291

Observation b5ac3c96-bf86-46e3-aa66-1b9d6e0a142d · outbound

This paper cites Gen- erative spoken dialogue language modeling.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues Gen- erative spoken dialogue language modeling

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:41:34.174437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T05:41:33.235122Z digest=sha256:25ba017ba4a3413fce9c9c05a970edaaaaae876424e91790ca5ad1d2124169d1

Observation d28c59e3-69d6-4a40-8262-5099d4ed297e · outbound

This paper cites Librispeech: an asr corpus based on public do- main audio books.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues Librispeech: an asr corpus based on public do- main audio books

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:41:34.157323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T05:41:33.240627Z digest=sha256:43f2d5f5144cc94f7f46e31de7fb2c9dda2df989f09a5875be1e614d08a4e49b

Observation cd831bb4-08ac-497c-8bc1-5c98edf90e94 · outbound

This paper cites Let's Go Real Talk: Spoken Dialogue Model for Face-to-Face Conversation.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues Let's Go Real Talk: Spoken Dialogue Model for Face-to-Face Conversation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T05:41:33.245811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:41:33.245811Z digest=sha256:400307071c3d5acf8d11ba5860023d56e3c7d4e39222e3998d9d27dd5fb852ef

Observation 888f3fa5-ca73-42f2-8aeb-0ca558f1a949 · outbound

This paper cites A Call for Clarity in Reporting BLEU Scores.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues A Call for Clarity in Reporting BLEU Scores

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T05:41:33.253158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:41:33.253158Z digest=sha256:dcd97a8d6c5089e0198d06dcb2effd90379060bffd9676b141786ba798f6af61

Observation e0bbf93c-7a3d-4862-954e-6e8abc9e4d89 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues Learning transferable visual models from natural language supervi- sion

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T05:41:33.259314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:41:33.259314Z digest=sha256:b3d0d5ef97b735788b6101aad79f6e6b6b13176f89c748e83279002cc2503e02

Observation 743ee57f-053e-45ce-b1d1-695af3a57bc1 · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues Robust speech recognition via large-scale weak supervision

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:41:34.131058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T05:41:33.264135Z digest=sha256:415413ac639a7ac8d7db6ac4a2e3ae2d6a215a981677155e40f0487a1338cfe5

Observation 56dcf946-7027-4f02-acaf-3cdb3ed3000d · outbound

This paper cites Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T05:41:33.268906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:41:33.268906Z digest=sha256:5bbaf510f9e75ff6ad74dc1a8815ada74cfaa79bc49a05a7fc72641db061fd36

Observation 603a6181-8037-4858-8586-39910ee2cfbb · outbound

This paper cites Deepface: Closing the gap to human-level per- formance in face verification.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues Deepface: Closing the gap to human-level per- formance in face verification

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:41:34.101671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T05:41:33.273920Z digest=sha256:6f99facdb16d6883317a01baddfdb3d90f9e5bd4746a4e4fc9b176f838bb4e18

Observation 0f5ec78b-f67e-4117-a13d-22e74e8b8cf3 · outbound

This paper cites SALMONN: Towards Generic Hearing Abilities for Large Language Models.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues SALMONN: Towards Generic Hearing Abilities for Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T05:41:33.279337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:41:33.279337Z digest=sha256:da0323186b4ffdc3fd818e031cacdaf3272406f5c33a82707f8fc70f807a330a

Observation a93e3875-9208-4627-a1ef-b2b7cd2e37cc · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues Gemini: A Family of Highly Capable Multimodal Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T05:41:33.284458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:41:33.284458Z digest=sha256:8bcea6f792b391486a040895ba45bc094686cfa8b14fd0f1b667b5df23f84350

Observation 4c6cbd7c-a95e-417c-be6e-c94589a3539c · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues LLaMA: Open and Efficient Foundation Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T05:41:33.289815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:41:33.289815Z digest=sha256:c705e8c519db0121a47a8f1187fb8d2eeb4d4f7d74306d98cdb9785c4e3cb50e

Observation a66f0a91-7e53-451e-9c10-df680246853f · outbound

This paper cites Nonverbal cues in human–robot interaction: A communication studies perspec- tive.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues Nonverbal cues in human–robot interaction: A communication studies perspec- tive

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:41:34.083757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T05:41:33.294830Z digest=sha256:4336a51bd3498f25154c6bf3210fb71b6b41df78d09f63052d57b312bcefd1fc

Observation 197bb5ce-6c80-4c14-a1ec-71a2e2db2c3a · outbound

This paper cites Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T05:41:33.300088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:41:33.300088Z digest=sha256:2cf0ed348233953cb9e0ecb9813612c342e2b3bf493b8fb97dc88c14b3584630

Observation ef9e7d27-876e-44ac-8e72-8338b337f9a1 · outbound

This paper cites NExT-GPT: Any-to-Any Multimodal LLM.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues NExT-GPT: Any-to-Any Multimodal LLM

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T05:41:33.305307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:41:33.305307Z digest=sha256:32d5f4c70a7568542c763383db1b2f04f7a1f58233081957e84ad7e4c52e4b73

Observation 232acf97-afc6-4654-9994-9bd2ff31758a · outbound

This paper cites E-chat: Emotion-sensitive Spoken Dialogue System with Large Language Models.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues E-chat: Emotion-sensitive Spoken Dialogue System with Large Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T05:41:33.310250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:41:33.310250Z digest=sha256:3d9f75cf93071ff684130354bedd8dd7c9d256deded919100a44fd3b34898404

Observation 992ada81-9b8d-49a7-87fa-bc54fdeb9c55 · outbound

This paper cites SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T05:41:33.316145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:41:33.316145Z digest=sha256:826a5eaa9994497b8517e70ba7286d819e9096987e4b838e6ced3b6e2d95502f

Observation 8a4ad359-db96-4a07-a166-3948da8617d7 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T05:41:33.321332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:41:33.321332Z digest=sha256:684d7adcf20b97d8de2e3365666b651bf59209601054b640ebba3c6aa5808e16

Observation 9af723fe-c657-40b5-999e-7fabb25417e7 · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues OPT: Open Pre-trained Transformer Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T05:41:33.326680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:41:33.326680Z digest=sha256:dfece437e98b0961bcad9673eca68a5147d622ae1fef4c2e60c9ca1495014d12

Observation dd18ea5b-f986-47ae-a955-b61271b4d3bf · outbound

This paper cites BERTScore: Evaluating Text Generation with BERT.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues BERTScore: Evaluating Text Generation with BERT

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T05:41:33.332586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:41:33.332586Z digest=sha256:020638026217f3d1b453604d465fa2321387c3fb04fbe64c142b190f4708261d

Observation 6ac453cb-0805-4ce5-9ad7-f2d2d0388e5e · outbound

This paper cites DialoGPT: Large-Scale Generative Pre-training for Conversational Response Generation.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues DialoGPT: Large-Scale Generative Pre-training for Conversational Response Generation

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T05:41:33.337719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:41:33.337719Z digest=sha256:72d84f68296ec40a816de39a8c43222be1888f18a8d703cd9d692d8890620c35

Observation b02e6bb9-6af8-4011-87e8-2eba579bab59 · outbound

This paper cites A higher BLEU score indicates a more natural and engaging dialogue model.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues A higher BLEU score indicates a more natural and engaging dialogue model

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:41:34.065645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T05:41:33.343532Z digest=sha256:f396cd708ea5ea2a6a0777e8c7b6b268003842a66fb99a862ed9685256d5dca8

Observation 5ea4968a-f7f4-4369-8911-42d70cb02f94 · outbound

This paper cites Fluency evaluates the grammatical correctness, smoothness, and natural flow of the response.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues Fluency evaluates the grammatical correctness, smoothness, and natural flow of the response

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:41:34.049017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T05:41:33.348630Z digest=sha256:4b6f6584210ffcfad7bf616ac1ed43bddb8dadfa1d2d31aefdd9180d8ea22653

Observation fae329b9-1064-4ffa-a20b-f112b6440eed · outbound

This paper cites an unresolved cited work.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-11T05:41:34.030674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T05:41:33.354052Z digest=sha256:878df29be5c39a031c90fc5ab4abe8a4c7941244189e14194dfb217e80920ef9

Observation dbc4d716-2063-4973-b899-f2669da73c0b · outbound

This paper cites It con- tains 2,615 hours of English speech from 92,325 voices with diverse genders, ages, and accents.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues It con- tains 2,615 hours of English speech from 92,325 voices with diverse genders, ages, and accents

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:41:34.013229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T05:41:33.359058Z digest=sha256:7f0a3747c4b285125f0509b167359c1be41d7c0b83350573c2a82f76b4384438

Observation aa9cfbb6-a444-4f41-9f1c-b09b7e64ca78 · outbound

This paper cites The prompt given to the GPT is: These are the frames in a video.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues The prompt given to the GPT is: These are the frames in a video

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:41:33.994769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T05:41:33.364317Z digest=sha256:576cac1543eddc4e40386e0a359ef8fb3d448fb19ef79aab9b908c15b69a479d

Observation e7b34924-928c-47b7-a190-3fd516d0764e · outbound

This paper cites Yeah, I think it’s unfair how the FD burns 6 tons of books.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues Yeah, I think it’s unfair how the FD burns 6 tons of books

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:41:33.976852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T05:41:33.369504Z digest=sha256:d158e967fb30f5a5ed90120453eef7b9107e145966c12659100a0d7b11d717f9

Observation cbdd5d48-a78e-4d9c-8bd0-c8f1cc0766b8 · outbound

This paper cites Understanding the context is crucial for a fair evaluation.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues Understanding the context is crucial for a fair evaluation

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:41:33.960759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T05:41:33.374722Z digest=sha256:af921c464eab3bd50f7b770257ed8bf5c0758fe3524efba6bd6f4c0e76491a5d

Observation f59e432a-3a1e-44cd-9012-1442a6576f8b · outbound

This paper cites Consider if the emotion expressed is suitable for the situation.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues Consider if the emotion expressed is suitable for the situation

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:41:33.943115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T05:41:33.379738Z digest=sha256:f2e1b9fd7b25bfb44759884a871d38dad95907f78ff21fa09b0000d2e519a433

Observation c9647bcc-93cc-485d-a050-452bec147607 · outbound

This paper cites an unresolved cited work.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-08-11T05:41:33.926595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T05:41:33.384727Z digest=sha256:5cdc30a6ccecf0516d48e3d20d703c2bc6b1fe7b0aebf23d4e5dbb4439aa19fe

Observation 8a3b6a30-15fb-463b-aba8-f6a6c29d223a · outbound

This paper cites I've been feeling really down lately. Nothing seems to cheer me up.

AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues I've been feeling really down lately. Nothing seems to cheer me up

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:41:33.910063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T05:41:33.390250Z digest=sha256:79f3b8430f4cc77b4d028b5e7a926e9faf8ec4f20bf4991272ca898e966d55f6

Pith citing papers

No inbound Pith citation observations are available.