Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T11:06:50.337605Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 1 inbound Pith citation observation for arXiv:2510.07355.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T11:06:50.337605Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-03T00:56:58.344964Z
A source-named dated measurement, never combined with another source.
Source: cited_works
57 of 57 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 8b51d3a7-2e73-4cf8-bf6d-d8a2ba21af06 · outbound
AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues Emotional communication in speech and music: The role of melodic and rhythmic contrasts,
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d4917ed-eef9-49dc-b21f-0e0d0879834c · outbound
AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues Effects of variation in emotional tone of voice on speech perception,
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56007e68-f4c7-4b26-8fd2-92161abff7f0 · outbound
AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues Analysis of emotion recognition using facial ex- pressions, speech and multimodal information,
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50d76174-2266-43d9-a942-edcaf98c9b63 · outbound
AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues Language models are few-shot learners,
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55ce28e0-4813-4353-8197-81b50a8f88de · outbound
AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues LLaMA-Omni: Seamless Speech Interaction with Large Language Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b8bbee8-5951-4f8c-a92d-3c152e36ae7c · outbound
AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 875ecdf8-0960-418d-b24a-1bc98129c738 · outbound
AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d8c3881-632a-40ca-9478-ee325d43010f · outbound
AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88a46cbb-0276-4a0c-b631-cfcac641ba3a · outbound
AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues Moshi: a speech-text foundation model for real-time dialogue
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d21cad1c-f787-4e84-98ac-87728cf598ea · outbound
AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues EMO-Reasoning: Benchmarking Emotional Rea- soning Capabilities in Spoken Dialogue Systems,
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39279745-9fba-4474-9f3e-652bed7a1193 · outbound
AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues Textually pretrained speech language mod- els,
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0f98697-da92-443a-b0aa-be673aa1316b · outbound
AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues GSQA: An End-to-End Model for Generative Spoken Question Answering
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da687337-5354-4610-8d33-81c741879d91 · outbound
AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues Align-SLM: Textless Spoken Language Models with Reinforcement Learning from AI Feedback
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 125e2bb9-743b-4c43-a38f-8fc9c0a22b07 · outbound
AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues Dynamic-superb: Towards a dynamic, collab- orative, and comprehensive instruction-tuning benchmark for speech,
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e7e6682-6c59-46bc-a742-1b2da70970a0 · outbound
AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues Dynamic-SUPERB Phase-2: A Collaboratively Expanding Benchmark for Measuring the Capabilities of Spoken Language Models with 180 Tasks
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5bcdd55d-b602-4683-9c78-67d02463c97b · outbound
AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues Sd-eval: A benchmark dataset for spoken dialogue understanding beyond words,
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 294f0709-7083-4820-8032-f79d87d4813c · outbound
AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues Advancing Large Language Models to Capture Varied Speaking Styles and Respond Properly in Spoken Conversations
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d79b70b3-246a-4063-ad4d-4050923f06a6 · outbound
AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues Can LLMs Understand the Implication of Emphasized Sentences in Dialogue?
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59b66f37-2dda-4843-acd9-3416ffaabd9e · outbound
AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues Paralinguistics-enhanced large language model- ing of spoken dialogue,
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8cf06290-04af-4d4b-b500-0508e399abb4 · outbound
AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues Full-Duplex-Bench: A Benchmark to Evaluate Full-duplex Spoken Dialogue Models on Turn-taking Capabilities
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 393e8571-d257-4c2c-85f1-4bf558dec92d · outbound
AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues Full-Duplex-Bench v1.5: Evaluating Overlap Handling for Full-Duplex Speech Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a3bd3b4-2cb7-481c-9ffb-82fbbd7114b7 · outbound
AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues EmoNet-Face: An Expert-Annotated Benchmark for Synthetic Emotion Recognition
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1952de1b-9e3c-48fb-83b2-c9ab0b44afeb · outbound
AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues Dfme: A new benchmark for dynamic facial micro- expression recognition,
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 650e59fa-b87e-4d33-996e-ec183875082e · outbound
AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues What comprises a good talking-head video generation?: A Survey and Benchmark
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e28c4f6-9261-415c-b69f-acb4787ddc3d · outbound
AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues Subjective and objective quality-of-experience assessment for 3d talking heads,
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6619faf-fe91-4fdc-91ca-efe844f7e6ac · outbound
AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues Av- data2vec: Self-supervised learning of audio-visual speech representa- tions with contextualized target representations,
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c84f1758-c514-4d74-adb5-c47ff4736a94 · outbound
AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues Jointly Learning From Unimodal and Multimodal-Rated Labels in Audio-Visual Emotion Recognition,
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e491b503-38be-46cd-ae0d-99b8fc7f6426 · outbound
AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues Stimulus Modality Matters: Impact of Perceptual Evaluations from Different Modalities on Speech Emotion Recognition System Performance,
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94344922-5168-4498-b185-ff6548bd7b27 · outbound
AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues Cross-modal incongruity aligning and collaborating for multi-modal sarcasm detection,
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 324591fa-222c-4683-aab5-a9d0f6d31fa5 · outbound
AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues Let’s Go Real Talk: Spoken Dialogue Model for Face-to-Face Conversation,
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5741f955-ca5f-4795-9766-ca92451055e3 · outbound
AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b2d7aaa-cd9f-4a57-8f0c-a142052a77cb · outbound
AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues Facial Expression Recognition with Adaptive Frame Rate based on Multiple Testing Correction,
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19302319-6d60-438c-ba27-2ed0c0b0ef8d · outbound
AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues Introducing the RECOLA multimodal corpus of remote collaborative and affective interactions,
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9d32696-507a-47d3-a9ea-25741e192fb7 · outbound
AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues Avec 2016: Depression, mood, and emotion recognition workshop and challenge,
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69f15386-e08b-434c-b053-4dfb0cccb9d9 · outbound
AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues Hicmae: Hierarchical contrastive masked autoencoder for self-supervised audio-visual emotion recognition,
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e33dd677-042a-4cc7-9867-0beb00476dde · outbound
AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues Baichuan-Omni-1.5 Technical Report,
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 158e03a8-2b26-406f-8e84-9d8f6cbc1791 · outbound
AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e3ad335-ff11-4fb1-966f-20765b90f8e2 · outbound
AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues Qwen2.5-Omni Technical Report
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b570d25-4ff0-4d74-915c-2a9b94f73a29 · outbound
AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues GPT-4 Technical Report
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21296364-8cff-4b23-8a30-2db2bb9e81d5 · outbound
AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d394c90c-79b4-45d5-9426-971267eb3242 · outbound
AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues AI-Face: A Million-Scale Demographically Annotated AI-Generated Face Dataset and Fairness Benchmark,
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18ba8698-1e52-45ca-8930-12de44413092 · outbound
AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues DreamTalk: When Emotional Talking Head Generation Meets Diffusion Probabilistic Models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92494f21-2696-4f33-8c6e-89b5d4088fb3 · outbound
AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues Unconstrained dysfluency modeling for dys- fluent speech transcription and detection,
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 981e0606-c161-4694-be0a-c9722b56dde7 · outbound
AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues Towards hierarchical spo- ken language disfluency modeling,
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e6ba329-a502-4cd3-aafb-36ff0892a010 · outbound
AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues Ssdm: Scalable speech dysfluency model- ing,
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ef1c0e1-47e9-4c16-8d7c-5b26e31ef006 · outbound
AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues Auto- matic detection of articulatory-based disfluencies in primary progres- sive aphasia,
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23c766c1-283b-400f-b708-8f41dd377b11 · outbound
AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues Yolo-stutter: End-to-end region-wise speech dysfluency detection,
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6c095a5-c59c-4296-90a6-f4270545c875 · outbound
AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues Stutter-solver: End-to-end multi- lingual dysfluency detection,
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0fd4617-7d21-4a44-9fa9-f104ec2ee1bf · outbound
AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues Time and tokens: Benchmarking end-to-end speech dys- fluency detection,
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f5d49ad-c691-4f74-946f-1d5f811bd2bb · outbound
AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues Towards accurate phonetic error detection through phoneme similarity modeling,
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee530a01-f84e-46d9-a9e6-15199c6855d0 · outbound
AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues Dysfluent wfst: A framework for zero-shot speech dysfluency transcription and detection,
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbea160c-b9e7-43e5-828d-747d8831208d · outbound
AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues Analysis and evaluation of synthetic data generation in speech dysfluency detec- tion,
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6fab458a-3d21-45c4-9b3e-731303dd4d5f · outbound
AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues LCS-CTC: Leveraging Soft Alignments to Enhance Phonetic Transcription Robustness
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d32a70c-a167-4f3d-a62f-1c516f3e9576 · outbound
AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues Seamless dysfluent speech text alignment for disordered speech analysis,
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f7edb46-3e57-4b63-9836-0d198d68fec6 · outbound
AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues K-function: Joint pronunciation transcription and feedback for evaluating kids language function,
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50249b44-c621-400b-b776-11c39a144355 · outbound
AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues Deep Neural Convolutive Matrix Factorization for Articulatory Representation Decomposition,
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7dea22d-a56e-4a87-a507-4492bc2313e8 · outbound
AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues Articulatory representation learn- ing via joint factor analysis and neural matrix factorization,
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4bbd6493-6c83-43e5-9f56-dba555af0d74 · inbound
S-MARC: Causal Streaming Reasoning for Full-Duplex Conversational Behavior Modeling AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.