Pith. sign in

Paper Citation Record · LEDGER

Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation

As of 17 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 0 inbound Pith citation observations for arXiv:2508.16188.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.16188 v2

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T17:34:02.234077Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

55 of 55 outbound references displayed

  • verified exact7
  • verified fuzzy2
  • unresolved40
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch6

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3cf0d4d2-2ebe-432c-9288-0463dec14467 · outbound

This paper cites LRS3-TED: a large-scale dataset for visual speech recognition.

Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation LRS3-TED: a large-scale dataset for visual speech recognition

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T17:33:56.229318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T17:33:56.229318Z digest=sha256:6c7f7002b4e03b40847a35fe88333fc6b48fb81c50a3928ff69d1430ec90b486

Observation 13b47d6a-81c6-486f-bf1e-89fc4ff8527e · outbound

This paper cites FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs.

Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T17:33:56.301967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T17:33:56.301967Z digest=sha256:91fd41ada4a69138f60c21f5177e4cee3d890a226764aae92cf3e58d27f48346

Observation 5555e9ba-1fed-44de-8d26-972ac3a39424 · outbound

This paper cites MuAViC: A Multilingual Audio-Visual Corpus for Robust Speech Recognition and Robust Speech-to-Text Translation.

Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation MuAViC: A Multilingual Audio-Visual Corpus for Robust Speech Recognition and Robust Speech-to-Text Translation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T17:33:56.409986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T17:33:56.409986Z digest=sha256:655b8f601bcf9d4d08065f5dd50e2f61c66db08218aa80c70936b656637b11d6

Observation 4ba25ac1-08ec-4c10-b491-99f59b2ecf7e · outbound

This paper cites Chang, Sungbok Lee, and Shrikanth S.

Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation Chang, Sungbok Lee, and Shrikanth S

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:34:06.833469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-05T17:33:56.531446Z digest=sha256:88764d6049931ea3186c11a190e08320d29993af6112a29d92ad06aa65e4c70c

Observation 75770544-4a51-430d-a404-736e0a4f1e12 · outbound

This paper cites an unresolved cited work.

Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation Unresolved cited work

Reference 5

Resolution
metadata mismatch
raw_fallback, observed 2026-08-05T17:34:05.497457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-05T17:33:56.630516Z digest=sha256:d718448adaea82eb5457706aa9e5b1e38c5bf4f569877a4bbdadeed21eb070ce

Observation a9401075-ec6d-4da4-9746-ba747df083ea · outbound

This paper cites Cooper, Michael K.

Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation Cooper, Michael K

Reference 6

Resolution
metadata mismatch
raw_fallback, observed 2026-08-05T17:34:05.244459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-05T17:33:56.754534Z digest=sha256:2a8de3f9db8fcde0865a0dd032e95a6fd670c8de5084769ddf8c6a9ebd1ed859

Observation 31389eec-c4ee-457b-9404-3ebfc61d891b · outbound

This paper cites VGGFace2: A dataset for recognising faces across pose and age.

Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation VGGFace2: A dataset for recognising faces across pose and age

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T17:33:56.813807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T17:33:56.813807Z digest=sha256:488f6347ba939a7269b8045d02697d8713bfa78cf2415dee3211ba6bc6314969

Observation f91372ad-c0f2-4ef3-9bd6-62f32321d44d · outbound

This paper cites Large Language Models are Strong Audio-Visual Speech Recognition Learners.

Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation Large Language Models are Strong Audio-Visual Speech Recognition Learners

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-05T17:34:05.071696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-05T17:33:56.937543Z digest=sha256:87028ff55fa481723e5977a8677f0ef8957972af8dc65ff30654c2e138b2166c

Observation 835e4410-3f65-4d47-bf3a-bba964b3caf9 · outbound

This paper cites MinMo: A Multimodal Large Language Model for Seamless Voice Interaction.

Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation MinMo: A Multimodal Large Language Model for Seamless Voice Interaction

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T17:33:57.014492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T17:33:57.014492Z digest=sha256:e19d7a93b07cac9ec235f742aa39bc5b0ef3c59fa0093b9358f4add339eebdd2

Observation d2ba0adf-2b9a-40a9-b9e7-0599c93a6649 · outbound

This paper cites Qwen2-Audio Technical Report.

Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation Qwen2-Audio Technical Report

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T17:33:57.132219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T17:33:57.132219Z digest=sha256:c583f481a2be0ef3f4e44616d5c9b7e2697e22fa76fb916c148d918b88992baa

Observation d293e247-636e-4aec-9648-e436efc70f62 · outbound

This paper cites an unresolved cited work.

Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T17:33:57.271812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T17:33:57.271812Z digest=sha256:348a3cdd1c1802d8777b3b63275015d321bc8483a75b14f1e8991470207bce0a

Observation 9b876059-c3de-432e-9dc2-a5cf68d9b25c · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T17:33:57.386825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T17:33:57.386825Z digest=sha256:bb313cae373bc2a9cd7da2b66fcd48a7fee5e4ae95a5e6c98ae9fb9a765bc9c8

Observation 6acc3141-b617-4c78-8483-ff80c4b596a9 · outbound

This paper cites an unresolved cited work.

Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-05T17:34:06.722664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-05T17:33:57.569543Z digest=sha256:85d2effa6da36a24efabc6857a13ff7686d2af4861b200b7e49f22a1ee792e9c

Observation 5caf015c-fbbf-4dcc-b68c-9331fc6c4e2a · outbound

This paper cites High Fidelity Neural Audio Compression.

Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation High Fidelity Neural Audio Compression

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T17:33:57.695783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T17:33:57.695783Z digest=sha256:8edceb200978d7829fbe4e55815ac0d1e07013c7425035a368626e5231661c75

Observation 70b15896-bed9-4e2a-8865-9c592fdb7958 · outbound

This paper cites an unresolved cited work.

Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T17:33:57.858524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T17:33:57.858524Z digest=sha256:8254b1e4d864addbdff0a61a58eebbb1d181f1149a0341d1752e002fa4a02b2f

Observation 817dc1c0-57e6-4a7d-b923-1c255ea4fc10 · outbound

This paper cites an unresolved cited work.

Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-05T17:34:06.553435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-05T17:33:57.996257Z digest=sha256:420aef40cb4c5b5dee7934fde4c782949c9142add8dcf2e57a117010f0a2c3f7

Observation 7597679a-f89b-461f-9337-07358291bb03 · outbound

This paper cites an unresolved cited work.

Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T17:33:58.204614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T17:33:58.204614Z digest=sha256:0a167e94718f0954a1a9d0296833aa58e204aaacb5eb01f6dd5271e1e74e32e6

Observation 2ae88353-80e5-4496-9b92-98f7dcf8b312 · outbound

This paper cites an unresolved cited work.

Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-05T17:34:06.383296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-05T17:33:58.384346Z digest=sha256:9e72b2450f583a311f83d897e2cea03ecb1d95137fadc568deb5b46594e352f0

Observation d20f6562-f71c-4c16-8d5f-bdf6236aaf4f · outbound

This paper cites HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units.

Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T17:33:58.485256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T17:33:58.485256Z digest=sha256:7c8d4636db61bc1b33e6dfc9ac07501bf6b7559d3778c18b6830dc2e473e51ad

Observation a826c117-363f-48f4-a952-f96380c88037 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation LoRA: Low-Rank Adaptation of Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T17:33:58.560527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T17:33:58.560527Z digest=sha256:b0816e15f5dc8525b87b3bedd123dab1380f8c2a6cc7d4a86e045de67118c8f7

Observation 3710c718-f85f-4bb9-8e13-35981f786184 · outbound

This paper cites Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction.

Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T17:33:58.719267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T17:33:58.719267Z digest=sha256:c034d7617245e9fd3661e9eaa5f17a80ff7a1f4de8ef3aa63a5ee9b4ebd8303c

Observation a4bf7298-392e-475e-9afa-343fcd91e00b · outbound

This paper cites LMCodec: A Low Bitrate Speech Codec With Causal Transformer Models.

Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation LMCodec: A Low Bitrate Speech Codec With Causal Transformer Models

Reference 22

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T17:34:04.833194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-05T17:33:58.848250Z digest=sha256:d905fe2f372faf4ce117bbb8540e6af4c08ca2387bcac3968a0686497e927516

Observation 97e80f69-33b7-42eb-b6dc-6a7bf78de943 · outbound

This paper cites an unresolved cited work.

Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation Unresolved cited work

Reference 23

Resolution
verified exact
doi, observed 2026-08-05T17:34:02.788127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-05T17:33:58.952506Z digest=sha256:d535863b0eba7ee75d0e2c3084fd0587085f055d59e7071c80cbcef1561514b0

Observation c5553f19-6b4a-4f6c-8e28-6bcc66770cb7 · outbound

This paper cites Kimi-Audio Technical Report.

Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation Kimi-Audio Technical Report

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T17:33:59.015664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T17:33:59.015664Z digest=sha256:4dc243c7f3397e245d1a31a613d8f2c641ca49bbc30cc92113ee0760f27aa8b5

Observation 7e1b7d1e-8efa-41bc-ab18-74194098f6d6 · outbound

This paper cites Nicolaou, Athanasios Papaioannou, Guoying Zhao, Björn Schuller, Irene Kotsia, and Stefanos Zafeiriou.

Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation Nicolaou, Athanasios Papaioannou, Guoying Zhao, Björn Schuller, Irene Kotsia, and Stefanos Zafeiriou

Reference 25

Resolution
verified exact
doi, observed 2026-08-05T17:34:02.533049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-05T17:33:59.105478Z digest=sha256:955fb44e5224e1fa6d0d8a4ed7db8e0c5150cc08810ffbd3fe513baf8f4008a0

Observation ffd84a08-b0cc-49a2-a138-7d58b13a8840 · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T17:33:59.183333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T17:33:59.183333Z digest=sha256:64c3edb2b738b3b7e6179549aa686b10cd353f6177968ef22e609bd4199a6914

Observation 4d09edf4-5188-44eb-922d-c41a2922ab78 · outbound

This paper cites an unresolved cited work.

Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T17:33:59.337707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T17:33:59.337707Z digest=sha256:cada2331474070c0cb99350cceca4b14d345aaeb1af2c07e68a33836176b80d4

Observation 663ea35a-21fb-4967-b434-f56e628483c7 · outbound

This paper cites an unresolved cited work.

Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-05T17:34:06.263699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-05T17:33:59.483783Z digest=sha256:efff93181d1c560ed0c30d8f25e8ac4d42870210c8f2ba577c5c25972fbbe88a

Observation 4fc79cfb-f55f-461a-9148-47ecfdc03d70 · outbound

This paper cites Open-MAGVIT2: An Open-Source Project Toward Democratizing Auto-regressive Visual Generation.

Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation Open-MAGVIT2: An Open-Source Project Toward Democratizing Auto-regressive Visual Generation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T17:33:59.610675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T17:33:59.610675Z digest=sha256:e17fa425569433d6e4bf11ae1f27ff845b29f76dee635ccdc2b0884a569e7fab

Observation b5d759a2-6f1a-4089-8172-5397774acf31 · outbound

This paper cites an unresolved cited work.

Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation Unresolved cited work

Reference 30

Resolution
metadata mismatch
raw_fallback, observed 2026-08-05T17:34:04.520513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-05T17:33:59.676356Z digest=sha256:6208a4a0d415e3af534879a204747bc2fafc4e72aa1b8a05983c3fc7df6809f3

Observation cb258ebd-d063-4120-b459-a5dddc0c6ee4 · outbound

This paper cites EmoBox: Multilingual Multi-corpus Speech Emotion Recognition Toolkit and Benchmark.

Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation EmoBox: Multilingual Multi-corpus Speech Emotion Recognition Toolkit and Benchmark

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T17:33:59.835425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T17:33:59.835425Z digest=sha256:a62734f0b1dca7c4135f34acc7ab2e47e9d29267a727ea0d91ff97e98f09e04a

Observation 3bdb0781-3549-47ad-99f6-0415d20c507a · outbound

This paper cites an unresolved cited work.

Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-05T17:34:06.095830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-05T17:33:59.908256Z digest=sha256:0e44857f29a0c3dd8dee5e87d520b46a728caaa176da8bd0d68879e419619896

Observation eda168d1-c2e7-434b-a1a9-5e0c7175fa91 · outbound

This paper cites Spirit LM: Interleaved Spoken and Written Language Model.

Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation Spirit LM: Interleaved Spoken and Written Language Model

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T17:33:59.985974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T17:33:59.985974Z digest=sha256:dd16828a9f72fb9bafb245f95e4f30312f7b0559529412f2edcb517a01580449

Observation aa51acb1-b330-48da-b312-6b86216bf9bc · outbound

This paper cites GPT-4 Technical Report.

Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation GPT-4 Technical Report

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T17:34:00.056533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T17:34:00.056533Z digest=sha256:f15e303ee1f52e4c2e90d5519fa76c4c2aceb86e270f24305eb1c4a1a338c83a

Observation 6f7ad7f9-b680-474e-8659-d0284b47f7ea · outbound

This paper cites Speech Resynthesis from Discrete Disentangled Self-Supervised Representations.

Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation Speech Resynthesis from Discrete Disentangled Self-Supervised Representations

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T17:34:00.126868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T17:34:00.126868Z digest=sha256:f50a85b5d52cb2209e704efa87210289526f161b347cc50f039551f314c37041

Observation 1b72384d-209a-43ae-9000-9f2ff11018fe · outbound

This paper cites Gnana Praveen, Patrick Cardinal, and Eric Granger.

Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation Gnana Praveen, Patrick Cardinal, and Eric Granger

Reference 36

Resolution
metadata mismatch
raw_fallback, observed 2026-08-05T17:34:04.284502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-05T17:34:00.197965Z digest=sha256:8572ef4cf283399e19661dcd72e46f5581571471151ce9d53872e08947db951b

Observation caf100db-4200-489f-8325-cfe907f14d70 · outbound

This paper cites Robust Speech Recognition via Large-Scale Weak Supervision.

Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation Robust Speech Recognition via Large-Scale Weak Supervision

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T17:34:00.308771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T17:34:00.308771Z digest=sha256:9e5b9ea9c64bb57ceb204a31e43476b8f4c016fbaae3b2376a5f5fb69d32f497

Observation b21e640d-fa8a-4b7c-a65a-5a19db3dba99 · outbound

This paper cites You Only Look Once: Unified, Real-Time Object Detection.

Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation You Only Look Once: Unified, Real-Time Object Detection

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T17:34:00.453514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T17:34:00.453514Z digest=sha256:16aab4a11712c4aa618f80e240ec31a7c123c46e22a88adf442cd780f72a8621

Observation 9e32739f-246c-40f6-a297-46a552020ec6 · outbound

This paper cites Filntisis, Radek Danecek, Victoria F.

Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation Filntisis, Radek Danecek, Victoria F

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:34:05.843914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-05T17:34:00.543333Z digest=sha256:4b529dac6fbfa0c0365d94a15cb97298414f9b9fad329a39ccd9ba70d8c41d7a

Observation 5c9913c5-d99d-4286-9b2d-6fb3bf05cf72 · outbound

This paper cites an unresolved cited work.

Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-05T17:34:05.679001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-05T17:34:00.618482Z digest=sha256:2c43cc8a25b793e83e5db992d46ae1e81fa03c192ce3001e196de2a6969bd802

Observation 110c9b3f-1c72-491a-a7ff-363128b915ad · outbound

This paper cites Savchenko.

Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation Savchenko

Reference 41

Resolution
metadata mismatch
raw_fallback, observed 2026-08-05T17:34:03.938367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-05T17:34:00.690614Z digest=sha256:a7227af980f96883c10deed8bff8983eb766cc20602057da194abe2fa71b3ea6

Observation 808bfc5e-bf43-4fc8-9399-8e9f2e1dc2de · outbound

This paper cites Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction.

Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T17:34:00.745331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T17:34:00.745331Z digest=sha256:8d223e2d83e8c7d01a20e5fc79f1715d14feccc03f23d2049734310d6e0a3f4c

Observation 0c3cf4c4-5941-4935-b7f2-d7a544aa5371 · outbound

This paper cites SSR: Alignment-Aware Modality Connector for Speech Language Models.

Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation SSR: Alignment-Aware Modality Connector for Speech Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T17:34:00.859965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T17:34:00.859965Z digest=sha256:801e54d4d37680d14e2e3d288d21aacc6548b29abb429fc8bd82c009e58d4802

Observation c4860a1a-7ece-4929-872b-b83f8f6dfb78 · outbound

This paper cites SALMONN: Towards Generic Hearing Abilities for Large Language Models.

Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation SALMONN: Towards Generic Hearing Abilities for Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T17:34:00.986813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T17:34:00.986813Z digest=sha256:f71fff734c349b4f52dcc13e4be7839a5712e1a9b1c6ebad3a121379ad67c4c8

Observation e6e0c795-db41-4b7d-b1a3-b682878165fa · outbound

This paper cites ED-TTS: Multi-Scale Emotion Modeling using Cross-Domain Emotion Diarization for Emotional Speech Synthesis.

Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation ED-TTS: Multi-Scale Emotion Modeling using Cross-Domain Emotion Diarization for Emotional Speech Synthesis

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-08-05T17:34:03.655927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-05T17:34:01.108489Z digest=sha256:7f64e3b0de1857dde94ad4a808fd82165f308a29935514eea69337093ccc2f5c

Observation 7c0cd412-a729-4de7-97f9-6016df520be0 · outbound

This paper cites an unresolved cited work.

Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation Unresolved cited work

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-05T17:34:01.213055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T17:34:01.213055Z digest=sha256:617a76a1c474040fd35bfbed580836bd31462d69ca40f33f24daf2fe2fcbca37

Observation dfd0e853-501a-417b-9ef9-669af933ccb3 · outbound

This paper cites Learning Emotional Representations from Imbalanced Speech Data for Speech Emotion Recognition and Emotional Text-to-Speech.

Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation Learning Emotional Representations from Imbalanced Speech Data for Speech Emotion Recognition and Emotional Text-to-Speech

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-08-05T17:34:03.399162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-05T17:34:01.368593Z digest=sha256:3c25daf1815a5bda34b9b0b6bbd5bafe805b1a111e57164a097e3645c6c3d70f

Observation e6774275-4465-4540-938b-a465f1af10e6 · outbound

This paper cites On decoder-only architecture for speech-to-text and large language model integration.

Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation On decoder-only architecture for speech-to-text and large language model integration

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-05T17:34:01.439343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T17:34:01.439343Z digest=sha256:ee94695158322d4ca3a70641261b924d0c498609a9615845d99715c35549c3ca

Observation 8ff5aa8c-7edf-46c0-b050-e3613060da67 · outbound

This paper cites Zero-AVSR: Zero-Shot Audio-Visual Speech Recognition with LLMs by Learning Language-Agnostic Speech Representations.

Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation Zero-AVSR: Zero-Shot Audio-Visual Speech Recognition with LLMs by Learning Language-Agnostic Speech Representations

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-08-05T17:34:03.190941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-05T17:34:01.561018Z digest=sha256:ec9b130f4da646805c40ef0c32dce52ea3246035839e8286f2690f488dfe015f

Observation 9f5fa781-25c0-46e3-9844-bdff186365df · outbound

This paper cites MMS-LLaMA: Efficient LLM-based Audio-Visual Speech Recognition with Minimal Multimodal Speech Tokens.

Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation MMS-LLaMA: Efficient LLM-based Audio-Visual Speech Recognition with Minimal Multimodal Speech Tokens

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-08-05T17:34:03.090588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-05T17:34:01.706260Z digest=sha256:cbb72eb8aab127d3e339112965dacd51c9848fce133d92d3cdcc3286bdfbd312

Observation aaea03a9-88b0-4faa-9aa9-509c2580afd1 · outbound

This paper cites Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation.

Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-05T17:34:01.775720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T17:34:01.775720Z digest=sha256:7307857064e58774589c6d4f918a034d128ab640b8a30c61d0958c5cbb59a58c

Observation 0838ca54-6fcb-4d9b-85f9-69feb110e7ab · outbound

This paper cites Connecting Speech Encoder and Large Language Model for ASR.

Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation Connecting Speech Encoder and Large Language Model for ASR

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-05T17:34:01.896810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T17:34:01.896810Z digest=sha256:d56b6b898a81a54c16476262ada06a2db6394a3b1b29ff786dff58ca812fd571

Observation abe02ded-932b-4f26-995d-3162e8dc2afa · outbound

This paper cites SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities.

Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-05T17:34:02.039158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T17:34:02.039158Z digest=sha256:431d16bbf3a7c723b7cca31e3819d607e5ceffeabb012cbbdd5ed3d8033abb30

Observation 31ccc46f-1fbb-4f3a-a6fc-d4040db498ac · outbound

This paper cites online" 'onlinestring :=.

Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation online" 'onlinestring :=

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-05T17:34:02.123009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T17:34:02.123009Z digest=sha256:e25911cd3bc465e14f24dc15c429c90deca5cf2d12ae0b2c1286db9af9720b21

Observation 712385bd-aedf-40f4-a670-0a09eabc81b9 · outbound

This paper cites write newline.

Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation write newline

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-05T17:34:02.234077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T17:34:02.234077Z digest=sha256:ae537476cfb7e17aa0ecc860b468c70afa2cab094d5fbcead823acc3828d280a

Pith citing papers

No inbound Pith citation observations are available.