Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T19:24:43.629693Z
Paper Citation Record · LEDGER
As of 23 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 0 inbound Pith citation observations for arXiv:2506.16716.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T19:24:43.629693Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
47 of 47 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation b8860e63-d9d5-451c-97d2-c2be9251fdfb · outbound
V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos Video Summarization Using Deep Neural Networks: A Survey,
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 158d775e-5328-4e1a-9c82-27ee2e104a6e · outbound
V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos Exploring Video Captioning Techniques: A Comprehensive Survey on Deep Learning Methods,
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 406f83ab-514b-45e2-bf0b-beeb70fe0fa2 · outbound
V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos Automatic Image and Video Caption Generation With Deep Learning: A Concise Review and Algorithmic Overlap,
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e9c05887-a454-4d89-a9d9-738dfb89e2c1 · outbound
V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos Paralinguistic and spectral feature extraction for speech emotion classification using machine learning techniques,
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 2a5d50ef-e210-4feb-9889-4cbb998f71bb · outbound
V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos V ocal communication of emotion: A review of research paradigms,
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 01bd08e2-d22d-40e3-a44d-0acfa3074792 · outbound
V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos Does speech rate influence intertemporal decisions? an experimental investigation,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 7ae59257-e3a4-467e-acef-53e88500a753 · outbound
V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos Rhythmic and speech rate effects in the perception of durational cues,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation a2ade0d4-3000-4419-baed-af51b68cd87e · outbound
V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos Emotion and Motivation,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 1206c63c-2edc-40d9-abe7-de2313d37242 · outbound
V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos Language and Emotion: Introduction to the Special Issue,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation eb766553-021e-4fd9-baef-4af8c773e48d · outbound
V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos Audio Description Generation in the Era of LLMs and VLMs: A Review of Transferable Generative AI Technologies,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 154ff93f-b5fc-4b23-800e-19f63f840a1a · outbound
V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos Audio Description in the UK: What works, what doesn’t, and understanding the need for personalising access,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e242ada3-45d2-4b6e-a2de-bea2dfc8a49d · outbound
V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos Audio description: The visual made verbal,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 7ae61579-eaf9-49c0-836d-540cf762688e · outbound
V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos Ambient Lights Influence Perception and Decision-Making,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 54bb42c2-df1a-4170-baae-cb96245cc62e · outbound
V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos Kobayasi,Colorist: a practical handbook for personal and profes- sional use
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation da44f779-e6a1-4dc9-b894-433c89c40ec4 · outbound
V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos Bordwell and K
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 30296089-67be-4a4f-b0a7-9c35ccd29167 · outbound
V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos Tacotron: Towards End-to-End Speech Synthesis
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc9c4cc1-59e0-4f63-a266-cf91ff424250 · outbound
V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos WaveNet: A Gener- ative Model for Raw Audio,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 8d6dbc91-1d9d-4bcf-b033-7731b722fa84 · outbound
V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos Styletts 2: Towards human-level text-to-speech through style diffusion and adversarial training with large speech language models,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 9fc1bf72-70ac-4baa-8d65-5f4eb52cfbc1 · outbound
V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos Yourtts: Towards zero-shot multi-speaker tts and zero-shot voice conversion for everyone,
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5c9a6db-24f0-45a6-b00d-10560b18e533 · outbound
V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos InstructTTS: Modelling Expressive TTS in Discrete Latent Space with Natural Language Style Prompt,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 1a85e798-4907-4c7d-9f46-97d327645299 · outbound
V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos V oxinstruct: Expressive human instruction-to-speech generation with unified multilingual codec language modelling,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation fbf54452-192a-4421-8cd2-3377ceb8c5af · outbound
V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos TextrolSpeech: A Text Style Control Speech Corpus with Codec Language Text-to-Speech Models,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation d6eec240-f6d7-4dea-b8db-02a3be46a599 · outbound
V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos PromptTTS 2: Describing and Generating V oices with Text Prompt,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 574dda41-3981-4538-892a-91d362b3df7f · outbound
V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos What Does Your Face Sound Like? 3D Face Shape towards V oice,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 05f92937-b0e1-4410-9732-ae36ae82d53c · outbound
V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos Multimodal Machine Learning: A Survey and Taxonomy,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 32efbc5e-5cac-43de-afb7-d41906ca3844 · outbound
V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos Multimodal Transformer for Unaligned Multimodal Language Sequences,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation ee20c200-d7fe-4b0d-8b73-2a133dc09926 · outbound
V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos MM-TTS: Multi-Modal Prompt Based Style Transfer for Expressive Text-to-Speech Synthesis,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 6a3d15c0-2bb8-4b13-8a99-002a745e76f8 · outbound
V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos Face2Speech: Towards Multi-Speaker Text-to-Speech Synthesis Using an Embedding Vector Predicted from a Face Image
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 737b1011-0653-4d9a-95b6-266124a944b8 · outbound
V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos Imaginary V oice: Face-Styled Diffusion Model for Text-to-Speech,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 69cb87c1-020e-4e9a-a8af-7ddcc8a71166 · outbound
V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos Face-based V oice Conversion: Learning the V oice behind a Face,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 8e8ae700-4c78-4e78-9ee8-315aa8d2c3bf · outbound
V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos EALD-MLLM: Emotion Analysis in Long-sequential and De-identity videos with Multi-modal Large Language Model,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 9d4793a2-8c95-464a-9375-920ad525d446 · outbound
V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos Prompt-to-Prompt Image Editing with Cross Attention Control,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 126278fd-f50f-43c1-b1eb-eb724913a4e3 · outbound
V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos On the Opportunities and Risks of Foundation Models,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 31ef3ae8-e367-4b71-b60a-6c4ade301bbb · outbound
V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos Learning Transferable Visual Models From Natural Language Super- vision,
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d06d6dc-79dc-4c42-9330-c2ae289fea9f · outbound
V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos ViLT: Vision-and-Language Transformer Without Convolution or Region Supervision,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 5c3ab2d0-8748-4e9b-bcc7-e1f0c1626e44 · outbound
V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos Video (language) modeling: a baseline for generative models of natural videos,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e1701d47-786d-4fa4-a8da-2256797fce40 · outbound
V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos ESCoT: Towards Interpretable Emotional Support Dialogue Systems,
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 7c76fb4c-2b38-4fa2-bf05-0294ed236ded · outbound
V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos Chain-of-thought prompting elicits reasoning in large language models,
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 497cc129-74de-4c4c-83d6-34f02c58d374 · outbound
V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos AutoFoley: Artificial Synthesis of Syn- chronized Sound Tracks for Silent Videos With Deep Learning,
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 9ba9a32c-6ca3-47eb-8882-12b75e67327d · outbound
V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos MM-Diffusion: Learning Multi-Modal Diffusion Models for Joint Audio and Video Generation,
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 594c3961-62c0-4484-bab1-3a1e523497aa · outbound
V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos V ocoder-Based Speech Synthesis from Silent Videos,
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 24afd26e-ed5a-4f7d-8684-ae19e4c93433 · outbound
V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos Sonicvisionlm: Playing sound with vision language models,
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 7b40f8ce-19fe-4f2d-a40a-3538288d9573 · outbound
V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos ‘The problem-centred expert interview’. Combining qual- itative interviewing approaches for investigating implicit expert knowl- edge,
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation d5c21d8a-1c4c-4770-8be6-0cf24d659f51 · outbound
V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos Pleasure-arousal-dominance: A general framework for describing and measuring individual differences in Temperament,
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 2a71155e-bb1e-4264-a9c3-f76e58c61d21 · outbound
V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos Pleasure, Arousal, Dominance: Mehrabian and Russell revisited,
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 75de29f2-1089-402a-a2d7-e7f90cc8894c · outbound
V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos Gemini: A Family of Highly Capable Multimodal Models,
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 1b1aa4ce-d4d3-4b2c-8712-18e69e1f7595 · outbound
V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks,
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
No inbound Pith citation observations are available.