Pith. sign in

Paper Citation Record · LEDGER

V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos

As of 23 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 0 inbound Pith citation observations for arXiv:2506.16716.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.16716 v1

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T19:24:43.629693Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

47 of 47 outbound references displayed

  • verified exact0
  • verified fuzzy44
  • unresolved3
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b8860e63-d9d5-451c-97d2-c2be9251fdfb · outbound

This paper cites Video Summarization Using Deep Neural Networks: A Survey,.

V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos Video Summarization Using Deep Neural Networks: A Survey,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:24:44.464110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T19:24:43.394542Z digest=sha256:fb47022e60264c77e0d0b72a53e258aab691cea8f43225dad0aa2e14fd81e337

Observation 158d775e-5328-4e1a-9c82-27ee2e104a6e · outbound

This paper cites Exploring Video Captioning Techniques: A Comprehensive Survey on Deep Learning Methods,.

V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos Exploring Video Captioning Techniques: A Comprehensive Survey on Deep Learning Methods,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:24:44.446713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T19:24:43.400358Z digest=sha256:5e91375289cdc1fcb6f8fe2c91e55eaf3be76868e91546b58ec7ca4366fc5489

Observation 406f83ab-514b-45e2-bf0b-beeb70fe0fa2 · outbound

This paper cites Automatic Image and Video Caption Generation With Deep Learning: A Concise Review and Algorithmic Overlap,.

V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos Automatic Image and Video Caption Generation With Deep Learning: A Concise Review and Algorithmic Overlap,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:24:44.428989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T19:24:43.405355Z digest=sha256:8c3942131656f752342def5319dc241e3e592fc4c25cd0908a41f84b5a0ac88e

Observation e9c05887-a454-4d89-a9d9-738dfb89e2c1 · outbound

This paper cites Paralinguistic and spectral feature extraction for speech emotion classification using machine learning techniques,.

V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos Paralinguistic and spectral feature extraction for speech emotion classification using machine learning techniques,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:24:44.412894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T19:24:43.410147Z digest=sha256:aef6d9533e77b0ecce64bfa8cfc3191351893d96eb0bf099ce9669f415a06682

Observation 2a5d50ef-e210-4feb-9889-4cbb998f71bb · outbound

This paper cites V ocal communication of emotion: A review of research paradigms,.

V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos V ocal communication of emotion: A review of research paradigms,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:24:44.396586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T19:24:43.415413Z digest=sha256:90529988b588cb088957e450e0c92757c8f485e85303913290408bc1d6cbcbc9

Observation 01bd08e2-d22d-40e3-a44d-0acfa3074792 · outbound

This paper cites Does speech rate influence intertemporal decisions? an experimental investigation,.

V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos Does speech rate influence intertemporal decisions? an experimental investigation,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:24:44.380378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T19:24:43.420385Z digest=sha256:0085bc892fb1fc163ac38f3a7d6f28b16b6141c09e90fe0b3d8428cfcfc60caa

Observation 7ae59257-e3a4-467e-acef-53e88500a753 · outbound

This paper cites Rhythmic and speech rate effects in the perception of durational cues,.

V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos Rhythmic and speech rate effects in the perception of durational cues,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:24:44.364329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T19:24:43.426198Z digest=sha256:60bb6986dfe5a2f42b9d02c518d149c08e08283f5d2479cc882e7a21f0c47402

Observation a2ade0d4-3000-4419-baed-af51b68cd87e · outbound

This paper cites Emotion and Motivation,.

V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos Emotion and Motivation,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:24:44.347065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T19:24:43.431109Z digest=sha256:beb70bab25532ae60c613e2a3e0ad51cd8930f5d54890a4d56f3ec02e64e327d

Observation 1206c63c-2edc-40d9-abe7-de2313d37242 · outbound

This paper cites Language and Emotion: Introduction to the Special Issue,.

V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos Language and Emotion: Introduction to the Special Issue,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:24:44.330583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T19:24:43.435971Z digest=sha256:add361f429e29384961cf98ac6ec7608c7dc0410a29df53d02c537682209411b

Observation eb766553-021e-4fd9-baef-4af8c773e48d · outbound

This paper cites Audio Description Generation in the Era of LLMs and VLMs: A Review of Transferable Generative AI Technologies,.

V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos Audio Description Generation in the Era of LLMs and VLMs: A Review of Transferable Generative AI Technologies,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:24:44.313997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T19:24:43.440950Z digest=sha256:379aa8632083fd4d4b1a835795b2f6ba5ed3749f38971e156dfdd656c6d9fd0a

Observation 154ff93f-b5fc-4b23-800e-19f63f840a1a · outbound

This paper cites Audio Description in the UK: What works, what doesn’t, and understanding the need for personalising access,.

V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos Audio Description in the UK: What works, what doesn’t, and understanding the need for personalising access,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:24:44.296960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T19:24:43.445798Z digest=sha256:091a09b6902392383e2cd0fb0012ed4823eec15408a2beae9ac0196ae6823435

Observation e242ada3-45d2-4b6e-a2de-bea2dfc8a49d · outbound

This paper cites Audio description: The visual made verbal,.

V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos Audio description: The visual made verbal,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:24:44.279056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T19:24:43.450609Z digest=sha256:08d1de770bb635918f8a144a8f55905c4850a9f37cdf28acdb93f76fb6b7e177

Observation 7ae61579-eaf9-49c0-836d-540cf762688e · outbound

This paper cites Ambient Lights Influence Perception and Decision-Making,.

V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos Ambient Lights Influence Perception and Decision-Making,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:24:44.260231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T19:24:43.455802Z digest=sha256:7341d69edf8389606050419161f6117f4ca8a9c68149f06ebcaaa76cdcd5ba03

Observation 54bb42c2-df1a-4170-baae-cb96245cc62e · outbound

This paper cites Kobayasi,Colorist: a practical handbook for personal and profes- sional use.

V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos Kobayasi,Colorist: a practical handbook for personal and profes- sional use

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:24:44.243014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T19:24:43.463527Z digest=sha256:9f5afabd351c020ea8d67f55821346fa5d293ed3f44964710f8c650f955b1e8f

Observation da44f779-e6a1-4dc9-b894-433c89c40ec4 · outbound

This paper cites Bordwell and K.

V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos Bordwell and K

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:24:44.226099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T19:24:43.468534Z digest=sha256:fb92cc80a8419b00ee87dff07fed7c30b5624a8d058093325ca41639d0b5ebef

Observation 30296089-67be-4a4f-b0a7-9c35ccd29167 · outbound

This paper cites Tacotron: Towards End-to-End Speech Synthesis.

V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos Tacotron: Towards End-to-End Speech Synthesis

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T19:24:43.473186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:24:43.473186Z digest=sha256:01773319339ace841e0495ef84e258f76e9028cab2498341a3f7bec16eb1d48a

Observation cc9c4cc1-59e0-4f63-a266-cf91ff424250 · outbound

This paper cites WaveNet: A Gener- ative Model for Raw Audio,.

V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos WaveNet: A Gener- ative Model for Raw Audio,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:24:44.209238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T19:24:43.478474Z digest=sha256:fc409495b91f8dce16234151857bc5b2d04f29388af6f127f32e0b7183173dee

Observation 8d6dbc91-1d9d-4bcf-b033-7731b722fa84 · outbound

This paper cites Styletts 2: Towards human-level text-to-speech through style diffusion and adversarial training with large speech language models,.

V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos Styletts 2: Towards human-level text-to-speech through style diffusion and adversarial training with large speech language models,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:24:44.192402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T19:24:43.483267Z digest=sha256:3a37d5dcbfa6417b4e05c862209ee53639f87ed6fa261a12bb82ab0717198646

Observation 9fc1bf72-70ac-4baa-8d65-5f4eb52cfbc1 · outbound

This paper cites Yourtts: Towards zero-shot multi-speaker tts and zero-shot voice conversion for everyone,.

V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos Yourtts: Towards zero-shot multi-speaker tts and zero-shot voice conversion for everyone,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T19:24:43.488222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:24:43.488222Z digest=sha256:b73ce132a29adadc6d23828ca57b256100d5e371c7164ac7bc6fd23150b3d8ae

Observation f5c9a6db-24f0-45a6-b00d-10560b18e533 · outbound

This paper cites InstructTTS: Modelling Expressive TTS in Discrete Latent Space with Natural Language Style Prompt,.

V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos InstructTTS: Modelling Expressive TTS in Discrete Latent Space with Natural Language Style Prompt,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:24:44.163673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T19:24:43.492914Z digest=sha256:1b1b62308a92e94b45e5eb7bbbd1d9b8c4ed6d96d36c45af9e043a0e1cfdbc97

Observation 1a85e798-4907-4c7d-9f46-97d327645299 · outbound

This paper cites V oxinstruct: Expressive human instruction-to-speech generation with unified multilingual codec language modelling,.

V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos V oxinstruct: Expressive human instruction-to-speech generation with unified multilingual codec language modelling,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:24:44.147043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T19:24:43.498094Z digest=sha256:3e439758ef4f3e4fcf553e97e1d0378ca00c4702a6cee5a796cb2298ce4bda12

Observation fbf54452-192a-4421-8cd2-3377ceb8c5af · outbound

This paper cites TextrolSpeech: A Text Style Control Speech Corpus with Codec Language Text-to-Speech Models,.

V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos TextrolSpeech: A Text Style Control Speech Corpus with Codec Language Text-to-Speech Models,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:24:44.130230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T19:24:43.503050Z digest=sha256:51c6d6ceaa317df4dfd7cba7033167cca5ddd70b45692fc60cfe9fec5d8f549c

Observation d6eec240-f6d7-4dea-b8db-02a3be46a599 · outbound

This paper cites PromptTTS 2: Describing and Generating V oices with Text Prompt,.

V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos PromptTTS 2: Describing and Generating V oices with Text Prompt,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:24:44.113137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T19:24:43.507944Z digest=sha256:f9fa461c6e32d174615f0bc7d8cc99014d08384b1896db7bcb5eacbbd3ff61e4

Observation 574dda41-3981-4538-892a-91d362b3df7f · outbound

This paper cites What Does Your Face Sound Like? 3D Face Shape towards V oice,.

V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos What Does Your Face Sound Like? 3D Face Shape towards V oice,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:24:44.096453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T19:24:43.512894Z digest=sha256:ddc160ad258c957a6214d235e95ea95aed3e7c8a8290dd27fa852eb9d8bcd527

Observation 05f92937-b0e1-4410-9732-ae36ae82d53c · outbound

This paper cites Multimodal Machine Learning: A Survey and Taxonomy,.

V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos Multimodal Machine Learning: A Survey and Taxonomy,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:24:44.078241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T19:24:43.517703Z digest=sha256:69c2a12c418e1ec17c2227f91eb54710ba0a35b308eef77bac71fb06375a8cc0

Observation 32efbc5e-5cac-43de-afb7-d41906ca3844 · outbound

This paper cites Multimodal Transformer for Unaligned Multimodal Language Sequences,.

V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos Multimodal Transformer for Unaligned Multimodal Language Sequences,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:24:44.060827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T19:24:43.522791Z digest=sha256:974e28d0c619940db72f7d8fca1a1a962ca96abcb6fe10f9e5b238ecd930b6f1

Observation ee20c200-d7fe-4b0d-8b73-2a133dc09926 · outbound

This paper cites MM-TTS: Multi-Modal Prompt Based Style Transfer for Expressive Text-to-Speech Synthesis,.

V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos MM-TTS: Multi-Modal Prompt Based Style Transfer for Expressive Text-to-Speech Synthesis,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:24:44.042914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T19:24:43.527689Z digest=sha256:16a36f1b442cddedd9bd38ade8187b3c93dacd1bd5941fe11f7be9f47857a1ba

Observation 6a3d15c0-2bb8-4b13-8a99-002a745e76f8 · outbound

This paper cites Face2Speech: Towards Multi-Speaker Text-to-Speech Synthesis Using an Embedding Vector Predicted from a Face Image.

V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos Face2Speech: Towards Multi-Speaker Text-to-Speech Synthesis Using an Embedding Vector Predicted from a Face Image

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:24:44.024907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T19:24:43.532663Z digest=sha256:3c68c13b56ef74965fae214cf921267d159d511dee18b9302e5f355b691bb9e0

Observation 737b1011-0653-4d9a-95b6-266124a944b8 · outbound

This paper cites Imaginary V oice: Face-Styled Diffusion Model for Text-to-Speech,.

V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos Imaginary V oice: Face-Styled Diffusion Model for Text-to-Speech,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:24:44.006035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T19:24:43.538046Z digest=sha256:da51374ad03a1e3cf9bd9d30538764e1b450a393832651fbd5854c0c8c7b801e

Observation 69cb87c1-020e-4e9a-a8af-7ddcc8a71166 · outbound

This paper cites Face-based V oice Conversion: Learning the V oice behind a Face,.

V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos Face-based V oice Conversion: Learning the V oice behind a Face,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:24:43.987917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T19:24:43.542951Z digest=sha256:ca123af1a59a071022f7f49b1f47817e4d7f98e9ddb20c27ba1ad32619e93ace

Observation 8e8ae700-4c78-4e78-9ee8-315aa8d2c3bf · outbound

This paper cites EALD-MLLM: Emotion Analysis in Long-sequential and De-identity videos with Multi-modal Large Language Model,.

V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos EALD-MLLM: Emotion Analysis in Long-sequential and De-identity videos with Multi-modal Large Language Model,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:24:43.970554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T19:24:43.547738Z digest=sha256:12e1bbb82cc8150c0875f58a06c2c2f3e1131c01f5a668ba0c431c16a2788833

Observation 9d4793a2-8c95-464a-9375-920ad525d446 · outbound

This paper cites Prompt-to-Prompt Image Editing with Cross Attention Control,.

V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos Prompt-to-Prompt Image Editing with Cross Attention Control,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:24:43.952962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T19:24:43.552971Z digest=sha256:65bc5e532342767e8924a76bb47220dc4560be11fcd6a1a895d4b66f1c4899d3

Observation 126278fd-f50f-43c1-b1eb-eb724913a4e3 · outbound

This paper cites On the Opportunities and Risks of Foundation Models,.

V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos On the Opportunities and Risks of Foundation Models,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:24:43.934970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T19:24:43.557812Z digest=sha256:2fc008761f70472fbc23d08a605044a71f7b49d46d5822e51abfe4ac62e0a4b1

Observation 31ef3ae8-e367-4b71-b60a-6c4ade301bbb · outbound

This paper cites Learning Transferable Visual Models From Natural Language Super- vision,.

V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos Learning Transferable Visual Models From Natural Language Super- vision,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T19:24:43.562721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:24:43.562721Z digest=sha256:23f73aa05ea2209971bdd32f8483a27f6aa4e50eb091620bfcc5344143f92ea0

Observation 3d06d6dc-79dc-4c42-9330-c2ae289fea9f · outbound

This paper cites ViLT: Vision-and-Language Transformer Without Convolution or Region Supervision,.

V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos ViLT: Vision-and-Language Transformer Without Convolution or Region Supervision,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:24:43.904770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T19:24:43.567538Z digest=sha256:3a607219668680680fee3f1fbe9ded1a9072a5f3f77fe604191612f462c95145

Observation 5c3ab2d0-8748-4e9b-bcc7-e1f0c1626e44 · outbound

This paper cites Video (language) modeling: a baseline for generative models of natural videos,.

V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos Video (language) modeling: a baseline for generative models of natural videos,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:24:43.886972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T19:24:43.572588Z digest=sha256:7a2f4058699003911641fc5810a1878908499fdde0f2d20ed2d500c817862d9c

Observation e1701d47-786d-4fa4-a8da-2256797fce40 · outbound

This paper cites ESCoT: Towards Interpretable Emotional Support Dialogue Systems,.

V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos ESCoT: Towards Interpretable Emotional Support Dialogue Systems,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:24:43.869201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T19:24:43.577944Z digest=sha256:f768c8f470633729160f2ce4f3f43e2b353a53434e5653f9684660d687070dce

Observation 7c76fb4c-2b38-4fa2-bf05-0294ed236ded · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models,.

V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos Chain-of-thought prompting elicits reasoning in large language models,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:24:43.851199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T19:24:43.582760Z digest=sha256:9ae77eb205a9943021e0731fc586fb0f3ad8345521259fcc47c6933600fade86

Observation 497cc129-74de-4c4c-83d6-34f02c58d374 · outbound

This paper cites AutoFoley: Artificial Synthesis of Syn- chronized Sound Tracks for Silent Videos With Deep Learning,.

V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos AutoFoley: Artificial Synthesis of Syn- chronized Sound Tracks for Silent Videos With Deep Learning,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:24:43.833608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T19:24:43.588104Z digest=sha256:7392df11dd0e13df62ead4281e6f0266cf3fad6903f9fd64d51ad381cb5d275d

Observation 9ba9a32c-6ca3-47eb-8882-12b75e67327d · outbound

This paper cites MM-Diffusion: Learning Multi-Modal Diffusion Models for Joint Audio and Video Generation,.

V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos MM-Diffusion: Learning Multi-Modal Diffusion Models for Joint Audio and Video Generation,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:24:43.816183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T19:24:43.593219Z digest=sha256:a12c566e093c0fd1bedb4e0f36bd0bbec56b414fcdb2beeb9732ff8d656561f7

Observation 594c3961-62c0-4484-bab1-3a1e523497aa · outbound

This paper cites V ocoder-Based Speech Synthesis from Silent Videos,.

V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos V ocoder-Based Speech Synthesis from Silent Videos,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:24:43.799106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T19:24:43.598464Z digest=sha256:6cdfe6e1136275fdb784d7c267bb90c1fcec4b350713d422e3ea67fdc9a220ae

Observation 24afd26e-ed5a-4f7d-8684-ae19e4c93433 · outbound

This paper cites Sonicvisionlm: Playing sound with vision language models,.

V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos Sonicvisionlm: Playing sound with vision language models,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:24:43.780714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T19:24:43.603410Z digest=sha256:7191336b9517a92132c60c9903c811c317948c19996d93b5213b28ad46461860

Observation 7b40f8ce-19fe-4f2d-a40a-3538288d9573 · outbound

This paper cites ‘The problem-centred expert interview’. Combining qual- itative interviewing approaches for investigating implicit expert knowl- edge,.

V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos ‘The problem-centred expert interview’. Combining qual- itative interviewing approaches for investigating implicit expert knowl- edge,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:24:43.761256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T19:24:43.608825Z digest=sha256:383c0333d2334c5df7fd7e3ac14ea99b54bccd00cd8641c160f3da2e748a77a6

Observation d5c21d8a-1c4c-4770-8be6-0cf24d659f51 · outbound

This paper cites Pleasure-arousal-dominance: A general framework for describing and measuring individual differences in Temperament,.

V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos Pleasure-arousal-dominance: A general framework for describing and measuring individual differences in Temperament,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:24:43.743018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T19:24:43.614045Z digest=sha256:fbebc9b6b467478cff6a7d83880112f38399ccdc86db736a8ab9335d3ade1a0c

Observation 2a71155e-bb1e-4264-a9c3-f76e58c61d21 · outbound

This paper cites Pleasure, Arousal, Dominance: Mehrabian and Russell revisited,.

V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos Pleasure, Arousal, Dominance: Mehrabian and Russell revisited,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:24:43.726248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T19:24:43.619728Z digest=sha256:b0073bba81c623872736a469536111ddca0bb75e696d1ba4322c6a67e7d33c21

Observation 75de29f2-1089-402a-a2d7-e7f90cc8894c · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models,.

V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos Gemini: A Family of Highly Capable Multimodal Models,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:24:43.708286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T19:24:43.624968Z digest=sha256:73b8098dd6dbdc55a311c933c6f20226661973e5d27280b86194ecd0a08f1ee1

Observation 1b1aa4ce-d4d3-4b2c-8712-18e69e1f7595 · outbound

This paper cites Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks,.

V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:24:43.690513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T19:24:43.629693Z digest=sha256:4167fffe5c2a64f3025f92d255288b5e35fa6159251afd1638de13790f3cddbc

Pith citing papers

No inbound Pith citation observations are available.