Pith. sign in

Paper Citation Record · LEDGER

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections

As of 18 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 1 inbound Pith citation observation for arXiv:2608.11576.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.11576 v1

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:36:42.599014Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:36:42.391185Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-16T00:36:42.953731Z

Reference resolution

54 of 54 outbound references displayed

  • verified exact6
  • verified fuzzy29
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9994dfab-e618-4b92-91d6-f403c9decf4b · outbound

This paper cites Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-16T00:36:42.957902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:36:42.391185Z digest=sha256:1426e182a8ee75e4d49fdbb8ada0bce9fb5641c09e42d171ee8b30fc834332a8

Observation 49f7049e-e796-46e3-9af6-b670310f557f · outbound

This paper cites Early work in modern video-to-music generation systems uses large-scale web music-video corpora and autoregressive modeling over semantic acoustic tokens to generate music [8].

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Early work in modern video-to-music generation systems uses large-scale web music-video corpora and autoregressive modeling over semantic acoustic tokens to generate music [8]

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:43.330732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:36:42.395669Z digest=sha256:7c257dce5404f6598e49f03302794921c5556f28c4f2da6703f563e93d31a40b

Observation d1f0f9bd-1d9e-4177-aebe-5ce875711696 · outbound

This paper cites trance music.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections trance music

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:43.319454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:36:42.399453Z digest=sha256:413da98c6e52697840de3c58b3a003ee1fdcf7ca741a3e267bed2323d5c13607

Observation f0743c4e-cba8-4a45-915f-35bb08426abd · outbound

This paper cites an unresolved cited work.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-16T00:36:43.308324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:36:42.403577Z digest=sha256:4fa9e2c9b84b05684dbb2e93f8a361dfa67789377ed7c448eca36300828bd336

Observation 1369da0c-923d-4a28-b521-da386bbf7aa4 · outbound

This paper cites First, we bench- mark video-to-music generation tasks by comparing existing video-to-music generation models trained and evaluated on identical data, using OSSL-v2.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections First, we bench- mark video-to-music generation tasks by comparing existing video-to-music generation models trained and evaluated on identical data, using OSSL-v2

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:43.298362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:36:42.408016Z digest=sha256:b8a45d4cf91950ac09bfebb81795d9da3b572b36d7e2c524a6c6a36421d21f1b

Observation 442247b2-c215-4420-b9d4-a54b69ada218 · outbound

This paper cites +Dialogue.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections +Dialogue

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:43.287922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:36:42.412530Z digest=sha256:e38b4fbc227dbaebe36007cd74105efbd771def49e055996b1c5ccbc3fe9c54a

Observation 4d97db7e-2296-4f1a-a2c9-6fff5d6fbc44 · outbound

This paper cites Be- cause the dataset is free from link rot and does not require separate web scraping, our dataset is suitable as a durable benchmark for the field.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Be- cause the dataset is free from link rot and does not require separate web scraping, our dataset is suitable as a durable benchmark for the field

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:43.275395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:36:42.416447Z digest=sha256:f52aaea8ca144af6b82ebf2fda5e6bc056a9a4ecdcd3fe03e889bf0bc18d7138

Observation df5e0c5b-906b-47e3-959a-cc4b271b19e6 · outbound

This paper cites Teaser Generation for Long Documentaries and Educational Videos.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Teaser Generation for Long Documentaries and Educational Videos

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:43.264105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:36:42.420783Z digest=sha256:6407bde9e360888f443c6c903490e73a429281bbc309bcee5400cc56d5790464

Observation 48d25803-3e24-4035-bc77-87021a8676c0 · outbound

This paper cites Attendaffectnet–emotion pre- diction of movie viewers using multimodal fusion with self-attention,.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Attendaffectnet–emotion pre- diction of movie viewers using multimodal fusion with self-attention,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:43.252296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:36:42.424057Z digest=sha256:fa11a58e05bcadab4955fc550dbae0af82326a963280bc26352e94cfea6d264b

Observation cefa76f8-82df-4414-a38f-80e911720855 · outbound

This paper cites Predicting emotion from music videos: exploring the relative contribution of visual and auditory information to affective responses.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Predicting emotion from music videos: exploring the relative contribution of visual and auditory information to affective responses

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:42.427623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:42.427623Z digest=sha256:4226bf7e0b4f239b85fb204d1f661f5c3c186ff9ffcf9cc91ecfbc3076fc336a

Observation aefa4075-821b-4a94-9155-3c07c6c43813 · outbound

This paper cites The cognitive processing of film and musical soundtracks,.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections The cognitive processing of film and musical soundtracks,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:43.240784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:36:42.431355Z digest=sha256:48e914ca2e6406e0d2f9757c0dc2bcf1843864e7162dbc4531d12255fa40e431

Observation 008a119a-ccf9-473b-b194-b9633b4b2f70 · outbound

This paper cites Multimodal deep models for predicting affec- tive responses evoked by movies.,.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Multimodal deep models for predicting affec- tive responses evoked by movies.,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:43.229494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:36:42.434869Z digest=sha256:bdeb55e9b293f05e8cf522887d382b74be19bbb7bae80ac253120913cd4f6b6b

Observation fa33de3d-8bb4-4105-b5cc-52ed53cdaad7 · outbound

This paper cites On music’s potential to convey meaning in film: A systematic review of empirical evi- dence,.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections On music’s potential to convey meaning in film: A systematic review of empirical evi- dence,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:43.219280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:36:42.438501Z digest=sha256:d4376c01e552cf2d48d704011cb3f15a04c2dafa36374ac741ba638af5c72fa7

Observation e4bb1dcc-a488-4e3f-b982-43011b6020d4 · outbound

This paper cites Emotion Embedding Spaces for Matching Music to Stories.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Emotion Embedding Spaces for Matching Music to Stories

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:42.442268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:42.442268Z digest=sha256:7fda1b446078a02966418b0f95caf6b9e86b106dcab056ea2eaabc902acfa630

Observation b24fc0fe-dfd4-463c-96d2-b5d1293e74d7 · outbound

This paper cites Foley music: Learning to generate music from videos,.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Foley music: Learning to generate music from videos,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:43.208400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:36:42.446296Z digest=sha256:91cacf03838520a3dc63fc42f20755ca02a81cdf1e20c9340b09fbbc38d6cae4

Observation 4b2cefab-443c-4b9e-8172-064e5352d578 · outbound

This paper cites V2meow: Meow- ing to the visual beat via video-to-music generation,.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections V2meow: Meow- ing to the visual beat via video-to-music generation,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:43.197380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:36:42.450042Z digest=sha256:1aa4c3c85280c4501343f6c3be0564f47c57fddf261e68118d92059938dd42db

Observation fd6a3ed6-191b-4b45-8089-1bf9bbd4218a · outbound

This paper cites Vidmuse: A simple video-to-music generation framework with long-short-term modeling,.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Vidmuse: A simple video-to-music generation framework with long-short-term modeling,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:43.185861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:36:42.453909Z digest=sha256:1b81ddfed619e3d82c86fb2710377c125b5073debcd8527a4c02190cd4edf733

Observation 1ceb4568-70ff-40a4-9f39-ff64d417ccc9 · outbound

This paper cites Vmas: Video-to-music generation via semantic alignment in web music videos,.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Vmas: Video-to-music generation via semantic alignment in web music videos,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:43.174223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:36:42.457562Z digest=sha256:fc783a2484a254cbf1cd1ff7c781bc4ddf7cb26c543b5a15b3509f7ceccb564e

Observation 0a616f9c-5105-4098-9803-76aed2621f3b · outbound

This paper cites Sonique: Video background music generation using unpaired audio- visual data,.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Sonique: Video background music generation using unpaired audio- visual data,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:43.163597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:36:42.461759Z digest=sha256:c2f1aa4253ebe8717aa6623ecc76d1b05e6268d84f90a7daea2594c759cf3d38

Observation 707f5136-b83a-41bc-bac4-8007b047d9ec · outbound

This paper cites GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:42.465952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:42.465952Z digest=sha256:4d7bd188388595ed0793e061a25ad08790eb0b8fb5fe31932541a5828213d69d

Observation 19f194b0-b420-4d16-8e89-42189f1c22a4 · outbound

This paper cites Vision-to-Music Generation: A Survey.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Vision-to-Music Generation: A Survey

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:42.470428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:42.470428Z digest=sha256:1ca67ad3a265c8eea024dd1571c4990e4169fca675d27a8a0e716dd09db9c18a

Observation a329da2a-9f37-45c7-a4d3-bad6ee9ea753 · outbound

This paper cites Ai- based chinese-style music generation from video con- tent: a study on cross-modal analysis and generation methods,.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Ai- based chinese-style music generation from video con- tent: a study on cross-modal analysis and generation methods,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:43.151434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:36:42.474275Z digest=sha256:fc4f58f6550c55bb6f4e15088ad104f54dd6d5db98ac0207b8db9b4035287f94

Observation 81b7990b-37ce-403d-83d9-2d9c84b6d211 · outbound

This paper cites V2M-Zero: Zero-Pair Time-Aligned Video-to-Music Generation.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections V2M-Zero: Zero-Pair Time-Aligned Video-to-Music Generation

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-08-16T00:36:42.893043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:36:42.478417Z digest=sha256:742289288f284978f366fccca7e2084ea84cd30d010bb3e0f22cfc70e97f7518

Observation aa944ece-843f-4cd3-9506-b9cc63eb9cc7 · outbound

This paper cites Diff- v2m: A hierarchical conditional diffusion model with explicit rhythmic modeling for video-to-music genera- tion,.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Diff- v2m: A hierarchical conditional diffusion model with explicit rhythmic modeling for video-to-music genera- tion,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:43.139456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:36:42.482121Z digest=sha256:e397415786663262d28de0a3d9f7700b7f65986d98dec6b5097597e94b7fb4b7

Observation 10022ec1-ca19-4aa1-bf0f-aad20ccbb1fe · outbound

This paper cites Quan- tized gan for complex music generation from dance videos,.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Quan- tized gan for complex music generation from dance videos,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:43.128702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:36:42.489105Z digest=sha256:8a7eb8e3afd2ee1170cb301f6e7b34102c7bb7d514ca7823535867ad6d7fab64

Observation b2d7c4b9-30bd-4184-86b6-a939bf58ac3e · outbound

This paper cites Video background music genera- tion: Dataset, method and evaluation,.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Video background music genera- tion: Dataset, method and evaluation,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:43.117905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:36:42.492941Z digest=sha256:3d43185dfcba771521918c2b3753a993112784c59569e3d99271c828c9ff09c4

Observation 990607ec-3e49-4516-874e-35c7c1f52409 · outbound

This paper cites Content-Based Video-Music Retrieval Using Soft Intra-Modal Structure Constraint.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Content-Based Video-Music Retrieval Using Soft Intra-Modal Structure Constraint

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:42.496612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:42.496612Z digest=sha256:63481b27cfa0bf6d506a72b989f891ba9ed0e0c1c6ec453e54f74f5bd2fd28db

Observation 8147f160-bd41-4192-82f0-c30b73745ab6 · outbound

This paper cites Video-Guided Text-to-Music Generation Using Public Domain Movie Collections.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Video-Guided Text-to-Music Generation Using Public Domain Movie Collections

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:42.500746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:42.500746Z digest=sha256:63c16a51e92e5c33a3033ed2547c32427df26841bffb0811cc22095bd4cfe81b

Observation fa6e8aab-9e17-454e-b054-7b958e060045 · outbound

This paper cites Acoustic profiles in vocal emotion expression.,.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Acoustic profiles in vocal emotion expression.,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:43.106258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:36:42.504700Z digest=sha256:0522a746c713b5bbe17334649e91d80a1729fc144fbca2bf85add26fb043ef2d

Observation 6abfe8f2-f194-4e2c-84c1-41efdb5a99a1 · outbound

This paper cites Background ducking to produce esthetically pleasing audio for tv with clear speech,.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Background ducking to produce esthetically pleasing audio for tv with clear speech,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:43.092776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:36:42.508491Z digest=sha256:ba5dbb364c31bebfb4d974a32f918adbbe4fa00f05c29a7c5d6dce35fe614884

Observation bee2f69a-b172-452f-ae3e-7047d87f4192 · outbound

This paper cites Improving dialogue intelligi- bility in streaming media,.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Improving dialogue intelligi- bility in streaming media,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:43.080172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:36:42.512290Z digest=sha256:f7e865f063d58fe75dfcbd862ce6930e12b84b2e3e00d4cba3e91446b3115517

Observation 120bc34a-93c5-4792-90e6-0256dc84d0c3 · outbound

This paper cites Creating a multitrack clas- sical music performance dataset for multimodal mu- sic analysis: Challenges, insights, and applications,.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Creating a multitrack clas- sical music performance dataset for multimodal mu- sic analysis: Challenges, insights, and applications,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:43.068775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:36:42.517678Z digest=sha256:81d8a7f73eaeab8532a70a61ab88c744a07e885651e4c747be2acb152926f59b

Observation ebc889d5-7b93-4f4d-9a5e-3aeb1039d244 · outbound

This paper cites Video2music: Suitable music generation from videos using an affective multimodal transformer model,.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Video2music: Suitable music generation from videos using an affective multimodal transformer model,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:43.058045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:36:42.521020Z digest=sha256:a59137c763aaf5fc3a4766aea30775f1849c3d581b095d5ba48992d4272c43da

Observation e1e7e50f-29fd-4335-a102-3ddf5172aec5 · outbound

This paper cites Extending Visual Dynamics for Video-to-Music Generation.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Extending Visual Dynamics for Video-to-Music Generation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:42.524714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:42.524714Z digest=sha256:7e35df43ecbf417b59f85a2cda2087a87ddcde47258aac95bbce2efc6f6ba8a2

Observation bd94a98a-47b9-4228-ba40-0b7c2eac11fb · outbound

This paper cites VidMusician: Video-to-Music Generation with Semantic-Rhythmic Alignment via Hierarchical Visual Features.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections VidMusician: Video-to-Music Generation with Semantic-Rhythmic Alignment via Hierarchical Visual Features

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:42.528088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:42.528088Z digest=sha256:150dc6807c1e2c03575cfeee279ffc072a42030679550ec2f87b78c675cd38df

Observation f06b666f-bace-4d58-a581-5d524fbabd79 · outbound

This paper cites MuVi: Video-to-Music Generation with Semantic Alignment and Rhythmic Synchronization.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections MuVi: Video-to-Music Generation with Semantic Alignment and Rhythmic Synchronization

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:42.531770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:42.531770Z digest=sha256:7b6bf1833fa3ce4cfa45e701e034a5afb7a96c1d84d184a3952f7b33765a94c3

Observation 6e1b4654-799c-4717-9f79-8ce89d3db55f · outbound

This paper cites Video echoed in music: Semantic, temporal, and rhythmic alignment for video-to-music generation,.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Video echoed in music: Semantic, temporal, and rhythmic alignment for video-to-music generation,

Reference 38

Resolution
verified exact
raw_fallback, observed 2026-08-16T00:36:42.818139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:36:42.535975Z digest=sha256:c8a9e552c4dee439f461045903039203915614e88676eacadc427400ad502485

Observation 7c0c2411-6d68-43cc-8757-d2ece80bb540 · outbound

This paper cites Video-Robin: Autoregressive Diffusion Planning for Intent-Grounded Video-to-Music Generation.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Video-Robin: Autoregressive Diffusion Planning for Intent-Grounded Video-to-Music Generation

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-08-16T00:36:42.743314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:36:42.539678Z digest=sha256:b8dc999d9f89e67e4f976c85a40ddcfae14d5d0e229962b96fa4f6bbc37c4eeb

Observation 21d31ded-83a0-4005-be0e-91e719cd9342 · outbound

This paper cites Harmonizing Pixels and Melodies: Maestro-Guided Film Score Generation and Composition Style Transfer.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Harmonizing Pixels and Melodies: Maestro-Guided Film Score Generation and Composition Style Transfer

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-08-16T00:36:42.726835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:36:42.544084Z digest=sha256:65a0fb06696db3880b0ac1e84d06af277f3a030d3dcd495fe7df8ea494372cd5

Observation 1209e8e0-5c05-4638-a493-3f2d2ff08163 · outbound

This paper cites FilmComposer: LLM-Driven Music Production for Silent Film Clips.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections FilmComposer: LLM-Driven Music Production for Silent Film Clips

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:42.548207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:42.548207Z digest=sha256:58d48dedd0ba744e0a2971933e52eb1bc7a8ab33edbe954ba7bb204dabd2d14b

Observation 72562688-b4de-4467-b3ec-8b104b63b241 · outbound

This paper cites M$^{2}$UGen: Multi-modal Music Understanding and Generation with the Power of Large Language Models.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections M$^{2}$UGen: Multi-modal Music Understanding and Generation with the Power of Large Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:42.551715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:42.551715Z digest=sha256:c5b653441d62e93e0d376b17124fa4cf14321648ecab8739d06eec1c25201560

Observation 16bf29af-c87e-4b33-9971-4c33d6f8d77d · outbound

This paper cites MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:42.555094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:42.555094Z digest=sha256:9820511f2a0f953990d6ed9c9c84365403194a5d066ffdd99c67d07bf95a5db3

Observation 8d8c3492-5c3a-4909-9fa0-35fc7b2ed464 · outbound

This paper cites Multimodal Music Generation with Explicit Bridges and Retrieval Augmentation.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Multimodal Music Generation with Explicit Bridges and Retrieval Augmentation

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:42.558498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:42.558498Z digest=sha256:498faf46bad9e2d17f0997f8836913c6c656b2367b6cd1ed6baa4c78afaec6d3

Observation b074ce61-01be-4387-8b04-d49dfe2f24fd · outbound

This paper cites Benchmarks and leaderboards for sound demixing tasks.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Benchmarks and leaderboards for sound demixing tasks

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-08-16T00:36:42.661931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:36:42.562317Z digest=sha256:b15fc915dc727a469993ce4df32040d83f89129982c8c32d38729e9c99b4e9e0

Observation eb3d6160-8e0d-4e0f-a8db-8702c047d2af · outbound

This paper cites Panns: Large- scale pretrained audio neural networks for audio pat- tern recognition,.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Panns: Large- scale pretrained audio neural networks for audio pat- tern recognition,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:43.046961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:36:42.566111Z digest=sha256:f86fce747686ec5b43a08fd2ca1b60045c93bf0cad3c969f7179f815792453c9

Observation 653a13bf-c98b-4d2a-83cf-051647f1e180 · outbound

This paper cites Film: Visual reasoning with a general conditioning layer,.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Film: Visual reasoning with a general conditioning layer,

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:42.569655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:42.569655Z digest=sha256:1cbd0b644e4f0f7217c3925abe3b8dcf28bba73f03566890b0fa1c9eaefc5a76

Observation da8bf902-2b3d-4bbe-9776-dc74b80f2ac7 · outbound

This paper cites Simple and controllable music generation,.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Simple and controllable music generation,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:43.028037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:36:42.573729Z digest=sha256:8d578a79572b1aadda99dd52eec9a42ca222b15acf32c5d4853ea0cea3a396ce

Observation 3ce23509-f9dd-4bc1-aec2-972436a940c0 · outbound

This paper cites Learning transferable visual models from natural lan- guage supervision,.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Learning transferable visual models from natural lan- guage supervision,

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:42.578021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:42.578021Z digest=sha256:f1a8a85cef4b9cc69b55410cc952864b18fd668e677242a3f88cfe2839128c53

Observation 3eee001c-eab3-4d51-b092-e28182dd0b42 · outbound

This paper cites Blip-2: Bootstrapping language-image pre- training with frozen image encoders and large language models,.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Blip-2: Bootstrapping language-image pre- training with frozen image encoders and large language models,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:43.008983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:36:42.581737Z digest=sha256:16055d7e6f64b1cc7ed1d97571c786dade8a173ff46aa09a4f6d93aa6232d549

Observation 8dba42b6-6f13-4d4f-ac83-62e87cb04ec0 · outbound

This paper cites Stable audio open,.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Stable audio open,

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:42.585657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:42.585657Z digest=sha256:41fdbb459417114cd1598b70b23ae4371843994b8cc2ab23e1b0fb6bb726bfe5

Observation 845fcd3d-3188-46eb-b21f-4d82c9bd4ff7 · outbound

This paper cites Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:42.589124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:42.589124Z digest=sha256:eb4175a71ca90eb474134ac7ca98ed08d4a84ab0c192e4727faacd2dd32b684e

Observation 033875e8-7627-4d8c-9f7d-d2ee6ffd8ca0 · outbound

This paper cites Reliable fidelity and diversity metrics for generative models,.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Reliable fidelity and diversity metrics for generative models,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:42.986288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:36:42.592522Z digest=sha256:9e6743aa598a5736b575b0b5543bc405a5c3ffbf78d58cfe768b4431f6623e33

Observation 94ed3f79-cf9d-4933-9880-15011d8c22bb · outbound

This paper cites Large-scale contrastive language-audio pretraining with feature fu- sion and keyword-to-caption augmentation,.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Large-scale contrastive language-audio pretraining with feature fu- sion and keyword-to-caption augmentation,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:42.972018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:36:42.595514Z digest=sha256:9587c3215185b96ab34fbe0379d56b893ddb81dccf411bbcf6757efb0033b5d0

Observation 8a87dcdc-06b6-4acc-8217-f75e2e86a756 · outbound

This paper cites Efficient Training of Audio Transformers with Patchout.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Efficient Training of Audio Transformers with Patchout

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:42.599014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:42.599014Z digest=sha256:b3aefd6757c939b3f7b7b391ba113625c44762c3b34e245a71fbe85c580cb1d5

Pith citing papers

Observation 9994dfab-e618-4b92-91d6-f403c9decf4b · inbound

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections cites this paper.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-16T00:36:42.957902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:36:42.391185Z digest=sha256:1426e182a8ee75e4d49fdbb8ada0bce9fb5641c09e42d171ee8b30fc834332a8