Pith. sign in

Paper Citation Record · LEDGER

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections

As of 18 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 1 inbound Pith citation observation for arXiv:2608.11576.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.11576 v1

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:36:42.599014Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:36:42.391185Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-16T00:36:42.953731Z

Reference resolution

54 of 54 outbound references displayed

  • verified exact6
  • verified fuzzy29
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9994dfab-e618-4b92-91d6-f403c9decf4b · outbound

This paper cites Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-16T00:36:42.957902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:36:42.391185Z digest=sha256:785427c40ff0d74f5f09e29652eed84786ca0d100fedde1adeaa71a6b89cb2e2

Observation 49f7049e-e796-46e3-9af6-b670310f557f · outbound

This paper cites Early work in modern video-to-music generation systems uses large-scale web music-video corpora and autoregressive modeling over semantic acoustic tokens to generate music [8].

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Early work in modern video-to-music generation systems uses large-scale web music-video corpora and autoregressive modeling over semantic acoustic tokens to generate music [8]

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:43.330732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:36:42.395669Z digest=sha256:e6e0a6c56d7fd284d2bcb0ac9805c9f9973c6e83a09a2272a27d69f0a4bdfab5

Observation d1f0f9bd-1d9e-4177-aebe-5ce875711696 · outbound

This paper cites trance music.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections trance music

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:43.319454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:36:42.399453Z digest=sha256:1a8ac734d46c4e9df0321d6156f8d82af464f7ab917a9ab51bc5feab1b4ba392

Observation f0743c4e-cba8-4a45-915f-35bb08426abd · outbound

This paper cites an unresolved cited work.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-16T00:36:43.308324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:36:42.403577Z digest=sha256:541dc0ccc9a3b9f8f68dcba6391f861ae7eb4f10e9802e5ec65fc32cfdf0347d

Observation 1369da0c-923d-4a28-b521-da386bbf7aa4 · outbound

This paper cites First, we bench- mark video-to-music generation tasks by comparing existing video-to-music generation models trained and evaluated on identical data, using OSSL-v2.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections First, we bench- mark video-to-music generation tasks by comparing existing video-to-music generation models trained and evaluated on identical data, using OSSL-v2

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:43.298362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:36:42.408016Z digest=sha256:7f4247d45a35a3a83ad5690868cc9882787d53fae4ac80496a37e86ffcee16a8

Observation 442247b2-c215-4420-b9d4-a54b69ada218 · outbound

This paper cites +Dialogue.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections +Dialogue

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:43.287922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:36:42.412530Z digest=sha256:002e624242f9358e4db75a2dfad67aa3ee26c5d8c436c8cbb25fda6bece2a41a

Observation 4d97db7e-2296-4f1a-a2c9-6fff5d6fbc44 · outbound

This paper cites Be- cause the dataset is free from link rot and does not require separate web scraping, our dataset is suitable as a durable benchmark for the field.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Be- cause the dataset is free from link rot and does not require separate web scraping, our dataset is suitable as a durable benchmark for the field

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:43.275395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:36:42.416447Z digest=sha256:2f69702cf7b4ab9b96c1cc587ed237aea13e847037c626884a706d7b22345f17

Observation df5e0c5b-906b-47e3-959a-cc4b271b19e6 · outbound

This paper cites Teaser Generation for Long Documentaries and Educational Videos.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Teaser Generation for Long Documentaries and Educational Videos

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:43.264105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:36:42.420783Z digest=sha256:60b08874573f6f6704da8670aa4eafecdbd0dab550f9aed201c707817ce6cf75

Observation 48d25803-3e24-4035-bc77-87021a8676c0 · outbound

This paper cites Attendaffectnet–emotion pre- diction of movie viewers using multimodal fusion with self-attention,.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Attendaffectnet–emotion pre- diction of movie viewers using multimodal fusion with self-attention,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:43.252296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:36:42.424057Z digest=sha256:5757369bbf55f06a2419adb69c2162c0ea606535670fb58dca3f38541ef5d875

Observation cefa76f8-82df-4414-a38f-80e911720855 · outbound

This paper cites Predicting emotion from music videos: exploring the relative contribution of visual and auditory information to affective responses.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Predicting emotion from music videos: exploring the relative contribution of visual and auditory information to affective responses

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:42.427623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:42.427623Z digest=sha256:a1b7883ed0a8b4f0949fb56a3d66a3f414d810868d762162558f4c00307130fe

Observation aefa4075-821b-4a94-9155-3c07c6c43813 · outbound

This paper cites The cognitive processing of film and musical soundtracks,.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections The cognitive processing of film and musical soundtracks,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:43.240784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:36:42.431355Z digest=sha256:67a00dc333f2d911208d7a0d11ced33a9a714af5bbbc310805d75bedafa416e2

Observation 008a119a-ccf9-473b-b194-b9633b4b2f70 · outbound

This paper cites Multimodal deep models for predicting affec- tive responses evoked by movies.,.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Multimodal deep models for predicting affec- tive responses evoked by movies.,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:43.229494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:36:42.434869Z digest=sha256:04286ad1d8afafc8fd8745d396d8f2ef34635cc5c955b5466a5c8f176bafb167

Observation fa33de3d-8bb4-4105-b5cc-52ed53cdaad7 · outbound

This paper cites On music’s potential to convey meaning in film: A systematic review of empirical evi- dence,.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections On music’s potential to convey meaning in film: A systematic review of empirical evi- dence,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:43.219280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:36:42.438501Z digest=sha256:2cd6c7861091336300a851dc51a3c19956e181342a4569ae9915513fcf81e09a

Observation e4bb1dcc-a488-4e3f-b982-43011b6020d4 · outbound

This paper cites Emotion Embedding Spaces for Matching Music to Stories.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Emotion Embedding Spaces for Matching Music to Stories

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:42.442268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:42.442268Z digest=sha256:eac27408657fac455cdf4eee29b32c0dd69ec70530885cd10a90d32cbe81e1b9

Observation b24fc0fe-dfd4-463c-96d2-b5d1293e74d7 · outbound

This paper cites Foley music: Learning to generate music from videos,.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Foley music: Learning to generate music from videos,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:43.208400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:36:42.446296Z digest=sha256:647d686101ec710c610f633ba3121bb7a88c0a70af788b7703ddf51a269c80f4

Observation 4b2cefab-443c-4b9e-8172-064e5352d578 · outbound

This paper cites V2meow: Meow- ing to the visual beat via video-to-music generation,.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections V2meow: Meow- ing to the visual beat via video-to-music generation,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:43.197380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:36:42.450042Z digest=sha256:07ad7930011a9029b2ef478204a5f2c16ef9416c64ab38be7ad6953e2e4ae2eb

Observation fd6a3ed6-191b-4b45-8089-1bf9bbd4218a · outbound

This paper cites Vidmuse: A simple video-to-music generation framework with long-short-term modeling,.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Vidmuse: A simple video-to-music generation framework with long-short-term modeling,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:43.185861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:36:42.453909Z digest=sha256:04aa458eb6096dd43bd050c487091cb05320c1fabe3f87265d3e00a5edc7fc3b

Observation 1ceb4568-70ff-40a4-9f39-ff64d417ccc9 · outbound

This paper cites Vmas: Video-to-music generation via semantic alignment in web music videos,.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Vmas: Video-to-music generation via semantic alignment in web music videos,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:43.174223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:36:42.457562Z digest=sha256:54be052a8e24de4e7cb37919a1e1ce229f1944b8e998972bf8e6f3cef266851b

Observation 0a616f9c-5105-4098-9803-76aed2621f3b · outbound

This paper cites Sonique: Video background music generation using unpaired audio- visual data,.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Sonique: Video background music generation using unpaired audio- visual data,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:43.163597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:36:42.461759Z digest=sha256:589436022a0b9ec6a25239e48c28ef61c98b2d13ffd6eed90889acb1e7b9c23f

Observation 707f5136-b83a-41bc-bac4-8007b047d9ec · outbound

This paper cites GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:42.465952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:42.465952Z digest=sha256:64919e3958007a062712020fe2593ab3ce230a81f69306fa7c1cf713aaa51343

Observation 19f194b0-b420-4d16-8e89-42189f1c22a4 · outbound

This paper cites Vision-to-Music Generation: A Survey.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Vision-to-Music Generation: A Survey

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:42.470428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:42.470428Z digest=sha256:a6c171a0339b958407a35f83f69327786eb025bd813a908f0c3d3cf62742324e

Observation a329da2a-9f37-45c7-a4d3-bad6ee9ea753 · outbound

This paper cites Ai- based chinese-style music generation from video con- tent: a study on cross-modal analysis and generation methods,.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Ai- based chinese-style music generation from video con- tent: a study on cross-modal analysis and generation methods,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:43.151434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:36:42.474275Z digest=sha256:e4ae4f8730f3a802426326cc67aa94250a65eb854b2b65ce2dc4000eb003eaf0

Observation 81b7990b-37ce-403d-83d9-2d9c84b6d211 · outbound

This paper cites V2M-Zero: Zero-Pair Time-Aligned Video-to-Music Generation.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections V2M-Zero: Zero-Pair Time-Aligned Video-to-Music Generation

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-08-16T00:36:42.893043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:36:42.478417Z digest=sha256:0a26d8730928d29e5adbc63520542f87b36c8359b841273889e56471ea3bf37d

Observation aa944ece-843f-4cd3-9506-b9cc63eb9cc7 · outbound

This paper cites Diff- v2m: A hierarchical conditional diffusion model with explicit rhythmic modeling for video-to-music genera- tion,.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Diff- v2m: A hierarchical conditional diffusion model with explicit rhythmic modeling for video-to-music genera- tion,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:43.139456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:36:42.482121Z digest=sha256:6e7b33dc31792d5428cf6674bbb208a7fef26327f167f13bd3798e7fac175594

Observation 10022ec1-ca19-4aa1-bf0f-aad20ccbb1fe · outbound

This paper cites Quan- tized gan for complex music generation from dance videos,.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Quan- tized gan for complex music generation from dance videos,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:43.128702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:36:42.489105Z digest=sha256:ca00aca63d01bfd52cef1ae295862b40a54e879505f27668644dcf7ff70f9ddd

Observation b2d7c4b9-30bd-4184-86b6-a939bf58ac3e · outbound

This paper cites Video background music genera- tion: Dataset, method and evaluation,.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Video background music genera- tion: Dataset, method and evaluation,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:43.117905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:36:42.492941Z digest=sha256:5031433ff8f5563365a19a62bd7a2cf47657b3b25acfd3180de6037b5499c6cd

Observation 990607ec-3e49-4516-874e-35c7c1f52409 · outbound

This paper cites Content-Based Video-Music Retrieval Using Soft Intra-Modal Structure Constraint.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Content-Based Video-Music Retrieval Using Soft Intra-Modal Structure Constraint

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:42.496612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:42.496612Z digest=sha256:873756689855aef556c1c2cc6a0c596af7057824a30f109ab098d25d66ffed0c

Observation 8147f160-bd41-4192-82f0-c30b73745ab6 · outbound

This paper cites Video-Guided Text-to-Music Generation Using Public Domain Movie Collections.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Video-Guided Text-to-Music Generation Using Public Domain Movie Collections

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:42.500746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:42.500746Z digest=sha256:718de42c84ef1d5b8f244b3def6445eac6eaab8077f9baeb6e1310ccc36a5999

Observation fa6e8aab-9e17-454e-b054-7b958e060045 · outbound

This paper cites Acoustic profiles in vocal emotion expression.,.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Acoustic profiles in vocal emotion expression.,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:43.106258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:36:42.504700Z digest=sha256:b389ab81b62468d8b98df49b0f1b314af2d87d3b45c6b968bdfc1999603a5e16

Observation 6abfe8f2-f194-4e2c-84c1-41efdb5a99a1 · outbound

This paper cites Background ducking to produce esthetically pleasing audio for tv with clear speech,.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Background ducking to produce esthetically pleasing audio for tv with clear speech,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:43.092776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:36:42.508491Z digest=sha256:060513f14a59fccee1656b0325807931781dd15003255f37dd07d4422dc5d679

Observation bee2f69a-b172-452f-ae3e-7047d87f4192 · outbound

This paper cites Improving dialogue intelligi- bility in streaming media,.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Improving dialogue intelligi- bility in streaming media,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:43.080172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:36:42.512290Z digest=sha256:fd33399da319ab90a7813c69fc28951315a8d591e3a30725048fedcdc33c276d

Observation 120bc34a-93c5-4792-90e6-0256dc84d0c3 · outbound

This paper cites Creating a multitrack clas- sical music performance dataset for multimodal mu- sic analysis: Challenges, insights, and applications,.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Creating a multitrack clas- sical music performance dataset for multimodal mu- sic analysis: Challenges, insights, and applications,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:43.068775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:36:42.517678Z digest=sha256:bd472f759a14ee826f32f1a77576be5a36b49a75b0190d7853fcbe6883573912

Observation ebc889d5-7b93-4f4d-9a5e-3aeb1039d244 · outbound

This paper cites Video2music: Suitable music generation from videos using an affective multimodal transformer model,.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Video2music: Suitable music generation from videos using an affective multimodal transformer model,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:43.058045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:36:42.521020Z digest=sha256:8b2d8c95245b75b8c5e14988476c87304cd867df488b1142043732d47d6f5749

Observation e1e7e50f-29fd-4335-a102-3ddf5172aec5 · outbound

This paper cites Extending Visual Dynamics for Video-to-Music Generation.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Extending Visual Dynamics for Video-to-Music Generation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:42.524714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:42.524714Z digest=sha256:9257be0892a8fa28e7bc275d7c8e9f988f46d03b3f2f991561bb3471b33f0620

Observation bd94a98a-47b9-4228-ba40-0b7c2eac11fb · outbound

This paper cites VidMusician: Video-to-Music Generation with Semantic-Rhythmic Alignment via Hierarchical Visual Features.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections VidMusician: Video-to-Music Generation with Semantic-Rhythmic Alignment via Hierarchical Visual Features

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:42.528088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:42.528088Z digest=sha256:e80c3851b1790e4068595a73b36d1c033d7c6a02787e8137c3ccead8e65b719f

Observation f06b666f-bace-4d58-a581-5d524fbabd79 · outbound

This paper cites MuVi: Video-to-Music Generation with Semantic Alignment and Rhythmic Synchronization.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections MuVi: Video-to-Music Generation with Semantic Alignment and Rhythmic Synchronization

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:42.531770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:42.531770Z digest=sha256:2f4d8d5b687f7828517f6217fdd90fe73b2edc5d3946142ecb5535e66d1ddd0c

Observation 6e1b4654-799c-4717-9f79-8ce89d3db55f · outbound

This paper cites Video echoed in music: Semantic, temporal, and rhythmic alignment for video-to-music generation,.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Video echoed in music: Semantic, temporal, and rhythmic alignment for video-to-music generation,

Reference 38

Resolution
verified exact
raw_fallback, observed 2026-08-16T00:36:42.818139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:36:42.535975Z digest=sha256:fa909c952ae754b3d92ff91950a7430aa36da4147dcd574ab0e9163d36813a77

Observation 7c0c2411-6d68-43cc-8757-d2ece80bb540 · outbound

This paper cites Video-Robin: Autoregressive Diffusion Planning for Intent-Grounded Video-to-Music Generation.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Video-Robin: Autoregressive Diffusion Planning for Intent-Grounded Video-to-Music Generation

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-08-16T00:36:42.743314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:36:42.539678Z digest=sha256:6ca5692d796d0c772be55336967f8ee7ae327e42fd3ed802571c4c7496565ca7

Observation 21d31ded-83a0-4005-be0e-91e719cd9342 · outbound

This paper cites Harmonizing Pixels and Melodies: Maestro-Guided Film Score Generation and Composition Style Transfer.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Harmonizing Pixels and Melodies: Maestro-Guided Film Score Generation and Composition Style Transfer

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-08-16T00:36:42.726835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:36:42.544084Z digest=sha256:859058ad5ed196fc0c78549034311d2fb87431199de8876594dfce45f2296a5a

Observation 1209e8e0-5c05-4638-a493-3f2d2ff08163 · outbound

This paper cites FilmComposer: LLM-Driven Music Production for Silent Film Clips.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections FilmComposer: LLM-Driven Music Production for Silent Film Clips

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:42.548207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:42.548207Z digest=sha256:44e29bdfed7cc3efcba400c24d9c2c406e6ab9ddfcf4750c591e8513ff543ed6

Observation 72562688-b4de-4467-b3ec-8b104b63b241 · outbound

This paper cites M$^{2}$UGen: Multi-modal Music Understanding and Generation with the Power of Large Language Models.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections M$^{2}$UGen: Multi-modal Music Understanding and Generation with the Power of Large Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:42.551715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:42.551715Z digest=sha256:57fd22086f1e749344d8c675805dbf422d0af545540977215de1a60a5deab1fb

Observation 16bf29af-c87e-4b33-9971-4c33d6f8d77d · outbound

This paper cites MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:42.555094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:42.555094Z digest=sha256:6632bb9503c30dac55d3aa65b85e4b80f4c773cf1c23569c3edfafe550b2ae1e

Observation 8d8c3492-5c3a-4909-9fa0-35fc7b2ed464 · outbound

This paper cites Multimodal Music Generation with Explicit Bridges and Retrieval Augmentation.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Multimodal Music Generation with Explicit Bridges and Retrieval Augmentation

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:42.558498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:42.558498Z digest=sha256:78ccbbbbbd19b3c07f2b66b49e9549db925d2d7f3fa75447f80d78de52605986

Observation b074ce61-01be-4387-8b04-d49dfe2f24fd · outbound

This paper cites Benchmarks and leaderboards for sound demixing tasks.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Benchmarks and leaderboards for sound demixing tasks

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-08-16T00:36:42.661931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:36:42.562317Z digest=sha256:0d2a43bfe41309744a9de3067de1ef495639d5521e1c4931f1394c5c99c96b7c

Observation eb3d6160-8e0d-4e0f-a8db-8702c047d2af · outbound

This paper cites Panns: Large- scale pretrained audio neural networks for audio pat- tern recognition,.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Panns: Large- scale pretrained audio neural networks for audio pat- tern recognition,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:43.046961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:36:42.566111Z digest=sha256:18ee06826402f3b6ad69486ea990634e08773976e0bff6a3b7161bfeafd3aa5e

Observation 653a13bf-c98b-4d2a-83cf-051647f1e180 · outbound

This paper cites Film: Visual reasoning with a general conditioning layer,.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Film: Visual reasoning with a general conditioning layer,

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:42.569655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:42.569655Z digest=sha256:63a1d454b909e7eaf1f86890d6b10e7b87e587c1c5271779d3fd2b6077b825fb

Observation da8bf902-2b3d-4bbe-9776-dc74b80f2ac7 · outbound

This paper cites Simple and controllable music generation,.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Simple and controllable music generation,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:43.028037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:36:42.573729Z digest=sha256:c629fde88eafb06b290c21b149110bb69f54487df4a6465ba8010cfeee2d42ad

Observation 3ce23509-f9dd-4bc1-aec2-972436a940c0 · outbound

This paper cites Learning transferable visual models from natural lan- guage supervision,.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Learning transferable visual models from natural lan- guage supervision,

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:42.578021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:42.578021Z digest=sha256:2afff8dc130bbaa85673d21d44241de1cb9aef02a8feba2f26020edebe6cca1e

Observation 3eee001c-eab3-4d51-b092-e28182dd0b42 · outbound

This paper cites Blip-2: Bootstrapping language-image pre- training with frozen image encoders and large language models,.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Blip-2: Bootstrapping language-image pre- training with frozen image encoders and large language models,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:43.008983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:36:42.581737Z digest=sha256:3e12137fa2c908a834dd142e3f67c445c39db5eef224b85a662f27932b801623

Observation 8dba42b6-6f13-4d4f-ac83-62e87cb04ec0 · outbound

This paper cites Stable audio open,.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Stable audio open,

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:42.585657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:42.585657Z digest=sha256:371dab7868ff23db3ec58aa35a55dc6ba6a0881c890f34161a3fa2f3d32e01fc

Observation 845fcd3d-3188-46eb-b21f-4d82c9bd4ff7 · outbound

This paper cites Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:42.589124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:42.589124Z digest=sha256:4917632f2b6c79b1210ad417eae0bc629b9f361ec7cd52c8480dfef0a12b6b71

Observation 033875e8-7627-4d8c-9f7d-d2ee6ffd8ca0 · outbound

This paper cites Reliable fidelity and diversity metrics for generative models,.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Reliable fidelity and diversity metrics for generative models,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:42.986288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:36:42.592522Z digest=sha256:54da586949757ac8b67e9c0e976ccbaf920bb5e225a000fe9627b9d7738aff0e

Observation 94ed3f79-cf9d-4933-9880-15011d8c22bb · outbound

This paper cites Large-scale contrastive language-audio pretraining with feature fu- sion and keyword-to-caption augmentation,.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Large-scale contrastive language-audio pretraining with feature fu- sion and keyword-to-caption augmentation,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:42.972018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:36:42.595514Z digest=sha256:e47f4551e2381b61a4966eff4bda9f9290ff616f6f16472b6e4547015c2e75d2

Observation 8a87dcdc-06b6-4acc-8217-f75e2e86a756 · outbound

This paper cites Efficient Training of Audio Transformers with Patchout.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Efficient Training of Audio Transformers with Patchout

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:42.599014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:42.599014Z digest=sha256:669ba9ea3ee2bd6c37d1853b3afa3b4c276db24a4f346a80dcd30e0d7a3b0c7b

Pith citing papers

Observation 9994dfab-e618-4b92-91d6-f403c9decf4b · inbound

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections cites this paper.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-16T00:36:42.957902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:36:42.391185Z digest=sha256:785427c40ff0d74f5f09e29652eed84786ca0d100fedde1adeaa71a6b89cb2e2