Pith. sign in

Paper Citation Record · LEDGER

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections

As of 7 August 2026, this Paper Citation Record lists 84 of 84 outbound references and 1 inbound Pith citation observation for arXiv:2506.12573.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.12573 v3

Coverage vector

measured 84 of 84 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:49:47.983763Z

measured 85 of 85 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:49:40.786176Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T00:49:50.392438Z

Reference resolution

84 of 84 outbound references displayed

  • verified exact6
  • verified fuzzy34
  • unresolved41
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5af4092a-926b-44ce-b64d-34ac5cdc75e5 · outbound

This paper cites Video-Guided Text-to-Music Generation Using Public Domain Movie Collections.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Video-Guided Text-to-Music Generation Using Public Domain Movie Collections

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:49:50.536973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:49:40.786176Z digest=sha256:d390917fea2aabdf690014473a223528cc020313cbc846110b20ee6863d898f7

Observation ad5fa3ec-dacf-498f-bb97-38acd2d4cd7b · outbound

This paper cites an unresolved cited work.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:40.864936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:40.864936Z digest=sha256:ad0d29580a4b09a8276db3ea793acfc779f4d770a57132a3fd265a855c2ec274

Observation 5a892216-9ee5-4236-be83-f39bfade8dd6 · outbound

This paper cites We provide an overview of the comparison of video-music datasets in Table 1 and an illustration of our dataset con- struction methodology in Figure 1.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections We provide an overview of the comparison of video-music datasets in Table 1 and an illustration of our dataset con- struction methodology in Figure 1

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:40.950159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:40.950159Z digest=sha256:6c6378c4b9ad00ae4d990babf32c3ad9ca0c939cb19e345a116ab6bae0af72c9

Observation 96cf67e9-1b5d-4273-8464-6b780233dcf4 · outbound

This paper cites an unresolved cited work.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:41.045157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:41.045157Z digest=sha256:5f6edb1f6e1dd5d7e3b356e226bd026a0aabbe965543b0dba8cc4c7f1c8ddd4a

Observation 2749241b-9bb3-4ded-8c1c-40b2dd661c27 · outbound

This paper cites S” and “M.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections S” and “M

Reference 5

Resolution
verified exact
raw_fallback, observed 2026-08-07T00:49:50.232025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:49:41.120789Z digest=sha256:62c23823f31fe1691cd6085dab9f4c858220dab9fbdd5ebf00244c997bee2cce

Observation 2ccb242a-283f-46be-99c7-458adbebe840 · outbound

This paper cites Distributional FidelityOur evaluation on OES-Com re- veals that fine-tuning on OSSL enhances distributional fi- delity.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Distributional FidelityOur evaluation on OES-Com re- veals that fine-tuning on OSSL enhances distributional fi- delity

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:41.198985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:41.198985Z digest=sha256:3238056c5776a9db5b80a67e1641fb9ac755e4116d3ff51bdcb71f8f1d0719f8

Observation 330770bb-000c-4f02-bfd3-07d4e36c92cf · outbound

This paper cites To show the effectiveness of our dataset, we adapted a text-to-music generation model with video conditions and fine-tuned it on our dataset.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections To show the effectiveness of our dataset, we adapted a text-to-music generation model with video conditions and fine-tuned it on our dataset

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:57.630279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:49:41.272353Z digest=sha256:17218430e7cfc183d6045b6877736db69b793890fca26ded90ce320641631431

Observation aa0b77f5-740f-4a30-94a9-5bef220bc6b6 · outbound

This paper cites an unresolved cited work.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:49:57.502642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:49:41.339642Z digest=sha256:34411b55dfbff2bd9342a7cc2f3e49813e1c09440fd88ef5839f22657eb14fad

Observation a2b075f8-ad4d-41ad-9f10-d69886d45ea0 · outbound

This paper cites Soundtrack design: The impact of music on visual attention and affective responses,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Soundtrack design: The impact of music on visual attention and affective responses,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:57.424060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:49:41.417326Z digest=sha256:52c994ac5c18102a6a76387151bf8f73f4eb6f6cd6b8a5c646266f2add0b163e

Observation 931c243d-a0be-450f-8fe9-b6abb48a5c18 · outbound

This paper cites Multimodal deep models for predicting affective responses evoked by movies.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Multimodal deep models for predicting affective responses evoked by movies

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:57.347541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:49:41.494318Z digest=sha256:d81de82f76763a002dee53044715cfd751eab89c8a3dc07c76f4d15513779808

Observation 65756565-9c31-4f5b-acc1-9927be466874 · outbound

This paper cites Emotion Embedding Spaces for Matching Music to Stories.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Emotion Embedding Spaces for Matching Music to Stories

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:49:49.863527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:49:41.584055Z digest=sha256:72ffa8d16f91f0f7142ded84494e70406690fec4a531b6a0e23c504c6575fa41

Observation 2d14e174-ec1a-42a6-8b9b-228347da9a6f · outbound

This paper cites Attendaffectnet–emotion prediction of movie viewers using multimodal fusion with self-attention,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Attendaffectnet–emotion prediction of movie viewers using multimodal fusion with self-attention,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:57.192060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:49:41.662550Z digest=sha256:f8bd7a90e49d139979fd628acd03b02daeefbd077e01e25fc526342851e28ca2

Observation 24dc64c5-acca-4af7-92cb-6137c395ebaa · outbound

This paper cites Predicting emotion from music videos: exploring the relative contribution of visual and auditory information to affective responses.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Predicting emotion from music videos: exploring the relative contribution of visual and auditory information to affective responses

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:49:49.551873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:49:41.739826Z digest=sha256:58dbc289118ee02985ba99dda3f595a0833638c3f0d3464f16f8e705197496e2

Observation 4b747be7-3c98-4c30-ba78-447e0aa5feb9 · outbound

This paper cites Analysis of the roles of film soundtracks in films,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Analysis of the roles of film soundtracks in films,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:57.116682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:49:41.820319Z digest=sha256:f8fcf71f9adc98b56957b71ab2df9756855f0cb12b8a5946f9be573d969fc02f

Observation 7f35f1ab-b34e-4616-b036-c2cfcc4bb411 · outbound

This paper cites Actions in context,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Actions in context,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:56.941092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:49:41.884556Z digest=sha256:30d1b6ca789e3abae792d16f4a6b532e36d6859d19df03776ac1437b25f876d6

Observation 57e6cdd2-bd3f-4ded-a1b7-431a494383cb · outbound

This paper cites Movieqa: Understanding stories in movies through question-answering,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Movieqa: Understanding stories in movies through question-answering,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:56.692466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:49:41.963832Z digest=sha256:4d1bd22e31858df0a5a9976843a6ecafe61c7c0da218aad638187a8a4c5dde7a

Observation 178f64c9-084c-4961-8dd7-9428a98bcc0f · outbound

This paper cites Movienet: A holistic dataset for movie understand- ing,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Movienet: A holistic dataset for movie understand- ing,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:56.534391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:49:42.035597Z digest=sha256:8655994fda92e10eb90b4a9ebb37ac9a7ebcbaf88bb41f238e388f6c14b4de11

Observation 8255ea0f-8c7d-4a2e-a083-578fb835f1e5 · outbound

This paper cites Mad: A scalable dataset for language grounding in videos from movie audio descriptions,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Mad: A scalable dataset for language grounding in videos from movie audio descriptions,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:56.367972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:49:42.103527Z digest=sha256:231124a8befcb81917f7fa98df5b46ee14019c27faaa4b420dfa842604f5001e

Observation 9f955c11-efdf-4d39-8a82-67e5f90dbc62 · outbound

This paper cites A dataset for movie description,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections A dataset for movie description,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:56.278451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:49:42.197524Z digest=sha256:7653b765c927a0d8b1d2db4b4aff0adb2589047fbb36875935272ae37681703d

Observation f61528fb-e050-4ddf-bbf7-9f2c67e5fc0b · outbound

This paper cites Moviegraphs: Towards understanding human-centric situations from videos,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Moviegraphs: Towards understanding human-centric situations from videos,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:56.112885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:49:42.279348Z digest=sha256:de88cbaf4c4ad6c9eec4cb06dde18903c32a718887e944ea694e6dbc3b42f9eb

Observation 9086561a-5ce8-448d-84b3-30760610658f · outbound

This paper cites Hlvu: A new challenge to test deep understanding of movies the way humans do,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Hlvu: A new challenge to test deep understanding of movies the way humans do,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:56.021400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:49:42.366444Z digest=sha256:bfcd8657d5c5c62213936b79a3e1aa22871139854b0a870327b70e5deabf62b0

Observation 3237553e-2891-420a-abb0-136945d11bf6 · outbound

This paper cites Condensed movies: Story based retrieval with contex- tual embeddings,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Condensed movies: Story based retrieval with contex- tual embeddings,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:55.920877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:49:42.445904Z digest=sha256:0ff908dc8336230901c1e975a04c7ba40527c13ecd86195991774b3db6e2187a

Observation 4799cefd-2193-40fc-a7d3-0a70dd655e0d · outbound

This paper cites TeaserGen: Generating Teasers for Long Documentaries.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections TeaserGen: Generating Teasers for Long Documentaries

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:42.511515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:42.511515Z digest=sha256:c17b0d0ae4c54049501be57534ea34a3d069c8aab7ee121c27659cf52b50d942

Observation a9178205-829d-4963-b251-0bb2ad2877d2 · outbound

This paper cites Simple and control- lable music generation,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Simple and control- lable music generation,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:55.764831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:49:42.599104Z digest=sha256:26a4fd6ded3c7a0a14773a9f415d7b9932ff34dd3c95e139f1f4b601865616c9

Observation f1a5fa92-d749-4ce4-9143-a20f91670f51 · outbound

This paper cites Attention is all you need,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Attention is all you need,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:42.675447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:42.675447Z digest=sha256:096d00e697904e04d14c3a1f1a3a50d6ed7de9b3adefb95f67e20720f8c8ab99

Observation f47ca879-4b18-487b-9530-1e89f691e266 · outbound

This paper cites MusicLM: Generating Music From Text.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections MusicLM: Generating Music From Text

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:42.750717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:42.750717Z digest=sha256:e309c428f5f5fa144230c8f6f99c21dc0878e826471f0030864e5676b145091f

Observation 744cb494-a203-4312-a521-fafac7c164da · outbound

This paper cites MusiConGen: Rhythm and Chord Control for Transformer-Based Text-to-Music Generation.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections MusiConGen: Rhythm and Chord Control for Transformer-Based Text-to-Music Generation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:42.827934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:42.827934Z digest=sha256:438cd6cf2bc631e5bd5594f4b7d8ae5d42c4349c62599cd72b933c101bc4764e

Observation 1450699d-fda3-4db1-be6d-12e8ad47a34d · outbound

This paper cites VidMuse: A Simple Video-to-Music Generation Framework with Long-Short-Term Modeling.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections VidMuse: A Simple Video-to-Music Generation Framework with Long-Short-Term Modeling

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:42.900461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:42.900461Z digest=sha256:84d6f92e80560167e8a9cdbb89d05a6ebd8a403a170d7868a3f4d5fd49bf4473

Observation f33b4d9b-4689-465d-bb03-a5353cfe780b · outbound

This paper cites V2meow: Meowing to the visual beat via video-to- music generation,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections V2meow: Meowing to the visual beat via video-to- music generation,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:55.665078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:49:42.977529Z digest=sha256:749cbe1abce6ce0e3cf52b3187781d58fe60c2b32f6d25fe1d53200e67c466db

Observation ab03e73b-0eb7-424e-86c1-3955da2955bd · outbound

This paper cites GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:49:49.247373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:49:43.064927Z digest=sha256:d7bc9fed96f3f481eae44b8da07f6ef574085cd80b36b8d90cbca0c70c2384bf

Observation 7f8710a4-62bd-4f3e-ab57-14e14f6e0164 · outbound

This paper cites Riffusion-stable diffusion for real-time music generation,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Riffusion-stable diffusion for real-time music generation,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:55.456481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:49:43.156770Z digest=sha256:4e5ddc346a7484eb3960d8af35b622349cc6f94590b123125b565ae6597e20d0

Observation 92b97fe1-1c0a-456e-a812-f804e48e40f9 · outbound

This paper cites Noise2Music: Text-conditioned Music Generation with Diffusion Models.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Noise2Music: Text-conditioned Music Generation with Diffusion Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:43.230631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:43.230631Z digest=sha256:b9fca492a11ba6044d5b47414b187cb9ca22c828aca556d7da1b3f959c83f5e4

Observation ff44a7fe-ccb8-437f-a06d-699d6dba7156 · outbound

This paper cites Mo\^usai: Text-to-Music Generation with Long-Context Latent Diffusion.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Mo\^usai: Text-to-Music Generation with Long-Context Latent Diffusion

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:43.294050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:43.294050Z digest=sha256:69da05a9a0274882a1d0155db53c24e678e39708c9c69eac006d66bc6ffcf8ab

Observation 41e83f31-e1a9-4b62-9f2b-e0d367cf7fda · outbound

This paper cites Efficient neu- ral music generation,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Efficient neu- ral music generation,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:55.286800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:49:43.347990Z digest=sha256:82064000e3a522b4b80984e13e324086c493203d971cb791b72debee708d4e69

Observation 712b6267-1f37-4ea1-94e7-9cfb31d71bfd · outbound

This paper cites Mustango: Toward Controllable Text-to-Music Generation.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Mustango: Toward Controllable Text-to-Music Generation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:43.437966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:43.437966Z digest=sha256:b7f038926fe9e870c1697b02424432dde7864d58325bcab7e7c8188a267c2d68

Observation c168ce5b-df36-4fc9-a7bb-213c000ebc30 · outbound

This paper cites Multi-Track MusicLDM: Towards Versatile Music Generation with Latent Diffusion Model.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Multi-Track MusicLDM: Towards Versatile Music Generation with Latent Diffusion Model

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:49:48.971726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:49:43.507489Z digest=sha256:dcea6f62ffc6ff7054973214797d9b150449b08ab7d84c10d304f7848c466179

Observation 7e3009df-3e6b-43a6-a441-32a41699f533 · outbound

This paper cites Stable Audio Open.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Stable Audio Open

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:43.573592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:43.573592Z digest=sha256:e0b4cdeef4798aed5c63a0c165fe7c69172bd3bbfd2637f27168eaa9ca50295d

Observation d9d5c618-8e40-4ceb-a172-1711575d2e26 · outbound

This paper cites VMAS: Video-to-Music Generation via Semantic Alignment in Web Music Videos.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections VMAS: Video-to-Music Generation via Semantic Alignment in Web Music Videos

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:43.639290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:43.639290Z digest=sha256:cc9cc2788021fd7d772b14c615cbc47195ec8ed9703194aa5fcd5ad88ab96c0d

Observation 92882141-7d2d-423f-9e38-e178bfac7bf1 · outbound

This paper cites MuVi: Video-to-Music Generation with Semantic Alignment and Rhythmic Synchronization.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections MuVi: Video-to-Music Generation with Semantic Alignment and Rhythmic Synchronization

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:43.716277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:43.716277Z digest=sha256:24a269e6eb8458bb20a9d053e725519a7acb0f7f69599c1fd325cae5785cdeba

Observation 9b963403-04a1-475c-bfe0-7f4522db4f95 · outbound

This paper cites Joint Audio and Symbolic Conditioning for Temporally Controlled Text-to-Music Generation.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Joint Audio and Symbolic Conditioning for Temporally Controlled Text-to-Music Generation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:43.791478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:43.791478Z digest=sha256:52182ca56fac9b589dc9f5f322e34d7dcfc4dec7d1d2319c1595304f821ab94e

Observation 7611343f-ec97-4eca-803e-ef17b1d82238 · outbound

This paper cites Music controlnet: Multiple time-varying controls for music generation,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Music controlnet: Multiple time-varying controls for music generation,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:55.092002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:49:43.880967Z digest=sha256:f2847205bcd674958b47f89a7806e4097a3c1facc29c98649796d4772d19d2e9

Observation 2e6b4824-b7d1-43fb-a24f-ddec9c5092af · outbound

This paper cites DITTO: Diffusion inference-time t- optimization for music generation,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections DITTO: Diffusion inference-time t- optimization for music generation,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:54.987129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:49:43.945081Z digest=sha256:afe23dbcd1f978e62b4cc555bdaa126a3a8f4753e0d59a2476496659601707b8

Observation 6d17aff5-b6c0-460a-bd6a-44a3e4baf800 · outbound

This paper cites DITTO-2: Distilled diffusion inference-time t-optimization for music generation,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections DITTO-2: Distilled diffusion inference-time t-optimization for music generation,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:54.851023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:49:44.026351Z digest=sha256:bedc8d932def206ae04f16f38bbe0a59c3b9e60937eb42f45cfa39a8fe6d5a30

Observation 29b928b4-861d-4832-8757-7750994f7a18 · outbound

This paper cites MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:44.099128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:44.099128Z digest=sha256:9c36653749ef92f24c7bfd5a896e0883579bdf9bde20e1bcd599e9a496322563

Observation b0d5dc2d-5a3c-431a-b6cc-59798dfe8b86 · outbound

This paper cites Multimodal Music Generation with Explicit Bridges and Retrieval Augmentation.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Multimodal Music Generation with Explicit Bridges and Retrieval Augmentation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:44.175395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:44.175395Z digest=sha256:1bc1fa8ad0f91796c11d266dd8f9bcc1006f5d803ab49aa7ac2edaf46858e650

Observation 6ff4f4fd-722b-4618-8a9d-efd043432fe7 · outbound

This paper cites Foley music: Learning to generate music from videos,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Foley music: Learning to generate music from videos,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:54.708344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:49:44.227354Z digest=sha256:0eaa0c370046631349e6a0d9158f9c908a4a13b0f3876ec2b3e3ce76635f5a55

Observation 333a7a68-750d-4f41-8e71-6270ce04784c · outbound

This paper cites Video background music generation with controllable music transformer,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Video background music generation with controllable music transformer,

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:44.292165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:44.292165Z digest=sha256:86cf9e734b08a616100a4ce281e602dbe655442da78080ca0558b39042589011

Observation 34a5a28c-d19a-423c-8526-5ea63f78a210 · outbound

This paper cites Vivit: A video vision transformer,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Vivit: A video vision transformer,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:54.286387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:49:44.438782Z digest=sha256:b5b5cfafdbf4b1c1e4d68bb61c8c83678696cd15262f61cdcb7d29934c93af35

Observation e1f7b1d7-f835-4570-92d1-e247baee4a2a · outbound

This paper cites Parameter-efficient transfer learning for nlp,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Parameter-efficient transfer learning for nlp,

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:44.522171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:44.522171Z digest=sha256:8d28d404c997576573adcb5e0dd18f7b4e28538cce35c0fe3f3a62b4617e17ab

Observation 7e28f1b3-58b5-4aa7-a2d5-331bf2c35972 · outbound

This paper cites AdapterFusion: Non-Destructive Task Composition for Transfer Learning.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections AdapterFusion: Non-Destructive Task Composition for Transfer Learning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:44.596943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:44.596943Z digest=sha256:9bd997afea8877f85ec96d6df2ee19e56d9696915724abbe0e73f219e511c459

Observation 98494540-e731-4a61-96b1-884d338d1334 · outbound

This paper cites AdapterHub: A Framework for Adapting Transformers.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections AdapterHub: A Framework for Adapting Transformers

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:44.690125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:44.690125Z digest=sha256:c8b7110983d8863bdc28ac81e4e1b3b57c89307450b44c9e6946783c70f5bf72

Observation 32c0adc6-f711-46ee-acfb-e46541056b6c · outbound

This paper cites IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:44.750532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:44.750532Z digest=sha256:ac862be7e52a51c7e24f2edb43ab86974eb3123e448bd879c326bec197ca032d

Observation 72702529-6ad2-48f3-9b62-3d1570f1714b · outbound

This paper cites T2i-adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections T2i-adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:54.107266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:49:44.833785Z digest=sha256:6db897a58ce24677dad35317d17d1a8bbf9f1fe34256f9cb21d4703c1a017f07

Observation ce65a2d3-06f7-4f66-9924-d8574cc2a9e5 · outbound

This paper cites Audio Prompt Adapter: Unleashing Music Editing Abilities for Text-to-Music with Lightweight Finetuning.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Audio Prompt Adapter: Unleashing Music Editing Abilities for Text-to-Music with Lightweight Finetuning

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:44.938203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:44.938203Z digest=sha256:389edec1466eca4c6733c02be6b81519caa1ea8c1ecac9551b5ee5a48206d718

Observation c0c004cf-0283-4c5d-9f20-94aaf665c911 · outbound

This paper cites Content-Based Video-Music Retrieval Using Soft Intra-Modal Structure Constraint.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Content-Based Video-Music Retrieval Using Soft Intra-Modal Structure Constraint

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:45.008013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:45.008013Z digest=sha256:ca39a39f6c673fb974fdf93450b143d8eb945115d187a1e70809d313f8440930

Observation 62e97fff-57f7-4bc8-8b5a-1aa173c5c3ae · outbound

This paper cites Creating a multitrack classical music performance dataset for multimodal music analysis: Challenges, insights, and applications,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Creating a multitrack classical music performance dataset for multimodal music analysis: Challenges, insights, and applications,

Reference 57

Resolution
malformed identifier
no resolver link, observed 2026-08-07T00:49:45.121557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:45.121557Z digest=sha256:647556aadd2635f7b0043607764efa936a76cff01aa1a1f8fad820c9dea996ed

Observation cb5d3735-9a17-41c5-9b44-bd2b4875388b · outbound

This paper cites Quantized GAN for Complex Music Generation from Dance Videos.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Quantized GAN for Complex Music Generation from Dance Videos

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:45.246773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:45.246773Z digest=sha256:be23463f1f6c6a6f59ad1bd38bf3d4f851fe53f86d8567fd5ef16851e0144d0f

Observation 50e3b075-93b3-4d87-b3f2-18ffe2df324f · outbound

This paper cites AI Choreographer: Music Conditioned 3D Dance Generation with AIST++.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections AI Choreographer: Music Conditioned 3D Dance Generation with AIST++

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:45.341616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:45.341616Z digest=sha256:2f906b1244aa11799d3a37a9b384a44216161c6bab8eb9e2557b009f2cad3562

Observation e25abf34-5f26-4e9a-8f49-7743e2818e36 · outbound

This paper cites Video back- ground music generation: Dataset, method and evalu- ation,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Video back- ground music generation: Dataset, method and evalu- ation,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:54.547185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:49:45.429434Z digest=sha256:d0b099f249327a6ebb484aa74b1c51ad0b6757bf478a6bac0f5cbff8d5c50007

Observation e3261b1c-afce-44ca-b1f9-9cf8749ed34e · outbound

This paper cites Video2music: Suitable music generation from videos using an affective multimodal transformer model,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Video2music: Suitable music generation from videos using an affective multimodal transformer model,

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:53.721432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:49:45.523930Z digest=sha256:93b1512a47885df52228bc758d2b01dbeabef9f20a1d8326be924a57750c542a

Observation 5c29d675-b417-483e-b7db-5049e09cea7e · outbound

This paper cites The Kinetics Human Action Video Dataset.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections The Kinetics Human Action Video Dataset

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:46.834553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:46.834553Z digest=sha256:ba7349dbb0d7237d1e29c735e397633fef83793fc4a89ff120759a3315ace46a

Observation b536ca68-8cc5-4f46-b632-2d4f528d1b24 · outbound

This paper cites Diff- bgm: A diffusion model for video background music generation,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Diff- bgm: A diffusion model for video background music generation,

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:53.453381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:49:45.792613Z digest=sha256:e11d7c9ccfc595126ae5cf698dafd573a8e7bb9d1851c2d293aa9e0b20cff4e5

Observation ed757164-55b8-47f6-a5e6-548e964111ed · outbound

This paper cites The nes video-music database: A dataset of symbolic video game music paired with gameplay videos,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections The nes video-music database: A dataset of symbolic video game music paired with gameplay videos,

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:53.215919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:49:45.876102Z digest=sha256:0ef34977e3fc8148499b8b8d3ea3a040d0776df639d56eaafe767c67e2164aa9

Observation 4a9c785d-d9ee-4467-81df-a3c8c7f3d811 · outbound

This paper cites an unresolved cited work.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Unresolved cited work

Reference 65

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T00:49:48.478189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:49:45.996145Z digest=sha256:7ea678711561b46aba24f7889f450e7a40cdbfa6043aeefd5f5021cf44f7ef7e

Observation e34b0e31-a537-43d3-a04c-e3180ff025b9 · outbound

This paper cites Benchmarks and leaderboards for sound demixing tasks,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Benchmarks and leaderboards for sound demixing tasks,

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:52.959182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:49:46.099559Z digest=sha256:00915fec0fe28a50a33738af280389cfd2e1319a2b7e2187b533820e2d0ab6c8

Observation a7b7110e-f6db-48df-8c8f-0f298f8ce98e · outbound

This paper cites pyaudioanalysis: An open-source python library for audio signal analysis,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections pyaudioanalysis: An open-source python library for audio signal analysis,

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:52.616912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:49:46.248153Z digest=sha256:13cff208972e9bb3b694015cdd6b1fce4f388c95f36a51923cb5e0c8b91f5659

Observation 9dfd5c2b-5a9e-4eef-a362-a5f81764df87 · outbound

This paper cites Panns: Large-scale pretrained audio neural networks for audio pattern recognition,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Panns: Large-scale pretrained audio neural networks for audio pattern recognition,

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:46.378501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:46.378501Z digest=sha256:8506e3ddb73edd4803d2d92660aa317346679107fa9c263d2c88e1ae4f853bed

Observation 50cc1ea6-278d-4b79-96b4-68bbdbe10c97 · outbound

This paper cites A circumplex model of affect.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections A circumplex model of affect

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:46.500264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:46.500264Z digest=sha256:04788a67f13dbccf5e915e9c3ebf343054fc965ab138b94f0290db810161998c

Observation 44613096-827c-4252-953e-7ace2a26ea60 · outbound

This paper cites EMOPIA: A Multi-Modal Pop Piano Dataset For Emotion Recognition and Emotion-based Music Generation.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections EMOPIA: A Multi-Modal Pop Piano Dataset For Emotion Recognition and Emotion-based Music Generation

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:46.604276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:46.604276Z digest=sha256:a38986b725a11ff6c115265bc6aad4970984425938dca0d2a519d9cb4b75956d

Observation 7fc4e7aa-0042-4bf6-bfa3-f9f85ef5d283 · outbound

This paper cites High Fidelity Neural Audio Compression.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections High Fidelity Neural Audio Compression

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:46.732431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:46.732431Z digest=sha256:93babac21c55fd31485279668c6fc89aec40ea188a8b529f0585093b70aba027

Observation 00b7d1d0-bfaa-4804-b2eb-2456a7ab899f · outbound

This paper cites Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:51.161900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:49:47.758112Z digest=sha256:857b7d11fc3a84057dd1f298ed6117073736cef7e2e5a4e68a070ef18f96ed76

Observation 4f9aa396-4221-48f1-89c8-704c1015dd39 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections LoRA: Low-Rank Adaptation of Large Language Models

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:46.967955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:46.967955Z digest=sha256:0d624836f1df3fc4f1e0447a0261a7bdb9cc638e084cca66e5a4473cfad4aeb1

Observation 0ebc6fab-03ae-4375-a19c-4f2db7f4fc0e · outbound

This paper cites Language models are few-shot learners,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Language models are few-shot learners,

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:52.186777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:49:47.082948Z digest=sha256:2ce9729645d7dcf7f03b78223178ece42f7bac98ec3302fb4a4c4fb14c9f28a4

Observation 69f6d0cd-4b1a-421a-8448-b2f4c0123638 · outbound

This paper cites Design guidelines for prompt engineering text-to-image generative models,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Design guidelines for prompt engineering text-to-image generative models,

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:51.823139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:49:47.186083Z digest=sha256:cc4d8a6e5fec1eb08159f084c255044d9f7789bb2aa23066787673b9a879f804

Observation 6e8f562d-5af2-4bc7-a4b9-8d521f795dd6 · outbound

This paper cites A Prompt Pattern Catalog to Enhance Prompt Engineering with ChatGPT.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections A Prompt Pattern Catalog to Enhance Prompt Engineering with ChatGPT

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:47.311232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:47.311232Z digest=sha256:235538e345a789d5ce99eca42a4c27540d24ee3d4592e5dcde93cda65a8af9e2

Observation 18e95974-ec41-4849-a687-64572867160f · outbound

This paper cites LP-MusicCaps: LLM-Based Pseudo Music Captioning.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections LP-MusicCaps: LLM-Based Pseudo Music Captioning

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:47.415707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:47.415707Z digest=sha256:bfe86b8e6f3c382556d336b98e24d11088ad89c7f434cc1fd44dc41072145b1a

Observation 81eda005-80ec-47dd-8f4c-298358cfd8da · outbound

This paper cites Decoupled Weight Decay Regularization.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Decoupled Weight Decay Regularization

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:47.484231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:47.484231Z digest=sha256:dbabc57a42bacf23a828fe70a2c68db4e1c00f10a8957ab9d0a4c6562496c6b4

Observation 5a4deb80-728e-4503-99cc-5da2dccca4ba · outbound

This paper cites SGDR: Stochastic Gradient Descent with Warm Restarts.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections SGDR: Stochastic Gradient Descent with Warm Restarts

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:47.550534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:47.550534Z digest=sha256:99d938a75a77d1db109a4227d43bd091c1af8abc39674cd8ca491e63a4fe637a

Observation 50ab875a-1595-4b31-bd65-6e8a2f65f51d · outbound

This paper cites Presto! distilling steps and layers for accelerating music generation.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Presto! distilling steps and layers for accelerating music generation

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:51.477011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:49:47.631070Z digest=sha256:c062d1417a5fc07402e50b0a46b6ef0e84ebc8d2413ff46ffba8016cc9de36d1

Observation 17565a16-a129-4653-ad4d-251f7255020c · outbound

This paper cites Diff-A-Riff: Musical Accompaniment Co-creation via Latent Diffusion Models.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Diff-A-Riff: Musical Accompaniment Co-creation via Latent Diffusion Models

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:47.697740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:47.697740Z digest=sha256:9e4ed6a0876d9ffc61cb2b3f02cbb82c3bfca82ec63eb66d7ef369acf1e91772

Observation 7bc0f0b4-575b-4092-9629-2b5e5b319de3 · outbound

This paper cites Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:47.832625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:47.832625Z digest=sha256:5ff18e0c3c6dc53ec6e84506cc390ef5b9b5055e66573b81b67da802de9cf71e

Observation 5851456f-0269-48f1-ae3c-d32dc18108bb · outbound

This paper cites Reliable fidelity and diversity metrics for generative models,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Reliable fidelity and diversity metrics for generative models,

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:50.882373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:49:47.896765Z digest=sha256:7aaea3b175ddff9157b168dc4959e259aa69a2b1f727f35a553782a30ee1c48b

Observation cc2e518c-34dc-4078-9479-835127505348 · outbound

This paper cites Efficient Training of Audio Transformers with Patchout.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Efficient Training of Audio Transformers with Patchout

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:47.983763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:47.983763Z digest=sha256:6b14ca84e0cf46e59709eabc6cfe7457620d044b986a641f732aa30ac5f607fb

Observation 91732895-4c5d-4970-93e6-4f4a5c4fe0e1 · outbound

This paper cites Available: http://dx.doi.org/10.1016/j.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Available: http://dx.doi.org/10.1016/j

Reference 2024

Resolution
malformed identifier
no resolver link, observed 2026-08-07T00:49:45.673911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:45.673911Z digest=sha256:f02cc30fb00aedfc2246bc272e8bd5e2cfe1017cbcfd1d794488fa9860567908

Pith citing papers

Observation 5af4092a-926b-44ce-b64d-34ac5cdc75e5 · inbound

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections cites this paper.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Video-Guided Text-to-Music Generation Using Public Domain Movie Collections

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:49:50.536973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:49:40.786176Z digest=sha256:d390917fea2aabdf690014473a223528cc020313cbc846110b20ee6863d898f7