Pith. sign in

Paper Citation Record · LEDGER

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections

As of 20 August 2026, this Paper Citation Record lists 84 of 84 outbound references and 2 inbound Pith citation observations for arXiv:2506.12573.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.12573 v3

Coverage vector

measured 84 of 84 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:49:47.983763Z

measured 86 of 86 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:36:42.500746Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T00:49:50.392438Z

Reference resolution

84 of 84 outbound references displayed

  • verified exact6
  • verified fuzzy34
  • unresolved41
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5af4092a-926b-44ce-b64d-34ac5cdc75e5 · outbound

This paper cites Video-Guided Text-to-Music Generation Using Public Domain Movie Collections.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Video-Guided Text-to-Music Generation Using Public Domain Movie Collections

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:49:50.536973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:49:40.786176Z digest=sha256:88842690ec1e3a984bde64038dcebf2eb82a417e90c68a7e41ad19bdac985b82

Observation ad5fa3ec-dacf-498f-bb97-38acd2d4cd7b · outbound

This paper cites an unresolved cited work.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:40.864936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:40.864936Z digest=sha256:99cf2f22c0dc6084d9fb6922eb74ac35e2a7f6fc3bdba3d785e354c252f4e64c

Observation 5a892216-9ee5-4236-be83-f39bfade8dd6 · outbound

This paper cites We provide an overview of the comparison of video-music datasets in Table 1 and an illustration of our dataset con- struction methodology in Figure 1.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections We provide an overview of the comparison of video-music datasets in Table 1 and an illustration of our dataset con- struction methodology in Figure 1

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:40.950159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:40.950159Z digest=sha256:281c75dc29fa2fc4cc003c0de8b2e8b6ec53a0f0ccd83c4a712f08468b68e4dc

Observation 96cf67e9-1b5d-4273-8464-6b780233dcf4 · outbound

This paper cites an unresolved cited work.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:41.045157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:41.045157Z digest=sha256:62336eb981ae5a2c9b5c8e75b43d7f814588a24a79c666be8a0a0f12bb94c487

Observation 2749241b-9bb3-4ded-8c1c-40b2dd661c27 · outbound

This paper cites S” and “M.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections S” and “M

Reference 5

Resolution
verified exact
raw_fallback, observed 2026-08-07T00:49:50.232025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:49:41.120789Z digest=sha256:93335599eabd863fa340790eabe6a269ba2fb0fa39c217e67510e886446ef7dc

Observation 2ccb242a-283f-46be-99c7-458adbebe840 · outbound

This paper cites Distributional FidelityOur evaluation on OES-Com re- veals that fine-tuning on OSSL enhances distributional fi- delity.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Distributional FidelityOur evaluation on OES-Com re- veals that fine-tuning on OSSL enhances distributional fi- delity

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:41.198985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:41.198985Z digest=sha256:e96532b23633b9a2d72ceb88dfbed1f1a632dbdf3a6bd0cc7a5adb625213be10

Observation 330770bb-000c-4f02-bfd3-07d4e36c92cf · outbound

This paper cites To show the effectiveness of our dataset, we adapted a text-to-music generation model with video conditions and fine-tuned it on our dataset.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections To show the effectiveness of our dataset, we adapted a text-to-music generation model with video conditions and fine-tuned it on our dataset

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:57.630279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:49:41.272353Z digest=sha256:750ac227bd2b7a45dc2f1691001e80867495bec88edfb05e42007ceba1c26a6a

Observation aa0b77f5-740f-4a30-94a9-5bef220bc6b6 · outbound

This paper cites an unresolved cited work.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:49:57.502642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:49:41.339642Z digest=sha256:7c7481937367acd7ea115338196335148b9d9706e7433a008db6393fb745688f

Observation a2b075f8-ad4d-41ad-9f10-d69886d45ea0 · outbound

This paper cites Soundtrack design: The impact of music on visual attention and affective responses,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Soundtrack design: The impact of music on visual attention and affective responses,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:57.424060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:49:41.417326Z digest=sha256:e3237f66e39ced32b4d86d333594240d7146056c511247c7820ad6b0db9f4ac0

Observation 931c243d-a0be-450f-8fe9-b6abb48a5c18 · outbound

This paper cites Multimodal deep models for predicting affective responses evoked by movies.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Multimodal deep models for predicting affective responses evoked by movies

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:57.347541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:49:41.494318Z digest=sha256:bdfbb96037613a601fb7ab934d4119e9fbe67ec06c5b98a7ee78a8b6728ccb68

Observation 65756565-9c31-4f5b-acc1-9927be466874 · outbound

This paper cites Emotion Embedding Spaces for Matching Music to Stories.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Emotion Embedding Spaces for Matching Music to Stories

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:49:49.863527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:49:41.584055Z digest=sha256:d2ef1d2ce340fe61d14a76b332bf11f5095696de1e78f3554ea88887bf96dac1

Observation 2d14e174-ec1a-42a6-8b9b-228347da9a6f · outbound

This paper cites Attendaffectnet–emotion prediction of movie viewers using multimodal fusion with self-attention,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Attendaffectnet–emotion prediction of movie viewers using multimodal fusion with self-attention,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:57.192060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:49:41.662550Z digest=sha256:ddcb2c0f191d0c4c2747ff9990241b735e95f7671eb533290a76513c151b2990

Observation 24dc64c5-acca-4af7-92cb-6137c395ebaa · outbound

This paper cites Predicting emotion from music videos: exploring the relative contribution of visual and auditory information to affective responses.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Predicting emotion from music videos: exploring the relative contribution of visual and auditory information to affective responses

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:49:49.551873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:49:41.739826Z digest=sha256:23f9a8d6b70b6040e352623b55bd1e164264129aa710828fa7a3f4af4e1e6d60

Observation 4b747be7-3c98-4c30-ba78-447e0aa5feb9 · outbound

This paper cites Analysis of the roles of film soundtracks in films,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Analysis of the roles of film soundtracks in films,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:57.116682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:49:41.820319Z digest=sha256:a3abb6d3ea54ba72cc70e0f6d323357b9a76f96e38014f3ccd4fa4666180ae55

Observation 7f35f1ab-b34e-4616-b036-c2cfcc4bb411 · outbound

This paper cites Actions in context,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Actions in context,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:56.941092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:49:41.884556Z digest=sha256:891d27a4029a215c62e2c871b2a255bdf7c2ad2455b515b2f84feedee28c2c13

Observation 57e6cdd2-bd3f-4ded-a1b7-431a494383cb · outbound

This paper cites Movieqa: Understanding stories in movies through question-answering,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Movieqa: Understanding stories in movies through question-answering,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:56.692466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:49:41.963832Z digest=sha256:098873c892cc9051c554bd0bb6384fa5ef9bc8b0ecd825a59f832aab6392357a

Observation 178f64c9-084c-4961-8dd7-9428a98bcc0f · outbound

This paper cites Movienet: A holistic dataset for movie understand- ing,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Movienet: A holistic dataset for movie understand- ing,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:56.534391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:49:42.035597Z digest=sha256:a67fa2800ab42fd17ee4e16ee7972e74fe36dad60e60a445b82fce32ad1f0b43

Observation 8255ea0f-8c7d-4a2e-a083-578fb835f1e5 · outbound

This paper cites Mad: A scalable dataset for language grounding in videos from movie audio descriptions,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Mad: A scalable dataset for language grounding in videos from movie audio descriptions,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:56.367972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:49:42.103527Z digest=sha256:8263538b823860ccec3865094a37daa43936839ad441de8ba889db7c1b295f21

Observation 9f955c11-efdf-4d39-8a82-67e5f90dbc62 · outbound

This paper cites A dataset for movie description,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections A dataset for movie description,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:56.278451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:49:42.197524Z digest=sha256:a24331a958970ac6810aec20863ad073a01c7db1c20bff5343415dadbb7919d0

Observation f61528fb-e050-4ddf-bbf7-9f2c67e5fc0b · outbound

This paper cites Moviegraphs: Towards understanding human-centric situations from videos,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Moviegraphs: Towards understanding human-centric situations from videos,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:56.112885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:49:42.279348Z digest=sha256:797f8c7ec6417fc5b68606de7a487798a744831e560e7c9219ed86a0d54f22f4

Observation 9086561a-5ce8-448d-84b3-30760610658f · outbound

This paper cites Hlvu: A new challenge to test deep understanding of movies the way humans do,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Hlvu: A new challenge to test deep understanding of movies the way humans do,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:56.021400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:49:42.366444Z digest=sha256:87d2465892e8a8f655ab422b4c5fa48446a124b8fc3bbf21c4cbc203a099c9ae

Observation 3237553e-2891-420a-abb0-136945d11bf6 · outbound

This paper cites Condensed movies: Story based retrieval with contex- tual embeddings,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Condensed movies: Story based retrieval with contex- tual embeddings,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:55.920877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:49:42.445904Z digest=sha256:ac2715e7e57bec71cf54bd022f3a50732f5f4eb140e3705671ae77da788ce760

Observation 4799cefd-2193-40fc-a7d3-0a70dd655e0d · outbound

This paper cites TeaserGen: Generating Teasers for Long Documentaries.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections TeaserGen: Generating Teasers for Long Documentaries

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:42.511515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:42.511515Z digest=sha256:d3975db37001458fabb8815ff39b307e9cbe3a84798e89f8697e8adfaa0b1e28

Observation a9178205-829d-4963-b251-0bb2ad2877d2 · outbound

This paper cites Simple and control- lable music generation,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Simple and control- lable music generation,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:55.764831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:49:42.599104Z digest=sha256:8786d3445999d92800a776b01114a3435cb7a0e7bc0d93c21a3526f3bdd0b8c6

Observation f1a5fa92-d749-4ce4-9143-a20f91670f51 · outbound

This paper cites Attention is all you need,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Attention is all you need,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:42.675447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:42.675447Z digest=sha256:224e86fb6062d8fbb642e8932a09195cedc122494a6dfd6a14bfa469e43c8972

Observation f47ca879-4b18-487b-9530-1e89f691e266 · outbound

This paper cites MusicLM: Generating Music From Text.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections MusicLM: Generating Music From Text

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:42.750717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:42.750717Z digest=sha256:762c0114ce54c50690af23f34c7f5d9b8de142883101884552d3ec93b773d4cf

Observation 744cb494-a203-4312-a521-fafac7c164da · outbound

This paper cites MusiConGen: Rhythm and Chord Control for Transformer-Based Text-to-Music Generation.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections MusiConGen: Rhythm and Chord Control for Transformer-Based Text-to-Music Generation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:42.827934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:42.827934Z digest=sha256:83042abbe05d35eda89c06ec7fb6b2e11b82ee3cfcc1328734b89e798ca66861

Observation 1450699d-fda3-4db1-be6d-12e8ad47a34d · outbound

This paper cites VidMuse: A Simple Video-to-Music Generation Framework with Long-Short-Term Modeling.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections VidMuse: A Simple Video-to-Music Generation Framework with Long-Short-Term Modeling

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:42.900461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:42.900461Z digest=sha256:2e44f794dcb7d77e1c9a14d8077ddef25854823cad1f46e41420e0c199b117f3

Observation f33b4d9b-4689-465d-bb03-a5353cfe780b · outbound

This paper cites V2meow: Meowing to the visual beat via video-to- music generation,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections V2meow: Meowing to the visual beat via video-to- music generation,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:55.665078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:49:42.977529Z digest=sha256:17f82386d85b5739c7fd44124e185411902e02b7f1113e6496ca49f3ea244e14

Observation ab03e73b-0eb7-424e-86c1-3955da2955bd · outbound

This paper cites GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:49:49.247373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:49:43.064927Z digest=sha256:4527659b1d363dab1b39130d0f6588a8aed87e7460f70dfcf4f321d045519f78

Observation 7f8710a4-62bd-4f3e-ab57-14e14f6e0164 · outbound

This paper cites Riffusion-stable diffusion for real-time music generation,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Riffusion-stable diffusion for real-time music generation,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:55.456481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:49:43.156770Z digest=sha256:091af405147073c373c8f6830f025e5aca54fea3209c29d54cf680f41a7cc61e

Observation 92b97fe1-1c0a-456e-a812-f804e48e40f9 · outbound

This paper cites Noise2Music: Text-conditioned Music Generation with Diffusion Models.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Noise2Music: Text-conditioned Music Generation with Diffusion Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:43.230631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:43.230631Z digest=sha256:7d591f84eff679cf4f1d264cb35a9d20988611cd1e2b6d2dc44153a607eb2257

Observation ff44a7fe-ccb8-437f-a06d-699d6dba7156 · outbound

This paper cites Mo\^usai: Text-to-Music Generation with Long-Context Latent Diffusion.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Mo\^usai: Text-to-Music Generation with Long-Context Latent Diffusion

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:43.294050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:43.294050Z digest=sha256:30d1e186ee66c7183d2f0bf521e86aae69fc2f514569f9437fa5010673b9e435

Observation 41e83f31-e1a9-4b62-9f2b-e0d367cf7fda · outbound

This paper cites Efficient neu- ral music generation,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Efficient neu- ral music generation,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:55.286800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:49:43.347990Z digest=sha256:5fdafc0bce62aad6fd62b20e64e55dd32802867665f5dfffd600545f7b19368b

Observation 712b6267-1f37-4ea1-94e7-9cfb31d71bfd · outbound

This paper cites Mustango: Toward Controllable Text-to-Music Generation.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Mustango: Toward Controllable Text-to-Music Generation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:43.437966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:43.437966Z digest=sha256:a75cbdac1584f1271e63d1f38a1d9268c970db6242538561f6fc707e8f76bdc2

Observation c168ce5b-df36-4fc9-a7bb-213c000ebc30 · outbound

This paper cites Multi-Track MusicLDM: Towards Versatile Music Generation with Latent Diffusion Model.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Multi-Track MusicLDM: Towards Versatile Music Generation with Latent Diffusion Model

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:49:48.971726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:49:43.507489Z digest=sha256:14d75ce579b2fa6cd5711342bd9503edab2d976283c12298b1b6caf96c2e5993

Observation 7e3009df-3e6b-43a6-a441-32a41699f533 · outbound

This paper cites Stable Audio Open.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Stable Audio Open

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:43.573592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:43.573592Z digest=sha256:d6897b6b7b64444b2c5863ad899698f6b01e4c844a821707292e8ad3fabe20c4

Observation d9d5c618-8e40-4ceb-a172-1711575d2e26 · outbound

This paper cites VMAS: Video-to-Music Generation via Semantic Alignment in Web Music Videos.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections VMAS: Video-to-Music Generation via Semantic Alignment in Web Music Videos

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:43.639290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:43.639290Z digest=sha256:1deb5e438474bce263e1a8f71c40e90541541e6a3cd0d92c17d2a4d8cbf66ac1

Observation 92882141-7d2d-423f-9e38-e178bfac7bf1 · outbound

This paper cites MuVi: Video-to-Music Generation with Semantic Alignment and Rhythmic Synchronization.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections MuVi: Video-to-Music Generation with Semantic Alignment and Rhythmic Synchronization

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:43.716277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:43.716277Z digest=sha256:7bdea18db7d6221111420c98eaefa6bd8dda15051501db584bb7055f76fe8611

Observation 9b963403-04a1-475c-bfe0-7f4522db4f95 · outbound

This paper cites Joint Audio and Symbolic Conditioning for Temporally Controlled Text-to-Music Generation.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Joint Audio and Symbolic Conditioning for Temporally Controlled Text-to-Music Generation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:43.791478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:43.791478Z digest=sha256:cad9e914aeeaf65d077ea8bed7523b68ae3a34a8e04e79bc8a531096e3cde109

Observation 7611343f-ec97-4eca-803e-ef17b1d82238 · outbound

This paper cites Music controlnet: Multiple time-varying controls for music generation,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Music controlnet: Multiple time-varying controls for music generation,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:55.092002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:49:43.880967Z digest=sha256:d24e276a87678df8621aab5363ed917677e1b5f8f293ef2342b2a3a90c9ed86f

Observation 2e6b4824-b7d1-43fb-a24f-ddec9c5092af · outbound

This paper cites DITTO: Diffusion inference-time t- optimization for music generation,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections DITTO: Diffusion inference-time t- optimization for music generation,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:54.987129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:49:43.945081Z digest=sha256:133a84a78da66f17cee8db895caba5854e28af6434d48c84d51c879578bfc7ba

Observation 6d17aff5-b6c0-460a-bd6a-44a3e4baf800 · outbound

This paper cites DITTO-2: Distilled diffusion inference-time t-optimization for music generation,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections DITTO-2: Distilled diffusion inference-time t-optimization for music generation,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:54.851023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:49:44.026351Z digest=sha256:451084437abf78c0feb9c65edeea931f935aa12588754f73fe93acdc3aaa9a42

Observation 29b928b4-861d-4832-8757-7750994f7a18 · outbound

This paper cites MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:44.099128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:44.099128Z digest=sha256:a6353873f7e6cdabc72d6bc23d2cee413471ad1a5a3be3f92b62f7cf680d04e6

Observation b0d5dc2d-5a3c-431a-b6cc-59798dfe8b86 · outbound

This paper cites Multimodal Music Generation with Explicit Bridges and Retrieval Augmentation.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Multimodal Music Generation with Explicit Bridges and Retrieval Augmentation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:44.175395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:44.175395Z digest=sha256:f962026d05d5319b5fba382ebb03600556761344e6175e6ded5ab11b67dbf897

Observation 6ff4f4fd-722b-4618-8a9d-efd043432fe7 · outbound

This paper cites Foley music: Learning to generate music from videos,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Foley music: Learning to generate music from videos,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:54.708344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:49:44.227354Z digest=sha256:58a002591470047734908154b3cef5e457b27a6a27fee236c56e950fb308ff2d

Observation 333a7a68-750d-4f41-8e71-6270ce04784c · outbound

This paper cites Video background music generation with controllable music transformer,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Video background music generation with controllable music transformer,

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:44.292165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:44.292165Z digest=sha256:e3938e3eddc78ca59c05209e24e07a0f83e4e8bf78d3a6cb8a5d89c1e990fdda

Observation 34a5a28c-d19a-423c-8526-5ea63f78a210 · outbound

This paper cites Vivit: A video vision transformer,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Vivit: A video vision transformer,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:54.286387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:49:44.438782Z digest=sha256:d1f394636b74fac09b0defa808675bec6971214851edf7c251250e5f3b4ebfb6

Observation e1f7b1d7-f835-4570-92d1-e247baee4a2a · outbound

This paper cites Parameter-efficient transfer learning for nlp,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Parameter-efficient transfer learning for nlp,

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:44.522171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:44.522171Z digest=sha256:a714f8bfb9b9adb9e4cc6dc3d4a6d76d18e030a3d963735d42eadb919e060865

Observation 7e28f1b3-58b5-4aa7-a2d5-331bf2c35972 · outbound

This paper cites AdapterFusion: Non-Destructive Task Composition for Transfer Learning.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections AdapterFusion: Non-Destructive Task Composition for Transfer Learning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:44.596943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:44.596943Z digest=sha256:36ea92073fd2854c47ac3c84aa1969b7a322dce80e982243bb6afbd8319506be

Observation 98494540-e731-4a61-96b1-884d338d1334 · outbound

This paper cites AdapterHub: A Framework for Adapting Transformers.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections AdapterHub: A Framework for Adapting Transformers

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:44.690125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:44.690125Z digest=sha256:b16445708729ad8e9fc7b5f56db75ae2eb3eca5bfc78ae1a20d0d451f82782d5

Observation 32c0adc6-f711-46ee-acfb-e46541056b6c · outbound

This paper cites IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:44.750532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:44.750532Z digest=sha256:bc287b9906bd818f2587ac4d5ec8d6a5691fc8eff5fce208ebe0518c9589c4b8

Observation 72702529-6ad2-48f3-9b62-3d1570f1714b · outbound

This paper cites T2i-adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections T2i-adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:54.107266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:49:44.833785Z digest=sha256:df4f6da0d4debcee424adff1f0a72d703c0cb75892e18d1d5900ca3a43d47372

Observation ce65a2d3-06f7-4f66-9924-d8574cc2a9e5 · outbound

This paper cites Audio Prompt Adapter: Unleashing Music Editing Abilities for Text-to-Music with Lightweight Finetuning.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Audio Prompt Adapter: Unleashing Music Editing Abilities for Text-to-Music with Lightweight Finetuning

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:44.938203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:44.938203Z digest=sha256:b02f839db9285798e21e705e5e903a3dcb468a54bdf357ea11aeff3809b5c4cf

Observation c0c004cf-0283-4c5d-9f20-94aaf665c911 · outbound

This paper cites Content-Based Video-Music Retrieval Using Soft Intra-Modal Structure Constraint.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Content-Based Video-Music Retrieval Using Soft Intra-Modal Structure Constraint

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:45.008013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:45.008013Z digest=sha256:6424117214d87a15527665b37c617ca86bdd0fbb3c1d66b826efe04ed76ad2d8

Observation 62e97fff-57f7-4bc8-8b5a-1aa173c5c3ae · outbound

This paper cites Creating a multitrack classical music performance dataset for multimodal music analysis: Challenges, insights, and applications,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Creating a multitrack classical music performance dataset for multimodal music analysis: Challenges, insights, and applications,

Reference 57

Resolution
malformed identifier
no resolver link, observed 2026-08-07T00:49:45.121557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:45.121557Z digest=sha256:ad4c023222e3478c505d6245ee1eae50b5312ba06c42b953d337e376f430cb5a

Observation cb5d3735-9a17-41c5-9b44-bd2b4875388b · outbound

This paper cites Quantized GAN for Complex Music Generation from Dance Videos.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Quantized GAN for Complex Music Generation from Dance Videos

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:45.246773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:45.246773Z digest=sha256:5448c87bb7fce96ba351cc13a6c776a7f58779de358f924a9afdc2d267407799

Observation 50e3b075-93b3-4d87-b3f2-18ffe2df324f · outbound

This paper cites AI Choreographer: Music Conditioned 3D Dance Generation with AIST++.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections AI Choreographer: Music Conditioned 3D Dance Generation with AIST++

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:45.341616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:45.341616Z digest=sha256:73c03c23ca2c72bbfb4cc9d7b8123c4d94d1907c95c59f1776bbee3f997a2555

Observation e25abf34-5f26-4e9a-8f49-7743e2818e36 · outbound

This paper cites Video back- ground music generation: Dataset, method and evalu- ation,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Video back- ground music generation: Dataset, method and evalu- ation,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:54.547185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:49:45.429434Z digest=sha256:f9fbc68596026d35703d31e9e274856f302ce2ec3f837059a8b6034caffbce07

Observation e3261b1c-afce-44ca-b1f9-9cf8749ed34e · outbound

This paper cites Video2music: Suitable music generation from videos using an affective multimodal transformer model,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Video2music: Suitable music generation from videos using an affective multimodal transformer model,

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:53.721432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:49:45.523930Z digest=sha256:94437d7d70889654e9db9873137d9f87fdb7156339efd5f14c58d4d8918a6dbd

Observation 5c29d675-b417-483e-b7db-5049e09cea7e · outbound

This paper cites The Kinetics Human Action Video Dataset.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections The Kinetics Human Action Video Dataset

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:46.834553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:46.834553Z digest=sha256:8047db8b4e1a552e3b3067519c7983236e7cb2aa9c546ace7e5d1e834b8f54a2

Observation b536ca68-8cc5-4f46-b632-2d4f528d1b24 · outbound

This paper cites Diff- bgm: A diffusion model for video background music generation,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Diff- bgm: A diffusion model for video background music generation,

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:53.453381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:49:45.792613Z digest=sha256:d8da2b9cf532191b927b46193564177023045af1c312ea5e283a40a7c4c13883

Observation ed757164-55b8-47f6-a5e6-548e964111ed · outbound

This paper cites The nes video-music database: A dataset of symbolic video game music paired with gameplay videos,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections The nes video-music database: A dataset of symbolic video game music paired with gameplay videos,

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:53.215919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:49:45.876102Z digest=sha256:aaeca3e006d91b1928e0a31413db269090b7164e97a723d4aa08f37237ac8998

Observation 4a9c785d-d9ee-4467-81df-a3c8c7f3d811 · outbound

This paper cites an unresolved cited work.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Unresolved cited work

Reference 65

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T00:49:48.478189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:49:45.996145Z digest=sha256:a1c079512dfeb14314f114dde22411af28d4d4616e2acfe2cc9017e19a6933ef

Observation e34b0e31-a537-43d3-a04c-e3180ff025b9 · outbound

This paper cites Benchmarks and leaderboards for sound demixing tasks,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Benchmarks and leaderboards for sound demixing tasks,

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:52.959182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:49:46.099559Z digest=sha256:b169d866cae025c7b07b4510ef55984eec3c43393d09b82e1e31c38a41766636

Observation a7b7110e-f6db-48df-8c8f-0f298f8ce98e · outbound

This paper cites pyaudioanalysis: An open-source python library for audio signal analysis,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections pyaudioanalysis: An open-source python library for audio signal analysis,

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:52.616912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:49:46.248153Z digest=sha256:4434a48b45f229c6f5d8eb8f8fcbd4941c7e70725067980adc44e982f6f25483

Observation 9dfd5c2b-5a9e-4eef-a362-a5f81764df87 · outbound

This paper cites Panns: Large-scale pretrained audio neural networks for audio pattern recognition,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Panns: Large-scale pretrained audio neural networks for audio pattern recognition,

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:46.378501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:46.378501Z digest=sha256:5f78ec4851a4292301eb6739bbf4e939d4734fe3eed58b6e43eaf1e177b059d8

Observation 50cc1ea6-278d-4b79-96b4-68bbdbe10c97 · outbound

This paper cites A circumplex model of affect.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections A circumplex model of affect

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:46.500264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:46.500264Z digest=sha256:54fb8c7b61a5517c8c2accc2ca8f254206ec7b3fdd0f27b60d3d249eb1adfde6

Observation 44613096-827c-4252-953e-7ace2a26ea60 · outbound

This paper cites EMOPIA: A Multi-Modal Pop Piano Dataset For Emotion Recognition and Emotion-based Music Generation.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections EMOPIA: A Multi-Modal Pop Piano Dataset For Emotion Recognition and Emotion-based Music Generation

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:46.604276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:46.604276Z digest=sha256:d9e1e136a26fa40b86c32dc4579f43157eef83284514cafdb7396fc2211d4c7a

Observation 7fc4e7aa-0042-4bf6-bfa3-f9f85ef5d283 · outbound

This paper cites High Fidelity Neural Audio Compression.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections High Fidelity Neural Audio Compression

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:46.732431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:46.732431Z digest=sha256:327385a48d7e54e11204d54992735a8c97eb4691a261ce6c575e873de592141a

Observation 00b7d1d0-bfaa-4804-b2eb-2456a7ab899f · outbound

This paper cites Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:51.161900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:49:47.758112Z digest=sha256:62737230507110cc0c0e94d56bd7849921094f30856c52cd51a9c2007555536e

Observation 4f9aa396-4221-48f1-89c8-704c1015dd39 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections LoRA: Low-Rank Adaptation of Large Language Models

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:46.967955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:46.967955Z digest=sha256:b0ebb6c45f9b9c3eb0ca34f35ab00fd9a21ed1f8d212755e219a48866cabeb5d

Observation 0ebc6fab-03ae-4375-a19c-4f2db7f4fc0e · outbound

This paper cites Language models are few-shot learners,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Language models are few-shot learners,

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:52.186777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:49:47.082948Z digest=sha256:99e8acfacf935e77553d3767411782f9affd3ee1757505b2d52c97cb98e5dc72

Observation 69f6d0cd-4b1a-421a-8448-b2f4c0123638 · outbound

This paper cites Design guidelines for prompt engineering text-to-image generative models,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Design guidelines for prompt engineering text-to-image generative models,

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:51.823139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:49:47.186083Z digest=sha256:76eb08c1eb5d47d161f0ebfdaacae4ef90e4a6dccdb2ee9c99af066e681c6111

Observation 6e8f562d-5af2-4bc7-a4b9-8d521f795dd6 · outbound

This paper cites A Prompt Pattern Catalog to Enhance Prompt Engineering with ChatGPT.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections A Prompt Pattern Catalog to Enhance Prompt Engineering with ChatGPT

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:47.311232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:47.311232Z digest=sha256:e7dc4d7132ea5cfee0f9fe3a7d86d98280c663e35467db005f7f3aaa6941a2dd

Observation 18e95974-ec41-4849-a687-64572867160f · outbound

This paper cites LP-MusicCaps: LLM-Based Pseudo Music Captioning.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections LP-MusicCaps: LLM-Based Pseudo Music Captioning

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:47.415707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:47.415707Z digest=sha256:26532deb1342a2d2428a9c3e35fc6146f4e0c34f1cd542bd3227270c16cbd050

Observation 81eda005-80ec-47dd-8f4c-298358cfd8da · outbound

This paper cites Decoupled Weight Decay Regularization.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Decoupled Weight Decay Regularization

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:47.484231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:47.484231Z digest=sha256:b5a376ac0299afb57358869f9b11813cc081c51520e4a8faae3e799ea0c20f60

Observation 5a4deb80-728e-4503-99cc-5da2dccca4ba · outbound

This paper cites SGDR: Stochastic Gradient Descent with Warm Restarts.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections SGDR: Stochastic Gradient Descent with Warm Restarts

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:47.550534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:47.550534Z digest=sha256:46a1eaa06d167121b0e3e3df28541eb02143bd32e5d2af9aaa52d680526fb48a

Observation 50ab875a-1595-4b31-bd65-6e8a2f65f51d · outbound

This paper cites Presto! distilling steps and layers for accelerating music generation.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Presto! distilling steps and layers for accelerating music generation

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:51.477011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:49:47.631070Z digest=sha256:ea7a8972761d4251503910dcdcaa0627e2ec7434857a6efc386574f8f25f4e99

Observation 17565a16-a129-4653-ad4d-251f7255020c · outbound

This paper cites Diff-A-Riff: Musical Accompaniment Co-creation via Latent Diffusion Models.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Diff-A-Riff: Musical Accompaniment Co-creation via Latent Diffusion Models

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:47.697740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:47.697740Z digest=sha256:c284409093cfc21615ce12d3b284322e33f9b472149ddc3f69ea2107e1c43c45

Observation 7bc0f0b4-575b-4092-9629-2b5e5b319de3 · outbound

This paper cites Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:47.832625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:47.832625Z digest=sha256:1522d5fa03fd7710017e237582cdaba22046563d21d5f5058cd89edc68b7483b

Observation 5851456f-0269-48f1-ae3c-d32dc18108bb · outbound

This paper cites Reliable fidelity and diversity metrics for generative models,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Reliable fidelity and diversity metrics for generative models,

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:50.882373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:49:47.896765Z digest=sha256:78f27b40e620c737076f821217aa79e4431676788113c9be4a91036956db1673

Observation cc2e518c-34dc-4078-9479-835127505348 · outbound

This paper cites Efficient Training of Audio Transformers with Patchout.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Efficient Training of Audio Transformers with Patchout

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:47.983763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:47.983763Z digest=sha256:6750f535ce045c1278c47b0ff569801e8ba136d49a5eda55fc50eacb89bdbd52

Observation 91732895-4c5d-4970-93e6-4f4a5c4fe0e1 · outbound

This paper cites Available: http://dx.doi.org/10.1016/j.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Available: http://dx.doi.org/10.1016/j

Reference 2024

Resolution
malformed identifier
no resolver link, observed 2026-08-07T00:49:45.673911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:45.673911Z digest=sha256:aa41ff998d4a6ac779d93393035aef6f89a137e4ed7afcdda4df1d89b543184d

Pith citing papers

Observation 5af4092a-926b-44ce-b64d-34ac5cdc75e5 · inbound

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections cites this paper.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Video-Guided Text-to-Music Generation Using Public Domain Movie Collections

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:49:50.536973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T00:49:40.786176Z digest=sha256:88842690ec1e3a984bde64038dcebf2eb82a417e90c68a7e41ad19bdac985b82

Observation 8147f160-bd41-4192-82f0-c30b73745ab6 · inbound

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections cites this paper.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Video-Guided Text-to-Music Generation Using Public Domain Movie Collections

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:42.500746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:42.500746Z digest=sha256:1207c1da355f0b3b7a692caf7838a671bfe224462c7dfd0031df2bd03183cee7