Pith. sign in

Paper Citation Record · LEDGER

GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions

As of 19 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 3 inbound Pith citation observations for arXiv:2501.09972.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.09972 v1

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T19:34:03.414251Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:36:42.465952Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T00:49:49.179124Z

Reference resolution

39 of 39 outbound references displayed

  • verified exact2
  • verified fuzzy4
  • unresolved32
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 00e8ac5b-22d9-457f-b6ed-053e4633baad · outbound

This paper cites , " * write output.state after.block = add.period write newline.

GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions , " * write output.state after.block = add.period write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T19:34:03.167864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:34:03.167864Z digest=sha256:d1e7ef15a7078ccebfd4813f6d88d13cf3cc290e597ee5c3cddd5c1667acc740

Observation cadff098-657a-421f-80d3-28d38c8bff45 · outbound

This paper cites write newline.

GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T19:34:03.176206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:34:03.176206Z digest=sha256:258f3ad7b4bedb3dee74b4de61b42e227c209e514e004f3517080f7b7390779e

Observation f052bbe5-0e7d-4ea7-bf14-9ced26425e68 · outbound

This paper cites MusicLM: Generating Music From Text.

GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions MusicLM: Generating Music From Text

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T19:34:03.184697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:34:03.184697Z digest=sha256:1d20b567b5bc7dedbaae63b56cf55843cdce826c1b195fa17d1a13493aa0a5c9

Observation b4696e85-fef2-4782-a263-d5d766faa2ec · outbound

This paper cites an unresolved cited work.

GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-10T19:34:04.240798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-10T19:34:03.196852Z digest=sha256:aa6b97256cc0b5478a92e18fa5ae75dd01aa7b39d48720587220e06da249437e

Observation edf73a6d-e1a9-4e8f-915d-7a840dd1a4ad · outbound

This paper cites MIDI-VAE: Modeling Dynamics and Instrumentation of Music with Applications to Style Transfer.

GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions MIDI-VAE: Modeling Dynamics and Instrumentation of Music with Applications to Style Transfer

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T19:34:03.203124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:34:03.203124Z digest=sha256:3d968e7c09c9f56b120fee9da41253268bd984f01c4a43c8c14ee1ae1514ed75

Observation e231e4d7-e71a-418f-82a2-1a9143eccb01 · outbound

This paper cites an unresolved cited work.

GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T19:34:03.209762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:34:03.209762Z digest=sha256:11918981c2b2734267e22a4f339d1bddeca8e8d7eb36a5357775690f4d433cc0

Observation 265024d9-a317-4f79-a9c6-4270aac04a06 · outbound

This paper cites an unresolved cited work.

GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-10T19:34:04.202919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-10T19:34:03.215644Z digest=sha256:18004a97f70cbd41841655082312bc2ba431ca22858b979a62d3de281621138e

Observation 68a52df8-4290-4d46-a90a-88e412f594a8 · outbound

This paper cites an unresolved cited work.

GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-10T19:34:04.178181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-10T19:34:03.221999Z digest=sha256:c2e27830883364a5050a0e49d6a25da1d17c6c7fe5e8a8f377eae0e6ef8c740c

Observation 86c69587-bf6d-4bf3-a99a-2723feed2881 · outbound

This paper cites an unresolved cited work.

GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-10T19:34:04.152744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-10T19:34:03.227349Z digest=sha256:c9b29daf8cc2bc592d543b5192c8d54b97749121880ae6001d092fcb0785a55b

Observation d7a752d1-77a2-4c28-9574-452d030352ba · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T19:34:03.232529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:34:03.232529Z digest=sha256:bf078636d56078ddc99f55864e8070777f9f75469d98c16f278b7622e00205f3

Observation 1a5065e4-013d-4c2a-b32d-57a51d81548b · outbound

This paper cites High Fidelity Neural Audio Compression.

GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions High Fidelity Neural Audio Compression

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T19:34:03.238946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:34:03.238946Z digest=sha256:6c60843800f2cb60d14901f7708ff8a2da55c76fd205d51ef12346384c0e3d11

Observation 3c44e8c3-6c4c-4c34-a6f0-dade0ee9bf5a · outbound

This paper cites Fast Timing-Conditioned Latent Audio Diffusion.

GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions Fast Timing-Conditioned Latent Audio Diffusion

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T19:34:03.245573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:34:03.245573Z digest=sha256:625565d3cc3413f771014737ee1dea4b8f05057df89ea71db926af4f4fac7636

Observation 4752cb0e-7e9e-42ec-927b-42ed9272fb3b · outbound

This paper cites an unresolved cited work.

GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-10T19:34:04.134277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-10T19:34:03.253138Z digest=sha256:2f1e66dae1e7207de5ee2150004ecd5d3260a9dea23294ea6de224f24ba1a7ca

Observation 32adf641-231b-42a7-b920-f04a5272417b · outbound

This paper cites B.; and Torralba, A.

GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions B.; and Torralba, A

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:34:04.115735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-10T19:34:03.259382Z digest=sha256:8e3f37de287f2f55deb5cfa69e5f27151787c0bc2bee562a9517e67f89bfdc0d

Observation fb0a8788-d61d-47bb-8c29-39e1227e0b4b · outbound

This paper cites F.; Ellis, D.

GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions F.; Ellis, D

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:34:04.095521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-10T19:34:03.264636Z digest=sha256:1c7821d1fe451bebce575d193131008b6599bd1a95a31df1d09757dc7eec0ee0

Observation 30ff54a8-c678-4ef1-8ec8-034a48461fe2 · outbound

This paper cites Music Transformer.

GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions Music Transformer

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T19:34:03.271239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:34:03.271239Z digest=sha256:61a703fdd3819ee6960ef47051e01a0a49694f022a115675052d9d3a5d59ed8c

Observation dee3761e-d304-46c4-8832-f40997bd36fd · outbound

This paper cites MuLan: A Joint Embedding of Music Audio and Natural Language.

GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions MuLan: A Joint Embedding of Music Audio and Natural Language

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T19:34:03.276199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:34:03.276199Z digest=sha256:d5c9ea24f4f7314cec6be4548b0c5145873611081bd8a4724f81c23046dc85c6

Observation 74667877-5b8d-48ec-a005-2542e1eb2cfb · outbound

This paper cites M$^{2}$UGen: Multi-modal Music Understanding and Generation with the Power of Large Language Models.

GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions M$^{2}$UGen: Multi-modal Music Understanding and Generation with the Power of Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T19:34:03.283850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:34:03.283850Z digest=sha256:146313e616ddc4c3de382ca8675c72433055bee06aa83a8a6e5f27612937bb64

Observation c3a916e8-be50-411d-93ed-d1cec6538fbb · outbound

This paper cites A Comprehensive Survey on Deep Music Generation: Multi-level Representations, Algorithms, Evaluations, and Future Directions.

GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions A Comprehensive Survey on Deep Music Generation: Multi-level Representations, Algorithms, Evaluations, and Future Directions

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T19:34:03.289152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:34:03.289152Z digest=sha256:92252ce35f3d0d90fc6d5a0b0fd7036bfb4f61f035f6934777421c33ce287996

Observation 45bc36ba-4079-46ac-909f-8097602241cf · outbound

This paper cites an unresolved cited work.

GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-10T19:34:04.072580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-10T19:34:03.296680Z digest=sha256:0fae7555750d646bd9f28cac9b8c312a90b6b08c305c2a555505a488e33ec748

Observation a1f01869-c170-4d74-9f71-ce9246a17082 · outbound

This paper cites an unresolved cited work.

GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T19:34:03.303016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:34:03.303016Z digest=sha256:be4080da6cba51bec86f98702d391f65e8e5780518defcc6e913f24d531f9519

Observation 0b114760-cc9e-43f5-93b9-15d7120d3723 · outbound

This paper cites A.; and Kanazawa, A.

GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions A.; and Kanazawa, A

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:34:04.033601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-10T19:34:03.308164Z digest=sha256:54347fe787b1a275188b25a775d33be8d7daef1f1530724bbbb0ddac1ecae7f7

Observation 81186434-3cb3-417c-bc66-13bcddd12ced · outbound

This paper cites an unresolved cited work.

GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-10T19:34:04.009027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-10T19:34:03.315617Z digest=sha256:ae90199b514d2f5cfb384b76686a7226d2c94c8598277f3630c7cca51bdb2765

Observation fd3a5470-c1a3-4e73-b2e5-2f6364715a57 · outbound

This paper cites an unresolved cited work.

GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-10T19:34:03.987942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-10T19:34:03.321889Z digest=sha256:b4dc886a6dbc2e203f89f8db82fdf61827f68734ca5e89ad7fd47daedddba715

Observation 2b9a702e-5cfa-4dee-bcea-8e9107ccdec3 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions Learning Transferable Visual Models From Natural Language Supervision

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T19:34:03.328175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:34:03.328175Z digest=sha256:c1d998d8aa913b0f67412db2313edc8b63ab0d54a84401e6f223b037d057d973

Observation 2c15d592-2566-45c5-a930-3d23e49cec74 · outbound

This paper cites an unresolved cited work.

GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T19:34:03.335266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:34:03.335266Z digest=sha256:f3a61bd8392c6a8d335f89ab381046eac0bbc13fb779013ebae091e2a54b0e1c

Observation 1ca21307-eeaf-4a13-8732-059285d8d76b · outbound

This paper cites an unresolved cited work.

GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-10T19:34:03.956749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-10T19:34:03.341047Z digest=sha256:20f17ad2700ddd4cf51e491c33ba8ac926e2b9f6daf1cde9065013c5d2364465

Observation 3cbf07c3-abb7-4df3-a6cf-a9634488c948 · outbound

This paper cites V2Meow: Meowing to the Visual Beat via Video-to-Music Generation.

GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions V2Meow: Meowing to the Visual Beat via Video-to-Music Generation

Reference 28

Resolution
metadata mismatch
local_arxiv, observed 2026-08-10T19:34:03.631176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-10T19:34:03.347758Z digest=sha256:ba764ffeaa6b5021a32acaa4851a60f0d4b05df77bc65664a2cea252bdc511b9

Observation 9ee946ed-7ba9-4bfd-ac61-7804205eddb2 · outbound

This paper cites an unresolved cited work.

GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-10T19:34:03.936955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-10T19:34:03.353921Z digest=sha256:f302416988a1e05b07bbf7eb9a8b7d45410b1b5fd00b1717801e6dca3147a589

Observation f8eaa8e8-f8fe-4987-b672-0ca2b7892cc6 · outbound

This paper cites C.; and Salamon, J.

GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions C.; and Salamon, J

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:34:03.919347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-10T19:34:03.362317Z digest=sha256:ef0ecf4a286fc44112a6c20bcef1fd64ba95651fd45c7e448bfdb60633d44e1e

Observation e409faa0-6990-439b-8bc4-fa809be1e61b · outbound

This paper cites an unresolved cited work.

GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-10T19:34:03.900827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-10T19:34:03.368082Z digest=sha256:5df302642a43d48d9271454abc1f131be42d5b5876d92c8e022c2d4ea2a014b0

Observation 463709a4-53bd-4274-87e4-aa20368b5d2c · outbound

This paper cites N.; Kaiser, .; and Polosukhin, I.

GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions N.; Kaiser, .; and Polosukhin, I

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T19:34:03.373965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:34:03.373965Z digest=sha256:85a4c82ad29524f7c20d37a8cc8618b5942b66de5892f22d483116d90b07e845

Observation 35656cb0-1090-4d42-9119-aff59f59a1b7 · outbound

This paper cites NExT-GPT: Any-to-Any Multimodal LLM.

GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions NExT-GPT: Any-to-Any Multimodal LLM

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T19:34:03.379560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:34:03.379560Z digest=sha256:ad9bd5940719555a90dabc6e512fff6c29c30ea6e48253f8c0950a793f029147

Observation c638963a-d613-4833-8a32-50babbfec5b3 · outbound

This paper cites VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding.

GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T19:34:03.386398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:34:03.386398Z digest=sha256:b08573b98ef89561364ede37b6b3c399afb97f13902c6d7d41acda4f848e3f02

Observation 81dafed8-a95e-4883-a0df-9d34b8bab202 · outbound

This paper cites Long-Term Rhythmic Video Soundtracker.

GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions Long-Term Rhythmic Video Soundtracker

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-08-10T19:34:03.563367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-10T19:34:03.393670Z digest=sha256:50dc58f556d871a47483268f98b9fadfe76ea9e77d865f650c6639dce35511f3

Observation a956133d-1c98-4525-bbcf-8bdc96119c82 · outbound

This paper cites SoundStream: An End-to-End Neural Audio Codec.

GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions SoundStream: An End-to-End Neural Audio Codec

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T19:34:03.399077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:34:03.399077Z digest=sha256:6eab70a603767e7a57bd87394ecc1189b4be901844f0553d7e1954fd5f104d01

Observation 90dd70bd-6c09-441f-8f38-24e0b0ee568d · outbound

This paper cites Quantized GAN for Complex Music Generation from Dance Videos.

GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions Quantized GAN for Complex Music Generation from Dance Videos

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T19:34:03.404036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:34:03.404036Z digest=sha256:f84474de99adf732256e00bd2b17d02be67c891446333adb94e8a0c25cfa119b

Observation 25c13174-86e4-407e-9422-7f8f0f2e4892 · outbound

This paper cites Discrete Contrastive Diffusion for Cross-Modal Music and Image Generation.

GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions Discrete Contrastive Diffusion for Cross-Modal Music and Image Generation

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-08-10T19:34:03.477083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-10T19:34:03.409304Z digest=sha256:20cd6331023dcbe4cf286705ae9a3da80e725a800737619e46f00defb660db15

Observation 5d127102-0d5e-43dd-8c7f-1367dd12317c · outbound

This paper cites an unresolved cited work.

GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-10T19:34:03.867018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-10T19:34:03.414251Z digest=sha256:592234c35a05ca1bdf7ddc63c0756ac3274274677229f96c9417c4482e260208

Pith citing papers

Observation 8dfead22-de67-4f01-b641-00b7cddd2890 · inbound

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation cites this paper.

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:58.967310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:21:58.967310Z digest=sha256:00bf4f73c8ee26f987c87c78d59e4939df9bcd21c966165b5e6e09102c7c788c

Observation ab03e73b-0eb7-424e-86c1-3955da2955bd · inbound

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections cites this paper.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:49:49.247373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:49:43.064927Z digest=sha256:8cdc080ab12825506141ed5f54ff062191020aa7a0b6f9597f01ad7091ef475a

Observation 707f5136-b83a-41bc-bac4-8007b047d9ec · inbound

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections cites this paper.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:42.465952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:42.465952Z digest=sha256:21c5072d43150a947fff7972ec344675dac723a41b753ef3b9199ab2cf29e5a5