Pith. sign in

Paper Citation Record · LEDGER

Semi-Supervised Contrastive Learning for Controllable Video-to-Music Retrieval

As of 20 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 0 inbound Pith citation observations for arXiv:2412.05831.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.05831 v2

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T20:21:45.248236Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

24 of 24 outbound references displayed

  • verified exact0
  • verified fuzzy21
  • unresolved3
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bb449eaf-2301-4268-9367-95989e861735 · outbound

This paper cites Cbvmr: Content-based video-music retrieval using soft intra-modal structure constraint,.

Semi-Supervised Contrastive Learning for Controllable Video-to-Music Retrieval Cbvmr: Content-based video-music retrieval using soft intra-modal structure constraint,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:21:45.697174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T20:21:45.120483Z digest=sha256:4ab77ed7bf9d9e52e795566b28d65137f3bad9df9703a8481550ed60fa68fd3c

Observation f5f4363e-4e16-439a-801a-a63c9e3840f3 · outbound

This paper cites Cross-modal music-video recommendation: A study of design choices,.

Semi-Supervised Contrastive Learning for Controllable Video-to-Music Retrieval Cross-modal music-video recommendation: A study of design choices,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:21:45.680434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T20:21:45.126828Z digest=sha256:0eec06a74d178a5ffc674f39b5a40f9d8448a31cc3c74cfa7239fe520b70426c

Observation c6af6665-5c07-4a42-82ee-5b0b736a97d4 · outbound

This paper cites Is there a.

Semi-Supervised Contrastive Learning for Controllable Video-to-Music Retrieval Is there a

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:21:45.663486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T20:21:45.132315Z digest=sha256:0a3fe8d353f9198bd57fbe3f13f75c7474e1dc4601e43753db615313a0f6bae0

Observation 27449000-beef-427e-9282-e4b6e83e1ee6 · outbound

This paper cites It’s time for artistic correspondence in music and video,.

Semi-Supervised Contrastive Learning for Controllable Video-to-Music Retrieval It’s time for artistic correspondence in music and video,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:21:45.645865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T20:21:45.141179Z digest=sha256:551efd2ede096c72dc365283e8b7197f421c8b576cbe616dfb400be931767b25

Observation 38a289ed-55e0-4c7a-adfc-7045d6dcdabc · outbound

This paper cites Attention is all you need,.

Semi-Supervised Contrastive Learning for Controllable Video-to-Music Retrieval Attention is all you need,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T20:21:45.146762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:21:45.146762Z digest=sha256:1e4a0642a1f44a935f8fc8bf77d257ab6a877f352e454fb642462f3add76cbcb

Observation 87c840bb-2d31-4861-a70e-572cd4012f7f · outbound

This paper cites Language-guided music recommendation for video via prompt analogies,.

Semi-Supervised Contrastive Learning for Controllable Video-to-Music Retrieval Language-guided music recommendation for video via prompt analogies,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:21:45.616596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T20:21:45.152079Z digest=sha256:c6b49690fdf421744bb2ec584ce919a84698c08695490a29837f1829157fc7de

Observation 9cd9c6ce-5e02-4100-a3df-2f65224f4608 · outbound

This paper cites Mulan: A joint embedding of music audio and natural language,.

Semi-Supervised Contrastive Learning for Controllable Video-to-Music Retrieval Mulan: A joint embedding of music audio and natural language,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:21:45.599208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T20:21:45.157783Z digest=sha256:d1e8c160a2fdd0efcbe34d84155bc4ded19272858f726dedf4892f00ed75e97e

Observation 839b6c45-13b0-48c1-aa9c-b6ca7a82b65d · outbound

This paper cites Contrastive audio- language learning for music,.

Semi-Supervised Contrastive Learning for Controllable Video-to-Music Retrieval Contrastive audio- language learning for music,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:21:45.582209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T20:21:45.162532Z digest=sha256:08c10d271c5539e4892f9fe86a21e1e3428da2f76a847304abf505974ead85ac

Observation 46f0e205-ad4a-4668-a746-d986377aff5c · outbound

This paper cites Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,.

Semi-Supervised Contrastive Learning for Controllable Video-to-Music Retrieval Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:21:45.564464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T20:21:45.167403Z digest=sha256:32ed82032db1b8821af7b9e7e44ae60737f973fc6cc0ab8ff8cf4ee00cf3a660

Observation ce8764b0-7571-4890-b2c9-5126d060ba8b · outbound

This paper cites Textless speech-to-music retrieval using emotion similarity,.

Semi-Supervised Contrastive Learning for Controllable Video-to-Music Retrieval Textless speech-to-music retrieval using emotion similarity,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:21:45.548239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T20:21:45.172472Z digest=sha256:8e67f9582592f0a091471d332a522474d91e7b442b5bf6ea9f023abe961047ca

Observation 246ff8b7-8352-43c2-a5d1-7b7e745d4d3e · outbound

This paper cites Emotion-aligned contrastive learning between images and music,.

Semi-Supervised Contrastive Learning for Controllable Video-to-Music Retrieval Emotion-aligned contrastive learning between images and music,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:21:45.532940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T20:21:45.177502Z digest=sha256:c777bc3b115e36d22cd8d1d096578cd3706a10882d59f74c0694499136da3d5a

Observation aa6e237d-28f0-4c98-b99e-ea898419faf1 · outbound

This paper cites Bridging high-quality audio and video via language for sound effects retrieval from visual queries,.

Semi-Supervised Contrastive Learning for Controllable Video-to-Music Retrieval Bridging high-quality audio and video via language for sound effects retrieval from visual queries,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:21:45.515253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T20:21:45.182409Z digest=sha256:a8e91752e43dfb50cb96a6dac36e0a3851fd218141ec6575895a2fab5722b35e

Observation 87f43bf2-7912-433e-92ed-3b47121c36dd · outbound

This paper cites Query by video: Cross-modal music retrieval,.

Semi-Supervised Contrastive Learning for Controllable Video-to-Music Retrieval Query by video: Cross-modal music retrieval,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:21:45.499166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T20:21:45.187500Z digest=sha256:21ab48abc910723ea65a3a6e61192bf087eb8c88b683b4f1b17e870ba1e95978

Observation 8a1d4109-64f7-419f-a3c5-fae3900c02f8 · outbound

This paper cites Mert: Acoustic music understanding model with large-scale self-supervised training,.

Semi-Supervised Contrastive Learning for Controllable Video-to-Music Retrieval Mert: Acoustic music understanding model with large-scale self-supervised training,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:21:45.481239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T20:21:45.192570Z digest=sha256:42a9073ba81e50abd424fc2f3ba5e8299c613b8dfff674904e23cbdf12e6479c

Observation 5315cc00-07e9-4ec6-bf7b-31b9f8fee674 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Semi-Supervised Contrastive Learning for Controllable Video-to-Music Retrieval Learning transferable visual models from natural language supervision,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:21:45.461955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T20:21:45.197738Z digest=sha256:915a1159a0255ff25e7c41bfbf91e80ffee8687e14a812a6abbb9bfbb40f73cb

Observation 9a90eb1a-149d-486b-b3c9-d111a5cbc16b · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

Semi-Supervised Contrastive Learning for Controllable Video-to-Music Retrieval Representation Learning with Contrastive Predictive Coding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T20:21:45.202873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:21:45.202873Z digest=sha256:07823845aa78073a8810d779bcf10a3c0d69e409970723307dbebd0c08dc2dee

Observation 4c4248c1-2fea-4f18-a459-441c48197366 · outbound

This paper cites Supervised contrastive learn- ing,.

Semi-Supervised Contrastive Learning for Controllable Video-to-Music Retrieval Supervised contrastive learn- ing,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:21:45.444341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T20:21:45.209163Z digest=sha256:92375739ae7921c3643bffedcef58d06fcc347d0694f6315d94fe6d0e53c9290

Observation dc021c04-d0be-413c-92b5-28d481b10248 · outbound

This paper cites Audio set: An ontology and human-labeled dataset for audio events,.

Semi-Supervised Contrastive Learning for Controllable Video-to-Music Retrieval Audio set: An ontology and human-labeled dataset for audio events,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:21:45.425481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T20:21:45.214156Z digest=sha256:0d709216dccda4b98bbb98211252a5ae244d761c86082bdcc231d3d9191fe464

Observation e1bb36fb-4eb3-4685-8bea-62196cbe658e · outbound

This paper cites YouTube-8M: A Large-Scale Video Classification Benchmark.

Semi-Supervised Contrastive Learning for Controllable Video-to-Music Retrieval YouTube-8M: A Large-Scale Video Classification Benchmark

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T20:21:45.219469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:21:45.219469Z digest=sha256:faf07a9ce2001b6fb956aeb7ea385c1aece219ebef4c49cbb75257e0210898d7

Observation d33b1025-da5c-431d-bb30-330bd096f8ac · outbound

This paper cites Decoupled weight decay regularization,.

Semi-Supervised Contrastive Learning for Controllable Video-to-Music Retrieval Decoupled weight decay regularization,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:21:45.406609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T20:21:45.225626Z digest=sha256:c66138969f0bdfe6857051c3def87787271939f2ecc75ed7fe4042ad9a7482d0

Observation cebd0518-2118-4e0b-907d-5f280b6fcbe6 · outbound

This paper cites Wav2clip: Learning robust audio representations from clip,.

Semi-Supervised Contrastive Learning for Controllable Video-to-Music Retrieval Wav2clip: Learning robust audio representations from clip,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:21:45.389454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T20:21:45.230673Z digest=sha256:cabfe6c0a15a69b4a361faa9f9d96bd4d53a7274cbdf9d5027917b5a235c4085

Observation e8927204-5729-4347-aa63-2d6b62883056 · outbound

This paper cites Audioclip: Extending clip to image, text and audio,.

Semi-Supervised Contrastive Learning for Controllable Video-to-Music Retrieval Audioclip: Extending clip to image, text and audio,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:21:45.371868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T20:21:45.236467Z digest=sha256:8a5a5c7cb35bd7d35d3fb16472ad17d2eb171c767f6104eb2f9a131ffe5e34ca

Observation 056ff467-4796-400f-9baa-23b753fcbc4d · outbound

This paper cites Emotion embedding spaces for matching music to stories,.

Semi-Supervised Contrastive Learning for Controllable Video-to-Music Retrieval Emotion embedding spaces for matching music to stories,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:21:45.353573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T20:21:45.243223Z digest=sha256:29808febd48eaf313d3160b9954d7885d360b7a81cfa191e73f3be02b7c6bbd3

Observation 27f1c0d7-b391-4ae3-8296-0644829dff74 · outbound

This paper cites Vggsound: A large- scale audio-visual dataset,.

Semi-Supervised Contrastive Learning for Controllable Video-to-Music Retrieval Vggsound: A large- scale audio-visual dataset,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:21:45.335525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T20:21:45.248236Z digest=sha256:ee6c3e557d2ad4f72a2517df7c102020882703ba9236a1412757ff3ce4f9f18d

Pith citing papers

No inbound Pith citation observations are available.