Pith. sign in

Paper Citation Record · LEDGER

Multimodal Video-to-Music Recommendation via Semantic Retrieval and Temporal Reranking

As of 23 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 0 inbound Pith citation observations for arXiv:2607.05971.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.05971 v1

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-08T19:18:56.334254Z

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

18 of 18 outbound references displayed

  • verified exact12
  • verified fuzzy2
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 259b61b3-d44e-4231-bb7a-c2add02cbdb5 · outbound

This paper cites Qwen3-VL Technical Report.

Multimodal Video-to-Music Recommendation via Semantic Retrieval and Temporal Reranking Qwen3-VL Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-08T19:25:32.389946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-08T19:18:56.334254Z digest=sha256:ce0acc4e15fbd51e4eb0b70ea72f80407501ebc351045557e3951890fd4e378d

Observation 7ebe0395-9509-48d8-92db-6cd7385b7699 · outbound

This paper cites PianoBind: A Multimodal Joint Embedding Model for Pop-piano Music.

Multimodal Video-to-Music Recommendation via Semantic Retrieval and Temporal Reranking PianoBind: A Multimodal Joint Embedding Model for Pop-piano Music

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-08T19:25:32.404599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-08T19:18:56.334254Z digest=sha256:5471ec73f56d6ef2cf622270a9ce8a50acd2b19d13f88e25b8100768ee968c88

Observation e97f020d-7ff6-49d1-bfc2-b5c8bb65b8ab · outbound

This paper cites MMTrail: A Multimodal Trailer Video Dataset with Language and Music Descriptions.

Multimodal Video-to-Music Recommendation via Semantic Retrieval and Temporal Reranking MMTrail: A Multimodal Trailer Video Dataset with Language and Music Descriptions

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-08T19:25:32.386908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-08T19:18:56.334254Z digest=sha256:9e1252218a147647fcd6093dcc05ce52e4a1ba82473344e58661e4d4b33d76de

Observation 0f43de9d-3a28-49f7-bcda-b161e1b38bc6 · outbound

This paper cites Listen, Read, and Identify: Multimodal Singing Language Identification of Music.

Multimodal Video-to-Music Recommendation via Semantic Retrieval and Temporal Reranking Listen, Read, and Identify: Multimodal Singing Language Identification of Music

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-08T19:25:32.392550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-08T19:18:56.334254Z digest=sha256:caaa4111d39ce9cea687419942404e4e9ee11aadaea2d2eb21c6050cae82e047

Observation 62e5769b-1db1-4242-b4a5-304f1d3eec2a · outbound

This paper cites LP-MusicCaps: LLM-Based Pseudo Music Captioning.

Multimodal Video-to-Music Recommendation via Semantic Retrieval and Temporal Reranking LP-MusicCaps: LLM-Based Pseudo Music Captioning

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-08T19:25:32.409284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-08T19:18:56.334254Z digest=sha256:b50b006557d1e86920b01b861a3639fdc214cfdd0187365a074eeabe1334119d

Observation df24de1f-78d8-4daf-97ae-00ce9be4a859 · outbound

This paper cites TALKPLAY: Multimodal Music Recommendation with Large Language Models.

Multimodal Video-to-Music Recommendation via Semantic Retrieval and Temporal Reranking TALKPLAY: Multimodal Music Recommendation with Large Language Models

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-08T19:25:32.416406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-08T19:18:56.334254Z digest=sha256:02a0446a66e6259105898090cd891a81ca83ad91aec5c2b2ec0f9c2fb1563d5b

Observation 106044c0-c970-42c3-b7b0-55d5f5e293d5 · outbound

This paper cites Music Flamingo: Scaling music understanding in audio language models.

Multimodal Video-to-Music Recommendation via Semantic Retrieval and Temporal Reranking Music Flamingo: Scaling music understanding in audio language models

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-07-08T19:25:32.412044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-08T19:18:56.334254Z digest=sha256:d22fe7005084af99426733d8bcf5942f51dd26dd8584c5fd3270b04b8b35dfcf

Observation 11b251ca-0c62-4049-bc12-f094ec43e4bc · outbound

This paper cites Audioclip: Extending clip to image, text and audio.

Multimodal Video-to-Music Recommendation via Semantic Retrieval and Temporal Reranking Audioclip: Extending clip to image, text and audio

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T21:15:39.624144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-08T19:18:56.334254Z digest=sha256:39c167f86e26d701af3f67977f7b0cd7b20c3b1a6ba6b60ec64cf1e81952484f

Observation 63d8c160-3af7-4b67-911a-0382a29a9226 · outbound

This paper cites Poly-encoders: Transformer Architectures and Pre-training Strategies for Fast and Accurate Multi-sentence Scoring.

Multimodal Video-to-Music Recommendation via Semantic Retrieval and Temporal Reranking Poly-encoders: Transformer Architectures and Pre-training Strategies for Fast and Accurate Multi-sentence Scoring

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-08T19:25:32.394949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-08T19:18:56.334254Z digest=sha256:0c246adb982b9e8d3887cb0e58b436e1293bb59c7ebe4e567959a6361e58c4d4

Observation 24cca5ac-b8cd-433b-a52c-17b234aa06ae · outbound

This paper cites Qwen3-ASR Technical Report.

Multimodal Video-to-Music Recommendation via Semantic Retrieval and Temporal Reranking Qwen3-ASR Technical Report

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-08T19:25:32.397341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-08T19:18:56.334254Z digest=sha256:3db07e5baf9a028bfdd2b1c3e746b6a19de3f9215d0fe1c3b72587958e09ae84

Observation 54100690-59be-4a42-be60-535360c2eac1 · outbound

This paper cites Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound.

Multimodal Video-to-Music Recommendation via Semantic Retrieval and Temporal Reranking Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-08T19:25:32.401997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-08T19:18:56.334254Z digest=sha256:8fdca03617fea9387d9571d064280e0a49d3055bee43bdab46b533c8d61e9b8d

Observation 6446a29d-988b-4712-9e70-91a565430759 · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

Multimodal Video-to-Music Recommendation via Semantic Retrieval and Temporal Reranking SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 12

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T19:25:32.419925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-08T19:18:56.334254Z digest=sha256:7ed92c923efaa08eca3bbdd28d34638e76fc2cf2ecfafbb0e96f1f75cc6b4eb5

Observation f3493831-43e7-471f-b24c-1c4ba103c559 · outbound

This paper cites Pushing the frontier of audiovisual perception with large-scale multimodal correspondence learning.

Multimodal Video-to-Music Recommendation via Semantic Retrieval and Temporal Reranking Pushing the frontier of audiovisual perception with large-scale multimodal correspondence learning

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-08T19:25:32.423599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-08T19:18:56.334254Z digest=sha256:fda5c5a3850b434b3d128f81183f6c856edea8a40d5521ff47e70746560caa28

Observation 93fc22e5-8c88-4a8e-8711-d2b5fcc2d0e9 · outbound

This paper cites an unresolved cited work.

Multimodal Video-to-Music Recommendation via Semantic Retrieval and Temporal Reranking Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-07-08T21:15:39.620897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-08T19:18:56.334254Z digest=sha256:1ab0fcf73bf6a2199534d2ec62daab6156a2b331577293cb58c6d2df27d9ecce

Observation 40586d8e-e555-4f52-864b-102e45a0dc0a · outbound

This paper cites CLaMP 3: Universal Music Information Retrieval Across Unaligned Modalities and Unseen Languages.

Multimodal Video-to-Music Recommendation via Semantic Retrieval and Temporal Reranking CLaMP 3: Universal Music Information Retrieval Across Unaligned Modalities and Unseen Languages

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-08T19:25:32.428257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-08T19:18:56.334254Z digest=sha256:7821ebdffbf4bba380de6bc727fecd05a6a2760e9f866a43e340e7248a0a31eb

Observation c1a56a38-f603-44a6-a0c4-d078f2299d2f · outbound

This paper cites Qwen3 Technical Report.

Multimodal Video-to-Music Recommendation via Semantic Retrieval and Temporal Reranking Qwen3 Technical Report

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-08T19:25:32.432408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-08T19:18:56.334254Z digest=sha256:f703a9fbe6eab9b82989e26a4656668035acbaaf984714f1d9c611084a876aa3

Observation afb22485-52a9-41b0-8329-d4f9cdbb7685 · outbound

This paper cites C-Pack: Packed Resources For General Chinese Embeddings.

Multimodal Video-to-Music Recommendation via Semantic Retrieval and Temporal Reranking C-Pack: Packed Resources For General Chinese Embeddings

Reference 17

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T19:25:32.438058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-08T19:18:56.334254Z digest=sha256:27fbaabe148981c7fc6c6095a05307f1a7f2fdffc967264cb2db0523f503b8f3

Observation 00597ec6-990a-441e-9ce9-3dc155acaa75 · outbound

This paper cites Languagebind: Extending video-language pretraining to n-modality by language-based semantic alignment.

Multimodal Video-to-Music Recommendation via Semantic Retrieval and Temporal Reranking Languagebind: Extending video-language pretraining to n-modality by language-based semantic alignment

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T21:15:39.622530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-08T19:18:56.334254Z digest=sha256:c48363f131f61753776791ef95d1edf502c4fcda77478d558599500abecb22ad

Pith citing papers

No inbound Pith citation observations are available.