Pith. sign in

Paper Citation Record · LEDGER

JenBridge: Adaptive Long-Form Video Soundtracking across Scene Transitions

As of 5 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 0 inbound Pith citation observations for arXiv:2606.01703.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.01703 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-28T13:10:29.213917Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

29 of 29 outbound references displayed

  • verified exact14
  • verified fuzzy0
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch4

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f61a8482-72d6-4445-a7ad-f9aa7142aef3 · outbound

This paper cites MusicLM: Generating Music From Text.

JenBridge: Adaptive Long-Form Video Soundtracking across Scene Transitions MusicLM: Generating Music From Text

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-02T00:56:24.676205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T13:10:29.213917Z digest=sha256:be2dcb452dc7f28014d1aac669dc4a8d30598ea44b2258b5454b93f4716c7881

Observation ddceba15-82dc-4022-9073-154217ac5ad3 · outbound

This paper cites Seed-Music: A Unified Framework for High Quality and Controlled Music Generation.

JenBridge: Adaptive Long-Form Video Soundtracking across Scene Transitions Seed-Music: A Unified Framework for High Quality and Controlled Music Generation

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-02T00:56:24.636627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T13:10:29.213917Z digest=sha256:91f3f56b2cd83d33163fe7d3102969c4a1d03c11675e8ef9be92db54399601bd

Observation cb41cb2c-d56d-442a-8bb9-eca7e396b313 · outbound

This paper cites Ke Chen, Yusong Wu, Haohe Liu, Marianna Nezhurina, Taylor Berg-Kirkpatrick, and Shlomo Dub- nov.

JenBridge: Adaptive Long-Form Video Soundtracking across Scene Transitions Ke Chen, Yusong Wu, Haohe Liu, Marianna Nezhurina, Taylor Berg-Kirkpatrick, and Shlomo Dub- nov

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-28T13:10:29.213917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T13:10:29.213917Z digest=sha256:c6e8fbb3bd362eaca4bf77a4ab0bcb0de42cf341333aa296bf1bc84524eabdcf

Observation 86268b62-024b-4408-829e-5de6a3adc9c4 · outbound

This paper cites MMTrail: A Multimodal Trailer Video Dataset with Language and Music Descriptions.

JenBridge: Adaptive Long-Form Video Soundtracking across Scene Transitions MMTrail: A Multimodal Trailer Video Dataset with Language and Music Descriptions

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-02T00:56:24.653940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T13:10:29.213917Z digest=sha256:b5876f844633adfee32b29869ad7dfc9f996d20563310c81bf0856e2d94febb4

Observation 1ff7d939-01d1-4d2c-b0ad-ad1ab179f453 · outbound

This paper cites High Fidelity Neural Audio Compression.

JenBridge: Adaptive Long-Form Video Soundtracking across Scene Transitions High Fidelity Neural Audio Compression

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-02T00:56:24.679750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T13:10:29.213917Z digest=sha256:b8d858036be149a988d16daffd42a81e23b38fdd35c1de43be22eb41b0e36491

Observation 4249648a-a684-458f-8ea1-dd3417311a96 · outbound

This paper cites Video background music generation with controllable music transformer.

JenBridge: Adaptive Long-Form Video Soundtracking across Scene Transitions Video background music generation with controllable music transformer

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-28T13:10:29.213917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T13:10:29.213917Z digest=sha256:490564e8cc1d5e71b4ebe3abcbae5229f016bc9620c834970c23b2c9fe882c8d

Observation a69dbd38-6052-4df7-a1a8-e5c018386d06 · outbound

This paper cites Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas M ¨uller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, et al.

JenBridge: Adaptive Long-Form Video Soundtracking across Scene Transitions Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas M ¨uller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, et al

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-28T13:10:29.213917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T13:10:29.213917Z digest=sha256:311c6f3d1ef964a567c01831cda4807016b2b4684ecd2553e09da1fafda2c8d2

Observation 669cfe5f-3763-43f4-bc93-f419d6045d16 · outbound

This paper cites Stable audio open.

JenBridge: Adaptive Long-Form Video Soundtracking across Scene Transitions Stable audio open

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-28T13:10:29.213917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T13:10:29.213917Z digest=sha256:e519bec33bd290fb31dff5cff26aff94a370f46551e726211f40e139349c80b7

Observation 2a59da7a-3d8f-4cc7-bd24-1a26e710e76d · outbound

This paper cites Cot-vtm: Visual-to-music genera- tion with chain-of-thought reasoning.

JenBridge: Adaptive Long-Form Video Soundtracking across Scene Transitions Cot-vtm: Visual-to-music genera- tion with chain-of-thought reasoning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-28T13:10:29.213917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T13:10:29.213917Z digest=sha256:24f179ca313fe0e633212beefbc1765ab7e759cb11e0e23997d3de0b74567298

Observation 6703c3df-9d62-4e89-bec4-759e24a71161 · outbound

This paper cites Noise2Music: Text-conditioned Music Generation with Diffusion Models.

JenBridge: Adaptive Long-Form Video Soundtracking across Scene Transitions Noise2Music: Text-conditioned Music Generation with Diffusion Models

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T00:56:24.705379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T13:10:29.213917Z digest=sha256:aa3e3b4f52466aca65df19d5fa66c959fbadbbcc2c1fc37cbe96ac1a70982e9e

Observation a115c7fe-2aa2-4387-b9eb-865217e755bb · outbound

This paper cites Analyzable Chain-of-Musical-Thought Prompting for High-Fidelity Music Generation.

JenBridge: Adaptive Long-Form Video Soundtracking across Scene Transitions Analyzable Chain-of-Musical-Thought Prompting for High-Fidelity Music Generation

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-02T00:56:24.670420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T13:10:29.213917Z digest=sha256:50bedbfff064f5520d56fc3c8bf93c2ad3b539238248e14d0c91807226ba8b9d

Observation 00bede47-b91c-4548-92a9-1b49e45abfc6 · outbound

This paper cites MuVi: Video-to-Music Generation with Semantic Alignment and Rhythmic Synchronization.

JenBridge: Adaptive Long-Form Video Soundtracking across Scene Transitions MuVi: Video-to-Music Generation with Semantic Alignment and Rhythmic Synchronization

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-02T00:56:24.684395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T13:10:29.213917Z digest=sha256:a2a355ed4de1546815748b86d462daf5ef5a518b3523bd039cdbd2514f8d912c

Observation b92ffc5f-ce7e-4359-adff-675c846ee8a6 · outbound

This paper cites Flow Matching for Generative Modeling.

JenBridge: Adaptive Long-Form Video Soundtracking across Scene Transitions Flow Matching for Generative Modeling

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-02T00:56:24.660049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T13:10:29.213917Z digest=sha256:466addb23cfb5664539db0e889739055babbe06bc27f8015b88fc589a879cfea

Observation 6f86704a-5c1a-4950-9d1b-daabb4d47bc4 · outbound

This paper cites MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models.

JenBridge: Adaptive Long-Form Video Soundtracking across Scene Transitions MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-02T00:56:24.675619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T13:10:29.213917Z digest=sha256:b4db86b02812fe123c9565d0910bc939386e2a53bb93ceb7fc50c5195b79b7fc

Observation 1e727a19-38b2-48e4-9157-e792c7716438 · outbound

This paper cites Extending Visual Dynamics for Video-to-Music Generation.

JenBridge: Adaptive Long-Form Video Soundtracking across Scene Transitions Extending Visual Dynamics for Video-to-Music Generation

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-02T00:56:24.696594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T13:10:29.213917Z digest=sha256:f55a5f7552f3953593b45659e0dda01f033218daa26d40408a4803b3df328cea

Observation 8b75afb5-e8bf-4f16-b92d-4f948b9b7f8d · outbound

This paper cites Mustango: Toward Controllable Text-to-Music Generation.

JenBridge: Adaptive Long-Form Video Soundtracking across Scene Transitions Mustango: Toward Controllable Text-to-Music Generation

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T00:56:24.665422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T13:10:29.213917Z digest=sha256:bf015f595f414379edb347d36fe8bca0c06cc46849c49ac9e78c3154d6991296

Observation 5ddaf623-e052-4bb6-b5c1-c152a59d9c27 · outbound

This paper cites Diff-A-Riff: Musical Accompaniment Co-creation via Latent Diffusion Models.

JenBridge: Adaptive Long-Form Video Soundtracking across Scene Transitions Diff-A-Riff: Musical Accompaniment Co-creation via Latent Diffusion Models

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-02T00:56:24.699439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T13:10:29.213917Z digest=sha256:eb585b8ab17d7823dbb7f21f0da778405e3d722cc3a9ff83decb37b00c074724

Observation 3dd3d0e3-df57-4f0d-be46-a97a7bd64f5f · outbound

This paper cites MusicFlow: Cascaded Flow Matching for Text Guided Music Generation.

JenBridge: Adaptive Long-Form Video Soundtracking across Scene Transitions MusicFlow: Cascaded Flow Matching for Text Guided Music Generation

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-02T00:56:24.701290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T13:10:29.213917Z digest=sha256:568a0e73ff726574e40a3ab17c27111fcdf56eff8ce7352d028aac6ad0be9763

Observation 576a224c-08a5-42c5-bb02-fa49935a6cad · outbound

This paper cites 11 Preprint.

JenBridge: Adaptive Long-Form Video Soundtracking across Scene Transitions 11 Preprint

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-28T13:10:29.213917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T13:10:29.213917Z digest=sha256:416f4fe6c8d8b3b861d5a086cc6b01e95c58e268cdeb500978f4d5a0f3e328f1

Observation 3f1586e7-f484-490f-b2f0-500017c520b1 · outbound

This paper cites AudioX: A Unified Framework for Anything-to-Audio Generation.

JenBridge: Adaptive Long-Form Video Soundtracking across Scene Transitions AudioX: A Unified Framework for Anything-to-Audio Generation

Reference 20

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T00:56:24.687730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T13:10:29.213917Z digest=sha256:6243b75649f6265bff236dce1ef18a906ba4ed2950c6c475e8d2ca4c7f2e83de

Observation 652bcbd5-0627-42ba-aff6-1e6c2c274e90 · outbound

This paper cites Zhifeng Xie, Qile He, Youjia Zhu, Qiwei He, and Mengtian Li.

JenBridge: Adaptive Long-Form Video Soundtracking across Scene Transitions Zhifeng Xie, Qile He, Youjia Zhu, Qiwei He, and Mengtian Li

Reference 21

Resolution
unresolved
no resolver link, observed 2026-06-28T13:10:29.213917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T13:10:29.213917Z digest=sha256:07cbfe5c82be40897bc624cc2d6042bde25d76d92ac1511d37d3c7714237bdc1

Observation 4a7d42f3-fa1e-4178-8676-8fdfaffcf970 · outbound

This paper cites Qwen3 Technical Report.

JenBridge: Adaptive Long-Form Video Soundtracking across Scene Transitions Qwen3 Technical Report

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-07-02T00:56:24.672571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T13:10:29.213917Z digest=sha256:6abbb5791f5ba531fb41f2059a6b1bbe5dc2c740d9146758c9dfe7b06443356d

Observation 96e38a82-1309-41eb-b199-282bd09eb9ae · outbound

This paper cites Yue: Scaling open foundation models for long-form music generation.arXiv:2503.08638.

JenBridge: Adaptive Long-Form Video Soundtracking across Scene Transitions Yue: Scaling open foundation models for long-form music generation.arXiv:2503.08638

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-02T00:56:24.668818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T13:10:29.213917Z digest=sha256:b1a9075d3f123cc2f9dae85c5d480fcd6f0377a2a07ea7c0dffc85d398b98458

Observation ef41e6cb-56ec-4f6c-94d8-4c1e4c9854ef · outbound

This paper cites LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models.

JenBridge: Adaptive Long-Form Video Soundtracking across Scene Transitions LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models

Reference 24

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T00:56:24.687977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T13:10:29.213917Z digest=sha256:f98516a92693a2ad3522f917f7f0ffb6e1f7a9c758d724ffdb518bfa82f9af83

Observation 1d24e822-2186-49e3-8437-1816b78d467f · outbound

This paper cites ERNIE-Music: Text-to-Waveform Music Generation with Diffusion Models.

JenBridge: Adaptive Long-Form Video Soundtracking across Scene Transitions ERNIE-Music: Text-to-Waveform Music Generation with Diffusion Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-02T00:56:24.695561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T13:10:29.213917Z digest=sha256:c8a5fdfd05a9a5b76871b016f715cbce45f4fa84671eafab56de68f8f103b4a7

Observation a5839b60-8cfb-4de2-9e54-f3a868dd356e · outbound

This paper cites A STATEMENTS ANDBROADERIMPACT A.1 ETHICSSTATEMENT Data Usage.Our work adheres to strict ethical guidelines regarding data usage.

JenBridge: Adaptive Long-Form Video Soundtracking across Scene Transitions A STATEMENTS ANDBROADERIMPACT A.1 ETHICSSTATEMENT Data Usage.Our work adheres to strict ethical guidelines regarding data usage

Reference 26

Resolution
unresolved
no resolver link, observed 2026-06-28T13:10:29.213917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T13:10:29.213917Z digest=sha256:dbbe3cfc06f56d39b8a7816c4caef2a3481f7f5919ea99cbf4e5f1b929c7072e

Observation c3cc47e3-f7ff-4e62-97fb-53aaf2d62485 · outbound

This paper cites However, to facilitate further research and application, we will provide public API ac- cess to our foundational text-to-music model.

JenBridge: Adaptive Long-Form Video Soundtracking across Scene Transitions However, to facilitate further research and application, we will provide public API ac- cess to our foundational text-to-music model

Reference 27

Resolution
unresolved
no resolver link, observed 2026-06-28T13:10:29.213917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T13:10:29.213917Z digest=sha256:da8e6fd904110c64bd5ab9fa50de4a0678da971db371c570b838056713abbc39

Observation d13a7461-08e3-4ecf-a794-3e881f4ed1d9 · outbound

This paper cites This configuration allows us to partition long videos into a sequence of meaningful, temporally substantial clips suitable for individual soundtracking.

JenBridge: Adaptive Long-Form Video Soundtracking across Scene Transitions This configuration allows us to partition long videos into a sequence of meaningful, temporally substantial clips suitable for individual soundtracking

Reference 28

Resolution
unresolved
no resolver link, observed 2026-06-28T13:10:29.213917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T13:10:29.213917Z digest=sha256:c770a6987e98023e72e5995686eaae0160584d0db85558f18668d70806fd8994

Observation 6764d0ca-b4f1-4a3f-976f-d2f72346f71b · outbound

This paper cites ground-truth.

JenBridge: Adaptive Long-Form Video Soundtracking across Scene Transitions ground-truth

Reference 29

Resolution
unresolved
no resolver link, observed 2026-06-28T13:10:29.213917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T13:10:29.213917Z digest=sha256:7cfce8f128e99272f049eae77b406891edac3f148a3d8aeeeb97c8fd0ba72a7e

Pith citing papers

No inbound Pith citation observations are available.