Pith. sign in

Paper Citation Record · LEDGER

Towards Video to Piano Music Generation with Chain-of-Perform Support Benchmarks

As of 8 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 0 inbound Pith citation observations for arXiv:2505.20038.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.20038 v1

Coverage vector

measured 25 of 25 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:05:12.150659Z

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

25 of 25 outbound references displayed

  • verified exact3
  • verified fuzzy8
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5001ff69-54ec-4cfe-8d99-1f5e3c17a8b5 · outbound

This paper cites Video-Guided Foley Sound Generation with Multimodal Controls.

Towards Video to Piano Music Generation with Chain-of-Perform Support Benchmarks Video-Guided Foley Sound Generation with Multimodal Controls

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:05:09.819840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:05:09.819840Z digest=sha256:2573980e06c205973546914d00f8642406aac7381fc9aa01f93527251512e46b

Observation bade64f8-cec2-4ab2-ab56-2e7f85d2b730 · outbound

This paper cites YingSound: Video-Guided Sound Effects Generation with Multi-modal Chain-of-Thought Controls.

Towards Video to Piano Music Generation with Chain-of-Perform Support Benchmarks YingSound: Video-Guided Sound Effects Generation with Multi-modal Chain-of-Thought Controls

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:05:12.786853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:05:09.874214Z digest=sha256:ab4257765cbe49af69ac2328935254f37f56e6b2166373859742326b6a0b8225

Observation 2309e642-f146-464a-a083-36fb2b903a66 · outbound

This paper cites MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis.

Towards Video to Piano Music Generation with Chain-of-Perform Support Benchmarks MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:05:09.943750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:05:09.943750Z digest=sha256:885817be3565a98f507d243ece1f841f88ebfe8e39c223d6ca1e8af4e4cd8403

Observation 0d41f6e1-c5b0-4aba-8b10-537ae6f9645d · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Towards Video to Piano Music Generation with Chain-of-Perform Support Benchmarks DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:05:10.009315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:05:10.009315Z digest=sha256:c07e74fd9820cb179e34d0264f0766aff116efb009adfe37cd27cd9a87d5dd42

Observation 26bf0cc3-e0af-410d-8e83-5b219d1d5a88 · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilib- rium.

Towards Video to Piano Music Generation with Chain-of-Perform Support Benchmarks Gans trained by a two time-scale update rule converge to a local nash equilib- rium

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:05:10.132451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:05:10.132451Z digest=sha256:c0fea969db1e02b37ff0926e573af9052b92b4fca1da55e7db325bbaebed6e6e

Observation be61c9ec-bfb7-4bdc-9141-01856861c8e0 · outbound

This paper cites Taming visually guided sound generation.

Towards Video to Piano Music Generation with Chain-of-Perform Support Benchmarks Taming visually guided sound generation

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:05:14.505860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:05:10.225285Z digest=sha256:5476f1acde7ccffb5185545b721182e496d4d11f2785d891865ebe8dcd98a67d

Observation 0562a702-3591-42da-91a1-9f6291dbbe66 · outbound

This paper cites Sophia Koepke, Olivia Wiles, Yael Moses, and Andrew Zisserman.

Towards Video to Piano Music Generation with Chain-of-Perform Support Benchmarks Sophia Koepke, Olivia Wiles, Yael Moses, and Andrew Zisserman

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:05:14.327614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:05:10.374623Z digest=sha256:b9942e803e0aabf93de9eea484e315c0a77568e451cf9b691dde51b7e1edccdf

Observation 84250236-4a2f-410b-8b7e-80c0110144ff · outbound

This paper cites Crandall, and Christopher Raphael.

Towards Video to Piano Music Generation with Chain-of-Perform Support Benchmarks Crandall, and Christopher Raphael

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:05:14.101676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:05:10.495877Z digest=sha256:acd74beb175965a073e965d5ecfcfcbc0e8966b567e70980b2a26c94157b933a

Observation 8e7bcbcc-767a-40d2-ae68-369c5ecf4d75 · outbound

This paper cites Tri-Ergon: Fine-grained Video-to-Audio Generation with Multi-modal Conditions and LUFS Control.

Towards Video to Piano Music Generation with Chain-of-Perform Support Benchmarks Tri-Ergon: Fine-grained Video-to-Audio Generation with Multi-modal Conditions and LUFS Control

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:05:12.592926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:05:10.608172Z digest=sha256:acbe89f0fceec92d2065c6f9d7ffbeb81edd62525059ab4c4f4dc455e6982515

Observation 58c5112c-09e0-4fee-8b63-9ea593e6c88e · outbound

This paper cites Diff-foley: Synchronized video-to-audio synthesis with la- tent diffusion models, 2023.

Towards Video to Piano Music Generation with Chain-of-Perform Support Benchmarks Diff-foley: Synchronized video-to-audio synthesis with la- tent diffusion models, 2023

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:05:13.840027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:05:10.693496Z digest=sha256:29ad8306d68105cf17cff6ee4102b111215a38b917dfa3178857248227d8ea1f

Observation 142adb7a-b282-4e1f-b92e-0b5caa682562 · outbound

This paper cites Foleygen: Visually-guided audio generation.

Towards Video to Piano Music Generation with Chain-of-Perform Support Benchmarks Foleygen: Visually-guided audio generation

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:05:13.604068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:05:10.818067Z digest=sha256:e343c284b4877eb32e404707c71b5041fe864246583458dca59ce86dd52801cc

Observation 5726a169-6270-466d-b1fa-5fbad09bc8d6 · outbound

This paper cites Qwen2.5 technical report, 2025.

Towards Video to Piano Music Generation with Chain-of-Perform Support Benchmarks Qwen2.5 technical report, 2025

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:05:13.403506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:05:10.960063Z digest=sha256:59cd9f3211845f177a772f167e34577409e645ab9f1bc1b27f7af1608ed15ee1

Observation 571820b7-73f0-4787-ac9a-5534892b862e · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Towards Video to Piano Music Generation with Chain-of-Perform Support Benchmarks Direct preference optimization: Your language model is secretly a reward model

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:05:11.035226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:05:11.035226Z digest=sha256:fbc5fdf78493fa6a80b69a66690d0571f22c38ebade3d5ce42205a60e5e124e8

Observation 0c40d3e8-45a1-4301-9083-8ca8d04573a2 · outbound

This paper cites Improved techniques for training gans.

Towards Video to Piano Music Generation with Chain-of-Perform Support Benchmarks Improved techniques for training gans

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:05:11.152150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:05:11.152150Z digest=sha256:2f2088f55df193cb50ce3a2c3238987d2a834579c6cc14e69c5c3fa5ad35a5d2

Observation 2fde6f11-e80f-45b7-9201-2879022d1fe5 · outbound

This paper cites Audeo: Audio Generation for a Silent Performance Video.

Towards Video to Piano Music Generation with Chain-of-Perform Support Benchmarks Audeo: Audio Generation for a Silent Performance Video

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:05:12.402036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:05:11.255279Z digest=sha256:0a0b932cfa607a8de929a8795480cf1fa81b2c2a838acaa95d4ab38297afc046

Observation bfde4dd2-da87-4953-a989-e1eb11484bd0 · outbound

This paper cites AudioX: A Unified Framework for Anything-to-Audio Generation.

Towards Video to Piano Music Generation with Chain-of-Perform Support Benchmarks AudioX: A Unified Framework for Anything-to-Audio Generation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:05:11.305127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:05:11.305127Z digest=sha256:785aecf8dd8fdd9258534d734a5076e3b35a5d627c76274d33692009009c48eb

Observation 8aab714f-f38e-46bd-abbc-3c3d4a717b46 · outbound

This paper cites Temporally Aligned Audio for Video with Autoregression.

Towards Video to Piano Music Generation with Chain-of-Perform Support Benchmarks Temporally Aligned Audio for Video with Autoregression

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:05:11.397850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:05:11.397850Z digest=sha256:9195ab7d15873ad406bf2d7067eadab1096c20f78962e8cee3380b4cbc97e135

Observation 4163fa1e-cbdb-4479-8a77-037d6e2d1217 · outbound

This paper cites V2a-mapper: A lightweight solution for vision-to-audio generation by connecting foun- dation models, 2023.

Towards Video to Piano Music Generation with Chain-of-Perform Support Benchmarks V2a-mapper: A lightweight solution for vision-to-audio generation by connecting foun- dation models, 2023

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:05:13.199507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:05:11.478150Z digest=sha256:f86009f63d2072689b979e451f09da7f0a9eb59368b6592c62b485d76c27d6d7

Observation c847c1af-9770-42ea-b268-32cd39ee6b86 · outbound

This paper cites Frieren: Efficient Video-to-Audio Generation Network with Rectified Flow Matching.

Towards Video to Piano Music Generation with Chain-of-Perform Support Benchmarks Frieren: Efficient Video-to-Audio Generation Network with Rectified Flow Matching

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:05:11.602718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:05:11.602718Z digest=sha256:5b3ee1b9d89a10c502f4b876b2c3d11cebc8aaef276b5092245c8c6393ebeaf6

Observation d01e4aff-f9f4-49a1-abdd-c810a9e3ff68 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large lan- guage models.

Towards Video to Piano Music Generation with Chain-of-Perform Support Benchmarks Chain-of-thought prompting elicits reasoning in large lan- guage models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:05:11.717744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:05:11.717744Z digest=sha256:acafe2796fb065e64ec89169b65f25885b240950807dc9d3b8c48dea3032bcc3

Observation 04a16f20-97f0-4ae9-ba51-3c4b66cfaf48 · outbound

This paper cites Large-scale con- trastive language-audio pretraining with feature fusion and keyword-to-caption augmentation.

Towards Video to Piano Music Generation with Chain-of-Perform Support Benchmarks Large-scale con- trastive language-audio pretraining with feature fusion and keyword-to-caption augmentation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:05:11.786319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:05:11.786319Z digest=sha256:06db733cae5e87bfdeee039c52d7eb95289ea2dcc4519afa5f4416ba32ab4e7f

Observation c82a12f4-a32d-4160-8063-9d5fe817a8ea · outbound

This paper cites LLaVA-CoT: Let Vision Language Models Reason Step-by-Step.

Towards Video to Piano Music Generation with Chain-of-Perform Support Benchmarks LLaVA-CoT: Let Vision Language Models Reason Step-by-Step

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:05:11.869274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:05:11.869274Z digest=sha256:7c54714a0c35c0b6e84cc955430269f9910301f51eb03ab51bf6f325f91b7297

Observation 408206cd-881b-4eb6-be8a-b9c0006f1583 · outbound

This paper cites Diverse and aligned audio-to-video genera- tion via text-to-video model adaptation.

Towards Video to Piano Music Generation with Chain-of-Perform Support Benchmarks Diverse and aligned audio-to-video genera- tion via text-to-video model adaptation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:05:12.973198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:05:11.940525Z digest=sha256:70d8e2648bb7002fde326a233cc55909bfbf58a4da53a370274dd8b18bda2ee8

Observation 10a5b411-b034-47a1-b8e3-b003ce232157 · outbound

This paper cites Improve Vision Language Model Chain-of-thought Reasoning.

Towards Video to Piano Music Generation with Chain-of-Perform Support Benchmarks Improve Vision Language Model Chain-of-thought Reasoning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:05:12.046711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:05:12.046711Z digest=sha256:a5549736d5f63e762d4096af014d041f0796bd17ee583954fb1b6d94be075170

Observation e5935e36-2c37-445e-81ae-9d1e08b05f4c · outbound

This paper cites FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds.

Towards Video to Piano Music Generation with Chain-of-Perform Support Benchmarks FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:05:12.150659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:05:12.150659Z digest=sha256:05fdf4357a6046224b9994c2d5a15d4a1771c4199741e9a1738647391e704880

Pith citing papers

No inbound Pith citation observations are available.