Pith. sign in

Paper Citation Record · LEDGER

Precise Video-to-Audio Generation with Cross-Modal Alignment in Latent Space

As of 4 August 2026, this Paper Citation Record lists 17 of 17 outbound references and 0 inbound Pith citation observations for arXiv:2607.06405.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.06405 v2

Coverage vector

measured 17 of 17 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T08:23:35.864870Z

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

17 of 17 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 88100169-bcc7-4dbb-aff8-d976a1520203 · outbound

This paper cites Qwen2.5-VL Technical Report.

Precise Video-to-Audio Generation with Cross-Modal Alignment in Latent Space Qwen2.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T08:23:34.359214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:23:34.359214Z digest=sha256:9783c7d09ba148add6d2196afae19895a654a67ee752c4b194580fed88e64332

Observation 9efdea3c-64cf-4fce-a7f2-22e83f54a36b · outbound

This paper cites In: ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP).

Precise Video-to-Audio Generation with Cross-Modal Alignment in Latent Space In: ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T08:23:34.515487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:23:34.515487Z digest=sha256:0e642a517678ee1a9e64e936bb5d540422deab1a276b16ce7ae1c3930347b964

Observation 986b53e8-81e5-4d66-b799-8dcce7ac779e · outbound

This paper cites an unresolved cited work.

Precise Video-to-Audio Generation with Cross-Modal Alignment in Latent Space Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T08:23:34.575467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:23:34.575467Z digest=sha256:3a72aa3cde0b9a86a23fc8bdca17ec334f1d4056955f39698c228baceda26760

Observation dd93d474-4552-464a-87cc-7365a6fbb5b0 · outbound

This paper cites an unresolved cited work.

Precise Video-to-Audio Generation with Cross-Modal Alignment in Latent Space Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T08:23:34.698219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:23:34.698219Z digest=sha256:003e5f63a875460c31451f261724b2de1e75518207f58be3b16f8cda50c3d3ea

Observation 27a2f274-2780-43fe-9a78-09fec54f6d58 · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).

Precise Video-to-Audio Generation with Cross-Modal Alignment in Latent Space In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T08:23:34.836848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:23:34.836848Z digest=sha256:017e2d641d457a8729a23075d2c91844909b32dbafb25183411e0d21e8b3f9f8

Observation 07c5f217-b934-4d1c-85f3-09be015a52ef · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).

Precise Video-to-Audio Generation with Cross-Modal Alignment in Latent Space In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T08:23:34.955335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:23:34.955335Z digest=sha256:975993abf434f5170d7ff9738876a652e7c8952b5ceba67935c855a66fd03825

Observation 740dc824-80f1-49f6-9798-fa92caf04a17 · outbound

This paper cites In: Proceedings of the Computer Vision and Pattern Recognition Conference (CVPR).

Precise Video-to-Audio Generation with Cross-Modal Alignment in Latent Space In: Proceedings of the Computer Vision and Pattern Recognition Conference (CVPR)

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T08:23:35.058989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:23:35.058989Z digest=sha256:048cdba1de70466880b0db15b783cc64f4615637df00d345b89d8524c2731adb

Observation b62be7cb-31ca-4094-bd2a-8f81f522614b · outbound

This paper cites In: International Conference on Learning Representations (2019),https://openreview.net/forum? id=Bkg6RiCqY7.

Precise Video-to-Audio Generation with Cross-Modal Alignment in Latent Space In: International Conference on Learning Representations (2019),https://openreview.net/forum? id=Bkg6RiCqY7

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T08:23:35.188003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:23:35.188003Z digest=sha256:6fc2428c181d763a20aab8788fec77f84cc6b5be1603cea71908e9e7201208cc

Observation c337599d-5ab4-4a47-9baf-579cacb0d55f · outbound

This paper cites In: Oh, A., Naumann, T., Globerson, A., Saenko, K., Hardt, M., Levine, S.

Precise Video-to-Audio Generation with Cross-Modal Alignment in Latent Space In: Oh, A., Naumann, T., Globerson, A., Saenko, K., Hardt, M., Levine, S

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T08:23:35.264243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:23:35.264243Z digest=sha256:9b1c0121806ef7335dd17e8752df06327a8e4077745f13fe900ee6083fea1fa0

Observation 096338fd-9d86-4e64-83bc-40114820ceed · outbound

This paper cites In: Meila, M., Zhang, T.

Precise Video-to-Audio Generation with Cross-Modal Alignment in Latent Space In: Meila, M., Zhang, T

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T08:23:35.310440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:23:35.310440Z digest=sha256:39ae63de245d68682ed3ba17c94956ea253f89cd71bddc8259f5c85d72a309cc

Observation 870ddedc-1cc3-4d0a-a2b5-97c3fa09eb7a · outbound

This paper cites In: The Thirteenth International Conference on Learning Representations (2025),https: //openreview.net/forum?id=yFEqYwgttJ.

Precise Video-to-Audio Generation with Cross-Modal Alignment in Latent Space In: The Thirteenth International Conference on Learning Representations (2025),https: //openreview.net/forum?id=yFEqYwgttJ

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T08:23:35.369837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:23:35.369837Z digest=sha256:bdecf9aa0948dd650c1430cdf91c74c4cd37596553230c501a89766a53ca3de7

Observation a2a3e85e-7492-48b6-bab2-9b7948240738 · outbound

This paper cites Movie Gen: A Cast of Media Foundation Models.

Precise Video-to-Audio Generation with Cross-Modal Alignment in Latent Space Movie Gen: A Cast of Media Foundation Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T08:23:35.411862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:23:35.411862Z digest=sha256:62b0fe0b09f15dbc9220dd97baa34a5c233b7ffad6266969873dcf0dbccc5dc4

Observation 3193c7ce-fe66-4d3d-ac33-69eefe42ce98 · outbound

This paper cites (eds.) Proceedings of the 41st International Conference on Machine Learning.

Precise Video-to-Audio Generation with Cross-Modal Alignment in Latent Space (eds.) Proceedings of the 41st International Conference on Machine Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T08:23:35.474923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:23:35.474923Z digest=sha256:08664d7a13e8bfa4c7061b30f89a52929a76d75584d58c745f694643993ab7f5

Observation 7a14b2d7-3fd6-47d1-b0e7-cfc164fa75c3 · outbound

This paper cites an unresolved cited work.

Precise Video-to-Audio Generation with Cross-Modal Alignment in Latent Space Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T08:23:35.567854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:23:35.567854Z digest=sha256:c04a39bbc186d8e7a56f0161ebd2e05d4d71c13f92ba0ac933b2154e629a2e6b

Observation d1e879e2-5562-4ad0-8afc-cc2bf0240421 · outbound

This paper cites In: Proceed- ings of the Computer Vision and Pattern Recognition Conference (CVPR).

Precise Video-to-Audio Generation with Cross-Modal Alignment in Latent Space In: Proceed- ings of the Computer Vision and Pattern Recognition Conference (CVPR)

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T08:23:35.658546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:23:35.658546Z digest=sha256:026223463ba2ff5adf607e939161f6cb5ba391b456212efd6095eec31facc0fd

Observation 76871a0a-0c5e-46a0-9431-0b56cf431eb9 · outbound

This paper cites In: Globerson, A., Mackey, L., Belgrave, D., Fan, A., Paquet, U., Tomczak, J., Zhang, C.

Precise Video-to-Audio Generation with Cross-Modal Alignment in Latent Space In: Globerson, A., Mackey, L., Belgrave, D., Fan, A., Paquet, U., Tomczak, J., Zhang, C

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T08:23:35.762401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:23:35.762401Z digest=sha256:15ff9637093420f1cc7894da1ac5160669705815ab9ae2e55f87064e212fa503

Observation bdea5d69-94cc-4cd0-a427-0d4cbc8e5e53 · outbound

This paper cites FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds.

Precise Video-to-Audio Generation with Cross-Modal Alignment in Latent Space FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T08:23:35.864870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:23:35.864870Z digest=sha256:a013affc87969b132d7d0e0648c3e881579a9f6680cf674a446abd7af7d862ea

Pith citing papers

No inbound Pith citation observations are available.