← back to paper
arxiv: 2607.06405 · 2 revisions
Precise Video-to-Audio Generation with Cross-Modal Alignment in Latent Space