Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T11:38:08.713130Z
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 100 of 105 outbound references and 1 inbound Pith citation observation for arXiv:2412.15191.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T11:38:08.713130Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-03T16:25:56.032167Z
A source-named dated measurement, never combined with another source.
Source: cited_works
100 of 105 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation d296dfca-7f78-4ed4-ad45-9d27d9470b9d · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation Vid- styleode: Disentangled video editing via stylegan and neu- ralodes
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9359a8d-e7c7-4f8d-ab45-c9513933f26d · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation Label-Efficient Semantic Segmentation with Diffusion Models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d83f8c8f-4a5a-4581-83f9-9467f2cdbec1 · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation SonicDiffusion: Audio-Driven Image Generation and Editing with Pretrained Diffusion Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d15ec61-52dc-4fb9-9dc3-41e1931aa447 · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c14217bb-83e5-40f8-8dfb-f277c4c48743 · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation Align your latents: High-resolution video synthesis with latent diffusion models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 190444a4-fb82-4e21-a2f1-f5700c7d9b94 · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation The mtg-jamendo dataset for au- tomatic music tagging
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7e5e277-678d-407e-9756-72b5901cfd16 · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation Video generation models as world simulators
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c65a41bc-86c4-4fb1-8037-5e0af3aaae0e · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation Ac- tion2sound: Ambient-aware generation of action sounds from egocentric videos
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ebe36cc-7eac-4a74-8eff-e592fb698c62 · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation Semantically consistent video-to-audio generation using multimodal language large model, 2024
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53c66c15-5a4e-4dfa-a5b5-8d8d33eadb2e · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation Vggsound: A large-scale audio-visual dataset
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8835410d-eec2-4e66-ad2a-9bdad8343a2e · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation Beats: audio pre-training with acoustic tokeniz- ers
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2bbfb5e-8eed-417f-803c-5a30e08840e1 · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation Panda-70m: Captioning 70m videos with multiple cross-modality teachers
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa0388dc-d1d6-4625-bbb7-c2d912a3ecc3 · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation Unrav- eling instance associations: A closer look for audio-visual segmentation
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0bf5ea14-9876-48ba-a115-d84c7c909ef7 · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation Video-Guided Foley Sound Generation with Multimodal Controls
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38893e5b-74f9-4975-874b-f70ba02afcbf · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87acbeb6-f1d4-46a6-8775-7cafb0fc1144 · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation LoVA: Long-form Video-to-Audio Generation
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4da5ae48-9c18-40cb-82d2-614b52eedfef · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation Visionllama: A unified llama backbone for vision tasks
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c916368b-5b23-488c-a859-6da8e0700625 · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation FMA: A dataset for music analysis
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86aac1ef-94e9-4e15-85ee-20240d518095 · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation Conditional generation of audio from 9 video via foley analogies
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9f203f6-1502-4a60-bcf1-3cf0536f07b2 · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation Clap learning audio concepts from natural language supervision
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58dad096-f52e-43dd-b2a4-f5763cc6b66d · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation Scaling recti- fied flow transformers for high-resolution image synthesis
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27cada57-4874-4724-b88c-288452bc6502 · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation Hawley, and Jordi Pons
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3df8024-a0d7-430d-b42f-f96de7d39de6 · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation Gemmeke, Daniel P
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 345df0a7-7b26-43f5-9dd4-d4e26beb1806 · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation Imagebind: One embedding space to bind them all
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9f57ae0-859d-4f53-97c0-5ea1ff617699 · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation Gotta Hear Them All: Towards Sound Source Aware Audio Generation
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1077e4aa-094b-498e-829b-60aa55850084 · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation Taming Data and Transformers for Audio Generation
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d3c1d2f-5b50-40de-90c7-60d7aa395fea · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation MMDisCo: Multi-Modal Discriminator-Guided Cooperative Diffusion for Joint Audio and Video Generation
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ffdbda98-bd85-47f2-bea0-ddef94a55ac9 · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation Unsupervised semantic correspondence using stable diffusion, 2023
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be0f8032-d1fe-4f88-9550-82f8162ab192 · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation Gans trained by a two time-scale update rule converge to a local nash equi- librium
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a14cfa0-667a-48ea-b424-a126158d8588 · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation Classifier-free diffusion guidance
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4a0bd31-d441-421b-984a-14c116f6da46 · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation Make-An-Audio 2: Temporal-Enhanced Text-to-Audio Generation
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ccb6364-d46c-4181-a344-f7ef90054b3e · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation Rhythmic Foley: A Framework For Seamless Audio-Visual Alignment In Video-to-Audio Synthesis
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe63b53a-86be-461c-b420-be6531c6fe01 · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation Epic-sounds: A large- scale dataset of actions that sound
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c33fc9b-ae33-4f4f-ab48-15a4240349d6 · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation Taming visually guided sound generation
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f26d0e5-5a5e-488e-8b79-6cf965eb5261 · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation A Simple but Strong Baseline for Sounding Video Generation: Effective Adaptation of Audio and Video Diffusion Models for Joint Generation
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c023b92-9c29-49ac-8357-c674e1a97113 · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation The power of sound (tpos): Audio reactive video generation with stable diffusion
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fdb49b99-ba46-41c2-b18e-c2843704a41d · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation Pyramidal flow matching for efficient video generative modeling
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81065381-d8b1-453a-acf5-eb516161f96e · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation Re- purposing diffusion-based image generators for monocular depth estimation
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb732a7a-d8e7-4279-98ef-4b1b9af7299d · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation A Versatile Diffusion Transformer with Mixture of Noise Levels for Audiovisual Generation
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation aee831d3-0c03-4500-8a2b-4639b39f53d8 · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation Hifi-gan: Generative adversarial networks for efficient and high fi- delity speech synthesis
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 52d52cb8-48a5-4be9-989e-07211047f4bf · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1da30b39-53eb-4a0e-b4a0-803257faffb9 · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation Mandel, Mert Bay, and J
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a492f3c0-dc1a-4e96-b4ae-85b486274e10 · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation Video-foley: Two-stage video-to-sound generation via tem- poral event condition for foley sound
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08fb4f88-1400-45e6-9001-162a68cd3ee4 · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation AADiff: Audio-Aligned Video Synthesis with Text-to-Image Diffusion
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f2d74c1-4fc9-4855-a314-b5d48a2cad39 · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation Sound-guided semantic video generation
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation cc3d2874-6914-43b0-8606-63c48b05cbaa · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation T2v-turbo-v2: Enhancing video generation model post- training through data, reward, and conditional guidance de- sign
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8665cc07-3bd3-474d-a636-d1a4e135c064 · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation Ross, and Angjoo Kanazawa
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ccc2f29f-7925-4ba2-853c-5c31c45ebc5a · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation MuVi: Video-to-Music Generation with Semantic Alignment and Rhythmic Synchronization
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6f3e263-0aaf-4035-863e-03ade21e1c34 · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation Language-guided joint audio-visual editing via one-shot adaptation
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5508d397-fdaa-414c-baa7-b7cd7b5b4e4f · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation VMAS: Video-to-Music Generation via Semantic Alignment in Web Music Videos
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81ff63fc-d68d-44a5-83c0-2d634979c4c9 · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation Flow Matching for Generative Modeling
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db5a51b4-dbb8-45b3-8626-b30f21889909 · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation Plumbley
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 836db613-10af-4d07-8f33-c24fcf3f0009 · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation AudioLDM 2: Learning Holistic Audio Generation with Self-supervised Pretraining
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d26fc6e-5f18-41e0-a3c4-a39b3e6bfd48 · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91f568f4-b8df-452d-91e9-eccbf1446a63 · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation Diffusion hyperfeatures: Searching through time and space for semantic correspon- dence
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ce2c0459-1686-4549-b415-966e4ad8bdab · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation Diff-foley: Synchronized video-to-audio synthesis with la- tent diffusion models
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4b19e8e2-9ce0-423b-a4dd-e5f6488c3d19 · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation Videofusion: Decomposed diffusion models for high-quality video generation
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ff0f5ce3-fe6e-47e0-821d-f298dac1f104 · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation Tango 2: Aligning diffusion-based text-to-audio genera- tions through direct preference optimization
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation bd0a535c-367d-49a1-ac21-060fd96d3ee0 · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation The song describer dataset: a corpus of audio captions for music-and-language evaluation
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 7f424f3d-e7c0-4b7b-9820-3c008dab700f · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation Tavg- bench: Benchmarking text to audible-video generation
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 1a3e5ef6-de52-4a2d-8a91-fd88d807b80d · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation Foleygen: Visually-guided audio generation
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e5b67b7a-156f-43d6-80cc-b52c39f9095a · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation Snap video: Scaled spatiotemporal transformers for text-to-video synthesis
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ee4a89d5-0dee-40fc-9b10-165b1395ad71 · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation Adelson, and William T
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2628a2a9-614f-473b-81b7-8e7698f9d5db · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation Masked generative video-to-audio transform- ers with enhanced synchronicity
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d968493c-f847-4c0d-b704-3601b99c10e2 · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation Scalable diffusion mod- els with transformers
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation be76b20f-bdc0-4cde-a924-2e3f00be19e0 · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation Film: Visual reasoning with a general conditioning layer
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9398dfab-a169-448d-92ea-f0287750c2da · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation Sampson, Shikai Li, Si- mone Parmeggiani, Steve Fine, Tara Fowler, Vladan Petro- 11 vic, and Yuming Du
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 27931671-cd45-4082-9b04-9390c66f625d · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation Learn- ing transferable visual models from natural language super- vision
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b941f193-824a-429a-920d-a8ae3afb5702 · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation Unresolved cited work
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 33d81531-961a-4009-a078-e41a8cbac970 · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation EgoSonics: Generating Synchronized Audio for Silent Egocentric Videos
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6572317c-31b1-4bd6-847d-e69010d1a213 · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation STA-V2A: Video-to-Audio Generation with Semantic and Temporal Alignment
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3855bbe9-cde8-4ab7-8d19-ebf778934d1b · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation Mm-diffusion: Learning multi-modal diffusion mod- els for joint audio and video generation
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e7452dcc-fe0c-43b7-9871-f9dded336bbd · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation Improved techniques for training gans
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c31ed0eb-ce9a-4e14-91ab-119022731ec6 · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation I hear your true colors: Im- age guided audio generation
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e339fad5-3435-45f4-8b7a-bfba2c375d31 · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation Make-a-video: Text-to-video generation without text-video data
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5f42b313-b954-406b-a07b-baf800e90425 · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation Roformer: Enhanced transformer with rotary position embedding
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5b295371-2952-460f-a006-c9e040a762ad · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation From Vision to Audio and Beyond: A Unified Model for Audio-Visual Representation and Generation
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c42c4b43-a187-4170-9b72-4d57bcdfa1a8 · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation Mm-ldm: Multi-modal latent diffusion model for sounding video generation
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e6436e64-3cc4-4283-a367-9968b2297196 · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation Motion to dance music generation using latent dif- fusion model
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 17f509ef-8cf4-48bf-be47-94fe8379f1bd · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation Sequential Contrastive Audio-Visual Learning
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d39e8293-8224-4efd-ac56-17fe3c3e9265 · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation Towards Accurate Generative Models of Video: A New Metric & Challenges
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17c6a2ae-05ee-4646-b9bd-c3365752c2e5 · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation Temporally Aligned Audio for Video with Autoregression
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d616061-eb7f-4e5d-8b83-f30196b59afc · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation Phenaki: Variable length video generation from open do- main textual descriptions
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c5402d15-6c29-4779-b6de-572cb59c2fc0 · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8cd47d0e-1072-475e-8ec1-2822c3882e28 · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation V2a-mapper: A lightweight solution for vision-to-audio generation by connecting foun- dation models
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c64c4b3a-1444-4aef-ae97-ee09b416cf1c · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation AV-DiT: Efficient Audio-Visual Diffusion Transformer for Joint Audio and Video Generation
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f09b0b5-3138-49b2-98a6-1bd920b09879 · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation Tiva: Time-aligned video-to-audio generation
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 0fe7c216-2d3c-46c7-baea-8628c70b65f4 · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation Frieren: Efficient video-to-audio generation with rectified flow matching
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 40c0b88e-52a7-4f86-bc78-9d057a59e20b · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation Large-scale con- trastive language-audio pretraining with feature fusion and keyword-to-caption augmentation
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f8b5c5dd-6edd-467d-9a8f-f05b3403ba3d · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation Son- icvisionlm: Playing sound with vision language models
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2aeb31dd-8372-4365-9f8a-27fef25138a5 · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation Seeing and hearing: Open-domain visual- audio generation with diffusion latent aligners
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9ba67d64-8a39-4605-988a-adb3d59d14a2 · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation Demystifying CLIP Data
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 783d744a-e4aa-40e3-9e06-2d623bf2308c · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation Open-vocabulary panop- tic segmentation with text-to-image diffusion models
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation be5f2385-ce48-4bff-a520-1b547a610d1d · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation Auf- fusion: Leveraging the power of diffusion and large lan- guage models for text-to-audio generation
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e8c34cb8-e5e7-49d9-b3fb-412f0dd77e56 · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation Draw an Audio: Leveraging Multi-Instruction for Video-to-Audio Synthesis
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f0a9c24-7b86-49fb-b758-a0fa974a616a · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation CMMD: Contrastive Multi-Modal Diffusion for Video-Audio Conditional Modeling
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a9d01660-7b03-48bc-982f-d3b2aa92b682 · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation Diffusion model as rep- resentation learner
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 90778f86-5f84-465a-b49b-884719eee09f · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d3b3dbf-0b50-495d-8127-ac13507d7c48 · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation Diverse and aligned audio-to- video generation via text-to-video model adaptation
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 59f13910-172e-40cf-b89e-3d0a29552419 · outbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation Momu-diffusion: On learning long-term motion-music synchronization and cor- respondence
Reference 101
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 16755c82-6b7f-4480-9a83-682402e864d4 · inbound
JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.