Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-15T07:35:14.257562Z
Paper Citation Record · LEDGER
As of 6 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 1 inbound Pith citation observation for arXiv:2603.19857.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-15T07:35:14.257562Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-10T19:11:43.172296Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-05-10T23:20:54.250212Z
62 of 62 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation bf2fa41c-3316-497d-8346-57860a849652 · outbound
FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Black forest labs; frontier ai lab
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e2ef8972-038b-4a97-8f9a-1dd14857ca62 · outbound
FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Vggsound: A large-scale audio-visual dataset
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3f815c42-d86f-4098-93a9-117b6b171e0d · outbound
FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Video-guided foley sound generation with multimodal con- trols
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5ea8d461-33a6-4d6f-a914-25eee82e763d · outbound
FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts MMAu- dio: Taming multimodal joint training for high-quality video-to-audio synthesis
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1b7adcd7-7798-4692-b383-766a874321c2 · outbound
FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Simple and controllable music generation
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 68f29590-7766-484f-937f-9f915678d1d5 · outbound
FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a5284398-a38e-4d22-a5e1-9b7318301cc6 · outbound
FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1233b08b-7529-4716-ac5c-2f268e99b882 · outbound
FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation cc2f9143-e0da-45f8-90c0-5b8af801aee5 · outbound
FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2d43860a-3b14-43e1-b7dc-c0ae64728f83 · outbound
FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Imagebind: One embedding space to bind them all
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation bbccbe17-8fe8-477c-824a-ab8db7d2321d · outbound
FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Unresolved cited work
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation abb910f8-0020-47b9-8c46-da7aa1a50d51 · outbound
FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Video-to-Audio Generation with Fine-grained Temporal Semantics
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 457c038e-c392-422f-9f65-d8a465ba0383 · outbound
FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Spotlighting partially visible cinematic language for video-to-audio gen- eration via self-distillation
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f7217cf6-c0cf-4606-8373-7a3cb4d4e8ad · outbound
FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Make-An-Audio: Text-To-Audio Generation with Prompt-Enhanced Diffusion Models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 97d8d566-39f0-4746-9368-5bd523152868 · outbound
FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Tangoflux: Super fast and faithful text to audio generation with flow matching and clap-ranked preference optimization
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ffd7fa4d-89fe-492e-86fd-be23e579153e · outbound
FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Zisserman
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 35ca8a3f-bd26-44f0-bba6-ff82304f9904 · outbound
FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Audiocaps: Generating captions for audios in the wild
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b71e4b0b-1717-401d-8050-8c79d3717216 · outbound
FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Plumbley
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5bd501e2-6f78-458e-9dd2-8c802731a476 · outbound
FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts HunyuanVideo: A Systematic Framework For Large Video Generative Models
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation bfad09d1-ddc4-438d-8640-35e45deb9fba · outbound
FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Efficient training of audio transformers with patchout
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 061b5dad-dfc2-42d9-bc6e-291fe3375a92 · outbound
FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts AudioGen: Textually Guided Audio Generation
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a159c0f6-1e21-4cc5-acc2-85e9791d5f20 · outbound
FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Video-foley: Two-stage video-to-sound generation via tem- poral event condition for foley sound.IEEE Transactions on Audio, Speech and Language Processing
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6217bfbf-1d4d-428b-9467-5285616db589 · outbound
FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Dreamfoley: Scalable vlms for high-fidelity video-to- audio generation
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 52e3cd3d-fc9b-450e-9da5-55576307077c · outbound
FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7cd9ed13-ecd7-4dda-b22c-b884fb57acc6 · outbound
FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Imagine and seek: Improv- ing composed image retrieval with an imagined proxy
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f3669e5b-77a1-4e2d-8d16-247117093d96 · outbound
FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Audi- oLDM: Text-to-audio generation with latent diffusion mod- els.Proceedings of the International Conference on Machine Learning, pages 21450–21474
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1262e0b1-5b49-429a-ad0d-01d01dc80a3b · outbound
FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Plumbley
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation db8b65f2-ce98-4903-b1ca-a53b327de6e6 · outbound
FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Flashaudio: Rectified flows for fast and high-fidelity text-to-audio generation
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a2f6a240-2213-472f-b25b-71904619a732 · outbound
FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Thinksound: Chain-of- thought reasoning in multimodal large language models for audio generation and editing
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 186059dd-547d-44dd-934f-04cb85f13f1f · outbound
FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Diff-foley: Synchronized video-to-audio synthesis with la- tent diffusion models
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 789584c7-8d71-4aab-8744-acdce79db322 · outbound
FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Vista-llama: Reducing hallucination in video language models via equal distance to visual tokens
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 37dafcf1-55e2-40b3-9020-d4fb2539e913 · outbound
FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 962cd879-8b12-4bd7-ae9c-07a0e6358d00 · outbound
FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts High-resolution image syn- thesis with latent diffusion models
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6fd8684f-17c5-4d59-8a0b-26a26dd7b6be · outbound
FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Hunyuanvideo- foley: Multimodal diffusion with representation alignment for high-fidelity foley audio generation
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d602ab2c-a437-453f-87a2-85b28799d44d · outbound
FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Seedream 4.0: Toward Next-generation Multimodal Image Generation
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 25ec8201-7247-4d1c-8d7c-56fdeb8c9bc5 · outbound
FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Temporally aligned audio for video with autoregression
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d96234ef-1141-4efd-bc5d-20b7acf2d289 · outbound
FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Audiobox: Unified audio generation with natural language prompts
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation dca90a6c-031a-45e1-8c6e-e1d6efd4e026 · outbound
FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Wan: Open and Advanced Large-Scale Video Generative Models
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b5574116-6960-495e-b3d9-0056cea8d7d1 · outbound
FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Kling-foley: Multimodal diffusion trans- former for high-quality video-to-audio generation
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation cc4140ad-b044-46e7-96d7-46e44db93cec · outbound
FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Frieren: Efficient video-to-audio generation network with rectified flow matching.Advances in Neural Information Processing Systems, 37:128118–128138
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f2debca2-c864-42a6-b137-d91ef1b07d80 · outbound
FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2971fe07-3228-4cfe-b5bb-920697271c66 · outbound
FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Qwen2.5-Omni Technical Report
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c37a2910-650a-4ecd-a11c-4524546b64d1 · outbound
FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Video-to-audio generation with hidden alignment
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 774d5e36-1922-40ba-b902-8ebb08dd0ce5 · outbound
FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Con- textgen: Contextual layout anchoring for identity-consistent multi-instance generation.arXiv preprint arXiv:2510.11000
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d246dbf2-cd28-4dbc-91ce-caf561cb1524 · outbound
FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Towards Weakly Supervised Text-to-Audio Grounding
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation aef1d58f-46b2-44d7-8072-76d77c14f918 · outbound
FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Diffsound: Discrete diffusion model for text-to-sound generation.IEEE/ACM Transactions on Audio, Speech, and Language Processing, 31:1720–1733
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ae11fb55-b5b7-4b1c-a873-e175d3004ed4 · outbound
FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8e6a628d-a948-48e8-a2da-d39b984a7bb1 · outbound
FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Foley- crafter: Bring silent videos to life with lifelike and synchro- nized sounds
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 23e0861c-4e7f-4ffb-9b64-9158d0bfeac1 · outbound
FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Migc++: Advanced multi-instance generation controller for image synthesis.IEEE Transactions on Pattern Analysis and Machine Intelligence
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3bbd6672-5cd1-4839-a0e9-53194717df96 · outbound
FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Migc: Multi-instance generation controller for text-to-image synthesis
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation be354d4c-014f-4cd5-b12b-2fef546f42df · outbound
FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts 3dis: Depth-driven decoupled instance synthesis for text-to-image generation
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 751d62f7-6b40-4bb3-8e95-522f0b2a79f8 · outbound
FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Bidedpo: Conditional image generation with simultaneous text and condition alignment
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d4f23135-9a7b-44b6-ad3d-e167deff07f3 · outbound
FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Dreamrenderer: Taming multi-instance attribute control in large-scale text-to-image models
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1df1957c-1839-496f-a69a-5e429aa669f1 · outbound
FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts 3DIS-FLUX: simple and efficient multi-instance generation with DiT rendering
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f9549235-6119-4bb1-9eae-3a7b6d610caa · outbound
FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Masked audio generation using a single non- autoregressive transformer
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 27f04722-67b9-453f-9995-336e20b12d4b · outbound
FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Unresolved cited work
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3616e69c-1865-4430-ba7b-925378572b7d · outbound
FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Unresolved cited work
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d7660209-014b-458c-974e-3ee463c969c5 · outbound
FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Unresolved cited work
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 84a277d4-a083-4cd7-902c-1d46bbae8096 · outbound
FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts (b) Segment-level Classification You are an audio-visual analysis expert
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8c1aac7c-efc7-4fb6-8e05-a6a96844735b · outbound
FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Yes") or absent (
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 52bd301c-bd8f-4fe0-849c-ca34994288a4 · outbound
FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts soft", "medium
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3fd4c9ce-1fa9-443e-88a1-6d2b784a805c · outbound
FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts During training, we randomly drop TSR features with a probability of 0.1 We set an initial learning rate of2.0×10 −5
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 14285130-a947-437c-b8d6-d0e0c0629475 · inbound
RefineAnything: Multimodal Region-Specific Refinement for Perfect Local Details FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.