Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T11:47:56.159595Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 75 of 75 outbound references and 4 inbound Pith citation observations for arXiv:2412.15023.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T11:47:56.159595Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T20:25:31.026317Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-21T22:20:42.066924Z
75 of 75 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 8aad8b79-bd5e-4838-8165-ade016158b30 · outbound
FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment A survey of multimodal deep generative models,
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f2f68887-ca34-44cf-82e0-29ea684fcc98 · outbound
FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment L3DAS21 challenge: Machine learning for 3D audio signal processing,
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation eb4ad66b-472c-4161-a37c-f26fc118b4de · outbound
FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment L3DAS22 challenge: Learning 3D audio sources in a real office environment,
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 4b36f5b0-9853-4bfa-ac12-5a12fc4821e3 · outbound
FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Video-LLaMA: An instruction- tuned audio-visual language model for video understanding,
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 13f67f3c-5a38-4ff2-b1b0-d9a6445560b6 · outbound
FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Empowering llms with pseudo-untrimmed videos for audio-visual temporal understanding,
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f22601e7-98ee-474a-955c-def900195879 · outbound
FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Soundnet: Learning sound representations from unlabeled video,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b94f8ae0-ac10-4a30-895c-b7b141128b00 · outbound
FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Visual to sound: Generating natural sound for videos in the wild,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 2bb8e3eb-2c7b-4688-bef9-d361c898fd18 · outbound
FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment I hear your true colors: Image guided audio generation,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 9a2b03e8-20c9-4dd6-b0f4-124607ae6942 · outbound
FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment AutoFoley: Artificial synthesis of synchronized sound tracks for silent videos with deep learning,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 7b9d8e22-02b2-446a-a372-a212a896b2f5 · outbound
FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment FoleyGAN: Visually guided generative adversarial network-based synchronous sound generation in silent videos,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 656f0442-9bd2-4657-8080-0f8687900829 · outbound
FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Video background music generation with controllable music transformer,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation fc4313d9-a5c0-4bbc-9464-37e2c0ea42c7 · outbound
FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Video background music generation: Dataset, method and evaluation,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e21430f2-83eb-4bc1-9cbf-c25bf59de308 · outbound
FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Semantically consistent Video-to-Audio Generation using Multimodal Language Large Model
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38a6abbb-8b86-4249-9a31-17ecbeb5aa08 · outbound
FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment An overview of visual sound synthesis generation tasks based on deep learning networks,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 9fe156ec-e9c5-4e25-b5bb-d4814317721f · outbound
FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Perception of synchrony between the senses,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 93a8e422-c641-4d41-8ce5-bb7ccdaf3614 · outbound
FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Visually indicated sound generation by perceptually optimized classification,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 15774c97-d334-4f8b-be49-ba5dcc2ae7d3 · outbound
FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Va- rietysound: Timbre-controllable video to sound generation via unsuper- vised information disentanglement,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 3c706d48-760b-4ec9-a4d5-9226bf83234e · outbound
FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Efficient Video to Audio Mapper with Visual Scene Detection
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7325ee3-c854-42b1-b248-3037bbaf9549 · outbound
FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment A Simple but Strong Baseline for Sounding Video Generation: Effective Adaptation of Audio and Video Diffusion Models for Joint Generation
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2fe971a-2f5e-44cd-9d87-9714613aab5a · outbound
FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment T-foley: A control- lable waveform-domain diffusion model for temporal-event-guided foley sound synthesis,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ea2415b0-28da-4c66-aea0-9c440c8b8ce5 · outbound
FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Visually indicated sounds,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 71e82d0f-afa2-4278-94e4-8c214cc69d1b · outbound
FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Neural synthesis of footsteps sound effects with generative adversarial networks,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5b8a4293-26cf-4bcf-b6d0-82be920927fb · outbound
FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Pix2Video: Video editing using image diffusion,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 91f3d3db-9785-4b52-ba3b-05d6f9a479f4 · outbound
FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Audio-visual contrastive learning with temporal self-supervision,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d680e219-0ddc-476b-9c06-47f9f1b0cb83 · outbound
FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Audio match cutting: Finding and creating matching audio transitions in movies and videos,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 357b3b92-4cd4-49ea-a147-2561e8383cc3 · outbound
FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment DubWise: Video-guided speech duration control in multimodal llm-based text-to-speech for dubbing,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d659b5d9-2fb6-491b-8c0a-077c63169a31 · outbound
FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment MambaFoley: Foley Sound Generation using Selective State-Space Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c06162d-e275-42be-8c2b-ec3183bc028b · outbound
FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Conditional sound generation using neural discrete time-frequency representation learning,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c0a833cc-20a8-4688-8e1d-a8b7c54e3fee · outbound
FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Latent Diffusion Model Based Foley Sound Generation System For DCASE Challenge 2023 Task 7
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 750e44d8-e380-4b89-bf27-3b0e290deeb5 · outbound
FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Real-time sound synthesis of audience applause,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e7cc9bb6-f25f-401c-a622-85dd07de54b9 · outbound
FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Learning transferable visual models from natural language supervision,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c6c6e856-8337-48b3-a496-5be54f44477b · outbound
FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Seeing and Hearing: Open-domain visual-audio generation with diffusion latent aligners,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5160599f-4700-4f92-b17a-6260e8e6f280 · outbound
FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment ImageBind one embedding space to bind them all,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d181e934-d13f-4f7f-b414-3d916a47ec03 · outbound
FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Generating visually aligned sound from videos,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 4b2b1a72-181d-4e10-aca4-0d3c74cc5200 · outbound
FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Taming visually guided sound generation,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation efda7b9a-a325-4cf1-9072-be216eb670d6 · outbound
FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Conditional generation of audio from video via foley analo- gies,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation fc329c8b-0373-4f4f-8616-123481671d3b · outbound
FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Diff-Foley: Synchronized video-to-audio synthesis with latent diffusion models,
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e8ad2b03-5882-428e-8437-b6c8e9dc4f35 · outbound
FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Sync- fusion: Multimodal onset-synchronized video-to-audio foley synthesis,
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b67f90d7-9eb0-46ff-a0a9-23ff0d421503 · outbound
FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment A closer look at spatiotemporal convolutions for action recognition,
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f9a6cccd-9f20-4693-a08c-d8fbe78660a6 · outbound
FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Mo\^usai: Text-to-Music Generation with Long-Context Latent Diffusion
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 736cdc61-d45e-4b48-b65d-478768145afe · outbound
FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Sta-v2a: Video-to-audio generation with semantic and temporal alignment,
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 436166ea-fea2-430f-b99b-f5b690ed6cac · outbound
FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Video-Foley: Two-stage video-to-sound generation via temporal event condition for foley sound,
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5dae2f3-25ee-4078-84f2-4e30ed5d97c9 · outbound
FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment AudioLDM: Text- to-audio generation with latent diffusion models,
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 0e2dc920-acfa-43a5-9bdb-6f823bd9b86c · outbound
FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Denoising diffusion probabilistic models,
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9bb2bd5e-8922-456c-8f5c-948ef63cb60f · outbound
FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Multi-source diffusion models for simultaneous music generation and separation,
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 396522d5-738e-4895-a2ec-e3d8fadadbc2 · outbound
FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment DiffWave: A versatile diffusion model for audio synthesis,
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 3540dd17-00c1-4ff2-b17a-007de300d6a7 · outbound
FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment High-resolution image synthesis with latent diffusion models,
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b9accb8f-a7db-4cbf-a07b-63e0aab9d3df · outbound
FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment AudioLDM 2: Learning holistic audio generation with self- supervised pretraining,
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 274ef97e-7239-4ecd-a233-54fe3799614d · outbound
FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Fast Timing-Conditioned Latent Audio Diffusion
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ff3bd98-15fe-40b3-aa88-26a6ca38b496 · outbound
FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Long-form music generation with latent diffusion
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b875930e-4111-40a9-a651-a7de0ea6bb63 · outbound
FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Adding conditional control to text-to-image diffusion models,
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c8383a2a-85e8-4600-acdb-38b9b05efd1a · outbound
FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment PixArt- α: Fast training of diffusion transformer for photorealistic 12 text-to-image synthesis,
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 300d38cb-aa3b-4750-bd97-3a006c3c5590 · outbound
FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Music ControlNet: Multiple time-varying controls for music gener- ation,
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation dde66191-fdc4-4580-9641-359ab999e317 · outbound
FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Discerning real from synthetic: analysis and perceptual evaluation of sound effects,
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 2257d4e4-38a2-4f3f-bf11-9e9c9baa1f46 · outbound
FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment a machine learning method to evaluate and improve sound effects synthesis model design,
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 0ca5f7fc-d1e3-4d10-bed2-67f6b19d3d6c · outbound
FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment WaveNet: A generative model for raw audio,
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation db6e4e98-995a-491e-a70a-4cb057015316 · outbound
FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Joint detection and classification of singing voice melody using convolutional recurrent neural networks,
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 068cd80b-60b6-4734-bff1-13b49e7bd7bd · outbound
FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment RAFT: Recurrent all-pairs field transforms for optical flow,
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 39a82b1b-fcc3-4930-954c-8775ae79cebf · outbound
FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Lever- aging temporal contextualization for video action recognition,
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation eb9b6a76-7a31-4880-9ef1-c568b88c1927 · outbound
FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Framewise phoneme classifica- tion with bidirectional LSTM networks,
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 209affa8-5626-4d4a-ba9e-655cbe034fc5 · outbound
FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Stable Audio Open
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation efe75b57-ddd6-4b43-843c-54f8a35a9524 · outbound
FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment BYOL for Audio: Self-supervised learning for general- purpose audio representation,
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a1cc6ba5-e531-4925-87f5-08e6dc816774 · outbound
FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Large-scale contrastive language- audio pretraining with feature fusion and keyword-to-caption augmen- tation,
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation dd118ad5-7255-409a-97c8-1e4b645258e1 · outbound
FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Separate What You Describe: Language-queried audio source separation,
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 1d0d1907-3ab2-4e8c-938b-44106dc9d413 · outbound
FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Classifier-free diffusion guidance,
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f1f65320-bac9-4e7a-8ab3-fbbaa8c7b2b2 · outbound
FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Fr´echet audio distance: A reference-free metric for evaluating music enhancement algorithms,
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation bd98a2bc-f8d9-402d-a874-9f8799e8ab33 · outbound
FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Correlation of fr´echet audio distance with human perception of environmental audio is embedding dependent,
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 0f499230-d784-4568-8271-710b404a0fe4 · outbound
FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment PANNs: Large-scale pretrained audio neural networks for audio pattern recognition,
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 59f414aa-5e20-4b73-bdb1-cb8e633c1d54 · outbound
FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment CLAP learning audio concepts from natural language supervision,
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c485989e-4a57-4ad6-9791-c92f919420e0 · outbound
FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Perceptual evaluation of audio-visual synchrony grounded in viewers’ opinion scores,
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c4159048-d2ca-4cce-ac3f-ca0e1730e2f1 · outbound
FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Quo vadis, action recognition? a new model and the kinetics dataset,
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 90645924-ade3-418d-872a-b5289a49db5f · outbound
FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment CNN architectures for large-scale audio classification,
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 9bf5f365-95b2-4bf0-a514-0232bbcbbc76 · outbound
FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Batch normalization: accelerating deep network training by reducing internal covariate shift,
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 362ddce1-cf8f-492c-83aa-549100912ae0 · outbound
FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment ImageNet: A large-scale hierarchical image database,
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ce3ccac1-d5a8-4c92-98e9-6380c20126b4 · outbound
FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment The Kinetics Human Action Video Dataset
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f584c6a1-2a11-49c5-8f75-1f9adfd55165 · inbound
Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6cb4d445-1e69-4e49-ba26-f9ebff2361a3 · inbound
SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72e809fd-2512-4928-8d4d-1335665e735b · inbound
SonicGauss: Position-Aware Physical Sound Synthesis for 3D Gaussian Representations FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 348a9a03-cc71-4736-bbfd-c8bbbd289a30 · inbound
Training-Free Multimodal Guidance for Video to Audio Generation FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.