Pith. sign in

Paper Citation Record · LEDGER

FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment

As of 18 August 2026, this Paper Citation Record lists 75 of 75 outbound references and 4 inbound Pith citation observations for arXiv:2412.15023.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.15023 v3

Coverage vector

measured 75 of 75 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T11:47:56.159595Z

measured 79 of 79 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:25:31.026317Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T22:20:42.066924Z

Reference resolution

75 of 75 outbound references displayed

  • verified exact1
  • verified fuzzy62
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8aad8b79-bd5e-4838-8165-ade016158b30 · outbound

This paper cites A survey of multimodal deep generative models,.

FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment A survey of multimodal deep generative models,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:47:57.305249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T11:47:55.885512Z digest=sha256:e9f0d8b8cf4586958a3ab76d6709e7a7a710212e4134d9cb327a250a4999554d

Observation f2f68887-ca34-44cf-82e0-29ea684fcc98 · outbound

This paper cites L3DAS21 challenge: Machine learning for 3D audio signal processing,.

FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment L3DAS21 challenge: Machine learning for 3D audio signal processing,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:47:57.295209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T11:47:55.890069Z digest=sha256:f6e238e6ddc0705523030b8e2e641c035110896c93d8ea6d28cbff49600c8407

Observation eb4ad66b-472c-4161-a37c-f26fc118b4de · outbound

This paper cites L3DAS22 challenge: Learning 3D audio sources in a real office environment,.

FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment L3DAS22 challenge: Learning 3D audio sources in a real office environment,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:47:57.284623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T11:47:55.893966Z digest=sha256:569ec26a78269fe5f2ce90c3e800e09d657d6bc73d61333e6f2bb05a27c5e1d6

Observation 4b36f5b0-9853-4bfa-ac12-5a12fc4821e3 · outbound

This paper cites Video-LLaMA: An instruction- tuned audio-visual language model for video understanding,.

FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Video-LLaMA: An instruction- tuned audio-visual language model for video understanding,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:47:57.269022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T11:47:55.898289Z digest=sha256:ea6f07ca9cd4faff0003e6cecf1f3be3b73d9ed10c52cd9581e9314b51e2296e

Observation 13f67f3c-5a38-4ff2-b1b0-d9a6445560b6 · outbound

This paper cites Empowering llms with pseudo-untrimmed videos for audio-visual temporal understanding,.

FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Empowering llms with pseudo-untrimmed videos for audio-visual temporal understanding,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T11:47:55.902001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:47:55.902001Z digest=sha256:7ba6e850a9fcc5daf7334cc03a5ec390177d2100d91c13eb2be5269cc7e7e6e6

Observation f22601e7-98ee-474a-955c-def900195879 · outbound

This paper cites Soundnet: Learning sound representations from unlabeled video,.

FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Soundnet: Learning sound representations from unlabeled video,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:47:57.256710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T11:47:55.905741Z digest=sha256:cb64af4d2459edcaccec58057f73bee32ccc86b2883d9a39d8d74338a4ee533a

Observation b94f8ae0-ac10-4a30-895c-b7b141128b00 · outbound

This paper cites Visual to sound: Generating natural sound for videos in the wild,.

FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Visual to sound: Generating natural sound for videos in the wild,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:47:57.243730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T11:47:55.910021Z digest=sha256:194aacb3a78ac11b7c7fe5d745bc7ef83f0e3bfd00af89e46153f153860a9401

Observation 2bb8e3eb-2c7b-4688-bef9-d361c898fd18 · outbound

This paper cites I hear your true colors: Image guided audio generation,.

FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment I hear your true colors: Image guided audio generation,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:47:57.233094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T11:47:55.913879Z digest=sha256:dfd45dbd92146c22d81037a46900ec536eea6092df01e3cf382d8e52179f4d87

Observation 9a2b03e8-20c9-4dd6-b0f4-124607ae6942 · outbound

This paper cites AutoFoley: Artificial synthesis of synchronized sound tracks for silent videos with deep learning,.

FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment AutoFoley: Artificial synthesis of synchronized sound tracks for silent videos with deep learning,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:47:57.222987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T11:47:55.917525Z digest=sha256:7a617a02371ac420d07747f88517314bfd7708c575acf5cba3dd75d5c941012f

Observation 7b9d8e22-02b2-446a-a372-a212a896b2f5 · outbound

This paper cites FoleyGAN: Visually guided generative adversarial network-based synchronous sound generation in silent videos,.

FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment FoleyGAN: Visually guided generative adversarial network-based synchronous sound generation in silent videos,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:47:57.212359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T11:47:55.921543Z digest=sha256:2a5de3709dc8931f4fedefa5f1f5aa2ea12cdeaed8b2b4de2e23d1a5b7c3056b

Observation 656f0442-9bd2-4657-8080-0f8687900829 · outbound

This paper cites Video background music generation with controllable music transformer,.

FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Video background music generation with controllable music transformer,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:47:57.200106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T11:47:55.925494Z digest=sha256:b73d43a27f44592caa5ec7954a95b933d98d401b1233624be5921ccedfc1ec1c

Observation fc4313d9-a5c0-4bbc-9464-37e2c0ea42c7 · outbound

This paper cites Video background music generation: Dataset, method and evaluation,.

FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Video background music generation: Dataset, method and evaluation,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:47:57.187387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T11:47:55.930222Z digest=sha256:0464586fc882a5f855bd571177df11f7122a2df880c037b85c3723256e6b85c0

Observation e21430f2-83eb-4bc1-9cbf-c25bf59de308 · outbound

This paper cites Semantically consistent Video-to-Audio Generation using Multimodal Language Large Model.

FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Semantically consistent Video-to-Audio Generation using Multimodal Language Large Model

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T11:47:55.934025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:47:55.934025Z digest=sha256:f6f207b9e83199fd46af697c10287997370d20ac1f2d03fe6dc319262d09806c

Observation 38a6abbb-8b86-4249-9a31-17ecbeb5aa08 · outbound

This paper cites An overview of visual sound synthesis generation tasks based on deep learning networks,.

FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment An overview of visual sound synthesis generation tasks based on deep learning networks,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:47:57.175526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T11:47:55.938075Z digest=sha256:79bc68ea1da1daa12f3fa56efcd047d09da89bfde00473c9bd3a31e08a2200d0

Observation 9fe156ec-e9c5-4e25-b5bb-d4814317721f · outbound

This paper cites Perception of synchrony between the senses,.

FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Perception of synchrony between the senses,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:47:57.162652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T11:47:55.941747Z digest=sha256:a56a69eec4c29abb1c1fa804190ca5054ddd6ee8bb32f0d676c0713050f9bbc2

Observation 93a8e422-c641-4d41-8ce5-bb7ccdaf3614 · outbound

This paper cites Visually indicated sound generation by perceptually optimized classification,.

FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Visually indicated sound generation by perceptually optimized classification,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:47:57.150878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T11:47:55.945333Z digest=sha256:68745ebb54950c10b2c748b621078e400731447c63874616dcd29c290b14fa32

Observation 15774c97-d334-4f8b-be49-ba5dcc2ae7d3 · outbound

This paper cites Va- rietysound: Timbre-controllable video to sound generation via unsuper- vised information disentanglement,.

FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Va- rietysound: Timbre-controllable video to sound generation via unsuper- vised information disentanglement,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:47:57.138199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T11:47:55.948357Z digest=sha256:7c3ddd1238ca490c50786f6c001e115b0957cd314e72b79c4f68b835b389f6bd

Observation 3c706d48-760b-4ec9-a4d5-9226bf83234e · outbound

This paper cites Efficient Video to Audio Mapper with Visual Scene Detection.

FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Efficient Video to Audio Mapper with Visual Scene Detection

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T11:47:55.951519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:47:55.951519Z digest=sha256:907c6b72c9f4fe59e3b730958489034d7b98726848024b7fe43cb477a4db9b36

Observation c7325ee3-c854-42b1-b248-3037bbaf9549 · outbound

This paper cites A Simple but Strong Baseline for Sounding Video Generation: Effective Adaptation of Audio and Video Diffusion Models for Joint Generation.

FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment A Simple but Strong Baseline for Sounding Video Generation: Effective Adaptation of Audio and Video Diffusion Models for Joint Generation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T11:47:55.955024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:47:55.955024Z digest=sha256:836ad16af5bc8112dae333f9fbb64423b39c684c348c53c96ef167ba71777b1c

Observation c2fe971a-2f5e-44cd-9d87-9714613aab5a · outbound

This paper cites T-foley: A control- lable waveform-domain diffusion model for temporal-event-guided foley sound synthesis,.

FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment T-foley: A control- lable waveform-domain diffusion model for temporal-event-guided foley sound synthesis,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:47:57.116672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T11:47:55.959296Z digest=sha256:46e49220e0aea36d5982bad5b8c0afb6a8ab23a71583b00dd7b7eb9c2d95bb6a

Observation ea2415b0-28da-4c66-aea0-9c440c8b8ce5 · outbound

This paper cites Visually indicated sounds,.

FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Visually indicated sounds,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:47:57.103728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T11:47:55.962392Z digest=sha256:d5e80e0903cd6297e8f033a310666db38c5cc23d4e8399399c471af85c243ef4

Observation 71e82d0f-afa2-4278-94e4-8c214cc69d1b · outbound

This paper cites Neural synthesis of footsteps sound effects with generative adversarial networks,.

FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Neural synthesis of footsteps sound effects with generative adversarial networks,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:47:57.084959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T11:47:55.965544Z digest=sha256:defcb70ae16e5f17753ca45aaa77bd956b7bef6b84d9a036299cbf779c939a5b

Observation 5b8a4293-26cf-4bcf-b6d0-82be920927fb · outbound

This paper cites Pix2Video: Video editing using image diffusion,.

FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Pix2Video: Video editing using image diffusion,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:47:57.074663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T11:47:55.968362Z digest=sha256:0a20bb8b13d301bc000296d5f1a77a89b94cc76c34422d7ad4fc1c7531b62499

Observation 91f3d3db-9785-4b52-ba3b-05d6f9a479f4 · outbound

This paper cites Audio-visual contrastive learning with temporal self-supervision,.

FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Audio-visual contrastive learning with temporal self-supervision,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:47:57.059134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T11:47:55.971911Z digest=sha256:8aa24ca69252c82b3f4d09983ff3daf30c68edc9e6868f6dfd8155f20c942f3e

Observation d680e219-0ddc-476b-9c06-47f9f1b0cb83 · outbound

This paper cites Audio match cutting: Finding and creating matching audio transitions in movies and videos,.

FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Audio match cutting: Finding and creating matching audio transitions in movies and videos,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:47:57.042836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T11:47:55.975489Z digest=sha256:94c665c85adb83761de2ced24ddade5fcdce920a639226a206a11626cf567f21

Observation 357b3b92-4cd4-49ea-a147-2561e8383cc3 · outbound

This paper cites DubWise: Video-guided speech duration control in multimodal llm-based text-to-speech for dubbing,.

FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment DubWise: Video-guided speech duration control in multimodal llm-based text-to-speech for dubbing,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:47:57.023981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T11:47:55.979056Z digest=sha256:d6a96a88751ef51dfac58069ee0c5a076d36439add2d491d3d3bed868ddb16be

Observation d659b5d9-2fb6-491b-8c0a-077c63169a31 · outbound

This paper cites MambaFoley: Foley Sound Generation using Selective State-Space Models.

FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment MambaFoley: Foley Sound Generation using Selective State-Space Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T11:47:55.982423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:47:55.982423Z digest=sha256:94f2a4ab3b43d213eb503e74ca2df5073e4b222950c4738eeb484b4daae37087

Observation 4c06162d-e275-42be-8c2b-ec3183bc028b · outbound

This paper cites Conditional sound generation using neural discrete time-frequency representation learning,.

FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Conditional sound generation using neural discrete time-frequency representation learning,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:47:57.008408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T11:47:55.986462Z digest=sha256:33b98a87363a3f9b9ea7c616be90ef163b61d8bd082258fb6faef605ab89ea0f

Observation c0a833cc-20a8-4688-8e1d-a8b7c54e3fee · outbound

This paper cites Latent Diffusion Model Based Foley Sound Generation System For DCASE Challenge 2023 Task 7.

FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Latent Diffusion Model Based Foley Sound Generation System For DCASE Challenge 2023 Task 7

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-08-11T11:47:56.402228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T11:47:55.990167Z digest=sha256:cc64407affbc2e41807c0d19ef74558713c34fd4e0cd4d69721998cba7888db0

Observation 750e44d8-e380-4b89-bf27-3b0e290deeb5 · outbound

This paper cites Real-time sound synthesis of audience applause,.

FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Real-time sound synthesis of audience applause,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:47:56.994821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T11:47:55.993954Z digest=sha256:b2f22a3d43c5af2f127fae509b069cc9d24461660b465e0e316ea41cd7d31900

Observation e7cc9bb6-f25f-401c-a622-85dd07de54b9 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Learning transferable visual models from natural language supervision,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:47:56.979429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T11:47:55.997427Z digest=sha256:9d58c96bae3ebfe7443ade86eff59ff67554e4354f6f7700570c0dd32cac9fec

Observation c6c6e856-8337-48b3-a496-5be54f44477b · outbound

This paper cites Seeing and Hearing: Open-domain visual-audio generation with diffusion latent aligners,.

FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Seeing and Hearing: Open-domain visual-audio generation with diffusion latent aligners,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:47:56.966953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T11:47:56.000868Z digest=sha256:075221e1ffafffa4ea9eb3747c224cb4179d217075e2cba5f631de2810a8f766

Observation 5160599f-4700-4f92-b17a-6260e8e6f280 · outbound

This paper cites ImageBind one embedding space to bind them all,.

FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment ImageBind one embedding space to bind them all,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:47:56.955015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T11:47:56.004277Z digest=sha256:f66adb1064226fc57d6b33deef42f0a267766f7df95ff7a25fee78795eac2d61

Observation d181e934-d13f-4f7f-b414-3d916a47ec03 · outbound

This paper cites Generating visually aligned sound from videos,.

FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Generating visually aligned sound from videos,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:47:56.941646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T11:47:56.007815Z digest=sha256:c662ceae7f9349f9281fd53fc237b9565e265287e9d1b5e43e1a80e0d97d297d

Observation 4b2b1a72-181d-4e10-aca4-0d3c74cc5200 · outbound

This paper cites Taming visually guided sound generation,.

FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Taming visually guided sound generation,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:47:56.928009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T11:47:56.011436Z digest=sha256:bebfc030dce7ee239fa13b75da448b1136a52b93dc9cfbf3db49dc74a9a07029

Observation efda7b9a-a325-4cf1-9072-be216eb670d6 · outbound

This paper cites Conditional generation of audio from video via foley analo- gies,.

FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Conditional generation of audio from video via foley analo- gies,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:47:56.915728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T11:47:56.015217Z digest=sha256:bdb9b5ba719c32f0f5a57c2061ba51d12cc30689076a7303dbd1ac07f5b2fbc6

Observation fc329c8b-0373-4f4f-8616-123481671d3b · outbound

This paper cites Diff-Foley: Synchronized video-to-audio synthesis with latent diffusion models,.

FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Diff-Foley: Synchronized video-to-audio synthesis with latent diffusion models,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:47:56.904586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T11:47:56.018350Z digest=sha256:e42a15da0e0f4c13a1080b847b45184fdeedb8aaa96dc8da517f4f71bf141c35

Observation e8ad2b03-5882-428e-8437-b6c8e9dc4f35 · outbound

This paper cites Sync- fusion: Multimodal onset-synchronized video-to-audio foley synthesis,.

FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Sync- fusion: Multimodal onset-synchronized video-to-audio foley synthesis,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:47:56.892151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T11:47:56.022247Z digest=sha256:5f5463f1d106d2cbc6da589272f699c7467a9d55e3c9f3e7ccbb634c818112f5

Observation b67f90d7-9eb0-46ff-a0a9-23ff0d421503 · outbound

This paper cites A closer look at spatiotemporal convolutions for action recognition,.

FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment A closer look at spatiotemporal convolutions for action recognition,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:47:56.880969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T11:47:56.025676Z digest=sha256:ed3b24aa1a4130fe1e818cacd77b5e7726c0850471498d1856ac7cd1fb8bb435

Observation f9a6cccd-9f20-4693-a08c-d8fbe78660a6 · outbound

This paper cites Mo\^usai: Text-to-Music Generation with Long-Context Latent Diffusion.

FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Mo\^usai: Text-to-Music Generation with Long-Context Latent Diffusion

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T11:47:56.029455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:47:56.029455Z digest=sha256:1d9239da2874c67288d189d4dc27bed6cb939ff66565e8a7608419f22e977171

Observation 736cdc61-d45e-4b48-b65d-478768145afe · outbound

This paper cites Sta-v2a: Video-to-audio generation with semantic and temporal alignment,.

FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Sta-v2a: Video-to-audio generation with semantic and temporal alignment,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:47:56.868826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T11:47:56.033540Z digest=sha256:fe727a5ca2bde54a14be230dae31217aabc0f556efcc8f562ed329f0ea43f8fa

Observation 436166ea-fea2-430f-b99b-f5b690ed6cac · outbound

This paper cites Video-Foley: Two-stage video-to-sound generation via temporal event condition for foley sound,.

FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Video-Foley: Two-stage video-to-sound generation via temporal event condition for foley sound,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T11:47:56.037163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:47:56.037163Z digest=sha256:eb8f7ef88855fcb2744fdf1a35559cbb023b5a1da9a5759280ff6364b7bb4f3d

Observation a5dae2f3-25ee-4078-84f2-4e30ed5d97c9 · outbound

This paper cites AudioLDM: Text- to-audio generation with latent diffusion models,.

FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment AudioLDM: Text- to-audio generation with latent diffusion models,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:47:56.857547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T11:47:56.040763Z digest=sha256:b904eb8a54e645c9ce3e07467fdfd148b3a93b4868230d4734c067c1e2fcb1ca

Observation 0e2dc920-acfa-43a5-9bdb-6f823bd9b86c · outbound

This paper cites Denoising diffusion probabilistic models,.

FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Denoising diffusion probabilistic models,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T11:47:56.044771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:47:56.044771Z digest=sha256:9762f3f086cce726fdfd96d3422584f1e0ba81c6dc7a9b4bb0118cc9be88ed2e

Observation 9bb2bd5e-8922-456c-8f5c-948ef63cb60f · outbound

This paper cites Multi-source diffusion models for simultaneous music generation and separation,.

FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Multi-source diffusion models for simultaneous music generation and separation,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:47:56.838724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T11:47:56.050119Z digest=sha256:8b3c6aa06893afa75b84224eea13784c615699710dd1f6a53f2b44668eab5be1

Observation 396522d5-738e-4895-a2ec-e3d8fadadbc2 · outbound

This paper cites DiffWave: A versatile diffusion model for audio synthesis,.

FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment DiffWave: A versatile diffusion model for audio synthesis,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:47:56.826047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T11:47:56.053042Z digest=sha256:3046e690987ba56f20249bcee7ed7c8b06851955de94001e89a7bcef8d85f227

Observation 3540dd17-00c1-4ff2-b17a-007de300d6a7 · outbound

This paper cites High-resolution image synthesis with latent diffusion models,.

FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment High-resolution image synthesis with latent diffusion models,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:47:56.814888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T11:47:56.056137Z digest=sha256:b810dcee9a2d75563f0a05d81372a0c206dec059d19dcddaec16ccfda314d446

Observation b9accb8f-a7db-4cbf-a07b-63e0aab9d3df · outbound

This paper cites AudioLDM 2: Learning holistic audio generation with self- supervised pretraining,.

FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment AudioLDM 2: Learning holistic audio generation with self- supervised pretraining,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:47:56.804736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T11:47:56.058919Z digest=sha256:c3ecdd82be5634d6ddfaedb3e55855a8a8f186f6c4b49ce167c783636e15260e

Observation 274ef97e-7239-4ecd-a233-54fe3799614d · outbound

This paper cites Fast Timing-Conditioned Latent Audio Diffusion.

FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Fast Timing-Conditioned Latent Audio Diffusion

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T11:47:56.061472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:47:56.061472Z digest=sha256:6042d1ecb7032c84b94f6562c908efbaf0701947df82c514e20fc758619de05a

Observation 9ff3bd98-15fe-40b3-aa88-26a6ca38b496 · outbound

This paper cites Long-form music generation with latent diffusion.

FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Long-form music generation with latent diffusion

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T11:47:56.064401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:47:56.064401Z digest=sha256:92ed2e573ed823f3fd6b0c70b6650f421fd45817aa230733d6e7b32f9faf7714

Observation b875930e-4111-40a9-a651-a7de0ea6bb63 · outbound

This paper cites Adding conditional control to text-to-image diffusion models,.

FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Adding conditional control to text-to-image diffusion models,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:47:56.795503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T11:47:56.068602Z digest=sha256:84b805d317608c119d6bf1970d77a85b20d2d78b8893a52322b094662ba9c04a

Observation c8383a2a-85e8-4600-acdb-38b9b05efd1a · outbound

This paper cites PixArt- α: Fast training of diffusion transformer for photorealistic 12 text-to-image synthesis,.

FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment PixArt- α: Fast training of diffusion transformer for photorealistic 12 text-to-image synthesis,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:47:56.784994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T11:47:56.072532Z digest=sha256:78dc83377967e40139f453da4b0e34c4869a58d27f1868a89510bcdb840c634a

Observation 300d38cb-aa3b-4750-bd97-3a006c3c5590 · outbound

This paper cites Music ControlNet: Multiple time-varying controls for music gener- ation,.

FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Music ControlNet: Multiple time-varying controls for music gener- ation,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:47:56.773716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T11:47:56.076127Z digest=sha256:67a424ec8d06c09a51ef021d950fcfd8edccbafe2d50e82bfa91f48d8f62d7a7

Observation dde66191-fdc4-4580-9641-359ab999e317 · outbound

This paper cites Discerning real from synthetic: analysis and perceptual evaluation of sound effects,.

FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Discerning real from synthetic: analysis and perceptual evaluation of sound effects,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:47:56.762864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T11:47:56.080021Z digest=sha256:4a5f5d71bad91ade28d31e963f7eeee4c6156075c35c85d598b720839783b156

Observation 2257d4e4-38a2-4f3f-bf11-9e9c9baa1f46 · outbound

This paper cites a machine learning method to evaluate and improve sound effects synthesis model design,.

FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment a machine learning method to evaluate and improve sound effects synthesis model design,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:47:56.751860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T11:47:56.084360Z digest=sha256:cd4b753df44a35a72ac5eeaf89abeb3550b5e8f2a12abd12ce5121ca0a3ade26

Observation 0ca5f7fc-d1e3-4d10-bed2-67f6b19d3d6c · outbound

This paper cites WaveNet: A generative model for raw audio,.

FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment WaveNet: A generative model for raw audio,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:47:56.740901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T11:47:56.088120Z digest=sha256:69f726851f6df3c89b332e3ffcac31c8a82ffe94abedcea6856bde707cbb330a

Observation db6e4e98-995a-491e-a70a-4cb057015316 · outbound

This paper cites Joint detection and classification of singing voice melody using convolutional recurrent neural networks,.

FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Joint detection and classification of singing voice melody using convolutional recurrent neural networks,

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:47:56.730092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T11:47:56.091792Z digest=sha256:0d00f24128ac632ed07515ddb2e4dee7a17bea3d8639b101ba8e14433be44bfc

Observation 068cd80b-60b6-4734-bff1-13b49e7bd7bd · outbound

This paper cites RAFT: Recurrent all-pairs field transforms for optical flow,.

FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment RAFT: Recurrent all-pairs field transforms for optical flow,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:47:56.719150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T11:47:56.095278Z digest=sha256:e5f335024237fdcf4b0cd04de8abd8b0ba7f036b2bc401bc128e16c7fc49012b

Observation 39a82b1b-fcc3-4930-954c-8775ae79cebf · outbound

This paper cites Lever- aging temporal contextualization for video action recognition,.

FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Lever- aging temporal contextualization for video action recognition,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:47:56.706966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T11:47:56.099133Z digest=sha256:bc3b0668c5dbab93c1319eda6a7eb24a991b81650da7cf20cbdfd6ca1b7e8dd8

Observation eb9b6a76-7a31-4880-9ef1-c568b88c1927 · outbound

This paper cites Framewise phoneme classifica- tion with bidirectional LSTM networks,.

FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Framewise phoneme classifica- tion with bidirectional LSTM networks,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:47:56.697200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T11:47:56.103550Z digest=sha256:0270b8b9c3c0dfc3ef80264de3bf8d13b890571f1f18fa577f95bc8d8e76073f

Observation 209affa8-5626-4d4a-ba9e-655cbe034fc5 · outbound

This paper cites Stable Audio Open.

FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Stable Audio Open

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-11T11:47:56.107155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:47:56.107155Z digest=sha256:e9d0450d0ee0603577334f56ea8393a9b0328e7fac6476f20d824145e08a4e44

Observation efe75b57-ddd6-4b43-843c-54f8a35a9524 · outbound

This paper cites BYOL for Audio: Self-supervised learning for general- purpose audio representation,.

FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment BYOL for Audio: Self-supervised learning for general- purpose audio representation,

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:47:56.685766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T11:47:56.111119Z digest=sha256:03ffe4cf817912a45c45bde830de1f75b80d5cdfb715682696ed259ce6bd12c2

Observation a1cc6ba5-e531-4925-87f5-08e6dc816774 · outbound

This paper cites Large-scale contrastive language- audio pretraining with feature fusion and keyword-to-caption augmen- tation,.

FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Large-scale contrastive language- audio pretraining with feature fusion and keyword-to-caption augmen- tation,

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:47:56.673646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T11:47:56.114764Z digest=sha256:80813d2d592b01a1c71d1f7fa73446b9e9383d7cfb20a1120538cf8dce58f0b1

Observation dd118ad5-7255-409a-97c8-1e4b645258e1 · outbound

This paper cites Separate What You Describe: Language-queried audio source separation,.

FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Separate What You Describe: Language-queried audio source separation,

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:47:56.662355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T11:47:56.118463Z digest=sha256:738a28da160a7d7e7b08f83db50d3b660bc657f0675630c10408fa117970efac

Observation 1d0d1907-3ab2-4e8c-938b-44106dc9d413 · outbound

This paper cites Classifier-free diffusion guidance,.

FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Classifier-free diffusion guidance,

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:47:56.651821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T11:47:56.122212Z digest=sha256:cff72c1b9b4d2dadddea9350e401b982e67e00454aadfe32621a4635d8f0eb5c

Observation f1f65320-bac9-4e7a-8ab3-fbbaa8c7b2b2 · outbound

This paper cites Fr´echet audio distance: A reference-free metric for evaluating music enhancement algorithms,.

FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Fr´echet audio distance: A reference-free metric for evaluating music enhancement algorithms,

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:47:56.640162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T11:47:56.126185Z digest=sha256:0017c760814c66320960526eba7e30fc3a9a929474f3ced61ca3ee2792d1763b

Observation bd98a2bc-f8d9-402d-a874-9f8799e8ab33 · outbound

This paper cites Correlation of fr´echet audio distance with human perception of environmental audio is embedding dependent,.

FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Correlation of fr´echet audio distance with human perception of environmental audio is embedding dependent,

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:47:56.629481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T11:47:56.129917Z digest=sha256:05ceef14c00a8dd95e535642fde88e222545756b64d99c9f92d87d0b1b35d905

Observation 0f499230-d784-4568-8271-710b404a0fe4 · outbound

This paper cites PANNs: Large-scale pretrained audio neural networks for audio pattern recognition,.

FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment PANNs: Large-scale pretrained audio neural networks for audio pattern recognition,

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:47:56.619061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T11:47:56.133866Z digest=sha256:04859f076294907d44696cde11d3ff0a873fd6f5abe78e974dd2356764d9aaef

Observation 59f414aa-5e20-4b73-bdb1-cb8e633c1d54 · outbound

This paper cites CLAP learning audio concepts from natural language supervision,.

FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment CLAP learning audio concepts from natural language supervision,

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:47:56.608322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T11:47:56.137568Z digest=sha256:a7707a8072b2e88a79d7e5db02726471f635dd24b2bac0353ee676d5b2cbef90

Observation c485989e-4a57-4ad6-9791-c92f919420e0 · outbound

This paper cites Perceptual evaluation of audio-visual synchrony grounded in viewers’ opinion scores,.

FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Perceptual evaluation of audio-visual synchrony grounded in viewers’ opinion scores,

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:47:56.595896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T11:47:56.141142Z digest=sha256:46abd59d962ab7b2632727b426899df8459a316485697b49337c33f6c247bf53

Observation c4159048-d2ca-4cce-ac3f-ca0e1730e2f1 · outbound

This paper cites Quo vadis, action recognition? a new model and the kinetics dataset,.

FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Quo vadis, action recognition? a new model and the kinetics dataset,

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:47:56.584247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T11:47:56.144589Z digest=sha256:704c36368419d08ca117171a9f63bba113a5e639278e16e424d70320e4e61cbb

Observation 90645924-ade3-418d-872a-b5289a49db5f · outbound

This paper cites CNN architectures for large-scale audio classification,.

FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment CNN architectures for large-scale audio classification,

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:47:56.573308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T11:47:56.148545Z digest=sha256:013334022cddc5016c5db6e342f3dfc277ae02769cc61a8919bae1e6ae434dcc

Observation 9bf5f365-95b2-4bf0-a514-0232bbcbbc76 · outbound

This paper cites Batch normalization: accelerating deep network training by reducing internal covariate shift,.

FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Batch normalization: accelerating deep network training by reducing internal covariate shift,

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:47:56.561331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T11:47:56.152011Z digest=sha256:f66131cd06158b5d5fb12676c2616864433c4d6227d7d3f81bbd18b9885cacff

Observation 362ddce1-cf8f-492c-83aa-549100912ae0 · outbound

This paper cites ImageNet: A large-scale hierarchical image database,.

FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment ImageNet: A large-scale hierarchical image database,

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:47:56.549411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T11:47:56.155940Z digest=sha256:5414269a944cd77d59e108279c0c70fe5662992f9a5e9d47a13f69cf35fd2cbb

Observation ce3ccac1-d5a8-4c92-98e9-6380c20126b4 · outbound

This paper cites The Kinetics Human Action Video Dataset.

FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment The Kinetics Human Action Video Dataset

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-11T11:47:56.159595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:47:56.159595Z digest=sha256:ebff5f473b2216ae6336c41f3a307145c4ce9528ec006acc2c484c3dba126e90

Pith citing papers

Observation f584c6a1-2a11-49c5-8f75-1f9adfd55165 · inbound

Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model cites this paper.

Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T20:25:31.026317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:25:31.026317Z digest=sha256:c145d65842a4311f156becc79b213ecbc39f8b1be590474c5ea839c59ae6057b

Observation 6cb4d445-1e69-4e49-ba26-f9ebff2361a3 · inbound

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet cites this paper.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:56.308460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:56.308460Z digest=sha256:d644c01d052c2cc2011d3c7c23486e9b20cd520feae2dca20f7d799f9b01d336

Observation 72e809fd-2512-4928-8d4d-1335665e735b · inbound

SonicGauss: Position-Aware Physical Sound Synthesis for 3D Gaussian Representations cites this paper.

SonicGauss: Position-Aware Physical Sound Synthesis for 3D Gaussian Representations FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T14:03:59.912983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:03:59.912983Z digest=sha256:cdba3e9cc30c168645a55ec0ef683aad5aedf8ec737c26b93801c62f03136455

Observation 348a9a03-cc71-4736-bbfd-c8bbbd289a30 · inbound

Training-Free Multimodal Guidance for Video to Audio Generation cites this paper.

Training-Free Multimodal Guidance for Video to Audio Generation FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-21T22:20:42.069249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-21T22:17:42.348945Z digest=sha256:b78634ca468c97dabd609f5b7957fc5e7f757c78411c03b024270d5d3b7aaa8a