Pith. sign in

Paper Citation Record · LEDGER

FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

As of 19 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 47 inbound Pith citation observations for arXiv:2407.01494.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.01494 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 47 of 47 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:43:41.374558Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T07:14:45.319897Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c22d073e-5443-45d2-a638-8d142880e04c · inbound

Gotta Hear Them All: Towards Sound Source Aware Audio Generation cites this paper.

Gotta Hear Them All: Towards Sound Source Aware Audio Generation FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-12T14:22:12.022083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:22:12.022083Z digest=sha256:0615a43008a64b9b6f3424c75f96fae0f9f7a8e0c9b296c24295c20a4b2162b1

Observation d0b00aba-8be9-4d19-b1ff-ccebc15b3b6c · inbound

Video-Guided Foley Sound Generation with Multimodal Controls cites this paper.

Video-Guided Foley Sound Generation with Multimodal Controls FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-12T12:00:01.822180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:00:01.822180Z digest=sha256:5762905a3cd2e568881c89c910dd78cd295bdac73ef97b902672c29dbedd130a

Observation 9acad6b1-eb4d-4d90-974b-aa7626f93fec · inbound

YingSound: Video-Guided Sound Effects Generation with Multi-modal Chain-of-Thought Controls cites this paper.

YingSound: Video-Guided Sound Effects Generation with Multi-modal Chain-of-Thought Controls FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-11T17:17:49.308893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:17:49.308893Z digest=sha256:52223edbad1a406d267e1ebb0d7a22536d23713327922053518b91a11856c9c5

Observation 1a8583da-a916-4c0f-9ccc-8a02c6ce23b0 · inbound

AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation cites this paper.

AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

Reference 104

Resolution
unresolved
no resolver link, observed 2026-08-11T11:38:08.729854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:38:08.729854Z digest=sha256:c26f14d62c8196544acf1b674f7838a20746f2f5b4890a8c31c0d136efcb1131

Observation 8e99a702-9bdf-4d01-bc9b-39811cf81fad · inbound

MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis cites this paper.

MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-11T11:35:57.857179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:35:57.857179Z digest=sha256:7eec90b758e2135e96892325ae5c7d929bc72e160fa037df9ce0bd5e6dc00f77

Observation 82f2665c-0c2a-40d2-aeef-9889bc041184 · inbound

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance cites this paper.

Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T05:01:38.651937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:01:38.651937Z digest=sha256:bb9136cca914c2ceac54b138d78ae3477081f08b03fe53d78b8022b3cfc71d97

Observation 24cb6520-edae-4e6e-ab85-24c89b42df46 · inbound

AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment cites this paper.

AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-10T00:02:26.312439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T00:02:26.312439Z digest=sha256:fbabc223ca26c608801642c4b94bd50c50a3d3c86ad6e9ea65fbad0ee2b33335

Observation 3650c39e-68d2-4981-9145-f8c25141abf9 · inbound

UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation cites this paper.

UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T00:23:08.188006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T00:23:08.188006Z digest=sha256:ce11394c5175c00c564bd07159ed3a1483d79d70ef095a202994240b63db60f6

Observation a469cb36-96d5-4c18-89f6-96ce8ead16a7 · inbound

Lumina-Video: Efficient and Flexible Video Generation with Multi-scale Next-DiT cites this paper.

Lumina-Video: Efficient and Flexible Video Generation with Multi-scale Next-DiT FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-08T14:26:46.883688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:26:46.883688Z digest=sha256:82c07456e0f291afaa2494b704d1bc33f529e24e0aa6f0f9be8b53769a622ef8

Observation 461833f9-d760-429b-aab9-61cd51286cff · inbound

OmniAudio: Generating Spatial Audio from 360-Degree Video cites this paper.

OmniAudio: Generating Spatial Audio from 360-Degree Video FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-16T11:43:41.374558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:43:41.374558Z digest=sha256:f8e3db02f7e6d64570e01af05a1eb6bbff854379824c2fb1b78e4c3166a875c9

Observation 8fb3c112-9498-4a3f-b0fe-a785e763ebd2 · inbound

Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model cites this paper.

Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T20:25:31.019247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:25:31.019247Z digest=sha256:ccaf615841917dd60737974e263229d0cceba4d9059dd64274e8cd5c9198cf2f

Observation 72a83755-a8c4-4920-b509-daf337c91156 · inbound

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet cites this paper.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:55.910310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:55.910310Z digest=sha256:fa5eadf532bf5aeee075863f50d5acb333558af9de6b68c9f223ae17feddeaeb

Observation e5935e36-2c37-445e-81ae-9d1e08b05f4c · inbound

Towards Video to Piano Music Generation with Chain-of-Perform Support Benchmarks cites this paper.

Towards Video to Piano Music Generation with Chain-of-Perform Support Benchmarks FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:05:12.150659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:05:12.150659Z digest=sha256:749d6ae49c57335578ec8a08d0160457846b5bc6571a5f2acf6f9173935d74a7

Observation b264ce73-22ed-4027-9115-481e8ed68cd8 · inbound

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation cites this paper.

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:58.600648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:21:58.600648Z digest=sha256:129476f1dd24d57d69cf15dbbb1bcd84a49fc4f7256bb87177f100dde9742e45

Observation d226c09d-d252-4134-a793-847f7f69ca39 · inbound

Sounding that Object: Interactive Object-Aware Image to Audio Generation cites this paper.

Sounding that Object: Interactive Object-Aware Image to Audio Generation FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T10:53:17.851996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:53:17.851996Z digest=sha256:88e7911beb2330f3b884a8339e3b2b4ef06e88efdedc4d96dd9872e27a37874c

Observation f86c4cf0-2d82-4547-8a55-783902287549 · inbound

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance cites this paper.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:08.730625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:08.730625Z digest=sha256:6722472351489fc0034b2b87f0b02c05db48eeee4f3633d325f0adecbebfab6a

Observation 1c65f088-3fe7-4bc2-9256-114872a13097 · inbound

Spotlighting Partially Visible Cinematic Language for Video-to-Audio Generation via Self-distillation cites this paper.

Spotlighting Partially Visible Cinematic Language for Video-to-Audio Generation via Self-distillation FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T20:37:19.460025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:37:19.460025Z digest=sha256:0dd95c12f55c6ec8c101b45cb917cba0e172745629709be951d35be70a18d0a1

Observation 6607a917-256c-4f2f-bd66-2a9b9f04c558 · inbound

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation cites this paper.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T19:40:02.722459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:40:02.722459Z digest=sha256:7c3dc51acf4e433d1f62b37a7cff972e326ee6c1f4ca93457ab16028e8dcf5d3

Observation bfc0c11b-f339-45d8-9cc3-b579320ba730 · inbound

Hear-Your-Click: Interactive Object-Specific Video-to-Audio Generation cites this paper.

Hear-Your-Click: Interactive Object-Specific Video-to-Audio Generation FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T19:39:57.158925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:39:57.158925Z digest=sha256:71a5f65a9a737ac8c7b06992aa06bf6bf6920dd52e1cf7e8a9605b8793149f6e

Observation 66285bcf-7fed-49f0-8900-4f4d5a0d744f · inbound

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis cites this paper.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:18.900460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:18.900460Z digest=sha256:8bf4f42b8166fa585875b2ae4c098e741b98a0a5772da42cbbf77fdfdbfa1cd4

Observation 7e2735b9-945d-4da9-bf83-a8d4ac21bf7a · inbound

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation cites this paper.

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T06:04:29.971825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T06:04:29.971825Z digest=sha256:ce40f16537386a579f4bf3be313e535a7740d592b2f5424febcea153136bce12

Observation 961eab08-0c24-4cff-80d1-e4afeacfda5f · inbound

LD-LAudio-V1: Video-to-Long-Form-Audio Generation Extension with Dual Lightweight Adapters cites this paper.

LD-LAudio-V1: Video-to-Long-Form-Audio Generation Extension with Dual Lightweight Adapters FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T17:43:11.900566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:43:11.900566Z digest=sha256:cb8fad70d546dbcf80984f26ad9d248ebe789d92ea528c41072f87b7de801af6

Observation 74f1ff39-134b-41a1-b0fa-820359ac3e42 · inbound

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper cites this paper.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-05T05:50:17.728841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:50:17.728841Z digest=sha256:66e7128f3003e9ab4dc6590a246af6d29d41677a5130573d3a515770b49cd6d7

Observation 0dcbd26a-08fd-4884-adeb-b9529976d844 · inbound

MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation cites this paper.

MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:32.481088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:46:32.481088Z digest=sha256:c0e463e0f1b12bb09943ba34a8f6c681a2197b149a027c89e98536c652ea1012

Observation 30b453a0-6a9a-4f6e-a2d1-f92dbce83eef · inbound

StereoFoley: Object-Aware Stereo Audio Generation from Video cites this paper.

StereoFoley: Object-Aware Stereo Audio Generation from Video FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:26:28.615296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T14:23:08.144599Z digest=sha256:c0d8455c38d366151afd788bec87601238973fec4278ae28041e6ac8cb847887

Observation bb145c55-3032-4ca0-b7ad-a7d9cd1ec4c1 · inbound

AudioMoG: Guiding Audio Generation with Mixture-of-Guidance cites this paper.

AudioMoG: Guiding Audio Generation with Mixture-of-Guidance FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:11:23.863177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-18T13:10:18.700497Z digest=sha256:6d2c6257ba4e2b2b1ff6db937c27e05abd92fdbef6b7f562111ff40782cb783b

Observation e9523196-8c24-4383-8b41-a806591b514e · inbound

NoiseShift: Resolution-Aware Noise Recalibration for Better Low-Resolution Image Generation cites this paper.

NoiseShift: Resolution-Aware Noise Recalibration for Better Low-Resolution Image Generation FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-21T21:44:22.914024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-21T21:41:29.502472Z digest=sha256:4bc01e8f408681614aa9b43b266a1acc2dee44a913a0b0c2208c7b7274327add

Observation efdf7318-8259-4684-9be2-2691000d5af7 · inbound

MMAudioSep: Taming Video-to-Audio Generative Model Towards Video/Text-Queried Sound Separation cites this paper.

MMAudioSep: Taming Video-to-Audio Generative Model Towards Video/Text-Queried Sound Separation FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-18T08:21:06.832653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T08:20:02.986562Z digest=sha256:6452065fd34d189bf58eb4779e095fa6c9ffd9d8156df038938639d954863f85

Observation 270cdb8c-cd4a-449e-b745-efa3bffbc071 · inbound

MVAD: A Benchmark Dataset for Multimodal AI-Generated Video-Audio Detection cites this paper.

MVAD: A Benchmark Dataset for Multimodal AI-Generated Video-Audio Detection FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:48:58.459651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-17T03:48:26.807495Z digest=sha256:1ce0f1ebea5c53234b2509ef382603d3fa575080286103ca9297fc780fcd7a63

Observation fd5b86f8-7e4b-42c5-aa90-1aaa05de5c4d · inbound

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing cites this paper.

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-03T16:26:00.172036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:26:00.172036Z digest=sha256:95c2d6c7c11b7e84d87f3eb074b1ea0e2a93a40135882bba8637df244f41e41f

Observation 38dca50c-d3e2-4ae5-8c0a-987523629705 · inbound

Aliasing-Free Neural Audio Synthesis cites this paper.

Aliasing-Free Neural Audio Synthesis FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-16T20:38:24.819824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-16T20:34:50.539351Z digest=sha256:825e71c799d5e389cdde5b1327110c800c16ad8d9d232d46eb4c297654cdf7fa

Observation aacf3a6c-c988-4296-baaf-0a0eb3aae606 · inbound

Aliasing-Free Neural Audio Synthesis cites this paper.

Aliasing-Free Neural Audio Synthesis FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-03T14:33:13.284553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:33:13.284553Z digest=sha256:1712a2d19c3a1446788fa67116949d18037d0de4e83e29c54480fe516cee067f

Observation 60b34967-c7dc-44ea-8326-832959ce304e · inbound

PhyAVBench: A Challenging Audio Physics-Sensitivity Benchmark for Physically Grounded Text-to-Audio-Video Generation cites this paper.

PhyAVBench: A Challenging Audio Physics-Sensitivity Benchmark for Physically Grounded Text-to-Audio-Video Generation FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-16T19:48:21.873213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-16T19:43:37.604351Z digest=sha256:99d10e108a0d54d6d91e85c9eb42f14f5f60a0a0457c655df6fd447eca2db075

Observation d368668f-7ac6-4ef1-af97-4dfc74e2fa66 · inbound

PhyAVBench: A Challenging Audio Physics-Sensitivity Benchmark for Physically Grounded Text-to-Audio-Video Generation cites this paper.

PhyAVBench: A Challenging Audio Physics-Sensitivity Benchmark for Physically Grounded Text-to-Audio-Video Generation FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-21T16:14:15.286466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-21T16:10:31.015783Z digest=sha256:e5748c553d700c32ac186db2ca38625e890f3cff011e0beb766fc9aa661fa9d7

Observation 024b790b-0065-4ec2-a200-819594430712 · inbound

EchoFoley: Event-Centric Hierarchical Control for Video Grounded Creative Sound Generation cites this paper.

EchoFoley: Event-Centric Hierarchical Control for Video Grounded Creative Sound Generation FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-03T13:21:11.469829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:21:11.469829Z digest=sha256:3096a4be063d3d971c562c18b7ae65f0a97616b6f020a9d0b26d38c0cc5cc63a

Observation 7c68a5e7-5887-4578-a2f9-067e55854cac · inbound

LTX-2: Efficient Joint Audio-Visual Foundation Model cites this paper.

LTX-2: Efficient Joint Audio-Visual Foundation Model FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:06:20.588495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-13T07:06:20.470686Z digest=sha256:e8b0bef1e32176424c35579373bda5665e2b8c5b5987f565e9bf1681882c543d

Observation 7ba0f5e0-f392-44e4-9998-45b1d7a2da2e · inbound

Echoes Over Time: Unlocking Length Generalization in Video-to-Audio Generation Models cites this paper.

Echoes Over Time: Unlocking Length Generalization in Video-to-Audio Generation Models FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-15T19:56:33.536685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T19:53:18.200223Z digest=sha256:c6a162a41ab46d7fb8475759b994cd32b021735dbe69c22dbc4b173d7005230d

Observation 652fc853-fccd-41c3-a7cc-f63e44889002 · inbound

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text cites this paper.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:00:47.523966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:ccecdd2fc93f84a382678a733bdf87112cd11c0c95a6314a38fea3510ccb074a

Observation dc06ec0a-a6f5-48ab-8001-6162eefc0fe6 · inbound

FoleyDesigner: Immersive Stereo Foley Generation with Precise Spatio-Temporal Alignment for Film Clips cites this paper.

FoleyDesigner: Immersive Stereo Foley Generation with Precise Spatio-Temporal Alignment for Film Clips FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:40:51.786766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T18:57:21.434793Z digest=sha256:f0551d825fafe60e455e6251550fe33407a1f92605b9f63e7ff3e235cd68383b

Observation 586985e3-390f-4ee7-8c32-fdf4ec9cb9eb · inbound

Geo2Sound: A Scalable Geo-Aligned Framework for Soundscape Generation from Satellite Imagery cites this paper.

Geo2Sound: A Scalable Geo-Aligned Framework for Soundscape Generation from Satellite Imagery FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:59:03.509778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T09:52:35.741400Z digest=sha256:df99f7823881015479d27690b74983fcdc1f27123402089407c70be673c8a47d

Observation db1c1ca5-c3f1-4f33-9b6d-21f3e19ff80a · inbound

MMAudio-LABEL: Audio Event Labeling via Audio Generation for Silent Video cites this paper.

MMAudio-LABEL: Audio Event Labeling via Audio Generation for Silent Video FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:01:22.064276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-09T18:53:07.364051Z digest=sha256:23d64f84bf5184af652a4c50af73ed52d2b0495227a4b4ebb225d945d2ea8e85

Observation c78bf827-1d7e-43b2-9537-4b9a10808678 · inbound

LongAV-Compass: Towards Unified Evaluation of Minute-Scale Audio-Visual Generation Across T2AV, I2AV, and V2AV cites this paper.

LongAV-Compass: Towards Unified Evaluation of Minute-Scale Audio-Visual Generation Across T2AV, I2AV, and V2AV FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:54:00.751695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-29T22:52:38.330851Z digest=sha256:ada56b8317b09f66fde26bf722391823b87b5b4d1189d4d7a6b5284863c1ca55

Observation ad1a69ed-29b4-4851-9192-a1e478a4ffd5 · inbound

Benchmarking Single-Factor Physical Video-to-Audio Generation cites this paper.

Benchmarking Single-Factor Physical Video-to-Audio Generation FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

Reference 78

Resolution
verified exact
arxiv_id, observed 2026-06-29T07:43:13.374668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-29T07:41:56.917119Z digest=sha256:346d8c32f5a0eb3f4f2006d91c50f461f7ce89ec7b8d43424e50e8f523eb6969

Observation b36f4850-85bc-4d41-ac98-0c0103d423e0 · inbound

AudioX-Turbo: A Unified Framework for Efficient Anything-to-Audio Generation cites this paper.

AudioX-Turbo: A Unified Framework for Efficient Anything-to-Audio Generation FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-03T13:28:18.788790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T08:04:48.283908Z digest=sha256:e702486ecb6aa7b2dc4ac022b63c79f084a2ce18f4583b8f1a5d1e31fe0326ea

Observation 25a17be2-65fb-4005-b514-56a14b719d22 · inbound

Precise Video-to-Audio Generation with Cross-Modal Alignment in Latent Space cites this paper.

Precise Video-to-Audio Generation with Cross-Modal Alignment in Latent Space FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-07-08T07:14:45.321352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T07:10:20.519266Z digest=sha256:383ef2b90fbbc25cb95d771340a5011a52ad39dd2f83300c21d1c2ca3b356445

Observation bdea5d69-94cc-4cd0-a427-0d4cbc8e5e53 · inbound

Precise Video-to-Audio Generation with Cross-Modal Alignment in Latent Space cites this paper.

Precise Video-to-Audio Generation with Cross-Modal Alignment in Latent Space FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T08:23:35.864870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:23:35.864870Z digest=sha256:6ad51f530a94585879e68a47c10dff1dde4a08f2d24b84222d6e6836b10842b8

Observation 52ae80db-ecd6-4720-bd3e-77e755cb7060 · inbound

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion cites this paper.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T00:27:56.735765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:27:56.735765Z digest=sha256:aec0a869cff5981ef0a68153b3b408f4eb44d650b7be49b86100599095a2cd91