Pith. sign in

Paper Citation Record · LEDGER

FoleyDesigner: Immersive Stereo Foley Generation with Precise Spatio-Temporal Alignment for Film Clips

As of 25 July 2026, this Paper Citation Record lists 51 of 51 outbound references and 0 inbound Pith citation observations for arXiv:2604.05731.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.05731 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T18:57:21.434793Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-07-24T06:31:00.690269+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

51 of 51 outbound references displayed

  • verified exact11
  • verified fuzzy38
  • unresolved0
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e051faa1-c25e-473d-8d22-09ff1e4fbc43 · outbound

This paper cites Qwen2.5-vl technical report.

FoleyDesigner: Immersive Stereo Foley Generation with Precise Spatio-Temporal Alignment for Film Clips Qwen2.5-vl technical report

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:52:49.886971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-10T18:57:21.434793Z digest=sha256:8618c81222aace18e5ad655c379a1ace8627a39faead796d0ba4c528866a1e1a

Observation 2365880e-7537-46ec-b788-2c6f5a3be32f · outbound

This paper cites An improved event- independent network for polyphonic sound event localiza- tion and detection.

FoleyDesigner: Immersive Stereo Foley Generation with Precise Spatio-Temporal Alignment for Film Clips An improved event- independent network for polyphonic sound event localiza- tion and detection

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:52:49.889152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-10T18:57:21.434793Z digest=sha256:3dcdef96457ab96045ca2cdf590eae183272d088d5fdb78e5b2e5ebac90fe3eb

Observation 1b426b2e-7c5a-4b8c-891b-a1ee1aab87c6 · outbound

This paper cites Video-guided foley sound generation with multimodal con- trols.

FoleyDesigner: Immersive Stereo Foley Generation with Precise Spatio-Temporal Alignment for Film Clips Video-guided foley sound generation with multimodal con- trols

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:52:49.891118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-10T18:57:21.434793Z digest=sha256:570831b9bbabe520ec8ec89881f570d7cca560c2427c9d12c561708ace1f58c5

Observation 7244bc48-d4f4-41f9-9ec9-caa3a12af7b1 · outbound

This paper cites MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis.

FoleyDesigner: Immersive Stereo Foley Generation with Precise Spatio-Temporal Alignment for Film Clips MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:40:51.764241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-10T18:57:21.434793Z digest=sha256:000f3e83955ca021d57868a2cd0b92ffd9f2f9091ef15edee8d71c6eb806b75d

Observation c879398c-5410-45c8-b011-e0edb2ffcbaa · outbound

This paper cites Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models.

FoleyDesigner: Immersive Stereo Foley Generation with Precise Spatio-Temporal Alignment for Film Clips Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:57:28.886773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-10T18:57:21.434793Z digest=sha256:fea224856dac1e6bef60914aef2ea1ca5cbcfccf50ee0b2891975efd683335ac

Observation 43236040-cd06-4176-a8a2-1fe3de51d711 · outbound

This paper cites Sim- ple and controllable music generation.Advances in Neural Information Processing Systems, 36:47704–47720.

FoleyDesigner: Immersive Stereo Foley Generation with Precise Spatio-Temporal Alignment for Film Clips Sim- ple and controllable music generation.Advances in Neural Information Processing Systems, 36:47704–47720

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:52:49.897292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-10T18:57:21.434793Z digest=sha256:b64703048d708660f67245dada272061ab375b74d02eb68148db39aacd256eba

Observation 6860ccf8-c406-4958-b2d8-4a62c6da381a · outbound

This paper cites SEE-2-SOUND: Zero-Shot Spatial Environment-to-Spatial Sound.

FoleyDesigner: Immersive Stereo Foley Generation with Precise Spatio-Temporal Alignment for Film Clips SEE-2-SOUND: Zero-Shot Spatial Environment-to-Spatial Sound

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:40:51.738420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-10T18:57:21.434793Z digest=sha256:058a1b20e7edc02715b72fafe5661173a1fb5db2b8cfba12655039a5e62c6d77

Observation e90e8e0d-e00c-4417-aac6-b5f7b68ff54a · outbound

This paper cites an unresolved cited work.

FoleyDesigner: Immersive Stereo Foley Generation with Precise Spatio-Temporal Alignment for Film Clips Unresolved cited work

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:52:49.895173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-10T18:57:21.434793Z digest=sha256:ce1de56c1b873fb2b5a4174e91603e91187d19056fb8e7fe2ff0b3ea23220f22

Observation 17149d09-7eb6-46c6-8ff7-6c8ca0a9c0b3 · outbound

This paper cites CLIPSonic: Text-to-audio synthesis with unlabeled videos and pretrained language- vision models.

FoleyDesigner: Immersive Stereo Foley Generation with Precise Spatio-Temporal Alignment for Film Clips CLIPSonic: Text-to-audio synthesis with unlabeled videos and pretrained language- vision models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:52:49.893437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-10T18:57:21.434793Z digest=sha256:5be5dcbf1bedfec70f3a396b3a19c056a42488374e9d09d9668cd1079283c43f

Observation 16e52a92-bad3-4496-8d24-364ba2409bc1 · outbound

This paper cites Conditional generation of audio from video via foley analogies.

FoleyDesigner: Immersive Stereo Foley Generation with Precise Spatio-Temporal Alignment for Film Clips Conditional generation of audio from video via foley analogies

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:52:49.830044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-10T18:57:21.434793Z digest=sha256:0d8b50847171771c350f26cb25b43809339dea1538f4f27c79e0377fa1c4fb5f

Observation 73787ab8-a8fb-43d0-b540-d0fadb9ef702 · outbound

This paper cites Stable audio open.

FoleyDesigner: Immersive Stereo Foley Generation with Precise Spatio-Temporal Alignment for Film Clips Stable audio open

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:52:49.836371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-10T18:57:21.434793Z digest=sha256:247027c06c53f56012c6df002994ec6e75e542c450b959c1f1e204343661cf48

Observation 2a1e3fa2-6d33-4fc3-a82c-d1331f3e74b5 · outbound

This paper cites Geometry-Aware Multi-Task Learning for Binaural Audio Generation from Video.

FoleyDesigner: Immersive Stereo Foley Generation with Precise Spatio-Temporal Alignment for Film Clips Geometry-Aware Multi-Task Learning for Binaural Audio Generation from Video

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:40:51.732710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-10T18:57:21.434793Z digest=sha256:52e38c528b36e0a7a2ae436598e21eae47afb186db9507174003e049eb6a3855

Observation 1ac95da9-2deb-45d0-a547-ae5de59698b3 · outbound

This paper cites Imagebind one embedding space to bind them all.

FoleyDesigner: Immersive Stereo Foley Generation with Precise Spatio-Temporal Alignment for Film Clips Imagebind one embedding space to bind them all

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:52:49.840578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-10T18:57:21.434793Z digest=sha256:eb289e19aa8f71c50a3e7683e3233e2da1fe13ee67f4ae64c82fc452d9d07b0d

Observation 894cc288-cee6-42f5-a815-a84fb33e6dd0 · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilib- rium.

FoleyDesigner: Immersive Stereo Foley Generation with Precise Spatio-Temporal Alignment for Film Clips Gans trained by a two time-scale update rule converge to a local nash equilib- rium

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:52:49.826310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-10T18:57:21.434793Z digest=sha256:81e5bb5d22c7f5613b6f27f8a8d54cb5c9bed5fda8bd861b026c7a568e45e845

Observation db0ac4a9-8b70-42ce-b06a-619205c133e6 · outbound

This paper cites Spotlight- ing partially visible cinematic language for video-to-audio generation via self-distillation.

FoleyDesigner: Immersive Stereo Foley Generation with Precise Spatio-Temporal Alignment for Film Clips Spotlight- ing partially visible cinematic language for video-to-audio generation via self-distillation

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:52:49.881310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-10T18:57:21.434793Z digest=sha256:08f5f553d6fda431d8d42d35e92c179cb098d468c239d1d6ee8016e5d3c2918c

Observation 68f9d1ff-0d88-4dbb-9185-470e3099b7b9 · outbound

This paper cites Make-An-Audio 2: Temporal-Enhanced Text-to-Audio Generation.

FoleyDesigner: Immersive Stereo Foley Generation with Precise Spatio-Temporal Alignment for Film Clips Make-An-Audio 2: Temporal-Enhanced Text-to-Audio Generation

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:40:51.769570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-10T18:57:21.434793Z digest=sha256:d9ace121d3429612c5651e5cfc7d8480223b8abbcb6c01c608b555c32d2b8a47

Observation e479394d-c728-4728-b2af-db770be515b0 · outbound

This paper cites Synchformer: Efficient synchronization from sparse cues.

FoleyDesigner: Immersive Stereo Foley Generation with Precise Spatio-Temporal Alignment for Film Clips Synchformer: Efficient synchronization from sparse cues

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:52:49.862138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-10T18:57:21.434793Z digest=sha256:dc432eb89acdde81a972f135a6b01972dfb5d3192d5792f1ec643df53fc97542

Observation 9fe76c6a-cfb7-4a13-97fe-35ab1e262f60 · outbound

This paper cites Multichannel stereophonic sound system with and without accompanying picture.

FoleyDesigner: Immersive Stereo Foley Generation with Precise Spatio-Temporal Alignment for Film Clips Multichannel stereophonic sound system with and without accompanying picture

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:52:49.824066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-10T18:57:21.434793Z digest=sha256:78ecca01e72b3eef2be8a1ec3b42d27d4112581b75033c9be44a3f2a936e765c

Observation f8be0e67-d50e-4bd8-b547-4173cd2fb180 · outbound

This paper cites Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms.

FoleyDesigner: Immersive Stereo Foley Generation with Precise Spatio-Temporal Alignment for Film Clips Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:40:51.749046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-10T18:57:21.434793Z digest=sha256:c36ee668225ffe0871d5f0a8be1aa4f4a5870d9316faa6fc9f490b0f6b403fc6

Observation 97460d33-4d3c-4005-8e54-fe74b6302896 · outbound

This paper cites AudioGen: Textually Guided Audio Generation.

FoleyDesigner: Immersive Stereo Foley Generation with Precise Spatio-Temporal Alignment for Film Clips AudioGen: Textually Guided Audio Generation

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:40:51.743617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-10T18:57:21.434793Z digest=sha256:641720cdf5c5c0b1a3fac503f6bd1a9b5dd8a0a48750aea031969bbe4a552ea7

Observation 91232027-99fe-4d63-861a-8c7e23de9ed5 · outbound

This paper cites Video-Foley: Two-Stage Video-To-Sound generation via temporal event condition for foley sound.

FoleyDesigner: Immersive Stereo Foley Generation with Precise Spatio-Temporal Alignment for Film Clips Video-Foley: Two-Stage Video-To-Sound generation via temporal event condition for foley sound

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:40:51.780133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-10T18:57:21.434793Z digest=sha256:c602b25801ac0397d8a349620d858c768a69b44df3be90240088a8fd104d6e6e

Observation 01351671-3543-4f01-a0f1-0ac538ec27e2 · outbound

This paper cites AudioLDM: Text-to-Audio Generation with Latent Diffusion Models.

FoleyDesigner: Immersive Stereo Foley Generation with Precise Spatio-Temporal Alignment for Film Clips AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:40:51.725801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-10T18:57:21.434793Z digest=sha256:8782ad0d679ba60e97b8ba703f1e6c13087a4547a3ff148e4b159f99574571e7

Observation 62bbf88a-41c0-4175-8a65-ac6ae1c9b5a6 · outbound

This paper cites Plumbley.

FoleyDesigner: Immersive Stereo Foley Generation with Precise Spatio-Temporal Alignment for Film Clips Plumbley

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:52:49.828075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-10T18:57:21.434793Z digest=sha256:befc761876f628a1e985772fe939634b927834ef9221c88c5743ec1e289d0d4e

Observation 76375638-67a5-4d7c-81d9-affa66357b59 · outbound

This paper cites OmniAudio: Generating Spatial Audio from 360-Degree Video.

FoleyDesigner: Immersive Stereo Foley Generation with Precise Spatio-Temporal Alignment for Film Clips OmniAudio: Generating Spatial Audio from 360-Degree Video

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:40:51.754194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-10T18:57:21.434793Z digest=sha256:9ca9a08a332909a691755e9bfbea48fc71d22f877356964a4201c538fa53bdf1

Observation f2735f0f-4ef8-4111-a327-cfd62df93fb0 · outbound

This paper cites Visu- ally guided binaural audio generation with cross-modal con- sistency.

FoleyDesigner: Immersive Stereo Foley Generation with Precise Spatio-Temporal Alignment for Film Clips Visu- ally guided binaural audio generation with cross-modal con- sistency

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:52:49.822189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-10T18:57:21.434793Z digest=sha256:bf19325c2b829d27c36f6701d5413156ef58f70f9d803b187dbd69449bec1c28

Observation 061ae5d4-fbd9-4589-a7a8-5ee9da2c3f6f · outbound

This paper cites Diff-Foley: Synchronized video-to-audio synthesis with la- tent diffusion models.Advances in Neural Information Pro- cessing Systems, 36:48855–48876.

FoleyDesigner: Immersive Stereo Foley Generation with Precise Spatio-Temporal Alignment for Film Clips Diff-Foley: Synchronized video-to-audio synthesis with la- tent diffusion models.Advances in Neural Information Pro- cessing Systems, 36:48855–48876

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:52:49.817932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-10T18:57:21.434793Z digest=sha256:62736e7bff32b82347bb358abc84f3c655e2d65dc070b8052a8f920f7277d4eb

Observation aa5850e8-aa30-416f-8631-5abe6785eb69 · outbound

This paper cites Tango 2: Aligning diffusion-based text-to-audio generations through direct preference optimization.

FoleyDesigner: Immersive Stereo Foley Generation with Precise Spatio-Temporal Alignment for Film Clips Tango 2: Aligning diffusion-based text-to-audio generations through direct preference optimization

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:52:49.820063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-10T18:57:21.434793Z digest=sha256:e12e9aa71b5d4be39f185b1058029a6e6a807f9ea04595222a9815ebb064ab7b

Observation 0c464d35-75b0-46d6-9661-2e6eb0cf9d09 · outbound

This paper cites A probabilistic model for robust localization based on a binau- ral auditory front-end.IEEE Transactions on Audio, Speech, and Language Processing, 19(1):1–13.

FoleyDesigner: Immersive Stereo Foley Generation with Precise Spatio-Temporal Alignment for Film Clips A probabilistic model for robust localization based on a binau- ral auditory front-end.IEEE Transactions on Audio, Speech, and Language Processing, 19(1):1–13

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:52:49.832581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-10T18:57:21.434793Z digest=sha256:620fb27e47dd91266ee01b9da3f5c506a322cbbdaf83fc2275655e6176ce8a36

Observation 2536962a-1811-4e94-aaec-db161e8fdbd3 · outbound

This paper cites Beyond mono to binaural: Generating binaural audio from mono audio with depth and cross modal atten- tion.

FoleyDesigner: Immersive Stereo Foley Generation with Precise Spatio-Temporal Alignment for Film Clips Beyond mono to binaural: Generating binaural audio from mono audio with depth and cross modal atten- tion

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:52:49.838430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-10T18:57:21.434793Z digest=sha256:c822bf02e1122025fa5a6444c1597985596da8d5a96a0bb18a65ffbef289346d

Observation bef8906c-96b7-4835-8882-8b208eaec83c · outbound

This paper cites Improved techniques for training gans.

FoleyDesigner: Immersive Stereo Foley Generation with Precise Spatio-Temporal Alignment for Film Clips Improved techniques for training gans

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:52:49.854007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-10T18:57:21.434793Z digest=sha256:b2a32fad2ea27637767dc4ca95d65f82ca71facc0f78b36f71dc42d0066f04b5

Observation 6b22086f-cf6f-437d-8c6f-2c1d720cd267 · outbound

This paper cites I hear your true colors: Image guided audio generation.

FoleyDesigner: Immersive Stereo Foley Generation with Precise Spatio-Temporal Alignment for Film Clips I hear your true colors: Image guided audio generation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:52:49.874976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-10T18:57:21.434793Z digest=sha256:d9eed7e9a14d3d7d4711727daf2a61415e958eb16021d2c9499a2db1c7aa9959

Observation 05ee09f0-d164-45c2-988a-941acd3d7775 · outbound

This paper cites Both Ears Wide Open: Towards Language-Driven Spatial Audio Generation.

FoleyDesigner: Immersive Stereo Foley Generation with Precise Spatio-Temporal Alignment for Film Clips Both Ears Wide Open: Towards Language-Driven Spatial Audio Generation

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:40:51.721552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-10T18:57:21.434793Z digest=sha256:f62dc75832ad8839f6d60aa32996618f215e365d829535f8b5a55229fa83f1fd

Observation 0561bca9-0b6a-43f1-870f-514ff56beb58 · outbound

This paper cites Fourier features let networks learn high frequency functions in low dimen- sional domains.Advances in Neural Information Processing Systems, 33:7537–7547.

FoleyDesigner: Immersive Stereo Foley Generation with Precise Spatio-Temporal Alignment for Film Clips Fourier features let networks learn high frequency functions in low dimen- sional domains.Advances in Neural Information Processing Systems, 33:7537–7547

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:52:49.883370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-10T18:57:21.434793Z digest=sha256:eb496c1b27cd893afc3f6189a496fd814498d775c8c7e39d8df6469047a9005e

Observation df1dadb4-0f33-4de7-8478-b12b5a56957b · outbound

This paper cites V2a-mapper: A lightweight solution for vision-to-audio generation by connecting foun- dation models.

FoleyDesigner: Immersive Stereo Foley Generation with Precise Spatio-Temporal Alignment for Film Clips V2a-mapper: A lightweight solution for vision-to-audio generation by connecting foun- dation models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:52:49.877076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-10T18:57:21.434793Z digest=sha256:d3b7e3dd27c8c35187d04a5c4aeab41aa4251e698d44b7ec95e00a3133465cb1

Observation ef8591c0-f59d-4d71-8bf3-0b78a1861bdb · outbound

This paper cites Frieren: Efficient video-to-audio generation network with rectified flow matching.Advances in Neural Information Processing Systems, 37:128118–128138.

FoleyDesigner: Immersive Stereo Foley Generation with Precise Spatio-Temporal Alignment for Film Clips Frieren: Efficient video-to-audio generation network with rectified flow matching.Advances in Neural Information Processing Systems, 37:128118–128138

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:52:49.860203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-10T18:57:21.434793Z digest=sha256:99baa5b42863374af11a60cfd3b1683eba4f79b6e93dd960e91ed302d6f3522b

Observation 76088393-10d0-4183-a976-54158d1db3e0 · outbound

This paper cites Large-scale con- trastive language-audio pretraining with feature fusion and keyword-to-caption augmentation.

FoleyDesigner: Immersive Stereo Foley Generation with Precise Spatio-Temporal Alignment for Film Clips Large-scale con- trastive language-audio pretraining with feature fusion and keyword-to-caption augmentation

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:52:49.866436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-10T18:57:21.434793Z digest=sha256:1435c7c316d7f58c6f7e482f1db7ed89c108f19481a3de6e11569ef210ac4c84

Observation 696d8ed2-5b2b-4b2c-9647-a23446dd455e · outbound

This paper cites Son- icVisionLM: Playing sound with vision language models.

FoleyDesigner: Immersive Stereo Foley Generation with Precise Spatio-Temporal Alignment for Film Clips Son- icVisionLM: Playing sound with vision language models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:52:49.834560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-10T18:57:21.434793Z digest=sha256:63d62daa43e1ece7a2708dabe0bc5f3bafb8be70d3fd25d42c9924f5bade35a8

Observation 2e8e7252-e221-4ba9-9800-b51d1a5904f9 · outbound

This paper cites Depth any- thing v2.Advances in Neural Information Processing Sys- tems, 37:21875–21911.

FoleyDesigner: Immersive Stereo Foley Generation with Precise Spatio-Temporal Alignment for Film Clips Depth any- thing v2.Advances in Neural Information Processing Sys- tems, 37:21875–21911

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:52:49.868396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-10T18:57:21.434793Z digest=sha256:c590f12384ddc4ad7ff620aee66a1cbd0d11cd21dbcf84c30be68d52691a765e

Observation 84d6b5ee-bbdf-4b58-bf3f-a2f69a88a80b · outbound

This paper cites Deepear: Sound local- ization with binaural microphones.IEEE Transactions on Mobile Computing, 23(1):359–375.

FoleyDesigner: Immersive Stereo Foley Generation with Precise Spatio-Temporal Alignment for Film Clips Deepear: Sound local- ization with binaural microphones.IEEE Transactions on Mobile Computing, 23(1):359–375

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:52:49.864370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-10T18:57:21.434793Z digest=sha256:7f05890fa237b1f64c77bfe0a84b17ece760c64bfb6a153cbb139a669c1639e1

Observation 5d92dd1e-ece9-4584-91e9-dd33c5066dfd · outbound

This paper cites Griffiths, Yuan Cao, and Karthik Narasimhan.

FoleyDesigner: Immersive Stereo Foley Generation with Precise Spatio-Temporal Alignment for Film Clips Griffiths, Yuan Cao, and Karthik Narasimhan

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:52:49.858035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-10T18:57:21.434793Z digest=sha256:3fba9dd1559597144e8ba02660041f949d0e1c540a818b8c1b8116e97d80959a

Observation dc06ec0a-a6f5-48ab-8001-6162eefc0fe6 · outbound

This paper cites FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds.

FoleyDesigner: Immersive Stereo Foley Generation with Precise Spatio-Temporal Alignment for Film Clips FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:40:51.786766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-10T18:57:21.434793Z digest=sha256:eef12d76f4835c45d1042cf1132f6bfecb28ab6b2af8ac0991c8fd89a82e55d4

Observation 045f7a90-4c7f-44d2-8c42-055e5f461c16 · outbound

This paper cites Sep-stereo: Visually guided stereophonic audio generation by associating source separation.

FoleyDesigner: Immersive Stereo Foley Generation with Precise Spatio-Temporal Alignment for Film Clips Sep-stereo: Visually guided stereophonic audio generation by associating source separation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:52:49.872972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-10T18:57:21.434793Z digest=sha256:82f82eae2a7b22a359d547dfcf341d6b9cc23647f34490a46e4fb1176c08b39a

Observation c06f1e28-ad7c-43a3-8d95-0460b36dfddd · outbound

This paper cites As summarized in Table 1, ex- isting datasets typically lack stereophonic recordings or pre- cise temporal annotations.

FoleyDesigner: Immersive Stereo Foley Generation with Precise Spatio-Temporal Alignment for Film Clips As summarized in Table 1, ex- isting datasets typically lack stereophonic recordings or pre- cise temporal annotations

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:52:49.879089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-10T18:57:21.434793Z digest=sha256:b557c1c32e491aba77ff3b5ab7168348148ad3b6209f72d471cdc1f3f0b02ca5

Observation 059d1100-85e1-49aa-bf3b-f74fb41e0fe6 · outbound

This paper cites Multi-Agent Refinement The complete pseudocode for our multi-agent Foley refine- ment pipeline is presented in Algorithm 2.

FoleyDesigner: Immersive Stereo Foley Generation with Precise Spatio-Temporal Alignment for Film Clips Multi-Agent Refinement The complete pseudocode for our multi-agent Foley refine- ment pipeline is presented in Algorithm 2

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:52:49.842697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-10T18:57:21.434793Z digest=sha256:570832ce58874d3f5df09de8f4a2b8b044486bdeb9a132f01fdefbefe66fa951

Observation 440d6669-ab40-4ccd-be21-4b7f85304f6b · outbound

This paper cites an unresolved cited work.

FoleyDesigner: Immersive Stereo Foley Generation with Precise Spatio-Temporal Alignment for Film Clips Unresolved cited work

Reference 45

Resolution
parse uncertain
raw_fallback, observed 2026-05-16T11:52:49.885153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-10T18:57:21.434793Z digest=sha256:6dbe0b711452e2de5194ffb4c443a5b94768b70381ae301036b841f5ef8bc40d

Observation a47067db-d282-4a67-97e1-9e1883691d9b · outbound

This paper cites an unresolved cited work.

FoleyDesigner: Immersive Stereo Foley Generation with Precise Spatio-Temporal Alignment for Film Clips Unresolved cited work

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:52:49.852072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-10T18:57:21.434793Z digest=sha256:ba8e77380d87cd0833815411285ff34a9d7e92b0bdc3af64cf59b3850280eedb

Observation 161a1251-b443-44a5-8259-91961a4d2123 · outbound

This paper cites an unresolved cited work.

FoleyDesigner: Immersive Stereo Foley Generation with Precise Spatio-Temporal Alignment for Film Clips Unresolved cited work

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:52:49.855740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-10T18:57:21.434793Z digest=sha256:75881c04e8a003faa30ffc6e87d8bc2c0919dacaf34cabe5bc0ebf89c4cf3ef5

Observation d9680298-a914-4ec6-9818-10b0e7e84e5c · outbound

This paper cites End-to-End (∼5s).

FoleyDesigner: Immersive Stereo Foley Generation with Precise Spatio-Temporal Alignment for Film Clips End-to-End (∼5s)

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:52:49.850361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-10T18:57:21.434793Z digest=sha256:d9fd79776059f209ddd89435ee459c9bd633066c804c9c94e7b0cbfa9b46ca44

Observation e557e389-c95f-4b3b-8b58-19874800f15a · outbound

This paper cites an unresolved cited work.

FoleyDesigner: Immersive Stereo Foley Generation with Precise Spatio-Temporal Alignment for Film Clips Unresolved cited work

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:52:49.846474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-10T18:57:21.434793Z digest=sha256:1da476bdbf2cafff79dc71011fe8c9f2765232aba4220bbccb588e224b09828e

Observation 0fd13472-f005-43f7-9dfb-14329caa78c0 · outbound

This paper cites Experimental Setup Our user study was conducted through both offline and on- line evaluations to comprehensively assess the perceived quality of generated foley audio.

FoleyDesigner: Immersive Stereo Foley Generation with Precise Spatio-Temporal Alignment for Film Clips Experimental Setup Our user study was conducted through both offline and on- line evaluations to comprehensively assess the perceived quality of generated foley audio

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:52:49.848488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-10T18:57:21.434793Z digest=sha256:f6c41953ec3d8519c0136c6b11170f23ad7985059564e18e8888b783a158cf96

Observation 9ad3407c-c257-46e9-91fc-35a949f66d5c · outbound

This paper cites an unresolved cited work.

FoleyDesigner: Immersive Stereo Foley Generation with Precise Spatio-Temporal Alignment for Film Clips Unresolved cited work

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:52:49.844457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-10T18:57:21.434793Z digest=sha256:8df8c03db0331d1a5677a7dc4e224b6eadf261ef29d21dfe59f04f73d90933ed

Pith citing papers

No inbound Pith citation observations are available.