Pith. sign in

Paper Citation Record · LEDGER

FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts

As of 6 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 1 inbound Pith citation observation for arXiv:2603.19857.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2603.19857 v2

Coverage vector

measured 62 of 62 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-15T07:35:14.257562Z

measured 63 of 63 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T19:11:43.172296Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-05-10T23:20:54.250212Z

Reference resolution

62 of 62 outbound references displayed

  • verified exact18
  • verified fuzzy38
  • unresolved4
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bf2fa41c-3316-497d-8346-57860a849652 · outbound

This paper cites Black forest labs; frontier ai lab.

FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Black forest labs; frontier ai lab

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T07:40:13.572047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T07:35:14.257562Z digest=sha256:eaae64e66ca1bbe78a98a533185d8355779a817a069dcf5e10f509a2d114f60a

Observation e2ef8972-038b-4a97-8f9a-1dd14857ca62 · outbound

This paper cites Vggsound: A large-scale audio-visual dataset.

FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Vggsound: A large-scale audio-visual dataset

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T07:40:13.541844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T07:35:14.257562Z digest=sha256:ee96d160e89c851b286f652d3aa8e6a96f3b6a1a7383b9e533f105e645d0d038

Observation 3f815c42-d86f-4098-93a9-117b6b171e0d · outbound

This paper cites Video-guided foley sound generation with multimodal con- trols.

FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Video-guided foley sound generation with multimodal con- trols

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T07:40:13.534196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T07:35:14.257562Z digest=sha256:5d11a3d8d7ea53b4f1ce90e164954b22821a45a44edd9414a868b0a9bbec626e

Observation 5ea8d461-33a6-4d6f-a914-25eee82e763d · outbound

This paper cites MMAu- dio: Taming multimodal joint training for high-quality video-to-audio synthesis.

FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts MMAu- dio: Taming multimodal joint training for high-quality video-to-audio synthesis

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T07:40:13.425198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T07:35:14.257562Z digest=sha256:e68d158c2de4ea87b5410a86c6c178ad8afd14dc9b0b7a8641379eb54d26281c

Observation 1b7adcd7-7798-4692-b383-766a874321c2 · outbound

This paper cites Simple and controllable music generation.

FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Simple and controllable music generation

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T07:40:13.468694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T07:35:14.257562Z digest=sha256:4afe4f0334c0a2d8b3db45d1ba42e4b4d3809e4e1ad2f6fba113a4289c135f36

Observation 68f29590-7766-484f-937f-9f915678d1d5 · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:01:21.074684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T07:35:14.257562Z digest=sha256:76e3d9fb061f0f68d4716638972ca0b593876e6e79776de0fc9449c187fc19b9

Observation a5284398-a38e-4d22-a5e1-9b7318301cc6 · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-15T07:39:50.507500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T07:35:14.257562Z digest=sha256:ce0cea47ad18ead1eab98f525630cafd40958067ebf5f5240d96366cc79e8379

Observation 1233b08b-7529-4716-ac5c-2f268e99b882 · outbound

This paper cites CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training.

FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:27:25.722747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T07:35:14.257562Z digest=sha256:bf780313e92ce9bd2c0e50174353b4af3d74720370b069f68ed7c135119c1182

Observation cc2f9143-e0da-45f8-90c0-5b8af801aee5 · outbound

This paper cites Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model.

FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-15T07:39:50.497266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T07:35:14.257562Z digest=sha256:ad13be19bf89a4a0dac96b98760c9ff137b5e1dc01b4601e68c07657674765c0

Observation 2d43860a-3b14-43e1-b7dc-c0ae64728f83 · outbound

This paper cites Imagebind: One embedding space to bind them all.

FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Imagebind: One embedding space to bind them all

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T07:40:13.507977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T07:35:14.257562Z digest=sha256:4252a23235ff1d5b642ba57a600fd903e365839df1acdfbe574292134345abca

Observation bbccbe17-8fe8-477c-824a-ab8db7d2321d · outbound

This paper cites an unresolved cited work.

FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-05-15T07:40:13.464534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T07:35:14.257562Z digest=sha256:2a5276a0df0c24fffd743043337ea97028d18891552acb19110a5a1c597b9131

Observation abb910f8-0020-47b9-8c46-da7aa1a50d51 · outbound

This paper cites Video-to-Audio Generation with Fine-grained Temporal Semantics.

FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Video-to-Audio Generation with Fine-grained Temporal Semantics

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-15T07:39:50.493731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T07:35:14.257562Z digest=sha256:3bb0bbf7af1ecbe6e54a579b66f52500790ef0b60c00788d2fc44c9cb426fa2c

Observation 457c038e-c392-422f-9f65-d8a465ba0383 · outbound

This paper cites Spotlighting partially visible cinematic language for video-to-audio gen- eration via self-distillation.

FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Spotlighting partially visible cinematic language for video-to-audio gen- eration via self-distillation

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T07:40:13.452566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T07:35:14.257562Z digest=sha256:153f77c293a0b42f849aa30775464108368d4e107ee3655a12b81b8537648d26

Observation f7217cf6-c0cf-4606-8373-7a3cb4d4e8ad · outbound

This paper cites Make-An-Audio: Text-To-Audio Generation with Prompt-Enhanced Diffusion Models.

FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Make-An-Audio: Text-To-Audio Generation with Prompt-Enhanced Diffusion Models

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-15T07:39:50.518487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T07:35:14.257562Z digest=sha256:b1be33a24a20c1088d675707e1e639be5fb517153dcaff64883e76c4c83369da

Observation 97d8d566-39f0-4746-9368-5bd523152868 · outbound

This paper cites Tangoflux: Super fast and faithful text to audio generation with flow matching and clap-ranked preference optimization.

FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Tangoflux: Super fast and faithful text to audio generation with flow matching and clap-ranked preference optimization

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T07:40:13.484427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T07:35:14.257562Z digest=sha256:d386e386ac3ef45beed04980206b636898e63637aa9e574357fd2297bca8c34a

Observation ffd7fa4d-89fe-492e-86fd-be23e579153e · outbound

This paper cites Zisserman.

FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Zisserman

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T07:40:13.472192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T07:35:14.257562Z digest=sha256:56c63e862c8a1649305989fdacad6c93838d13ab39e1ff3f51cfe70843777b9e

Observation 35ca8a3f-bd26-44f0-bba6-ff82304f9904 · outbound

This paper cites Audiocaps: Generating captions for audios in the wild.

FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Audiocaps: Generating captions for audios in the wild

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T07:40:13.503752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T07:35:14.257562Z digest=sha256:dae5a09150126e3c6d71f055247a114582cb9e6b541bd47f926d9a4eec15f2e6

Observation b71e4b0b-1717-401d-8050-8c79d3717216 · outbound

This paper cites Plumbley.

FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Plumbley

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T07:40:13.552825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T07:35:14.257562Z digest=sha256:0fa8af7ae8980d6642ba91f83f0389ed48393204ccc9072f58090e8dbecb77e9

Observation 5bd501e2-6f78-458e-9dd2-8c802731a476 · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-15T07:39:50.514212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T07:35:14.257562Z digest=sha256:172098327fc1cfd8dcad945957be55373ac36a6e078b225102547b972b049a4f

Observation bfad09d1-ddc4-438d-8640-35e45deb9fba · outbound

This paper cites Efficient training of audio transformers with patchout.

FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Efficient training of audio transformers with patchout

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T07:40:13.564810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T07:35:14.257562Z digest=sha256:279ebe2c821dcf97759594311df4ca58f89043d14338cb471f9a9165959a44c7

Observation 061b5dad-dfc2-42d9-bc6e-291fe3375a92 · outbound

This paper cites AudioGen: Textually Guided Audio Generation.

FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts AudioGen: Textually Guided Audio Generation

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T07:39:50.489770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T07:35:14.257562Z digest=sha256:057b84e60e6818c40cac40f8546f487080a93a49c36d4bf9c45bb1d728720b97

Observation a159c0f6-1e21-4cc5-acc2-85e9791d5f20 · outbound

This paper cites Video-foley: Two-stage video-to-sound generation via tem- poral event condition for foley sound.IEEE Transactions on Audio, Speech and Language Processing.

FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Video-foley: Two-stage video-to-sound generation via tem- poral event condition for foley sound.IEEE Transactions on Audio, Speech and Language Processing

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T07:40:13.480302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T07:35:14.257562Z digest=sha256:2a160fdaf2c0548c82ac01e002e5f931c532cb0084d6342be0779cdf1276215e

Observation 6217bfbf-1d4d-428b-9467-5285616db589 · outbound

This paper cites Dreamfoley: Scalable vlms for high-fidelity video-to- audio generation.

FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Dreamfoley: Scalable vlms for high-fidelity video-to- audio generation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T07:40:13.518866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T07:35:14.257562Z digest=sha256:3ce54eda0053ab3117f79330164cfa9676313b5d8233c707ebfc194905cce18c

Observation 52e3cd3d-fc9b-450e-9da5-55576307077c · outbound

This paper cites AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks.

FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-15T07:39:50.480308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T07:35:14.257562Z digest=sha256:9960f1bb15811c60f59b79da56ec7c04139ae5f64b1f02bfebb4c532c5770858

Observation 7cd9ed13-ecd7-4dda-b22c-b884fb57acc6 · outbound

This paper cites Imagine and seek: Improv- ing composed image retrieval with an imagined proxy.

FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Imagine and seek: Improv- ing composed image retrieval with an imagined proxy

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T07:40:13.545160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T07:35:14.257562Z digest=sha256:44861046ce223f7ff74852e881e1a6dfe5470c8e88277f4b22060fc1680fbd7e

Observation f3669e5b-77a1-4e2d-8d16-247117093d96 · outbound

This paper cites Audi- oLDM: Text-to-audio generation with latent diffusion mod- els.Proceedings of the International Conference on Machine Learning, pages 21450–21474.

FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Audi- oLDM: Text-to-audio generation with latent diffusion mod- els.Proceedings of the International Conference on Machine Learning, pages 21450–21474

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T07:40:13.560910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T07:35:14.257562Z digest=sha256:b6f585db75f0aa745df156436820b0a26edbb5650dd3ac7f709d49985fa16a13

Observation 1262e0b1-5b49-429a-ad0d-01d01dc80a3b · outbound

This paper cites Plumbley.

FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Plumbley

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T07:40:13.511192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T07:35:14.257562Z digest=sha256:fe30d65ae79ca62d80dcd2a0a3351a5cb2b128d7726ab7c51e8fc1d3a13ea9f7

Observation db8b65f2-ce98-4903-b1ca-a53b327de6e6 · outbound

This paper cites Flashaudio: Rectified flows for fast and high-fidelity text-to-audio generation.

FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Flashaudio: Rectified flows for fast and high-fidelity text-to-audio generation

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T07:40:13.460614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T07:35:14.257562Z digest=sha256:d1445f148be39cbd71d5d81fff880142982e3c53adfeb9122111b90d3cf709e6

Observation a2f6a240-2213-472f-b25b-71904619a732 · outbound

This paper cites Thinksound: Chain-of- thought reasoning in multimodal large language models for audio generation and editing.

FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Thinksound: Chain-of- thought reasoning in multimodal large language models for audio generation and editing

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T07:40:13.437306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T07:35:14.257562Z digest=sha256:602586a7f03e9e594985de595453f1c121ddb44458fcc71f1db04431a80872d1

Observation 186059dd-547d-44dd-934f-04cb85f13f1f · outbound

This paper cites Diff-foley: Synchronized video-to-audio synthesis with la- tent diffusion models.

FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Diff-foley: Synchronized video-to-audio synthesis with la- tent diffusion models

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T07:40:13.579079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T07:35:14.257562Z digest=sha256:250f205373c8e4a07307ac1a14a01fe1e821e8e976432542aa0087bfb7260b51

Observation 789584c7-8d71-4aab-8744-acdce79db322 · outbound

This paper cites Vista-llama: Reducing hallucination in video language models via equal distance to visual tokens.

FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Vista-llama: Reducing hallucination in video language models via equal distance to visual tokens

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T07:40:13.488577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T07:35:14.257562Z digest=sha256:afa7fbb7faa9ad8378ae41742469d26947ecf18c7306d20608f606e6e1220941

Observation 37dafcf1-55e2-40b3-9020-d4fb2539e913 · outbound

This paper cites Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever.

FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T07:40:13.448983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T07:35:14.257562Z digest=sha256:207976177e80669ee4f560be3385824515e89aba2811cc6619835faaa83c57d8

Observation 962cd879-8b12-4bd7-ae9c-07a0e6358d00 · outbound

This paper cites High-resolution image syn- thesis with latent diffusion models.

FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts High-resolution image syn- thesis with latent diffusion models

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T07:40:13.440095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T07:35:14.257562Z digest=sha256:d322422cd68be254e79b4d0aafa393b253fc2e131e63f7d092ef0a52e4eb4f85

Observation 6fd8684f-17c5-4d59-8a0b-26a26dd7b6be · outbound

This paper cites Hunyuanvideo- foley: Multimodal diffusion with representation alignment for high-fidelity foley audio generation.

FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Hunyuanvideo- foley: Multimodal diffusion with representation alignment for high-fidelity foley audio generation

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T07:40:13.434039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T07:35:14.257562Z digest=sha256:b044f6d8956d27a586fde1feb5d0f2efd614bba29722a66fe0919a52914fbcdb

Observation d602ab2c-a437-453f-87a2-85b28799d44d · outbound

This paper cites Seedream 4.0: Toward Next-generation Multimodal Image Generation.

FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Seedream 4.0: Toward Next-generation Multimodal Image Generation

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-15T07:39:50.455277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T07:35:14.257562Z digest=sha256:dcc1ddf99ecdd9e77d9074dcff6fd9e2653c1bd40ab55c15426555f3aeec7f3e

Observation 25ec8201-7247-4d1c-8d7c-56fdeb8c9bc5 · outbound

This paper cites Temporally aligned audio for video with autoregression.

FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Temporally aligned audio for video with autoregression

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T07:40:13.475959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T07:35:14.257562Z digest=sha256:356cb5e03a6687090c7bb7f4050dd458aae5d372404269df9b37946794bc42ae

Observation d96234ef-1141-4efd-bc5d-20b7acf2d289 · outbound

This paper cites Audiobox: Unified audio generation with natural language prompts.

FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Audiobox: Unified audio generation with natural language prompts

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T07:40:13.522610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T07:35:14.257562Z digest=sha256:80c54b83a3ca0b18c2c51cc335e6575040739cbce3f71bc06513267744d7a759

Observation dca90a6c-031a-45e1-8c6e-e1d6efd4e026 · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Wan: Open and Advanced Large-Scale Video Generative Models

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-15T07:39:50.451319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T07:35:14.257562Z digest=sha256:e576acab3344cf986c5269d1be0bc06cce71eb261719cc8bd42d0d629063d63f

Observation b5574116-6960-495e-b3d9-0056cea8d7d1 · outbound

This paper cites Kling-foley: Multimodal diffusion trans- former for high-quality video-to-audio generation.

FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Kling-foley: Multimodal diffusion trans- former for high-quality video-to-audio generation

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T07:40:13.445490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T07:35:14.257562Z digest=sha256:93782805ee2e428aafa8407e0d1998253edb311059f8d1b9badd79781f06ed9f

Observation cc4140ad-b044-46e7-96d7-46e44db93cec · outbound

This paper cites Frieren: Efficient video-to-audio generation network with rectified flow matching.Advances in Neural Information Processing Systems, 37:128118–128138.

FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Frieren: Efficient video-to-audio generation network with rectified flow matching.Advances in Neural Information Processing Systems, 37:128118–128138

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T07:40:13.568458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T07:35:14.257562Z digest=sha256:e4449a032d65ed3971317e1193a2d4d41304590127c7397843828a38c1a20667

Observation f2debca2-c864-42a6-b137-d91ef1b07d80 · outbound

This paper cites OmniGen2: Towards Instruction-Aligned Multimodal Generation.

FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-05-15T07:39:50.485071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T07:35:14.257562Z digest=sha256:7a82c8776686ce153267861dcedcfbe8caf5622bed9b677508e6c5be5cff82f9

Observation 2971fe07-3228-4cfe-b5bb-920697271c66 · outbound

This paper cites Qwen2.5-Omni Technical Report.

FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Qwen2.5-Omni Technical Report

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-05-15T07:39:50.484652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T07:35:14.257562Z digest=sha256:09675af5293f52c13a96c5958c7cdf508372309c8b6c64efebfdee6a13c7a94e

Observation c37a2910-650a-4ecd-a11c-4524546b64d1 · outbound

This paper cites Video-to-audio generation with hidden alignment.

FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Video-to-audio generation with hidden alignment

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T07:40:13.496034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T07:35:14.257562Z digest=sha256:6e40f55fdb13b7592d5cc636adf3207c305c3b0e1d9f31739692860a6ef22d58

Observation 774d5e36-1922-40ba-b902-8ebb08dd0ce5 · outbound

This paper cites Con- textgen: Contextual layout anchoring for identity-consistent multi-instance generation.arXiv preprint arXiv:2510.11000.

FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Con- textgen: Contextual layout anchoring for identity-consistent multi-instance generation.arXiv preprint arXiv:2510.11000

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-15T07:39:50.511077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T07:35:14.257562Z digest=sha256:b090fc58d94c16d44d580b5ad7e7cc34b836c64e8fb90a26806992c971f421d9

Observation d246dbf2-cd28-4dbc-91ce-caf561cb1524 · outbound

This paper cites Towards Weakly Supervised Text-to-Audio Grounding.

FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Towards Weakly Supervised Text-to-Audio Grounding

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-15T07:39:50.464018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T07:35:14.257562Z digest=sha256:e22d2e7fd7a2bd4eac748d255fab076bee4aeb35bb06a5ecac68015332560e97

Observation aef1d58f-46b2-44d7-8072-76d77c14f918 · outbound

This paper cites Diffsound: Discrete diffusion model for text-to-sound generation.IEEE/ACM Transactions on Audio, Speech, and Language Processing, 31:1720–1733.

FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Diffsound: Discrete diffusion model for text-to-sound generation.IEEE/ACM Transactions on Audio, Speech, and Language Processing, 31:1720–1733

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T07:40:13.515099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T07:35:14.257562Z digest=sha256:04eed7eae93e5d6e107b4034a297ec1449cd3c160df6eac66f8e3eba0c484915

Observation ae11fb55-b5b7-4b1c-a873-e175d3004ed4 · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-05-15T07:39:50.468836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T07:35:14.257562Z digest=sha256:b8535a555226ce2d02a51fc4a2491038f4645d1a20470048897efc766f705031

Observation 8e6a628d-a948-48e8-a2da-d39b984a7bb1 · outbound

This paper cites Foley- crafter: Bring silent videos to life with lifelike and synchro- nized sounds.

FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Foley- crafter: Bring silent videos to life with lifelike and synchro- nized sounds

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T07:40:13.526854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T07:35:14.257562Z digest=sha256:a355ea4a0b0902b53ea4b264674c991184ba4f15b35135d94d928d40510565c6

Observation 23e0861c-4e7f-4ffb-9b64-9158d0bfeac1 · outbound

This paper cites Migc++: Advanced multi-instance generation controller for image synthesis.IEEE Transactions on Pattern Analysis and Machine Intelligence.

FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Migc++: Advanced multi-instance generation controller for image synthesis.IEEE Transactions on Pattern Analysis and Machine Intelligence

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T07:40:13.548824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T07:35:14.257562Z digest=sha256:1839c68c915f804ac504791c2856b33d2f5d51ade21261fbb63b249ed774571d

Observation 3bbd6672-5cd1-4839-a0e9-53194717df96 · outbound

This paper cites Migc: Multi-instance generation controller for text-to-image synthesis.

FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Migc: Multi-instance generation controller for text-to-image synthesis

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T07:40:13.538370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T07:35:14.257562Z digest=sha256:dd0d41c6cd652f3109a2ed17cf0fb0e9cf87dc86b33d2d8e6a23e71bd93bed30

Observation be354d4c-014f-4cd5-b12b-2fef546f42df · outbound

This paper cites 3dis: Depth-driven decoupled instance synthesis for text-to-image generation.

FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts 3dis: Depth-driven decoupled instance synthesis for text-to-image generation

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-15T07:39:50.522770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T07:35:14.257562Z digest=sha256:7f4f2cf3fb0c854f1d0094bf8b0781f0936ba4556874466899d9187fab625ba4

Observation 751d62f7-6b40-4bb3-8e95-522f0b2a79f8 · outbound

This paper cites Bidedpo: Conditional image generation with simultaneous text and condition alignment.

FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Bidedpo: Conditional image generation with simultaneous text and condition alignment

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-15T07:39:50.445057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T07:35:14.257562Z digest=sha256:4ac6a15fd9be3c17e224511f44837d674f160a8096c75b76b2936331c56167ac

Observation d4f23135-9a7b-44b6-ad3d-e167deff07f3 · outbound

This paper cites Dreamrenderer: Taming multi-instance attribute control in large-scale text-to-image models.

FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Dreamrenderer: Taming multi-instance attribute control in large-scale text-to-image models

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T07:40:13.442908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T07:35:14.257562Z digest=sha256:09516bf9be6cd6b0696297cf0d22ef2f041b59cb766a42bc8ecc3326ee62a787

Observation 1df1957c-1839-496f-a69a-5e429aa669f1 · outbound

This paper cites 3DIS-FLUX: simple and efficient multi-instance generation with DiT rendering.

FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts 3DIS-FLUX: simple and efficient multi-instance generation with DiT rendering

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-15T07:39:50.500931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T07:35:14.257562Z digest=sha256:8141be44cfee23cd03e5521994cae6332f98cb1232f08f15a4d6abd812754771

Observation f9549235-6119-4bb1-9eae-3a7b6d610caa · outbound

This paper cites Masked audio generation using a single non- autoregressive transformer.

FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Masked audio generation using a single non- autoregressive transformer

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T07:40:13.456310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T07:35:14.257562Z digest=sha256:dde3730306de017e317f0c07e8f9b497c5c0de8670fbe058f7689fa343bc5aac

Observation 27f04722-67b9-453f-9995-336e20b12d4b · outbound

This paper cites an unresolved cited work.

FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-05-15T07:40:13.575457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T07:35:14.257562Z digest=sha256:cd62af40f46cfe24dc46f761ef362343b44beb0c811780250b0be83f55fc0169

Observation 3616e69c-1865-4430-ba7b-925378572b7d · outbound

This paper cites an unresolved cited work.

FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-05-15T07:40:13.530268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T07:35:14.257562Z digest=sha256:79bd0e30ef056020db13a55187cf788398935676623e9b112a642e3012c6b03d

Observation d7660209-014b-458c-974e-3ee463c969c5 · outbound

This paper cites an unresolved cited work.

FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-05-15T07:40:13.499432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T07:35:14.257562Z digest=sha256:68a240a455eaf94318143a77bc68b6080c9cd7153d0b21ea6396b3e10c206efc

Observation 84a277d4-a083-4cd7-902c-1d46bbae8096 · outbound

This paper cites (b) Segment-level Classification You are an audio-visual analysis expert.

FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts (b) Segment-level Classification You are an audio-visual analysis expert

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T07:40:13.556738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T07:35:14.257562Z digest=sha256:4ecfc7ac818da14233ba97794c407743f6b092937e1984f08e6427e2dcfc0d8e

Observation 8c1aac7c-efc7-4fb6-8e05-a6a96844735b · outbound

This paper cites Yes") or absent (.

FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Yes") or absent (

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T07:40:13.427714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T07:35:14.257562Z digest=sha256:2c975d7ada358bddc04e5973f32cbee11d9e484b7130b4b286c33a32ac211786

Observation 52bd301c-bd8f-4fe0-849c-ca34994288a4 · outbound

This paper cites soft", "medium.

FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts soft", "medium

Reference 61

Resolution
malformed identifier
raw_fallback, observed 2026-05-15T07:40:13.492118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T07:35:14.257562Z digest=sha256:b0390b1f16d2cba19ca347c6d0d58dd2f934f125999b199b221ab601a641b638

Observation 3fd4c9ce-1fa9-443e-88a1-6d2b784a805c · outbound

This paper cites During training, we randomly drop TSR features with a probability of 0.1 We set an initial learning rate of2.0×10 −5.

FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts During training, we randomly drop TSR features with a probability of 0.1 We set an initial learning rate of2.0×10 −5

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T07:40:13.430949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T07:35:14.257562Z digest=sha256:77d68a6191a389d823812ffb8a9b9be550fe66bab2ea94bdbc2035e28f0e307f

Pith citing papers

Observation 14285130-a947-437c-b8d6-d0e0c0629475 · inbound

RefineAnything: Multimodal Region-Specific Refinement for Perfect Local Details cites this paper.

RefineAnything: Multimodal Region-Specific Refinement for Perfect Local Details FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts

Reference 24

Resolution
metadata mismatch
local_arxiv, observed 2026-05-10T23:20:54.259356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T19:11:43.172296Z digest=sha256:5fb874666117d963e1d356c70fdbae27feaf8787fe6ace90e40bb7aeba71a244