Pith. sign in

Paper Citation Record · LEDGER

MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation

As of 19 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 1 inbound Pith citation observation for arXiv:2509.06389.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.06389 v1

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T23:46:32.585912Z

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T23:46:32.435054Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-04T23:46:32.951229Z

Reference resolution

35 of 35 outbound references displayed

  • verified exact1
  • verified fuzzy26
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3600fb9b-cce2-4bb8-a054-c5e367bbf70a · outbound

This paper cites MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation.

MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-08-04T23:46:32.956166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-04T23:46:32.435054Z digest=sha256:50f772f9300fb3681b73be936aaf5ee7b1d1cb9753a530fbbfcd369df7bb1a10

Observation 03f5461a-ad39-404e-a58c-f6d106d9c3ad · outbound

This paper cites On top of this backbone, MeanFlow formulation is introduced which directly models the average velocity to enable native one-step generation.

MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation On top of this backbone, MeanFlow formulation is introduced which directly models the average velocity to enable native one-step generation

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:46:33.363058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-04T23:46:32.440419Z digest=sha256:d12b682181d4baa85ee0316879317ce3f9e7ac323e04404ac7ec00b85df82aa2

Observation 3fd227ec-169f-479e-8daf-41354811e4b9 · outbound

This paper cites Multimodal Dataset The proposed MF-MJT is trained on multimodal datasets comprising both audio-video-text and audio-text pairs.

MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation Multimodal Dataset The proposed MF-MJT is trained on multimodal datasets comprising both audio-video-text and audio-text pairs

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:46:33.348985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-04T23:46:32.444989Z digest=sha256:8548d7f502eaab8c77feac5b3ed1654b716db2552540a7d2b9c3e4b4c8a5184c

Observation 5f345c19-dbae-41e8-8956-f0fc65fd41f9 · outbound

This paper cites Comparison with Baselines Table 1 summarizes the performance of the proposed MF-MJT against representative VTA synthesis baselines on the VGGSound test set.

MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation Comparison with Baselines Table 1 summarizes the performance of the proposed MF-MJT against representative VTA synthesis baselines on the VGGSound test set

Reference 4

Resolution
verified exact
raw_fallback, observed 2026-08-04T23:46:32.934088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-04T23:46:32.449742Z digest=sha256:660fe226d6cb882354fa3b2d5ae75dba3ff9cb5f6b11fecfbdc01ecc07ac7184

Observation 5f5c05e2-f170-4adf-a37d-ab118d5c469f · outbound

This paper cites an unresolved cited work.

MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-04T23:46:33.334130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-04T23:46:32.455488Z digest=sha256:ef805b13cceb94cf67ea5bf58eeaff54b8477cc75e56805f933028248217fc04

Observation a67df04d-e494-46b0-a727-aed4387b61e6 · outbound

This paper cites Frieren: Efficient video- to-audio generation network with rectified flow matching,.

MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation Frieren: Efficient video- to-audio generation network with rectified flow matching,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:46:33.319821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-04T23:46:32.459826Z digest=sha256:9b73e29ea2a3cf1ef40b503bf953a9bb6b77bc08fe8df351807d3ece1ce485f9

Observation b2172030-e743-4194-8873-e560323e78b9 · outbound

This paper cites LoV A: Long-form video- to-audio generation,.

MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation LoV A: Long-form video- to-audio generation,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:46:33.304827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-04T23:46:32.464269Z digest=sha256:fbc7dc430278144e5dd58573cc75358f3489d1d9aa64a30ab62fe1d8c836c734

Observation 042929a9-f966-4f2d-a360-4236041ad7a7 · outbound

This paper cites AudioLDM: Text-to-audio generation with latent diffusion models,.

MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation AudioLDM: Text-to-audio generation with latent diffusion models,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:46:33.289701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-04T23:46:32.468413Z digest=sha256:6cb7a36692dbfe8022d18d968cc0f719a659b5e7824369de52d5ee5f147f35e6

Observation 47b97ae1-b526-4cfb-a4a5-55c4653ba99c · outbound

This paper cites ImageBind: One embedding space to bind them all,.

MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation ImageBind: One embedding space to bind them all,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:46:33.274510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-04T23:46:32.472472Z digest=sha256:42bcee35b39fb563708864ddd0a3d2a135d7ea8bc76acaa6822bd2bf33da4600

Observation 84baceec-4072-4255-a76d-e439ee6b142b · outbound

This paper cites Seeing and Hearing: Open- domain visual-audio generation with diffusion latent aligners,.

MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation Seeing and Hearing: Open- domain visual-audio generation with diffusion latent aligners,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:46:33.257882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-04T23:46:32.476656Z digest=sha256:17dc324c76cc25a21cf753c67ad86771e5f638fdc657fc58167731d6f60de423

Observation 0dcbd26a-08fd-4884-adeb-b9529976d844 · outbound

This paper cites FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds.

MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:32.481088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:46:32.481088Z digest=sha256:c0e463e0f1b12bb09943ba34a8f6c681a2197b149a027c89e98536c652ea1012

Observation 6b767ada-39d8-4d89-adaf-751159b082f7 · outbound

This paper cites TA-V2A: Textually assisted video- to-audio generation,.

MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation TA-V2A: Textually assisted video- to-audio generation,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:46:33.242030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-04T23:46:32.485945Z digest=sha256:f266204f31835cce60d1ce48430ca33ea5079d000234333126e976b44f0d4823

Observation 41ae4ce0-e219-4ab1-9cbe-e6fbf2058de3 · outbound

This paper cites MMAudio: Tam- ing multimodal joint training for high-quality video-to-audio synthesis,.

MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation MMAudio: Tam- ing multimodal joint training for high-quality video-to-audio synthesis,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:46:33.227680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-04T23:46:32.489971Z digest=sha256:95017669519827b2190eaead195c62fdd490e5c14ff94aada9bbe9e221057161

Observation 30f5277f-ba65-4eef-bdb6-289bf375842d · outbound

This paper cites Kling-Foley: Multimodal diffusion transformer for high-quality video-to-audio genera- tion,.

MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation Kling-Foley: Multimodal diffusion transformer for high-quality video-to-audio genera- tion,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:32.494196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:46:32.494196Z digest=sha256:c551be55b9d89d77ae645c24b5e44ca5f98f671c38b7890917add29e2e21c7b7

Observation 4e0faab3-cb83-47e5-9dbc-ff112b26a7bd · outbound

This paper cites Denoising diffusion probabilis- tic models,.

MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation Denoising diffusion probabilis- tic models,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:46:33.212191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-04T23:46:32.498259Z digest=sha256:faaa8dd24d92c050b4c85da27cb37025169bc6d9cdb328e7db991fe5a06fc438

Observation 2bb2793f-c683-49cd-bbe6-6032eff702ec · outbound

This paper cites Flow straight and fast: Learn- ing to generate and transfer data with rectified flow,.

MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation Flow straight and fast: Learn- ing to generate and transfer data with rectified flow,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:46:33.197545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-04T23:46:32.502378Z digest=sha256:42e5ba8306a0851820bdf1dcc7aca1f5f9e28e9e259e118f2474d9ad05a5534f

Observation a9688b03-26c1-46ab-b37e-091350255898 · outbound

This paper cites Instaflow: One step is enough for high-quality diffusion-based text-to-image generation,.

MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation Instaflow: One step is enough for high-quality diffusion-based text-to-image generation,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:46:33.183440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-04T23:46:32.506431Z digest=sha256:3dc60da8cbefe17297335eb8a8ffc5b7b2dc108edb0d9b190f4dd01ed470a9dc

Observation 896d1435-b761-40ad-a5be-728234741047 · outbound

This paper cites Mean Flows for One-step Generative Modeling.

MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation Mean Flows for One-step Generative Modeling

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:32.510295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:46:32.510295Z digest=sha256:c47ab9d5600714d1b8a364ce61bbb2289da97f9888e3bd061c81abbe65dd6e49

Observation 1c1debeb-1c45-4715-bfe5-e3a9505e361b · outbound

This paper cites Classifier-Free Diffusion Guidance.

MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation Classifier-Free Diffusion Guidance

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:32.514602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:46:32.514602Z digest=sha256:2d098b56ac94465a4365758eff57c5063d907c3e0308e5f074eb10da404e75f8

Observation 992acc49-5a1e-412d-86bd-554aa35a69ab · outbound

This paper cites CFG-Zero*: Improved Classifier-Free Guidance for Flow Matching Models.

MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation CFG-Zero*: Improved Classifier-Free Guidance for Flow Matching Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:32.518893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:46:32.518893Z digest=sha256:3a3285cdf60740d7872b7062481fc52724ef6ad08471f69df1546c2569746be9

Observation d78d1673-8707-46fe-967c-4e491df61e79 · outbound

This paper cites Scaling rectified flow transformers for high-resolution image synthesis,.

MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation Scaling rectified flow transformers for high-resolution image synthesis,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:46:33.169343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-04T23:46:32.523361Z digest=sha256:642b3e9a6921044ba5ac1f225a0e1bc715a511076ff89f703419821e5c48c7b2

Observation 56f499bf-339a-48f6-a196-eb479603b0a3 · outbound

This paper cites Scalable diffusion models with trans- formers,.

MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation Scalable diffusion models with trans- formers,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:46:33.154757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-04T23:46:32.527401Z digest=sha256:7be624ee7044f3bbce52a8e0efa3409847c9caff4e01210a5d09feeab496c1ff

Observation a1303cec-94e9-40a7-8d4f-5d6781368649 · outbound

This paper cites Learning transfer- able visual models from natural language supervision,.

MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation Learning transfer- able visual models from natural language supervision,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:46:33.140434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-04T23:46:32.531552Z digest=sha256:675b97b37a0dae3ff3fcc7169da58a781beb678d837171bf6baa8d65a51fe27f

Observation 5079f05a-0c9a-4aeb-b6fc-e0f069be909e · outbound

This paper cites A versatile diffusion transformer with mixture of noise levels for audiovisual gen- eration,.

MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation A versatile diffusion transformer with mixture of noise levels for audiovisual gen- eration,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:46:33.126195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-04T23:46:32.535534Z digest=sha256:d79927a6735d1d01e84582ff3b28541cd218eeb366ca95ce7b8de976fb7c0695

Observation 358179d5-ffae-4870-b6d1-f346144c9806 · outbound

This paper cites Synchformer: Efficient syn- chronization from sparse cues,.

MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation Synchformer: Efficient syn- chronization from sparse cues,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:46:33.112032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-04T23:46:32.542615Z digest=sha256:cd6c780350ef0180a161ca46fc0373afa53980cd5981a4bf5a1b0f5e2d64daa0

Observation 72ba48a8-6683-4bfb-93da-3286d88a10e2 · outbound

This paper cites Score-based generative modeling through stochastic differential equations,.

MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation Score-based generative modeling through stochastic differential equations,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:46:33.097241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-04T23:46:32.546543Z digest=sha256:7aab70ab83d0f9df8e2c29c9aaf29a345e1bec54b1a4887bfbbafd5779d3108d

Observation f9c9b6c0-aecb-4e90-bd51-744f0d1fa999 · outbound

This paper cites VGGSound: A large-scale audio-visual dataset,.

MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation VGGSound: A large-scale audio-visual dataset,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:46:33.083256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-04T23:46:32.550612Z digest=sha256:856a64b3db152bdf7caf36f0bd86bf41894310673aaf3f5672109de614f813a5

Observation c7ca1149-03fd-48b4-b82f-a105d0d9e8b4 · outbound

This paper cites AudioCaps: Generating cap- tions for audios in the wild,.

MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation AudioCaps: Generating cap- tions for audios in the wild,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:46:33.067953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-04T23:46:32.555203Z digest=sha256:51d92ec9ea126eeaa4b469434b25b436d87487f0e314e7fa3bb76c023f9b3b2d

Observation ba099949-7db9-43f2-9403-2a65f35fabf5 · outbound

This paper cites WavCaps: A ChatGPT- assisted weakly-labelled audio captioning dataset for audio- language multimodal research,.

MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation WavCaps: A ChatGPT- assisted weakly-labelled audio captioning dataset for audio- language multimodal research,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:46:33.052749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-04T23:46:32.559318Z digest=sha256:f95b59380e6217a5a75c9ed593483ef0ba4f301e1b375331faf03b004057ea71

Observation 11b40170-8da9-4314-948a-d1aadc9f118e · outbound

This paper cites Decoupled Weight Decay Regularization.

MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation Decoupled Weight Decay Regularization

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:32.563516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:46:32.563516Z digest=sha256:ae85c882f757c12e769ce7df8f3a78f79c885522825442671ab48ce6633a3a5a

Observation d42521d5-2ffb-4f22-a5c0-9e33a9c06816 · outbound

This paper cites Audio Set: An ontology and human-labeled dataset for audio events,.

MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation Audio Set: An ontology and human-labeled dataset for audio events,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:46:33.038213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-04T23:46:32.568945Z digest=sha256:983c4f0ae4f418c7ecebb8ffaa17939ef8bcb322c41106137e15fa87a55633e4

Observation b4f63100-b62e-409a-9b3b-bb3b38fa46af · outbound

This paper cites PANNs: Large-scale pre- trained audio neural networks for audio pattern recognition,.

MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation PANNs: Large-scale pre- trained audio neural networks for audio pattern recognition,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:46:33.020868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-04T23:46:32.573072Z digest=sha256:5f15b6a6a9237cd9387f2a0ec72e3efdd426fd54e6790218f0ac5d156be2d6c7

Observation b1ebeab0-251d-4093-bf90-e8ce74829df5 · outbound

This paper cites Efficient train- ing of audio transformers with patchout,.

MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation Efficient train- ing of audio transformers with patchout,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:46:33.005535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-04T23:46:32.577230Z digest=sha256:9646226c9bd7bead3dc6256f750c99c44b4f9ad391a0a10a9dd4e33c7f450fbe

Observation d9a819d7-f750-4095-bbfc-7dc47566a11d · outbound

This paper cites CLAP: Learn- ing audio concepts from natural language supervision,.

MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation CLAP: Learn- ing audio concepts from natural language supervision,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:46:32.989932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-04T23:46:32.581541Z digest=sha256:2f839ae874db18b7e1183911ced855d3173b47e2ec4ab1b69d230f2822d80874

Observation a7287d51-c426-474d-a050-35f35fa93fdf · outbound

This paper cites AudioLCM: Efficient and high-quality text-to-audio generation with minimal inference steps,.

MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation AudioLCM: Efficient and high-quality text-to-audio generation with minimal inference steps,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:46:32.973426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-04T23:46:32.585912Z digest=sha256:8433a6592d49e83c836188e4236298c08bd4148fa327e640f7c41c55c987536e

Pith citing papers

Observation 3600fb9b-cce2-4bb8-a054-c5e367bbf70a · inbound

MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation cites this paper.

MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-08-04T23:46:32.956166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-04T23:46:32.435054Z digest=sha256:50f772f9300fb3681b73be936aaf5ee7b1d1cb9753a530fbbfcd369df7bb1a10