Pith. sign in

Paper Citation Record · LEDGER

Sound Scene Synthesis at the DCASE 2024 Challenge

As of 22 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 3 inbound Pith citation observations for arXiv:2501.08587.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.08587 v1

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T20:27:03.800014Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:24:02.180525Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-10T20:27:04.247482Z

Reference resolution

43 of 43 outbound references displayed

  • verified exact3
  • verified fuzzy23
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 24fce64e-375e-45ed-bee2-ad952755a333 · outbound

This paper cites an unresolved cited work.

Sound Scene Synthesis at the DCASE 2024 Challenge Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:27:04.699769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T20:27:03.583398Z digest=sha256:6a9cb37717e779ce6abe289ca1193cb870cc02101d1331784b9ea7e2e90b32dc

Observation 77ceb030-86f3-478e-9d32-62ed4ca956c6 · outbound

This paper cites This is a more flexible setup than the category- based generation used in the last year [2].

Sound Scene Synthesis at the DCASE 2024 Challenge This is a more flexible setup than the category- based generation used in the last year [2]

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:27:04.682834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T20:27:03.589088Z digest=sha256:46352293cf1768a302adbe58e6c7ef3748d3492f0946a9d172f7333f3402398c

Observation 2f7e9928-baca-466e-bc81-f6ec4e9c8ef6 · outbound

This paper cites Sound Scene Synthesis at the DCASE 2024 Challenge.

Sound Scene Synthesis at the DCASE 2024 Challenge Sound Scene Synthesis at the DCASE 2024 Challenge

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-08-10T20:27:04.253045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T20:27:03.594360Z digest=sha256:3fce4ee3b5f6f92606d5ee7ed30f06499cd51d2910d51959a66367e8b998925e

Observation c19976b5-8a2d-4213-a290-7d4f0a482791 · outbound

This paper cites Objective Evaluation We employed the Fr ´echet Audio Distance (FAD) [6] with PANN- Wavegram-Logmel [7] embeddings as our primary objective metric.

Sound Scene Synthesis at the DCASE 2024 Challenge Objective Evaluation We employed the Fr ´echet Audio Distance (FAD) [6] with PANN- Wavegram-Logmel [7] embeddings as our primary objective metric

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:27:04.665436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T20:27:03.600077Z digest=sha256:fda5def2d61cb910107f6e89147494e1c4cc79d6fe9c4c3d714da8d8fa5f6cce

Observation d022a7f9-bc37-4c29-a1d0-fd48f1a9fa16 · outbound

This paper cites System Performance Table 1 summarizes the evaluation results.

Sound Scene Synthesis at the DCASE 2024 Challenge System Performance Table 1 summarizes the evaluation results

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:27:04.648589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T20:27:03.605674Z digest=sha256:d9be987df82db999d7bfe1b4dc2bb195098211f8ab0586cee67be33b9fd39b04

Observation 01984bcb-c90b-4f23-995f-a26ac8cb3f52 · outbound

This paper cites First, the generative aspect of organizing this challenge has been costly and labor intensive.

Sound Scene Synthesis at the DCASE 2024 Challenge First, the generative aspect of organizing this challenge has been costly and labor intensive

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:27:04.627655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T20:27:03.611429Z digest=sha256:42fe94b67fcd134f59efb87deaa513a758838c8e08ac8cf15f54009ef4d10f13

Observation d6d2267f-0b13-465c-9cca-5ebaddf896d5 · outbound

This paper cites While the submit- ted systems demonstrated promising capabilities, the significant gap between synthetic and reference audio quality indicates substantial room for improvement.

Sound Scene Synthesis at the DCASE 2024 Challenge While the submit- ted systems demonstrated promising capabilities, the significant gap between synthetic and reference audio quality indicates substantial room for improvement

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:27:04.611139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T20:27:03.616819Z digest=sha256:68b7c8091c42919b79105611ee05c1864dac590b8753080905f40f1470b32af3

Observation 896b9a5b-301a-4fa0-bb13-5126a48df30a · outbound

This paper cites an unresolved cited work.

Sound Scene Synthesis at the DCASE 2024 Challenge Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:27:04.594740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T20:27:03.621939Z digest=sha256:b0b593727ff26dbbad1d9c347011f205202af4e11ec3cfb458dcd72e16f6002d

Observation a2bfb407-3084-446d-ae3c-33759162b60a · outbound

This paper cites A Proposal for Foley Sound Synthesis Challenge.

Sound Scene Synthesis at the DCASE 2024 Challenge A Proposal for Foley Sound Synthesis Challenge

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-10T20:27:04.219052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T20:27:03.627264Z digest=sha256:ff355a26e1d339ab36d392445a829a3683531386e14b4de641dbeb82ffd7b94b

Observation c0476e70-cdb1-4b4c-bc49-f867c67df3e0 · outbound

This paper cites Foley sound synthesis at the dcase 2023 challenge,.

Sound Scene Synthesis at the DCASE 2024 Challenge Foley sound synthesis at the dcase 2023 challenge,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:27:04.577548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T20:27:03.633090Z digest=sha256:0cd15b126afeef4a442c893e0c0ff8e54a7187c9e3707cc814281e2d319bfd9a

Observation 3ea3db78-67ed-4cba-86e0-0e8a7cee0414 · outbound

This paper cites AudioLDM: Text-to-Audio Generation with Latent Diffusion Models.

Sound Scene Synthesis at the DCASE 2024 Challenge AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T20:27:03.638263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:27:03.638263Z digest=sha256:c3d4aca421be4c0766f77612418073181f16f1bbf83c8418c29a6142db66fa3d

Observation fdb65355-b319-4204-a4ba-312a42b89132 · outbound

This paper cites Audiocaps: Gen- erating captions for audios in the wild,.

Sound Scene Synthesis at the DCASE 2024 Challenge Audiocaps: Gen- erating captions for audios in the wild,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:27:04.561227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T20:27:03.643353Z digest=sha256:9bd0b01e29b81bddf10db68f2627090ff2820f36b4f29d1bab12e8a1a6d21edc

Observation 6a829246-5d5c-43f3-bdd3-6d6bfaff6bd4 · outbound

This paper cites Audio set: An ontology and human-labeled dataset for audio events,.

Sound Scene Synthesis at the DCASE 2024 Challenge Audio set: An ontology and human-labeled dataset for audio events,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:27:04.541639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T20:27:03.648007Z digest=sha256:bb2e6767f9c554897260bb6654ea2e8f0ef424a03a930a36734cde615499b4cf

Observation b25d28e6-1320-4048-9135-74800fd6d976 · outbound

This paper cites Fr ´echet audio distance: A reference-free metric for evaluating music enhancement algorithms.

Sound Scene Synthesis at the DCASE 2024 Challenge Fr ´echet audio distance: A reference-free metric for evaluating music enhancement algorithms

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:27:04.523789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T20:27:03.652982Z digest=sha256:f1e2833ff974a1c6e00faaa0796773a2d2f629525763ef01e59bc5a2b213ebbf

Observation ae928ba1-6638-4f65-a538-1356c3f45ac3 · outbound

This paper cites Panns: Large-scale pretrained audio neural net- works for audio pattern recognition,.

Sound Scene Synthesis at the DCASE 2024 Challenge Panns: Large-scale pretrained audio neural net- works for audio pattern recognition,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:27:04.506246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T20:27:03.657324Z digest=sha256:b03153ea631b4e1ea6993f5710a9b8a7bf46472920964983e35bd7ded89a8a02

Observation ae6d1d63-b213-49e9-9052-efc718594ebb · outbound

This paper cites Correlation of fr´echet audio dis- tance with human perception of environmental audio is em- bedding dependent,.

Sound Scene Synthesis at the DCASE 2024 Challenge Correlation of fr´echet audio dis- tance with human perception of environmental audio is em- bedding dependent,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:27:04.489672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T20:27:03.661793Z digest=sha256:479c35c02f5826133f217b8401c06ce418e511fe7b6f3a5583a629c1e58ee503

Observation 65f02729-3396-4aa0-8e56-43c264ec4f13 · outbound

This paper cites Sound scene synthesis with audioldm and tango2 for dcase 2024 task7,.

Sound Scene Synthesis at the DCASE 2024 Challenge Sound scene synthesis with audioldm and tango2 for dcase 2024 task7,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:27:04.473310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T20:27:03.666375Z digest=sha256:83723036ec8977c4fcd41106bf4c761c036f009acf1a81d5ae4915e0070d1e40

Observation 5dc700f2-4240-4786-b517-66c3ffcf2c77 · outbound

This paper cites Sound scene synthesis based on gan using contrastive learning and effective time-frequency swap cross attention mechanism,.

Sound Scene Synthesis at the DCASE 2024 Challenge Sound scene synthesis based on gan using contrastive learning and effective time-frequency swap cross attention mechanism,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:27:04.456868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T20:27:03.671176Z digest=sha256:b92b9fb14636ecd96af3d9df9bc0abcd9f3fa57542b72dcb31a77b6480a027d9

Observation 2de08bad-89f7-49d7-bd1e-af93a9bebf24 · outbound

This paper cites Dif- fusion based sound scene synthesis for dcase challenge 2024 task 7,.

Sound Scene Synthesis at the DCASE 2024 Challenge Dif- fusion based sound scene synthesis for dcase challenge 2024 task 7,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:27:04.439364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T20:27:03.676077Z digest=sha256:96fd8e7ccabe001b0bd70393a99fccb8e7051dcf95c1aede65fc21aac0656e30

Observation 51529f4f-4c18-46b9-a575-eae1a0b49120 · outbound

This paper cites Sound scene synthesis based on fine-tuned latent diffusion model for dcase challenge 2024 task 7,.

Sound Scene Synthesis at the DCASE 2024 Challenge Sound scene synthesis based on fine-tuned latent diffusion model for dcase challenge 2024 task 7,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:27:04.421853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T20:27:03.680779Z digest=sha256:2560910c2fbdbd49dcf5c2a87858f685299985b0b447e827bde0c4bda7027e5e

Observation 6aff4a8f-529d-4cb2-b20a-dc6f31b2ab39 · outbound

This paper cites Challenge on sound scene synthesis: Evaluating text-to-audio generation,.

Sound Scene Synthesis at the DCASE 2024 Challenge Challenge on sound scene synthesis: Evaluating text-to-audio generation,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:27:04.405134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T20:27:03.685468Z digest=sha256:0a8f1d4671c7777f8a3a3823449d7a6085f67de54595f6713d16531a10532079

Observation 6bad9403-92a0-4be7-898a-fc683ea926f7 · outbound

This paper cites T-foley: A controllable waveform-domain diffusion model for temporal-event-guided foley sound synthesis,.

Sound Scene Synthesis at the DCASE 2024 Challenge T-foley: A controllable waveform-domain diffusion model for temporal-event-guided foley sound synthesis,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:27:04.388081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T20:27:03.690229Z digest=sha256:96afcf5bd4a456abaad6024fa29c66f5cbefb0b1caa9285933529da90671a437

Observation 9b085ddf-f36b-4e14-8b47-932f52b08be0 · outbound

This paper cites MambaFoley: Foley Sound Generation using Selective State-Space Models.

Sound Scene Synthesis at the DCASE 2024 Challenge MambaFoley: Foley Sound Generation using Selective State-Space Models

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-08-10T20:27:04.168186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T20:27:03.694927Z digest=sha256:3e84bfcb0008371b138e553ca50390ae374f00e77d6d665683196883ac9aaad6

Observation 4792880d-b8f9-4c3e-b55c-e6a4eb4f081a · outbound

This paper cites Audio generation with multiple conditional diffu- sion model,.

Sound Scene Synthesis at the DCASE 2024 Challenge Audio generation with multiple conditional diffu- sion model,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:27:04.372130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T20:27:03.699980Z digest=sha256:5d23fbe02531da7f061242146b9e4ca72e154fee86e3f5917715783c5d9baf9f

Observation c442015e-633a-4f80-b227-07b2e24fd41f · outbound

This paper cites PicoAudio: Enabling Precise Timestamp and Frequency Controllability of Audio Events in Text-to-audio Generation.

Sound Scene Synthesis at the DCASE 2024 Challenge PicoAudio: Enabling Precise Timestamp and Frequency Controllability of Audio Events in Text-to-audio Generation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T20:27:03.704933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:27:03.704933Z digest=sha256:209d65e6626a7240bce211dc3c7500369bc26ed55a7ff73198d22225ffff227d

Observation b8ad09ea-ca13-482f-b2c5-23ecf02bc56f · outbound

This paper cites Audioldm 2: Learn- ing holistic audio generation with self-supervised pretraining,.

Sound Scene Synthesis at the DCASE 2024 Challenge Audioldm 2: Learn- ing holistic audio generation with self-supervised pretraining,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:27:04.353850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T20:27:03.710664Z digest=sha256:4782447f7fc3cb9d6f38dd52c004edc6083c4f6f72e4dec3a5a0ea3aa9f9064f

Observation 310b16f1-67d9-462d-8ce4-46edf18bd4ea · outbound

This paper cites Auffusion: Leveraging the Power of Diffusion and Large Language Models for Text-to-Audio Generation.

Sound Scene Synthesis at the DCASE 2024 Challenge Auffusion: Leveraging the Power of Diffusion and Large Language Models for Text-to-Audio Generation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T20:27:03.716112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:27:03.716112Z digest=sha256:a8a1ecfb39cd5e758f6636261a4159b5ba4e42667921d29427f51f9ed6b4dc8a

Observation 629502ad-7164-4f64-b0bd-b10ca7ec0d24 · outbound

This paper cites EzAudio: Enhancing Text-to-Audio Generation with Efficient Diffusion Transformer.

Sound Scene Synthesis at the DCASE 2024 Challenge EzAudio: Enhancing Text-to-Audio Generation with Efficient Diffusion Transformer

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T20:27:03.721679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:27:03.721679Z digest=sha256:cb7a594a494def1b34c41c36fc461c0b2e4cf900dc0e4e36ebaed76d01a728dd

Observation 6fce6b11-ee5f-457c-ab71-a405aeb3c26d · outbound

This paper cites Fugatto 1: Foundational generative audio transformer opus 1,.

Sound Scene Synthesis at the DCASE 2024 Challenge Fugatto 1: Foundational generative audio transformer opus 1,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:27:04.334553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T20:27:03.727108Z digest=sha256:8ebade746e6bf5d6b6d244f509f43863422b192dba4b86c8929d1c5f782874ab

Observation 603edb4f-0b3c-4112-bda5-bf8f02ab24c6 · outbound

This paper cites Stable Audio Open.

Sound Scene Synthesis at the DCASE 2024 Challenge Stable Audio Open

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T20:27:03.732934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:27:03.732934Z digest=sha256:4f99b57f5da39aac306719086458bdf5840e9eaccd57ba45f6eb734c80376faa

Observation 6257149f-256c-4648-b54d-2de605277656 · outbound

This paper cites Improving Text-To-Audio Models with Synthetic Captions.

Sound Scene Synthesis at the DCASE 2024 Challenge Improving Text-To-Audio Models with Synthetic Captions

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T20:27:03.738918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:27:03.738918Z digest=sha256:a4ca47380275a49338b95080f12b6751044681694f4b0e85bdf1c06e0b925ddb

Observation c0cf9b71-9bd3-40d9-b780-85d5570f34b5 · outbound

This paper cites Syncfusion: Multi- modal onset-synchronized video-to-audio foley synthesis,.

Sound Scene Synthesis at the DCASE 2024 Challenge Syncfusion: Multi- modal onset-synchronized video-to-audio foley synthesis,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:27:04.314130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T20:27:03.744932Z digest=sha256:d58471a4e04c5069579826205e934e5a035f120008e7f6dc62b669cc8908a4b9

Observation 04198b01-4384-4d8c-a34d-6f06b9467c6c · outbound

This paper cites Sonicvisionlm: Play- ing sound with vision language models,.

Sound Scene Synthesis at the DCASE 2024 Challenge Sonicvisionlm: Play- ing sound with vision language models,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:27:04.289175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T20:27:03.749590Z digest=sha256:7706150aa03e27763c6722836e4c2a46d740a63f7420953f2f84a29f7a4e8e36

Observation f75205d1-a832-4e17-a211-7090fc7e4ca9 · outbound

This paper cites Video-foley: Two-stage video-to-sound generation via temporal event condition for fo- ley sound,.

Sound Scene Synthesis at the DCASE 2024 Challenge Video-foley: Two-stage video-to-sound generation via temporal event condition for fo- ley sound,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T20:27:03.754175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:27:03.754175Z digest=sha256:2a6e401c570da91a6b0633dd23888652dc4405be458b1e43e2b860b3c95d6c8c

Observation 0fe1a964-75ef-4ecf-808b-b314da4bc19c · outbound

This paper cites Read, Watch and Scream! Sound Generation from Text and Video.

Sound Scene Synthesis at the DCASE 2024 Challenge Read, Watch and Scream! Sound Generation from Text and Video

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T20:27:03.758711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:27:03.758711Z digest=sha256:28aee607d74735018728dc8efa7823ea9fb8e7c052f6ebf912630de43699eab7

Observation ae0ef7f6-0387-4fa6-b606-d8f6fae36ee0 · outbound

This paper cites Movie Gen: A Cast of Media Foundation Models.

Sound Scene Synthesis at the DCASE 2024 Challenge Movie Gen: A Cast of Media Foundation Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T20:27:03.763247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:27:03.763247Z digest=sha256:670bd9e61695481ab24c63cca37fa376cece7ce304a69c81f58c1eb08c42be71

Observation 2dd8e76a-9b05-4864-af2c-d36c1d7586b2 · outbound

This paper cites Video-Guided Foley Sound Generation with Multimodal Controls.

Sound Scene Synthesis at the DCASE 2024 Challenge Video-Guided Foley Sound Generation with Multimodal Controls

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T20:27:03.768169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:27:03.768169Z digest=sha256:6671ed8dd36c80cd0e63fa1e66d1411bc0ad4253250737000bd40fb248bb9a56

Observation eaa0d4c0-e44a-4942-bc74-9a9d04cb9093 · outbound

This paper cites VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation.

Sound Scene Synthesis at the DCASE 2024 Challenge VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-08-10T20:27:03.921470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T20:27:03.773063Z digest=sha256:3e536487510c038d044e430c74fa679332a993d935c63df10cc4cd0eaaa32988

Observation d656700e-3452-4743-a68e-9de0cdf232db · outbound

This paper cites Frieren: Efficient Video-to-Audio Generation Network with Rectified Flow Matching.

Sound Scene Synthesis at the DCASE 2024 Challenge Frieren: Efficient Video-to-Audio Generation Network with Rectified Flow Matching

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T20:27:03.778565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:27:03.778565Z digest=sha256:40ace654c0cd224738cb00e1ef201222e503f9fb01d8766587725eb15ce3514f

Observation 3f8e29d9-2b0f-4cb6-987b-c9d8ec0fee72 · outbound

This paper cites Masked gener- ative video-to-audio transformers with enhanced synchronic- ity,.

Sound Scene Synthesis at the DCASE 2024 Challenge Masked gener- ative video-to-audio transformers with enhanced synchronic- ity,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:27:04.271468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T20:27:03.783936Z digest=sha256:38903f86d067d271df1c7d72a38f316524311c460adff6a31af6ff835234da6d

Observation 7ca429fd-5320-47ba-add0-d721791b7105 · outbound

This paper cites Temporally Aligned Audio for Video with Autoregression.

Sound Scene Synthesis at the DCASE 2024 Challenge Temporally Aligned Audio for Video with Autoregression

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T20:27:03.788957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:27:03.788957Z digest=sha256:7c75998d1a3eec0aee461e7b58d92823c6b4a2f2ebffac910cf5fdfec9db2f37

Observation e762b05d-a28f-45fe-a7d5-0e7fd3fc0cdf · outbound

This paper cites Gotta Hear Them All: Towards Sound Source Aware Audio Generation.

Sound Scene Synthesis at the DCASE 2024 Challenge Gotta Hear Them All: Towards Sound Source Aware Audio Generation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T20:27:03.794177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:27:03.794177Z digest=sha256:eedc91b9dceda6c408924ec3209804872a27f03c7a459d8ff5a6aba96c56a31c

Observation 5cecb5b1-87b6-46a2-932a-2cc195f06d3d · outbound

This paper cites MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis.

Sound Scene Synthesis at the DCASE 2024 Challenge MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T20:27:03.800014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:27:03.800014Z digest=sha256:d7a322c7de0a8cab2e81f07f0701d2b06199f98e5b99a5fb536027b82139b376

Pith citing papers

Observation 2f7e9928-baca-466e-bc81-f6ec4e9c8ef6 · inbound

Sound Scene Synthesis at the DCASE 2024 Challenge cites this paper.

Sound Scene Synthesis at the DCASE 2024 Challenge Sound Scene Synthesis at the DCASE 2024 Challenge

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-08-10T20:27:04.253045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T20:27:03.594360Z digest=sha256:3fce4ee3b5f6f92606d5ee7ed30f06499cd51d2910d51959a66367e8b998925e

Observation c708edc0-fc08-4182-8948-978c0913df5e · inbound

SoundscapeAgent: Agentic Soundscape Construction for Controllable Synthesis and Scalable Audio-Language Supervision cites this paper.

SoundscapeAgent: Agentic Soundscape Construction for Controllable Synthesis and Scalable Audio-Language Supervision Sound Scene Synthesis at the DCASE 2024 Challenge

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T06:28:46.124972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:28:46.124972Z digest=sha256:fe08d6f8f74d41a0881e196b52ccf82ca9bcfc4cf3d1b3fc1bc1ce586d2c0edc

Observation 58a07c9c-cbe5-4c96-96df-705333699a58 · inbound

Scene2Sound: Auditory-Grounded Soundscape Generation for 3D Gaussian Worlds cites this paper.

Scene2Sound: Auditory-Grounded Soundscape Generation for 3D Gaussian Worlds Sound Scene Synthesis at the DCASE 2024 Challenge

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T15:24:02.180525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:24:02.180525Z digest=sha256:04b656b26cd78c6b564cbbcf68d570b27b22657f43fc2284acb3468a5661bcaf