Pith. sign in

Paper Citation Record · LEDGER

Sounding that Object: Interactive Object-Aware Image to Audio Generation

As of 8 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 0 inbound Pith citation observations for arXiv:2506.04214.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.04214 v1

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:53:18.476414Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

41 of 41 outbound references displayed

  • verified exact0
  • verified fuzzy18
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e79acf52-0658-4e85-9d0d-4c7491b25a18 · outbound

This paper cites S., and Zisserman, A.

Sounding that Object: Interactive Object-Aware Image to Audio Generation S., and Zisserman, A

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:22.979644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:53:14.166587Z digest=sha256:bda26134e711db2a26660280f867f53f003a11288a94d0bad9abe944da7a3fea

Observation 690029ba-d601-4347-8455-815102b30820 · outbound

This paper cites Effect of positional encoding.We assess the impact of positional encoding (PE) on our model’s performance.

Sounding that Object: Interactive Object-Aware Image to Audio Generation Effect of positional encoding.We assess the impact of positional encoding (PE) on our model’s performance

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:19.495273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:53:18.261352Z digest=sha256:1cb80e3c6da8a06d6904a37fc5ed20358b6c82141f132988a3397c29b952411e

Observation f34a7d41-b1ed-4f67-8430-9330dc8265c9 · outbound

This paper cites L., Wu, H.-H., Salamon, J., and Bello, J.

Sounding that Object: Interactive Object-Aware Image to Audio Generation L., Wu, H.-H., Salamon, J., and Bello, J

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T10:53:15.135183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:53:15.135183Z digest=sha256:46af5ffb149d2a3a2bbda4772f509b18a60df46bdedd1580741912d22db941d4

Observation 24645505-7d44-4ab5-ab66-924f51f1f9f0 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Sounding that Object: Interactive Object-Aware Image to Audio Generation BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T10:53:15.258122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:53:15.258122Z digest=sha256:ff7d0af2086adaf09c842e3139130edacd377719dd95e89d830f7a6942520b82

Observation 2ab79958-5864-400b-b386-d8ed413de65b · outbound

This paper cites CLIPSep: Learning Text-queried Sound Separation with Noisy Unlabeled Videos.

Sounding that Object: Interactive Object-Aware Image to Audio Generation CLIPSep: Learning Text-queried Sound Separation with Noisy Unlabeled Videos

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T10:53:15.412519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:53:15.412519Z digest=sha256:570436f07d70dfd7b5611b6a3162b6c37153fcd77ce639e2b36420996ab456d6

Observation b84b5353-d282-4f1d-b3e2-4fb45c42dbfe · outbound

This paper cites We randomly selected 100 samples for evaluation, each rated by 50 unique participants to ensure reliability.

Sounding that Object: Interactive Object-Aware Image to Audio Generation We randomly selected 100 samples for evaluation, each rated by 50 unique participants to ensure reliability

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:19.678359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:53:18.190175Z digest=sha256:e34cec2791dc52d1aa48a25c2c2044733bbd7040686ffd57551899ee9af5ff58

Observation 306f4733-c98d-4a9a-a73c-6f422635d477 · outbound

This paper cites B., and Tor- ralba, A.

Sounding that Object: Interactive Object-Aware Image to Audio Generation B., and Tor- ralba, A

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:22.041748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:53:15.903051Z digest=sha256:765f94b564840ed8a5b95bf984d640bc9edbc57b82d11f4f6dbc411fea5cdaf7

Observation 03ba7b75-97c6-4683-9c90-352d36f6924d · outbound

This paper cites Gotta Hear Them All: Towards Sound Source Aware Audio Generation.

Sounding that Object: Interactive Object-Aware Image to Audio Generation Gotta Hear Them All: Towards Sound Source Aware Audio Generation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T10:53:16.014440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:53:16.014440Z digest=sha256:2b4017aa88458f844da58998eff84cd31afd1874e0a0d53410bd58fcd66585d9

Observation 15194fdf-c0ab-4eaf-b1f4-3b133068b7b8 · outbound

This paper cites Classifier-Free Diffusion Guidance.

Sounding that Object: Interactive Object-Aware Image to Audio Generation Classifier-Free Diffusion Guidance

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T10:53:16.085528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:53:16.085528Z digest=sha256:a2e841d94079d0f5d217c0bfe34d5521ade5a3a3d2675f8164a285c04cc2a949

Observation 352ceb62-d23f-43df-85a7-88a5ea815f5b · outbound

This paper cites D., Kim, B., Lee, H., and Kim, G.

Sounding that Object: Interactive Object-Aware Image to Audio Generation D., Kim, B., Lee, H., and Kim, G

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T10:53:16.285900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:53:16.285900Z digest=sha256:b8351fb3d4049d9d793e45553a16b3f2944f302c3336bf3d97878d54a3aaeebb

Observation 96613bab-25e7-4572-a390-b5f75b4c23af · outbound

This paper cites Hifi-gan: Generative ad- versarial networks for efficient and high fidelity speech synthesis.Advances in Neural Information Processing Systems, 33:17022–17033, 2020a.

Sounding that Object: Interactive Object-Aware Image to Audio Generation Hifi-gan: Generative ad- versarial networks for efficient and high fidelity speech synthesis.Advances in Neural Information Processing Systems, 33:17022–17033, 2020a

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:21.367734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:53:16.567043Z digest=sha256:c6e41bd400d63cf8fe096da7cea46f3dca674db1310559a5beff2de48ac8fbea

Observation 1ae440de-da3b-4066-a714-4e99af60fb63 · outbound

This paper cites Soundini: Sound-Guided Diffusion for Natural Video Editing.

Sounding that Object: Interactive Object-Aware Image to Audio Generation Soundini: Sound-Guided Diffusion for Natural Video Editing

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T10:53:16.636613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:53:16.636613Z digest=sha256:8fe411347368405e976915c31053f326b07e2b5ccef50221d128bfca26a38a42

Observation a6df55b8-a4a8-4bb8-9fb1-8b679d97c704 · outbound

This paper cites Decoupled Weight Decay Regularization.

Sounding that Object: Interactive Object-Aware Image to Audio Generation Decoupled Weight Decay Regularization

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T10:53:16.702062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:53:16.702062Z digest=sha256:3b807b930464acdd4d846545e0ef1c60488e1b2f6abf8b3f6d22a93f0b412963

Observation e044c547-1dba-4001-add8-0f90575d7d5b · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

Sounding that Object: Interactive Object-Aware Image to Audio Generation Representation Learning with Contrastive Predictive Coding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T10:53:16.890076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:53:16.890076Z digest=sha256:4d0e0ca79afba516315ec41a7db8f7c6152dd54d83e6b6f5bf9d215c8d1c540c

Observation fab05a56-1416-46ff-9cb6-7079b351df7b · outbound

This paper cites A., Zhang, R., and Zhu, J.-Y.

Sounding that Object: Interactive Object-Aware Image to Audio Generation A., Zhang, R., and Zhu, J.-Y

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:21.187264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:53:17.029469Z digest=sha256:5f6559abbe2439088d88c2c2003ef352e18393ace593d5cb32d84cd515df48a2

Observation d66c4616-7dd0-47b8-ab14-93dde8116703 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

Sounding that Object: Interactive Object-Aware Image to Audio Generation SAM 2: Segment Anything in Images and Videos

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T10:53:17.123505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:53:17.123505Z digest=sha256:4612eb888e5f3d1861d08293193cac6d82b423e81d0b3b15c16114526ea701cb

Observation d1fc9ea9-2aa6-43f3-a4b8-fe17c98098bb · outbound

This paper cites Self-supervised audio-visual co- segmentation.

Sounding that Object: Interactive Object-Aware Image to Audio Generation Self-supervised audio-visual co- segmentation

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:20.993648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:53:17.231123Z digest=sha256:82e91976e929bfe9b18fe02419e563fe93bc24fbda7b99d23698c51eee1336b0

Observation 88f551ae-306b-4a73-9f74-e5bdc202b4ad · outbound

This paper cites and Adi, Y.

Sounding that Object: Interactive Object-Aware Image to Audio Generation and Adi, Y

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:20.773082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:53:17.324515Z digest=sha256:9611ee9f087f9037fabf62b1c4ba3d54bbf1e95564e622e5ce9375c1287fba27

Observation f18cc79b-1cf5-483d-8215-d94d3a14a6e3 · outbound

This paper cites Denoising Diffusion Implicit Models.

Sounding that Object: Interactive Object-Aware Image to Audio Generation Denoising Diffusion Implicit Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T10:53:17.375318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:53:17.375318Z digest=sha256:80bf4640a03db152b6fb997dde8c2f99c6c817a7777ec87ea58511f0fe5954cf

Observation 924826e6-0172-4273-b1ae-d6e2d83287d7 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Sounding that Object: Interactive Object-Aware Image to Audio Generation Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T10:53:17.445595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:53:17.445595Z digest=sha256:4faea1a9d2be1ace0bde7c93b98362213d00e395e9196c26b10c01541a4c3245

Observation 59579319-00f7-4000-b8af-ca7706245cb1 · outbound

This paper cites Analyzing Multi-Head Self-Attention: Specialized Heads Do the Heavy Lifting, the Rest Can Be Pruned.

Sounding that Object: Interactive Object-Aware Image to Audio Generation Analyzing Multi-Head Self-Attention: Specialized Heads Do the Heavy Lifting, the Rest Can Be Pruned

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T10:53:17.538154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:53:17.538154Z digest=sha256:fd310081f32c54b1c466d95bc7df181635c921bf84d4716d36de4b381c92bc78

Observation eb6f9f6b-836a-4797-b06a-4a3e7a96640c · outbound

This paper cites P., and Salamon, J.

Sounding that Object: Interactive Object-Aware Image to Audio Generation P., and Salamon, J

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:20.586856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:53:17.603114Z digest=sha256:41073d35ca851761d7b5d07eb72594b3a29c0bf1dd1502f4a6c0ac4b8d52b6c9

Observation 428750bf-740f-4dbf-b041-10c95fbbbbac · outbound

This paper cites Auffusion: Leveraging the Power of Diffusion and Large Language Models for Text-to-Audio Generation.

Sounding that Object: Interactive Object-Aware Image to Audio Generation Auffusion: Leveraging the Power of Diffusion and Large Language Models for Text-to-Audio Generation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T10:53:17.684335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:53:17.684335Z digest=sha256:f77ccebddad810cc0b3f55ba1d135c99c914f0f3a61a5b1f2cb7a2def5ef8ced

Observation f491f5a3-11f0-4614-94de-1a71b8fca4f4 · outbound

This paper cites Sound-VECaps: Improving Audio Generation with Visual Enhanced Captions.

Sounding that Object: Interactive Object-Aware Image to Audio Generation Sound-VECaps: Improving Audio Generation with Visual Enhanced Captions

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T10:53:17.755251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:53:17.755251Z digest=sha256:946d325cc89dae9771928476e4d9cb7544e29284617bfac06f85cfd7b3827f74

Observation d226c09d-d252-4134-a793-847f7f69ca39 · outbound

This paper cites FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds.

Sounding that Object: Interactive Object-Aware Image to Audio Generation FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T10:53:17.851996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:53:17.851996Z digest=sha256:e78beb81af7f47563ca7696d9892fa36d355d58c4a27b4642a76f16b7830d7e1

Observation 014a0f98-9611-4d99-9212-b584e52053f6 · outbound

This paper cites Results Video We provide a results video on the project webpage, which showcases our model’s ability to generate sounds based on the masked object prompts.

Sounding that Object: Interactive Object-Aware Image to Audio Generation Results Video We provide a results video on the project webpage, which showcases our model’s ability to generate sounds based on the masked object prompts

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:20.349641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:53:17.911045Z digest=sha256:30d12c042234dd2d5331fedfc5af79bdd63d770b68790f0e87cc1f2a60761fb8

Observation f2e512bb-446e-45db-9b36-ab6e5a2b0c89 · outbound

This paper cites The original dataset comprises 4,616 hours of video clips, each paired with corresponding labels and captions.

Sounding that Object: Interactive Object-Aware Image to Audio Generation The original dataset comprises 4,616 hours of video clips, each paired with corresponding labels and captions

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:20.122666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:53:17.979108Z digest=sha256:eb4dbf37d40e15acf0bdbeb315f34012abcb008cd1fadbd82f3be2b27ac9992b

Observation 89c7e839-6d02-4d91-b3e5-4d8946b2b60b · outbound

This paper cites Speech" and “Music.

Sounding that Object: Interactive Object-Aware Image to Audio Generation Speech" and “Music

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:19.887206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:53:18.061160Z digest=sha256:1ac5d49a9286d89f7186de81de7f4ca69aa9690cdc3306b20edc30ff825ca3b5

Observation c4b27ee5-ee27-4ffb-95ef-e47c2b6e99f3 · outbound

This paper cites By extracting features from both modalities and computing cosine similarity, we show in Table 9 that our method consistently outperforms baselines on this metric.

Sounding that Object: Interactive Object-Aware Image to Audio Generation By extracting features from both modalities and computing cosine similarity, we show in Table 9 that our method consistently outperforms baselines on this metric

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:19.281192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:53:18.389315Z digest=sha256:a038e2eaf3598211f5b939af21ecf15e8246205d795e279db2dc0d0ab2951016

Observation eb734ed4-bcef-4f39-9e5b-731ed4ad6714 · outbound

This paper cites video clips with better audio-visual synchronization, for test- ing.

Sounding that Object: Interactive Object-Aware Image to Audio Generation video clips with better audio-visual synchronization, for test- ing

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:19.118786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:53:18.476414Z digest=sha256:164ded94a9068e39ed22f7fab4bdf31942bcfcb7e6a4afbd76aaf4d5171b85ec

Observation 05235672-3c85-44af-928a-a5ffbf6dde8a · outbound

This paper cites MONet: Unsupervised Scene Decomposition and Representation.

Sounding that Object: Interactive Object-Aware Image to Audio Generation MONet: Unsupervised Scene Decomposition and Representation

Reference 1994

Resolution
unresolved
no resolver link, observed 2026-08-07T10:53:14.410618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:53:14.410618Z digest=sha256:62bfb0843b4dd6309b09ded90968fc324aba5ba507f751495c8ea97d37bab160

Observation 36292d4d-164e-47c1-a47d-b26706e60b98 · outbound

This paper cites Auto-Encoding Variational Bayes.

Sounding that Object: Interactive Object-Aware Image to Audio Generation Auto-Encoding Variational Bayes

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-07T10:53:16.344543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:53:16.344543Z digest=sha256:a940e1061f0af597595e58fecff5af48c4d56be9debbbd72a37d830192c4e304

Observation 10f7d14b-2f4b-4782-bd2e-6c614d482d2d · outbound

This paper cites D., Carr, C., Zukowski, Z., Taylor, J., and Pons, J.

Sounding that Object: Interactive Object-Aware Image to Audio Generation D., Carr, C., Zukowski, Z., Taylor, J., and Pons, J

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-07T10:53:15.739919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:53:15.739919Z digest=sha256:4b212160c4aad35c99575fc2939876dd1682575ea022a3df94327439a2796be2

Observation a35530dd-e4d0-41b2-9fe0-2674cf50081d · outbound

This paper cites Diff-Foley: Synchronized Video-to-Audio Synthesis with Latent Diffusion Models.

Sounding that Object: Interactive Object-Aware Image to Audio Generation Diff-Foley: Synchronized Video-to-Audio Synthesis with Latent Diffusion Models

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-07T10:53:16.772565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:53:16.772565Z digest=sha256:da504d9a85980b329ee8ae351e7a4fac99ed9b8e5a964875a62fe728733e40eb

Observation 9b5802cf-c42f-4ba2-a2ef-a519db155eee · outbound

This paper cites Neural Machine Translation by Jointly Learning to Align and Translate.

Sounding that Object: Interactive Object-Aware Image to Audio Generation Neural Machine Translation by Jointly Learning to Align and Translate

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-07T10:53:14.283844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:53:14.283844Z digest=sha256:8c72b0a2f65df50761d8944e6dc1b476233d1ffe50b4f5e0c21d31e9ca121105

Observation 75e97090-91e3-411d-88d0-3d00fd96cdb5 · outbound

This paper cites Visual acoustic matching.

Sounding that Object: Interactive Object-Aware Image to Audio Generation Visual acoustic matching

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:22.719167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:53:14.542529Z digest=sha256:900fc65636e53689a626df721e6096916dfe6256d9cbe7bb4808b5b64dfa4665

Observation 8ac5f492-375f-4c66-a64e-f9ad9a2bc2a7 · outbound

This paper cites Audio-Visual Synchronisation in the wild.

Sounding that Object: Interactive Object-Aware Image to Audio Generation Audio-Visual Synchronisation in the wild

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T10:53:14.750169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:53:14.750169Z digest=sha256:6fd28371b80e8f8c7da16eedc80773cbcc7227178f4f7da0cab3ef0f9a75bd1f

Observation 92982670-0cc1-4f22-9e56-f5e5bfa7c2e8 · outbound

This paper cites Synch- former: Efficient synchronization from sparse cues.

Sounding that Object: Interactive Object-Aware Image to Audio Generation Synch- former: Efficient synchronization from sparse cues

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:21.806357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:53:16.157720Z digest=sha256:ce27137684eb068449654c4263f0fdc278b187212bd0933a5a859a8649bf62dc

Observation ba9961c9-9d3c-42a4-9dfd-ff2d9a87981d · outbound

This paper cites On uni-modal feature learning in supervised multi-modal learning.

Sounding that Object: Interactive Object-Aware Image to Audio Generation On uni-modal feature learning in supervised multi-modal learning

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:22.306427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:53:15.611374Z digest=sha256:267877287521ffe1321d0c7ba0a71cd7bd110bd1538488b023106a8c828921c3

Observation f344cc3e-447f-4383-8639-dbe578e98bd8 · outbound

This paper cites S., Wiles, O., Moses, Y ., and Zisserman, A.

Sounding that Object: Interactive Object-Aware Image to Audio Generation S., Wiles, O., Moses, Y ., and Zisserman, A

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:53:21.536182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:53:16.461969Z digest=sha256:33efa97a2167726cdd9fde911badfc5d785fa5f7eb5afcbc34ba458c74c26f45

Observation 14e806fc-87e0-48fa-9c31-79ec9b6b96c0 · outbound

This paper cites MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis.

Sounding that Object: Interactive Object-Aware Image to Audio Generation MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T10:53:14.895706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:53:14.895706Z digest=sha256:b0352aa9d01ec8f9916fd831285f7b0039cfd20bebb5ecf269d045f607a2ab25

Pith citing papers

No inbound Pith citation observations are available.