Pith. sign in

Paper Citation Record · LEDGER

Learning from Silence and Noise for Visual Sound Source Localization

As of 7 August 2026, this Paper Citation Record lists 69 of 69 outbound references and 0 inbound Pith citation observations for arXiv:2508.21761.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.21761 v1

Coverage vector

measured 69 of 69 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T14:03:22.242986Z

measured 69 of 69 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

69 of 69 outbound references displayed

  • verified exact6
  • verified fuzzy56
  • unresolved6
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3e4a7fab-55d5-457f-86a4-8f3593837185 · outbound

This paper cites Adobe audition sound effects, 2023.

Learning from Silence and Noise for Visual Sound Source Localization Adobe audition sound effects, 2023

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:34.971984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:03:15.918896Z digest=sha256:9aa581eb35223d3c1bcbd2ccca10a7dabf5b12ce010470b2ed20e18aa2a27eb1

Observation 2487fb81-ca58-4a6e-9bbc-3c103ff6edd9 · outbound

This paper cites Self- supervised learning of audio-visual objects from video.

Learning from Silence and Noise for Visual Sound Source Localization Self- supervised learning of audio-visual objects from video

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:34.794728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:03:16.037998Z digest=sha256:942d418693e5bc72346e0f6abaa838fa75933e2ab349e2a8092ff665871b6211

Observation 167a7fe3-becb-4103-88ac-23a71fd954ee · outbound

This paper cites Look, listen and learn.

Learning from Silence and Noise for Visual Sound Source Localization Look, listen and learn

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:34.604618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:03:16.212958Z digest=sha256:72300c96503f7d0435f501787529cdef99aff79efa0bb30cafc6f3954a910db5

Observation 36d7af9e-35fc-4dd2-a2c6-11d63e7e929c · outbound

This paper cites Objects that sound.

Learning from Silence and Noise for Visual Sound Source Localization Objects that sound

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:34.384136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:03:16.317154Z digest=sha256:579c0c1475ee68c4685d2bce51434e49283833eb3cceb635dcbf1071e0a62bde

Observation 06da2763-a6f0-434b-915e-20afb4a4cf7f · outbound

This paper cites Vggsound: A large-scale audio-visual dataset.

Learning from Silence and Noise for Visual Sound Source Localization Vggsound: A large-scale audio-visual dataset

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:34.185269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:03:16.462096Z digest=sha256:e17555ab8f0c093e7a11a5b5be06c8c3fbd2aa9360e577d1bad7e960717422cb

Observation 8f8dccc0-9e7d-48aa-8820-e48257bbecff · outbound

This paper cites Localizing visual sounds the hard way.

Learning from Silence and Noise for Visual Sound Source Localization Localizing visual sounds the hard way

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:33.877814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:03:16.662025Z digest=sha256:940cb2e7883919870a6d6aa08c3c373976616c01645e686580db08a20f84f23c

Observation ce4379dd-7053-4613-9993-a857258bcb61 · outbound

This paper cites A simple framework for contrastive learning of visual representations.

Learning from Silence and Noise for Visual Sound Source Localization A simple framework for contrastive learning of visual representations

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T14:03:16.836987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:03:16.836987Z digest=sha256:ffd75ef921646e8bd9ad4067c7dbb87088ac8331807273a463f39001802504df

Observation 9612ced4-e39e-4d8f-98e7-88a8bb288676 · outbound

This paper cites Exploring simple siamese representation learning.

Learning from Silence and Noise for Visual Sound Source Localization Exploring simple siamese representation learning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:33.654480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:03:16.989821Z digest=sha256:ee012bea3991f0eadd2cfd5c06ab82784bfc63241c6aec6915ac1988979919e7

Observation 18f1a1bf-66a1-46a6-84d2-e8af7158bfb4 · outbound

This paper cites Integrating Audio, Visual, and Semantic Information for Enhanced Multimodal Speaker Diarization.

Learning from Silence and Noise for Visual Sound Source Localization Integrating Audio, Visual, and Semantic Information for Enhanced Multimodal Speaker Diarization

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-05T14:03:23.762550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:03:17.100791Z digest=sha256:f3e5091bcc4cb879b729ef8d48ee4302dbec8094b573b5e075cc67fb69e129a1

Observation 0a716c15-8ebd-4fd0-88bb-34b2ee5bbf71 · outbound

This paper cites Learning a similarity metric discrimi- natively, with application to face verification.

Learning from Silence and Noise for Visual Sound Source Localization Learning a similarity metric discrimi- natively, with application to face verification

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:33.487670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:03:17.193287Z digest=sha256:6adc9886ab459d52df2a363c977e56aeddcca22c422a69dce5ac6a5fbd59549f

Observation e14f1e04-fc7c-4f20-a678-de813f45fdb2 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

Learning from Silence and Noise for Visual Sound Source Localization Imagenet: A large-scale hierarchical image database

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:33.265198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:03:17.318838Z digest=sha256:2fc0bc75ca420eb038394a278bd9380765ddcb2ad899447d27bbd812b012b5ef

Observation d975b22f-4ca3-4589-a2e4-72ee818513cb · outbound

This paper cites Condi- tional generation of audio from video via foley analogies.

Learning from Silence and Noise for Visual Sound Source Localization Condi- tional generation of audio from video via foley analogies

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:33.105208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:03:17.370379Z digest=sha256:ffef8106859e23a271aaa3c88ab159a004e536f7260b2030329f6d940c2a24f2

Observation 289e305d-b66e-4ba4-bce0-e02e77410ff9 · outbound

This paper cites Audio-Visual Approach For Multimodal Concurrent Speaker Detection.

Learning from Silence and Noise for Visual Sound Source Localization Audio-Visual Approach For Multimodal Concurrent Speaker Detection

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-05T14:03:23.586547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:03:17.449504Z digest=sha256:7ee506a4aaacbf0cc2e17cb9b4116f8c202bf4146180060f881a17cd738e6b24

Observation 4ab82633-9016-4ffb-81d5-3db50d0ec276 · outbound

This paper cites Effect of acoustic scene complexity and visual scene representation on auditory perception in virtual audio-visual environments.

Learning from Silence and Noise for Visual Sound Source Localization Effect of acoustic scene complexity and visual scene representation on auditory perception in virtual audio-visual environments

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:32.961414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:03:17.494864Z digest=sha256:0877bf156274f2dc5fa4d9b85610f8565092fee8138c42fb678b99be294f5c33

Observation 7035228b-f75d-47ad-8bb6-095d71fefd7f · outbound

This paper cites Learning joint sta- tistical models for audio-visual fusion and segregation.

Learning from Silence and Noise for Visual Sound Source Localization Learning joint sta- tistical models for audio-visual fusion and segregation

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:32.769252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:03:17.608532Z digest=sha256:3e47d54606fc3450ec37c033c0120a5a5cd8a3c9abb8548f7ce2a08a88bbaff5

Observation 21b16037-084b-48f8-9727-8994e6d63e22 · outbound

This paper cites Visualvoice: Audio-visual speech separation with cross-modal consistency.

Learning from Silence and Noise for Visual Sound Source Localization Visualvoice: Audio-visual speech separation with cross-modal consistency

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:32.569106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:03:17.717130Z digest=sha256:c846bbc35435a2b99d95c58d9752705f8d9677d1f063088e2173cadd7e34cafb

Observation 7a8dccc9-a42b-4a45-9b98-5a25c45fe29f · outbound

This paper cites Cyclip: Cyclic contrastive language-image pretraining.

Learning from Silence and Noise for Visual Sound Source Localization Cyclip: Cyclic contrastive language-image pretraining

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:32.373662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:03:17.817540Z digest=sha256:6db710d10769d82eb70d8268b43be0da230d6b78a55f3c97666bcdbcd88bc374

Observation 555accc2-bd1e-4443-be62-406557953cf7 · outbound

This paper cites Bootstrap your own latent-a new approach to self-supervised learning.

Learning from Silence and Noise for Visual Sound Source Localization Bootstrap your own latent-a new approach to self-supervised learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T14:03:17.872247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:03:17.872247Z digest=sha256:0f3af2e8e9642a589a847effa6d63dc6f9ae17d0a77d25e74e9d88b14217a1f6

Observation 9117824f-d757-446f-a13a-c5ca6b004c8a · outbound

This paper cites chirp" from the.

Learning from Silence and Noise for Visual Sound Source Localization chirp" from the

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:32.139912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:03:17.970801Z digest=sha256:3e9832102bd3f78bf52bac4cc864658c9458870934f3c1bd37effc45ef984528

Observation c1f857d6-a920-4f0c-8c65-9a2f577b3ec1 · outbound

This paper cites Canonical correlation analysis: An overview with application to learning methods.

Learning from Silence and Noise for Visual Sound Source Localization Canonical correlation analysis: An overview with application to learning methods

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:31.989916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:03:18.046880Z digest=sha256:5bf353feb823ec4166a757870724ab16c23fd2f35503e559212a4b4a0ae5744d

Observation f9ac9d16-c864-4f9a-874c-84270023d2d5 · outbound

This paper cites Deep residual learning for image recognition.

Learning from Silence and Noise for Visual Sound Source Localization Deep residual learning for image recognition

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:31.716044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:03:18.151815Z digest=sha256:3f2271541907a2835e6782741e26b6af53cef1d5e06a7827148a39ca2897f4f1

Observation bdcaa123-6481-48ef-8d32-f6d0a08e3848 · outbound

This paper cites Momentum contrast for unsupervised visual representation learning.

Learning from Silence and Noise for Visual Sound Source Localization Momentum contrast for unsupervised visual representation learning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:31.579954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:03:18.223211Z digest=sha256:30be1a898cabe0ce968431c419ee47adfd93620c44bae1f6053299c0ffd253b8

Observation 1531a6c8-67b1-4cd1-b540-fb97086735e4 · outbound

This paper cites Audio vision: Using audio-visual synchrony to locate sounds.

Learning from Silence and Noise for Visual Sound Source Localization Audio vision: Using audio-visual synchrony to locate sounds

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:31.401708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:03:18.309446Z digest=sha256:1c5ebade9a3dd18c5ddee2a5d53ff5bbe596860af31323bb836fdeb7b272ab43

Observation 2f85e6e0-8b9c-4b19-b9d8-e3c329afbaa8 · outbound

This paper cites Discriminative sounding objects localization via self-supervised audiovisual matching.

Learning from Silence and Noise for Visual Sound Source Localization Discriminative sounding objects localization via self-supervised audiovisual matching

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:31.135238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:03:18.380386Z digest=sha256:95969b252cd4965f96ed261d7f7e0fe1199c46a9c72bea0336085667f36fe799

Observation d41ebcae-edae-4131-a152-077352edea82 · outbound

This paper cites Mix and localize: Localizing sound sources in mixtures.

Learning from Silence and Noise for Visual Sound Source Localization Mix and localize: Localizing sound sources in mixtures

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:30.939933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:03:18.493766Z digest=sha256:8fa42fb8bc53dcbe2eab166a207dd5cddd25daf6a0b2fb9c9a18e10fa572f9a5

Observation f24628e5-4deb-4bd6-a161-2ec939372d7b · outbound

This paper cites You said that?: Synthesising talking faces from audio.

Learning from Silence and Noise for Visual Sound Source Localization You said that?: Synthesising talking faces from audio

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:30.689850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:03:18.591004Z digest=sha256:13079fcae3c9fe4a3987de7da2ea1121b8be3930b3389fc035e05116ba89d01b

Observation fe8c4062-3b87-4099-b58b-ab27916b5b14 · outbound

This paper cites A critical assessment of visual sound source localization models including negative audio.

Learning from Silence and Noise for Visual Sound Source Localization A critical assessment of visual sound source localization models including negative audio

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:30.496216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:03:18.664703Z digest=sha256:af5d1f0e2d1386120d794d560f6092a79b3d89f2d9c9528ba13efc7767796d46

Observation 3c20823b-4950-455a-9f2d-e625d405f647 · outbound

This paper cites Pixels that sound.

Learning from Silence and Noise for Visual Sound Source Localization Pixels that sound

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:30.279438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:03:18.762759Z digest=sha256:5f65dcb86f5ea8e523b815c9e5b863958674253bb2489479266480528e17a5c3

Observation 86a4e8a1-aca1-4a30-a092-b47e19d647af · outbound

This paper cites Learning to visually localize sound sources from mixtures without prior source knowledge.

Learning from Silence and Noise for Visual Sound Source Localization Learning to visually localize sound sources from mixtures without prior source knowledge

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:30.115536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:03:18.830545Z digest=sha256:55c302534f4dbc7c9ea2b4b7c083f1150fe03a6b3334544856cf6f4a0130ad65

Observation f74c8c59-b5b1-4b15-9869-a16ad8230e4a · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Learning from Silence and Noise for Visual Sound Source Localization Adam: A Method for Stochastic Optimization

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T14:03:18.902349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:03:18.902349Z digest=sha256:9a96786b73cd8e3caff07180cf6095200b3a9cce13f08cc664c9006d569294aa

Observation 26e5a637-3c5a-4d02-83c3-de343d8cab76 · outbound

This paper cites Cooperative learning of audio and video models from self-supervised synchronization.

Learning from Silence and Noise for Visual Sound Source Localization Cooperative learning of audio and video models from self-supervised synchronization

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:29.932477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:03:19.044555Z digest=sha256:de6c6b2a22c0e07759764dde10a28c9a6eaf0aa43a4ff79378865d6cd15fe532

Observation c3270f41-9b6a-4e61-a63d-3ae7c5b4fc9e · outbound

This paper cites Recent Advances in Multi-modal 3D Intelligence: A Comprehensive Survey and Evaluation.

Learning from Silence and Noise for Visual Sound Source Localization Recent Advances in Multi-modal 3D Intelligence: A Comprehensive Survey and Evaluation

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-05T14:03:23.396154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:03:19.177165Z digest=sha256:0b8f7dfe6b867b7da3105adb5e959de2b6c65281bf00401ea472d4bfb66eb532

Observation 5ac2cb49-56c7-419a-97c4-702c5577b400 · outbound

This paper cites Do Audio-Visual Segmentation Models Truly Segment Sounding Objects?.

Learning from Silence and Noise for Visual Sound Source Localization Do Audio-Visual Segmentation Models Truly Segment Sounding Objects?

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T14:03:19.216337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:03:19.216337Z digest=sha256:89a5f4be0159cb32d486a82351427b43cc2ddb9d315c9bdc0c008f470ae28e6c

Observation 5726eb08-efbe-48d2-8b9a-57ff590efcdd · outbound

This paper cites Av-nerf: Learning neural fields for real-world audio-visual scene synthesis.

Learning from Silence and Noise for Visual Sound Source Localization Av-nerf: Learning neural fields for real-world audio-visual scene synthesis

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:29.749676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:03:19.287441Z digest=sha256:fb5fb27b2d0386e1dfa0a2d17f316cfcc5802c336eff5def119e70d6f093007e

Observation 545b8489-112a-4a7c-acb4-3bf5445ac3b9 · outbound

This paper cites Exploiting transformation invariance and equivariance for self-supervised sound localisation.

Learning from Silence and Noise for Visual Sound Source Localization Exploiting transformation invariance and equivariance for self-supervised sound localisation

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:29.444007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:03:19.364018Z digest=sha256:d8819c650002f34493c67891864eafe222052f13e050d5508f920e5682ea520d

Observation d6efb336-eb53-4d2b-8eb8-ead267617474 · outbound

This paper cites Visual sound localization in the wild by cross-modal interference erasing.

Learning from Silence and Noise for Visual Sound Source Localization Visual sound localization in the wild by cross-modal interference erasing

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:29.268918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:03:19.430163Z digest=sha256:e51e76f46cac3edc75b46e3ec434bd1126227cd0bc305665ba900035d2180230

Observation dfadd034-3502-4d21-8367-823a8361e709 · outbound

This paper cites Image segmentation using text and image prompts.

Learning from Silence and Noise for Visual Sound Source Localization Image segmentation using text and image prompts

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:29.062912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:03:19.524109Z digest=sha256:746e877251afa85ff8290ec348670cd1f600ca0a24b6c8539ebb3e9106852c45

Observation f3b6196e-4d12-4e59-9b43-ef4eb58d925e · outbound

This paper cites T-vsl: Text-guided visual sound source localization in mixtures.

Learning from Silence and Noise for Visual Sound Source Localization T-vsl: Text-guided visual sound source localization in mixtures

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:28.730076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:03:19.613935Z digest=sha256:95dd16dbb265b0edd78547ff2e44dc25eeda7e552d7f11403978511e440f3ed6

Observation 1b55d64d-b16d-409d-bd20-36164e3dab4a · outbound

This paper cites Localizing visual sounds the easy way.

Learning from Silence and Noise for Visual Sound Source Localization Localizing visual sounds the easy way

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:28.457229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:03:19.694552Z digest=sha256:0c97f8cadda8fe18ddb506642d51126f3b7234082ab1a59d08f68c4c145ec754

Observation 5198aa3c-400b-4ebf-a71e-8d6cee7cc734 · outbound

This paper cites A closer look at weakly-supervised audio-visual source localization.

Learning from Silence and Noise for Visual Sound Source Localization A closer look at weakly-supervised audio-visual source localization

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:28.241052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:03:19.728379Z digest=sha256:f9b765b89541e984e1cbb1264a963108697b22c41e1239f3a65f60a89242d89c

Observation 29abf2f1-28ad-4e4d-9c21-79903fa57789 · outbound

This paper cites Audio-visual grouping network for sound localization from mixtures.

Learning from Silence and Noise for Visual Sound Source Localization Audio-visual grouping network for sound localization from mixtures

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:28.027267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:03:19.819750Z digest=sha256:9f9204283e0b929d95545599492f4e43366b86d951f328059d86b4ad7f28d59c

Observation f6d1524a-ddb1-46a0-8027-bc1a17cdaaee · outbound

This paper cites V ovit: Low latency graph-based audio-visual voice separation transformer.

Learning from Silence and Noise for Visual Sound Source Localization V ovit: Low latency graph-based audio-visual voice separation transformer

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:27.867430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:03:19.877524Z digest=sha256:36cabc1b3f93143b289e470857bbce4bd39cdaaf7f4f88acfc02e32097202a68

Observation 4212176c-8216-4b62-9e64-87aa0313ae39 · outbound

This paper cites Speech inpainting: Context-based speech synthesis guided by video.

Learning from Silence and Noise for Visual Sound Source Localization Speech inpainting: Context-based speech synthesis guided by video

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:27.640497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:03:19.963920Z digest=sha256:7230ff1b7be07cd92dde6671d225f85a9e82f2600c2a874ed8be695d59fa0bf2

Observation 35c1b0bd-8031-4bfc-87ec-decfcf16a85b · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

Learning from Silence and Noise for Visual Sound Source Localization Representation Learning with Contrastive Predictive Coding

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T14:03:20.022923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:03:20.022923Z digest=sha256:045b25eb97bfb3689470f367d2b7d5539bd6d065a197a0418df297e1445d57db

Observation ccb7fd8a-a7cb-4ea2-8729-807d1485a669 · outbound

This paper cites Audio-visual scene analysis with self-supervised multisensory features.

Learning from Silence and Noise for Visual Sound Source Localization Audio-visual scene analysis with self-supervised multisensory features

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:27.447874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:03:20.134878Z digest=sha256:e01a80b830a649fe666955b207f95680cffc20f5c83efc8bfc1b1c6ede8e84f2

Observation 8bff9481-9821-4878-a193-e6c244a8ef4b · outbound

This paper cites Do we need sound for sound source localization? In Asian Conference on Computer Vision, 2020.

Learning from Silence and Noise for Visual Sound Source Localization Do we need sound for sound source localization? In Asian Conference on Computer Vision, 2020

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:27.224197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:03:20.231901Z digest=sha256:e9b30b4a009b109e72cf436b9dcc411d04ed4cb6c04d2deff6b1897451d9990d

Observation 40558691-05dd-42d0-b7bb-56127c08ea6e · outbound

This paper cites Marginnce: Robust sound localization with a negative margin.

Learning from Silence and Noise for Visual Sound Source Localization Marginnce: Robust sound localization with a negative margin

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:27.021849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:03:20.299684Z digest=sha256:53cec735c07658f237da5ab266b0284ccb6d0ebc6928516b6530b946bed98903

Observation b267725d-7f77-42e1-88e6-f3f085c22a92 · outbound

This paper cites Can clip help sound source localization? In IEEE/CVF Winter Conference on Applications of Computer Vision, pages 5711–5720, 2024.

Learning from Silence and Noise for Visual Sound Source Localization Can clip help sound source localization? In IEEE/CVF Winter Conference on Applications of Computer Vision, pages 5711–5720, 2024

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:26.681627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:03:20.368803Z digest=sha256:2468f3e59a1f7a665936f67221a13922f7972a87a0749be2dfe2713f72f5eeb9

Observation 41058908-543f-4c88-a706-d013f88c9a81 · outbound

This paper cites Multiple sound sources localization from coarse to fine.

Learning from Silence and Noise for Visual Sound Source Localization Multiple sound sources localization from coarse to fine

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:26.217672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:03:20.433210Z digest=sha256:4fa716b23654d14bd69cb63b6afc3f0963a915f9bffcd67b9fede3c8f37a1c59

Observation b70c6e83-750c-47ff-bb69-fae5b234ca61 · outbound

This paper cites See the sound, hear the pixels.

Learning from Silence and Noise for Visual Sound Source Localization See the sound, hear the pixels

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:26.068140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:03:20.496314Z digest=sha256:4c6208edf97cee60de672b4a11e7f646294a140d8240fd2cbccdcc6094f07725

Observation c89f614e-64eb-4afe-9485-56ad4c7a6dc3 · outbound

This paper cites Sound source localization.

Learning from Silence and Noise for Visual Sound Source Localization Sound source localization

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:25.968208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:03:20.563495Z digest=sha256:ddbe2c22dbdb14e91cfb3bff107ea38b4548c8640df274b58016d590329f1be2

Observation d1015605-cbf7-49cd-b70f-aeb0766ec065 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Learning from Silence and Noise for Visual Sound Source Localization High-resolution image synthesis with latent diffusion models

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:25.822302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:03:20.665585Z digest=sha256:c351da275939bcbbe33007b368945d88e827486863657c1d9cd8c34c3da0fb36

Observation 7bdcae6c-96c0-4e71-94d6-f754b9f90e0f · outbound

This paper cites Multimodal emotion recognition based on a fusion of audiovi- sual information with temporal dynamics.

Learning from Silence and Noise for Visual Sound Source Localization Multimodal emotion recognition based on a fusion of audiovi- sual information with temporal dynamics

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:25.641943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:03:20.753635Z digest=sha256:3fa9d108ed824fb80f303fd81ce72e264042a9e95bfe748a41bf324b4b9a13c7

Observation 76cf4767-5cb6-4e57-95b5-ec764ceea4c5 · outbound

This paper cites Learn- ing to localize sound source in visual scenes.

Learning from Silence and Noise for Visual Sound Source Localization Learn- ing to localize sound source in visual scenes

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:25.491772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:03:20.867559Z digest=sha256:46b5c2a3e7ce6f20d14ae7923dcbaa59a4c4154f7c3d42d4c0b96e5e051f175d

Observation fa0e209e-fb30-4268-a906-a1141e833c0f · outbound

This paper cites Learn- ing to localize sound sources in visual scenes: Analysis and applications.

Learning from Silence and Noise for Visual Sound Source Localization Learn- ing to localize sound sources in visual scenes: Analysis and applications

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:25.322500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:03:20.939327Z digest=sha256:47dfc81c596c899d617384dda0cd32ed569799e061bd62660c815157d80f0a8f

Observation c0ed37e3-6dda-4b6b-845c-8fd6163ece78 · outbound

This paper cites Learning sound localization better from semantically similar samples.

Learning from Silence and Noise for Visual Sound Source Localization Learning sound localization better from semantically similar samples

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:25.181472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:03:21.045335Z digest=sha256:ce66137d152eb138b100de8b47253aae0e719b312a75d74f7e99bf148bad07eb

Observation c3ee7642-6288-4b1d-a719-e0eea20dadac · outbound

This paper cites Less can be more: Sound source localization with a classification model.

Learning from Silence and Noise for Visual Sound Source Localization Less can be more: Sound source localization with a classification model

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:25.058136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:03:21.119429Z digest=sha256:af73b2d8e55e03effa84465bf96fa621ae4b5d8d8afc265470d00ebb0ffcd10f

Observation a86842d6-bc6f-4693-9f2b-08d3fc350837 · outbound

This paper cites Aligning Sight and Sound: Advanced Sound Source Localization Through Audio-Visual Alignment.

Learning from Silence and Noise for Visual Sound Source Localization Aligning Sight and Sound: Advanced Sound Source Localization Through Audio-Visual Alignment

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-08-05T14:03:23.177348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:03:21.201865Z digest=sha256:3c6365e80db2c7cab35cf07e7ea772458a9a58c15133d8cf6679406c2538c24e

Observation 3c7d422b-5187-4053-8ea8-18ee6b274945 · outbound

This paper cites A Survey on Audio Synthesis and Audio-Visual Multimodal Processing.

Learning from Silence and Noise for Visual Sound Source Localization A Survey on Audio Synthesis and Audio-Visual Multimodal Processing

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-05T14:03:21.327224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:03:21.327224Z digest=sha256:07863fc1a24fa66f7fbe3ebd80b390ef19789136d675e8f9d5a3e334e0ee8e85

Observation c6f32e40-f962-4dbf-a376-12ab955ba608 · outbound

This paper cites En- hancing sound source localization via false negative elimination.

Learning from Silence and Noise for Visual Sound Source Localization En- hancing sound source localization via false negative elimination

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:24.933449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:03:21.398564Z digest=sha256:ae04bb8cfd26818467b1894c31abf64cf964f378d2d2421abdf68f3eca1db0dc

Observation b90020e4-a3a1-4aaf-a769-4fc824b1fef4 · outbound

This paper cites Learning audio-visual source localization via false negative aware contrastive learning.

Learning from Silence and Noise for Visual Sound Source Localization Learning audio-visual source localization via false negative aware contrastive learning

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:24.767706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:03:21.503132Z digest=sha256:3c69997f4a7c5b3b72183886ff881c73bcd9f39b80b02e3870d0d2a2d5518481

Observation 582a0c3a-1945-42fb-b321-9d4ed159de83 · outbound

This paper cites Sound to visual scene generation by audio-to-visual latent alignment.

Learning from Silence and Noise for Visual Sound Source Localization Sound to visual scene generation by audio-to-visual latent alignment

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:24.635351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:03:21.569517Z digest=sha256:cd03b1292179e51526f59aa35cf02785fe38b25aa807d1738a2d6a82b2aa010b

Observation fd6bf942-0d75-4daa-8780-f3e02d52b7b9 · outbound

This paper cites Sound2Vision: Generating Diverse Visuals from Audio through Cross-Modal Latent Alignment.

Learning from Silence and Noise for Visual Sound Source Localization Sound2Vision: Generating Diverse Visuals from Audio through Cross-Modal Latent Alignment

Reference 63

Resolution
verified exact
local_arxiv, observed 2026-08-05T14:03:23.027866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:03:21.674724Z digest=sha256:6e3d665ada4ce9e32761b5ce9c13a5180e4ce859d541a8d76a36d3207a9676d1

Observation d030f199-b42e-4bd6-a4d8-bf6c0aef12c9 · outbound

This paper cites Audio-visual event localization in unconstrained videos.

Learning from Silence and Noise for Visual Sound Source Localization Audio-visual event localization in unconstrained videos

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:24.453977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:03:21.770318Z digest=sha256:5f0311ec7863ed4551707b7b7ce557acd4ccf740cece56509795d86c3ffceabf

Observation ff4b3a39-ebe0-402b-b3e6-63d44945da3b · outbound

This paper cites Phrasecut: Language-based image segmentation in the wild.

Learning from Silence and Noise for Visual Sound Source Localization Phrasecut: Language-based image segmentation in the wild

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:24.279956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:03:21.859711Z digest=sha256:279e575adad41d64f7f410252ccc27c9d941b5b48063fad7e75c2fee33289a91

Observation 913503cf-93d5-4e6b-a859-7063cfbc141a · outbound

This paper cites How to listen? rethinking visual sound localization.

Learning from Silence and Noise for Visual Sound Source Localization How to listen? rethinking visual sound localization

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:24.077187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:03:21.931296Z digest=sha256:5c5020b871dcd6477aa282a5131ecc2938603a49df9f14f84c989d442d29c21b

Observation 36113071-3274-493b-80ff-fc83234719e8 · outbound

This paper cites Acoustic and visual knowledge distillation for contrastive audio-visual localization.

Learning from Silence and Noise for Visual Sound Source Localization Acoustic and visual knowledge distillation for contrastive audio-visual localization

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:23.900431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:03:22.024661Z digest=sha256:2565bd6cc72fee2881ec7e1f92dbd219f5a27575908f8f67f68b7fa8019b8946

Observation 5472dad6-6e1f-401b-a811-31d910bf3686 · outbound

This paper cites Diagnosing and Rectifying Vision Models using Language.

Learning from Silence and Noise for Visual Sound Source Localization Diagnosing and Rectifying Vision Models using Language

Reference 68

Resolution
verified exact
local_arxiv, observed 2026-08-05T14:03:22.863662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:03:22.115862Z digest=sha256:54cc1528a5a0455cabb4098266485de8285add409d8f7cdd56a9f503063f7ea3

Observation 0e808f95-19fd-484e-813b-c4a6e281eee7 · outbound

This paper cites chicken clucking.

Learning from Silence and Noise for Visual Sound Source Localization chicken clucking

Reference 69

Resolution
malformed identifier
raw_fallback, observed 2026-08-05T14:03:22.624081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T14:03:22.242986Z digest=sha256:9bc75302d493273806ddb00cba3a0db4912f471d2ef8ddc1f6d8a5d8880044df

Pith citing papers

No inbound Pith citation observations are available.