Pith. sign in

Paper Citation Record · LEDGER

Learning from Silence and Noise for Visual Sound Source Localization

As of 20 August 2026, this Paper Citation Record lists 69 of 69 outbound references and 0 inbound Pith citation observations for arXiv:2508.21761.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.21761 v1

Coverage vector

measured 69 of 69 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T14:03:22.242986Z

measured 69 of 69 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

69 of 69 outbound references displayed

  • verified exact6
  • verified fuzzy56
  • unresolved6
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3e4a7fab-55d5-457f-86a4-8f3593837185 · outbound

This paper cites Adobe audition sound effects, 2023.

Learning from Silence and Noise for Visual Sound Source Localization Adobe audition sound effects, 2023

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:34.971984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T14:03:15.918896Z digest=sha256:3b1559cfc03dc2c632abc77d984e614984227935ba0567a2061db5069d3a1cdc

Observation 2487fb81-ca58-4a6e-9bbc-3c103ff6edd9 · outbound

This paper cites Self- supervised learning of audio-visual objects from video.

Learning from Silence and Noise for Visual Sound Source Localization Self- supervised learning of audio-visual objects from video

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:34.794728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T14:03:16.037998Z digest=sha256:b1b587d319a41352fa7f6526cf216b7cb388c4ac2331f603a314d4e7030f64ab

Observation 167a7fe3-becb-4103-88ac-23a71fd954ee · outbound

This paper cites Look, listen and learn.

Learning from Silence and Noise for Visual Sound Source Localization Look, listen and learn

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:34.604618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T14:03:16.212958Z digest=sha256:baead3cab8e5d4f5e0c0c52b85d1ae266b7f594fc988418b1eb0cb86e1fc249a

Observation 36d7af9e-35fc-4dd2-a2c6-11d63e7e929c · outbound

This paper cites Objects that sound.

Learning from Silence and Noise for Visual Sound Source Localization Objects that sound

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:34.384136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T14:03:16.317154Z digest=sha256:695b2ee778f0a30067de3ed0ba46b4b1eea943b79213061dd046c2339f3678e6

Observation 06da2763-a6f0-434b-915e-20afb4a4cf7f · outbound

This paper cites Vggsound: A large-scale audio-visual dataset.

Learning from Silence and Noise for Visual Sound Source Localization Vggsound: A large-scale audio-visual dataset

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:34.185269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T14:03:16.462096Z digest=sha256:0584c81c84472d779c47aa1949de13bd421dadcd5dda01566dda91f31264938f

Observation 8f8dccc0-9e7d-48aa-8820-e48257bbecff · outbound

This paper cites Localizing visual sounds the hard way.

Learning from Silence and Noise for Visual Sound Source Localization Localizing visual sounds the hard way

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:33.877814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T14:03:16.662025Z digest=sha256:c66ec46c988d9cce453e326fb2417c00d1f76c469ce68ff1f5a3b772c0af0bcb

Observation ce4379dd-7053-4613-9993-a857258bcb61 · outbound

This paper cites A simple framework for contrastive learning of visual representations.

Learning from Silence and Noise for Visual Sound Source Localization A simple framework for contrastive learning of visual representations

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T14:03:16.836987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:03:16.836987Z digest=sha256:62c4b32ad6e11a963cd88404224dcce0b305f1dc9ef3558d4dc7b20536ce4a5a

Observation 9612ced4-e39e-4d8f-98e7-88a8bb288676 · outbound

This paper cites Exploring simple siamese representation learning.

Learning from Silence and Noise for Visual Sound Source Localization Exploring simple siamese representation learning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:33.654480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T14:03:16.989821Z digest=sha256:1b915ed016fd7e1ef8405cd4f55aaa639728e3f200830e601e24b8f6dd50eb5a

Observation 18f1a1bf-66a1-46a6-84d2-e8af7158bfb4 · outbound

This paper cites Integrating Audio, Visual, and Semantic Information for Enhanced Multimodal Speaker Diarization.

Learning from Silence and Noise for Visual Sound Source Localization Integrating Audio, Visual, and Semantic Information for Enhanced Multimodal Speaker Diarization

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-05T14:03:23.762550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T14:03:17.100791Z digest=sha256:8b4e5b67b319e663fd91c200b57fd320cb2dfc4b857079b839abdac835c7524c

Observation 0a716c15-8ebd-4fd0-88bb-34b2ee5bbf71 · outbound

This paper cites Learning a similarity metric discrimi- natively, with application to face verification.

Learning from Silence and Noise for Visual Sound Source Localization Learning a similarity metric discrimi- natively, with application to face verification

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:33.487670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T14:03:17.193287Z digest=sha256:58976ee34944db75dc674317035a7fb6b0fd4ea489e5245b8ebd95b75b539998

Observation e14f1e04-fc7c-4f20-a678-de813f45fdb2 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

Learning from Silence and Noise for Visual Sound Source Localization Imagenet: A large-scale hierarchical image database

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:33.265198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T14:03:17.318838Z digest=sha256:5fc5b89506e82f8e85634a6f4e699961cae25752f0e002534a5e110110e1a9bf

Observation d975b22f-4ca3-4589-a2e4-72ee818513cb · outbound

This paper cites Condi- tional generation of audio from video via foley analogies.

Learning from Silence and Noise for Visual Sound Source Localization Condi- tional generation of audio from video via foley analogies

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:33.105208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T14:03:17.370379Z digest=sha256:134d5fd71ac571ef99628a9e909b8b0370d7e0baccbff5be1fa7bc2fd0402097

Observation 289e305d-b66e-4ba4-bce0-e02e77410ff9 · outbound

This paper cites Audio-Visual Approach For Multimodal Concurrent Speaker Detection.

Learning from Silence and Noise for Visual Sound Source Localization Audio-Visual Approach For Multimodal Concurrent Speaker Detection

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-05T14:03:23.586547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T14:03:17.449504Z digest=sha256:9e61623ed639f598c524c4ddfa32535ee393e0f58aca383d281f0f21728fdf96

Observation 4ab82633-9016-4ffb-81d5-3db50d0ec276 · outbound

This paper cites Effect of acoustic scene complexity and visual scene representation on auditory perception in virtual audio-visual environments.

Learning from Silence and Noise for Visual Sound Source Localization Effect of acoustic scene complexity and visual scene representation on auditory perception in virtual audio-visual environments

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:32.961414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T14:03:17.494864Z digest=sha256:bcf4fc97b1b08d8e76bcfb2b69fb847da1391e9a1bfaad175d7af1292ded1089

Observation 7035228b-f75d-47ad-8bb6-095d71fefd7f · outbound

This paper cites Learning joint sta- tistical models for audio-visual fusion and segregation.

Learning from Silence and Noise for Visual Sound Source Localization Learning joint sta- tistical models for audio-visual fusion and segregation

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:32.769252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T14:03:17.608532Z digest=sha256:a923485fe1f04570d639f0789624816e32b1f5657806d48e366b423c5df620ae

Observation 21b16037-084b-48f8-9727-8994e6d63e22 · outbound

This paper cites Visualvoice: Audio-visual speech separation with cross-modal consistency.

Learning from Silence and Noise for Visual Sound Source Localization Visualvoice: Audio-visual speech separation with cross-modal consistency

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:32.569106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T14:03:17.717130Z digest=sha256:315930268c739c0784432c67ed0d21d4e3e7cc0a769fe90c1b15e62d636d8e25

Observation 7a8dccc9-a42b-4a45-9b98-5a25c45fe29f · outbound

This paper cites Cyclip: Cyclic contrastive language-image pretraining.

Learning from Silence and Noise for Visual Sound Source Localization Cyclip: Cyclic contrastive language-image pretraining

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:32.373662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T14:03:17.817540Z digest=sha256:62f2ee173937bbc7a06180ee9b4150af2bad270a3635f8984461e395e0944880

Observation 555accc2-bd1e-4443-be62-406557953cf7 · outbound

This paper cites Bootstrap your own latent-a new approach to self-supervised learning.

Learning from Silence and Noise for Visual Sound Source Localization Bootstrap your own latent-a new approach to self-supervised learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T14:03:17.872247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:03:17.872247Z digest=sha256:60b44323e89bb6d767b2bc6119780dbc3addc7a575c874016d681e130eb1084b

Observation 9117824f-d757-446f-a13a-c5ca6b004c8a · outbound

This paper cites chirp" from the.

Learning from Silence and Noise for Visual Sound Source Localization chirp" from the

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:32.139912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T14:03:17.970801Z digest=sha256:99cbcc096654edab766d52cdbd1377f0f87b667259e2e9939ce3f4e0ca7ecd39

Observation c1f857d6-a920-4f0c-8c65-9a2f577b3ec1 · outbound

This paper cites Canonical correlation analysis: An overview with application to learning methods.

Learning from Silence and Noise for Visual Sound Source Localization Canonical correlation analysis: An overview with application to learning methods

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:31.989916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T14:03:18.046880Z digest=sha256:638a6f86ab3ebf87f58cfe9920fc0a59a62192c7568f466146941448423c74fe

Observation f9ac9d16-c864-4f9a-874c-84270023d2d5 · outbound

This paper cites Deep residual learning for image recognition.

Learning from Silence and Noise for Visual Sound Source Localization Deep residual learning for image recognition

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:31.716044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T14:03:18.151815Z digest=sha256:1feed85eeb10344836987407f8927f01760c9286089e5bc6b05233b956478668

Observation bdcaa123-6481-48ef-8d32-f6d0a08e3848 · outbound

This paper cites Momentum contrast for unsupervised visual representation learning.

Learning from Silence and Noise for Visual Sound Source Localization Momentum contrast for unsupervised visual representation learning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:31.579954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T14:03:18.223211Z digest=sha256:8df4fbd338e6a53ef2f620b8b747c5875b65ac132b646efad4242e27d18acd1b

Observation 1531a6c8-67b1-4cd1-b540-fb97086735e4 · outbound

This paper cites Audio vision: Using audio-visual synchrony to locate sounds.

Learning from Silence and Noise for Visual Sound Source Localization Audio vision: Using audio-visual synchrony to locate sounds

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:31.401708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T14:03:18.309446Z digest=sha256:737e95e736438edc611dc2f3f45fc55af684b160deb5ca4a3c3e5472a49d4503

Observation 2f85e6e0-8b9c-4b19-b9d8-e3c329afbaa8 · outbound

This paper cites Discriminative sounding objects localization via self-supervised audiovisual matching.

Learning from Silence and Noise for Visual Sound Source Localization Discriminative sounding objects localization via self-supervised audiovisual matching

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:31.135238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T14:03:18.380386Z digest=sha256:a488d68352d92a721edd793f39506fa14969da8b375abaa6ff8c0c1141d937c1

Observation d41ebcae-edae-4131-a152-077352edea82 · outbound

This paper cites Mix and localize: Localizing sound sources in mixtures.

Learning from Silence and Noise for Visual Sound Source Localization Mix and localize: Localizing sound sources in mixtures

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:30.939933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T14:03:18.493766Z digest=sha256:88e5c22afb2ef6bf2c0c00799f7f761c206d6ac2c37007269f98b939571d0093

Observation f24628e5-4deb-4bd6-a161-2ec939372d7b · outbound

This paper cites You said that?: Synthesising talking faces from audio.

Learning from Silence and Noise for Visual Sound Source Localization You said that?: Synthesising talking faces from audio

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:30.689850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T14:03:18.591004Z digest=sha256:299bfd5ff3987df7a501791341fdf3b1b8dbfbedf4f4ee01bfb8fcc05672ce98

Observation fe8c4062-3b87-4099-b58b-ab27916b5b14 · outbound

This paper cites A critical assessment of visual sound source localization models including negative audio.

Learning from Silence and Noise for Visual Sound Source Localization A critical assessment of visual sound source localization models including negative audio

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:30.496216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T14:03:18.664703Z digest=sha256:37c68763e6f69a4993d258d99c1731fef23cb0d45e381dad10cb5803f084b5b6

Observation 3c20823b-4950-455a-9f2d-e625d405f647 · outbound

This paper cites Pixels that sound.

Learning from Silence and Noise for Visual Sound Source Localization Pixels that sound

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:30.279438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T14:03:18.762759Z digest=sha256:180ce1b395a2693bb1fdd633043193265f5dd8295dfec03f1182d3fbf75bed5a

Observation 86a4e8a1-aca1-4a30-a092-b47e19d647af · outbound

This paper cites Learning to visually localize sound sources from mixtures without prior source knowledge.

Learning from Silence and Noise for Visual Sound Source Localization Learning to visually localize sound sources from mixtures without prior source knowledge

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:30.115536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T14:03:18.830545Z digest=sha256:a06fc09a242fc74076a1bfa526ad73a91c8571bb4839315cb37879ce988ef874

Observation f74c8c59-b5b1-4b15-9869-a16ad8230e4a · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Learning from Silence and Noise for Visual Sound Source Localization Adam: A Method for Stochastic Optimization

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T14:03:18.902349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:03:18.902349Z digest=sha256:63795a949ab3f8650ba77409df2aaf22ea7e79abc4c0dcdd78cb26ce74176ebe

Observation 26e5a637-3c5a-4d02-83c3-de343d8cab76 · outbound

This paper cites Cooperative learning of audio and video models from self-supervised synchronization.

Learning from Silence and Noise for Visual Sound Source Localization Cooperative learning of audio and video models from self-supervised synchronization

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:29.932477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T14:03:19.044555Z digest=sha256:1ee3652f26a7fb97ada7ca7ba02a750956d09dfb4fd18265ca7d923c3df10b23

Observation c3270f41-9b6a-4e61-a63d-3ae7c5b4fc9e · outbound

This paper cites Recent Advances in Multi-modal 3D Intelligence: A Comprehensive Survey and Evaluation.

Learning from Silence and Noise for Visual Sound Source Localization Recent Advances in Multi-modal 3D Intelligence: A Comprehensive Survey and Evaluation

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-05T14:03:23.396154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T14:03:19.177165Z digest=sha256:6928499882c35daed5b9bb1765b721d2df095d93f2d9cf806b155b80fe99403a

Observation 5ac2cb49-56c7-419a-97c4-702c5577b400 · outbound

This paper cites Do Audio-Visual Segmentation Models Truly Segment Sounding Objects?.

Learning from Silence and Noise for Visual Sound Source Localization Do Audio-Visual Segmentation Models Truly Segment Sounding Objects?

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T14:03:19.216337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:03:19.216337Z digest=sha256:314b9c8974cecd11bdaf9d2fb0eba06314d534596b4ff4adc2f9d4f4d24a11c9

Observation 5726eb08-efbe-48d2-8b9a-57ff590efcdd · outbound

This paper cites Av-nerf: Learning neural fields for real-world audio-visual scene synthesis.

Learning from Silence and Noise for Visual Sound Source Localization Av-nerf: Learning neural fields for real-world audio-visual scene synthesis

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:29.749676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T14:03:19.287441Z digest=sha256:e6591e29aed364b3ca68b55572cec03da6ea455b086e68a850a6bda25bfbdda6

Observation 545b8489-112a-4a7c-acb4-3bf5445ac3b9 · outbound

This paper cites Exploiting transformation invariance and equivariance for self-supervised sound localisation.

Learning from Silence and Noise for Visual Sound Source Localization Exploiting transformation invariance and equivariance for self-supervised sound localisation

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:29.444007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T14:03:19.364018Z digest=sha256:c870da83aa710cf3b213638a238a8be2d0dab401d3d7bf014ee7a8b506d29993

Observation d6efb336-eb53-4d2b-8eb8-ead267617474 · outbound

This paper cites Visual sound localization in the wild by cross-modal interference erasing.

Learning from Silence and Noise for Visual Sound Source Localization Visual sound localization in the wild by cross-modal interference erasing

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:29.268918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T14:03:19.430163Z digest=sha256:c5161599238d4b6c504df6c66c28508804c330f2070356375546698104681fdd

Observation dfadd034-3502-4d21-8367-823a8361e709 · outbound

This paper cites Image segmentation using text and image prompts.

Learning from Silence and Noise for Visual Sound Source Localization Image segmentation using text and image prompts

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:29.062912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T14:03:19.524109Z digest=sha256:9606d1791ade809e9dbdb3327a7addf145ab8a218ad81799c33dda3f6de0abb3

Observation f3b6196e-4d12-4e59-9b43-ef4eb58d925e · outbound

This paper cites T-vsl: Text-guided visual sound source localization in mixtures.

Learning from Silence and Noise for Visual Sound Source Localization T-vsl: Text-guided visual sound source localization in mixtures

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:28.730076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T14:03:19.613935Z digest=sha256:1538860a0069114c723000735f12795cdeae3c30793b1c5a629bb2185da1b1bd

Observation 1b55d64d-b16d-409d-bd20-36164e3dab4a · outbound

This paper cites Localizing visual sounds the easy way.

Learning from Silence and Noise for Visual Sound Source Localization Localizing visual sounds the easy way

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:28.457229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T14:03:19.694552Z digest=sha256:31a954f868818d9b965e2c4ba7787055107c44b8264187b7360860c80535ae70

Observation 5198aa3c-400b-4ebf-a71e-8d6cee7cc734 · outbound

This paper cites A closer look at weakly-supervised audio-visual source localization.

Learning from Silence and Noise for Visual Sound Source Localization A closer look at weakly-supervised audio-visual source localization

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:28.241052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T14:03:19.728379Z digest=sha256:be552a98b8d8d08d566515861a7a84a06670e8ff8ab77059b0cae4991978a10c

Observation 29abf2f1-28ad-4e4d-9c21-79903fa57789 · outbound

This paper cites Audio-visual grouping network for sound localization from mixtures.

Learning from Silence and Noise for Visual Sound Source Localization Audio-visual grouping network for sound localization from mixtures

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:28.027267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T14:03:19.819750Z digest=sha256:f5c9f5ddcaa26daf7b5d3f867f9161f848d1144495a4f7877d2b6cde310d6708

Observation f6d1524a-ddb1-46a0-8027-bc1a17cdaaee · outbound

This paper cites V ovit: Low latency graph-based audio-visual voice separation transformer.

Learning from Silence and Noise for Visual Sound Source Localization V ovit: Low latency graph-based audio-visual voice separation transformer

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:27.867430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T14:03:19.877524Z digest=sha256:e1e74965892f99535445ef79c94e95093756ad659664151880e93d525de2ebc3

Observation 4212176c-8216-4b62-9e64-87aa0313ae39 · outbound

This paper cites Speech inpainting: Context-based speech synthesis guided by video.

Learning from Silence and Noise for Visual Sound Source Localization Speech inpainting: Context-based speech synthesis guided by video

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:27.640497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T14:03:19.963920Z digest=sha256:9851f43d945d7803fbfdafa11b17de5bf32b2c971292b8904253f2e2daff19a2

Observation 35c1b0bd-8031-4bfc-87ec-decfcf16a85b · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

Learning from Silence and Noise for Visual Sound Source Localization Representation Learning with Contrastive Predictive Coding

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T14:03:20.022923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:03:20.022923Z digest=sha256:b8f3bb65712e6bb7417237653ed2905ab29fe4aae8c57e5b1f4154cf0df10d70

Observation ccb7fd8a-a7cb-4ea2-8729-807d1485a669 · outbound

This paper cites Audio-visual scene analysis with self-supervised multisensory features.

Learning from Silence and Noise for Visual Sound Source Localization Audio-visual scene analysis with self-supervised multisensory features

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:27.447874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T14:03:20.134878Z digest=sha256:3e32a625b2f63d8f01cd7085a503be961239b92bf8009cc978f46e068547a1ff

Observation 8bff9481-9821-4878-a193-e6c244a8ef4b · outbound

This paper cites Do we need sound for sound source localization? In Asian Conference on Computer Vision, 2020.

Learning from Silence and Noise for Visual Sound Source Localization Do we need sound for sound source localization? In Asian Conference on Computer Vision, 2020

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:27.224197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T14:03:20.231901Z digest=sha256:c63a44383e994bba7aeff0201af8200eec762463b1b4b149c34ceb9db3373432

Observation 40558691-05dd-42d0-b7bb-56127c08ea6e · outbound

This paper cites Marginnce: Robust sound localization with a negative margin.

Learning from Silence and Noise for Visual Sound Source Localization Marginnce: Robust sound localization with a negative margin

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:27.021849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T14:03:20.299684Z digest=sha256:1ca0ad56dd4968be061f516ea471e1d4e266d5ed0a229ebc874bafc62e8eefb9

Observation b267725d-7f77-42e1-88e6-f3f085c22a92 · outbound

This paper cites Can clip help sound source localization? In IEEE/CVF Winter Conference on Applications of Computer Vision, pages 5711–5720, 2024.

Learning from Silence and Noise for Visual Sound Source Localization Can clip help sound source localization? In IEEE/CVF Winter Conference on Applications of Computer Vision, pages 5711–5720, 2024

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:26.681627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T14:03:20.368803Z digest=sha256:d93b2923ea45f1986701ed782d33e33e63740ac054545bd2eb2999abef355606

Observation 41058908-543f-4c88-a706-d013f88c9a81 · outbound

This paper cites Multiple sound sources localization from coarse to fine.

Learning from Silence and Noise for Visual Sound Source Localization Multiple sound sources localization from coarse to fine

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:26.217672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T14:03:20.433210Z digest=sha256:ed4989f5e94a906e66257d1f69bb9352b3add3b411d6373b151a723bd37f4441

Observation b70c6e83-750c-47ff-bb69-fae5b234ca61 · outbound

This paper cites See the sound, hear the pixels.

Learning from Silence and Noise for Visual Sound Source Localization See the sound, hear the pixels

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:26.068140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T14:03:20.496314Z digest=sha256:6b837a83c4dc9d536651bd49e07a94a3b3ca048e9fe8a550329cd70ff1d45fd0

Observation c89f614e-64eb-4afe-9485-56ad4c7a6dc3 · outbound

This paper cites Sound source localization.

Learning from Silence and Noise for Visual Sound Source Localization Sound source localization

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:25.968208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T14:03:20.563495Z digest=sha256:fa0b55eeef9be2373553184ef8b2e65df0143ab145232865c3786cce13b21a61

Observation d1015605-cbf7-49cd-b70f-aeb0766ec065 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Learning from Silence and Noise for Visual Sound Source Localization High-resolution image synthesis with latent diffusion models

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:25.822302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T14:03:20.665585Z digest=sha256:8e3ac297050d8db7ef1fdb0998926240a5d4137de46fe1d29e9e97fe202870d4

Observation 7bdcae6c-96c0-4e71-94d6-f754b9f90e0f · outbound

This paper cites Multimodal emotion recognition based on a fusion of audiovi- sual information with temporal dynamics.

Learning from Silence and Noise for Visual Sound Source Localization Multimodal emotion recognition based on a fusion of audiovi- sual information with temporal dynamics

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:25.641943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T14:03:20.753635Z digest=sha256:3df763009c33790e14019d3f9de14a47125bff8d40d99a8d5615501f7ada400d

Observation 76cf4767-5cb6-4e57-95b5-ec764ceea4c5 · outbound

This paper cites Learn- ing to localize sound source in visual scenes.

Learning from Silence and Noise for Visual Sound Source Localization Learn- ing to localize sound source in visual scenes

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:25.491772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T14:03:20.867559Z digest=sha256:2f3d9049aa8af49cf9b36b51f727a464f37b75ee95e19d3005383f3575ffede0

Observation fa0e209e-fb30-4268-a906-a1141e833c0f · outbound

This paper cites Learn- ing to localize sound sources in visual scenes: Analysis and applications.

Learning from Silence and Noise for Visual Sound Source Localization Learn- ing to localize sound sources in visual scenes: Analysis and applications

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:25.322500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T14:03:20.939327Z digest=sha256:3bcefe50af1f645369543eb281f892be575b1de499978f3ef6e2577754fcf2ce

Observation c0ed37e3-6dda-4b6b-845c-8fd6163ece78 · outbound

This paper cites Learning sound localization better from semantically similar samples.

Learning from Silence and Noise for Visual Sound Source Localization Learning sound localization better from semantically similar samples

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:25.181472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T14:03:21.045335Z digest=sha256:369f6eee8ef9c49b9855086fa049974ea4f40b031ca665179c391b7a352cf6a7

Observation c3ee7642-6288-4b1d-a719-e0eea20dadac · outbound

This paper cites Less can be more: Sound source localization with a classification model.

Learning from Silence and Noise for Visual Sound Source Localization Less can be more: Sound source localization with a classification model

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:25.058136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T14:03:21.119429Z digest=sha256:52cf0894f3c6b29435fd5343e980ec9d956f550a71f775e1790438f8992b8850

Observation a86842d6-bc6f-4693-9f2b-08d3fc350837 · outbound

This paper cites Aligning Sight and Sound: Advanced Sound Source Localization Through Audio-Visual Alignment.

Learning from Silence and Noise for Visual Sound Source Localization Aligning Sight and Sound: Advanced Sound Source Localization Through Audio-Visual Alignment

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-08-05T14:03:23.177348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T14:03:21.201865Z digest=sha256:51bfa74c65fe01f66ba3031880bff8e5f4fd0aea99bc2b14a5702a687e7d6bc2

Observation 3c7d422b-5187-4053-8ea8-18ee6b274945 · outbound

This paper cites A Survey on Audio Synthesis and Audio-Visual Multimodal Processing.

Learning from Silence and Noise for Visual Sound Source Localization A Survey on Audio Synthesis and Audio-Visual Multimodal Processing

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-05T14:03:21.327224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:03:21.327224Z digest=sha256:d629a20acc455237f7ea9e58b3e21e4d18067855ddb5fc1bfef7809948de26e6

Observation c6f32e40-f962-4dbf-a376-12ab955ba608 · outbound

This paper cites En- hancing sound source localization via false negative elimination.

Learning from Silence and Noise for Visual Sound Source Localization En- hancing sound source localization via false negative elimination

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:24.933449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T14:03:21.398564Z digest=sha256:3c5f6918b98f58e12e6fd70a07e47bfd3a4137cca5e3c63cd26dd9e453327415

Observation b90020e4-a3a1-4aaf-a769-4fc824b1fef4 · outbound

This paper cites Learning audio-visual source localization via false negative aware contrastive learning.

Learning from Silence and Noise for Visual Sound Source Localization Learning audio-visual source localization via false negative aware contrastive learning

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:24.767706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T14:03:21.503132Z digest=sha256:c52a180c8f6f3472ab96ddff09c644b698652dadc0fa559e017fa86062d8929c

Observation 582a0c3a-1945-42fb-b321-9d4ed159de83 · outbound

This paper cites Sound to visual scene generation by audio-to-visual latent alignment.

Learning from Silence and Noise for Visual Sound Source Localization Sound to visual scene generation by audio-to-visual latent alignment

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:24.635351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T14:03:21.569517Z digest=sha256:713ab4152ef530002ee4b75c00fafd7afbcc81f4e58ee2043ded7d4f67de6106

Observation fd6bf942-0d75-4daa-8780-f3e02d52b7b9 · outbound

This paper cites Sound2Vision: Generating Diverse Visuals from Audio through Cross-Modal Latent Alignment.

Learning from Silence and Noise for Visual Sound Source Localization Sound2Vision: Generating Diverse Visuals from Audio through Cross-Modal Latent Alignment

Reference 63

Resolution
verified exact
local_arxiv, observed 2026-08-05T14:03:23.027866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T14:03:21.674724Z digest=sha256:d71072b877bbf7fb08a1041fb140542a688500e794ab2167cb71690ec4dd65f5

Observation d030f199-b42e-4bd6-a4d8-bf6c0aef12c9 · outbound

This paper cites Audio-visual event localization in unconstrained videos.

Learning from Silence and Noise for Visual Sound Source Localization Audio-visual event localization in unconstrained videos

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:24.453977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T14:03:21.770318Z digest=sha256:fbcf6659bb2501468e27aea8c0f7b36e15f0616e5d8504737605f9beeb55d4e3

Observation ff4b3a39-ebe0-402b-b3e6-63d44945da3b · outbound

This paper cites Phrasecut: Language-based image segmentation in the wild.

Learning from Silence and Noise for Visual Sound Source Localization Phrasecut: Language-based image segmentation in the wild

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:24.279956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T14:03:21.859711Z digest=sha256:70f11c5e4467403032d8662a1283f551a5953c179be54dde674d7f374eda1e35

Observation 913503cf-93d5-4e6b-a859-7063cfbc141a · outbound

This paper cites How to listen? rethinking visual sound localization.

Learning from Silence and Noise for Visual Sound Source Localization How to listen? rethinking visual sound localization

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:24.077187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T14:03:21.931296Z digest=sha256:5d7c37ca665b66a38664e01eb000eeec21971dafdcb4f95fecfe742b9ec15bf3

Observation 36113071-3274-493b-80ff-fc83234719e8 · outbound

This paper cites Acoustic and visual knowledge distillation for contrastive audio-visual localization.

Learning from Silence and Noise for Visual Sound Source Localization Acoustic and visual knowledge distillation for contrastive audio-visual localization

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:23.900431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T14:03:22.024661Z digest=sha256:c99f0a5ba50d82179bc1b29a086518a96132917bc70fc524cca0710e4ca4323b

Observation 5472dad6-6e1f-401b-a811-31d910bf3686 · outbound

This paper cites Diagnosing and Rectifying Vision Models using Language.

Learning from Silence and Noise for Visual Sound Source Localization Diagnosing and Rectifying Vision Models using Language

Reference 68

Resolution
verified exact
local_arxiv, observed 2026-08-05T14:03:22.863662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T14:03:22.115862Z digest=sha256:f5583d403b97ff54a51946670d5e3266de016f55356ed091edb97a49e30f3d26

Observation 0e808f95-19fd-484e-813b-c4a6e281eee7 · outbound

This paper cites chicken clucking.

Learning from Silence and Noise for Visual Sound Source Localization chicken clucking

Reference 69

Resolution
malformed identifier
raw_fallback, observed 2026-08-05T14:03:22.624081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T14:03:22.242986Z digest=sha256:fa2e3909128a3abf8d8077bbe7ecc8ce103a600f2e603779041bab4da66fc756

Pith citing papers

No inbound Pith citation observations are available.