Pith. sign in

Paper Citation Record · LEDGER

Object-aware Sound Source Localization via Audio-Visual Scene Understanding

As of 10 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 0 inbound Pith citation observations for arXiv:2506.18557.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.18557 v2

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:20:49.374967Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

49 of 49 outbound references displayed

  • verified exact1
  • verified fuzzy40
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cc21691c-a3a3-4237-830a-168b45d996dc · outbound

This paper cites write newline.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:44.623191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:20:44.623191Z digest=sha256:a1f899f80ec7c98aafe10e967ca71b2e254414ddaad4a1e23a22a96c8d49fd14

Observation e009438d-2da3-44a7-937e-2cab930c2b7a · outbound

This paper cites Objects that sound.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Objects that sound

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:59.016524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T23:20:44.739030Z digest=sha256:1519f6c3a604969e42a1c4f5c1f69a02fd0c19ffbf96bb440c4e9dcb19c83b57

Observation dcfd5eab-c51b-4ab0-9a80-ba0fed71de2a · outbound

This paper cites Vggsound: A large-scale audio-visual dataset.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Vggsound: A large-scale audio-visual dataset

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:58.812481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T23:20:44.827505Z digest=sha256:cd6a66970ca6aedf6331c0e8bc2a881bd2ee8fcf73fbe65fee6979cbe409084b

Observation f152976f-0ba2-4acb-954e-ec303d1b96a0 · outbound

This paper cites Localizing visual sounds the hard way.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Localizing visual sounds the hard way

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:58.588619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T23:20:44.963367Z digest=sha256:bc00096390787063a06a4ba0f04c6544e2501867f6c8efda9a7bf273946b8f08

Observation dd4af06e-17c5-4d08-8ab4-0589e05ecb75 · outbound

This paper cites Exploring simple siamese representation learning.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Exploring simple siamese representation learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:45.067294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:20:45.067294Z digest=sha256:45fba23b97cc310cdc8310562813c80f9eb4910c76c07d4f38e156233c4af4cb

Observation 9d7937df-f7d3-4751-b852-a6663328b554 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:58.416200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T23:20:45.124745Z digest=sha256:e01eae1ee9bfec23e0b7c295f530e048d1a6dcd3c86742cf17cbeedeaeb8fb5a

Observation 71495029-6d41-4882-920a-4fd2ee8da924 · outbound

This paper cites Sinkhorn distances: Lightspeed computation of optimal transport.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Sinkhorn distances: Lightspeed computation of optimal transport

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:45.241334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:20:45.241334Z digest=sha256:dcb09762ba8dae92721fd24e3a7e750a9e6ff0669ff7330fee77a8b110d8eaeb

Observation cb5fff76-df51-49c0-8092-58bc5216a9f1 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Imagenet: A large-scale hierarchical image database

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:45.352669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:20:45.352669Z digest=sha256:2b5d88f317cfe90317fc2f7d24e15046adcf1cd6f539cd6c9112197f5067a84f

Observation 100fedaa-026b-4114-907f-c8f5cbcad8dd · outbound

This paper cites Cross-modal prompts: Adapting large pre-trained models for audio-visual downstream tasks.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Cross-modal prompts: Adapting large pre-trained models for audio-visual downstream tasks

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:58.265975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T23:20:45.494963Z digest=sha256:e84b48e3aee8fc1fb0e62b51e65df0edab74842880c166c70eb981eeb7b8b2ac

Observation 12d58a27-c3a6-42b7-879b-6a065254e04c · outbound

This paper cites With a little help from my friends: Nearest-neighbor contrastive learning of visual representations.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding With a little help from my friends: Nearest-neighbor contrastive learning of visual representations

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:58.116336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T23:20:45.594337Z digest=sha256:86adf8560731e18ca681901d77e6a98bfdd73b940f3080c3cabfb1639dbaa3ae

Observation ffcba95a-21df-47a5-8d80-a8e8ed04626e · outbound

This paper cites Hear the flow: Optical flow-based self-supervised visual sound source localization.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Hear the flow: Optical flow-based self-supervised visual sound source localization

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:57.831456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T23:20:45.678394Z digest=sha256:6ab75dfe5f94197134ed15220011a6e714806c7a56d9fd8e808e9a433e6a74b1

Observation 686ef128-6d45-4c11-a4da-ae432dc4a78a · outbound

This paper cites Audioclip: Extending clip to image, text and audio.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Audioclip: Extending clip to image, text and audio

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:57.654611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T23:20:45.810588Z digest=sha256:99bc27f7f5f4092e7b3cbda018f498d3687510e97f364a013ff585824ae6b7ce

Observation 314c11cc-f68e-4daa-97a6-0aad424811eb · outbound

This paper cites Deep residual learning for image recognition.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Deep residual learning for image recognition

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:45.882461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:20:45.882461Z digest=sha256:ca84487ad4f4457c88e14a1c0f90f7688b35e3dd0681ee4e4c541d514f4cda3d

Observation 45025512-b2a6-4da2-b930-10160292978c · outbound

This paper cites Deep multimodal clustering for unsupervised audiovisual learning.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Deep multimodal clustering for unsupervised audiovisual learning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:57.394754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T23:20:45.999412Z digest=sha256:ae0e62fd714f80fe6aa7c69fc37a4809eed3434aad21cd5f14701c38d02280d7

Observation 64deb004-f7ae-4e9b-abbe-a48c365b3762 · outbound

This paper cites Discriminative sounding objects localization via self-supervised audiovisual matching.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Discriminative sounding objects localization via self-supervised audiovisual matching

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:57.153525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T23:20:46.075610Z digest=sha256:4a832e9a09a63d5d7780658d94fb624b2cab6bf094b5ce048bf6c8f6f249071b

Observation cc069196-7209-4cd1-bb72-3dc76df8ebdf · outbound

This paper cites Mix and localize: Localizing sound sources in mixtures.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Mix and localize: Localizing sound sources in mixtures

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:56.914740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T23:20:46.152099Z digest=sha256:bb767bc4d1ecdf66c1df659417fbbdf9ddaa032ea32d5525dd4a5b853ad86733

Observation 9faa0021-b426-4b9a-85af-cd5d45bafa24 · outbound

This paper cites Boosting contrastive self-supervised learning with false negative cancellation.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Boosting contrastive self-supervised learning with false negative cancellation

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:56.674748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T23:20:46.260804Z digest=sha256:0c542e604f9d7f53f3a941bcce1528242c7fcddd481e452303525459cb5f013d

Observation f11684f8-189a-4455-bcfd-dfee9d4c5ca1 · outbound

This paper cites A review of recent advances on deep learning methods for audio-visual speech recognition.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding A review of recent advances on deep learning methods for audio-visual speech recognition

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:56.385786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T23:20:46.326449Z digest=sha256:dd353b62266ac467f0e12e1724fdf788a49fb96423cddcc146a3fe6d1e71f869

Observation ac83f528-b373-4ce5-94d9-41338915fbbb · outbound

This paper cites Learning to visually localize sound sources from mixtures without prior source knowledge.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Learning to visually localize sound sources from mixtures without prior source knowledge

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:56.134749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T23:20:46.404760Z digest=sha256:d9a8ed512ebcc23f1425340cb03a8528ef134b6e31dd9dfb7a9ced1ea8a11d83

Observation 085df064-f62c-4d87-8d26-60ae9757c918 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Adam: A Method for Stochastic Optimization

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:46.466395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:20:46.466395Z digest=sha256:0de9a6bfab9b9b7e1eaef91907af99e328428d253962c820b0809f822d34af2e

Observation 572067d7-d2fc-498e-968e-d1122aeb18d6 · outbound

This paper cites Unsupervised sound localization via iterative contrastive learning.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Unsupervised sound localization via iterative contrastive learning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:55.942335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T23:20:46.543786Z digest=sha256:25554d93df5d19942080a44072c6a2e4fde6c8515eb6d716221ebe99b598bef0

Observation e9005eba-beb0-4a48-98b5-bd5048d15ffb · outbound

This paper cites Exploiting transformation invariance and equivariance for self-supervised sound localisation.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Exploiting transformation invariance and equivariance for self-supervised sound localisation

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:55.749982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T23:20:46.716945Z digest=sha256:edee63ebcfc56f6e502894238a5be9c2dcfbaa07e9f16a51aa2ce056dba5b9e5

Observation 5ee4ac45-00ea-492c-aee1-293e2cc937d9 · outbound

This paper cites Generalized video anomaly event detection: Systematic taxonomy and comparison of deep models.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Generalized video anomaly event detection: Systematic taxonomy and comparison of deep models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:55.434742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T23:20:46.825427Z digest=sha256:9ba3d42a78aaac80a1f0c7b71cd135f18a4291710b57eda3400a47dc05a60816

Observation dab2a57e-d48f-462e-ae71-3874f91a96fa · outbound

This paper cites T-vsl: Text-guided visual sound source localization in mixtures.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding T-vsl: Text-guided visual sound source localization in mixtures

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:55.207310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T23:20:47.179413Z digest=sha256:d7dd691ab6e86c61d03d432615e9718c3ea5f08a01d649a88599ac13082d01e8

Observation d6076cb8-c9fa-45a4-a79f-0e9f64d5713c · outbound

This paper cites Localizing visual sounds the easy way.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Localizing visual sounds the easy way

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:54.989857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T23:20:47.319708Z digest=sha256:f6137a6565fd803ad8bac8faeac0a2780b1a3c5c741e5a194c80bd87c53ef144

Observation 510a2c9a-c178-450c-9083-5d6387ce1c88 · outbound

This paper cites A closer look at weakly-supervised audio-visual source localization.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding A closer look at weakly-supervised audio-visual source localization

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:54.793986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T23:20:47.399677Z digest=sha256:60cd7e48224232bc0a35d67effb5e8100f77fc4cb1c2e43bfa215f950b6a8d2b

Observation 7ead8897-a468-4d84-9723-bbdd899358ea · outbound

This paper cites Audio-visual grouping network for sound localization from mixtures.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Audio-visual grouping network for sound localization from mixtures

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:54.602287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T23:20:47.445298Z digest=sha256:a3b515f3e075a464b11df0b93908482a738f373be626712dbc9e0f87b6d4c506

Observation a7f2f281-70a1-44d5-a856-c0cdf82229f3 · outbound

This paper cites Audio-visual scene analysis with self-supervised multisensory features.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Audio-visual scene analysis with self-supervised multisensory features

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:54.380654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T23:20:47.474313Z digest=sha256:130881c23926d90b883f4f25fb08899f2fd73e785b8174d95209ec3f9f7fa9d5

Observation 78868aa1-7c8c-4746-98c3-374d3e091bc9 · outbound

This paper cites Multiple sound sources localization from coarse to fine.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Multiple sound sources localization from coarse to fine

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:54.184834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T23:20:47.545828Z digest=sha256:886fb3faa3bd7d8c2a1cb7aedf64c85cd31318e48a95fc540204b85c7637e661

Observation a8698035-bfd8-4a4e-993b-51396e7e26a4 · outbound

This paper cites Multimodal Open-Vocabulary Video Classification via Pre-Trained Vision and Language Models.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Multimodal Open-Vocabulary Video Classification via Pre-Trained Vision and Language Models

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-06T23:20:49.591703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T23:20:47.607798Z digest=sha256:a129a38717b23dd8646610e50197551986ddac284c69a4ad0f2d9df6bb19807f

Observation 9b0ae4b5-47ac-48ad-9615-77b480b67d4a · outbound

This paper cites Mm-diffusion: Learning multi-modal diffusion models for joint audio and video generation.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Mm-diffusion: Learning multi-modal diffusion models for joint audio and video generation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:53.964747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T23:20:47.662872Z digest=sha256:b85de496b0f27dd23736342f4c86e5b09dca94757364b9837e2fa6355d7e6957

Observation d51aa6db-5f85-46d9-ad3c-9af470c889bc · outbound

This paper cites Learning to localize sound source in visual scenes.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Learning to localize sound source in visual scenes

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:53.797386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T23:20:47.742209Z digest=sha256:ed8aeddb2f938140810b44af441ed0507ad8987dc1547ac29a50ed37c3a148bc

Observation fc4440a1-9511-4a3e-a515-fb31081f8a81 · outbound

This paper cites Learning sound localization better from semantically similar samples.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Learning sound localization better from semantically similar samples

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:53.536440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T23:20:47.780015Z digest=sha256:99867542066052145efe6df5b305c83b915058b8045bd8c86e25316ebd52fbbe

Observation c3c94bfc-8ce8-4f7c-99ff-cfaa9bf03581 · outbound

This paper cites Sound source localization is all about cross-modal alignment.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Sound source localization is all about cross-modal alignment

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:53.104744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T23:20:47.952858Z digest=sha256:471d6de5ab4926bf20c2931d373dd125e34d0486c355936e7c9aedfbf296185d

Observation edc0cf83-fb89-4bc0-8973-1cce97dd04a5 · outbound

This paper cites Unsupervised sounding object localization with bottom-up and top-down attention.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Unsupervised sounding object localization with bottom-up and top-down attention

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:52.876261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T23:20:47.985284Z digest=sha256:3a60e5ac68fe417eab708cf4d0054d7b6ea28f65483e0bb03a97a7e1a6d4e3d6

Observation e821573b-01b8-42d4-b2ee-4af3b572fda0 · outbound

This paper cites Flowgrad: Using motion for visual sound source localization.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Flowgrad: Using motion for visual sound source localization

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:52.534945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T23:20:48.061595Z digest=sha256:dbf89906ae6a0baa80d510ccae830fcb49cf14ae355d9da7a0355b683aa28d96

Observation 9b7a432f-dd00-4e2e-880b-98c6df771e1f · outbound

This paper cites Self-supervised predictive learning: A negative-free method for sound source localization in visual scenes.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Self-supervised predictive learning: A negative-free method for sound source localization in visual scenes

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:52.215295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T23:20:48.196924Z digest=sha256:0bca4bbaaa12c721b4e24d04ce6e201f0d3967678e1f90f69a115c7b9c9b650d

Observation 0099b1eb-197c-4982-bbdd-470db32f7dbe · outbound

This paper cites Learning audio-visual source localization via false negative aware contrastive learning.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Learning audio-visual source localization via false negative aware contrastive learning

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:51.998121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T23:20:48.316561Z digest=sha256:beda3a46cc4eadf451fa69e6f3ff24df659ed7e6e8ee3cb2f6b6f45ab6ae64ed

Observation 04793bae-3c14-47f6-b257-7b79e6ed984b · outbound

This paper cites Audio-visual spatial integration and recursive attention for robust sound source localization.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Audio-visual spatial integration and recursive attention for robust sound source localization

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:51.739907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T23:20:48.413073Z digest=sha256:1194a6b26fac6903cdef8ed1660c6189023fb9eb9767a1ebb734add368bf1d2d

Observation a4793b17-fe85-4b43-8de0-53d48099bcae · outbound

This paper cites Watch video, catch keyword: Context-aware keyword attention for moment retrieval and highlight detection.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Watch video, catch keyword: Context-aware keyword attention for moment retrieval and highlight detection

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:51.514758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T23:20:48.481956Z digest=sha256:04ad9683e2c4ae4508c6b1ad41703a545f417fd93a396e78cd3498d65e3a38a9

Observation de45388e-df01-4cc3-aafa-63a089495525 · outbound

This paper cites V2a-mapper: A lightweight solution for vision-to-audio generation by connecting foundation models.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding V2a-mapper: A lightweight solution for vision-to-audio generation by connecting foundation models

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:51.283695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T23:20:48.538040Z digest=sha256:005c8cd26b4d73c2793cd5a60fe0943225e4c73dc06ed35be08f8620043b5606

Observation 02f9acb8-8e93-4110-856a-999f7b07683d · outbound

This paper cites Multimodal large language models: A survey.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Multimodal large language models: A survey

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:51.046232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T23:20:48.669587Z digest=sha256:9acfc50d6f0b59b85a90b85f76978c826f400709608e3a73ff9a78d0eeee396f

Observation caf9bdfc-63e5-467e-a050-368ba2d6297b · outbound

This paper cites Sonicvisionlm: Playing sound with vision language models.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Sonicvisionlm: Playing sound with vision language models

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:50.814775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T23:20:48.722137Z digest=sha256:103f4a88bc937ffe957a7c1ecc0f3f393ef1b10b1e25373af1cdc69a4e67ba34

Observation e09ecbac-9ff0-495b-9354-2b1ed2b00458 · outbound

This paper cites A proposal-based paradigm for self-supervised sound source localization in videos.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding A proposal-based paradigm for self-supervised sound source localization in videos

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:50.588081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T23:20:48.915324Z digest=sha256:9751d803901d97f930d4ec1c6a61a0e9fe7b6211b6bb7711f1b77fdc55a0f423

Observation 0ae44f15-eda2-4597-b370-81865df6e06d · outbound

This paper cites A Survey on Multimodal Large Language Models.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding A Survey on Multimodal Large Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:48.985249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:20:48.985249Z digest=sha256:2002a5aaff38ac594a53142d4162aecc707cebdb4f745beb7c81e712fd36a190

Observation 52a526d2-17b8-4015-8ead-64026d200fcc · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:49.075435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:20:49.075435Z digest=sha256:fce0eb1d55683fe14612e7efba8927ab3b718bbd8b6783a1b99706a8ec4e5771

Observation 7d765c18-5d41-445c-a565-addc16c9924e · outbound

This paper cites The sound of pixels.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding The sound of pixels

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:50.348537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T23:20:49.160956Z digest=sha256:a4f4b6f55b808d4f87019bcc0cd539fefa26c1fcd25c6112b252e10344412ad1

Observation 37e29985-6257-4e8f-af2e-e41c0dfa0027 · outbound

This paper cites Weakly supervised contrastive learning.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Weakly supervised contrastive learning

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:50.052232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T23:20:49.253720Z digest=sha256:0fafce8862e98731e0559e0be1239e71a5e2b54b515a024e58440d59cb802e51

Observation 69480215-0ca9-4260-b091-d46be833e310 · outbound

This paper cites Exploiting visual context semantics for sound source localization.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Exploiting visual context semantics for sound source localization

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:49.795672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T23:20:49.374967Z digest=sha256:2c84b45d7b4921369553ec6a63aff2a83046af49af31dd022fd038c0721313fa

Pith citing papers

No inbound Pith citation observations are available.