Pith. sign in

Paper Citation Record · LEDGER

What's Making That Sound Right Now? Video-centric Audio-Visual Localization

As of 7 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 0 inbound Pith citation observations for arXiv:2507.04667.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.04667 v2

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:47:28.634151Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

37 of 37 outbound references displayed

  • verified exact0
  • verified fuzzy28
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5d25d9bc-8170-472e-af95-cf5808dfbd8f · outbound

This paper cites Self-supervised learning of audio-visual objects from video.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Self-supervised learning of audio-visual objects from video

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:29.045291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:47:28.120371Z digest=sha256:a87bd1c9406febee7a132f39c484d451639533908a84240f06070afe57419296

Observation 1aa1ec69-56e9-4da3-90c5-fc901871e907 · outbound

This paper cites Vivit: A video vision transformer.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Vivit: A video vision transformer

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T19:47:28.262239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:47:28.262239Z digest=sha256:e8b242bbe1538cc2cbcb7346e0e2a9697fbc856560f81432e0da9a0ec4f6e74f

Observation 3320ce3b-6bf2-4ec7-b40e-b04bbcb804e0 · outbound

This paper cites Soundnet: Learning sound representations from unlabeled video.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Soundnet: Learning sound representations from unlabeled video

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:29.025881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:47:28.379125Z digest=sha256:9a4cdcc45c1287131dc619b1e2a4c6eaf7763d64b0e597418be0bf82291a7cf3

Observation 6077323f-fcd4-43c7-ae78-702acd364b30 · outbound

This paper cites Is space-time attention all you need for video understanding? In ICML, page 4, 2021.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Is space-time attention all you need for video understanding? In ICML, page 4, 2021

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:29.013865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:47:28.442925Z digest=sha256:e79674cd72603ebd1739b6b15f223723102784b27b2f36ecdf4298f5ec1aae52

Observation 4b2c3113-78f9-4010-b99d-9ab5c5d32334 · outbound

This paper cites Vggsound: A large-scale audio-visual dataset.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Vggsound: A large-scale audio-visual dataset

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:29.001333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:47:28.516909Z digest=sha256:cf38ddecb8f261c4e2b632198a285d2d41eb418a1ccef69dfd66766f7be2c940

Observation 6db0cbf5-aea7-41de-a057-ff38cf9d5ef5 · outbound

This paper cites Localizing visual sounds the hard way.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Localizing visual sounds the hard way

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:28.989452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:47:28.520441Z digest=sha256:18e22040800aa5058ec84025c107b00f9f7b9a935b6ebd893b0118652a4e1390

Observation dbff7220-d553-448d-a284-ba15e0d503bc · outbound

This paper cites an unresolved cited work.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:47:28.978673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:47:28.523621Z digest=sha256:7f0bdfecc1b1010a47d7c7d95fa090e8f2eb76f63de543f4f4362d9484a9d2d7

Observation 4bc77b1f-e344-41a9-a1ee-cecfed6faeef · outbound

This paper cites Gemmeke, Daniel P.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Gemmeke, Daniel P

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:28.966826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:47:28.526956Z digest=sha256:5ef1d93a3ff3b746a3f212858653ae134a041507bcffd555c44c0ebdae1c766a

Observation 3cf20378-63ed-4945-8d0a-db61d337984e · outbound

This paper cites Liu, David Harwath, Leonid Karlinsky, Hilde Kuehne, and James R.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Liu, David Harwath, Leonid Karlinsky, Hilde Kuehne, and James R

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:28.955716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:47:28.530439Z digest=sha256:27eb5f96e18cd9e6f23168525803a803353e6cd4a087af72e202a90eeb0cd984

Observation 2d615601-776e-4ec3-9ab3-5d54eb5462bb · outbound

This paper cites Dual mean-teacher: An unbiased semi-supervised framework for audio-visual source localization.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Dual mean-teacher: An unbiased semi-supervised framework for audio-visual source localization

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:28.945502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:47:28.533809Z digest=sha256:e3178f529b8231c4acc1d5f56059823a913bb134931c158624141374be0de25b

Observation ace7c10c-b03b-4f78-b764-a66dcc45fe87 · outbound

This paper cites Deep residual learning for image recognition.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Deep residual learning for image recognition

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T19:47:28.536881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:47:28.536881Z digest=sha256:d22b22cb5559bc2c67fffb44b710c41db7650caefec635c5641aebb87efc1a90

Observation ccd609c0-720a-4ce9-b980-5a02641883a7 · outbound

This paper cites Deep multimodal clus- tering for unsupervised audiovisual learning.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Deep multimodal clus- tering for unsupervised audiovisual learning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:28.928999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:47:28.540100Z digest=sha256:3fed4c7536cf36138dffe1fc44e5a3d8bbeebab63cc64d13797e8d7b04e47896

Observation 729f0c5a-698e-4cbc-9517-57dffab53d7d · outbound

This paper cites Mix and local- ize: Localizing sound sources in mixtures.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Mix and local- ize: Localizing sound sources in mixtures

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T19:47:28.543459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:47:28.543459Z digest=sha256:d57cb39d284df08f9dafeebcae124eb597972ade4ce1776558496ec49bc68357

Observation 6523c2d7-6bed-4c0b-9da2-d0bda742bb2b · outbound

This paper cites Egocentric audio-visual object localization.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Egocentric audio-visual object localization

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:28.910711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:47:28.546983Z digest=sha256:685e627668c38aa5d0a91f0b453459c41aa946e7e427c96e429e4e41b78cdf6f

Observation b556ca4e-d0e4-4e44-8542-91f5fe1c0163 · outbound

This paper cites Learning to visually localize sound sources from mix- tures without prior source knowledge.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Learning to visually localize sound sources from mix- tures without prior source knowledge

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:28.899326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:47:28.550853Z digest=sha256:0af1577b3d953aff360e60cf36ddefce957377269d4e96811d04dd3c97c3b5ba

Observation 7a83fb99-9282-440f-aa18-6bb13256a02c · outbound

This paper cites Segment any- thing.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Segment any- thing

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T19:47:28.555315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:47:28.555315Z digest=sha256:316490da6b1d7fb95d3cf9bc2fa9a944a5f2012f1806ac7bc5fa76a22e86ca4a

Observation dd8d6a7a-de6a-4da8-aa46-360b7106627c · outbound

This paper cites Openimages: A public dataset for large-scale multi-label and multi-class image classification.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Openimages: A public dataset for large-scale multi-label and multi-class image classification

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:28.882386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:47:28.558796Z digest=sha256:9300020acddc59a8ed7b3012b519b34ad71fe2960dddb02c4809ba3dc0cb7ba0

Observation 756d42fb-3d7f-48ff-805f-e39fab60c70a · outbound

This paper cites Ex- ploiting transformation invariance and equivariance for self- supervised sound localisation.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Ex- ploiting transformation invariance and equivariance for self- supervised sound localisation

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:28.872036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:47:28.562746Z digest=sha256:c71a3d87dbf472746dbe6871aa4748de323b0870703e99ebc3351eb0ac10c4fc

Observation 64222c37-fb0c-4454-8815-d7f71efe56a5 · outbound

This paper cites A framework for multiple-instance learning.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization A framework for multiple-instance learning

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:28.861737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:47:28.566583Z digest=sha256:846089fa2075de2674928517d84138ac00d45a5dfafaa19625d3a69b1023b527

Observation 4b09be43-d2d4-4ca5-9ca1-702280eeaded · outbound

This paper cites A closer look at weakly- supervised audio-visual source localization.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization A closer look at weakly- supervised audio-visual source localization

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:28.850639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:47:28.570189Z digest=sha256:a73f410c06298fadd342a581ef98ebf7d21e328d6c32c1a329640049cad0aae3

Observation 1f88f78c-6ba2-4fe3-8716-cefef78857d1 · outbound

This paper cites Localizing visual sounds the easy way.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Localizing visual sounds the easy way

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:28.839103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:47:28.574111Z digest=sha256:9192ec4325cb8a445eadc7217912edbb5387d4edfe06095c70376a1131a3b0c9

Observation 4f6053e6-4114-40b8-a6a6-35d4f6652d1e · outbound

This paper cites Audio-visual grouping net- work for sound localization from mixtures.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Audio-visual grouping net- work for sound localization from mixtures

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:28.827331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:47:28.577704Z digest=sha256:119b085b9f7ec20020327632d55af9b26649ed452098a3a09ea09274e4964199

Observation bae5192a-2750-4212-9aaf-42fa629ceb0d · outbound

This paper cites Learn- ing representations from audio-visual spatial alignment.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Learn- ing representations from audio-visual spatial alignment

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:28.816423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:47:28.581769Z digest=sha256:285ee420fc69abf10c36ded5bdddbc5c32fb20670171c314838b464b8d837d2d

Observation 834fcbfb-1546-4393-b36b-9a6975cd24c2 · outbound

This paper cites Can clip help sound source localization? In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 5711–5720, 2024.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Can clip help sound source localization? In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 5711–5720, 2024

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:28.805247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:47:28.585329Z digest=sha256:dff2d88dbfe5ea02f142122fbefc2faec941c92a72b60344ca166e48c48c9431

Observation b500f5a6-4eff-41d1-bf5b-fb9e77c5dedd · outbound

This paper cites Audio-visual object localization and separation using low- rank and sparsity.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Audio-visual object localization and separation using low- rank and sparsity

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:28.794982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:47:28.588633Z digest=sha256:9a0b31ddc1861f0ffc173c3b651c88e79fc28b49c4a723305ad5309d19ac0357

Observation 99621e61-5379-41c4-8f43-73d3fa554fcc · outbound

This paper cites Multiple sound sources localization from coarse to fine.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Multiple sound sources localization from coarse to fine

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:28.784403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:47:28.592354Z digest=sha256:e071030d75947e88e9f7bf226935b652ffdcf4e7f848ba96ba2b050ecc8f5e00

Observation 73338dac-fbe3-46c6-bbfb-0bdb26c5117e · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Learning transferable visual models from natural language supervi- sion

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T19:47:28.595754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:47:28.595754Z digest=sha256:e50f7857ab05a105d3a6cf6d739b480e317fe66c3833fe068423b767bc870b74

Observation 1bed4dee-3972-4944-a7e4-65461141bebf · outbound

This paper cites Real-Time Flying Object Detection with YOLOv8.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Real-Time Flying Object Detection with YOLOv8

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T19:47:28.600083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:47:28.600083Z digest=sha256:0b9b1d5eba8262645c48ed10423cb954b492753a17a064a8dfa9d595a6f99459

Observation 0bc6ec16-b7bf-4a6e-bba4-366b0b56ffc8 · outbound

This paper cites Learning to localize sound source in visual scenes.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Learning to localize sound source in visual scenes

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:28.767085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:47:28.604552Z digest=sha256:bf682dc0e87156051d96b783e11f4b8f756166866c7ff6e43067ebb785e4f12f

Observation 46019f8d-bf6f-498f-bbea-5173633a346d · outbound

This paper cites Sound source local- ization is all about cross-modal alignment.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Sound source local- ization is all about cross-modal alignment

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:28.756066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:47:28.608218Z digest=sha256:9e84e9e4e0b1c3d8273d8fd17d68e2b69efcd667f86994bd647e30ededbe226d

Observation 1a1ba15e-929a-455c-a454-bce4a1f66423 · outbound

This paper cites Aligning Sight and Sound: Advanced Sound Source Localization Through Audio-Visual Alignment.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Aligning Sight and Sound: Advanced Sound Source Localization Through Audio-Visual Alignment

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T19:47:28.611835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:47:28.611835Z digest=sha256:174959b4062e97ccd1c36447a3a9904e27ac3e7330ae65c9cb22b25660788307

Observation 1108011b-256e-48b9-bb8f-624531372e34 · outbound

This paper cites Learning audio-visual source localization via false negative aware contrastive learning.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Learning audio-visual source localization via false negative aware contrastive learning

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:28.743921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:47:28.615704Z digest=sha256:2430ea1608d2efc1f31e7edb920110ba86c18110dac836708a78935531b3e95b

Observation 8da0e691-4b13-4b6b-b5b5-c3adc29ce8f7 · outbound

This paper cites Audio-visual event localization in unconstrained videos.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Audio-visual event localization in unconstrained videos

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:28.732696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:47:28.619505Z digest=sha256:f0f5dd1045eee309493ff2aab50bd1902b37764bfe9b6b74102ca3f46f1d778d

Observation d421adb9-2239-4eb8-b456-106d4c5313dc · outbound

This paper cites Scaling autoregressive video models.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Scaling autoregressive video models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:28.722139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:47:28.622770Z digest=sha256:915ec07d0016d845e61e266eeb1b19e1dabf67cc8e266fd1d17a2f6def6eb840

Observation ebbd4606-7474-40bd-ada4-3083f496b229 · outbound

This paper cites The sound of pixels.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization The sound of pixels

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:28.710135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:47:28.626178Z digest=sha256:59618e3c5e97dfb8304dd45e2736ea8385f666a5ff1425b1c78c59b69c3179ea

Observation b9dcf0b8-4232-4fb7-b00c-f6ec9afee9f1 · outbound

This paper cites Audio-visual segmentation.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Audio-visual segmentation

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:28.698811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:47:28.630375Z digest=sha256:91d3e7a58e57ddfea56f7ddbb59ec64657413f4fa61a53da659fdd39e6853443

Observation 4f421f1c-d8a9-4b1f-b122-40798004a8a2 · outbound

This paper cites Audio-Visual Segmentation with Semantics.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Audio-Visual Segmentation with Semantics

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T19:47:28.634151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:47:28.634151Z digest=sha256:5a9d3fd1d5e00d768a80a36be0a67a8d4527530f5f87a337e6757a28f00a8e63

Pith citing papers

No inbound Pith citation observations are available.