Pith. sign in

Paper Citation Record · LEDGER

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects?

As of 23 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 2 inbound Pith citation observations for arXiv:2502.00358.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.00358 v2

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T19:24:20.464475Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T14:03:19.216337Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T23:06:21.321184Z

Reference resolution

54 of 54 outbound references displayed

  • verified exact6
  • verified fuzzy31
  • unresolved16
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7590ae97-f740-46c7-9be8-1ea125844558 · outbound

This paper cites Look, listen and learn.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Look, listen and learn

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:24:21.108505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-09T19:24:20.277575Z digest=sha256:624abfce446c2497c6758c7d1061d5e5ae6ec3aa6bd575403e3a586081cf4834

Observation f289fa68-93f3-4d34-bcf9-ab473b252dda · outbound

This paper cites Objects that sound.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Objects that sound

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:24:21.098943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-09T19:24:20.282196Z digest=sha256:dfb1d1d4bff9f80eb9b62ffef7357de1db5bd96dd1ad45945492352eabc74409

Observation df9469b7-b5f7-4678-9a25-de7ab25123dc · outbound

This paper cites Multimodal syn- chronization in musical ensembles: Investigating audio and visual cues.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Multimodal syn- chronization in musical ensembles: Investigating audio and visual cues

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:24:21.088660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-09T19:24:20.286016Z digest=sha256:9d46b730f29918e3c19c1bdc6f65168c2752c20320b56a38ebf57c8ad368f53b

Observation 1e6d503c-2281-4486-97d5-1353591a9210 · outbound

This paper cites an unresolved cited work.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-09T19:24:21.077744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-09T19:24:20.290327Z digest=sha256:03d8c82316518c2509785b6d4914c61197c2e86fb69a5a359799652dee661cb6

Observation 069c42ab-714b-4596-92a4-4fdb5ee3e258 · outbound

This paper cites Vggsound: A large-scale audio-visual dataset.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Vggsound: A large-scale audio-visual dataset

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:20.294182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:20.294182Z digest=sha256:f3d6574156f498867a67c9e8b9a39f5c5cd6f12fb6cd43bb1da5d19b1d42d26c

Observation b0faeb93-dd04-482f-bb6f-8cea463edf83 · outbound

This paper cites Localizing visual sounds the hard way.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Localizing visual sounds the hard way

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:24:21.059677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-09T19:24:20.298792Z digest=sha256:2263cf304443eccb3da55551049e900e5e651c79f82e66a4cb92d89539962e1e

Observation 250e5ccd-80b4-49fa-87e0-7ad5cea08e18 · outbound

This paper cites Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolu- tion, and fully connected crfs.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolu- tion, and fully connected crfs

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:24:21.049165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-09T19:24:20.302763Z digest=sha256:5d0686e7c260e6d5478ab458f8a9bad6ff934e877ac92ba032986e0bfba55e30

Observation 403ffee5-05a0-49ee-b7e8-5077c94e0489 · outbound

This paper cites Unraveling in- stance associations: A closer look for audio-visual segmenta- tion.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Unraveling in- stance associations: A closer look for audio-visual segmenta- tion

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:20.306658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:20.306658Z digest=sha256:9bebed56c6a7493c66b0515695ca73a70d423880d84041317abf9d4152c2237c

Observation dc9b2cb6-1412-48c9-a838-bfe33c019e85 · outbound

This paper cites Masked-attention mask transformer for universal image segmentation.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Masked-attention mask transformer for universal image segmentation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:20.310828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:20.310828Z digest=sha256:64ccc554544d90e93992032c9b94f015ac6f25cb0ba1f2226d08f6542097387c

Observation 6ad4a3f6-4f28-485a-808c-cc706cdd7143 · outbound

This paper cites Self-awareness for autonomous systems.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Self-awareness for autonomous systems

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:24:21.025332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-09T19:24:20.314450Z digest=sha256:d9e328f5323fdd3ce8c5c8f6025f127f5773ac6d9bca25e39eee853647c330c1

Observation 78b6a000-990a-419b-a67a-47308e4220f5 · outbound

This paper cites Design of intelligent human-computer inter- action system for hard of hearing and non-disabled people.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Design of intelligent human-computer inter- action system for hard of hearing and non-disabled people

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:24:21.014811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-09T19:24:20.317908Z digest=sha256:1b4439b1ce2987d5de578b69f9481c39d8f35bcb4bb6fe227aa2121789ba8a8d

Observation bf269e69-6846-44f7-9c99-a8ad2ed65804 · outbound

This paper cites Avsegformer: Audio-visual segmentation with trans- former.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Avsegformer: Audio-visual segmentation with trans- former

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:24:21.004497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-09T19:24:20.321551Z digest=sha256:1c1e2fc1f43fc88b57550d4cc2d699a0ca79725e4c105a5cf0cd9bb508fef4c2

Observation 6fc51821-cfa7-497a-99d6-a60706b9127c · outbound

This paper cites Audio set: An ontology and human- labeled dataset for audio events.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Audio set: An ontology and human- labeled dataset for audio events

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:24:20.994065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-09T19:24:20.324911Z digest=sha256:b287a1569f72735917e15d1dd0595bb85569101ce895fc23ab847670de92a82e

Observation f63d4afd-c2b0-454e-94a5-27e305641c1c · outbound

This paper cites Open-Vocabulary Audio-Visual Semantic Segmentation.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Open-Vocabulary Audio-Visual Semantic Segmentation

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-08-09T19:24:20.671139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-09T19:24:20.328347Z digest=sha256:bc1fbc373f86631539a54f86d7a9c1707fb2f2ff6db07dbddd2f8fa653662cdb

Observation 5d722fc3-def5-40ea-9c2c-83845017fc62 · outbound

This paper cites Multi-modal instruction tuned llms with fine-grained visual perception.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Multi-modal instruction tuned llms with fine-grained visual perception

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:24:20.983016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-09T19:24:20.332139Z digest=sha256:00dcf74759db81039c5a724a47732aa0656c4b15e56c8ac5af5eca5e950c9337

Observation e19bd29a-cf44-4012-8854-4a31d3ab6e4a · outbound

This paper cites Cnn archi- tectures for large-scale audio classification.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Cnn archi- tectures for large-scale audio classification

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:24:20.971597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-09T19:24:20.335980Z digest=sha256:914a542f2d463269559e78ee0ef05bee24033c64ec3fc66eb817240231582e67

Observation ce6648f2-77f2-4db1-bc78-df2fd710fecb · outbound

This paper cites Discriminative sounding objects localization via self-supervised audiovisual match- ing.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Discriminative sounding objects localization via self-supervised audiovisual match- ing

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:24:20.960336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-09T19:24:20.339591Z digest=sha256:281a970ebaa1595947db66368044e1fb5770c4cbae32030dd34a418b743eb7e7

Observation 88c8a958-b0eb-496f-9f24-16dd9d0e50b6 · outbound

This paper cites Mix and local- ize: Localizing sound sources in mixtures.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Mix and local- ize: Localizing sound sources in mixtures

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:24:20.950095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-09T19:24:20.343114Z digest=sha256:f8bda11e0ca3ed209f23f07f9717d8f8f75612856c492000667d577cd84e1aae

Observation 65436e9e-d69b-4047-b0cb-9ca97860d307 · outbound

This paper cites Char- acterising soundscape research in human-computer interac- tion.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Char- acterising soundscape research in human-computer interac- tion

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:24:20.941001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-09T19:24:20.346602Z digest=sha256:2a17a544a3579d137ecad142f5591f0ff1744b708fe7ffc041ed2587056f23c9

Observation 1afa3391-9c82-4eff-b4af-bfeab31b4a5f · outbound

This paper cites A Critical Assessment of Visual Sound Source Localization Models Including Negative Audio.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? A Critical Assessment of Visual Sound Source Localization Models Including Negative Audio

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-09T19:24:20.645939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-09T19:24:20.350125Z digest=sha256:44f2f089fd1197ce5286c96bf24f867490ad539beae8296731409fce3b43cfd6

Observation ce65fb0d-0146-4d7f-bcbb-b6981be59b81 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Adam: A Method for Stochastic Optimization

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:20.353719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:20.353719Z digest=sha256:342ccb478091c7ca844aef1238a2f22694109dff8b8d28b40a9af39aef8a18a9

Observation c6ab8375-0aec-4086-9608-c6667db94dee · outbound

This paper cites Panoptic feature pyramid networks.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Panoptic feature pyramid networks

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:24:20.931140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-09T19:24:20.357125Z digest=sha256:a6aa84ecc498b117c28564726affffda4c8e8845c1729d5bd2a28e4224dc8297

Observation 202a8954-99e9-420a-878d-3bae15edb23d · outbound

This paper cites Segment any- thing.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Segment any- thing

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:20.360276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:20.360276Z digest=sha256:b16e5a53b7355ed4aee36b8cb3210e8a564b0b7ecbcdd47bee4640e72def3349

Observation d0c00e84-ce2b-47f1-9ab0-9ab5363234cf · outbound

This paper cites Catr: Combinatorial-dependence audio-queried transformer for audio-visual video segmentation.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Catr: Combinatorial-dependence audio-queried transformer for audio-visual video segmentation

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:24:20.915529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-09T19:24:20.363453Z digest=sha256:7310ecdc87e35312529167bf71d404af4347744ba00a849a651f22cefcd888bb

Observation ea1d684b-ae15-4ee4-9100-e6f73176d1f8 · outbound

This paper cites Qdformer: Towards ro- bust audiovisual segmentation in complex environments with quantization-based semantic decomposition.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Qdformer: Towards ro- bust audiovisual segmentation in complex environments with quantization-based semantic decomposition

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:24:20.905024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-09T19:24:20.366364Z digest=sha256:e4ba03ce62dfdb72ede645ca422e2ddfe400e0996820be4db0e6fdf731dfeddf

Observation 5e9f3fc5-d5cb-4f4c-9f29-88eaf79d2a51 · outbound

This paper cites Audio-aware Query-enhanced Transformer for Audio-Visual Segmentation.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Audio-aware Query-enhanced Transformer for Audio-Visual Segmentation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:20.369849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:20.369849Z digest=sha256:dd896fffa8f701c4885c56c6952a5a99af0a83f10d95006dd1cf8a9b148ad50b

Observation 6650e041-06a1-4230-9ba1-b5e8094d9522 · outbound

This paper cites Annotation-free audio-visual segmentation.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Annotation-free audio-visual segmentation

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:24:20.894454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-09T19:24:20.374215Z digest=sha256:798f8d9f9ad453ca0cccaba3df13ac7c32101ac8e4d3d9d8001ce7113a0a348c

Observation 22b3d3ef-30b4-4719-b152-91e9622b1339 · outbound

This paper cites Decoupled Weight Decay Regularization.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Decoupled Weight Decay Regularization

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:20.377738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:20.377738Z digest=sha256:7dcf0ba5e13a5680f802a2ec69b1c5699250a5b6e7d7f35309394e594e9ef89f

Observation 23c13131-eb53-4202-8298-3281ddda682d · outbound

This paper cites Deep learning for intelligent human–computer in- teraction.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Deep learning for intelligent human–computer in- teraction

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:24:20.883756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-09T19:24:20.381448Z digest=sha256:40a19a061819c3f98763de59ce701fe1909dca86ad2cf532d67470955ac70341

Observation 8aec4654-2427-443b-8e48-7723b5de656a · outbound

This paper cites Stepping Stones: A Progressive Training Strategy for Audio-Visual Semantic Segmentation.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Stepping Stones: A Progressive Training Strategy for Audio-Visual Semantic Segmentation

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-09T19:24:20.600464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-09T19:24:20.384721Z digest=sha256:1b3ed90cebed6c1a502005adefc769df2d349b7170aacd1b0c3bbb9f5c8ad95c

Observation e1f277c6-4b03-44bb-b872-4c440f75739e · outbound

This paper cites T-vsl: Text-guided visual sound source localization in mixtures.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? T-vsl: Text-guided visual sound source localization in mixtures

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:20.388365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:20.388365Z digest=sha256:532ae61e865e0a7f9d42e3a843e9747488b12b523f4919b8377220de1062398b

Observation 5a4d91c5-2570-482e-a667-a3d74f5351d0 · outbound

This paper cites A closer look at weakly- supervised audio-visual source localization.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? A closer look at weakly- supervised audio-visual source localization

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:24:20.863671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-09T19:24:20.391561Z digest=sha256:52c1464cf413220aec94a0e23c5bc1ac321713d39732755bdf5e5dd78a4150a1

Observation 5e6e0b2f-196d-471b-8e61-9aa5d2eea53b · outbound

This paper cites AV-SAM: Segment Anything Model Meets Audio-Visual Localization and Segmentation.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? AV-SAM: Segment Anything Model Meets Audio-Visual Localization and Segmentation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:20.394989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:20.394989Z digest=sha256:e2a82ac8c39f7942b0565624eb09053bc07acde677a84c2e530cb97da7bcf385

Observation 70a85fd3-07f4-43e8-8e23-f392a5eeadad · outbound

This paper cites Multi-scale Multi-instance Visual Sound Localization and Segmentation.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Multi-scale Multi-instance Visual Sound Localization and Segmentation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:20.398584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:20.398584Z digest=sha256:2d2c1cadee1ee86451012597db3223d371163bf55a1bd7644e0453cb919edc05

Observation 59d920ce-cba4-45de-93d0-df03dcf96615 · outbound

This paper cites Balanced multimodal learning via on-the-fly gradient modulation.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Balanced multimodal learning via on-the-fly gradient modulation

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:24:20.850493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-09T19:24:20.402033Z digest=sha256:716a52f4c52100f57fba43186f3797a5070242cbe470ca24980f247f9211080a

Observation 72e4aa2f-0220-42b7-9ed8-d2ad24c307f5 · outbound

This paper cites Multiple sound sources localization from coarse to fine.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Multiple sound sources localization from coarse to fine

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:24:20.837026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-09T19:24:20.406084Z digest=sha256:0628be507a54c196d1c3e035740059619a940566b3dbec374cc0a4782dab1f02

Observation 9e795381-6bd6-47f5-87ad-53956781f52e · outbound

This paper cites Imagenet large scale visual recognition challenge.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Imagenet large scale visual recognition challenge

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:20.409475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:20.409475Z digest=sha256:74814edd8f75d3558600098dbfbe13c2fb3d40407cd73f0a1ea101da57a73a86

Observation 6242c99e-24e2-457f-b4df-af9dc4d793a9 · outbound

This paper cites Acoustic self-awareness of autonomous sys- tems in a world of sounds.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Acoustic self-awareness of autonomous sys- tems in a world of sounds

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:24:20.808735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-09T19:24:20.412783Z digest=sha256:2c1743be184a839eeb7f5650c1398b519b82464c91ed4b76af9954a293153c44

Observation 7c6af101-8642-46d1-96d2-f8c22e40c9d9 · outbound

This paper cites Learning to localize sound source in visual scenes.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Learning to localize sound source in visual scenes

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:24:20.796559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-09T19:24:20.416214Z digest=sha256:fdf80aa89db5fc0a9aa3fd6fbaf8fa692f514530991f3cf7fdd7b0ef1f57db01

Observation 05e60b53-f984-4953-aa14-636976fbe689 · outbound

This paper cites Odor/taste integration and the perception of flavor.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Odor/taste integration and the perception of flavor

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:24:20.783592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-09T19:24:20.419549Z digest=sha256:44796e7a0fe7aaacd189eb83262cd229fc15d9fd3d30c46851d5081bc3356ad0

Observation 624ff189-17f9-4269-a392-e5f4fcf6ac16 · outbound

This paper cites Unveiling and Mitigating Bias in Audio Visual Segmentation.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Unveiling and Mitigating Bias in Audio Visual Segmentation

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-08-09T19:24:20.564339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-09T19:24:20.422869Z digest=sha256:d81e93ac541d7d8837ea82d9ff789cc034d4b43aacf2b22fc7ac6fb039162f4f

Observation c664f050-f538-472c-a8fd-9484947e3cea · outbound

This paper cites Unified mul- tisensory perception: Weakly-supervised audio-visual video parsing.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Unified mul- tisensory perception: Weakly-supervised audio-visual video parsing

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:24:20.770675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-09T19:24:20.426645Z digest=sha256:13c07c0a80fea8e7e365691eb1d53e526bfb81adbb0ddc13e823f870e66f611c

Observation 6011ac71-824c-45f6-a37b-5b3ddd781939 · outbound

This paper cites Assured Autonomy: Path Toward Living With Autonomous Systems We Can Trust.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Assured Autonomy: Path Toward Living With Autonomous Systems We Can Trust

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:20.430094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:20.430094Z digest=sha256:f33159a0a4e775549efad6ef9ea485774c192511babd248d52cc181151d01d82

Observation 6d0e6881-6005-4f12-aecb-dac4844dd551 · outbound

This paper cites AudioBench: A Universal Benchmark for Audio Large Language Models.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? AudioBench: A Universal Benchmark for Audio Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:20.433964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:20.433964Z digest=sha256:1a420ee1ed81100bd29a39b9a4b7069954f2313f1bb63343b4da0f646b92bf8a

Observation 001c15d9-596c-4bfa-a0c0-462ff0717e40 · outbound

This paper cites What makes train- ing multi-modal classification networks hard? In Proceed- ings of the IEEE/CVF conference on computer vision and pattern recognition, pages 12695–12705, 2020.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? What makes train- ing multi-modal classification networks hard? In Proceed- ings of the IEEE/CVF conference on computer vision and pattern recognition, pages 12695–12705, 2020

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:24:20.757431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-09T19:24:20.437710Z digest=sha256:cc88557c292bdbf8c963250b8adaaecb3d74362bc4cfc3b1fc355a358857c34b

Observation e47230cf-78ed-4edf-8fbd-e72631751744 · outbound

This paper cites Pvt v2: Improved baselines with pyramid vision transformer.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Pvt v2: Improved baselines with pyramid vision transformer

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:24:20.745520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-09T19:24:20.441068Z digest=sha256:b6cdf39ad4478de5751506992d7138f0b874f6323b79141098223298b071b6fc

Observation 8a03ae3e-458a-4034-9434-c9856878e4ac · outbound

This paper cites Prompting segmentation with sound is gen- eralizable audio-visual source localizer.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Prompting segmentation with sound is gen- eralizable audio-visual source localizer

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:24:20.733937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-09T19:24:20.444378Z digest=sha256:6ea6110f9cf3c28e8752f730a862823b39708279cf2833ecd32d871d44b85f3f

Observation bd44ef52-0ab8-49ce-8156-768eb8918bc1 · outbound

This paper cites Can Textual Semantics Mitigate Sounding Object Segmentation Preference?.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Can Textual Semantics Mitigate Sounding Object Segmentation Preference?

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-08-09T19:24:20.529588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-09T19:24:20.447630Z digest=sha256:ee653a9895bc878fa68633e6cbbf4c2f9c6af3a89620fa18e3eb8d8d0d409c6f

Observation fc4dd78f-a199-4c1d-8114-ca097ae91fe3 · outbound

This paper cites Ref-AVS: Refer and Segment Objects in Audio-Visual Scenes.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Ref-AVS: Refer and Segment Objects in Audio-Visual Scenes

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-08-09T19:24:20.513951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-09T19:24:20.450615Z digest=sha256:d6b2904f87ec3933ff91a327d8148833dfa713d6da83e813cd36a0d7abae9e8b

Observation 5f904dfc-2706-4170-9a55-646f8d6fc03e · outbound

This paper cites MMPareto: Boosting Multimodal Learning with Innocent Unimodal Assistance.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? MMPareto: Boosting Multimodal Learning with Innocent Unimodal Assistance

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:20.453394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:20.453394Z digest=sha256:8cef59bbbf99455ccef514fb8d50c23f7a4d2f0ef59573340e6a4b901b86f796

Observation 9f9efe18-145c-4fdc-aa98-12120b0235c0 · outbound

This paper cites Analyzing audiovisual data for understanding user’s emotion in human- computer interaction environment.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Analyzing audiovisual data for understanding user’s emotion in human- computer interaction environment

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:24:20.720184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-09T19:24:20.456185Z digest=sha256:7aae555b088e81f90eabf16f074ec1f0fa80d65f4166a670b3b045444c17b1ca

Observation a4c5a844-0fa7-4fae-ba4a-a4ce09f3492d · outbound

This paper cites Cooperation does matter: Exploring multi-order bilateral relations for audio- visual segmentation.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Cooperation does matter: Exploring multi-order bilateral relations for audio- visual segmentation

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:24:20.707155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-09T19:24:20.459070Z digest=sha256:39e1f255c5108e37e001d49fbcbeed52da805e920f0a4f8b39bf06a2f347a94c

Observation f416126d-e8c2-43e7-9803-d90922dd7128 · outbound

This paper cites Audio–visual segmentation.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Audio–visual segmentation

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:20.461698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:20.461698Z digest=sha256:f469a169dca97d85a8e465c0f30b611a1c98fb59cd1b0b0b0144d6f1a8b88c58

Observation 6628f8ac-7e2f-4c0e-9575-3549b0176a7c · outbound

This paper cites 1, 2, 3, 4, 6, 7, 8, 14 Supplementary Material A.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? 1, 2, 3, 4, 6, 7, 8, 14 Supplementary Material A

Reference 403

Resolution
malformed identifier
raw_fallback, observed 2026-08-09T19:24:20.684763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-09T19:24:20.464475Z digest=sha256:122cb1d4c5b769534830bb6d12719b25ad01e29e4f4ad2eab141ff2fd0f27a6a

Pith citing papers

Observation 5ac2cb49-56c7-419a-97c4-702c5577b400 · inbound

Learning from Silence and Noise for Visual Sound Source Localization cites this paper.

Learning from Silence and Noise for Visual Sound Source Localization Do Audio-Visual Segmentation Models Truly Segment Sounding Objects?

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T14:03:19.216337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:03:19.216337Z digest=sha256:314b9c8974cecd11bdaf9d2fb0eba06314d534596b4ff4adc2f9d4f4d24a11c9

Observation a7d7f17a-5f35-478f-9cd9-388ff19e134c · inbound

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs cites this paper.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs Do Audio-Visual Segmentation Models Truly Segment Sounding Objects?

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-07-01T23:06:21.323631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:a4cedddaa520a934da40412fd5368d79a94a94b6f949ee5fb42a0060dc41eca7