Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T21:41:36.060867Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 0 inbound Pith citation observations for arXiv:2506.23623.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T21:41:36.060867Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
44 of 44 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation aa8f3d36-a08a-421c-8fdf-47fd3b0c4810 · outbound
Revisiting Audio-Visual Segmentation with Vision-Centric Transformer YouTube-8M: A Large-Scale Video Classification Benchmark
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23cd1ae3-f838-4f6e-bfd7-6a2f85723205 · outbound
Revisiting Audio-Visual Segmentation with Vision-Centric Transformer Look, listen and learn
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b1e50aeb-48e1-43bd-bf45-e4008fa20df0 · outbound
Revisiting Audio-Visual Segmentation with Vision-Centric Transformer Objects that sound
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 09f6cf4b-f046-4e8f-82c2-bc55882d4daa · outbound
Revisiting Audio-Visual Segmentation with Vision-Centric Transformer End-to- end object detection with transformers
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e5201e53-e6e6-48ff-ad7b-cea907d754ce · outbound
Revisiting Audio-Visual Segmentation with Vision-Centric Transformer Unraveling in- stance associations: A closer look for audio-visual segmen- tation
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation dc189158-aeb2-4d76-a8cc-187179750181 · outbound
Revisiting Audio-Visual Segmentation with Vision-Centric Transformer Cpm: Class-conditional prompting ma- chine for audio-visual segmentation
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation efcd098a-21a3-40c1-975d-f9d168260952 · outbound
Revisiting Audio-Visual Segmentation with Vision-Centric Transformer Masked-attention mask transformer for universal image segmentation
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29abd25d-36ff-4733-a9bb-ffe5ea9ffba4 · outbound
Revisiting Audio-Visual Segmentation with Vision-Centric Transformer Imagenet: A large-scale hierarchical image database
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7b0c203d-9e91-49f0-b601-91f4e7141e6a · outbound
Revisiting Audio-Visual Segmentation with Vision-Centric Transformer Avsegformer: Audio-visual segmentation with trans- former
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e420913e-a1d9-40cc-af95-7490e3fdee29 · outbound
Revisiting Audio-Visual Segmentation with Vision-Centric Transformer Improving audio-visual segmentation with bidirectional generation
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7361d6a7-0cd5-4279-9834-3c7393f42224 · outbound
Revisiting Audio-Visual Segmentation with Vision-Centric Transformer Deep residual learning for image recognition
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac069957-c7ae-472b-89a3-ad52f747f997 · outbound
Revisiting Audio-Visual Segmentation with Vision-Centric Transformer Cnn ar- chitectures for large-scale audio classification
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8c3f7eca-2434-46e9-aab8-232ffd9aca7b · outbound
Revisiting Audio-Visual Segmentation with Vision-Centric Transformer Discriminative sounding objects localization via self-supervised audiovisual match- ing
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c82fe159-1178-473d-aedb-5bb85bc7d67f · outbound
Revisiting Audio-Visual Segmentation with Vision-Centric Transformer Mix and local- ize: Localizing sound sources in mixtures
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 40b57b97-ca7c-4cee-a191-ed46c83c5091 · outbound
Revisiting Audio-Visual Segmentation with Vision-Centric Transformer Discovering sound- ing objects by audio queries for audio visual segmentation
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d25fb5fe-93bb-4656-9b93-1d1cff10c754 · outbound
Revisiting Audio-Visual Segmentation with Vision-Centric Transformer Unleashing the temporal-spatial reasoning capacity of gpt for training-free audio and language referenced video object segmentation
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ee624052-5cc3-4b98-9a52-18deccb58ea6 · outbound
Revisiting Audio-Visual Segmentation with Vision-Centric Transformer Categorical Reparameterization with Gumbel-Softmax
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41ad0a72-2ffe-45a9-aa0f-8e4e1e4ddf6f · outbound
Revisiting Audio-Visual Segmentation with Vision-Centric Transformer Learning to visually localize sound sources from mixtures without prior source knowledge
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e0850671-1124-4d09-9b43-ad461b3e9f85 · outbound
Revisiting Audio-Visual Segmentation with Vision-Centric Transformer Segment and recognize anything at any granularity
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 557eaabc-3bc0-4ac5-b9b6-64bc4c705ff8 · outbound
Revisiting Audio-Visual Segmentation with Vision-Centric Transformer Selm: Selective mechanism based audio-visual segmentation
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 594137aa-e7b9-45a3-b15e-c931e8c2cb97 · outbound
Revisiting Audio-Visual Segmentation with Vision-Centric Transformer Catr: Combinatorial-dependence audio-queried transformer for audio-visual video segmentation
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation bc2fa17b-80a6-4743-8d1d-ba532f346af0 · outbound
Revisiting Audio-Visual Segmentation with Vision-Centric Transformer Qdformer: Towards robust audiovisual segmentation in complex environments with quantization-based semantic decomposition
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 99edfce9-9f0a-4e00-8f3a-90cd10fa66f3 · outbound
Revisiting Audio-Visual Segmentation with Vision-Centric Transformer Audio-visual seg- mentation by exploring cross-modal mutual semantics
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5fd01971-599b-4e35-8b0d-6118fd06e6f5 · outbound
Revisiting Audio-Visual Segmentation with Vision-Centric Transformer Bavs: bootstrapping audio- visual segmentation by integrating foundation knowledge
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5594549c-6997-45d0-be52-b6ec97324966 · outbound
Revisiting Audio-Visual Segmentation with Vision-Centric Transformer Audio-visual segmentation via unlabeled frame exploitation
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 64d8d45c-f16c-46e5-adc9-45f0bd51b49b · outbound
Revisiting Audio-Visual Segmentation with Vision-Centric Transformer Swin transformer: Hierarchical vision transformer using shifted windows
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f4a473fa-928b-4c7b-9581-f898b61efaf2 · outbound
Revisiting Audio-Visual Segmentation with Vision-Centric Transformer Step- ping stones: A progressive training strategy for audio-visual semantic segmentation
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 93a05f2a-da34-4a01-a35e-c31f1cd5d779 · outbound
Revisiting Audio-Visual Segmentation with Vision-Centric Transformer The Concrete Distribution: A Continuous Relaxation of Discrete Random Variables
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a683ffc-aa9d-49a0-b4b7-9f8a1414b584 · outbound
Revisiting Audio-Visual Segmentation with Vision-Centric Transformer Multimodal variational auto-encoder based audio-visual segmentation
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c896ac0a-91cd-4d05-8e86-02e4e2a07110 · outbound
Revisiting Audio-Visual Segmentation with Vision-Centric Transformer V-net: Fully convolutional neural networks for volumetric medical image segmentation
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12acaf33-1cc4-4f26-bc45-2b3916bb233d · outbound
Revisiting Audio-Visual Segmentation with Vision-Centric Transformer Localizing visual sounds the easy way
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 61e08455-777a-485e-922f-b50865cd6976 · outbound
Revisiting Audio-Visual Segmentation with Vision-Centric Transformer ImageNet-21K Pretraining for the Masses
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e212d0b-4b98-4770-9859-b7f7d524c2a8 · outbound
Revisiting Audio-Visual Segmentation with Vision-Centric Transformer Learning to localize sound source in visual scenes
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8558b1c3-c7cd-4b42-a45b-986f175d4a75 · outbound
Revisiting Audio-Visual Segmentation with Vision-Centric Transformer Unveiling and mitigating bias in audio visual segmentation
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4b9407e4-f987-47c8-bfa1-3e9993290f55 · outbound
Revisiting Audio-Visual Segmentation with Vision-Centric Transformer Attention is all you need
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4e5570ef-adf9-4359-b489-04f566873b41 · outbound
Revisiting Audio-Visual Segmentation with Vision-Centric Transformer Pvt v2: Improved baselines with pyramid vision transformer
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 68b63946-3b18-4703-a312-e51b8b6b3182 · outbound
Revisiting Audio-Visual Segmentation with Vision-Centric Transformer Prompting segmentation with sound is gen- eralizable audio-visual source localizer
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 654e324a-4466-4d0d-ba48-662b261e815e · outbound
Revisiting Audio-Visual Segmentation with Vision-Centric Transformer Can textual semantics mitigate sounding object segmentation preference? In ECCV, 2024
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 1a3152aa-e2d3-4363-a01a-37eeae646bef · outbound
Revisiting Audio-Visual Segmentation with Vision-Centric Transformer Language as queries for referring video object segmen- tation
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 694d6013-9516-4e6b-aefd-7e19eb6e14d7 · outbound
Revisiting Audio-Visual Segmentation with Vision-Centric Transformer Groupvit: Semantic segmentation emerges from text supervision
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7b97016a-c9d5-4421-b092-c2d49bdb3fca · outbound
Revisiting Audio-Visual Segmentation with Vision-Centric Transformer Cooperation does matter: Exploring multi-order bilateral relations for audio- visual segmentation
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8f542558-3947-4313-aae6-06267c83d111 · outbound
Revisiting Audio-Visual Segmentation with Vision-Centric Transformer DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae275b49-d7fd-433b-bef5-0c9faecb429b · outbound
Revisiting Audio-Visual Segmentation with Vision-Centric Transformer Audio–visual segmentation
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 0970f7ae-fe0a-4f3b-9879-55ca2b4724c6 · outbound
Revisiting Audio-Visual Segmentation with Vision-Centric Transformer Audio-visual segmentation with semantics
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
No inbound Pith citation observations are available.