Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T04:23:41.635307Z
Paper Citation Record · LEDGER
As of 19 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 0 inbound Pith citation observations for arXiv:2412.01488.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T04:23:41.635307Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
52 of 52 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation e6e7773a-7197-48d7-94fb-c29b0d7da8fb · outbound
TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Emerging properties in self-supervised vision transformers
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 2da26d9d-7167-467b-a525-caf0e041a8ed · outbound
TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Localizing visual sounds the hard way
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 67e6dbe0-b1d5-4eee-be42-36d42a8b93b6 · outbound
TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Vggsound: A large-scale audio-visual dataset
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 8f790430-b373-4c93-a839-ff6924b20b6d · outbound
TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Vast: A vision-audio-subtitle-text omni-modality foundation model and dataset, 2023
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation be7389c0-8488-4d44-91d3-bbaa887a0bb1 · outbound
TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Reproducible scaling laws for contrastive language-image learning
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 76edd49a-a693-4939-9fcc-f3d76e00b591 · outbound
TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Meerkat: Audio-visual large language model for grounding in space and time
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ac52431f-cea6-4b2b-85a6-6cc1bcd83f86 · outbound
TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Multilayer nonnegative matrix factorisation
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 9bfa13ba-454b-4d59-a8d4-e989254fe9c9 · outbound
TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Deep feature factorization for concept discovery
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 43db80ca-88b3-45ec-8b0e-0852439267b8 · outbound
TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Neural Network Matrix Factorization
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2066cef2-5c37-433d-8131-9aae3db6beb2 · outbound
TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Clap learning audio concepts from natural language supervision
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation e77553b9-a849-4cd7-99f6-c81de089b320 · outbound
TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Avsegformer: Audio-visual segmentation with transformer
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation f3cbbf35-8fa8-4d06-bc28-55fa4df9cf14 · outbound
TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Audio set: An ontology and human-labeled dataset for audio events
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 394d892c-586e-4ea6-92fa-e73e57718423 · outbound
TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization The why and how of nonnegative matrix factorization
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ab4b0125-4112-4438-8aa5-9e52508fe2c7 · outbound
TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Imagebind: One embedding space to bind them all
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a983b572-ac08-48bc-a575-5ef48fcf5c18 · outbound
TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Contrastive Audio-Visual Masked Autoencoder
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09c971e0-ce71-45dd-9998-ccf2ff70be83 · outbound
TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Single channel speech music separation using nonnegative matrix factorization and spectral masks
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation a70468c1-e73e-4d8a-b01c-0f6802c3f190 · outbound
TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Non-negative matrix factorization for face recognition
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 07f012fd-eed4-4236-9d3d-2f0d2baff75c · outbound
TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization chirp" from the
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 7e272c82-839e-4054-ac26-32fcd7b74da5 · outbound
TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Hubert: Self-supervised speech representation learning by masked prediction of hidden units
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation f537bdda-6134-4bdd-8480-0dd85db28d83 · outbound
TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Transfer Learning from Audio-Visual Grounding to Speech Recognition
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cba18e4f-a2b0-4137-b677-8a65b7e2ad6a · outbound
TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization A Critical Assessment of Visual Sound Source Localization Models Including Negative Audio
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd5de8c9-5f6d-4680-9409-9d29238e6539 · outbound
TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Algorithms for non-negative matrix factorization
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 55ab5fa2-739a-4b4c-ba38-ce50bc4d74fc · outbound
TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Unsupervised sound localization via iterative contrastive learning
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation c174fe3b-9ac8-49ea-9e27-4352b71c6c08 · outbound
TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Audio-visual segmentation by exploring cross-modal mutual semantics
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation adea6297-dde1-4f77-ae3c-3ed2af783d06 · outbound
TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Audio-aware Query-enhanced Transformer for Audio-Visual Segmentation
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03102f43-292a-420b-836e-129cfb75311a · outbound
TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Exploiting transformation invariance and equivariance for self-supervised sound localisation
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation dea00248-ac9b-4fdf-a823-d51478db0ec6 · outbound
TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Image segmentation using text and image prompts
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation a8cd910e-e30e-4e4e-a3c8-d8d16a8f0f6b · outbound
TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization An eye for an ear: zero-shot audio description leveraging an image captioner with audio-visual token distribution matching
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 4eea906a-1d55-4d2a-9cb1-ffbf7d517a40 · outbound
TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization A closer look at weakly-supervised audio-visual source localization
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 0e90062f-60ed-447b-97a6-977b09e920cf · outbound
TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization A Concept-Based Explainability Framework for Large Multimodal Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aac9c43f-6969-48e9-89ad-4defc724afdc · outbound
TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Listen to interpret: Post-hoc interpretability for audio networks with nmf
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation b38c4eb1-38f8-459f-ba47-2a794b9aad4d · outbound
TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Guiding audio source separation by video object information
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 31bf4e73-ac6d-4ed4-b01a-842b831e4de3 · outbound
TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Marginnce: Robust sound localization with a negative margin
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 049cf943-4e67-4b62-b51c-603333264557 · outbound
TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Can clip help sound source localization? In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), January 2024
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation e1bd8688-db34-4e0d-b068-cd945dadc870 · outbound
TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Learning transferable visual models from natural language supervision
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation b96c824a-9c57-4100-ad82-ff70fedabe54 · outbound
TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Laion- 5b: An open large-scale dataset for training next generation image-text models
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation d2d773af-542d-4b50-8b11-ddbe30127d4e · outbound
TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Soft nonnegative matrix co-factorization
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ade26ffd-c0f6-48ad-9fb0-7d20d6860cfa · outbound
TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Learning to localize sound source in visual scenes
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ba10f421-b590-41a8-a293-dea8d5040411 · outbound
TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Learning to localize sound sources in visual scenes: Analysis and applications
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 724eaf77-0934-4d99-847f-4bbdb6051a6e · outbound
TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Learning sound localization better from semantically similar samples
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 6d4a439d-eaff-4ac3-b460-6904964ce2af · outbound
TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Sound source localization is all about cross-modal alignment
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 6e811ac2-eba3-4bb7-a1c1-1576ac0e8b8a · outbound
TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Increasing importance of joint analysis of audio and video in computer vision: A survey
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation bd27ecd2-f5e5-4248-a694-64ffe7e5c8c2 · outbound
TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Learning audio-visual source localization via false negative aware contrastive learning
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation affd8a84-9597-4234-8128-c7bf6f43aea3 · outbound
TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Audio-visual event localization in unconstrained videos
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation d5c4fbac-c93b-4f84-92c5-798058cb14da · outbound
TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Combining non-negative matrix factoriza- tion and deep neural networks for speech enhancement and automatic speech recognition
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation d38866b8-2cd0-4dce-a59b-2038c55374e8 · outbound
TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Document clustering based on non-negative matrix factorization
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 0bf2f4e7-82f8-4ff4-9c6c-21629b283753 · outbound
TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Coupled nonnegative matrix factorization unmixing for hyperspectral and multispectral data fusion
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation cea3375a-ad04-4c25-8af0-163da9acb579 · outbound
TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Matrix co-factorization on compressed sensing
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation fad3c349-078c-4638-9a83-0e7c9a13e8d8 · outbound
TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Convolutions die hard: Open-vocabulary segmentation with single frozen convolutional clip
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation d70059dc-7c51-4dee-8833-81b812226705 · outbound
TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Semantic understanding of scenes through the ade20k dataset
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 212c30af-bc19-468b-8a4b-0497ba2bbe45 · outbound
TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Audio-visual segmentation with semantics
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation eadfaf95-d7b5-4b70-a9d6-2ed3f29471da · outbound
TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Audio–visual segmentation
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
No inbound Pith citation observations are available.