Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:44:12.200891Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 0 inbound Pith citation observations for arXiv:2506.01558.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:44:12.200891Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
47 of 47 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 5f657d51-1ebb-49b7-a1c1-a31055a06513 · outbound
SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Self-calibrated clip for training-free open-vocabulary segmentation
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4106bdea-8ebf-42b7-b522-168cf3635064 · outbound
SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Unraveling in- stance associations: A closer look for audio-visual segmenta- tion
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28ddbb4c-1743-43ca-a927-991312d9307f · outbound
SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9dbd80ee-966b-44cf-9bd8-158540a1080f · outbound
SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Avsegformer: Audio-visual segmentation with trans- former
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 20ec1905-41b2-4500-bd1c-255311445b5a · outbound
SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Audio set: An ontology and human- labeled dataset for audio events
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 36ef3bfc-767c-4f91-9234-e25c73b26324 · outbound
SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Cnn archi- tectures for large-scale audio classification
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0bd5aab6-3c94-4e93-9d15-3be7573625d0 · outbound
SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Referitgame: Referring to objects in pho- tographs of natural scenes
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 112943e8-e328-4570-85bf-0f711cd5c1a1 · outbound
SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Segment any- thing
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44aa1544-d441-432f-a99f-8a2339efc465 · outbound
SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Lisa: Reasoning segmentation via large language model
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8e998dc-9bfe-4169-a895-9b7f0e2d6e10 · outbound
SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Learning to answer questions in dynamic audio-visual scenarios
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dd7a614f-769c-40c2-a826-1465021b808b · outbound
SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Robust referring video object segmentation with cyclic structural consensus
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68ac84f6-97a6-4d9e-9619-a4d26330f4de · outbound
SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Libero: Benchmarking knowl- edge transfer for lifelong robot learning
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b7696e5c-1259-4a90-a136-bc97a26f253b · outbound
SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes GRES: Gen- eralized referring expression segmentation
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation abedcc9d-ea78-471c-b193-3feea0248855 · outbound
SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Improved baselines with visual instruction tuning
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e108d5a6-4276-4322-8943-95722e94cc8f · outbound
SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Visual instruction tuning
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2c1f8d2-72ed-4b08-804b-9a28108aa5a3 · outbound
SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes RoBERTa: A Robustly Optimized BERT Pretraining Approach
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1580b59-e6a2-4b27-adb5-7382a9c7bf06 · outbound
SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Open-vocabulary segmentation with semantic-assisted calibration
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d3be685d-3edd-4cb2-a54e-cce2bfe266c4 · outbound
SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Universal segmentation at arbi- trary granularity with language instruction
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f94e348d-ca8a-47fc-ace8-4688e635702d · outbound
SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Decoupled weight de- cay regularization
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09e0cbf6-df4a-49e9-94f7-25215f04a837 · outbound
SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Soc: Semantic-assisted object cluster for referring video object segmentation
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 16f510cc-11dc-4c5a-a067-d6c1aef562ba · outbound
SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Mod- eling context between objects for referring expression un- derstanding
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b645e32b-6c58-4403-94c1-1563de57e8aa · outbound
SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes SAM 2: Segment Anything in Images and Videos
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 247f7548-6d8e-430a-9382-b699f0dee7a6 · outbound
SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Pixellm: Pixel reasoning with large multimodal model
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e51c9301-f233-417f-95a4-7afa4d1929a7 · outbound
SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3028eb4f-1c46-4812-92eb-c73560c00a93 · outbound
SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Efficient attention: Attention with lin- ear complexities
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5878c5f0-72c0-41ed-a211-55eac3e2713c · outbound
SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Roformer: Enhanced transformer with rotary position embedding
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc52a43d-4679-4847-ad56-7721d435448c · outbound
SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Auto- acd: A large-scale dataset for audio-language representation learning
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f31b91df-7213-47e0-aca7-bd32ddac1e79 · outbound
SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Prompting Segmentation with Sound Is Generalizable Audio-Visual Source Localizer
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64f558e1-147b-468f-878c-483ff6fa7bbe · outbound
SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Efficient remote sensing transformer for coastline detection with sentinel-2 satellite imagery
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 58ab5cbc-bd89-45e4-b97f-f47dc87a431e · outbound
SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Prompting segmentation with sound is gen- eralizable audio-visual source localizer
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2e111730-5219-4b15-8ebf-ad71dd6e9371 · outbound
SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Ref-avs: Refer and segment objects in audio-visual scenes
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 493dc4ad-5486-4df8-804b-0238b2ed74bc · outbound
SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Convolution meets trans- former: Efficient hybrid transformer for semantic segmenta- tion with very high resolution imagery
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5cb92975-e83f-4199-a288-d262b10ca69c · outbound
SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes IteRPrimE: Zero-shot Referring Image Segmentation with Iterative Grad-CAM Refinement and Primary Word Emphasis
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4138baf4-d4f3-423a-bb44-b3b7c22d4c7c · outbound
SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Onlinerefer: A simple online baseline for referring video object segmentation
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4aace26-db3c-4949-8b24-73b9f780121b · outbound
SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Language as queries for referring video object seg- mentation
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0882748c-073e-47fa-aa91-48331e4454a4 · outbound
SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Language as queries for referring video object seg- mentation
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4d8b0770-911d-4830-955e-49fbd5ba5748 · outbound
SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Gsva: Generalized segmentation via multimodal large language models
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d7575162-9125-4f72-a778-1645e4c6b5a8 · outbound
SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Efficientsam: Leveraged masked image pretraining for efficient segment anything
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 338baf62-16e1-439a-bd25-27965cef8bea · outbound
SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Avqa: A dataset for audio- visual question answering on videos
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9556304f-7048-4819-9f84-f36361c06f50 · outbound
SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Lavt: Language-aware vision transformer for referring image segmentation
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 16a50437-db4c-45f7-9e15-127705cd219f · outbound
SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Language- aware vision transformer for referring segmentation
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fd632782-eadc-4606-8d53-4ee658adf5c9 · outbound
SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Faster Segment Anything: Towards Lightweight SAM for Mobile Applications
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7f2f6e5-a0fd-4d05-91d9-d1a3c3424bb8 · outbound
SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d85d712-5c57-410d-8660-822dd4315a3c · outbound
SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Fast Segment Anything
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98aa6ac5-3a2e-4830-9d09-ed2199079c42 · outbound
SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Audio-visual segmentation
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3321ace5-26c3-49e0-8a45-7ab64cb4f444 · outbound
SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Audio-Visual Segmentation with Semantics
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 518e8f9d-885c-4021-83d6-3e19bb0a794f · outbound
SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Generalized decoding for pixel, image, and lan- guage
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.