Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T04:42:30.341132Z
Paper Citation Record · LEDGER
As of 15 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 2 inbound Pith citation observations for arXiv:2412.01136.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T04:42:30.341132Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T12:08:16.065856Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T05:09:31.997736Z
29 of 29 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 72b3f13b-d469-49cc-af3f-03b302024ed0 · outbound
Referring Video Object Segmentation via Language-aligned Track Selection GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7aedcaaa-2530-4ba2-8bfd-3eca2b38a316 · outbound
Referring Video Object Segmentation via Language-aligned Track Selection End-to-end referring video object segmentation with mul- timodal transformers
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 2e5a6799-73e8-4236-97bb-42b1e1f4627f · outbound
Referring Video Object Segmentation via Language-aligned Track Selection Vision-language transformer and query generation for refer- ring segmentation
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 021f977b-b978-4745-96c1-a0bf0c45470d · outbound
Referring Video Object Segmentation via Language-aligned Track Selection Mevis: A large-scale benchmark for video segmentation with motion expressions
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation e8af9f4e-02b6-4310-813e-d1c9fe1ff8ca · outbound
Referring Video Object Segmentation via Language-aligned Track Selection Language-bridged spatial-temporal interaction for referring video object segmentation
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation ee750bb6-3a96-4573-bdf1-84b2f0fed86c · outbound
Referring Video Object Segmentation via Language-aligned Track Selection The pascal visual object classes (voc) challenge
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 2e25aaf5-ba70-495e-9386-7bcf1ceb2269 · outbound
Referring Video Object Segmentation via Language-aligned Track Selection Actor and action video segmentation from a sen- tence
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation ce2302a8-491a-4ef8-8d50-72b4a0c643a7 · outbound
Referring Video Object Segmentation via Language-aligned Track Selection Html: Hybrid temporal-scale mul- timodal learning framework for referring video object seg- mentation
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 5ba59491-4c49-46a1-8788-8e9f46a22f65 · outbound
Referring Video Object Segmentation via Language-aligned Track Selection Decoupling static and hier- archical motion perception for referring video segmentation
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 72a94c27-7bf6-4408-8487-2d83cf546cdf · outbound
Referring Video Object Segmentation via Language-aligned Track Selection Benchmarking Neural Network Robustness to Common Corruptions and Perturbations
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41c97bf8-dcb4-4d16-ac06-e2152538a441 · outbound
Referring Video Object Segmentation via Language-aligned Track Selection Unleashing the Temporal-Spatial Reasoning Capacity of GPT for Training-Free Audio and Language Referenced Video Object Segmentation
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b34fb5a-31c7-4a10-a5f1-1511cd520211 · outbound
Referring Video Object Segmentation via Language-aligned Track Selection Scaling up visual and vision-language representa- tion learning with noisy text supervision
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 19fee968-32f3-468c-a432-a56361293f7d · outbound
Referring Video Object Segmentation via Language-aligned Track Selection Video object segmentation with language referring expressions
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 5c37430b-569c-4528-b8ca-035f1ecc81c4 · outbound
Referring Video Object Segmentation via Language-aligned Track Selection Video object segmentation with language referring expressions
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a9ab9048-f549-47f7-aca3-3a4c9a07bc0a · outbound
Referring Video Object Segmentation via Language-aligned Track Selection Segment any- thing
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2102bee9-3c31-493a-a017-c3540bfab65e · outbound
Referring Video Object Segmentation via Language-aligned Track Selection Referring image seg- mentation via recurrent refinement networks
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation c5ebdb6e-1be1-4721-acbe-1f04f1af376b · outbound
Referring Video Object Segmentation via Language-aligned Track Selection Robust referring video object segmentation with cyclic structural consensus
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation c5366fd6-a2d9-4667-8ec4-f8cfbac51261 · outbound
Referring Video Object Segmentation via Language-aligned Track Selection RefSAM: Efficiently Adapting Segmenting Anything Model for Referring Video Object Segmentation
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c146b9a3-2a06-4a09-b7ce-677443afd363 · outbound
Referring Video Object Segmentation via Language-aligned Track Selection Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70c4a044-c2b8-473a-85f6-7cd0364320c3 · outbound
Referring Video Object Segmentation via Language-aligned Track Selection RoBERTa: A Robustly Optimized BERT Pretraining Approach
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af574789-eae5-4493-ab62-1d0c2fe1bb50 · outbound
Referring Video Object Segmentation via Language-aligned Track Selection Temporally consistent referring video object segmentation with hybrid memory
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 74a8702f-3397-492b-9705-167cb268bcbe · outbound
Referring Video Object Segmentation via Language-aligned Track Selection Efficient non- maximum suppression
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 6f332e0b-e8e8-422c-8616-623c1c5c75c7 · outbound
Referring Video Object Segmentation via Language-aligned Track Selection A benchmark dataset and evaluation methodology for video object segmentation
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation deaf41a7-9c1d-4c4d-90c4-d639b97bbee4 · outbound
Referring Video Object Segmentation via Language-aligned Track Selection SAM 2: Segment Anything in Images and Videos
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a8dcfc2-c299-46a1-9f6d-86faa7caf827 · outbound
Referring Video Object Segmentation via Language-aligned Track Selection Urvos: Unified referring video object segmentation network with a large-scale benchmark
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 1617197a-0fdd-4668-afc2-5ef4222d579f · outbound
Referring Video Object Segmentation via Language-aligned Track Selection Attention is all you need
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77670a91-7112-4f11-89a8-e3b7ef0d0497 · outbound
Referring Video Object Segmentation via Language-aligned Track Selection Data-efficient mul- timodal fusion on a single gpu
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 4a000abf-1e12-427a-b2a3-5943585bfade · outbound
Referring Video Object Segmentation via Language-aligned Track Selection Language as queries for referring video object segmen- tation
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation c50724b1-ee13-4332-949a-76b6a6232ef8 · outbound
Referring Video Object Segmentation via Language-aligned Track Selection Going right
Reference 2022
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation ec8a01c8-6981-4e8b-8e64-4b8fa22f65bb · inbound
MOVE: Motion-Guided Few-Shot Video Object Segmentation Referring Video Object Segmentation via Language-aligned Track Selection
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cff9ce97-b9ba-401e-82fe-fa7684435c61 · inbound
Unleashing Hierarchical Reasoning: An LLM-Driven Framework for Training-Free Referring Video Object Segmentation Referring Video Object Segmentation via Language-aligned Track Selection
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.