Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T21:28:55.043436Z
Paper Citation Record · LEDGER
As of 19 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 1 inbound Pith citation observation for arXiv:2412.04434.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T21:28:55.043436Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T06:01:19.889564Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-16T06:01:20.333852Z
61 of 61 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 3e94c777-44a9-4294-8980-602f10de40f4 · outbound
Towards Real-Time Open-Vocabulary Video Instance Segmentation GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a29efce4-5d58-4c6f-989a-397d5e083c5f · outbound
Towards Real-Time Open-Vocabulary Video Instance Segmentation Tarvis: A unified approach for target-based video segmentation
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation c84a1983-93bf-4801-929e-80b38bbb6581 · outbound
Towards Real-Time Open-Vocabulary Video Instance Segmentation Burst: A benchmark for unifying object recognition, segmentation and tracking in video
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 1d621219-4170-4760-8369-45afe20c6487 · outbound
Towards Real-Time Open-Vocabulary Video Instance Segmentation Simple online and realtime tracking
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation fc2e99c9-e58c-4ee5-8419-5c6e734f8232 · outbound
Towards Real-Time Open-Vocabulary Video Instance Segmentation Efficientvit: Lightweight multi-scale attention for high- resolution dense prediction
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 7c28fd84-81e1-4892-a348-4f2a5f99f6b9 · outbound
Towards Real-Time Open-Vocabulary Video Instance Segmentation End-to- end object detection with transformers
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85cdf484-b9d5-4cc2-a366-7b6e369d4e48 · outbound
Towards Real-Time Open-Vocabulary Video Instance Segmentation Xmem: Long- term video object segmentation with an atkinson-shiffrin memory model
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation f7fa137d-5f9a-4511-977b-38c42d7d3f10 · outbound
Towards Real-Time Open-Vocabulary Video Instance Segmentation Yolo-world: Real-time open-vocabulary object detection
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation b4e42626-faaa-4785-bcc7-01cf2ff2a6ae · outbound
Towards Real-Time Open-Vocabulary Video Instance Segmentation Xception: Deep learning with depthwise separable convolutions
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7cc4a43-ba9d-4841-95e5-ed1d385f2b77 · outbound
Towards Real-Time Open-Vocabulary Video Instance Segmentation Tao: A large-scale benchmark for tracking any object
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 2e78778b-0129-4ca9-a24f-9a61a6957124 · outbound
Towards Real-Time Open-Vocabulary Video Instance Segmentation BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 781acbd7-3fdd-4ba4-8e52-fbb51a83a4de · outbound
Towards Real-Time Open-Vocabulary Video Instance Segmentation EVA-02: A Visual Representation for Neon Genesis
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ddc3d53-07c3-4754-98ef-38deda2fc02b · outbound
Towards Real-Time Open-Vocabulary Video Instance Segmentation Ultralyt- ics yolov8
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation a1ff6b3e-ff4a-44e5-a23e-c0318f47520b · outbound
Towards Real-Time Open-Vocabulary Video Instance Segmentation Lvis: A dataset for large vocabulary instance segmentation
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation f595326f-757b-458a-8f42-e9b126249d96 · outbound
Towards Real-Time Open-Vocabulary Video Instance Segmentation Masked autoencoders are scalable vision learners
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd8713f0-25dd-4c26-85a0-1df7c81f1cd2 · outbound
Towards Real-Time Open-Vocabulary Video Instance Segmentation Deep residual learning for image recognition
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a656057d-3610-46ef-b34c-9b7f9707ea22 · outbound
Towards Real-Time Open-Vocabulary Video Instance Segmentation Scaling Laws for Neural Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b71396c-6a87-4142-a623-c2e7337fbf99 · outbound
Towards Real-Time Open-Vocabulary Video Instance Segmentation Segment any- thing
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation b75216db-7462-48ae-890b-2a3efddfe9a2 · outbound
Towards Real-Time Open-Vocabulary Video Instance Segmentation Visual genome: Connecting language and vision using crowdsourced dense image annotations
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 543a8b65-88e5-42a7-a03a-de63292f208e · outbound
Towards Real-Time Open-Vocabulary Video Instance Segmentation The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 589522ef-17e3-4bc1-aa10-f092658650f7 · outbound
Towards Real-Time Open-Vocabulary Video Instance Segmentation Mask dino: Towards a unified transformer-based framework for object detection and segmentation
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 159d557d-cd6e-4e99-bf8f-5923a404fdac · outbound
Towards Real-Time Open-Vocabulary Video Instance Segmentation Grounded language-image pre-training
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fec45a30-ffa7-442c-870d-76937071e0ac · outbound
Towards Real-Time Open-Vocabulary Video Instance Segmentation Microsoft coco: Common objects in context
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c459ebd0-5c65-4ad1-a996-b5df0b545310 · outbound
Towards Real-Time Open-Vocabulary Video Instance Segmentation Decoupled weight decay regularization
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4421389a-56da-44ab-a5f0-be0f8efefa08 · outbound
Towards Real-Time Open-Vocabulary Video Instance Segmentation Hota: A higher order metric for evaluating multi-object tracking
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 07451f5e-d3f5-4529-a935-28dc8f951fcc · outbound
Towards Real-Time Open-Vocabulary Video Instance Segmentation Simple open-vocabulary object detection with vision transformers
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation a58e9310-84e4-4cba-bc00-795b1ce87eab · outbound
Towards Real-Time Open-Vocabulary Video Instance Segmentation Mod- eling context between objects for referring expression under- standing
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 14687239-aa48-4659-8102-819bcc475362 · outbound
Towards Real-Time Open-Vocabulary Video Instance Segmentation Unresolved cited work
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 499ed06f-3d6b-4bd7-a01a-a205707cc40f · outbound
Towards Real-Time Open-Vocabulary Video Instance Segmentation Kosmos-2: Grounding Multimodal Large Language Models to the World
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90f47f70-3228-4574-8ddf-a352cd0132ca · outbound
Towards Real-Time Open-Vocabulary Video Instance Segmentation Occluded video instance segmentation: A bench- mark
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation c9c2d132-4a86-4125-ac12-fcdaee5589c5 · outbound
Towards Real-Time Open-Vocabulary Video Instance Segmentation Learning transferable visual models from natural language supervi- sion
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e7912d9-9ca6-486e-b049-9e19c9cb377e · outbound
Towards Real-Time Open-Vocabulary Video Instance Segmentation SAM 2: Segment Anything in Images and Videos
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62526b83-200d-4dad-9747-52774a27aefa · outbound
Towards Real-Time Open-Vocabulary Video Instance Segmentation Grounding DINO 1.5: Advance the "Edge" of Open-Set Object Detection
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3be39989-6f43-4632-94ad-1e02bf43a07d · outbound
Towards Real-Time Open-Vocabulary Video Instance Segmentation Urvos: Unified referring video object segmentation network with a large-scale benchmark
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3982d794-fca4-4bfc-887f-5bbdee313914 · outbound
Towards Real-Time Open-Vocabulary Video Instance Segmentation Objects365: A large-scale, high-quality dataset for object detection
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68076fcd-b3a6-445f-aee0-b07a2c02b502 · outbound
Towards Real-Time Open-Vocabulary Video Instance Segmentation Aligning and prompting everything all at once for univer- sal visual perception
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ddeb0bf-fb3c-4bac-89e6-3ab96e5ef369 · outbound
Towards Real-Time Open-Vocabulary Video Instance Segmentation Mobile- clip: Fast image-text models through multi-modal reinforced training
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2468e4d-2f0f-48de-a663-ccbcc89386c2 · outbound
Towards Real-Time Open-Vocabulary Video Instance Segmentation Towards open-vocabulary video instance segmentation
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 3bf861b6-1381-45e4-a839-585c3026939b · outbound
Towards Real-Time Open-Vocabulary Video Instance Segmentation Unidentified video objects: A benchmark for dense, open- world segmentation
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation eecf87ff-dbfe-443a-8e45-b59ece3621c8 · outbound
Towards Real-Time Open-Vocabulary Video Instance Segmentation General object foundation model for images and videos at scale
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 708f9edf-2dc7-4dc8-8c98-f5a173fb2d5a · outbound
Towards Real-Time Open-Vocabulary Video Instance Segmentation Tinyclip: Clip dis- tillation via affinity mimicking and weight inheritance
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62984b1b-28ce-4406-9ae9-7a45a2649f45 · outbound
Towards Real-Time Open-Vocabulary Video Instance Segmentation Efficientsam: Leveraged masked image pretraining for efficient segment anything
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f0406ce-7348-4173-ba9a-7d2466ac8fe5 · outbound
Towards Real-Time Open-Vocabulary Video Instance Segmentation Unresolved cited work
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 7007955f-20c0-49f0-88bd-1668fed77baa · outbound
Towards Real-Time Open-Vocabulary Video Instance Segmentation Towards grand unification of object tracking
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation db941ade-d092-48a5-af24-cf9f1753c7e9 · outbound
Towards Real-Time Open-Vocabulary Video Instance Segmentation Universal instance perception as object discovery and retrieval
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 34639510-a8e1-4a42-acca-2838e1b48694 · outbound
Towards Real-Time Open-Vocabulary Video Instance Segmentation Video instance seg- mentation
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 3ad41a81-dc22-4820-ba08-b47c73db1561 · outbound
Towards Real-Time Open-Vocabulary Video Instance Segmentation Detclip: Dictionary-enriched visual-concept paralleled pre- training for open-world detection
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 0f5cccf9-b0e7-44e3-a958-50c6770a29f3 · outbound
Towards Real-Time Open-Vocabulary Video Instance Segmentation Modeling context in referring expres- sions
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation cd55fee2-5ded-4457-b71d-d187a4268cd3 · outbound
Towards Real-Time Open-Vocabulary Video Instance Segmentation Open-vocabulary object detection using captions
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation b9459cb0-f3fd-49a2-9f1a-e75ad850283d · outbound
Towards Real-Time Open-Vocabulary Video Instance Segmentation Faster Segment Anything: Towards Lightweight SAM for Mobile Applications
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 681f3d4a-f930-4ef8-8ae6-be68c7466da3 · outbound
Towards Real-Time Open-Vocabulary Video Instance Segmentation Dino: Detr with improved denoising anchor boxes for end-to-end object detection
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation acfc595e-9e74-477c-9e65-b6a25a57e134 · outbound
Towards Real-Time Open-Vocabulary Video Instance Segmentation A simple framework for open-vocabulary segmentation and detection
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 354ff347-eaf0-4477-9e7c-81c8eb486e71 · outbound
Towards Real-Time Open-Vocabulary Video Instance Segmentation Mobileinst: Video in- stance segmentation on the mobile
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation fd06e973-41b9-4ee3-9f55-d8689cac4e7e · outbound
Towards Real-Time Open-Vocabulary Video Instance Segmentation Bridging the gap between anchor-based and anchor-free detection via adaptive training sample selection
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 85b4331f-947e-4f74-84bc-55d042a08cd9 · outbound
Towards Real-Time Open-Vocabulary Video Instance Segmentation EfficientViT-SAM: Accelerated Segment Anything Model Without Accuracy Loss
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc8f0053-2b15-458b-acad-3b21d14fae45 · outbound
Towards Real-Time Open-Vocabulary Video Instance Segmentation Real-time Transformer-based Open-Vocabulary Detection with Efficient Fusion Head
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 437f8fc6-39ae-42a8-97d9-1023dfa6b882 · outbound
Towards Real-Time Open-Vocabulary Video Instance Segmentation Detrs beat yolos on real-time object detection
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8db961c-cbfa-4413-90f4-f58632584898 · outbound
Towards Real-Time Open-Vocabulary Video Instance Segmentation Detecting twenty-thousand classes using image-level supervision
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation d1a1aa1c-08fb-4fac-a53f-8ace31c28fbc · outbound
Towards Real-Time Open-Vocabulary Video Instance Segmentation Deformable detr: Deformable transformers for end-to-end object detection
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 987506cd-d1fd-4508-9785-c5d1b9180110 · outbound
Towards Real-Time Open-Vocabulary Video Instance Segmentation Generalized decoding for pixel, image, and lan- guage
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cac69651-d2c3-48db-83f1-c4c4b25cd614 · outbound
Towards Real-Time Open-Vocabulary Video Instance Segmentation Unresolved cited work
Reference 368
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ba96b6eb-c730-4041-940b-ce05dfa59553 · inbound
OpenFusion++: An Open-vocabulary Real-time Scene Understanding System Towards Real-Time Open-Vocabulary Video Instance Segmentation
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.