Pith. sign in

Paper Citation Record · LEDGER

Towards Real-Time Open-Vocabulary Video Instance Segmentation

As of 19 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 1 inbound Pith citation observation for arXiv:2412.04434.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.04434 v1

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T21:28:55.043436Z

measured 62 of 62 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T06:01:19.889564Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-16T06:01:20.333852Z

Reference resolution

61 of 61 outbound references displayed

  • verified exact0
  • verified fuzzy31
  • unresolved29
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3e94c777-44a9-4294-8980-602f10de40f4 · outbound

This paper cites GPT-4 Technical Report.

Towards Real-Time Open-Vocabulary Video Instance Segmentation GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:54.761939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:54.761939Z digest=sha256:6105cb486d5f902d9f1b1745e14c359b64437741426facc4479fe459f7c858f6

Observation a29efce4-5d58-4c6f-989a-397d5e083c5f · outbound

This paper cites Tarvis: A unified approach for target-based video segmentation.

Towards Real-Time Open-Vocabulary Video Instance Segmentation Tarvis: A unified approach for target-based video segmentation

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:56.286659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:28:54.768390Z digest=sha256:593e2e900548d74c3de60d06b71a39cb93ad4aa60b9c2f85b81619dd9bf9e089

Observation c84a1983-93bf-4801-929e-80b38bbb6581 · outbound

This paper cites Burst: A benchmark for unifying object recognition, segmentation and tracking in video.

Towards Real-Time Open-Vocabulary Video Instance Segmentation Burst: A benchmark for unifying object recognition, segmentation and tracking in video

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:56.269868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:28:54.773331Z digest=sha256:b66b19205b649fee7b9f39e9731a59533414a2e8e37f2105c551419e9ed8b7af

Observation 1d621219-4170-4760-8369-45afe20c6487 · outbound

This paper cites Simple online and realtime tracking.

Towards Real-Time Open-Vocabulary Video Instance Segmentation Simple online and realtime tracking

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:56.253565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:28:54.778521Z digest=sha256:345a76e78bec7f03048a839e23a67c9f919e9e8ba9d4d74e0e2f034fc429052e

Observation fc2e99c9-e58c-4ee5-8419-5c6e734f8232 · outbound

This paper cites Efficientvit: Lightweight multi-scale attention for high- resolution dense prediction.

Towards Real-Time Open-Vocabulary Video Instance Segmentation Efficientvit: Lightweight multi-scale attention for high- resolution dense prediction

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:56.236364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:28:54.783547Z digest=sha256:0de10fffb280e8d9815343a5d86427531a2a598e94a5383aec030c5c10502a45

Observation 7c28fd84-81e1-4892-a348-4f2a5f99f6b9 · outbound

This paper cites End-to- end object detection with transformers.

Towards Real-Time Open-Vocabulary Video Instance Segmentation End-to- end object detection with transformers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:54.789052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:54.789052Z digest=sha256:b607d3f0acd657fd169cd3a63db7af7cc802309b372c20252cf580f7655167ec

Observation 85cdf484-b9d5-4cc2-a366-7b6e369d4e48 · outbound

This paper cites Xmem: Long- term video object segmentation with an atkinson-shiffrin memory model.

Towards Real-Time Open-Vocabulary Video Instance Segmentation Xmem: Long- term video object segmentation with an atkinson-shiffrin memory model

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:56.208807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:28:54.794234Z digest=sha256:f2ba53d7e3a4745dbb87d6e907d1eb980fe968679e4bea0a4918a491f62a31a3

Observation f7fa137d-5f9a-4511-977b-38c42d7d3f10 · outbound

This paper cites Yolo-world: Real-time open-vocabulary object detection.

Towards Real-Time Open-Vocabulary Video Instance Segmentation Yolo-world: Real-time open-vocabulary object detection

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:56.192660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:28:54.799665Z digest=sha256:ee1aab9557d2d6d05bbd9ef68dca80b5123ebb090ca7fde211266efc72457dd1

Observation b4e42626-faaa-4785-bcc7-01cf2ff2a6ae · outbound

This paper cites Xception: Deep learning with depthwise separable convolutions.

Towards Real-Time Open-Vocabulary Video Instance Segmentation Xception: Deep learning with depthwise separable convolutions

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:54.804105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:54.804105Z digest=sha256:5b0726febf3f5387b145fda0d60d3e5fe552fad1182d7ae870cb93a9de8a1f66

Observation a7cc4a43-ba9d-4841-95e5-ed1d385f2b77 · outbound

This paper cites Tao: A large-scale benchmark for tracking any object.

Towards Real-Time Open-Vocabulary Video Instance Segmentation Tao: A large-scale benchmark for tracking any object

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:56.168140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:28:54.808717Z digest=sha256:e0231868805f58bedb3dc6731d314e9d2d859f2b628ca65a61ac5bdd1bea746b

Observation 2e78778b-0129-4ca9-a24f-9a61a6957124 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Towards Real-Time Open-Vocabulary Video Instance Segmentation BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:54.813272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:54.813272Z digest=sha256:2df45433cc0238e001e5ab36a4ddde74e18945203f83b5b22362273c52faac23

Observation 781acbd7-3fdd-4ba4-8e52-fbb51a83a4de · outbound

This paper cites EVA-02: A Visual Representation for Neon Genesis.

Towards Real-Time Open-Vocabulary Video Instance Segmentation EVA-02: A Visual Representation for Neon Genesis

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:54.818470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:54.818470Z digest=sha256:08c2d2c5a5391af2d30ad43afca18af4e4ccac88c3e74594c0c22342d63fa981

Observation 8ddc3d53-07c3-4754-98ef-38deda2fc02b · outbound

This paper cites Ultralyt- ics yolov8.

Towards Real-Time Open-Vocabulary Video Instance Segmentation Ultralyt- ics yolov8

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:56.153642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:28:54.823357Z digest=sha256:6c44ba9651d62654ce86c4f1d376fd8431ee9ab173b9f582366915afaf2b7bb8

Observation a1ff6b3e-ff4a-44e5-a23e-c0318f47520b · outbound

This paper cites Lvis: A dataset for large vocabulary instance segmentation.

Towards Real-Time Open-Vocabulary Video Instance Segmentation Lvis: A dataset for large vocabulary instance segmentation

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:56.139497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:28:54.827789Z digest=sha256:60def874c5e836574130676d912997d2def7d17a9cd9a8779397c4983ed1733b

Observation f595326f-757b-458a-8f42-e9b126249d96 · outbound

This paper cites Masked autoencoders are scalable vision learners.

Towards Real-Time Open-Vocabulary Video Instance Segmentation Masked autoencoders are scalable vision learners

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:54.833305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:54.833305Z digest=sha256:e9a4cd5c87e15b24270bbc4d78e5dd5133b013562b5da5dbbdaea66af0e3830c

Observation cd8713f0-25dd-4c26-85a0-1df7c81f1cd2 · outbound

This paper cites Deep residual learning for image recognition.

Towards Real-Time Open-Vocabulary Video Instance Segmentation Deep residual learning for image recognition

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:54.837996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:54.837996Z digest=sha256:1576d9060718999c3754fea029f968ba21de4935549768ff393049662ce57e75

Observation a656057d-3610-46ef-b34c-9b7f9707ea22 · outbound

This paper cites Scaling Laws for Neural Language Models.

Towards Real-Time Open-Vocabulary Video Instance Segmentation Scaling Laws for Neural Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:54.842846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:54.842846Z digest=sha256:782dba51a756daeb89b00532a3ef4de4e8c5a2cdfe0e76fa54edc50b78a56ac8

Observation 1b71396c-6a87-4142-a623-c2e7337fbf99 · outbound

This paper cites Segment any- thing.

Towards Real-Time Open-Vocabulary Video Instance Segmentation Segment any- thing

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:56.105035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:28:54.848250Z digest=sha256:21b54c0b4635dd9e25636248d530719886c6d8d130630271fc02bcb4e2314842

Observation b75216db-7462-48ae-890b-2a3efddfe9a2 · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense image annotations.

Towards Real-Time Open-Vocabulary Video Instance Segmentation Visual genome: Connecting language and vision using crowdsourced dense image annotations

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:54.853078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:54.853078Z digest=sha256:70b7079001c8f7457d364c9ab5e595c69f2da7ddc8b2d1b228dd87b3d42dd478

Observation 543a8b65-88e5-42a7-a03a-de63292f208e · outbound

This paper cites The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale.

Towards Real-Time Open-Vocabulary Video Instance Segmentation The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:56.080782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:28:54.858059Z digest=sha256:8f04ee9fd8794462bc4cacfb723bb28102393e2551e2287277b943fec58c5a23

Observation 589522ef-17e3-4bc1-aa10-f092658650f7 · outbound

This paper cites Mask dino: Towards a unified transformer-based framework for object detection and segmentation.

Towards Real-Time Open-Vocabulary Video Instance Segmentation Mask dino: Towards a unified transformer-based framework for object detection and segmentation

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:56.065644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:28:54.862894Z digest=sha256:691e183cde14050e9e7d75db0463ab3a0c84e18d84217b9ee0ddfdeaed910271

Observation 159d557d-cd6e-4e99-bf8f-5923a404fdac · outbound

This paper cites Grounded language-image pre-training.

Towards Real-Time Open-Vocabulary Video Instance Segmentation Grounded language-image pre-training

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:54.867565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:54.867565Z digest=sha256:beb4e64e396d14a2ee49c90776bc3d4acb0d17372d7214841dcd44d674c6d453

Observation fec45a30-ffa7-442c-870d-76937071e0ac · outbound

This paper cites Microsoft coco: Common objects in context.

Towards Real-Time Open-Vocabulary Video Instance Segmentation Microsoft coco: Common objects in context

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:54.872535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:54.872535Z digest=sha256:28b8b31c3787d3cd815b757aa00916122a0706b91bc3287ac91be7e3a734923a

Observation c459ebd0-5c65-4ad1-a996-b5df0b545310 · outbound

This paper cites Decoupled weight decay regularization.

Towards Real-Time Open-Vocabulary Video Instance Segmentation Decoupled weight decay regularization

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:54.876618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:54.876618Z digest=sha256:73047927cf597d43676eed5f8fac830c94aa43db535cd34b277f47a7b16937cf

Observation 4421389a-56da-44ab-a5f0-be0f8efefa08 · outbound

This paper cites Hota: A higher order metric for evaluating multi-object tracking.

Towards Real-Time Open-Vocabulary Video Instance Segmentation Hota: A higher order metric for evaluating multi-object tracking

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:56.016861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:28:54.880807Z digest=sha256:240c74bd7efd177a1923b59161ce97c543c6e00b93bf227d4b0b092c6b5d891e

Observation 07451f5e-d3f5-4529-a935-28dc8f951fcc · outbound

This paper cites Simple open-vocabulary object detection with vision transformers.

Towards Real-Time Open-Vocabulary Video Instance Segmentation Simple open-vocabulary object detection with vision transformers

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:55.999831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:28:54.884963Z digest=sha256:e1e95e2fc742914dd5526a8f635d29b69891dd5338694fa49d059ba9b90c2077

Observation a58e9310-84e4-4cba-bc00-795b1ce87eab · outbound

This paper cites Mod- eling context between objects for referring expression under- standing.

Towards Real-Time Open-Vocabulary Video Instance Segmentation Mod- eling context between objects for referring expression under- standing

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:55.954424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:28:54.889282Z digest=sha256:0638970148f040c45cb3103e0810d8678c942a88e4314eacf441e773e0d1deb1

Observation 14687239-aa48-4659-8102-819bcc475362 · outbound

This paper cites an unresolved cited work.

Towards Real-Time Open-Vocabulary Video Instance Segmentation Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-11T21:28:55.895004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:28:54.893659Z digest=sha256:5b769379a1c3fe3bc0df2db5d4aaeb743f4e32c4d789c89d74a942a922577d41

Observation 499ed06f-3d6b-4bd7-a01a-a205707cc40f · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

Towards Real-Time Open-Vocabulary Video Instance Segmentation Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:54.897763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:54.897763Z digest=sha256:852b10fae5fe19b6f0f6467e8cfc0dd57bc23ee8a06e13b977ac43fdd43a8bd9

Observation 90f47f70-3228-4574-8ddf-a352cd0132ca · outbound

This paper cites Occluded video instance segmentation: A bench- mark.

Towards Real-Time Open-Vocabulary Video Instance Segmentation Occluded video instance segmentation: A bench- mark

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:55.838610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:28:54.902315Z digest=sha256:0f9ac15bac9090b601ba48d631f734059ecc74221fd052e586ba69d20f7c893f

Observation c9c2d132-4a86-4125-ac12-fcdaee5589c5 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Towards Real-Time Open-Vocabulary Video Instance Segmentation Learning transferable visual models from natural language supervi- sion

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:54.907132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:54.907132Z digest=sha256:6bd4b8defd79b877544f309e4abbd167addfe1b143e5a5c910d0491e0dbcba99

Observation 3e7912d9-9ca6-486e-b049-9e19c9cb377e · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

Towards Real-Time Open-Vocabulary Video Instance Segmentation SAM 2: Segment Anything in Images and Videos

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:54.911708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:54.911708Z digest=sha256:946f6f6e910d0b96a7292949705b08b7cdfe7419db7334675ea8415c1cfea8e5

Observation 62526b83-200d-4dad-9747-52774a27aefa · outbound

This paper cites Grounding DINO 1.5: Advance the "Edge" of Open-Set Object Detection.

Towards Real-Time Open-Vocabulary Video Instance Segmentation Grounding DINO 1.5: Advance the "Edge" of Open-Set Object Detection

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:54.916967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:54.916967Z digest=sha256:e2519a4ad680ce29b30e0860f07b22cae70ffbcf45c315420dc67faa2cb45a9e

Observation 3be39989-6f43-4632-94ad-1e02bf43a07d · outbound

This paper cites Urvos: Unified referring video object segmentation network with a large-scale benchmark.

Towards Real-Time Open-Vocabulary Video Instance Segmentation Urvos: Unified referring video object segmentation network with a large-scale benchmark

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:54.921908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:54.921908Z digest=sha256:789f56b669d05cb6316b7dea5762b1b3ac3e0ded61d2e3c08efbd4f7bcda98ea

Observation 3982d794-fca4-4bfc-887f-5bbdee313914 · outbound

This paper cites Objects365: A large-scale, high-quality dataset for object detection.

Towards Real-Time Open-Vocabulary Video Instance Segmentation Objects365: A large-scale, high-quality dataset for object detection

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:54.926722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:54.926722Z digest=sha256:acd2a32b4ee6392b24bdd989eb50915133bf0c8f65edf53f4e708ed4d671b3c1

Observation 68076fcd-b3a6-445f-aee0-b07a2c02b502 · outbound

This paper cites Aligning and prompting everything all at once for univer- sal visual perception.

Towards Real-Time Open-Vocabulary Video Instance Segmentation Aligning and prompting everything all at once for univer- sal visual perception

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:54.931766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:54.931766Z digest=sha256:ff9f1443106240fb23fbfd49ee9b0f422ec42477398c615a1d49d7b0d68dd55b

Observation 1ddeb0bf-fb3c-4bac-89e6-3ab96e5ef369 · outbound

This paper cites Mobile- clip: Fast image-text models through multi-modal reinforced training.

Towards Real-Time Open-Vocabulary Video Instance Segmentation Mobile- clip: Fast image-text models through multi-modal reinforced training

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:54.936305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:54.936305Z digest=sha256:9611cf7f2985a07c825356df798219e1699cce06fe64a5386912a242a2befd1a

Observation b2468e4d-2f0f-48de-a663-ccbcc89386c2 · outbound

This paper cites Towards open-vocabulary video instance segmentation.

Towards Real-Time Open-Vocabulary Video Instance Segmentation Towards open-vocabulary video instance segmentation

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:55.769769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:28:54.940955Z digest=sha256:008c82525c7f34f5d6c160c20bc3fc230b28334919dee2f6ed0e8ba99e98ab6f

Observation 3bf861b6-1381-45e4-a839-585c3026939b · outbound

This paper cites Unidentified video objects: A benchmark for dense, open- world segmentation.

Towards Real-Time Open-Vocabulary Video Instance Segmentation Unidentified video objects: A benchmark for dense, open- world segmentation

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:55.754142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:28:54.945339Z digest=sha256:0042e7bc3dcb13fcd34bd7f10438e48b27ca167f8ce882ef7be2f623b210872b

Observation eecf87ff-dbfe-443a-8e45-b59ece3621c8 · outbound

This paper cites General object foundation model for images and videos at scale.

Towards Real-Time Open-Vocabulary Video Instance Segmentation General object foundation model for images and videos at scale

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:55.738618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:28:54.949698Z digest=sha256:45e605b2e429f07745a8ba1c2b29de2caf53e2db785970f492de356b6fbf31f4

Observation 708f9edf-2dc7-4dc8-8c98-f5a173fb2d5a · outbound

This paper cites Tinyclip: Clip dis- tillation via affinity mimicking and weight inheritance.

Towards Real-Time Open-Vocabulary Video Instance Segmentation Tinyclip: Clip dis- tillation via affinity mimicking and weight inheritance

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:54.954002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:54.954002Z digest=sha256:b5ff34de562a9a80340e4aee6addc7f28d3cfda155d074b19bcc4b410f6939ff

Observation 62984b1b-28ce-4406-9ae9-7a45a2649f45 · outbound

This paper cites Efficientsam: Leveraged masked image pretraining for efficient segment anything.

Towards Real-Time Open-Vocabulary Video Instance Segmentation Efficientsam: Leveraged masked image pretraining for efficient segment anything

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:54.958007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:54.958007Z digest=sha256:dfaf26a2c3d54d6be36d72e64823c40dca8d20a3baa23c55a5cc144c0f3867a3

Observation 6f0406ce-7348-4173-ba9a-7d2466ac8fe5 · outbound

This paper cites an unresolved cited work.

Towards Real-Time Open-Vocabulary Video Instance Segmentation Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-11T21:28:55.700465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:28:54.962201Z digest=sha256:acae0c6a4d38191875f5519fe10d9d3e3d6b3d90e8e8f659953ff87a595e5658

Observation 7007955f-20c0-49f0-88bd-1668fed77baa · outbound

This paper cites Towards grand unification of object tracking.

Towards Real-Time Open-Vocabulary Video Instance Segmentation Towards grand unification of object tracking

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:55.685614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:28:54.966760Z digest=sha256:0def9a5bbed99b3952672a510c90002abea53cd56fbc78d913f62137b4c824b5

Observation db941ade-d092-48a5-af24-cf9f1753c7e9 · outbound

This paper cites Universal instance perception as object discovery and retrieval.

Towards Real-Time Open-Vocabulary Video Instance Segmentation Universal instance perception as object discovery and retrieval

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:55.670114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:28:54.971168Z digest=sha256:c27aedf92db1f69d1645ab11d8805699e7ea65948d6b2422ac6f807296aa0a20

Observation 34639510-a8e1-4a42-acca-2838e1b48694 · outbound

This paper cites Video instance seg- mentation.

Towards Real-Time Open-Vocabulary Video Instance Segmentation Video instance seg- mentation

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:55.653838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:28:54.975232Z digest=sha256:2ef62e1f8fb87d28373bf79abae664b80631e17b95ed8e2152ff187c7594abc7

Observation 3ad41a81-dc22-4820-ba08-b47c73db1561 · outbound

This paper cites Detclip: Dictionary-enriched visual-concept paralleled pre- training for open-world detection.

Towards Real-Time Open-Vocabulary Video Instance Segmentation Detclip: Dictionary-enriched visual-concept paralleled pre- training for open-world detection

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:55.616550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:28:54.979749Z digest=sha256:4e82c69fcc56f9ef1891f6f8f053bb02d7d650f397dab33a7ec113a868a84176

Observation 0f5cccf9-b0e7-44e3-a958-50c6770a29f3 · outbound

This paper cites Modeling context in referring expres- sions.

Towards Real-Time Open-Vocabulary Video Instance Segmentation Modeling context in referring expres- sions

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:55.466578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:28:54.984216Z digest=sha256:0dd21f310e19a52a769b2a87edab6d8f057f41588a750465be015a4143f68175

Observation cd55fee2-5ded-4457-b71d-d187a4268cd3 · outbound

This paper cites Open-vocabulary object detection using captions.

Towards Real-Time Open-Vocabulary Video Instance Segmentation Open-vocabulary object detection using captions

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:55.379460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:28:54.988836Z digest=sha256:5d5c39f9897494c167bb90f996b5c7e11714daec31ecae0de08550dec4a4e9b7

Observation b9459cb0-f3fd-49a2-9f1a-e75ad850283d · outbound

This paper cites Faster Segment Anything: Towards Lightweight SAM for Mobile Applications.

Towards Real-Time Open-Vocabulary Video Instance Segmentation Faster Segment Anything: Towards Lightweight SAM for Mobile Applications

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:54.993311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:54.993311Z digest=sha256:8870e896e6d84e39b72099c6aa6e3f3c0f6ea61244d4e95b6df228de9d67547e

Observation 681f3d4a-f930-4ef8-8ae6-be68c7466da3 · outbound

This paper cites Dino: Detr with improved denoising anchor boxes for end-to-end object detection.

Towards Real-Time Open-Vocabulary Video Instance Segmentation Dino: Detr with improved denoising anchor boxes for end-to-end object detection

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:55.364331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:28:54.998071Z digest=sha256:948076d9c17af10d47a9188f1f3f2af2377168c6188a66c933a39bf6911168ba

Observation acfc595e-9e74-477c-9e65-b6a25a57e134 · outbound

This paper cites A simple framework for open-vocabulary segmentation and detection.

Towards Real-Time Open-Vocabulary Video Instance Segmentation A simple framework for open-vocabulary segmentation and detection

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:55.348767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:28:55.003046Z digest=sha256:9117c40c1a25af2e0eaeffa6bcb39a8a4a5d5c86bfe09f52b49baf7b97ace4f3

Observation 354ff347-eaf0-4477-9e7c-81c8eb486e71 · outbound

This paper cites Mobileinst: Video in- stance segmentation on the mobile.

Towards Real-Time Open-Vocabulary Video Instance Segmentation Mobileinst: Video in- stance segmentation on the mobile

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:55.333912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:28:55.007766Z digest=sha256:ed645f227bd8afa2e33af4c0e2ace7992751821d945744891db3f8f8b5934509

Observation fd06e973-41b9-4ee3-9f55-d8689cac4e7e · outbound

This paper cites Bridging the gap between anchor-based and anchor-free detection via adaptive training sample selection.

Towards Real-Time Open-Vocabulary Video Instance Segmentation Bridging the gap between anchor-based and anchor-free detection via adaptive training sample selection

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:55.319777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:28:55.012213Z digest=sha256:73ef4f861caeb8c983cd3a53a8765ff262820340094750bb334e1dd0a6ab32a2

Observation 85b4331f-947e-4f74-84bc-55d042a08cd9 · outbound

This paper cites EfficientViT-SAM: Accelerated Segment Anything Model Without Accuracy Loss.

Towards Real-Time Open-Vocabulary Video Instance Segmentation EfficientViT-SAM: Accelerated Segment Anything Model Without Accuracy Loss

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:55.016917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:55.016917Z digest=sha256:c7be6e08e0901df1e4f89b97b1a3fbb4915aca2b88246b38189745774d3eef90

Observation dc8f0053-2b15-458b-acad-3b21d14fae45 · outbound

This paper cites Real-time Transformer-based Open-Vocabulary Detection with Efficient Fusion Head.

Towards Real-Time Open-Vocabulary Video Instance Segmentation Real-time Transformer-based Open-Vocabulary Detection with Efficient Fusion Head

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:55.021270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:55.021270Z digest=sha256:0c6237eb958de80098540c00dacbc09c5cd1cd652ff27ebc85dc64b5fc973389

Observation 437f8fc6-39ae-42a8-97d9-1023dfa6b882 · outbound

This paper cites Detrs beat yolos on real-time object detection.

Towards Real-Time Open-Vocabulary Video Instance Segmentation Detrs beat yolos on real-time object detection

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:55.026043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:55.026043Z digest=sha256:7e99f0e51d4c3a3803b217632878ce17e95ac31df45ab2b3139a9ca68ad71395

Observation f8db961c-cbfa-4413-90f4-f58632584898 · outbound

This paper cites Detecting twenty-thousand classes using image-level supervision.

Towards Real-Time Open-Vocabulary Video Instance Segmentation Detecting twenty-thousand classes using image-level supervision

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:55.294472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:28:55.030130Z digest=sha256:d7324e1ab67761f4f7e48d6b3fcccbc2fa7c2acbf4d276c430bc325636d7f4f5

Observation d1a1aa1c-08fb-4fac-a53f-8ace31c28fbc · outbound

This paper cites Deformable detr: Deformable transformers for end-to-end object detection.

Towards Real-Time Open-Vocabulary Video Instance Segmentation Deformable detr: Deformable transformers for end-to-end object detection

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:28:55.263466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:28:55.039208Z digest=sha256:881c7b6f5e400014185e67836941293ff35906173eff83829d212fb650848d36

Observation 987506cd-d1fd-4508-9785-c5d1b9180110 · outbound

This paper cites Generalized decoding for pixel, image, and lan- guage.

Towards Real-Time Open-Vocabulary Video Instance Segmentation Generalized decoding for pixel, image, and lan- guage

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:55.043436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:55.043436Z digest=sha256:b34fbe37507d3de787fbdb2f20a394d9e561fc04827c00819f17f665a2001a5e

Observation cac69651-d2c3-48db-83f1-c4c4b25cd614 · outbound

This paper cites an unresolved cited work.

Towards Real-Time Open-Vocabulary Video Instance Segmentation Unresolved cited work

Reference 368

Resolution
parse uncertain
raw_fallback, observed 2026-08-11T21:28:55.279097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:28:55.034788Z digest=sha256:235eecdb67c9234a1f091776ad4e172877333d4fc6221aaf40dd71124ec66950

Pith citing papers

Observation ba96b6eb-c730-4041-940b-ce05dfa59553 · inbound

OpenFusion++: An Open-vocabulary Real-time Scene Understanding System cites this paper.

OpenFusion++: An Open-vocabulary Real-time Scene Understanding System Towards Real-Time Open-Vocabulary Video Instance Segmentation

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-16T06:01:20.339228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T06:01:19.889564Z digest=sha256:04c09b665305bb6acf2b175fab4c57133a6d9af33567bb3148ea7de96141e444