Pith. sign in

Paper Citation Record · LEDGER

Autoregressive Universal Video Segmentation Model

As of 18 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 0 inbound Pith citation observations for arXiv:2508.19242.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.19242 v1

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T15:52:16.689829Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

58 of 58 outbound references displayed

  • verified exact0
  • verified fuzzy44
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 86b416ae-2ac7-42d1-906c-40a7941ad941 · outbound

This paper cites Just read twice: closing the recall gap for recurrent language models.

Autoregressive Universal Video Segmentation Model Just read twice: closing the recall gap for recurrent language models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T15:52:16.392102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:52:16.392102Z digest=sha256:cf3c4cc8c5d59850cda85ea4bb9fce52685797fdd54cf005d2176a86df09b1b8

Observation ae6d0bf0-1b02-4538-841f-12baaedb2af1 · outbound

This paper cites Tarvis: A unified approach for target-based video segmentation.

Autoregressive Universal Video Segmentation Model Tarvis: A unified approach for target-based video segmentation

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:52:17.722317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:52:16.398111Z digest=sha256:9edf5f515f95b8bcc0ea9d7b6c1cf230cf78e926ece29e5dd3c03884a8a648bb

Observation ca615088-376e-49bc-9125-482cc92f49a8 · outbound

This paper cites End-to-end object detection with transformers.

Autoregressive Universal Video Segmentation Model End-to-end object detection with transformers

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:52:17.701666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:52:16.403264Z digest=sha256:71753d2940b4ada376887b128c03df6863a22628f8ca4cc0d9659bb7b396700e

Observation 97f0ece7-c232-4467-a313-477f1bf25b58 · outbound

This paper cites Per-pixel classification is not all you need for semantic segmentation.

Autoregressive Universal Video Segmentation Model Per-pixel classification is not all you need for semantic segmentation

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:52:17.684671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:52:16.408380Z digest=sha256:6dcab62163964fc06f2191a9fbb650b17bfb8d35fb4c32297ab872a0d687e5df

Observation b9172dd7-93d9-4e20-bdb8-7d22ae9c4c11 · outbound

This paper cites Masked-attention mask transformer for universal image segmentation.

Autoregressive Universal Video Segmentation Model Masked-attention mask transformer for universal image segmentation

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:52:17.666164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:52:16.413889Z digest=sha256:ad6a575977c6805a29b3f9c3f0818aa8941a6e74e0e21393422a14f9f935627f

Observation 7913f846-7918-44f7-a6cc-7bcf07e5d27c · outbound

This paper cites Xmem: Long-term video object segmentation with an atkinson- shiffrin memory model.

Autoregressive Universal Video Segmentation Model Xmem: Long-term video object segmentation with an atkinson- shiffrin memory model

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:52:17.649604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:52:16.419859Z digest=sha256:721627ad7ccba9dcb038724815cf36650f0f84419dcb2cdf7aa0530495a8ee2c

Observation f6c0c31c-4410-4c91-932d-c6b70941c826 · outbound

This paper cites The cityscapes dataset for semantic urban scene understanding.

Autoregressive Universal Video Segmentation Model The cityscapes dataset for semantic urban scene understanding

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:52:17.632470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:52:16.427098Z digest=sha256:29b03ee47986e077f691646bae0230b7ec21c4c09ec314ca170cc3d4b7adba98

Observation 4731e06c-f2f2-47e9-bae1-045836d56a9c · outbound

This paper cites Mose: A new dataset for video object segmentation in complex scenes.

Autoregressive Universal Video Segmentation Model Mose: A new dataset for video object segmentation in complex scenes

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:52:17.614476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:52:16.431862Z digest=sha256:773fd4603fa9040371a22c984b53df6cdfc77b81d25fd8ccb0d5ed96795c835e

Observation bb0d6999-317a-475b-9b9a-693a1fadbb33 · outbound

This paper cites The Llama 3 Herd of Models.

Autoregressive Universal Video Segmentation Model The Llama 3 Herd of Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T15:52:16.436505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:52:16.436505Z digest=sha256:90917b6de34d7f952ebc1b5f2651ff1150f6ec1f15846e198a6f83823cb64a1c

Observation 2649edb2-2207-4798-9572-e6178d0e3a6c · outbound

This paper cites Mamba: Linear-Time Sequence Modeling with Selective State Spaces.

Autoregressive Universal Video Segmentation Model Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T15:52:16.441780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:52:16.441780Z digest=sha256:5244b6bce340c142993f23fb6f08b47536a66b5ec51aa6f560705973a6944abc

Observation d91cf4ff-0e1e-4bb8-bc3b-92b85d64b33e · outbound

This paper cites Efficiently Modeling Long Sequences with Structured State Spaces.

Autoregressive Universal Video Segmentation Model Efficiently Modeling Long Sequences with Structured State Spaces

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T15:52:16.447019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:52:16.447019Z digest=sha256:7042ca876865c5c927bb2ad3949836022023ee401cd8d8e4615d9d16bd2e998e

Observation 57c9f109-2d46-4f4e-bbcf-e234a9c7b385 · outbound

This paper cites On the parameterization and initialization of diagonal state space models.

Autoregressive Universal Video Segmentation Model On the parameterization and initialization of diagonal state space models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:52:17.598398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:52:16.453019Z digest=sha256:8e92790830c5284defee037d390e9f2d013373c8dc4a29ce9a5249665b2488c6

Observation 2d6e736f-e66b-42c0-9932-5169ac518a0f · outbound

This paper cites Vita: Video instance segmentation via object token association.

Autoregressive Universal Video Segmentation Model Vita: Video instance segmentation via object token association

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:52:17.582327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:52:16.457809Z digest=sha256:1e086eea9b021a83ccc125751f03aefed4db69ba8f9b1751ac4bff454aa141c7

Observation 4a4bd33f-9d05-4b97-861b-8397af59a42d · outbound

This paper cites A generalized framework for video instance segmentation.

Autoregressive Universal Video Segmentation Model A generalized framework for video instance segmentation

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:52:17.565401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:52:16.462408Z digest=sha256:3d46e851eda35c641459a545633e5eb676d6760dd4bfe54a30d4147221e32c2a

Observation d08d82a7-5b73-4a83-ba97-14940e838a40 · outbound

This paper cites Omni-rgpt: Unifyingimageandvideoregion-levelunderstanding via token marks.

Autoregressive Universal Video Segmentation Model Omni-rgpt: Unifyingimageandvideoregion-levelunderstanding via token marks

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:52:17.549091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:52:16.467534Z digest=sha256:85e9aa9123c5e0b19779e9fdf15f6dced9035c5f8688fe30b2c13290bc2f33a0

Observation e5c62569-62d9-46f9-9be6-c0416591949a · outbound

This paper cites Robust and consistent online video instance segmentation via instance mask propagation.

Autoregressive Universal Video Segmentation Model Robust and consistent online video instance segmentation via instance mask propagation

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:52:17.532803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:52:16.472268Z digest=sha256:bba8847b0bc76096eb05140bbbfdb5c63c35816e8b50f47e9d8bfa69128dfef4

Observation cd06ca60-2ec1-4a76-b8b5-03085f6dc174 · outbound

This paper cites RULER: What's the Real Context Size of Your Long-Context Language Models?.

Autoregressive Universal Video Segmentation Model RULER: What's the Real Context Size of Your Long-Context Language Models?

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T15:52:16.476792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:52:16.476792Z digest=sha256:57f414b177f115d040e775da6cea404b4463802e3d1e3266a1a202f8d63acf91

Observation c0b308c0-6515-4f70-9e8f-71e5b060850d · outbound

This paper cites Minvis: A minimal video instance segmentation framework without video-based training.

Autoregressive Universal Video Segmentation Model Minvis: A minimal video instance segmentation framework without video-based training

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:52:17.514993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:52:16.482625Z digest=sha256:dc2dc6db34927223a4f1f423281a3dc36d0707c4a9faef870fc0a96e754924c9

Observation 7b45a9d9-c36c-44d2-ad63-e64170754081 · outbound

This paper cites Video instance segmentation using inter-frame communication transformers.

Autoregressive Universal Video Segmentation Model Video instance segmentation using inter-frame communication transformers

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:52:17.497831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:52:16.488015Z digest=sha256:446baad1466ffff20e200b9539fefb3cda58f5ec7747f0f8ffaf89f44ac4c376

Observation 4825247a-88c9-482d-93e6-d78cb83ab75b · outbound

This paper cites Video panoptic segmentation.

Autoregressive Universal Video Segmentation Model Video panoptic segmentation

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:52:17.480792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:52:16.492725Z digest=sha256:ad5bd0adc8b7e106f505d8fb15f5d838b814d7a181a42d949b52eaaa7d389720

Observation b8594234-570c-40c2-aac4-daf8e1bb48b6 · outbound

This paper cites Tubeformer-deeplab: Video mask transformer.

Autoregressive Universal Video Segmentation Model Tubeformer-deeplab: Video mask transformer

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:52:17.464204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:52:16.497752Z digest=sha256:14b250ed796d927fda1fbf77f1a58299fd39063c3cace06e8d536e1f00be06ba

Observation e0771a88-4d8d-4fe6-ac32-a5f5d94b833d · outbound

This paper cites Visage: Video instance segmentation with appearance-guided enhancement.

Autoregressive Universal Video Segmentation Model Visage: Video instance segmentation with appearance-guided enhancement

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:52:17.447234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:52:16.502950Z digest=sha256:fbc471cf9cd65f40693fdd03343b4e4bb3c21efdc5b46cdd3bdcf51cb2f452c4

Observation 1461a5e3-f4d3-4260-8759-8a302446bdf3 · outbound

This paper cites Berg, Wan-Yen Lo, Piotr Dollar, and Ross Girshick.

Autoregressive Universal Video Segmentation Model Berg, Wan-Yen Lo, Piotr Dollar, and Ross Girshick

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:52:17.425917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:52:16.508058Z digest=sha256:bb63eae6b8ee88007d07d09d9f492b08694332a8add4cdd380fddbbac2ad1fb0

Observation afc8ab21-2689-439d-b37f-0d5b105190ea · outbound

This paper cites Univs: Unified and universal video segmentation with prompts as queries.

Autoregressive Universal Video Segmentation Model Univs: Unified and universal video segmentation with prompts as queries

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:52:17.405124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:52:16.513868Z digest=sha256:d5b5152416ecd4714fcbc4237134e39d031c9a245bd3823f591ae960ff8cb8da

Observation fbd9a8f8-6014-45d1-b787-f737545eb36d · outbound

This paper cites Video k-net: A simple, strong, and unified baseline for video segmentation.

Autoregressive Universal Video Segmentation Model Video k-net: A simple, strong, and unified baseline for video segmentation

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:52:17.387623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:52:16.520167Z digest=sha256:b6af6665d5f231f818a693ca0cf1a7a97242959612511a06de53b0fc303caf4d

Observation 5bd90744-5653-4830-812d-1e2451df1eb7 · outbound

This paper cites Microsoft coco: Common objects in context.

Autoregressive Universal Video Segmentation Model Microsoft coco: Common objects in context

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:52:17.368698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:52:16.526647Z digest=sha256:f06c2e24eee89d012180acb9062cbe20824334ff93cee67928aa5656d2be8894

Observation 51e38b0e-7f0e-4874-bb39-8c9463bf3a93 · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows.

Autoregressive Universal Video Segmentation Model Swin transformer: Hierarchical vision transformer using shifted windows

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:52:17.352602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:52:16.531289Z digest=sha256:252d823ea8fd3dc1ad1c4bbd810e74878ec185c0027e5e0063f1c12f2b1da631

Observation 891f8809-64ee-4748-857a-a44bfab5c832 · outbound

This paper cites Decoupled weight decay regularization.

Autoregressive Universal Video Segmentation Model Decoupled weight decay regularization

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T15:52:16.536513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:52:16.536513Z digest=sha256:9b07adf17bed853b3870f61e6ce123dfd4cd200a518602fd6b322d9cacedd4f7

Observation e0e281c9-c11e-4333-bc78-8b1c9d2113c8 · outbound

This paper cites Language Models are Few-Shot Learners.

Autoregressive Universal Video Segmentation Model Language Models are Few-Shot Learners

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T15:52:16.541193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:52:16.541193Z digest=sha256:4967213cfb6da987625ab48c127ed19d8f89899b76dc39e539a045b9faf2c49b

Observation d723adf5-6981-48f6-a702-07d497284ddd · outbound

This paper cites Trackformer: Multi- object tracking with transformers.

Autoregressive Universal Video Segmentation Model Trackformer: Multi- object tracking with transformers

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:52:17.326155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:52:16.546111Z digest=sha256:bd0f471a8f300d39c117b3a45f407762e0c9662e73b99a5db8698fe3201a3289

Observation 39084838-b895-49db-b5e9-23042f0ea2ea · outbound

This paper cites MOT16: A Benchmark for Multi-Object Tracking.

Autoregressive Universal Video Segmentation Model MOT16: A Benchmark for Multi-Object Tracking

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T15:52:16.551066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:52:16.551066Z digest=sha256:c2225bf624db5122a34797304ccd1f3f7c7431546582c729d58c573e494e2553

Observation 96b0c5b9-0c8b-46d3-a322-5775d6c38d19 · outbound

This paper cites Videoobjectsegmentationusingspace-time memory networks.

Autoregressive Universal Video Segmentation Model Videoobjectsegmentationusingspace-time memory networks

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:52:17.309936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:52:16.556170Z digest=sha256:fe4c4c2170e2537ce86c73b70ea944c4dfb6e664b54bf922ae8700e1507ae85e

Observation ffe03473-a0b6-406e-9123-b4b74133be84 · outbound

This paper cites The 2017 DAVIS Challenge on Video Object Segmentation.

Autoregressive Universal Video Segmentation Model The 2017 DAVIS Challenge on Video Object Segmentation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T15:52:16.560810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:52:16.560810Z digest=sha256:0e5fdd9291b32b3f61fe7fe2817dbd79aea63a0bd1cfcfacf4dbc8f1ed8f6324

Observation 17033fd4-ea4f-40b0-a618-7550bc36d650 · outbound

This paper cites Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation.

Autoregressive Universal Video Segmentation Model Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T15:52:16.566222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:52:16.566222Z digest=sha256:626d765e5a23d6c5779a005a6d10660f227db259e49bdb0a39401dfa1becb5ca

Observation 97f69a1b-300c-4796-bfc5-f0a32e6de72c · outbound

This paper cites Occluded video instance segmentation: A benchmark.IJCV, 2022.

Autoregressive Universal Video Segmentation Model Occluded video instance segmentation: A benchmark.IJCV, 2022

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:52:17.292854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:52:16.571184Z digest=sha256:0d64ec411635adece7e09aee3dafba95dcad03b3dfb30f34fddcdde4577197f4

Observation a27ccbf2-f2a7-4e03-9083-c16f18856fc8 · outbound

This paper cites Language models are unsupervised multitask learners.OpenAI blog, 1(8):9, 2019.

Autoregressive Universal Video Segmentation Model Language models are unsupervised multitask learners.OpenAI blog, 1(8):9, 2019

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:52:17.277081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:52:16.576816Z digest=sha256:c8f6f1724666926f6b7ecbea1aad832b488d94325f331a7563f47028c8cc1ee0

Observation ddc4f406-bde8-4145-93bb-a8f4f4d39be0 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Autoregressive Universal Video Segmentation Model Learning transferable visual models from natural language supervision

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:52:17.261577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:52:16.581668Z digest=sha256:4d447c7dea19e3562497ca9e7a0e7c312532be0e82bd57e90958289f35482007

Observation 9cfa00d0-639a-4296-a769-b3d9dbd209c6 · outbound

This paper cites Sam 2: Segment anything in images and videos.

Autoregressive Universal Video Segmentation Model Sam 2: Segment anything in images and videos

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:52:17.244625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:52:16.586678Z digest=sha256:1c685ca7f6414bf1eb708700e29263830c941bd3a0d91ba2be18de3ab24bcaab

Observation 5689917a-dba1-45b7-8618-18c96592be45 · outbound

This paper cites Urvos: Unified referring video object segmentation network with a large-scale benchmark.

Autoregressive Universal Video Segmentation Model Urvos: Unified referring video object segmentation network with a large-scale benchmark

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:52:17.227593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:52:16.591657Z digest=sha256:4ed524c2e319f418f02ed48ed9805c2ff165bc1d81e6ceab18e21d88e91ce9f5

Observation 8b14a890-e01d-43b7-a6a8-ee613d1820c0 · outbound

This paper cites Repetition Improves Language Model Embeddings.

Autoregressive Universal Video Segmentation Model Repetition Improves Language Model Embeddings

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T15:52:16.597131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:52:16.597131Z digest=sha256:99f4078bbddc1d29a47a20946dd7de7aca8589981eb814ef3e2795429260b868

Observation 57c5c273-9677-46b0-bf01-f88a32608fa7 · outbound

This paper cites Sequence to sequence learning with neural networks.

Autoregressive Universal Video Segmentation Model Sequence to sequence learning with neural networks

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:52:17.209269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:52:16.602369Z digest=sha256:84af4875bf6f72ce2acd63ed3464ac296de6e946fe3f2412065485af27826973

Observation 58d6caf9-96d9-4f37-8235-aa98213c07f8 · outbound

This paper cites Attention is all you need.

Autoregressive Universal Video Segmentation Model Attention is all you need

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:52:17.192463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:52:16.607296Z digest=sha256:35b36b9bc45a929f18a66c217751a7b36be222028ddcc95a8830a08600eb140f

Observation 471ab83e-d319-43c1-b713-0b7ad505bc3d · outbound

This paper cites Mots: Multi-object tracking and segmentation.

Autoregressive Universal Video Segmentation Model Mots: Multi-object tracking and segmentation

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:52:17.174807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:52:16.612009Z digest=sha256:093d5fcf8b2fa4efb8ae83bf7ed82247f8d927338b843a606ce17d8718bb34ec

Observation fe02fd67-d773-4f2d-8c45-b940e4c0ba9f · outbound

This paper cites Max-deeplab: End-to-end panoptic segmentation with mask transformers.

Autoregressive Universal Video Segmentation Model Max-deeplab: End-to-end panoptic segmentation with mask transformers

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:52:17.158195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:52:16.617095Z digest=sha256:3a4099810b2f52d4523fd91c1a052727ddb972bd1efe729a51db01edafbd6453

Observation 1b47c41b-c718-4f3a-9541-ffefbc1331ad · outbound

This paper cites End-to-end video instance segmentation with transformers.

Autoregressive Universal Video Segmentation Model End-to-end video instance segmentation with transformers

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:52:17.141749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:52:16.622055Z digest=sha256:0f1e2bef2c39dd6c97283da389413c78c32eb879660caa3c5a40381cbf5bfbcd

Observation 88c1a9fb-1e2e-42e3-93d4-4ee22bdb26ff · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Autoregressive Universal Video Segmentation Model Chain-of-thought prompting elicits reasoning in large language models

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:52:17.122859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:52:16.627092Z digest=sha256:4860c02adb92d776e931d58fa5167853cb09ffe2b4045a25006b22993d896b80

Observation 54cc4cc7-8dac-482e-8275-4e8218799bdf · outbound

This paper cites Segment every reference object in spatial and temporal spaces.

Autoregressive Universal Video Segmentation Model Segment every reference object in spatial and temporal spaces

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:52:17.105300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:52:16.632189Z digest=sha256:54ef8151cb0958abb449803b669cbea05707b0645b0e1758877818024f09b9c6

Observation 12536e1d-f8d0-46ec-b7a0-6669087dc48b · outbound

This paper cites UniRef++: Segment Every Reference Object in Spatial and Temporal Spaces.

Autoregressive Universal Video Segmentation Model UniRef++: Segment Every Reference Object in Spatial and Temporal Spaces

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-05T15:52:16.637223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:52:16.637223Z digest=sha256:2249096f8dede38c3b0d9dc1554a1436a3c406921f86c875c29d5f20f1319c89

Observation b8c805a8-9f0d-4803-b494-4103fd5a4c78 · outbound

This paper cites In defense of online models for video instance segmentation.

Autoregressive Universal Video Segmentation Model In defense of online models for video instance segmentation

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:52:17.086960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:52:16.642391Z digest=sha256:72ea9f36470bffdaa2ee1275fb4b0aa1b281462cfc58211e673642acc30a3b40

Observation 03fabcbb-02d6-453d-814e-9316fd8407af · outbound

This paper cites Online object tracking: A benchmark.

Autoregressive Universal Video Segmentation Model Online object tracking: A benchmark

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:52:17.067808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:52:16.648003Z digest=sha256:eb545fb33cadcd1cf2b394630bef30bf0f7df11bf3bc4bc94aef9487fa30a717

Observation 99af8edc-eff1-4366-8650-221d764ec550 · outbound

This paper cites Efficient Streaming Language Models with Attention Sinks.

Autoregressive Universal Video Segmentation Model Efficient Streaming Language Models with Attention Sinks

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-05T15:52:16.653696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:52:16.653696Z digest=sha256:420bf534017ba6ad24b1eed953e883bffe40bbe7f70811941f621615e1fd59e8

Observation 69fb682e-a55a-4b1f-acab-f281f0b2c7b8 · outbound

This paper cites YouTube-VOS: A Large-Scale Video Object Segmentation Benchmark.

Autoregressive Universal Video Segmentation Model YouTube-VOS: A Large-Scale Video Object Segmentation Benchmark

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-05T15:52:16.659221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:52:16.659221Z digest=sha256:9006393832ee323a18d807360a4f61d1ba1d02313a2dc2db00aaaa8860f05de2

Observation 5eb3c589-7670-4912-8c20-f706e56b01f4 · outbound

This paper cites Universal instance perception as object discovery and retrieval.

Autoregressive Universal Video Segmentation Model Universal instance perception as object discovery and retrieval

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:52:17.050220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:52:16.664517Z digest=sha256:bd92cedae74456d21e246ce71c7dce3b5ccf801c254503c702a34051a7b6c8c0

Observation 41f27a09-3114-4a91-aca5-58a7a9d5f584 · outbound

This paper cites Video instance segmentation.

Autoregressive Universal Video Segmentation Model Video instance segmentation

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:52:17.030729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:52:16.669522Z digest=sha256:0e6647cfcdcb6870441bce3987dd646fcbeff633d0661af133d2be41cf44daa9

Observation 7ca2f19f-1bc9-49ea-b5be-af608e24fff5 · outbound

This paper cites Decoupling features in hierarchical propagation for video object segmentation.

Autoregressive Universal Video Segmentation Model Decoupling features in hierarchical propagation for video object segmentation

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:52:17.011198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:52:16.674355Z digest=sha256:eab05b9fbdfde92f0d881afd3b96e686b74f7ec564e7f9b386ee041e1b50b68d

Observation 2e1593db-fa4d-45ad-a308-0788bd75f222 · outbound

This paper cites Associating objects with transformers for video object segmen- tation.

Autoregressive Universal Video Segmentation Model Associating objects with transformers for video object segmen- tation

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:52:16.992951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:52:16.680179Z digest=sha256:0e3bb69f16b5dc20f6f99d086136a6119ed778f1079fc6dbbe5f32801bd36d4c

Observation d813b1bd-be76-44b6-85a5-45ccd77f5dd2 · outbound

This paper cites Ctvis: Consistent training for online video instance segmentation.

Autoregressive Universal Video Segmentation Model Ctvis: Consistent training for online video instance segmentation

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:52:16.975710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:52:16.684828Z digest=sha256:49cfa1a021c03ade9c1b7962a4522727b98e86f3b19c9aefe540f0765f95423a

Observation aaf548a7-c725-4cef-932b-04afca00c196 · outbound

This paper cites Dvis: Decoupled video instance segmentation framework.

Autoregressive Universal Video Segmentation Model Dvis: Decoupled video instance segmentation framework

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:52:16.957792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:52:16.689829Z digest=sha256:447d6a0ba2575f5193bfcd27b00a677cbe0388597ac68ed0c35fe008df197a39

Pith citing papers

No inbound Pith citation observations are available.