Pith. sign in

Paper Citation Record · LEDGER

YouTube-VOS: A Large-Scale Video Object Segmentation Benchmark

As of 5 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 27 inbound Pith citation observations for arXiv:1809.03327.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1809.03327 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 27 of 27 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T05:28:44.443964Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T02:04:26.310872Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 2837349d-c4e8-4805-a500-0a3923168857 · inbound

Understanding Deep Learning Techniques for Image Segmentation cites this paper.

Understanding Deep Learning Techniques for Image Segmentation YouTube-VOS: A Large-Scale Video Object Segmentation Benchmark

Reference 209

Resolution
verified exact
local_arxiv, observed 2026-05-24T21:46:24.412506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T21:46:17.736097Z digest=sha256:7de1068f59354e21af22fa7ad3fa516baa9e7a075080f823e41ea7240cbb2a66

Observation 04a1d9c5-c254-41c2-a2bf-a600c88c39cf · inbound

SAM 2: Segment Anything in Images and Videos cites this paper.

SAM 2: Segment Anything in Images and Videos YouTube-VOS: A Large-Scale Video Object Segmentation Benchmark

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:56:25.466836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T13:56:25.331304Z digest=sha256:624c9cc00e40e4e980ed5cd927f894eb18b72d4cefa3ddcd04e7a86426421b0a

Observation 2acf683e-3971-46fc-aa0c-9921b207226e · inbound

Towards a General-Purpose Zero-Shot Synthetic Low-Light Image and Video Pipeline cites this paper.

Towards a General-Purpose Zero-Shot Synthetic Low-Light Image and Video Pipeline YouTube-VOS: A Large-Scale Video Object Segmentation Benchmark

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-22T20:22:03.834071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T20:18:35.764575Z digest=sha256:1dbd4df0e38beb4c648cef8ea8de8347b41ee24bbc297cc1da2f3d0bb971119b

Observation 828246b3-e0a2-46a8-9800-d6a586a05c89 · inbound

SIV-Bench: A Video Benchmark for Social Interaction Understanding and Reasoning cites this paper.

SIV-Bench: A Video Benchmark for Social Interaction Understanding and Reasoning YouTube-VOS: A Large-Scale Video Object Segmentation Benchmark

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-05-19T11:37:15.726369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:36:36.687324Z digest=sha256:99a506cec373bc78252755be61b7def632c36027953df80db58915d82999972e

Observation 71a37c2a-352c-4fdc-a98e-6d621cf2c18d · inbound

FlowSeek: Optical Flow Made Easier with Depth Foundation Models and Motion Bases cites this paper.

FlowSeek: Optical Flow Made Easier with Depth Foundation Models and Motion Bases YouTube-VOS: A Large-Scale Video Object Segmentation Benchmark

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-05T05:28:44.443964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:28:44.443964Z digest=sha256:fbf4cb4f04acd86f7ff0793a17969728e47e64032c8a51e82c561ae6c28eb054

Observation d0042d06-6297-4269-8993-09d056fa8d71 · inbound

SAM 2++: Tracking Anything at Any Granularity cites this paper.

SAM 2++: Tracking Anything at Any Granularity YouTube-VOS: A Large-Scale Video Object Segmentation Benchmark

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-05-21T19:54:20.151923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T19:51:56.047515Z digest=sha256:03247c8767201af9b3a48ea373a75d85deb26b2dcb808467dc39e29724e44d80

Observation a3a3a651-84f5-40b2-8a68-ded2354c8081 · inbound

SAM 3: Segment Anything with Concepts cites this paper.

SAM 3: Segment Anything with Concepts YouTube-VOS: A Large-Scale Video Object Segmentation Benchmark

Reference 145

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:25:11.536644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:76e7cf927ccc981dfe71c94c18f949f8d1ced113738f42f6fd8b4c0ff48de9f5

Observation 54a755de-dd47-4e30-9ba1-9fcb5002d239 · inbound

Recurrent Video Masked Autoencoders cites this paper.

Recurrent Video Masked Autoencoders YouTube-VOS: A Large-Scale Video Object Segmentation Benchmark

Reference 75

Resolution
verified exact
local_arxiv, observed 2026-05-16T21:58:35.816176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T21:55:42.555679Z digest=sha256:88acf8a7d414a5ff7d0e44a7d5e4431cc325672c67f819706f7bad9e3ca546fa

Observation ccd8cf70-8809-4803-b303-04f07816ea3f · inbound

Backdoor Attacks on Prompt-Driven Video Segmentation Foundation Models cites this paper.

Backdoor Attacks on Prompt-Driven Video Segmentation Foundation Models YouTube-VOS: A Large-Scale Video Object Segmentation Benchmark

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-05-16T19:11:11.665827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T19:08:37.460895Z digest=sha256:14bca602ebf35c7db22cb366de7a4664d0a7e80919e9c0869ee6e705e8c76f69

Observation a3327468-b513-46c3-aa64-ded663dcb682 · inbound

3AM: 3egment Anything with Geometric Consistency in Videos cites this paper.

3AM: 3egment Anything with Geometric Consistency in Videos YouTube-VOS: A Large-Scale Video Object Segmentation Benchmark

Reference 92

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T14:21:01.642414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T14:19:34.641108Z digest=sha256:579d085af1bdc182c21d864fa4cd25fa3949870532a16fa7c04bde71f36ebdff

Observation 5d7c2b68-6663-4bd6-9907-afd4a0d487fc · inbound

FlowIt: Global Matching via Hierarchical Transformers and Optimal Transport for Optical Flow cites this paper.

FlowIt: Global Matching via Hierarchical Transformers and Optimal Transport for Optical Flow YouTube-VOS: A Large-Scale Video Object Segmentation Benchmark

Reference 78

Resolution
unresolved
no resolver link, observed 2026-07-13T16:08:30.031334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:08:30.031334Z digest=sha256:59a8b9a28048fbde2ba1ac832b8076d2417dc462717858daea9b1f6d0961ffcd

Observation fea0e949-68d5-4197-bd0c-53d3439aed35 · inbound

PanoSAM2: Lightweight Distortion- and Memory-aware Adaptions of SAM2 for 360 Video Object Segmentation cites this paper.

PanoSAM2: Lightweight Distortion- and Memory-aware Adaptions of SAM2 for 360 Video Object Segmentation YouTube-VOS: A Large-Scale Video Object Segmentation Benchmark

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T00:30:55.265816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:27:39.914591Z digest=sha256:04f92aee78963f138a2003a7d6cee9dcead1aa688b8daf1c7f86ab086ea470c9

Observation 1b41aa70-6c34-4efd-b843-4fb0c7de53ad · inbound

CMTM: Cross-Modal Token Modulation for Unsupervised Video Object Segmentation cites this paper.

CMTM: Cross-Modal Token Modulation for Unsupervised Video Object Segmentation YouTube-VOS: A Large-Scale Video Object Segmentation Benchmark

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:51:01.038342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T12:32:11.911775Z digest=sha256:544b42d80094402d9c870e4f47fcb355687e56a9721274d08ad7ad799079899c

Observation 6c9c2355-208c-4954-b29e-859f2564987a · inbound

Featurising Pixels from Dynamic 3D Scenes with Linear In-Context Learners cites this paper.

Featurising Pixels from Dynamic 3D Scenes with Linear In-Context Learners YouTube-VOS: A Large-Scale Video Object Segmentation Benchmark

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:51:24.798782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-07T13:37:39.954299Z digest=sha256:e1badff48b13b73d4f3cb218d294e535de5e45450b364904696c51947a391862

Observation 060bc62b-c9bf-4756-a4a0-2f6b6e69e453 · inbound

YOSE: You Only Select Essential Tokens for Efficient DiT-based Video Object Removal cites this paper.

YOSE: You Only Select Essential Tokens for Efficient DiT-based Video Object Removal YouTube-VOS: A Large-Scale Video Object Segmentation Benchmark

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:46:27.935892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T09:12:13.460169Z digest=sha256:e7cda458e6d09708f1d23caa5e48eddcd4432b9f72cb09298479bd135bcbb4f2

Observation 2cb0fd27-26fd-4a0c-a948-00e046bcd75b · inbound

X2SAM: Any Segmentation in Images and Videos cites this paper.

X2SAM: Any Segmentation in Images and Videos YouTube-VOS: A Large-Scale Video Object Segmentation Benchmark

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:01:05.428048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-09T20:47:06.698475Z digest=sha256:b0b7de6414d8f484fec14299f85d5973b9719be6de6ed595c5e583bab56c17ef

Observation 8dc0410c-3933-4020-96a4-8c8d736e2788 · inbound

VEBench:Benchmarking Large Multimodal Models for Real-World Video Editing cites this paper.

VEBench:Benchmarking Large Multimodal Models for Real-World Video Editing YouTube-VOS: A Large-Scale Video Object Segmentation Benchmark

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:06:22.357494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T01:30:19.531699Z digest=sha256:5263984e2fb1d645049eef623181e6797b7bef9f4e49755a56b5b050688708cf

Observation 19766d70-a65c-4db4-90d9-718cd387a5f0 · inbound

VEBench:Benchmarking Large Multimodal Models for Real-World Video Editing cites this paper.

VEBench:Benchmarking Large Multimodal Models for Real-World Video Editing YouTube-VOS: A Large-Scale Video Object Segmentation Benchmark

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-12T01:46:13.949720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:43:34.898639Z digest=sha256:fb564a01e579422f2cefc762f98820592a6f5967d48126228ba885dc6c08a6e6

Observation 73302bf2-b071-4a48-9048-9338e3dbac4c · inbound

Robust Promptable Video Object Segmentation cites this paper.

Robust Promptable Video Object Segmentation YouTube-VOS: A Large-Scale Video Object Segmentation Benchmark

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:22:29.355364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T07:18:20.281383Z digest=sha256:e679370507aec3de6bb2ee616ccfd4115bbebd78411c5b3fe910218793f242e2

Observation f75d5859-c275-486b-a021-7cdec648e0e4 · inbound

Functionalization via Structure Completion and Motion Rectification cites this paper.

Functionalization via Structure Completion and Motion Rectification YouTube-VOS: A Large-Scale Video Object Segmentation Benchmark

Reference 281

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T12:28:17.069803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-20T12:25:07.157086Z digest=sha256:3231892531532b3343d4d18cc41861fc3705f7b976285f8c1a76afc41c731120

Observation 4ec12bb2-e568-476f-aa53-3df3d84aa9b5 · inbound

A Trajectory-Driven Spatio-Temporal Refinement Solution for CVPR 2026 8th UG2+ Challenge Track 3: DOST cites this paper.

A Trajectory-Driven Spatio-Temporal Refinement Solution for CVPR 2026 8th UG2+ Challenge Track 3: DOST YouTube-VOS: A Large-Scale Video Object Segmentation Benchmark

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-06-28T19:42:35.941042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T18:55:17.222924Z digest=sha256:6be1d581f3cc4e0b283c701afd7844f5c2a310891ae6864bebfd24ba5b66df04

Observation 93fef265-98ed-4c22-85ce-b1e4f2291e93 · inbound

MG-RWKV: Multi-Grained Context-Aware RWKV for Temporal Forgery Localization cites this paper.

MG-RWKV: Multi-Grained Context-Aware RWKV for Temporal Forgery Localization YouTube-VOS: A Large-Scale Video Object Segmentation Benchmark

Reference 46

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T14:07:02.278608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-02T14:03:46.585001Z digest=sha256:38eb212104ac40ddb1abce2dc88e9e98d33e37297566d183ce6df112703cdb5b

Observation ba25135d-c025-4de1-b363-f149987afac8 · inbound

STAC: Selective Spatiotemporal Aggregation and Compression for Video Reasoning Segmentation cites this paper.

STAC: Selective Spatiotemporal Aggregation and Compression for Video Reasoning Segmentation YouTube-VOS: A Large-Scale Video Object Segmentation Benchmark

Reference 77

Resolution
unresolved
no resolver link, observed 2026-07-12T06:06:47.233814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:06:47.233814Z digest=sha256:6681c3b719c26709462d7f1f5892c8ce55a8ef5cc556fa6b790f5ee9c0eaf7dd

Observation 5faf9b26-e55b-462f-b2cc-1f560eaf1362 · inbound

Vision Pretraining for Dense Spatial Perception cites this paper.

Vision Pretraining for Dense Spatial Perception YouTube-VOS: A Large-Scale Video Object Segmentation Benchmark

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-07-07T22:24:11.145551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-07T22:20:17.456484Z digest=sha256:8d90387e7ad382f90be7f5072cf160ea45eb781397c3136d28e7aaa09afca07f

Observation 9c3047a0-4581-474b-adb6-3e7fdf800106 · inbound

Vision as Unified Multimodal Generation cites this paper.

Vision as Unified Multimodal Generation YouTube-VOS: A Large-Scale Video Object Segmentation Benchmark

Reference 197

Resolution
verified exact
local_arxiv, observed 2026-07-08T02:04:26.312361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:f1e0afc2910b4290f1f299c21d45ab0db2d144f7b57f677161ab9ed02f7cffd0

Observation f6a1186f-f177-43b6-91e8-07412c4b0314 · inbound

PercepCap: Video Captioner with Structured Spatio-Temporal Perception cites this paper.

PercepCap: Video Captioner with Structured Spatio-Temporal Perception YouTube-VOS: A Large-Scale Video Object Segmentation Benchmark

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T10:02:03.252658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:02:03.252658Z digest=sha256:8477655448d6bc800c4cb0bd0302af4613969963457b4efe5af4989399d610d6

Observation 35b55820-2081-4f19-8147-b4dd6e62b70f · inbound

Self-Supervised Learning of Structured Dynamics from Videos cites this paper.

Self-Supervised Learning of Structured Dynamics from Videos YouTube-VOS: A Large-Scale Video Object Segmentation Benchmark

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:43.997932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:43.997932Z digest=sha256:405e15dcb5a5ac679086acc7b03f01399aef4cdf970988b329869633d1004010