Pith. sign in

Paper Citation Record · LEDGER

Mask2Former for Video Instance Segmentation

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 20 inbound Pith citation observations for arXiv:2112.10764.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2112.10764 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 20 of 20 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:25:44.060493Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T03:06:43.601051Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 164a91ab-6c2d-4928-8067-bfaaf7be41f5 · inbound

Primus: Enforcing Attention Usage for 3D Medical Image Segmentation cites this paper.

Primus: Enforcing Attention Usage for 3D Medical Image Segmentation Mask2Former for Video Instance Segmentation

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-23T01:25:16.506146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-23T01:22:29.901174Z digest=sha256:c3895476b444b486604aec757fbfa7da19444d3f075c4f5ad6cf399b79fc8c58

Observation 10f86999-c4bc-44b4-bc1c-5005dc1bc195 · inbound

OneIG-Bench: Omni-dimensional Nuanced Evaluation for Image Generation cites this paper.

OneIG-Bench: Omni-dimensional Nuanced Evaluation for Image Generation Mask2Former for Video Instance Segmentation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:44.060493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:44.060493Z digest=sha256:6453b07f12c47645c0121cfa8a4f288d60433f09dbef51e588534b15d5afeefb

Observation cab951f1-9f45-4213-978b-f68eea78264e · inbound

Stepping Out of Similar Semantic Space for Open-Vocabulary Segmentation cites this paper.

Stepping Out of Similar Semantic Space for Open-Vocabulary Segmentation Mask2Former for Video Instance Segmentation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T23:50:23.286782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:50:23.286782Z digest=sha256:ec98596c4433d990290da6ba96b47297ba8a615ef8ae07745db50738c8c74bf5

Observation 87b8979b-f055-487d-850c-6ce6b719beea · inbound

GeoMag: A Vision-Language Model for Pixel-level Fine-Grained Remote Sensing Image Parsing cites this paper.

GeoMag: A Vision-Language Model for Pixel-level Fine-Grained Remote Sensing Image Parsing Mask2Former for Video Instance Segmentation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T19:21:13.164766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:21:13.164766Z digest=sha256:863396fcf8de12a533b93c12b85790a23fdb4deffcc3b8be6f6191e61cf55eb4

Observation bd7d29b2-c6db-4d8c-b6fc-7cec9ca92856 · inbound

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation cites this paper.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Mask2Former for Video Instance Segmentation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T19:17:31.737828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:17:31.737828Z digest=sha256:71f6dc005244320fe7e4e69ce9fea4168e28c5d1b4a7ddba462148530dd6ec16

Observation d6d28bd7-fe06-4d42-8334-568ec688b40d · inbound

GVCCS: A Dataset for Contrail Identification and Tracking on Visible Whole Sky Camera Sequences cites this paper.

GVCCS: A Dataset for Contrail Identification and Tracking on Visible Whole Sky Camera Sequences Mask2Former for Video Instance Segmentation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T14:37:32.089540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:37:32.089540Z digest=sha256:8e6a0d246206db661a0c8b80ca9b50a93d1aa086b17051944955da9a2173cf76

Observation 2489c505-f001-449f-97bc-6756f5153e56 · inbound

GS-Occ3D: Scaling Vision-only Occupancy Reconstruction with Gaussian Splatting cites this paper.

GS-Occ3D: Scaling Vision-only Occupancy Reconstruction with Gaussian Splatting Mask2Former for Video Instance Segmentation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T14:22:13.254653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:22:13.254653Z digest=sha256:61689589d209e1f08d9a739f075d577c74fff5b4a91db756c6be07fa0d84aaea

Observation c6ff474e-4aed-44d8-83bf-b09ba6be3630 · inbound

Latest Object Memory Management for Temporally Consistent Video Instance Segmentation cites this paper.

Latest Object Memory Management for Temporally Consistent Video Instance Segmentation Mask2Former for Video Instance Segmentation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T14:08:43.698403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:08:43.698403Z digest=sha256:2b51c9ae5a3d4ed088bf3eb359b5c7903f53ec75756d31fa5cfd8c2a913992ea

Observation 6f339dd2-819a-4472-a54d-7544ee0bdf16 · inbound

Temporal Cluster Assignment for Efficient Real-Time Video Segmentation cites this paper.

Temporal Cluster Assignment for Efficient Real-Time Video Segmentation Mask2Former for Video Instance Segmentation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T23:11:37.909813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:11:37.909813Z digest=sha256:afbb3e1edaa0ec9cda7150961e684602d4e136f6d473db255bd1536bbbcc92fd

Observation a89c39a0-f55f-4a9b-8ba0-b528648f14c8 · inbound

CObL: Toward Zero-Shot Ordinal Layering without User Prompting cites this paper.

CObL: Toward Zero-Shot Ordinal Layering without User Prompting Mask2Former for Video Instance Segmentation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T21:35:35.401431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:35:35.401431Z digest=sha256:7d106e07a46bc5bae09d4bd9a842a8da850c7ae46081b825d7aeb7c61a84b197

Observation 453aadd4-07b8-4fe4-8e23-54c911b8f2fb · inbound

Interleaved Transceiver Design for a Continuous- Transmission MIMO-OFDM ISAC System cites this paper.

Interleaved Transceiver Design for a Continuous- Transmission MIMO-OFDM ISAC System Mask2Former for Video Instance Segmentation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T20:30:02.984463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:30:02.984463Z digest=sha256:f6702e6c68317dde3407ec4ec9166d88b95ca4752526a6293bc118d1f0b1e1c7

Observation 500da900-2e9b-4573-94d9-6a9d882cd9a5 · inbound

AutoQ-VIS: Improving Unsupervised Video Instance Segmentation via Automatic Quality Assessment cites this paper.

AutoQ-VIS: Improving Unsupervised Video Instance Segmentation via Automatic Quality Assessment Mask2Former for Video Instance Segmentation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T15:33:15.913482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:33:15.913482Z digest=sha256:104e59b2cd2571233c149fda8c44a7ced0f02f1cba9cbcc40c7daa982173dfb9

Observation 10c61aa5-734c-4aec-b0a5-643ae63bd371 · inbound

CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects cites this paper.

CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects Mask2Former for Video Instance Segmentation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T09:31:55.211242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:31:55.211242Z digest=sha256:978f363b8b89bae7a881186e54c4021138ea5ba5daff2b8d5cfbfea5ec6fa700

Observation be6a2d55-45cc-495b-a708-0f4ef7620b64 · inbound

PAT-VCM: Plug-and-Play Auxiliary Tokens for Video Coding for Machines cites this paper.

PAT-VCM: Plug-and-Play Auxiliary Tokens for Video Coding for Machines Mask2Former for Video Instance Segmentation

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T09:56:01.846052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:45:09.913980Z digest=sha256:b9b911b28d5cc15855e7ea01eb8c271453cebaa069b7867c651e18a0828b4360

Observation adf30736-1412-4539-9207-ce3599cc819d · inbound

GOLD-BEV: GrOund and aeriaL Data for Dense Semantic BEV Mapping of Dynamic Scenes cites this paper.

GOLD-BEV: GrOund and aeriaL Data for Dense Semantic BEV Mapping of Dynamic Scenes Mask2Former for Video Instance Segmentation

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:41:01.637082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T03:20:24.032405Z digest=sha256:77642e0a1cec45672b4f840a6bf88f15679d6f6d777bfeb19a64c68e26c9ee5e

Observation fc1c1c04-981c-4149-8ea9-74bbae5d0e56 · inbound

Mind the Gap: Disentangling Performance Bottlenecks in Video Instance Segmentation cites this paper.

Mind the Gap: Disentangling Performance Bottlenecks in Video Instance Segmentation Mask2Former for Video Instance Segmentation

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:07:12.925722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T22:10:36.011508Z digest=sha256:c0378ea81cc06594c2af7a8e977934f42e2d002634373257ed19762ceed5a9dd

Observation 7f9d0ede-c8c2-4449-8697-5369740011c7 · inbound

SA-VIS: Sparse frame Annotations for training Video Instance Segmentation cites this paper.

SA-VIS: Sparse frame Annotations for training Video Instance Segmentation Mask2Former for Video Instance Segmentation

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T03:09:30.493585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T18:19:56.997875Z digest=sha256:721367ce512758cf034cfff135cceabfb4fa393516edffd8403ce2f0c32a6dc0

Observation 46758175-4b06-49e8-a5a7-1e1fe1319b2f · inbound

SA-VIS: Sparse frame Annotations for training Video Instance Segmentation cites this paper.

SA-VIS: Sparse frame Annotations for training Video Instance Segmentation Mask2Former for Video Instance Segmentation

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T10:54:37.012591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T10:45:35.568212Z digest=sha256:137cec9ecd3ff286c518611704c64d88d8ae2027e16e9c347b07236d1fa18080

Observation aacbb5a3-64a2-4e90-bd87-ffdedc95b691 · inbound

SAM-MT: Real-Time Interactive Multi-Target Video Segmentation cites this paper.

SAM-MT: Real-Time Interactive Multi-Target Video Segmentation Mask2Former for Video Instance Segmentation

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-07-10T03:06:43.602664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-10T03:04:57.109924Z digest=sha256:8f7ebedf46a14501356d37da3821f99b01adbddb2b5cfb37948444ceaa0ddc6e

Observation c631030b-3c73-4f78-a06e-796a12b477db · inbound

Hear to See: Discerning Stateful Listening for Audio-Visual Instance Segmentation cites this paper.

Hear to See: Discerning Stateful Listening for Audio-Visual Instance Segmentation Mask2Former for Video Instance Segmentation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T22:20:46.977826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:20:46.977826Z digest=sha256:8de0ff9733814c8bbd572876d736a26bfd2c80ffb6c8e93ff565e9cfd6782240