Pith. sign in

Paper Citation Record · LEDGER

Video Swin Transformer

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 25 inbound Pith citation observations for arXiv:2106.13230.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2106.13230 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 25 of 25 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T15:50:32.514966Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T19:08:49.804199Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f9c80622-dd4a-4296-8644-c964521a19a2 · inbound

Florence: A New Foundation Model for Computer Vision cites this paper.

Florence: A New Foundation Model for Computer Vision Video Swin Transformer

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:38:09.533761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T09:38:09.427509Z digest=sha256:384689072ef90ba4022b769bd665e2ab9a2a37b6ad6050719809152ad9601079

Observation f61f62ae-f6e4-414c-9216-2a8e3757797d · inbound

Video-Language Understanding: A Survey from Model Architecture, Model Training, and Data Perspectives cites this paper.

Video-Language Understanding: A Survey from Model Architecture, Model Training, and Data Perspectives Video Swin Transformer

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-24T00:23:39.590980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-24T00:22:35.635679Z digest=sha256:d1f4dbfec2daae621c822a9cab5f4242693cee15c4bd757f20314ae28afaafb4

Observation 268ae948-2c9d-46bb-9d91-28360e4305ab · inbound

Rapid Reconstruction of Extremely Accelerated Liver 4D MRI via Chained Iterative Refinement cites this paper.

Rapid Reconstruction of Extremely Accelerated Liver 4D MRI via Chained Iterative Refinement Video Swin Transformer

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T15:50:32.514966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:50:32.514966Z digest=sha256:066ac50e8221f108672508fc56128959a07780049916a9d6a094487566eb1e6e

Observation 366c63c6-aa71-4bb7-a88c-1f8bda6db2b2 · inbound

Mask-RadarNet: Enhancing Transformer With Spatial-Temporal Semantic Context for Radar Object Detection in Autonomous Driving cites this paper.

Mask-RadarNet: Enhancing Transformer With Spatial-Temporal Semantic Context for Radar Object Detection in Autonomous Driving Video Swin Transformer

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T11:20:27.090105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:20:27.090105Z digest=sha256:5a7ea9ccdd1a149aa51377f7020f701bbcc444f30f384f8fea1c22687105eca5

Observation 5bee6127-0dcc-43df-a926-b519a3dd18a4 · inbound

Visual WetlandBirds Dataset: Bird Species Identification and Behavior Recognition in Videos cites this paper.

Visual WetlandBirds Dataset: Bird Species Identification and Behavior Recognition in Videos Video Swin Transformer

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T20:17:36.366468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:17:36.366468Z digest=sha256:f07f1ef3b651d31135eef604a4d43bcd3a19b3c44b03117567bdf6b5f9d1a29d

Observation de12811d-cc0d-4bcc-a084-55f0f3072b63 · inbound

Data-Efficient Challenges in Visual Inductive Priors: A Retrospective cites this paper.

Data-Efficient Challenges in Visual Inductive Priors: A Retrospective Video Swin Transformer

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-07T05:10:59.434114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:10:59.434114Z digest=sha256:84acd46e9247a563b36b5d1e9972edb63a62fc9db2560de18fcb4d7a8963196b

Observation 31bac54f-710e-42bd-9f64-fc4ae6435b56 · inbound

Feature Hallucination for Self-supervised Action Recognition cites this paper.

Feature Hallucination for Self-supervised Action Recognition Video Swin Transformer

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-06T22:55:23.881825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:55:23.881825Z digest=sha256:b5438ca4c589a4b208ed23b6ca5bd42c45c66775a172af695763c43734dca0db

Observation 08074c3b-45e3-4666-b1fb-f80f23217d01 · inbound

Structured Spectral Graph Learning for Anomaly Classification in 3D Chest CT Scans cites this paper.

Structured Spectral Graph Learning for Anomaly Classification in 3D Chest CT Scans Video Swin Transformer

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T05:58:27.648814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:58:27.648814Z digest=sha256:44142304abef05dece01ff6278acc4490c7591e00b9afe3ad44918514e7e87d1

Observation 33867df1-31d6-4f57-977a-8e13d95fb8cb · inbound

T-MASK: Temporal Masking for Probing Foundation Models across Camera Views in Driver Monitoring cites this paper.

T-MASK: Temporal Masking for Probing Foundation Models across Camera Views in Driver Monitoring Video Swin Transformer

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T17:29:03.792173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T17:29:03.792173Z digest=sha256:7720c865ffa4930cbefa6119161efd62ba05ade97b7161b28d726184511a4428

Observation 45a90db6-05b5-4f1c-aa47-26c4cf7f6668 · inbound

Every Subtlety Counts: Fine-grained Person Independence Micro-Action Recognition via Distributionally Robust Optimization cites this paper.

Every Subtlety Counts: Fine-grained Person Independence Micro-Action Recognition via Distributionally Robust Optimization Video Swin Transformer

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T13:41:25.582613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-18T13:39:40.050017Z digest=sha256:0b763a1434c670150c0a85b04cbbf552c704059c6cd39bfe5cf406b33af76a9f

Observation fca90970-7504-4e69-8218-5e7a1f3c5984 · inbound

Structured Spectral Graph Representation Learning for Multi-label Abnormality Analysis from 3D CT Scans cites this paper.

Structured Spectral Graph Representation Learning for Multi-label Abnormality Analysis from 3D CT Scans Video Swin Transformer

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T10:20:00.972236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:20:00.972236Z digest=sha256:dcaa0fffb1cec72d819551376533a5a4aac8d7ebdf4a9c904be5537da5b2b4ee

Observation dad599ad-5493-4212-86c6-1541f35a956a · inbound

RobustSora: De-Watermarked Benchmark for Robust AI-Generated Video Detection cites this paper.

RobustSora: De-Watermarked Benchmark for Robust AI-Generated Video Detection Video Swin Transformer

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T23:13:39.366912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T23:13:02.040618Z digest=sha256:1a9b1d0bf74dc2b9982d608523d152c3e6681790da3dc3929177c6de5fcaf42f

Observation b7182760-3804-4eba-a166-9ffc836ea0fc · inbound

Multimodal Anomaly Detection for Human-Robot Interaction cites this paper.

Multimodal Anomaly Detection for Human-Robot Interaction Video Swin Transformer

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T06:56:03.705513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T17:22:13.877296Z digest=sha256:ec661fd2b7282cc7e8f0540dca682659a2f2883f911312ad8170f6a4cdaa5fa8

Observation a2f356bf-7d7a-4703-b1fe-d420fcf31f3f · inbound

ConvFormer3D-TAP: Phase/Uncertainty-Aware Front-End Fusion for Cine CMR View Classification Pipelines cites this paper.

ConvFormer3D-TAP: Phase/Uncertainty-Aware Front-End Fusion for Cine CMR View Classification Pipelines Video Swin Transformer

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T09:36:02.132180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T15:56:32.221330Z digest=sha256:fbb0b0b39e96953eb1c9233830ff26452799d9581003053d4c139ec2fa313d91

Observation e485d4af-f542-482a-8768-5f3a1d18db1f · inbound

DVAR: Adversarial Multi-Agent Debate for Video Authenticity Detection cites this paper.

DVAR: Adversarial Multi-Agent Debate for Video Authenticity Detection Video Swin Transformer

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:52:13.661109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T07:51:04.580617Z digest=sha256:adf3ed44c0af40c4ddf1972b348f82fa16bb28b32ccd3a924e8b761ec5fae017

Observation 58128543-0302-4e1b-b32f-20cc9b8ad66a · inbound

SignMAE: Segmentation-Driven Self-Supervised Learning for Sign Language Recognition cites this paper.

SignMAE: Segmentation-Driven Self-Supervised Learning for Sign Language Recognition Video Swin Transformer

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-09T05:55:30.886907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T19:23:07.258575Z digest=sha256:ab7ff7209f13057d9ade21d1c97b75a7489b7fd024c8c674181971c5785b6d5d

Observation 4d1c6c00-887c-4f63-87b8-1ab91b108f48 · inbound

From Articulated Kinematics to Routed Visual Control for Action-Conditioned Surgical Video Generation cites this paper.

From Articulated Kinematics to Routed Visual Control for Action-Conditioned Surgical Video Generation Video Swin Transformer

Reference 54

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T03:16:18.777601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-12T03:15:40.944411Z digest=sha256:4205e21e0745f72ac7775446f60e66bd8b864fd7ae1e9ae61f03d556638f72cc

Observation 4fd0229e-93a6-4b99-b806-09946d6b93b6 · inbound

MegaScale-Omni: A Hyper-Scale, Workload-Resilient System for MultiModal LLM Training in Production cites this paper.

MegaScale-Omni: A Hyper-Scale, Workload-Resilient System for MultiModal LLM Training in Production Video Swin Transformer

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-12T02:06:15.037782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-12T02:04:07.344134Z digest=sha256:4b779b3a8d3b1013fe0d21e178f17d137d008a7790a5bedbf5dd835097f6aeaf

Observation 6a28b164-8843-4293-bf48-29a75376b793 · inbound

CAM-VFD: Cross-Attention Multimodal Video Forgery Detection cites this paper.

CAM-VFD: Cross-Attention Multimodal Video Forgery Detection Video Swin Transformer

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T15:13:25.065817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T15:08:25.309094Z digest=sha256:4bb310b873566561f68cce662e13adbc33e162cd763f556ad495ad62191cc6d9

Observation 5aaa64a0-7d81-4001-be48-c7f848a98c88 · inbound

Tensor Memory: Fixed-Size Recurrent State for Long-Horizon Transformers cites this paper.

Tensor Memory: Fixed-Size Recurrent State for Long-Horizon Transformers Video Swin Transformer

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:13:48.726915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-29T18:08:42.533804Z digest=sha256:5b18bd4f09f2682bfa87841bcb48558c9268e95dea0d117463eb1b595d8f54b2

Observation 66797a37-6d28-4a73-99d2-a0a9d73ece4d · inbound

VTI-CoT: Visual-Textual Interleaved Chain of Thought for Video Reasoning cites this paper.

VTI-CoT: Visual-Textual Interleaved Chain of Thought for Video Reasoning Video Swin Transformer

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:46:56.768986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T01:52:44.785582Z digest=sha256:49ea985dbb1b99f5b62d750bf2f13063fc4409f71a98777e558851ddf0b82f69

Observation b2f6fa34-a672-4bb2-8fe8-f559dbbf406b · inbound

MLT-Dedup: Efficient Large-Scale Online Video Deduplication via Multi-Level Representations and Spatial-Temporal Matching cites this paper.

MLT-Dedup: Efficient Large-Scale Online Video Deduplication via Multi-Level Representations and Spatial-Temporal Matching Video Swin Transformer

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T10:58:03.528046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-27T09:43:22.054789Z digest=sha256:b5f23ffed37bc6a507c8af516763a0f6e1aa879f401cd21074e26a313484a1c4

Observation aefef589-a3d6-4877-a7fd-db1cbd7fc10e · inbound

Spatio-Temporal Fusion Model for Standard View Classification of Echocardiographic Videos cites this paper.

Spatio-Temporal Fusion Model for Standard View Classification of Echocardiographic Videos Video Swin Transformer

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-07-03T19:08:49.805619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-27T02:11:37.740576Z digest=sha256:5057d1da3f82a05beb9a84bdeabd56eedba8f898f18bbaf4095f9363d19613be

Observation c0767247-40c6-4882-a33a-79aed324b7e0 · inbound

Spatio-Temporal Wildfire Spread Prediction in Canada using a Video Swin-Hybrid-U-Net and Satellite Imagery cites this paper.

Spatio-Temporal Wildfire Spread Prediction in Canada using a Video Swin-Hybrid-U-Net and Satellite Imagery Video Swin Transformer

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-03T17:38:43.730098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-27T03:59:57.906013Z digest=sha256:32e1139b36a00ceaab92fb3a8233e70535fd77b0fc6f36f07fc85010d786acce

Observation c8c7c708-419f-498a-94fa-83116a766446 · inbound

GMoT: Gated Motion-Aware Tokenization for Fine-Grained Micro-Gesture Video Reasoning with Multimodal LLMs cites this paper.

GMoT: Gated Motion-Aware Tokenization for Fine-Grained Micro-Gesture Video Reasoning with Multimodal LLMs Video Swin Transformer

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T04:39:23.642262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:39:23.642262Z digest=sha256:e75ff76011a0fd72fc85c26525ceedcac4c815d82afed7c5c571ba9eb4c6ad3c