Pith. sign in

Paper Citation Record · LEDGER

UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 18 inbound Pith citation observations for arXiv:2211.09552.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2211.09552 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:37:36.711194Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T19:08:49.809225Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation df6acf2d-ce08-4040-9031-ec1f5337f94f · inbound

InternVideo: General Video Foundation Models via Generative and Discriminative Learning cites this paper.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:36:53.339682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:ef7db6fbf5e9fe2ee4bd12b486caaf7da887d9a8feb1fee24659933313ff38e3

Observation a5b874da-8fef-43d5-8329-ae9d9f30a5a9 · inbound

VideoChat: Chat-Centric Video Understanding cites this paper.

VideoChat: Chat-Centric Video Understanding UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:30:00.561864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T23:30:00.457974Z digest=sha256:af75b2d5f58d86a74f2d4226b45eca5425a15eb8c152ed3217bfda17be063049

Observation 82c9ad0b-52d7-4917-8e8f-d6e77fa45beb · inbound

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation cites this paper.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:30:22.561008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:6eb94daf948a2740c9da9e58d42663fe8f3983abb27e3acf75ff8fbc66c647fd

Observation eb5d0769-acbf-4ce3-9e46-8105f8c30d1f · inbound

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark cites this paper.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:22:35.091018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:60be4234786c4b438d5fffa44b5098a2cc5e90a9df974cc78cc7973205cf16da

Observation 215434de-9888-454e-8479-a6ce6fd57807 · inbound

CogVLM2: Visual Language Models for Image and Video Understanding cites this paper.

CogVLM2: Visual Language Models for Image and Video Understanding UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-16T20:10:27.735657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:1fb30daa40601b3c0faca99bf2198c649d7f54d5d31ab98ced816e8f3be2fc29

Observation 79847b2b-b338-4975-b4fe-70ff0311c6c4 · inbound

MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models cites this paper.

MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-23T05:45:28.360703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T05:44:31.546843Z digest=sha256:80439b4c4462a7792578676e77aa0951162967fe5f8a8b53ee53758c4e7c3df3

Observation 2a772835-48f9-4514-9aab-aaccfa86327c · inbound

IPFormer-VideoLLM: Enhancing Multi-modal Video Understanding for Multi-shot Scenes cites this paper.

IPFormer-VideoLLM: Enhancing Multi-modal Video Understanding for Multi-shot Scenes UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T22:37:36.711194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:37:36.711194Z digest=sha256:a28ac98f49c3a500f852e5c6b2c7d378a4442dfe81b057a906ec8e08f1defab2

Observation 55fee3db-b3f0-49d2-a136-629b6b46e647 · inbound

Uncertainty-aware Diffusion and Reinforcement Learning for Joint Plane Localization and Anomaly Diagnosis in 3D Ultrasound cites this paper.

Uncertainty-aware Diffusion and Reinforcement Learning for Joint Plane Localization and Anomaly Diagnosis in 3D Ultrasound UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T21:44:13.171863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:44:13.171863Z digest=sha256:940a990ffeb85ccf7b96cf030b4766a9b6c5bd5567b7105c15ce7e4921fe07ef

Observation f1a474c0-6c44-45d3-9420-44286c7b83b3 · inbound

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning cites this paper.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:19.646948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:19.646948Z digest=sha256:45dc5ad2aa867c65e9ba9f494f86206d0898293e7ecc2ce3b56786dd59d9715d

Observation 93f59791-a2f0-4a2b-8256-9388ba18df40 · inbound

ConvFormer3D-TAP: Phase/Uncertainty-Aware Front-End Fusion for Cine CMR View Classification Pipelines cites this paper.

ConvFormer3D-TAP: Phase/Uncertainty-Aware Front-End Fusion for Cine CMR View Classification Pipelines UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:36:02.125281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:56:32.221330Z digest=sha256:55fa3708e947e701a95423635e81f911925562c3481feb9862226fc45194b721

Observation b8ab7530-db06-4175-9376-079e41b6119b · inbound

V-Nutri: Dish-Level Nutrition Estimation from Egocentric Cooking Videos cites this paper.

V-Nutri: Dish-Level Nutrition Estimation from Egocentric Cooking Videos UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:21:02.439765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:00:21.173667Z digest=sha256:24bf5a1f973974d5b58ee30c307cd7c2d95fd34ed475d373a2b9d66914353620

Observation 0db39fe4-456c-482c-a3ac-7d217f11adb4 · inbound

NeuroLip: An Event-driven Spatiotemporal Learning Framework for Cross-Scene Lip-Motion-based Visual Speaker Recognition cites this paper.

NeuroLip: An Event-driven Spatiotemporal Learning Framework for Cross-Scene Lip-Motion-based Visual Speaker Recognition UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:28:38.693323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T09:27:32.335117Z digest=sha256:56964ecb7cb665e92038a6b70a9831bac4e573f9f67c6f9959709be0e5cbe222

Observation 06a5d465-ba89-40c1-9c3e-bc18cd7bca3f · inbound

DVAR: Adversarial Multi-Agent Debate for Video Authenticity Detection cites this paper.

DVAR: Adversarial Multi-Agent Debate for Video Authenticity Detection UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:52:13.618993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T07:51:04.580617Z digest=sha256:bdc620e9ddb04b88c3926712f14a25f33a655515054703a9eacdbc3a5a7ce3be

Observation ca830749-d822-408a-975e-5b5946051a1d · inbound

CAM-VFD: Cross-Attention Multimodal Video Forgery Detection cites this paper.

CAM-VFD: Cross-Attention Multimodal Video Forgery Detection UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-20T15:13:25.074915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T15:08:25.309094Z digest=sha256:f0cb7359d6b530745a1f3975ad16fe5613e87d0d466d36116d83e341fbe4a4ed

Observation 6ceaa5fe-b749-42d1-a426-f45c7a9b4b9d · inbound

Spatio-Temporal Fusion Model for Standard View Classification of Echocardiographic Videos cites this paper.

Spatio-Temporal Fusion Model for Standard View Classification of Echocardiographic Videos UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-07-03T19:08:49.810953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T02:11:37.740576Z digest=sha256:0d218b8f2ebb9080f0bd4eab7834d018a6fbc89a17ade47bb5b17f6b8ac611cb

Observation 09b3d675-5426-4733-88da-c9b72119243f · inbound

Rethinking the Readout: Unlocking Video Backbones for AI-Generated Video Detection cites this paper.

Rethinking the Readout: Unlocking Video Backbones for AI-Generated Video Detection UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T01:25:16.561201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:25:16.561201Z digest=sha256:4786ac62345abb1e37f74776559e393a995c52aad19cbb602107d4ef897d3f6b

Observation 1661bf28-b79b-4f43-941b-0ed9a937da40 · inbound

PhiZero: A World Model Built Around Physical Language cites this paper.

PhiZero: A World Model Built Around Physical Language UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-31T01:50:30.642683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T01:50:30.642683Z digest=sha256:9d3836d45563c514df2962d69d645e1cdfe141f0496a223da8c3ebf58bacd45a

Observation 623012d6-b50d-49de-be5d-6ead28915273 · inbound

Decoding Children's Gait Behavior cites this paper.

Decoding Children's Gait Behavior UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-04T00:41:55.526214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:41:55.526214Z digest=sha256:8a5e315f1055611ea1616eaecd368dc86cc0b8cdf6d29db57accf938aba9f286