Pith. sign in

Paper Citation Record · LEDGER

UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 21 inbound Pith citation observations for arXiv:2211.09552.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2211.09552 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 21 of 21 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:46:52.113307Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T19:08:49.809225Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation df6acf2d-ce08-4040-9031-ec1f5337f94f · inbound

InternVideo: General Video Foundation Models via Generative and Discriminative Learning cites this paper.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:36:53.339682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:35f63743279b1d699302ebc389e685563fd6f5d20fa33c941edba85393d9ff25

Observation a5b874da-8fef-43d5-8329-ae9d9f30a5a9 · inbound

VideoChat: Chat-Centric Video Understanding cites this paper.

VideoChat: Chat-Centric Video Understanding UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:30:00.561864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T23:30:00.457974Z digest=sha256:202d4291bff20f39a86a20bc5f698b460a31940c677f55c79baaf03b052e47f2

Observation 82c9ad0b-52d7-4917-8e8f-d6e77fa45beb · inbound

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation cites this paper.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:30:22.561008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:d8faa498cb0c1e5c0740b765fdaf5f0418d41e2b25cff4579304d83095a9af43

Observation eb5d0769-acbf-4ce3-9e46-8105f8c30d1f · inbound

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark cites this paper.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:22:35.091018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:8f72c457caa14a545b2637aabb26a6647385c9f98544856986c95187fb1f42bc

Observation 215434de-9888-454e-8479-a6ce6fd57807 · inbound

CogVLM2: Visual Language Models for Image and Video Understanding cites this paper.

CogVLM2: Visual Language Models for Image and Video Understanding UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-16T20:10:27.735657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:774698269f259d08892723a387d3bd72b6aba6a35be6051d237eb38d1f9e9d74

Observation 79847b2b-b338-4975-b4fe-70ff0311c6c4 · inbound

MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models cites this paper.

MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-23T05:45:28.360703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T05:44:31.546843Z digest=sha256:cb7605fac017988854c432a9830f10a6ff69e09b99e5de67eff6bb614cba3221

Observation f966cc04-5513-4b7b-ab90-7dd8676b441e · inbound

CA3D: Convolutional-Attentional 3D Nets for Efficient Video Activity Recognition on the Edge cites this paper.

CA3D: Convolutional-Attentional 3D Nets for Efficient Video Activity Recognition on the Edge UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:09:32.639342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:09:32.639342Z digest=sha256:d84795065385984fab7474122d7c405c603058fe620a869b2d269036066174c8

Observation 64e43334-38e5-4c57-8386-dc24cde323e4 · inbound

Large-scale Self-supervised Video Foundation Model for Intelligent Surgery cites this paper.

Large-scale Self-supervised Video Foundation Model for Intelligent Surgery UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T11:22:59.602264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:22:59.602264Z digest=sha256:6d1f27cf61b7f1c1dc094f3d9d0766b2650ca5ee01151c0e6dfe332d68e69960

Observation 0df613e5-a734-41fe-955c-6a86b907c704 · inbound

FuseMamba-VD: Dual Branch VideoMamba with Gated Class Token Fusion for Violence Detection cites this paper.

FuseMamba-VD: Dual Branch VideoMamba with Gated Class Token Fusion for Violence Detection UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:46:52.113307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:46:52.113307Z digest=sha256:c29b9c731483c5b3abe399ca14a6e10dbd1fcc3f6f216922e28914dfd59e6497

Observation 2a772835-48f9-4514-9aab-aaccfa86327c · inbound

IPFormer-VideoLLM: Enhancing Multi-modal Video Understanding for Multi-shot Scenes cites this paper.

IPFormer-VideoLLM: Enhancing Multi-modal Video Understanding for Multi-shot Scenes UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T22:37:36.711194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:37:36.711194Z digest=sha256:31c518efaf1919128bfba197058e22b97b720fa57dbc5f6b5372d83e3f997f46

Observation 55fee3db-b3f0-49d2-a136-629b6b46e647 · inbound

Uncertainty-aware Diffusion and Reinforcement Learning for Joint Plane Localization and Anomaly Diagnosis in 3D Ultrasound cites this paper.

Uncertainty-aware Diffusion and Reinforcement Learning for Joint Plane Localization and Anomaly Diagnosis in 3D Ultrasound UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T21:44:13.171863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:44:13.171863Z digest=sha256:dcf1b85e55f9f924de82b962fddf39287942d514777135508e444ce5d683c1cd

Observation f1a474c0-6c44-45d3-9420-44286c7b83b3 · inbound

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning cites this paper.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:19.646948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:19.646948Z digest=sha256:2e54457775bacc1cf62066793b841bc141f20d4364878e886f0d2265a23c50c7

Observation 93f59791-a2f0-4a2b-8256-9388ba18df40 · inbound

ConvFormer3D-TAP: Phase/Uncertainty-Aware Front-End Fusion for Cine CMR View Classification Pipelines cites this paper.

ConvFormer3D-TAP: Phase/Uncertainty-Aware Front-End Fusion for Cine CMR View Classification Pipelines UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:36:02.125281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T15:56:32.221330Z digest=sha256:7111cb1affb4df0c7ecadfd89d7d13bee5ee886b89ad2ea50d17d52f0a954819

Observation b8ab7530-db06-4175-9376-079e41b6119b · inbound

V-Nutri: Dish-Level Nutrition Estimation from Egocentric Cooking Videos cites this paper.

V-Nutri: Dish-Level Nutrition Estimation from Egocentric Cooking Videos UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:21:02.439765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T15:00:21.173667Z digest=sha256:39733bd7c854adea10a3b7b5981ac57c02a290796c61976383790a390706565a

Observation 0db39fe4-456c-482c-a3ac-7d217f11adb4 · inbound

NeuroLip: An Event-driven Spatiotemporal Learning Framework for Cross-Scene Lip-Motion-based Visual Speaker Recognition cites this paper.

NeuroLip: An Event-driven Spatiotemporal Learning Framework for Cross-Scene Lip-Motion-based Visual Speaker Recognition UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:28:38.693323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T09:27:32.335117Z digest=sha256:2393590721f3892af876c3d9e4638c83063c5594de36034a5159611ccea228ba

Observation 06a5d465-ba89-40c1-9c3e-bc18cd7bca3f · inbound

DVAR: Adversarial Multi-Agent Debate for Video Authenticity Detection cites this paper.

DVAR: Adversarial Multi-Agent Debate for Video Authenticity Detection UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:52:13.618993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T07:51:04.580617Z digest=sha256:c1483776af801be8361360c53a9dc4fbf0aee13497397bf12eb732c239101f07

Observation ca830749-d822-408a-975e-5b5946051a1d · inbound

CAM-VFD: Cross-Attention Multimodal Video Forgery Detection cites this paper.

CAM-VFD: Cross-Attention Multimodal Video Forgery Detection UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-20T15:13:25.074915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T15:08:25.309094Z digest=sha256:3f69f859638a31da59910447a37ceb7dd644f757a6767a943327bb03969c21d2

Observation 6ceaa5fe-b749-42d1-a426-f45c7a9b4b9d · inbound

Spatio-Temporal Fusion Model for Standard View Classification of Echocardiographic Videos cites this paper.

Spatio-Temporal Fusion Model for Standard View Classification of Echocardiographic Videos UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-07-03T19:08:49.810953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T02:11:37.740576Z digest=sha256:d5c0e752dfd30bbcb49514407ba3d6db0435441f6925c0c050c2ab5a4821be2a

Observation 09b3d675-5426-4733-88da-c9b72119243f · inbound

Rethinking the Readout: Unlocking Video Backbones for AI-Generated Video Detection cites this paper.

Rethinking the Readout: Unlocking Video Backbones for AI-Generated Video Detection UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T01:25:16.561201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:25:16.561201Z digest=sha256:e1c02e153b62714058479fa279636779a4613d569f513720c561461f99ea8b22

Observation 1661bf28-b79b-4f43-941b-0ed9a937da40 · inbound

PhiZero: A World Model Built Around Physical Language cites this paper.

PhiZero: A World Model Built Around Physical Language UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-31T01:50:30.642683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T01:50:30.642683Z digest=sha256:5b487d48a5bba1e359889340c1c18d1a4e6280738ab8fc9114c72c2ec9fe2753

Observation 623012d6-b50d-49de-be5d-6ead28915273 · inbound

Decoding Children's Gait Behavior cites this paper.

Decoding Children's Gait Behavior UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-04T00:41:55.526214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:41:55.526214Z digest=sha256:fc6df8753521f58507da2c9b87be12bb26ca58078bb8108ab75d228838ce44c8