Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 18 inbound Pith citation observations for arXiv:2211.09552.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T22:37:36.711194Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T19:08:49.809225Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation df6acf2d-ce08-4040-9031-ec1f5337f94f · inbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a5b874da-8fef-43d5-8329-ae9d9f30a5a9 · inbound
VideoChat: Chat-Centric Video Understanding UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 82c9ad0b-52d7-4917-8e8f-d6e77fa45beb · inbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation eb5d0769-acbf-4ce3-9e46-8105f8c30d1f · inbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 215434de-9888-454e-8479-a6ce6fd57807 · inbound
CogVLM2: Visual Language Models for Image and Video Understanding UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 79847b2b-b338-4975-b4fe-70ff0311c6c4 · inbound
MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2a772835-48f9-4514-9aab-aaccfa86327c · inbound
IPFormer-VideoLLM: Enhancing Multi-modal Video Understanding for Multi-shot Scenes UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55fee3db-b3f0-49d2-a136-629b6b46e647 · inbound
Uncertainty-aware Diffusion and Reinforcement Learning for Joint Plane Localization and Anomaly Diagnosis in 3D Ultrasound UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1a474c0-6c44-45d3-9420-44286c7b83b3 · inbound
OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93f59791-a2f0-4a2b-8256-9388ba18df40 · inbound
ConvFormer3D-TAP: Phase/Uncertainty-Aware Front-End Fusion for Cine CMR View Classification Pipelines UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b8ab7530-db06-4175-9376-079e41b6119b · inbound
V-Nutri: Dish-Level Nutrition Estimation from Egocentric Cooking Videos UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0db39fe4-456c-482c-a3ac-7d217f11adb4 · inbound
NeuroLip: An Event-driven Spatiotemporal Learning Framework for Cross-Scene Lip-Motion-based Visual Speaker Recognition UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 06a5d465-ba89-40c1-9c3e-bc18cd7bca3f · inbound
DVAR: Adversarial Multi-Agent Debate for Video Authenticity Detection UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ca830749-d822-408a-975e-5b5946051a1d · inbound
CAM-VFD: Cross-Attention Multimodal Video Forgery Detection UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6ceaa5fe-b749-42d1-a426-f45c7a9b4b9d · inbound
Spatio-Temporal Fusion Model for Standard View Classification of Echocardiographic Videos UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 09b3d675-5426-4733-88da-c9b72119243f · inbound
Rethinking the Readout: Unlocking Video Backbones for AI-Generated Video Detection UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1661bf28-b79b-4f43-941b-0ed9a937da40 · inbound
PhiZero: A World Model Built Around Physical Language UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 623012d6-b50d-49de-be5d-6ead28915273 · inbound
Decoding Children's Gait Behavior UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.