Pith. sign in

Paper Citation Record · LEDGER

Is Space-Time Attention All You Need for Video Understanding?

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 25 inbound Pith citation observations for arXiv:2102.05095.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2102.05095 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 25 of 25 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:55:28.412053Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1359
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 140aac23-27fa-41c3-8ea3-ac9d6efb6e82 · inbound

Video Diffusion Models cites this paper.

Video Diffusion Models Is Space-Time Attention All You Need for Video Understanding?

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-13T14:38:27.964923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T14:38:27.919104Z digest=sha256:1901b334571fc0e7e7a5f3db82cb450ce5f6683f0fea3ce756f22dcfe8bf369b

Observation 5ab418a0-6bfc-4f97-990e-a8b928f6f186 · inbound

MOOSE: Pay Attention to Temporal Dynamics for Video Understanding via Optical Flows cites this paper.

MOOSE: Pay Attention to Temporal Dynamics for Video Understanding via Optical Flows Is Space-Time Attention All You Need for Video Understanding?

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:28.412053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:28.412053Z digest=sha256:0d699e06014a0630d58b10c9f77f3d765a3455d08b94dcbf992e410f124fcdb8

Observation eeb68259-9d9e-44d4-b995-f92cc6f0856e · inbound

Fine-Tuning Video Transformers for Word-Level Bangla Sign Language: A Comparative Analysis for Classification Tasks cites this paper.

Fine-Tuning Video Transformers for Word-Level Bangla Sign Language: A Comparative Analysis for Classification Tasks Is Space-Time Attention All You Need for Video Understanding?

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T10:46:14.159369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:46:14.159369Z digest=sha256:ef10dc7b9d3b063cdc910fe9b355bd28350b4749eddd5779500071f95a5b2675

Observation 9e0c1bef-2930-49ac-872f-316afaa219da · inbound

Data-Efficient Challenges in Visual Inductive Priors: A Retrospective cites this paper.

Data-Efficient Challenges in Visual Inductive Priors: A Retrospective Is Space-Time Attention All You Need for Video Understanding?

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:10:59.021704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:10:59.021704Z digest=sha256:68d2f7f7841da0f964459142dcd91bb335b3a2e6007035a895187b8441c0df9b

Observation a890a18d-7361-412a-8960-615401f929aa · inbound

Vision Transformer-Based Time-Series Image Reconstruction for Cloud-Filling Applications cites this paper.

Vision Transformer-Based Time-Series Image Reconstruction for Cloud-Filling Applications Is Space-Time Attention All You Need for Video Understanding?

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-19T08:02:10.626908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T07:59:15.937439Z digest=sha256:b93de89d2f4bca7ce3c6f2453ffbc44f01eac54b28c800fd8745934bb18a988f

Observation bd7e80c6-f1a6-4542-be12-6e0b683c42a7 · inbound

Comparing Learning Paradigms for Egocentric Video Summarization cites this paper.

Comparing Learning Paradigms for Egocentric Video Summarization Is Space-Time Attention All You Need for Video Understanding?

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T22:23:41.597659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:23:41.597659Z digest=sha256:55d45ef912dd4492e53bd5517139c0e30cf29d58048291bbc0d18d3ebeadbdbf

Observation 1c41e3d9-773f-4c9d-bf1e-460f39b99941 · inbound

MVP: Winning Solution to SMP Challenge 2025 Video Track cites this paper.

MVP: Winning Solution to SMP Challenge 2025 Video Track Is Space-Time Attention All You Need for Video Understanding?

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:56.108327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:56.108327Z digest=sha256:bb58cdb00bfdfc89541a70377a93ebe38f7b46fd6907daa6031c86b73783d559

Observation d9c232d4-c6b1-404e-8f17-5e6da800d8ae · inbound

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes cites this paper.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Is Space-Time Attention All You Need for Video Understanding?

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T17:13:01.481907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:13:01.481907Z digest=sha256:b2562560e041fb196bad95f4ec221482581703318465d923030e42f7935f4aa5

Observation 82bce4bf-2976-4f3a-a159-93b903fc27f8 · inbound

CascadeFormer: A Family of Two-stage Cascading Transformers for Skeleton-based Human Action Recognition cites this paper.

CascadeFormer: A Family of Two-stage Cascading Transformers for Skeleton-based Human Action Recognition Is Space-Time Attention All You Need for Video Understanding?

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:13.584772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:13.584772Z digest=sha256:523b3aa4a94fb3dd9cc88bd6bf4a02e115bf80e4a2232b41d7837dfcca4c5d8b

Observation 2a5d2dfb-a2a1-4b8a-bbeb-84eca88c0abd · inbound

Cataract-LMM Large-Scale Multi-Source Multi-Task Benchmark for Deep Learning in Surgical Video Analysis cites this paper.

Cataract-LMM Large-Scale Multi-Source Multi-Task Benchmark for Deep Learning in Surgical Video Analysis Is Space-Time Attention All You Need for Video Understanding?

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:40:55.918865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T05:37:51.737038Z digest=sha256:ebc0432535c652791f88a0af5a696f5df94d05b3198c6cf11ed70664c2322bdc

Observation 76e0ead7-bfbe-4fa7-9d3f-13c125b192eb · inbound

A Space-Time Transformer for Precipitation Nowcasting cites this paper.

A Space-Time Transformer for Precipitation Nowcasting Is Space-Time Attention All You Need for Video Understanding?

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T22:19:23.932491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:19:23.932491Z digest=sha256:ef59974accc44ccbb56052aab6e646f7a81c9d5b93c0ed8b85025bd3e78aac11

Observation 45670124-a771-4715-b6df-cf104cb965f5 · inbound

RobustSora: De-Watermarked Benchmark for Robust AI-Generated Video Detection cites this paper.

RobustSora: De-Watermarked Benchmark for Robust AI-Generated Video Detection Is Space-Time Attention All You Need for Video Understanding?

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T23:13:39.376251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T23:13:02.040618Z digest=sha256:0518e94894d992f228bc54425df4e4b002d69de83563d5e0ed0425d23d4ebbff

Observation 0c425cb9-266e-4991-aa6e-391b8754c243 · inbound

Explainable Fall Detection for Elderly Monitoring via Temporally Stable SHAP in Skeleton-Based Human Activity Recognition cites this paper.

Explainable Fall Detection for Elderly Monitoring via Temporally Stable SHAP in Skeleton-Based Human Activity Recognition Is Space-Time Attention All You Need for Video Understanding?

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-10T15:55:33.481242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:55:17.256153Z digest=sha256:c89d517e59340a555121f3005cdf77d79ad01d0a93119031e20e38c1afca33cf

Observation 8fd48442-8300-4968-afb9-7ebebff1f7dc · inbound

Seeing Further and Wider: Joint Spatio-Temporal Enlargement for Micro-Video Popularity Prediction cites this paper.

Seeing Further and Wider: Joint Spatio-Temporal Enlargement for Micro-Video Popularity Prediction Is Space-Time Attention All You Need for Video Understanding?

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-09T23:14:36.516021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-09T23:13:34.852488Z digest=sha256:69da8ccc7a93a171d6e8e1dae587aba25523710712144d4327a225716c61190f

Observation 8a4b40fe-1be0-441a-bd1e-de4a4bf32084 · inbound

Exploring High-Order Self-Similarity for Video Understanding cites this paper.

Exploring High-Order Self-Similarity for Video Understanding Is Space-Time Attention All You Need for Video Understanding?

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:59:49.399058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T00:59:03.890135Z digest=sha256:0f3cc067116f8886008e1461dad973d001ee97822cd221d46467503288470983

Observation 539c8429-561d-4ffc-b7fa-1f728c33144f · inbound

Only Brains Align with Brains: Cross-Region Alignment Patterns Expose Limits of Normative Models cites this paper.

Only Brains Align with Brains: Cross-Region Alignment Patterns Expose Limits of Normative Models Is Space-Time Attention All You Need for Video Understanding?

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:56:07.521561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T13:13:03.129195Z digest=sha256:fd6660cd177e3cdd66eb92daae2297558b23e654b61a072a3e0a4ee3fc22a86e

Observation 6d788f70-2247-45f5-90fb-baff01f58a14 · inbound

Parameter-Efficient Multi-View Proficiency Estimation: From Discriminative Classification to Generative Feedback cites this paper.

Parameter-Efficient Multi-View Proficiency Estimation: From Discriminative Classification to Generative Feedback Is Space-Time Attention All You Need for Video Understanding?

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:16:24.059130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-07T17:44:28.109134Z digest=sha256:154e316451aceaa385366d909bee59a63803e06331e908cd4a7c220c91f384b9

Observation e67b4d71-56db-49d5-85a0-7d6e499a8a29 · inbound

LookWhen? Fast Video Recognition by Learning When, Where, and What to Compute cites this paper.

LookWhen? Fast Video Recognition by Learning When, Where, and What to Compute Is Space-Time Attention All You Need for Video Understanding?

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:40:52.109484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T01:36:50.406353Z digest=sha256:542b620ad798f90729637820283f1bde49df34cc0690c360a92511b9e8e5dc99

Observation edb86348-ec6e-48fb-9f45-737820589425 · inbound

CAM-VFD: Cross-Attention Multimodal Video Forgery Detection cites this paper.

CAM-VFD: Cross-Attention Multimodal Video Forgery Detection Is Space-Time Attention All You Need for Video Understanding?

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-20T15:13:25.080824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T15:08:25.309094Z digest=sha256:a4859dd21f661da726bcfbed55412c5c193e13b527dd52d225b9f0d1afcb2c21

Observation a2303355-69b9-4e82-a690-322acc4b77b5 · inbound

Tensor Memory: Fixed-Size Recurrent State for Long-Horizon Transformers cites this paper.

Tensor Memory: Fixed-Size Recurrent State for Long-Horizon Transformers Is Space-Time Attention All You Need for Video Understanding?

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:13:48.619377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T18:08:42.533804Z digest=sha256:336ac77948a0c79d21285da5afac02446183b3a78cc56f9ae83893c3c8921f3e

Observation daa30a4d-0e4f-4095-b886-38e9453a7445 · inbound

Signed Dual Attention: Capturing Signed Dependencies in Time Series Forecasting cites this paper.

Signed Dual Attention: Capturing Signed Dependencies in Time Series Forecasting Is Space-Time Attention All You Need for Video Understanding?

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-02T07:26:45.540511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T07:00:04.541121Z digest=sha256:f6a1d3977266f0ea316c1a8677fad3018eb7669f2a4c64ba97a30c2bc0429ede

Observation 681dc00c-6bde-453d-a614-ff4caba6f2f5 · inbound

A multi-task spatiotemporal deep neural network for predicting penetration depth and morphology in laser welding cites this paper.

A multi-task spatiotemporal deep neural network for predicting penetration depth and morphology in laser welding Is Space-Time Attention All You Need for Video Understanding?

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-06-26T01:38:50.693618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T01:37:58.322339Z digest=sha256:de81ccde8da6049bac99e564c952f6101e0a8c6949943ba8210f8637011b4643

Observation 520529b3-782c-46a2-a61c-cdd8a937fe75 · inbound

Incentivizing Vision Language Models to Search for Long Video Question Answering cites this paper.

Incentivizing Vision Language Models to Search for Long Video Question Answering Is Space-Time Attention All You Need for Video Understanding?

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-12T05:50:16.895740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T05:50:16.895740Z digest=sha256:0b6bf432db089f006d4937fed429c618d95ab4a3d83bdc3dc61fe52b2927587f

Observation 93a2de9b-9a1a-4845-874b-f4b23582ace7 · inbound

The evolution of AI from image interpretation toward scientific inference in nanoparticle electron microscopy cites this paper.

The evolution of AI from image interpretation toward scientific inference in nanoparticle electron microscopy Is Space-Time Attention All You Need for Video Understanding?

Reference 104

Resolution
unresolved
no resolver link, observed 2026-07-14T12:06:28.342674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:06:28.342674Z digest=sha256:102fd99ae9e6ebbe48fc4e8d8c8bb4843f2df1fed04c410211e1b23d8ccead91

Observation ba800377-3a00-4419-9e01-e6dda56e0606 · inbound

GMoT: Gated Motion-Aware Tokenization for Fine-Grained Micro-Gesture Video Reasoning with Multimodal LLMs cites this paper.

GMoT: Gated Motion-Aware Tokenization for Fine-Grained Micro-Gesture Video Reasoning with Multimodal LLMs Is Space-Time Attention All You Need for Video Understanding?

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T04:39:19.873848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:39:19.873848Z digest=sha256:7901be98f286b4ec9c25cbf7237e82996c2aa2b4e2b1e1c7ebfe9abc79f2cc92