Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T16:38:09.895402Z
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 0 inbound Pith citation observations for arXiv:2411.13626.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T16:38:09.895402Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
54 of 54 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 5a99891b-2e49-4489-8005-51c9dc53d51a · outbound
Principles of Visual Tokens for Efficient Video Understanding Vivit: A video vi- sion transformer
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation c9348b78-6a67-4449-bb64-609cc30b4912 · outbound
Principles of Visual Tokens for Efficient Video Understanding Is Space-Time Attention All You Need for Video Understanding?
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49cabee7-aefe-4090-8476-4ad7f9bdc92f · outbound
Principles of Visual Tokens for Efficient Video Understanding Is space-time attention all you need for video understanding? In ICML, page 4, 2021
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81fd44e9-98bc-428e-9fa7-c3e18c494187 · outbound
Principles of Visual Tokens for Efficient Video Understanding Token merging: Your ViT but faster
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 42b7f1f4-3f60-4fac-a742-73a4fe5dc6ef · outbound
Principles of Visual Tokens for Efficient Video Understanding Revisiting the” video” in video-language understanding
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 9651a4fb-f220-45f0-8ef8-75e05b8d0472 · outbound
Principles of Visual Tokens for Efficient Video Understanding Space-time mixing attention for video transformer
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation fabe6ce3-449a-4b92-b86a-857a4caa921e · outbound
Principles of Visual Tokens for Efficient Video Understanding Activitynet: A large-scale video benchmark for human activity understanding
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9612554-9565-471e-920a-d940a518b9df · outbound
Principles of Visual Tokens for Efficient Video Understanding Quo vadis, action recognition? a new model and the kinetics dataset
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation a852fa36-f5f0-458f-bc78-b053c27ab8f5 · outbound
Principles of Visual Tokens for Efficient Video Understanding An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models, 2024
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 25b26d8c-8dc3-44cb-bad3-b381833ec000 · outbound
Principles of Visual Tokens for Efficient Video Understanding Unresolved cited work
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 47cf182b-374d-48fe-93a6-9457ea518760 · outbound
Principles of Visual Tokens for Efficient Video Understanding Prune spatio-temporal tokens by semantic-aware temporal accumulation
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ff94fd5d-3e5c-40e1-83de-4cce007a8738 · outbound
Principles of Visual Tokens for Efficient Video Understanding An image is worth 16x16 words: Trans- formers for image recognition at scale
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b345fdf-3da9-4331-a4cc-dd588465739c · outbound
Principles of Visual Tokens for Efficient Video Understanding Multiscale vision transformers
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 48f3a647-1921-41d7-8c5e-859bf389090c · outbound
Principles of Visual Tokens for Efficient Video Understanding X3d: Expanding architectures for efficient video recognition
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 0bd73605-a781-43b2-ac83-0abdbf5e436a · outbound
Principles of Visual Tokens for Efficient Video Understanding Slowfast networks for video recognition
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7ea3224-d89b-494d-a436-a9a769696a5a · outbound
Principles of Visual Tokens for Efficient Video Understanding Efficient video transformers via spatial-temporal token merging for action recognition
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ec917cda-c365-402b-a165-4e86a39f60ab · outbound
Principles of Visual Tokens for Efficient Video Understanding Smart frame selection for action recognition
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3eeda23-3d4a-4dea-9fd3-c64001b2df71 · outbound
Principles of Visual Tokens for Efficient Video Understanding Watt For What: Rethinking Deep Learning's Energy-Performance Relationship
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb986a7b-b58f-402c-a25d-1fc5628338a2 · outbound
Principles of Visual Tokens for Efficient Video Understanding Optimizing factorized encoder models: Time and memory reduction for scalable and efficient action recognition
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 2edfc8e9-b2ce-4bce-a2eb-53498af96b52 · outbound
Principles of Visual Tokens for Efficient Video Understanding something something
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation c0a9017a-ad20-4001-8630-4e55eef1602e · outbound
Principles of Visual Tokens for Efficient Video Understanding Ava: A video dataset of spatio-temporally localized atomic visual actions
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation adfa02c1-ed9d-4c8d-90f7-a8970e545584 · outbound
Principles of Visual Tokens for Efficient Video Understanding Unresolved cited work
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation cfa34fc8-2a72-4cc2-8ae8-ac1bf3f41028 · outbound
Principles of Visual Tokens for Efficient Video Understanding LookupViT: Compressing visual information to a limited number of tokens
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ecbcfb01-bc9f-473c-92a5-b1b7cdfe0ec4 · outbound
Principles of Visual Tokens for Efficient Video Understanding Revisiting token pruning for object detection and instance segmentation
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 602013d9-90ab-49bc-873f-38010dd65c0a · outbound
Principles of Visual Tokens for Efficient Video Understanding Swin transformer: 9 Hierarchical vision transformer using shifted windows
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 1340edae-8416-4d56-8865-ae5b6f5ede89 · outbound
Principles of Visual Tokens for Efficient Video Understanding Video swin transformer
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 59cbb268-d70c-42cd-975b-dc68b59c9c99 · outbound
Principles of Visual Tokens for Efficient Video Understanding Video transformer network
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 6bb2d98a-286b-487b-8c46-1544f41c56a2 · outbound
Principles of Visual Tokens for Efficient Video Understanding Expanding language-image pretrained models for gen- eral video recognition
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0797f54-8feb-45a7-8fdf-0511431ca87b · outbound
Principles of Visual Tokens for Efficient Video Understanding St-adapter: Parameter-efficient image-to-video transfer learning
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a513123d-c351-407e-95d4-8c80f29e1fa2 · outbound
Principles of Visual Tokens for Efficient Video Understanding K-centered patch sampling for efficient video recognition
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 34fbc06d-8cd9-44a6-8a19-3d578edff604 · outbound
Principles of Visual Tokens for Efficient Video Understanding Asano, Is- han Misra Florian Metze, Christoph Feichtenhofer, Andrea Vedaldi, and Jo ˜ao F
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation d5be688a-fbbc-48ed-862b-48795aed311f · outbound
Principles of Visual Tokens for Efficient Video Understanding So, Maud Texier, and Jeff Dean
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 89f85019-3d55-4d11-a005-a117a8eef86b · outbound
Principles of Visual Tokens for Efficient Video Understanding How does the primate brain combine generative and discriminative computations in vision?
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation afecf784-5a76-4a59-b954-f984930b3757 · outbound
Principles of Visual Tokens for Efficient Video Understanding Dynamicvit: Efficient vision transformers with dynamic token sparsification
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6e191aa-4563-456b-9d38-200a2507f493 · outbound
Principles of Visual Tokens for Efficient Video Understanding TokenLearner: What Can 8 Learned Tokens Do for Images and Videos?
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c64056be-71b2-4f73-a879-28f300e8441f · outbound
Principles of Visual Tokens for Efficient Video Understanding Selvaraju, Abhishek Das, Ramakrishna Vedantam, Michael Cogswell, Devi Parikh, and Dhruv Ba- tra
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 253eb58c-3783-48b2-8a40-03be771481f6 · outbound
Principles of Visual Tokens for Efficient Video Understanding Only time can tell: Discovering temporal data for temporal modeling
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation d826674c-9cc0-4f29-8b22-2b105315db20 · outbound
Principles of Visual Tokens for Efficient Video Understanding UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74e6b19c-f94e-4ecb-9e89-7e5ed0031801 · outbound
Principles of Visual Tokens for Efficient Video Understanding VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db30876b-61c8-4e75-900b-c2afe0611f5f · outbound
Principles of Visual Tokens for Efficient Video Understanding Training data-efficient image transformers & distillation through at- tention
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42879cc2-add6-46c6-ae1a-77b6447084ba · outbound
Principles of Visual Tokens for Efficient Video Understanding Attention is all you need
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 723c88fb-5569-4fa7-b021-9037e90ce588 · outbound
Principles of Visual Tokens for Efficient Video Understanding Efficient video transformers with spatial- temporal token selection
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation dffe762e-8325-4ff8-993b-c2834b93bc77 · outbound
Principles of Visual Tokens for Efficient Video Understanding Actionclip: Adapting language-image pretrained models for video action recognition
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 17e7c3c5-5d9f-40c0-8763-a95ada95aafc · outbound
Principles of Visual Tokens for Efficient Video Understanding Vila: Efficient video-language alignment for video question answering
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 84b220ef-6547-4687-ac53-421d9391f825 · outbound
Principles of Visual Tokens for Efficient Video Understanding Video-focalnets: Spatio-temporal focal modu- lation for video action recognition
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 665eb6e4-ab17-43a3-b74d-afbf74dd7766 · outbound
Principles of Visual Tokens for Efficient Video Understanding Manmatha, Alex Smola, and Philipp Kr¨ahenb¨uhl
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 7b6373a5-13ef-4083-b0ba-fab7f0b5b684 · outbound
Principles of Visual Tokens for Efficient Video Understanding Memvit: Memory-augmented multiscale vision transformer for efficient long-term video recognition
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 79bb0272-bd07-4e33-93ce-2f1be485168c · outbound
Principles of Visual Tokens for Efficient Video Understanding Can i trust your answer? visually grounded video question answering
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5201b10-1d8b-4aa0-a730-2072546fec75 · outbound
Principles of Visual Tokens for Efficient Video Understanding Aim: Adapting image models for effi- cient video action recognition
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 7768a25e-84b1-41fe-a56b-15c65398e260 · outbound
Principles of Visual Tokens for Efficient Video Understanding A-vit: Adap- tive tokens for efficient vision transformer
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation a7ba7476-5d2f-4e15-8c85-82037d865916 · outbound
Principles of Visual Tokens for Efficient Video Understanding Self-chained image-language model for video localization and question answering
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 6f62fb58-6e13-44f0-b29d-15c2176a3f41 · outbound
Principles of Visual Tokens for Efficient Video Understanding Pyramid feature attention net- work for saliency detection
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation c2c53224-c922-44fb-bbfc-34ab9845defc · outbound
Principles of Visual Tokens for Efficient Video Understanding How can objects help action recognition? 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 2353–2362, 2023
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 95ba41bf-3a83-443d-8bac-6c62795b021e · outbound
Principles of Visual Tokens for Efficient Video Understanding ECO: Efficient Convolutional Network for Online Video Understanding
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
No inbound Pith citation observations are available.