Pith. sign in

Paper Citation Record · LEDGER

Principles of Visual Tokens for Efficient Video Understanding

As of 13 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 0 inbound Pith citation observations for arXiv:2411.13626.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.13626 v2

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T16:38:09.895402Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

54 of 54 outbound references displayed

  • verified exact3
  • verified fuzzy32
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5a99891b-2e49-4489-8005-51c9dc53d51a · outbound

This paper cites Vivit: A video vi- sion transformer.

Principles of Visual Tokens for Efficient Video Understanding Vivit: A video vi- sion transformer

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:11.107604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:38:09.538275Z digest=sha256:50f21bcb0c2bc0527c9fa2fa29906573d3896070b2397d21d8bb9f1a34157b12

Observation c9348b78-6a67-4449-bb64-609cc30b4912 · outbound

This paper cites Is Space-Time Attention All You Need for Video Understanding?.

Principles of Visual Tokens for Efficient Video Understanding Is Space-Time Attention All You Need for Video Understanding?

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:09.546390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:09.546390Z digest=sha256:36ddba769428188e280694bc0c9c6c92e74a592a8d2fac353c5c36b1b0fbd294

Observation 49cabee7-aefe-4090-8476-4ad7f9bdc92f · outbound

This paper cites Is space-time attention all you need for video understanding? In ICML, page 4, 2021.

Principles of Visual Tokens for Efficient Video Understanding Is space-time attention all you need for video understanding? In ICML, page 4, 2021

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:09.554368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:09.554368Z digest=sha256:a62d12b26e5b66994481bf2597cc595b290d1b855bb32467002d880d119b0dac

Observation 81fd44e9-98bc-428e-9fa7-c3e18c494187 · outbound

This paper cites Token merging: Your ViT but faster.

Principles of Visual Tokens for Efficient Video Understanding Token merging: Your ViT but faster

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:11.075015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:38:09.561593Z digest=sha256:7bff1ccc3f128dbbeae1bd880f75fcf8392924c6615470fecb0b6d3a5144df36

Observation 42b7f1f4-3f60-4fac-a742-73a4fe5dc6ef · outbound

This paper cites Revisiting the” video” in video-language understanding.

Principles of Visual Tokens for Efficient Video Understanding Revisiting the” video” in video-language understanding

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:11.052252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:38:09.568188Z digest=sha256:b567a8264092f706a160c2564fd84dc34fa4bb41af347476b35cbe12cd519983

Observation 9651a4fb-f220-45f0-8ef8-75e05b8d0472 · outbound

This paper cites Space-time mixing attention for video transformer.

Principles of Visual Tokens for Efficient Video Understanding Space-time mixing attention for video transformer

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:11.031143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:38:09.584966Z digest=sha256:7a1990e8528c4fbb5eb29d0e77928dc877813ea16c94e6c6b6877fbdf6ac9ddd

Observation fabe6ce3-449a-4b92-b86a-857a4caa921e · outbound

This paper cites Activitynet: A large-scale video benchmark for human activity understanding.

Principles of Visual Tokens for Efficient Video Understanding Activitynet: A large-scale video benchmark for human activity understanding

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:09.595147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:09.595147Z digest=sha256:1c2ddf7a4e23e6b5a81d9080a2830c7120053ff9372a7a13674b0875f7323c50

Observation e9612554-9565-471e-920a-d940a518b9df · outbound

This paper cites Quo vadis, action recognition? a new model and the kinetics dataset.

Principles of Visual Tokens for Efficient Video Understanding Quo vadis, action recognition? a new model and the kinetics dataset

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:11.001220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:38:09.609017Z digest=sha256:92ef202b87f3851295d6329bce5378f506d30b4525939f695c0c727058b70113

Observation a852fa36-f5f0-458f-bc78-b053c27ab8f5 · outbound

This paper cites An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models, 2024.

Principles of Visual Tokens for Efficient Video Understanding An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models, 2024

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:10.981335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:38:09.615828Z digest=sha256:bef4f15e23f9f5e78a71836e69c66053037cadd9300333d525beac373fb719fa

Observation 25b26d8c-8dc3-44cb-bad3-b381833ec000 · outbound

This paper cites an unresolved cited work.

Principles of Visual Tokens for Efficient Video Understanding Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-12T16:38:10.960743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:38:09.622145Z digest=sha256:bbe126a9d8f1b8841c405bc1c7c2062b7339e19028dce42461b47e32e80544b7

Observation 47cf182b-374d-48fe-93a6-9457ea518760 · outbound

This paper cites Prune spatio-temporal tokens by semantic-aware temporal accumulation.

Principles of Visual Tokens for Efficient Video Understanding Prune spatio-temporal tokens by semantic-aware temporal accumulation

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:10.941052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:38:09.629775Z digest=sha256:803e0e0d7551764148c609db909bb8516bd48cf03de4ee08945d9754832b8a2f

Observation ff94fd5d-3e5c-40e1-83de-4cce007a8738 · outbound

This paper cites An image is worth 16x16 words: Trans- formers for image recognition at scale.

Principles of Visual Tokens for Efficient Video Understanding An image is worth 16x16 words: Trans- formers for image recognition at scale

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:09.635930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:09.635930Z digest=sha256:2c766a77af023cf2d4a468af522d761060e1bae2b087d144eef4e61b185b10ed

Observation 5b345fdf-3da9-4331-a4cc-dd588465739c · outbound

This paper cites Multiscale vision transformers.

Principles of Visual Tokens for Efficient Video Understanding Multiscale vision transformers

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:10.910305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:38:09.643481Z digest=sha256:17902fdd2fc1b3ac4029da7245d9bebdde24519938360345d32aba25956faf32

Observation 48f3a647-1921-41d7-8c5e-859bf389090c · outbound

This paper cites X3d: Expanding architectures for efficient video recognition.

Principles of Visual Tokens for Efficient Video Understanding X3d: Expanding architectures for efficient video recognition

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:10.890972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:38:09.649118Z digest=sha256:ac3148e7cf9eaff348eb23f9fda024320b05b2446af5d6ac89e2474f3b1d6fe1

Observation 0bd73605-a781-43b2-ac83-0abdbf5e436a · outbound

This paper cites Slowfast networks for video recognition.

Principles of Visual Tokens for Efficient Video Understanding Slowfast networks for video recognition

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:09.658326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:09.658326Z digest=sha256:71921d4fd5952bb24f8a80162bb52ccc3d168e1170eaf5b7475d902df5abc12d

Observation a7ea3224-d89b-494d-a436-a9a769696a5a · outbound

This paper cites Efficient video transformers via spatial-temporal token merging for action recognition.

Principles of Visual Tokens for Efficient Video Understanding Efficient video transformers via spatial-temporal token merging for action recognition

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:10.851935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:38:09.664529Z digest=sha256:6e9f1938c68970cf4318c0d93d72f8bdf70b446adf77f4eda8da55cc746d4f43

Observation ec917cda-c365-402b-a165-4e86a39f60ab · outbound

This paper cites Smart frame selection for action recognition.

Principles of Visual Tokens for Efficient Video Understanding Smart frame selection for action recognition

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:09.670309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:09.670309Z digest=sha256:5742cf4316e3d07f475ef3c04558ecb86f728aa1cf97f6d30ea93c74c03dd214

Observation a3eeda23-3d4a-4dea-9fd3-c64001b2df71 · outbound

This paper cites Watt For What: Rethinking Deep Learning's Energy-Performance Relationship.

Principles of Visual Tokens for Efficient Video Understanding Watt For What: Rethinking Deep Learning's Energy-Performance Relationship

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:09.676924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:09.676924Z digest=sha256:ed08c78f24542b707d82ee27eb8a6012553588ef16ec85f8a766e16351dd63f3

Observation bb986a7b-b58f-402c-a25d-1fc5628338a2 · outbound

This paper cites Optimizing factorized encoder models: Time and memory reduction for scalable and efficient action recognition.

Principles of Visual Tokens for Efficient Video Understanding Optimizing factorized encoder models: Time and memory reduction for scalable and efficient action recognition

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:10.812740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:38:09.684927Z digest=sha256:adc9bdfedb5e6ee210496aabdf0e125e20a79db05e1bfa5673b9ba471ee2548d

Observation 2edfc8e9-b2ce-4bce-a2eb-53498af96b52 · outbound

This paper cites something something.

Principles of Visual Tokens for Efficient Video Understanding something something

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:10.787929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:38:09.691332Z digest=sha256:541781721ff086b8b63c6fe4a20bf68023cf929897aa469e250aba994dc48c13

Observation c0a9017a-ad20-4001-8630-4e55eef1602e · outbound

This paper cites Ava: A video dataset of spatio-temporally localized atomic visual actions.

Principles of Visual Tokens for Efficient Video Understanding Ava: A video dataset of spatio-temporally localized atomic visual actions

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:09.697107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:09.697107Z digest=sha256:5db1364443bc041f817a55bfaa7dbd73bf141448c3519d7ca220c22af61d9a77

Observation adfa02c1-ed9d-4c8d-90f7-a8970e545584 · outbound

This paper cites an unresolved cited work.

Principles of Visual Tokens for Efficient Video Understanding Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-12T16:38:10.753008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:38:09.702748Z digest=sha256:4d47585326ffe838053dc5d31e521624c24dbc02a2c908339654a9154d811547

Observation cfa34fc8-2a72-4cc2-8ae8-ac1bf3f41028 · outbound

This paper cites LookupViT: Compressing visual information to a limited number of tokens.

Principles of Visual Tokens for Efficient Video Understanding LookupViT: Compressing visual information to a limited number of tokens

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-08-12T16:38:10.119977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:38:09.707821Z digest=sha256:524cf63cfbda3afc2cbdaf92b6cc0807aed1cda2c30f9a53e1c977e23ac1f349

Observation ecbcfb01-bc9f-473c-92a5-b1b7cdfe0ec4 · outbound

This paper cites Revisiting token pruning for object detection and instance segmentation.

Principles of Visual Tokens for Efficient Video Understanding Revisiting token pruning for object detection and instance segmentation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:09.714750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:09.714750Z digest=sha256:afd4a25fa8debcb3bcf3273354df193f1dc7e6b7c4d07890ef41ff2a44e30575

Observation 602013d9-90ab-49bc-873f-38010dd65c0a · outbound

This paper cites Swin transformer: 9 Hierarchical vision transformer using shifted windows.

Principles of Visual Tokens for Efficient Video Understanding Swin transformer: 9 Hierarchical vision transformer using shifted windows

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:10.718586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:38:09.720206Z digest=sha256:f7973c9fbb540f7cc1f3c263e2b975447f2641dcbbd594627c08bb8894f81018

Observation 1340edae-8416-4d56-8865-ae5b6f5ede89 · outbound

This paper cites Video swin transformer.

Principles of Visual Tokens for Efficient Video Understanding Video swin transformer

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:10.695342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:38:09.725399Z digest=sha256:ef5c97abeb90df3eb5b7cc7ab0a27eb98f26512f51af9d5adc0b1f4900f5f359

Observation 59cbb268-d70c-42cd-975b-dc68b59c9c99 · outbound

This paper cites Video transformer network.

Principles of Visual Tokens for Efficient Video Understanding Video transformer network

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:10.671573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:38:09.732078Z digest=sha256:329e9a0a52d8b23356b4d94ac15ff087ab4ede0b903385065fc0a183f3eb06dc

Observation 6bb2d98a-286b-487b-8c46-1544f41c56a2 · outbound

This paper cites Expanding language-image pretrained models for gen- eral video recognition.

Principles of Visual Tokens for Efficient Video Understanding Expanding language-image pretrained models for gen- eral video recognition

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:09.737194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:09.737194Z digest=sha256:19ccd974cb2787d9e3ac973ac61fd2ecdca6b498c194cd64d884c5ba66cb1337

Observation d0797f54-8feb-45a7-8fdf-0511431ca87b · outbound

This paper cites St-adapter: Parameter-efficient image-to-video transfer learning.

Principles of Visual Tokens for Efficient Video Understanding St-adapter: Parameter-efficient image-to-video transfer learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:09.744700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:09.744700Z digest=sha256:dd27a5bd6b88c8bcfb7745c8065414eea25125d75f76cfaaaf502a497cc80595

Observation a513123d-c351-407e-95d4-8c80f29e1fa2 · outbound

This paper cites K-centered patch sampling for efficient video recognition.

Principles of Visual Tokens for Efficient Video Understanding K-centered patch sampling for efficient video recognition

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:10.624875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:38:09.750699Z digest=sha256:be7322019ec6a4cd686483c74455008db792a68f1a27b61e2ff86504b2851f14

Observation 34fbc06d-8cd9-44a6-8a19-3d578edff604 · outbound

This paper cites Asano, Is- han Misra Florian Metze, Christoph Feichtenhofer, Andrea Vedaldi, and Jo ˜ao F.

Principles of Visual Tokens for Efficient Video Understanding Asano, Is- han Misra Florian Metze, Christoph Feichtenhofer, Andrea Vedaldi, and Jo ˜ao F

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:10.602725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:38:09.756350Z digest=sha256:e4548bc739ca6ab2408b5be4db0077fafed1743fed489f56963ed77e9de6fdef

Observation d5be688a-fbbc-48ed-862b-48795aed311f · outbound

This paper cites So, Maud Texier, and Jeff Dean.

Principles of Visual Tokens for Efficient Video Understanding So, Maud Texier, and Jeff Dean

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:10.578776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:38:09.761549Z digest=sha256:4e367dbf0636741235e8dc4ddc587680de25a8aa8f23402ebbb8c0273db63655

Observation 89f85019-3d55-4d11-a005-a117a8eef86b · outbound

This paper cites How does the primate brain combine generative and discriminative computations in vision?.

Principles of Visual Tokens for Efficient Video Understanding How does the primate brain combine generative and discriminative computations in vision?

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-08-12T16:38:10.086545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:38:09.767230Z digest=sha256:c421f57e042af26b9687fdf8fdfa5ed0dcd42dad4ab4ab689f9248d900e4c0a5

Observation afecf784-5a76-4a59-b954-f984930b3757 · outbound

This paper cites Dynamicvit: Efficient vision transformers with dynamic token sparsification.

Principles of Visual Tokens for Efficient Video Understanding Dynamicvit: Efficient vision transformers with dynamic token sparsification

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:09.773955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:09.773955Z digest=sha256:faaf450c6ca6c4cdaa9cdb64ecd7b9d62971ed9151d773d5ac8df76c0342864c

Observation c6e191aa-4563-456b-9d38-200a2507f493 · outbound

This paper cites TokenLearner: What Can 8 Learned Tokens Do for Images and Videos?.

Principles of Visual Tokens for Efficient Video Understanding TokenLearner: What Can 8 Learned Tokens Do for Images and Videos?

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:09.781891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:09.781891Z digest=sha256:7e560ab052dc629d836f813197a1339a2370396f535a2f111bcecb914895ee2c

Observation c64056be-71b2-4f73-a879-28f300e8441f · outbound

This paper cites Selvaraju, Abhishek Das, Ramakrishna Vedantam, Michael Cogswell, Devi Parikh, and Dhruv Ba- tra.

Principles of Visual Tokens for Efficient Video Understanding Selvaraju, Abhishek Das, Ramakrishna Vedantam, Michael Cogswell, Devi Parikh, and Dhruv Ba- tra

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:10.531612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:38:09.788899Z digest=sha256:0c47e20560fc86cc04bdf2f8786feeda54845507ad69f2dfc803ba4e2e37ad62

Observation 253eb58c-3783-48b2-8a40-03be771481f6 · outbound

This paper cites Only time can tell: Discovering temporal data for temporal modeling.

Principles of Visual Tokens for Efficient Video Understanding Only time can tell: Discovering temporal data for temporal modeling

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:10.507034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:38:09.795722Z digest=sha256:3909c055332bd7d8f6b30acfb014d1bc4e8bd9e8bfb6ab524de67ee0eec5c3e6

Observation d826674c-9cc0-4f29-8b22-2b105315db20 · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

Principles of Visual Tokens for Efficient Video Understanding UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:09.801167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:09.801167Z digest=sha256:7b8d76c300f83bc077ae98dc393b0d223b4bf7940cc7df374d446e5e8e56ec49

Observation 74e6b19c-f94e-4ecb-9e89-7e5ed0031801 · outbound

This paper cites VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training.

Principles of Visual Tokens for Efficient Video Understanding VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:09.806583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:09.806583Z digest=sha256:84214f367d5c473ff1d9c89c956778bfc40e4196753f92996b691750fe05e316

Observation db30876b-61c8-4e75-900b-c2afe0611f5f · outbound

This paper cites Training data-efficient image transformers & distillation through at- tention.

Principles of Visual Tokens for Efficient Video Understanding Training data-efficient image transformers & distillation through at- tention

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:09.812884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:09.812884Z digest=sha256:d7881a780dd10df02c4cb045c35b08f0473c2a559532fec6d135ac58941cd1f4

Observation 42879cc2-add6-46c6-ae1a-77b6447084ba · outbound

This paper cites Attention is all you need.

Principles of Visual Tokens for Efficient Video Understanding Attention is all you need

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:10.463112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:38:09.818358Z digest=sha256:74fc174d9765c6f53d85f44fe14d76925d90ca10d4939a64db5b0c6377bb989e

Observation 723c88fb-5569-4fa7-b021-9037e90ce588 · outbound

This paper cites Efficient video transformers with spatial- temporal token selection.

Principles of Visual Tokens for Efficient Video Understanding Efficient video transformers with spatial- temporal token selection

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:10.430041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:38:09.823324Z digest=sha256:8a9f074be084b9706ea776bd230b3059459d2d6754aabef4653f571d07d44fe8

Observation dffe762e-8325-4ff8-993b-c2834b93bc77 · outbound

This paper cites Actionclip: Adapting language-image pretrained models for video action recognition.

Principles of Visual Tokens for Efficient Video Understanding Actionclip: Adapting language-image pretrained models for video action recognition

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:10.405228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:38:09.829073Z digest=sha256:e640fbe80557ee76110da128961e65f3c7396385b20221df6dd50651cfffe401

Observation 17e7c3c5-5d9f-40c0-8763-a95ada95aafc · outbound

This paper cites Vila: Efficient video-language alignment for video question answering.

Principles of Visual Tokens for Efficient Video Understanding Vila: Efficient video-language alignment for video question answering

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:10.378103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:38:09.833947Z digest=sha256:d1aa8f36cb7b4e0bf5ab200d465af226d7cc4c4080ea7ab57fd64cb360e1f556

Observation 84b220ef-6547-4687-ac53-421d9391f825 · outbound

This paper cites Video-focalnets: Spatio-temporal focal modu- lation for video action recognition.

Principles of Visual Tokens for Efficient Video Understanding Video-focalnets: Spatio-temporal focal modu- lation for video action recognition

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:10.351817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:38:09.841034Z digest=sha256:e103efbada5c480d1fe1df0506538485f6b35fb584227b7591f99c42a1c9eb11

Observation 665eb6e4-ab17-43a3-b74d-afbf74dd7766 · outbound

This paper cites Manmatha, Alex Smola, and Philipp Kr¨ahenb¨uhl.

Principles of Visual Tokens for Efficient Video Understanding Manmatha, Alex Smola, and Philipp Kr¨ahenb¨uhl

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:10.325086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:38:09.847893Z digest=sha256:95ef58482f37ee2bd1592d445e29c3cbf4951a5fe9868b854cf1a85fa2c642e9

Observation 7b6373a5-13ef-4083-b0ba-fab7f0b5b684 · outbound

This paper cites Memvit: Memory-augmented multiscale vision transformer for efficient long-term video recognition.

Principles of Visual Tokens for Efficient Video Understanding Memvit: Memory-augmented multiscale vision transformer for efficient long-term video recognition

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:10.306355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:38:09.853463Z digest=sha256:3b8eae8fa49c1348c25e4abcb75eaa824bb5359319d4992c718d1bbd20755cf5

Observation 79bb0272-bd07-4e33-93ce-2f1be485168c · outbound

This paper cites Can i trust your answer? visually grounded video question answering.

Principles of Visual Tokens for Efficient Video Understanding Can i trust your answer? visually grounded video question answering

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:09.859186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:09.859186Z digest=sha256:2cd3b667c972d08bf1421286212c5a1bd9a1899b85b2039cd4ebcf6ea9d5dde9

Observation d5201b10-1d8b-4aa0-a730-2072546fec75 · outbound

This paper cites Aim: Adapting image models for effi- cient video action recognition.

Principles of Visual Tokens for Efficient Video Understanding Aim: Adapting image models for effi- cient video action recognition

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:10.271496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:38:09.865766Z digest=sha256:e006e1b4affa5bfb817c622e7448fba71999866aaadc63dc063e13d6a5653db2

Observation 7768a25e-84b1-41fe-a56b-15c65398e260 · outbound

This paper cites A-vit: Adap- tive tokens for efficient vision transformer.

Principles of Visual Tokens for Efficient Video Understanding A-vit: Adap- tive tokens for efficient vision transformer

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:10.253123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:38:09.872069Z digest=sha256:904521b39797ee476f2c5f83cb2073748e9a0a2d50254c76144bcc77920ef96b

Observation a7ba7476-5d2f-4e15-8c85-82037d865916 · outbound

This paper cites Self-chained image-language model for video localization and question answering.

Principles of Visual Tokens for Efficient Video Understanding Self-chained image-language model for video localization and question answering

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:10.235125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:38:09.877725Z digest=sha256:d562d1b9b1aeaab4e0fd6e418f5b8647ab6a77e0dbd2feef1276c06cc0acc736

Observation 6f62fb58-6e13-44f0-b29d-15c2176a3f41 · outbound

This paper cites Pyramid feature attention net- work for saliency detection.

Principles of Visual Tokens for Efficient Video Understanding Pyramid feature attention net- work for saliency detection

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:10.217473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:38:09.883184Z digest=sha256:cfee1e7d40560c36bc9461c628370d9b902cc7db8267f975df9f7a5da3784a94

Observation c2c53224-c922-44fb-bbfc-34ab9845defc · outbound

This paper cites How can objects help action recognition? 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 2353–2362, 2023.

Principles of Visual Tokens for Efficient Video Understanding How can objects help action recognition? 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 2353–2362, 2023

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:10.197334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:38:09.889971Z digest=sha256:fc42377e82035c815674d6879e83543a43f4b9f06f7818ddacc8de0f164feb29

Observation 95ba41bf-3a83-443d-8bac-6c62795b021e · outbound

This paper cites ECO: Efficient Convolutional Network for Online Video Understanding.

Principles of Visual Tokens for Efficient Video Understanding ECO: Efficient Convolutional Network for Online Video Understanding

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-08-12T16:38:09.954883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:38:09.895402Z digest=sha256:c5fc0dec2cd72fe9f2dcc6411a913f7578d7d3546f60ebd8f519665a2bc8bef3

Pith citing papers

No inbound Pith citation observations are available.