Pith. sign in

Paper Citation Record · LEDGER

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition

As of 8 August 2026, this Paper Citation Record lists 95 of 95 outbound references and 0 inbound Pith citation observations for arXiv:2507.12426.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.12426 v2

Coverage vector

measured 95 of 95 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T16:52:55.019874Z

measured 95 of 95 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

95 of 95 outbound references displayed

  • verified exact2
  • verified fuzzy70
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7d6fa9a7-95b3-4e9b-b0c8-bfc6deb3b292 · outbound

This paper cites Quo vadis, action recognition? a new model and the kinetics dataset,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Quo vadis, action recognition? a new model and the kinetics dataset,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:46.543670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:46.543670Z digest=sha256:a78ef456bd5be686c64e554e251a1a5a4941ae1b1e9e54c1bae013187f1c2ecb

Observation f6dfecc4-6bac-46de-8600-043905dbef6d · outbound

This paper cites Spatiotemporal residual networks for video action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Spatiotemporal residual networks for video action recognition,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:46.603264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:46.603264Z digest=sha256:b828cee742fd765f20bf9c7fd07a36cb1fa07748608a64f9aea92966a21bd3b5

Observation da7f7c26-20f2-4582-829d-229133615bd3 · outbound

This paper cites Learning spatiotemporal features with 3d convolutional networks,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Learning spatiotemporal features with 3d convolutional networks,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:46.762590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:46.762590Z digest=sha256:eadf9790a9259f0258eafb939ef86efff135139939c01a3adda37bdd77f9431d

Observation 1139d00a-d086-4239-9ee1-f7fbbe824386 · outbound

This paper cites Large-scale video classification with convolutional neural networks,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Large-scale video classification with convolutional neural networks,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:46.834049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:46.834049Z digest=sha256:6a707f36d5585b8a3cfc8a0b9c6639289affca1c2e2d3dfea068a2d5e667caad

Observation 81df5fdd-3704-4da7-864d-36ff4ef50b15 · outbound

This paper cites Beyond short snippets: Deep networks for video classification,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Beyond short snippets: Deep networks for video classification,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:46.912945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:46.912945Z digest=sha256:4fffb9a973078a07d9934fff3c08a5dee5457f04d6ef73b06c3a3458f3187320

Observation 30e07942-23e6-46a0-9685-6d2274f38ca8 · outbound

This paper cites Two-stream convolutional networks for action recognition in videos,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Two-stream convolutional networks for action recognition in videos,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:46.999396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:46.999396Z digest=sha256:93dca2750c81c38e8177768eb978eff037b39e7bfb73a04dabcd4c9ce9de8884

Observation 2ab53b6f-85c6-446e-b41c-2ad3340312f1 · outbound

This paper cites Action recognition using deep 3d cnns with sequential feature aggregation and attention,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Action recognition using deep 3d cnns with sequential feature aggregation and attention,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:47.061463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:47.061463Z digest=sha256:ad221a6560a79d8bbbab2cdf51db35d54b64b1c56759d6a81caf50a2592bef6a

Observation befd3be9-d7eb-4f72-a1a3-e92c65ab6b2a · outbound

This paper cites A closer look at spatiotemporal convolutions for action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition A closer look at spatiotemporal convolutions for action recognition,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:47.139621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:47.139621Z digest=sha256:a793117bbb769a24d89a1e418a4597d1465fe908cf838feda965ceaeb683a2d6

Observation 89171f78-0d81-4a0f-8b44-9d16f4498847 · outbound

This paper cites Human action recognition in videos using convolution long short-term memory network with spatio-temporal networks,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Human action recognition in videos using convolution long short-term memory network with spatio-temporal networks,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:47.215422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:47.215422Z digest=sha256:b624b068795f08423469ac9609f265a652480ea06c7de2c466522ecebd9ae9c4

Observation ea61037a-f432-42f1-8c1e-5bc5d0ed4871 · outbound

This paper cites Action recognition in videos using pre-trained 2d convolutional neural networks,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Action recognition in videos using pre-trained 2d convolutional neural networks,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:47.293116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:47.293116Z digest=sha256:76f071c1245604b0c1d7dcf8ef1b2da0eb8a0e5d17b41d2cd013f734857f9970

Observation 57809d8d-af21-4dad-a9c4-caa9ce0ce57b · outbound

This paper cites Video-focalnets: Spatio-temporal focal modulation for video action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Video-focalnets: Spatio-temporal focal modulation for video action recognition,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:47.365973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:47.365973Z digest=sha256:1a267fab3363c838b5f36aa6a456db28fc17a774924a34f5481063afc2dbc4a9

Observation cbd63ea4-ac98-4630-b0d5-2f34a2a50e21 · outbound

This paper cites Attention is all you need,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Attention is all you need,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:47.458192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:47.458192Z digest=sha256:081a86e147d74163034ed256dacfac110cdc7dd2dcd083aa0da160ee383ba6c3

Observation 94573c9d-c9d9-405f-a809-205fd9385d3b · outbound

This paper cites Morph: flexible acceleration for 3d cnn-based video understanding,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Morph: flexible acceleration for 3d cnn-based video understanding,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:47.549508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:47.549508Z digest=sha256:9eaa73fbbbf5541a4c8e9fabdbe0f1813a77f94f8b6ccb0e689c0259bc446799

Observation 7dacace1-b6ab-4d7f-a2f6-134e5c83ccb9 · outbound

This paper cites Video swin transformer,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Video swin transformer,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:47.646960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:47.646960Z digest=sha256:aac01132818716b9240327d30483201bf0169850cead72368b2f866d8c2646cf

Observation f0cf805b-b70b-45ea-9f7b-4c0d38c27a59 · outbound

This paper cites Is space-time attention all you need for video understanding?.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Is space-time attention all you need for video understanding?

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:47.720373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:47.720373Z digest=sha256:a79f433a61f16b1cb526391a93c587e3f243ba9920f2b6b2b55506acb206ae69

Observation 3eb7b69d-ddd7-49e0-9c4f-fe022bf11c61 · outbound

This paper cites Vivit: A video vision transformer,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Vivit: A video vision transformer,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:06.347041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:52:47.802678Z digest=sha256:3545278de140da42d50ae392eafd4a061be37e04a6ea7d6d0b9a5e21e74287d7

Observation 82e5e3f7-b6c3-41a8-9fc9-f6b66ec19b23 · outbound

This paper cites Video swin transformer,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Video swin transformer,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:06.256262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:52:47.915356Z digest=sha256:745a2ba65a66d99c3d12509d6c04e59b5975c86f3ac9a1bf3d8b505016d6d47e

Observation 05571d0b-4696-43a4-a4c0-f3ef2b565420 · outbound

This paper cites Multiview transformers for video recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Multiview transformers for video recognition,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:06.027968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:52:48.002422Z digest=sha256:a93dd91a39960a820339644f43685fcdd4f5f583fdd585d76f84d36576ea3af3

Observation d19db6ac-64a3-4d5b-8bc4-38a75d42c01d · outbound

This paper cites Vivit: a video vision transformer,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Vivit: a video vision transformer,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:05.838377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:52:48.113860Z digest=sha256:fa89913ec9d3677ec1a0e4956e047649c995d260f0395b6c51117a26d6f8dff8

Observation b1bcdfe6-4cdc-44f8-aad7-61da12e91435 · outbound

This paper cites A Short Note on the Kinetics-700 Human Action Dataset.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition A Short Note on the Kinetics-700 Human Action Dataset

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:48.257823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:48.257823Z digest=sha256:7b4aa1b61e953dda178e2f824c6eb1d4d5216a8f89f349044e83cf37254692ba

Observation a818e4a1-4b5e-4a94-b1fa-ad39061bc3c9 · outbound

This paper cites The" something something.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition The" something something

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:05.619678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:52:48.310977Z digest=sha256:9883e4951cb1639f875d14a5be9e556482d014caeb6bf2ee7def25d1f57edbd1

Observation 619217a0-41a1-4d5f-bb13-e281e3ab75a4 · outbound

This paper cites Tsnet: token sparsification for efficient video transformer,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Tsnet: token sparsification for efficient video transformer,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:05.374484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:52:48.397443Z digest=sha256:18382e34995753070952e68ca51a3b90f9da17c1bda83e478b2ed1f23128ca4b

Observation 14426831-847c-42f1-a119-ab7de1c633fd · outbound

This paper cites Dualformer: local- global stratified transformer for efficient video recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Dualformer: local- global stratified transformer for efficient video recognition,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:05.198144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:52:48.495776Z digest=sha256:d91a1aae123dc595e712a733c63590423b1b2c72d399ca4e7d8b4a9fbc29502e

Observation 9e639f13-90b9-4e10-890f-1c9f41288763 · outbound

This paper cites Aerobics action recognition algorithm based on three-dimensional convolutional neural network and multilabel clas- sification,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Aerobics action recognition algorithm based on three-dimensional convolutional neural network and multilabel clas- sification,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:05.032271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:52:48.645264Z digest=sha256:26547503738a7b62b5a8beeb020b13b9fd92181f21484e9d0a0662b802c6a13f

Observation ecd963af-485a-4dc1-8fed-b62f79566714 · outbound

This paper cites Tsm: temporal shift module for efficient video understanding,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Tsm: temporal shift module for efficient video understanding,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:04.810289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:52:48.745524Z digest=sha256:c4673561302cd16de6f7fccbd193b9d8c5e41574e70403f62442f1b9f147095a

Observation 02dadde9-ea1b-43a4-9339-e281a66574ed · outbound

This paper cites Multi-stream interaction networks for human action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Multi-stream interaction networks for human action recognition,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:48.903466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:48.903466Z digest=sha256:d7c6a77a53cd001ad08d62f3a154dab45b28c1bd17e26fca7a4d31ad3aa44b6e

Observation f4125143-b5a1-47c0-ad04-f863252f597f · outbound

This paper cites Spatio- temporal adaptive network with bidirectional temporal difference for action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Spatio- temporal adaptive network with bidirectional temporal difference for action recognition,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:04.577825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:52:48.988513Z digest=sha256:0a0eea4d1b2705df3724166e8f49ff6d2f3960ad2a09903bf23a32785e622fe7

Observation 7239b067-039d-4a18-b7ed-b876d7c032a5 · outbound

This paper cites Agpn: Action granularity pyramid network for video action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Agpn: Action granularity pyramid network for video action recognition,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:04.413558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:52:49.052816Z digest=sha256:f406ec39be7f1c062859e3c385c4baf5ac75e9fe03bb3591743ca49b10ec1dde

Observation 10739f60-4c9f-4362-bc5d-71672cacce9c · outbound

This paper cites Mawkdn: A multimodal fusion wavelet knowledge distillation approach based on cross-view attention for action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Mawkdn: A multimodal fusion wavelet knowledge distillation approach based on cross-view attention for action recognition,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:04.180393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:52:49.130816Z digest=sha256:94627c2535d1d7301a2d142b260f7d8e4505f805daf48f45d9242af99dc148b6

Observation a3b87581-b7aa-4cce-8d76-78178a971912 · outbound

This paper cites Convolutional neural networks or vision transformers: who will win the race for action recognitions in visual data?.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Convolutional neural networks or vision transformers: who will win the race for action recognitions in visual data?

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:04.004396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:52:49.203358Z digest=sha256:5cf26f3f54b1c23f8ef701b4b7c097c498a9fd54896cb7f345f66c95b512b190

Observation dc449c7c-7aab-4408-9730-7df66158556e · outbound

This paper cites Decoupled knowledge distillation,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Decoupled knowledge distillation,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:03.788026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:52:49.291727Z digest=sha256:b4d983e882e01c8a4f82381c9983f360a1d303f30cf247b0f913c567ea1d4f9d

Observation 6c54a01c-3a35-469e-8b2d-214a8e7009c1 · outbound

This paper cites Knowledge distil- lation in video-based human action recognition: an intuitive approach to efficient and flexible model training,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Knowledge distil- lation in video-based human action recognition: an intuitive approach to efficient and flexible model training,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:03.616547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:52:49.359338Z digest=sha256:87033fddb85cfe66579be965da42ae3f0ef7ff3de923ad90a6965890a3c6ee03

Observation 0c4740cb-6e71-425d-b39b-8bb395344d67 · outbound

This paper cites Tomato leaf disease recognition based on multi-task distillation learning,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Tomato leaf disease recognition based on multi-task distillation learning,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:03.441174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:52:49.445207Z digest=sha256:41b91c6ca17aa25ee2b297e3e6c533044589d9c6bb5c2195ceaa9c3cfce0557f

Observation abc9a7cf-cff3-428b-86d2-4a8ec455593e · outbound

This paper cites Videoadviser: video knowledge distillation for multimodal transfer learning,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Videoadviser: video knowledge distillation for multimodal transfer learning,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:03.264063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:52:49.510135Z digest=sha256:ae6558b967244350e3bf8bf0d1c98aeb375428823cefc0a195dc33df69a24ed7

Observation f0e5042d-7bdc-44c7-ad1e-eb7de09209f5 · outbound

This paper cites Generative model- based feature knowledge distillation for action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Generative model- based feature knowledge distillation for action recognition,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:03.099709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:52:49.627318Z digest=sha256:153dadba8297924c336be70df9245ba9babef39ed61a3a76e7d60d68f0b6c316

Observation 7ac67c90-6540-4d09-9d3f-49fb1a495643 · outbound

This paper cites Distillation of human-object interaction contexts for action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Distillation of human-object interaction contexts for action recognition,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:02.880266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:52:49.680694Z digest=sha256:994f73c9f201a026b7abed35488185835e590c5299ee6f164b8ba43d29391b6a

Observation 51a66055-0d5c-4e05-ae76-efc69f911c32 · outbound

This paper cites Gaussian Error Linear Units (GELUs).

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Gaussian Error Linear Units (GELUs)

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:49.779462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:49.779462Z digest=sha256:5e95cfc4b037d44bd8fc09ac6676eca6524e786d57f8bfdee7a8f4eaa4af376c

Observation 42c9bee4-121f-49f0-bb40-e03ca6c20299 · outbound

This paper cites Recognizing 50 human action categories of web videos,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Recognizing 50 human action categories of web videos,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:02.753222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:52:49.860074Z digest=sha256:86b72afb7f5b686fc123ce9b763af1417e7e0b6f7a27f464c5f110bcc8f612b4

Observation 9a6faf9b-3c84-4091-854b-7615a44fba10 · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:50.009362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:50.009362Z digest=sha256:aa0e77cec42f98aa6b8c34cd955fc21aca218a7b5be44fcaf3e99851e166d403

Observation 3f17c9e8-b6e2-4c9c-bddf-97a6c44e5f88 · outbound

This paper cites Hmdb: a large video database for human motion recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Hmdb: a large video database for human motion recognition,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:02.650847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:52:50.106813Z digest=sha256:defec0ba9d3858149f9a94116890f17b45fd4ec76ee5b950e51cdb6b7169ca67

Observation c75510fe-4399-4f99-b634-5292bf4a285f · outbound

This paper cites The" something something.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition The" something something

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:02.556260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:52:50.174304Z digest=sha256:de39fb166a53ed94d928c9bd4215d5d50b8b66a0334ac5b26d2797421a89cb7a

Observation b3a55fc2-d47e-44ba-bc91-dbe64eb2618d · outbound

This paper cites The Kinetics Human Action Video Dataset.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition The Kinetics Human Action Video Dataset

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:50.256104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:50.256104Z digest=sha256:9eb297ccd8f6b19d769c9ad9bcb92024c96973c8be61514b0f6deef277f872fa

Observation 1c01ee1e-7001-44a3-a8f4-b282d3eed1db · outbound

This paper cites UniFormer: Unified Transformer for Efficient Spatiotemporal Representation Learning.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition UniFormer: Unified Transformer for Efficient Spatiotemporal Representation Learning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:50.390492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:50.390492Z digest=sha256:f4837fd8f5df69a906693371a143c72c7a5aa9689a3fc137ae2e8955330abdb3

Observation 771bb7c5-6e0d-495d-8627-a7fe6aec94b5 · outbound

This paper cites Going deeper with convolutions,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Going deeper with convolutions,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:02.389286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:52:50.471227Z digest=sha256:27a5cee355a6efb921c02bee1859a10f26fe06d48110b79dac69114b99b7e34b

Observation 24febacd-fbf9-4fa0-95f4-43b6defeb682 · outbound

This paper cites Making sense of neuromorphic event data for human action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Making sense of neuromorphic event data for human action recognition,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:02.263934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:52:50.520535Z digest=sha256:b68f46b580a3114ec97864dc89a5b1ad8e20f35dfc31b07e353073e5ef99dc55

Observation d670189b-3b5f-453a-9bb4-193e9c814505 · outbound

This paper cites Human action recognition using dis- tance transform and entropy based features,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Human action recognition using dis- tance transform and entropy based features,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:02.200658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:52:50.601096Z digest=sha256:a7171704db6799aa2de6628bae8eedd1793d0cb1b6448e84d1ea325941d12101

Observation 121473c8-a8a4-402a-bb27-b0ab77aab99d · outbound

This paper cites Human action recognition using hybrid deep evolving neural networks,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Human action recognition using hybrid deep evolving neural networks,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:02.116040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:52:50.719001Z digest=sha256:58c437546d6e9aa87f5a10aa6ea2740d8d086aa339b4d6ad5d8c756c410e7d64

Observation 36e68b1f-41af-4cf8-b42b-2f5eebc74e94 · outbound

This paper cites Simple-action-guided dictionary learning for complex action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Simple-action-guided dictionary learning for complex action recognition,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:02.035273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:52:50.830577Z digest=sha256:72da065f7dc853cabc640fe6e056e43410e1be72c3051f5d631450ac9e83041b

Observation 6a77b057-5bb1-4b19-9eeb-6b4c9707ba87 · outbound

This paper cites Human activity classification using the 3dcnn architecture,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Human activity classification using the 3dcnn architecture,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:01.921818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:52:50.937287Z digest=sha256:6290ac00f45407140ac569d1c569b45f902ba3d9af40cea1828b0011acc0c227

Observation 46e3ecc0-c91d-4f86-8982-338625a5b2c6 · outbound

This paper cites Fast classification and action recognition with event-based imaging,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Fast classification and action recognition with event-based imaging,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:01.770390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:52:51.096512Z digest=sha256:577774d7dec8301559158ec08003f997062529863e8186f6bf0dd94416513d5d

Observation 04678e70-a280-4266-b2bf-1cbc78009620 · outbound

This paper cites Spatio-temporal features based human action recognition using convolutional long short-term deep neural network,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Spatio-temporal features based human action recognition using convolutional long short-term deep neural network,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:01.610341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:52:51.214457Z digest=sha256:e49e6ee6654159a4cd2299d893d68a854028dfdb6eea776adb645daf96d2da88

Observation 26c54948-d8b3-443e-b9b7-c1fb8ddcacac · outbound

This paper cites Human action recognition using multi-stream fusion and hybrid deep neural networks,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Human action recognition using multi-stream fusion and hybrid deep neural networks,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:01.479218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:52:51.327664Z digest=sha256:104a05b0c3da2eeb693d49f3842ab19b642e94ffb5824bb0f88765e5441bd6d9

Observation 51710b76-8eed-4762-b895-c9825289873f · outbound

This paper cites Self-supervised video representation learning by uncovering spatio-temporal statistics,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Self-supervised video representation learning by uncovering spatio-temporal statistics,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:01.420896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:52:51.425880Z digest=sha256:8db9a0a4a222bad3d6aae5030d2a55e11fb8d10376002792b9c20d0741f2442b

Observation 7057970c-2ce6-4cfe-a130-d6867f5205b4 · outbound

This paper cites Enhancing self-supervised video representation learning via multi-level feature optimization,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Enhancing self-supervised video representation learning via multi-level feature optimization,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:01.332744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:52:51.505885Z digest=sha256:7f0d9a3242bd36bb453f6ef23f2339d0c0d50ca560134bba3cce88d0bcb65fe8

Observation d59234f2-5c29-47a9-83f8-a177742866da · outbound

This paper cites Videomoco: Contrastive video representation learning with temporally adversarial examples,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Videomoco: Contrastive video representation learning with temporally adversarial examples,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:01.169505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:52:51.608989Z digest=sha256:fb9ac872e779b916bc3ee053e692569ef171e7a8b559fe7b01e14e661ac57139

Observation ef665412-a01e-4b14-ae2b-f6664f0c0d74 · outbound

This paper cites Action recognition from a single coded image,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Action recognition from a single coded image,

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:01.043969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:52:51.708625Z digest=sha256:f076d88b4c65c25065ed630317c08b59df8acdd620f6ba1133f298b7f7d9e084

Observation 7c3b9fb0-aafb-4894-bebe-e9dd1baf0f02 · outbound

This paper cites Tclr: Temporal contrastive learning for video representation,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Tclr: Temporal contrastive learning for video representation,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:00.949242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:52:51.782435Z digest=sha256:f896ab981c9465496f1594c34f5f72aad3af80f4c0670a5357e75cc2008084b2

Observation 8e4ccfc8-44d5-4f05-b48a-8474961d34c8 · outbound

This paper cites Learn2augment: learning to composite videos for data augmentation in action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Learn2augment: learning to composite videos for data augmentation in action recognition,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:00.857703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:52:51.887584Z digest=sha256:6fb244ece5b0839d4da088e192e82564b962c73cb17064590cf2eb16299f0ea3

Observation 5e96c946-27dc-4cae-b725-e151ef84c116 · outbound

This paper cites Learning from temporal gradient for semi-supervised action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Learning from temporal gradient for semi-supervised action recognition,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:00.657387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:52:51.993613Z digest=sha256:b75bd0ae35be63212476f034f18233a1f7521edd8eb9efc37fc7b4523d9d6e3e

Observation 627c3103-b39a-40f0-aac2-395d19dc5380 · outbound

This paper cites Preserve Pre-trained Knowledge: Transfer Learning With Self-Distillation For Action Recognition.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Preserve Pre-trained Knowledge: Transfer Learning With Self-Distillation For Action Recognition

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-08-06T16:52:55.326931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:52:52.098718Z digest=sha256:b2f368166a0462272b6e834371d4dde1b262bb12470aa9ff81bd9d4ffa6f39c3

Observation 64a00eee-4644-447b-b2c3-22bcf3d2d01a · outbound

This paper cites Extreme low- resolution action recognition with confident spatial-temporal attention transfer,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Extreme low- resolution action recognition with confident spatial-temporal attention transfer,

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:00.478916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:52:52.195968Z digest=sha256:77e4d9ed9ad0adfb23eabe46f84e9b3b27a5b51761f54d5a584e826f6e68510f

Observation 0a6f9449-9838-4d17-a28b-95f3b333cdee · outbound

This paper cites Self-supervised video-based action recognition with disturbances,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Self-supervised video-based action recognition with disturbances,

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:00.286153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:52:52.315336Z digest=sha256:72e1456755aaf00dd102b139bd714d1dcb77ba4d0ae6a8a3a4f6e71504d0397c

Observation 24aaf547-bc02-420f-8f6f-ae968884f75f · outbound

This paper cites Spatial-temporal exclusive capsule network for open set action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Spatial-temporal exclusive capsule network for open set action recognition,

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:00.146160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:52:52.428855Z digest=sha256:39ce82939c385cd165c9fe65906ed89853de81ea8a68b25ecf611b89e6f7e883

Observation 7cff12ac-649e-46be-b717-016115aaf030 · outbound

This paper cites Sv- former: Semi-supervised video transformer for action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Sv- former: Semi-supervised video transformer for action recognition,

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:59.986980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:52:52.509199Z digest=sha256:125948fdc8ad61167ad2faacaa4ee3c4d280c1425c8a509582eea429cf591438

Observation 6e74282b-f332-4229-bd4c-a2008662ed7a · outbound

This paper cites ActionHub: A Large-scale Action Video Description Dataset for Zero-shot Action Recognition.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition ActionHub: A Large-scale Action Video Description Dataset for Zero-shot Action Recognition

Reference 66

Resolution
verified exact
local_arxiv, observed 2026-08-06T16:52:55.174880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:52:52.604016Z digest=sha256:f3dec5e5269eb25a06da7abe8f426f52c37462d75d326e8b44be3122aa04a901

Observation 3ac35da3-01a0-4f25-8917-be0cde4a5013 · outbound

This paper cites Self-supervised learning via multi-transformation classification for action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Self-supervised learning via multi-transformation classification for action recognition,

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:59.802805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:52:52.730793Z digest=sha256:2a531afa07323f3e2f808430116a94cc9d0338bf12749f7fbf9363f4c6ed47eb

Observation 48d46bd4-9904-490f-9eaa-43495d8f3c7f · outbound

This paper cites Semi-supervised action recog- nition with dynamic temporal information fusion,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Semi-supervised action recog- nition with dynamic temporal information fusion,

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:59.610844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:52:52.788124Z digest=sha256:454442516465b1db40c8ebf00db19014fb24be51d3d9c42e390acfcaa148b5a6

Observation fa98add2-c079-44e4-aac9-8e0bc6c79754 · outbound

This paper cites Spatiotemporal contrastive video representation learning,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Spatiotemporal contrastive video representation learning,

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:59.422918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:52:52.852312Z digest=sha256:8b16f9bf4f06d17bab3a44648c616a436578f6e70b00d41da7157cf61647c45e

Observation 9dd942c3-2298-4b33-bc85-24eb0ffdef8a · outbound

This paper cites Representation learning for compressed video action recognition via attentive cross- modal interaction with motion enhancement,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Representation learning for compressed video action recognition via attentive cross- modal interaction with motion enhancement,

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:59.292874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:52:52.921579Z digest=sha256:2ae6163d17a3d521a323c445ec5fe31bc12739e18254c2fa6166b7ac5d9319db

Observation 001c1faf-2447-4562-8b39-aec022990205 · outbound

This paper cites Motion-driven visual tempo learn- ing for video-based action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Motion-driven visual tempo learn- ing for video-based action recognition,

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:59.149388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:52:52.991457Z digest=sha256:ad046b25f3abe53fb4ce00417132202003eefddefa116b76183ab3d8e3346fd6

Observation 1fea8405-db0e-418c-a4ed-095cc6fcfc9b · outbound

This paper cites Learning spatiotemporal and motion features in a unified 2d network for action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Learning spatiotemporal and motion features in a unified 2d network for action recognition,

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:58.958930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:52:53.056544Z digest=sha256:9d619f3e285e1e5cec6c1cfb81c263e7de1a3e033d3d1c365e84c43d7dbab2c4

Observation 72f9b3c9-d38f-4a75-b885-c9606e8b1c55 · outbound

This paper cites Vit-ret: Vision and recurrent transformer neural networks for human activity recognition in videos,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Vit-ret: Vision and recurrent transformer neural networks for human activity recognition in videos,

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:58.805940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:52:53.175022Z digest=sha256:b1a4980fb050619cb02cf7b702e23678b7c8ffc851b179acfc89edf10a953b3c

Observation 6274f1a9-a5e3-4101-ace2-d631abf33154 · outbound

This paper cites Spatial-temporal interleaved net- work for efficient action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Spatial-temporal interleaved net- work for efficient action recognition,

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:58.668315Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:52:53.283537Z digest=sha256:5421694abacc3c26dbef11b48b9b6835d02e3ec2c4c56daa7ea8feffe15f0048

Observation 35e04df3-1d2c-4336-90ce-67622692a76b · outbound

This paper cites A hybrid transformer framework for efficient activity recog- nition using consumer electronics,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition A hybrid transformer framework for efficient activity recog- nition using consumer electronics,

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:58.469065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:52:53.396459Z digest=sha256:ba88ccfdb256a766b814d2199d27462f5296702e0ea02e737e5f9e6cfbb5636f

Observation affa9eb6-feaf-49ce-ac9c-4e00473fd572 · outbound

This paper cites A knowledge-based hierarchical causal inference network for video action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition A knowledge-based hierarchical causal inference network for video action recognition,

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:58.287590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:52:53.500055Z digest=sha256:508714fb3be8b0dab8faabd9558730d487170adb44618af0735830344990ad9d

Observation 158182ec-519a-4c5b-9884-f96baf7d8dfb · outbound

This paper cites Is space-time attention all you need for video understanding?.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Is space-time attention all you need for video understanding?

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:53.559512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:53.559512Z digest=sha256:a6a0718002d8f243761d79f0e9481eb7ce7b8494d5c9938da8971f4eabf23898

Observation 070b335d-9239-4dfb-b0cb-f10f38c64ee1 · outbound

This paper cites Vidtr: Video transformer without convolutions,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Vidtr: Video transformer without convolutions,

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:58.141635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:52:53.632514Z digest=sha256:690e5fd9922e6797a676409451b84e6b9b42f6a8b19860e6ca79ddc2df1899e9

Observation cc5a6118-9795-47ed-af13-2a4b5f8d00de · outbound

This paper cites Keeping your eye on the ball: Tra- jectory attention in video transformers,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Keeping your eye on the ball: Tra- jectory attention in video transformers,

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:57.961367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:52:53.714960Z digest=sha256:2fe5b1bf422f6a4fcf3cc78b737f68d3d1664ade3af5fdeb36e25eee3ca00c2e

Observation 7be14330-1fd3-4783-b88b-c3567693fad5 · outbound

This paper cites Multiscale vision transformers,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Multiscale vision transformers,

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:57.813054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:52:53.837040Z digest=sha256:7ce05981e3cb9e8ee5151f8aec97140b407fc23967d882cd2b8bfbe98d89e0b8

Observation 30bfd305-887a-4872-9621-edd153208baf · outbound

This paper cites Multiview transformers for video recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Multiview transformers for video recognition,

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:57.677629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:52:53.978259Z digest=sha256:fe56e1e99c4871bc29ed4b4b5cf6b3524db85f0d3b4d8bff6e51d01dbd5990c9

Observation f796a309-6d28-429e-ab81-1870fb80b295 · outbound

This paper cites Video swin transformer,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Video swin transformer,

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:57.567092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:52:54.063447Z digest=sha256:98316416d8f553d7577c6456e3f6bf320dfbad362e7b1f4ce8a3f04293c56d35

Observation 7af3d5bd-983f-4ee8-a0c6-9cd773fde5e9 · outbound

This paper cites Mvitv2: Improved multiscale vision transformers for classification and detection,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Mvitv2: Improved multiscale vision transformers for classification and detection,

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:57.452633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:52:54.121446Z digest=sha256:1152c616463c5914c05756dd672d7f86b528d6990eb7a245057f3b704d9abea9

Observation 1ac19103-74cc-4d7a-b7f4-eef175cdbd75 · outbound

This paper cites A novel spatio-temporal-wise network for action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition A novel spatio-temporal-wise network for action recognition,

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:57.285877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:52:54.207067Z digest=sha256:2b051e6fbe4d2ce39d909418a810b5825dabe3cc285a8c88a8a56f5eeefc91d4

Observation bf3879a6-426f-4218-82e6-3ee66704101c · outbound

This paper cites D-tsm: Discriminative temporal shift module for action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition D-tsm: Discriminative temporal shift module for action recognition,

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:57.163489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:52:54.282421Z digest=sha256:107fe6c8f8464cd26016bb30954b71c2ca72e87244f4c12548c3ee9732bf8f0a

Observation e352fb0d-f249-40ef-b2c8-c164288a4a29 · outbound

This paper cites Scene adaptive mechanism for action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Scene adaptive mechanism for action recognition,

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:57.021801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:52:54.344367Z digest=sha256:859a06741dc25227e2c05f8a20cb1febc942519b6a6ef0f7b0ff22c2c9cf2dff

Observation 8affa414-54cd-4ff1-b837-4f01b33aaef0 · outbound

This paper cites Sta+: Spatiotemporal adaptation with adaptive model selection for video action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Sta+: Spatiotemporal adaptation with adaptive model selection for video action recognition,

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:56.855411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:52:54.396120Z digest=sha256:f3222e0cc73a2291877a9ecfc3fb56cf81cfbd62d1daf82a02caafe7d3b52a32

Observation bab42be6-c522-4b76-95fa-665b474a8812 · outbound

This paper cites Short-term action learning for video action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Short-term action learning for video action recognition,

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:56.720156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:52:54.487275Z digest=sha256:b14f6ad8dcaa9481f2bad3badf646aca72438c6c371ce3cc6618c40b24508d74

Observation 75ed9bf1-ed5b-4bae-a3e7-d754c807b0ba · outbound

This paper cites Tea: Temporal excitation and aggregation for action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Tea: Temporal excitation and aggregation for action recognition,

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:56.585739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:52:54.566351Z digest=sha256:5929b5d09ff0cd2a6218c2476f9dd21245f2374d5952e67e321fbf6c71055694

Observation 1c59ee0d-d2d0-4368-8eab-645b74f76a7e · outbound

This paper cites Movinets: Mobile video networks for efficient video recogni- tion,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Movinets: Mobile video networks for efficient video recogni- tion,

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:56.394162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:52:54.657092Z digest=sha256:e544745507e9014a5648a723009a571ba67eb5bf7fd911d654f5bb0945739804

Observation 3901d6cc-435a-4bb1-87dc-5176f6e73f41 · outbound

This paper cites Timebal- ance: Temporally-invariant and temporally-distinctive video represen- tations for semi-supervised action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Timebal- ance: Temporally-invariant and temporally-distinctive video represen- tations for semi-supervised action recognition,

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:56.180372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:52:54.746931Z digest=sha256:0998562c2afea34cb9bce74e2f891702d43df565c7d6e65fa1cbc2013d0db599

Observation 89d04c6f-4153-4f0e-b987-06a169856d17 · outbound

This paper cites Dilated multi-temporal modeling for action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Dilated multi-temporal modeling for action recognition,

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:56.035863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:52:54.862860Z digest=sha256:2fbc03aa8d93dbcfefe82e211e4fc49f18cf36d66c395ce000cd018727f63533

Observation 2d9693e9-22a7-4efb-9789-44932768db2c · outbound

This paper cites Learning Discriminative Spatio-temporal Representations for Semi-supervised Action Recognition.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Learning Discriminative Spatio-temporal Representations for Semi-supervised Action Recognition

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:54.894113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:54.894113Z digest=sha256:364b764baf034ef3a62bff73374527a5f917ec4bd6f35978f5b87bebe8c2276a

Observation a4bdf994-b313-41e5-a2f3-029c2ad88fd7 · outbound

This paper cites Discrimina- tive segment focus network for fine-grained video action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Discrimina- tive segment focus network for fine-grained video action recognition,

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:55.869769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:52:54.944965Z digest=sha256:ebf2ababb17fa273abaa89081d13569f8b8d6ba37e50c8bc65e82bf55aea38f2

Observation 37f17f06-1d35-4c0e-ad1f-0514af327267 · outbound

This paper cites Temporal difference attention for action recog- nition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Temporal difference attention for action recog- nition,

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:55.666237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:52:54.978639Z digest=sha256:17ae03d30ba9dfaec4cbc43ada9e7999e669c602f137459ef5505b386b09576d

Observation 62891965-1003-479c-9435-0d893005c4e6 · outbound

This paper cites An efficient motion visual learning method for video action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition An efficient motion visual learning method for video action recognition,

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:55.515616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:52:55.019874Z digest=sha256:373a77135ca80f6b8a1d88432e19bdffa0d77976dd28004b48ec6941b04af535

Pith citing papers

No inbound Pith citation observations are available.