Pith. sign in

Paper Citation Record · LEDGER

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition

As of 20 August 2026, this Paper Citation Record lists 95 of 95 outbound references and 0 inbound Pith citation observations for arXiv:2507.12426.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.12426 v2

Coverage vector

measured 95 of 95 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T16:52:55.019874Z

measured 95 of 95 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

95 of 95 outbound references displayed

  • verified exact2
  • verified fuzzy70
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7d6fa9a7-95b3-4e9b-b0c8-bfc6deb3b292 · outbound

This paper cites Quo vadis, action recognition? a new model and the kinetics dataset,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Quo vadis, action recognition? a new model and the kinetics dataset,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:46.543670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:46.543670Z digest=sha256:d20c661726d430b1d594aa6381c26a9f93e24d278177b80e51bb9728687f84fd

Observation f6dfecc4-6bac-46de-8600-043905dbef6d · outbound

This paper cites Spatiotemporal residual networks for video action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Spatiotemporal residual networks for video action recognition,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:46.603264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:46.603264Z digest=sha256:e392fd7a45c49ab39953392ca96c62202e846dbdc055dac3367640246b2df84f

Observation da7f7c26-20f2-4582-829d-229133615bd3 · outbound

This paper cites Learning spatiotemporal features with 3d convolutional networks,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Learning spatiotemporal features with 3d convolutional networks,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:46.762590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:46.762590Z digest=sha256:0208f58e43fd22059d491da0433d347841609bea993c07987f0463a5cf3031ab

Observation 1139d00a-d086-4239-9ee1-f7fbbe824386 · outbound

This paper cites Large-scale video classification with convolutional neural networks,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Large-scale video classification with convolutional neural networks,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:46.834049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:46.834049Z digest=sha256:82f5cd0ea864f5350f3893eca1da5ab40522362ce44e46939740c3a363aea89c

Observation 81df5fdd-3704-4da7-864d-36ff4ef50b15 · outbound

This paper cites Beyond short snippets: Deep networks for video classification,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Beyond short snippets: Deep networks for video classification,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:46.912945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:46.912945Z digest=sha256:b36654bb864ebd60c6557888cd255df1fdc6eaea111001b3a1883ffff334d7df

Observation 30e07942-23e6-46a0-9685-6d2274f38ca8 · outbound

This paper cites Two-stream convolutional networks for action recognition in videos,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Two-stream convolutional networks for action recognition in videos,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:46.999396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:46.999396Z digest=sha256:def9304eafac12b587408f67fd65a84c65a0ef02ea762416323c388f6c33f196

Observation 2ab53b6f-85c6-446e-b41c-2ad3340312f1 · outbound

This paper cites Action recognition using deep 3d cnns with sequential feature aggregation and attention,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Action recognition using deep 3d cnns with sequential feature aggregation and attention,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:47.061463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:47.061463Z digest=sha256:6ca7ee0200268144b588b2f91cf299a06b10f99f0b2ebc0df8c76bdd6b297f9d

Observation befd3be9-d7eb-4f72-a1a3-e92c65ab6b2a · outbound

This paper cites A closer look at spatiotemporal convolutions for action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition A closer look at spatiotemporal convolutions for action recognition,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:47.139621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:47.139621Z digest=sha256:a5246a7d77974672a8b495d65e8c575c37863179f67000ce1fe5120844a7e524

Observation 89171f78-0d81-4a0f-8b44-9d16f4498847 · outbound

This paper cites Human action recognition in videos using convolution long short-term memory network with spatio-temporal networks,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Human action recognition in videos using convolution long short-term memory network with spatio-temporal networks,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:47.215422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:47.215422Z digest=sha256:3deae609d769abe98ed7f6443c142c8537531f1ff4bf8a58ae2ef7e9ae015b54

Observation ea61037a-f432-42f1-8c1e-5bc5d0ed4871 · outbound

This paper cites Action recognition in videos using pre-trained 2d convolutional neural networks,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Action recognition in videos using pre-trained 2d convolutional neural networks,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:47.293116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:47.293116Z digest=sha256:162874b0f924d1b72609120561b0b1bfaf314ac552796cb18bbc47113696abe4

Observation 57809d8d-af21-4dad-a9c4-caa9ce0ce57b · outbound

This paper cites Video-focalnets: Spatio-temporal focal modulation for video action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Video-focalnets: Spatio-temporal focal modulation for video action recognition,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:47.365973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:47.365973Z digest=sha256:c56a1d51f40fc03ae93cefc1697fcbe8cf82ec8bb9b323d238cc7ba9df8a035c

Observation cbd63ea4-ac98-4630-b0d5-2f34a2a50e21 · outbound

This paper cites Attention is all you need,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Attention is all you need,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:47.458192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:47.458192Z digest=sha256:edd178298a7a9598d0cbabf6cdd5535566a44965e3036b7398c8facc7208d8e1

Observation 94573c9d-c9d9-405f-a809-205fd9385d3b · outbound

This paper cites Morph: flexible acceleration for 3d cnn-based video understanding,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Morph: flexible acceleration for 3d cnn-based video understanding,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:47.549508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:47.549508Z digest=sha256:84dd825a79e891b519bd1bd8a325aecffa269a08421e2f34d9768955c90cf8cd

Observation 7dacace1-b6ab-4d7f-a2f6-134e5c83ccb9 · outbound

This paper cites Video swin transformer,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Video swin transformer,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:47.646960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:47.646960Z digest=sha256:e408815bac548c5b6f17087435867c5b118d630d48924008545a568061d60eb4

Observation f0cf805b-b70b-45ea-9f7b-4c0d38c27a59 · outbound

This paper cites Is space-time attention all you need for video understanding?.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Is space-time attention all you need for video understanding?

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:47.720373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:47.720373Z digest=sha256:c6dd7fa276fa785d661f80d6ce7e80391cd0d061316f8eea0dd3f517fa5461ca

Observation 3eb7b69d-ddd7-49e0-9c4f-fe022bf11c61 · outbound

This paper cites Vivit: A video vision transformer,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Vivit: A video vision transformer,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:06.347041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T16:52:47.802678Z digest=sha256:94e909ceec26cd84fc3edfefe58360b71fd0e302694c00451a896d293728d492

Observation 82e5e3f7-b6c3-41a8-9fc9-f6b66ec19b23 · outbound

This paper cites Video swin transformer,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Video swin transformer,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:06.256262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T16:52:47.915356Z digest=sha256:435d4b0edd77ee1b3e13e3c832dd11be32093339b19e3497e06652ca586b53d0

Observation 05571d0b-4696-43a4-a4c0-f3ef2b565420 · outbound

This paper cites Multiview transformers for video recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Multiview transformers for video recognition,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:06.027968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T16:52:48.002422Z digest=sha256:79bd98c0fc8d52c986085f0921df04fe4bb2830bf5fc273f073f64c8fdf002c4

Observation d19db6ac-64a3-4d5b-8bc4-38a75d42c01d · outbound

This paper cites Vivit: a video vision transformer,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Vivit: a video vision transformer,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:05.838377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T16:52:48.113860Z digest=sha256:57ac49609cc9f1bc9209c3467722770ae5fee5f7ff70198ee57258beefb4b27f

Observation b1bcdfe6-4cdc-44f8-aad7-61da12e91435 · outbound

This paper cites A Short Note on the Kinetics-700 Human Action Dataset.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition A Short Note on the Kinetics-700 Human Action Dataset

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:48.257823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:48.257823Z digest=sha256:7ff0cf40340cddac7a3b866e91bcfc6b039b7c1d56bdab225f25d3c0d3d2ca56

Observation a818e4a1-4b5e-4a94-b1fa-ad39061bc3c9 · outbound

This paper cites The" something something.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition The" something something

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:05.619678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T16:52:48.310977Z digest=sha256:db73ed77001253ab7f2b5552e9899e33c4e4145992dcef97cdc5dd1987b680f9

Observation 619217a0-41a1-4d5f-bb13-e281e3ab75a4 · outbound

This paper cites Tsnet: token sparsification for efficient video transformer,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Tsnet: token sparsification for efficient video transformer,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:05.374484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T16:52:48.397443Z digest=sha256:0a4345ec2212435368430d280ed0c64607c8aaa762a46edeb4b5303edabf7ae3

Observation 14426831-847c-42f1-a119-ab7de1c633fd · outbound

This paper cites Dualformer: local- global stratified transformer for efficient video recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Dualformer: local- global stratified transformer for efficient video recognition,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:05.198144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T16:52:48.495776Z digest=sha256:33b8ebc9ef12554d1e957081f3ae0bb6251de8acb9a61fa42c73e4df755f61e0

Observation 9e639f13-90b9-4e10-890f-1c9f41288763 · outbound

This paper cites Aerobics action recognition algorithm based on three-dimensional convolutional neural network and multilabel clas- sification,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Aerobics action recognition algorithm based on three-dimensional convolutional neural network and multilabel clas- sification,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:05.032271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T16:52:48.645264Z digest=sha256:5eeacf96defd90eb07f32e60f15cf484ab2beb444a9edfd3910bd16f5ebeaa74

Observation ecd963af-485a-4dc1-8fed-b62f79566714 · outbound

This paper cites Tsm: temporal shift module for efficient video understanding,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Tsm: temporal shift module for efficient video understanding,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:04.810289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T16:52:48.745524Z digest=sha256:414d46a25ae2010d647562cd98c5832faef5c08c2c17839fe98189ad5a5e9218

Observation 02dadde9-ea1b-43a4-9339-e281a66574ed · outbound

This paper cites Multi-stream interaction networks for human action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Multi-stream interaction networks for human action recognition,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:48.903466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:48.903466Z digest=sha256:8cee01b30481283e1e8380097cc2d1e7568a09692f69f0922a9772e63edfaada

Observation f4125143-b5a1-47c0-ad04-f863252f597f · outbound

This paper cites Spatio- temporal adaptive network with bidirectional temporal difference for action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Spatio- temporal adaptive network with bidirectional temporal difference for action recognition,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:04.577825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T16:52:48.988513Z digest=sha256:137a6a25d2a10268b797c4baa428f9d366d503949b0d98d7f0933deb9ba31e38

Observation 7239b067-039d-4a18-b7ed-b876d7c032a5 · outbound

This paper cites Agpn: Action granularity pyramid network for video action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Agpn: Action granularity pyramid network for video action recognition,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:04.413558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T16:52:49.052816Z digest=sha256:84a3bb185c82eac98bb5172f51e72d9d7ca91d17208a44893646527b328704fa

Observation 10739f60-4c9f-4362-bc5d-71672cacce9c · outbound

This paper cites Mawkdn: A multimodal fusion wavelet knowledge distillation approach based on cross-view attention for action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Mawkdn: A multimodal fusion wavelet knowledge distillation approach based on cross-view attention for action recognition,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:04.180393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T16:52:49.130816Z digest=sha256:cb4feb016cc4d9113de8144600a2b9365867d898dfb21591bf86f2cb58712131

Observation a3b87581-b7aa-4cce-8d76-78178a971912 · outbound

This paper cites Convolutional neural networks or vision transformers: who will win the race for action recognitions in visual data?.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Convolutional neural networks or vision transformers: who will win the race for action recognitions in visual data?

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:04.004396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T16:52:49.203358Z digest=sha256:2488b1bdf5ffeace05b8b735829ad4e3e1270d517b37d5121e476a36d9161d91

Observation dc449c7c-7aab-4408-9730-7df66158556e · outbound

This paper cites Decoupled knowledge distillation,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Decoupled knowledge distillation,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:03.788026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T16:52:49.291727Z digest=sha256:6ac1e29f9c9e5f115c54ed6af55e4648de557158b35b37e7a0c100aeaa0a25b7

Observation 6c54a01c-3a35-469e-8b2d-214a8e7009c1 · outbound

This paper cites Knowledge distil- lation in video-based human action recognition: an intuitive approach to efficient and flexible model training,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Knowledge distil- lation in video-based human action recognition: an intuitive approach to efficient and flexible model training,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:03.616547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T16:52:49.359338Z digest=sha256:eb01e753a37eb9ce0b1418bb4f0e7c44f7ea2e6a821d29651a2ece04047921f9

Observation 0c4740cb-6e71-425d-b39b-8bb395344d67 · outbound

This paper cites Tomato leaf disease recognition based on multi-task distillation learning,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Tomato leaf disease recognition based on multi-task distillation learning,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:03.441174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T16:52:49.445207Z digest=sha256:1a42fa1bd852bbd0dcfbe608f727b34172931721bf9e3bfef38528866c34b45a

Observation abc9a7cf-cff3-428b-86d2-4a8ec455593e · outbound

This paper cites Videoadviser: video knowledge distillation for multimodal transfer learning,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Videoadviser: video knowledge distillation for multimodal transfer learning,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:03.264063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T16:52:49.510135Z digest=sha256:800fc7b0e253731f519d7511f5c3b45f4cc115d2bbf83e9c5e2837b2263ae695

Observation f0e5042d-7bdc-44c7-ad1e-eb7de09209f5 · outbound

This paper cites Generative model- based feature knowledge distillation for action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Generative model- based feature knowledge distillation for action recognition,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:03.099709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T16:52:49.627318Z digest=sha256:0160dd32ef34b06e25165c1d75aa89059c749390b3ff8f77c16d97f0ce5be236

Observation 7ac67c90-6540-4d09-9d3f-49fb1a495643 · outbound

This paper cites Distillation of human-object interaction contexts for action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Distillation of human-object interaction contexts for action recognition,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:02.880266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T16:52:49.680694Z digest=sha256:a53b93e02bbfb8baed71c8867277263ad81a390ca33b0db288a7d53d9001ff00

Observation 51a66055-0d5c-4e05-ae76-efc69f911c32 · outbound

This paper cites Gaussian Error Linear Units (GELUs).

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Gaussian Error Linear Units (GELUs)

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:49.779462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:49.779462Z digest=sha256:fd70860692d5594e6779a11e9bd6e809d5b90d8531fd83e0491f6ae242167c97

Observation 42c9bee4-121f-49f0-bb40-e03ca6c20299 · outbound

This paper cites Recognizing 50 human action categories of web videos,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Recognizing 50 human action categories of web videos,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:02.753222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T16:52:49.860074Z digest=sha256:ec88915671fd8f6da983faa83fb7481b430959e87abc91abbd049fc9c96d4827

Observation 9a6faf9b-3c84-4091-854b-7615a44fba10 · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:50.009362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:50.009362Z digest=sha256:6fb1567d3c6ca1d4ebe1aacbe52513630ddb465d38e8e25c3ea51e5ab30cd7e1

Observation 3f17c9e8-b6e2-4c9c-bddf-97a6c44e5f88 · outbound

This paper cites Hmdb: a large video database for human motion recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Hmdb: a large video database for human motion recognition,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:02.650847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T16:52:50.106813Z digest=sha256:be813435a02032b8365463aa6ed19880b45495eea12ace233eeeb15a9fc03d83

Observation c75510fe-4399-4f99-b634-5292bf4a285f · outbound

This paper cites The" something something.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition The" something something

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:02.556260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T16:52:50.174304Z digest=sha256:6a0c03c609966c5b5ef8715cbea5aadb09c6bdd27d51ce42771809813531cdef

Observation b3a55fc2-d47e-44ba-bc91-dbe64eb2618d · outbound

This paper cites The Kinetics Human Action Video Dataset.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition The Kinetics Human Action Video Dataset

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:50.256104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:50.256104Z digest=sha256:a1c24dc369faed79284b921c578771d05e74cdb4bce0b9fee2c511bdf071b3b8

Observation 1c01ee1e-7001-44a3-a8f4-b282d3eed1db · outbound

This paper cites UniFormer: Unified Transformer for Efficient Spatiotemporal Representation Learning.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition UniFormer: Unified Transformer for Efficient Spatiotemporal Representation Learning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:50.390492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:50.390492Z digest=sha256:9a5476fab27b6c7333da39e70b4bae9fde0513af2167fa822c1d5c6dbe2ea00c

Observation 771bb7c5-6e0d-495d-8627-a7fe6aec94b5 · outbound

This paper cites Going deeper with convolutions,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Going deeper with convolutions,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:02.389286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T16:52:50.471227Z digest=sha256:085fc5e2fa44fa85429bea4e7cddc8c8dd89c505f1fe3c34d8ae70a369cfbf74

Observation 24febacd-fbf9-4fa0-95f4-43b6defeb682 · outbound

This paper cites Making sense of neuromorphic event data for human action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Making sense of neuromorphic event data for human action recognition,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:02.263934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T16:52:50.520535Z digest=sha256:3a1431c2fc3b7ee846484de52ce08054d7ae733862a4b2a6f3c0b75f6c8d9768

Observation d670189b-3b5f-453a-9bb4-193e9c814505 · outbound

This paper cites Human action recognition using dis- tance transform and entropy based features,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Human action recognition using dis- tance transform and entropy based features,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:02.200658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T16:52:50.601096Z digest=sha256:b5f34e3d688d2aac1b632ff2b61f48890a19c7209bf65e820219411bd38424ef

Observation 121473c8-a8a4-402a-bb27-b0ab77aab99d · outbound

This paper cites Human action recognition using hybrid deep evolving neural networks,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Human action recognition using hybrid deep evolving neural networks,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:02.116040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T16:52:50.719001Z digest=sha256:787175e58988df5d454bfd06a89bf3b00418dc6ad6e7791cc213f1d8520abe58

Observation 36e68b1f-41af-4cf8-b42b-2f5eebc74e94 · outbound

This paper cites Simple-action-guided dictionary learning for complex action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Simple-action-guided dictionary learning for complex action recognition,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:02.035273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T16:52:50.830577Z digest=sha256:a4da3cb1b41a125e12a7824a72f5fbff439e3cc70e6dd098125d680de9787b63

Observation 6a77b057-5bb1-4b19-9eeb-6b4c9707ba87 · outbound

This paper cites Human activity classification using the 3dcnn architecture,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Human activity classification using the 3dcnn architecture,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:01.921818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T16:52:50.937287Z digest=sha256:d216132394f729671b597ea197eb31986b3693d799d78684dabc18b74e262972

Observation 46e3ecc0-c91d-4f86-8982-338625a5b2c6 · outbound

This paper cites Fast classification and action recognition with event-based imaging,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Fast classification and action recognition with event-based imaging,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:01.770390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T16:52:51.096512Z digest=sha256:c8fa272f2c21b591a299122e7aeaa5f9e137675b5ab4e2501fae20066f9d9292

Observation 04678e70-a280-4266-b2bf-1cbc78009620 · outbound

This paper cites Spatio-temporal features based human action recognition using convolutional long short-term deep neural network,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Spatio-temporal features based human action recognition using convolutional long short-term deep neural network,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:01.610341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T16:52:51.214457Z digest=sha256:4c18fac11a1b6195dd50242f0855e89573966d7bf21a12c36fe9db0bcbf14cdb

Observation 26c54948-d8b3-443e-b9b7-c1fb8ddcacac · outbound

This paper cites Human action recognition using multi-stream fusion and hybrid deep neural networks,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Human action recognition using multi-stream fusion and hybrid deep neural networks,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:01.479218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T16:52:51.327664Z digest=sha256:09522c90e9a5eea39ed813b15fa052a5d156f96484ddb3b0670f343cf6bb17ff

Observation 51710b76-8eed-4762-b895-c9825289873f · outbound

This paper cites Self-supervised video representation learning by uncovering spatio-temporal statistics,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Self-supervised video representation learning by uncovering spatio-temporal statistics,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:01.420896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T16:52:51.425880Z digest=sha256:176ee20039b1011f2a4639f5d24dd0ddbe1e6037663b56a3f6b86b3789479bbc

Observation 7057970c-2ce6-4cfe-a130-d6867f5205b4 · outbound

This paper cites Enhancing self-supervised video representation learning via multi-level feature optimization,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Enhancing self-supervised video representation learning via multi-level feature optimization,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:01.332744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T16:52:51.505885Z digest=sha256:5e7e912f45738f04e78bbdb2024414815c82021637808c697a83a6e739762bf2

Observation d59234f2-5c29-47a9-83f8-a177742866da · outbound

This paper cites Videomoco: Contrastive video representation learning with temporally adversarial examples,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Videomoco: Contrastive video representation learning with temporally adversarial examples,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:01.169505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T16:52:51.608989Z digest=sha256:3490957255170d3d57a15321d0f7e8ae39a44ab0d2061175bd673bbada179ccf

Observation ef665412-a01e-4b14-ae2b-f6664f0c0d74 · outbound

This paper cites Action recognition from a single coded image,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Action recognition from a single coded image,

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:01.043969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T16:52:51.708625Z digest=sha256:9b78229fbe86fa9e75907ee6ac109110388d48caab2d5d6ef0262964bf2d4ea9

Observation 7c3b9fb0-aafb-4894-bebe-e9dd1baf0f02 · outbound

This paper cites Tclr: Temporal contrastive learning for video representation,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Tclr: Temporal contrastive learning for video representation,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:00.949242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T16:52:51.782435Z digest=sha256:26c9a359c24af1ab5c3a907631c5ce88245eca79005d2319bf01a51b0ac98b39

Observation 8e4ccfc8-44d5-4f05-b48a-8474961d34c8 · outbound

This paper cites Learn2augment: learning to composite videos for data augmentation in action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Learn2augment: learning to composite videos for data augmentation in action recognition,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:00.857703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T16:52:51.887584Z digest=sha256:398adb555cfe0277f51adc64522248218f661c72b93f879ada482201ccf25401

Observation 5e96c946-27dc-4cae-b725-e151ef84c116 · outbound

This paper cites Learning from temporal gradient for semi-supervised action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Learning from temporal gradient for semi-supervised action recognition,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:00.657387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T16:52:51.993613Z digest=sha256:374570313e505f424a380877feda48823d31e9bdc1ee77a5408b1b49f985c252

Observation 627c3103-b39a-40f0-aac2-395d19dc5380 · outbound

This paper cites Preserve Pre-trained Knowledge: Transfer Learning With Self-Distillation For Action Recognition.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Preserve Pre-trained Knowledge: Transfer Learning With Self-Distillation For Action Recognition

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-08-06T16:52:55.326931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T16:52:52.098718Z digest=sha256:11db3505de4e313d0e0e18ac82e3a6bd648ef9c5b405256f15749ece9e28bfe8

Observation 64a00eee-4644-447b-b2c3-22bcf3d2d01a · outbound

This paper cites Extreme low- resolution action recognition with confident spatial-temporal attention transfer,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Extreme low- resolution action recognition with confident spatial-temporal attention transfer,

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:00.478916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T16:52:52.195968Z digest=sha256:7f4f2a8ca810fbcdfb841909416d96477dd44d07e8e17bb4eb609197527d43de

Observation 0a6f9449-9838-4d17-a28b-95f3b333cdee · outbound

This paper cites Self-supervised video-based action recognition with disturbances,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Self-supervised video-based action recognition with disturbances,

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:00.286153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T16:52:52.315336Z digest=sha256:b4761c355b47e13b29f1e2da4908044674b11cbba635d30f6751df94ab25dc31

Observation 24aaf547-bc02-420f-8f6f-ae968884f75f · outbound

This paper cites Spatial-temporal exclusive capsule network for open set action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Spatial-temporal exclusive capsule network for open set action recognition,

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:00.146160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T16:52:52.428855Z digest=sha256:1908a4992b8c4c52c9026f16410244f6c7b08673d8f7bb42f3f9ea5d5e88f4b9

Observation 7cff12ac-649e-46be-b717-016115aaf030 · outbound

This paper cites Sv- former: Semi-supervised video transformer for action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Sv- former: Semi-supervised video transformer for action recognition,

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:59.986980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T16:52:52.509199Z digest=sha256:44094029695967acb5a1964c4829410f02ef285cbc8e650f9cc678ceefd6f108

Observation 6e74282b-f332-4229-bd4c-a2008662ed7a · outbound

This paper cites ActionHub: A Large-scale Action Video Description Dataset for Zero-shot Action Recognition.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition ActionHub: A Large-scale Action Video Description Dataset for Zero-shot Action Recognition

Reference 66

Resolution
verified exact
local_arxiv, observed 2026-08-06T16:52:55.174880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T16:52:52.604016Z digest=sha256:fe17a47ae7a3fc58e89551b852067439a35dae7544767f94a8507014597f4a95

Observation 3ac35da3-01a0-4f25-8917-be0cde4a5013 · outbound

This paper cites Self-supervised learning via multi-transformation classification for action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Self-supervised learning via multi-transformation classification for action recognition,

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:59.802805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T16:52:52.730793Z digest=sha256:915403e3a2402e14dcae3a504c6347378ccd3b91afdf7bfd59cc854793db687f

Observation 48d46bd4-9904-490f-9eaa-43495d8f3c7f · outbound

This paper cites Semi-supervised action recog- nition with dynamic temporal information fusion,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Semi-supervised action recog- nition with dynamic temporal information fusion,

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:59.610844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T16:52:52.788124Z digest=sha256:74be7beabaae64698c61bb501d3170bcfbd175f12c5b3b7fc6d4847198a30e99

Observation fa98add2-c079-44e4-aac9-8e0bc6c79754 · outbound

This paper cites Spatiotemporal contrastive video representation learning,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Spatiotemporal contrastive video representation learning,

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:59.422918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T16:52:52.852312Z digest=sha256:9b27b098b977efe6dcaa30bacaa31cec0f134166f173efbeb68d1511418cbc9a

Observation 9dd942c3-2298-4b33-bc85-24eb0ffdef8a · outbound

This paper cites Representation learning for compressed video action recognition via attentive cross- modal interaction with motion enhancement,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Representation learning for compressed video action recognition via attentive cross- modal interaction with motion enhancement,

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:59.292874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T16:52:52.921579Z digest=sha256:e4bb68e0cd717c8f96b6129b36275ce19cac2102e9ca9b139bde3f311e7ad912

Observation 001c1faf-2447-4562-8b39-aec022990205 · outbound

This paper cites Motion-driven visual tempo learn- ing for video-based action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Motion-driven visual tempo learn- ing for video-based action recognition,

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:59.149388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T16:52:52.991457Z digest=sha256:90c03701fcd4426abf4ea40e3e74357c5379d58583d5158b8737bef147c10acc

Observation 1fea8405-db0e-418c-a4ed-095cc6fcfc9b · outbound

This paper cites Learning spatiotemporal and motion features in a unified 2d network for action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Learning spatiotemporal and motion features in a unified 2d network for action recognition,

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:58.958930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T16:52:53.056544Z digest=sha256:0d063a726ab22b9ae403b5f961300c1ff71eba4d6f29ec31342e5b3f1a45bd16

Observation 72f9b3c9-d38f-4a75-b885-c9606e8b1c55 · outbound

This paper cites Vit-ret: Vision and recurrent transformer neural networks for human activity recognition in videos,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Vit-ret: Vision and recurrent transformer neural networks for human activity recognition in videos,

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:58.805940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T16:52:53.175022Z digest=sha256:61abb827f486b3d0bdad4e11c71ebcd01007188e4fd929ab26b6428e9974d476

Observation 6274f1a9-a5e3-4101-ace2-d631abf33154 · outbound

This paper cites Spatial-temporal interleaved net- work for efficient action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Spatial-temporal interleaved net- work for efficient action recognition,

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:58.668315Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T16:52:53.283537Z digest=sha256:10adfb804fda923274735ae24bbc9a965b2afee0a89a6a72e82f76f8200d9bf2

Observation 35e04df3-1d2c-4336-90ce-67622692a76b · outbound

This paper cites A hybrid transformer framework for efficient activity recog- nition using consumer electronics,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition A hybrid transformer framework for efficient activity recog- nition using consumer electronics,

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:58.469065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T16:52:53.396459Z digest=sha256:0ba1c301004673001fca8d8a29c5fc29df5102d9bb62f53876a7d3f356cf3221

Observation affa9eb6-feaf-49ce-ac9c-4e00473fd572 · outbound

This paper cites A knowledge-based hierarchical causal inference network for video action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition A knowledge-based hierarchical causal inference network for video action recognition,

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:58.287590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T16:52:53.500055Z digest=sha256:aea8313d1aca17eab2251b1d5b4e77fea5f5a9d4b7a3cd62a4068a4e015da3f4

Observation 158182ec-519a-4c5b-9884-f96baf7d8dfb · outbound

This paper cites Is space-time attention all you need for video understanding?.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Is space-time attention all you need for video understanding?

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:53.559512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:53.559512Z digest=sha256:6675a91c15650005ffb7a281189714997cddc108537a1715b133b7d7fadd6eda

Observation 070b335d-9239-4dfb-b0cb-f10f38c64ee1 · outbound

This paper cites Vidtr: Video transformer without convolutions,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Vidtr: Video transformer without convolutions,

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:58.141635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T16:52:53.632514Z digest=sha256:acc8b76563d6cf6c94d27c82fd09a2bf595fdc94fc5c1a9e76eac39b9d71cebf

Observation cc5a6118-9795-47ed-af13-2a4b5f8d00de · outbound

This paper cites Keeping your eye on the ball: Tra- jectory attention in video transformers,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Keeping your eye on the ball: Tra- jectory attention in video transformers,

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:57.961367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T16:52:53.714960Z digest=sha256:b2c7220fa9ea6047aa2c48f37ff814b052df665d01c140b82ebf8cb15cf2ee68

Observation 7be14330-1fd3-4783-b88b-c3567693fad5 · outbound

This paper cites Multiscale vision transformers,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Multiscale vision transformers,

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:57.813054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T16:52:53.837040Z digest=sha256:2bab5d5cc8f596b7a4846cfda06082270027cf2a30feb37bf6cc42c6a45e9b1d

Observation 30bfd305-887a-4872-9621-edd153208baf · outbound

This paper cites Multiview transformers for video recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Multiview transformers for video recognition,

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:57.677629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T16:52:53.978259Z digest=sha256:1e4f11d7978641356d4d95342d6aba1ed386f2b4f601df07568c5a476fba79c7

Observation f796a309-6d28-429e-ab81-1870fb80b295 · outbound

This paper cites Video swin transformer,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Video swin transformer,

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:57.567092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T16:52:54.063447Z digest=sha256:5b7a93cb901236aa062ef226b525ad00ab415215967096e2c610ebe8f18492d1

Observation 7af3d5bd-983f-4ee8-a0c6-9cd773fde5e9 · outbound

This paper cites Mvitv2: Improved multiscale vision transformers for classification and detection,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Mvitv2: Improved multiscale vision transformers for classification and detection,

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:57.452633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T16:52:54.121446Z digest=sha256:4521fbf530e5a4c1f2c1e7825ecd82578079ee26cba2247919bf2daadef8902e

Observation 1ac19103-74cc-4d7a-b7f4-eef175cdbd75 · outbound

This paper cites A novel spatio-temporal-wise network for action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition A novel spatio-temporal-wise network for action recognition,

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:57.285877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T16:52:54.207067Z digest=sha256:42abf3bddb3c4eb65742f1169a3303dfca221d2347dea4e1a4d6ff788d430778

Observation bf3879a6-426f-4218-82e6-3ee66704101c · outbound

This paper cites D-tsm: Discriminative temporal shift module for action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition D-tsm: Discriminative temporal shift module for action recognition,

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:57.163489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T16:52:54.282421Z digest=sha256:5a31eeeecb42c97c40ac16196b349228be813f985b1f58d437e1b0c5d901aebc

Observation e352fb0d-f249-40ef-b2c8-c164288a4a29 · outbound

This paper cites Scene adaptive mechanism for action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Scene adaptive mechanism for action recognition,

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:57.021801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T16:52:54.344367Z digest=sha256:5dbe92e6db365c83cffb49862e3cd6772ef6838689f97946796574619bc43f3a

Observation 8affa414-54cd-4ff1-b837-4f01b33aaef0 · outbound

This paper cites Sta+: Spatiotemporal adaptation with adaptive model selection for video action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Sta+: Spatiotemporal adaptation with adaptive model selection for video action recognition,

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:56.855411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T16:52:54.396120Z digest=sha256:4dd4a531e3769dbacd57cb3b7291b0a4851e6afce8605c3151f6cbbdc63f0bc3

Observation bab42be6-c522-4b76-95fa-665b474a8812 · outbound

This paper cites Short-term action learning for video action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Short-term action learning for video action recognition,

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:56.720156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T16:52:54.487275Z digest=sha256:4ba8fc544ec8722393d7dde4ba724c0ee644a64ee3193fe4dcfe2fdf3424072f

Observation 75ed9bf1-ed5b-4bae-a3e7-d754c807b0ba · outbound

This paper cites Tea: Temporal excitation and aggregation for action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Tea: Temporal excitation and aggregation for action recognition,

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:56.585739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T16:52:54.566351Z digest=sha256:05377a3061b20155dafbbaaf0b29f89645d73ddebb0ddd36d352a8c7b37f9602

Observation 1c59ee0d-d2d0-4368-8eab-645b74f76a7e · outbound

This paper cites Movinets: Mobile video networks for efficient video recogni- tion,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Movinets: Mobile video networks for efficient video recogni- tion,

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:56.394162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T16:52:54.657092Z digest=sha256:634f13a59a72e18fe76a619516d391e47d3c70f595b81d0056eafe327cc9c727

Observation 3901d6cc-435a-4bb1-87dc-5176f6e73f41 · outbound

This paper cites Timebal- ance: Temporally-invariant and temporally-distinctive video represen- tations for semi-supervised action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Timebal- ance: Temporally-invariant and temporally-distinctive video represen- tations for semi-supervised action recognition,

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:56.180372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T16:52:54.746931Z digest=sha256:e4f63062ac9b75ee6d9caf5d047e211a9e964c6d3f3f58bd37aa17bbb58038b8

Observation 89d04c6f-4153-4f0e-b987-06a169856d17 · outbound

This paper cites Dilated multi-temporal modeling for action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Dilated multi-temporal modeling for action recognition,

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:56.035863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T16:52:54.862860Z digest=sha256:30cba73ca6806e99a96604166348e2fa6768feea2c75810bb7cbb23acaffa2f7

Observation 2d9693e9-22a7-4efb-9789-44932768db2c · outbound

This paper cites Learning Discriminative Spatio-temporal Representations for Semi-supervised Action Recognition.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Learning Discriminative Spatio-temporal Representations for Semi-supervised Action Recognition

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:54.894113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:54.894113Z digest=sha256:8f01df31db41a6cd880085bb7216935f9f5fa0eae3de6724ff70a1a3cfb8f676

Observation a4bdf994-b313-41e5-a2f3-029c2ad88fd7 · outbound

This paper cites Discrimina- tive segment focus network for fine-grained video action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Discrimina- tive segment focus network for fine-grained video action recognition,

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:55.869769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T16:52:54.944965Z digest=sha256:fd18879ab4bd9f4266f0291a0632b85c4572c8cd9d09af10bde91dac778b73e9

Observation 37f17f06-1d35-4c0e-ad1f-0514af327267 · outbound

This paper cites Temporal difference attention for action recog- nition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Temporal difference attention for action recog- nition,

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:55.666237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T16:52:54.978639Z digest=sha256:80cb8006a50db8ba1c7f6d8be43166ecbe58e9eb4824217137843e6f508db52d

Observation 62891965-1003-479c-9435-0d893005c4e6 · outbound

This paper cites An efficient motion visual learning method for video action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition An efficient motion visual learning method for video action recognition,

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:55.515616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T16:52:55.019874Z digest=sha256:4ae1d861ac67e877607c3a5667755e040b2d4a139d79286d47bd63891f5c30e7

Pith citing papers

No inbound Pith citation observations are available.