Pith. sign in

Paper Citation Record · LEDGER

Multiscaled Multi-Head Attention-based Video Transformer Network for Hand Gesture Recognition

As of 11 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 0 inbound Pith citation observations for arXiv:2501.00935.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.00935 v1

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T22:44:42.468224Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

30 of 30 outbound references displayed

  • verified exact2
  • verified fuzzy23
  • unresolved5
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 01016162-a430-49d0-a8cc-0666398c37ef · outbound

This paper cites Exploiting recurrent neural networks and leap motion controller for the recognition of sign language and semaphoric hand gestures,.

Multiscaled Multi-Head Attention-based Video Transformer Network for Hand Gesture Recognition Exploiting recurrent neural networks and leap motion controller for the recognition of sign language and semaphoric hand gestures,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:44:42.898970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:44:42.338049Z digest=sha256:c4c47bbd451f962c827e403c71e9701e75e3a1ee6334569ef7e75eb2d6cfafcf

Observation 8ea96a26-a8af-4e17-81f0-99f4260994e0 · outbound

This paper cites Attention in convolutional LSTM for gesture recognition,.

Multiscaled Multi-Head Attention-based Video Transformer Network for Hand Gesture Recognition Attention in convolutional LSTM for gesture recognition,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:44:42.886256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:44:42.343199Z digest=sha256:eb8bc2a12fb55383659d29cb8e9a7a4d68a485a3e2e7ff57e9434c3e656312d7

Observation d7ec6c33-debc-4041-a47a-d7518ab44d3b · outbound

This paper cites Attention-based gated recurrent unit for gesture recognition,.

Multiscaled Multi-Head Attention-based Video Transformer Network for Hand Gesture Recognition Attention-based gated recurrent unit for gesture recognition,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:44:42.873000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:44:42.347751Z digest=sha256:c9e15b290411ed73dd3b9be8407c31f8332c7f9482c97782015197fe04f4a274

Observation 17dd7c9e-576d-4da0-9294-a81fa22911ba · outbound

This paper cites Online detection and classification of dynamic hand gestures with recurrent 3D convolutional neural network,.

Multiscaled Multi-Head Attention-based Video Transformer Network for Hand Gesture Recognition Online detection and classification of dynamic hand gestures with recurrent 3D convolutional neural network,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:44:42.860177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:44:42.352633Z digest=sha256:8d75d02209164b95a042c0e86ea882927181e93c632d4176d635360c5e681e53

Observation 22bf9b31-3bf1-4c4d-babb-d9c810005100 · outbound

This paper cites Attention is all you need,.

Multiscaled Multi-Head Attention-based Video Transformer Network for Hand Gesture Recognition Attention is all you need,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:44:42.848104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:44:42.357398Z digest=sha256:9ccf5505422dbeccc748b3dbd2719e11ebf6eaab537b1579b413457dd6521a9e

Observation bfbb28b8-bd59-4ae7-8b28-33d1cd3b9961 · outbound

This paper cites An image is worth 16 ×16 words: Transformers for image recognition at scale,.

Multiscaled Multi-Head Attention-based Video Transformer Network for Hand Gesture Recognition An image is worth 16 ×16 words: Transformers for image recognition at scale,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:44:42.835575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:44:42.362231Z digest=sha256:fc366fcbb1486fe6c2a8464e471bfa4ef4bd2bb986e17c395ede2d609f2f0ac4

Observation 7164164d-1986-427c-98b1-cab1b9334db1 · outbound

This paper cites Training data-efficient image transformers & distillation through attention.

Multiscaled Multi-Head Attention-based Video Transformer Network for Hand Gesture Recognition Training data-efficient image transformers & distillation through attention

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T22:44:42.366245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:44:42.366245Z digest=sha256:2f9b856b1e513319e91b747b378b6f14be617593c903851653b770b27818b954

Observation 422b1919-b928-445e-a91f-9bc6b623814f · outbound

This paper cites Video Transformer Network.

Multiscaled Multi-Head Attention-based Video Transformer Network for Hand Gesture Recognition Video Transformer Network

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-10T22:44:42.580679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:44:42.370070Z digest=sha256:83b1dafa54240a6d667f12cace7e737d9b39174de71671d3a6b0657300f61604

Observation cd7e5ee1-2f35-433d-ab50-f22063b446a8 · outbound

This paper cites CrossViT: Cross-Attention Multi-Scale Vision Transformer for Image Classification.

Multiscaled Multi-Head Attention-based Video Transformer Network for Hand Gesture Recognition CrossViT: Cross-Attention Multi-Scale Vision Transformer for Image Classification

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T22:44:42.374513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:44:42.374513Z digest=sha256:0b83ff954a0a62db98aa2bf0c7c7e93d8d5d77528cdcd127a1a4c3a126044f5d

Observation 3f4d0af7-febb-42fc-8b62-73b8ddfc7ac5 · outbound

This paper cites Multiscale Vision Transformers.

Multiscaled Multi-Head Attention-based Video Transformer Network for Hand Gesture Recognition Multiscale Vision Transformers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T22:44:42.379339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:44:42.379339Z digest=sha256:0aeb9d16c9ee6dcac93612a0ff95e76bc9b92d9412f54d101da873d4f9671787

Observation 30f1d953-db89-4e5d-841f-44dda78cb4fe · outbound

This paper cites MViTv2: Improved Multiscale Vision Transformers for Classification and Detection.

Multiscaled Multi-Head Attention-based Video Transformer Network for Hand Gesture Recognition MViTv2: Improved Multiscale Vision Transformers for Classification and Detection

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T22:44:42.384108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:44:42.384108Z digest=sha256:45c2f0de7a53b6d4ba95e95ad8c5bb52a0e571962edf4376754975bf42c36df0

Observation 68c2f107-f4f3-42a9-9fb1-678ec968b632 · outbound

This paper cites A Transformer-based network for dynamic hand gesture recognition,.

Multiscaled Multi-Head Attention-based Video Transformer Network for Hand Gesture Recognition A Transformer-based network for dynamic hand gesture recognition,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:44:42.822679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:44:42.388888Z digest=sha256:532d0918a804d7269967d29d0138e762845d5b8811abf663688bed102ca773f1

Observation 2cd495e4-8898-44c3-a2a9-c69fba8f1568 · outbound

This paper cites Incorporating relative position information in transformer-based sign language recognition and translation,.

Multiscaled Multi-Head Attention-based Video Transformer Network for Hand Gesture Recognition Incorporating relative position information in transformer-based sign language recognition and translation,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:44:42.811127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:44:42.393489Z digest=sha256:4ae38dbeb64e0ce4c3eee653e9db94b132d187080d06f30212315b3488efb4fc

Observation 2790838a-6db1-4ad2-8fb6-9e6bbf06c04b · outbound

This paper cites Searching multi-rate and multi-modal temporal enhanced networks for gesture recognition,.

Multiscaled Multi-Head Attention-based Video Transformer Network for Hand Gesture Recognition Searching multi-rate and multi-modal temporal enhanced networks for gesture recognition,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:44:42.799286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:44:42.397778Z digest=sha256:b40a7795c0127c88594f5ac25d6e82569033b6a4053348949b74f713b0b9a35f

Observation 68c51c78-318f-4110-a498-055d440cd0c6 · outbound

This paper cites TMMF: Temporal Multi-Modal Fusion for single-stage continuous gesture recog- nition,.

Multiscaled Multi-Head Attention-based Video Transformer Network for Hand Gesture Recognition TMMF: Temporal Multi-Modal Fusion for single-stage continuous gesture recog- nition,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:44:42.787116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:44:42.403181Z digest=sha256:4e34b85438e12fd8bf30a1e5eddc3bf680597fd68fbcb2dcd294c30a5451f502

Observation 504e1112-2760-4982-ac4f-02749289d282 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

Multiscaled Multi-Head Attention-based Video Transformer Network for Hand Gesture Recognition Imagenet: A large-scale hierarchical image database

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:44:42.774288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:44:42.408169Z digest=sha256:9689216af758c6a2198ae4a268c2f9c2a215a07a955a6d0c57bd0f89e0c53684

Observation 8366ec62-9116-4e03-88a6-6ea0424ff9a6 · outbound

This paper cites Two-Stream Convolutional Networks for Action Recognition in Videos.

Multiscaled Multi-Head Attention-based Video Transformer Network for Hand Gesture Recognition Two-Stream Convolutional Networks for Action Recognition in Videos

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T22:44:42.412647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:44:42.412647Z digest=sha256:036bf6edad69e93c836809b387799f1237b1657c19f5dd9a12190613076d0b59

Observation 733ce869-316a-48c3-9fb7-6d46535b37ee · outbound

This paper cites A robust and efficient video representation for action recognition,.

Multiscaled Multi-Head Attention-based Video Transformer Network for Hand Gesture Recognition A robust and efficient video representation for action recognition,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:44:42.761668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:44:42.417345Z digest=sha256:719634b0f806337c43455e1e218483ff3a5608846b7dae96774e2c0f7dfad36c

Observation 24fd4795-6661-4413-b6f8-35b5d5c504b8 · outbound

This paper cites Res3atn-deep 3D residual attention network for hand gesture recognition in videos,.

Multiscaled Multi-Head Attention-based Video Transformer Network for Hand Gesture Recognition Res3atn-deep 3D residual attention network for hand gesture recognition in videos,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:44:42.747742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:44:42.421408Z digest=sha256:12bc7a5cf0cf0dadc751d1ed9497f56a33312643dac84d0db1add3735c148512

Observation 88a65e5d-42fb-488a-a23c-6af0e4a9156f · outbound

This paper cites Learning spatiotemporal features with 3d convolutional networks,.

Multiscaled Multi-Head Attention-based Video Transformer Network for Hand Gesture Recognition Learning spatiotemporal features with 3d convolutional networks,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:44:42.734802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:44:42.425533Z digest=sha256:e9dad27d52321b152ae6b1d93c2ca55083111aed231c41c93da987a02f17fe23

Observation be70412e-52b7-462c-969d-c291b3adf076 · outbound

This paper cites Multi-task and multi-modal learning for RGB dynamic gesture recognition,.

Multiscaled Multi-Head Attention-based Video Transformer Network for Hand Gesture Recognition Multi-task and multi-modal learning for RGB dynamic gesture recognition,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:44:42.721546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:44:42.430604Z digest=sha256:ae5a9cacbb67bd43d914f5a8e6569dab2d77c4361b7ba574ece09199d3ca5287

Observation 0bed9be7-8c61-4583-b2b8-247cc5b08c87 · outbound

This paper cites Making convolutional networks recurrent for visual sequence learning,.

Multiscaled Multi-Head Attention-based Video Transformer Network for Hand Gesture Recognition Making convolutional networks recurrent for visual sequence learning,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:44:42.707408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:44:42.435169Z digest=sha256:a2aecd1e0d7e089f236270e1eae7f3aaf7b8bf743166e577f6a69c692cac4ece

Observation 8304e128-bf4a-4640-8d07-8373307a7945 · outbound

This paper cites Quo vadis, action recognition? A new model and the kinetics dataset,.

Multiscaled Multi-Head Attention-based Video Transformer Network for Hand Gesture Recognition Quo vadis, action recognition? A new model and the kinetics dataset,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:44:42.693721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:44:42.439168Z digest=sha256:d93cfa078fa4260d3faa901ba1b8cdb1505d8a96402432b5586121e8bcda3cef

Observation 0bc6a973-ed18-4152-88e0-7b23947773c7 · outbound

This paper cites Real-time hand ges- ture detection and classification using convolutional neural networks,.

Multiscaled Multi-Head Attention-based Video Transformer Network for Hand Gesture Recognition Real-time hand ges- ture detection and classification using convolutional neural networks,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:44:42.680593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:44:42.444421Z digest=sha256:c784854c26069c0b391024aef828d5b58ad826cdea34ed834778abc51a47bd7f

Observation f502ad87-70f5-4758-9f93-0e6ab8316c40 · outbound

This paper cites Improving the perfor- mance of unimodal dynamic hand-gesture recognition with multimodal training,.

Multiscaled Multi-Head Attention-based Video Transformer Network for Hand Gesture Recognition Improving the perfor- mance of unimodal dynamic hand-gesture recognition with multimodal training,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:44:42.667176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:44:42.448724Z digest=sha256:e8e222ef104c4f11bd0ab6dffa18d7c56f6794888a30a182dfc02f7bffa9451f

Observation 62664a0e-456d-4296-b70d-c4380f177282 · outbound

This paper cites Super normal vector for activity recognition using depth sequences,.

Multiscaled Multi-Head Attention-based Video Transformer Network for Hand Gesture Recognition Super normal vector for activity recognition using depth sequences,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:44:42.654082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:44:42.452677Z digest=sha256:85b448e42b6bfd902a6b8b04ff3818da9eed65f16358aaa301fd5d086eaeb34c

Observation 5a8e84e2-f1d3-4ed4-b213-bd0204b02d81 · outbound

This paper cites Motion fused frames: Data level fusion strategy for hand gesture recognition,.

Multiscaled Multi-Head Attention-based Video Transformer Network for Hand Gesture Recognition Motion fused frames: Data level fusion strategy for hand gesture recognition,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:44:42.632428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:44:42.456529Z digest=sha256:fd0cbc11bcdf6e627f2aa0c024c6a52ad8a003c6f5625d04c0bfe7976b9d815f

Observation 21cce0f5-8e5b-4067-af29-96b6fb0330bd · outbound

This paper cites Dynamic hand gesture recognition based on short-term sampling neural networks,.

Multiscaled Multi-Head Attention-based Video Transformer Network for Hand Gesture Recognition Dynamic hand gesture recognition based on short-term sampling neural networks,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:44:42.619923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:44:42.460704Z digest=sha256:76b79236689998f6217646f67a70194f062c8a175d9beab516a828a78fef4135

Observation 6b435dcb-e792-435c-828e-c5a8e9bb572e · outbound

This paper cites Hand gestures for the human-car interaction: The Briareo dataset,.

Multiscaled Multi-Head Attention-based Video Transformer Network for Hand Gesture Recognition Hand gestures for the human-car interaction: The Briareo dataset,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:44:42.606536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:44:42.464869Z digest=sha256:1cef8791b99ebf1b159973ef842fc895dab87b92d304235edb1a9b1b2c7adc6c

Observation c38f8d55-585f-4dc1-8916-f6425ccb2cf9 · outbound

This paper cites Multimodal hand gesture classification for the human–car interaction,.

Multiscaled Multi-Head Attention-based Video Transformer Network for Hand Gesture Recognition Multimodal hand gesture classification for the human–car interaction,

Reference 30

Resolution
verified exact
doi, observed 2026-08-10T22:44:42.505375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:44:42.468224Z digest=sha256:b23620c2627d1ba223cf9592c47d40d9b6797e5ad22ed733ef0d68ed4212dd1d

Pith citing papers

No inbound Pith citation observations are available.