Pith. sign in

Paper Citation Record · LEDGER

Cross-Modal Action Recognition in Egocentric Video Using Mamba: Integrating RGB and Hand Skeleton Streams via CLS Token Fusion Strategies

As of 19 August 2026, this Paper Citation Record lists 11 of 11 outbound references and 0 inbound Pith citation observations for arXiv:2605.24302.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.24302 v2

Coverage vector

measured 11 of 11 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-30T14:30:47.643204Z

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

11 of 11 outbound references displayed

  • verified exact0
  • verified fuzzy10
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3820128b-af82-4a91-ab75-fcfd589bdcf7 · outbound

This paper cites Video mamba suite: State space model as a versatile alternative for video understanding, 2024.

Cross-Modal Action Recognition in Egocentric Video Using Mamba: Integrating RGB and Hand Skeleton Streams via CLS Token Fusion Strategies Video mamba suite: State space model as a versatile alternative for video understanding, 2024

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T22:55:41.067679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T14:30:47.643204Z digest=sha256:ea47ce2d6ddae6ac92d2d48f8a01a91685f3ff2a2c86ece9820d6ef957365c65

Observation 6c58db89-72a4-4458-92a9-be26e8248bbe · outbound

This paper cites an unresolved cited work.

Cross-Modal Action Recognition in Egocentric Video Using Mamba: Integrating RGB and Hand Skeleton Streams via CLS Token Fusion Strategies Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-07-08T22:55:41.070125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T14:30:47.643204Z digest=sha256:0a60991d409d5b8a698b0dd7240bbf2b50f7034ae974827b269ada9294db0baa

Observation dac8d7ac-b43a-485b-a216-f7a2a4dfcd8a · outbound

This paper cites Gigahands: A massive annotated dataset of bimanual hand activities, 2025.

Cross-Modal Action Recognition in Egocentric Video Using Mamba: Integrating RGB and Hand Skeleton Streams via CLS Token Fusion Strategies Gigahands: A massive annotated dataset of bimanual hand activities, 2025

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T22:55:41.071998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T14:30:47.643204Z digest=sha256:52dacb4f46ac6785bf4f94fed5ebf5c33286e950011cfe9e93a95f881c590b67

Observation ad02f11e-00a3-40c3-8a8c-5aa44616ac2e · outbound

This paper cites The ”something something” video database for learning and evaluating visual common sense,.

Cross-Modal Action Recognition in Egocentric Video Using Mamba: Integrating RGB and Hand Skeleton Streams via CLS Token Fusion Strategies The ”something something” video database for learning and evaluating visual common sense,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T22:55:41.073993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T14:30:47.643204Z digest=sha256:f042523b9989a5276763566edde9b5a26001c1ee09890910e750643d26a379c6

Observation 0c52213a-4fa5-4310-b6c8-25c28f9146da · outbound

This paper cites Mamba: Linear-time sequence mod- eling with selective state spaces, 2023.

Cross-Modal Action Recognition in Egocentric Video Using Mamba: Integrating RGB and Hand Skeleton Streams via CLS Token Fusion Strategies Mamba: Linear-time sequence mod- eling with selective state spaces, 2023

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T22:55:41.062339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T14:30:47.643204Z digest=sha256:8089882db6d0548353b84900d5553a67b7281c85419b107e0d1c51729441c9ec

Observation a1f1ead7-f9a0-4faa-832a-3719238eb5a5 · outbound

This paper cites H2o: Two hands manipulating objects for first person interaction recognition.

Cross-Modal Action Recognition in Egocentric Video Using Mamba: Integrating RGB and Hand Skeleton Streams via CLS Token Fusion Strategies H2o: Two hands manipulating objects for first person interaction recognition

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T22:55:41.064102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T14:30:47.643204Z digest=sha256:bbfe8c000f2449ee28f5d481c8d7d6d0809b9e0c40ae9e52f722eda22aadf1a8

Observation cc3b3d89-1488-4e08-9fd5-68546ffaac27 · outbound

This paper cites Videomamba: State space model for efficient video understanding.

Cross-Modal Action Recognition in Egocentric Video Using Mamba: Integrating RGB and Hand Skeleton Streams via CLS Token Fusion Strategies Videomamba: State space model for efficient video understanding

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T22:55:41.065872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T14:30:47.643204Z digest=sha256:163c9de62363a6f564257cd74d25c3d506a0c33e9b8b78389dbfb2c8d8409f82

Observation 9ef56a6e-9c70-4e40-80e2-166ad32da6c5 · outbound

This paper cites Gomez, Lukasz Kaiser, and Illia Polosukhin.

Cross-Modal Action Recognition in Egocentric Video Using Mamba: Integrating RGB and Hand Skeleton Streams via CLS Token Fusion Strategies Gomez, Lukasz Kaiser, and Illia Polosukhin

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T22:55:41.075769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T14:30:47.643204Z digest=sha256:9871afb1b1fdb1df61198e1f4e17762a15fd42c083aebd744b49a1ad85e19614

Observation 4949dae5-dc5d-486f-8065-1d741caacfd0 · outbound

This paper cites Actionmamba: Action spatial-temporal aggregation network based on mamba and gcn for skeleton-based action recognition.Electronics, 14 (18):3610, 2025.

Cross-Modal Action Recognition in Egocentric Video Using Mamba: Integrating RGB and Hand Skeleton Streams via CLS Token Fusion Strategies Actionmamba: Action spatial-temporal aggregation network based on mamba and gcn for skeleton-based action recognition.Electronics, 14 (18):3610, 2025

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T22:55:41.077721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T14:30:47.643204Z digest=sha256:383e0c7edae2ea4d6d7271f11b066b525ae0ba905b016c9545f20dd0566c56ca

Observation 549ab09b-7168-4154-9387-efb43551ef67 · outbound

This paper cites Oakink2: A dataset of bimanual hands-object manipulation in complex task completion, 2024.

Cross-Modal Action Recognition in Egocentric Video Using Mamba: Integrating RGB and Hand Skeleton Streams via CLS Token Fusion Strategies Oakink2: A dataset of bimanual hands-object manipulation in complex task completion, 2024

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T22:55:41.058199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T14:30:47.643204Z digest=sha256:11e2df89cc91989de838a933bbb7508de5c4f98d2d746a7bc2b88e66a26a7447

Observation 6e25f894-d5a5-4a40-a42e-cec55d7a796e · outbound

This paper cites Vision mamba: Efficient visual representation learning with bidirectional state space model, 2024.

Cross-Modal Action Recognition in Egocentric Video Using Mamba: Integrating RGB and Hand Skeleton Streams via CLS Token Fusion Strategies Vision mamba: Efficient visual representation learning with bidirectional state space model, 2024

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T22:55:41.060451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T14:30:47.643204Z digest=sha256:14bc03cde6c637fb71ec9881adf741abd289d0afb02a8ec596203b81a34f3b80

Pith citing papers

No inbound Pith citation observations are available.