Pith. sign in

Paper Citation Record · LEDGER

Cross-Modal Action Recognition in Egocentric Video Using Mamba: Integrating RGB and Hand Skeleton Streams via CLS Token Fusion Strategies

As of 22 August 2026, this Paper Citation Record lists 11 of 11 outbound references and 0 inbound Pith citation observations for arXiv:2605.24302.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.24302 v2

Coverage vector

measured 11 of 11 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-30T14:30:47.643204Z

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

11 of 11 outbound references displayed

  • verified exact0
  • verified fuzzy10
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3820128b-af82-4a91-ab75-fcfd589bdcf7 · outbound

This paper cites Video mamba suite: State space model as a versatile alternative for video understanding, 2024.

Cross-Modal Action Recognition in Egocentric Video Using Mamba: Integrating RGB and Hand Skeleton Streams via CLS Token Fusion Strategies Video mamba suite: State space model as a versatile alternative for video understanding, 2024

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T22:55:41.067679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T14:30:47.643204Z digest=sha256:6b6e132efdac82f524d700f48294dfde58828e4cba0f26cad783eff65f7aad48

Observation 6c58db89-72a4-4458-92a9-be26e8248bbe · outbound

This paper cites an unresolved cited work.

Cross-Modal Action Recognition in Egocentric Video Using Mamba: Integrating RGB and Hand Skeleton Streams via CLS Token Fusion Strategies Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-07-08T22:55:41.070125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T14:30:47.643204Z digest=sha256:764b04d01b542fd9334396da92a7fd6a545b3084f6a13bc6bd2faab8f47957b2

Observation dac8d7ac-b43a-485b-a216-f7a2a4dfcd8a · outbound

This paper cites Gigahands: A massive annotated dataset of bimanual hand activities, 2025.

Cross-Modal Action Recognition in Egocentric Video Using Mamba: Integrating RGB and Hand Skeleton Streams via CLS Token Fusion Strategies Gigahands: A massive annotated dataset of bimanual hand activities, 2025

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T22:55:41.071998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T14:30:47.643204Z digest=sha256:302de8cd1f14e01e7cae7dfba749b39b49bd69c74237f1f614f6e2b55d862f1e

Observation ad02f11e-00a3-40c3-8a8c-5aa44616ac2e · outbound

This paper cites The ”something something” video database for learning and evaluating visual common sense,.

Cross-Modal Action Recognition in Egocentric Video Using Mamba: Integrating RGB and Hand Skeleton Streams via CLS Token Fusion Strategies The ”something something” video database for learning and evaluating visual common sense,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T22:55:41.073993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T14:30:47.643204Z digest=sha256:800428bd7cdda93f6f15fbfe76436b99d049147c938b510d6bed4e69d63e7f46

Observation 0c52213a-4fa5-4310-b6c8-25c28f9146da · outbound

This paper cites Mamba: Linear-time sequence mod- eling with selective state spaces, 2023.

Cross-Modal Action Recognition in Egocentric Video Using Mamba: Integrating RGB and Hand Skeleton Streams via CLS Token Fusion Strategies Mamba: Linear-time sequence mod- eling with selective state spaces, 2023

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T22:55:41.062339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T14:30:47.643204Z digest=sha256:f4a66ef10288feedc4609bf955a9c5a026993b27e5662c53614e788e709f3ecc

Observation a1f1ead7-f9a0-4faa-832a-3719238eb5a5 · outbound

This paper cites H2o: Two hands manipulating objects for first person interaction recognition.

Cross-Modal Action Recognition in Egocentric Video Using Mamba: Integrating RGB and Hand Skeleton Streams via CLS Token Fusion Strategies H2o: Two hands manipulating objects for first person interaction recognition

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T22:55:41.064102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T14:30:47.643204Z digest=sha256:1f1e469df9f807d5acdaf7d5585e052863f8bc0fb758987471ba88198ec35afb

Observation cc3b3d89-1488-4e08-9fd5-68546ffaac27 · outbound

This paper cites Videomamba: State space model for efficient video understanding.

Cross-Modal Action Recognition in Egocentric Video Using Mamba: Integrating RGB and Hand Skeleton Streams via CLS Token Fusion Strategies Videomamba: State space model for efficient video understanding

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T22:55:41.065872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T14:30:47.643204Z digest=sha256:897149c7a3ff6cee8c303d5e29b9e3b291d2e025c5ebd1a82e9bfa97c40dea4e

Observation 9ef56a6e-9c70-4e40-80e2-166ad32da6c5 · outbound

This paper cites Gomez, Lukasz Kaiser, and Illia Polosukhin.

Cross-Modal Action Recognition in Egocentric Video Using Mamba: Integrating RGB and Hand Skeleton Streams via CLS Token Fusion Strategies Gomez, Lukasz Kaiser, and Illia Polosukhin

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T22:55:41.075769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T14:30:47.643204Z digest=sha256:82389afd210e3a86d8b6feea267b63eee0fc89cc55ae5f53fd5d588c73f60f1d

Observation 4949dae5-dc5d-486f-8065-1d741caacfd0 · outbound

This paper cites Actionmamba: Action spatial-temporal aggregation network based on mamba and gcn for skeleton-based action recognition.Electronics, 14 (18):3610, 2025.

Cross-Modal Action Recognition in Egocentric Video Using Mamba: Integrating RGB and Hand Skeleton Streams via CLS Token Fusion Strategies Actionmamba: Action spatial-temporal aggregation network based on mamba and gcn for skeleton-based action recognition.Electronics, 14 (18):3610, 2025

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T22:55:41.077721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T14:30:47.643204Z digest=sha256:dfcbe1f63483a6bc8667dfc1812e40c195a30eefd5c4913dddc1cadd011542df

Observation 549ab09b-7168-4154-9387-efb43551ef67 · outbound

This paper cites Oakink2: A dataset of bimanual hands-object manipulation in complex task completion, 2024.

Cross-Modal Action Recognition in Egocentric Video Using Mamba: Integrating RGB and Hand Skeleton Streams via CLS Token Fusion Strategies Oakink2: A dataset of bimanual hands-object manipulation in complex task completion, 2024

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T22:55:41.058199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T14:30:47.643204Z digest=sha256:de22f4f1b6b6a7f6d9305d63e4360bd5250445f1ff56d118bdd38f507acf61a0

Observation 6e25f894-d5a5-4a40-a42e-cec55d7a796e · outbound

This paper cites Vision mamba: Efficient visual representation learning with bidirectional state space model, 2024.

Cross-Modal Action Recognition in Egocentric Video Using Mamba: Integrating RGB and Hand Skeleton Streams via CLS Token Fusion Strategies Vision mamba: Efficient visual representation learning with bidirectional state space model, 2024

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T22:55:41.060451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T14:30:47.643204Z digest=sha256:3c871a17cb93ab41db2a0108e7d11958cfa016ef08d44cae365f0f6daab703aa

Pith citing papers

No inbound Pith citation observations are available.