Pith. sign in

REVIEW 3 cited by

Exploiting Ensemble Learning for Cross-View Isolated Sign Language Recognition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.02196 v1 pith:QDZYMTP5 submitted 2025-02-04 cs.CV cs.AI

classification cs.CVcs.AI
keywords languagerecognitionsigncross-viewislrensembleisolatedangles
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this paper, we present our solution to the Cross-View Isolated Sign Language Recognition (CV-ISLR) challenge held at WWW 2025. CV-ISLR addresses a critical issue in traditional Isolated Sign Language Recognition (ISLR), where existing datasets predominantly capture sign language videos from a frontal perspective, while real-world camera angles often vary. To accurately recognize sign language from different viewpoints, models must be capable of understanding gestures from multiple angles, making cross-view recognition challenging. To address this, we explore the advantages of ensemble learning, which enhances model robustness and generalization across diverse views. Our approach, built on a multi-dimensional Video Swin Transformer model, leverages this ensemble strategy to achieve competitive performance. Finally, our solution ranked 3rd in both the RGB-based ISLR and RGB-D-based ISLR tracks, demonstrating the effectiveness in handling the challenges of cross-view recognition. The code is available at: https://github.com/Jiafei127/CV_ISLR_WWW2025.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Task-Generalized Adaptive Cross-Domain Learning for Multimodal Image Fusion

    cs.CV 2025-08 conditional novelty 4.0 of 10

    AdaSFFuse combines a learnable wavelet transform and a spatial-frequency Mamba block to report state-of-the-art fusion scores on infrared-visible, multi-exposure, multi-focus, and medical image pairs.

  2. MM-Gesture: Towards Precise Micro-Gesture Recognition through Multimodal Fusion

    cs.CV 2025-07 conditional novelty 4.0 of 10

    Combining joint, limb, RGB, Taylor-video, optical-flow, and depth streams with two video backbones and a validation-tuned weighted ensemble reaches 73.213% top-1 accuracy on iMiGUE, the best MiGA challenge result to date.

  3. Online Micro-gesture Recognition Using Data Augmentation and Spatial-Temporal Attention

    cs.CV 2025-07 reject novelty 3.0 of 10

    The paper claims a first-place micro-gesture detection result from data augmentation and spatial-temporal attention, but its own table shows the winning F1 comes from the unmodified AdaTAD baseline, while the proposed...

Pith tools