Pith. sign in

REVIEW 1 cited by

Enhancing Inertial Hand based HAR through Joint Representation of Language, Pose and Synthetic IMUs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.01316 v2 pith:7SRYFUQF submitted 2024-06-03 cs.CV cs.AI

classification cs.CVcs.AI
keywords datasyntheticvideosactivitiesapproachfine-grainedinertialjoint
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Due to the scarcity of labeled sensor data in HAR, prior research has turned to video data to synthesize Inertial Measurement Units (IMU) data, capitalizing on its rich activity annotations. However, generating IMU data from videos presents challenges for HAR in real-world settings, attributed to the poor quality of synthetic IMU data and its limited efficacy in subtle, fine-grained motions. In this paper, we propose Multi$^3$Net, our novel multi-modal, multitask, and contrastive-based framework approach to address the issue of limited data. Our pretraining procedure uses videos from online repositories, aiming to learn joint representations of text, pose, and IMU simultaneously. By employing video data and contrastive learning, our method seeks to enhance wearable HAR performance, especially in recognizing subtle activities.Our experimental findings validate the effectiveness of our approach in improving HAR performance with IMU data. We demonstrate that models trained with synthetic IMU data generated from videos using our method surpass existing approaches in recognizing fine-grained activities.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SImpHAR: Advancing impedance-based human activity recognition using 3D simulation and text-to-motion models

    cs.CV 2025-07 conditional novelty 6.0 of 10

    SImpHAR simulates bio-impedance signals from 3D motion and text, then uses contrastive pretraining and fine-tuning to improve impedance-based human activity recognition on two of three datasets.

Pith tools