Pith. sign in

Transformer-Based Approaches for Sensor-Based Human Activity Recognition: Opportunities and Challenges

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

Transformers have excelled in natural language processing and computer vision, paving their way to sensor-based Human Activity Recognition (HAR). Previous studies show that transformers outperform their counterparts exclusively when they harness abundant data or employ compute-intensive optimization algorithms. However, neither of these scenarios is viable in sensor-based HAR due to the scarcity of data in this field and the frequent need to perform training and inference on resource-constrained devices. Our extensive investigation into various implementations of transformer-based versus non-transformer-based HAR using wearable sensors, encompassing more than 500 experiments, corroborates these concerns. We observe that transformer-based solutions pose higher computational demands, consistently yield inferior performance, and experience significant performance degradation when quantized to accommodate resource-constrained devices. Additionally, transformers demonstrate lower robustness to adversarial attacks, posing a potential threat to user trust in HAR.

fields

cs.LG 1

years

2025 1

verdicts

CONDITIONAL 1

representative citing papers

Hierarchical Motion Captioning Utilizing External Text Data Source

cs.LG · 2025-09-01 · conditional · novelty 5.0

This paper introduces a hierarchical motion captioning system that generates low-level descriptions with an LLM and retrieves high-level captions from a database, reporting large gains over prior methods on three datasets.

citing papers explorer

Showing 1 of 1 citing paper.

  • Hierarchical Motion Captioning Utilizing External Text Data Source cs.LG · 2025-09-01 · conditional · none · ref 7 · internal anchor

    This paper introduces a hierarchical motion captioning system that generates low-level descriptions with an LLM and retrieves high-level captions from a database, reporting large gains over prior methods on three datasets.