Pith. sign in

REVIEW 4 major objections 7 minor 85 references

TxP: Reciprocal Generation of Ground Pressure Dynamics and Activity Descriptions for Improving Human Activity Recognition

T0 review · 4 major / 7 minor · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read The paper shows that a two-way text-pressure model, trained on 81,100 text-pressure pairs, generates pressure maps from text and classifies them via language, lifting activity recognition by up to 12.4% macro F1.

desk verdict Genuinely new pressure-to-text pipeline and a large synthetic corpus, but the headline accuracy gain is a post-hoc best cell and the transfer test is partly in-distribution. read the letter →

arxiv 2505.02052 v1 pith:LVUSCOXM submitted 2025-05-04 cs.AI cs.CV

classification cs.AIcs.CV
keywords humanactivityrecognitionpressuresensorsgrounddynamicstext-to-pressuregenerationpressure-to-textvectorquantizationlargelanguagemodelssyntheticdataaugmentation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tries to establish that natural language can act as the interchange medium for pressure-sensor data: a model trained on synthetic text-pressure pairs can turn activity descriptions into realistic ground-pressure dynamics, and turn real pressure sequences back into activity descriptions and class labels. If true, this would let researchers generate training data for pressure-based activity recognition from text alone, bypassing slow and costly real sensor collection, and would replace fixed activity labels with open-ended, interpretable descriptions. The headline result is that the two directions combined, with a 50/50 mix of real and synthetic data, raise macro F1 on real mattress datasets by up to 12.4% over prior state-of-the-art classifiers. The paper is explicit that the gain is conditional: on sleeping-posture data, augmentation actually lowers performance because the synthetic corpus lacks lying-down motions.

What carries the argument

The load-bearing mechanism is the pressure codebook: PressureRQVAE, a residual vector-quantized autoencoder that slices dynamic pressure maps into two-second windows and quantizes each into a small set of discrete tokens including an end token, so that continuous sensor streams become a finite vocabulary like words. This discrete tokenization is what lets a CLIP text encoder condition an autoregressive transformer to generate pressure sequences from text, and what lets a frozen LLaMA model, via a learned projection head, generate descriptions from pressure. The second pillar is the PressLang corpus itself—81,100 text-pressure pairs simulated by converting 3D SMPL poses into pressure maps with body-shape variations—since both generators are trained end-to-end on this synthetic data. The machinery fails for motion classes absent from the corpus, as the sleeping-posture results show.

What would settle it

Run the same Text2Pressure augmentation on a real mattress dataset whose activity classes all appear in the PressLang motion corpus; if the 50/50-mix macro F1 does not exceed the real-only baseline by roughly the reported margin, or turns negative, the claim that simulated text-to-pressure data transfers to real recordings would be falsified. A cheaper, direct check: generate Text2Pressure maps for the 11 PmatData sleeping classes and measure per-pixel overlap with real PmatData maps—the paper's own numbers already indicate the synthetic maps are poor proxies.

Watch

Extended reading notes

Core claim

The central claim is that pressure dynamics can be tokenized into a discrete codebook and then aligned with frozen text models in both directions, making pressure a language-compatible modality. TxP's PressureRQVAE compresses variable-length pressure maps into residual-quantized codebook tokens; a CLIP-conditioned autoregressive transformer (Text2Pressure) predicts those tokens from activity descriptions, and a projection head feeding the same token sequence into a frozen 13-billion-parameter LLM (Pressure2Text) generates atomic-motion descriptions that a prompt-engineered classifier maps to activity labels. Trained on PressLang—81,100 motions simulated from 3D body poses with five body-shape variations per motion—the system beats the previous text-to-pressure generator on generation fidelity and, at the best mixing ratio, raises macro F1 by 11.6% through augmentation alone, by 8.2% through grounded classification alone, and by up to 12.4% when both are combined on the TMD daily-activity dataset. The authors present this as an advance for pressure-based HAR, with the caveat that recognition collapses on sleeping-posture data because such motions are absent from the training corpus.

Load-bearing premise

The load-bearing premise is that pressure maps simulated by PresSim from 3D body poses faithfully reproduce real ground-pressure dynamics for the target activities and sensor layout, so that classifiers trained with text-generated synthetic data transfer to real recordings; the paper itself shows this premise fails for sleeping postures, where augmentation drops macro F1 from 0.765 to 0.544.

Editorial extensions

If this is right

  • Text-based augmentation can cut the cost of pressure-HAR dataset collection: with a 50/50 real-to-synthetic mix, the paper reports up to 12.4% macro-F1 gains over prior state-of-the-art on a daily-activity dataset.
  • The bridge between synthetic and real pressure data holds only for activities covered by the simulation corpus; expanding to new activity families requires adding their motion-text pairs to the training data, not just fine-tuning the classifier.
  • Pressure2Text turns classification into a language task, so a single model can output free-form atomic-action descriptions rather than a fixed label set, enabling open-vocabulary recognition and human-readable explanations of a pressure sequence.
  • The best augmentation ratio is roughly balanced (50% real, 50% synthetic); too much synthetic data degrades performance because the model drifts from real sensor statistics.
  • LLM-based classification is computationally heavier than a 3D-CNN baseline, but the paper reports it can still run at interactive rates on a consumer GPU.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Because the synthetic corpus defines the vocabulary of the codebook, the failure on sleeping postures implies that the bottleneck is corpus coverage rather than the quantization or alignment architecture; a testable extension is to add recumbent motion-text pairs to PressLang and watch PmatData F1 recover.
  • Since generation is tied to one fixed sensor geometry (an 80x28 SensingTex mat), reported gains may be partly geometry-specific; the paper's own adaptation recipe assumes the target array is a crop or resample of the original, so larger or differently shaped mats are an untested edge case.
  • The same token-plus-LLM recipe could extend to other spatially rich modalities, but the PID4TC insole result warns that the transfer is not automatic because Pressure2Text was trained on mattress data with environment-anchored positions.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 7 minor

Summary. The paper proposes TxP, a bidirectional Text×Pressure framework for pressure-based HAR. Text2Pressure maps activity descriptions to dynamic pressure sequences via a PressureRQVAE tokenizer and a CLIP-conditioned autoregressive transformer, trained on a new synthetic corpus (PressLang) built by simulating Motion-X SMPL poses with PresSim and re-annotating them with LLaMA. Pressure2Text maps pressure token sequences to activity descriptions through a trainable projection head and a frozen LLaMA 2 13B Chat model, enabling an LLM-grounded classifier. The authors evaluate pressure-map reconstruction, text generation, and downstream HAR on PresSim, TMD, PmatData, MeX, and PID4TC, and report that Text2Pressure augmentation plus Pressure2Text classification improves macro F1 by up to 12.4% over state-of-the-art. The paper includes extensive ablations over codebook size, window size, quantization dropout, real/synthetic ratios, and LLM backbones, and it candidly documents failures on PmatData and on insole data.

Significance. If the central claim held as stated, the contribution would be significant: a text-conditioned synthetic data generator for pressure maps, a language-grounded classifier, and an 81.1K-pair corpus would give the community a new tool for addressing pressure-data scarcity and for interpretable HAR. The architecture is reasonable and the component-level evaluations are thorough, including the useful ablation of LLM backbones and the ethically motivated removal of gender attributes in Pressure2Text (footnote 1). The significance is weakened, however, by two factors that the manuscript itself partly acknowledges: the headline 12.4% is the best cell of a per-dataset ratio sweep with no validation-based selection rule, and the favorable results concentrate on datasets whose sensor geometry is the one PresSim was designed to simulate, while the genuinely different PmatData degrades and the insole dataset gains only within noise. The paper is therefore a promising systems contribution whose broad claims need to be recalibrated and re-evaluated.

major comments (4)
  1. [§4.5, Tables 3 and 5] The conclusion that the 50/50 real/synthetic ratio 'achieved the highest performance across datasets' is contradicted by the paper's own data: for PmatData in Table 3, macro F1 falls from 0.765 at 100% real to 0.544 at 50/50, and every augmentation ratio including synthetic data is worse than real-only data. The headline 12.4% gain is the best cell of an eight-ratio sweep (TMD, Pressure2Text, 50/50), not a configuration selected by a described validation procedure. Please provide a validation-based rule for choosing the mix ratio, or explicitly present the sweep as exploratory and avoid selecting the maximum on the test sets.
  2. [§5.3, Table 8] The transfer claim is load-bearing and currently supported only for datasets whose sensing hardware matches the PresSim simulation target. PresSim, TMD, and MeX use SensingTex mats (80×28, or crops/resamples of that grid), while PmatData (Vista Medical 32×64) degrades under augmentation and PID4TC (insole) improves only from 0.731 to 0.744, within the reported ±0.035 standard deviation. Section 5.3 attributes the PmatData collapse to the absence of lying-down motions in PressLang, which means the corpus is not yet sufficient for the paper's broad 'advancing pressure-based HAR' claim. Please scope the central claim to the simulated mattress geometry or add evidence on additional sensor configurations.
  3. [§4.5, Table 5] The comparison against 'SOTA' is not controlled: the SOTA rows are numbers taken from the original dataset papers with different classifiers and evaluation protocols, while the TxP rows use the authors' 3D CNN baseline or the Pressure2Text classifier. Since the abstract's 12.4% is framed as a gain over state-of-the-art, the comparison should be re-run under a common protocol, or the claim should be limited to gains over the paper's own baseline. The differences in several cells are also within one standard deviation (e.g., PresSim 0.912±0.024 vs PressureTransferNet 0.911±0.015), so significance testing or confidence intervals should accompany the claim.
  4. [§4.5, §5.4] The paper's own limitation statements in §5.4 describe unaddressed LLM bias in re-annotating descriptions and in translating hard labels to activity descriptions, and the proposed mitigation is manual checking rather than a implemented filter. This is an honest disclosure, but it should be reflected in the conclusions: the synthetic corpus is not yet a verified resource for activities outside a narrow set, and the reported gains may partly reflect bias shared between PressLang and the evaluation datasets' label vocabularies.
minor comments (7)
  1. [§3, first paragraph] The paragraph beginning 'This section presents the comprehensive approach...' is duplicated verbatim; please remove one instance.
  2. [Appendix B] The appendix states that the residual quantization content 'originates initially from RQ-VAE [20]', but reference [20] is MoMask; the original RQ-VAE citation is [34] (Lee et al.).
  3. [§4.5 and Table 3] The text refers to 'TDM [54]' in several places while the table and reference list use 'TMD'; please unify the notation.
  4. [§4.5] 'PreSim' should be 'PresSim' in the sentence 'it exceeded real-only data on PreSim, TMD dataset, and MeX...'.
  5. [§2 and Table 2] In §4.3, 'different matrices' should read 'different metrics'.
  6. [§3.2, Figure 2] The label 'LRQVAE' in Figure 2 appears to be a typo for 'RQVAE' or 'PressureRQVAE'.
  7. [Abstract and §1] The abstract says '81,100 text-pressure pairs' while the contributions say '81.1K unique motions and 78 million individual pressure frames'; please clarify whether 81.1K counts motions, text-pressure pairs, or both.

Circularity Check

1 steps flagged · score 4.0 of 10

No definitional circularity; main caveat is that the headline SOTA gain is partly in-distribution, resting on the authors' own PresSim simulator, own datasets, and self-cited baselines.

  1. self citation load bearing [Section 3.1 (Pressure Dynamics Simulation) and Section 4.2 (Evaluation Datasets); Table 5 note]
    "We employ the PresSim framework [55], which combines physics simulations with neural networks, to generate pressure profiles from SMPL pose sequences designed explicitly for a SensingTex pressure-sensing mattress with 80×28 sensor array. ... It has already been validated in PressureTransferNet [53] that can generate very realistic Pressure maps that can, in turn, improve HAR by 4.8%. PresSim dataset [55] ... captured by SensingTex pressure mattress with 80×28 sensor array. TMD dataset [54] ... also captured by SensingTex pressure mattress with 80×28 sensor array."

    The synthetic PressLang data are generated with PresSim [55], a simulator built for the same 80×28 SensingTex mattress on which two of the four evaluation datasets (PresSim [55] and TMD [54]) were recorded by the same group. The realism premise of that simulator is itself supported by citing the authors' own PressureTransferNet [53], and the SOTA numbers used for the headline comparison in Table 5 come from the authors' own original papers. The 12.4% gain is therefore measured inside the sensor geometry the generator was designed for, so the 'real-world validation' is partly in-distribution rather than an independent transfer test.

full rationale

The paper does not exhibit a self-definitional reduction: Text2Pressure and Pressure2Text are trained on PressLang synthetic data and tested on real recordings, and the loss functions (Eqs. 4-7) do not encode the test labels. No ansatz or uniqueness theorem is smuggled in via self-citation. The main circularity concern is a load-bearing self-citation chain: the synthetic data source (PresSim [55]), the validation of that source (PressureTransferNet [53]), two evaluation datasets (PresSim [55], TMD [54]), and the SOTA values for those datasets all come from the same authors, and the simulator and those datasets share the same SensingTex 80×28 sensor geometry. This makes the headline 12.4% claim an in-distribution result rather than evidence of general transfer. However, the paper is not wholly self-referential: MeX is an external dataset and shows a 10-point gain, PmatData and PID4TC provide independent negative evidence, and the central architecture is compared against external VideoLLaMA baselines. The reported ratio sweep (Section 4.5) is a statistical caveat about selecting the best cell on test data, but it is not a derivation-level circularity. Weighing the self-citation load with the independent external content gives a score of 4.

Assumptions & free parameters 6 free parameters · 6 assumptions · 0 invented entities

The central claim rests on a chain of learned or simulated components (Motion-X, PresSim, LLaMA re-annotation, CLIP, and the RQVAE codebook) and on several numeric design choices (window size, codebook size, mix ratio) that are validated only on the authors' own benchmarks. No external physical constants or independent measurements anchor the pipeline.

free parameters (6)
  • Pressure tokenization window size = 2 s (selected from 1 s, 2 s, 4 s)
    Table 1 sweeps window sizes and picks 2 s as optimal; the downstream generation and augmentation depend on this choice.
  • Codebook size = 1024 (selected from 128, 256, 512, 1024)
    Table 1 shows larger codebooks improve FID and R-Precision; 1024 is used for the downstream tasks.
  • Real/synthetic mix ratio = 50%/50% (selected as the best of eight ratios)
    Tables 3 and 5 report the improvement claim using the ratio that peaks on the same real datasets where accuracy is reported; PmatData peaks at no synthetic data.
  • Quantization dropout probability q = not specified
    Section 3.2 states q in [0,1] and says dropout improves results, but the exact value used for the experiments is never given.
  • Random masking fraction tau = not specified
    Section 3.3 replaces a percentage tau of ground-truth indices with random indices during training; the value is not reported.
  • SMPL body-shape categories = 5 categories after merging light male and light female
    Appendix D defines six body categories, then merges two into one; each seed motion is simulated with five random shape variations, affecting the synthetic data distribution.
assumptions (6)
  • domain assumption Motion-X SMPL pose sequences correspond to the activities named in their text annotations.
    Section 3.1 uses Motion-X motions as seeds for PressLang; if poses and text are mismatched, the paired text-pressure training data is corrupted.
  • domain assumption PresSim simulation of pressure from SMPL poses faithfully represents real pressure-mattress readings for the target activities.
    Section 3.1 generates all synthetic pressure data with PresSim, and the paper later states TxP carries over PresSim's imperfections (Sections 5.3 and 6).
  • domain assumption LLaMA 2 re-annotation preserves the ground-body interaction content of Motion-X descriptions while removing irrelevant detail.
    Section 3.1 uses LLaMA 2 prompt engineering to rewrite annotations; if rewrites drop or distort contact information, Text2Pressure learns a wrong text-pressure mapping. Section 5.4 acknowledges LLM-induced biases.
  • domain assumption CLIP text embeddings provide a sufficient conditioning signal for generating pressure token sequences.
    Section 3.3 uses a frozen CLIP text encoder and no ablation tests alternative text encoders.
  • domain assumption FID computed on pressure maps is a valid proxy for the downstream utility of generated pressure data.
    Section 4.3 uses FID to rank reconstruction and generation quality, but downstream HAR gains are the central claim and FID does not predict the PmatData failure.
  • domain assumption The 3D CNN baseline from PressureTransferNet is representative of state-of-the-art pressure-HAR classifiers.
    Section 4.5 uses this baseline for augmentation comparisons, and the strongest claims are relative to it and to SOTA values from unspecified original papers.

how reviews work

0 comments
Cite this review

Pith. "Pith review of TxP: Reciprocal Generation of Ground Pressure Dynamics and Activity Descriptions for Improving Human Activity Recognition." pith.science (2026). https://pith.science/paper/LVUSCOXM

@misc{pith2026250502052,
  author       = {Pith},
  title        = {Pith review of: TxP: Reciprocal Generation of Ground Pressure Dynamics and Activity Descriptions for Improving Human Activity Recognition},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/LVUSCOXM}},
  note         = {Machine review of arXiv:2505.02052}
}
abstract

Sensor-based human activity recognition (HAR) has predominantly focused on Inertial Measurement Units and vision data, often overlooking the capabilities unique to pressure sensors, which capture subtle body dynamics and shifts in the center of mass. Despite their potential for postural and balance-based activities, pressure sensors remain underutilized in the HAR domain due to limited datasets. To bridge this gap, we propose to exploit generative foundation models with pressure-specific HAR techniques. Specifically, we present a bidirectional Text$\times$Pressure model that uses generative foundation models to interpret pressure data as natural language. TxP accomplishes two tasks: (1) Text2Pressure, converting activity text descriptions into pressure sequences, and (2) Pressure2Text, generating activity descriptions and classifications from dynamic pressure maps. Leveraging pre-trained models like CLIP and LLaMA 2 13B Chat, TxP is trained on our synthetic PressLang dataset, containing over 81,100 text-pressure pairs. Validated on real-world data for activities such as yoga and daily tasks, TxP provides novel approaches to data augmentation and classification grounded in atomic actions. This consequently improved HAR performance by up to 12.4\% in macro F1 score compared to the state-of-the-art, advancing pressure-based HAR with broader applications and deeper insights into human movement.

Figures

Figures reproduced from arXiv: 2505.02052 by the authors.

Figure 1
Figure 1. The TxP framework consists of two key components: (i) Text2Pressure, which generates dynamic pressure sequences [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. We used PressureRQVAE to quantize continuous dynamic length pressure time series data into discrete NLP-like [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗
Figure 3
Figure 3. Text2Pressure uses a frozen CLIP text encoder to convert text into vector embeddings, which are then given as input [PITH_FULL_IMAGE:figures/full_fig_p009_3.png] view at source ↗
Figures from the paper (7 more)
Figure 4
Figure 4. Figure 4: Pressure2Text utilizes a pre-trained PressureRQVAE encoder to transform pressure dynamics into NLP-like tokens. [PITH_FULL_IMAGE:figures/full_fig_p010_4.png]
Figure 5
Figure 5. Figure 5: We showcase how both components of TxP can be utilized to create a better HAR system where LLaMA 2, along with [PITH_FULL_IMAGE:figures/full_fig_p011_5.png]
Figure 6
Figure 6. Figure 6: Qualitative results showcasing generating the same activity with different variations (sitting) as well as different [PITH_FULL_IMAGE:figures/full_fig_p014_6.png]
Figure 7
Figure 7. Figure 7: Figure depicting some instances of incorrect data generation by Text2Pressure. [PITH_FULL_IMAGE:figures/full_fig_p020_7.png]
Figure 8
Figure 8. Figure 8: Figure depicting incorrect data generation by Text2Pressure for a particular activity class from PmatData, along with [PITH_FULL_IMAGE:figures/full_fig_p023_8.png]
Figure 9
Figure 9. Figure 9: Figure depicting variations generated by Text2Pressure for an identical prompt about the activity, but changing the [PITH_FULL_IMAGE:figures/full_fig_p030_9.png]
Figure 10
Figure 10. Figure 10: Figure depicting adaptation of TxP for different sensor configurations. [PITH_FULL_IMAGE:figures/full_fig_p031_10.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

85 extracted references · 58 canonical work pages

  1. [1]

    Mir Mushhood Afsar, Shizza Saqib, Mohammad Aladfaj, Mohammed Hamad Alatiyyah, Khaled Alnowaiser, Hanan Aljuaid, Ahmad Jalal, and Jeongmin Park. 2023. Body-worn sensors for recognizing physical sports activities in Exergaming via deep learning model. IEEE Access 11 (2023), 12460–12473

  2. [2]

    Moran Amit, Leanne Chukoskie, Andrew J Skalsky, Harinath Garudadri, and Tse Nga Ng. 2020. Flexible pressure sensors for objective assessment of motor disorders. Advanced Functional Materials 30, 20 (2020), 1905241

  3. [3]

    Shilpa Ankalaki. 2024. Simple to Complex, Single to Concurrent Sensor based Human Activity Recognition: Perception and Open Challenges. IEEE Access (2024)

  4. [4]

    Maxwell Fordjour Antwi-Afari, Heng Li, Waleed Umer, Yantao Yu, and Xuejiao Xing. 2020. Construction activity recognition and ergonomic risk assessment using a wearable insole pressure system. Journal of Construction Engineering and Management 146, 7 (2020), 04020077

  5. [5]

    Hymalai Bello, Sungho Suh, Daniel Geißler, Lala Shakti Swarup Ray, Bo Zhou, and Paul Lukowicz. 2023. CaptAinGlove: Capacitive and inertial fusion-based glove for real-time on edge hand gesture recognition for drone control. In Adjunct Proceedings of the 2023 ACM International Joint Conference on Pervasive and Ubiquitous Computing & the 2023 ACM Internatio...

  6. [6]

    Geetanjali Bhola and Dinesh Kumar Vishwakarma. 2024. A review of vision-based indoor HAR: state-of-the-art, challenges, and future prospects. Multimedia Tools and Applications 83, 1 (2024), 1965–2005

  7. [7]

    Vytautas Bucinskas, Andrius Dzedzickis, Juste Rozene, Jurga Subaciute-Zemaitiene, Igoris Satkauskas, Valentinas Uvarovas, Rokas Bobina, and Inga Morkvenaite-Vilkonciene. 2021. Wearable feet pressure sensor for human gait and falling diagnosis. Sensors 21, 15 (2021), 5240

  8. [8]

    Zhongang Cai, Daxuan Ren, Ailing Zeng, Zhengyu Lin, Tao Yu, Wenjia Wang, Xiangyu Fan, Yang Gao, Yifan Yu, Liang Pan, et al. 2022. Humman: Multi-modal 4d human dataset for versatile sensing and modeling. In European Conference on Computer Vision . Springer, 557–577

Show all 85 references
  1. [9]

    Wenqiang Chen, Yexin Hu, Wei Song, Yingcheng Liu, Antonio Torralba, and Wojciech Matusik. 2024. CAvatar: Real-time Human Activity Mesh Reconstruction via Tactile Carpets. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies 7, 4 (2024), 1–24

  2. [10]

    Zihan Chen, Yaojia Qian, Yuxi Wang, and Yinfeng Fang. 2022. Deep convolutional generative adversarial network-based EMG data enhancement for hand motion classification. Frontiers in Bioengineering and Biotechnology 10 (2022), 909653

  3. [11]

    Jihoon Chung, Cheng-hsin Wuu, Hsuan-ru Yang, Yu-Wing Tai, and Chi-Keung Tang. 2021. Haa500: Human-centric atomic action dataset with curated videos. In Proceedings of the IEEE/CVF international conference on computer vision . 13465–13474

  4. [12]

    Henry M Clever, Patrick L Grady, Greg Turk, and Charles C Kemp. 2022. Bodypressure-inferring body pose and contact pressure from a depth image. IEEE Transactions on Pattern Analysis and Machine Intelligence 45, 1 (2022), 137–153

  5. [13]

    Giovanni Diraco, Gabriele Rescio, Pietro Siciliano, and Alessandro Leone. 2023. Review on human action recognition in smart living: Sensing technology, multimodality, real-time processing, interoperability, and resource-constrained processing. Sensors 23, 11 (2023), 5281

  6. [14]

    Danny Driess, Fei Xia, Mehdi SM Sajjadi, Corey Lynch, Aakanksha Chowdhery, Brian Ichter, Ayzaan Wahid, Jonathan Tompson, Quan Vuong, Tianhe Yu, et al. 2023. Palm-e: An embodied multimodal language model. arXiv preprint arXiv:2303.03378 (2023)

  7. [15]

    Luigi D’Arco, Haiying Wang, and Huiru Zheng. 2022. Assessing impact of sensors and feature selection in smart-insole-based human activity recognition. Methods and Protocols 5, 3 (2022), 45

  8. [16]

    Vitor Fortes Rey, Lala Shakti Swarup Ray, Qingxin Xia, Kaishun Wu, and Paul Lukowicz. 2024. Enhancing Inertial Hand based HAR through Joint Representation of Language, Pose and Synthetic IMUs. In Proceedings of the 2024 ACM International Symposium on Wearable Computers. 25–31

  9. [17]

    Rohit Girdhar, Alaaeldin El-Nouby, Zhuang Liu, Mannat Singh, Kalyan Vasudev Alwala, Armand Joulin, and Ishan Misra. 2023. Imagebind: One embedding space to bind them all. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 15180–15190. Proc. ...

  10. [18]

    R Gnanavel, P Anjana, KS Nappinnai, and N Pavithra Sahari. 2016. Smart home system using a Wireless Sensor Network for elderly care. In 2016 Second International Conference on Science Technology Engineering and Management (ICONSTEM) . IEEE, 51–55

  11. [19]

    Albert Gu and Tri Dao. 2023. Mamba: Linear-time sequence modeling with selective state spaces. arXiv preprint arXiv:2312.00752 (2023)

  12. [20]

    Chuan Guo, Yuxuan Mu, Muhammad Gohar Javed, Sen Wang, and Li Cheng. 2024. Momask: Generative masked modeling of 3d human motions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 1900–1910

  13. [21]

    Neha Gupta, Suneet K Gupta, Rajesh K Pathak, Vanita Jain, Parisa Rashidi, and Jasjit S Suri. 2022. Human activity recognition in artificial intelligence framework: a narrative review. Artificial intelligence review 55, 6 (2022), 4755–4808

  14. [22]

    Foad Hamidi, Morgan Klaus Scheuerman, and Stacy M Branham. 2018. Gender recognition or gender reductionism? The social implications of embedded gender recognition systems. In Proceedings of the 2018 chi conference on human factors in computing systems . 1–13

  15. [23]

    Isaac Han, Seoyoung Lee, Sangyeon Park, Ecehan Akan, Yiyue Luo, and Kyung-Joong Kim. [n. d.]. Smart Insole: Predicting 3D human pose from foot pressure. In 2nd NeurIPS Workshop on Touch Processing: From Data to Knowledge

  16. [24]

    Jiaming Han, Kaixiong Gong, Yiyuan Zhang, Jiaqi Wang, Kaipeng Zhang, Dahua Lin, Yu Qiao, Peng Gao, and Xiangyu Yue. 2024. Onellm: One framework to align all modalities with language. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 26584–26595

  17. [25]

    Yuanfeng Han, Aadith Varadarajan, Taekyoung Kim, Gang Zheng, Kris Kitani, Aisling Kelliher, Thanassis Rikakis, and Yong-Lae Park

  18. [26]

    Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2021. Lora: Low-rank adaptation of large language models. arXiv preprint arXiv:2106.09685 (2021)

  19. [27]

    Yifan Hu. 2023. BSDGAN: Balancing Sensor Data Generative Adversarial Networks for Human Activity Recognition. In2023 International Joint Conference on Neural Networks (IJCNN) . IEEE, 1–8

  20. [28]

    Md Milon Islam, Sheikh Nooruddin, Fakhri Karray, and Ghulam Muhammad. 2023. Multi-level feature fusion for multimodal human activity recognition in Internet of Healthcare Things. Information Fusion 94 (2023), 17–31

  21. [29]

    Eun-tae Jeon and Hwi-young Cho. 2020. A novel method for gait analysis on center of pressure excursion based on a pressure-sensitive mat. International Journal of Environmental Research and Public Health 17, 21 (2020), 7845

  22. [30]

    Yongrok Jeong, Jimin Gu, Jaiyeul Byun, Junseong Ahn, Jaebum Byun, Kyuyoung Kim, Jaeho Park, Jiwoo Ko, Jun-ho Jeong, Morteza Amjadi, et al. 2021. Ultra-wide range pressure sensor based on a microstructured conductive nanocomposite for wearable workout monitoring. Advanced Healt...

  23. [31]

    Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al. 2023. Mistral 7B. arXiv preprint arXiv:2310.06825 (2023)

  24. [32]

    Os Keyes. 2018. The misgendering machines: Trans/HCI implications of automatic gender recognition. Proceedings of the ACM on human-computer interaction 2, CSCW (2018), 1–22

  25. [33]

    Hyeokhyen Kwon, Catherine Tong, Harish Haresamudram, Yan Gao, Gregory D Abowd, Nicholas D Lane, and Thomas Ploetz. 2020. Imutube: Automatic extraction of virtual on-body accelerometry from video for human activity recognition. Proceedings of the ACM on Interactive, Mobile, Wea...

  26. [34]

    Doyup Lee, Chiheon Kim, Saehoon Kim, Minsu Cho, and Wook-Shin Han. 2022. Autoregressive image generation using residual quantization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 11523–11532

  27. [35]

    Zikang Leng, Amitrajit Bhattacharjee, Hrudhai Rajasekhar, Lizhe Zhang, Elizabeth Bruda, Hyeokhyen Kwon, and Thomas Plötz. 2024. Imugpt 2.0: Language-based cross modality transfer for sensor-based human activity recognition. Proceedings of the ACM on Interactive, Mobile, Wearab...

  28. [36]

    Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. 2023. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In International conference on machine learning . PMLR, 19730–19742

  29. [37]

    Jing Lin, Ailing Zeng, Shunlin Lu, Yuanhao Cai, Ruimao Zhang, Haoqian Wang, and Lei Zhang. 2024. Motion-x: A large-scale 3d expressive whole-body human motion dataset. Advances in Neural Information Processing Systems 36 (2024)

  30. [38]

    Mengxi Liu, Vitor Fortes Rey, Yu Zhang, Lala Shakti Swarup Ray, Bo Zhou, and Paul Lukowicz. 2024. iMove: Exploring Bio-impedance Sensing for Fitness Activity Recognition. In 2024 IEEE International Conference on Pervasive Computing and Communications (PerCom) . IEEE, 194–205

  31. [39]

    Ruonan Liu, Yiying Liu, Yugui Cheng, He Liu, Simian Fu, Kaiming Jin, Deliang Li, Zhiwei Fu, Yixuan Han, Yanpeng Wang, et al. 2023. Aloe Inspired Special Structure Hydrogel Pressure Sensor for Real-Time Human-Computer Interaction and Muscle Rehabilitation System. Advanced Funct...

  32. [40]

    Matthew Loper, Naureen Mahmood, Javier Romero, Gerard Pons-Moll, and Michael J Black. 2023. SMPL: A skinned multi-person linear model. In Seminal Graphics Papers: Pushing the Boundaries, Volume 2 . 851–866

  33. [41]

    Naureen Mahmood, Nima Ghorbani, Nikolaus F Troje, Gerard Pons-Moll, and Michael J Black. 2019. AMASS: Archive of motion capture as surface shapes. In Proceedings of the IEEE/CVF international conference on computer vision . 5442–5451. Proc. ACM Interact. Mob. Wearable Ubiquito...

  34. [42]

    Andrea Marcante, Roberto Di Marco, Giovanni Gentile, Clelia Pellicano, Francesca Assogna, Francesco Ernesto Pontieri, Gianfranco Spalletta, Lucia Macchiusi, Dimitris Gatsios, Alexandros Giannakis, et al . 2020. Foot pressure wearable sensors for freezing of gait detection in P...

  35. [43]

    Keyu Meng, Xiao Xiao, Wenxin Wei, Guorui Chen, Ardo Nashalian, Sophia Shen, Xiao Xiao, and Jun Chen. 2022. Wearable pressure sensors for pulse wave monitoring. Advanced Materials 34, 21 (2022), 2109357

  36. [44]

    Seungwhan Moon, Andrea Madotto, Zhaojiang Lin, Aparajita Saraf, Amy Bearman, and Babak Damavandi. 2023. IMU2CLIP: Language- grounded Motion Sensor Translation with Multimodal Contrastive Learning. In Findings of the Association for Computational Linguistics: EMNLP 2023. 13246–13253

  37. [45]

    Jianyuan Ni, Hao Tang, Syed Tousiful Haque, Yan Yan, and Anne HH Ngu. 2024. A Survey on Multimodal Wearable Sensor-based Human Action Recognition. arXiv preprint arXiv:2404.15349 (2024)

  38. [46]

    Patricia O’Sullivan, Matteo Menolotto, Andrea Visentin, Brendan O’Flynn, and Dimitrios-Sokratis Komaris. 2024. Ai-based task classification with pressure insoles for occupational safety. IEEE Access 12 (2024), 21347–21357

  39. [47]

    Georgios Pavlakos, Vasileios Choutas, Nima Ghorbani, Timo Bolkart, Ahmed AA Osman, Dimitrios Tzionas, and Michael J Black. 2019. Expressive body capture: 3d hands, face, and body from a single image. In Proceedings of the IEEE/CVF conference on computer vision and pattern reco...

  40. [48]

    Mathis Petrovich, Michael J Black, and Gül Varol. 2023. TMR: Text-to-motion retrieval using contrastive 3D human motion synthesis. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 9488–9497

  41. [49]

    M Baran Pouyan, Javad Birjandtalab, Mehrdad Heydarzadeh, Mehrdad Nourani, and Sarah Ostadabbas. 2017. A pressure map dataset for posture and subject analytics. In 2017 IEEE EMBS international conference on biomedical & health informatics (BHI) . IEEE, 65–68

  42. [50]

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. In International conference on machine learni...

  43. [51]

    Gundala Jhansi Rani, Mohammad Farukh Hashmi, and Aditya Gupta. 2023. Surface electromyography and artificial intelligence for human activity recognition-A systematic review on methods, emerging trends applications, challenges, and future implementation. IEEE Access (2023)

  44. [52]

    Lala Shakti Swarup Ray, Daniel Geißler, Mengxi Liu, Bo Zhou, Sungho Suh, and Paul Lukowicz. 2024. ALS-HAR: Harnessing Wearable Ambient Light Sensors to Enhance IMU-based HAR. arXiv preprint arXiv:2408.09527 (2024)

  45. [53]

    Lala Shakti Swarup Ray, Vitor Fortes Rey, Bo Zhou, Sungho Suh, and Paul Lukowicz. 2023. PressureTransferNet: Human Attribute Guided Dynamic Ground Pressure Profile Transfer using 3D simulated Pressure Maps. arXiv preprint arXiv:2308.00538 (2023)

  46. [54]

    Lala Shakti Swarup Ray, Bo Zhou, Sungho Suh, Lars Krupp, Vitor Fortes Rey, and Paul Lukowicz. 2024. Text me the data: Generating Ground Pressure Sequence from Textual Descriptions for HAR. In 2024 IEEE International Conference on Pervasive Computing and Communications Workshop...

  47. [55]

    Lala Shakti Swarup Ray, Bo Zhou, Sungho Suh, and Paul Lukowicz. 2023. Pressim: An end-to-end framework for dynamic ground pressure profile generation from monocular videos using physics-based 3d simulation. In 2023 IEEE International Conference on Pervasive Computing and Commu...

  48. [56]

    Lala Shakti Swarup Ray, Bo Zhou, Sungho Suh, and Paul Lukowicz. 2025. OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models. In ICASSP 2025 - 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (...

  49. [57]

    Javier Romero, Dimitrios Tzionas, and Michael J Black. 2022. Embodied hands: Modeling and capturing hands and bodies together. arXiv preprint arXiv:2201.02610 (2022)

  50. [58]

    Ozell Sanders, Bin Wang, and Kimberly Kontson. 2024. Concurrent Validity Evidence for Pressure-Sensing Walkways Measuring Spatiotemporal Features of Gait: A Systematic Review and Meta-Analysis. Sensors 24, 14 (2024), 4537

  51. [59]

    Jesse Scott, Bharadwaj Ravichandran, Christopher Funk, Robert T Collins, and Yanxi Liu. 2020. From image to stability: Learning dynamics from human pose. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XXIII 16. Spring...

  52. [60]

    Davinder Pal Singh, Lala Shakti Swarup Ray, Bo Zhou, Sungho Suh, and Paul Lukowicz. 2024. A Novel Local-Global Feature Fusion Framework for Body-Weight Exercise Recognition with Pressure Mapping Sensors. In ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech an...

  53. [61]

    Ahmed Snoun, Tahani Bouchrika, and Olfa Jemai. 2023. Deep-learning-based human activity recognition for Alzheimer’s patients’ daily life activities assistance. Neural Computing and Applications 35, 2 (2023), 1777–1802

  54. [62]

    Vaibhav Soni, Himanshu Yadav, Vijay Bhaskar Semwal, Bholanath Roy, Dilip Kumar Choubey, and Dheeresh K Mallick. 2023. A novel smartphone-based human activity recognition using deep learning in health care. In Machine Learning, Image Processing, Network Security and Data Scienc...

  55. [63]

    Jordan Tabor, Talha Agcayazi, Aaron Fleming, Brendan Thompson, Ashish Kapoor, Ming Liu, Michael Y Lee, He Huang, Alper Bozkurt, and Tushar K Ghosh. 2021. Textile-based pressure sensors for monitoring prosthetic-socket interfaces. IEEE sensors journal 21, 7 (2021), Proc. ACM In...

  56. [64]

    Omid Taheri, Nima Ghorbani, Michael J Black, and Dimitrios Tzionas. 2020. GRAB: A dataset of whole-body human grasping of objects. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part IV 16 . Springer, 581–600

  57. [65]

    Gemma Team, Thomas Mesnard, Cassidy Hardin, Robert Dadashi, Surya Bhupatiraju, Shreya Pathak, Laurent Sifre, Morgane Rivière, Mi- hir Sanjay Kale, Juliette Love, et al. 2024. Gemma: Open models based on gemini research and technology.arXiv preprint arXiv:2403.08295 (2024)

  58. [66]

    Catherine Tong, Jinchen Ge, and Nicholas D Lane. 2021. Zero-shot learning for imu-based activity recognition using video embeddings. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies 5, 4 (2021), 1–23

  59. [67]

    Phuc Huu Truong, Sujeong You, Sang-Hoon Ji, and Gu-Min Jeong. 2019. Wearable system for daily activity recognition using inertial and pressure sensors of a smart band and smart shoes. International Journal of Computers Communications & Control 14, 6 (2019), 726–742

  60. [68]

    Shuhei Tsuchida, Satoru Fukayama, Masahiro Hamasaki, and Masataka Goto. 2019. AIST Dance Video Database: Multi-Genre, Multi- Dancer, and Multi-Camera Database for Dance Information Processing.. In ISMIR, Vol. 1. 6

  61. [69]

    Repuri Mohan Vamsi, Neha Adapa, Dinesh Yelamanchili, Nurul Amin Choudhury, and Badal Soni. 2024. An Efficient and Optimized CNN-LSTM Framework for Complex Human Activity Recognition System Using Surface EMG Physiological Sensors and Feature Engineering. In 2024 IEEE Students C...

  62. [70]

    Jiwei Wang, Yiqiang Chen, Yang Gu, Yunlong Xiao, and Haonan Pan. 2018. Sensorygans: An effective generative adversarial framework for sensor-based human activity recognition. In 2018 International Joint Conference on Neural Networks (IJCNN) . IEEE, 1–8

  63. [71]

    Anjana Wijekoon, Nirmalie Wiratunga, and Kay Cooper. 2019. Mex: Multi-modal exercises dataset for human activity recognition. arXiv preprint arXiv:1908.08992 (2019)

  64. [72]

    Erwin Wu, Rawal Khirodkar, Hideki Koike, and Kris Kitani. 2024. SolePoser: Full Body Pose Estimation using a Single Pair of Insole Sensor. In Proceedings of the 37th Annual ACM Symposium on User Interface Software and Technology . 1–9

  65. [73]

    Yi Wu, Luis Alonso González Villalobos, Zhenning Yang, Gregory Thomas Croisdale, Çağdaş Karataş, and Jian Liu. 2023. SmarCyPad: A Smart Seat Pad for Cycling Fitness Tracking Leveraging Low-cost Conductive Fabric Sensors. Proceedings of the ACM on Interactive, Mobile, Wearable ...

  66. [74]

    Ziyu Wu, Fangting Xie, Yiran Fang, Zhen Liang, Quan Wan, Yufan Xiong, and Xiaohui Cai. 2024. Seeing through the Tactile: 3D Human Shape Estimation from Temporal In-Bed Pressure Images. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies 8, 2 (20...

  67. [75]

    Lei Xiao, Yang Cao, Yihe Gai, Edris Khezri, Juntong Liu, and Mingzhu Yang. 2023. Recognizing sports activities from video frames using deformable convolution and adaptive multiscale features. Journal of Cloud Computing 12, 1 (2023), 167

  68. [76]

    Sara Zhalehpour, Onur Onder, Zahid Akhtar, and Cigdem Eroglu Erdem. 2016. BAUM-1: A spontaneous audio-visual face database of affective and mental states. IEEE Transactions on Affective Computing 8, 3 (2016), 300–313

  69. [77]

    Hang Zhang, Xin Li, and Lidong Bing. 2023. Video-llama: An instruction-tuned audio-visual language model for video understanding. arXiv preprint arXiv:2306.02858 (2023)

  70. [78]

    Jianrong Zhang, Yangsong Zhang, Xiaodong Cun, Yong Zhang, Hongwei Zhao, Hongtao Lu, Xi Shen, and Ying Shan. 2023. Generating human motion from textual descriptions with discrete representations. In Proceedings of the IEEE/CVF conference on computer vision and pattern recogniti...

  71. [79]

    Siwei Zhang, Qianli Ma, Yan Zhang, Zhiyin Qian, Taein Kwon, Marc Pollefeys, Federica Bogo, and Siyu Tang. 2022. Egobody: Human body shape and motion of interacting people from head-mounted devices. In European conference on computer vision . Springer, 180–200

  72. [80]

    Yiyuan Zhang, Kaixiong Gong, Kaipeng Zhang, Hongsheng Li, Yu Qiao, Wanli Ouyang, and Xiangyu Yue. 2023. Meta-transformer: A unified framework for multimodal learning. arXiv preprint arXiv:2307.10802 (2023)

  73. [81]

    Junwen Zhong, Zhaoyang Li, Masahito Takakuwa, Daishi Inoue, Daisuke Hashizume, Zhi Jiang, Yujun Shi, Lexiang Ou, Md Osman Goni Nayeem, Shinjiro Umezu, et al. 2022. Smart face mask based on an ultrathin pressure sensor for wireless monitoring of breath conditions. Advanced Mate...

  74. [82]

    Mengjuan Zhong, Lijuan Zhang, Xu Liu, Yaning Zhou, Maoyi Zhang, Yangjian Wang, Lu Yang, and Di Wei. 2021. Wide linear range and highly sensitive flexible pressure sensor based on multistage sensing process for health monitoring and human-machine interfaces. Chemical Engineerin...

  75. [83]

    Bo Zhou, Sungho Suh, Vitor Fortes Rey, Carlos Andres Velez Altamirano, and Paul Lukowicz. 2022. Quali-mat: Evaluating the quality of execution in body-weight exercises with a pressure sensitive sports mat. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous ...

  76. [84]

    Parham Zolfaghari, Vitor Fortes Rey, Lala Ray, Hyun Kim, Sungho Suh, and Paul Lukowicz. 2024. Sensor data augmentation from skeleton pose sequences for improving human activity recognition. In 2024 International Conference on Activity and Behavior Computing (ABC). IEEE, 1–8. P...

  77. [2022]

    Soft Robotics 9, 3 (2022), 473–485

    Smart skin: Vision-based soft pressure sensing system for in-home hand rehabilitation. Soft Robotics 9, 3 (2022), 473–485

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.