Pith. sign in

REVIEW 4 major objections 5 minor 67 references

A 13.2k-parameter 1D CNN can classify affective touch from a plush toy's sensors, reaching 85% accuracy on unseen users and a 20 Hz real-time budget on a low-power microcontroller.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-02 16:09 UTC pith:QOP4JPNJ

load-bearing objection The public dataset and open pipeline are the real contributions; the 20 Hz embedded feasibility claim rests on a benchmark transfer that likely overstates real ESP32 performance. the 4 major comments →

arxiv 2607.16196 v1 pith:QOP4JPNJ submitted 2026-04-16 cs.AI

Design and Validation of a Lightweight 1D CNN for Affective Touch Classification in Soft Plush Companions

classification cs.AI
keywords affective touch1D CNNembedded machine learninggesture recognitionsoft interactive companionleave-one-subject-out cross-validationdilated convolutiontactile sensing
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper aims to show that affective touch—gentle stroking, scratching, holding, hitting—can be classified from the multichannel capacitive and accelerometer sensors inside a soft plush companion by a compact 1D convolutional network. It reports that a three-layer dilated network with about 13,200 parameters reaches 75% accuracy on a held-out test split and a mean 85% accuracy on entirely unseen users via leave-one-subject-out cross-validation. It further argues, from published microcontroller benchmark data, that quantized inference costs about 3.2 million multiply-accumulates per 2.5-second window, which fits a 20 Hz real-time budget on the target low-power microcontroller. The practical claim is that emotionally meaningful touch interpretation can run entirely on-device, without cloud or external compute, and that a hybrid of a fast threshold filter for energetic negative gestures plus the CNN for subtle ones is the right embedded design.

Core claim

On its own terms, the paper's central discovery is that a deliberately small dilated 1D CNN—layer filter counts 14, 39, 41 with dilation rates 1, 2, 4—matches or beats far larger networks on this tactile gesture task. Contrary to the usual scaling assumption, models with tens of thousands of parameters performed in the same accuracy band as models with hundreds of thousands, and dilated layers consistently led the rankings. Subject-independent evaluation (LOSO-CV, 25 folds) gave a mean accuracy of 0.850 ± 0.072 and macro-F1 of 0.687 ± 0.114, with confusions concentrated in gestures that have similar signal profiles rather than across emotional categories. The authors interpret this as eviden

What carries the argument

The load-bearing object is the '14d1_39d2_41d4' architecture: three 1D convolutional layers with 14, 39, and 41 filters and dilation rates 1, 2, and 4. In the symbolic notation, a number is the filter count and 'd' a dilation coefficient. Dilated convolutions widen the temporal receptive field—covering the full 2.5-second, 250-sample window—without adding parameters, which is what lets a 13.2k-parameter model match much larger ones. The second piece of machinery is the sliding-window protocol: 250-sample windows with a 50-sample step, loop-padded rather than zero-padded for short sequences, turning roughly 1,300 raw records into about 5,600 training windows. The third is the hybrid inference

Load-bearing premise

The real-time feasibility rests on published microcontroller benchmark speeds, not on a measured run of the actual quantized model; if the real device runs even two or three times slower than those benchmarks, the 20 Hz target is missed.

What would settle it

Measure wall-clock inference time of the fully quantized 14d1_39d2_41d4 model with an 11×250 input window on the target microcontroller. If per-window latency exceeds 50 ms (that is, throughput falls below 20 Hz) under realistic memory and battery conditions, the paper's embedded-feasibility claim is refuted. A secondary falsifier: run a larger subject-independent evaluation with many more unseen participants; if mean LOSO accuracy drops well below 75%, the generalization claim weakens.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • On-device affective touch recognition becomes plausible for soft therapeutic toys: no cloud link, no laptop, and privacy-preserving local interpretation of a child's touch.
  • The open dataset and pipeline let other groups train and compare affective-touch models without re-collecting subject data.
  • The accuracy-complexity analysis suggests that for this sensor and data regime, very small dilated CNNs are sufficient, so deployment cost and power budget can be low.
  • The hybrid heuristic-plus-CNN pattern gives immediate reaction for high-force negative gestures (0.1–0.2 s) while the CNN resolves subtle gestures the old heuristic missed (e.g., activation in ~6% vs ~100% of subtle-gesture trials).
  • Confusion clusters are within same-emotion families, implying that merging similar gesture classes could push accuracy higher without collecting new data.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The 75% random-split figure versus the 85% LOSO figure is puzzle for a strict 'random split is an upper bound' reading; the two protocols are not directly comparable, and window-level leakage can inflate training variance in unexpected ways.
  • The real-time estimate hinges on benchmark-derived throughput; until the quantized model is actually run on the target microcontroller, a two-to-threefold slowdown from memory access or scheduling would push inference past the 50 ms frame budget.
  • The 'merge similar classes' suggestion points toward a valence-only (positive/negative/neutral) classifier as a more robust product, which the confusion patterns already support.
  • A testable extension: train the same 14d1_39d2_41d4 architecture on other soft-sensor systems with different sensor layouts to see whether the dilated 1D-CNN result transfers or is specific to this 11-channel capacitive-plus-accelerometer setup.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper presents a framework for affective touch classification in soft plush companions, built around a newly collected FAIR-compliant dataset of 1326 labelled gesture sequences from 25 participants and an open-source MATLAB pipeline. Through systematic exploration of 468 CNN variants, the authors select a compact dilated 1D CNN (14d1_39d2_41d4, 13.2k parameters) that achieves 75% test accuracy on a random window-level split and 85% mean accuracy under leave-one-subject-out cross-validation. The paper further estimates a quantized inference cost of 3.2 MMAC per 2.5 s window and, using published ESP32 DSP benchmarks, derives a 23 ms inference time, concluding that 20 Hz real-time operation on an ESP32-WROVER-E is feasible. A PC-based real-time simulation with the physical toy is used to qualitatively compare the CNN against a prior heuristic algorithm, leading to a proposed hybrid pipeline. The paper is positioned as a model-development and validation study, with physical hardware integration explicitly deferred to future work.

Significance. If the central results hold, this paper demonstrates that lightweight 1D CNNs can bring affective touch classification to microcontroller-class hardware, which is relevant for privacy-preserving and autonomous social assistive devices. The strongest concrete contributions are the public, FAIR-compliant dataset, the complete open-source MATLAB pipeline, and the systematic 468-model architecture search; these are reproducible and reusable assets. The LOSO-CV evaluation is a meaningful step beyond the random-split results, and the qualitative real-time comparison with the heuristic baseline provides useful engineering insight. However, the headline 20 Hz real-time claim rests on unmeasured benchmark transfer with a likely double-counting of quantization speedup, and the generalization claim is weakened by a macro-F1 score near 0.69 despite 85% mean accuracy. The paper is honestly qualified in places, but those qualifications need to be more consistently reflected in the abstract and conclusions.

major comments (4)
  1. [§4.7] The 23 ms inference estimate and the resulting 20 Hz real-time claim are not supported by the cited evidence. The 137 MMAC/s Int16 figure already corresponds to quantized int16 operations; applying an additional '4–5x quantization speedup' from [65] on top of it double-counts the benefit, and [65] is for Arm Cortex-M CMSIS-NN kernels, not the ESP32's Xtensa LX6. The ESP-DSP benchmarks are also DSP-library peak rates, not end-to-end CNN inference; memory access, activation, batch-norm folding, softmax, and firmware overhead are omitted. The paper acknowledges these are estimates, but the abstract and §5.1 present 20 Hz operation as a demonstrated capability. The authors should either measure inference on the target MCU or explicitly relabel the 20 Hz claim as an idealized projection pending direct measurement.
  2. [§4.4 vs §3.6] The LOSO-CV protocol uses a different, improved training schedule (initial LR 0.003, piecewise decay, early stopping, up to 150 epochs) than the architecture-search protocol (Adam, LR 0.001, no early stopping). The 85% LOSO accuracy therefore does not evaluate the same model configuration that produced the 75% random-split test accuracy, and the architecture itself was selected on the random window-level split with acknowledged subject leakage. The paper should state explicitly that the 85% figure applies to a retrained variant, and should ideally evaluate the final architecture under both protocols with identical training schedules to permit a clean comparison. As written, the distinction is buried in §4.4 while the abstract highlights 85% without this caveat.
  3. [§4.4/§5.1] The headline LOSO metric is mean accuracy 0.850, but macro F1 is 0.687 and macro recall 0.717. For 18 imbalanced classes, accuracy overstates per-class performance; the claim in §5.1 that the model 'generalizes well to unseen users' is weakened by per-class F1 scores below 0.5 for Tail hit 2, Tail scratch 2, and Tail tap 2. The paper should report class-balanced metrics prominently, explain why accuracy is the primary headline, and temper the generalization conclusion to reflect the macro-F1 result.
  4. [§4.7 / §3.6.2] The 3.2 MMAC calculation is asserted without the intermediate arithmetic. The number of input channels is ambiguous: §3.1 replaces the accelerometer X/Y/Z with a magnitude (11 channels), while §3.6.1 refers to '13-channel signals'. Include a reproducible table or formula specifying input channels, kernel sizes, dilations, padding, stride, and layer-wise MACs, so that the memory (14–53 KB) and MAC totals can be verified by readers.
minor comments (5)
  1. [§3.6.1] The text says '13-channel signals' but §3.1 and Table 1 imply 11 channels after replacing accelerometer X/Y/Z with a magnitude. Please reconcile.
  2. [§4.4] The sentence 'LOSO-CV on the selected CNN architecture 14d1_39d2_41d4 with 8, 12, and 16 filters in the three convolutional layers' is ambiguous: clarify that the 13.2k-parameter configuration uses 8, 12, and 16 filters, and specify which configuration produced the reported LOSO results.
  3. [§4.7] Reference [64] should specify whether the 58 MMAC/s and 137 MMAC/s figures apply to the ESP32-WROVER-E specifically or to a generic ESP32, and under what clock/optimization settings.
  4. [§4.6] Reaction times were measured manually with ±0.3 s uncertainty. Consider automated timestamp logging from the simulator to make these latency comparisons more reliable and reproducible.
  5. [§5.2] The phrase '50ms steps, exactly as the models were tested in the real-time simulator' would be clearer if the simulator's prediction-rate setting in §3.5 is explicitly reported for the experiments in §4.6.

Circularity Check

0 steps flagged

No significant circularity: the accuracy numbers are empirical evaluations, the MAC/20 Hz estimate is externally benchmarked arithmetic explicitly flagged as unmeasured, and self-citations provide data/context rather than proof.

full rationale

The paper's headline claims are measured results, not definitional equivalences. The 75% test accuracy and 85% LOSO-CV accuracy are obtained by training on splits/folds and evaluating on held-out windows/subjects; they are not defined as functions of fitted parameters. The 3.2 MMAC figure follows from the selected architecture's layer sizes and dilation settings, and the 20 Hz inference estimate is arithmetic from published ESP-DSP benchmark MAC/s rates plus the cited CMSIS-NN quantization speedup. Section 4.7 explicitly says 'these are benchmark evaluations' and 'the reported inference times are estimates based on ESP32 performance benchmarks, rather than direct measurements,' so the feasibility claim is presented as a theoretical transfer, not as a measured result. Any weakness there is benchmark-transfer risk, not circular reasoning. Self-citations [18], [53], [54], [55] are used for prior-device context, a companion behavioral study, and the open dataset/software; none is invoked as a uniqueness theorem or as the proof of the model's accuracy. The paper also acknowledges that the random-window test accuracies are 'upper-bound estimates' and supplies LOSO-CV as a stricter subject-independent check. No prediction in the paper reduces by construction to its own input.

Axiom & Free-Parameter Ledger

5 free parameters · 3 axioms · 0 invented entities

The accuracy claims rest on user-selected hyperparameters (window length, step, architecture, LOSO schedule, QA threshold) and on domain assumptions about sensor informativeness and label validity. No new physical entities are postulated. The CNN weights themselves are learned, not free parameters in the derivation sense, but the training choices directly modulate the headline numbers.

free parameters (5)
  • Window length (250 samples / 2.5 s) = 250 samples (2.5 s)
    Chosen from initial manual tests as the shortest time needed to distinguish the 18 gesture classes; determines the receptive field and all reported accuracies (Section 3.2, 5.1).
  • Sliding window step (50 samples / 0.5 s) = 50 samples
    Augmentation step that expands ~1300 samples to ~5600 windows; influences training data density and classification cadence (Section 3.2).
  • Architecture hyperparameters for 14d1_39d2_41d4 = filters [8,12,16], dilation [1,2,4], 250 epochs, mini-batch 64
    Selected as the top compact model on the random-split test set among 468 trials; the headline 75% accuracy depends on this selection (Section 4.2).
  • LOSO training schedule = LR 0.003, decay 0.5 every 30 epochs, patience 10, max 150 epochs
    Used for the 85% LOSO-CV result and differs from the architecture-search schedule; the LOSO numbers are tied to these hyperparameters (Section 4.4).
  • QA outlier threshold = ±40% of mean duration
    Manual threshold for flagging duration outliers before training; affects dataset composition (Section 3.2).
axioms (3)
  • domain assumption Capacitive and accelerometer-magnitude signals at 100 Hz contain sufficient information to distinguish the 18 affective gesture classes.
    The whole supervised pipeline assumes the chosen sensor channels and sampling rate encode label information; no information-theoretic check is provided.
  • domain assumption Instructed gesture labels are a valid ground truth for affective touch categories.
    Labels come from participants following a gesture protocol, not from independent affect ratings or real therapeutic interaction; Section 3.2 defines the gestures.
  • domain assumption Sliding windows from the same subject can be treated as independent samples in random-split experiments.
    This assumption is used for architecture search and baseline 5-fold CV; the paper acknowledges it causes subject leakage and addresses it separately with LOSO-CV.

pith-pipeline@v1.3.0-alltime-deepseek · 22703 in / 11759 out tokens · 434986 ms · 2026-08-02T16:09:03.601716+00:00 · methodology

0 comments
read the original abstract

Soft, sensorized companions offer a physically safe and emotionally intuitive interface for socially assistive technologies, yet their deformability and multichannel tactile sensing complicate the robust interpretation of human affect. This study presents a complete open-source MATLAB-based framework for the development and validation of compact deep learning models for affective touch recognition in soft interactive companions. As a primary contribution, a diverse FAIR-compliant dataset of 1326 labelled gesture sequences collected from 25 participants spanning children, teenagers, and adults is made publicly available, providing a reusable resource for future research in affective touch recognition. Through systematic architecture and hyperparameter exploration across 468 CNN models, the study identifies compact dilated one-dimensional convolutional neural networks (1D CNNs) as the most effective solution, with a 13.2k-parameter model achieving 75% test accuracy and 85% mean leave-one-subject-out cross-validation accuracy. Theoretical inference-time analysis shows that quantized deployment requires 3.2 MMAC per window, compatible with 20 Hz real-time operation on the target microcontroller. PC-based real-time simulation with the physical toy streaming sensor data demonstrates that the CNN resolves subtle social touches that the previous heuristic system failed to detect, whereas high-force negative interactions are captured more reliably by trivial threshold-based logic. The resulting hybrid inference pipeline - instantaneous heuristic filtering followed by CNN-based nuanced gesture classification - is proposed as the embedded deployment strategy. The study demonstrates that emotionally meaningful, privacy-preserving touch interpretation is computationally feasible for direct embedding within soft therapeutic companions, with hardware integration addressed in a forthcoming study.

Figures

Figures reproduced from arXiv: 2607.16196 by Airisa \v{S}teinberga, Aleksandrs Okss, Aleksandrs Vali\v{s}evskis, Aleksejs Kata\v{s}evs, Anete Hofmane, Dina Bethere, Inese T\=i\c{g}ere, Lucie Matou\v{s}kov\'a, Santa Me\c{l}\c{k}e, Und\=ine Gavri\c{l}enko.

Figure 1
Figure 1. Figure 1: Soft plush companions with integrated capacitive sensors. (a) – Soft plush companions; (b) – Sensors placement However, all the variations share the same hardware part. The toys have several capacitive touch sensors that are integrated under the plush layer. These are located in the extremities, ears, head, neck, nose, tail, back, belly, see [PITH_FULL_IMAGE:figures/full_fig_p005_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Gesture Recorder tool used to collect training data [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Example of a QA routine output 3. Truncation and alignment – The script c_training_data_truncation.m provides an interactive tool for trimming initial latency caused by human reaction time, see [PITH_FULL_IMAGE:figures/full_fig_p008_3.png] view at source ↗
Figure 5
Figure 5. Figure 5: Sliding window method for training data segmentation [PITH_FULL_IMAGE:figures/full_fig_p009_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: GUI of a real-time gesture classifier The simulator establishes a real-time asynchronous data link with the interactive toy, which is continuously streaming sensor data. The simulator displays classification probabilities across gesture classes in real time. The GUI shows both the incoming signals and the current prediction with the highest ranking, as well as likelihood evaluations for all the classes. Th… view at source ↗
Figure 7
Figure 7. Figure 7: Accuracy of the trained CNNs As can be seen, models with dilated layers are clearly in the lead both regarding the maximum accuracy and the average accuracy, which is calculated across all the hyperparameter variations. Another observation is that bigger models do not necessarily perform better. The number of filters influences greatly the size of the model. By testing various configurations indicated in S… view at source ↗
Figure 8
Figure 8. Figure 8: Performance of models, depending on the number of filters in each layer: (a) [32 48 64 64]; (b) [16 24 32 32]; (c) [8 12 16 16]; (d) [4 6 8 8] While only the bigger models with number of filters [32 48 64 64] were close to accuracy of 80% - architecture 14d1_39d2_41d4 demonstrated accuracy of 79.03%, their size makes them unfit for embedded systems with scarce computational power and limited memory capacit… view at source ↗
Figure 9
Figure 9. Figure 9: Accuracy and size of top performers Based on these data, 14d1_39d2_41d4 architecture was selected for further analysis, as it demonstrates a solid performance with accuracy on unseen test data of 75% with only a fraction of learnable parameters, which is 13.2k. This level of accuracy was reached after 250 epochs, learning algorithm used MiniBatchSize=64. Due to its compactness, it has a high ACR value of 1… view at source ↗
Figure 10
Figure 10. Figure 10: Confusion matrix for 14d1_39d2_41d4 It is important to mention that the matrix shows no such confusion among classes that belong to different emotional classes (negative, positive, neutral). Thus, these confusions do not influence the overall performance of the model and in order to improve quantitative performance of the model, our suggestion would be to merge some of the similar classes in future model … view at source ↗
Figure 11
Figure 11. Figure 11: Confusion matrix for 19d1_19d2_19d4_17d8 Confusion matrix for the model 19d1_19d2_19d4_17d8 share similar characteristics and is shown in [PITH_FULL_IMAGE:figures/full_fig_p018_11.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

67 extracted references · 26 canonical work pages

  1. [1]

    Scassellati, H

    B. Scassellati, H. Admoni, M. Matarić, Robots for Use in Autism Research, Annu. Rev. Biomed. Eng. 14 (2012) 275–294. https://doi.org/10.1146/annurev-bioeng-071811-150036

  2. [2]

    Hertenstein, D

    M.J. Hertenstein, D. Keltner, B. App, B.A. Bulleit, A.R. Jaskolka, Touch communicates distinct emotions., Emotion 6 (2006) 528–533. https://doi.org/10.1037/1528-3542.6.3.528

  3. [3]

    McGlone, J

    F . McGlone, J. Wessberg, H. Olausson, Discriminative and Affective Touch: Sensing and Feeling, Neuron 82 (2014) 737–755. https://doi.org/10.1016/j.neuron.2014.05.001

  4. [4]

    Mastergeorge, C

    A.M. Mastergeorge, C. Kahathuduwa, J. Blume, Eye-Tracking in Infants and Young Children at Risk for Autism Spectrum Disorder: A Systematic Review of Visual Stimuli in Experimental Paradigms, J Autism Dev Disord 51 (2021) 2578–2599. https://doi.org/10.1007/s10803-020-04731-w

  5. [5]

    McLaughlin, H.E

    C.S. McLaughlin, H.E. Grosman, S.B. Guillory, E.L. Isenstein, E. Wilkinson, M.D.P . Trelles, D.B. Halpern, P .M. Siper, A. Kolevzon, J.D. Buxbaum, A.T. Wang, J.H. Foss-Feig, Reduced engagement of visual attention in children with autism spectrum disorder, Autism 25 (2021) 2064–2073. https://doi.org/10.1177/13623613211010072. 25

  6. [6]

    James, E

    P . James, E. Schafer, J. Wolfe, L. Matthews, S. Browning, J. Oleson, E. Sorensen, G. Rance, L. Shiels, A. Dunn, Increased rate of listening difficulties in autistic children, Journal of Communication Disorders 99 (2022) 106252. https://doi.org/10.1016/j.jcomdis.2022.106252

  7. [7]

    Kushniruk, W

    A.W. Kushniruk, W. Kuang, E.M. Borycki, The NAO Robot in Healthcare and Education: A Scoping Review, in: J. Mantas, A. Hasman, E. Zoulias, K. Karitis, P . Gallos, M. Diomidous, S. Zogas, M. Charalampidou (Eds.), Studies in Health Technology and Informatics, IOS Press, 2025. https://doi.org/10.3233/SHTI250115

  8. [8]

    Mishra, G.A

    D. Mishra, G.A. Romero, A. Pande, B. Nachenahalli Bhuthegowda, D. Chaskopoulos, B. Shrestha, An Exploration of the Pepper Robot’s Capabilities: Unveiling Its Potential, Applied Sciences 14 (2023) 110. https://doi.org/10.3390/app14010110

  9. [9]

    L.J. Wood, A. Zaraki, B. Robins, K. Dautenhahn, Developing Kaspar: A Humanoid Robot for Children with Autism, Int J of Soc Robotics 13 (2021) 491–508. https://doi.org/10.1007/s12369-019-00563-6

  10. [10]

    L. Hung, C. Liu, E. Woldum, A. Au-Yeung, A. Berndt, C. Wallsworth, N. Horne, M. Gregorio, J. Mann, H. Chaudhury, The benefits of and barriers to using a social robot PARO in care settings: a scoping review, BMC Geriatr 19 (2019) 232. https://doi.org/10.1186/s12877-019-1244-6

  11. [11]

    Bennett, C

    C.C. Bennett, C. Stanojevic, S. Kim, J. Lee, J. Yu, J. Oh, S. Sabanovic, J.A. Piatt, Comparison of In- home Robotic Companion Pet Use in South Korea and the United States: A Case Study, in: 2022 9th IEEE RAS/EMBS International Conference for Biomedical Robotics and Biomechatronics (BioRob), IEEE, Seoul, Korea, Republic of, 2022: pp. 01–07. https://doi.org...

  12. [12]

    Nadeem, J.M.H

    M. Nadeem, J.M.H. Barakat, D. Daas, A. Potams, A Review of Socially Assistive Robotics in Supporting Children with Autism Spectrum Disorder, MTI 9 (2025) 98. https://doi.org/10.3390/mti9090098

  13. [13]

    K. Wada, T. Shibata, Living With Seal Robots—Its Sociopsychological and Physiological Influences on the Elderly at a Care House, IEEE Trans. Robot. 23 (2007) 972–980. https://doi.org/10.1109/TRO.2007.906261

  14. [14]

    N. Geva, F . Uzefovsky, S. Levy-Tzedek, Touching the social robot PARO reduces pain perception and salivary oxytocin levels, Sci Rep 10 (2020) 9814. https://doi.org/10.1038/s41598-020-66982-y

  15. [15]

    Fogelson, C

    D.M. Fogelson, C. Rutledge, K.S. Zimbro, The Impact of Robotic Companion Pets on Depression and Loneliness for Older Adults with Dementia During the COVID-19 Pandemic, J Holist Nurs 40 (2022) 397–

  16. [16]

    Landowska, A

    A. Landowska, A. Karpus, T. Zawadzka, B. Robins, D. Erol Barkana, H. Kose, T. Zorcec, N. Cummins, Automatic Emotion Recognition in Children with Autism: A Systematic Literature Review, Sensors 22 (2022) 1649. https://doi.org/10.3390/s22041649

  17. [17]

    Garcia, F

    S. Garcia, F . Gomez-Donoso, M. Cazorla, Enhancing Human–Robot Interaction: Development of Multimodal Robotic Assistant for User Emotion Recognition, Applied Sciences 14 (2024) 11914. https://doi.org/10.3390/app142411914

  18. [18]

    Bethere, I

    D. Bethere, I. Tīģere, A. Hofmane, A. Šteinberga, U. Gavriļenko, S. Meļķe, A. Okss, A. Kataševs, A. Vališevskis, AI-Powered Plush Robots for Children with ASD in Education, Rehabilitation: Expert Evaluation, Ije 13 (2025) 61–86. https://doi.org/10.22492/ije.13.2.03

  19. [19]

    Breazeal, Emotion and sociable humanoid robots, International Journal of Human-Computer Studies 59 (2003) 119–155

    C. Breazeal, Emotion and sociable humanoid robots, International Journal of Human-Computer Studies 59 (2003) 119–155. https://doi.org/10.1016/S1071-5819(03)00018-1

  20. [20]

    Nomura, T

    T. Nomura, T. Uratani, T. Kanda, K. Matsumoto, H. Kidokoro, Y . Suehiro, S. Yamada, Why Do Children Abuse Robots?, in: Proceedings of the Tenth Annual ACM/IEEE International Conference on Human- Robot Interaction Extended Abstracts, ACM, Portland Oregon USA, 2015: pp. 63–64. https://doi.org/10.1145/2701973.2701977. 26

  21. [21]

    Salvini, G

    P . Salvini, G. Ciaravella, W. Yu, G. Ferri, A. Manzi, B. Mazzolai, C. Laschi, S.R. Oh, P . Dario, How safe are service robots in urban environments? Bullying a robot, in: 19th International Symposium in Robot and Human Interactive Communication, IEEE, Viareggio, 2010: pp. 1–7. https://doi.org/10.1109/ROMAN.2010.5654677

  22. [22]

    Brščić, H

    D. Brščić, H. Kidokoro, Y . Suehiro, T. Kanda, Escaping from Children’s Abuse of Social Robots, in: Proceedings of the Tenth Annual ACM/IEEE International Conference on Human-Robot Interaction, ACM, Portland Oregon USA, 2015: pp. 59–66. https://doi.org/10.1145/2696454.2696468

  23. [23]

    Coeckelbergh, Artificial Companions: Empathy and Vulnerability Mirroring in Human-Robot Relations, Studies in Ethics, Law, and Technology 4 (2011)

    M. Coeckelbergh, Artificial Companions: Empathy and Vulnerability Mirroring in Human-Robot Relations, Studies in Ethics, Law, and Technology 4 (2011). https://doi.org/10.2202/1941-6008.1126

  24. [24]

    Ibrahim Ahmed Al-mashhadani, T

    M. Ibrahim Ahmed Al-mashhadani, T. H. H. Aldhyani, M. Hmoud Al-Adhaileh, A. M. Bamhdi, M. Y . Alzahrani, F . Waselallah Alsaade, H. Alkahtani, Human-Animal Affective Robot Touch Classification Using Deep Neural Network, Computer Systems Science and Engineering 38 (2021) 25–37. https://doi.org/10.32604/csse.2021.014992

  25. [25]

    Cascio, D

    C.J. Cascio, D. Moore, F . McGlone, Social touch and human development, Developmental Cognitive Neuroscience 35 (2019) 5–11. https://doi.org/10.1016/j.dcn.2018.04.009

  26. [26]

    C. Y . Zheng, Y . Chen, A. Latupeirissa, G. Andrikopoulos, A. Ståhl, M. Balaam, Towards Caring Touch From Technologies: Knowledge From Healthcare Practitioners, in: Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems, ACM, Yokohama Japan, 2025: pp. 1–16. https://doi.org/10.1145/3706598.3713736

  27. [27]

    Hertenstein, J.M

    M.J. Hertenstein, J.M. Verkamp, A.M. Kerestes, R.M. Holmes, The Communicative Functions of Touch in Humans, Nonhuman Primates, and Rats: A Review and Synthesis of the Empirical Research, Genetic, Social, and General Psychology Monographs 132 (2006) 5–94. https://doi.org/10.3200/MONO.132.1.5-94

  28. [28]

    Mazzei, C

    D. Mazzei, C. De Maria, G. Vozzi, Touch sensor for social robots and interactive objects affective interaction, Sensors and Actuators A: Physical 251 (2016) 92–99. https://doi.org/10.1016/j.sna.2016.10.006

  29. [29]

    Bennett, S

    C.C. Bennett, S. Sabanovic, C. Stanojevic, Z. Henkel, S. Kim, J. Lee, K. Baugus, J.A. Piatt, J. Yu, J. Oh, S. Collins, C.L. Bethel, Enabling Robotic Pets to Autonomously Adapt Their Own Behaviors to Enhance Therapeutic Effects: A Data-Driven Approach* , in: 2023 32nd IEEE International Conference on Robot and Human Interactive Communication (RO-MAN), IEEE...

  30. [30]

    Burns, H

    R.B. Burns, H. Lee, H. Seifi, R. Faulkner, K.J. Kuchenbecker, Endowing a NAO Robot With Practical Social-Touch Perception, Front. Robot. AI 9 (2022) 840335. https://doi.org/10.3389/frobt.2022.840335

  31. [31]

    E. Y . Zhang, Z. Pan, A.D. Cheok, Emotion Recognition Using Affective Touch: A Survey, IEEE Trans. Affective Comput. (2025) 1–20. https://doi.org/10.1109/TAFFC.2025.3592197

  32. [32]

    Picard, Affective Computing, The MIT Press, 1997

    R.W. Picard, Affective Computing, The MIT Press, 1997. https://doi.org/10.7551/mitpress/1140.001.0001

  33. [33]

    Cowie, E

    R. Cowie, E. Douglas-Cowie, N. Tsapatsoulis, G. Votsis, S. Kollias, W. Fellenz, J.G. Taylor, Emotion recognition in human-computer interaction, IEEE Signal Process. Mag. 18 (2001) 32–80. https://doi.org/10.1109/79.911197

  34. [34]

    Calvo, S

    R.A. Calvo, S. D’Mello, Affect Detection: An Interdisciplinary Review of Models, Methods, and Their Applications, IEEE Trans. Affective Comput. 1 (2010) 18–37. https://doi.org/10.1109/T-AFFC.2010.1

  35. [35]

    Tapus, M

    A. Tapus, M. Mataric, B. Scassellati, Socially assistive robotics [Grand Challenges of Robotics], IEEE Robot. Automat. Mag. 14 (2007) 35–42. https://doi.org/10.1109/MRA.2007.339605

  36. [36]

    Cansev, A.J

    M.E. Cansev, A.J. Miller, J.D. Brown, P . Beckerle, Implementing social and affective touch to enhance user experience in human-robot interaction, Front. Robot. AI 11 (2024) 1403679. https://doi.org/10.3389/frobt.2024.1403679. 27

  37. [37]

    Ramaswamy, S

    M.P .A. Ramaswamy, S. Palaniswamy, Multimodal emotion recognition: A comprehensive review, trends, and challenges, WIREs Data Min & Knowl 14 (2024) e1563. https://doi.org/10.1002/widm.1563

  38. [38]

    Dahiya, P

    R.S. Dahiya, P . Mittendorfer, M. Valle, G. Cheng, V .J. Lumelsky, Directions Toward Effective Utilization of Tactile Skin: A Review, IEEE Sensors J. 13 (2013) 4121–4138. https://doi.org/10.1109/JSEN.2013.2279056

  39. [39]

    Alenljung, R

    B. Alenljung, R. Andreasson, R. Lowe, E. Billing, J. Lindblom, Conveying Emotions by Touch to the Nao Robot: A User Experience Perspective, MTI 2 (2018) 82. https://doi.org/10.3390/mti2040082

  40. [40]

    Vagnetti, A

    R. Vagnetti, A. Di Nuovo, M. Mazza, M. Valenti, Social Robots: A Promising Tool to Support People with Autism. A Systematic Review of Recent Research and Critical Analysis from the Clinical Perspective, Rev J Autism Dev Disord (2024). https://doi.org/10.1007/s40489-024-00434-5

  41. [41]

    An Emotional Support Animal, Without the Animal

    S. Collins, K. Baugus Henkel, Z. Henkel, C.C. Bennett, C. Stanojevic, J.A. Piatt, C.L. Bethel, S. Sabanović, “An Emotional Support Animal, Without the Animal”: Design Guidelines for a Social Robot to Address Symptoms of Depression, in: Proceedings of the 2024 ACM/IEEE International Conference on Human-Robot Interaction, ACM, Boulder CO USA, 2024: pp. 147–...

  42. [42]

    Silva-Plata, C

    C. Silva-Plata, C. Rosel, B.G. Cangan, H. Alagi, B. Hein, R.K. Katzschmann, R. Fernández, Y . Mojtahedi, S.E. Navarro, Model-Based Capacitive Touch Sensing in Soft Robotics: Achieving Robust Tactile Interactions for Artistic Applications, IEEE Robot. Autom. Lett. 10 (2025) 4596–4603. https://doi.org/10.1109/LRA.2025.3539940

  43. [43]

    Burns, H

    R.B. Burns, H. Seifi, H. Lee, K.J. Kuchenbecker, Getting in touch with children with autism: Specialist guidelines for a touch-perceiving robot, Paladyn, Journal of Behavioral Robotics 12 (2020) 115–135. https://doi.org/10.1515/pjbr-2021-0010

  44. [44]

    Cabibihan, H

    J.-J. Cabibihan, H. Javed, M. Ang, S.M. Aljunied, Why Robots? A Survey on the Roles and Benefits of Social Robots in the Therapy of Children with Autism, Int J of Soc Robotics 5 (2013) 593 –618. https://doi.org/10.1007/s12369-013-0202-2

  45. [45]

    Teyssier, G

    M. Teyssier, G. Bailly, C. Pelachaud, E. Lecolinet, Conveying Emotions Through Device-Initiated Touch, IEEE Trans. Affective Comput. 13 (2022) 1477–1488. https://doi.org/10.1109/TAFFC.2020.3008693

  46. [46]

    Albawi, O

    S. Albawi, O. Bayat, S. Al-Azawi, O.N. Ucan, Social Touch Gesture Recognition Using Convolutional Neural Network, Computational Intelligence and Neuroscience 2018 (2018) 1–10. https://doi.org/10.1155/2018/6973103

  47. [47]

    P . Wang, J. Liu, F . Hou, D. Chen, Z. Xia, S. Guo, Organization and Understanding of a Tactile Information Dataset TacAct For Physical Human-Robot Interaction, in: 2021 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), IEEE, Prague, Czech Republic, 2021: pp. 7328–

  48. [48]

    Kiranyaz, T

    S. Kiranyaz, T. Ince, M. Gabbouj, Real-Time Patient-Specific ECG Classification by 1-D Convolutional Neural Networks, IEEE Trans. Biomed. Eng. 63 (2016) 664–675. https://doi.org/10.1109/TBME.2015.2468589

  49. [49]

    Kiranyaz, O

    S. Kiranyaz, O. Avci, O. Abdeljaber, T. Ince, M. Gabbouj, D.J. Inman, 1D convolutional neural networks and applications: A survey, Mechanical Systems and Signal Processing 151 (2021) 107398. https://doi.org/10.1016/j.ymssp.2020.107398

  50. [50]

    F . Yu, V . Koltun, Multi-Scale Context Aggregation by Dilated Convolutions, (2016). https://doi.org/10.48550/arXiv.1511.07122

  51. [51]

    Li, Q.-H

    Y.-K. Li, Q.-H. Meng, Y .-X. Wang, H.-R. Hou, MMFN: Emotion recognition by fusing touch gesture and facial expression information, Expert Systems with Applications 228 (2023) 120469. https://doi.org/10.1016/j.eswa.2023.120469. 28

  52. [52]

    Ghiglino, P

    D. Ghiglino, P . Chevalier, F . Floris, T. Priolo, A. Wykowska, Follow the white robot: Efficacy of robot- assistive training for children with autism spectrum disorder, Research in Autism Spectrum Disorders 86 (2021) 101822. https://doi.org/10.1016/j.rasd.2021.101822

  53. [53]

    Hofmane, I

    A. Hofmane, I. Tīģere, A. Šteinberga, D. Bethere, S. Meļķe, U. Gavriļenko, A. Okss, A. Kataševs, A. Vališevskis, Developing a Psychological Research Methodology for Evaluating AI-Powered Plush Robots in Education and Rehabilitation, Behavioral Sciences 15 (2025) 1310. https://doi.org/10.3390/bs15101310

  54. [54]

    Vališevskis, D

    A. Vališevskis, D. Bethere, I. Tīģere, A. Hofmane, A. Šteinberga, U. Gavriļenko, S. Meļķe, A. Okss, A. Kataševs (2025). Labelled training data for gesture recognition in smart plush companions [dataset], Zenodo, v1.0, https://doi.org/10.5281/zenodo.17650600

  55. [55]

    Vališevskis

    A. Vališevskis. End-to-end pipeline for training CNN for affective touch recognition in Smart Plush Companions v1.0 [software], Zenodo, November 19, 2025, https://doi.org/10.5281/zenodo.17650546

  56. [56]

    Yohanan, K.E

    S. Yohanan, K.E. MacLean, The Role of Affective Touch in Human-Robot Interaction: Human Intent and Expectations in Touching the Haptic Creature, Int J of Soc Robotics 4 (2012) 163–180. https://doi.org/10.1007/s12369-011-0126-7

  57. [57]

    Wilkinson, et al

    M.D. Wilkinson, et al. The FAIR Guiding Principles for scientific data management and stewardship, Sci Data 3 (2016) 160018. https://doi.org/10.1038/sdata.2016.18

  58. [58]

    Froehlich

    J.E. Froehlich. Gesture Recorder [software]. GitHub repository. https://github.com/makeabilitylab/arduino/tree/master/Processing/GestureRecorder

  59. [59]

    Goodfellow, Y

    I. Goodfellow, Y. Bengio, A. Courville. Deep Learning. 2016. MIT Press. https://www.deeplearningbook.org/

  60. [60]

    Dessai, H

    A. Dessai, H. Virani, Lightweight 1D CNN model for emotion classification using GSR signal, Ijirss 8 (2025) 5100–5116. https://doi.org/10.53894/ijirss.v8i3.7709

  61. [61]

    Graves, A

    A. Graves, A. Mohamed, G. Hinton, Speech recognition with deep recurrent neural networks, in: 2013 IEEE International Conference on Acoustics, Speech and Signal Processing, IEEE, Vancouver, BC, Canada, 2013: pp. 6645–6649. https://doi.org/10.1109/ICASSP .2013.6638947

  62. [62]

    T.-W. Chen, D. Wang, W. Tao, D. Wen, L. Yin, T. Ito, K. Osa, M. Kato, CASSOD-Net: Cascaded and Separable Structures of Dilated Convolution for Embedded Vision Systems and Applications, in: 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), IEEE, Nashville, TN, USA, 2021: pp. 3176–3184. https://doi.org/10.1109/CVPRW53098...

  63. [63]

    J. Hu, L. Shen, G. Sun, Squeeze-and-Excitation Networks, in: 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, IEEE, Salt Lake City, UT, 2018: pp. 7132–7141. https://doi.org/10.1109/CVPR.2018.00745

  64. [64]

    Espressif. (2023). ESP-DSP library benchmarks. Read the Docs. https://espressif-docs.readthedocs- hosted.com/_/downloads/esp-dsp/en/latest/pdf/

  65. [65]

    L. Lai, N. Suda, V . Chandra, CMSIS-NN: Efficient Neural Network Kernels for Arm Cortex-M CPUs, (2018). https://doi.org/10.48550/arXiv.1801.06601

  66. [409]

    https://doi.org/10.1177/08980101211064605

  67. [7333]

    https://doi.org/10.1109/IROS51168.2021.9636389