Pith. sign in

REVIEW 4 major objections 5 minor 147 references

RF-HOI: Recognize Human-Object Interaction with Radio Frequency Signals

T0 review · 4 major / 5 minor · reviewed 2026-08-04 · deepseek-v4-flash

Pith's one-line read This paper claims that a person's full interaction—the action and the object being acted on—can be recognized from radio signals alone, with no camera, by fusing mmWave radar and RFID data.

desk verdict Solid multimodal wireless-sensing systems paper undercut by an overstated 'first RF-only HOI' claim and an unvalidated simulator assumption about static target tags. read the letter →

arxiv 2608.00289 v1 pith:BH32GU6W submitted 2026-07-31 cs.AI cs.HCcs.RO

classification cs.AIcs.HCcs.RO
keywords human-objectinteractionRFsensingmmWaveradarRFIDmodalityfusionwirelesssyntheticdataprivacy-preserving
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

RF-HOI is the paper's proposal that a person's interaction with an object—both the action and which object—can be determined without any camera, using only radio signals. The system combines two commercial sensing modalities: mmWave radar captures the body's motion, while RFID tags stuck on candidate objects report which tag is moving. A neural network fuses these streams and outputs action and target object in separate branches. The paper reports 96.97% average HOI accuracy across five users, three rooms, and 17 HOI categories, within 1.96 percentage points of a vision baseline. A physics-based simulator generates synthetic RF training data, and pre-training on it boosts accuracy when real-world data is scarce.

What carries the argument

The load-bearing signal is the RFID phase difference Δφ(t): each tag's phase change between consecutive frames indicates movement direction, with static non-target tags assigned Δφ=0 and moving targets using the hand's trajectory as a proxy for the object's trajectory. Around this, mmWave heatmaps encode range, Doppler, and angle; separate transformer blocks pool temporal features; the features are concatenated; and two cross-entropy branches output action and target index. A ray-tracing-based simulator for mmWave and a physical model for RFID phase generation produce synthetic pretraining data.

What would settle it

Take the trained system and test it on interactions where a non-target object moves at the same time (for example, a second person sets down a cup) or where the target moves differently from the hand (sliding a drawer, swinging a cabinet door, wiping a table). If object accuracy drops toward chance while action accuracy stays high, the reported 96.97% figure rests on the static-background assumption rather than on general RF-based object identification.

Watch

Extended reading notes

Core claim

The central claim is that fusing mmWave radar and RFID makes RF-only HOI recognition feasible and nearly competitive with cameras: radar supplies fine-grained motion features, RFID supplies per-object identity through phase differences between successive reads of each tag, and the fused features are decoded by two branches. The paper also claims the decoupled design lets the system generalize to unseen action-object pairs and even to unseen object categories when interaction dynamics are similar, and that synthetic data generated from 3D mesh motion capture shrinks the need for real-world collection.

Load-bearing premise

Object identification relies on the assumption that only the target object moves during an interaction, so all other tags can be treated as phase-static and the target's motion can be approximated by the user's hand; if a non-target object moves or an object moves without following the hand, the RFID cue that carries object identity is no longer reliable.

Editorial extensions

If this is right

  • Cameras can be replaced by RF sensors in privacy-sensitive or poorly lit spaces, with HOI accuracy within about two points of vision.
  • The decoupled action and object outputs enable compositional generalization: unseen action-object pairs can be recognized when both components have been seen in training.
  • Pre-training on synthetic RF data reduces the amount of real-world data needed for deployment, with average accuracy gains around 12 percentage points in low-data settings.
  • Downstream goal inference improves when the system provides full action-object tuples instead of actions or objects alone.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial extension: the static-non-target assumption suggests the reported accuracy is an upper bound for orderly single-user scenes; a fair benchmark with moving background objects would likely lower it.
  • Editorial extension: since two RFID antennas capture most of the accuracy gain, a leaner two-antenna deployment may be viable in practice despite the paper's default four-antenna setup.
  • Editorial extension: the same simulator-plus-fine-tuning recipe could transfer to other RF modalities or to multi-user HOI if phase signatures of multiple moving tags can be disentangled.
  • Editorial extension: environments that already RFID-tag retail products or medicine containers are the most immediately deployable niche for this approach.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. RF-HOI is an RF-only human-object interaction recognition system that fuses mmWave radar and RFID tags on objects. The paper contributes a modality-fusion neural network with decoupled action and object branches, a physics-based simulator that synthesizes mmWave and RFID data from the TRUMANS mesh dataset, and a real-world dataset of 3,615 samples across 78 setups, 5 users, and 3 environments. The main reported result is 96.97% HOI accuracy with a 1.96% gap to a vision baseline, with additional experiments on fusion ablations, synthetic-data benefits, unseen HOI categories, occlusion robustness, dynamic multipath, latency, and a goal-inference case study.

Significance. If the central claim holds, RF-HOI would be a meaningful privacy-preserving alternative to camera-based HOI recognition, with concrete strengths: a substantial real-world dataset collected on commercial hardware, systematic ablations (fusion, antennas, encoders, loss weight, synthetic data scale), a synthetic pretraining pipeline that transfers to real data, and an evaluation spanning unseen object categories, occlusion, and dynamic environments. The paper is an empirical systems contribution rather than a mathematical derivation, and its internal comparisons are largely consistent. The main risk is external validity: the RFID synthesis assumptions and the absence of a comparison to the closest prior RF-only HOI system leave the generality of the 96.97% accuracy claim under-supported.

major comments (4)
  1. [§3.4.3, §5.6.1] The synthetic RFID branch is generated under two strong assumptions: non-target tags have Δφ=0, and target-object motion is approximated by hand motion. This produces a clean phase-contrast prior for target identification. The paper's own occlusion experiment (Fig. 26) shows non-target tags behind the user exhibit 'noticeable phase fluctuations,' and the taxonomy includes <wipe, table>, where the target tag is static and thus also has Δφ≈0. No quantitative sim-to-real comparison of Δφ statistics is provided beyond the qualitative Fig. 10, and no per-category object accuracy or confusion matrix is reported for static-target categories (wipe, open, close). Without such evidence, the reported accuracy may reflect the simulator's 'largest phase excursion = target' prior rather than a general capability. Please report object accuracy per HOI category and a quantitative sim-to-real Δφ distribu
  2. [§5.2, §6.2] The abstract and §5.2 claim 'the first framework that only uses radio frequency signals for HOI recognition' and that RF-HOI 'outperforms all baselines.' However, no quantitative comparison is made to RF-Diary [27], which is cited in §6.2 as an RF-based HOI system combining radar with object-location maps. The only empirical baselines are internal ablations (w/o fusion, RFID-only, mmWave-only) and a vision model. The 'first' claim and the practical advantage over prior RF-only HOI work require either a head-to-head comparison or an explicit, justified reason why RF-Diary is not a valid baseline in this setting.
  3. [§5.5.4] The hyperparameter α in Eq. (1) is selected by sweeping on the test set: Fig. 23 reports HOI accuracy versus α evaluated on real-world test samples, and α=0.5 is then used for all reported results, including the 96.97% headline number. This is test-set leakage and can inflate the reported accuracies. Use a held-out validation split for hyperparameter selection, or report results with nested cross-validation.
  4. [§5.4.2, Fig. 17] The unseen-object generalization evaluation is based on only 145 real samples, five new objects, and three orientations, with no per-object breakdown, confidence intervals, or raw counts. The enhanced synthetic pretraining includes the unseen object names, so Fig. 17 isolates category expansion but does not fully establish generalization to new physical objects or interaction dynamics. Given that unseen-object generalization is a claimed contribution, please report per-object accuracy and variability.
minor comments (5)
  1. [Fig. 10] The real/synthetic Δφ comparison is qualitative and appears to show a single example. Please state the number of samples, report distribution statistics (e.g., mean/percentiles), and label both panels consistently.
  2. [Abstract, §5.2] The abstract reports 96.97% as an average accuracy without noting that this is the 100%-real-training-data condition. State the conditioning explicitly.
  3. [Fig. 29] Typo: 'Accutacy' should be 'Accuracy'.
  4. [§5.3] The vision baseline uses RGB video and a ResNet backbone; it would be helpful to state whether it also receives object-name embeddings, given that RF-HOI's object branch uses spaCy name embeddings. A one-sentence clarification would improve comparability.
  5. [§2.1, Table 1] The table row labels for mmWave 'Tags on objects' are ambiguous: the second column appears to contain both device-assistance and action-capture text. Reformat for clarity.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: RF-HOI's headline accuracy is a measured held-out result on real-world data; simulator assumptions and the test-set alpha sweep are external-validity concerns, not input-output equivalences.

full rationale

RF-HOI is an empirical systems paper rather than a formal derivation, and its central claim (96.97% HOI accuracy, 1.96% gap to vision) is a measured result on a held-out real-world test split (70% of 3,615 samples), not a consequence of the simulator or of any equation. The synthetic-data pretraining is evaluated by fine-tuning on a disjoint real training subset and testing on real samples, so the test labels do not reduce to the simulator's generation rule. The RFID simulator's simplifying assumptions—'For static objects, we set Δφ=0' and using hand location as a proxy for object location (§3.4.3)—are potential external-validity limitations, and the paper itself acknowledges related limits (§2.1 on RFID-only approaches failing when objects do not move; §7.1 on category coverage being constrained by TRUMANS). These affect generalization claims, but they do not make the reported accuracy analytically derivable from the inputs. The choice of α=0.5 in §5.5.4 is a hyperparameter sweep performed on the test set, which is a data-leakage/overfitting concern rather than a circular reduction; the same is true of the qualitative sim-to-real phase comparison in Fig. 10. Self-citations (e.g., [67], [97], [135]) are used for background, design motivation, or downstream baseline construction, not as load-bearing justification of the core recognition result, and no invoked uniqueness theorem forces the architecture. External anchors (vision baseline, XRF-55, CubeLearn, unseen-object generalization) provide independent calibration. No equation or fitted parameter is renamed as a prediction, so no circularity step can be exhibited.

Assumptions & free parameters 5 free parameters · 5 assumptions · 0 invented entities

All free parameters were chosen by hand or selected from the paper's own experiments; none are derived from external theory. The main domain assumptions concern the RFID simulator and the transferability of min-max normalized ray-traced heatmaps. No new physical entities are introduced.

free parameters (5)
  • α (loss weight) = 0.5
    Balances action and target losses (Eq. 1); chosen after a sweep in §5.5.4 on the real-world test set, not on a strictly held-out validation set.
  • Learning rates = 1e-5 pre-train, 1e-6 fine-tune
    Hand-set in §4.4; affect sim-to-real transfer but are conventional choices.
  • Architecture dimensions (N_map, N_M_dim, N_R_dim, heads) = 256, 1024, 128, 2
    Hand-chosen model capacity in §4.4; no sensitivity analysis reported.
  • Simulator augmentation parameters = not reported
    Packet-loss drop rate, Gaussian noise std, and normalization details for synthetic RFID data are described qualitatively in §3.4.3 but values are not given; they affect the sim-to-real gap.
  • TRUMANS filtering threshold = ~20% removed
    Samples too short or too long are removed before synthesis (§4.3); the cutoffs are not specified.
assumptions (5)
  • domain assumption Only the target object moves during an interaction; non-target tags produce Δφ=0
    Used for synthetic RFID data (§3.4.3) and assumed in preprocessing; if non-target objects move, RFID target identification loses its clean contrast.
  • domain assumption The hand trajectory is an adequate proxy for target-object trajectory
    Used to synthesize Δφ from TRUMANS hand coordinates (§3.4.3, Fig. 9); fails for interactions where the object moves differently from the hand (e.g., hinged doors, sliding objects) or does not move (wipe).
  • domain assumption RFID phase difference Δφ is a reliable, multipath-robust kinematic feature and can be interpolated/denoised without losing discriminative information
    Stated in §3.2.2; RFID phase is known to be sensitive to multipath and reader frequency hopping (calibrated per [121]).
  • domain assumption Ray-traced, min-max-normalized mmWave heatmaps transfer to real radar without RCS modeling
    Stated in §3.4.3; relies on normalization to hide absolute intensity mismatch. No quantitative domain gap metric is provided.
  • domain assumption spaCy word embeddings capture object semantics sufficiently for unseen-object generalization
    Object name embeddings are a model input (§3.2.2, §3.5.2); unseen-object performance partly measures semantic generalization of the embedding space, not RF sensing of new objects.

how reviews work

0 comments
Cite this review

Pith. "Pith review of RF-HOI: Recognize Human-Object Interaction with Radio Frequency Signals." pith.science (2026). https://pith.science/paper/BH32GU6W

@misc{pith2026260800289,
  author       = {Pith},
  title        = {Pith review of: RF-HOI: Recognize Human-Object Interaction with Radio Frequency Signals},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/BH32GU6W}},
  note         = {Machine review of arXiv:2608.00289}
}
read the original abstract

Recognizing Human-Object Interactions (HOI) is essential for intelligent systems, underpinning applications in virtual and augmented reality, embodied AI, and assistive robotics. However, vision-based HOI methods face challenges in privacy concerns and poor light conditions. In this work, we introduce RF-HOI, the first framework that only uses radio frequency (RF) signals for HOI recognition. A key challenge of RF-HOI is that single-modality RF sensing is insufficient to recognize both actions and the objects being interacted with. RF-HOI addresses this through a novel modality fusion that combines mmWave radar and RFID, enabling simultaneous action recognition and target identification. Another challenge is limited training data across diverse setups, which impairs the generalizability of the recognition model. To overcome this, we develop a simulator that synthesizes multimodal RF data for diverse HOIs at scale, allowing us to fine-tune with only a small amount of real-world data. Experiment results show that RF-HOI outperforms all baselines, approaching vision model performance, and that our diverse synthetic training data can significantly boost our system's performance on real-world scenarios. These results highlight the potential of multimodal RF sensing for robust and privacy-preserving HOI recognition as well as the effectiveness of our RF data synthesis.

Figures

Figures reproduced from arXiv: 2608.00289 by the authors.

Figure 1
Figure 1. Using RF signals, RF-HOI can understand the user better in a privacy-preserving way by recognizing the user’s interaction with objects. Compared with HAR, HOI recognition considers both human actions and target objects. As is shown in the figure, one of the applications of HOI recognition are the goal inference task. The agent can infer a more appropriate goal of the user based on HOI sequences, and then provide ass… view at source ↗
Figure 2
Figure 2. Overview of RF-HOI Framework [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Photos and corresponding RFID signal phases of the tag attached to the target object in two different interactions. [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (19 more)
Figure 4
Figure 4. Figure 4: Factors that may affect model performance. [PITH_FULL_IMAGE:figures/full_fig_p006_4.png]
Figure 6
Figure 6. Figure 6: Examples for mmWave heatmaps. samples, thereby reducing data collection effort. We detail the simulator design and the integration of synthetic data in Sec. 3.4. 3 METHODS 3.1 Overview As shown in [PITH_FULL_IMAGE:figures/full_fig_p007_6.png]
Figure 7
Figure 7. Figure 7: Structure of the modality fusion network. We carefully designs spatial and temporal feature extraction for each [PITH_FULL_IMAGE:figures/full_fig_p008_7.png]
Figure 8
Figure 8. Figure 8: Overview of the RF signal simulator. Our simulator can synthesize both mmWave and RFID data given a mesh sample. [PITH_FULL_IMAGE:figures/full_fig_p010_8.png]
Figure 9
Figure 9. Figure 9: The coordinates in three axes of the object and the hand in a ‘pick up’ action. [PITH_FULL_IMAGE:figures/full_fig_p011_9.png]
Figure 10
Figure 10. Figure 10: Comparing normalized tag phase difference in a ‘pick up’ action. [PITH_FULL_IMAGE:figures/full_fig_p012_10.png]
Figure 11
Figure 11. Figure 11: Hardware platform and experiment setups. [PITH_FULL_IMAGE:figures/full_fig_p014_11.png]
Figure 12
Figure 12. Figure 12: Real-world sample durations of 17 HOI Categories. The proposed modality fusion network is designed to support [PITH_FULL_IMAGE:figures/full_fig_p015_12.png]
Figure 13
Figure 13. Figure 13: Effectiveness of modality fusion. when using all real-world training samples, RF-HOI reaches 96.97% HOI accuracy, which is 19.87%, 36.07%, and 80.06% higher than RF-HOI w/o Fusion, RFID-only, and mmWave-only, respectively. To further elucidate the advantages of RF-HOI…
Figure 14
Figure 14. Figure 14: Compare RF-HOI with the vision model in terms of HOI accuracy. 1 2 3 4 5 Number of Unseen Categories 40 50 60 70 80 90 100 HOI Accuracy (%) RF-HOI w/o Synthetic Data RF-HOI [PITH_FULL_IMAGE:figures/full_fig_p018_14.png]
Figure 16
Figure 16. Figure 16: Five unseen objects: jar, lunchbox, bowl, cup, and eyeglass case. Main (real) Unseen objects (real) Evaluation setting 0 20 40 60 80 100 HOI Accuracy (%) Main synthetic pretraining Enhanced synthetic pretraining [PITH_FULL_IMAGE:figures/full_fig_p018_16.png]
Figure 18
Figure 18. Figure 18: Impact of training set size. times to reduce sampling bias [PITH_FULL_IMAGE:figures/full_fig_p019_18.png]
Figure 19
Figure 19. Figure 19: Comparison between different numbers of unseen setups. [PITH_FULL_IMAGE:figures/full_fig_p020_19.png]
Figure 21
Figure 21. Figure 21: Comparison for different mmWave signal encoders. 1 2 3 4 Number of Antennas 50 60 70 80 90 100 HOI Accuracy (%) [PITH_FULL_IMAGE:figures/full_fig_p021_21.png]
Figure 23
Figure 23. Figure 23: Impact of weight coefficient 𝛼 of loss func￾tion. 20 40 60 80 100 % of Synthetic Data 50 60 70 80 90 100 HOI Accuracy (%) [PITH_FULL_IMAGE:figures/full_fig_p022_23.png]
Figure 25
Figure 25. Figure 25: Impact of hand occlusion on RFID RSSI. 0 8 16 24 32 Frames −2 −1 0 1 2 3 Phase (rad.) Phone (target) Pen (occluded) Book (occluded) Others [PITH_FULL_IMAGE:figures/full_fig_p023_25.png]
Figure 27
Figure 27. Figure 27: Impact of human occlusion. Action Accuracy Object Accuracy HOI Accuracy Category of Accuracy 50 60 70 80 90 100 Accuracy (%) w/o Dynamic Interference with Dynamic Interference [PITH_FULL_IMAGE:figures/full_fig_p023_27.png]
Figure 29
Figure 29. Figure 29: Impact of fine-tuning du￾ration. L40S A100 RTX 4050 GPU Type 0.0 0.5 1.0 1.5 2.0 Preprocessing Latency (s) 0 50 100 150 200 Inference Latency (ms) [PITH_FULL_IMAGE:figures/full_fig_p025_29.png]
Figure 32
Figure 32. Figure 32: Streaming pipeline of RF-HOI for real-time processing. To further assess the system’s responsiveness in an online setting, we analyze the frame-level preprocessing latency of each sensing modality in [PITH_FULL_IMAGE:figures/full_fig_p025_32.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

147 extracted references · 2 canonical work pages

  1. [27]

    Lijie Fan, Tianhong Li, Yuan Yuan, and Dina Katabi. 2020. In-home daily-life captioning using radio signals. InEuropean Conference on Computer Vision. Springer, Springer, Berlin, Germany, 105–123

  2. [1]

    Aakriti Adhikari, Hem Regmi, Sanjib Sur, and Srihari Nelakuditi. 2022. Mishape: Accurate human silhouettes and body joints from commodity millimeter-wave devices.Proc. ACM Interact. Mob. Wearable Ubiquitous Technol.6, 3 (2022), 1–31

  3. [2]

    Karan Ahuja, Yue Jiang, Mayank Goel, and Chris Harrison. 2021. Vid2doppler: Synthesizing doppler radar data from videos for training privacy-preserving activity recognition. InProceedings of the 2021 CHI Conference on Human Factors in Computing Systems. Association for Computing Machinery, New York, NY, USA, 1–10

  4. [3]

    Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton. 2016. Layer normalization.arXiv preprint arXiv:1607.064501, 1 (2016), 1. Proc. ACM Interact. Mob. Wearable Ubiquitous Technol., Vol. 10, No. 3, Article 159. Publication date: September 2026. 159:30•Wang et al

  5. [4]

    Kang Min Bae, Namjo Ahn, Yoon Chae, Parth Pathak, Sung-Min Sohn, and Song Min Kim. 2022. OmniScatter: extreme sensitivity mmWave backscattering using commodity FMCW radar. InThe 20th Annual International Conference on Mobile Systems, Applications and Services. Association for Computing Machinery, New York, NY, USA, 316–329

  6. [5]

    Kang Min Bae, Hankyeol Moon, and Song Min Kim. 2024. SuperSight: Sub-cm NLOS Localization for mmWave Backscatter. InThe 22nd Annual International Conference on Mobile Systems, Applications and Services. Association for Computing Machinery, New York, NY, USA, 278–291

  7. [6]

    Bluesight. 2025. RFID Medication Management and Tracking System. https://bluesight.com/rfid-medication-management/

  8. [7]

    Yanling Bu, Lei Xie, Yinyin Gong, Chuyu Wang, Lei Yang, Jia Liu, and Sanglu Lu. 2020. RF-dial: Rigid motion tracking and touch gesture detection for interaction via RFID tags.IEEE Transactions on Mobile Computing21, 3 (2020), 1061–1080

Show all 147 references
  1. [8]

    Ryan Canales and Sophie Jörg. 2020. Performance is not everything: Audio feedback preferred over visual feedback for grasping task in virtual reality. InProceedings of the 13th ACM SIGGRAPH Conference on Motion, Interaction and Games. Association for Computing Machinery, New Y...

  2. [9]

    Qiming Cao, Hongfei Xue, Tianci Liu, Xingchen Wang, Haoyu Wang, Xincheng Zhang, and Lu Su. 2024. mmCLIP: Boosting mmWave- based Zero-shot HAR via Signal-Text Alignment. InThe 22nd ACM Conference on Embedded Networked Sensor Systems. Association for Computing Machinery, New Yor...

  3. [10]

    Chainway. 2026. MC51 5G Built-in RFID Reader. https://www.chainway.net/Products/Info/172

  4. [11]

    Chainway. 2026. MR20 Wearable RFID Reader. https://www.chainway.net/mobile/Products/Info/98

  5. [12]

    Yu-Wei Chao, Zhan Wang, Yugeng He, Jiaxuan Wang, and Jia Deng. 2015. Hico: A benchmark for recognizing human-object interactions in images. InIEEE International Conference on Computer Vision. IEEE, Piscataway, NJ, USA, 1017–1025

  6. [13]

    Jian Chen, Lan Du, Guanbo Guo, Linwei Yin, and Di Wei. 2022. Target-attentional CNN for radar automatic target recognition with HRRP.Signal Processing196 (2022), 108497

  7. [14]

    Xingyu Chen and Xinyu Zhang. 2023. RF genesis: Zero-shot generalization of mmwave sensing through simulation-based data synthesis and generative diffusion models. InThe 21st ACM Conference on Embedded Networked Sensor Systems. Association for Computing Machinery, New York, NY,...

  8. [15]

    Yue Chen, Yalong Bai, Wei Zhang, and Tao Mei. 2019. Destruction and construction learning for fine-grained image recognition. In IEEE/CVF Conference on Computer Vision and Pattern Recognition. IEEE, Piscataway, NJ, USA, 5157–5166

  9. [16]

    Guoxuan Chi, Zheng Yang, Chenshu Wu, Jingao Xu, Yuchong Gao, Yunhao Liu, and Tony Xiao Han. 2024. RF-diffusion: Radio signal generation via time-frequency diffusion. InThe 30th Annual International Conference on Mobile Computing and Networking. Association for Computing Machin...

  10. [17]

    Han Cui, Shu Zhong, Jiacheng Wu, Zichao Shen, Naim Dahnoun, and Yiren Zhao. 2023. Milipoint: A point cloud dataset for mmwave radar.Advances in Neural Information Processing Systems36 (2023), 62713–62726

  11. [18]

    Shenghong Dai, Shiqi Jiang, Yifan Yang, Ting Cao, Mo Li, Suman Banerjee, and Lili Qiu. 2025. Babel: A scalable pre-trained model for multi-modal sensing via expandable modality alignment. InThe 23rd ACM Conference on Embedded Networked Sensor Systems. Association for Computing...

  12. [19]

    Decathlon. 2025. Product traceability and RFID technology at DECATHLON. https://sustainability.decathlon.com/product-traceability- and-rfid-technology-at-decathlon

  13. [20]

    DeepSeek AI. 2025. DeepSeek 3.2 model page. https://chat.deepseek.com/

  14. [21]

    Dell. 2025. Dell G15 5535 laptop. https://www.dell.com/en-us/shop/dell-laptops/g15-gaming-laptop/spd/g-series-15-5535-laptop

  15. [22]

    Kaikai Deng, Dong Zhao, Qiaoyue Han, Zihan Zhang, Shuyue Wang, Anfu Zhou, and Huadong Ma. 2023. Midas: Generating mmWave radar data from videos for training pervasive and privacy-preserving human sensing tasks.Proc. ACM Interact. Mob. Wearable Ubiquitous Technol.7, 1 (2023), 1–26

  16. [23]

    Kaikai Deng, Dong Zhao, Zihan Zhang, Shuyue Wang, Wenxin Zheng, and Huadong Ma. 2023. Midas++: generating training data of mmWave radars from videos for privacy-preserving human sensing with mobility.IEEE Transactions on Mobile Computing23, 6 (2023), 6650–6666

  17. [24]

    Christian Diller and Angela Dai. 2024. Cg-hoi: Contact-guided 3d human-object interaction generation. InIEEE/CVF Conference on Computer Vision and Pattern Recognition. IEEE, Piscataway, NJ, USA, 19888–19901

  18. [25]

    e tag. 2025. RFID In Smart Kitchens: Enhancing Efficiency and Reducing Waste. https://e-tagrfid.com/rfid-in-smart-kitchens-enhancing- efficiency-reducing-waste

  19. [26]

    Junqiao Fan, Haocong Rao, Jiarui Zhang, Jianfei Yang, and Lihua Xie. 2026. MMPred: Radar-Based Human Motion Prediction in the Dark. InProceedings of the AAAI conference on artificial intelligence, Vol. 1. AAAI Press, Palo Alto, California, USA, 1

  20. [28]

    Chao Feng, Jiashen Chen, Shuo Liang, Xiaopeng Peng, Baizhou Yang, Xuan Wang, Zexuan Huang, Xianjia Meng, and Xiaojiang Chen

  21. [29]

    Yixu Feng, Shuo Hou, Haotian Lin, Yu Zhu, Peng Wu, Wei Dong, Jinqiu Sun, Qingsen Yan, and Yanning Zhang. 2024. Difflight: integrating content and detail for low-light image enhancement. InIEEE/CVF Conference on Computer Vision and Pattern Recognition. IEEE, Piscataway, NJ, USA...

  22. [30]

    Chuhan Gao, Yilong Li, and Xinyu Zhang. 2019. Livetag: Sensing human-object interaction through passive chipless wi-fi tags. GetMobile: Mobile Computing and Communications22, 3 (2019), 32–35

  23. [31]

    Georgia Gkioxari, Ross Girshick, Piotr Dollár, and Kaiming He. 2018. Detecting and recognizing human-object interactions. InIEEE/CVF Conference on Computer Vision and Pattern Recognition. IEEE, Piscataway, NJ, USA, 8359–8367

  24. [32]

    Google DeepMind. 2025. Gemini 2.5 Flash model page. https://gemini.google.com/app

  25. [33]

    Abhinav Gupta, Aniruddha Kembhavi, and Larry S Davis. 2009. Observing human-object interactions: Using spatial and functional compatibility for recognition.IEEE Transactions on Pattern Analysis and Machine Intelligence31, 10 (2009), 1775–1789

  26. [34]

    Guorong He, Shaojie Chen, Dan Xu, Xiaojiang Chen, Yaxiong Xie, Xinhuai Wang, and Dingyi Fang. 2023. Fusang: Graph-inspired robust and accurate object recognition on commodity mmWave devices. InThe 21st Annual International Conference on Mobile Systems, Applications and Service...

  27. [35]

    Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep residual learning for image recognition. InIEEE/CVF Conference on Computer Vision and Pattern Recognition. IEEE, Piscataway, NJ, USA, 770–778

  28. [36]

    H&M. 2025. H&M Group: Taking steps to grow the customer experience through tech. https://hmgroup.com/our-stories/hm-group- taking-steps-to-grow-the-customer-experience-through-tech/

  29. [37]

    Sepp Hochreiter and Jürgen Schmidhuber. 1997. Long short-term memory.Neural computation9, 8 (1997), 1735–1780

  30. [38]

    Markus Höll, Markus Oberweger, Clemens Arth, and Vincent Lepetit. 2018. Efficient physics-based implementation for realistic hand-object interaction in virtual reality. InIEEE Conference on Virtual Reality and 3D User Interfaces. IEEE, IEEE, Piscataway, NJ, USA, 175–182

  31. [39]

    Matthew Honnibal, Ines Montani, Sofie Van Landeghem, Adriane Boyd, et al . 2020. spaCy: Industrial-strength Natural Language Processing in Python. https://spacy.io/

  32. [40]

    Han Hu, Jiayuan Gu, Zheng Zhang, Jifeng Dai, and Yichen Wei. 2018. Relation networks for object detection. InIEEE/CVF Conference on Computer Vision and Pattern Recognition. IEEE, Piscataway, NJ, USA, 3588–3597

  33. [41]

    Anna Huang, Dong Wang, Run Zhao, and Qian Zhang. 2019. Au-id: Automatic user identification and authentication through the motions captured from sequential human activities using rfid.Proc. ACM Interact. Mob. Wearable Ubiquitous Technol.3, 2 (2019), 1–26

  34. [42]

    Mingzhen Huang, Fu-Jen Chu, Bugra Tekin, Kevin J Liang, Haoyu Ma, Weiyao Wang, Xingyu Chen, Pierre Gleize, Hongfei Xue, Siwei Lyu, et al. 2025. HOIGPT: Learning Long-Sequence Hand-Object Interaction with Language Models. InIEEE/CVF Conference on Computer Vision and Pattern Rec...

  35. [43]

    Impinj. 2025. Impinj R420 reader. https://www.impinj.com/products/readers/impinj-speedway

  36. [44]

    Naveed Imran, Jian Zhang, Jehad Ali, Sana Hameed, Houbing Herbert Song, and Byeong-hee Roh. 2025. SwinLSTM-EmoRec: A Robust Dual-Modal Emotion Recognition Framework Combining mmWave Radar and Camera for IoT-Enabled Multimedia Applications.IEEE Internet of Things Journal1 (2025), 1

  37. [45]

    Texas Instruments. 2025. AWR1843AOPEVM. https://www.ti.com/tool/AWR1843AOPEVM

  38. [46]

    Sergey Ioffe and Christian Szegedy. 2015. Batch normalization: Accelerating deep network training by reducing internal covariate shift. InInternational Conference on Machine Learning. pmlr, PMLR, San Diego, CA, USA, 448–456

  39. [47]

    Cesar Iovescu and Sandeep Rao. 2020. The fundamentals of millimeter wave radar sensors. 7 pages

  40. [48]

    Rahul Jain, Jingyu Shi, Runlin Duan, Zhengzhe Zhu, Xun Qian, and Karthik Ramani. 2023. Ubi-TOUCH: Ubiquitous Tangible Object Utilization through Consistent Hand-object interaction in Augmented Reality. InProceedings of the 36th Annual ACM Symposium on User Interface Software a...

  41. [49]

    Nan Jiang, Zimo He, Zi Wang, Hongjie Li, Yixin Chen, Siyuan Huang, and Yixin Zhu. 2024. Autonomous character-scene interaction synthesis from text instruction. InSIGGRAPH Asia 2024 Conference Papers. Association for Computing Machinery, New York, NY, USA, 1–11

  42. [50]

    Nan Jiang, Tengyu Liu, Zhexuan Cao, Jieming Cui, Zhiyuan Zhang, Yixin Chen, He Wang, Yixin Zhu, and Siyuan Huang. 2023. Full-body articulated human-object interaction. InIEEE/CVF International Conference on Computer Vision. IEEE, Piscataway, NJ, USA, 9365–9376

  43. [51]

    Nan Jiang, Zhiyuan Zhang, Hongjie Li, Xiaoxuan Ma, Zan Wang, Yixin Chen, Tengyu Liu, Yixin Zhu, and Siyuan Huang. 2024. Scaling up dynamic human-scene interaction modeling. InIEEE/CVF Conference on Computer Vision and Pattern Recognition. IEEE, Piscataway, NJ, USA, 1737–1747

  44. [52]

    Haojian Jin, Zhijian Yang, Swarun Kumar, and Jason I Hong. 2018. Towards wearable everyday body-frame tracking using passive RFIDs.Proc. ACM Interact. Mob. Wearable Ubiquitous Technol.1, 4 (2018), 1–23

  45. [53]

    Vinay Joshi, Shengkai Xu, Qiming Cao, Yi Zhu, Pu Wang, and Hongfei Xue. 2024. Towards Robust mmWave-based Human Activity Recognition using Large Simulated Dataset for Model Pretraining. In2024 IEEE International Conference on Big Data (BigData). IEEE, IEEE, Piscataway, NJ, USA...

  46. [54]

    Belal Korany, Chitra R Karanam, Hong Cai, and Yasamin Mostofi. 2019. XModal-ID: Using WiFi for through-wall person identification from candidate video footage. InThe 25th Annual International Conference on Mobile Computing and Networking. Association for Computing Machinery, N...

  47. [55]

    Bo Lan, Pei Li, Jiaxi Yin, Yunpeng Song, Ge Wang, Han Ding, Jinsong Han, and Fei Wang. 2025. XRF V2: A Dataset for Action Summarization with Wi-Fi Signals, and IMUs in Phones, Watches, Earbuds, and Glasses.Proc. ACM Interact. Mob. Wearable Ubiquitous Technol.9, 3, Article 98 (...

  48. [56]

    Gierad Laput and Chris Harrison. 2019. Sensing fine-grained hand activity with smartwatches. InProceedings of the 2019 CHI Conference on Human Factors in Computing Systems. Association for Computing Machinery, New York, NY, USA, 1–13

  49. [57]

    Chi-Jung Lee, Jiaxin Li, Tianhong Catherine Yu, Ruidong Zhang, Vipin Gunda, François Guimbretière, and Cheng Zhang. 2025. Grab-n-Go: On-the-Go Microgesture Recognition with Objects in Hand.Proc. ACM Interact. Mob. Wearable Ubiquitous Technol.9, 3 (2025), 1–27

  50. [58]

    Chi-Jung Lee, Ruidong Zhang, Devansh Agarwal, Tianhong Catherine Yu, Vipin Gunda, Oliver Lopez, James Kim, Sicheng Yin, Boao Dong, Ke Li, et al. 2024. Echowrist: Continuous hand pose tracking and hand-object interaction recognition using low-power active acoustic sensing on a ...

  51. [59]

    Chenglong Li, Emmeric Tanghe, David Plets, Pieter Suanet, Jeroen Hoebeke, Eli De Poorter, and Wout Joseph. 2020. ReLoc: Hybrid RSSI-and phase-based relative UHF-RFID tag localization with COTS devices.IEEE Transactions on Instrumentation and Measurement 69, 10 (2020), 8613–8627

  52. [60]

    Hanchuan Li, Can Ye, and Alanson P Sample. 2015. IDSense: A human object interaction detection system based on passive UHF RFID. InProceedings of the 33rd Annual ACM Conference on Human Factors in Computing Systems. Association for Computing Machinery, New York, NY, USA, 2555–2564

  53. [61]

    Jiamu Li, Dongheng Zhang, Zhi Wu, Cong Yu, Yadong Li, Qi Chen, Yang Hu, Qibin Sun, and Yan Chen. 2024. SBRF: A fine-grained radar signal generator for human sensing.IEEE Transactions on Mobile Computing1-1, 1-1 (2024), 1–1

  54. [62]

    Rong Li, Tao Deng, Siwei Feng, Mingjie Sun, and Juncheng Jia. 2025. ConSense: Continually Sensing Human Activity with WiFi via Growing and Picking. InThe AAAI Conference on Artificial intelligence, Vol. 39. AAAI Press, Palo Alto, California, USA, 14292–14300

  55. [63]

    Xinyu Li, Yanyi Zhang, Ivan Marsic, Aleksandra Sarcevic, and Randall S Burd. 2016. Deep learning for rfid-based activity recognition. InProceedings of the 14th ACM Conference on Embedded Network Sensor Systems. Association for Computing Machinery, New York, NY, USA, 164–175

  56. [64]

    Zhengxiong Li, Baicheng Chen, Zhuolin Yang, Huining Li, Chenhan Xu, Xingyu Chen, Kun Wang, and Wenyao Xu. 2019. Ferrotag: A paper-based mmwave-scannable tagging infrastructure. InThe 17th Conference on Embedded Networked Sensor Systems. Association for Computing Machinery, New...

  57. [65]

    Zhe Li, Jiakun Li, Mingqi Gao, Jinyu Yang, Wei Wang, and Feng Zheng. 2026. ContA-HOI: Towards Physically Plausible Human-Object Interaction Generation via Contact-Aware Modeling. InThirteenth International Conference on 3D Vision. IEEE, Piscataway, NJ, USA, 1–1

  58. [66]

    Zhuolong Li, Xingao Li, Changxing Ding, and Xiangmin Xu. 2024. Disentangled pre-training for human-object interaction detection. InIEEE/CVF Conference on Computer Vision and Pattern Recognition. IEEE, Piscataway, NJ, USA, 28191–28201

  59. [67]

    2023.{RF-Chord}: Towards deployable{RFID} localization system for logistic networks

    Bo Liang, Purui Wang, Renjie Zhao, Heyu Guo, Pengyu Zhang, Junchen Guo, Shunmin Zhu, Hongqiang Harry Liu, Xinyu Zhang, and Chenren Xu. 2023.{RF-Chord}: Towards deployable{RFID} localization system for logistic networks. In20th USENIX Symposium on Networked Systems Design and I...

  60. [68]

    Yue Liao, Si Liu, Fei Wang, Yanjie Chen, Chen Qian, and Jiashi Feng. 2020. Ppdm: Parallel point detection and matching for real-time human-object interaction detection. InIEEE/CVF Conference on Computer Vision and Pattern Recognition. IEEE, Piscataway, NJ, USA, 482–490

  61. [69]

    Mingzhi Lin, Teng Huang, Han Ding, Cui Zhao, Fei Wang, Ge Wang, and Wei Xi. 2026. Active Domain Adaptation for Mmwave-Based HAR Via Rényi Entropy-Based Uncertainty Estimation.IEEE Transactions on Mobile Computing1-1, 1-1 (2026), 1–1

  62. [70]

    Haipeng Liu, Kening Cui, Kaiyuan Hu, Yuheng Wang, Anfu Zhou, Liang Liu, and Huadong Ma. 2022. mTransSee: Enabling environment- independent mmWave sensing based gesture recognition via transfer learning.Proc. ACM Interact. Mob. Wearable Ubiquitous Technol.6, 1 (2022), 1–28

  63. [71]

    Ting Liu, Yao Zhao, Yunchao Wei, Yufeng Zhao, and Shikui Wei. 2019. Concealed object detection for activate millimeter wave image. IEEE Transactions on Industrial Electronics66, 12 (2019), 9909–9917

  64. [72]

    Xiulong Liu, Dongdong Liu, Jiuwu Zhang, Tao Gu, and Keqiu Li. 2021. RFID and camera fusion for recognition of human-object interactions. InThe 27th Annual International Conference on Mobile Computing and Networking. Association for Computing Machinery, New York, NY, USA, 296–308

  65. [73]

    Logitech. 2025. Logitech Spot: Presence and Environmental Sensor. https://www.logitech.com/en-us/products/video-conferencing/ accessories/spot-sensor.950-000107.html

  66. [74]

    Ilya Loshchilov and Frank Hutter. 2017. Decoupled weight decay regularization.arXiv preprint arXiv:1711.051011-1, 1-1 (2017), 1–1. Proc. ACM Interact. Mob. Wearable Ubiquitous Technol., Vol. 10, No. 3, Article 159. Publication date: September 2026. RF-HOI: Recognize Human-Obje...

  67. [75]

    Haofan Lu, Mohammad Mazaheri, Reza Rezvani, and Omid Abari. 2023. A millimeter wave backscatter network for two-way communication and localization. InProceedings of the ACM SIGCOMM 2023 Conference. Association for Computing Machinery, New York, NY, USA, 49–61

  68. [76]

    Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. 2019. Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks.Advances in neural information processing systems32 (2019), 1

  69. [77]

    Jiaxin Lu, Chun-Hao Paul Huang, Uttaran Bhattacharya, Qixing Huang, and Yi Zhou. 2025. HUMOTO: A 4D Dataset of Mocap Human Object Interactions.arXiv preprint arXiv:2504.104141-1, 1-1 (2025), 1–1

  70. [78]

    Tianlun Luo, Steven Guan, Rui Yang, and Jeremy Smith. 2023. From detection to understanding: A survey on representation learning for human-object interaction.Neurocomputing543 (2023), 126243

  71. [79]

    Andrew L Maas, Awni Y Hannun, Andrew Y Ng, et al. 2013. Rectifier nonlinearities improve neural network acoustic models. InProc. icml, Vol. 30. Atlanta, GA, PMLR, San Diego, CA, USA, 3

  72. [80]

    Valerio Magnago, Luigi Palopoli, Alice Buffi, Bernardo Tellini, Andrea Motroni, Paolo Nepa, David Macii, and Daniele Fontanelli. 2019. Ranging-free UHF-RFID robot positioning through phase measurements of passive tags.IEEE Transactions on Instrumentation and Measurement69, 5 (...

  73. [81]

    Esteve Valls Mascaro, Daniel Sliwowski, and Dongheui Lee. 2023. HOI4ABOT: Human-Object Interaction Anticipation for Human Intention Reading Assistive roBOTs. In7th Annual Conference on Robot Learning, Vol. 1-1. PMLR, San Diego, CA, USA, 1–1

  74. [82]

    Fairview Microwave. 2025. FM51FP1006 high performance antenna. https://www.fairviewmicrowave.com/product/antennas/directional- antennas/panel-antennas/wide-antenna-9-dbi-gain-n-fm51fp1006.html

  75. [83]

    Mahathir Monjur and Shahriar Nirjon. 2026. mmWEAVER: Environment-Specific mmWave Signal Synthesis from a Photo and Activity Description. InProceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision. IEEE, Piscataway, NJ, USA, 1875–1884

  76. [84]

    Lingzhou Mu, Wang Qiang, Fan Jiang, Mengchao Wang, Mu Xu, and Kai Zhang. 2026. FantasyHSI: Video-Generation-Centric 4D Human Synthesis in Any Scene Through a Graph-Based Multi-Agent Framework. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 40. AAAI Press...

  77. [85]

    RMS Omega. 2026. Track Critical Healthcare Supplies with RFID Technology. https://rmsomega.com/healthcare/track/rfid-in- healthcare/

  78. [86]

    Omnicell. 2025. Medical RFID Cabinet. https://www.omnicell.me/medical-supply-automation/rfid-cabinet/

  79. [87]

    OpenAI. 2025. ChatGPT (GPT-5) model page. https://chat.openai.com/

  80. [88]

    ORBBEC. 2025. Femto Bolt. https://www.orbbec.com/products/tof-camera/femto-bolt/

  81. [89]

    ORBBEC. 2025. Python Bindings for Orbbec SDK. https://github.com/orbbec/pyorbbecsdk

  82. [90]

    OSSIYGAR. 2025. UHF RFID tag 9662 Long Distance Passive Alien H3 Adhesive Inlay Markable Lable. https://www.amazon.com/dp/ B0B1H7Z6KZ

  83. [91]

    A Paszke. 2019. Pytorch: An imperative style, high-performance deep learning library.arXiv preprint arXiv:1912.017031, 1 (2019), 1

  84. [92]

    Georgios Pavlakos, Vasileios Choutas, Nima Ghorbani, Timo Bolkart, Ahmed AA Osman, Dimitrios Tzionas, and Michael J Black

  85. [93]

    Mingtao Pei, Yunde Jia, and Song-Chun Zhu. 2011. Parsing video events with goal inference and intent prediction. InInternational Conference on Computer Vision. IEEE, IEEE, Piscataway, NJ, USA, 487–494

  86. [94]

    Xiaogang Peng, Yiming Xie, Zizhao Wu, Varun Jampani, Deqing Sun, and Huaizu Jiang. 2025. Hoi-diff: Text-driven synthesis of 3d human-object interactions using diffusion models. InIEEE/CVF Conference on Computer Vision and Pattern Recognition. IEEE, Piscataway, NJ, USA, 2878–2888

  87. [95]

    Ethan Perez, Florian Strub, Harm De Vries, Vincent Dumoulin, and Aaron Courville. 2018. Film: Visual reasoning with a general conditioning layer. InProceedings of the AAAI conference on artificial intelligence, Vol. 32. AAAI Press, Palo Alto, California, USA, 1

  88. [96]

    Swadhin Pradhan, Eugene Chai, Karthikeyan Sundaresan, Lili Qiu, Mohammad A Khojastepour, and Sampath Rangarajan. 2017. RIO: A pervasive RFID-based touch gesture interface. InThe 23rd Annual International Conference on Mobile Computing and Networking. Association for Computing ...

  89. [97]

    Xavier Puig, Tianmin Shu, Joshua B Tenenbaum, and Antonio Torralba. 2023. Nopa: Neurally-guided online probabilistic assistance for building socially intelligent home assistants. InIEEE International Conference on Robotics and Automation. IEEE, IEEE, Piscataway, NJ, USA, 7628–7634

  90. [98]

    Siyuan Qi, Wenguan Wang, Baoxiong Jia, Jianbing Shen, and Song-Chun Zhu. 2018. Learning human-object interactions by graph parsing neural networks. InEuropean conference on computer vision. Springer, Berlin, Germany, 401–417

  91. [99]

    RFIDlabel. 2025. Upgrade Your Kitchen with Smart Microwaves and RFID Tagging in The Future! https://www.rfidlabel.com/upgrade- your-kitchen-source-smart-microwaves-with-rfid-tagging-today/

  92. [100]

    2005.Fundamentals of radar signal processing

    Mark A Richards et al. 2005.Fundamentals of radar signal processing. Vol. 1. Mcgraw-hill New York, New York, NY, USA. Proc. ACM Interact. Mob. Wearable Ubiquitous Technol., Vol. 10, No. 3, Article 159. Publication date: September 2026. 159:34•Wang et al

  93. [101]

    Manolis Savva, Abhishek Kadian, Oleksandr Maksymets, Yili Zhao, Erik Wijmans, Bhavana Jain, Julian Straub, Jia Liu, Vladlen Koltun, Jitendra Malik, et al. 2019. Habitat: A platform for embodied ai research. InIEEE/CVF International Conference on Computer Vision. IEEE, Piscataw...

  94. [102]

    Ralph Schmidt. 1986. Multiple emitter location and signal parameter estimation.IEEE Transactions on Antennas and Propagation34, 3 (1986), 276–280

  95. [103]

    Akash Deep Singh, Sandeep Singh Sandha, Luis Garcia, and Mani Srivastava. 2019. Radhar: Human activity recognition from point clouds generated through a millimeter-wave radar. InProceedings of the 3rd ACM Workshop on Millimeter-wave Networks and Sensing Systems. Association fo...

  96. [104]

    sllurp. 2025. A Python library to interface with RFID readers. https://github.com/sllurp/sllurp

  97. [105]

    Elahe Soltanaghaei, Akarsh Prabhakara, Artur Balanuta, Matthew Anderson, Jan M Rabaey, Swarun Kumar, and Anthony Rowe. 2021. Millimetro: mmWave retro-reflective tags for accurate, long range localization. InThe 27th Annual International Conference on Mobile Computing and Netwo...

  98. [106]

    Andrew Spielberg, Alanson Sample, Scott E Hudson, Jennifer Mankoff, and James McCann. 2016. RapID: A framework for fabricating low-latency interactive objects with RFID tags. InProceedings of the 2016 CHI Conference on Human Factors in Computing Systems. Association for Comput...

  99. [107]

    Ke Sun, Chunyu Xia, Xinyu Zhang, Hao Chen, and Charlie Jianzhong Zhang. 2024. Multimodal Daily-Life Logging in Free-living Environment Using Non-Visual Egocentric Sensors on a Smartphone.Proc. ACM Interact. Mob. Wearable Ubiquitous Technol.8, 1, Article 17 (March 2024), 32 pag...

  100. [108]

    Xiaoqi Sun, Yanwen Wang, Chenwei Zhang, Zheng Wang, Xiaokang Shi, and Yuanqing Zheng. 2024. LoDiHAR: A Low-Cost Distributed Human Activity Recognition System Based on RFID. InThe 21st Annual IEEE International Conference on Sensing, Communication, and Networking. IEEE, IEEE, P...

  101. [109]

    L Talbi. 2001. Simulation of indoor UHF propagation using numerical technique. InCanadian Conference on Electrical and Computer Engineering 2001. Conference Proceedings (Cat. No. 01TH8555), Vol. 2. IEEE, IEEE, Piscataway, NJ, USA, 1357–1362

  102. [110]

    UROVO. 2026. Enhancing Retail with Handheld RFID Readers. https://en.urovo.com/blog/RFID-handheld-readers/enhancing-retail- with-handheld-rfid-readers.html

  103. [111]

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need.Advances in Neural Information Processing Systems30 (2017), 1–1

  104. [112]

    Anandghan Waghmare, Youssef Ben Taleb, Ishan Chatterjee, Arjun Narendra, and Shwetak Patel. 2023. Z-ring: Single-point bio- impedance sensing for gesture, touch, object and user recognition. InProceedings of the 2023 CHI Conference on Human Factors in Computing Systems. Associ...

  105. [113]

    Fei Wang, Yizhe Lv, Mengdie Zhu, Han Ding, and Jinsong Han. 2024. Xrf55: A radio frequency dataset for human indoor action analysis. Proc. ACM Interact. Mob. Wearable Ubiquitous Technol.8, 1 (2024), 1–34

  106. [114]

    Hengbo Wang, Lvqing Yang, Yishu Qiu, Zhipeng Liu, Kai Li, Bo Yu, and Shihui Guo. 2026. RFGAT: Generative Adversarial Teacher for Cross-Domain RFID Activity Recognition. InICASSP 2026-2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, ...

  107. [115]

    Liang Wang, Tao Gu, Xianping Tao, and Jian Lu. 2016. Toward a wearable RFID system for real-time activity recognition using radio patterns.IEEE Transactions on Mobile Computing16, 1 (2016), 228–242

  108. [116]

    Wei Wang, Alex X Liu, Muhammad Shahzad, Kang Ling, and Sanglu Lu. 2015. Understanding and modeling of wifi signal based human activity recognition. InThe 21st Annual International Conference on Mobile Computing and Networking. Association for Computing Machinery, New York, NY,...

  109. [117]

    Yuheng Wang, Haipeng Liu, Kening Cui, Anfu Zhou, Wensheng Li, and Huadong Ma. 2021. m-activity: Accurate and real-time human activity recognition via millimeter wave radar. InIEEE International Conference on Acoustics, Speech and Signal Processing. IEEE, Association for Comput...

  110. [118]

    Yida Wang, Yu Lu, Yuxuan Zhou, Yifei Shen, Lili Qiu, Zeyuan Lai, Yi-Chao Chen, Hao Pan, Juntao Zhou, Dian Ding, et al . 2025. High-resolution mmWave Imaging using Metasurface and Diffusion. InThe 23rd Annual International Conference on Mobile Systems, Applications, and Service...

  111. [119]

    Yanwen Wang and Yuanqing Zheng. 2018. Modeling RFID signal reflection for contact-free activity recognition.Proc. ACM Interact. Mob. Wearable Ubiquitous Technol.2, 4 (2018), 1–22

  112. [120]

    Ziqi Wang and Shiwen Mao. 2025. Generative AI-Empowered RFID Sensing for 3D Human Pose Augmentation and Completion.IEEE Open Journal of the Communications Society1-1, 1-1 (2025), 1–1

  113. [121]

    Teng Wei and Xinyu Zhang. 2016. Gyro in the air: tracking 3d orientation of batteryless internet-of-things. InThe 22nd Annual International Conference on Mobile Computing and Networking. Association for Computing Machinery, New York, NY, USA, 55–68

  114. [122]

    Jiahao Wu, Yishu Qiu, Lvqing Yang, Shaoqin Shen, Kai Li, Bo Yu, and Shihui Guo. 2026. RF-OnlineHAR: A Real-Time RFID Activity Recognition Framework with Memory-Augmented Transformer.IEEE Internet of Things Journal1-1, 1-1 (2026), 1–1. Proc. ACM Interact. Mob. Wearable Ubiquito...

  115. [123]

    Xiaomi. 2025. Xiaomi Mijia Human Presence Sensor Detector for Smart Wireless Home. https://www.aliexpress.us/item/ 3256807256719195.html

  116. [124]

    Xin Xie, Xiulong Liu, Xibin Zhao, Weilian Xue, Bin Xiao, Heng Qi, Keqiu Li, and Jie Wu. 2019. Implementation of differential tag sampling for COTS RFID systems.IEEE Transactions on Mobile Computing19, 8 (2019), 1848–1861

  117. [125]

    Sirui Xu, Dongting Li, Yucheng Zhang, Xiyan Xu, Qi Long, Ziyin Wang, Yunzhi Lu, Shuchang Dong, Hezi Jiang, Akshat Gupta, et al

  118. [126]

    Yongchao Xu, Jiawei Liu, Sen Tao, Qiang Zhang, and Zheng-Jun Zha. 2025. HOIMamba: Efficient Mamba-based Disentangled Progressive Learning for HOI Detection. InThe AAAI Conference on Artificial Intelligence, Vol. 39. AAAI Press, Palo Alto, California, USA, 8987–8995

  119. [127]

    Hongfei Xue, Qiming Cao, Chenglin Miao, Yan Ju, Haochen Hu, Aidong Zhang, and Lu Su. 2023. Towards generalized mmwave-based human pose estimation through signal augmentation. InThe 29th Annual International Conference on Mobile Computing and Networking. Association for Computi...

  120. [128]

    Hongfei Xue, Yan Ju, Chenglin Miao, Yijiang Wang, Shiyang Wang, Aidong Zhang, and Lu Su. 2021. mmMesh: Towards 3D real-time dynamic human mesh construction using millimeter-wave. InThe 19th Annual International Conference on Mobile Systems, Applications, and Services. Associat...

  121. [129]

    Yifan Yan, Shuai Yang, Xiuzhen Guo, Xiangguang Wang, Wei Chow, Yuanchao Shu, and Shibo He. 2025. mmExpert: Integrating Large Language Models for Comprehensive mmWave Data Synthesis and Understanding. InProceedings of the Twenty-sixth International Symposium on Theory, Algorith...

  122. [130]

    Chao Yang, Lingxiao Wang, Xuyu Wang, and Shiwen Mao. 2022. Environment adaptive RFID-based 3D human pose tracking with a meta-learning approach.IEEE Journal of Radio Frequency Identification6 (2022), 413–425

  123. [131]

    Jie Yang, Bingliang Li, Ailing Zeng, Lei Zhang, and Ruimao Zhang. 2024. Open-world human-object interaction detection via multi-modal prompts. InIEEE/CVF Conference on Computer Vision and Pattern Recognition. IEEE, Piscataway, NJ, USA, 16954–16964

  124. [132]

    Zheng Yang, Yi Zhang, Kun Qian, and Chenshu Wu. 2023. {SLNet}: A Spectrogram Learning Neural Network for Deep Wireless Sensing. In20th USENIX Symposium on Networked Systems Design and Implementation. USENIX Association, Berkeley, CA, USA, 1221–1236

  125. [133]

    Bangpeng Yao and Li Fei-Fei. 2010. Modeling mutual context of object and human pose in human-object interaction activities. In2010 IEEE Computer Society Conference on Computer Vision and Pattern Recognition. IEEE, IEEE, Piscataway, NJ, USA, 17–24

  126. [134]

    Wei Yao, Yunlian Sun, Hongwen Zhang, Yebin Liu, and Jinhui Tang. 2026. Hosig: Full-body human-object-scene interaction generation with hierarchical scene perception. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 40. AAAI Press, Palo Alto, California, US...

  127. [135]

    Tenenbaum, and Tianmin Shu

    Lance Ying, Xinyi Li, Shivam Aarya, Yizirui Fang, Yifan Yin, Jason Xinyu Liu, Stefanie Tellex, Joshua B. Tenenbaum, and Tianmin Shu. 2025. Pragmatic Embodied Spoken Instruction Following in Human-Robot Collaboration with Theory of Mind.arXiv preprint arXiv:2409.108491-1, 1-1 (...

  128. [136]

    Bo Zhang, Zimeng Zhou, Boyu Jiang, and Rong Zheng. 2024. SUPER: Seated Upper Body Pose Estimation using mmWave Radars. In IEEE/ACM 9th International Conference on Internet-of-Things Design and Implementation. IEEE, Association for Computing Machinery, New York, NY, USA, 181–191

  129. [137]

    Feng Zhang, Chenshu Wu, Beibei Wang, and KJ Ray Liu. 2020. mmEye: Super-resolution millimeter wave imaging.IEEE Internet of Things Journal8, 8 (2020), 6995–7008

  130. [138]

    Juze Zhang, Jingyan Zhang, Zining Song, Zhanhe Shi, Chengfeng Zhao, Ye Shi, Jingyi Yu, Lan Xu, and Jingya Wang. 2024. HOI-Mˆ 3: Capture Multiple Humans and Objects Interaction within Contextual Environment. InProceedings of the IEEE/CVF Conference on Computer Vision and Patter...

  131. [139]

    Qian Zhang, Dong Wang, Run Zhao, Yufeng Deng, and Yinggang Yu. 2019. Shopeye: fusing rfid and smartwatch for multi-relation excavation in physical stores. InThe 24th International Conference on Intelligent User Interfaces. Association for Computing Machinery, New York, NY, USA, 86–95

  132. [140]

    Yi Zhang, Yue Zheng, Kun Qian, Guidong Zhang, Yunhao Liu, Chenshu Wu, and Zheng Yang. 2021. Widar3. 0: Zero-effort cross-domain gesture recognition with Wi-Fi.IEEE Transactions on Pattern Analysis and Machine Intelligence44, 11 (2021), 8671–8688

  133. [141]

    Zhining Zhang, Chuanyang Jin, Mung Yao Jia, and Tianmin Shu. 2025. Autotom: Automated bayesian inverse planning and model discovery for open-ended theory of mind.arXiv preprint arXiv:2502.156761-1, 1-1 (2025), 1–1

  134. [142]

    Peijun Zhao, Chris Xiaoxuan Lu, Bing Wang, Niki Trigoni, and Andrew Markham. 2023. Cubelearn: End-to-end learning for human motion recognition from raw mmwave radar signals.IEEE Internet of Things Journal10, 12 (2023), 10236–10249

  135. [143]

    Ruoyu Zhao, Yushu Zhang, Tao Wang, Wenying Wen, Yong Xiang, and Xiaochun Cao. 2025. Visual content privacy protection: A survey.Comput. Surveys57, 5 (2025), 1–36. Proc. ACM Interact. Mob. Wearable Ubiquitous Technol., Vol. 10, No. 3, Article 159. Publication date: September 20...

  136. [144]

    Desen Zhou, Zhichao Liu, Jian Wang, Leshan Wang, Tao Hu, Errui Ding, and Jingdong Wang. 2022. Human-object interaction detection via disentangled transformer. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. Association for Computing Machine...

  137. [2019]

    InIEEE/CVF Conference on Computer Vision and Pattern Recognition

    Expressive body capture: 3d hands, face, and body from a single image. InIEEE/CVF Conference on Computer Vision and Pattern Recognition. IEEE, Piscataway, NJ, USA, 10975–10985

  138. [2025]

    InIEEE/CVF Conference on Computer Vision and Pattern Recognition

    Interact: Advancing large-scale versatile 3d human-object interaction generation. InIEEE/CVF Conference on Computer Vision and Pattern Recognition. IEEE, Piscataway, NJ, USA, 7048–7060

  139. [2026]

    RFusion: Dynamic Multimodal RF Fusion for Few-Shot Human Activity Recognition.IEEE Transactions on Mobile Computing1-1, 1-1 (2026), 1–1. Proc. ACM Interact. Mob. Wearable Ubiquitous Technol., Vol. 10, No. 3, Article 159. Publication date: September 2026. RF-HOI: Recognize Huma...

Pith tools

Reviewed August 4, 2026 · model on record in the stance chip above.