Pith. sign in

REVIEW 4 major objections 5 minor 205 references

Vision-Based Multimodal Interfaces: A Survey and Taxonomy for Enhanced Context-Aware System Design

T0 review · 4 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read A systematic review of 109 papers claims that a data-modality-driven taxonomy, the Macro-Micro-Macro framework, can organize the design of vision-based multimodal interfaces.

desk verdict A useful, carefully organized taxonomy of vision-based multimodal interfaces; the coding reliability is unverified but the framework is a genuine synthesis worth engaging. read the letter →

arxiv 2501.13443 v6 pith:IPAKQ5O5 submitted 2025-01-23 cs.HC

classification cs.HC
keywords vision-basedmultimodalinterfacescontextawarenessdataintegrationtaxonomysystemdesignframeworksystematicsurveyvisualmodalitycontext-awaresystems
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to establish that the scattered work on vision-based multimodal interfaces can be organized into a single, data-driven design taxonomy that would help practitioners build context-aware systems. It surveys 109 recent papers and classifies them along dimensions ranging from context factors and visual-data types to integration stages, processing approaches, evaluation strategies, application domains, and open challenges. The organizing device is a Macro-Micro-Macro (3M) framework that moves from whole-context considerations to system details and back to whole-system synthesis. If the taxonomy holds up, a designer could use it as an iterative manual for choosing which modalities to fuse, where to fuse them, and how to evaluate the result.

What carries the argument

The central object is the Macro-Micro-Macro (3M) system design framework, a whole-to-details-to-whole structure that organizes the taxonomy into three passes: first, the contextual factors that shape what a system should perceive; second, the concrete building blocks of input modalities, data integration stages, processing approaches, and evaluation strategies; third, the synthesis of these choices into application domains and design challenges. The taxonomy's categories do the work of making individual systems comparable, while the 3M ordering turns those categories into a step-by-step design procedure and a way to trace information flow across the design space.

What would settle it

Re-run the authors' coding on a fresh, independently drawn sample of vision-based multimodal papers from the same period: if many papers cannot be classified or if the category proportions differ sharply from those reported, the taxonomy is not capturing a stable design space.

Watch

Extended reading notes

Core claim

The paper's central claim is that a data modality-driven perspective is the right lens for organizing vision-based multimodal interfaces (VMIs), defined as interfaces that pair visual input with at least one non-visual modality or with two or more distinct visual dimensions. Based on a systematic review of 109 papers, it proposes the Macro-Micro-Macro (3M) framework: macro-level contextual factors (human, environment, system), micro-level system foundations (visual dimensions, other sensing modalities, data integration stages, processing approaches, evaluation strategies), and a return to macro-level design synthesis (application domains, design considerations, key challenges). The claim is that this taxonomy, accompanied by a Sankey diagram and an interactive website, provides an actionable reference for developing context-aware systems and reveals trends and gaps in the existing design space.

Load-bearing premise

The whole taxonomy stands on the assumption that the 109 papers chosen for review fairly represent the variety of vision-based multimodal interfaces and that the authors' manual coding of those papers is consistent and unbiased.

Editorial extensions

If this is right

  • A designer starting a new context-aware system can walk through the framework's sections in order, from context-source factors to evaluation, as a step-by-step manual.
  • The taxonomy makes modality choices comparable across applications, so a team can see, for example, that audio is the most commonly paired modality and haptics is nearly absent.
  • The reported statistics and Sankey diagram give practitioners a way to locate underexplored combinations, such as beyond-human-vision inputs paired with only a few integration stages.
  • If the framework is stable, future reviews of vision-based multimodal interfaces can use its categories as a common vocabulary, making individual system papers easier to compare.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper does not say this, but the same data-modality lens could be inverted into a generative design tool: given a target context, the taxonomy could suggest candidate modalities and fusion stages.
  • One testable extension would be to apply the taxonomy to a second, independently selected corpus published after 2024; stable category proportions would strengthen the claim that the design space is well captured.
  • The near absence of haptic input in the surveyed literature suggests that novel interaction concepts combining vision with touch remain an open opportunity, though the paper itself only reports the count.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper presents a systematic literature review of Vision-Based Multimodal Interfaces (VMIs) for context-aware systems, based on 109 papers selected via a PRISMA-style process from ACM and IEEE. It proposes a Macro-Micro-Macro (3M) framework, organizes the reviewed work into a taxonomy with dimensions covering context source factors, context categories, visual and other sensing modalities, data integration stages, multimodal processing strategies, evaluation strategies, application domains, and design challenges, and supports the taxonomy with an interactive Sankey diagram and an appendix of per-category literature statistics. The central claim is that the 3M framework and taxonomy constitute an actionable, data-modality-driven reference for designing context-aware VMI systems.

Significance. If the taxonomy is reliable and representative, this would be a useful synthesis for HCI practitioners: it is one of the few surveys that explicitly organizes VMI design around data modalities rather than tasks or scenarios, and it ships practical resources (interactive Sankey, Appendix B statistics, a section-by-section design manual). The PRISMA-style selection and independent coding by four co-authors with consolidation are methodological strengths, and the paper is candid about several limitations, including the non-orthogonality of visual dimensions and the use of a 10% subset for final validation. However, the actionability of the 3M framework depends on the stability of the coding categories and counts, and the manuscript currently does not provide quantitative inter-rater reliability or the full coding dataset, so the high-stakes statistics (e.g., Section 8.1's 106/109 System Factors count) cannot yet be independently verified.

major comments (4)
  1. [Section 4.1, Appendix A] The visual-dimension taxonomy is applied with a subjective 'most prominent feature' rule for non-orthogonal dimensions, as acknowledged by the LiDAR (Spatial vs. Beyond-Human-Vision) and grayscale microscope image (Standard-Vision vs. Scale) examples, yet no operational definition of 'most prominent' is given and the reported 10% inter-coder validation in Appendix A is not quantified. Because the subsequent design guidance (Section 8.1, the Sankey pathway analysis, and the challenge counts in Section 7) rests on these single-label assignments, I ask the authors to report inter-rater reliability statistics (e.g., Cohen's kappa) per dimension and to provide either the complete coding rubric or a per-paper category-confusion matrix.
  2. [Section 2.3.1, Appendix A] The corpus is selected by a narrow boolean query ('vision-based' AND 'multimodal' AND 'context aware' plus unspecified synonyms) in ACM and IEEE since 2018, with exclusion criteria that removed 831 works, so the claim that the resulting 109 papers constitute a representative basis for a VMI design-space taxonomy is not yet established. Please add a sensitivity check (e.g., searching with alternative query formulations or without the 'vision-based' conjunct) and demonstrate that the excluded categories do not materially change the main counts and pathway structure.
  3. [Section 2.2, Section 7, Appendix B] The 'actionable reference' claim depends on the exact challenge and category counts (e.g., Privacy and Security 46, Ethics 27, Cognitive Load 58, System Factors 106), but the coding categories overlap and the per-paper assignment data are not published, so these counts are not independently auditable. The authors should release the full coding dataset (or at least per-category disagreement statistics) alongside the interactive Sankey so that designers can assess how much the numbers would shift under alternative coding.
  4. [Section 2.1.3, Section 2.3.1] The VMI definition is author-defined and simultaneously used as the inclusion criterion for the corpus from which the taxonomy is induced, which is a mild self-referential loop. This is not fatal because the taxonomy is explicitly presented as a design tool, but the authors should state clearly that the taxonomy is relative to this definition and should discuss how the framework would change if the VMI definition were widened (e.g., to include single-modality vision systems or non-HCI multimodal systems).
minor comments (5)
  1. [Section 2.3.1] The search description refers to 'related synonyms' but does not list them, which limits reproducibility; please provide the exact synonym set and query strings used in ACM and IEEE.
  2. [Figure 7, Section 8] The Sankey diagram and Appendix B counts are multi-labeled (a paper can appear in multiple categories), but neither the figure caption nor Section 8 explains this explicitly, so a reader may incorrectly sum columns to 109; please add an explicit note that papers are counted in multiple categories.
  3. [Section 7, Challenge 11] The text contains the typo 'scability' in the scalable architecture challenge; please correct it to 'scalability'.
  4. [Section 5.2] The heading 'Developing A Dedicated ML Models' mixes singular and plural; it should be 'Developing a Dedicated ML Model' or 'Developing Dedicated ML Models'.
  5. [Section 8.1, Appendix A] The interactive resources are hosted on a Google Drive folder; for long-term accessibility and citation, please deposit the Sankey source, coding materials, and literature statistics in a DOI-backed repository.

Circularity Check

1 steps flagged · score 2.0 of 10

No significant circularity: only a mild self-referential loop between the VMI definition and the Section 4 visual-dimension scheme; no headline count is forced by it, and the authors' self-citations are not load-bearing.

  1. self definitional [Section 2.1.3 (VMI definition) + Section 4.1 (dimension coding rule) + Appendix A (corpus filter)]
    "Specifically, a VMI is a subset of VI, characterized by the inclusion of at least one non-visual modality in addition to the visual modality or a combination of two or more distinct visual dimensions within its input data (detailed in Section 4). It is important to note that the defined dimensions are not strictly orthogonal; a single image type may exhibit characteristics spanning multiple dimensions. For instance, LiDAR data [33], categorized under the spatial dimension, also possesses properties beyond human vision. However, classification is guided by the most prominent feature."

    The VMI scope definition makes corpus membership depend, for dimension-mixed inputs, on the Section 4 visual-dimension scheme: any input combining 'two or more distinct visual dimensions (detailed in Section 4)' is a VMI, and Section 4.1 resolves overlaps by the subjective 'most prominent feature' rule; Appendix A excludes 65 works as 'Misaligned with the definition of VMIs.' The same Section 4 scheme then supplies the reported dimension frequencies (Appendix B: Spatial 65, Scale 25, Beyond-Human-Vision 22) discussed in Section 8.1, so those per-dimension counts are partly definition-enforced rather than independently discovered (e.g., a LiDAR+RGB system's inclusion and its dimension label both derive from the 'most prominent feature' coding).

full rationale

This is a survey whose derivation chain is an inductive coding exercise (Section 2.3.2, Appendix A), not a quantitative derivation; there are no predictions, fitted parameters, or equations whose outputs could reduce to their inputs. The central claims, that the 3M framework organizes the VMI design space and that the counts (e.g., System Factors in 106/109 studies, 97%) identify priorities, are descriptive statistics over the hand-coded corpus. I checked the seven circularity patterns. (1) Self-definition: the VMI definition's second disjunct invokes the Section 4 visual dimensions, and Section 4's dimension counts describe the corpus filtered by that same definition; this is a genuine but mild loop (flagged as a step) that does not force any headline number. (2) Fitted input called prediction: none, since no quantitative prediction is made from the corpus. (3) Self-citation load-bearing: the authors' own systems (MicroCam [67], RadarFoot [41], Auth+Track [100], SpeCam [187], Video2Haptics [25], [65]) appear as corpus examples, but the VMI definition is justified by external sources ([141], [145], [161]), the context taxonomy is explicitly credited to Dey and Abowd [3] and Grubert et al. [51], and removing the self-cited examples would not change the taxonomy's structure. (4) Uniqueness imported from authors: none. (5) Ansatz smuggled via citation: none; the visual dimensions are stated as conceptual image-based categories and the 'most prominent feature' tie-breaker is disclosed in-text. (6) Renaming known result: the paper explicitly declines to create a new context framework ('Rather than creating a new framework, our work adapts these taxonomies'), so the reorganization is credited rather than disguised. Per the review rule on flagged limitations: Appendix A asserts that inter-coder reliability and alignment were verified on a 10% random subset but reports no statistic (no Cohen's kappa or agreement coefficient), and Section 4.1's 'most prominent feature' rule lacks an operational definition; these are reliability and validity gaps that could shift the count-based design guidance (the skeptic's attack), but they are correctness risks, not circularity, and do not raise the score. Score 2 reflects the one mild self-referential loop; the paper is otherwise self-contained against external benchmarks and makes no prediction that reduces to its inputs.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The paper is a survey and taxonomy; it introduces no free parameters or invented entities. It rests on domain assumptions about the representativeness of the selected literature and the reliability of the coding process.

assumptions (3)
  • domain assumption The selected 109 papers are representative of the VMI literature
    The taxonomy is derived solely from this corpus; a biased selection would invalidate the framework's generality.
  • domain assumption The coding categories and their assignments are reliable
    Coding was done by four co-authors but inter-rater reliability is not reported.
  • domain assumption The definition of VMI (visual plus at least one non-visual modality) captures the essential scope
    This definition shapes which papers are included; alternative definitions would change the taxonomy.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Vision-Based Multimodal Interfaces: A Survey and Taxonomy for Enhanced Context-Aware System Design." pith.science (2026). https://pith.science/paper/IPAKQ5O5

@misc{pith2026250113443,
  author       = {Pith},
  title        = {Pith review of: Vision-Based Multimodal Interfaces: A Survey and Taxonomy for Enhanced Context-Aware System Design},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/IPAKQ5O5}},
  note         = {Machine review of arXiv:2501.13443}
}
read the original abstract

The recent surge in artificial intelligence, particularly in multimodal processing technology, has advanced human-computer interaction, by altering how intelligent systems perceive, understand, and respond to contextual information (i.e., context awareness). Despite such advancements, there is a significant gap in comprehensive reviews examining these advances, especially from a multimodal data perspective, which is crucial for refining system design. This paper addresses a key aspect of this gap by conducting a systematic survey of data modality-driven Vision-based Multimodal Interfaces (VMIs). VMIs are essential for integrating multimodal data, enabling more precise interpretation of user intentions and complex interactions across physical and digital environments. Unlike previous task- or scenario-driven surveys, this study highlights the critical role of the visual modality in processing contextual information and facilitating multimodal interaction. Adopting a design framework moving from the whole to the details and back, it classifies VMIs across dimensions, providing insights for developing effective, context-aware systems.

Figures

Figures reproduced from arXiv: 2501.13443 by the authors.

Figure 1
Figure 1. We review and categorize VMIs aimed at enhancing context awareness. Our key contribution is a [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. The publication growth trend for vision-based multimodal interface and context awareness in the ACM Digital Library [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Examples of context source factors in VMIs with descriptions and citations (illustrative references (a): [ [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figures from the paper (5 more)
Figure 4
Figure 4. Figure 4: Examples of context categories in VMIs with descriptions and citations (illustrative references (a): [ [PITH_FULL_IMAGE:figures/full_fig_p005_4.png]
Figure 5
Figure 5. Figure 5: Examples of application domains for VMIs (illustrative references (a): [ [PITH_FULL_IMAGE:figures/full_fig_p014_5.png]
Figure 6
Figure 6. Figure 6: Examples of design considerations and key challenges for VMIs (illustrative references (a): [ [PITH_FULL_IMAGE:figures/full_fig_p017_6.png]
Figure 7
Figure 7. Figure 7: A Sankey diagram summarizing the overall literature counts across critical dimensions of our taxonomy. From top to [PITH_FULL_IMAGE:figures/full_fig_p018_7.png]
Figure 8
Figure 8. Figure 8: A PRISMA-style flowchart of the selection of studies for the systematic review and meta-analysis. [PITH_FULL_IMAGE:figures/full_fig_p026_8.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

205 extracted references · 44 canonical work pages

  1. [1]

    The PRISMA Statement for Reporting Systematic Reviews and Meta- Analyses of Studies That Evaluate Health Care Interventions: Explanation and Elaboration

    2009. The PRISMA Statement for Reporting Systematic Reviews and Meta- Analyses of Studies That Evaluate Health Care Interventions: Explanation and Elaboration. Annals of Internal Medicine 151, 4 (2009), W–65–W–94. https: //doi.org/10.7326/0003-4819-151-4-200908180-00136 PMID: 19622512

  2. [2]

    Abbas, Bie Tong, and Raid Abdulla

    Maythem K. Abbas, Bie Tong, and Raid Abdulla. 2018. A Hybrid Alert System for Deaf People using Context-Aware Computing and Image Processing. In2018 4th International Conference on Computer and Information Sciences (ICCOINS) . IEEE, 1–6. https://doi.org/10.1109/ICCOINS.2018.8510584

  3. [3]

    Abowd, Anind K

    Gregory D. Abowd, Anind K. Dey, Peter J. Brown, Nigel Davies, Mark Smith, and Pete Steggles. 1999. Towards a Better Understanding of Context and Context- Awareness. In Handheld and Ubiquitous Computing , Hans-W. Gellersen (Ed.). Springer Berlin Heidelberg, Berlin, Heidelberg, 304–307

  4. [4]

    Karan Ahuja, Sven Mayer, Mayank Goel, and Chris Harrison. 2021. Pose-on- the-Go: Approximating User Pose with Smartphone Sensor Fusion and Inverse Kinematics. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems (Yokohama, Japan) (CHI ’21). Association for Computing Machinery, New York, NY, USA, Article 9, 12 pages. https://doi....

  5. [5]

    Sheeraz Athar, Gaurav Patel, Zhengtong Xu, Qiang Qiu, and Yu She. 2023. VisTac Toward a Unified Multimodal Sensing Finger for Robotic Manipulation. IEEE Sensors Journal 23 (2023), 25440–25450. https://api.semanticscholar.org/ CorpusID:261599688

  6. [6]

    Tadas Baltrušaitis, Chaitanya Ahuja, and Louis-Philippe Morency. 2018. Multi- modal machine learning: A survey and taxonomy. IEEE transactions on pattern analysis and machine intelligence 41, 2 (2018), 423–443

  7. [7]

    Kimin Ban and Eui S Jung. 2020. Ear shape categorization for ergonomic product design. International Journal of Industrial Ergonomics 80 (2020), 102962

  8. [8]

    Hyuntae Bang, Jiyoung Min, and Haemin Jeon. 2021. Deep Learning-Based Con- crete Surface Damage Monitoring Method Using Structured Lights and Depth Camera. Sensors (Basel, Switzerland) 21 (2021). https://api.semanticscholar.org/ CorpusID:233396089

Show all 205 references
  1. [9]

    Abdelkareem Bedri, Richard Li, Malcolm Haynes, Raj Prateek Kosaraju, Ishaan Grover, Temiloluwa Prioleau, Min Yan Beh, Mayank Goel, Thad Starner, and Gregory Abowd. 2017. EarBit: using wearable sensors to detect eating episodes in unconstrained environments. Proceedings of the ...

  2. [12]

    Anuraag Bodi, Samuel Berweger, Raied Caromi, Jihoon Bang, Jelena Senic, and Camillo Gentile. 2024. AI-Based Environment Segmentation Using a Context- Aware Channel Sounder. 2024 18th European Conference on Antennas and Propagation (EuCAP) (2024), 1–5. https://api.semanticschol...

  3. [13]

    Panagiotis-Alexandros Bokaris, Benjamin Askenazi, and Michael Haddad. 2019. Light me up: An augmented-reality projection system. In SIGGRAPH Asia 2019 XR. 21–22

  4. [14]

    Cristiana Bolchini, Carlo A Curino, Elisa Quintarelli, Fabio A Schreiber, and Letizia Tanca. 2007. A data-oriented survey of context models. ACM Sigmod Record 36, 4 (2007), 19–26

  5. [15]

    Eran Borenstein, Eitan Sharon, and Shimon Ullman. 2004. Combining top-down and bottom-up segmentation. In 2004 Conference on Computer Vision and Pattern Recognition Workshop. IEEE, 46–46

  6. [16]

    Jacob Bouchard-Roy, Aidin Delnavaz, and Jérémie Voix. 2020. In-ear energy harvesting: Evaluation of the power capability of the temporomandibular joint. IEEE Sensors Journal 20, 12 (2020), 6338–6345

  7. [17]

    John Brooke. 2013. SUS: a retrospective. Journal of Usability Studies 8 (01 2013), 29–40

  8. [18]

    Qiong Cai, Hao Wang, Zhenmin Li, and Xiao Liu. 2019. A survey on multimodal data-driven smart healthcare systems: approaches and applications. IEEE Access 7 (2019), 133583–133599

  9. [20]

    Yanpeng Cao, Baobei Xu, Zhangyu Ye, Jiangxin Yang, Yanlong Cao, Christel- Loïc Tisse, and Xin Li. 2018. Depth and thermal sensor fusion to enhance 3D thermographic reconstruction. Optics express 26 7 (2018), 8179–8193. https: //api.semanticscholar.org/CorpusID:25437618

  10. [21]

    Nitesh V Chawla, Kevin W Bowyer, Lawrence O Hall, and W Philip Kegelmeyer

  11. [22]

    Chen Chen, Cuong Nguyen, Jane Hoffswell, Jennifer Healey, Trung Bui, and Nadir Weibel. 2023. PaperToPlace: Transforming Instruction Documents into Spatialized and Context-Aware Mixed Reality Experiences. In Proceedings of the 36th Annual ACM Symposium on User Interface Softwar...

  12. [23]

    Liuqing Chen, Yu Cai, Ruyue Wang, Shixian Ding, Yilin Tang, Preben Hansen, and Lingyun Sun. 2024. Supporting Text Entry in Virtual Reality with Large Language Models. In 2024 IEEE Conference Virtual Reality and 3D User Interfaces (VR). IEEE, 524–534

  13. [24]

    Tuochao Chen, Yaxuan Li, Songyun Tao, Hyunchul Lim, Mose Sakashita, Ruidong Zhang, François Guimbretière, and Cheng Zhang. 2021. NeckFace: Continuously Tracking Full Facial Expressions on Neck-mounted Wearables. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiqu...

  14. [25]

    Xiaoming Chen, Zeke Zexi Hu, Guangxin Zhao, Haisheng Li, Vera Chung, and Aaron Quigley. 2024. Video2Haptics: Converting Video Motion to Dynamic Haptic Feedback with Bio-Inspired Event Processing. IEEE Transactions on Visualization and Computer Graphics (2024)

  15. [27]

    Youngjun Cho, Nadia Bianchi-Berthouze, Nicolai Marquardt, and Simon J. Julier

  16. [28]

    Chieh Chou, Haifeng Li, and Dezhen Song. 2020. Encoder-Camera-Ground Penetrating Radar Sensor Fusion: Bimodal Calibration and Subsurface Mapping. IEEE Transactions on Robotics 37 (2020), 67–81. https://api.semanticscholar.org/ CorpusID:225506468

  17. [29]

    Michael Compton, Cory Henson, Laurent Lefort, Holger Neuhaus, and Amit Sheth. 2009. A survey of the semantic specification of sensors. In Proceedings of the 2nd International Conference on Semantic Sensor Networks - Volume 522 (Washington DC) (SSN’09). CEUR-WS.org, Aachen, DEU, 17–32

  18. [30]

    Fernanda De La Torre, Cathy Mengying Fang, Han Huang, Andrzej Banburski- Fahey, Judith Amores Fernandez, and Jaron Lanier. 2024. Llmr: Real-time prompt- ing of interactive worlds using large language models. In Proceedings of the CHI Conference on Human Factors in Computing Sy...

  19. [31]

    Aidin Delnavaz and Jérémie Voix. 2013. Piezo-earpiece for micro-power genera- tion from ear canal dynamic motion. Journal of Micromechanics and microengi- neering 23, 11 (2013), 114001

  20. [32]

    Pengchao Deng, Chenyang Ge, Hao Wei, Yuan Sun, and Xin Qiao. 2023. Attention-Aware Dual-Stream Network for Multimodal Face Anti-Spoofing. IEEE Transactions on Information Forensics and Security 18 (2023), 4258–4271. https://api.semanticscholar.org/CorpusID:259612172

  21. [33]

    Yuanzhi Deng, Cheng Chi, Huajie Wen, Yang Zhou, Gang Xu, and Jianhao Shen

  22. [34]

    Anind K Dey, Gregory D Abowd, and Daniel Salber. 2001. A conceptual frame- work and a toolkit for supporting the rapid prototyping of context-aware appli- cations. Human–Computer Interaction 16, 2-4 (2001), 97–166

  23. [35]

    Anind K Dey and Jennifer Mankoff. 2005. Designing mediation for context- aware applications. ACM Transactions on Computer-Human Interaction (TOCHI) 12, 1 (2005), 53–80

  24. [36]

    Alan Dix. 2004. Human-computer interaction. Vol. 1. Pearson Education

  25. [37]

    Mustafa Doga Dogan, Steven Vidal Acevedo Colon, Varnika Sinha, Kaan Akşit, and Stefanie Mueller. 2021. SensiCut: Material-Aware Laser Cutting Using Speckle Sensing and Deep Learning. InThe 34th Annual ACM Symposium on User Interface Software and Technology (Virtual Event, USA)...

  26. [38]

    Messaoud Doudou, Abdelmadjid Bouabdallah, and Véronique Berge-Cherfaoui

  27. [39]

    Paul Dourish. 2004. What we talk about when we talk about context. Personal and ubiquitous computing 8 (2004), 19–30

  28. [40]

    Bruno Dumas, Denis Lalanne, and Sharon Oviatt. 2009. Multimodal Interfaces: A Survey of Principles, Models and Frameworks . Vol. 5440. Springer, 3–26. https: //doi.org/10.1007/978-3-642-00437-7_1

  29. [41]

    Don Samitha Elvitigala, Yunfan Wang, Yongquan Hu, and Aaron J Quigley. 2023. RadarFoot: Fine-grain Ground Surface Context Awareness for Smart Shoes. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology (San Francisco, CA, USA) (UIST ’23). A...

  30. [42]

    Morris, Chun-Cheng Chang, Xuhai Orson Xu, Lianhui Qin, Daniel McDuff, Xin Liu, Shwetak N

    Zachary Englhardt, Chengqian Ma, Margaret E. Morris, Chun-Cheng Chang, Xuhai Orson Xu, Lianhui Qin, Daniel McDuff, Xin Liu, Shwetak N. Patel, and Vikram Iyer. 2023. From Classification to Clinical Insights. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous...

  31. [43]

    Eva Eriksson, Thomas Riisgaard Hansen, and Andreas Lykke-Olesen. 2007. Movement-based interaction in camera spaces: a conceptual framework. Per- sonal and Ubiquitous Computing 11 (2007), 621–632. Vision-Based Multimodal Interfaces: A Survey and Taxonomy for Enhanced Context-Aw...

  32. [45]

    Andreas Fender and Jörg Müller. 2018. Velt: A Framework for Multi RGB-D Camera Systems. In Proceedings of the 2018 ACM International Conference on In- teractive Surfaces and Spaces (Tokyo, Japan) (ISS ’18). Association for Computing Machinery, New York, NY, USA, 73–83. https:/...

  33. [46]

    What’s Happening at that Hip?

    Hasan Shahid Ferdous, Thuong Hoang, Zaher Joukhadar, Martin N Reinoso, Frank Vetere, David Kelly, and Louisa Remedios. 2019. " What’s Happening at that Hip?" Evaluating an On-body Projection based Augmented Reality System for Physiotherapy Classroom. In Proceedings of the 2019...

  34. [47]

    https://api.semanticscholar.org/CorpusID:265351685

  35. [48]

    Fang Fu and Yan Luximon. 2020. A systematic review on ear anthropome- try and its industrial design applications. Human Factors and Ergonomics in Manufacturing & Service Industries 30, 3 (2020), 176–194

  36. [49]

    Mana Fukasawa and Yu Nakayama. 2022. Spatial Augmented Reality Assistance System with Accelerometer and Projection Mapping at Cleaning Activities. In ACM SIGGRAPH 2022 Posters. 1–2

  37. [50]

    Raghuraman Gopalan and Behzad Dariush. 2009. Toward a vision based hand gesture interface for robotic grasping. In Proceedings of the 2009 IEEE/RSJ Inter- national Conference on Intelligent Robots and Systems (IROS’09) . IEEE Press, St. Louis, MO, USA, 1452–1459

  38. [51]

    Jens Grubert, Tobias Langlotz, Stefanie Zollmann, and Holger Regenbrecht

  39. [52]

    David Fleer and Christian Leichsenring. 2012. MISO: a context-sensitive mul- timodal interface for smart objects based on hand gestures and finger snaps. In Adjunct Proceedings of the 25th Annual ACM Symposium on User Interface Software and Technology (Cambridge, Massachusetts...

  40. [53]

    Kunal Gupta, Yuewei Zhang, Tamil Selvan Gunasekaran, Prasanth Sasikumar, Nanditha Krishna, Philip Pits, Conor Russomanno, and Mark Billinghurst. 2023. SensoryScape: Context-Aware Empathic VR Photography. InSIGGRAPH Asia 2023 XR (Sydney, NSW, Australia)(SA ’23). Association for...

  41. [54]

    Spencer Hallyburton, Yupei Liu, Yulong Cao, Z

    R. Spencer Hallyburton, Yupei Liu, Yulong Cao, Z. Morley Mao, and Miroslav Pajic. 2022. Security Analysis of Camera-LiDAR Fusion Against Black-Box Attacks on Autonomous Vehicles. In 31st USENIX Security Symposium (USENIX Security 22). USENIX Association, Boston, MA, 1903–1920....

  42. [55]

    Albert Haque, Arnold Milstein, and Li Fei-Fei. 2020. Illuminating the dark spaces of healthcare with ambient intelligence. Nature 585, 7824 (2020), 193–202

  43. [56]

    Chris Harrison and Scott E. Hudson. 2008. Lightweight material detection for placement-aware mobile computing. In Proceedings of the 21st Annual ACM Symposium on User Interface Software and Technology (Monterey, CA, USA) (UIST ’08). Association for Computing Machinery, New Yor...

  44. [57]

    Sandra G. Hart. 2006. Nasa-Task Load Index (NASA-TLX); 20 Years Later. Proceedings of the Human Factors and Ergonomics Society Annual Meet- ing 50, 9 (2006), 904–908. https://doi.org/10.1177/154193120605000909 arXiv:https://doi.org/10.1177/154193120605000909

  45. [58]

    Renan Guarese, João Becker, Henrique Fensterseifer, Marcelo Walter, Carla Freitas, Luciana Nedel, and Anderson Maciel. 2020. Augmented Situated Vi- sualization for Spatial and Context-Aware Decision-Making. In Proceedings of the 2020 International Conference on Advanced Visual...

  46. [59]

    Abdul Kareem

    Haitham Sabah Hasan and S. Abdul Kareem. 2012. Human Computer Interaction for Vision Based Hand Gesture Recognition: A Survey. In 2012 International Conference on Advanced Computer Science Applications and Technologies (ACSAT). IEEE, 55–60. https://doi.org/10.1109/ACSAT.2012.37

  47. [61]

    Haibo He and Edwardo A Garcia. 2009. Learning from imbalanced data. IEEE Transactions on knowledge and data engineering 21, 9 (2009), 1263–1284

  48. [62]

    Liwen He, Yifan Li, Mingming Fan, Liang He, and Yuhang Zhao. 2023. A Multi- modal Toolkit to Support DIY Assistive Technology Creation for Blind and Low Vision People. In Adjunct Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology (San Franci...

  49. [63]

    Tom Heath, Christian Bizer, and J Hendler. 2011. Synthesis lectures on the Semantic Web: theory and technology. Linked data: Evolving the Web into a global data space 1 (2011), 1–136

  50. [64]

    Andy Harter, Andy Hopper, Pete Steggles, Andy Ward, and Paul Webster. 1999. The anatomy of a context-aware application. In Proceedings of the 5th Annual ACM/IEEE International Conference on Mobile Computing and Networking (Seat- tle, Washington, USA) (MobiCom ’99). Association...

  51. [65]

    Yongquan Hu, Wen Hu, and Aaron J. Quigley. 2024. Towards Enhanced Context Awareness with Vision-based Multimodal Interfaces. In Adjunct Proceedings of the 26th International Conference on Mobile Human-Computer Interaction (Melbourne, VIC, Australia) (MobileHCI ’24 Adjunct). As...

  52. [66]

    Yongquan Hu, Black Sun, Pengcheng An, Zhuying Li, Wen Hu, and Aaron J Quigley. 2024. MultiSurf-GPT: Facilitating Context-Aware Reasoning with Large-Scale Language Models for Multimodal Surface Sensing. arXiv preprint arXiv:2408.07311 (2024)

  53. [67]

    Yongquan Hu, Hui-Shyong Yeo, Mingyue Yuan, Haoran Fan, Don Samitha Elvitigala, Wen Hu, and Aaron Quigley. 2023. Microcam: Leveraging smartphone microscope camera for context-aware contact surface sensing. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous T...

  54. [68]

    Yongquan Hu, Mingyue Yuan, Kaiqi Xian, Don Samitha Elvitigala, and Aaron Quigley. 2023. Exploring the design space of employing ai-generated content for augmented reality display. arXiv preprint arXiv:2303.16593 (2023)

  55. [69]

    Salim, Wen Hu, and Aaron J

    Yongquan Hu, Shuning Zhang, Ting Dang, Hong Jia, Flora D. Salim, Wen Hu, and Aaron J. Quigley. 2024. Exploring Large-Scale Language Models to Evaluate EEG-Based Multimodal Data for Mental Health. In Companion of the 2024 on ACM International Joint Conference on Pervasive and U...

  56. [70]

    Trong-Vu Hoang, Quang-Binh Nguyen, Duy-Nam Ly, Khanh-Duy Le, Tam Nguyen, Minh-Triet Tran, and Trung-Nghia Le. 2024. ARtVista: Gateway To Empower Anyone Into Artist. In Extended Abstracts of the CHI Conference on Human Factors in Computing Systems . 1–8

  57. [71]

    Ming Huang, Toshiyo Tamura, Takumi Yoshimura, Tadahiro Tsuchikawa, and Shigehiko Kanaya. 2016. Wearable deep body thermometers and their uses in continuous monitoring for daily healthcare. In 2016 38th Annual International Conference of the IEEE Engineering in Medicine and Bio...

  58. [72]

    Shaoshuai Huang, Xuandong Zhao, Dapeng Wei, Xinheng Song, and Yuanbo Sun. 2024. Chatbot and Fatigued Driver: Exploring the Use of LLM-Based Voice Assistants for Driving Fatigue. In Extended Abstracts of the 2024 CHI Conference on Human Factors in Computing Systems (CHI EA ’24)...

  59. [74]

    Ajune Wanis Ismail, Mark Billinghurst, and Mohd Shahrizal Sunar. 2015. Vision- Based Technique and Issues for Multimodal Interaction in Augmented Reality. In Proceedings of the 8th International Symposium on Visual Information Communi- cation and Interaction (Tokyo, AA, Japan)...

  60. [75]

    Amiruzzaman

    Suphanut Jamonnak, Ye Zhao, Xinyi Huang, and Md. Amiruzzaman. 2021. Geo- Context Aware Study of Vision-Based Autonomous Driving Models and Spatial Video Data. IEEE Transactions on Visualization and Computer Graphics PP (2021), 1–1. https://api.semanticscholar.org/CorpusID:237592808

  61. [76]

    Gang Hua and Matthew Turk. 2022. Vision-based interaction. Springer Nature

  62. [77]

    Qiao Jin, Yu Liu, Ruixuan Sun, Chen Chen, Puqi Zhou, Bo Han, Feng Qian, and Svetlana Yarosh. 2023. Collaborative Online Learning with VR Video: Roles of Collaborative Tools and Shared Video Control. In Proceedings of the 2023 CHI Conference on Human Factors in Computing System...

  63. [78]

    Ankur Joshi, Saket Kale, Satish Chandel, and D Kumar Pal. 2015. Likert scale: Explored and explained. British journal of applied science & technology 7, 4 CHI ’25, April 26-May 1, 2025, Yokohama, Japan Hu et al. (2015), 396–403

  64. [79]

    Kernchen, P.P

    R. Kernchen, P.P. Boda, K. Moessner, B. Mrohs, M. Boussard, and G. Giuliani

  65. [80]

    Mina Khan, Glenn Fernandes, and Pattie Maes. 2021. PAL: Wearable and Per- sonalized Habit-support Interventions in Egocentric Visual and Physiological Contexts. In Proceedings of the Augmented Humans International Conference 2021 (Rovaniemi, Finland) (AHs ’21). Association for...

  66. [81]

    Mohammad Kianpisheh, Alex Mariakakis, and Khai-Nghi Truong. 2024. exHAR: An Interface for Helping Non-Experts Develop and Debug Knowledge-based Human Activity Recognition Systems. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies 8, 1 (2024), ...

  67. [82]

    Mingyu Jin, Qinkai Yu, Dong Shu, Chong Zhang, Lizhou Fan, Wenyue Hua, Suiyuan Zhu, Yanda Meng, Zhenting Wang, Mengnan Du, and Yongfeng Zhang

  68. [84]

    Kolsch, M

    M. Kolsch, M. Turk, and T. Hollerer. 2004. Vision-based interfaces for mobility. In The First Annual International Conference on Mobile and Ubiquitous Systems: Networking and Services, 2004. MOBIQUITOUS 2004. IEEE, 86–94. https://doi. org/10.1109/MOBIQ.2004.1331713

  69. [85]

    Andy Kong, Karan Ahuja, Mayank Goel, and Chris Harrison. 2021. EyeMU Interactions: Gaze + IMU Gestures on Mobile Devices. In Proceedings of the 2021 International Conference on Multimodal Interaction (Montréal, QC, Canada) (ICMI ’21). Association for Computing Machinery, New Y...

  70. [86]

    Yuki Kubo, Ryosuke Takada, Buntarou Shizuki, and Shin Takahashi. 2017. Syn- Cro: context-aware user interface system for smartphone-smartwatch cross- device interaction. In Proceedings of the 2017 CHI Conference Extended Abstracts on Human Factors in Computing Systems . 1794–1801

  71. [87]

    Y. Kuno, M. Sakamoto, K. Sakata, and Y. Shirai. 1994. Vision-based human interface with user-centered frame. In Proceedings of IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS’94) , Vol. 3. IEEE, 2023–2029 vol.3. https://doi.org/10.1109/IROS.1994.407586

  72. [88]

    Summerskill, Russell Marshall, and Ashleigh J

    Alexander Kunze, Steve J. Summerskill, Russell Marshall, and Ashleigh J. Filtness

  73. [89]

    Gierad Laput, Xiang ’Anthony’ Chen, and Chris Harrison. 2016. SweepSense: Ad Hoc Configuration Sensing Using Reflected Swept-Frequency Ultrasonics. In Proceedings of the 21st International Conference on Intelligent User Interfaces (Sonoma, California, USA) (IUI ’16). Associati...

  74. [91]

    Jeungchan Lee, Ishtiaq Mawla, Jieun Kim, Marco L Loggia, Ana Ortiz, Changjin Jung, Suk-Tak Chan, Jessica Gerber, Vincent J Schmithorst, Robert R Edwards, et al. 2019. Machine learning–based prediction of clinical pain using multimodal neuroimaging and autonomic metrics. pain 1...

  75. [93]

    John D Lee, Dary Fiorentino, Michelle L Reyes, Timothy L Brown, Omar Ahmad, James Fell, Nic Ward, and Robert Dufour. 2010. Assessing the feasibility of vehicle-based sensors to detect alcohol impairment. Washington, DC: National Highway Traffic Safety Administration 1, 2 (2010), 7

  76. [94]

    Wonsup Lee, Xiaopeng Yang, Hayoung Jung, Ilgeun Bok, Chulwoo Kim, Ochae Kwon, and Heecheon You. 2018. Anthropometric analysis of 3D ear scans of Koreans and Caucasians for ear product design. Ergonomics 61, 11 (2018), 1480–1495

  77. [95]

    Shengyu Li, Xingxing Li, Shuolong Chen, Yuxuan Zhou, and Shiwen Wang

  78. [96]

    Yifang Li, Nishant Vishwamitra, Hongxin Hu, and Kelly Caine. 2020. Towards a taxonomy of content sensitivity and sharing preferences for photos. In Pro- ceedings of the 2020 CHI Conference on Human Factors in Computing Systems . 1–14

  79. [97]

    Ergonomics 62 (2019), 345 –

    Automation transparency: implications of uncertainty communication for human-automation interaction and interfaces. Ergonomics 62 (2019), 345 –

  80. [98]

    Zhipeng Li, Yi Fei Cheng, Yukang Yan, and David Lindlbauer. 2024. Predicting the Noticeability of Dynamic Virtual Elements in Virtual Reality. In Proceedings of the CHI Conference on Human Factors in Computing Systems . 1–17

  81. [99]

    Yuanfeng Lian, Xu Shi, ShaoChen Shen, and Jing Hua. 2024. Multitask learning for image translation and salient object detection from multimodal remote sensing images. The Visual Computer 40, 3 (2024), 1395–1414

  82. [100]

    Danh Le-Phuoc and Manfred Hauswirth. 2009. Linked open data in sensor data mashups. In Proceedings of the 2nd International Conference on Semantic Sensor Networks - Volume 522 (Washington DC) (SSN’09). CEUR-WS.org, Aachen, DEU, 1–16

  83. [101]

    Haicheng Liao, Huanming Shen, Zhenning Li, Chengyue Wang, Guofa Li, Yiming Bie, and Chengzhong Xu. 2024. GPT-4 enhanced multimodal grounding for autonomous driving: Leveraging cross-modal attention with large language models. Communications in Transportation Research 4 (2024),...

  84. [102]

    Jian Liao, Kevin Van, Zhijie Xia, and Ryo Suzuki. 2024. RealityEffects: Augment- ing 3D Volumetric Videos with Object-Centric Annotation and Dynamic Visual Effects. In Proceedings of the 2024 ACM Designing Interactive Systems Conference . 1248–1261

  85. [104]

    Qing Lin and Youngjoon Han. 2014. A Context-Aware-Based Audio Guidance System for Blind People Using a Multimodal Profile Model. Sensors (Basel, Switzerland) 14 (2014), 18670 – 18700. https://api.semanticscholar.org/CorpusID: 1897887

  86. [105]

    Zihan Lin, Francisco Cruz, and Eduardo Benitez Sandoval. 2024. Self context-aware emotion perception on human-robot interaction. arXiv:2401.10946 [cs.HC] https://arxiv.org/abs/2401.10946

  87. [106]

    IEEE Transactions on Industrial Electronics 71 (2023), 3182–3191

    Two-Step LiDAR/Camera/IMU Spatial and Temporal Calibration Based on Continuous-Time Trajectory Estimation. IEEE Transactions on Industrial Electronics 71 (2023), 3182–3191. https://api.semanticscholar.org/CorpusID: 258452547

  88. [107]

    Yiting Liu, Liang Li, Beichen Zhang, Shan Huang, Zheng-Jun Zha, and Qing- ming Huang. 2023. MaTCR: Modality-Aligned Thought Chain Reasoning for Multimodal Task-Oriented Dialogue Generation. In Proceedings of the 31st ACM International Conference on Multimedia (Ottawa ON, Canad...

  89. [108]

    Yaxuan Li, Yongjae Yoo, Antoine Weill-Duflos, and Jeremy Cooperstock. 2021. Towards Context-aware Automatic Haptic Effect Generation for Home Theatre Environments. In Proceedings of the 27th ACM Symposium on Virtual Reality Software and Technology (Osaka, Japan) (VRST ’21). As...

  90. [109]

    2009.Context A wareness in Human-Computer Interaction

    Jarmo Makkonen, Ivan Avdouevski, Riitta Kerminen, and Ari Visa. 2009.Context A wareness in Human-Computer Interaction . IntechOpen, Rijeka, Chapter 1. https://doi.org/10.5772/7743

  91. [110]

    Arnav Vaibhav Malawade, Trier Mortlock, and Mohammad Abdullah Al Faruque

  92. [111]

    Chen Liang, Chun Yu, Xiaoying Wei, Xuhai Xu, Yongquan Hu, Yuntao Wang, and Yuanchun Shi. 2021. Auth+Track: Enabling Authentication Free Interaction on Smartphone by Continuous User Tracking. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems (Yokoh...

  93. [112]

    Yuki Matsuda, Dmitrii Fedotov, Yuta Takahashi, Yutaka Arakawa, Keiichi Ya- sumoto, and Wolfgang Minker. 2018. EmoTour: Multimodal Emotion Recogni- tion using Physiological and Audio-Visual Features. In Proceedings of the 2018 ACM International Joint Conference and 2018 Interna...

  94. [113]

    Daniel McDuff, Kael Rowan, Piali Choudhury, Jessica Wolk, ThuVan Pham, and Mary Czerwinski. 2019. A Multimodal Emotion Sensing Platform for Building Emotion-Aware Applications. arXiv:1903.12133 [cs.HC] https://arxiv.org/abs/ 1903.12133

  95. [114]

    Liyu Meng, Yuchen Liu, Xiaolong Liu, Zhaopei Huang, Wenqiang Jiang, Tenggan Zhang, Chuanhe Liu, and Qin Jin. 2022. Valence and Arousal Estimation based on Multimodal Temporal-Aware Features for Videos in the Wild. In 2022 IEEE/CVF Conference on Computer Vision and Pattern Reco...

  96. [115]

    Johannes Meyer, Adrian Frank, Thomas Schlebusch, and Enkelejda Kasneci. 2021. A CNN-based Human Activity Recognition System Combining a Laser Feedback Interferometry Eye Movement Sensor and an IMU for Context-aware Smart Glasses. Proceedings of the ACM on Interactive, Mobile, ...

  97. [116]

    Johannes Meyer, Adrian Frank, Thomas Schlebusch, and Enkelejda Kasneci. 2022. U-har: A convolutional approach to human activity recognition combining head and eye movements for context-aware smart glasses. Proceedings of the ACM on Human-Computer Interaction 6, ETRA (2022), 1–...

  98. [117]

    Tang, Mark D

    Xubo Liu, Qiushi Huang, Xinhao Mei, Haohe Liu, Qiuqiang Kong, Jianyuan Sun, Shengchen Li, Tom Ko, Yu Zhang, Lilian H. Tang, Mark D. Plumbley, Volkan Kılıç, and Wenwu Wang. 2023. Visually-Aware Audio Captioning With Adaptive Audio-Visual Attention. arXiv:2210.16428 [eess.AS] ht...

  99. [118]

    Trisha Mittal, Aniket Bera, and Dinesh Manocha. 2021. Multimodal and Context- Aware Emotion Perception Model With Multiplicative Fusion. IEEE MultiMedia 28 (2021), 67–75. https://api.semanticscholar.org/CorpusID:234228590

  100. [119]

    Camillo Lugaresi, Jiuqiang Tang, Hadon Nash, Chris McClanahan, Esha Uboweja, Michael Hays, Fan Zhang, Chuo-Ling Chang, Ming Guang Yong, Juhyun Lee, et al. 2019. Mediapipe: A framework for building perception pipelines. arXiv preprint arXiv:1906.08172 (2019)

  101. [120]

    Colver Ken Howe Ne, Jameel Muzaffar, Aakash Amlani, and Manohar Bance

  102. [121]

    Hall- man, Juha Kostamovaara, and Rauno Heikkilä

    Ilpo Niskanen, Guoyong Duan, Erik Vartiainen, Matti Immonen, Lauri W. Hall- man, Juha Kostamovaara, and Rauno Heikkilä. 2024. Enhancing point cloud data fusion through 2D thermal infrared camera and 2D lidar scanning. In- frared Physics & Technology (2024). https://api.semanti...

  103. [122]

    Mina Nouredanesh, Alan Godfrey, Dylan Powell, and James Tung. 2022. Ego- centric vision-based detection of surfaces: towards context-aware free-living digital biomarkers for gait and fall risk assessment. Journal of neuroengineering and rehabilitation 19, 1 (2022), 79

  104. [123]

    Buxton, and Ken Hinckley

    Nicolai Marquardt, Nathalie Henry Riche, Christian Holz, Hugo Romat, Michel Pahud, Frederik Brudy, David Ledo, Chunjong Park, Molly Jane Nicholas, Teddy Seyed, Eyal Ofek, Bongshin Lee, William A.S. Buxton, and Ken Hinckley

  105. [124]

    José Ramón Padilla-López, Alexandros Andre Chaaraoui, and Francisco Flórez- Revuelta. 2015. Visual privacy protection methods: A survey. Expert Systems with Applications 42, 9 (2015), 4177–4195

  106. [125]

    Frederik Pahde, Mihai Puscas, Tassilo Klein, and Moin Nabi. 2020. Multimodal Prototypical Networks for Few-shot Learning. arXiv:2011.08899 [cs.CV] https: //arxiv.org/abs/2011.08899

  107. [126]

    Matthias Peissner, Dagmar Häbe, Doris Janssen, and Thomas Sellner. 2012. MyUI: generating accessible user interfaces from multimodal design patterns. In Proceedings of the 4th ACM SIGCHI Symposium on Engineering Interactive Com- puting Systems (Copenhagen, Denmark) (EICS ’12)....

  108. [127]

    Charith Perera, Prem Jayaraman, Arkady Zaslavsky, Peter Christen, and Dim- itrios Georgakopoulos. 2013. Dynamic configuration of sensors using mobile sensor hub in internet of things paradigm. In 2013 IEEE Eighth International Conference on Intelligent Sensors, Sensor Networks...

  109. [128]

    Pavan Kartheek Rachabatuni, Filippo Principi, Paolo Mazzanti, and Marco Bertini. 2024. Context-aware chatbot using MLLMs for Cultural Heritage. In Proceedings of the 15th ACM Multimedia Systems Conference (Bari, Italy) (MM- Sys ’24). Association for Computing Machinery, New Yo...

  110. [129]

    Kankanhalli

    Yogesh Singh Rawat and M. Kankanhalli. 2017. ClickSmart: A Context-Aware Viewpoint Recommendation System for Mobile Photography. IEEE Transactions on Circuits and Systems for Video Technology 27 (2017), 149–158. https://api. semanticscholar.org/CorpusID:9415762

  111. [130]

    Chulhong Min, Akhil Mathur, Alessandro Montanari, and Fahim Kawsar. 2019. An early characterisation of wearing variability on motion signals for wearables. In Proceedings of the 2019 ACM International Symposium on Wearable Computers (London, United Kingdom) (ISWC ’19). Associa...

  112. [131]

    Vítor Sá, Cornelius Malerczyk, and Michael Schnaider. 2001. Vision-Based Interaction within a Multimodal Framework. In Proceedings of the 10th Con- ference of the Eurographics Portuguese Chapter . The Eurographics Association. https://doi.org/10.2312/pt.20011318

  113. [132]

    Gloria Anahi Molina-Barron, Rebeca Elizabeth Alvarado-Ramirez, and Angeles Aguirre-Acosta. 2023. Storytelling: Digital Narration Enhanced by Artificial Intelligence in the Metaverse. In Proceedings of the 2023 7th International Con- ference on Education and E-Learning . 33–40

  114. [133]

    Daniel Salber. 2000. Context-awareness and multimodality. In Colloque sur la multimodalité

  115. [134]

    Expert Review of Medical Devices 18, sup1 (2021), 95–128

    Hearables, in-ear sensing devices for bio-signal acquisition: a narrative review. Expert Review of Medical Devices 18, sup1 (2021), 95–128

  116. [135]

    Schilit, N

    B. Schilit, N. Adams, and R. Want. 1994. Context-Aware Computing Applications. In 1994 First Workshop on Mobile Computing Systems and Applications . IEEE, 85–90. https://doi.org/10.1109/WMCSA.1994.16

  117. [136]

    Albrecht Schmidt. 2000. Implicit human computer interaction through context. Personal technologies 4 (2000), 191–199

  118. [137]

    Sharon Oviatt. 2002. Multimodal interfaces. L. Erlbaum Associates Inc., USA, 286–304

  119. [138]

    Schuller, Tuomas Virtanen, Maria Riveiro, Georgios Rizos, Jing Han, Annamaria Mesaros, and Konstantinos Drossos

    Björn W. Schuller, Tuomas Virtanen, Maria Riveiro, Georgios Rizos, Jing Han, Annamaria Mesaros, and Konstantinos Drossos. 2021. Towards Sonification in Multimodal and User-friendlyExplainable Artificial Intelligence. In Proceedings of the 2021 International Conference on Multi...

  120. [139]

    Tristan Schwörer, Jonathan Eichild Schmidt, and Dimitrios Chrysostomou. 2023. Nav2CAN: Achieving Context Aware Navigation in ROS2 Using Nav2 and RGB-D sensing. In 2023 IEEE International Conference on Imaging Systems and Techniques (IST). IEEE Signal Processing Society, United...

  121. [140]

    Omer Berat Sezer, Erdogan Dogdu, and Ahmet Murat Ozbayoglu. 2017. Context- aware computing, learning, and big data in internet of things: a survey. IEEE Internet of Things Journal 5, 1 (2017), 1–27

  122. [141]

    Rajeev Sharma, Vladimir I Pavlovic, and Thomas S Huang. 1998. Toward multi- modal human-computer interface. Proc. IEEE 86, 5 (1998), 853–869

  123. [142]

    Mali Shen, Yun Gu, Ning Liu, and Guang-Zhong Yang. 2019. Context-Aware Depth and Pose Estimation for Bronchoscopic Navigation. IEEE Robotics and Automation Letters 4 (2019), 732–739. https://api.semanticscholar.org/CorpusID: 59619567

  124. [144]

    Natalie Ruiz, Fang Chen, and Sharon Oviatt. 2010. Chapter 12 - Multimodal Input. In Multimodal Signal Processing, Jean-Philippe Thiran, Ferran Marqués, and Hervé Bourlard (Eds.). Academic Press, Oxford, 231–255. https://doi.org/ 10.1016/B978-0-12-374825-6.00010-1

  125. [145]

    Gihan Shin and Junchul Chun. 2007. Vision-Based Multimodal Human Com- puter Interface Based on Parallel Tracking of Eye and Hand Motion. In 2007 International Conference on Convergence Information Technology (ICCIT 2007) . IEEE, 2443–2448. https://doi.org/10.1109/ICCIT.2007.142

  126. [146]

    Alia Saad, Kian Izadi, Anam Ahmad Khan, Pascal Knierim, Stefan Schneegass, Florian Alt, and Yomna Abdelrahman. 2023. HotFoot: Foot-Based User Iden- tification Using Thermal Imaging. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (Hamburg, Germa...

  127. [147]

    Muhammad Hameed Siddiqi, Nabil Almashfi, Amjad Ali, Madallah Alruwaili, Yousef Alhwaiti, Saad Awadh Alanazi, and M. M. Kamruzzaman. 2021. A Unified Approach for Patient Activity Recognition in Healthcare Using Depth Camera. IEEE Access 9 (2021), 92300–92317. https://api.semant...

  128. [148]

    Javad Sameri, Sam Van Damme, Susanna Schwarzmann, Qing Wei, Riccardo Trivisonno, Filip De Turck, and Maria Torres Vega. 2024. Collaborative Cooking in VR: Effects of Network Distortion in Multi-User Virtual Environments. In Proceedings of the 15th ACM Multimedia Systems Confer...

  129. [149]

    Rukshani Somarathna, Don Samitha Elvitigala, Yijun Yan, Aaron J Quigley, and Gelareh Mohammadi. 2023. Exploring User Engagement in Immersive Virtual Reality Games through Multimodal Body Movements. In Proceedings of the 29th ACM Symposium on Virtual Reality Software and Techno...

  130. [150]

    Dimitrios-Emmanuel Spanos, Periklis Stavrou, Nikolas Mitrou, and Nikolas Konstantinou. 2012. SensorStream: A semantic real–time stream management system. International Journal of Ad Hoc and Ubiquitous Computing 11, 2-3 (2012), 178–193

  131. [151]

    Maximilian Schrapel, Philipp Etgeton, and Michael Rohs. 2021. SpectroPhone: Enabling Material Surface Sensing with Rear Camera and Flashlight LEDs. In Extended Abstracts of the 2021 CHI Conference on Human Factors in Computing Systems (Yokohama, Japan) (CHI EA ’21). Associatio...

  132. [152]

    Xia Su, Eunyee Koh, and Chang Xiao. 2024. SonifyAR: Context-Aware Sound Effect Generation in Augmented Reality. In Extended Abstracts of the 2024 CHI Conference on Human Factors in Computing Systems (CHI EA ’24) . Association for Computing Machinery, New York, NY, USA, Article...

  133. [153]

    Froehlich

    Xia Su, Han Zhang, Kaiming Cheng, Jaewook Lee, Qiaochu Liu, Wyatt Olson, and Jon E. Froehlich. 2024. RASSAR: Room Accessibility and Safety Scanning in Augmented Reality. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI ’2...

  134. [154]

    Zixiong Su, Xinlei Zhang, Naoki Kimura, and Jun Rekimoto. 2021. Gaze+ Lip: rapid, precise and expressive interactions combining gaze input and silent speech commands for hands-free smart TV control. In ACM symposium on eye tracking research and applications . 1–6

  135. [155]

    Fengyuan Sun, Sezer Karaoglu, and Theo Gevers. 2023. Temporally Consis- tent Semantic Segmentation using Spatially Aware Multi-view Semantic Fusion for Indoor RGB-D videos. In 2023 IEEE/CVF International Conference on Com- puter Vision Workshops (ICCVW) . IEEE, 4250–4259. http...

  136. [156]

    Xin Suo, Minye Wu, Yanshun Zhang, Yingliang Zhang, Lan Xu, Qiang Hu, and Jingyi Yu. 2020. Neural3D: Light-weight Neural Portrait Scanning via Context-aware Correspondence Learning. In Proceedings of the 28th ACM In- ternational Conference on Multimedia (Seattle, WA, USA) (MM ’...

  137. [157]

    Hemant Bhaskar Surale, Aakar Gupta, Mark Hancock, and Daniel Vogel. 2019. TabletInVR: Exploring the Design Space for Using a Multi-Touch Tablet in Virtual Reality. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems (Glasgow, Scotland Uk)(CHI ’19). ...

  138. [158]

    Jingyu Shi, Rahul Jain, Hyungjun Doh, Ryo Suzuki, and Karthik Ramani. 2024. An HCI-Centric Survey and Taxonomy of Human-Generative-AI Interactions. arXiv:2310.07127 [cs.HC] https://arxiv.org/abs/2310.07127

  139. [159]

    Ryosuke Suzuki, Tadachika Ozono, and Toramatsu Shintani. 2019. An Offline Mahjong Support System Based on Augmented Reality with Context-aware Image Recognition. In 2019 8th International Congress on Advanced Applied Informatics (IIAI-AAI). IEEE, 127–132. https://doi.org/10.11...

  140. [160]

    Alon Shoa, Ramon Oliva, Mel Slater, and Doron Friedman. 2023. Sushi with Einstein: Enhancing Hybrid Live Events with LLM-Based Virtual Humans. In Proceedings of the 23rd ACM International Conference on Intelligent Virtual Agents. 1–6

  141. [161]

    Shengbang Tong, Ellis Brown, Penghao Wu, Sanghyun Woo, Manoj Mid- depogu, Sai Charitha Akula, Jihan Yang, Shusheng Yang, Adithya Iyer, Xi- chen Pan, Austin Wang, Rob Fergus, Yann LeCun, and Saining Xie. 2024. Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal L...

  142. [162]

    Amit Kumar Sikder, Leonardo Babun, Z Berkay Celik, Hidayet Aksu, Patrick McDaniel, Engin Kirda, and A Selcuk Uluagac. 2022. Who’s controlling my device? Multi-user multi-device-aware access control system for shared smart home environment. ACM Transactions on Internet of Thing...

  143. [163]

    Mikael Uimonen, Paul Kemppi, and Taru Hakanen. 2023. A Gesture-based Mul- timodal Interface for Human-Robot Interaction. In 2023 32nd IEEE International Conference on Robot and Human Interactive Communication (RO-MAN) . IEEE, IEEE, 165–170. https://doi.org/10.1109/RO-MAN57019....

  144. [164]

    Chongyang Wang, Yuan Feng, Lin Xiao Zhong, Siyi Zhu, Chi Zhang, Siqi Zheng, Chen Liang, Yuntao Wang, Chen-Jun He, Chun Yu, and Yuanchun Shi. 2023. UbiPhysio: Support Daily Functioning, Fitness, and Rehabilitation with Action Understanding and Feedback in Natural Language. Proc...

  145. [165]

    Thad Starner. 1995. Visual recognition of american sign language using hidden markov models. Ph. D. Dissertation. Massachusetts Institute of Technology

  146. [167]

    Tianyi Wang, Xun Qian, Fengming He, Xiyun Hu, Ke Huo, Yuanzhi Cao, and Karthik Ramani. 2020. CAPturAR: An Augmented Reality Tool for Authoring Human-Involved Context-Aware Applications. InProceedings of the 33rd Annual ACM Symposium on User Interface Software and Technology (V...

  147. [169]

    Wang, D.Q

    X.H. Wang, D.Q. Zhang, T. Gu, and H.K. Pung. 2004. Ontology based context modeling and reasoning using OWL. In IEEE Annual Conference on Pervasive Computing and Communications Workshops, 2004. Proceedings of the Second . IEEE, 18–22. https://doi.org/10.1109/PERCOMW.2004.1276898

  148. [170]

    Yikai Wang, Wenbing Huang, Bin Fang, Fuchun Sun, and Chang Li. 2021. Elastic tactile simulation towards tactile-visual perception. In Proceedings of the 29th ACM International Conference on Multimedia . 2690–2698

  149. [171]

    Zeyu Wang, Yuanchun Shi, Yuntao Wang, Yuchen Yao, Kun Yan, Yuhan Wang, Lei Ji, Xuhai Xu, and Chun Yu. 2024. G-VOILA: Gaze-Facilitated Information Querying in Daily Scenarios.Proceedings of the ACM on Interactive, Mobile, Wear- able and Ubiquitous Technologies 8 (2024), 1 – 33....

  150. [172]

    Ryo Suzuki, Adnan Karim, Tian Xia, Hooman Hedayati, and Nicolai Marquardt

  151. [173]

    In Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems (New Orleans, LA, USA) (CHI ’22)

    Augmented Reality and Robotics: A Survey and Taxonomy for AR- enhanced Human-Robot Interaction and Robotic Interfaces. In Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems (New Orleans, LA, USA) (CHI ’22). Association for Computing Machinery, New Yor...

  152. [174]

    Mark Weiser and John Seely Brown. 1996. Designing calm technology.PowerGrid Journal 1, 1 (1996), 75–85

  153. [175]

    Next-Gen Vehicle Safety: The Futuristic Approach to Auto-Stop in Modern Vehicles

    Anil Kumar Thandlam, T. Vignesh, A Saravanan, Prakash Subramani, Rama- ganesh Marimuthu, and Sandeep Gupta. 2024. "Next-Gen Vehicle Safety: The Futuristic Approach to Auto-Stop in Modern Vehicles". In2024 10th International Conference on Communication and Signal Processing (IC...

  154. [176]

    Shaoyue Wen, Songming Ping, Jialin Wang, Hai-Ning Liang, Xuhai Xu, and Yukang Yan. 2024. AdaptiveVoice: Cognitively Adaptive Voice Interface for Driving Assistance. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI ’24). A...

  155. [177]

    Hsin-Ruey Tsai, Shih-Kang Chiu, and Bryan Wang. 2024. GazeNoter: Co-Piloted AR Note-Taking via Gaze Selection of LLM Suggestions to Match Users’ Inten- tions. arXiv:2407.01161 [cs.HC] https://arxiv.org/abs/2407.01161

  156. [178]

    Erwin Wu, Chen-Chieh Liao, Ruofan Liu, and Hideki Koike. 2022. Context- aware Risk Degree Prediction for Smartphone Zombies. InACM SIGGRAPH 2022 Posters (Vancouver, BC, Canada) (SIGGRAPH ’22). Association for Computing Machinery, New York, NY, USA, Article 48, 2 pages. https:/...

  157. [179]

    Xinyu Xie, Xiaozhi Zhang, Dongping Xiong, and Lijun Ouyang. 2023. MFA- DAF: Unsupervised Multimodal Medical Image Fusion via Multiscale Fourier Attention and Detail-Aware Fusion Strategy. 2023 International Conference on Image Processing, Computer Vision and Machine Learning (...

  158. [180]

    Chongyang Wang, Siqi Zheng, Lingxiao Zhong, Chun Yu, Chen Liang, Yuntao Wang, Yuan Gao, Tin Lun Lam, and Yuanchun Shi. 2024. PepperPose: Full- Body Pose Estimation with a Companion Robot. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems (Honolulu...

  159. [181]

    Dey, and Dakuo Wang

    Xuhai Xu, Bingsheng Yao, Yuanzhe Dong, Saadia Gabriel, Hong Yu, James Hendler, Marzyeh Ghassemi, Anind K. Dey, and Dakuo Wang. 2024. Mental- LLM: Leveraging Large Language Models for Mental Health Prediction via Online Text Data. Proc. ACM Interact. Mob. Wearable Ubiquitous Te...

  160. [182]

    Yating Xu, Conghui Hu, and Gim Hee Lee. 2022. Motion and Context-Aware Audio-Visual Conditioned Video Prediction. InBritish Machine Vision Conference. https://api.semanticscholar.org/CorpusID:254536032

  161. [183]

    Zhenyu Xu, Hailin Xu, Zhouyang Lu, Yingying Zhao, Rui Zhu, Yujiang Wang, Mingzhi Dong, Yuhu Chang, Qin Lv, Robert P Dick, et al . 2024. Can Large Language Models Be Good Companions? An LLM-Based Eyewear System with Conversational Common Ground. Proceedings of the ACM on Intera...

  162. [184]

    Songlin Yang, Wei Wang, Jun Ling, Bo Peng, Xu Tan, and Jing Dong. 2023. Context-Aware Talking-Head Video Editing. In Proceedings of the 31st ACM International Conference on Multimedia (Ottawa ON, Canada) (MM ’23). As- sociation for Computing Machinery, New York, NY, USA, 7718–...

  163. [185]

    Xing-Dong Yang, Tovi Grossman, Daniel Wigdor, and George Fitzmaurice. 2012. Magic finger: always-available input through finger instrumentation. InProceed- ings of the 25th Annual ACM Symposium on User Interface Software and Technol- ogy (Cambridge, Massachusetts, USA)(UIST ’1...

  164. [186]

    Mengmei Ye, Zhongze Tang, Huy Phan, Yi Xie, Bo Yuan, and Sheng Wei. 2022. Visual privacy protection in mobile image recognition using protective pertur- bation. In Proceedings of the 13th ACM Multimedia Systems Conference . 164–176

  165. [187]

    Zhan Wang, Lin-Ping Yuan, Liangwei Wang, Bingchuan Jiang, and Wei Zeng

  166. [188]

    In Proceedings of the CHI conference on human factors in computing systems

    Virtuwander: Enhancing multi-modal interaction for virtual tour guidance through large language models. In Proceedings of the CHI conference on human factors in computing systems . 1–20

  167. [189]

    Mark Weiser. 1999. The computer for the 21st century. ACM SIGMOBILE mobile computing and communications review 3, 3 (1999), 3–11

  168. [190]

    So-In, and Paramate Horkaew

    Watcharaphong Yookwan, Krisana Chinnasarn, C. So-In, and Paramate Horkaew

  169. [191]

    Linda Yilin Wen, Cecily Morrison, Martin Grayson, Rita Faia Marques, Daniela Massiceti, Camilla Longden, and Edward Cutrell. 2024. Find My Things: Per- sonalized Accessibility through Teachable AI for People who are Blind or Low Vision. Extended Abstracts of the CHI Conference...

  170. [192]

    I Know What You Mean

    Nima Zargham, Mohamed Lamine Fetni, Laura Spillner, Thomas Muender, and Rainer Malaka. 2024. " I Know What You Mean": Context-Aware Recognition to Enhance Speech-Based Games. In Proceedings of the CHI Conference on Human Factors in Computing Systems . 1–18

  171. [193]

    Johann Wentzel, Fraser Anderson, George Fitzmaurice, Tovi Grossman, and Daniel Vogel. 2024. SwitchSpace: Understanding Context-Aware Peeking Be- tween VR and Desktop Interfaces. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA)...

  172. [194]

    Xiyuxing Zhang, Yuntao Wang, Yuxuan Han, Chen Liang, Ishan Chatterjee, Jiankai Tang, Xin Yi, Shwetak Patel, and Yuanchun Shi. 2024. The EarSAVAS Dataset: Enabling Subject-Aware Vocal Activity Sensing on Earables.Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiqu...

  173. [195]

    Fei Zhao, Chengcui Zhang, and Baocheng Geng. 2024. Deep Multimodal Data Fusion. ACM Comput. Surv. 56, 9, Article 216 (April 2024), 36 pages. https: //doi.org/10.1145/3649447

  174. [196]

    Anran Xu, Shitao Fang, Huan Yang, Simo Hosio, and Koji Yatani. 2024. Examin- ing Human Perception of Generative Content Replacement in Image Privacy Protection. In Proceedings of the CHI Conference on Human Factors in Computing Systems. 1–16. Vision-Based Multimodal Interfaces...

  175. [197]

    Zheng, Yuejie Zhang, Rui Feng, Tao Zhang, and Weiguo Fan

    Y. Zheng, Yuejie Zhang, Rui Feng, Tao Zhang, and Weiguo Fan. 2021. Stacked Multimodal Attention Network for Context-Aware Video Captioning. IEEE Transactions on Circuits and Systems for Video Technology 32 (2021), 31–42. https://api.semanticscholar.org/CorpusID:236657677

  176. [198]

    Dingfu Zhou, Xibin Song, Jin Fang, Yuchao Dai, Hongdong Li, and Liangjun Zhang. 2022. Context-Aware 3D Object Detection From a Single Image in Autonomous Driving. IEEE Transactions on Intelligent Transportation Systems 23 (2022), 18568–18580. https://api.semanticscholar.org/Co...

  177. [199]

    Zhongyi Zhou, Jing Jin, Vrushank Phadnis, Xiuxiu Yuan, Jun Jiang, Xun Qian, Jingtao Zhou, Yiyi Huang, Zheng Xu, Yinda Zhang, et al . 2023. InstructPipe: Building Visual Programming Pipelines with Human Instructions.arXiv preprint arXiv:2312.09672 (2023)

  178. [200]

    Zhongyi Zhou and Koji Yatani. 2022. Gesture-aware Interactive Machine Teach- ing with In-situ Object Annotations. In Proceedings of the 35th Annual ACM Symposium on User Interface Software and Technology (Bend, OR, USA) (UIST ’22). Association for Computing Machinery, New York...

  179. [201]

    Rongrong Zhu, Liang Shi, Yunpeng Song, and Zhongmin Cai. 2023. Integrating Gaze and Mouse Via Joint Cross-Attention Fusion Net for Students’ Activity Recognition in E-learning. Proceedings of the ACM on Interactive, Mobile, Wear- able and Ubiquitous Technologies 7 (2023), 1 – ...

  180. [202]

    vision-based

    Chris Zimmerer, Martin Fischbach, and Marc Erich Latoschik. 2022. A Case Study on the Rapid Development of Natural and Synergistic Multimodal Interfaces for XR Use-Cases. In Extended Abstracts of the 2022 CHI Conference on Human Factors in Computing Systems (New Orleans, LA, U...

  181. [203]

    Hui-Shyong Yeo, Juyoung Lee, Andrea Bianchi, David Harris-Birtill, and Aaron Quigley. 2017. SpeCam: sensing surface color and material with the front-facing camera of a mobile device. In Proceedings of the 19th International Conference on Human-Computer Interaction with Mobile...

  182. [204]

    Nur Yildirim, Hannah Richardson, Maria Teodora Wetscherek, Junaid Bajwa, Joseph Jacob, Mark Ames Pinnock, Stephen Harris, Daniel Coelho De Castro, Shruthi Bannur, Stephanie Hyland, Pratik Ghosh, Mercy Ranjit, Kenza Bouzid, Anton Schwaighofer, Fernando Pérez-García, Harshita Sh...

  183. [205]

    Zhizhuo Yin, Yuyang Wang, Theodoros Papatheodorou, and Pan Hui. 2024. Text2VRScene: Exploring the Framework of Automated Text-driven Generation System for VR Experience. In 2024 IEEE Conference Virtual Reality and 3D User Interfaces (VR). IEEE, 701–711

  184. [207]

    IEEE Access 10 (2022), 77123–77136

    Multimodal Fusion of Deeply Inferred Point Clouds for 3D Scene Re- construction Using Cross-Entropy ICP. IEEE Access 10 (2022), 77123–77136. https://api.semanticscholar.org/CorpusID:250976713

  185. [208]

    Xenophon Zabulis, Haris Baltzakis, and Antonis A Argyros. 2009. Vision-Based Hand Gesture Recognition for Human-Computer Interaction. The universal access handbook 34 (2009), 30

  186. [210]

    Chao Zhang, Xuechen Liu, Katherine Ziska, Soobin Jeon, Chi-Lin Yu, and Ying Xu. 2024. Mathemyths: Leveraging Large Language Models to Teach Mathe- matical Language through Child-AI Co-Creative Storytelling. In Proceedings of the 2024 CHI Conference on Human Factors in Computin...

  187. [213]

    Paradiso

    Nan Zhao, Elena Kodama, and Joseph A. Paradiso. 2022. Mediated Atmosphere Table (MAT): Adaptive Multimodal Media System for Stress Restoration. IEEE Internet of Things Journal 9 (2022), 23614–23625. https://api.semanticscholar. org/CorpusID:250564158

  188. [360]

    https://api.semanticscholar.org/CorpusID:54507824

  189. [2002]

    Journal of artificial intelligence research 16 (2002), 321–357

    SMOTE: synthetic minority over-sampling technique. Journal of artificial intelligence research 16 (2002), 321–357

  190. [2005]

    In 2005 IEEE 16th International Symposium on Personal, Indoor and Mobile Radio Communications, Vol

    Multimodal user interfaces for context-aware mobile applications. In 2005 IEEE 16th International Symposium on Personal, Indoor and Mobile Radio Communications, Vol. 4. IEEE, 2268–2273 Vol. 4. https://doi.org/10.1109/PIMRC. 2005.1651849

  191. [2017]

    IEEE Transactions on Visualization and Computer Graphics 23 (2017), 1706–1724

    Towards Pervasive Augmented Reality: Context-Awareness in Augmented Reality. IEEE Transactions on Visualization and Computer Graphics 23 (2017), 1706–1724. https://api.semanticscholar.org/CorpusID:2560516

  192. [2018]

    In Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems (Montreal QC, Canada) (CHI ’18)

    Deep Thermal Imaging: Proximate Material Type Recognition in the Wild through Deep Learning of Spatial Surface Temperature Patterns. In Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems (Montreal QC, Canada) (CHI ’18). Association for Computing Machi...

  193. [2019]

    International Journal of Intelligent Transportation Systems Research 18 (2019), 297 – 319

    Driver Drowsiness Measurement Technologies: Current Research, Market Solutions, and Challenges. International Journal of Intelligent Transportation Systems Research 18 (2019), 297 – 319. https://api.semanticscholar.org/CorpusID: 203081957

  194. [2021]

    AirConstellations: In-Air Device Formations for Cross-Device Interac- tion via Multiple Spatially-Aware Armatures. In The 34th Annual ACM Sym- posium on User Interface Software and Technology (Virtual Event, USA) (UIST Vision-Based Multimodal Interfaces: A Survey and Taxonomy ...

  195. [2022]

    In 2022 ACM/IEEE 13th International Conference on Cyber-Physical Systems (ICCPS)

    HydraFusion: Context-Aware Selective Sensor Fusion for Robust and Efficient Autonomous Vehicle Perception. In 2022 ACM/IEEE 13th International Conference on Cyber-Physical Systems (ICCPS) . IEEE, 68–79. https://doi.org/10. 1109/ICCPS54341.2022.00013

  196. [2023]

    In 2023 4th International Conference on Computer Vision, Image and Deep Learning (CVIDL)

    Context-Aware Fusion for 3D Object Detection in LiDAR-Camera Systems. In 2023 4th International Conference on Computer Vision, Image and Deep Learning (CVIDL). IEEE, 601–608. https://doi.org/10.1109/CVIDL58838.2023.10166260

  197. [2024]

    arXiv:2402.00746 [cs.CL] https://arxiv.org/abs/2402.00746

    Health-LLM: Personalized Retrieval-Augmented Disease Prediction Sys- tem. arXiv:2402.00746 [cs.CL] https://arxiv.org/abs/2402.00746

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.