REVIEW 2 major objections 5 minor 84 references
Gensors: Authoring Personalized Visual Sensors with Multimodal Foundation Models and Reasoning
T0 review · 2 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read Users feel more control, understanding, and clarity defining camera-based AI sensors when they work from testable criteria than when they iterate on a single prompt.
desk verdict Worth a serious referee: Gensors' criteria-decomposition workflow is a real step forward for end-user MLLM sensors, but the baseline model is under-specified and the conclusion overclaims robustness. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the criterion, an atomic unit of reasoning: a minimal natural-language question targeting one aspect of the sensing task, evaluated by an MLLM that returns a valence and a short explanation. Gensors runs all active criteria in parallel on recent camera frames, then has a second MLLM, or a user-chosen boolean rule, combine the individual results into a final verdict. This divergent-then-convergent pipeline, together with the Examples-Diff feature that converts user-labeled positive and negative image frames into new criteria, carries the argument by making each requirement independently inspectable and by surfacing blind spots.
What would settle it
Run a pre-registered study with a larger, less technically oriented sample in which users must make a sensor meet five fixed behavioral specifications using either Gensors or a single-prompt baseline, then measure the fraction of specification checks passed; the central claim would not hold if the criteria interface fails to outperform the baseline on objective correctness or if its control and clarity advantages disappear once age, technical background, and prior experience are controlled.
Extended reading notes
Core claim
The central claim is that decomposing a high-level sensing goal into atomic criteria, each evaluated separately by a multimodal large language model and shown as a color-coded result, gives end users more control, understanding, and communicative clarity than refining a single free-text prompt. In a within-subjects study with 12 participants, Gensors significantly outperformed a prompt-editor baseline on the self-reported measures of control, understanding of underlying model capabilities, and ability to think through and communicate requirements, while the difference on testing was directionally positive but not statistically significant. Qualitatively, participants could localize failures to one criterion, disable or revise it without touching others, and use automatically generated criteria or image-based difference examples to uncover conditions they had not considered. The paper presents this as evidence that MLLM-powered sensing can be made customizable by non-experts when requirements are treated as first-class objects.
Load-bearing premise
The 12 tech-savvy, self-selected participants and their self-reported Likert ratings fairly represent the everyday users the system is meant to serve; the paper itself flags that this sample may limit generalizability.
Editorial extensions
If this is right
- Users can specify high-level, subjective sensing tasks by editing criteria instead of wrestling with one long prompt, lowering the barrier for non-technical people.
- Debugging becomes systematic because a misbehaving sensor can be traced to a single criterion and fixed without risking unintended changes to other behavior.
- Automatically generated criteria and image-difference reasoning reduce omissions, prompting users to consider edge cases and hazards they would not anticipate on their own.
- MLLM stochasticity becomes visible as flickering outputs, and users can treat flickering as a signal that a criterion or prompt needs refinement rather than only as a reliability problem.
- The design goals should remain relevant as models improve, because the bottleneck shifts from raw model capability to requirement articulation and testing.
Reading between the lines
- The criteria-based authoring pattern likely transfers beyond visual sensing to any MLLM task where a long prompt is hard to debug, such as document triage or email classification, since atomic testable conditions provide the same isolation benefits.
- A larger, less tech-savvy sample would clarify whether the interface helps or adds overhead: managing many criteria may become a new burden for users who are not comfortable with parallel, modular thinking.
- One testable extension is to apply the Examples-Diff mechanism to non-binary tasks and non-visual modalities, such as audio or sensor streams, converting labeled examples into criteria in those domains.
- An objective measure, such as the number of edits needed to reach a fixed behavioral specification or the accuracy of final verdicts on held-out scenes, would test whether the self-reported gains in control and understanding translate into better downstream sensor behavior.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces Gensors, a system that lets end-users create personalized AI-powered visual sensors by decomposing high-level sensing tasks into individual criteria, with support for automatically generated criteria, example-based criteria discovery, test-case suggestions, and configurable verdict logic. The authors report a formative study (n=6) that motivates four design goals, then present a within-subjects user study (n=12) comparing Gensors to a prompt-editor baseline. They find significantly higher self-reported ratings for control, understanding of model capabilities, and ability to communicate requirements, while the testing metric did not reach significance. Qualitative analysis describes how criteria-level debugging, auto-generated criteria, and example-diff features helped users handle blind spots and MLLM idiosyncrasies.
Significance. If valid, the central comparative result is a useful contribution to end-user authoring of intelligent sensors: it provides evidence that criteria-based decomposition, rather than single-prompt iteration, can improve perceived control and understanding for non-expert users. The paper is transparent in reporting that the Test metric was not significant, applies full Bonferroni correction, and grounds its claims in both quantitative and qualitative data. The main threat to the central claim is that the baseline condition's underlying model configuration is not specified, leaving open a potential confound between the interface manipulation and model capability. This issue is load-bearing and needs to be addressed before the central claim is fully established.
major comments (2)
- [Sections 4.3, 5.3, and Figure 3] The baseline condition's MLLM configuration is unspecified. Section 4.3 states that Gensors runs individual criteria with Gemini 1.5 Flash at temperature 0 and the final verdict with Gemini 1.5 Pro at temperature 0.8, while Section 5 describes the baseline only as 'a prompt-editor version, similar to the one used in the formative study' and Figure 3 does not state which model or temperature evaluated the baseline prompt. If the baseline prompt was evaluated by Gemini 1.5 Flash alone, or by a different model than the one used for Gensors' final verdict, the significant differences in Control, Understand Capabilities, and Communicate Requirements reported in Section 6.1 could be explained by the stronger reasoning and explanation quality of Gemini 1.5 Pro rather than by the criteria-based authoring workflow. The paper must specify the exact model, temperature, and inference pipeline used in the baseline, and ideally demonstrate that both conditions used comparable model capability. Section 8's limitation list does not mention this potential confound.
- [Sections 6.1 and 6.3] The self-reported Test metric did not reach statistical significance (Section 6.1 reports no p-value, just 'not statistically significant'), yet Section 6.3 makes strong qualitative claims that Gensors enabled 'robustly test and debug' and that users 'could robustly test and debug with Gensors.' Given that the quantitative evidence for testing is weak, the qualitative claims should be tempered or accompanied by effect sizes and exact p-values for all four metrics so that readers can assess the strength of the evidence. The non-significant result also does not rule out the model-capability confound raised above, since perceived control and understanding may be more sensitive to the quality of the model's explanations than to the interface alone.
minor comments (5)
- [Section 6.2] There is a duplicated word: 'feeling feeling that her efforts were counterproductive' should read 'feeling that her efforts were counterproductive.'
- [Section 6.3.1] There is a duplicated phrase: 'changed this “and” to “or”, which then caused the sensor to briefly classify the desk as cluttered, though not consistently so' reads fine, but elsewhere the text has 'described as 'chaotic and disorganized ” to be triggered...' with a stray space before the quotation mark.
- [Section 6.3.2] The phrase 'assessed the the sensor's overall effectiveness' contains a duplicated 'the.'
- [Figure 2] The figure labels are dense and small; consider enlarging the UI screenshots and using higher-contrast labels for the numeric callouts, as some callouts (e.g., k1–k3) are difficult to distinguish in the printed version.
- [Section 5.1] The study procedure says participants spent 50 minutes creating two sensors with 25 minutes per condition, but the earlier description says a 90-minute total commitment; consider clarifying how the 30-minute tutorial, 50-minute building, questionnaire, and interview fit into 90 minutes.
Circularity Check
No significant circularity: the paper's central claims rest on a within-subjects user study with an independent baseline comparison, not on fitted parameters, definitional identities, or a self-citation chain.
full rationale
The paper is an empirical human-centered computing study rather than a mathematical derivation, so the standard circularity patterns do not apply. The central claim, that participants reported significantly greater control, understanding, and ease of communication with Gensors than with a prompt-editor baseline, is supported by paired questionnaire ratings analyzed with Wilcoxon tests (Section 5.3, Section 6.1). No equation, fitted parameter, or defined quantity is renamed as a prediction; the outcome measures are self-reports collected after two counterbalanced conditions, and the statistical analysis does not force the reported differences by construction. The Gensors implementation details in Section 4.3 (Gemini 1.5 Flash for criteria at temperature 0 and Gemini 1.5 Pro for verdicts at temperature 0.8) describe the system under test, not a fitted input that predetermines the outcome. Self-citations in the related work are informative but not load-bearing for the study's conclusions, and no uniqueness theorem or prior author-derived ansatz is invoked to forbid alternatives. The identified limitations, such as the unspecified baseline model configuration (a potential internal validity concern) and the small tech-savvy sample acknowledged in Section 8, are threats to generalizability and confound control, not circularity. Because the empirical comparison is externally collected and independent of any assumed conclusion, the appropriate circularity score is 0.
Assumptions & free parameters
assumptions (3)
- domain assumption Self-reported Likert ratings are valid proxies for the constructs of control, understanding, and communication.
- domain assumption The prompt-editor baseline is a fair representation of the unaided alternative.
- domain assumption The 12-participant tech-savvy sample generalizes to the target population of everyday users.
Cite this review
Pith. "Pith review of Gensors: Authoring Personalized Visual Sensors with Multimodal Foundation Models and Reasoning." pith.science (2026). https://pith.science/paper/3WR4QRCB
@misc{pith2026250115727,
author = {Pith},
title = {Pith review of: Gensors: Authoring Personalized Visual Sensors with Multimodal Foundation Models and Reasoning},
year = {2026},
howpublished = {\url{https://pith.science/paper/3WR4QRCB}},
note = {Machine review of arXiv:2501.15727}
}
read the original abstract
Multimodal large language models (MLLMs), with their expansive world knowledge and reasoning capabilities, present a unique opportunity for end-users to create personalized AI sensors capable of reasoning about complex situations. A user could describe a desired sensing task in natural language (e.g., "alert if my toddler is getting into mischief"), with the MLLM analyzing the camera feed and responding within seconds. In a formative study, we found that users saw substantial value in defining their own sensors, yet struggled to articulate their unique personal requirements and debug the sensors through prompting alone. To address these challenges, we developed Gensors, a system that empowers users to define customized sensors supported by the reasoning capabilities of MLLMs. Gensors 1) assists users in eliciting requirements through both automatically-generated and manually created sensor criteria, 2) facilitates debugging by allowing users to isolate and test individual criteria in parallel, 3) suggests additional criteria based on user-provided images, and 4) proposes test cases to help users "stress test" sensors on potentially unforeseen scenarios. In a user study, participants reported significantly greater sense of control, understanding, and ease of communication when defining sensors using Gensors. Beyond addressing model limitations, Gensors supported users in debugging, eliciting requirements, and expressing unique personal requirements to the sensor through criteria-based reasoning; it also helped uncover users' "blind spots" by exposing overlooked criteria and revealing unanticipated failure modes. Finally, we discuss how unique characteristics of MLLMs--such as hallucinations and inconsistent responses--can impact the sensor-creation process. These findings contribute to the design of future intelligent sensing systems that are intuitive and customizable by everyday users.
Figures
Reference graph
Works this paper leans on
-
[1]
Amazon. 2024. Amazon Mechanical Turk. https://www.mturk.com/
2024
-
[2]
Bach, Victor Sanh, Zheng-Xin Yong, Albert Webson, Colin Raffel, Nihal V
Stephen H. Bach, Victor Sanh, Zheng-Xin Yong, Albert Webson, Colin Raffel, Nihal V. Nayak, Abheesht Sharma, Taewoon Kim, M. Saiful Bari, Thibault Fevry, Zaid Alyafeai, Manan Dey, Andrea Santilli, Zhiqing Sun, Srulik Ben-David, Can- wen Xu, Gunjan Chhablani, Han Wang, Jason Alan Fries, Maged S. Al-shaibani, Shanya Sharma, Urmish Thakker, Khalid Almubarak, ...
-
[3]
Joshua Bell. 2023. Indexed Database API . Technical Report. World Wide Web Consortium (W3C). https://www.w3.org/TR/IndexedDB/
2023
-
[4]
Bigham, Chandrika Jayant, Hanjie Ji, Greg Little, Andrew Miller, Robert C
Jeffrey P. Bigham, Chandrika Jayant, Hanjie Ji, Greg Little, Andrew Miller, Robert C. Miller, Robin Miller, Aubrey Tatarowicz, Brandyn White, Samual White, and Tom Yeh. 2010. VizWiz: nearly real-time answers to visual questions. In Proceedings of the 23nd annual ACM symposium on User interface software and technology (UIST ’10). Association for Computing ...
arXiv 2010
-
[5]
Svetlin Bostandjiev, John O’Donovan, and Tobias Höllerer. 2012. TasteWeights: a visual interactive hybrid recommender system. In Proceedings of the sixth ACM conference on Recommender systems (RecSys ’12) . Association for Computing Machinery, New York, NY, USA, 35–42. doi:10.1145/2365952.2365964
-
[6]
Virginia Braun and Victoria Clarke. 2006. Using thematic analysis in psychology. Qualitative research in psychology 3, 2 (2006), 77–101. Publisher: Taylor & Francis
2006
-
[7]
Michelle Carney, Barron Webster, Irene Alvarado, Kyle Phillips, Noura Howell, Jordan Griffith, Jonas Jongejan, Amit Pitaru, and Alexander Chen. 2020. Teach- able Machine: Approachable Web-Based Tool for Exploring Machine Learning Classification. In Extended Abstracts of the 2020 CHI Conference on Human Factors in Computing Systems (CHI EA ’20) . Associati...
arXiv 2020
-
[8]
Qiaochu Chen, Xinyu Wang, Xi Ye, Greg Durrett, and Isil Dillig. 2020. Multi- modal synthesis of regular expressions. In Proceedings of the 41st ACM SIGPLAN Conference on Programming Language Design and Implementation (PLDI 2020) . Association for Computing Machinery, New York, NY, USA, 487–502. doi:10. 1145/3385412.3385988
arXiv 2020
Show all 84 references
-
[9]
Bernstein
Justin Cheng and Michael S. Bernstein. 2015. Flock: Hybrid Crowd-Machine Learning Classifiers. In Proceedings of the 18th ACM Conference on Computer Supported Cooperative Work & Social Computing (CSCW ’15). Association for Com- puting Machinery, New York, NY, USA, 600–611. doi...
2015
-
[10]
Becker, and Brent N
Paul Denny, Juho Leinonen, James Prather, Andrew Luxton-Reilly, Thezyrie Amarouche, Brett A. Becker, and Brent N. Reeves. 2024. Prompt Problems: A New Programming Exercise for the Generative AI Era. In Proceedings of the 55th ACM Technical Symposium on Computer Science Educati...
2024
- [11]
-
[12]
Cooper, Paul F
Dan Ding, Rory A. Cooper, Paul F. Pasquina, and Lavinia Fici-Pasquina. 2011. Sensor technology for smart homes. Maturitas 69, 2 (June 2011), 131–136. doi:10. 1016/j.maturitas.2011.03.016
2011
-
[13]
Nanyi Fei, Zhiwu Lu, Yizhao Gao, Guoxing Yang, Yuqi Huo, Jingyuan Wen, Haoyu Lu, Ruihua Song, Xin Gao, Tao Xiang, Hao Sun, and Ji-Rong Wen. 2022. Towards artificial general intelligence via a multimodal foundation model. Nature Com- munications 13, 1 (June 2022), 3094. doi:10....
2022 doi
-
[14]
Google DeepMind. 2024. Gemini 1.5 Flash. https://deepmind.google/ technologies/gemini/flash/ Published: Online; accessed October 10, 2024
2024
-
[15]
Google DeepMind. 2024. Gemini 1.5 Pro. https://deepmind.google/technologies/ gemini/pro/ Published: Online; accessed October 10, 2024
2024
-
[16]
Anhong Guo, Xiang ’Anthony’ Chen, Haoran Qi, Samuel White, Suman Ghosh, Chieko Asakawa, and Jeffrey P. Bigham. 2016. VizLens: A Robust and Interactive Screen Reader for Interfaces in the Real World. In Proceedings of the 29th Annual Symposium on User Interface Software and Tec...
2016 doi
-
[17]
Anhong Guo, Anuraag Jain, Shomiron Ghose, Gierad Laput, Chris Harrison, and Jeffrey P. Bigham. 2018. Crowd-AI Camera Sensing in the Real World. Proc. ACM Interact. Mob. Wearable Ubiquitous Technol. 2, 3 (Sept. 2018). doi:10.1145/3264921 Place: New York, NY, USA Publisher: Asso...
2018 doi
-
[18]
Xin Hong, Chris Nugent, Maurice Mulvenna, Sally McClean, Bryan Scotney, and Steven Devlin. 2009. Evidential fusion of sensor data for activity recognition in smart homes. Pervasive and Mobile Computing 5, 3 (June 2009), 236–252. doi:10.1016/j.pmcj.2008.05.002
2009 doi
-
[19]
Myers, and Aniket Kittur
Jane Hsieh, Michael Xieyang Liu, Brad A. Myers, and Aniket Kittur. 2018. An Exploratory Study of Web Foraging to Understand and Support Programming Decisions. In 2018 IEEE Symposium on Visual Languages and Human-Centric Computing (VL/HCC). 305–306. doi:10.1109/VLHCC.2018.85065...
2018
-
[20]
Amy Hwang and Jesse Hoey. 2012. Smart home, the next generation: Closing the gap between users and technology. In 2012 AAAI Fall Symposium Series
2012
-
[21]
Timo Jakobi, Gunnar Stevens, Nico Castelli, Corinna Ogonowski, Florian Schaub, Nils Vindice, Dave Randall, Peter Tolmie, and Volker Wulf. 2018. Evolving Needs in IoT Control and Accountability: A Longitudinal Study on Smart Home Intelligibility. Proc. ACM Interact. Mob. Wearab...
2018 doi
-
[22]
Ellen Jiang, Kristen Olson, Edwin Toh, Alejandra Molina, Aaron Donsbach, Michael Terry, and Carrie J Cai. 2022. PromptMaker: Prompt-based Prototyping with Large Language Models. In Extended Abstracts of the 2022 CHI Conference on Human Factors in Computing Systems (CHI EA ’22)...
2022
-
[23]
Ellen Jiang, Edwin Toh, Alejandra Molina, Aaron Donsbach, Carrie J Cai, and Michael Terry. 2021. GenLine and GenForm: Two Tools for Interacting with Generative Language Models in a Code Editor. In The Adjunct Publication of the 34th Annual ACM Symposium on User Interface Softw...
2021
- [24]
-
[25]
Minsuk Kahng, Ian Tenney, Mahima Pushkarna, Michael Xieyang Liu, James Wexler, Emily Reif, Krystal Kallarackal, Minsuk Chang, Michael Terry, and Lucas Dixon. 2024. LLM Comparator: Interactive Analysis of Side-by-Side Evaluation of Large Language Models. IEEE Transactions on Vi...
2024
-
[27]
V Kastrinaki, M Zervakis, and K Kalaitzakis. 2003. A survey of video processing techniques for traffic applications. Image and Vision Computing 21, 4 (April 2003), 359–381. doi:10.1016/S0262-8856(03)00004-0
2003 doi
-
[29]
Neil Klingensmith, Joseph Bomber, and Suman Banerjee. 2014. Hot, cold and in between: enabling fine-grained environmental control in homes for efficiency and comfort. In Proceedings of the 5th international conference on Future energy systems (e-Energy ’14) . Association for C...
2014
-
[30]
Amy J. Ko, Robin Abraham, Laura Beckwith, Alan Blackwell, Margaret Burnett, Martin Erwig, Chris Scaffidi, Joseph Lawrance, Henry Lieberman, Brad Myers, Mary Beth Rosson, Gregg Rothermel, Mary Shaw, and Susan Wiedenbeck. 2011. The State of the Art in End-user Software Engineeri...
2011
-
[31]
Johannes Kunkel, Benedikt Loepp, and Jürgen Ziegler. 2017. A 3D Item Space Visualization for Presenting and Manipulating User Preferences in Collaborative Filtering. In Proceedings of the 22nd International Conference on Intelligent User Interfaces (IUI ’17) . Association for ...
2017
-
[32]
Stacey Kuznetsov and Eric Paulos. 2010. UpStream: motivating water conser- vation with low-cost water flow sensing and persuasive displays. In Proceed- ings of the SIGCHI Conference on Human Factors in Computing Systems (CHI ’10). Association for Computing Machinery, New York,...
2010
-
[33]
van Lamsweerde
A. van Lamsweerde. 2009. Requirements engineering : from system goals to UML models to software specifications . John Wiley & Sons, Ltd. https://thuvienso. hoasen.edu.vn/handle/123456789/9362
2009
-
[34]
Gierad Laput, Karan Ahuja, Mayank Goel, and Chris Harrison. 2018. Ubicoustics: Plug-and-Play Acoustic Activity Recognition. In Proceedings of the 31st Annual ACM Symposium on User Interface Software and Technology (UIST ’18). Association for Computing Machinery, New York, NY, ...
2018 doi
-
[35]
Lasecki, Jason Wiese, Robert Xiao, Jeffrey P
Gierad Laput, Walter S. Lasecki, Jason Wiese, Robert Xiao, Jeffrey P. Bigham, and Chris Harrison. 2015. Zensors: Adaptive, Rapidly Deployable, Human-Intelligent Sensor Feeds. In Proceedings of the 33rd Annual ACM Conference on Human Factors in Computing Systems (CHI ’15) . Ass...
2015
-
[36]
Gierad Laput, Yang Zhang, and Chris Harrison. 2017. Synthetic Sensors: Towards General-Purpose Sensing. In Proceedings of the 2017 CHI Conference on Human Factors in Computing Systems (CHI ’17) . Association for Computing Machin- ery, New York, NY, USA, 3986–3999. doi:10.1145/...
2017
-
[37]
Peter Lee, Sebastien Bubeck, and Joseph Petro. 2023. Benefits, Limits, and Risks of GPT-4 as an AI Chatbot for Medicine. New England Journal of Medicine 388, 13 (March 2023), 1233–1239. doi:10.1056/NEJMsr2214184 Publisher: Massachusetts Medical Society _eprint: https://www.nej...
2023 doi
-
[38]
Chunyuan Li, Zhe Gan, Zhengyuan Yang, Jianwei Yang, Linjie Li, Lijuan Wang, and Jianfeng Gao. 2024. Multimodal Foundation Models: From Specialists to General-Purpose Assistants. Found. Trends. Comput. Graph. Vis. 16, 1-2 (May 2024), 1–214. doi:10.1561/0600000110
2024 doi
-
[39]
Kane, and Patrick Carrington
Franklin Mingzhe Li, Michael Xieyang Liu, Shaun K. Kane, and Patrick Carrington
-
[40]
Lit. 2024. Lit Web Components. https://lit.dev/
2024
-
[41]
Michael Xieyang Liu. 2023. Tool Support for Knowledge Foraging, Structuring, and Transfer during Online Sensemaking . Ph. D. Dissertation. Carnegie Mellon University. http://reports-archive.adm.cs.cmu.edu/anon/anon/usr0/ftp/usr/ftp/ hcii/abstracts/23-105.html
2023
-
[42]
Michael Xieyang Liu, Jane Hsieh, Nathan Hahn, Angelina Zhou, Emily Deng, Shaun Burley, Cynthia Taylor, Aniket Kittur, and Brad A. Myers. 2019. Unakite: Scaffolding Developers’ Decision-Making Using the Web. In Proceedings of the 32Nd Annual ACM Symposium on User Interface Soft...
2019
-
[43]
Michael Xieyang Liu, Aniket Kittur, and Brad A. Myers. 2021. To Reuse or Not To Reuse? A Framework and System for Evaluating Summarized Knowledge. Proceedings of the ACM on Human-Computer Interaction 5, CSCW1 (April 2021), 166:1–166:35. doi:10.1145/3449240
2021 doi
-
[44]
Michael Xieyang Liu, Aniket Kittur, and Brad A. Myers. 2022. Crystalline: Low- ering the Cost for Developers to Collect and Organize Information for Decision Making. In Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems (CHI ’22). Association for Comp...
2022
-
[45]
We Need Structured Output
Michael Xieyang Liu, Frederick Liu, Alexander J. Fiannaca, Terry Koo, Lucas Dixon, Michael Terry, and Carrie J. Cai. 2024. "We Need Structured Output": Towards User-centered Constraints on Large Language Model Output. doi:10. 1145/3613905.3650756 arXiv:2404.07362 [cs]
2024
-
[46]
What It Wants Me To Say
Michael Xieyang Liu, Advait Sarkar, Carina Negreanu, Benjamin Zorn, Jack Williams, Neil Toronto, and Andrew D. Gordon. 2023. “What It Wants Me To Say”: Bridging the Abstraction Gap Between End-User Programmers and Code- Generating Large Language Models. In Proceedings of the 2...
2023
-
[47]
Michael Xieyang Liu, Tongshuang Wu, Tianying Chen, Franklin Mingzhe Li, Aniket Kittur, and Brad A. Myers. 2024. Selenite: Scaffolding Online Sense- making with Comprehensive Overviews Elicited from Large Language Models. In Proceedings of the 2024 CHI Conference on Human Facto...
2024
-
[48]
Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig. 2023. Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing. ACM Comput. Surv. 55, 9 (Jan. 2023), 195:1–195:35. doi:10.1145/3560815
2023 doi
-
[49]
Bertalan Meskó. 2023. Prompt Engineering as an Important Emerging Skill for Medical Professionals: Tutorial. Journal of Medical Internet Research 25, 1 (Oct. 2023), e50638. doi:10.2196/50638 Company: Journal of Medical Internet Research Distributor: Journal of Medical Internet...
2023 doi
-
[50]
Millett, M
L.I. Millett, M. Thomas, and D. Jackson. 2007. Software for Dependable Systems: Sufficient Evidence? National Academies Press. https://books.google.com/books? id=_aScAgAAQBAJ
2007
-
[51]
Changmo Nam, Jun-Cheol Park, and Dae-Shik Kim. 2016. Indoor Human Activity Recognition with Contextual Cues in Videos. In Proceedings of the 10th Interna- tional Conference on Ubiquitous Information Management and Communication (IMCOM ’16). Association for Computing Machinery,...
2016
- [52]
-
[53]
Sydney Nguyen, Hannah McLean Babe, Yangtian Zi, Arjun Guha, Carolyn Jane Anderson, and Molly Q Feldman. 2024. How Beginning Programmers and Code LLMs (Mis)read Each Other. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems (CHI ’24) . Association f...
2024
-
[54]
Mahda Noura, Sebastian Heil, and Martin Gaedke. 2020. Natural language goal understanding for smart home environments. In Proceedings of the 10th Interna- tional Conference on the Internet of Things (IoT ’20) . Association for Computing Machinery, New York, NY, USA, 1–8. doi:1...
2020
- [55]
-
[56]
Savvas Petridis, Nediyana Daskalova, Sarah Mennicken, Samuel F Way, Paul Lamere, and Jennifer Thom. 2022. TastePaths: Enabling Deeper Exploration and Understanding of Personal Preferences in Recommender Systems. In Pro- ceedings of the 27th International Conference on Intellig...
2022
-
[57]
Savvas Petridis, Nicholas Diakopoulos, Kevin Crowston, Mark Hansen, Keren Henderson, Stan Jastrzebski, Jeffrey V Nickerson, and Lydia B Chilton. 2023. AngleKindling: Supporting Journalistic Angle Ideation with Large Language Models. In Proceedings of the 2023 CHI Conference on...
2023
-
[58]
Fiannaca, Vivian Tsai, Michael Terry, and Carrie J
Savvas Petridis, Michael Xieyang Liu, Alexander J. Fiannaca, Vivian Tsai, Michael Terry, and Carrie J. Cai. 2024. In Situ AI Prototyping: Infusing Multimodal Prompts into Mobile Settings with MobileMaker. In 2024 IEEE Symposium on Visual Languages and Human-Centric Computing (...
2024
-
[59]
Savvas Petridis, Michael Terry, and Carrie Jun Cai. 2023. PromptInfuser: Bring- ing User Interface Mock-ups to Life with Large Language Models. In Extended Abstracts of the 2023 CHI Conference on Human Factors in Computing Systems (CHI EA ’23). Association for Computing Machin...
2023
- [60]
-
[61]
Savvas Petridis, Benjamin D Wedin, James Wexler, Mahima Pushkarna, Aaron Donsbach, Nitesh Goyal, Carrie J Cai, and Michael Terry. 2024. Constitution- Maker: Interactively Critiquing Large Language Models by Converting Feedback into Principles. In Proceedings of the 29th Intern...
2024
-
[62]
Wong, and Nick Merrill
James Pierce, Richmond Y. Wong, and Nick Merrill. 2020. Sensor Illumination: Exploring Design Qualities and Ethical Implications of Smart Cameras and Im- age/Video Analytics. In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems (CHI ’20) . Associatio...
2020
-
[63]
Cai, Michael Terry, and Minsuk Kahng
Crystal Qian, Michael Xieyang Liu, Emily Reif, Grady Simon, Nada Hussein, Nathan Clement, James Wexler, Carrie J. Cai, Michael Terry, and Minsuk Kahng
-
[64]
Ring. 2024. Security Cameras - Outdoor Wireless & Wired
2024
- [65]
-
[66]
Schmidt, Jesse Spencer-Smith, Quchen Fu, and Jules White
Douglas C. Schmidt, Jesse Spencer-Smith, Quchen Fu, and Jules White. 2024. Towards a Catalog of Prompt Patterns to Enhance the Discipline of Prompt Engineering. Ada Lett. 43, 2 (June 2024), 43–51. doi:10.1145/3672359.3672364
2024
-
[67]
Russell, Laura Koesten, Aniket Kittur, Nitesh Goyal, and Michael Xieyang Liu
Daniel M. Russell, Laura Koesten, Aniket Kittur, Nitesh Goyal, and Michael Xieyang Liu. 2024. Sensemaking: What is it today?. In Extended Ab- stracts of the 2024 CHI Conference on Human Factors in Computing Systems (CHI EA ’24) . Association for Computing Machinery, New York, ...
2024
-
[68]
doi:10.1145/3613905.3636322
- [69]
-
[70]
Bernheim Brush, John Krumm, Brian Meyers, Michael Hazas, Stephen Hodges, and Nicolas Villar
James Scott, A.J. Bernheim Brush, John Krumm, Brian Meyers, Michael Hazas, Stephen Hodges, and Nicolas Villar. 2011. PreHeat: controlling home heating using occupancy prediction. In Proceedings of the 13th international conference on Ubiquitous computing (UbiComp ’11) . Associ...
2011
- [71]
-
[72]
Mojtaba Vaismoradi, Hannele Turunen, and Terese Bondas. 2013. Content analy- sis and thematic analysis: Implications for conducting a qualitative descriptive study. Nursing & Health Sciences 15, 3 (2013), 398–405. doi:10.1111/nhs.12048 _eprint: https://onlinelibrary.wiley.com/...
2013 doi
-
[73]
Lei Shi, Maryam Ashoori, Yunfeng Zhang, and Shiri Azenkot. 2018. Knock knock, what’s there: converting passive objects into customizable smart controllers. In Proceedings of the 20th International Conference on Human-Computer Interaction with Mobile Devices and Services (Mobil...
2018
-
[74]
Song, Stephan J
Jean Y. Song, Stephan J. Lemmer, Michael Xieyang Liu, Shiyan Yan, Juho Kim, Jason J. Corso, and Walter S. Lasecki. 2019. Popup: reconstructing 3D video using particle filtering to aggregate crowd responses. In Proceedings of the 24th International Conference on Intelligent Use...
2019 doi
-
[75]
Jong-bum Woo and Youn-kyung Lim. 2015. User experience in do-it-yourself- style smart homes. In Proceedings of the 2015 ACM International Joint Conference on Pervasive and Ubiquitous Computing (UbiComp ’15). Association for Computing Machinery, New York, NY, USA, 779–790. doi:...
2015
-
[76]
Semantic programming by example with pre-trained models
Gust Verbruggen, Vu Le, and Sumit Gulwani. 2021. Semantic programming by example with pre-trained models. Replication Package for Article: "Semantic programming by example with pre-trained models" 5, OOPSLA (Oct. 2021), 100:1– 100:25. doi:10.1145/3485477
2021 doi
-
[77]
Carl Vondrick, Donald Patterson, and Deva Ramanan. 2013. Efficiently Scaling up Crowdsourced Video Annotation. International Journal of Computer Vision 101, 1 (Jan. 2013), 184–204. doi:10.1007/s11263-012-0564-1
2013 doi
-
[78]
Zamfirescu-Pereira, Heather Wei, Amy Xiao, Kitty Gu, Grace Jung, Matthew G Lee, Bjoern Hartmann, and Qian Yang
J.D. Zamfirescu-Pereira, Heather Wei, Amy Xiao, Kitty Gu, Grace Jung, Matthew G Lee, Bjoern Hartmann, and Qian Yang. 2023. Herding AI Cats: Lessons from De- signing a Chatbot by Prompting GPT-3. InProceedings of the 2023 ACM Designing Interactive Systems Conference (DIS ’23) ....
2023
-
[79]
Hui Ye and Hongbo Fu. 2022. ProGesAR: Mobile AR Prototyping for Proxemic and Gestural Interactions with Real-world IoT Enhanced Spaces. In Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems (CHI ’22) . Association for Computing Machinery, New York, NY...
2022
-
[80]
Sojeong Yun and Youn-Kyung Lim. 2023. Potential and Challenges of DIY Smart Homes with an ML-intensive Camera Sensor. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (CHI ’23) . Association for Computing Machinery, New York, NY, USA. doi:10.1145...
2023
- [81]
-
[82]
Zamfirescu-Pereira, Richmond Y
J.D. Zamfirescu-Pereira, Richmond Y. Wong, Bjoern Hartmann, and Qian Yang
-
[84]
Glassman
Tianyi Zhang, London Lowmanstone, Xinyu Wang, and Elena L. Glassman
-
[2020]
In Proceedings of the 33rd Annual ACM Symposium on User Interface Software and Technology (UIST ’20)
Interactive Program Synthesis by Augmented Examples. In Proceedings of the 33rd Annual ACM Symposium on User Interface Software and Technology (UIST ’20). Association for Computing Machinery, New York, NY, USA, 627–648. doi:10.1145/3379337.3415900
-
[2023]
In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (CHI ’23)
Why Johnny Can’t Prompt: How Non-AI Experts Try (and Fail) to Design LLM Prompts. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (CHI ’23) . Association for Computing Machinery, New York, NY, USA, 1–21. doi:10.1145/3544548.3581388
2023
-
[2024]
doi:10.1145/3613904.3642233 arXiv:2402.15108 [cs]
A Contextual Inquiry of People with Vision Impairments in Cooking. doi:10.1145/3613904.3642233 arXiv:2402.15108 [cs]
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.