REVIEW 4 major objections 6 minor 115 references
GLOSS: Group of LLMs for Open-Ended Sensemaking of Passive Sensing Data for Health and Wellbeing
T0 review · 4 major / 6 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read The paper claims that GLOSS, a group of collaborating LLM agents that generate and execute code over raw sensor databases, can answer open-ended health and wellbeing questions from passive sensing data far more accurately than standard…
desk verdict Interesting system, but the headline accuracy numbers rest on a self-scored evaluation that needs independent verification before they can be trusted. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing machinery is the two-loop agent architecture with a code-generation core rather than a single retrieval circuit. In the information-seeking loop, an agent that generates executable Python or bash code runs inside a sandbox, calls predefined helper functions that query raw per-stream databases (location, activity, app usage, steps, Wi-Fi, calls, battery, heart rate, and a stress model), and then triangulates across streams on demand. Because code, not the language model, does the arithmetic and joins, the system can answer computation-heavy queries accurately. In the sensemaking loop, a local agent converts raw outputs into natural language, a global agent maintains a cumulative "understanding" and notes what additional data is needed, and a next-step agent checks whether the understanding answers the query, with a five-iteration cutoff to avoid infinite loops. The final answer is formatted by a presentation agent that can tailor output to human-readable or machine-readable needs.
What would settle it
Have two independent researchers, blind to GLOSS's and RAG's outputs, write ground-truth code for the same 60 queries or a fresh sample and compare; if the independently computed accuracy for GLOSS does not remain far above RAG's, the claimed 87.93% versus 29.31% margin is not reproducible.
Extended reading notes
Core claim
The central claim is that open-ended sensemaking over passive sensing data can be automated by a group of cooperating LLM agents that fetch raw data, generate and execute code to process it, and iteratively build an understanding, rather than by retrieving pre-chunked data and generating text. GLOSS is the instantiation: an Action Plan agent decides whether and how the query can be answered, an Information Seeking agent formulates data requests, a Database/Model Manager selects retrieval helper functions, a Code Generation agent writes and runs Python or bash code that can triangulate multiple sensor streams, a Local Sensemaking agent renders results in natural language, a Global Sensemaking agent updates the accumulated understanding, a Next Step agent decides when to stop, and a Presentation agent formats the answer. On a crowd-sourced set of 244 objective, subjective, and mixed queries (with two participants per original query), GLOSS reports 87.93% accuracy on a random 60-query sample and 66.19% consistency across 210 queries, versus RAG's 29.31% accuracy and 52.85% consistency, with the accuracy and consistency differences reported as statistically significant. The authors attribute the gap to GLOSS performing calculations and triangulation in code rather than relying on the LLM's arithmetic, and they note that RAG frequently produced plausible but made-up numbers.
Load-bearing premise
The evaluation's ground truth for the 60 sampled objective and mixed queries was written by the first author alone as hand-written Python code, so the central accuracy comparison rests on that code being correct and unbiased.
Editorial extensions
If this is right
- For systems that answer health queries from raw sensor streams, code-generating agents over raw data should be preferred over text-retrieval RAG whenever queries involve arithmetic, thresholds, or joins across streams.
- Non-programming stakeholders, including clinicians, behavioral scientists, and self-trackers, could run open-ended analyses that previously required a data scientist or a custom dashboard.
- The intermediate action plans, memory, and understanding shown to users give GLOSS a transparency advantage over end-to-end LLM answers, supporting trust in research and clinical settings.
- The same core system can be embedded in automated pipelines: generating context-aware EMA questions, adding interpretability to black-box anomaly detectors, and producing personalized reflection narratives.
Reading between the lines
- Because the 60-query accuracy sample was ground-truthed by a single annotator, an independent replication with pre-registered ground-truth code is the natural next step before treating 87.93% as a stable number.
- The 66.19% consistency suggests a concrete improvement the paper mentions but does not test: inserting a clarification step for ambiguous queries could resolve cases where GLOSS answers correctly via different valid paths.
- The RAG baseline may be conservative, since it retrieves text chunks of raw data; a stronger RAG variant with schema-aware retrieval or precomputed aggregates might shrink the gap, and that comparison is not yet in the literature.
- If the result generalizes, the same agent-group pattern could be applied to other high-dimensional personal data, including images and audio, as the paper itself anticipates.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces GLOSS, a multi-agent LLM system for open-ended sensemaking of raw passive sensing data. GLOSS combines action-plan generation, an information-seeking/sensemaking loop, code generation over raw sensor databases, and a presentation agent. The authors evaluate GLOSS against a RAG baseline on 244 crowd-sourced queries, reporting that GLOSS achieves 87.93% accuracy and 66.19% consistency on objective/mixed queries versus RAG's 29.31% accuracy and 52.85% consistency, and they also report subjective ratings on relevance, interpretation, domain knowledge, logic, and clarity. Four use cases are presented: supporting non-CS researchers, generating reflection narratives, AI-triggered EMA prompting, and enhancing interpretability of black-box models.
Significance. If the central comparison is sound, GLOSS is a meaningful step toward making passive sensing data accessible to non-programmers, and the system design is well grounded in sensemaking theory. The paper ships public code and queries, which is a strength for reproducibility. However, the accuracy evaluation rests on ground truth written by the first author after seeing model outputs, and the RAG baseline is underspecified; these limitations directly affect the headline claim. The use cases are illustrative rather than fully evaluated, as the authors acknowledge, so the contribution should be judged primarily on the core system and its comparison.
major comments (4)
- [§4.2, Table 5] The accuracy ground truth is not independent: the first author wrote Python code 'based on the logic model followed' by GLOSS after the model runs. This can ratify plausible-but-wrong interpretations of ambiguous queries, such as defining 'mobility' as step count rather than GPS displacement. With only 60 queries and a reported gap of 58.6 percentage points, even a modest systematic scoring bias could materially change the conclusion. Please provide the scoring code, have at least one independent annotator re-derive ground truth for all 60 queries or a pre-registered subset, report inter-annotator agreement, and state how ambiguous queries were adjudicated.
- [§5.1, §3.2] Consistency is measured as 'identical across all three runs' using GPT-4o with temperature=1 and top_p=1. It is not clear whether 'identical' means verbatim string equality, normalized equality, or semantic equivalence. Under stochastic sampling, verbatim matching likely understates consistency for both systems. Please specify the matching criterion, report the distribution of consistency failures (e.g., changed wording versus changed numerical answer), and consider reporting semantic consistency with multiple samples.
- [§4.2] The RAG baseline is underspecified. The text states that raw sensor data were converted to natural language and stored in a Chroma database with LangChain's RAG framework, but it does not report chunk sizes, retrieval top-k, the prompt template, or whether retrieval was per sensor stream or over a combined corpus. Because the central claim is that GLOSS outperforms 'commonly used RAG,' the baseline needs enough detail to be reproduced and to be recognizable as a reasonable instantiation of RAG for this task.
- [§5.2] The subjective evaluation does not report the number of responses per condition, inter-coder reliability, or the full contingency tables behind the chi-square tests. The description that 'both response 1 and response 2 included a random number of GLOSS and RAG-generated responses' is ambiguous about whether each annotator saw exactly one response per query from each system. Please report Cohen's kappa, sample sizes, and collapsed category definitions; with 34 subjective queries, chi-square tests with df=1 may be fragile if categories were collapsed.
minor comments (6)
- [Title page / author block] The author block contains a typo: 'le.ha1@northeastern.edu,' has a stray comma, and Kaixin Ji's email is listed as 'li.jiachen4@northeastern.edu,' which appears to be a copy of Jiachen Li's email rather than a distinct address.
- [§1] The word 'anixety' should be 'anxiety' in the sentence 'passive mobile sensing has demonstrated significant potential in monitoring and assessing various health and well-being outcomes, including depression [27,99,104], stress [25,71,73,74,90,99], and anixety [96].'
- [§4.2] The phrase 'process & analyze the date to calculate accuracy' should read 'the data' rather than 'the date.'
- [§5.2.4] The heading 'Domain Knowlegde' is misspelled; it should be 'Domain Knowledge.'
- [Table 2] The legend for Table 2 is incomplete: the sentence fragment 'where indicates the presence of a particular capability in the given approach' is cut off and does not explain the symbol used in the 'Multi data streams' column.
- [Figure 4] Figure 4 shows subjective ratings without error bars or per-condition sample sizes; adding these would clarify the comparison and support the reported chi-square tests.
Circularity Check
The objective-query accuracy metric is self-referential: Section 4.2 defines accuracy as matching 'the underlying logic it followed' and builds the ground-truth code from that same logic, so the headline 87.93% vs. 29.31% gap is not an independent measurement.
-
self definitional
[Section 4.2, Evaluation Metrics]
"Accuracy is measured by evaluating whether the model retrieved the correct answer based on the underlying logic it followed. ... Thus, to evaluate the accuracy of models, we needed to manually write Python code and analyze the data based on the logic model followed."
The correctness criterion is defined by the very logic the system chose to follow. The first author writes the scoring code 'based on the logic model followed,' so any answer that faithfully executes the system's own interpretation (e.g., mobility as step count versus GPS displacement) is marked correct by construction rather than checked against an independent ground truth. The paper itself acknowledges that objective queries can have 'multiple valid logical paths' and that GLOSS 'followed different logical paths in different runs.' Consequently, the reported 87.93% accuracy measures internal agreement between the system's output and a re-implementation of that same logic, not agreement with an external answer.
full rationale
Most of this systems paper is not circular: there is no fitted equation, no derivation chain, and no imported uniqueness theorem. The four use cases are demonstrations, not predictions. The self-citations ([19], [61], [74]) are used for background, related work, and a stress ML model that is a data stream, not a load-bearing justification of the headline claim. The consistency metric is independently defined as identical responses across three runs, and the subjective evaluation uses blinded annotators with predefined criteria. However, the objective-query accuracy evaluation, which carries the central quantitative claim, is self-referential: Section 4.2 defines accuracy as whether the model retrieved the correct answer 'based on the underlying logic it followed' and then has the first author write the ground-truth code 'based on the logic model followed.' This means the scoring code is derived from the system's own chosen interpretation, so the accuracy number is partly guaranteed by construction rather than by independent verification. Because the headline comparison with RAG depends on this accuracy metric, the paper exhibits partial circularity in its central evaluation, though the subjective evaluation and consistency results retain independent content.
Assumptions & free parameters
assumptions (4)
- domain assumption Converting code outputs to natural language preserves enough information for the global sensemaking agent to reason correctly.
- ad hoc to paper The five-iteration loop cutoff is sufficient for the sensemaking process to reach accurate answers.
- domain assumption The RAG baseline is a representative implementation of prior retrieval-based passive-sensing systems.
- domain assumption Manually coded ground truth by the first author is correct for the 60 sampled objective queries.
Cite this review
Pith. "Pith review of GLOSS: Group of LLMs for Open-Ended Sensemaking of Passive Sensing Data for Health and Wellbeing." pith.science (2026). https://pith.science/paper/UPVMFPM4
@misc{pith2026250705461,
author = {Pith},
title = {Pith review of: GLOSS: Group of LLMs for Open-Ended Sensemaking of Passive Sensing Data for Health and Wellbeing},
year = {2026},
howpublished = {\url{https://pith.science/paper/UPVMFPM4}},
note = {Machine review of arXiv:2507.05461}
}
read the original abstract
The ubiquitous presence of smartphones and wearables has enabled researchers to build prediction and detection models for various health and behavior outcomes using passive sensing data from these devices. Achieving a high-level, holistic understanding of an individual's behavior and context, however, remains a significant challenge. Due to the nature of passive sensing data, sensemaking -- the process of interpreting and extracting insights -- requires both domain knowledge and technical expertise, creating barriers for different stakeholders. Existing systems designed to support sensemaking are either not open-ended or cannot perform complex data triangulation. In this paper, we present a novel sensemaking system, Group of LLMs for Open-ended Sensemaking (GLOSS), capable of open-ended sensemaking and performing complex multimodal triangulation to derive insights. We demonstrate that GLOSS significantly outperforms the commonly used Retrieval-Augmented Generation (RAG) technique, achieving 87.93% accuracy and 66.19% consistency, compared to RAG's 29.31% accuracy and 52.85% consistency. Furthermore, we showcase the promise of GLOSS through four use cases inspired by prior and ongoing work in the UbiComp and HCI communities. Finally, we discuss the potential of GLOSS, its broader implications, and the limitations of our work.
Figures
Figures from the paper (5 more)
Reference graph
Works this paper leans on
-
[1]
Sensing technologies for monitoring serious mental illnesses
Saeed Abdullah and Tanzeem Choudhury. Sensing technologies for monitoring serious mental illnesses. IEEE MultiMedia, 25(1):61–75,
-
[2]
Adler, Yuewen Yang, Thalia Viranda, Xuhai Xu, David C
Daniel A. Adler, Yuewen Yang, Thalia Viranda, Xuhai Xu, David C. Mohr, Anna R. Van Meter, Julia C. Tartaglia, Nicholas C. Jacobson, Fei Wang, Deborah Estrin, and Tanzeem Choudhury. Beyond detection: Towards actionable sensing research in clinical mental healthcare. Proc. ACM Interact. Mob. Wearable Ubiquitous Technol. , 8(4), November 2024. Number of page...
2024
-
[3]
Mindful-RAG: A Study of Points of Failure in Retrieval Augmented Generation, October 2024
Garima Agrawal, Tharindu Kumarage, Zeyad Alghamdi, and Huan Liu. Mindful-RAG: A Study of Points of Failure in Retrieval Augmented Generation, October 2024. arXiv:2407.12216 [cs]
arXiv 2024
-
[4]
Mobile and wearable sensors for data-driven health monitoring system: State-of-the-art and future prospect
Chioma Virginia Anikwe, Henry Friday Nweke, Anayo Chukwu Ikegwu, Chukwunonso Adolphus Egwuonwu, Fergus Uchenna Onu, Uzoma Rita Alo, and Ying Wah Teh. Mobile and wearable sensors for data-driven health monitoring system: State-of-the-art and future prospect. Expert Systems with Applications, 202:117362, 2022. Publisher: Elsevier
2022
-
[5]
Frustrated and confused: the American public rates its cancer-related information-seeking experiences
Neeraj K Arora, Bradford W Hesse, Barbara K Rimer, Kasisomayajula Viswanath, Marla L Clayman, and Robert T Croyle. Frustrated and confused: the American public rates its cancer-related information-seeking experiences. Journal of general internal medicine , 23:223–228, 2008. Publisher: Springer
2008
-
[6]
Abandonment of personal quantification: A review and empirical study investigating reasons for wearable activity tracking attrition
Christiane Attig and Thomas Franke. Abandonment of personal quantification: A review and empirical study investigating reasons for wearable activity tracking attrition. Computers in Human Behavior , 102:223–237, 2020. Publisher: Elsevier
2020
-
[7]
Ling Bao and Stephen S. Intille. Activity recognition from user-annotated acceleration data. In Takeo Kanade, Josef Kittler, Jon M. Kleinberg, Friedemann Mattern, John C. Mitchell, Oscar Nierstrasz, C. Pandu Rangan, Bernhard Steffen, Demetri Terzopoulos, Dough , Vol. 1, No. 1, Article . Publication date: September 2025. GLOSS: Group of LLMs for Open-Ended...
2025
-
[8]
The interpretation and exploitation of information in criminal investigations
Emma Caroline Barrett. The interpretation and exploitation of information in criminal investigations . phd, University of Birmingham, 2009
2009
Show all 115 references
-
[9]
Reviewing reflection: on the use of reflection in interactive system design
Eric PS Baumer, Vera Khovanskaya, Mark Matthews, Lindsay Reynolds, Victoria Schwanda Sosik, and Geri Gay. Reviewing reflection: on the use of reflection in interactive system design. In Proceedings of the 2014 conference on Designing interactive systems , pages 93–102, 2014
2014
-
[10]
Revisiting reflection in hci: Four design resources for technologies that support reflection
Marit Bentvelzen, Paweł W Woźniak, Pia SF Herbes, Evropi Stefanidi, and Jasmin Niess. Revisiting reflection in hci: Four design resources for technologies that support reflection. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies , 6(1):1–27, ...
2022
-
[11]
Detecting stress in college freshman from wearable sleep data
Laura Bloomfield, Mikaela Fudoligb, Peter Doddsb, and Llorinb Jordan. Detecting stress in college freshman from wearable sleep data
-
[12]
Go and grow: Mapping personal data to a living plant
Fadi Botros, Charles Perin, Bon Adriel Aseniero, and Sheelagh Carpendale. Go and grow: Mapping personal data to a living plant. In Proceedings of the international working conference on advanced visual interfaces , pages 112–119, 2016
2016
-
[13]
Patient engagement in a multimodal digital phenotyping study of opioid use disorder
Cynthia I Campbell, Ching-Hua Chen, Sara R Adams, Asma Asyyed, Ninad R Athale, Monique B Does, Saeed Hassanpour, Emily Hichborn, Melanie Jackson-Morris, Nicholas C Jacobson, and others. Patient engagement in a multimodal digital phenotyping study of opioid use disorder. Journa...
2023
-
[14]
Challenges and recommendations for wearable devices in digital health: Data quality, interoperability, health equity, fairness
Stefano Canali, Viola Schiaffonati, and Andrea Aliverti. Challenges and recommendations for wearable devices in digital health: Data quality, interoperability, health equity, fairness. PLOS Digital Health, 1(10):e0000104, 2022. Publisher: Public Library of Science San Francisc...
2022
-
[15]
Methods in predictive techniques for mental health status on social media: a critical review
Stevie Chancellor and Munmun De Choudhury. Methods in predictive techniques for mental health status on social media: a critical review. NPJ digital medicine, 3(1):43, 2020. Publisher: Nature Publishing Group UK London
2020
-
[16]
Sensor2Text: Enabling natural language interactions for daily activity tracking using wearable sensors
Wenqiang Chen, Jiaxuan Cheng, Leyao Wang, Wei Zhao, and Wojciech Matusik. Sensor2Text: Enabling natural language interactions for daily activity tracking using wearable sensors. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies , 8(4):1–26, 20...
2024
-
[17]
Understanding client support strategies to improve clinical outcomes in an online mental health intervention
Prerna Chikersal, Danielle Belgrave, Gavin Doherty, Angel Enrique, Jorge E Palacios, Derek Richards, and Anja Thieme. Understanding client support strategies to improve clinical outcomes in an online mental health intervention. In Proceedings of the 2020 CHI conference on huma...
2020
-
[18]
Understanding quantified-selfers’ practices in collecting and exploring personal data
Eun Kyoung Choe, Nicole B Lee, Bongshin Lee, Wanda Pratt, and Julie A Kientz. Understanding quantified-selfers’ practices in collecting and exploring personal data. In Proceedings of the SIGCHI conference on human factors in computing systems , pages 1143–1152, 2014
2014
-
[19]
Imputation Matters: A Deeper Look into an Overlooked Step in Longitudinal Health and Behavior Sensing Research, December 2024
Akshat Choube, Rahul Majethia, Sohini Bhattacharya, Vedant Das Swain, Jiachen Li, and Varun Mishra. Imputation Matters: A Deeper Look into an Overlooked Step in Longitudinal Health and Behavior Sensing Research, December 2024. arXiv:2412.06018 [stat]
2024 arXiv
-
[20]
SeSaMe: a framework to simulate self-reported ground truth for mental health sensing studies
Akshat Choube, Vedant Das Swain, and Varun Mishra. SeSaMe: a framework to simulate self-reported ground truth for mental health sensing studies. arXiv preprint arXiv:2403.17219, 2024
2024 arXiv
-
[21]
Health information needs, sources, and barriers of primary care patients to achieve patient-centered care: A literature review
Martina A Clarke, Joi L Moore, Linsey M Steege, Richelle J Koopman, Jeffery L Belden, Shannon M Canfield, Susan E Meadows, Susan G Elliott, and Min Soon Kim. Health information needs, sources, and barriers of primary care patients to achieve patient-centered care: A literature...
2016
-
[22]
Grounded theory research: Procedures, canons, and evaluative criteria
Juliet M Corbin and Anselm Strauss. Grounded theory research: Procedures, canons, and evaluative criteria. Qualitative sociology, 13(1):3–21, 1990. Publisher: Springer
1990
-
[23]
Furlotte, Zhun Yang, Chace Lee, Erik Schenck, Yojan Patel, Jian Cui, Logan Douglas Schneider, Robby Bryant, Ryan G
Justin Cosentino, Anastasiya Belyaeva, Xin Liu, Nicholas A. Furlotte, Zhun Yang, Chace Lee, Erik Schenck, Yojan Patel, Jian Cui, Logan Douglas Schneider, Robby Bryant, Ryan G. Gomes, Allen Jiang, Roy Lee, Yun Liu, Javier Perez, Jameson K. Rogers, Cathy Speed, Shyam Tailor, Meg...
2024 arXiv
-
[24]
Data sensemaking in self-tracking: Towards a new generation of self-tracking tools
Aykut Coşkun and Armağan Karahanoğlu. Data sensemaking in self-tracking: Towards a new generation of self-tracking tools. International Journal of Human–Computer Interaction , 39(12):2339–2360, 2023. Publisher: Taylor & Francis
2023
-
[25]
Semantic gap in predicting mental wellbeing through passive sensing
Vedant Das Swain, Victor Chen, Shrija Mishra, Stephen M Mattingly, Gregory D Abowd, and Munmun De Choudhury. Semantic gap in predicting mental wellbeing through passive sensing. In Proceedings of the 2022 CHI conference on human factors in computing systems , 2022
2022
-
[26]
Benchmarks for automated commonsense reasoning: A survey
Ernest Davis. Benchmarks for automated commonsense reasoning: A survey. ACM Computing Surveys, 56(4):1–41, 2023. Publisher: ACM New York, NY, USA
2023
-
[27]
Predicting depression via social media
Munmun De Choudhury, Michael Gamon, Scott Counts, and Eric Horvitz. Predicting depression via social media. Proceedings of the International AAAI Conference on Web and Social Media , 7(1):128–137, August 2021
2021
-
[28]
From the mind’s eye of the “user”: The sense-making qualitative-quantitative methodology
Brenda Dervin. From the mind’s eye of the “user”: The sense-making qualitative-quantitative methodology. Qualitative Research in Information Management/Libraries Unlimited, 1992. , Vol. 1, No. 1, Article . Publication date: September 2025. 28 • Akshat Choube, Ha Le, Jiachen Li...
1992
-
[29]
An overview of sense-making research: Concepts, methods, and results to date
Brenda Dervin and others. An overview of sense-making research: Concepts, methods, and results to date. 1983. Publisher: Author
1983
-
[30]
The role of feedback in the process of health behavior change
Carlo C DiClemente, Angela S Marinilli, Manu Singh, and Lori E Bellino. The role of feedback in the process of health behavior change. American journal of health behavior , 25(3):217–227, 2001. Publisher: PNG Publications and Scientific Research Limited
2001
-
[31]
Deep neural networks in psychiatry
Daniel Durstewitz, Georgia Koppe, and Andreas Meyer-Lindenberg. Deep neural networks in psychiatry. Molecular psychiatry, 24(11):1583–1598, 2019. Publisher: Nature Publishing Group UK London
2019
-
[32]
Imputation of missing longitudinal data: a comparison of methods
J Engels. Imputation of missing longitudinal data: a comparison of methods. Journal of Clinical Epidemiology , 56(10):968–976, October 2003
2003
-
[33]
Epstein, Clara Caldeira, Mayara Costa Figueiredo, Xi Lu, Lucas M
Daniel A. Epstein, Clara Caldeira, Mayara Costa Figueiredo, Xi Lu, Lucas M. Silva, Lucretia Williams, Jong Ho Lee, Qingyang Li, Simran Ahuja, Qiuer Chen, Payam Dowlatyari, Craig Hilby, Sazeda Sultana, Elizabeth V. Eikey, and Yunan Chen. Mapping and Taking Stock of the Personal...
2020
-
[34]
Physiollm: Supporting personalized health insights with wearables and large language models
Cathy Mengying Fang, Valdemar Danry, Nathan Whitmore, Andria Bao, Andrew Hutchison, Cayden Pierce, and Pattie Maes. Physiollm: Supporting personalized health insights with wearables and large language models. arXiv preprint arXiv:2406.19283, 2024
2024 arXiv
-
[35]
Mobile sensing for behavioral research: A component-based approach for rapid deployment of sensing campaigns
Ivan R Felix, Luis A Castro, Luis-Felipe Rodriguez, and Oresti Banos. Mobile sensing for behavioral research: A component-based approach for rapid deployment of sensing campaigns. International Journal of Distributed Sensor Networks , 15(9):1550147719874186,
-
[36]
Blue versus gray: A metaphor constraining sensemaking around a restructuring
Danna N Greenberg. Blue versus gray: A metaphor constraining sensemaking around a restructuring. Group & Organization Management, 20(2):183–209, 1995. Publisher: Sage Publications Sage CA: Thousand Oaks, CA
1995
-
[37]
Fostering engagement with personal informatics systems
Rebecca Gulotta, Jodi Forlizzi, Rayoung Yang, and Mark Wah Newman. Fostering engagement with personal informatics systems. In Proceedings of the 2016 ACM conference on designing interactive systems , pages 286–300, 2016
2016
-
[38]
Describing the elephant: a foundational model of human needs, motivation, behaviour, and wellbeing, December 2020
Andreas Habermacher, Argang Ghadiri, and Theo Peters. Describing the elephant: a foundational model of human needs, motivation, behaviour, and wellbeing, December 2020
2020
-
[39]
Systematic Evaluation of Personalized Deep Learning Models for Affect Recognition
Yunjo Han, Panyu Zhang, Minseo Park, and Uichin Lee. Systematic Evaluation of Personalized Deep Learning Models for Affect Recognition. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies , 8(4):1–35, November 2024
2024
-
[40]
Design of digital workplace stress-reduction intervention systems: Effects of intervention type and timing
Esther Howe, Jina Suh, Mehrab Bin Morshed, Daniel McDuff, Kael Rowan, Javier Hernandez, Marah Ihab Abdin, Gonzalo Ramos, Tracy Tran, and Mary P Czerwinski. Design of digital workplace stress-reduction intervention systems: Effects of intervention type and timing. In Proceeding...
2022
-
[41]
Large language models in mental health care: a scoping review
Yining Hua, Fenglin Liu, Kailai Yang, Zehan Li, Hongbin Na, Yi-han Sheu, Peilin Zhou, Lauren V Moran, Sophia Ananiadou, Andrew Beam, and others. Large language models in mental health care: a scoping review. arXiv preprint arXiv:2401.02984, 2024
2024 arXiv
-
[42]
Insufficient effort responding: examining an insidious confound in survey data
Jason L Huang, Mengqiao Liu, and Nathan A Bowling. Insufficient effort responding: examining an insidious confound in survey data. Journal of Applied Psychology, 100(3):828, 2015. Publisher: American Psychological Association
2015
-
[43]
Jacobson, Damien Lekkas, Raphael Huang, and Natalie Thomas
Nicholas C. Jacobson, Damien Lekkas, Raphael Huang, and Natalie Thomas. Deep learning paired with wearable passive sensing data predicts deterioration in anxiety disorder symptoms across 17–18 years. Journal of Affective Disorders , 282:104–111, March 2021
2021
-
[44]
Multimodal Autoencoder: A Deep Learning Approach to Filling In Missing Sensor Data and Enabling Better Mood Prediction
Natasha Jaques, Sara Taylor, Akane Sano, and Rosalind Picard. Multimodal Autoencoder: A Deep Learning Approach to Filling In Missing Sensor Data and Enabling Better Mood Prediction
-
[45]
HARGPT: Are LLMs Zero-Shot Human Activity Recognizers? pages 38–43
Sijie Ji, Xinzhe Zheng, and Chenshu Wu. HARGPT: Are LLMs Zero-Shot Human Activity Recognizers? pages 38–43. IEEE Computer Society, May 2024
2024
-
[46]
DT2VIS: A Focus+Context Answer Generation System to Facilitate Visual Exploration of Tabular Data
Qi Jiang, Guodao Sun, Yue Dong, and Ronghua Liang. DT2VIS: A Focus+Context Answer Generation System to Facilitate Visual Exploration of Tabular Data
-
[47]
Designing for data sensemaking practices: a complex challenge
Armağan Karahanoğlu and Aykut Coşkun. Designing for data sensemaking practices: a complex challenge. Interactions, 31(4):28–31,
-
[48]
Drawbacks of artificial intelligence and their potential solutions in the healthcare sector.Biomedical Materials & Devices, 1(2):731–738, 2023
Bangul Khan, Hajira Fatima, Ayatullah Qureshi, Sanjay Kumar, Abdul Hanan, Jawad Hussain, and Saad Abdullah. Drawbacks of artificial intelligence and their potential solutions in the healthcare sector.Biomedical Materials & Devices, 1(2):731–738, 2023. Publisher: Springer
2023
-
[49]
Health-LLM: Large Language Models for Health Prediction via Wearable Sensor Data, April 2024
Yubin Kim, Xuhai Xu, Daniel McDuff, Cynthia Breazeal, and Hae Won Park. Health-LLM: Large Language Models for Health Prediction via Wearable Sensor Data, April 2024. arXiv:2401.06866 [cs]
2024 arXiv
-
[50]
A data–frame theory of sensemaking
Gary Klein, Jennifer K Phillips, Erica L Rall, and Deborah A Peluso. A data–frame theory of sensemaking. In Expertise out of context , pages 118–160. Psychology Press, 2007
2007
-
[51]
Intelligence essentials for everyone
Lisa Krizan. Intelligence essentials for everyone . Joint Military Intelligence College, 1999. Number: 6
1999
-
[52]
Properties and Challenges of LLM-Generated Explanations, February 2024
Jenny Kunz and Marco Kuhlmann. Properties and Challenges of LLM-Generated Explanations, February 2024. arXiv:2402.10532 [cs]
2024 arXiv
-
[53]
Why we use and abandon smart devices
Amanda Lazar, Christian Koehler, Theresa Jean Tanenbaum, and David H Nguyen. Why we use and abandon smart devices. In Proceedings of the 2015 ACM international joint conference on pervasive and ubiquitous computing , pages 635–646, 2015
2015
-
[54]
Collecting self-reported physical activity and posture data using audio-based ecological momentary assessment
Ha Le, Rithika Lakshminarayanan, Jixin Li, Varun Mishra, and Stephen Intille. Collecting self-reported physical activity and posture data using audio-based ecological momentary assessment. Proc. ACM Interact. Mob. Wearable Ubiquitous Technol. , 8(3), September 2024. Number of ...
2024
-
[55]
Feasibility and Utility of Multimodal Micro Ecological Momentary Assessment on a Smartwatch
Ha Le, Veronika Potter, Rithika Lakshminarayanan, Varun Mishra, and Stephen Intille. Feasibility and Utility of Multimodal Micro Ecological Momentary Assessment on a Smartwatch. CHI Conference on Human Factors in Computing Systems (CHI ’25) , 2025
2025
-
[56]
Rossi, and Xiang Chen
Younghun Lee, Sungchul Kim, Tong Yu, Ryan A. Rossi, and Xiang Chen. Learning to Reduce: Optimal Representations of Structured Data in Prompting Large Language Models, February 2024. arXiv:2402.14195 [cs]
2024 arXiv
-
[57]
Retrieval-augmented generation for knowledge-intensive NLP tasks
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. Retrieval-augmented generation for knowledge-intensive NLP tasks. In Proceedings of the...
2020
-
[58]
Solving Quantitative Reasoning Problems with Language Models
Aitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer, Henryk Michalewski, Vinay Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman-Solo, Yuhuai Wu, Behnam Neyshabur, Guy Gur-Ari, and Vedant Misra. Solving Quantitative Reasoning Problems with Language Models. ...
2022
-
[59]
A stage-based model of personal informatics systems
Ian Li, Anind Dey, and Jodi Forlizzi. A stage-based model of personal informatics systems. In Proceedings of the SIGCHI conference on human factors in computing systems , pages 557–566, 2010
2010
-
[60]
Understanding my data, myself: supporting self-reflection with ubicomp technologies
Ian Li, Anind K Dey, and Jodi Forlizzi. Understanding my data, myself: supporting self-reflection with ubicomp technologies. In Proceedings of the 13th international conference on Ubiquitous computing , pages 405–414, 2011
2011
-
[61]
Vital insight: Assisting experts’ sensemaking process of multi-modal personal tracking data using visualization and LLM
Jiachen Li, Justin Steinberg, Xiwen Li, Akshat Choube, Bingsheng Yao, Dakuo Wang, Elizabeth Mynatt, and Varun Mishra. Vital insight: Assisting experts’ sensemaking process of multi-modal personal tracking data using visualization and LLM. arXiv preprint arXiv:2410.14879, 2024
-
[62]
Starcoder: may the source be with you! arXiv preprint arXiv:2305.06161, 2023
Raymond Li, Loubna Ben Allal, Yangtian Zi, Niklas Muennighoff, Denis Kocetkov, Chenghao Mou, Marc Marone, Christopher Akiki, Jia Li, Jenny Chim, and others. Starcoder: may the source be with you! arXiv preprint arXiv:2305.06161, 2023
2023 arXiv
-
[63]
Just-in-time but not too much: Determining treatment timing in mobile health.Proceedings of the ACM on interactive, mobile, wearable and ubiquitous technologies, 2(4):1–21, 2018
Peng Liao, Walter Dempsey, Hillol Sarker, Syed Monowar Hossain, Mustafa Al’Absi, Predrag Klasnja, and Susan Murphy. Just-in-time but not too much: Determining treatment timing in mobile health.Proceedings of the ACM on interactive, mobile, wearable and ubiquitous technologies,...
2018
-
[64]
Starcoder 2 and the stack v2: The next generation
Anton Lozhkov, Raymond Li, Loubna Ben Allal, Federico Cassano, Joel Lamy-Poirier, Nouamane Tazi, Ao Tang, Dmytro Pykhtar, Jiawei Liu, Yuxiang Wei, and others. Starcoder 2 and the stack v2: The next generation. arXiv preprint arXiv:2402.19173, 2024
2024 arXiv
-
[65]
Joint Modeling of Heterogeneous Sensing Data for Depression Assessment via Multi-task Learning
Jin Lu, Chao Shang, Chaoqun Yue, Reynaldo Morillo, Shweta Ware, Jayesh Kamath, Athanasios Bamis, Alexander Russell, Bing Wang, and Jinbo Bi. Joint Modeling of Heterogeneous Sensing Data for Depression Assessment via Multi-task Learning. Proceedings of the ACM on Interactive, M...
2018
-
[66]
Breathing as Input Modality in a Gameful Breathing Training App: Development and Evaluation of Breeze 2 (Preprint)
Yanick Xavier Lukic, Gisbert Wilhelm Teepe, Elgar Fleisch, and Tobias Kowatsch. Breathing as Input Modality in a Gameful Breathing Training App: Development and Evaluation of Breeze 2 (Preprint). preprint, JMIR Serious Games, May 2022
2022
-
[67]
Accessible visualization via natural language descriptions: A four-level model of semantic content
Alan Lundgard and Arvind Satyanarayan. Accessible visualization via natural language descriptions: A four-level model of semantic content. IEEE Transactions on Visualization and Computer Graphics , 28(1):1073–1083, January 2022
2022
-
[68]
Language models of code are few-shot commonsense learners
Aman Madaan, Shuyan Zhou, Uri Alon, Yiming Yang, and Graham Neubig. Language models of code are few-shot commonsense learners. arXiv preprint arXiv:2210.07128, 2022
2022 arXiv
-
[69]
Personal discovery in diabetes self-management: discovering cause and effect using self-monitoring data
Lena Mamykina, Elizabeth M Heitkemper, Arlene M Smaldone, Rita Kukafka, Heather J Cole-Lewis, Patricia G Davidson, Elizabeth D Mynatt, Andrea Cassells, Jonathan N Tobin, and George Hripcsak. Personal discovery in diabetes self-management: discovering cause and effect using sel...
2017
-
[70]
Transforming wearable data into health insights using large language model agents.arXiv preprint arXiv:2406.06464, 2024
Mike A Merrill, Akshay Paruchuri, Naghmeh Rezaei, Geza Kovacs, Javier Perez, Yun Liu, Erik Schenck, Nova Hammerquist, Jake Sunshine, Shyam Tailor, and others. Transforming wearable data into health insights using large language model agents.arXiv preprint arXiv:2406.06464, 2024
2024 arXiv
-
[71]
Walter, Marion J
Varun Mishra, Tian Hao, Si Sun, Kimberly N. Walter, Marion J. Ball, Ching-Hua Chen, and Xinxin Zhu. Investigating the Role of Context in Perceived Stress Detection in the Wild. In Proceedings of the 2018 ACM International Joint Conference and 2018 International Symposium on Pe...
2018
-
[72]
Detecting receptivity for mHealth interventions in the natural environment
Varun Mishra, Florian Künzler, Jan-Niklas Kramer, Elgar Fleisch, Tobias Kowatsch, and David Kotz. Detecting receptivity for mHealth interventions in the natural environment. Proceedings of the ACM on interactive, mobile, wearable and ubiquitous technologies , 5(2):1–24,
-
[73]
Continuous Detection of Physiological Stress with Commodity Hardware
Varun Mishra, Gunnar Pope, Sarah Lord, Stephanie Lewia, Byron Lowens, Kelly Caine, Sougata Sen, Ryan Halter, and David Kotz. Continuous Detection of Physiological Stress with Commodity Hardware. ACM Transactions on Computing for Healthcare (HEALTH) , 1(2):1–30, April 2020. Pub...
2020
-
[74]
Evaluating the reproducibility of physiological stress detection models
Varun Mishra, Sougata Sen, Grace Chen, Tian Hao, Jeffrey Rogers, Ching-Hua Chen, and David Kotz. Evaluating the reproducibility of physiological stress detection models. Proceedings of the ACM on interactive, mobile, wearable and ubiquitous technologies , 4(4):1–29,
-
[75]
Personal sensing: understanding mental health using ubiquitous sensors and machine learning
David C Mohr, Mi Zhang, and Stephen M Schueller. Personal sensing: understanding mental health using ubiquitous sensors and machine learning. Annual review of clinical psychology, 13:23–47, 2017. Publisher: Annual Reviews. , Vol. 1, No. 1, Article . Publication date: September...
2017
-
[76]
Momentary stressor logging and reflective visualizations: Implications for stress management with wearables
Sameer Neupane, Mithun Saha, Nasir Ali, Timothy Hnat, Shahin Alan Samiei, Anandatirtha Nandugudi, David M Almeida, and Santosh Kumar. Momentary stressor logging and reflective visualizations: Implications for stress management with wearables. In Proceedings of the CHI conferen...
2024
-
[77]
Supporting meaningful personal fitness: The tracker goal evolution model
Jasmin Niess and Paweł W Woźniak. Supporting meaningful personal fitness: The tracker goal evolution model. In Proceedings of the 2018 CHI conference on human factors in computing systems , pages 1–12, 2018
2018
-
[78]
On the role of active memory processes in perception and cognition
Donald A Norman and Daniel Gureasko Bobrow. On the role of active memory processes in perception and cognition . Center for Human Information Processing, Department of Psychology . . . , 1975
1975
-
[79]
OpenAI documentation: Prompt engineering
OpenAI. OpenAI documentation: Prompt engineering
-
[80]
GPT-4o System Card, October 2024
OpenAI. GPT-4o System Card, October 2024. arXiv:2410.21276 [cs]
2024 arXiv
-
[81]
LLMSense: Harnessing LLMs for High-level Reasoning Over Spatiotemporal Sensor Traces, March 2024
Xiaomin Ouyang and Mani Srivastava. LLMSense: Harnessing LLMs for High-level Reasoning Over Spatiotemporal Sensor Traces, March 2024. arXiv:2403.19857 [cs]
2024 arXiv
-
[82]
The sensemaking process and leverage points for analyst technology as identified through cognitive task analysis
Peter Pirolli and Stuart Card. The sensemaking process and leverage points for analyst technology as identified through cognitive task analysis. In Proceedings of international conference on intelligence analysis , volume 5, pages 2–4. McLean, VA, USA, 2005
2005
-
[83]
Behavioral indicators on a mobile sensing platform predict clinically validated psychiatric symptoms of mood and anxiety disorders
Skyler Place, Danielle Blanch-Hartigan, Channah Rubin, Cristina Gorrostieta, Caroline Mead, John Kane, Brian P Marx, Joshua Feast, Thilo Deckersbach, Alex “Sandy” Pentland, and others. Behavioral indicators on a mobile sensing platform predict clinically validated psychiatric ...
2017
-
[84]
Clear, and Peter Wright
Aare Puussaar, Adrian K. Clear, and Peter Wright. Enhancing Personal Informatics Through Social Sensemaking. In Proceedings of the 2017 CHI Conference on Human Factors in Computing Systems , CHI ’17, pages 6936–6942, New York, NY, USA, May 2017. Association for Computing Machinery
2017
-
[85]
Building shared understanding in collaborative sensemaking
Yan Qu and Derek L Hansen. Building shared understanding in collaborative sensemaking. In Proceedings of CHI 2008 sensemaking workshop, 2008
2008
-
[86]
Lee, Ashley Garrity, and Mark W
Shriti Raj, Joyce M. Lee, Ashley Garrity, and Mark W. Newman. Clinical data in context: Towards sensemaking tools for interpreting personal health data. Proc. ACM Interact. Mob. Wearable Ubiquitous Technol., 3(1):1–20, 2019. Publisher: ACM New York, NY, USA
2019
-
[87]
CataractBot: An LLM-powered expert-in-the-loop chatbot for cataract patients
Pragnya Ramjee, Bhuvan Sachdeva, Satvik Golechha, Shreyas Kulkarni, Geeta Fulari, Kaushik Murali, and Mohit Jain. CataractBot: An LLM-powered expert-in-the-loop chatbot for cataract patients. arXiv preprint arXiv:2402.04620, 2024
2024 arXiv
-
[88]
Personal tracking as lived informatics
John Rooksby, Mattias Rost, Alistair Morrison, and Matthew Chalmers. Personal tracking as lived informatics. In Proceedings of the SIGCHI conference on human factors in computing systems , pages 1163–1172, 2014
2014
-
[89]
The cost structure of sensemaking
Daniel M Russell, Mark J Stefik, Peter Pirolli, and Stuart K Card. The cost structure of sensemaking. In Proceedings of the INTERACT’93 and CHI’93 conference on Human factors in computing systems , pages 269–276, 1993
1993
-
[90]
Stress recognition using wearable sensors and mobile phones
Akane Sano and Rosalind W Picard. Stress recognition using wearable sensors and mobile phones. In 2013 Humaine association conference on affective computing and intelligent interaction , pages 671–676. IEEE, 2013
2013
-
[91]
Introducing WESAD, a multimodal dataset for wearable stress and affect detection
Philip Schmidt, Attila Reiss, Robert Dürichen, Claus Marberger, and Kristof Van Laerhoven. Introducing WESAD, a multimodal dataset for wearable stress and affect detection
-
[92]
Ecological momentary assessment
Saul Shiffman, Arthur A Stone, and Michael R Hufford. Ecological momentary assessment. 4(1):1–32, 2008. Publisher: Annual Reviews
2008
-
[93]
Towards expert-level medical question answering with large language models
Karan Singhal, Tao Tu, Juraj Gottweis, Rory Sayres, Ellery Wulczyn, Le Hou, Kevin Clark, Stephen Pfohl, Heather Cole-Lewis, Darlene Neal, and others. Towards expert-level medical question answering with large language models. arXiv preprint arXiv:2305.09617, 2023
2023 arXiv
-
[94]
Narrating fitness: Leveraging large language models for reflective fitness tracker data interpretation
Konstantin R Strömel, Stanislas Henry, Tim Johansson, Jasmin Niess, and Paweł W Woźniak. Narrating fitness: Leveraging large language models for reflective fitness tracker data interpretation. In Proceedings of the CHI conference on human factors in computing systems, pages 1–16, 2024
2024
-
[95]
Table Meets LLM: Can Large Language Models Understand Structured Table Data? A Benchmark and Empirical Study
Yuan Sui, Mengyu Zhou, Mingjie Zhou, Shi Han, and Dongmei Zhang. Table Meets LLM: Can Large Language Models Understand Structured Table Data? A Benchmark and Empirical Study. In Proceedings of the 17th ACM International Conference on Web Search and Data Mining, WSDM ’24, pages...
2024
-
[96]
DepreST-CAT: Retrospective smartphone call and text logs collected during the covid-19 pandemic to screen for mental illnesses
ML Tlachac, Ricardo Flores, Miranda Reisch, Katie Houskeeper, and Elke A Rundensteiner. DepreST-CAT: Retrospective smartphone call and text logs collected during the covid-19 pandemic to screen for mental illnesses. Proceedings of the ACM on Interactive, Mobile, Wearable and U...
2022
-
[97]
Language Models Don’t Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting
Miles Turpin, Julian Michael, Ethan Perez, and Samuel Bowman. Language Models Don’t Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting. Advances in Neural Information Processing Systems , 36:74952–74965, December 2023
2023
-
[98]
Design factors of longitudinal smartphone-based health surveys
Sudip Vhaduri and Christian Poellabauer. Design factors of longitudinal smartphone-based health surveys. Journal of Healthcare Informatics Research, 1:52–91, 2017. Publisher: Springer
2017
-
[99]
StudentLife: Assessing mental health, academic performance and behavioral trends of college students using smartphones
Rui Wang, Fanglin Chen, Zhenyu Chen, Tianxing Li, Gabriella Harari, Stefanie Tignor, Xia Zhou, Dror Ben-Zeev, and Andrew T Campbell. StudentLife: Assessing mental health, academic performance and behavioral trends of college students using smartphones. In Proceedings of the 20...
2014
-
[100]
Towards Natural Language-Based Visualization Authoring
Yun Wang, Zhitao Hou, Leixian Shen, Tongshuang Wu, Jiaqi Wang, He Huang, Haidong Zhang, and Dongmei Zhang. Towards Natural Language-Based Visualization Authoring. IEEE Transactions on Visualization and Computer Graphics , pages 1–11, 2022. , Vol. 1, No. 1, Article . Publicatio...
2022
-
[101]
MindShift: Leveraging large language models for mental-states-based problematic smartphone use intervention
Ruolan Wu, Chun Yu, Xiaole Pan, Yujia Liu, Ningning Zhang, Yue Fu, Yuhan Wang, Zhi Zheng, Li Chen, Qiaolei Jiang, and others. MindShift: Leveraging large language models for mental-states-based problematic smartphone use intervention. In Proceedings of the CHI conference on hu...
2024
-
[102]
Natural Language based Context Modeling and Reasoning for Ubiquitous Computing with Large Language Models: A Tutorial, December 2023
Haoyi Xiong, Jiang Bian, Sijia Yang, Xiaofei Zhang, Linghe Kong, and Daqing Zhang. Natural Language based Context Modeling and Reasoning for Ubiquitous Computing with Large Language Models: A Tutorial, December 2023. arXiv:2309.15074 [cs]
2023 arXiv
-
[103]
Dutcher, Yasaman S
Xuhai Xu, Prerna Chikersal, Janine M. Dutcher, Yasaman S. Sefidgar, Woosuk Seo, Michael J. Tumminia, Daniella K. Villalba, Sheldon Cohen, Kasey G. Creswell, J. David Creswell, Afsaneh Doryab, Paula S. Nurius, Eve Riskin, Anind K. Dey, and Jennifer Mankoff. Leveraging Collabora...
2021
-
[104]
GLOBEM: Cross-dataset generalization of longitudinal human behavior modeling
Xuhai Xu, Xin Liu, Han Zhang, Weichen Wang, Subigya Nepal, Yasaman Sefidgar, Woosuk Seo, Kevin S Kuehn, Jeremy F Huckins, Margaret E Morris, and others. GLOBEM: Cross-dataset generalization of longitudinal human behavior modeling. Proceedings of the ACM on Interactive, Mobile,...
2023
-
[105]
Understanding practices and needs of researchers in human state modeling by passive mobile sensing
Xuhai Xu, Jennifer Mankoff, and Anind K Dey. Understanding practices and needs of researchers in human state modeling by passive mobile sensing. CCF Transactions on Pervasive Computing and Interaction , 3:344–366, 2021. Publisher: Springer
2021
-
[106]
GLOBEM dataset: Multi-year datasets for longitudinal human behavior modeling generalization
Xuhai Xu, Han Zhang, Yasaman Sefidgar, Yiyi Ren, Xin Liu, Woosuk Seo, Jennifer Brown, Kevin Kuehn, Mike Merrill, Paula Nurius, and others. GLOBEM dataset: Multi-year datasets for longitudinal human behavior modeling generalization. Advances in Neural Information Processing Sys...
2022
-
[107]
DrHouse: An LLM-empowered Diagnostic Reasoning System through Harnessing Outcomes from Sensor Data and Expert Knowledge
Bufang Yang, Siyang Jiang, Lilin Xu, Kaiwei Liu, Hai Li, Guoliang Xing, Hongkai Chen, Xiaofan Jiang, and Zhenyu Yan. DrHouse: An LLM-empowered Diagnostic Reasoning System through Harnessing Outcomes from Sensor Data and Expert Knowledge. Proc. ACM Interact. Mob. Wearable Ubiqu...
2024
-
[108]
Bitterman, Jasmine Chiat Ling Ong, Daniel Shu Wei Ting, and Nan Liu
Rui Yang, Yilin Ning, Emilia Keppo, Mingxuan Liu, Chuan Hong, Danielle S. Bitterman, Jasmine Chiat Ling Ong, Daniel Shu Wei Ting, and Nan Liu. Retrieval-augmented generation for generative artificial intelligence in health care. npj Health Systems, 2(1):1–5, January
-
[109]
Why Johnny can’t prompt: How non-AI experts try (and fail) to design LLM prompts
JD Zamfirescu-Pereira, Richmond Y Wong, Bjoern Hartmann, and Qian Yang. Why Johnny can’t prompt: How non-AI experts try (and fail) to design LLM prompts. In Proceedings of the 2023 CHI conference on human factors in computing systems , pages 1–21, 2023
2023
-
[110]
Towards a comprehensive model of the cognitive process and mechanisms of individual sensemaking
Pengyi Zhang and Dagobert Soergel. Towards a comprehensive model of the cognitive process and mechanisms of individual sensemaking. Journal of the Association for Information Science and Technology , 65(9):1733–1756, 2014. Publisher: Wiley Online Library
2014
-
[111]
A survey on neural network interpretability
Yu Zhang, Peter Tiňo, Aleš Leonardis, and Ke Tang. A survey on neural network interpretability. IEEE Transactions on Emerging Topics in Computational Intelligence, 5(5):726–742, 2021. Publisher: IEEE. , Vol. 1, No. 1, Article . Publication date: September 2025
2021
-
[2019]
Publisher: SAGE Publications Sage UK: London, England
-
[2020]
Publisher: ACM New York, NY, USA
-
[2024]
Publisher: ACM New York, NY, USA tex.issue_date: July - August 2024
2024
-
[2025]
Publisher: Nature Publishing Group
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.