REVIEW 4 major objections 4 minor 81 references
Graphing the Everyday: A Neurosymbolic Approach to Eliciting Routines for Just-In-Time Adaptive Interventions
T0 review · 4 major / 4 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read A neurosymbolic voice agent that turns everyday talk into schedule knowledge fails precisely where linear extraction meets hierarchical human storytelling.
desk verdict Solid qualitative HCI study with a genuinely new evaluation method; the 'mental-model gap' attribution to inherent LLM linearity is untested and should not survive review as stated. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is a neurosymbolic conversation loop. A state machine drives a voice agent through schedule elicitation, and at each turn an LLM extractor converts the latest transcript into typed entities (events, locations, free-time windows) and temporal relations, which a persistent knowledge graph accumulates across the conversation. Transition rules fire on quantitative conditions read from the graph or semantic conditions judged by a second LLM, letting the agent clarify ambiguous details, confirm understood events, and return to earlier states when the user revises something. Afterward, two log-derived artifacts—the state-machine trace and the knowledge-graph trace—are shown to the user; their job is to make the system's internal model inspectable so that participants can point to the exact node, missing event, or transition where their story was distorted. This combination of symbolic persistence and visual traceability is what turns an abstract 'gap' into a catalog of concrete, user-identified failures.
What would settle it
Run the same elicitation task on matched transcripts with several different LLM extractors and with a variant that allows hierarchical nesting and clarification; if duplicate nodes, missing events, and ordering errors largely disappear or change character across conditions, the 'mental-model gap' is an implementation artifact rather than a structural property of linear extraction.
Extended reading notes
Core claim
The paper's central claim is that the friction between conversational routine elicitation and structured schedule data is not random noise but a systematic collision of two mental models. People narrate their day top-down and out of order, treating 'heading home' as one event that contains walking and taking the train, and assuming the listener already knows that 'office' means 'workplace' or that weekday patterns differ from Saturdays. The LLM extractor, by contrast, works bottom-up and linearly, so identical places become duplicate nodes, contained activities are split into unrelated events, stated commitments vanish or are reordered, and vague times are pinned to exact timestamps. The paper grounds this in a 16-participant member-checking study in which users audited the generated knowledge graph and state-machine trace, and it argues that the same collision extends from facts to feelings: a schedule can show free time that the user has no energy or motivation to use, which the authors call ecological mismatch. The proposed resolution is a set of design heuristics rather than a new algorithm: hierarchical extraction, clarification over assumption, contextual anchoring of nudges, adaptive negotiation, and scalable transparency.
Load-bearing premise
The core bet is that the observed failures come from the inherent linearity of LLM extraction rather than from this particular extractor, these prompts, and this state machine, which are never varied in the study.
Editorial extensions
If this is right
- Routine-elicitation systems for JITAIs should store and reason over hierarchical event structures with parent events and sub-events, so that 'heading home' can contain 'walking' and 'taking the train' without duplication.
- Extraction modules should treat vague time references as fuzzy bounds and respond to missing context by asking a clarifying question instead of assigning an arbitrary exact timestamp.
- Intervention timing should piggyback on existing behavioral transitions, such as lengthening a commute walk, rather than proposing isolated high-effort activities, so nudges align with real energy levels.
- Conversational agents should recognize user pushback, fatigue, or boundary-setting as a signal to pause their extraction agenda and negotiate, rather than looping to complete the state-machine goal.
- Transparency should be role-scaled: developers and researchers get the full knowledge graph for debugging, while end users get natural-language playbacks that let them correct the model without managing graph data.
Reading between the lines
- Because the extractor, prompts, and state machine are never varied, a comparison across different LLM extractors and conversation flows would settle whether the 'mental-model gap' is intrinsic to linear extraction or specific to this implementation.
- The fact that users eagerly corrected the graph when shown after the conversation suggests a live, in-chat editing surface—merging duplicate nodes or re-parenting sub-events while talking—could close the gap in real time, though the paper only tests post-hoc review.
- The same hierarchy-versus-linearity friction should appear in any proactive agent that must act on user-described routines, such as calendar assistants or reminder systems, so the design heuristics may transfer beyond health interventions.
- Combining this elicitation layer with wearable or physiological signals would let future work measure the ecological mismatch quantitatively, learning when a user's free time and receptive time diverge rather than only documenting the divergence qualitatively.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports a qualitative lab study of a neurosymbolic conversational agent (LLM extractor paired with a Neo4j knowledge graph) that elicits users' daily routines through voice dialogue and converts them into structured schedule data for Just-In-Time Adaptive Interventions (JITAIs). Sixteen participants interacted with the agent, then reviewed state-machine traces, knowledge graphs, and natural-language summaries in a member-checking phase. Thematic analysis of 221 coded units yields three reported themes: a 'mental-model gap' between non-linear, hierarchical human narration and 'linear' LLM extraction; positive conversational naturalness but rigidity and looping in the state-machine-driven dialogue; and users' willingness to bridge the gap when given structural transparency. The paper argues these findings reveal fundamental structural friction and derives design heuristics (hierarchical extraction, clarification over assumption, routine piggybacking, adaptive negotiation, scalable transparency).
Significance. If the central claim held, the work would be a significant empirical contribution to JITAI design, showing that routine elicitation requires adaptive, hierarchical, and negotiable representations rather than flat extraction. The paper has clear strengths: it follows a codebook thematic-analysis procedure with reported inter-rater reliability (Cohen's kappa improving from 0.32 to 0.71 to 0.79), uses human-centric evaluation via natural-language playbacks, and offers concrete, actionable design heuristics. The transparency artifacts (state-machine trace, knowledge-graph trace) are a useful methodological contribution. However, the paper's central causal attribution—that observed extraction failures are caused by the inherent linearity of LLM extraction—is not supported by the single-system study design. The contribution is better characterized as a detailed single-system case study of conversational routine elicitation; the broader 'fundamental' claims require comparative evidence or substantial reframing.
major comments (4)
- [Section 6.1 and Theme 1 (Section 5.1)] The central claim that 'standard zero-shot LLM extractors operate bottom-up and linearly' and that this 'strict linear architecture causes severe node duplication and fragmentation' is an attribution to inherent properties of linear extraction, but the study only evaluates one specific pipeline: LangExtract with gpt-4o-mini, a particular state-machine prompt regime, a graph-merge implementation, and an ASR front-end. The paper never varies the extractor, the prompt templates, the conversation flow, or the merge logic, nor does it compare against a non-linear or human extraction baseline. The reported failures—co-referent duplicates (n=12), missing information (n=20), wrong order (n=20), and even speech-recognition errors such as 'gym captured as dream' (P6)—could stem from model capability, prompt design, merge rules, or ASR errors rather than from a fundamental property of linearity. Because the paper's title, abstract, and design heuristics generalize beyond this system, this untested attribution is load-bearing and needs either direct comparative evidence or a clear reframing to single-system findings.
- [Section 4 (Study Design)] The study design has no baseline or control condition and reports no objective error counts. All evidence of the 'mental-model gap' comes from participants' subjective identification of errors during the member-checking phase, after they have been shown the system's own representations. Additionally, the codebook was developed by the same authors who designed the system, so the themes (especially Theme 1) may be shaped by the system's design stance. To support the claim that the gap is a property of human-algorithm translation rather than of this particular implementation, the paper would need a comparative condition (e.g., form-based elicitation, a human transcription baseline, or a different extractor) or a substantially more cautious interpretation.
- [Section 6.2 and Theme 2 (Section 5.2)] The state machine's rigid goal-driven behavior—looping (n=12), insistence (n=14), and perceived lack of adaptivity—is described as evidence of a general 'ecological mismatch' and is used to motivate the 'adaptive negotiation' heuristic. However, this behavior is a designed property of the particular state machine and prompt templates, not a general property of LLM extraction or neurosymbolic systems. The state machine itself may even induce the non-linear corrections and looping users exhibited, partly manufacturing the observed 'gap.' Without varying the state machine (e.g., testing an adaptive variant) or at least acknowledging this as a design choice rather than an inherent limitation, the recommendation for adaptive negotiation is not empirically grounded by this study.
- [Section 5 (Qualitative Analysis)] The text states that the 27 codes were grouped into 'five themes reported below,' but then says 'We report three themes' and presents only three subsections (5.1, 5.2, 5.3). This inconsistency is confusing and should be corrected, either by reporting all five themes or by revising the earlier sentence to match the actual number.
minor comments (4)
- [Section 5] There is a typo: 'independetly' should be 'independently.'
- [References] Reference [23] contains the incomplete placeholder 'Accessed: [Insert access date here].' This should be filled in or removed.
- [Section 6.1 and Conclusion] The words 'fundamental structural paradox' and 'fundamental mental-model gap' are stronger than the single-system evidence supports; consider tempering the language unless comparative evidence is added.
- [Section 5.1] Reporting the distribution of the 27 codes across the themes in a summary table would improve transparency and help readers assess the relative weight of the reported frequencies.
Circularity Check
No circularity: the qualitative evaluation is self-contained, with no fitted inputs, load-bearing self-citations, or definitional reductions.
full rationale
The paper makes no quantitative predictions and fits no parameters, so the high-inversion circularity patterns do not apply. The central 'mental-model gap' is an interpretive construct grounded in participant reports and thematic analysis (Section 5.1), not a quantity derived from an input that presupposes it. The codebook was developed transparently with the study's framing acknowledged (Section 5), which is standard thematic-analysis practice rather than a circular reduction. No load-bearing self-citations or author-imported uniqueness claims appear; the neurosymbolic design is justified by external references [15, 52, 75], and the paper explicitly disclaims the novelty of merely observing extraction imperfection (Section 2.3). The limitations section appropriately acknowledges the controlled lab setting and small sample (Section 6.5). The skeptic's concern about extractor, prompt, and state-machine confounds is an external-validity or causal-attribution risk, not circularity, and therefore falls outside this pass's scope.
Assumptions & free parameters
assumptions (3)
- domain assumption Human routine narratives are fundamentally hierarchical and non-linear.
- domain assumption User self-report during member checking is a valid ground truth for representational fidelity.
- standard math Conventional Cohen's kappa thresholds indicate sufficient inter-rater reliability.
Cite this review
Pith. "Pith review of Graphing the Everyday: A Neurosymbolic Approach to Eliciting Routines for Just-In-Time Adaptive Interventions." pith.science (2026). https://pith.science/paper/QEI4YV2G
@misc{pith2026260809294,
author = {Pith},
title = {Pith review of: Graphing the Everyday: A Neurosymbolic Approach to Eliciting Routines for Just-In-Time Adaptive Interventions},
year = {2026},
howpublished = {\url{https://pith.science/paper/QEI4YV2G}},
note = {Machine review of arXiv:2608.09294}
}
read the original abstract
Just-In-Time Adaptive Interventions (JITAIs) increasingly rely on conversational agents to elicit user routines, yet translating fluid human dialogue into rigid schedule data remains a significant challenge. We conducted a qualitative investigation of a neurosymbolic pipeline, combining Large Language Models (LLMs) with a Neo4j knowledge graph, to map unstructured verbal narratives into actionable interventions. Through human-centric evaluation using natural-language playbacks, we identified a critical "mental-model gap," where the linear extraction of LLMs clashes with hierarchical, non-linear human storytelling, causing severe entity fragmentation. Furthermore, we articulate an "ecological mismatch," demonstrating that algorithmic schedule availability frequently ignores the user's fluctuating psychological receptivity and physical energy levels. To resolve these tensions, we propose actionable design heuristics, including routine piggybacking, adaptive negotiation, and scalable transparency. Ultimately, these guidelines provide a foundational framework for evolving rigid schedule-trackers into empathetic, context-aware proactive agents capable of supporting long-term health behavior change.
Figures
Reference graph
Works this paper leans on
-
[72]
Yuqing Wang and Yun Zhao. 2024. TRAM: Benchmarking Temporal Reasoning for Large Language Models. InFindings of the Association for Computational Linguistics: ACL 2024. Association for Computational Linguistics, Bangkok, Thailand, 6389–6415. doi:10.18653/v1/2024.findings-acl.382
-
[1]
James F. Allen. 1983. Maintaining knowledge about temporal intervals.Commun. ACM26, 11 (Nov. 1983), 832–843. doi:10.1145/182.358434
-
[2]
Krisztian Balog and Tom Kenter. 2019. Personal Knowledge Graphs: A Research Agenda. InProceedings of the 2019 ACM SIGIR International Conference on Theory of Information Retrieval(Santa Clara, CA, USA)(ICTIR ’19). Association for Computing Machinery, New York, NY, USA, 217–220. doi:10.1145/3341981.3344241
arXiv 2019
-
[3]
Samuel L. Battalio, David E. Conroy, Walter Dempsey, Peng Liao, Marianne Menictas, Susan Murphy, Inbal Nahum-Shani, Tianchen Qian, Santosh Kumar, and Bonnie Spring. 2021. Sense2Stop: A Micro-randomized Trial Using Wearable Sensors to Optimize a Just-in-Time-Adaptive Stress Management Intervention for Smoking Relapse Prevention.Contemporary Clinical Trials...
arXiv 2021
-
[4]
Caterina Bérubé, Marcia Nißen, Rasita Vinay, Alexa Geiger, Tobias Budig, Aashish Bhandari, Catherine Rachel Pe Benito, Nathan Ibarcena, Olivia Pistolese, Pan Li, Abdullah Bin Sawad, Elgar Fleisch, Christoph Stettler, Bronwyn Hemsley, Shlomo Berkovsky, Tobias Kowatsch, and A. Baki Kocaballi. 2024. Proactive behavior in voice assistants: A systematic review...
-
[5]
Timothy W. Bickmore, Suzanne E. Mitchell, Brian W. Jack, Michael K. Paasche-Orlow, Laura M. Pfeifer, and Julie O’Donnell. 2010. Response to a Relational Agent by Hospital Patients with Depressive Symptoms.Interacting with Computers22, 4 (2010), 289–298. doi:10.1016/j.intcom.2009.12.001
-
[6]
Boyatzis
Richard E. Boyatzis. 1998.Transforming Qualitative Information: Thematic Analysis and Code Development. Sage Publications, Thousand Oaks, CA
1998
-
[7]
Virginia Braun and Victoria Clarke. 2006. Using Thematic Analysis in Psychology.Qualitative Research in Psychology3, 2 (2006), 77–101. doi:10.1191/1478088706qp063oa
Show all 81 references
-
[8]
Kelly Caine. 2016. Local Standards for Sample Size at CHI. InProceedings of the 2016 CHI Conference on Human Factors in Computing Systems(San Jose, California, USA)(CHI ’16). Association for Computing Machinery, New York, NY, USA, 981–992. doi:10.1145/2858036.2858498
2016
-
[9]
Alonso-Moral, Alejandro Catala, and Alberto Bugarín-Diz
Mariña Canabal-Juanatey, Jose M. Alonso-Moral, Alejandro Catala, and Alberto Bugarín-Diz. 2024. Enriching interactive explanations with fuzzy temporal constraint networks.International Journal of Approximate Reasoning171 (2024), 109128. doi:10.1016/j.ijar.2024.109128
2024
-
[10]
Narae Cha, Auk Kim, Cheul Young Park, Soowon Kang, Minkyu Park, Jae-Gil Lee, Sangsu Lee, and Uichin Lee. 2020. Hello There! Is Now a Good Time to Talk? Opportune Moments for Proactive Interactions with Smart Speakers.Proc. ACM Interact. Mob. Wearable Ubiquitous Technol.4, 3, A...
2020 doi
-
[11]
Prantika Chakraborty and Debarshi Kumar Sanyal. 2023. A comprehensive survey of personal knowledge graphs.Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery13, 6 (2023), e1513
2023
-
[12]
Zheng Chu, Jingchang Chen, Qianglong Chen, Weijiang Yu, Tao He, Haotian Wang, Weihua Peng, Ming Liu, Bing Qin, and Ting Liu. 2024. TimeBench: A Comprehensive Evaluation of Temporal Reasoning Abilities in Large Language Models. InProceedings of the 62nd Annual Meeting of the As...
2024 doi
-
[13]
Clark and Susan E
Herbert H. Clark and Susan E. Brennan. 1991. Grounding in communication. InPerspectives on socially shared cognition, L. B. Resnick, J. M. Levine, and S. D. Teasley (Eds.). American Psychological Association, 127–149. doi:10.1037/10096-006
1991 doi
-
[14]
Rudresh Dwivedi, Devam Dave, Het Naik, Smiti Singhal, Rana Omer, Pankesh Patel, Bin Qian, Zhenyu Wen, Tejal Shah, Graham Morgan, and Rajiv Ranjan. 2023. Explainable AI (XAI): Core Ideas, Techniques, and Solutions.ACM Comput. Surv.55, 9, Article 194 (Jan. 2023), 33 pages. doi:1...
2023 doi
-
[15]
Darren Edge, Ha Trinh, Newman Cheng, Joshua Bradley, Alex Chao, Apurva Mody, Steven Truitt, and Jonathan Larson. 2024. From Local to Global: A Graph RAG Approach to Query-Focused Summarization. arXiv:2404.16130 [cs.CL] https://arxiv.org/abs/2404.16130
2024 arXiv
-
[16]
Doyle, Holly P
Justin Edwards, Philip R. Doyle, Holly P. Branigan, and Benjamin R. Cowan. 2024. Comparing Perceptions of Static and Adaptive Proactive Speech Agents. InProceedings of the 6th ACM Conference on Conversational User Interfaces(Luxembourg, Luxembourg)(CUI ’24). Association for Co...
2024
-
[17]
Robin Emsley. 2023. ChatGPT: these are not hallucinations – they’re fabrications and falsifications.Schizophrenia9, 1 (2023), 52. doi:10.1038/s41537- 023-00379-4
2023 doi
-
[18]
OpenAI et al. 2023. GPT-4 Technical Report. arXiv:2303.08774 [cs.CL] https://arxiv.org/abs/2303.08774
2023 arXiv
-
[19]
Katayoun Farrahi and Daniel Gatica-Perez. 2011. Discovering routines from large-scale human locations using probabilistic topic models.ACM Transactions on Intelligent Systems and Technology (TIST)2, 1 (2011), 1–27. doi:10.1145/1889681.1889684
2011
-
[20]
Kathleen Kara Fitzpatrick, Alison Darcy, and Molly Vierhile. 2017. Delivering Cognitive Behavior Therapy to Young Adults With Symptoms of Depression and Anxiety Using a Fully Automated Conversational Agent (Woebot): A Randomized Controlled Trial.JMIR Mental Health4, 2 (2017), ...
2017 doi
-
[21]
Caroline Free, Gemma Phillips, Leandro Galli, Louise Watson, Lambert Felix, Phil Edwards, Vikram Patel, and Andy Haines. 2013. The Effectiveness of Mobile-Health Technology-Based Health Behaviour Change or Disease Management Interventions for Health Care Consumers: A Systemati...
2013 doi
- [22]
-
[23]
Akshay Goel and Atilla Kiraly. 2025. Introducing LangExtract: A Gemini powered information extraction library. Google Developers Blog. https://developers.googleblog.com/introducing-langextract-a-gemini-powered-information-extraction-library/ Accessed: [Insert access date here]
2025
-
[24]
MacQueen, and Emily E
Greg Guest, Kathleen M. MacQueen, and Emily E. Namey. 2012.Applied Thematic Analysis. Sage Publications, Thousand Oaks, CA. doi:10.4135/ 9781483384436
2012
-
[25]
Hofer, Mahdi Sareban, Gunnar Treff, Josef Niebauer, Christopher N Bull, Albrecht Schmidt, and Jan David Smeddinck
David Haag, Devender Kumar, Sebastian Gruber, Dominik P. Hofer, Mahdi Sareban, Gunnar Treff, Josef Niebauer, Christopher N Bull, Albrecht Schmidt, and Jan David Smeddinck. 2025. The Last JITAI? Exploring Large Language Models for Issuing Just-in-Time Adaptive Interventions: Fo...
2025
-
[26]
Wendy Hardeman, Jane Houghton, Katie Lane, Anne Beeke, Vase, Lydia, Johnston, Marie, Sutton, and Stephen. 2019. A systematic review of just-in-time adaptive interventions (JITAIs) to promote physical activity.International Journal of Behavioral Nutrition and Physical Activity1...
2019 doi
-
[27]
Henry, Morkeh Blay-Tofey, Clara E
Lauren M. Henry, Morkeh Blay-Tofey, Clara E. Haeffner, Cassandra N. Raymond, Elizabeth Tandilashvili, Nancy Terry, Miryam Kiderman, Olivia Metcalf, Melissa A. Brotman, and Silvia Lopez-Guzman. 2025. Just-In-Time Adaptive Interventions to Promote Behavioral Health: Protocol for...
2025 doi
-
[28]
Rashid, Anisa Rula, Lukas Schmelzeisen, Juan Sequeda, Steffen Staab, and Antoine Zimmermann
Aidan Hogan, Eva Blomqvist, Michael Cochez, Claudia D’amato, Gerard De Melo, Claudio Gutierrez, Sabrina Kirrane, José Emilio Labra Gayo, Roberto Navigli, Sebastian Neumaier, Axel-Cyrille Ngonga Ngomo, Axel Polleres, Sabbir M. Rashid, Anisa Rula, Lukas Schmelzeisen, Juan Sequed...
2021 doi
-
[29]
Becky Inkster, Shubhankar Sarda, and Vinod Subramanian. 2018. An Empathy-Driven, Conversational Artificial Intelligence Agent (Wysa) for Digital Mental Well-Being: Real-World Data Evaluation Mixed-Methods Study.JMIR mHealth and uHealth6, 11 (2018), e12106. doi:10.2196/12106
2018 doi
-
[30]
Iqbal and Brian P
Shamsi T. Iqbal and Brian P. Bailey. 2005. Investigating the effectiveness of mental workload as a predictor of opportune moments for interruption. InCHI ’05 Extended Abstracts on Human Factors in Computing Systems(Portland, OR, USA)(CHI EA ’05). Association for Computing Mach...
2005
-
[31]
Epstein, Hyunhoon Jung, and Young-Ho Kim
Eunkyung Jo, Daniel A. Epstein, Hyunhoon Jung, and Young-Ho Kim. 2023. Understanding the Benefits and Challenges of Deploying Conversational AI Leveraging Large Language Models for Public Health Intervention. InProceedings of the 2023 CHI Conference on Human Factors in Computi...
2023
-
[32]
Myung Ho Kim. 2025. Structured Cognitive Loop for Behavioral Intelligence in Large Language Model Agents. arXiv:2510.05107 [cs.AI] https: //arxiv.org/abs/2510.05107
2025 arXiv
-
[33]
Myung Ho Kim. 2026. Bridging Symbolic Control and Neural Reasoning in LLM Agents: Structured Cognitive Loop with a Governance Layer. arXiv:2511.17673 [cs.AI] https://arxiv.org/abs/2511.17673
2026 arXiv
-
[34]
Seewald, Andy Lee, Kelly Hall, Brook Luers, Eric B
Predrag Klasnja, Shawna Smith, Nicholas J. Seewald, Andy Lee, Kelly Hall, Brook Luers, Eric B. Hekler, and Susan A. Murphy. 2019. Efficacy of Contextually Tailored Suggestions for Physical Activity: A Micro-randomized Optimization Trial of HeartSteps.Annals of Behavioral Medic...
2019 doi
-
[35]
Rafal Kocielnik, Lillian Xiao, Daniel Avrahami, and Gary Hsieh. 2018. Reflection Companion: A Conversational System for Engaging Users in Reflection on Physical Activity.Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies2, 2, Article 70 (July 2...
2018 doi
-
[36]
Todd Kulesza, Margaret Burnett, Weng-Keen Wong, and Simone Stumpf. 2015. Principles of Explanatory Debugging to Personalize Interactive Machine Learning(IUI ’15). Association for Computing Machinery, New York, NY, USA, 126–137. doi:10.1145/2678025.2701399
2015
-
[37]
Florian Künzler, Varun Mishra, Jan-Niklas Kramer, David Kotz, Elgar Fleisch, and Tobias Kowatsch. 2019. Exploring the State-of-Receptivity for mHealth Interventions.Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies3, 4, Article 140 (Dec. 2019)...
2019 doi
-
[38]
Richard Landis and Gary G
J. Richard Landis and Gary G. Koch. 1977. The Measurement of Observer Agreement for Categorical Data.Biometrics33, 1 (1977), 159–174. doi:10.2307/2529310 Manuscript submitted to ACM Graphing the Everyday: A Neurosymbolic Approach to Eliciting Routines for Just-In-Time Adaptive...
1977 doi
-
[39]
Lane, Emiliano Miluzzo, Hong Lu, Daniel Peebles, Tanzeem Choudhury, and Andrew T
Nicholas D. Lane, Emiliano Miluzzo, Hong Lu, Daniel Peebles, Tanzeem Choudhury, and Andrew T. Campbell. 2010. A Survey of Mobile Phone Sensing.IEEE Communications Magazine48, 9 (2010), 140–150. doi:10.1109/MCOM.2010.5560598
2010
-
[40]
Dunn, Huong Ly Tong, Ahmet Baki Kocaballi, Jessica Chen, Rabia Bashir, Didi Surian, Blanca Gallego, Farah Magrabi, Annie Y
Liliana Laranjo, Adam G. Dunn, Huong Ly Tong, Ahmet Baki Kocaballi, Jessica Chen, Rabia Bashir, Didi Surian, Blanca Gallego, Farah Magrabi, Annie Y. S. Lau, and Enrico Coiera. 2018. Conversational Agents in Healthcare: A Systematic Review.Journal of the American Medical Inform...
2018 doi
-
[41]
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. 2020. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. InProceedings o...
2020
-
[42]
Vera Liao, Daniel Gruen, and Sarah Miller
Q. Vera Liao, Daniel Gruen, and Sarah Miller. 2020. Questioning the AI: Informing Design Practices for Explainable AI User Experiences. In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems(Honolulu, HI, USA)(CHI ’20). Association for Computing Machin...
2020
-
[43]
MacQueen, Eleanor McLellan, Kelly Kay, and Bobby Milstein
Kathleen M. MacQueen, Eleanor McLellan, Kelly Kay, and Bobby Milstein. 1998. Codebook Development for Team-Based Qualitative Analysis. CAM Journal10, 2 (1998), 31–36. doi:10.1177/1525822X980100020301
1998 doi
-
[44]
Raju Maharjan, Kevin Doherty, Darius Adam Rohani, Per Bækgaard, and Jakob E. Bardram. 2022. Experiences of a Speech-Enabled Conversational Agent for the Self-Report of Well-Being Among People Living with Affective Disorders: An In-the-Wild Study.ACM Transactions on Interactive...
2022 doi
-
[45]
Negar Maleki, Balaji Padmanabhan, and Kaushik Dutta. 2024. AI Hallucinations: A Misnomer Worth Clarifying. In2024 IEEE Conference on Artificial Intelligence (CAI). 133–138. doi:10.1109/CAI59869.2024.00033
2024
-
[46]
Abhinav Mehrotra, Veljko Pejovic, Jo Vermeulen, Robert Hendley, and Mirco Musolesi. 2016. My Phone and Me: Understanding People’s Receptivity to Mobile Notifications. (2016), 1021–1032. doi:10.1145/2858036.2858566
2016
-
[47]
Mikhail Menschikov, Dmitry Evseev, Victoria Dochkina, Ruslan Kostoev, Ilia Perepechkin, Petr Anokhin, Nikita Semenov, and Evgeny Burnaev. 2026. PersonalAI: A Systematic Comparison of Knowledge Graph Storage and Retrieval Approaches for Personalized LLM agents. arXiv:2506.17001...
2026 arXiv
-
[48]
Varun Mishra, Florian Künzler, Jan-Niklas Kramer, Elgar Fleisch, Tobias Kowatsch, and David Kotz. 2023. Detecting Receptivity for mHealth Interventions.GetMobile: Mobile Comp. and Comm.27, 2 (Aug. 2023), 23–28. doi:10.1145/3614214.3614221
2023
-
[49]
Andre Matthias Müller, Ann Blandford, and Lucy Yardley. 2017. The conceptualization of a Just-In-Time Adaptive Intervention (JITAI) for the reduction of sedentary behavior in older adults.MHealth3 (Sep 2017), 37. doi:10.21037/mhealth.2017.08.05
2017 doi
-
[50]
Inbal Nahum-Shani, Shawn N Smith, Bonnie J Spring, Linda M Collins, Katie Witkiewitz, Ambuj Tewari, and Susan A Murphy. 2018. Just-in-Time Adaptive Interventions (JITAIs) in Mobile Health: Key Components and Design Principles for Ongoing Health Behavior Support.Annals of Behav...
2018 doi
-
[51]
Patil, Ion Stoica, and Joseph E
Charles Packer, Sarah Wooders, Kevin Lin, Vivian Fang, Shishir G. Patil, Ion Stoica, and Joseph E. Gonzalez. 2023. MemGPT: Towards LLMs as Operating Systems. arXiv:2310.08560 [cs.AI] https://arxiv.org/abs/2310.08560
2023 arXiv
-
[52]
Shirui Pan, Linhao Luo, Yufei Wang, Chen Chen, Jiapu Wang, and Xindong Wu. 2024. Unifying Large Language Models and Knowledge Graphs: A Roadmap.IEEE Transactions on Knowledge and Data Engineering36, 7 (2024), 3580–3599. doi:10.1109/TKDE.2024.3352100
2024
-
[53]
Bernstein
Joon Sung Park, Joseph O’Brien, Carrie Jun Cai, Meredith Ringel Morris, Percy Liang, and Michael S. Bernstein. 2023. Generative Agents: Interactive Simulacra of Human Behavior. InProceedings of the 36th Annual ACM Symposium on User Interface Software and Technology(San Francis...
2023
-
[54]
Martin Pielot, Bruno Cardoso, Kleomenis Katevas, Joan Serrà, Aleksandar Matic, and Nuria Oliver. 2017. Beyond Interruptibility: Predicting Opportune Moments to Engage Mobile Phone Users.Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies1, 3, Ar...
2017 doi
-
[55]
Martin Pielot, Rodrigo de Oliveira, Haewoon Kwak, and Nuria Oliver. 2014. Didn’t You See My Message? Predicting Attentiveness to Mobile Instant Messages. InProceedings of the SIGCHI Conference on Human Factors in Computing Systems(Toronto, Ontario, Canada)(CHI ’14). Associatio...
2014
-
[56]
Cowan, and Russell Beale
Charlie Pinder, Jo Vermeulen, Benjamin R. Cowan, and Russell Beale. 2018. Digital Behaviour Change Interventions to Break and Form Habits.ACM Trans. Comput.-Hum. Interact.25, 3, Article 15 (June 2018), 66 pages. doi:10.1145/3196830
2018 doi
-
[57]
R. J. Planer. 2023. The evolution of hierarchically structured communication.Frontiers in Psychology14 (2023), 1224324. doi:10.3389/fpsyg.2023.1224324
2023
-
[58]
Mashfiqui Rabbi, Min Hane Aung, Mi Zhang, and Tanzeem Choudhury. 2015. MyBehavior: Automatic Personalized Health Feedback from User Behaviors and Preferences Using Smartphones. InProceedings of the 2015 ACM International Joint Conference on Pervasive and Ubiquitous Computing (...
2015
-
[59]
Reinhartz-Berger, S
I. Reinhartz-Berger, S. J. Ali, and D. Bork. 2025. Leveraging LLMs for Domain Modeling: The Impact of Granularity and Strategy on Quality. In Advanced Information Systems Engineering. CAiSE 2025 (Lecture Notes in Computer Science, Vol. 15701), J. Krogstie, S. Rinderle-Ma, G. K...
2025 doi
-
[60]
Josh Rosen and Seth Rosen. 2026. From Agent Loops to Deterministic Graphs: Execution Lineage for Reproducible AI-Native Work. arXiv:2605.06365 [cs.AI] https://arxiv.org/abs/2605.06365 Manuscript submitted to ACM 18 Jayasiriwardene & Mountford et al
2026 arXiv
-
[61]
Michele Salvagno, Fabio Silvio Taccone, and Alberto Giovanni Gerli. 2023. Artificial intelligence hallucinations.Critical Care27, 1 (2023), 180. doi:10.1186/s13054-023-04473-y
2023 doi
-
[62]
Husniya Salwa, Ntombifuthi Burhan, and Ernest Rahel. 2025. Continual Learning: Overcoming Catastrophic Forgetting for Adaptive AI Systems. doi:10.36227/techrxiv.173886426.63028528/v1
2025
-
[63]
Mahbubur Rahman, Rummana Bari, Syed Monowar Hossain, and Santosh Kumar
Hillol Sarker, Moushumi Sharmin, Amin Ahsan Ali, Md. Mahbubur Rahman, Rummana Bari, Syed Monowar Hossain, and Santosh Kumar. 2014. Assessing the Availability of Users to Engage in Just-in-Time Intervention in the Natural Environment. InProceedings of the 2014 ACM International...
2014
-
[64]
Jessica Schroeder, Chelsey Wilkes, Kael Rowan, Arturo Toledo, Ann Paradiso, Mary Czerwinski, Gloria Mark, and Marsha M. Linehan. 2018. Pocket Skills: A Conversational Mobile Web App To Support Dialectical Behavioral Therapy. InProceedings of the 2018 CHI Conference on Human Fa...
2018
-
[65]
Stone, and Michael R
Saul Shiffman, Arthur A. Stone, and Michael R. Hufford. 2008. Ecological Momentary Assessment.Annual Review of Clinical Psychology4 (2008), 1–32. doi:10.1146/annurev.clinpsy.3.022806.091415
2008
-
[66]
Karan Singhal, Shekoofeh Azizi, Tao Tu, S. Sara Mahdavi, Jason Wei, Hyung Won Chung, Nathan Scales, Ajay Tanwani, Heather Cole-Lewis, Stephen Pfohl, Perry Payne, Martin Seneviratne, Paul Gamble, Chris Kelly, Abubakr Babiker, Nathanael Schärli, Aakanksha Chowdhery, Philip Mansf...
2023
-
[67]
Stone and Saul Shiffman
Arthur A. Stone and Saul Shiffman. 1994. Ecological Momentary Assessment (EMA) in Behavioral Medicine.Annals of Behavioral Medicine16, 3 (1994), 199–202. doi:10.1093/abm/16.3.199
1994 doi
-
[68]
Lorainne Tudor Car, Dhakshenya Ardhithy Dhinagaran, Bhone Myint Kyaw, Tobias Kowatsch, Shafiq Joty, Yin-Leng Theng, and Rifat Atun
-
[69]
Niels van Berkel, Denzil Ferreira, and Vassilis Kostakos. 2017. The Experience Sampling Method on Mobile Devices.Comput. Surveys50, 6, Article 93 (Dec. 2017), 40 pages. doi:10.1145/3123988
2017 doi
-
[70]
van Dantzig, G
S. van Dantzig, G. Geleijnse, and A. T. van Halteren. 2013. Toward a persuasive mobile application to reduce sedentary behavior.Personal and Ubiquitous Computing17, 6 (Aug 2013), 1237–1246. doi:10.1007/s00779-012-0588-0
2013 doi
-
[71]
David Vela, Andrew Sharp, Ruizhe Zhang, et al. 2022. Temporal quality degradation in AI models.Scientific Reports12, 1 (2022), 11654. doi:10.1038/ s41598-022-15245-z
2022
-
[73]
Xiang Wei, Xingyu Cui, Ning Cheng, Xiaobin Wang, Xin Zhang, Shen Huang, Pengjun Xie, Jinan Xu, Yufeng Chen, Meishan Zhang, Yong Jiang, and Wenjuan Han. 2024. ChatIE: Zero-Shot Information Extraction via Chatting with ChatGPT. arXiv:2302.10205 [cs.CL] https://arxiv.org/abs/2302.10205
2024 arXiv
-
[74]
F. Xu, H. Uszkoreit, Y. Du, W. Fan, D. Zhao, and J. Zhu. 2019. Explainable AI: A Brief Survey on History, Research Areas, Approaches and Challenges. InNatural Language Processing and Chinese Computing. NLPCC 2019 (Lecture Notes in Computer Science, Vol. 11839), J. Tang, M. Y. ...
2019 doi
-
[75]
Michihiro Yasunaga, Hongyu Ren, Antoine Bosselut, Percy Liang, and Jure Leskovec. 2021. QA-GNN: Reasoning with Language Models and Knowledge Graphs for Question Answering. InProceedings of the 2021 Conference of the North American Chapter of the Association for Computational L...
2021 doi
-
[76]
Steve Young, Milica Gašić, Blaise Thomson, and Jason D Williams. 2013. Pomdp-based statistical spoken dialog systems: A review.Proc. IEEE101, 5 (2013), 1160–1179
2013
-
[77]
Zacks and Barbara Tversky
Jeffrey M. Zacks and Barbara Tversky. 2001. Event structure in perception and conception.Psychological Bulletin127, 1 (2001), 3–21. doi:10.1037/0033- 2909.127.1.3
2001 doi
-
[78]
Xing, Hao Zhang, Joseph E
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 2023. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. arXiv:2306.05685 [cs.CL] https://ar...
2023 arXiv
-
[79]
Sizhe Zhou and Jiawei Han. 2025. A Simple Yet Strong Baseline for Long-Term Conversational Memory of LLM Agents. arXiv:2511.17208 [cs.CL] https://arxiv.org/abs/2511.17208
2025
-
[80]
Yi Zhu, Zhaojun Yang, Helen Meng, Baichuan Li, Gina Levow, and Irwin King. 2010. Using finite state machines for evaluating spoken dialog systems. In2010 IEEE Spoken Language Technology Workshop. IEEE, 478–483. Manuscript submitted to ACM
2010
-
[2020]
doi:10.2196/17158
Conversational Agents in Health Care: Scoping Review and Conceptual Analysis.Journal of Medical Internet Research22, 8 (2020), e17158. doi:10.2196/17158
2020 doi
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.