REVIEW 4 cited by
Transforming Wearable Data into Personal Health Insights using Large Language Model Agents
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Deriving personalized insights from popular wearable trackers requires complex numerical reasoning that challenges standard LLMs, necessitating tool-based approaches like code generation. Large language model (LLM) agents present a promising yet largely untapped solution for this analysis at scale. We introduce the Personal Health Insights Agent (PHIA), a system leveraging multistep reasoning with code generation and information retrieval to analyze and interpret behavioral health data. To test its capabilities, we create and share two benchmark datasets with over 4000 health insights questions. A 650-hour human expert evaluation shows that PHIA significantly outperforms a strong code generation baseline, achieving 84% accuracy on objective, numerical questions and, for open-ended ones, earning 83% favorable ratings while being twice as likely to achieve the highest quality rating. This work can advance behavioral health by empowering individuals to understand their data, enabling a new era of accessible, personalized, and data-driven wellness for the wider population.
Forward citations
Cited by 4 Pith papers
-
AnnoSense: A Framework for Physiological Emotion Data Collection in Everyday Settings for AI
The authors propose AnnoSense, a set of 15 expert-reviewed guidelines for everyday emotion data collection, derived from survey, interview, and focus group insights from 119 stakeholders.
-
GLOSS: Group of LLMs for Open-Ended Sensemaking of Passive Sensing Data for Health and Wellbeing
A group of LLM agents that collaboratively generate code for raw passive sensing data outperforms RAG on objective query accuracy, while remaining only moderately consistent across repeated runs.
-
SensorLM: Learning the Language of Wearable Sensors
SensorLM is a sensor-language foundation model trained on 59.7M hours of wearable data with template-generated captions, reporting strong zero-shot, few-shot, and retrieval performance.
-
SePA: A Search-enhanced Predictive Agent for Personalized Health Coaching
SePA combines personalized wearable-data risk prediction with a whitelisted web search pipeline to give cited, context-aware health coaching.
Discussion (0). Sign in to comment.