Pith. sign in

REVIEW 8 cited by

Health-LLM: Large Language Models for Health Prediction via Wearable Sensor Data

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.06866 v2 pith:T3UZTKCG submitted 2024-01-12 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords healthcontextperformancedataknowledgelanguagellmsmodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large language models (LLMs) are capable of many natural language tasks, yet they are far from perfect. In health applications, grounding and interpreting domain-specific and non-linguistic data is crucial. This paper investigates the capacity of LLMs to make inferences about health based on contextual information (e.g. user demographics, health knowledge) and physiological data (e.g. resting heart rate, sleep minutes). We present a comprehensive evaluation of 12 state-of-the-art LLMs with prompting and fine-tuning techniques on four public health datasets (PMData, LifeSnaps, GLOBEM and AW_FB). Our experiments cover 10 consumer health prediction tasks in mental health, activity, metabolic, and sleep assessment. Our fine-tuned model, HealthAlpaca exhibits comparable performance to much larger models (GPT-3.5, GPT-4 and Gemini-Pro), achieving the best performance in 8 out of 10 tasks. Ablation studies highlight the effectiveness of context enhancement strategies. Notably, we observe that our context enhancement can yield up to 23.8% improvement in performance. While constructing contextually rich prompts (combining user context, health knowledge and temporal information) exhibits synergistic improvement, the inclusion of health knowledge context in prompts significantly enhances overall performance.

Discussion (0). Sign in to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 29 citations worldwide. Full citation record

  1. Auditing medical multi-agent AI reveals risks of false consensus

    cs.CL 2025-10 conditional novelty 6.0 of 10

    Medical AI doctor-teams often reach 'consensus' by repeating initial views and suppressing correct minorities, so high accuracy hides broken reasoning processes.

  2. Dynamic Uncertainty-aware Multimodal Fusion for Outdoor Health Monitoring

    cs.NI 2025-08 unverdicted novelty 6.0 of 10

    DUAL-Health is an uncertainty-aware multimodal fusion framework that quantifies sensor noise, customizes fusion weights accordingly, and aligns modality distributions to improve outdoor health monitoring.

  3. AnnoSense: A Framework for Physiological Emotion Data Collection in Everyday Settings for AI

    cs.HC 2025-07 conditional novelty 6.0 of 10

    The authors propose AnnoSense, a set of 15 expert-reviewed guidelines for everyday emotion data collection, derived from survey, interview, and focus group insights from 119 stakeholders.

  4. GLOSS: Group of LLMs for Open-Ended Sensemaking of Passive Sensing Data for Health and Wellbeing

    cs.HC 2025-07 conditional novelty 6.0 of 10

    A group of LLM agents that collaboratively generate code for raw passive sensing data outperforms RAG on objective query accuracy, while remaining only moderately consistent across repeated runs.

  5. DynamiCare: A Dynamic Multi-Agent Framework for Interactive and Open-Ended Medical Decision-Making

    cs.AI 2025-07 conditional novelty 6.0 of 10

    DynamiCare is a multi-agent LLM framework that runs multi-round diagnostic dialogues with a dynamically adjusted specialist team, evaluated on a new 500-patient benchmark built from MIMIC-III.

  6. SePA: A Search-enhanced Predictive Agent for Personalized Health Coaching

    cs.HC 2025-09 conditional novelty 5.0 of 10

    SePA combines personalized wearable-data risk prediction with a whitelisted web search pipeline to give cited, context-aware health coaching.

  7. TILES-2018 Sleep Benchmark Dataset: A Longitudinal Wearable Sleep Data Set of Hospital Workers for Modeling and Understanding Sleep Behaviors

    cs.HC 2025-07 conditional novelty 5.0 of 10

    A public benchmark dataset of ten weeks of Fitbit sleep data from 139 hospital workers, with sleep analyses and machine learning baselines for sleep quality, demographics, and sleep stages.

  8. From Text to Discovery: How Large Language Models Are Reshaping Research Across Scientific and Humanistic Disciplines

    cs.DL 2026-06 unverdicted novelty 3.0 of 10

    LLMs accelerate research workflows from idea generation to writing but introduce challenges like hallucination, bias, opacity, and ten systemic risks requiring new governance frameworks.

Pith tools