REVIEW 3 major objections 5 minor 1 cited by
SensorChat: Answering Qualitative and Quantitative Questions during Long-Term Multimodal Sensor Interactions
T0 review · 3 major / 5 minor · reviewed 2026-08-09 · deepseek-v4-flash
Pith's one-line read SensorChat claims to be the first end-to-end QA system that answers both qualitative and quantitative questions from days of raw, high-frequency multimodal sensor data, using a three-stage pipeline with a dedicated query stage.
desk verdict Solid systems paper with a genuinely new three-stage pipeline, but the headline 93% gain is not established because the sensor encoder is pretrained on the same subjects that appear in the SensorQA test split. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the query stage built on a partial-context contrastive pretraining loss, defined as Eq. (2) in the paper. For each 20-second window, the loss pushes the sensor embedding close to every positive short label phrase in that window, such as "sitting", "working on computers", and "at school", and away from all other phrases, instead of aligning to one full sentence the way CLIP-style losses do. At query time the text encoder encodes the extracted context, a similarity function scores every stored sensor embedding, and a set of small query functions, including CalculateDuration, CountingDays, and DetectFirstTime, convert the thresholded hits into a one-sentence textual summary. This is what lets the system compress days of 40Hz multimodal data into a few sentences that the answer LLM can reason over reliably.
What would settle it
A subject-disjoint evaluation would settle it: pretrain the sensor encoder only on a subset of subjects, say the first 48 of the 60 in SensorQA/ExtraSensory, then measure short-answer accuracy only on questions about the remaining subjects; if accuracy falls back toward the 0.28 baseline, the encoder is not transferring. The paper's Appendix I already shows query accuracy degrading for unseen users under such a split, making this an experiment the authors can run directly.
Extended reading notes
Core claim
On its own terms, SensorChat's central claim is that a dedicated sensor data query stage, placed between an LLM that turns a question into structured queries and an LLM that writes the answer, is what makes accurate question answering over long-duration, high-frequency sensor data possible. The sensor encoder and a text encoder are pretrained offline with a partial-context contrastive loss so that a raw 20-second window aligns with short phrases like "sitting" or "at school"; online, the query stage encodes the user's phrase, searches the stored sensor embeddings, and applies query functions such as CalculateDuration or CountingDays to emit a short textual summary. Only that summary plus the original question reaches the answer LLM, avoiding the token-limit and long-context failures that the paper documents in existing approaches. Evaluations on SensorQA report short-answer accuracy of 0.54 versus 0.28 for the best state-of-the-art baseline, and a user study with eight volunteers is offered as evidence that qualitative, open-ended questions can be answered with the same pipeline.
Load-bearing premise
The load-bearing premise is that the sensor encoder, pretrained on the ExtraSensory dataset, transfers to the people who will actually use the system, but the main benchmark is not a subject-disjoint transfer test because SensorQA's raw data comes from the same dataset the encoder was pretrained on, so real-world gains for new users are the assumption that carries the reported 93% improvement.
Editorial extensions
If this is right
- If SensorChat works as claimed, long-term personal monitoring becomes queryable in natural language: users can ask about activities, locations, and durations from raw IMU, audio, and phone-state data spanning multiple days and get precise numerical answers.
- Quantitative accuracy no longer depends on fitting the whole sensor history into an LLM context; a decomposed query, similarity search, and textual summary let the system run on edge devices, with 4-bit quantization costing only modest accuracy (SensorChatE 0.49 vs. SensorChatC 0.54 on short answers).
- Because the query stage converts sensor data into short text, answer assembly becomes a pure language task, which explains why a quantized LLM retains most of the performance and why the system can reuse general-purpose LLM reasoning.
- The same three-stage pipeline extends to qualitative questions by querying a list of relevant contexts, such as sitting, exercising, and talking, and letting the answer LLM combine the summary with world knowledge; the user study gives initial evidence for this.
- Duration breakdowns show SensorChat is stronger on multi-day queries (0.58 short-answer accuracy) than single-day ones (0.46), the opposite of prior baselines, suggesting the system is most valuable exactly where existing methods collapse.
Reading between the lines
- Beyond the paper's claims: the query-stage design should transfer to other long, high-frequency modalities such as ECG, continuous audio, or environmental sensors, since everything downstream depends only on aligned embeddings and threshold-based counting once a modality encoder is swapped in.
- Beyond the paper's claims: the clearest unrealized risk is encoder transfer; a subject-disjoint evaluation is the direct test, and the paper's Appendix I already hints that query accuracy drops for unseen users even when the answer-assembly stage generalizes well.
- Beyond the paper's claims: if the pipeline is correct, per-question compute can be bounded by the number of relevant sensor segments, which points toward personal assistants that answer on-device with a modest LLM and pre-indexed sensor embeddings, a direction the paper leaves for future latency optimization.
- Beyond the paper's claims: for qualitative questions, one could run a planted-behavior test in which a participant performs a scripted sequence while wearing sensors and then asks about it; comparing the answer to the script would separate query-stage retrieval errors from LLM reasoning errors.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents SensorChat, an end-to-end question-answering system for long-duration, high-frequency multimodal sensor data. The system uses a three-stage pipeline: an LLM-based question decomposition stage, an embedding-based sensor data query stage that retrieves relevant time windows via a contrastively pretrained sensor/text encoder, and a LoRA-finetuned LLM answer assembly stage. The authors evaluate on the SensorQA quantitative benchmark, report up to 93% higher short-answer accuracy than state-of-the-art baselines, and complement this with an eight-participant user study covering qualitative questions. Two deployment variants are described: SensorChatC on a cloud GPU and SensorChatE on an edge device with quantized LLMs.
Significance. If the quantitative results hold, SensorChat's central architectural idea—inserting an explicit sensor-data query stage between an LLM question decomposer and an LLM answer assembler—is a genuinely useful response to the token-limit and long-context problems that plague prior sensor-QA systems. The paper includes careful ablations (Table 7), sensitivity analyses (Appendix G), latency measurements on both cloud and edge platforms (Sec. 6.4), and a real-user deployment with an IRB-approved study. These are real strengths. However, the headline quantitative claim is currently not supported as stated because the sensor encoder is pretrained on the same ExtraSensory data that supplies the SensorQA test users, without a subject-wise split, and Appendix I shows that query accuracy degrades on unseen users. The qualitative claim rests on a small, subjective user study. The work is promising but needs a substantially strengthened evaluation before the central claims can be accepted.
major comments (3)
- [Sec. 6.1 / Sec. 6.2 / Appendix I] The main quantitative claim is compromised by the sensor-encoder pretraining setup. Section 6.1 states that the encoders are pretrained on ExtraSensory, the same dataset that supplies the raw sensor data for SensorQA, and Section 6.2 states that the QA pairs are split 80/20 randomly rather than by subject. Consequently, the SensorChat sensor encoder has seen the raw sensor signals of the users who appear in the test set, while none of the baselines in Table 5 receive this additional exposure. This is not a hypothetical concern: Appendix I and Fig. 14b show that when pretraining is limited to users 1-48, online querying accuracy degrades on unseen users, and Sec. 7 attributes real-user errors to the encoder's limited ability to generalize. The reported 93% improvement in Table 5 therefore may largely reflect test-user exposure rather than a generalizable query stage. The evaluation should be rerun with a subject-disjoint split (or an encoder pretrained without any test-user data), and the seen-user and unseen-user results should be reported separately.
- [Sec. 4.2 / Appendix C] The SensorQA benchmark is not an independent test of the system's generality. The benchmark was created by several co-authors of this manuscript, and the question-decomposition templates in Sec. 4.2 and Appendix C are manually designed with two solution templates for each of the six SensorQA question categories. This means the main benchmark evaluates the system on a taxonomy that the system was explicitly built to handle. The claim that SensorChat handles 'user-defined' or 'arbitrary' questions needs support from a benchmark or evaluation constructed independently of the system's design; the current eight-participant user study is too small to provide that support.
- [Sec. 7] The qualitative-claim evidence in Sec. 7 is too weak to carry the title's second half. The study has eight volunteers, each contributing one to three days of data, no objective ground truth or comparison to alternative systems, and the average answer-content rating is 3.12 out of 5. The authors themselves note that participants reported 'some numbers were a little off' and 'it mentioned activities I never did,' and attribute this to the sensor encoder's limited generalization. These results are consistent with a useful prototype but do not demonstrate that SensorChat provides accurate qualitative answers. This needs either a larger, quantitatively evaluated study or a clear reframing as a feasibility demonstration.
minor comments (5)
- [Table 3] The phrase 'The shorted version of the answers' should be 'The shortened version of the answers.'
- [Sec. 6.1] The sentence 'which servers as the sensor data source for SensorQA' contains a typo: 'servers' should be 'serves.'
- [Sec. 4.3.2] The phrase 'including time quries, activity quries' contains a typo: 'quries' should be 'queries.'
- [Table 5 / Sec. 6.3] The Rouge-1 and Rouge-L gains over text-only LLaMA2-7B are small (0.77 vs. 0.72 and 0.75 vs. 0.72, respectively); the 'outperforms' wording based on full answers should be qualified, since the more striking claim rests on the strict short-answer metric.
- [Sec. 6.1] The multiple-choice variant is generated by GPT-3.5-Turbo from the short answers; this should be described explicitly as a synthetic extension of the original benchmark, and the possibility of answer-order or phrasing bias should be acknowledged.
Circularity Check
Headline quantitative gain is not a held-out prediction: the sensor encoder is pretrained on the same ExtraSensory subjects that generate the SensorQA test queries, and the paper's own Appendix I admits degraded accuracy on unseen users.
-
fitted input called prediction
[Sec. 6.1 (Dataset and Metrics), Sec. 3.2 (Study Setup), Appendix I (Generalization and Robustness)]
"To ensure the best alignment between the QA pairs and sensor information, we conduct offline encoders pretraining on the ExtraSensory multimodal sensor dataset [53], which servers as the sensor data source for SensorQA. ... We randomly split the data in SensorQA [48] into an 80/20 train-test set. ... limiting sensor data to the first 48 users leads to accuracy degradation on unseen users due to variations in data distributions."
The paper's central quantitative claim, the 93% short-answer accuracy improvement in Table 5, depends on the sensor query stage retrieving correct segments. That stage's sensor encoder is pretrained on the ExtraSensory dataset, and SensorQA's raw sensor data comes from the same ExtraSensory dataset. The 80/20 split is over QA pairs, not over subjects, so sensor segments used for test queries were very likely included in the encoder's pretraining corpus. Thus the 'prediction' on test users is partly a measure of the encoder's exposure to those same users' sensor data, not a generalization to new users. The paper's Appendix I provides the internal check: when pretraining is restricted to users 1-48, query accuracy degrades on unseen users, showing the query stage does not transfer.
full rationale
The SensorChat system itself is an engineering contribution with a three-stage pipeline, and several components are evaluated in a self-contained way: the answer-assembly LLM is LoRA-finetuned on an 80/20 split of SensorQA QA pairs, the latency measurements are independent, and the ablation and sensitivity studies provide internal evidence about design choices. However, the headline quantitative result is not a clean out-of-sample evaluation. The sensor encoder used in the query stage is pretrained on the ExtraSensory dataset, and SensorQA's raw sensor data is drawn from that same dataset. Because the QA-pair split is not subject-wise, the test queries concern sensor segments from users whose data participated in encoder pretraining. Appendix I explicitly documents that the query stage degrades on unseen users, so the 93% improvement reflects, at least in part, the encoder having been fitted on the test users' sensor data. This is a fitted-input-called-prediction pattern: a parameter-heavy component is trained on the evaluation population and then the system's accuracy on that population is reported as a predicted capability. The user study with eight volunteers is real but is non-comparative and does not quantify accuracy against baselines, so it does not rescue the Table 5 claim. I do not see other circular steps: there is no self-definitional reduction, no uniqueness theorem imported from the authors, and no renamed known result. Score 6 rather than higher because the answer-assembly finetuning is held out and the pipeline's other evaluations stand on their own.
Assumptions & free parameters
free parameters (5)
- query threshold h =
0.5
- contrastive temperature tau =
0.1
- LoRA rank =
16
- LoRA learning rate =
2e-4
- number of solution templates per question category =
2
assumptions (5)
- domain assumption ExtraSensory context labels are correct and sufficient for aligning sensor signals with user-relevant activities and locations.
- domain assumption The cosine similarity between sensor and text embeddings is a reliable indicator of semantic relevance for arbitrary user questions.
- domain assumption SensorQA is a valid measure of practical quantitative QA performance.
- domain assumption LLMs can reliably decompose user questions into a context phrase, time range, and query function.
- domain assumption The 20-second time window is fine enough to capture activities of interest.
Cite this review
Pith. "Pith review of SensorChat: Answering Qualitative and Quantitative Questions during Long-Term Multimodal Sensor Interactions." pith.science (2026). https://pith.science/paper/D4UVMMKX
@misc{pith2026250202883,
author = {Pith},
title = {Pith review of: SensorChat: Answering Qualitative and Quantitative Questions during Long-Term Multimodal Sensor Interactions},
year = {2026},
howpublished = {\url{https://pith.science/paper/D4UVMMKX}},
note = {Machine review of arXiv:2502.02883}
}
read the original abstract
Natural language interaction with sensing systems is crucial for addressing users' personal concerns and providing health-related insights into their daily lives. When a user asks a question, the system automatically analyzes the full history of sensor data, extracts relevant information, and generates an appropriate response. However, existing systems are limited to short-duration (e.g., one minute) or low-frequency (e.g., daily step count) sensor data. In addition, they struggle with quantitative questions that require precise numerical answers. In this work, we introduce SensorChat, the first end-to-end QA system designed for daily life monitoring using long-duration, high-frequency time series data. Given raw sensor signals spanning multiple days and a user-defined natural language question, SensorChat generates semantically meaningful responses that directly address user concerns. SensorChat effectively handles both quantitative questions that require numerical precision and qualitative questions that require high-level reasoning to infer subjective insights. To achieve this, SensorChat uses an innovative three-stage pipeline including question decomposition, sensor data query, and answer assembly. The first and third stages leverage Large Language Models (LLMs) to interpret human queries and generate responses. The intermediate querying stage extracts relevant information from the complete sensor data history. Real-world implementations demonstrate SensorChat's capability for real-time interactions on a cloud server while also being able to run entirely on edge platforms after quantization. Comprehensive QA evaluations show that SensorChat achieves 93% higher answer accuracy than the best performing state-of-the-art systems on quantitative questions. Furthermore, a user study with eight volunteers highlights SensorChat's effectiveness in answering qualitative questions.
Figures
Figures from the paper (12 more)
Forward citations
Cited by 1 Pith paper
-
SensorLM: Learning the Language of Wearable Sensors
SensorLM is a sensor-language foundation model trained on 59.7M hours of wearable data with template-generated captions, reporting strong zero-shot, few-shot, and retrieval performance.
Reference graph
Works this paper leans on
- [1]
-
[2]
2025. Jetson Orin NX Module. https://developer.nvidia.com/embedded/jetson-modules#jetson_orin_nx. [Online]
work page 2025
-
[3]
2025. NVIDIA A100 Tensor Core GPU. https://www.nvidia.com/en-us/data-center/a100/. [Online]
work page 2025
-
[4]
2025. OpenAI o3-mini. https://openai.com/index/openai-o3-mini/. [Online]
work page 2025
-
[5]
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774 (2023)
arXiv 2023
-
[6]
Deepa Tavargeri Adiga, Maitry Bhavsar, Unnati Palan, and Sachin Patel. 2020. Daily Journals: Extracting Insights for Well-being. In Proceedings of the 14th EAI International Conference on Pervasive Computing Technologies for Healthcare . 305–315
work page 2020
-
[7]
Kiyoharu Aizawa, Datchakorn Tancharoen, Shinya Kawasaki, and Toshihiko Yamasaki. 2004. Efficient retrieval of life log based on context and content. In Proceedings of the the 1st ACM workshop on Continuous archival and retrieval of personal experiences . 22–31
work page 2004
-
[8]
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al. 2022. Flamingo: a visual language model for few-shot learning. Advances in neural information processing systems 35 (2022), 23716–23736
2022
Show all 72 references
-
[9]
Riku Arakawa, Jill Fain Lehman, and Mayank Goel. 2024. PrISM-Q&A: Step-Aware Voice Assistant on a Smartwatch Enabled by Multimodal Procedure Tracking and Large Language Models. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies 8, 4 (2024), 1–26
2024
-
[10]
Qiming Cao, Hongfei Xue, Tianci Liu, Xingchen Wang, Haoyu Wang, Xincheng Zhang, and Lu Su. 2024. mmCLIP: Boosting mmWave- based Zero-shot HAR via Signal-Text Alignment. In Proceedings of the 22nd ACM Conference on Embedded Networked Sensor Systems . 184–197
2024
-
[11]
Wenqiang Chen, Jiaxuan Cheng, Leyao Wang, Wei Zhao, and Wojciech Matusik. 2024. Sensor2Text: Enabling Natural Language Interactions for Daily Activity Tracking Using Wearable Sensors. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies 8, 4 (2024), 1–26
2024
-
[12]
Gonzalez, Ion Stoica, and Eric P
Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. 2023. Vicuna: An Open-Source Chatbot Impressing GPT-4 with 90%* ChatGPT Quality. https://lmsys.org/blog/2023-...
2023
-
[13]
Wonyoung Choi, Jisu Kim, SangEun Lee, and Eunil Park. 2021. Smart home and internet of things: A bibliometric study. Journal of Cleaner Production 301 (2021), 126908. , Vol. 1, No. 1, Article . Publication date: September 2025. SensorChat: Answering Qualitative and Quantitativ...
2021
-
[14]
Ranak Roy Chowdhury, Ritvik Kapila, Ameya Panse, Xiyuan Zhang, Diyan Teng, Rashmi Kulkarni, Dezhi Hong, Rajesh K Gupta, and Jingbo Shang. 2025. ZeroHAR: Sensor Context Augments Zero-Shot Wearable Action Recognition. InProceedings of the AAAI Conference on Artificial Intelligen...
2025
-
[15]
Zheng Chu, Jingchang Chen, Qianglong Chen, Weijiang Yu, Tao He, Haotian Wang, Weihua Peng, Ming Liu, Bing Qin, and Ting Liu
-
[16]
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805 (2018)
2018 arXiv
-
[17]
Zachary Englhardt, Chengqian Ma, Margaret E Morris, Chun-Cheng Chang, Xuhai" Orson" Xu, Lianhui Qin, Daniel McDuff, Xin Liu, Shwetak Patel, and Vikram Iyer. 2024. From Classification to Clinical Insights: Towards Analyzing and Reasoning About Mobile and Behavioral Health Data ...
2024
-
[18]
Matan Eyal, Tal Baumel, and Michael Elhadad. [n. d.]. Question Answering as an Automatic Evaluation Metric for News Article Summarization. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Techn...
2019
-
[19]
Yonggan Fu, Yongan Zhang, Zhongzhi Yu, Sixu Li, Zhifan Ye, Chaojian Li, Cheng Wan, and Yingyan Celine Lin. 2023. Gpt4aigchip: Towards next-generation ai accelerator design automation via large language models. In 2023 IEEE/ACM International Conference on Computer Aided Design ...
2023
-
[20]
Albert Gu and Tri Dao. 2023. Mamba: Linear-time sequence modeling with selective state spaces. arXiv preprint arXiv:2312.00752 (2023)
2023 arXiv
-
[21]
Jiaming Han, Kaixiong Gong, Yiyuan Zhang, Jiaqi Wang, Kaipeng Zhang, Dahua Lin, Yu Qiao, Peng Gao, and Xiangyu Yue. 2023. Onellm: One framework to align all modalities with language. arXiv preprint arXiv:2312.03700 (2023)
2023 arXiv
-
[22]
Jiaming Han, Kaixiong Gong, Yiyuan Zhang, Jiaqi Wang, Kaipeng Zhang, Dahua Lin, Yu Qiao, Peng Gao, and Xiangyu Yue. 2024. Onellm: One framework to align all modalities with language. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
2024
-
[23]
Gerhard Henrich and Peter Herschbach. 2000. Questions on Life Satisfaction (FLZM): a short questionnaire for assessing subjective quality of life. European Journal of Psychological Assessment 16, 3 (2000), 150
2000
-
[24]
Aritra Hota, Soumyajit Chatterjee, and Sandip Chakraborty. 2024. Evaluating large language models as virtual annotators for time-series physical sensing data. arXiv preprint arXiv:2403.01133 (2024)
2024 arXiv
-
[25]
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2021. Lora: Low-rank adaptation of large language models. arXiv preprint arXiv:2106.09685 (2021)
2021 arXiv
-
[26]
Lei Huang, Weijiang Yu, Weitao Ma, Weihong Zhong, Zhangyin Feng, Haotian Wang, Qianglong Chen, Weihua Peng, Xiaocheng Feng, Bing Qin, et al. 2023. A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions. arXiv preprint arXiv:231...
2023 arXiv
-
[27]
Keum-Sung Hwang and Sung-Bae Cho. 2014. A Lifelog Browser for Visualization and Search of Mobile Everyday-Life.Mobile Information Systems 10, 3 (2014), 243–258
2014
-
[28]
Sijie Ji, Xinzhe Zheng, Jiawei Sun, Renqi Chen, Wei Gao, and Mani Srivastava. 2024. MindGuard: Towards Accessible and Sitgma-free Mental Health First Aid via Edge LLM. arXiv preprint arXiv:2409.10064 (2024)
2024 arXiv
-
[29]
Sijie Ji, Xinzhe Zheng, and Chenshu Wu. 2024. Hargpt: Are llms zero-shot human activity recognizers?. In 2024 IEEE International Workshop on Foundation Models for Cyber-Physical Systems & Internet of Things (FMSys) . IEEE, 38–43
2024
-
[30]
Prannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna, Yonglong Tian, Phillip Isola, Aaron Maschinot, Ce Liu, and Dilip Krishnan
-
[31]
Yubin Kim, Xuhai Xu, Daniel McDuff, Cynthia Breazeal, and Hae Won Park. 2024. Health-llm: Large language models for health prediction via wearable sensor data. arXiv preprint arXiv:2401.06866 (2024)
2024 arXiv
-
[32]
Tianle Li, Ge Zhang, Quy Duc Do, Xiang Yue, and Wenhu Chen. 2024. Long-context llms struggle with long in-context learning. arXiv preprint arXiv:2404.02060 (2024)
2024 arXiv
-
[33]
Ji Lin, Jiaming Tang, Haotian Tang, Shang Yang, Wei-Ming Chen, Wei-Chen Wang, Guangxuan Xiao, Xingyu Dang, Chuang Gan, and Song Han. 2024. AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration. InMLSys
2024
-
[34]
Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. 2024. Improved baselines with visual instruction tuning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 26296–26306
2024
-
[35]
Xin Liu, Daniel McDuff, Geza Kovacs, Isaac Galatzer-Levy, Jacob Sunshine, Jiening Zhan, Ming-Zher Poh, Shun Liao, Paolo Di Achille, and Shwetak Patel. 2023. Large language models are few-shot health learners. arXiv preprint arXiv:2305.15525 (2023)
2023 arXiv
-
[36]
Zechun Liu, Changsheng Zhao, Forrest Iandola, Chen Lai, Yuandong Tian, Igor Fedorov, Yunyang Xiong, Ernie Chang, Yangyang Shi, Raghuraman Krishnamoorthi, et al. 2024. Mobilellm: Optimizing sub-billion parameter language models for on-device use cases. arXiv preprint arXiv:2402...
2024 arXiv
-
[37]
Pan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu, Kai-Wei Chang, Song-Chun Zhu, Oyvind Tafjord, Peter Clark, and Ashwin Kalyan
-
[38]
Shentong Mo, Russ Salakhutdinov, Louis-Philippe Morency, and Paul Pu Liang. 2024. Iot-lm: Large multisensory language models for the internet of things. arXiv preprint arXiv:2407.09801 (2024)
2024 arXiv
-
[39]
Seungwhan Moon, Andrea Madotto, Zhaojiang Lin, Tushar Nagarajan, Matt Smith, Shashank Jain, Chun-Fu Yeh, Prakash Murugesan, Peyman Heidari, Yue Liu, et al . 2023. Anymal: An efficient and scalable any-modality augmented language model. arXiv preprint arXiv:2309.16058 (2023)
2023 arXiv
-
[40]
Seungwhan Moon, Andrea Madotto, Zhaojiang Lin, Aparajita Saraf, Amy Bearman, and Babak Damavandi. 2023. IMU2CLIP: Language- grounded Motion Sensor Translation with Multimodal Contrastive Learning. In Findings of the Association for Computational Linguistics: EMNLP 2023. 13246–13253
2023
-
[41]
Jingping Nie, Minghui Zhao, Stephen Xia, Xinghua Sun, Hanya Shao, Yuang Fan, Matthias Preindl, and Xiaofan Jiang. 2022. Ai therapist for daily functioning assessment and intervention using smart home devices. In Proceedings of the 20th ACM Conference on Embedded Networked Sens...
2022
-
[42]
Xiaomin Ouyang, Xian Shuai, Jiayu Zhou, Ivy Wang Shi, Zhiyuan Xie, Guoliang Xing, and Jianwei Huang. 2022. Cosmo: contrastive fusion learning with small data for multimodal human activity recognition. In Proceedings of the 28th Annual International Conference on Mobile Computi...
2022
-
[43]
Xiaomin Ouyang and Mani Srivastava. 2024. LLMSense: Harnessing LLMs for high-level reasoning over spatiotemporal sensor traces. In 2024 IEEE 3rd Workshop on Machine Learning on Edge in Sensor Systems (SenSys-ML) . IEEE, 9–14
2024
-
[44]
Xiaomin Ouyang, Zhiyuan Xie, Heming Fu, Sitong Cheng, Li Pan, Neiwen Ling, Guoliang Xing, Jiayu Zhou, and Jianwei Huang. 2023. Harmony: Heterogeneous Multi-Modal Federated Learning through Disentangled Model Training. In Proceedings of the 21st Annual International Conference ...
2023
-
[45]
Adam Paszke et al. 2019. Pytorch: An imperative style, high-performance deep learning library.Advances in neural information processing systems 32 (2019)
2019
-
[46]
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. In International conference on machine learni...
2021
-
[47]
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer. Journal of Machine Learning Research 21, 140 (2020), 1–67. htt...
2020
-
[48]
Benjamin Reichman, Xiaofan Yu, Lanxiang Hu, Jack Truxal, Atishay Jain, Rushil Chandrupatla, Tajana Šimunić Rosing, and Larry Heck
-
[49]
Zhenwei Shao, Zhou Yu, Meng Wang, and Jun Yu. 2023. Prompting large language models with answer heuristics for knowledge-based visual question answering. In Proceedings of the IEEE/CVF Conference on computer vision and pattern recognition . 14974–14983
2023
-
[50]
Kohei Takahashi, Shinsuke Matsumoto, Sachio Saiki, and Masahide Nakamura. 2013. Design and evaluation of lifelog mashup platform with NoSQL database. In Proceedings of International Conference on Information Integration and Web-based Applications & Services . 133–139
2013
-
[51]
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971 (2023)
2023 arXiv
-
[52]
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, W...
2023 arXiv
-
[53]
Yonatan Vaizman, Katherine Ellis, and Gert Lanckriet. 2017. Recognizing detailed human context in the wild from smartphones and smartwatches. IEEE pervasive computing 16, 4 (2017), 62–74
2017
-
[54]
Yonatan Vaizman, Katherine Ellis, Gert Lanckriet, and Nadir Weibel. 2018. Extrasensory app: Data collection in-the-wild with rich user interface to self-report behavior. In Proceedings of the 2018 CHI conference on human factors in computing systems . 1–12
2018
-
[55]
A Vaswani. 2017. Attention is all you need. Advances in Neural Information Processing Systems (2017). , Vol. 1, No. 1, Article . Publication date: September 2025. SensorChat: Answering Qualitative and Quantitative Questions during Long-Term Multimodal Sensor Interactions • 23
2017
-
[56]
Yuxuan Weng, Guoquan Wu, Tianyue Zheng, Yanbing Yang, and Jun Luo. 2024. Large Model for Small Data: Foundation Model for Cross-Modal RF Human Activity Recognition. In Proceedings of the 22nd ACM Conference on Embedded Networked Sensor Systems . 436–449
2024
-
[57]
Tianwei Xing, Luis Garcia, Federico Cerutti, Lance Kaplan, Alun Preece, and Mani Srivastava. 2021. DeepSQA: Understanding Sensor Data via Question Answering. In Proceedings of the International Conference on Internet-of-Things Design and Implementation . 106–118
2021
-
[58]
Huatao Xu, Liying Han, Qirui Yang, Mo Li, and Mani Srivastava. 2024. Penetrative ai: Making llms comprehend the physical world. In Proceedings of the 25th International Workshop on Mobile Computing Systems and Applications . 1–7
2024
-
[59]
Huatao Xu, Panrong Tong, Mo Li, and Mani Srivastava. 2024. AutoLife: Automatic Life Journaling with Smartphones and LLMs. arXiv preprint arXiv:2412.15714 (2024)
2024 arXiv
-
[60]
Huatao Xu, Pengfei Zhou, Rui Tan, and Mo Li. 2023. Practically Adopting Human Activity Recognition. InProceedings of the 29th Annual International Conference on Mobile Computing and Networking . 1–15
2023
-
[61]
Lilin Xu, Chaojie Gu, Rui Tan, Shibo He, and Chen Jiming. 2023. MESEN: Exploit Multimodal Data to Design Unimodal Human Activity Recognition with Few Labels. In Proceedings of the 21st ACM Conference on Embedded Networked Sensor Systems
2023
-
[62]
Bufang Yang, Siyang Jiang, Lilin Xu, Kaiwei Liu, Hai Li, Guoliang Xing, Hongkai Chen, Xiaofan Jiang, and Zhenyu Yan. 2024. DrHouse: An LLM-empowered Diagnostic Reasoning System through Harnessing Outcomes from Sensor Data and Expert Knowledge. arXiv preprint arXiv:2405.12541 (2024)
2024 arXiv
-
[63]
Junjie Ye, Xuanting Chen, Nuo Xu, Can Zu, Zekai Shao, Shichun Liu, Yuhan Cui, Zeyang Zhou, Chao Gong, Yang Shen, et al. 2023. A comprehensive capability analysis of gpt-3 and gpt-3.5 series models. arXiv preprint arXiv:2303.10420 (2023)
2023 arXiv
-
[64]
Hyungjun Yoon, Biniyam Aschalew Tolera, Taesik Gong, Kimin Lee, and Sung-Ju Lee. 2024. By my eyes: Grounding multimodal large language models with sensor data via visual prompting. arXiv preprint arXiv:2407.10385 (2024)
2024 arXiv
-
[65]
Renrui Zhang, Jiaming Han, Chris Liu, Peng Gao, Aojun Zhou, Xiangfei Hu, Shilin Yan, Pan Lu, Hongsheng Li, and Yu Qiao. 2023. Llama-adapter: Efficient fine-tuning of language models with zero-init attention. arXiv preprint arXiv:2303.16199 (2023)
2023 arXiv
-
[66]
Penghao Zhao, Hailin Zhang, Qinhan Yu, Zhengren Wang, Yunteng Geng, Fangcheng Fu, Ling Yang, Wentao Zhang, and Bin Cui. 2024. Retrieval-augmented generation for ai-generated content: A survey. arXiv preprint arXiv:2402.19473 (2024)
2024 arXiv
-
[67]
Xuanhe Zhou, Xinyang Zhao, and Guoliang Li. 2024. LLM-Enhanced Data Management. arXiv preprint arXiv:2402.02643 (2024)
2024 arXiv
-
[68]
How long was I at school?
Yan Zhuang, Zhenzhe Zheng, Fan Wu, and Guihai Chen. 2024. LiteMoE: Customizing On-device LLM Serving via Proxy Submodel Tuning. In Proceedings of the 22nd ACM Conference on Embedded Networked Sensor Systems . 521–534. A DIFFERENT WAYS OF FEEDING SENSOR DATA TO LLMS Existing wo...
2024
-
[2020]
Advances in neural information processing systems 33 (2020), 18661–18673
Supervised contrastive learning. Advances in neural information processing systems 33 (2020), 18661–18673
2020
-
[2022]
Advances in Neural Information Processing Systems 35 (2022), 2507–2521
Learn to explain: Multimodal reasoning via thought chains for science question answering. Advances in Neural Information Processing Systems 35 (2022), 2507–2521
2022
-
[2024]
In The 62nd Annual Meeting of the Association for Computational Linguistics: ACL 2024, Bangkok, Thailand, August 11–16, 2024
Navigate through Enigmatic Labyrinth A Survey of Chain of Thought Reasoning: Advances, Frontiers and Future. In The 62nd Annual Meeting of the Association for Computational Linguistics: ACL 2024, Bangkok, Thailand, August 11–16, 2024 . Association for Computational Linguistics...
2024 arXiv
-
[2025]
arXiv preprint arXiv:2501.04974 (2025)
SensorQA: A Question Answering Benchmark for Daily-Life Monitoring. arXiv preprint arXiv:2501.04974 (2025). ACM Conference on Embedded Networked Sensor Systems (2025)
2025 arXiv
Reviewed August 9, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.