Pith. sign in

REVIEW 5 major objections 4 minor 1 cited by

SensorQA: A Question Answering Benchmark for Daily-Life Monitoring

T0 review · 5 major / 4 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read SensorQA is the first human-created question-answering benchmark for long-term time-series sensor data, and existing AI models answer at most 28% of its questions exactly.

desk verdict A genuinely new human-authored QA benchmark for long-term sensor data, with thoughtful collection design; but unvalidated ground-truth answers and a missing human baseline make the headline accuracy numbers provisional. read the letter →

arxiv 2501.04974 v3 pith:YZCAESWI submitted 2025-01-09 cs.CL cs.AI

classification cs.CLcs.AI
keywords questionansweringsensordatabenchmarkhuman-createddatasetdaily-lifemonitoringlargelanguagemodelsmultimodalreasoningwearablesensors
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper introduces SensorQA, which it presents as the first human-created benchmark for question answering over long-term time-series sensor data in daily-life monitoring. The dataset contains 5,648 question-answer pairs built by showing Amazon Mechanical Turk workers activity graphs derived from the ExtraSensory dataset and asking them to write first-person questions and second-person answers. Benchmarking state-of-the-art models shows a large accuracy gap: the best baseline, Llama-Adapter, reaches 28% exact-match accuracy, and sensor-plus-text models do worse than text-only models. The authors argue this gap, together with impractical latency on edge devices, shows that current AI is not ready to serve as a conversational personal sensing assistant.

What carries the argument

The mechanism that carries the argument is the QA-collection protocol: workers receive activity graphs, which are color-coded, Gantt-style visualizations of activity and context labels over time, and are instructed to role-play as the device owner, generating a first-person question and a second-person answer. Multi-time-scale graphs steer workers toward both short quantitative queries and long-horizon qualitative queries, while 14 label subsets cover postures, location, diet, sleep, commute, exercise, electronics, social life, work, and work-life balance. This protocol converts raw sensor streams into human-readable ground truth, which is then used to train and evaluate models.

What would settle it

Take a random sample of SensorQA answer pairs and re-derive the answers directly from the raw ExtraSensory sensor streams, such as accelerometer, location, and audio data, rather than from the pre-computed activity labels; if a substantial fraction disagree, the benchmark's ground truth is the weak link.

Watch

Extended reading notes

Core claim

On its own terms, the paper's central discovery is that today's AI models cannot reliably answer realistic human questions about weeks of wearable sensing data. By visualizing ExtraSensory activity and context labels as multi-time-scale Gantt-like graphs (daily graphs with a timetable, and multi-day graphs) and using 14 label subsets, the authors collect 5,648 human-authored QA pairs spanning six question categories and seven answer categories. They then benchmark text-only, vision+text, and sensor+text models and find the best exact-match accuracy is only 28%, with sensor+text fusion underperforming text-only baselines and time-related questions being hardest. This is taken as evidence that the missing ingredient is a realistic, long-duration QA benchmark and that new model architectures or training methods are needed.

Load-bearing premise

The load-bearing premise is that the answers in the 5,648 QA pairs are correct because they are read, by crowd workers, from activity graphs built on ExtraSensory labels; if those labels or readings are wrong, the benchmark numbers collapse.

Editorial extensions

If this is right

  • Any deployable personal AI assistant over wearable data must handle multi-day temporal context, not just short clips.
  • New sensor-text fusion methods are needed because feeding sensor information to models currently hurts performance relative to text-only prompting.
  • Time-related questions, such as durations, timestamps, and day comparisons, are the hardest and most distinctive category, so future work should target temporal reasoning.
  • Fine-tuning on the target data matters: Llama-Adapter goes from 0% exact-match accuracy without fine-tuning to 28% with it.
  • Billion-parameter LLMs are impractical for edge deployment, with average answer generation latency over 57 seconds on a Jetson TX2.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Because the ground truth is derived from ExtraSensory labels, a label-quality audit or a second round of human verification against raw sensor streams would materially strengthen the benchmark; the paper does not report one.
  • The paper's result that sensor+text fusion hurts performance suggests current adapters compress long time series into lossy text, so models that reason over event intervals or temporal graphs might beat the 28% ceiling.
  • The same graph-plus-label-subset recipe could create QA benchmarks for other long-term time-series domains, such as energy meters, sleep clinics, or vehicle telematics.
  • A practical corollary the authors leave implicit is that a viable edge assistant may need to run small models on distilled summaries rather than billion-parameter LLMs to meet latency and memory constraints.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 4 minor

Summary. The manuscript introduces SensorQA, a benchmark of 5,648 question-answer pairs built by AMT workers from activity/context-label graphs derived from the ExtraSensory dataset. The authors benchmark text-only, vision+text, and sensor+text baselines on both full-answer and short-answer versions, reporting a best exact-match accuracy of 28% (Llama-Adapter), and measure memory and latency on the Jetson TX2. The paper's main claims are that SensorQA is the first human-created QA benchmark for long-term time-series sensor data and that current AI models exhibit a substantial gap to optimal QA performance and efficiency.

Significance. If the reference answers and evaluation splits are validated, SensorQA would fill a real gap: existing sensor QA datasets are either template-generated (DeepSQA) or limited to short, fixed-duration signals (AnyMAL, OneLLM). The dataset is open-sourced, covers multiple timescales and diverse question/answer categories, and includes both accuracy and edge-deployment measurements, which are useful for the community. However, the quantitative conclusions currently rest on unvalidated reference answers, a potentially leaking train/test split, and no human performance baseline, so the significance of the reported numbers is not yet established.

major comments (5)
  1. [§3.2 and §5.2] The ground-truth answers are not validated: the paper reports no inter-annotator agreement, no second-pass review, and no consistency check against the label intervals in the activity graphs. Workers were instructed to ask questions that 'would interest someone' with a smart device, which invites subjective and evaluative questions such as the work-life balance example in Section 1, whose answers are not uniquely determined by the graph. Because every reported metric is computed against these reference answers, any noise in the ExtraSensory labels (self-reported and only described as 'after cleaning' in Section 3.1, without a protocol) or any worker error propagates directly into the benchmark numbers. Please add an annotation-quality study, including agreement metrics, a sample of questions verified against the raw label timestamps, and a discussion of rejected or corrected pairs.
  2. [§5.1] The random 80/20 train/test split is performed at the QA-pair level, not at the user or activity-graph level. Since multiple QA pairs are generated from the same graph and the same user can appear in multiple graphs, the same underlying activity pattern can appear in both training and testing, potentially inflating the fine-tuned model accuracies. Please evaluate on a user-disjoint or graph-disjoint split, or at minimum report results for both a random split and a disjoint split and discuss any performance difference.
  3. [Abstract and §5.2] The claim that the results reveal 'a gap between current models and optimal QA performance' is not supported by the evidence reported in the paper, because no human performance baseline is provided. A human baseline on a representative sample of test questions, scored with the same exact-match protocol, would calibrate what 'optimal' means for this benchmark and is needed to quantify how large the gap actually is. Please add such a baseline or substantially soften the claim.
  4. [§1, §3, and Table 4] The abstract states that answers are 'derived from sensor data,' but the QA pairs are actually generated from activity/context label graphs, not from raw sensor signals. This distinction changes what the benchmark tests: reading summarized activity labels versus reasoning over raw multimodal time-series data. It also affects the interpretation of the Sensor+Text baselines, which are evaluated on label-derived answers rather than raw signal inference. Please state this precisely, describe the label-cleaning protocol in Section 3.1, and discuss the implications for the claimed sensor-based QA contribution.
  5. [§5.1 and §5.2] The exact-match accuracy is computed from keywords extracted by GPT-3.5-Turbo from the full answers, with no validation of the extraction quality. The reported accuracy therefore depends on an unverified intermediate step, and the metric name 'exact-match' is misleading because the check is keyword containment rather than exact string or semantic equivalence. Please validate the distilled short answers on a human-annotated sample, report the extraction accuracy, and either rename or redefine the metric to match what is actually computed.
minor comments (4)
  1. [Table 4] DeepSQA reports an exact-match accuracy of 27.4% while achieving a Bleu score of 0.0; this unusual combination should be explained, since it is not typical for n-gram metrics to be zero when keywords are present.
  2. [Figure 3] The statement that Llama performs 'only slightly better than random guessing, with an accuracy of 58%' needs a defined random baseline; for Yes/No questions the chance level depends on the label distribution, which should be reported.
  3. [§5.1] The paper refers to the GitHub repository for 'more details' on splits, prompts, and hyperparameters; for reproducibility these details should be included in the paper or a supplementary document, including seeds and the exact few-shot prompts.
  4. [§3.2] The AMT collection description lacks basic quality-control information, such as the number of workers, qualification requirements, payment, and how many collected pairs were discarded before the final 5,648; reporting these numbers would help assess the reliability of the crowdsourced content.

Circularity Check

0 steps flagged · score 0.0 of 10

No circular derivation: SensorQA's QA pairs are human annotations over activity graphs, and the benchmark scores are computed on a held-out test split.

full rationale

SensorQA is a dataset and benchmark, not a fitted model or a theorem derivation. The claimed outputs are the 5,648 QA pairs, created by AMT workers reading activity graphs built from ExtraSensory activity labels, and the benchmark accuracies of several baseline models. No equation in the paper equates a prediction with its input, no parameter is fitted to the test answers, and no result is justified by a load-bearing self-citation. The only self-citation, [39], appears in the sentence 'AMT is extensively utilized in the field of NLP for dataset creation [13, 39]', where it is merely an illustrative example of the crowdsourcing platform, not evidence for any central claim. The benchmark evaluation randomly splits 80% for training and 20% for testing, and the reported 28% best accuracy is a measured outcome on that held-out split, not a quantity forced by construction. The reader's concern that ExtraSensory labels are self-reported and noisy, and that AMT answers were not validated, bears on dataset quality and the trustworthiness of the references, but it is not circularity: the task is defined as matching the human-written references, so noisy references would degrade validity without making the derivation self-referential. The claim 'accurate answers derived from sensor data' is imprecise because answers come from the activity labels rather than raw signals, but this is an accuracy-of-description issue, not a circular derivation. Overall, there is no circular step to exhibit.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The paper's contribution is empirical, so there are no fitted free parameters or invented entities. The benchmark rests on two domain assumptions: the ExtraSensory labels are accurate enough to ground answers, and the AMT workers produce correct questions and answers without validation. A further methodological assumption concerns the GPT-3.5-distilled short answers used for exact-match scoring.

assumptions (3)
  • domain assumption ExtraSensory activity/context labels are accurate enough to serve as ground truth for QA answers.
    Section 3.1 states the dataset is "tagged by 51 activity or context labels after cleaning" and Section 3.2 has workers write ground-truth answers from activity graphs; the paper never validates label accuracy.
  • domain assumption AMT workers reliably interpret the activity graphs and produce correct questions and answers without validation.
    Section 3.2 describes the AMT protocol and six author-provided examples but reports no filtering, inter-annotator agreement, or correctness checks on the 5,648 collected pairs.
  • domain assumption GPT-3.5-based distillation of short answers preserves the key semantics needed for exact-match evaluation.
    Section 5.1 introduces the short-answer version by prompting GPT-3.5 to extract 1-2 keywords; no manual check of distillation quality is reported.

how reviews work

0 comments
Cite this review

Pith. "Pith review of SensorQA: A Question Answering Benchmark for Daily-Life Monitoring." pith.science (2026). https://pith.science/paper/YZCAESWI

@misc{pith2026250104974,
  author       = {Pith},
  title        = {Pith review of: SensorQA: A Question Answering Benchmark for Daily-Life Monitoring},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/YZCAESWI}},
  note         = {Machine review of arXiv:2501.04974}
}
read the original abstract

With the rapid growth in sensor data, effectively interpreting and interfacing with these data in a human-understandable way has become crucial. While existing research primarily focuses on learning classification models, fewer studies have explored how end users can actively extract useful insights from sensor data, often hindered by the lack of a proper dataset. To address this gap, we introduce SensorQA, the first human-created question-answering (QA) dataset for long-term time-series sensor data for daily life monitoring. SensorQA is created by human workers and includes 5.6K diverse and practical queries that reflect genuine human interests, paired with accurate answers derived from sensor data. We further establish benchmarks for state-of-the-art AI models on this dataset and evaluate their performance on typical edge devices. Our results reveal a gap between current models and optimal QA performance and efficiency, highlighting the need for new contributions. The dataset and code are available at: https://github.com/benjamin-reichman/SensorQA.

Figures

Figures reproduced from arXiv: 2501.04974 by the authors.

Figure 1
Figure 1. Visualizations of existing QA datasets using time series IMU sensor data [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Example QA pairs in SensorQA. (a) and (b) are generated from daily graph, while (c) and (d) are generated [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Exact-match accuracy of Llama [43] displayed by question category (left) and answer category (right) [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: Model memory size (left) and average answer gen￾erating latency (right) on Jetson TX2 [2]. Footnote 𝑄 denotes models after quantization. QA performance per category Different Q&A types present varying difficulties for the models [PITH_FULL_IMAGE:figures/full_fig_p006_4.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SensorChat: Answering Qualitative and Quantitative Questions during Long-Term Multimodal Sensor Interactions

    cs.AI 2025-02 conditional novelty 6.0 of 10

    A three-stage pipeline with LLM decomposition, pretrained embedding retrieval, and LLM assembly outperforms prior sensor QA systems on long-duration, high-frequency data, with caveats on evaluation leakage.

Reference graph

Works this paper leans on

64 extracted references · 32 canonical work pages · cited by 1 Pith paper

  1. [1]

    Amazon Mechanical Turk

    2024. Amazon Mechanical Turk. https://www.mturk.com/. [Online]

  2. [2]

    Jetson TX2 Module

    2024. Jetson TX2 Module. https://developer.nvidia.com/embedded/ jetson-tx2. [Online]

  3. [3]

    Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al . 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774 (2023)

  4. [4]

    Satanjeev Banerjee and Alon Lavie. 2005. METEOR: An automatic metric for MT evaluation with improved correlation with human judgments. In Proceedings of the acl workshop on intrinsic and extrinsic evaluation measures for machine translation and/or summarization . 65–72

  5. [5]

    Asma Ben Abacha and Dina Demner-Fushman. 2019. A Question- Entailment Approach to Question Answering. BMC Bioinform. 20, 1 (2019), 511:1–511:23. https://bmcbioinformatics.biomedcentral.com/ articles/10.1186/s12859-019-3119-4

  6. [6]

    Qiming Cao, Hongfei Xue, Tianci Liu, Xingchen Wang, Haoyu Wang, Xincheng Zhang, and Lu Su. 2024. mmCLIP: Boosting mmWave-based Zero-shot HAR via Signal-Text Alignment. In Proceedings of the 22nd ACM Conference on Embedded Networked Sensor Systems . 184–197

  7. [7]

    Ricardo Chavarriaga, Hesam Sagha, Alberto Calatroni, Sundara Te- jaswi Digumarti, Gerhard Tröster, José del R Millán, and Daniel Roggen

  8. [8]

    Zhiyu Chen, Wenhu Chen, Charese Smiley, Sameena Shah, Iana Borova, Dylan Langdon, Reema Moussa, Matt Beane, Ting-Hao Huang, Bryan Routledge, and William Yang Wang. 2021. FinQA: A Dataset of Numerical Reasoning over Financial Data. In Proceed- ings of the 2021 Conference on Empirical Methods in Natural Language Processing, Marie-Francine Moens, Xuanjing Hu...

Show all 64 references
  1. [9]

    Gonzalez, Ion Stoica, and Eric P

    Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. 2023. Vicuna: An Open-Source Chatbot Impressing GPT-4 with 90%* ChatGPT Quality. https://lmsys. org/blog/2023...

  2. [10]

    Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova. 2019. BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Com...

  3. [11]

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova

  4. [12]

    Zachary Englhardt, Chengqian Ma, Margaret E Morris, Chun-Cheng Chang, Xuhai" Orson" Xu, Lianhui Qin, Daniel McDuff, Xin Liu, Shwe- tak Patel, and Vikram Iyer. 2024. From Classification to Clinical In- sights: Towards Analyzing and Reasoning About Mobile and Behav- ioral Health...

  5. [13]

    Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. 2017. Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering. In Conference on Computer Vision and Pattern Recognition (CVPR)

  6. [14]

    Kristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang, Miao Liu, Xingyu Liu, et al. 2022. Ego4d: Around the world in 3,000 hours of egocentric video. In Proceedings of the IEEE/CVF Conference on Computer V...

  7. [15]

    Jiaming Han, Kaixiong Gong, Yiyuan Zhang, Jiaqi Wang, Kaipeng Zhang, Dahua Lin, Yu Qiao, Peng Gao, and Xiangyu Yue. 2023. Onellm: One framework to align all modalities with language. arXiv preprint arXiv:2312.03700 (2023)

  8. [16]

    Jiaming Han, Kaixiong Gong, Yiyuan Zhang, Jiaqi Wang, Kaipeng Zhang, Dahua Lin, Yu Qiao, Peng Gao, and Xiangyu Yue. 2024. Onellm: One framework to align all modalities with language. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

  9. [17]

    Xing, and Pengtao Xie

    Xuehai He, Yichen Zhang, Luntian Mou, Eric P. Xing, and Pengtao Xie. 2020. PathVQA: 30000+ Questions for Medical Visual Question Answering. ArXiv abs/2003.10286 (2020). https://api.semanticscholar. org/CorpusID:214612106

  10. [18]

    Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2021. Lora: Low-rank adaptation of large language models. arXiv preprint arXiv:2106.09685 (2021)

  11. [19]

    Mengkang Hu, Haoyu Dong, Ping Luo, Shi Han, and Dongmei Zhang

  12. [20]

    Mandar Joshi, Eunsol Choi, Daniel Weld, and Luke Zettlemoyer

  13. [21]

    Yubin Kim, Xuhai Xu, Daniel McDuff, Cynthia Breazeal, and Hae Won Park. 2024. Health-llm: Large language models for health prediction via wearable sensor data. arXiv preprint arXiv:2401.06866 (2024)

  14. [22]

    Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov

    Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc Le, and Slav...

  15. [23]

    Chin-Yew Lin. 2004. Rouge: A package for automatic evaluation of summaries. In Text summarization branches out. 74–81

  16. [24]

    Ji Lin, Jiaming Tang, Haotian Tang, Shang Yang, Wei-Ming Chen, Wei- Chen Wang, Guangxuan Xiao, Xingyu Dang, Chuang Gan, and Song Han. 2024. AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration. In MLSys

  17. [25]

    Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. 2024. Im- proved baselines with visual instruction tuning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 26296–26306

  18. [26]

    Junling Liu, Peilin Zhou, Yining Hua, Dading Chong, Zhongyu Tian, Andrew Liu, Helin Wang, Chenyu You, Zhenhua Guo, Lei Zhu, et al

  19. [27]

    Sylvain Lobry, Diego Marcos, Jesse Murray, and Devis Tuia. 2020. RSVQA: Visual question answering for remote sensing data. IEEE Transactions on Geoscience and Remote Sensing 58, 12 (2020), 8555– 8566

  20. [28]

    Pan Lu, Swaroop Mishra, Tony Xia, Liang Qiu, Kai-Wei Chang, Song- Chun Zhu, Oyvind Tafjord, Peter Clark, and Ashwin Kalyan. 2022. Learn to Explain: Multimodal Reasoning via Thought Chains for Sci- ence Question Answering. In The 36th Conference on Neural Informa- tion Processi...

  21. [29]

    Seungwhan Moon, Andrea Madotto, Zhaojiang Lin, Tushar Nagarajan, Matt Smith, Shashank Jain, Chun-Fu Yeh, Prakash Murugesan, Peyman Heidari, Yue Liu, et al. 2023. Anymal: An efficient and scalable any- modality augmented language model. arXiv preprint arXiv:2309.16058 (2023)

  22. [30]

    Seungwhan Moon, Andrea Madotto, Zhaojiang Lin, Aparajita Saraf, Amy Bearman, and Babak Damavandi. 2023. IMU2CLIP: Language- grounded Motion Sensor Translation with Multimodal Contrastive Learning. In Findings of the Association for Computational Linguistics: EMNLP 2023. 13246–13253

  23. [31]

    Advances in Neural Information Processing Systems 36 (2024)

    Benchmarking large language models on cmexam-a comprehen- sive chinese medical exam dataset. Advances in Neural Information Processing Systems 36 (2024)

  24. [32]

    Xiaomin Ouyang, Xian Shuai, Jiayu Zhou, Ivy Wang Shi, Zhiyuan Xie, Guoliang Xing, and Jianwei Huang. 2022. Cosmo: contrastive fusion learning with small data for multimodal human activity recognition. In Proceedings of the 28th Annual International Conference on Mobile Computi...

  25. [33]

    Xiaomin Ouyang, Zhiyuan Xie, Heming Fu, Sitong Cheng, Li Pan, Neiwen Ling, Guoliang Xing, Jiayu Zhou, and Jianwei Huang. 2023. Harmony: Heterogeneous Multi-Modal Federated Learning through Disentangled Model Training. In Proceedings of the 21st Annual In- ternational Conferenc...

  26. [34]

    Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: a method for automatic evaluation of machine translation. InPro- ceedings of the 40th annual meeting of the Association for Computational Linguistics. 311–318

  27. [35]

    Tianwen Qian, Jingjing Chen, Linhai Zhuo, Yang Jiao, and Yu-Gang Jiang. 2024. Nuscenes-qa: A multi-modal visual question answering benchmark for autonomous driving scenario. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 38. 4542–4550

  28. [36]

    Jingping Nie, Hanya Shao, Minghui Zhao, Stephen Xia, Matthias Preindl, and Xiaofan Jiang. 2022. Conversational ai therapist for daily function screening in home environments. In Proceedings of the 1st ACM International Workshop on Intelligent Acoustic Systems and Applications. 31–36

  29. [37]

    Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer. Journal of Machine Learning Research 21, 140 (2020), 1–67. htt...

  30. [38]

    Pranav Rajpurkar, Robin Jia, and Percy Liang. 2018. Know What You Don’t Know: Unanswerable Questions for SQuAD. InProceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers) , Iryna Gurevych and Yusuke Miyao (Eds.). Associat...

  31. [39]

    Benjamin Z. Reichman, Anirudh Sundar, Christopher Richardson, Tamara Zubatiy, Prithwijit Chowdhury, Aaryan Shah, Jack Truxal, Micah Grimes, Dristi Shah, Woo Ju Chee, Saif Punjwani, Atishay Jain, and Larry Heck. 2023. Outside Knowledge Visual Question Answering Version 2.0. In ...

  32. [40]

    Argho Sarkar, Tashnim Chowdhury, Robin Murphy, Aryya Gan- gopadhyay, and Maryam Rahnemoonfar. 2023. Sam-vqa: Supervised attention-based visual question answering model for post-disaster damage assessment on remote sensing imagery. IEEE Transactions on Geoscience and Remote Sen...

  33. [41]

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al . 2021. Learning transferable visual mod- els from natural language supervision. In International conference on machine lea...

  34. [42]

    Mélisande Teng, Amna Elmustafa, Benjamin Akera, Yoshua Bengio, Hager Radi, Hugo Larochelle, and David Rolnick. 2023. SatBird: a Dataset for Bird Species Distribution Modeling using Remote Sens- ing and Citizen Science Data. In Thirty-seventh Conference on Neu- ral Information ...

  35. [43]

    Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie- Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971 (2023)

  36. [44]

    Yonatan Vaizman, Katherine Ellis, and Gert Lanckriet. 2017. Recog- nizing detailed human context in the wild from smartphones and smartwatches. IEEE pervasive computing 16, 4 (2017), 62–74

  37. [45]

    Yonatan Vaizman, Nadir Weibel, and Gert Lanckriet. 2018. Context recognition in-the-wild: Unified model for multi-modal sensors and multi-label classification. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies 1, 4 (2018), 1–22

  38. [46]

    Junjue Wang, Zhuo Zheng, Zihang Chen, Ailong Ma, and Yanfei Zhong

  39. [47]

    Satyajit Sinha. 2023. State of IoT 2024: Number of connected IoT devices growing 13% to 18.8 billion globally. https://iot-analytics.com/ number-connected-iot-devices/. [Online]. SensorQA: A Question Answering Benchmark for Daily-Life Monitoring

  40. [48]

    Tianwei Xing, Luis Garcia, Federico Cerutti, Lance Kaplan, Alun Preece, and Mani Srivastava. 2021. DeepSQA: Understanding Sen- sor Data via Question Answering. In Proceedings of the International Conference on Internet-of-Things Design and Implementation . 106–118

  41. [49]

    Huatao Xu, Pengfei Zhou, Rui Tan, and Mo Li. 2023. Practically Adopting Human Activity Recognition. In Proceedings of the 29th Annual International Conference on Mobile Computing and Networking . 1–15

  42. [50]

    Lilin Xu, Chaojie Gu, Rui Tan, Shibo He, and Chen Jiming. 2023. MESEN: Exploit Multimodal Data to Design Unimodal Human Ac- tivity Recognition with Few Labels. In Proceedings of the 21st ACM Conference on Embedded Networked Sensor Systems

  43. [51]

    Bufang Yang, Siyang Jiang, Lilin Xu, Kaiwei Liu, Hai Li, Guoliang Xing, Hongkai Chen, Xiaofan Jiang, and Zhenyu Yan. 2024. DrHouse: An LLM-empowered Diagnostic Reasoning System through Harnessing Outcomes from Sensor Data and Expert Knowledge. arXiv preprint arXiv:2405.12541 (2024)

  44. [52]

    Jianfei Yang, He Huang, Yunjiao Zhou, Xinyan Chen, Yuecong Xu, Shenghai Yuan, Han Zou, Chris Xiaoxuan Lu, and Lihua Xie. 2024. Mm- fi: Multi-modal non-intrusive 4d human dataset for versatile wireless sensing. Advances in Neural Information Processing Systems 36 (2024)

  45. [53]

    In Proceedings of the AAAI Conference on Artificial Intelligence , Vol

    EarthVQA: Towards Queryable Earth via Relational Reasoning- Based Remote Sensing Visual Question Answering. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 38. 5481–5489

  46. [54]

    Yuxuan Weng, Guoquan Wu, Tianyue Zheng, Yanbing Yang, and Jun Luo. 2024. Large Model for Small Data: Foundation Model for Cross- Modal RF Human Activity Recognition. In Proceedings of the 22nd ACM Conference on Embedded Networked Sensor Systems . 436–449

  47. [55]

    Renrui Zhang, Jiaming Han, Chris Liu, Peng Gao, Aojun Zhou, Xi- angfei Hu, Shilin Yan, Pan Lu, Hongsheng Li, and Yu Qiao. 2023. Llama-adapter: Efficient fine-tuning of language models with zero-init attention. arXiv preprint arXiv:2303.16199 (2023)

  48. [56]

    Haoxi Zhong, Chaojun Xiao, Cunchao Tu, Tianyang Zhang, Zhiyuan Liu, and Maosong Sun. 2019. JEC-QA: A Legal-Domain Question Answering Dataset. arXiv:1911.12011 [cs.CL]

  49. [57]

    Yunjiao Zhou, Jianfei Yang, Han Zou, and Lihua Xie. 2023. TENT: Connect Language Models with IoT Sensors for Zero-Shot Activity Recognition. arXiv preprint arXiv:2311.08245 (2023)

  50. [60]

    Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William Cohen, Ruslan Salakhutdinov, and Christopher D. Manning. 2018. HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language...

  51. [61]

    Junjie Ye, Xuanting Chen, Nuo Xu, Can Zu, Zekai Shao, Shichun Liu, Yuhan Cui, Zeyang Zhou, Chao Gong, Yang Shen, et al . 2023. A comprehensive capability analysis of gpt-3 and gpt-3.5 series models. arXiv preprint arXiv:2303.10420 (2023)

  52. [789]

    https://doi.org/10.18653/v1/P18-2124

  53. [2013]

    Pattern Recognition Letters 34, 15 (2013), 2033–2042

    The Opportunity challenge: A benchmark database for on-body sensor-based activity recognition. Pattern Recognition Letters 34, 15 (2013), 2033–2042

  54. [2017]

    In Proceedings of the 55th An- nual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , Regina Barzilay and Min-Yen Kan (Eds.)

    TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension. In Proceedings of the 55th An- nual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , Regina Barzilay and Min-Yen Kan (Eds.). Associa- tion for Computatio...

  55. [2018]

    arXiv preprint arXiv:1810.04805 (2018)

    Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805 (2018)

  56. [2024]

    arXiv preprint arXiv:2405.08099 (2024)

    KET-QA: A Dataset for Knowledge Enhanced Table Question Answering. arXiv preprint arXiv:2405.08099 (2024)

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.