REVIEW 5 major objections 7 minor 2 cited by
Foundation Models for CPS-IoT: Opportunities and Challenges
T0 review · 5 major / 7 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read Current foundation models are not yet viable for CPS-IoT applications, and the gap will close only through domain-specific architectures and community-built sensor data resources.
desk verdict A useful roadmap with four suggestive but confounded experiments; the agenda is plausible, the 'must' outruns the evidence. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The argument is carried by the Perception-Cognition-Communication-Action (PCCA) loop as the organizing frame, together with four distinguishing characteristics of CPS-IoT systems that the authors treat as foundational: tight resource/quality trade-offs, spatial embodiment, historical context, and structural constraints. These four characteristics do the analytical work: each is paired with a preliminary experiment that exposes a failure mode of current FMs, and each translates directly into a desideratum (resource-feasible inference, world-state representations, long-context state compression, and neurosymbolic knowledge injection). The inverse problem — recovering the underlying physical state from distributed sensor projections — and the notion of complex events as temporally extended patterns are the concrete mechanisms the paper proposes for turning sensor data into FM training objectives.
What would settle it
Scale a generic multimodal foundation model on a large corpus of raw sensor streams and test it on the paper's own benchmarks (ECG/PPG extrapolation-imputation, multi-vantage vehicle tracking, and 5-minute complex event detection); if it matches or beats the domain-specific models at deployable resource cost, the structural-gap claim is falsified.
Extended reading notes
Core claim
On the paper's own terms, the central discovery is that the gap between current foundation models (FMs) and large language models (LLMs) and the requirements of CPS-IoT applications is real, structural, and measurable. In head-to-head tests on mobile electrocardiogram and photoplethysmogram signals, the largest time-series foundation model tested (MOMENT-L) beat classical ARIMA baselines only at the price of roughly 1.2 gigabytes of memory and a 2.3-second cold-start latency, while a 1-billion-parameter LLM failed to beat ARIMA at all. A multi-vantage masked autoencoder with relative geolocation embeddings recovered physical trajectories better than contrastive baselines, but only because it was designed to learn world state rather than sensor projections. LLMs scored poorly on complex event detection even when given perfect atomic-activity labels, while state-based and neurosymbolic models held up better; and an LLM-based knowledge graph question-answering pipeline struggled to map vague CPS concepts like 'energy' onto building schema. From these results the authors conclude that CPS-IoT FMs must generalize across sensor configurations and tasks, tokenize continuous sensor streams without information loss, handle long and continuous streams, incorporate human knowledge through neurosymbolic layers, expose a rich language channel, and become composable system services rather than standalone per-application models.
Load-bearing premise
The whole agenda rests on the premise that the observed failures come from missing CPS-IoT-specific design rather than from current models simply being too small or not pretrained on enough sensor data.
Editorial extensions
If this is right
- General LLMs will not serve as drop-in processors of raw sensor time series; their reliable role is limited to metadata, high-level reasoning, and language-based analytics over sensor data.
- CPS-IoT FMs should be pretrained to reconstruct masked signals across vantage points and time, with relative geolocation embeddings, so that their internal representations track physical world state rather than sensor-specific projections.
- Long-horizon tasks such as complex event detection require state-compressing architectures (state-space models, neurosymbolic finite state machines) rather than attention over unbounded context windows.
- For edge devices, foundation models must become runtime composable services shared across applications, and the community should pursue application-specific micro foundation models rather than a single universal CPS-IoT FM.
- A shared ecosystem of raw sensor datasets, benchmarks, and ontologies is a prerequisite for progress; without it, the desiderata cannot be validated or iterated upon.
Reading between the lines
- The paper's framing implies that next-sample prediction benchmarks systematically flatter general time-series models, because they reward matching sensor projections rather than recovering world state; a benchmark that scores tracking or inverse-problem accuracy would likely widen the measured gap.
- If historical context is the bottleneck, a fixed-size state compressor (e.g., a small state-space encoder) could be tested as a drop-in front end to any transformer FM, converting unbounded sensor history into a bounded state vector before attention; such an experiment would isolate the context-length effect.
- The resource-efficiency results suggest a concrete engineering target: a CPS-IoT FM that beats ARIMA on accuracy and matches its memory and latency would be the practical threshold for mobile health deployment, and today no published model meets both.
- The KGQA failure pattern hints that ontology standardization (e.g., consistent naming of 'energy' concepts across buildings) is as important as model capability; a testable extension is measuring query success rate against schema variants, which would tell the community where to invest.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper argues that foundation models (FMs) and large language models (LLMs), despite their success in language and vision, are not yet viable for CPS-IoT systems because the domain has four distinguishing characteristics: tight resource/quality trade-offs, spatial embodiment, historical context, and structural constraints. After a survey of perception-focused CPS-IoT FMs organized by sensor modality (unimodal, multimodal, flexible) and task flexibility (fixed, configurable, selectable, run-time specifiable), the authors present four preliminary experiments probing each characteristic: an edge-resource comparison of MOMENT and Llama-3.2-1B against ARIMA on ECG/PPG forecasting and imputation, a multi-vantage masked autoencoder for acoustic/seismic vehicle tracking, LLM and neurosymbolic complex-event detection on a simulated smart-health dataset, and a KGQA study on a Brick-schema building knowledge graph. On this basis, Section 4 lists seven desiderata for CPS-IoT FMs (e.g., sensor-configuration generalizability, sensor tokenization, long-stream models, neurosymbolic structure, language channel, system-service abstractions) and calls for community datasets, benchmarks, and application-specific 'micro foundation models.' The central thesis is that domain-specific FM designs, not direct adaptation of general FMs, are needed to close the gap.
Significance. If the central thesis is correct, the paper provides a useful roadmap: the survey taxonomy is informative, the four characteristics are a plausible organizing framework, and the desiderata in Section 4 are concrete enough to guide research. The call for shared raw-sensor datasets and benchmarks is timely, and the proposal of application-specific micro foundation models is a reasonable middle path between fully general and task-specific models. The authors are candid in labeling their experiments as preliminary and in acknowledging that the spatial dataset is self-collected, which helps calibrate the strength of the claims. However, the empirical support for the strong 'must account for unique characteristics' conclusion is currently confounded: scale, tokenization, and sensor-data pretraining are not controlled, and the released evidence is narrow (one physiology-signal resource study, one six-node tracking study without error bars, one simulated CED study, and one building KGQA study). If the paper is reframed as a research agenda with illustrative case studies, its significance is solid; as an empirical demonstration of domain-specific FM necessity, it is not yet established.
major comments (5)
- [§3.1.1–3.1.2, Table 1] The resource-tradeoff experiment conflates architecture with scale and adaptation. Llama-3.2-1B is a 1B general-purpose LLM given raw 50 Hz digitized waveforms with no sensor-specific tokenizer or sensor pretraining, while MOMENT-L is a large time-series FM; neither comparison controls for parameter count, context length, or training data. The observation that Llama performs worse than ARIMA and MOMENT is consistent with the paper's resource-feasibility argument, but it does not show that a general FM with sufficient scale and sensor-aware tokenization would fail. The conclusion in §3.1.2 that this 'highlights the limitations of today's general LLMs in processing low-level sensory data directly' should be softened to a statement about the specific models tested, or supplemented with a matched comparison.
- [§3.3.2, Table 2 and Figure 5] The CED experiment's conclusion is stronger than the evidence. LLMs are given perfect atomic labels and are evaluated on a metric (conditional F1) that may be dominated by output-format brittleness; the low scores do not isolate a failure of temporal reasoning from a failure to produce the exact output format. The comparison to Mamba and AE+FSM also conflates model class with hand-coded rules, since AE+FSM receives the true complex-event rules and the other models must learn them. The manuscript should either add controls (e.g., constrained decoding, structured output parsing, or a learned model with similar state-compression capacity) or explicitly limit the claim to 'the specific prompting protocol we used.'
- [§3.2.3, Figure 3] The spatial-embedding result, as reported, does not support the strong claim. The evaluation uses a self-collected six-node dataset, reports no error bars or significance tests, and the data and code are not released. The statement that the method 'consistently achieves the highest accuracy' cannot be verified from the figure. For a paper whose central claim is that spatial embeddings are a required domain-specific component, this single proof-of-concept should be presented as an illustrative case study, with the corresponding claims in the text adjusted.
- [§3.4.2–3.4.3, Table 3] The KGQA study's own analysis undermines the architectural conclusion. The failures in Table 3 are attributed in the text to subgraph extraction (FAISS relevance), entity/relation linking, and one-to-many mappings, not to the LLM backbone; indeed, §3.4.3 states that the 'main challenge lies in supplying LLMs with the appropriate relevant subgraph.' As written, the study supports the need for better CPS-IoT KG retrieval and schema design, not the paper's conclusion that CPS-IoT FMs must incorporate structural constraints through neurosymbolic architecture. The paper should separate these two claims and not use this experiment as direct evidence for the latter.
- [Abstract and §5] The conclusion that CPS-IoT FMs 'must account for the unique characteristics' is presented as an established finding, while the evidence in Section 3 is explicitly preliminary and, as noted above, does not rule out the alternative that a sufficiently scaled general FM with sensor-aware tokenization and longer context would close the same gap. The manuscript should either provide such a control or clearly label the domain-specific requirements as hypotheses to be tested. This distinction is load-bearing because the paper's proposed research program (domain-specific architectures, muFMs, and new benchmarks) depends on it.
minor comments (7)
- [Figure 1 caption and §3.1.1] The figure caption contains the typo 'for ecasting' and the prompt text in the figure contains 'ret ur n' and 'for ecast'; these should be corrected.
- [§3.3.2–3.3.3] There is a duplicated 'the' in 'determines if the the complex event labels' and 'few-short' should be 'few-shot' in the discussion.
- [§4.1.2] 'embeddedings' should be 'embeddings'.
- [Header and references] The manuscript contains placeholder metadata ('Make sure to enter the correct conference title', 'XXXXXXX', 'Woodstock, NY', fake DOI) and reference [100] includes '<today>'; these must be completed before submission.
- [Table 1 and Table 3] The header of Table 1 is garbled (e.g., '𝑬 𝑪 𝑮𝑰 𝒎𝒑 𝑬 𝑪 𝑮𝑬𝒙𝒕 𝑷 𝑷 𝑮𝑰 𝒎𝒑 𝑷 𝑷 𝑮𝑬𝒙𝒕') and Table 3's ✓/× symbols are not defined in the caption; both tables need cleanup.
- [Table 2] The metrics 'Length Acc.', 'Coarse F1', and 'Conditional F1' are defined only in the text; the table caption should include brief definitions.
- [References] Reference [81] lists 'Ozan Baris Mulayim' while the paper's first author is 'Ozan Baris'; please check the name consistency.
Circularity Check
No significant circularity: the paper is a survey-and-experiment research agenda whose central claim is grounded in independent third-party model evaluations, not in a derivation that reduces to its inputs.
full rationale
The paper makes no formal derivation chain that could collapse into its own inputs. Its central assertion—that a gap persists between current FM/LLM capabilities and CPS-IoT requirements, and that domain-specific design is needed—is supported by (i) a survey of external and prior work and (ii) preliminary experiments. The experiments evaluate third-party or externally published models: MOMENT and Llama-3.2 on ECG/PPG data, Qwen2.5/GPT-4o on synthetic complex-event traces built from WISDM and ESC50, and AutoKGQA on the Mortar/Brick building KG. These are out-of-sample evaluations, not fits renamed as predictions. The authors do cite their own prior models (FOCAL, FreqMAE, PhyMask, LLMSense, IoT-LM), but those citations serve as examples of prior art or as comparison baselines; the conclusion does not depend on accepting any self-cited result as an unverified premise. The spatial-MAE 'Ours' result is an empirical experiment whose architecture encodes the paper's proposed spatial-embedding idea; that supports the agenda but is not circular because the comparison is to external baselines (SimCLR, CMC, Cocoa, GMC, FOCAL) on a collected dataset. The skeptical concern that Section 3 comparisons are confounded by model scale, tokenization, and adaptation is a methodological validity objection, not a circularity objection: the confound is not a case where the 'prediction' is equivalent to the input by construction. There is no self-definitional equation, no fitted-input-called-prediction step, no imported uniqueness theorem, and no renaming of a known result presented as derivation. Under the hard rule that circularity must be exhibited by quotation and specific reduction, none is present.
Assumptions & free parameters
assumptions (4)
- domain assumption Sensor data are discretized samples of continuous spatiotemporal physical signals, not ordered symbolic sequences like text.
- ad hoc to paper The four characteristics listed in Section 3 (resource/quality trade-offs, spatial embodiment, historical context, structural constraints) are the salient axes that make CPS-IoT FMs fundamentally different.
- ad hoc to paper A single application-agnostic CPS-IoT FM is impractical, so application-specific micro foundation models (muFMs) are a viable path.
- domain assumption Larger context windows alone will not solve long-horizon sensor reasoning; models need mechanisms to remember selected important events.
Cite this review
Pith. "Pith review of Foundation Models for CPS-IoT: Opportunities and Challenges." pith.science (2026). https://pith.science/paper/QB2GBCU5
@misc{pith2026250116368,
author = {Pith},
title = {Pith review of: Foundation Models for CPS-IoT: Opportunities and Challenges},
year = {2026},
howpublished = {\url{https://pith.science/paper/QB2GBCU5}},
note = {Machine review of arXiv:2501.16368}
}
read the original abstract
Methods from machine learning (ML) have transformed the implementation of Perception-Cognition-Communication-Action loops in Cyber-Physical Systems (CPS) and the Internet of Things (IoT), replacing mechanistic and basic statistical models with those derived from data. However, the first generation of ML approaches, which depend on supervised learning with annotated data to create task-specific models, faces significant limitations in scaling to the diverse sensor modalities, deployment configurations, application tasks, and operating dynamics characterizing real-world CPS-IoT systems. The success of task-agnostic foundation models (FMs), including multimodal large language models (LLMs), in addressing similar challenges across natural language, computer vision, and human speech has generated considerable enthusiasm for and exploration of FMs and LLMs as flexible building blocks in CPS-IoT analytics pipelines, promising to reduce the need for costly task-specific engineering. Nonetheless, a significant gap persists between the current capabilities of FMs and LLMs in the CPS-IoT domain and the requirements they must meet to be viable for CPS-IoT applications. In this paper, we analyze and characterize this gap through a thorough examination of the state of the art and our research, which extends beyond it in various dimensions. Based on the results of our analysis and research, we identify essential desiderata that CPS-IoT domain-specific FMs and LLMs must satisfy to bridge this gap. We also propose actions by CPS-IoT researchers to collaborate in developing key community resources necessary for establishing FMs and LLMs as foundational tools for the next generation of CPS-IoT systems.
Figures
Figures from the paper (3 more)
Forward citations
Cited by 2 Pith papers
-
RARR : Robust Real-World Activity Recognition with Vibration by Scavenging Near-Surface Audio Online
Pretraining a multitask VAE on online ASMR audio and finetuning only the activity head on vibration data improves cross-user activity recognition accuracy over baselines in a 4-participant study.
-
Zero-Trust Foundation Models: A New Paradigm for Secure and Collaborative Artificial Intelligence for Internet of Things
The paper defines the ZTFM concept, identifies four zero-trust principles, reviews enabling technologies and threats, and lays out open research challenges for AI-driven IoT security.
Reference graph
Works this paper leans on
-
[1]
Karan Ahuja, Paul Streli, and Christian Holz. 2021. TouchPose: Hand Pose Prediction, Depth Estimation, and Touch Classification from Capacitive Images. In Proceedings of the 34th Annual ACM Symposium on User Interface Software and Technology. Association for Computing Machinery, New York, NY, USA, 13 pages. https://doi.org/10.1145/ 3472749.3474801
arXiv 2021
-
[2]
Samaneh Aminikhanghahi and Diane J Cook. 2017. A survey of methods for time series change point detection. Knowledge and information systems 51, 2 (2017), 339–367
2017
-
[3]
Tuo An, Yunjiao Zhou, Han Zou, and Jianfei Yang. 2024. IoT-LLM: Enhancing Real-World IoT Task Reasoning with Large Language Models. arXiv preprint arXiv:2410.02429 (2024)
arXiv 2024
-
[4]
Abdul Fatir Ansari, Lorenzo Stella, Caner Turkmen, Xiyuan Zhang, Pedro Mercado, Huibin Shen, Oleksandr Shchur, Syama Sundar Ran- gapuram, Sebastian Pineda Arango, Shubham Kapoor, et al . 2024. Chronos: Learning the language of time series. Transactions on Ma- chine Learning Research (2024)
2024
-
[5]
Riku Arakawa, Mayank Goel, Chris Harrison, and Karan Ahuja. 2022. RGBDGaze: Gaze Tracking on Smartphones with RGB and Depth Data. In International Conference on Multimodal Interaction, ICMI 2022, Bengaluru, India, November 7-11, 2022 . ACM, New York, 329–
2022
-
[6]
Caio Viktor S. Avila, Vânia M.P. Vidal, Wellington Franco, and Marco A. Casanova. 2024. Experiments with text-to-SPARQL based on ChatGPT. In 2024 IEEE 18th International Conference on Semantic Computing (ICSC). IEEE, Laguna Hills, CA, USA, 277–284. https: //doi.org/10.1109/ICSC59802.2024.00050
arXiv 2024
-
[7]
Bharathan Balaji, Arka Bhattacharya, Gabriel Fierro, Jingkun Gao, Joshua Gluck, Dezhi Hong, Aslak Johansen, Jason Koh, Joern Ploen- nigs, Yuvraj Agarwal, Mario Berges, David Culler, Rajesh Gupta, Mikkel Baun Kjærgaard, Mani Srivastava, and Kamin Whitehouse
-
[8]
Iz Beltagy, Matthew E Peters, and Arman Cohan. 2020. Longformer: The long-document transformer. arXiv preprint arXiv:2004.05150 (2020)
arXiv 2020
Show all 147 references
-
[9]
Assaf Ben-Kish, Itamar Zimerman, Shady Abu-Hussein, Nadav Cohen, Amir Globerson, Lior Wolf, and Raja Giryes. 2024. DeciMamba: Ex- ploring the Length Extrapolation Potential of Mamba. arXiv preprint arXiv:2406.14528 (2024)
2024 arXiv
-
[10]
Imane Lahmam Bennani, Anand Krishnan Prakash, Marina Zafiris, Lazlo Paul, Carlos Duarte Roa, Paul Raftery, Marco Pritoni, and Gabe Fierro. 2021. Query relaxation for portable brick-based applications. In Proceedings of the 8th ACM International Conference on Systems for Energy...
2021
-
[11]
Sathya Kamesh Bhethanabhotla, Omar Swelam, Julien Siems, David Salinas, and Frank Hutter. 2024. Mamba4Cast: Efficient Zero-Shot Time Series Forecasting with State Space Models. arXiv preprint arXiv:2410.09385 (2024)
2024 arXiv
-
[12]
Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al. 2021. On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258 (2021)
2021 arXiv
-
[13]
Anthony Brohan, Noah Brown, Justice Carbajal, Yevgen Chebotar, Xi Chen, Krzysztof Choromanski, Tianli Ding, Danny Driess, Avinava Dubey, Chelsea Finn, et al . 2023. RT-2: Vision-language-action models transfer web knowledge to robotic control. arXiv preprint arXiv:2307.15818 (2023)
2023 arXiv
-
[14]
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffr...
2020
-
[15]
Aydar Bulatov, Yury Kuratov, and Mikhail Burtsev. 2022. Recurrent memory transformer. Advances in Neural Information Processing Systems 35 (2022), 11079–11091
2022
-
[16]
Sameep Chattopadhyay, Pulkit Paliwal, Sai Shankar Narasimhan, Shubhankar Agarwal, and Sandeep P Chinchali. 2024. Context Mat- ters: Leveraging Contextual Features for Time Series Forecasting. arXiv preprint arXiv:2410.12672 (2024)
2024 arXiv
-
[17]
Huajun Chen. 2023. Large knowledge model: Perspectives and chal- lenges. arXiv preprint arXiv:2312.02706 (2023)
2023 arXiv
-
[18]
Si-An Chen, Chun-Liang Li, Nate Yoder, Sercan O Arik, and Tomas Pfister. 2023. Tsmixer: An all-mlp architecture for time series fore- casting. arXiv preprint arXiv:2303.06053 (2023). XXX ’25, June 03–05, 2018, Woodstock, NY Baris et al
2023 arXiv
-
[19]
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. 2020. A simple framework for contrastive learning of visual representations. In International Conference on Machine Learning (ICML)
2020
-
[20]
Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever. 2019. Generating long sequences with sparse transformers. arXiv preprint arXiv:1904.10509 (2019)
2019 arXiv
-
[21]
Daniel Cunnington, Mark Law, Jorge Lobo, and Alessandra Russo
-
[22]
Shenghong Dai, Shiqi Jiang, Yifan Yang, Ting Cao, Mo Li, Suman Banerjee, and Lili Qiu. 2024. Advancing Multi-Modal Sens- ing Through Expandable Modality Alignment. arXiv preprint arXiv:2407.17777 (2024)
2024 arXiv
-
[23]
Tri Dao and Albert Gu. 2024. Transformers are SSMs: Generalized models and efficient algorithms through structured state space duality. arXiv preprint arXiv:2405.21060 (2024)
2024 arXiv
-
[24]
Abhimanyu Das, Weihao Kong, Rajat Sen, and Yichen Zhou. 2023. A decoder-only foundation model for time-series forecasting. arXiv preprint arXiv:2310.10688 (2023)
2023 arXiv
-
[25]
Robert Day. 2019. Accelerating Autonomous Vehicle Technology What challenges do we need to consider for the safe deployment of autonomous vehicles at scale? IEEE Spectrum (2019)
2019
-
[26]
Shohreh Deldari, Hao Xue, Aaqib Saeed, Daniel V Smith, and Flora D Salim. 2022. Cocoa: Cross modality contrastive learning for sensor data. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies 6, 3 (2022), 1–28
2022
-
[27]
Jiaxiang Dong, Haixu Wu, Haoran Zhang, Li Zhang, Jianmin Wang, and Mingsheng Long. 2023. Simmtm: A simple pre-training frame- work for masked time-series modeling. Advances in Neural Informa- tion Processing Systems 36 (2023)
2023
-
[28]
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weis- senborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al . 2020. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. In Internat...
2020
-
[29]
Danny Driess, Fei Xia, Mehdi SM Sajjadi, Corey Lynch, Aakanksha Chowdhery, Brian Ichter, Ayzaan Wahid, Jonathan Tompson, Quan Vuong, Tianhe Yu, et al. 2023. PaLM-E: An embodied multimodal language model. arXiv preprint arXiv:2303.03378 (2023)
2023 arXiv
-
[30]
Vijay Ekambaram, Arindam Jati, Nam Nguyen, Phanwadee Sinthong, and Jayant Kalagnanam. 2023. TSMixer: Lightweight MLP-Mixer Model for Multivariate Time Series Forecasting. In Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD ’23). 459–469...
2023
-
[31]
Gabe Fierro, Marco Pritoni, Moustafa Abdelbaky, Daniel Lengyel, John Leyden, Anand Prakash, Pranav Gupta, Paul Raftery, Therese Peffer, Greg Thomson, and David E. Culler. 2019. Mortar: An Open Testbed for Portable Building Analytics. ACM Transactions on Sensor Networks 16, 1 (...
2019 doi
-
[32]
Roya Firoozi, Johnathan Tucker, Stephen Tian, Anirudha Majumdar, Jiankai Sun, Weiyu Liu, Yuke Zhu, Shuran Song, Ashish Kapoor, Karol Hausman, et al. 2023. Foundation models in robotics: Applications, challenges, and the future. The International Journal of Robotics Research (2...
2023
-
[33]
Zhe Gan, Linjie Li, Chunyuan Li, Lijuan Wang, Zicheng Liu, Jianfeng Gao, et al. 2022. Vision-language pre-training: Basics, recent advances, and future trends. Foundations and Trends ® in Computer Graphics and Vision 14, 3–4 (2022), 163–352
2022
-
[34]
Shanghua Gao, Teddy Koker, Owen Queen, Thomas Hartvigsen, Theodoros Tsiligkaridis, and Marinka Zitnik. 2024. UniTS: A unified multi-task time series model. In The Thirty-eighth Annual Conference on Neural Information Processing Systems
2024
-
[35]
Artur d’Avila Garcez, Marco Gori, Luis C Lamb, Luciano Serafini, Michael Spranger, and Son N Tran. 2019. Neural-symbolic comput- ing: An effective methodology for principled integration of machine learning and reasoning. arXiv preprint arXiv:1905.06088 (2019)
2019 arXiv
-
[36]
Azul Garza and Max Mergenthaler-Canseco. 2023. TimeGPT-1. arXiv preprint arXiv:2310.03589 (2023)
2023 arXiv
-
[37]
A Geiger, P Lenz, C Stiller, and R Urtasun. 2013. Vision Meets Robotics: The KITTI Dataset. Int. J. Rob. Res. 32, 11 (sep 2013), 1231–1237. https://doi.org/10.1177/0278364913491297
2013 doi
-
[38]
Rohit Girdhar, Alaaeldin El-Nouby, Zhuang Liu, Mannat Singh, Kalyan Vasudev Alwala, Armand Joulin, and Ishan Misra. 2023. Im- agebind: One embedding space to bind them all. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 15180–15190
2023
-
[39]
Mononito Goswami, Konrad Szafer, Arjun Choudhry, Yifu Cai, Shuo Li, and Artur Dubrawski. 2024. Moment: A family of open time-series foundation models. arXiv preprint arXiv:2402.03885 (2024)
2024 arXiv
-
[40]
Nate Gruver, Marc Finzi, Shikai Qiu, and Andrew G Wilson. 2024. Large language models are zero-shot time series forecasters.Advances in Neural Information Processing Systems 36 (2024)
2024
-
[41]
Albert Gu and Tri Dao. 2023. Mamba: Linear-time sequence modeling with selective state spaces. arXiv preprint arXiv:2312.00752 (2023)
2023 arXiv
-
[42]
Yu Gu, Boyuan Zheng, Boyu Gou, Kai Zhang, Cheng Chang, Sanjari Srivastava, Yanan Xie, Peng Qi, Huan Sun, and Yu Su. 2024. Is your LLM secretly a world model of the internet? model-based planning for web agents. arXiv preprint arXiv:2411.06559 (2024)
2024 arXiv
-
[43]
Atharva Gundawar, Karthik Valmeekam, Mudit Verma, and Subbarao Kambhampati. 2024. Robust Planning with Compound LLM Archi- tectures: An LLM-Modulo Approach. arXiv preprint arXiv:2411.14484 (2024)
2024 arXiv
-
[44]
Tuomas Haarnoja, Sehoon Ha, Aurick Zhou, Jie Tan, George Tucker, and Sergey Levine. 2018. Learning to walk via deep reinforcement learning. arXiv preprint arXiv:1812.11103 (2018)
2018 arXiv
-
[45]
Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick. 2022. Masked autoencoders are scalable vision learners. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
2022
-
[46]
Aritra Hota, Soumyajit Chatterjee, and Sandip Chakraborty. 2024. Evaluating Large Language Models as Virtual Annotators for Time- series Physical Sensing Data. arXiv preprint arXiv:2403.01133 (2024)
2024 arXiv
-
[47]
Yafei Hu, Quanting Xie, Vidhi Jain, Jonathan Francis, Jay Patrikar, Nikhil Keetha, Seungchan Kim, Yaqi Xie, Tianyi Zhang, Hao-Shu Fang, et al . 2023. Toward general-purpose robots via foundation models: A survey and meta-analysis. arXiv preprint arXiv:2312.08782 (2023)
2023 arXiv
-
[48]
Po-Yao Huang, Hu Xu, Juncheng Li, Alexei Baevski, Michael Auli, Wojciech Galuba, Florian Metze, and Christoph Feichtenhofer. 2022. Masked autoencoders that listen. Advances in Neural Information Processing Systems 35 (2022), 28708–28720
2022
-
[49]
Xiao Huang, Jingyuan Zhang, Dingcheng Li, and Ping Li. 2019. Knowl- edge graph embedding based question answering. In Proceedings of the twelfth ACM international conference on web search and data min- ing. 105–113
2019
-
[50]
Sheikh Asif Imran, Mohammad Nur Hossain Khan, Subrata Biswas, and Bashima Islam. 2024. LLaSA: Large Multimodal Agent for Hu- man Activity Analysis Through Wearable Sensors. arXiv preprint arXiv:2406.14498 (2024)
2024
-
[51]
Andrew Jaegle, Felix Gimeno, Andy Brock, Oriol Vinyals, Andrew Zis- serman, and Joao Carreira. 2021. Perceiver: General perception with iterative attention. In International conference on machine learning . Foundation Models for CPS-IoT: Opportunities and Challenges XXX ’25, J...
2021
-
[52]
Ming Jin, Shiyu Wang, Lintao Ma, Zhixuan Chu, James Y Zhang, Xiaoming Shi, Pin-Yu Chen, Yuxuan Liang, Yuan-Fang Li, Shirui Pan, et al. 2023. Time-LLM: Time series forecasting by reprogramming large language models. arXiv preprint arXiv:2310.01728 (2023)
2023 arXiv
-
[53]
Tseng, Yu Zheng, Lei Chen, and Hui Xiong
Ming Jin, Qingsong Wen, Yuxuan Liang, Chaoli Zhang, Siqiao Xue, Xue Wang, James Zhang, Yi Wang, Haifeng Chen, Xiaoli Li, Shirui Pan, Vincent S. Tseng, Yu Zheng, Lei Chen, and Hui Xiong. 2023. Large Models for Time Series and Spatio-Temporal Data: A Survey and Outlook. http://a...
2023 arXiv
-
[54]
Denizhan Kara, Tomoyoshi Kimura, Yatong Chen, Jinyang Li, Rui- jie Wang, Yizhuo Chen, Tianshi Wang, Shengzhong Liu, and Tarek Abdelzaher. 2024. PhyMask: An Adaptive Masking Paradigm for Efficient Self-Supervised Learning in IoT. In Proceedings of the 22nd ACM Conference on Emb...
2024
-
[55]
Denizhan Kara, Tomoyoshi Kimura, Shengzhong Liu, Jinyang Li, Dongxin Liu, Tianshi Wang, Ruijie Wang, Yizhuo Chen, Yigong Hu, and Tarek Abdelzaher. 2024. FreqMAE: Frequency-Aware Masked Autoencoder for Multi-Modal IoT Sensing. In Proceedings of the ACM on Web Conference 2024. 2795–2806
2024
-
[56]
Moo Jin Kim, Karl Pertsch, Siddharth Karamcheti, Ted Xiao, Ashwin Balakrishna, Suraj Nair, Rafael Rafailov, Ethan Foster, Grace Lam, Pannag Sanketi, et al . 2024. OpenVLA: An Open-Source Vision- Language-Action Model. arXiv preprint arXiv:2406.09246 (2024)
2024 arXiv
-
[57]
Tomoyoshi Kimura, Jinyang Li, Tianshi Wang, Yizhuo Chen, Rui- jie Wang, Denizhan Kara, Maggie Wigness, Joydeep Bhattacharyya, Mudhakar Srivatsa, Shengzhong Liu, Mani Srivastava, Suhas Dig- gavi, and Tarek Abdelzaher. 2024. VibroFM: Towards Micro Foun- dation Models for Robust ...
2024
-
[58]
B Ravi Kiran, Ibrahim Sobh, Victor Talpaert, Patrick Mannion, Ah- mad A Al Sallab, Senthil Yogamani, and Patrick Pérez. 2021. Deep reinforcement learning for autonomous driving: A survey.IEEE Trans- actions on Intelligent Transportation Systems 23, 6 (2021), 4909–4926
2021
-
[59]
Andy Kong, Karan Ahuja, Mayank Goel, and Chris Harrison. 2021. EyeMU Interactions: Gaze + IMU Gestures on Mobile Devices. In Proceedings of the 2021 International Conference on Multimodal Inter- action (Montréal, QC, Canada) (ICMI ’21). Association for Computing Machinery, New...
2021
-
[60]
Yi-Lun Lee, Yi-Hsuan Tsai, Wei-Chen Chiu, and Chen-Yu Lee. 2023. Multimodal prompting with missing modalities for visual recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 14943–14952
2023
-
[61]
DaiFeng Li and Fan Xu. 2024. Synergizing knowledge graphs with large language models: a comprehensive review and future prospects. arXiv preprint arXiv:2407.18470 (2024)
2024 arXiv
-
[62]
Jacky Liang, Wenlong Huang, Fei Xia, Peng Xu, Karol Hausman, Brian Ichter, Pete Florence, and Andy Zeng. 2023. Code as poli- cies: Language model programs for embodied control. In 2023 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 9493–9500
2023
-
[63]
Chen, and Zihao Deng
Paul Pu Liang, Yun Cheng, Xiang Fan, Chun Kai Ling, Suzanne Nie, Richard J. Chen, and Zihao Deng. 2023. Quantifying & Modeling Multimodal Interactions: An Information Decomposition Framework. In NeurIPS
2023
-
[64]
Paul Pu Liang, Yiwei Lyu, Xiang Fan, Jeffrey Tsaw, Yudong Liu, Shentong Mo, Dani Yogatama, Louis-Philippe Morency, and Rus- lan Salakhutdinov. 2022. High-modality multimodal transformer: Quantifying modality & interaction heterogeneity for high-modality representation learning...
2022 arXiv
-
[65]
Paul Pu Liang, Amir Zadeh, and Louis-Philippe Morency. 2024. Foun- dations & trends in multimodal machine learning: Principles, chal- lenges, and open questions. Comput. Surveys 56, 10 (2024), 1–42
2024
-
[66]
Opher Lieber, Barak Lenz, Hofit Bata, Gal Cohen, Jhonathan Osin, Itay Dalmedigos, Erez Safahi, Shaked Meirom, Yonatan Belinkov, Shai Shalev-Shwartz, et al. 2024. Jamba: A hybrid transformer-mamba language model. arXiv preprint arXiv:2403.19887 (2024)
2024 arXiv
-
[67]
Kaiwei Liu, Bufang Yang, Lilin Xu, Yunqi Guo, Neiwen Ling, Zhihe Zhao, Guoliang Xing, Xian Shuai, Xiaozhe Ren, Xin Jiang, et al. 2024. Tasking Heterogeneous Sensor Systems with LLMs. In Proceedings of the 22nd ACM Conference on Embedded Networked Sensor Systems . 901–902
2024
-
[68]
Shengzhong Liu, Tomoyoshi Kimura, Dongxin Liu, Ruijie Wang, Jinyang Li, Suhas Diggavi, Mani Srivastava, and Tarek Abdelza- her. 2023. FOCAL: Contrastive Learning for Multimodal Time- Series Sensing Signals in Factorized Orthogonal Latent Space. arXiv:2310.20071 [cs.AI] https:/...
2023 arXiv
-
[69]
Xu Liu, Junfeng Hu, Yuan Li, Shizhe Diao, Yuxuan Liang, Bryan Hooi, and Roger Zimmermann. 2024. Unitime: A language-empowered unified model for cross-domain time series forecasting. InProceedings of the ACM on Web Conference 2024 . 4095–4106
2024
-
[70]
Ya Liu, Yingjie Zhou, Kai Yang, and Xin Wang. 2023. Unsupervised deep learning for IoT time series. IEEE Internet of Things Journal 10, 16 (2023), 14285–14306
2023
-
[71]
Kevin Lu, Aditya Grover, Pieter Abbeel, and Igor Mordatch. 2022. Frozen pretrained transformers as universal computation engines. In Proceedings of the AAAI conference on artificial intelligence , Vol. 36. 7628–7636
2022
-
[72]
Mengmeng Ma, Jian Ren, Long Zhao, Davide Testuggine, and Xi Peng
-
[73]
Yueen Ma, Zixing Song, Yuzheng Zhuang, Jianye Hao, and Irwin King
-
[74]
Robin Manhaeve, Sebastijan Dumancic, Angelika Kimmig, Thomas Demeester, and Luc De Raedt. 2018. DeepProbLog: Neural Probabilis- tic Logic Programming.CoRR abs/1805.10872 (2018). arXiv:1805.10872 http://arxiv.org/abs/1805.10872
2018 arXiv
-
[75]
Meta-AI. 2024. Executorch, 2024a. https://pytorch.org/executorch- overview
2024
-
[76]
Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng. 2021. Nerf: Representing scenes as neural radiance fields for view synthesis. Commun. ACM 65, 1 (2021), 99–106
2021
-
[77]
Suvir Mirchandani, Fei Xia, Pete Florence, Brian Ichter, Danny Driess, Montserrat Gonzalez Arenas, Kanishka Rao, Dorsa Sadigh, and Andy Zeng. 2023. Large Language Models as General Pattern Machines. In Conference on Robot Learning . PMLR, 2498–2518
2023
-
[78]
arXiv preprint arXiv:2405.14093 (2024)
A Survey on Vision-Language-Action Models for Embodied AI. arXiv preprint arXiv:2405.14093 (2024)
2024 arXiv
-
[79]
Shentong Mo, Russ Salakhutdinov, Louis-Philippe Morency, and Paul Pu Liang. 2024. IoT-LM: Large multisensory language mod- els for the internet of things. arXiv preprint arXiv:2407.09801 (2024)
2024 arXiv
-
[80]
Vimal Mollyn, Karan Ahuja, Dhruv Verma, Chris Harrison, and Mayank Goel. 2022. SAMoSA: Sensing Activities with Motion and Subsampled Audio. Proc. ACM Interact. Mob. Wearable Ubiquitous Technol. 6, 3, Article 132 (sep 2022), 19 pages. https://doi.org/10.1145/ 3550284 XXX ’25, J...
2022
-
[81]
Ozan Baris Mulayim, Pengrui Quan, Liying Han, Xiaomin Ouyang, Dezhi Hong, Mario Bergés, and Mani Srivastava. 2024. Are Time Series Foundation Models Ready to Revolutionize Predictive Building Analytics?. In Proceedings of the 11th ACM International Conference on Systems for En...
2024
-
[82]
Girish Narayanswamy, Xin Liu, Kumar Ayush, Yuzhe Yang, Xuhai Xu, Shun Liao, Jake Garrison, Shyam Tailor, Jake Sunshine, Yun Liu, et al. 2024. Scaling Wearable Foundation Models. arXiv preprint arXiv:2410.13638 (2024)
2024 arXiv
-
[83]
Ludovico Mitchener, David Tuckey, Matthew Crosby, and Alessandra Russo. 2022. Detect, understand, act: A neuro-symbolic hierarchical reinforcement learning framework. Machine Learning 111, 4 (2022), 1523–1549
2022
-
[84]
Abby O’Neill, Abdul Rehman, Abhinav Gupta, Abhiram Mad- dukuri, Abhishek Gupta, Abhishek Padalkar, Abraham Lee, Acorn Pooley, Agrim Gupta, Ajay Mandlekar, et al . 2023. Open X- Embodiment: Robotic learning datasets and rt-x models. arXiv preprint arXiv:2310.08864 (2023)
2023 arXiv
-
[85]
Xiaomin Ouyang, Xian Shuai, Jiayu Zhou, Ivy Wang Shi, Zhiyuan Xie, Guoliang Xing, and Jianwei Huang. 2022. Cosmo: Contrastive Fusion Learning with Small Data for Multimodal Human Activity Recognition. In International Conference on Mobile Computing And Networking (MobiCom)
2022
-
[86]
Xiaomin Ouyang and Mani Srivastava. 2024. LLMSense: Harnessing LLMs for High-level Reasoning Over Spatiotemporal Sensor Traces . In 2024 IEEE 3rd Workshop on Machine Learning on Edge in Sensor Systems (SenSys-ML). IEEE Computer Society, Los Alamitos, CA, USA, 9–14. https://doi...
2024
-
[87]
Xiaomin Ouyang, Jason Wu, Tomoyoshi Kimura, Yihan Lin, Gunjan Verma, Tarek Abdelzaher, and Mani Srivastava. 2024. MMBind: Un- leashing the Potential of Distributed and Heterogeneous Data for Multimodal Learning in IoT. arXiv preprint arXiv:2411.12126 (2024)
2024 arXiv
-
[88]
Yuqi Nie, Nam H Nguyen, Phanwadee Sinthong, and Jayant Kalagnanam. 2022. A time series is worth 64 words: Long-term fore- casting with transformers. arXiv preprint arXiv:2211.14730 (2022)
2022 arXiv
-
[89]
Srinivas Parthasarathy and Shiva Sundaram. 2021. Training Strategies to Handle Missing Modalities for Audio-Visual Expression Recogni- tion. In Companion Publication of the 2020 International Conference on Multimodal Interaction (Virtual Event, Netherlands) (ICMI ’20 Com- pani...
2021
-
[90]
Karol J Piczak. 2015. ESC: Dataset for environmental sound classi- fication. In Proceedings of the 23rd ACM international conference on Multimedia. 1015–1018
2015
-
[91]
Petra Poklukar, Miguel Vasco, Hang Yin, Francisco S Melo, Ana Paiva, and Danica Kragic. 2022. Geometric Multimodal Contrastive Repre- sentation Learning. In International Conference on Machine Learning (ICML)
2022
-
[92]
Shuwei Qian and Chongjun Wang. 2023. COM: Contrastive Masked- attention model for incomplete multimodal learning.Neural Networks 162 (2023), 443–455
2023
-
[93]
Shirui Pan, Linhao Luo, Yufei Wang, Chen Chen, Jiapu Wang, and Xindong Wu. 2024. Unifying large language models and knowledge graphs: A roadmap. IEEE Transactions on Knowledge and Data Engi- neering (2024)
2024
-
[94]
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of machine learning research 21, 140 (2020), 1–67
2020
-
[95]
Kashif Rasul, Arjun Ashok, Andrew Robert Williams, Hena Ghonia, Rishika Bhagwatkar, Arian Khorasani, Mohammad Javad Darvishi Bayazi, George Adamopoulos, Roland Riachi, Nadhir Hassen, Marin Biloš, Sahil Garg, Anderson Schneider, Nicolas Chapados, Alexandre Drouin, Valentina Zan...
-
[96]
Haoyu Ren, Darko Anicic, and Thomas A Runkler. 2021. The synergy of complex event processing and tiny machine learning in indus- trial IoT. In Proceedings of the 15th ACM International Conference on Distributed and Event-based Systems . 126–135
2021
-
[97]
Sandeep Singh Sandha, Bharathan Balaji, Luis Garcia, and Mani Srivastava. 2023. Eagle: End-to-end Deep Reinforcement Learning based Autonomous Control of PTZ Cameras. InProceedings of the 8th ACM/IEEE Conference on Internet of Things Design and Implementation. 144–157
2023
-
[98]
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. In International Conference on Machine Learni...
2021
-
[99]
Siavash Shams, Sukru Samet Dindar, Xilin Jiang, and Nima Mesgarani
-
[100]
Smith et al
Taylor G. Smith et al. 2017–. pmdarima: ARIMA estimators for Python. http://www.alkaline-ml.com/pmdarima [Online; accessed <today>]
2017
-
[101]
http://arxiv.org/abs/2310.08278 arXiv:2310.08278 [cs]
Lag-Llama: Towards Foundation Models for Probabilistic Time Series Forecasting. http://arxiv.org/abs/2310.08278 arXiv:2310.08278 [cs]
-
[102]
Lei Song, Chuheng Zhang, Li Zhao, and Jiang Bian. 2023. Pre- trained large language models for industrial control. arXiv preprint arXiv:2308.03028 (2023)
2023 arXiv
-
[103]
Dimitris Spathis and Fahim Kawsar. 2024. The first step is the hardest: Pitfalls of representing and tokenizing temporal data for large lan- guage models. Journal of the American Medical Informatics Association 31, 9 (2024), 2151–2158
2024
-
[104]
Tomoya Sawada, Takaomi Hasegawa, Keiichi Yokoyama, and Masahiro Mizuno. 2024. Office-in-the-Loop for Building HVAC Con- trol with Multimodal Foundation Models. In Proceedings of the 11th ACM International Conference on Systems for Energy-Efficient Build- ings, Cities, and Tran...
2024
-
[105]
Alexander G Tartakovsky, Aleksey S Polunchenko, and Grigory Sokolov. 2012. Efficient computer network anomaly detection by changepoint detection methods. IEEE Journal of Selected Topics in Signal Processing 7, 1 (2012), 4–11
2012
-
[106]
arXiv preprint arXiv:2405.11831 (2024)
SSAMBA: Self-supervised audio representation learning with MAMBA state space model. arXiv preprint arXiv:2405.11831 (2024)
2024 arXiv
-
[107]
Rahul Thapa, Bryan He, Magnus Ruud Kjaer, Hyatt Moore, Gauri Ganjoo, Emmanuel Mignot, and James Zou. 2024. SleepFM: Multi- modal Representation Learning for Sleep Across Brain Activity, ECG and Respiratory Signals. arXiv preprint arXiv:2405.17766 (2024)
2024 arXiv
-
[108]
Chan Hee Song, Jiaman Wu, Clayton Washington, Brian M Sadler, Wei-Lun Chao, and Yu Su. 2023. LLM-Planner: Few-shot grounded planning for embodied agents with large language models. In Pro- ceedings of the IEEE/CVF International Conference on Computer Vision . 2998–3009
2023
-
[109]
Sana Tonekaboni, Danny Eytan, and Anna Goldenberg. 2021. Un- supervised Representation Learning for Time Series with Temporal Neighborhood Coding. In International Conference on Learning Repre- sentations (ICLR)
2021
-
[110]
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023. Llama: Open and effi- cient foundation language models. arXiv preprint arXiv:2302.13971 Foundation Model...
2023 arXiv
-
[111]
Albert Tarantola. 2005. Inverse problem theory and methods for model parameter estimation. SIAM
2005
-
[112]
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. InAdvances in neural information processing systems. 5998–6008
2017
-
[113]
Workday’s Syman Team. 2017. ARIMA Java library. https://github. com/Workday/timeseries-forecast
2017
-
[114]
Brian Wang, Luis Antonio Garcia, and Mani Srivastava. 2024. Priva- cyOracle: Configuring Sensor Privacy Firewalls with Large Language Models in Smart Built Environments. In2024 IEEE Security and Privacy Workshops (SPW). IEEE, 239–245
2024
-
[115]
Yonglong Tian, Dilip Krishnan, and Phillip Isola. 2020. Con- trastive multiview coding. InEuropean Conference on Computer Vision (ECCV)
2020
-
[116]
Gary Weiss. 2019. WISDM Smartphone and Smartwatch Activity and Biometrics Dataset . UCI Machine Learning Repository. DOI: https://doi.org/10.24432/C5HK59
2019 doi
-
[117]
Gerald Woo, Chenghao Liu, Akshat Kumar, Caiming Xiong, Silvio Savarese, and Doyen Sahoo. 2024. Unified Training of Universal Time Series Forecasting Transformers. (May 2024). http://arxiv.org/abs/ 2402.02592 arXiv:2402.02592
2024 arXiv
-
[118]
Yao-Hung Hubert Tsai, Shaojie Bai, Paul Pu Liang, J Zico Kolter, Louis-Philippe Morency, and Ruslan Salakhutdinov. 2019. Multimodal Transformer for Unaligned Multimodal Language Sequences. InACL
2019
-
[119]
Haixu Wu, Tengge Hu, Yong Liu, Hang Zhou, Jianmin Wang, and Mingsheng Long. 2022. Timesnet: Temporal 2d-variation modeling for general time series analysis. arXiv preprint arXiv:2210.02186 (2022)
2022 arXiv
-
[120]
Sai Vemprala, Rogerio Bonatti, Arthur Bucker, and Ashish Kapoor. 2023. ChatGPT for Robotics: Design Principles and Model Abilities . Technical Report MSR-TR-2023-8. Microsoft. https://www.microsoft.com/en-us/research/publication/chatgpt- for-robotics-design-principles-and-mode...
2023
-
[121]
Yike Wu, Nan Hu, Sheng Bi, Guilin Qi, Jie Ren, Anhuan Xie, and Wei Song. 2023. Retrieve-Rewrite-Answer: A KG-to-Text Enhanced LLMs Framework for Knowledge Graph Question Answering. http: //arxiv.org/abs/2309.11206 arXiv:2309.11206 [cs]
2023 arXiv
-
[122]
Chao Wang, Jian Wang, Yuan Shen, and Xudong Zhang. 2019. Au- tonomous navigation of UAVs in large-scale complex environments: A deep reinforcement learning approach.IEEE Transactions on Vehicular Technology 68, 3 (2019), 2124–2136
2019
-
[123]
Huatao Xu, Liying Han, Qirui Yang, Mo Li, and Mani Srivastava
-
[124]
Huatao Xu, Pengfei Zhou, Rui Tan, Mo Li, and Guobin Shen. 2021. LIMU-BERT: Unleashing the potential of unlabeled data for IMU sensing applications. In Proceedings of the 19th ACM Conference on Embedded Networked Sensor Systems . 220–233
2021
-
[125]
Sangmin Woo, Sumin Lee, Yeonju Park, Muhammad Adi Nugroho, and Changick Kim. 2023. Towards good practices for missing modality robust action recognition. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 37. 2776–2784
2023
-
[126]
Shengzhe Xu, Christo Kurisummoottil Thomas, Omar Hashash, Nikhil Muralidhar, Walid Saad, and Naren Ramakrishnan. 2024. Large Language Models as Optimizers. arXiv preprint arXiv:2402.01748 (2024)
2024 arXiv
-
[127]
Jason Wu, Ziqi Wang, Xiaomin Ouyang, Ho Lyun Jeong, Colin Sam- plawski, Lance Kaplan, Benjamin Marlin, and Mani Srivastava. 2024. FlexLoc: Conditional Neural Networks for Zero-Shot Sensor Perspec- tive Invariance in Object Localization with Distributed Multimodal Sensors. arXi...
2024 arXiv
-
[128]
Hua Yan, Heng Tan, Yi Ding, Pengfei Zhou, Vinod Namboodiri, and Yu Yang. 2024. Language-centered Human Activity Recognition. arXiv preprint arXiv:2410.00003 (2024)
2024
-
[129]
Tianwei Xing, Luis Garcia, Marc Roig Vilamala, Federico Cerutti, Lance Kaplan, Alun Preece, and Mani Srivastava. 2020. Neuroplex: learning to detect complex events in sensor networks through knowl- edge injection. In Proceedings of the 18th conference on embedded networked sen...
2020
-
[130]
Liang Yu, Shuqi Qin, Meng Zhang, Chao Shen, Tao Jiang, and Xiao- hong Guan. 2021. A review of deep reinforcement learning for smart building energy management. IEEE Internet of Things Journal 8, 15 (2021), 12046–12063
2021
-
[131]
In Proceedings of the 25th International Workshop on Mobile Computing Systems and Applications
Penetrative ai: Making llms comprehend the physical world. In Proceedings of the 25th International Workshop on Mobile Computing Systems and Applications. 1–7
-
[132]
Jinliang Yuan, Chen Yang, Dongqi Cai, Shihe Wang, Xin Yuan, Zeling Zhang, Xiang Li, Dingge Zhang, Hanzi Mei, Xianqing Jia, et al. 2024. Mobile Foundation Model as Firmware. In Proceedings of the 30th Annual International Conference on Mobile Computing and Networking. 279–295
2024
-
[133]
Maxwell A Xu, Jaya Narain, Gregory Darnell, Haraldur Hallgrimsson, Hyewon Jeong, Darren Forde, Richard Fineman, Karthik J Raghuram, James M Rehg, and Shirley Ren. 2024. RelCon: Relative Contrastive Learning for a Motion Foundation Model for Wearable Data. arXiv preprint arXiv:...
2024 arXiv
-
[134]
Chaolv Zeng, Zhanyu Liu, Guanjie Zheng, and Linghe Kong. 2024. C-Mamba: Channel Correlation Enhanced State Space Models for Multivariate Time Series Forecasting. arXiv preprint arXiv:2406.05316 (2024)
2024 arXiv
-
[135]
Hao Xue and Flora D Salim. 2023. Promptcast: A new prompt-based learning paradigm for time series forecasting. IEEE Transactions on Knowledge and Data Engineering (2023)
2023
-
[136]
Xiyuan Zhang, Diyan Teng, Ranak Roy Chowdhury, Shuheng Li, Dezhi Hong, Rajesh K Gupta, and Jingbo Shang. 2024. UniMTS: Unified Pre-training for Motion Time Series. arXiv preprint arXiv:2410.19818 (2024)
2024 arXiv
-
[137]
Zhun Yang, Adam Ishay, and Joohyung Lee. 2023. NeurASP: Embracing Neural Networks into Answer Set Programming. arXiv:2307.07700 [cs.AI] https://arxiv.org/abs/2307.07700
2023 arXiv
-
[138]
Yu Zhang, Xiusi Chen, Bowen Jin, Sheng Wang, Shuiwang Ji, Wei Wang, and Jiawei Han. 2024. A Comprehensive Survey of Scien- tific Large Language Models and Their Applications in Scientific Discovery. arXiv preprint arXiv:2406.10833 (2024)
2024 arXiv
-
[139]
Hang Yuan, Shing Chan, Andrew P Creagh, Catherine Tong, Aidan Acquah, David A Clifton, and Aiden Doherty. 2024. Self-supervised learning for human activity recognition using 700,000 person-days of wearable data. NPJ digital medicine 7, 1 (2024), 91
2024
-
[141]
Zhihan Yue, Yujing Wang, Juanyong Duan, Tianmeng Yang, Congrui Huang, Yunhai Tong, and Bixiong Xu. 2022. TS2Vec: Towards Uni- versal Representation of Time Series. InAAAI Conference on Artificial Intelligence (AAAI)
2022
-
[143]
Xiyuan Zhang, Ranak Roy Chowdhury, Jiayun Zhang, Dezhi Hong, Rajesh K Gupta, and Jingbo Shang. 2023. Unleashing the Power of Shared Label Structures for Human Activity Recognition. In Proceed- ings of the 32nd ACM International Conference on Information and Knowledge Managemen...
2023
-
[145]
Xiang Zhang, Ziyuan Zhao, Theodoros Tsiligkaridis, and Marinka Zitnik. 2022. Self-supervised contrastive pre-training for time series via time-frequency consistency. Advances in Neural Information Processing Systems 35 (2022), 3988–4003
2022
-
[147]
Lianghui Zhu, Bencheng Liao, Qian Zhang, Xinlong Wang, Wenyu Liu, and Xinggang Wang. 2024. Vision mamba: Efficient visual repre- sentation learning with bidirectional state space model.arXiv preprint arXiv:2401.09417 (2024)
2024 arXiv
-
[336]
https://doi.org/10.1145/3536221.3556568
-
[2016]
In Proceedings of the 3rd ACM International Conference on Systems for Energy-Efficient Built Environments
Brick: Towards a Unified Metadata Schema For Buildings. In Proceedings of the 3rd ACM International Conference on Systems for Energy-Efficient Built Environments. ACM, Palo Alto CA USA, 41–50. https://doi.org/10.1145/2993422.2993577
-
[2022]
In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
Are multimodal transformers robust to missing modality?. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 18177–18186
-
[2024]
In International Conference on Neural-Symbolic Learning and Reasoning
The role of foundation models in neuro-symbolic learning and reasoning. In International Conference on Neural-Symbolic Learning and Reasoning. Springer, 84–100
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.