REVIEW 3 major objections 4 minor 218 references
SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information
T0 review · 3 major / 4 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read This paper argues that the main obstacle for LLM-based mobile assistants is not tool use but knowing which piece of scattered personal information to retrieve, and it supports this claim with a new 250-task benchmark.
desk verdict Solid new benchmark for scattered personal information; the 'commit-too-early' failure story is not actually pinned down by the evidence. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the SPIEval benchmark itself: a fully controllable simulated phone environment with 10 app schemas, 4,335 linked personal records, and 21 tools (11 retrieval, 10 execution), plus a shared user profile so instructions like "Call my manager" require resolving identity across apps. Each task has a manually constructed gold answer with exact execution calls and parameters, verified by cross-validation, so outcomes are determined by matching final tool calls, not by a judge model. Two controlled ablations, "full" (unpaginated results) and "no-search" (records in the prompt), isolate the retrieval bottleneck.
What would settle it
Compare accuracy under the standard protocol vs. the no-search protocol while systematically varying query specificity, or run the same tasks on a real phone using a GUI agent harness. If injecting progressively more specific hints does not push accuracy toward the human ceiling, or if models that solve SPIEval tasks fail real-device tasks with the same structure, then the localization-bottleneck claim would be weakened and other bottlenecks (planning, decomposition, interface understanding) would need to carry the explanation.
Extended reading notes
Core claim
According to the paper, state-of-the-art LLMs fail most mobile-assistant tasks that require scattered personal information, and the failure is concentrated in retrieval. Removing retrieval and handing all relevant records to the model raises average accuracy from 35.5% to 66.8%, while simply returning all matches unpaginated raises it only to 36.0%, implying that the bottleneck is the query, not the search results. Across the three strongest models, 79% of errors are wrong parameter values, and models perform fewer retrievals on failed tasks than on successful ones, suggesting early commitment to plausible but incorrect records. The paper also finds that fewer than 2% of 126,279 retrieval calls use advanced modes and that preference inference and multi-intent decomposition are the hardest capabilities.
Load-bearing premise
The paper's conclusions about real mobile assistants assume that performing well on its simulated, fictional phone with synthetic records tells us how models would behave on real phones with real personal data and real interfaces, since the benchmark is not validated against actual device use.
Editorial extensions
If this is right
- If information localization is the bottleneck, then improving LLMs' query reformulation, field-targeted search, and verification of whether evidence is sufficient should yield larger gains than adding more tools or longer contexts.
- The paper's no-search result implies models can reason about preferences once evidence is present; the hard part is knowing when and where to look.
- Multi-intent decomposition lags even with all records available, so progress there requires better task decomposition and planning, not just better retrieval.
- The underuse of regex and fuzzy matching suggests a concrete training or prompting target: teaching models to exploit structured retrieval modes, including field-specific and fuzzy search.
- With the best model below 60%, the paper's numbers imply that current assistants cannot be trusted to act on ambiguous personal-data requests without verification.
Reading between the lines
- Inference: If the simulated environment transfers, real-world assistants need a "verify before acting" mechanism—when search returns a plausible match among multiple candidates, the assistant should gather corroborating evidence before calling execution tools.
- Inference: The benchmark's finding about search efficiency suggests a useful standardized metric—task accuracy normalized by number of retrievals—that would separate models that search smartly from models that search a lot; the paper reports the underlying numbers but does not define such a metric.
- Inference: A testable extension would be to vary the surface form of records (typos, abbreviations, partial names) while keeping gold answers fixed; if accuracy drops sharply, that would confirm that over-reliance on exact substring matching is a measurable failure mode, not just a stylistic preference.
- Inference: Because all personal data is synthetic, the benchmark is safe to distribute but cannot capture privacy-sensitive behaviors like refusing to expose or misuse personal data; that remains a separate evaluation axis.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces SPIEval, a human-curated benchmark for evaluating large language models as mobile assistants that must gather scattered personal information across multiple apps. The benchmark contains 250 tasks covering five cognitive capabilities, 4,335 synthetic personal records across 10 simulated apps, and 21 tools for retrieval and execution. The authors evaluate nine LLMs under high- and low-reasoning settings and report that the best model, GPT-5.5 (xhigh), achieves only 57.3% accuracy, while the weakest reaches 16.4%. They further report that 79% of failures arise from inaccurate information localization, that models often commit to plausible but incorrect information instead of continuing retrieval, that fewer than 2% of retrieval actions use advanced search modes, and that successful tasks involve more retrieval calls than failed ones.
Significance. If the benchmark is valid, SPIEval fills a real gap: prior mobile-assistant benchmarks either provide all necessary information in the instruction or restrict personal data to single-document retrieval. The construction process is rigorous and clearly described: human annotators create tasks, independent annotators cross-validate without access to gold answers, gold answers are executed against the tool implementation, and the environment is deterministic with verifiable outcomes. The evaluation protocol is also solid, with three runs per setting, reported standard deviations, a paginated retrieval condition, and controlled 'full' and 'no-search' conditions that isolate the role of retrieval. The headline finding—that top LLMs solve only about half of these tasks and that performance improves substantially when all relevant records are provided—is a useful, falsifiable result. However, the paper's more specific causal claim about the dominant failure mechanism, namely that models prematurely commit to plausible but wrong information rather than continuing to search, is not supported by the current analyses and needs additional evidence or careful rewording.
major comments (3)
- [Section 5 (Figure 6) and Section 6 (Figure 7)] The central conclusion that 79% of failures stem from 'inaccurate information localization' and that LLMs 'commit to plausible but incorrect information instead of continuing retrieval for verification' is an inference, not a demonstrated mechanism. The error-type taxonomy in Figure 6 lumps together at least three distinct failure processes: (a) the gold record never appeared in any retrieved result because the model's query was insufficient; (b) the gold record was retrieved but the model acted on a distractor; and (c) the correct record was retrieved and selected but a parameter value was misread or inferred incorrectly. Figure 7's finding that failed tasks involve fewer retrievals than solved tasks is consistent with premature stopping, but it is equally consistent with the model failing to formulate an effective query and ceasing search after fruitless attempts. The no-search condition in Figure 5 removes all retrieval-related causes at once, so it cannot disentangle query formulation failure from premature commitment. Please add a trajectory-level analysis that records, for each failed task, whether the gold record appeared in any retrieved context; this would separate 'never retrieved' from 'retrieved but ignored.' If such an analysis is not feasible, the abstract and Section 6 should be reworded to present premature commitment as one plausible hypothesis rather than the established bottleneck.
- [Abstract and Section 5] The '79% of failures' statistic is presented without a qualifier in the abstract and conclusion, but the Figure 6 analysis is based on only the three strongest models (GPT-5.5 xhigh, Gemini 3.1 Pro high, Claude Opus 4.8 max), as stated in the Figure 6 caption and Section 5. If the claim is meant to apply to all nine evaluated models, the failure-mode analysis should be extended to the full set; otherwise, the abstract should specify that the 79% figure is for the three best models. As written, the abstract overstates the generality of the result.
- [Sections 1, 5, and 7] The paper repeatedly generalizes to 'current LLM-based mobile assistants' and 'mobile assistants,' but the evaluation environment is a simulated, text-based tool API with no GUI, no device-specific latencies, and synthetic personal records. No evidence is provided that performance in this environment transfers to real mobile assistants. Please add an explicit limitations paragraph and qualify statements such as 'the primary bottleneck for current mobile assistants' (Section 5) and 'fundamental limitations of current LLM-based mobile assistants' (abstract), or provide a small real-device validation study to support the transfer.
minor comments (4)
- [Section 4 and Appendices A/C] The paper does not state the language(s) in which the user instructions and system prompts were administered. The system prompt template is shown in both Chinese and English, and the tool schemas are in Chinese. Please specify the evaluation language and any translation/QA procedures for the English version, since accuracy may be language-dependent.
- [Section 3.4, Figure 3] The statement that 'one instruction consists of only 20 characters yet requires 45 execution parameters' should clarify the metric (characters in which language? including spaces?) to be reproducible.
- [Section 5, Figure 7] The differences in average retrieval counts between solved and failed tasks are reported for each model but without any significance testing or confidence intervals; please add statistical tests or confidence intervals to support the claim of a systematic pattern.
- [Appendix D, Table 5] The binary 'Human-Curated' column used to compare with existing benchmarks is not defined in the text; a one-sentence definition and references for the negative entries would help the reader interpret the comparison.
Circularity Check
No significant circularity: benchmark performance is an external measurement against human-annotated gold answers, and the bottleneck analysis rests on controlled protocol ablations rather than fitted or self-referential claims.
full rationale
SPIEval's construction chain is self-contained: tasks are manually authored, gold answers are human-annotated and cross-validated (Section 3.3), and models are scored by exact match of final execution tool calls to those gold answers (Section 4). None of the reported accuracy numbers is fitted, and no parameter is inferred from the data it later explains. The central analytical claim, that retrieval/localization is the primary bottleneck, is supported by independent protocol manipulations: removing retrieval entirely raises average accuracy from 35.5% to 66.8%, while de-pagination alone gives only 35.5% to 36.0% (Section 5, Figure 5). This is an A/B comparison over the same models and tasks, not a definitional equivalence. The '79% of failures' figure is a coded error-type distribution (Figure 6) over actual trajectories, and the 'commit to plausible but incorrect information' mechanism is an interpretation of that distribution together with retrieval-count comparisons (Figure 7); even if that mechanistic inference is weaker than the authors claim, it is an evidentiary gap rather than a circular derivation. The only self-citations (e.g., Ye et al. 2023; Ye et al. 2025; Tencent Hunyuan Team 2026) are background references for instruction-following, tool-use, and the Hy3 model card; none is load-bearing for the benchmark's validity or for the main results.
Assumptions & free parameters
assumptions (5)
- domain assumption The simulated apps and structured schemas are faithful abstractions of real mobile applications.
- domain assumption The five cognitive capabilities (reasoning, disambiguation, integration, preference inference, multi-intent decomposition) are the essential capabilities for completing tasks over scattered personal information.
- domain assumption The cross-validation protocol ensures that every task has a unique and verifiable gold answer and that human accuracy is 100%.
- domain assumption All information needed to complete a task is available in the constructed records and can be retrieved through the provided tools.
- domain assumption LLM performance in the simulated environment predicts performance on real mobile assistants.
Cite this review
Pith. "Pith review of SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information." pith.science (2026). https://pith.science/paper/ENCFUWTB
@misc{pith2026260810692,
author = {Pith},
title = {Pith review of: SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information},
year = {2026},
howpublished = {\url{https://pith.science/paper/ENCFUWTB}},
note = {Machine review of arXiv:2608.10692}
}
read the original abstract
Large language models (LLMs) are increasingly deployed as mobile assistants, where a key challenge is leveraging personal information scattered across multiple applications (apps) to complete user instructions. However, due to the lack of dedicated benchmarks, their capabilities remain poorly understood. To address this gap, we introduce SPIEval, a human-curated benchmark grounded in five cognitive capabilities (i.e., reasoning, disambiguation, integration, preference inference, and multi-intent decomposition). SPIEval comprises 250 tasks spanning 4,335 personal records distributed across 10 apps and supports multi-turn interaction through 21 tools. Analysis shows that the benchmark exhibits diverse scenarios, challenging tasks, scattered information, controllable environments, and verifiable outcomes. We evaluate nine representative LLMs and find substantial room for improvement. The best-performing model, GPT-5.5 (xhigh), achieves only 57.3% accuracy, while the weakest achieves just 16.4%. Further analysis reveals that 79% of failures stem from inaccurate information localization, as LLMs often commit to plausible but incorrect information instead of continuing retrieval for verification. We also find that fewer than 2% of retrieval actions employ advanced search methods and observe substantial variation in search efficiency across models. These findings expose fundamental limitations of current LLM-based mobile assistants and motivate future research in this direction. Data and code are available at https://huggingface.co/datasets/Junjie-Ye/SPIEval.
Figures
Figures from the paper (5 more)
Reference graph
Works this paper leans on
-
[1]
Langley , title =
P. Langley , title =. Proceedings of the 17th International Conference on Machine Learning (ICML 2000) , address =. 2000 , pages =
2000
-
[2]
Xiao Liu and Bo Qin and Dongzhu Liang and Guang Dong and Hanyu Lai and Hanchen Zhang and Hanlin Zhao and Iat Long Iong and Jiadai Sun and Jiaqi Wang and Junjie Gao and Junjun Shan and Kangning Liu and Shudan Zhang and Shuntian Yao and Siyi Cheng and Wentao Yao and Wenyi Zhao and Xinghan Liu and Xinyi Liu and Xinying Chen and Xinyue Yang and Yang Yang and ...
-
[3]
Jisoo Mok and Ik. Exploring the Potential of LLMs as Personalized Assistants: Dataset, Evaluation, and Analysis , booktitle =. 2025 , url =. doi:10.18653/V1/2025.ACL-LONG.504 , timestamp =
-
[4]
Juntao Tan and Liangwei Yang and Zuxin Liu and Zhiwei Liu and Rithesh R. N. and Tulika Manoj Awalgaonkar and Jianguo Zhang and Weiran Yao and Ming Zhu and Shirley Kokane and Silvio Savarese and Huan Wang and Caiming Xiong and Shelby Heinecke , editor =. PersonaBench: Evaluating. Findings of the Association for Computational Linguistics,. 2025 , url =. doi...
-
[5]
Romain Froger and Pierre Andrews and Matteo Bettini and Amar Budhiraja and Ricardo Silveira Cabral and Virginie Do and Emilien Garreau and Jean. Gaia2: Benchmarking. CoRR , volume =. 2026 , url =. doi:10.48550/ARXIV.2602.11964 , eprinttype =. 2602.11964 , timestamp =
-
[6]
The Thirteenth International Conference on Learning Representations,
Jingxuan Chen and Derek Yuen and Bin Xie and Yuhao Yang and Gongwei Chen and Zhihao Wu and Li Yixing and Xurui Zhou and Weiwen Liu and Shuai Wang and Kaiwen Zhou and Rui Shao and Liqiang Nie and Yasheng Wang and Jianye Hao and Jun Wang and Kun Shao , title =. The Thirteenth International Conference on Learning Representations,. 2025 , url =
2025
-
[7]
AndroidLab: Training and Systematic Benchmarking of Android Autonomous Agents , booktitle =
Yifan Xu and Xiao Liu and Xueqiao Sun and Siyi Cheng and Hao Yu and Hanyu Lai and Shudan Zhang and Dan Zhang and Jie Tang and Yuxiao Dong , editor =. AndroidLab: Training and Systematic Benchmarking of Android Autonomous Agents , booktitle =. 2025 , url =. doi:10.18653/V1/2025.ACL-LONG.107 , timestamp =
-
[8]
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments , booktitle =
Tianbao Xie and Danyang Zhang and Jixuan Chen and Xiaochuan Li and Siheng Zhao and Ruisheng Cao and Toh Jing Hua and Zhoujun Cheng and Dongchan Shin and Fangyu Lei and Yitao Liu and Yiheng Xu and Shuyan Zhou and Silvio Savarese and Caiming Xiong and Victor Zhong and Tao Yu , editor =. OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Co...
2024
Show all 218 references
-
[9]
Mind the Third Eye! Benchmarking Privacy Awareness in MLLM-powered Smartphone Agents , booktitle =
Zhixin Lin and Jungang Li and Shidong Pan and Yibo Shi and Yue Yao and Dongliang Xu , editor =. Mind the Third Eye! Benchmarking Privacy Awareness in MLLM-powered Smartphone Agents , booktitle =. 2026 , url =. doi:10.1609/AAAI.V40I42.40874 , timestamp =
2026 doi
-
[10]
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers),
Xueyu Hu and Tao Xiong and Biao Yi and Zishu Wei and Ruixuan Xiao and Yurun Chen and Jiasheng Ye and Meiling Tao and Xiangxin Zhou and Ziyu Zhao and Yuhuai Li and Shengze Xu and Shenzhi Wang and Xinchen Xu and Shuofei Qiao and Zhaokai Wang and Kun Kuang and Tieyong Zeng and Li...
2025
-
[11]
Guangyi Liu and Pengxiang Zhao and Yaozhen Liang and Liang Liu and Yaxuan Guo and Han Xiao and Weifeng Lin and Yuxiang Chai and Yue Han and Shuai Ren and Hao Wang and Xiaoyu Liang and WenHao Wang and Tianze Wu and Zhengxi Lu and Siheng Chen and LiLinghao and Hao Wang and Guanj...
2025
-
[12]
AppAgent: Multimodal Agents as Smartphone Users , booktitle =
Chi Zhang and Zhao Yang and Jiaxuan Liu and Yanda Li and Yucheng Han and Xin Chen and Zebiao Huang and Bin Fu and Gang Yu , editor =. AppAgent: Multimodal Agents as Smartphone Users , booktitle =. 2025 , url =. doi:10.1145/3706598.3713600 , timestamp =
2025
- [13]
-
[14]
CoRR , volume =
DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence , author=. CoRR , volume =. 2026 , url =. doi:10.48550/ARXIV.2606.19348 , eprinttype =
2026 doi
- [15]
-
[16]
From Complex to Simple: Enhancing Multi-Constraint Complex Instruction Following Ability of Large Language Models , booktitle =
Qianyu He and Jie Zeng and Qianxi He and Jiaqing Liang and Yanghua Xiao , editor =. From Complex to Simple: Enhancing Multi-Constraint Complex Instruction Following Ability of Large Language Models , booktitle =. 2024 , url =. doi:10.18653/V1/2024.FINDINGS-EMNLP.637 , timestamp =
2024 doi
-
[17]
The Thirteenth International Conference on Learning Representations,
Guanting Dong and Keming Lu and Chengpeng Li and Tingyu Xia and Bowen Yu and Chang Zhou and Jingren Zhou , title =. The Thirteenth International Conference on Learning Representations,. 2025 , url =
2025
-
[18]
The Twelfth International Conference on Learning Representations,
Bill Yuchen Lin and Abhilasha Ravichander and Ximing Lu and Nouha Dziri and Melanie Sclar and Khyathi Raghavi Chandu and Chandra Bhagavatula and Yejin Choi , title =. The Twelfth International Conference on Learning Representations,. 2024 , url =
2024
-
[19]
The Thirteenth International Conference on Learning Representations,
Hao Zhao and Maksym Andriushchenko and Francesco Croce and Nicolas Flammarion , title =. The Thirteenth International Conference on Learning Representations,. 2025 , url =
2025
-
[20]
FanOutQA:
Andrew Zhu and Alyssa Hwang and Liam Dugan and Chris Callison. FanOutQA:. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers),. 2024 , url =. doi:10.18653/V1/2024.ACL-SHORT.2 , timestamp =
2024 doi
-
[21]
Evaluating Design Decisions for Dual Encoder-based Entity Disambiguation , booktitle =
Susanna R. Evaluating Design Decisions for Dual Encoder-based Entity Disambiguation , booktitle =. 2025 , url =. doi:10.18653/V1/2025.ACL-LONG.764 , timestamp =
2025 doi
-
[22]
Cohen and Ruslan Salakhutdinov and Christopher D
Zhilin Yang and Peng Qi and Saizheng Zhang and Yoshua Bengio and William W. Cohen and Ruslan Salakhutdinov and Christopher D. Manning , editor =. HotpotQA:. Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Brussels, Belgium, October 31 - ...
2018 doi
- [23]
-
[24]
On the Multi-turn Instruction Following for Conversational Web Agents , booktitle =
Yang Deng and Xuan Zhang and Wenxuan Zhang and Yifei Yuan and See. On the Multi-turn Instruction Following for Conversational Web Agents , booktitle =. 2024 , url =. doi:10.18653/V1/2024.ACL-LONG.477 , timestamp =
2024 doi
-
[25]
Toward Verifiable Instruction-Following Alignment for Retrieval Augmented Generation , booktitle =
Guanting Dong and Xiaoshuai Song and Yutao Zhu and Runqi Qiao and Zhicheng Dou and Ji. Toward Verifiable Instruction-Following Alignment for Retrieval Augmented Generation , booktitle =. 2025 , url =. doi:10.1609/AAAI.V39I22.34551 , timestamp =
2025 doi
-
[26]
TL-Training:
Junjie Ye and Yilong Wu and Sixian Li and Yuming Yang and Zhiheng Xi and Tao Gui and Qi Zhang and Xuanjing Huang and Peng Wang and Zhongchao Shi and Jianping Fan and Zhengyin Du , editor =. TL-Training:. Findings of the Association for Computational Linguistics:. 2025 , url =
2025
- [27]
- [28]
-
[29]
Tool learning with large language models: a survey , journal =
Changle Qu and Sunhao Dai and Xiaochi Wei and Hengyi Cai and Shuaiqiang Wang and Dawei Yin and Jun Xu and Ji. Tool learning with large language models: a survey , journal =. 2025 , url =. doi:10.1007/S11704-024-40678-2 , timestamp =
2025 doi
-
[30]
T. M. Mitchell. The Need for Biases in Learning Generalizations. 1980
1980
-
[31]
M. J. Kearns , title =
-
[32]
Machine Learning: An Artificial Intelligence Approach, Vol. I. 1983
1983
-
[33]
R. O. Duda and P. E. Hart and D. G. Stork. Pattern Classification. 2000
2000
-
[34]
Suppressed for Anonymity , author=
-
[35]
Newell and P
A. Newell and P. S. Rosenbloom. Mechanisms of Skill Acquisition and the Law of Practice. Cognitive Skills and Their Acquisition. 1981
1981
-
[36]
A. L. Samuel. Some Studies in Machine Learning Using the Game of Checkers. IBM Journal of Research and Development. 1959
1959
-
[37]
Toolformer: Language Models Can Teach Themselves to Use Tools , booktitle =
Timo Schick and Jane Dwivedi. Toolformer: Language Models Can Teach Themselves to Use Tools , booktitle =. 2023 , url =
2023
-
[38]
and Stoica, Ion and Xing, Eric P
Chiang, Wei-Lin and Li, Zhuohan and Lin, Zi and Sheng, Ying and Wu, Zhanghao and Zhang, Hao and Zheng, Lianmin and Zhuang, Siyuan and Zhuang, Yonghao and Gonzalez, Joseph E. and Stoica, Ion and Xing, Eric P. , month =. Vicuna: An Open-Source Chatbot Impressing GPT-4 with 90\ url =
-
[39]
Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback , journal =
Yuntao Bai and Andy Jones and Kamal Ndousse and Amanda Askell and Anna Chen and Nova DasSarma and Dawn Drain and Stanislav Fort and Deep Ganguli and Tom Henighan and Nicholas Joseph and Saurav Kadavath and Jackson Kernion and Tom Conerly and Sheer El Showk and Nelson Elhage an...
-
[40]
Brown and Benjamin Chess and Rewon Child and Scott Gray and Alec Radford and Jeffrey Wu and Dario Amodei , title =
Jared Kaplan and Sam McCandlish and Tom Henighan and Tom B. Brown and Benjamin Chess and Rewon Child and Scott Gray and Alec Radford and Jeffrey Wu and Dario Amodei , title =. CoRR , volume =. 2020 , url =. 2001.08361 , timestamp =
2020 arXiv
-
[41]
2011 , publisher=
Animal tool behavior: the use and manufacture of tools by animals , author=. 2011 , publisher=
2011
- [42]
-
[43]
Improving Language Models by Retrieving from Trillions of Tokens , booktitle =
Sebastian Borgeaud and Arthur Mensch and Jordan Hoffmann and Trevor Cai and Eliza Rutherford and Katie Millican and George van den Driessche and Jean. Improving Language Models by Retrieving from Trillions of Tokens , booktitle =. 2022 , url =
2022
-
[44]
Chameleon: Plug-and-Play Compositional Reasoning with Large Language Models , journal =
Pan Lu and Baolin Peng and Hao Cheng and Michel Galley and Kai. Chameleon: Plug-and-Play Compositional Reasoning with Large Language Models , journal =. 2023 , url =. doi:10.48550/ARXIV.2304.09842 , eprinttype =. 2304.09842 , timestamp =
-
[45]
LaMDA: Language Models for Dialog Applications , journal =
Romal Thoppilan and Daniel De Freitas and Jamie Hall and Noam Shazeer and Apoorv Kulshreshtha and Heng. LaMDA: Language Models for Dialog Applications , journal =. 2022 , url =. 2201.08239 , timestamp =
2022 arXiv
- [46]
- [47]
- [48]
- [49]
- [50]
-
[51]
CoRR , volume =
Jingqing Ruan and Yihong Chen and Bin Zhang and Zhiwei Xu and Tianpeng Bao and Guoqing Du and Shiwei Shi and Hangyu Mao and Xingyu Zeng and Rui Zhao , title =. CoRR , volume =. 2023 , url =. doi:10.48550/ARXIV.2308.03427 , eprinttype =. 2308.03427 , timestamp =
2023 doi
-
[52]
CoRR , volume =
Reiichiro Nakano and Jacob Hilton and Suchir Balaji and Jeff Wu and Long Ouyang and Christina Kim and Christopher Hesse and Shantanu Jain and Vineet Kosaraju and William Saunders and Xu Jiang and Karl Cobbe and Tyna Eloundou and Gretchen Krueger and Kevin Button and Matthew Kn...
2021 arXiv
-
[53]
Zhao and Yanping Huang and Andrew M
Hyung Won Chung and Le Hou and Shayne Longpre and Barret Zoph and Yi Tay and William Fedus and Eric Li and Xuezhi Wang and Mostafa Dehghani and Siddhartha Brahma and Albert Webson and Shixiang Shane Gu and Zhuyun Dai and Mirac Suzgun and Xinyun Chen and Aakanksha Chowdhery and...
-
[54]
Chi and Tatsunori Hashimoto and Oriol Vinyals and Percy Liang and Jeff Dean and William Fedus , title =
Jason Wei and Yi Tay and Rishi Bommasani and Colin Raffel and Barret Zoph and Sebastian Borgeaud and Dani Yogatama and Maarten Bosma and Denny Zhou and Donald Metzler and Ed H. Chi and Tatsunori Hashimoto and Oriol Vinyals and Percy Liang and Jeff Dean and William Fedus , titl...
2022
-
[55]
Narasimhan and Yuan Cao , title =
Shunyu Yao and Jeffrey Zhao and Dian Yu and Nan Du and Izhak Shafran and Karthik R. Narasimhan and Yuan Cao , title =. The Eleventh International Conference on Learning Representations,. 2023 , url =
2023
-
[56]
Chi and Quoc V
Jason Wei and Xuezhi Wang and Dale Schuurmans and Maarten Bosma and Brian Ichter and Fei Xia and Ed H. Chi and Quoc V. Le and Denny Zhou , editor =. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models , booktitle =. 2022 , url =
2022
-
[57]
CoRR , volume =
Jianlin Su and Yu Lu and Shengfeng Pan and Bo Wen and Yunfeng Liu , title =. CoRR , volume =. 2021 , url =. 2104.09864 , timestamp =
2021 arXiv
- [58]
- [59]
-
[60]
GPT4Tools: Teaching Large Language Model to Use Tools via Self-instruction , booktitle =
Rui Yang and Lin Song and Yanwei Li and Sijie Zhao and Yixiao Ge and Xiu Li and Ying Shan , editor =. GPT4Tools: Teaching Large Language Model to Use Tools via Self-instruction , booktitle =. 2023 , url =
2023
-
[61]
, author=
The generalisation of student's problems when several different population variances are involved. , author=. Biometrika , volume=. 1947 , url=
1947
-
[62]
Patil and Tianjun Zhang and Xin Wang and Joseph E
Shishir G. Patil and Tianjun Zhang and Xin Wang and Joseph E. Gonzalez , editor =. Gorilla: Large Language Model Connected with Massive APIs , booktitle =. 2024 , url =
2024
-
[63]
Saravia, Elvis , journal =
-
[64]
Label Words are Anchors: An Information Flow Perspective for Understanding In-Context Learning , booktitle =
Lean Wang and Lei Li and Damai Dai and Deli Chen and Hao Zhou and Fandong Meng and Jie Zhou and Xu Sun , editor =. Label Words are Anchors: An Information Flow Perspective for Understanding In-Context Learning , booktitle =. 2023 , url =. doi:10.18653/V1/2023.EMNLP-MAIN.609 , ...
2023 doi
-
[65]
Long Ouyang and Jeffrey Wu and Xu Jiang and Diogo Almeida and Carroll L. Wainwright and Pamela Mishkin and Chong Zhang and Sandhini Agarwal and Katarina Slama and Alex Ray and John Schulman and Jacob Hilton and Fraser Kelton and Luke Miller and Maddie Simens and Amanda Askell ...
2022
-
[66]
Rethinking the Role of Demonstrations: What Makes In-Context Learning Work? , booktitle =
Sewon Min and Xinxi Lyu and Ari Holtzman and Mikel Artetxe and Mike Lewis and Hannaneh Hajishirzi and Luke Zettlemoyer , editor =. Rethinking the Role of Demonstrations: What Makes In-Context Learning Work? , booktitle =. 2022 , url =. doi:10.18653/V1/2022.EMNLP-MAIN.759 , timestamp =
2022 doi
-
[67]
Nye and Maarten Bosma and Henryk Michalewski and David Dohan and Ellen Jiang and Carrie J
Jacob Austin and Augustus Odena and Maxwell I. Nye and Maarten Bosma and Henryk Michalewski and David Dohan and Ellen Jiang and Carrie J. Cai and Michael Terry and Quoc V. Le and Charles Sutton , title =. CoRR , volume =. 2021 , url =. 2108.07732 , timestamp =
2021 arXiv
-
[68]
Measuring Mathematical Problem Solving With the
Dan Hendrycks and Collin Burns and Saurav Kadavath and Akul Arora and Steven Basart and Eric Tang and Dawn Song and Jacob Steinhardt , editor =. Measuring Mathematical Problem Solving With the. Proceedings of the Neural Information Processing Systems Track on Datasets and Benc...
2021
-
[69]
Daya Guo and Dejian Yang and Haowei Zhang and Junxiao Song and Peiyi Wang and Qihao Zhu and Runxin Xu and Ruoyu Zhang and Shirong Ma and Xiao Bi and Xiaokang Zhang and Xingkai Yu and Yu Wu and Z. F. Wu and Zhibin Gou and Zhihong Shao and Zhuoshu Li and Ziyi Gao and Aixin Liu a...
2025
- [70]
-
[71]
InFoBench: Evaluating Instruction Following Ability in Large Language Models , booktitle =
Yiwei Qin and Kaiqiang Song and Yebowen Hu and Wenlin Yao and Sangwoo Cho and Xiaoyang Wang and Xuansheng Wu and Fei Liu and Pengfei Liu and Dong Yu , editor =. InFoBench: Evaluating Instruction Following Ability in Large Language Models , booktitle =. 2024 , url =. doi:10.186...
2024 doi
- [72]
-
[73]
CIF-Bench:
Yizhi Li and Ge Zhang and Xingwei Qu and Jiali Li and Zhaoqun Li and Noah Wang and Hao Li and Ruibin Yuan and Yinghao Ma and Kai Zhang and Wangchunshu Zhou and Yiming Liang and Lei Zhang and Lei Ma and Jiajun Zhang and Zuowen Li and Wenhao Huang and Chenghua Lin and Jie Fu , e...
2024 doi
-
[74]
Manning and Stefano Ermon and Chelsea Finn , editor =
Rafael Rafailov and Archit Sharma and Eric Mitchell and Christopher D. Manning and Stefano Ermon and Chelsea Finn , editor =. Direct Preference Optimization: Your Language Model is Secretly a Reward Model , booktitle =. 2023 , url =
2023
- [75]
-
[76]
From Complex to Simple: Enhancing Multi-Constraint Complex Instruction Following Ability of Large Language Models , booktitle =
Qianyu He and Jie Zeng and Qianxi He and Jiaqing Liang and Yanghua Xiao , editor =. From Complex to Simple: Enhancing Multi-Constraint Complex Instruction Following Ability of Large Language Models , booktitle =. 2024 , url =
2024
- [77]
- [78]
- [79]
-
[80]
CoRR , volume =
Yun He and Di Jin and Chaoqi Wang and Chloe Bi and Karishma Mandyam and Hejia Zhang and Chen Zhu and Ning Li and Tengyu Xu and Hongjiang Lv and Shruti Bhosale and Chenguang Zhu and Karthik Abinav Sankararaman and Eryk Helenowski and Melanie Kambadur and Aditya Tayade and Hao M...
-
[81]
FollowBench:
Yuxin Jiang and Yufei Wang and Xingshan Zeng and Wanjun Zhong and Liangyou Li and Fei Mi and Lifeng Shang and Xin Jiang and Qun Liu and Wei Wang , editor =. FollowBench:. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Pa...
2024 doi
- [82]
-
[83]
CoRR , volume =
An Yang and Anfeng Li and Baosong Yang and Beichen Zhang and Binyuan Hui and Bo Zheng and Bowen Yu and Chang Gao and Chengen Huang and Chenxu Lv and Chujie Zheng and Dayiheng Liu and Fan Zhou and Fei Huang and Feng Hu and Hao Ge and Haoran Wei and Huan Lin and Jialong Tang and...
-
[84]
Findings of the Association for Computational Linguistics,
Bo Li and Gexiang Fang and Wei Ye and Zhenghua Xu and Jinglei Zhang and Hao Cheng and Shikun Zhang , editor =. Findings of the Association for Computational Linguistics,. 2025 , url =
2025
-
[85]
CoRR , volume =
Xiao Wang and Weikang Zhou and Can Zu and Han Xia and Tianze Chen and Yuansen Zhang and Rui Zheng and Junjie Ye and Qi Zhang and Tao Gui and Jihua Kang and Jingsheng Yang and Siyuan Li and Chunsai Du , title =. CoRR , volume =. 2023 , url =. doi:10.48550/ARXIV.2304.08085 , epr...
-
[86]
Long Context vs
Xinze Li and Yushi Bai and Bowen Jin and Fengbin Zhu and Liangming Pan and Yixin Cao , editor =. Long Context vs. Proceedings of the 48th International. 2025 , url =. doi:10.1145/3726302.3731690 , timestamp =
2025
-
[87]
CoRR , volume =
Junjie Ye and Caishuang Huang and Zhuohan Chen and Wenjie Fu and Chenyuan Yang and Leyi Yang and Yilong Wu and Peng Wang and Meng Zhou and Xiaolong Yang and Tao Gui and Qi Zhang and Zhongchao Shi and Jianping Fan and Xuanjing Huang , title =. CoRR , volume =. 2025 , url =. doi...
-
[88]
The Thirteenth International Conference on Learning Representations,
Yanzhao Qin and Tao Zhang and Tao Zhang and Yanjun Shen and Wenjing Luo and Haoze Sun and Yan Zhang and Yujing Qiao and Weipeng Chen and Zenan Zhou and Wentao Zhang and Bin Cui , title =. The Thirteenth International Conference on Learning Representations,. 2025 , url =
2025
-
[89]
Hanjie and Runzhe Yang and Karthik R
Shunyu Yao and Howard Chen and Austin W. Hanjie and Runzhe Yang and Karthik R. Narasimhan , title =. The Twelfth International Conference on Learning Representations,. 2024 , url =
2024
- [90]
- [91]
- [92]
-
[93]
CoRR , volume =
An Yang and Baosong Yang and Beichen Zhang and Binyuan Hui and Bo Zheng and Bowen Yu and Chengyuan Li and Dayiheng Liu and Fei Huang and Haoran Wei and Huan Lin and Jian Yang and Jianhong Tu and Jianwei Zhang and Jianxin Yang and Jiaxi Yang and Jingren Zhou and Junyang Lin and...
-
[94]
2024 , url=
Introducing Llama 3.1: Our most capable models to date , author=. 2024 , url=
2024
-
[95]
2023 , url=
NexusRaven-V2: Surpassing GPT-4 for Zero-shot Function Calling , author=. 2023 , url=
2023
- [96]
-
[97]
ToolHop:
Junjie Ye and Zhengyin Du and Xuesong Yao and Weijian Lin and Yufei Xu and Zehui Chen and Zaiyuan Wang and Sining Zhu and Zhiheng Xi and Siyu Yuan and Tao Gui and Qi Zhang and Xuanjing Huang and Jiecao Chen , editor =. ToolHop:. Proceedings of the 63rd Annual Meeting of the As...
2025
-
[98]
2023 , url=
Fortify the Shortest Stave in Attention: Enhancing Context Awareness of Large Language Models for Effective Tool Use , author=. 2023 , url=
2023
-
[99]
T-Eval: Evaluating the Tool Utilization Capability of Large Language Models Step by Step , booktitle =
Zehui Chen and Weihua Du and Wenwei Zhang and Kuikun Liu and Jiangning Liu and Miao Zheng and Jingming Zhuo and Songyang Zhang and Dahua Lin and Kai Chen and Feng Zhao , editor =. T-Eval: Evaluating the Tool Utilization Capability of Large Language Models Step by Step , bookti...
2024 doi
-
[100]
TaskBench: Benchmarking Large Language Models for Task Automation , booktitle =
Yongliang Shen and Kaitao Song and Xu Tan and Wenqi Zhang and Kan Ren and Siyu Yuan and Weiming Lu and Dongsheng Li and Yueting Zhuang , editor =. TaskBench: Benchmarking Large Language Models for Task Automation , booktitle =. 2024 , url =
2024
-
[101]
API-Bank:
Minghao Li and Yingxiu Zhao and Bowen Yu and Feifan Song and Hangyu Li and Haiyang Yu and Zhoujun Li and Fei Huang and Yongbin Li , editor =. API-Bank:. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing,. 2023 , url =
2023
- [102]
-
[103]
Select-Then-Decompose: From Empirical Analysis to Adaptive Selection Strategy for Task Decomposition in Large Language Models , booktitle =
Shuodi Liu and Yingzhuo Liu and Zi Wang and Yusheng Wang and Huijia Wu and Liuyu Xiang and Zhaofeng He , editor =. Select-Then-Decompose: From Empirical Analysis to Adaptive Selection Strategy for Task Decomposition in Large Language Models , booktitle =. 2025 , url =. doi:10....
2025 doi
-
[104]
Zhehao Zhang and Ryan A. Rossi and Branislav Kveton and Yijia Shao and Diyi Yang and Hamed Zamani and Franck Dernoncourt and Joe Barrow and Tong Yu and Sungchul Kim and Ruiyi Zhang and Jiuxiang Gu and Tyler Derr and Hongjie Chen and Junda Wu and Xiang Chen and Zichao Wang and ...
2025
- [105]
-
[106]
Budget-Constrained Tool Learning with Planning , booktitle =
Yuanhang Zheng and Peng Li and Ming Yan and Ji Zhang and Fei Huang and Yang Liu , editor =. Budget-Constrained Tool Learning with Planning , booktitle =. 2024 , url =. doi:10.18653/V1/2024.FINDINGS-ACL.536 , timestamp =
2024 doi
-
[107]
CoRR , volume =
Zekun Li and Shinda Huang and Jiangtian Wang and Nathan Zhang and Antonis Antoniades and Wenyue Hua and Kaijie Zhu and Sirui Zeng and William Yang Wang and Xifeng Yan , title =. CoRR , volume =. 2025 , url =. doi:10.48550/ARXIV.2503.08669 , eprinttype =. 2503.08669 , timestamp =
-
[108]
Renze Lou and Kai Zhang and Wenpeng Yin , title =. Comput. Linguistics , volume =. 2024 , url =. doi:10.1162/COLI\_A\_00523 , timestamp =
2024 doi
-
[109]
AppWorld:
Harsh Trivedi and Tushar Khot and Mareike Hartmann and Ruskin Manku and Vinty Dong and Edward Li and Shashank Gupta and Ashish Sabharwal and Niranjan Balasubramanian , editor =. AppWorld:. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics ...
2024 doi
-
[110]
CoRR , volume =
Ryan Wong and Jiawei Wang and Junjie Zhao and Li Chen and Yan Gao and Long Zhang and Xuan Zhou and Zuo Wang and Kai Xiang and Ge Zhang and Wenhao Huang and Yang Wang and Ke Wang , title =. CoRR , volume =. 2025 , url =. doi:10.48550/ARXIV.2508.07999 , eprinttype =. 2508.07999 ...
-
[111]
Narasimhan , title =
Shunyu Yao and Noah Shinn and Pedram Razavi and Karthik R. Narasimhan , title =. The Thirteenth International Conference on Learning Representations,. 2025 , url =
2025
- [112]
-
[113]
InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated Large Language Model Agents , booktitle =
Qiusi Zhan and Zhixiang Liang and Zifan Ying and Daniel Kang , editor =. InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated Large Language Model Agents , booktitle =. 2024 , url =. doi:10.18653/V1/2024.FINDINGS-ACL.624 , timestamp =
2024 doi
-
[114]
Pei Wang and Yanan Wu and Noah Wang and Jiaheng Liu and Xiaoshuai Song and Z. Y. Peng and Ken Deng and Chenchen Zhang and Jiakai Wang and Junran Peng and Ge Zhang and Hangyu Guo and Zhaoxiang Zhang and Wenbo Su and Bo Zheng , title =. The Thirteenth International Conference on...
2025
-
[115]
CoRR , volume =
Chen Chen and Xinlong Hao and Weiwen Liu and Xu Huang and Xingshan Zeng and Shuai Yu and Dexun Li and Shuai Wang and Weinan Gan and Yuefeng Huang and Wulong Liu and Xinzhi Wang and Defu Lian and Baoqun Yin and Yasheng Wang and Wu Liu , title =. CoRR , volume =. 2025 , url =. d...
2025 doi
-
[116]
Yuchen Zhuang and Yue Yu and Kuan Wang and Haotian Sun and Chao Zhang , editor =. ToolQA:. Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 2023 , year =
2023
- [117]
- [118]
- [119]
- [120]
-
[121]
The Twelfth International Conference on Learning Representations,
Yue Huang and Jiawen Shi and Yuan Li and Chenrui Fan and Siyuan Wu and Qihui Zhang and Yixin Liu and Pan Zhou and Yao Wan and Neil Zhenqiang Gong and Lichao Sun , title =. The Twelfth International Conference on Learning Representations,. 2024 , url =
2024
-
[122]
Gemini 3 Pro Model Card , url=
Google , year=. Gemini 3 Pro Model Card , url=
-
[123]
Gemini 3.1 Pro Model Card , url=
Google , year=. Gemini 3.1 Pro Model Card , url=
-
[124]
GPT-5.1 Instant and GPT-5.1 Thinking System Card Addendum , url=
OpenAI , year=. GPT-5.1 Instant and GPT-5.1 Thinking System Card Addendum , url=
-
[125]
Update to GPT-5 System Card: GPT-5.2 , url=
OpenAI , year=. Update to GPT-5 System Card: GPT-5.2 , url=
-
[126]
GPT-5.5 System Card , url=
OpenAI , year=. GPT-5.5 System Card , url=
-
[127]
System Card: Claude Opus 4.6 , url=
Anthropic , year=. System Card: Claude Opus 4.6 , url=
-
[128]
System Card: Claude Opus 4.8 , url=
Anthropic , year=. System Card: Claude Opus 4.8 , url=
-
[129]
Qwen3.5: Towards Native Multimodal Agents , url=
Qwen Team , year=. Qwen3.5: Towards Native Multimodal Agents , url=
-
[130]
Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity , url=
Bytedance Seed , year=. Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity , url=
-
[131]
Seed2.1 Model Card: Agentic Intelligence for Productivity , url=
-
[132]
OpenAI o3 and o4-mini System Card , url=
OpenAI , year=. OpenAI o3 and o4-mini System Card , url=
-
[133]
Seed1.6 Tech Introduction , url=
ByteDance Seed , year=. Seed1.6 Tech Introduction , url=
-
[134]
Kimi K2.6: Advancing Open-Source Coding , url=
Kimi , year=. Kimi K2.6: Advancing Open-Source Coding , url=
-
[135]
GLM-5.2: Built for Long-Horizon Tasks , url=
-
[136]
Introducing Hy3 , url=
- [137]
- [138]
-
[139]
MarioGPT: Open-Ended Text2Level Generation through Large Language Models , booktitle =
Shyam Sudhakaran and Miguel Gonz. MarioGPT: Open-Ended Text2Level Generation through Large Language Models , booktitle =. 2023 , url =
2023
-
[140]
Forty-second International Conference on Machine Learning,
Qinggang Zhang and Hao Chen and Junnan Dong and Shengyuan Chen and Feiran Huang and Xiao Huang , title =. Forty-second International Conference on Machine Learning,. 2025 , url =
2025
-
[141]
2025 , url =
Gianpaolo Ghiani and Emanuele Manni and Sandro Zacchino , title =. 2025 , url =. doi:10.1109/ACCESS.2025.3601518 , timestamp =
2025
-
[142]
Forty-second International Conference on Machine Learning,
Simin Chen and Pranav Pusarla and Baishakhi Ray , title =. Forty-second International Conference on Machine Learning,. 2025 , url =
2025
-
[143]
Madhavi Kumari and Rohit Chauhan and Rashi Jain and Prabha Garg , title =. Inf. Fusion , volume =. 2026 , url =. doi:10.1016/J.INFFUS.2025.103902 , timestamp =
2026
-
[144]
Find the Intention of Instruction: Comprehensive Evaluation of Instruction Understanding for Large Language Models , booktitle =
Hyeonseok Moon and Jaehyung Seo and Seungyoon Lee and Chanjun Park and Heuiseok Lim , editor =. Find the Intention of Instruction: Comprehensive Evaluation of Instruction Understanding for Large Language Models , booktitle =. 2025 , url =. doi:10.18653/V1/2025.FINDINGS-NAACL.3...
2025 doi
-
[145]
Large Language Models Meet Open-World Intent Discovery and Recognition: An Evaluation of ChatGPT , booktitle =
Xiaoshuai Song and Keqing He and Pei Wang and Guanting Dong and Yutao Mou and Jingang Wang and Yunsen Xian and Xunliang Cai and Weiran Xu , editor =. Large Language Models Meet Open-World Intent Discovery and Recognition: An Evaluation of ChatGPT , booktitle =. 2023 , url =. d...
2023 doi
-
[146]
International Joint Conference on Neural Networks,
Pengfei Yin and Yubing Ren and Fang Fang and Boxiang Hu and Xu Bai , title =. International Joint Conference on Neural Networks,. 2025 , url =. doi:10.1109/IJCNN64981.2025.11227926 , timestamp =
2025
-
[147]
FinDABench: Benchmarking Financial Data Analysis Ability of Large Language Models , booktitle =
Shu Liu and Shangqing Zhao and Chenghao Jia and Xinlin Zhuang and Zhaoguang Long and Jie Zhou and Aimin Zhou and Man Lan and Yang Chong , editor =. FinDABench: Benchmarking Financial Data Analysis Ability of Large Language Models , booktitle =. 2025 , url =
2025
-
[148]
CoRR , volume =
Zijun Liu and Peiyi Wang and Runxin Xu and Shirong Ma and Chong Ruan and Peng Li and Yang Liu and Yu Wu , title =. CoRR , volume =. 2025 , url =. doi:10.48550/ARXIV.2504.02495 , eprinttype =. 2504.02495 , timestamp =
2025 doi
-
[149]
Proceedings of the Twentieth European Conference on Computer Systems, EuroSys 2025, Rotterdam, The Netherlands, 30 March 2025 - 3 April 2025 , pages =
Guangming Sheng and Chi Zhang and Zilingfeng Ye and Xibin Wu and Wang Zhang and Ru Zhang and Yanghua Peng and Haibin Lin and Chuan Wu , title =. Proceedings of the Twentieth European Conference on Computer Systems, EuroSys 2025, Rotterdam, The Netherlands, 30 March 2025 - 3 Ap...
2025
- [150]
-
[151]
Xinyu Pan and Weibin Zhuang and Sijie Wen and Weigang Yu and Jinsong Bao and Xinyu Li , title =. Adv. Eng. Informatics , volume =. 2025 , url =. doi:10.1016/J.AEI.2025.103422 , timestamp =
2025
-
[152]
Analyzing the Effects of Supervised Fine-Tuning on Model Knowledge from Token and Parameter Levels
Ye, Junjie and Yang, Yuming and Nan, Yang and Li, Shuo and Zhang, Qi and Gui, Tao and Huang, Xuanjing and Wang, Peng and Shi, Zhongchao and Fan, Jianping. Analyzing the Effects of Supervised Fine-Tuning on Model Knowledge from Token and Parameter Levels. Proceedings of the 202...
2025 doi
-
[153]
Patil and Qi Qi and Shengyu Feng and Julian Katz
Yun He and Wenzhe Li and Hejia Zhang and Songlin Li and Karishma Mandyam and Sopan Khosla and Yuanhao Xiong and Nanshu Wang and Xiaoliang Peng and Beibin Li and Shengjie Bi and Shishir G. Patil and Qi Qi and Shengyu Feng and Julian Katz. AdvancedIF: Rubric-Based Benchmarking a...
2025 doi
-
[154]
Ross and Katherine L
Hongkai Zheng and Wenda Chu and Bingliang Zhang and Zihui Wu and Austin Wang and Berthy Feng and Caifeng Zou and Yu Sun and Nikola Borislavov Kovachki and Zachary E. Ross and Katherine L. Bouman and Yisong Yue , title =. The Thirteenth International Conference on Learning Repr...
2025
-
[155]
MultiChallenge:
Kaustubh Deshpande and Ved Sirdeshmukh and Johannes Baptist Mols and Lifeng Jin and Ed. MultiChallenge:. Findings of the Association for Computational Linguistics,. 2025 , url =
2025
-
[156]
Lillicrap and Jean
Machel Reid and Nikolay Savinov and Denis Teplyashin and Dmitry Lepikhin and Timothy P. Lillicrap and Jean. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context , journal =. 2024 , url =. doi:10.48550/ARXIV.2403.05530 , eprinttype =. 2403.05530 ,...
- [157]
-
[158]
CoRR , volume =
Weiwen Liu and Xu Huang and Xingshan Zeng and Xinlong Hao and Shuai Yu and Dexun Li and Shuai Wang and Weinan Gan and Zhengying Liu and Yuanqing Yu and Zezhong Wang and Yuxian Wang and Wu Ning and Yutai Hou and Bin Wang and Chuhan Wu and Xinzhi Wang and Yong Liu and Yasheng Wa...
-
[159]
The Twelfth International Conference on Learning Representations,
Yujia Qin and Shihao Liang and Yining Ye and Kunlun Zhu and Lan Yan and Yaxi Lu and Yankai Lin and Xin Cong and Xiangru Tang and Bill Qian and Sihan Zhao and Lauren Hong and Runchu Tian and Ruobing Xie and Jie Zhou and Mark Gerstein and Dahai Li and Zhiyuan Liu and Maosong Sun...
2024
- [160]
-
[161]
Yujia Qin and Shengding Hu and Yankai Lin and Weize Chen and Ning Ding and Ganqu Cui and Zheni Zeng and Xuanhe Zhou and Yufei Huang and Chaojun Xiao and Chi Han and Yi R. Fung and Yusheng Su and Huadong Wang and Cheng Qian and Runchu Tian and Kunlun Zhu and Shihao Liang and Xi...
2025
- [162]
-
[163]
CoRR , volume =
Junjie Ye and Xuanting Chen and Nuo Xu and Can Zu and Zekai Shao and Shichun Liu and Yuhan Cui and Zeyang Zhou and Chao Gong and Yang Shen and Jie Zhou and Siming Chen and Tao Gui and Qi Zhang and Xuanjing Huang , title =. CoRR , volume =. 2023 , url =. doi:10.48550/ARXIV.2303...
- [164]
-
[165]
Jacobs and Michael I
Robert A. Jacobs and Michael I. Jordan and Steven J. Nowlan and Geoffrey E. Hinton , title =. Neural Comput. , volume =. 1991 , url =. doi:10.1162/NECO.1991.3.1.79 , timestamp =
1991 doi
-
[166]
Constitutional
Yuntao Bai and Saurav Kadavath and Sandipan Kundu and Amanda Askell and Jackson Kernion and Andy Jones and Anna Chen and Anna Goldie and Azalia Mirhoseini and Cameron McKinnon and Carol Chen and Catherine Olsson and Christopher Olah and Danny Hernandez and Dawn Drain and Deep ...
-
[167]
Llama 2: Open Foundation and Fine-Tuned Chat Models , journal =
Hugo Touvron and Louis Martin and Kevin Stone and Peter Albert and Amjad Almahairi and Yasmine Babaei and Nikolay Bashlykov and Soumya Batra and Prajjwal Bhargava and Shruti Bhosale and Dan Bikel and Lukas Blecher and Cristian Canton. Llama 2: Open Foundation and Fine-Tuned Ch...
-
[168]
LLaMA: Open and Efficient Foundation Language Models , journal =
Hugo Touvron and Thibaut Lavril and Gautier Izacard and Xavier Martinet and Marie. LLaMA: Open and Efficient Foundation Language Models , journal =. 2023 , url =. doi:10.48550/ARXIV.2302.13971 , eprinttype =. 2302.13971 , timestamp =
- [169]
-
[170]
Tom B. Brown and Benjamin Mann and Nick Ryder and Melanie Subbiah and Jared Kaplan and Prafulla Dhariwal and Arvind Neelakantan and Pranav Shyam and Girish Sastry and Amanda Askell and Sandhini Agarwal and Ariel Herbert. Language Models are Few-Shot Learners , booktitle =. 202...
2020
-
[171]
Aho and Jeffrey D
Alfred V. Aho and Jeffrey D. Ullman , title =. 1972
1972
-
[172]
Publications Manual , year = "1983", publisher =
1983
-
[173]
Chandra and Dexter C
Ashok K. Chandra and Dexter C. Kozen and Larry J. Stockmeyer , year = "1981", title =. doi:10.1145/322234.322243
1981
-
[174]
Zhiheng Xi and Wenxiang Chen and Xin Guo and Wei He and Yiwen Ding and Boyang Hong and Ming Zhang and Junzhe Wang and Senjie Jin and Enyu Zhou and Rui Zheng and Xiaoran Fan and Xiao Wang and Limao Xiong and Yuhao Zhou and Weiran Wang and Changhao Jiang and Yicheng Zou and Xian...
2025
- [175]
-
[176]
CoRR , volume =
John Schulman and Filip Wolski and Prafulla Dhariwal and Alec Radford and Oleg Klimov , title =. CoRR , volume =. 2017 , url =. 1707.06347 , timestamp =
2017 arXiv
-
[177]
CoRR , volume =
Zhiheng Xi and Wenxiang Chen and Boyang Hong and Senjie Jin and Rui Zheng and Wei He and Yiwen Ding and Shichun Liu and Xin Guo and Junzhe Wang and Honglin Guo and Wei Shen and Xiaoran Fan and Yuhao Zhou and Shihan Dou and Xiao Wang and Xinbo Zhang and Peng Sun and Tao Gui and...
- [178]
-
[179]
CoRR , volume =
Jinze Bai and Shuai Bai and Yunfei Chu and Zeyu Cui and Kai Dang and Xiaodong Deng and Yang Fan and Wenbin Ge and Yu Han and Fei Huang and Binyuan Hui and Luo Ji and Mei Li and Junyang Lin and Runji Lin and Dayiheng Liu and Gao Liu and Chengqiang Lu and Keming Lu and Jianxin M...
-
[180]
Self-Distillation Bridges Distribution Gap in Language Model Fine-Tuning , booktitle =
Zhaorui Yang and Tianyu Pang and Haozhe Feng and Han Wang and Wei Chen and Minfeng Zhu and Qian Liu , editor =. Self-Distillation Bridges Distribution Gap in Language Model Fine-Tuning , booktitle =. 2024 , url =. doi:10.18653/V1/2024.ACL-LONG.58 , timestamp =
2024 doi
-
[181]
and Deepali Aneja and Zeyu Jin and Ramani Duraiswami and Dinesh Manocha , title =
Sreyan Ghosh and Chandra Kiran Reddy Evuru and Sonal Kumar and Ramaneswaran S. and Deepali Aneja and Zeyu Jin and Ramani Duraiswami and Dinesh Manocha , title =. Forty-first International Conference on Machine Learning,. 2024 , url =
2024
-
[182]
9th International Conference on Learning Representations,
Dan Hendrycks and Collin Burns and Steven Basart and Andy Zou and Mantas Mazeika and Dawn Song and Jacob Steinhardt , title =. 9th International Conference on Learning Representations,. 2021 , url =
2021
-
[183]
CoRR , volume =
Karl Cobbe and Vineet Kosaraju and Mohammad Bavarian and Mark Chen and Heewoo Jun and Lukasz Kaiser and Matthias Plappert and Jerry Tworek and Jacob Hilton and Reiichiro Nakano and Christopher Hesse and John Schulman , title =. CoRR , volume =. 2021 , url =. 2110.14168 , timestamp =
2021 arXiv
-
[184]
Evaluating Large Language Models Trained on Code , journal =
Mark Chen and Jerry Tworek and Heewoo Jun and Qiming Yuan and Henrique Pond. Evaluating Large Language Models Trained on Code , journal =. 2021 , url =. 2107.03374 , timestamp =
2021 arXiv
-
[185]
ToolSword: Unveiling Safety Issues of Large Language Models in Tool Learning Across Three Stages , booktitle =
Junjie Ye and Sixian Li and Guanyu Li and Caishuang Huang and Songyang Gao and Yilong Wu and Qi Zhang and Tao Gui and Xuanjing Huang , editor =. ToolSword: Unveiling Safety Issues of Large Language Models in Tool Learning Across Three Stages , booktitle =. 2024 , url =. doi:10...
2024 doi
-
[186]
Knowledge Mechanisms in Large Language Models:
Mengru Wang and Yunzhi Yao and Ziwen Xu and Shuofei Qiao and Shumin Deng and Peng Wang and Xiang Chen and Jia. Knowledge Mechanisms in Large Language Models:. CoRR , volume =. 2024 , url =. doi:10.48550/ARXIV.2407.15017 , eprinttype =. 2407.15017 , timestamp =
-
[187]
Learning or Self-aligning? Rethinking Instruction Fine-tuning , booktitle =
Mengjie Ren and Boxi Cao and Hongyu Lin and Cao Liu and Xianpei Han and Ke Zeng and Guanglu Wan and Xunliang Cai and Le Sun , editor =. Learning or Self-aligning? Rethinking Instruction Fine-tuning , booktitle =. 2024 , url =. doi:10.18653/V1/2024.ACL-LONG.330 , timestamp =
2024 doi
-
[188]
Self-Refine: Iterative Refinement with Self-Feedback , booktitle =
Aman Madaan and Niket Tandon and Prakhar Gupta and Skyler Hallinan and Luyu Gao and Sarah Wiegreffe and Uri Alon and Nouha Dziri and Shrimai Prabhumoye and Yiming Yang and Shashank Gupta and Bodhisattwa Prasad Majumder and Katherine Hermann and Sean Welleck and Amir Yazdanbakh...
2023
-
[189]
Reflexion: language agents with verbal reinforcement learning , booktitle =
Noah Shinn and Federico Cassano and Ashwin Gopinath and Karthik Narasimhan and Shunyu Yao , editor =. Reflexion: language agents with verbal reinforcement learning , booktitle =. 2023 , url =
2023
- [190]
-
[191]
7th International Conference on Learning Representations,
Ilya Loshchilov and Frank Hutter , title =. 7th International Conference on Learning Representations,. 2019 , url =
2019
-
[192]
CoRR , volume =
An Yang and Baosong Yang and Binyuan Hui and Bo Zheng and Bowen Yu and Chang Zhou and Chengpeng Li and Chengyuan Li and Dayiheng Liu and Fei Huang and Guanting Dong and Haoran Wei and Huan Lin and Jialong Tang and Jialin Wang and Jian Yang and Jianhong Tu and Jianwei Zhang and...
-
[193]
CoRR , volume =
Aohan Zeng and Bin Xu and Bowen Wang and Chenhui Zhang and Da Yin and Diego Rojas and Guanyu Feng and Hanlin Zhao and Hanyu Lai and Hao Yu and Hongning Wang and Jiadai Sun and Jiajie Zhang and Jiale Cheng and Jiayi Gui and Jie Tang and Jing Zhang and Juanzi Li and Lei Zhao and...
-
[194]
RoTBench:
Junjie Ye and Yilong Wu and Songyang Gao and Caishuang Huang and Sixian Li and Guanyu Li and Xiaoran Fan and Qi Zhang and Tao Gui and Xuanjing Huang , editor =. RoTBench:. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing,. 2024 , url =
2024
-
[195]
Rossi and Somdeb Sarkhel and Chao Zhang , title =
Yuchen Zhuang and Xiang Chen and Tong Yu and Saayan Mitra and Victor Bursztyn and Ryan A. Rossi and Somdeb Sarkhel and Chao Zhang , title =. CoRR , volume =. 2023 , url =. doi:10.48550/ARXIV.2310.13227 , eprinttype =. 2310.13227 , timestamp =
-
[196]
CoRR , volume =
Zhaoyang Liu and Zeqiang Lai and Zhangwei Gao and Erfei Cui and Zhiheng Li and Xizhou Zhu and Lewei Lu and Qifeng Chen and Yu Qiao and Jifeng Dai and Wenhai Wang , title =. CoRR , volume =. 2023 , url =. doi:10.48550/ARXIV.2310.17796 , eprinttype =. 2310.17796 , timestamp =
-
[197]
ToolEyes: Fine-Grained Evaluation for Tool Learning Capabilities of Large Language Models in Real-world Scenarios , booktitle =
Junjie Ye and Guanyu Li and Songyang Gao and Caishuang Huang and Yilong Wu and Sixian Li and Xiaoran Fan and Shihan Dou and Tao Ji and Qi Zhang and Tao Gui and Xuanjing Huang , editor =. ToolEyes: Fine-Grained Evaluation for Tool Learning Capabilities of Large Language Models ...
2025
-
[198]
Scalable training of
Andrew, Galen and Gao, Jianfeng , booktitle=. Scalable training of
-
[199]
Dan Gusfield , title =. 1997
1997
-
[200]
Tetreault , title =
Mohammad Sadegh Rasooli and Joel R. Tetreault , title =. Computing Research Repository , volume =. 2015 , url =
2015
-
[201]
A Framework for Learning Predictive Structures from Multiple Tasks and Unlabeled Data , Volume =
Ando, Rie Kubota and Zhang, Tong , Issn =. A Framework for Learning Predictive Structures from Multiple Tasks and Unlabeled Data , Volume =. Journal of Machine Learning Research , Month = dec, Numpages =
-
[202]
ToolSandbox:
Jiarui Lu and Thomas Holleis and Yizhe Zhang and Bernhard Aumayer and Feng Nan and Haoping Bai and Shuang Ma and Shen Ma and Mengyu Li and Guoli Yin and Zirui Wang and Ruoming Pang , editor =. ToolSandbox:. Findings of the Association for Computational Linguistics:. 2025 , url...
2025 doi
-
[203]
Mobile-Bench: An Evaluation Benchmark for LLM-based Mobile Agents , booktitle =
Shihan Deng and Weikai Xu and Hongda Sun and Wei Liu and Tao Tan and Jianfeng Liu and Ang Li and Jian Luan and Bin Wang and Rui Yan and Shuo Shang , editor =. Mobile-Bench: An Evaluation Benchmark for LLM-based Mobile Agents , booktitle =. 2024 , url =. doi:10.18653/V1/2024.AC...
2024 doi
-
[204]
In Prospect and Retrospect: Reflective Memory Management for Long-term Personalized Dialogue Agents , booktitle =
Zhen Tan and Jun Yan and I. In Prospect and Retrospect: Reflective Memory Management for Long-term Personalized Dialogue Agents , booktitle =. 2025 , url =. doi:10.18653/V1/2025.ACL-LONG.413 , timestamp =
2025 doi
-
[205]
Enabling Personalized Long-term Interactions in LLM-based Agents through Persistent Memory and User Profiles , journal =
Rebecca Westh. Enabling Personalized Long-term Interactions in LLM-based Agents through Persistent Memory and User Profiles , journal =. 2025 , url =. doi:10.48550/ARXIV.2510.07925 , eprinttype =. 2510.07925 , timestamp =
2025 doi
-
[206]
Zollo and Andrew Wei Tung Siah and Naimeng Ye and Ang Li and Hongseok Namkoong , title =
Thomas P. Zollo and Andrew Wei Tung Siah and Naimeng Ye and Ang Li and Hongseok Namkoong , title =. The Thirteenth International Conference on Learning Representations,. 2025 , url =
2025
-
[207]
StableToolBench: Towards Stable Large-Scale Benchmarking on Tool Learning of Large Language Models , booktitle =
Zhicheng Guo and Sijie Cheng and Hao Wang and Shihao Liang and Yujia Qin and Peng Li and Zhiyuan Liu and Maosong Sun and Yang Liu , editor =. StableToolBench: Towards Stable Large-Scale Benchmarking on Tool Learning of Large Language Models , booktitle =. 2024 , url =. doi:10....
2024 doi
-
[208]
Patil and Huanzhi Mao and Fanjia Yan and Charlie Cheng
Shishir G. Patil and Huanzhi Mao and Fanjia Yan and Charlie Cheng. The Berkeley Function Calling Leaderboard. Forty-second International Conference on Machine Learning,. 2025 , url =
2025
-
[209]
Jize Wang and Zerun Ma and Yining Li and Songyang Zhang and Cailian Chen and Kai Chen and Xinyi Le , editor =. Advances in Neural Information Processing Systems 37: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024, Vancouver, BC, Canada, December 1...
2024
-
[210]
Cohen and Emine Yilmaz , editor =
Zheng Zhao and Clara Vania and Subhradeep Kayal and Naila Khan and Shay B. Cohen and Emine Yilmaz , editor =. PersonaLens:. Findings of the Association for Computational Linguistics,. 2025 , url =. doi:10.18653/V1/2025.FINDINGS-ACL.927 , timestamp =
2025 doi
-
[211]
Mem-PAL: Towards Memory-based Personalized Dialogue Assistants for Long-term User-Agent Interaction , booktitle =
Zhaopei Huang and Qifeng Dai and Guozheng Wu and Xiaopeng Wu and Xubin Li and Tiezheng Ge and Wenxuan Wang and Qin Jin , editor =. Mem-PAL: Towards Memory-based Personalized Dialogue Assistants for Long-term User-Agent Interaction , booktitle =. 2026 , url =. doi:10.1609/AAAI....
2026 doi
- [212]
-
[213]
Personalized Safety in LLMs:
Yuchen Wu and Edward Sun and Kaijie Zhu and Jianxun Lian and Jos. Personalized Safety in LLMs:. CoRR , volume =. 2025 , url =. doi:10.48550/ARXIV.2505.18882 , eprinttype =. 2505.18882 , timestamp =
2025 doi
-
[214]
CoRR , volume =
Jiazheng Sun and Te Yang and Jiayang Niu and Mingxuan Li and Yongyong Lu and Ruimeng Yang and Xin Peng , title =. CoRR , volume =. 2025 , url =. doi:10.48550/ARXIV.2509.20729 , eprinttype =. 2509.20729 , timestamp =
2025 doi
- [215]
-
[216]
Mobile-Agent-v2: Mobile Device Operation Assistant with Effective Navigation via Multi-Agent Collaboration , booktitle =
Junyang Wang and Haiyang Xu and Haitao Jia and Xi Zhang and Ming Yan and Weizhou Shen and Ji Zhang and Fei Huang and Jitao Sang , editor =. Mobile-Agent-v2: Mobile Device Operation Assistant with Effective Navigation via Multi-Agent Collaboration , booktitle =. 2024 , url =
2024
- [217]
-
[218]
AutoDroid: LLM-powered Task Automation in Android , booktitle =
Hao Wen and Yuanchun Li and Guohong Liu and Shanhui Zhao and Tao Yu and Toby Jia. AutoDroid: LLM-powered Task Automation in Android , booktitle =. 2024 , url =. doi:10.1145/3636534.3649379 , timestamp =
2024
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.