Pith. sign in

REVIEW 3 major objections 4 minor 218 references

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information

T0 review · 3 major / 4 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read This paper argues that the main obstacle for LLM-based mobile assistants is not tool use but knowing which piece of scattered personal information to retrieve, and it supports this claim with a new 250-task benchmark.

desk verdict Solid new benchmark for scattered personal information; the 'commit-too-early' failure story is not actually pinned down by the evidence. read the letter →

arxiv 2608.10692 v1 pith:ENCFUWTB submitted 2026-08-11 cs.CL cs.AI

classification cs.CLcs.AI
keywords SPIEvalmobileassistantslargelanguagemodelsscatteredpersonalinformationbenchmarklocalizationtoolusecognitivecapabilities
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper introduces SPIEval, a benchmark that asks whether LLM-based mobile assistants can complete instructions by finding personal information scattered across multiple phone apps. It builds 250 tasks over 4,335 fictional records in 10 apps, each requiring one of five cognitive capabilities: reasoning, disambiguation, integration, preference inference, and multi-intent decomposition. Nine LLMs are evaluated through 21 tools. The paper's central empirical claim is that current assistants are far from reliable: the strongest model reaches only 57.3% accuracy, and the main failure is not tool selection but committing to plausible-but-wrong information instead of continuing to search. It argues that information localization—formulating search queries that identify the correct record—is the primary bottleneck.

What carries the argument

The central object is the SPIEval benchmark itself: a fully controllable simulated phone environment with 10 app schemas, 4,335 linked personal records, and 21 tools (11 retrieval, 10 execution), plus a shared user profile so instructions like "Call my manager" require resolving identity across apps. Each task has a manually constructed gold answer with exact execution calls and parameters, verified by cross-validation, so outcomes are determined by matching final tool calls, not by a judge model. Two controlled ablations, "full" (unpaginated results) and "no-search" (records in the prompt), isolate the retrieval bottleneck.

What would settle it

Compare accuracy under the standard protocol vs. the no-search protocol while systematically varying query specificity, or run the same tasks on a real phone using a GUI agent harness. If injecting progressively more specific hints does not push accuracy toward the human ceiling, or if models that solve SPIEval tasks fail real-device tasks with the same structure, then the localization-bottleneck claim would be weakened and other bottlenecks (planning, decomposition, interface understanding) would need to carry the explanation.

Watch

Extended reading notes

Core claim

According to the paper, state-of-the-art LLMs fail most mobile-assistant tasks that require scattered personal information, and the failure is concentrated in retrieval. Removing retrieval and handing all relevant records to the model raises average accuracy from 35.5% to 66.8%, while simply returning all matches unpaginated raises it only to 36.0%, implying that the bottleneck is the query, not the search results. Across the three strongest models, 79% of errors are wrong parameter values, and models perform fewer retrievals on failed tasks than on successful ones, suggesting early commitment to plausible but incorrect records. The paper also finds that fewer than 2% of 126,279 retrieval calls use advanced modes and that preference inference and multi-intent decomposition are the hardest capabilities.

Load-bearing premise

The paper's conclusions about real mobile assistants assume that performing well on its simulated, fictional phone with synthetic records tells us how models would behave on real phones with real personal data and real interfaces, since the benchmark is not validated against actual device use.

Editorial extensions

If this is right

  • If information localization is the bottleneck, then improving LLMs' query reformulation, field-targeted search, and verification of whether evidence is sufficient should yield larger gains than adding more tools or longer contexts.
  • The paper's no-search result implies models can reason about preferences once evidence is present; the hard part is knowing when and where to look.
  • Multi-intent decomposition lags even with all records available, so progress there requires better task decomposition and planning, not just better retrieval.
  • The underuse of regex and fuzzy matching suggests a concrete training or prompting target: teaching models to exploit structured retrieval modes, including field-specific and fuzzy search.
  • With the best model below 60%, the paper's numbers imply that current assistants cannot be trusted to act on ambiguous personal-data requests without verification.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Inference: If the simulated environment transfers, real-world assistants need a "verify before acting" mechanism—when search returns a plausible match among multiple candidates, the assistant should gather corroborating evidence before calling execution tools.
  • Inference: The benchmark's finding about search efficiency suggests a useful standardized metric—task accuracy normalized by number of retrievals—that would separate models that search smartly from models that search a lot; the paper reports the underlying numbers but does not define such a metric.
  • Inference: A testable extension would be to vary the surface form of records (typos, abbreviations, partial names) while keeping gold answers fixed; if accuracy drops sharply, that would confirm that over-reliance on exact substring matching is a measurable failure mode, not just a stylistic preference.
  • Inference: Because all personal data is synthetic, the benchmark is safe to distribute but cannot capture privacy-sensitive behaviors like refusing to expose or misuse personal data; that remains a separate evaluation axis.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper introduces SPIEval, a human-curated benchmark for evaluating large language models as mobile assistants that must gather scattered personal information across multiple apps. The benchmark contains 250 tasks covering five cognitive capabilities, 4,335 synthetic personal records across 10 simulated apps, and 21 tools for retrieval and execution. The authors evaluate nine LLMs under high- and low-reasoning settings and report that the best model, GPT-5.5 (xhigh), achieves only 57.3% accuracy, while the weakest reaches 16.4%. They further report that 79% of failures arise from inaccurate information localization, that models often commit to plausible but incorrect information instead of continuing retrieval, that fewer than 2% of retrieval actions use advanced search modes, and that successful tasks involve more retrieval calls than failed ones.

Significance. If the benchmark is valid, SPIEval fills a real gap: prior mobile-assistant benchmarks either provide all necessary information in the instruction or restrict personal data to single-document retrieval. The construction process is rigorous and clearly described: human annotators create tasks, independent annotators cross-validate without access to gold answers, gold answers are executed against the tool implementation, and the environment is deterministic with verifiable outcomes. The evaluation protocol is also solid, with three runs per setting, reported standard deviations, a paginated retrieval condition, and controlled 'full' and 'no-search' conditions that isolate the role of retrieval. The headline finding—that top LLMs solve only about half of these tasks and that performance improves substantially when all relevant records are provided—is a useful, falsifiable result. However, the paper's more specific causal claim about the dominant failure mechanism, namely that models prematurely commit to plausible but wrong information rather than continuing to search, is not supported by the current analyses and needs additional evidence or careful rewording.

major comments (3)
  1. [Section 5 (Figure 6) and Section 6 (Figure 7)] The central conclusion that 79% of failures stem from 'inaccurate information localization' and that LLMs 'commit to plausible but incorrect information instead of continuing retrieval for verification' is an inference, not a demonstrated mechanism. The error-type taxonomy in Figure 6 lumps together at least three distinct failure processes: (a) the gold record never appeared in any retrieved result because the model's query was insufficient; (b) the gold record was retrieved but the model acted on a distractor; and (c) the correct record was retrieved and selected but a parameter value was misread or inferred incorrectly. Figure 7's finding that failed tasks involve fewer retrievals than solved tasks is consistent with premature stopping, but it is equally consistent with the model failing to formulate an effective query and ceasing search after fruitless attempts. The no-search condition in Figure 5 removes all retrieval-related causes at once, so it cannot disentangle query formulation failure from premature commitment. Please add a trajectory-level analysis that records, for each failed task, whether the gold record appeared in any retrieved context; this would separate 'never retrieved' from 'retrieved but ignored.' If such an analysis is not feasible, the abstract and Section 6 should be reworded to present premature commitment as one plausible hypothesis rather than the established bottleneck.
  2. [Abstract and Section 5] The '79% of failures' statistic is presented without a qualifier in the abstract and conclusion, but the Figure 6 analysis is based on only the three strongest models (GPT-5.5 xhigh, Gemini 3.1 Pro high, Claude Opus 4.8 max), as stated in the Figure 6 caption and Section 5. If the claim is meant to apply to all nine evaluated models, the failure-mode analysis should be extended to the full set; otherwise, the abstract should specify that the 79% figure is for the three best models. As written, the abstract overstates the generality of the result.
  3. [Sections 1, 5, and 7] The paper repeatedly generalizes to 'current LLM-based mobile assistants' and 'mobile assistants,' but the evaluation environment is a simulated, text-based tool API with no GUI, no device-specific latencies, and synthetic personal records. No evidence is provided that performance in this environment transfers to real mobile assistants. Please add an explicit limitations paragraph and qualify statements such as 'the primary bottleneck for current mobile assistants' (Section 5) and 'fundamental limitations of current LLM-based mobile assistants' (abstract), or provide a small real-device validation study to support the transfer.
minor comments (4)
  1. [Section 4 and Appendices A/C] The paper does not state the language(s) in which the user instructions and system prompts were administered. The system prompt template is shown in both Chinese and English, and the tool schemas are in Chinese. Please specify the evaluation language and any translation/QA procedures for the English version, since accuracy may be language-dependent.
  2. [Section 3.4, Figure 3] The statement that 'one instruction consists of only 20 characters yet requires 45 execution parameters' should clarify the metric (characters in which language? including spaces?) to be reproducible.
  3. [Section 5, Figure 7] The differences in average retrieval counts between solved and failed tasks are reported for each model but without any significance testing or confidence intervals; please add statistical tests or confidence intervals to support the claim of a systematic pattern.
  4. [Appendix D, Table 5] The binary 'Human-Curated' column used to compare with existing benchmarks is not defined in the text; a one-sentence definition and references for the negative entries would help the reader interpret the comparison.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: benchmark performance is an external measurement against human-annotated gold answers, and the bottleneck analysis rests on controlled protocol ablations rather than fitted or self-referential claims.

full rationale

SPIEval's construction chain is self-contained: tasks are manually authored, gold answers are human-annotated and cross-validated (Section 3.3), and models are scored by exact match of final execution tool calls to those gold answers (Section 4). None of the reported accuracy numbers is fitted, and no parameter is inferred from the data it later explains. The central analytical claim, that retrieval/localization is the primary bottleneck, is supported by independent protocol manipulations: removing retrieval entirely raises average accuracy from 35.5% to 66.8%, while de-pagination alone gives only 35.5% to 36.0% (Section 5, Figure 5). This is an A/B comparison over the same models and tasks, not a definitional equivalence. The '79% of failures' figure is a coded error-type distribution (Figure 6) over actual trajectories, and the 'commit to plausible but incorrect information' mechanism is an interpretation of that distribution together with retrieval-count comparisons (Figure 7); even if that mechanistic inference is weaker than the authors claim, it is an evidentiary gap rather than a circular derivation. The only self-citations (e.g., Ye et al. 2023; Ye et al. 2025; Tencent Hunyuan Team 2026) are background references for instruction-following, tool-use, and the Hy3 model card; none is load-bearing for the benchmark's validity or for the main results.

Assumptions & free parameters 0 free parameters · 5 assumptions · 0 invented entities

No free parameters are fitted in this paper; it is a benchmark and empirical evaluation, not a parametric model or derivation. The key assumptions are domain assumptions about the representativeness of the simulated environment, the completeness of the constructed tasks, and the validity of the accuracy metric. No new scientific entities (particles, forces, conserved quantities) are introduced; the simulated apps and tools are measurement apparatus, not postulated physical entities.

assumptions (5)
  • domain assumption The simulated apps and structured schemas are faithful abstractions of real mobile applications.
    Introduced in Section 3.3 (Application Construction), this assumption underlies the benchmark's external validity for measuring real mobile assistant capability.
  • domain assumption The five cognitive capabilities (reasoning, disambiguation, integration, preference inference, multi-intent decomposition) are the essential capabilities for completing tasks over scattered personal information.
    The benchmark is organized around this taxonomy (Section 3.2); if other capabilities dominate real usage, the benchmark may miss them.
  • domain assumption The cross-validation protocol ensures that every task has a unique and verifiable gold answer and that human accuracy is 100%.
    Section 3.3 Quality Control; the soundness of using exact-match accuracy as the evaluation metric depends on this.
  • domain assumption All information needed to complete a task is available in the constructed records and can be retrieved through the provided tools.
    Section 3.3 and the system prompt (Appendix A) both state that all required information is retrievable; this is a design guarantee that must hold for every task.
  • domain assumption LLM performance in the simulated environment predicts performance on real mobile assistants.
    This assumption is necessary for the paper's significance claim but is not tested. It is implicit in Sections 1 and 5, where the authors generalize to 'current LLM-based mobile assistants.'

how reviews work

0 comments
Cite this review

Pith. "Pith review of SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information." pith.science (2026). https://pith.science/paper/ENCFUWTB

@misc{pith2026260810692,
  author       = {Pith},
  title        = {Pith review of: SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/ENCFUWTB}},
  note         = {Machine review of arXiv:2608.10692}
}
read the original abstract

Large language models (LLMs) are increasingly deployed as mobile assistants, where a key challenge is leveraging personal information scattered across multiple applications (apps) to complete user instructions. However, due to the lack of dedicated benchmarks, their capabilities remain poorly understood. To address this gap, we introduce SPIEval, a human-curated benchmark grounded in five cognitive capabilities (i.e., reasoning, disambiguation, integration, preference inference, and multi-intent decomposition). SPIEval comprises 250 tasks spanning 4,335 personal records distributed across 10 apps and supports multi-turn interaction through 21 tools. Analysis shows that the benchmark exhibits diverse scenarios, challenging tasks, scattered information, controllable environments, and verifiable outcomes. We evaluate nine representative LLMs and find substantial room for improvement. The best-performing model, GPT-5.5 (xhigh), achieves only 57.3% accuracy, while the weakest achieves just 16.4%. Further analysis reveals that 79% of failures stem from inaccurate information localization, as LLMs often commit to plausible but incorrect information instead of continuing retrieval for verification. We also find that fewer than 2% of retrieval actions employ advanced search methods and observe substantial variation in search efficiency across models. These findings expose fundamental limitations of current LLM-based mobile assistants and motivate future research in this direction. Data and code are available at https://huggingface.co/datasets/Junjie-Ye/SPIEval.

Figures

Figures reproduced from arXiv: 2608.10692 by the authors.

Figure 1
Figure 1. An example of a mobile assistant completing a user instruction by leveraging personal [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Framework of SPIEVAL. Bottom: SPIEVAL comprises 10 commonly used apps, together with 10 execution tools and 11 retrieval tools for accessing records across these apps. All records are associated with a unified user profile. Middle: Each task is designed to evaluate a specific cognitive capability. To ensure quality, the user instruction, personal records, reasoning process, and gold answer are manually constructed a… view at source ↗
Figure 3
Figure 3. Relationship between instruction length and the total number of execution parameters required for each task, with marginal distributions shown alongside. 0 20 40 60 80 100 10 13 16 19 22 25 28 31 34 37 40 mean 17.3 records 0 10 20 30 40 50 2 3 4 5 6 7 8 9 10 mean 6.2 apps [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figures from the paper (5 more)
Figure 5
Figure 5. Figure 5: Accuracy under the standard pag￾inated protocol and two controlled settings, averaged over all model configurations. Wrong value 79% Wrong count 13% Missing param 6% Others 2% [PITH_FULL_IMAGE:figures/full_fig_p008_5.png]
Figure 7
Figure 7. Figure 7: Average number of retrievals on solved versus failed tasks, reported for each model under its highest reasoning effort. Retrieve mode Retrieve scope Substring 98.5% Regex 0.9% Fuzzy 0.6% Full-text 90.5% Field-specific 9.5% [PITH_FULL_IMAGE:figures/full_fig_p009_7.png]
Figure 9
Figure 9. Figure 9: System prompt template used for experiments. [PITH_FULL_IMAGE:figures/full_fig_p016_9.png]
Figure 10
Figure 10. Figure 10: Specification documents for all retrieval tools. [PITH_FULL_IMAGE:figures/full_fig_p027_10.png]
Figure 11
Figure 11. Figure 11: Specification documents for all execution tools. [PITH_FULL_IMAGE:figures/full_fig_p037_11.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

218 extracted references · 13 canonical work pages

  1. [1]

    Langley , title =

    P. Langley , title =. Proceedings of the 17th International Conference on Machine Learning (ICML 2000) , address =. 2000 , pages =

  2. [2]

    CoRR , volume =

    Xiao Liu and Bo Qin and Dongzhu Liang and Guang Dong and Hanyu Lai and Hanchen Zhang and Hanlin Zhao and Iat Long Iong and Jiadai Sun and Jiaqi Wang and Junjie Gao and Junjun Shan and Kangning Liu and Shudan Zhang and Shuntian Yao and Siyi Cheng and Wentao Yao and Wenyi Zhao and Xinghan Liu and Xinyi Liu and Xinying Chen and Xinyue Yang and Yang Yang and ...

  3. [3]

    Exploring the Potential of LLMs as Personalized Assistants: Dataset, Evaluation, and Analysis , booktitle =

    Jisoo Mok and Ik. Exploring the Potential of LLMs as Personalized Assistants: Dataset, Evaluation, and Analysis , booktitle =. 2025 , url =. doi:10.18653/V1/2025.ACL-LONG.504 , timestamp =

  4. [4]

    Juntao Tan and Liangwei Yang and Zuxin Liu and Zhiwei Liu and Rithesh R. N. and Tulika Manoj Awalgaonkar and Jianguo Zhang and Weiran Yao and Ming Zhu and Shirley Kokane and Silvio Savarese and Huan Wang and Caiming Xiong and Shelby Heinecke , editor =. PersonaBench: Evaluating. Findings of the Association for Computational Linguistics,. 2025 , url =. doi...

  5. [5]

    Gaia2: Benchmarking

    Romain Froger and Pierre Andrews and Matteo Bettini and Amar Budhiraja and Ricardo Silveira Cabral and Virginie Do and Emilien Garreau and Jean. Gaia2: Benchmarking. CoRR , volume =. 2026 , url =. doi:10.48550/ARXIV.2602.11964 , eprinttype =. 2602.11964 , timestamp =

  6. [6]

    The Thirteenth International Conference on Learning Representations,

    Jingxuan Chen and Derek Yuen and Bin Xie and Yuhao Yang and Gongwei Chen and Zhihao Wu and Li Yixing and Xurui Zhou and Weiwen Liu and Shuai Wang and Kaiwen Zhou and Rui Shao and Liqiang Nie and Yasheng Wang and Jianye Hao and Jun Wang and Kun Shao , title =. The Thirteenth International Conference on Learning Representations,. 2025 , url =

  7. [7]

    AndroidLab: Training and Systematic Benchmarking of Android Autonomous Agents , booktitle =

    Yifan Xu and Xiao Liu and Xueqiao Sun and Siyi Cheng and Hao Yu and Hanyu Lai and Shudan Zhang and Dan Zhang and Jie Tang and Yuxiao Dong , editor =. AndroidLab: Training and Systematic Benchmarking of Android Autonomous Agents , booktitle =. 2025 , url =. doi:10.18653/V1/2025.ACL-LONG.107 , timestamp =

  8. [8]

    OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments , booktitle =

    Tianbao Xie and Danyang Zhang and Jixuan Chen and Xiaochuan Li and Siheng Zhao and Ruisheng Cao and Toh Jing Hua and Zhoujun Cheng and Dongchan Shin and Fangyu Lei and Yitao Liu and Yiheng Xu and Shuyan Zhou and Silvio Savarese and Caiming Xiong and Victor Zhong and Tao Yu , editor =. OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Co...

Show all 218 references
  1. [9]

    Mind the Third Eye! Benchmarking Privacy Awareness in MLLM-powered Smartphone Agents , booktitle =

    Zhixin Lin and Jungang Li and Shidong Pan and Yibo Shi and Yue Yao and Dongliang Xu , editor =. Mind the Third Eye! Benchmarking Privacy Awareness in MLLM-powered Smartphone Agents , booktitle =. 2026 , url =. doi:10.1609/AAAI.V40I42.40874 , timestamp =

  2. [10]

    Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers),

    Xueyu Hu and Tao Xiong and Biao Yi and Zishu Wei and Ruixuan Xiao and Yurun Chen and Jiasheng Ye and Meiling Tao and Xiangxin Zhou and Ziyu Zhao and Yuhuai Li and Shengze Xu and Shenzhi Wang and Xinchen Xu and Shuofei Qiao and Zhaokai Wang and Kun Kuang and Tieyong Zeng and Li...

  3. [11]

    Guangyi Liu and Pengxiang Zhao and Yaozhen Liang and Liang Liu and Yaxuan Guo and Han Xiao and Weifeng Lin and Yuxiang Chai and Yue Han and Shuai Ren and Hao Wang and Xiaoyu Liang and WenHao Wang and Tianze Wu and Zhengxi Lu and Siheng Chen and LiLinghao and Hao Wang and Guanj...

  4. [12]

    AppAgent: Multimodal Agents as Smartphone Users , booktitle =

    Chi Zhang and Zhao Yang and Jiaxuan Liu and Yanda Li and Yucheng Han and Xin Chen and Zebiao Huang and Bin Fu and Gang Yu , editor =. AppAgent: Multimodal Agents as Smartphone Users , booktitle =. 2025 , url =. doi:10.1145/3706598.3713600 , timestamp =

  5. [13]

    DeepSeek-V3 Technical Report , journal =

    DeepSeek. DeepSeek-V3 Technical Report , journal =. 2024 , url =. doi:10.48550/ARXIV.2412.19437 , eprinttype =. 2412.19437 , timestamp =

  6. [14]

    CoRR , volume =

    DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence , author=. CoRR , volume =. 2026 , url =. doi:10.48550/ARXIV.2606.19348 , eprinttype =

  7. [15]

    CoRR , volume =

    Yunjia Qi and Hao Peng and Xiaozhi Wang and Amy Xin and Youfeng Liu and Bin Xu and Lei Hou and Juanzi Li , title =. CoRR , volume =. 2025 , url =. doi:10.48550/ARXIV.2505.16944 , eprinttype =. 2505.16944 , timestamp =

  8. [16]

    From Complex to Simple: Enhancing Multi-Constraint Complex Instruction Following Ability of Large Language Models , booktitle =

    Qianyu He and Jie Zeng and Qianxi He and Jiaqing Liang and Yanghua Xiao , editor =. From Complex to Simple: Enhancing Multi-Constraint Complex Instruction Following Ability of Large Language Models , booktitle =. 2024 , url =. doi:10.18653/V1/2024.FINDINGS-EMNLP.637 , timestamp =

  9. [17]

    The Thirteenth International Conference on Learning Representations,

    Guanting Dong and Keming Lu and Chengpeng Li and Tingyu Xia and Bowen Yu and Chang Zhou and Jingren Zhou , title =. The Thirteenth International Conference on Learning Representations,. 2025 , url =

  10. [18]

    The Twelfth International Conference on Learning Representations,

    Bill Yuchen Lin and Abhilasha Ravichander and Ximing Lu and Nouha Dziri and Melanie Sclar and Khyathi Raghavi Chandu and Chandra Bhagavatula and Yejin Choi , title =. The Twelfth International Conference on Learning Representations,. 2024 , url =

  11. [19]

    The Thirteenth International Conference on Learning Representations,

    Hao Zhao and Maksym Andriushchenko and Francesco Croce and Nicolas Flammarion , title =. The Thirteenth International Conference on Learning Representations,. 2025 , url =

  12. [20]

    FanOutQA:

    Andrew Zhu and Alyssa Hwang and Liam Dugan and Chris Callison. FanOutQA:. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers),. 2024 , url =. doi:10.18653/V1/2024.ACL-SHORT.2 , timestamp =

  13. [21]

    Evaluating Design Decisions for Dual Encoder-based Entity Disambiguation , booktitle =

    Susanna R. Evaluating Design Decisions for Dual Encoder-based Entity Disambiguation , booktitle =. 2025 , url =. doi:10.18653/V1/2025.ACL-LONG.764 , timestamp =

  14. [22]

    Cohen and Ruslan Salakhutdinov and Christopher D

    Zhilin Yang and Peng Qi and Saizheng Zhang and Yoshua Bengio and William W. Cohen and Ruslan Salakhutdinov and Christopher D. Manning , editor =. HotpotQA:. Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Brussels, Belgium, October 31 - ...

  15. [23]

    CoRR , volume =

    Valentina Pyatkin and Saumya Malik and Victoria Graf and Hamish Ivison and Shengyi Huang and Pradeep Dasigi and Nathan Lambert and Hannaneh Hajishirzi , title =. CoRR , volume =. 2025 , url =. doi:10.48550/ARXIV.2507.02833 , eprinttype =. 2507.02833 , timestamp =

  16. [24]

    On the Multi-turn Instruction Following for Conversational Web Agents , booktitle =

    Yang Deng and Xuan Zhang and Wenxuan Zhang and Yifei Yuan and See. On the Multi-turn Instruction Following for Conversational Web Agents , booktitle =. 2024 , url =. doi:10.18653/V1/2024.ACL-LONG.477 , timestamp =

  17. [25]

    Toward Verifiable Instruction-Following Alignment for Retrieval Augmented Generation , booktitle =

    Guanting Dong and Xiaoshuai Song and Yutao Zhu and Runqi Qiao and Zhicheng Dou and Ji. Toward Verifiable Instruction-Following Alignment for Retrieval Augmented Generation , booktitle =. 2025 , url =. doi:10.1609/AAAI.V39I22.34551 , timestamp =

  18. [26]

    TL-Training:

    Junjie Ye and Yilong Wu and Sixian Li and Yuming Yang and Zhiheng Xi and Tao Gui and Qi Zhang and Xuanjing Huang and Peng Wang and Zhongchao Shi and Jianping Fan and Zhengyin Du , editor =. TL-Training:. Findings of the Association for Computational Linguistics:. 2025 , url =

  19. [27]

    CoRR , volume =

    Jeffrey Zhou and Tianjian Lu and Swaroop Mishra and Siddhartha Brahma and Sujoy Basu and Yi Luan and Denny Zhou and Le Hou , title =. CoRR , volume =. 2023 , url =. doi:10.48550/ARXIV.2311.07911 , eprinttype =. 2311.07911 , timestamp =

  20. [28]

    Robinson and Keren Gu

    Tyna Eloundou and Alex Beutel and David G. Robinson and Keren Gu. First-Person Fairness in Chatbots , journal =. 2024 , url =. doi:10.48550/ARXIV.2410.19803 , eprinttype =. 2410.19803 , timestamp =

  21. [29]

    Tool learning with large language models: a survey , journal =

    Changle Qu and Sunhao Dai and Xiaochi Wei and Hengyi Cai and Shuaiqiang Wang and Dawei Yin and Jun Xu and Ji. Tool learning with large language models: a survey , journal =. 2025 , url =. doi:10.1007/S11704-024-40678-2 , timestamp =

  22. [30]

    T. M. Mitchell. The Need for Biases in Learning Generalizations. 1980

  23. [31]

    M. J. Kearns , title =

  24. [32]

    Machine Learning: An Artificial Intelligence Approach, Vol. I. 1983

  25. [33]

    R. O. Duda and P. E. Hart and D. G. Stork. Pattern Classification. 2000

  26. [34]

    Suppressed for Anonymity , author=

  27. [35]

    Newell and P

    A. Newell and P. S. Rosenbloom. Mechanisms of Skill Acquisition and the Law of Practice. Cognitive Skills and Their Acquisition. 1981

  28. [36]

    A. L. Samuel. Some Studies in Machine Learning Using the Game of Checkers. IBM Journal of Research and Development. 1959

  29. [37]

    Toolformer: Language Models Can Teach Themselves to Use Tools , booktitle =

    Timo Schick and Jane Dwivedi. Toolformer: Language Models Can Teach Themselves to Use Tools , booktitle =. 2023 , url =

  30. [38]

    and Stoica, Ion and Xing, Eric P

    Chiang, Wei-Lin and Li, Zhuohan and Lin, Zi and Sheng, Ying and Wu, Zhanghao and Zhang, Hao and Zheng, Lianmin and Zhuang, Siyuan and Zhuang, Yonghao and Gonzalez, Joseph E. and Stoica, Ion and Xing, Eric P. , month =. Vicuna: An Open-Source Chatbot Impressing GPT-4 with 90\ url =

  31. [39]

    Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback , journal =

    Yuntao Bai and Andy Jones and Kamal Ndousse and Amanda Askell and Anna Chen and Nova DasSarma and Dawn Drain and Stanislav Fort and Deep Ganguli and Tom Henighan and Nicholas Joseph and Saurav Kadavath and Jackson Kernion and Tom Conerly and Sheer El Showk and Nelson Elhage an...

  32. [40]

    Brown and Benjamin Chess and Rewon Child and Scott Gray and Alec Radford and Jeffrey Wu and Dario Amodei , title =

    Jared Kaplan and Sam McCandlish and Tom Henighan and Tom B. Brown and Benjamin Chess and Rewon Child and Scott Gray and Alec Radford and Jeffrey Wu and Dario Amodei , title =. CoRR , volume =. 2020 , url =. 2001.08361 , timestamp =

  33. [41]

    2011 , publisher=

    Animal tool behavior: the use and manufacture of tools by animals , author=. 2011 , publisher=

  34. [42]

    CoRR , volume =

    Yifan Song and Weimin Xiong and Dawei Zhu and Cheng Li and Ke Wang and Ye Tian and Sujian Li , title =. CoRR , volume =. 2023 , url =. doi:10.48550/ARXIV.2306.06624 , eprinttype =. 2306.06624 , timestamp =

  35. [43]

    Improving Language Models by Retrieving from Trillions of Tokens , booktitle =

    Sebastian Borgeaud and Arthur Mensch and Jordan Hoffmann and Trevor Cai and Eliza Rutherford and Katie Millican and George van den Driessche and Jean. Improving Language Models by Retrieving from Trillions of Tokens , booktitle =. 2022 , url =

  36. [44]

    Chameleon: Plug-and-Play Compositional Reasoning with Large Language Models , journal =

    Pan Lu and Baolin Peng and Hao Cheng and Michel Galley and Kai. Chameleon: Plug-and-Play Compositional Reasoning with Large Language Models , journal =. 2023 , url =. doi:10.48550/ARXIV.2304.09842 , eprinttype =. 2304.09842 , timestamp =

  37. [45]

    LaMDA: Language Models for Dialog Applications , journal =

    Romal Thoppilan and Daniel De Freitas and Jamie Hall and Noam Shazeer and Apoorv Kulshreshtha and Heng. LaMDA: Language Models for Dialog Applications , journal =. 2022 , url =. 2201.08239 , timestamp =

  38. [46]

    CoRR , volume =

    Aaron Parisi and Yao Zhao and Noah Fiedel , title =. CoRR , volume =. 2022 , url =. doi:10.48550/ARXIV.2205.12255 , eprinttype =. 2205.12255 , timestamp =

  39. [47]

    CoRR , volume =

    Yongliang Shen and Kaitao Song and Xu Tan and Dongsheng Li and Weiming Lu and Yueting Zhuang , title =. CoRR , volume =. 2023 , url =. doi:10.48550/ARXIV.2303.17580 , eprinttype =. 2303.17580 , timestamp =

  40. [48]

    CoRR , volume =

    Shibo Hao and Tianyang Liu and Zhen Wang and Zhiting Hu , title =. CoRR , volume =. 2023 , url =. doi:10.48550/ARXIV.2305.11554 , eprinttype =. 2305.11554 , timestamp =

  41. [49]

    CoRR , volume =

    Qiantong Xu and Fenglu Hong and Bo Li and Changran Hu and Zhengyu Chen and Jian Zhang , title =. CoRR , volume =. 2023 , url =. doi:10.48550/ARXIV.2305.16504 , eprinttype =. 2305.16504 , timestamp =

  42. [50]

    CoRR , volume =

    Zihao Wang and Shaofei Cai and Anji Liu and Xiaojian Ma and Yitao Liang , title =. CoRR , volume =. 2023 , url =. doi:10.48550/ARXIV.2302.01560 , eprinttype =. 2302.01560 , timestamp =

  43. [51]

    CoRR , volume =

    Jingqing Ruan and Yihong Chen and Bin Zhang and Zhiwei Xu and Tianpeng Bao and Guoqing Du and Shiwei Shi and Hangyu Mao and Xingyu Zeng and Rui Zhao , title =. CoRR , volume =. 2023 , url =. doi:10.48550/ARXIV.2308.03427 , eprinttype =. 2308.03427 , timestamp =

  44. [52]

    CoRR , volume =

    Reiichiro Nakano and Jacob Hilton and Suchir Balaji and Jeff Wu and Long Ouyang and Christina Kim and Christopher Hesse and Shantanu Jain and Vineet Kosaraju and William Saunders and Xu Jiang and Karl Cobbe and Tyna Eloundou and Gretchen Krueger and Kevin Button and Matthew Kn...

  45. [53]

    Zhao and Yanping Huang and Andrew M

    Hyung Won Chung and Le Hou and Shayne Longpre and Barret Zoph and Yi Tay and William Fedus and Eric Li and Xuezhi Wang and Mostafa Dehghani and Siddhartha Brahma and Albert Webson and Shixiang Shane Gu and Zhuyun Dai and Mirac Suzgun and Xinyun Chen and Aakanksha Chowdhery and...

  46. [54]

    Chi and Tatsunori Hashimoto and Oriol Vinyals and Percy Liang and Jeff Dean and William Fedus , title =

    Jason Wei and Yi Tay and Rishi Bommasani and Colin Raffel and Barret Zoph and Sebastian Borgeaud and Dani Yogatama and Maarten Bosma and Denny Zhou and Donald Metzler and Ed H. Chi and Tatsunori Hashimoto and Oriol Vinyals and Percy Liang and Jeff Dean and William Fedus , titl...

  47. [55]

    Narasimhan and Yuan Cao , title =

    Shunyu Yao and Jeffrey Zhao and Dian Yu and Nan Du and Izhak Shafran and Karthik R. Narasimhan and Yuan Cao , title =. The Eleventh International Conference on Learning Representations,. 2023 , url =

  48. [56]

    Chi and Quoc V

    Jason Wei and Xuezhi Wang and Dale Schuurmans and Maarten Bosma and Brian Ichter and Fei Xia and Ed H. Chi and Quoc V. Le and Denny Zhou , editor =. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models , booktitle =. 2022 , url =

  49. [57]

    CoRR , volume =

    Jianlin Su and Yu Lu and Shengfeng Pan and Bo Wen and Yunfeng Liu , title =. CoRR , volume =. 2021 , url =. 2104.09864 , timestamp =

  50. [58]

    CoRR , volume =

    Yining Lu and Haoping Yu and Daniel Khashabi , title =. CoRR , volume =. 2023 , url =. doi:10.48550/ARXIV.2307.08775 , eprinttype =. 2307.08775 , timestamp =

  51. [59]

    CoRR , volume =

    Chenfei Wu and Shengming Yin and Weizhen Qi and Xiaodong Wang and Zecheng Tang and Nan Duan , title =. CoRR , volume =. 2023 , url =. doi:10.48550/ARXIV.2303.04671 , eprinttype =. 2303.04671 , timestamp =

  52. [60]

    GPT4Tools: Teaching Large Language Model to Use Tools via Self-instruction , booktitle =

    Rui Yang and Lin Song and Yanwei Li and Sijie Zhao and Yixiao Ge and Xiu Li and Ying Shan , editor =. GPT4Tools: Teaching Large Language Model to Use Tools via Self-instruction , booktitle =. 2023 , url =

  53. [61]

    , author=

    The generalisation of student's problems when several different population variances are involved. , author=. Biometrika , volume=. 1947 , url=

  54. [62]

    Patil and Tianjun Zhang and Xin Wang and Joseph E

    Shishir G. Patil and Tianjun Zhang and Xin Wang and Joseph E. Gonzalez , editor =. Gorilla: Large Language Model Connected with Massive APIs , booktitle =. 2024 , url =

  55. [63]

    Saravia, Elvis , journal =

  56. [64]

    Label Words are Anchors: An Information Flow Perspective for Understanding In-Context Learning , booktitle =

    Lean Wang and Lei Li and Damai Dai and Deli Chen and Hao Zhou and Fandong Meng and Jie Zhou and Xu Sun , editor =. Label Words are Anchors: An Information Flow Perspective for Understanding In-Context Learning , booktitle =. 2023 , url =. doi:10.18653/V1/2023.EMNLP-MAIN.609 , ...

  57. [65]

    Long Ouyang and Jeffrey Wu and Xu Jiang and Diogo Almeida and Carroll L. Wainwright and Pamela Mishkin and Chong Zhang and Sandhini Agarwal and Katarina Slama and Alex Ray and John Schulman and Jacob Hilton and Fraser Kelton and Luke Miller and Maddie Simens and Amanda Askell ...

  58. [66]

    Rethinking the Role of Demonstrations: What Makes In-Context Learning Work? , booktitle =

    Sewon Min and Xinxi Lyu and Ari Holtzman and Mikel Artetxe and Mike Lewis and Hannaneh Hajishirzi and Luke Zettlemoyer , editor =. Rethinking the Role of Demonstrations: What Makes In-Context Learning Work? , booktitle =. 2022 , url =. doi:10.18653/V1/2022.EMNLP-MAIN.759 , timestamp =

  59. [67]

    Nye and Maarten Bosma and Henryk Michalewski and David Dohan and Ellen Jiang and Carrie J

    Jacob Austin and Augustus Odena and Maxwell I. Nye and Maarten Bosma and Henryk Michalewski and David Dohan and Ellen Jiang and Carrie J. Cai and Michael Terry and Quoc V. Le and Charles Sutton , title =. CoRR , volume =. 2021 , url =. 2108.07732 , timestamp =

  60. [68]

    Measuring Mathematical Problem Solving With the

    Dan Hendrycks and Collin Burns and Saurav Kadavath and Akul Arora and Steven Basart and Eric Tang and Dawn Song and Jacob Steinhardt , editor =. Measuring Mathematical Problem Solving With the. Proceedings of the Neural Information Processing Systems Track on Datasets and Benc...

  61. [69]

    Daya Guo and Dejian Yang and Haowei Zhang and Junxiao Song and Peiyi Wang and Qihao Zhu and Runxin Xu and Ruoyu Zhang and Shirong Ma and Xiao Bi and Xiaokang Zhang and Xingkai Yu and Yu Wu and Z. F. Wu and Zhibin Gou and Zhihong Shao and Zhuoshu Li and Ziyi Gao and Aixin Liu a...

  62. [70]

    Zhihong Shao and Peiyi Wang and Qihao Zhu and Runxin Xu and Junxiao Song and Mingchuan Zhang and Y. K. Li and Y. Wu and Daya Guo , title =. CoRR , volume =. 2024 , url =. doi:10.48550/ARXIV.2402.03300 , eprinttype =. 2402.03300 , timestamp =

  63. [71]

    InFoBench: Evaluating Instruction Following Ability in Large Language Models , booktitle =

    Yiwei Qin and Kaiqiang Song and Yebowen Hu and Wenlin Yao and Sangwoo Cho and Xiaoyang Wang and Xuansheng Wu and Fei Liu and Pengfei Liu and Dong Yu , editor =. InFoBench: Evaluating Instruction Following Ability in Large Language Models , booktitle =. 2024 , url =. doi:10.186...

  64. [72]

    CoRR , volume =

    Xinghua Zhang and Haiyang Yu and Cheng Fu and Fei Huang and Yongbin Li , title =. CoRR , volume =. 2024 , url =. doi:10.48550/ARXIV.2411.06208 , eprinttype =. 2411.06208 , timestamp =

  65. [73]

    CIF-Bench:

    Yizhi Li and Ge Zhang and Xingwei Qu and Jiali Li and Zhaoqun Li and Noah Wang and Hao Li and Ruibin Yuan and Yinghao Ma and Kai Zhang and Wangchunshu Zhou and Yiming Liang and Lei Zhang and Lei Ma and Jiajun Zhang and Zuowen Li and Wenhao Huang and Chenghua Lin and Jie Fu , e...

  66. [74]

    Manning and Stefano Ermon and Chelsea Finn , editor =

    Rafael Rafailov and Archit Sharma and Eric Mitchell and Christopher D. Manning and Stefano Ermon and Chelsea Finn , editor =. Direct Preference Optimization: Your Language Model is Secretly a Reward Model , booktitle =. 2023 , url =

  67. [75]

    CoRR , volume =

    Yimin Jing and Renren Jin and Jiahao Hu and Huishi Qiu and Xiaohua Wang and Peng Wang and Deyi Xiong , title =. CoRR , volume =. 2023 , url =. doi:10.48550/ARXIV.2311.09829 , eprinttype =. 2311.09829 , timestamp =

  68. [76]

    From Complex to Simple: Enhancing Multi-Constraint Complex Instruction Following Ability of Large Language Models , booktitle =

    Qianyu He and Jie Zeng and Qianxi He and Jiaqing Liang and Yanghua Xiao , editor =. From Complex to Simple: Enhancing Multi-Constraint Complex Instruction Following Ability of Large Language Models , booktitle =. 2024 , url =

  69. [77]

    CoRR , volume =

    Haoran Sun and Lixin Liu and Junjie Li and Fengyu Wang and Baohua Dong and Ran Lin and Ruohui Huang , title =. CoRR , volume =. 2024 , url =. doi:10.48550/ARXIV.2404.02823 , eprinttype =. 2404.02823 , timestamp =

  70. [78]

    CoRR , volume =

    Guanting Dong and Keming Lu and Chengpeng Li and Tingyu Xia and Bowen Yu and Chang Zhou and Jingren Zhou , title =. CoRR , volume =. 2024 , url =. doi:10.48550/ARXIV.2406.13542 , eprinttype =. 2406.13542 , timestamp =

  71. [79]

    CoRR , volume =

    Jiale Cheng and Xiao Liu and Cunxiang Wang and Xiaotao Gu and Yida Lu and Dan Zhang and Yuxiao Dong and Jie Tang and Hongning Wang and Minlie Huang , title =. CoRR , volume =. 2024 , url =. doi:10.48550/ARXIV.2412.11605 , eprinttype =. 2412.11605 , timestamp =

  72. [80]

    CoRR , volume =

    Yun He and Di Jin and Chaoqi Wang and Chloe Bi and Karishma Mandyam and Hejia Zhang and Chen Zhu and Ning Li and Tengyu Xu and Hongjiang Lv and Shruti Bhosale and Chenguang Zhu and Karthik Abinav Sankararaman and Eryk Helenowski and Melanie Kambadur and Aditya Tayade and Hao M...

  73. [81]

    FollowBench:

    Yuxin Jiang and Yufei Wang and Xingshan Zeng and Wanjun Zhong and Liangyou Li and Fei Mi and Lifeng Shang and Xin Jiang and Qun Liu and Wei Wang , editor =. FollowBench:. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Pa...

  74. [82]

    CoRR , volume =

    Elliot Nelson and Georgios Kollias and Payel Das and Subhajit Chaudhury and Soham Dan , title =. CoRR , volume =. 2024 , url =. doi:10.48550/ARXIV.2407.01437 , eprinttype =. 2407.01437 , timestamp =

  75. [83]

    CoRR , volume =

    An Yang and Anfeng Li and Baosong Yang and Beichen Zhang and Binyuan Hui and Bo Zheng and Bowen Yu and Chang Gao and Chengen Huang and Chenxu Lv and Chujie Zheng and Dayiheng Liu and Fan Zhou and Fei Huang and Feng Hu and Hao Ge and Haoran Wei and Huan Lin and Jialong Tang and...

  76. [84]

    Findings of the Association for Computational Linguistics,

    Bo Li and Gexiang Fang and Wei Ye and Zhenghua Xu and Jinglei Zhang and Hao Cheng and Shikun Zhang , editor =. Findings of the Association for Computational Linguistics,. 2025 , url =

  77. [85]

    CoRR , volume =

    Xiao Wang and Weikang Zhou and Can Zu and Han Xia and Tianze Chen and Yuansen Zhang and Rui Zheng and Junjie Ye and Qi Zhang and Tao Gui and Jihua Kang and Jingsheng Yang and Siyuan Li and Chunsai Du , title =. CoRR , volume =. 2023 , url =. doi:10.48550/ARXIV.2304.08085 , epr...

  78. [86]

    Long Context vs

    Xinze Li and Yushi Bai and Bowen Jin and Fengbin Zhu and Liangming Pan and Yixin Cao , editor =. Long Context vs. Proceedings of the 48th International. 2025 , url =. doi:10.1145/3726302.3731690 , timestamp =

  79. [87]

    CoRR , volume =

    Junjie Ye and Caishuang Huang and Zhuohan Chen and Wenjie Fu and Chenyuan Yang and Leyi Yang and Yilong Wu and Peng Wang and Meng Zhou and Xiaolong Yang and Tao Gui and Qi Zhang and Zhongchao Shi and Jianping Fan and Xuanjing Huang , title =. CoRR , volume =. 2025 , url =. doi...

  80. [88]

    The Thirteenth International Conference on Learning Representations,

    Yanzhao Qin and Tao Zhang and Tao Zhang and Yanjun Shen and Wenjing Luo and Haoze Sun and Yan Zhang and Yujing Qiao and Weipeng Chen and Zenan Zhou and Wentao Zhang and Bin Cui , title =. The Thirteenth International Conference on Learning Representations,. 2025 , url =

  81. [89]

    Hanjie and Runzhe Yang and Karthik R

    Shunyu Yao and Howard Chen and Austin W. Hanjie and Runzhe Yang and Karthik R. Narasimhan , title =. The Twelfth International Conference on Learning Representations,. 2024 , url =

  82. [90]

    CoRR , volume =

    Haoran Wei and Yaofeng Sun and Yukun Li , title =. CoRR , volume =. 2025 , url =. doi:10.48550/ARXIV.2510.18234 , eprinttype =. 2510.18234 , timestamp =

  83. [91]

    CoRR , volume =

    Junjie Ye and Changhao Jiang and Zhengyin Du and Yufei Xu and Xuesong Yao and Zhiheng Xi and Xiaoran Fan and Qi Zhang and Xuanjing Huang and Jiecao Chen , title =. CoRR , volume =. 2025 , url =. doi:10.48550/ARXIV.2508.08791 , eprinttype =. 2508.08791 , timestamp =

  84. [92]

    DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models , journal =

    DeepSeek. DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models , journal =. 2025 , url =. doi:10.48550/ARXIV.2512.02556 , eprinttype =. 2512.02556 , timestamp =

  85. [93]

    CoRR , volume =

    An Yang and Baosong Yang and Beichen Zhang and Binyuan Hui and Bo Zheng and Bowen Yu and Chengyuan Li and Dayiheng Liu and Fei Huang and Haoran Wei and Huan Lin and Jian Yang and Jianhong Tu and Jianwei Zhang and Jianxin Yang and Jiaxi Yang and Jingren Zhou and Junyang Lin and...

  86. [94]

    2024 , url=

    Introducing Llama 3.1: Our most capable models to date , author=. 2024 , url=

  87. [95]

    2023 , url=

    NexusRaven-V2: Surpassing GPT-4 for Zero-shot Function Calling , author=. 2023 , url=

  88. [96]

    Code Llama: Open Foundation Models for Code , journal =

    Baptiste Rozi. Code Llama: Open Foundation Models for Code , journal =. 2023 , url =. doi:10.48550/ARXIV.2308.12950 , eprinttype =. 2308.12950 , timestamp =

  89. [97]

    ToolHop:

    Junjie Ye and Zhengyin Du and Xuesong Yao and Weijian Lin and Yufei Xu and Zehui Chen and Zaiyuan Wang and Sining Zhu and Zhiheng Xi and Siyu Yuan and Tao Gui and Qi Zhang and Xuanjing Huang and Jiecao Chen , editor =. ToolHop:. Proceedings of the 63rd Annual Meeting of the As...

  90. [98]

    2023 , url=

    Fortify the Shortest Stave in Attention: Enhancing Context Awareness of Large Language Models for Effective Tool Use , author=. 2023 , url=

  91. [99]

    T-Eval: Evaluating the Tool Utilization Capability of Large Language Models Step by Step , booktitle =

    Zehui Chen and Weihua Du and Wenwei Zhang and Kuikun Liu and Jiangning Liu and Miao Zheng and Jingming Zhuo and Songyang Zhang and Dahua Lin and Kai Chen and Feng Zhao , editor =. T-Eval: Evaluating the Tool Utilization Capability of Large Language Models Step by Step , bookti...

  92. [100]

    TaskBench: Benchmarking Large Language Models for Task Automation , booktitle =

    Yongliang Shen and Kaitao Song and Xu Tan and Wenqi Zhang and Kan Ren and Siyu Yuan and Weiming Lu and Dongsheng Li and Yueting Zhuang , editor =. TaskBench: Benchmarking Large Language Models for Task Automation , booktitle =. 2024 , url =

  93. [101]

    API-Bank:

    Minghao Li and Yingxiu Zhao and Bowen Yu and Feifan Song and Hangyu Li and Haiyang Yu and Zhoujun Li and Fei Huang and Yongbin Li , editor =. API-Bank:. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing,. 2023 , url =

  94. [102]

    CoRR , volume =

    Qiao Jin and Yifan Yang and Qingyu Chen and Zhiyong Lu , title =. CoRR , volume =. 2023 , url =. doi:10.48550/ARXIV.2304.09667 , eprinttype =. 2304.09667 , timestamp =

  95. [103]

    Select-Then-Decompose: From Empirical Analysis to Adaptive Selection Strategy for Task Decomposition in Large Language Models , booktitle =

    Shuodi Liu and Yingzhuo Liu and Zi Wang and Yusheng Wang and Huijia Wu and Liuyu Xiang and Zhaofeng He , editor =. Select-Then-Decompose: From Empirical Analysis to Adaptive Selection Strategy for Task Decomposition in Large Language Models , booktitle =. 2025 , url =. doi:10....

  96. [104]

    Zhehao Zhang and Ryan A. Rossi and Branislav Kveton and Yijia Shao and Diyi Yang and Hamed Zamani and Franck Dernoncourt and Joe Barrow and Tong Yu and Sungchul Kim and Ruiyi Zhang and Jiuxiang Gu and Tyler Derr and Hongjie Chen and Junda Wu and Xiang Chen and Zichao Wang and ...

  97. [105]

    CoRR , volume =

    Xu Huang and Yuefeng Huang and Weiwen Liu and Xingshan Zeng and Yasheng Wang and Ruiming Tang and Hong Xie and Defu Lian , title =. CoRR , volume =. 2025 , url =. doi:10.48550/ARXIV.2505.04072 , eprinttype =. 2505.04072 , timestamp =

  98. [106]

    Budget-Constrained Tool Learning with Planning , booktitle =

    Yuanhang Zheng and Peng Li and Ming Yan and Ji Zhang and Fei Huang and Yang Liu , editor =. Budget-Constrained Tool Learning with Planning , booktitle =. 2024 , url =. doi:10.18653/V1/2024.FINDINGS-ACL.536 , timestamp =

  99. [107]

    CoRR , volume =

    Zekun Li and Shinda Huang and Jiangtian Wang and Nathan Zhang and Antonis Antoniades and Wenyue Hua and Kaijie Zhu and Sirui Zeng and William Yang Wang and Xifeng Yan , title =. CoRR , volume =. 2025 , url =. doi:10.48550/ARXIV.2503.08669 , eprinttype =. 2503.08669 , timestamp =

  100. [108]

    Renze Lou and Kai Zhang and Wenpeng Yin , title =. Comput. Linguistics , volume =. 2024 , url =. doi:10.1162/COLI\_A\_00523 , timestamp =

  101. [109]

    AppWorld:

    Harsh Trivedi and Tushar Khot and Mareike Hartmann and Ruskin Manku and Vinty Dong and Edward Li and Shashank Gupta and Ashish Sabharwal and Niranjan Balasubramanian , editor =. AppWorld:. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics ...

  102. [110]

    CoRR , volume =

    Ryan Wong and Jiawei Wang and Junjie Zhao and Li Chen and Yan Gao and Long Zhang and Xuan Zhou and Zuo Wang and Kai Xiang and Ge Zhang and Wenhao Huang and Yang Wang and Ke Wang , title =. CoRR , volume =. 2025 , url =. doi:10.48550/ARXIV.2508.07999 , eprinttype =. 2508.07999 ...

  103. [111]

    Narasimhan , title =

    Shunyu Yao and Noah Shinn and Pedram Razavi and Karthik R. Narasimhan , title =. The Thirteenth International Conference on Learning Representations,. 2025 , url =

  104. [112]

    CoRR , volume =

    Chengrui Huang and Zhengliang Shi and Yuntao Wen and Xiuying Chen and Peng Han and Shen Gao and Shuo Shang , title =. CoRR , volume =. 2024 , url =. doi:10.48550/ARXIV.2407.03007 , eprinttype =. 2407.03007 , timestamp =

  105. [113]

    InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated Large Language Model Agents , booktitle =

    Qiusi Zhan and Zhixiang Liang and Zifan Ying and Daniel Kang , editor =. InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated Large Language Model Agents , booktitle =. 2024 , url =. doi:10.18653/V1/2024.FINDINGS-ACL.624 , timestamp =

  106. [114]

    Pei Wang and Yanan Wu and Noah Wang and Jiaheng Liu and Xiaoshuai Song and Z. Y. Peng and Ken Deng and Chenchen Zhang and Jiakai Wang and Junran Peng and Ge Zhang and Hangyu Guo and Zhaoxiang Zhang and Wenbo Su and Bo Zheng , title =. The Thirteenth International Conference on...

  107. [115]

    CoRR , volume =

    Chen Chen and Xinlong Hao and Weiwen Liu and Xu Huang and Xingshan Zeng and Shuai Yu and Dexun Li and Shuai Wang and Weinan Gan and Yuefeng Huang and Wulong Liu and Xinzhi Wang and Defu Lian and Baoqun Yin and Yasheng Wang and Wu Liu , title =. CoRR , volume =. 2025 , url =. d...

  108. [116]

    Yuchen Zhuang and Yue Yu and Kuan Wang and Haotian Sun and Chao Zhang , editor =. ToolQA:. Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 2023 , year =

  109. [117]

    CoRR , volume =

    Sherry Yang and Ofir Nachum and Yilun Du and Jason Wei and Pieter Abbeel and Dale Schuurmans , title =. CoRR , volume =. 2023 , url =. doi:10.48550/ARXIV.2303.04129 , eprinttype =. 2303.04129 , timestamp =

  110. [118]

    Augmented Language Models: a Survey , journal =

    Gr. Augmented Language Models: a Survey , journal =. 2023 , url =. doi:10.48550/ARXIV.2302.07842 , eprinttype =. 2302.07842 , timestamp =

  111. [119]

    CoRR , volume =

    Han Han and Tong Zhu and Xiang Zhang and Mengsong Wu and Hao Xiong and Wenliang Chen , title =. CoRR , volume =. 2024 , url =. doi:10.48550/ARXIV.2410.11805 , eprinttype =. 2410.11805 , timestamp =

  112. [120]

    CoRR , volume =

    Yuanqing Yu and Zhefan Wang and Weizhi Ma and Zhicheng Guo and Jingtao Zhan and Shuai Wang and Chuhan Wu and Zhiqiang Guo and Min Zhang , title =. CoRR , volume =. 2024 , url =. doi:10.48550/ARXIV.2410.07745 , eprinttype =. 2410.07745 , timestamp =

  113. [121]

    The Twelfth International Conference on Learning Representations,

    Yue Huang and Jiawen Shi and Yuan Li and Chenrui Fan and Siyuan Wu and Qihui Zhang and Yixin Liu and Pan Zhou and Yao Wan and Neil Zhenqiang Gong and Lichao Sun , title =. The Twelfth International Conference on Learning Representations,. 2024 , url =

  114. [122]

    Gemini 3 Pro Model Card , url=

    Google , year=. Gemini 3 Pro Model Card , url=

  115. [123]

    Gemini 3.1 Pro Model Card , url=

    Google , year=. Gemini 3.1 Pro Model Card , url=

  116. [124]

    GPT-5.1 Instant and GPT-5.1 Thinking System Card Addendum , url=

    OpenAI , year=. GPT-5.1 Instant and GPT-5.1 Thinking System Card Addendum , url=

  117. [125]

    Update to GPT-5 System Card: GPT-5.2 , url=

    OpenAI , year=. Update to GPT-5 System Card: GPT-5.2 , url=

  118. [126]

    GPT-5.5 System Card , url=

    OpenAI , year=. GPT-5.5 System Card , url=

  119. [127]

    System Card: Claude Opus 4.6 , url=

    Anthropic , year=. System Card: Claude Opus 4.6 , url=

  120. [128]

    System Card: Claude Opus 4.8 , url=

    Anthropic , year=. System Card: Claude Opus 4.8 , url=

  121. [129]

    Qwen3.5: Towards Native Multimodal Agents , url=

    Qwen Team , year=. Qwen3.5: Towards Native Multimodal Agents , url=

  122. [130]

    Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity , url=

    Bytedance Seed , year=. Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity , url=

  123. [131]

    Seed2.1 Model Card: Agentic Intelligence for Productivity , url=

  124. [132]

    OpenAI o3 and o4-mini System Card , url=

    OpenAI , year=. OpenAI o3 and o4-mini System Card , url=

  125. [133]

    Seed1.6 Tech Introduction , url=

    ByteDance Seed , year=. Seed1.6 Tech Introduction , url=

  126. [134]

    Kimi K2.6: Advancing Open-Source Coding , url=

    Kimi , year=. Kimi K2.6: Advancing Open-Source Coding , url=

  127. [135]

    GLM-5.2: Built for Long-Horizon Tasks , url=

  128. [136]

    Introducing Hy3 , url=

  129. [137]

    CoRR , volume =

    Kimi K2.5: Visual Agentic Intelligence , author=. CoRR , volume =. 2025 , url =. doi:10.48550/ARXIV.2602.02276 , eprinttype =

  130. [138]

    Yifan Bai and Yiping Bao and Guanduo Chen and Jiahao Chen and Ningxin Chen and Ruijue Chen and Yanru Chen and Yuankun Chen and Yutian Chen and Zhuofu Chen and Jialei Cui and Hao Ding and Mengnan Dong and Angang Du and Chenzhuang Du and Dikang Du and Yulun Du and Yu Fan and Yic...

  131. [139]

    MarioGPT: Open-Ended Text2Level Generation through Large Language Models , booktitle =

    Shyam Sudhakaran and Miguel Gonz. MarioGPT: Open-Ended Text2Level Generation through Large Language Models , booktitle =. 2023 , url =

  132. [140]

    Forty-second International Conference on Machine Learning,

    Qinggang Zhang and Hao Chen and Junnan Dong and Shengyuan Chen and Feiran Huang and Xiao Huang , title =. Forty-second International Conference on Machine Learning,. 2025 , url =

  133. [141]

    2025 , url =

    Gianpaolo Ghiani and Emanuele Manni and Sandro Zacchino , title =. 2025 , url =. doi:10.1109/ACCESS.2025.3601518 , timestamp =

  134. [142]

    Forty-second International Conference on Machine Learning,

    Simin Chen and Pranav Pusarla and Baishakhi Ray , title =. Forty-second International Conference on Machine Learning,. 2025 , url =

  135. [143]

    Madhavi Kumari and Rohit Chauhan and Rashi Jain and Prabha Garg , title =. Inf. Fusion , volume =. 2026 , url =. doi:10.1016/J.INFFUS.2025.103902 , timestamp =

  136. [144]

    Find the Intention of Instruction: Comprehensive Evaluation of Instruction Understanding for Large Language Models , booktitle =

    Hyeonseok Moon and Jaehyung Seo and Seungyoon Lee and Chanjun Park and Heuiseok Lim , editor =. Find the Intention of Instruction: Comprehensive Evaluation of Instruction Understanding for Large Language Models , booktitle =. 2025 , url =. doi:10.18653/V1/2025.FINDINGS-NAACL.3...

  137. [145]

    Large Language Models Meet Open-World Intent Discovery and Recognition: An Evaluation of ChatGPT , booktitle =

    Xiaoshuai Song and Keqing He and Pei Wang and Guanting Dong and Yutao Mou and Jingang Wang and Yunsen Xian and Xunliang Cai and Weiran Xu , editor =. Large Language Models Meet Open-World Intent Discovery and Recognition: An Evaluation of ChatGPT , booktitle =. 2023 , url =. d...

  138. [146]

    International Joint Conference on Neural Networks,

    Pengfei Yin and Yubing Ren and Fang Fang and Boxiang Hu and Xu Bai , title =. International Joint Conference on Neural Networks,. 2025 , url =. doi:10.1109/IJCNN64981.2025.11227926 , timestamp =

  139. [147]

    FinDABench: Benchmarking Financial Data Analysis Ability of Large Language Models , booktitle =

    Shu Liu and Shangqing Zhao and Chenghao Jia and Xinlin Zhuang and Zhaoguang Long and Jie Zhou and Aimin Zhou and Man Lan and Yang Chong , editor =. FinDABench: Benchmarking Financial Data Analysis Ability of Large Language Models , booktitle =. 2025 , url =

  140. [148]

    CoRR , volume =

    Zijun Liu and Peiyi Wang and Runxin Xu and Shirong Ma and Chong Ruan and Peng Li and Yang Liu and Yu Wu , title =. CoRR , volume =. 2025 , url =. doi:10.48550/ARXIV.2504.02495 , eprinttype =. 2504.02495 , timestamp =

  141. [149]

    Proceedings of the Twentieth European Conference on Computer Systems, EuroSys 2025, Rotterdam, The Netherlands, 30 March 2025 - 3 April 2025 , pages =

    Guangming Sheng and Chi Zhang and Zilingfeng Ye and Xibin Wu and Wang Zhang and Ru Zhang and Yanghua Peng and Haibin Lin and Chuan Wu , title =. Proceedings of the Twentieth European Conference on Computer Systems, EuroSys 2025, Rotterdam, The Netherlands, 30 March 2025 - 3 Ap...

  142. [150]

    CoRR , volume =

    OpenAI , title =. CoRR , volume =. 2025 , url =. doi:10.48550/ARXIV.2508.10925 , eprinttype =. 2508.10925 , timestamp =

  143. [151]

    Xinyu Pan and Weibin Zhuang and Sijie Wen and Weigang Yu and Jinsong Bao and Xinyu Li , title =. Adv. Eng. Informatics , volume =. 2025 , url =. doi:10.1016/J.AEI.2025.103422 , timestamp =

  144. [152]

    Analyzing the Effects of Supervised Fine-Tuning on Model Knowledge from Token and Parameter Levels

    Ye, Junjie and Yang, Yuming and Nan, Yang and Li, Shuo and Zhang, Qi and Gui, Tao and Huang, Xuanjing and Wang, Peng and Shi, Zhongchao and Fan, Jianping. Analyzing the Effects of Supervised Fine-Tuning on Model Knowledge from Token and Parameter Levels. Proceedings of the 202...

  145. [153]

    Patil and Qi Qi and Shengyu Feng and Julian Katz

    Yun He and Wenzhe Li and Hejia Zhang and Songlin Li and Karishma Mandyam and Sopan Khosla and Yuanhao Xiong and Nanshu Wang and Xiaoliang Peng and Beibin Li and Shengjie Bi and Shishir G. Patil and Qi Qi and Shengyu Feng and Julian Katz. AdvancedIF: Rubric-Based Benchmarking a...

  146. [154]

    Ross and Katherine L

    Hongkai Zheng and Wenda Chu and Bingliang Zhang and Zihui Wu and Austin Wang and Berthy Feng and Caifeng Zou and Yu Sun and Nikola Borislavov Kovachki and Zachary E. Ross and Katherine L. Bouman and Yisong Yue , title =. The Thirteenth International Conference on Learning Repr...

  147. [155]

    MultiChallenge:

    Kaustubh Deshpande and Ved Sirdeshmukh and Johannes Baptist Mols and Lifeng Jin and Ed. MultiChallenge:. Findings of the Association for Computational Linguistics,. 2025 , url =

  148. [156]

    Lillicrap and Jean

    Machel Reid and Nikolay Savinov and Denis Teplyashin and Dmitry Lepikhin and Timothy P. Lillicrap and Jean. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context , journal =. 2024 , url =. doi:10.48550/ARXIV.2403.05530 , eprinttype =. 2403.05530 ,...

  149. [157]

    CoRR , volume =

    Julian Schnitzler and Xanh Ho and Jiahao Huang and Florian Boudin and Saku Sugawara and Akiko Aizawa , title =. CoRR , volume =. 2024 , url =. doi:10.48550/ARXIV.2406.13397 , eprinttype =. 2406.13397 , timestamp =

  150. [158]

    CoRR , volume =

    Weiwen Liu and Xu Huang and Xingshan Zeng and Xinlong Hao and Shuai Yu and Dexun Li and Shuai Wang and Weinan Gan and Zhengying Liu and Yuanqing Yu and Zezhong Wang and Yuxian Wang and Wu Ning and Yutai Hou and Bin Wang and Chuhan Wu and Xinzhi Wang and Yong Liu and Yasheng Wa...

  151. [159]

    The Twelfth International Conference on Learning Representations,

    Yujia Qin and Shihao Liang and Yining Ye and Kunlun Zhu and Lan Yan and Yaxi Lu and Yankai Lin and Xin Cong and Xiangru Tang and Bill Qian and Sihan Zhao and Lauren Hong and Runchu Tian and Ruobing Xie and Jie Zhou and Mark Gerstein and Dahai Li and Zhiyuan Liu and Maosong Sun...

  152. [160]

    CoRR , volume =

    Qiaoyu Tang and Ziliang Deng and Hongyu Lin and Xianpei Han and Qiao Liang and Le Sun , title =. CoRR , volume =. 2023 , url =. doi:10.48550/ARXIV.2306.05301 , eprinttype =. 2306.05301 , timestamp =

  153. [161]

    Yujia Qin and Shengding Hu and Yankai Lin and Weize Chen and Ning Ding and Ganqu Cui and Zheni Zeng and Xuanhe Zhou and Yufei Huang and Chaojun Xiao and Chi Han and Yi R. Fung and Yusheng Su and Huadong Wang and Cheng Qian and Runchu Tian and Kunlun Zhu and Shihao Liang and Xi...

  154. [162]

    CoRR , volume =

    Xuanting Chen and Junjie Ye and Can Zu and Nuo Xu and Rui Zheng and Minlong Peng and Jie Zhou and Tao Gui and Qi Zhang and Xuanjing Huang , title =. CoRR , volume =. 2023 , url =. doi:10.48550/ARXIV.2303.00293 , eprinttype =. 2303.00293 , timestamp =

  155. [163]

    CoRR , volume =

    Junjie Ye and Xuanting Chen and Nuo Xu and Can Zu and Zekai Shao and Shichun Liu and Yuhan Cui and Zeyang Zhou and Chao Gong and Yang Shen and Jie Zhou and Siming Chen and Tao Gui and Qi Zhang and Xuanjing Huang , title =. CoRR , volume =. 2023 , url =. doi:10.48550/ARXIV.2303...

  156. [164]

    CoRR , volume =

    Zishan Guo and Renren Jin and Chuang Liu and Yufei Huang and Dan Shi and Supryadi and Linhao Yu and Yan Liu and Jiaxuan Li and Bojian Xiong and Deyi Xiong , title =. CoRR , volume =. 2023 , url =. doi:10.48550/ARXIV.2310.19736 , eprinttype =. 2310.19736 , timestamp =

  157. [165]

    Jacobs and Michael I

    Robert A. Jacobs and Michael I. Jordan and Steven J. Nowlan and Geoffrey E. Hinton , title =. Neural Comput. , volume =. 1991 , url =. doi:10.1162/NECO.1991.3.1.79 , timestamp =

  158. [166]

    Constitutional

    Yuntao Bai and Saurav Kadavath and Sandipan Kundu and Amanda Askell and Jackson Kernion and Andy Jones and Anna Chen and Anna Goldie and Azalia Mirhoseini and Cameron McKinnon and Carol Chen and Catherine Olsson and Christopher Olah and Danny Hernandez and Dawn Drain and Deep ...

  159. [167]

    Llama 2: Open Foundation and Fine-Tuned Chat Models , journal =

    Hugo Touvron and Louis Martin and Kevin Stone and Peter Albert and Amjad Almahairi and Yasmine Babaei and Nikolay Bashlykov and Soumya Batra and Prajjwal Bhargava and Shruti Bhosale and Dan Bikel and Lukas Blecher and Cristian Canton. Llama 2: Open Foundation and Fine-Tuned Ch...

  160. [168]

    LLaMA: Open and Efficient Foundation Language Models , journal =

    Hugo Touvron and Thibaut Lavril and Gautier Izacard and Xavier Martinet and Marie. LLaMA: Open and Efficient Foundation Language Models , journal =. 2023 , url =. doi:10.48550/ARXIV.2302.13971 , eprinttype =. 2302.13971 , timestamp =

  161. [169]

    CoRR , volume =

    OpenAI , title =. CoRR , volume =. 2023 , url =. doi:10.48550/ARXIV.2303.08774 , eprinttype =. 2303.08774 , timestamp =

  162. [170]

    Tom B. Brown and Benjamin Mann and Nick Ryder and Melanie Subbiah and Jared Kaplan and Prafulla Dhariwal and Arvind Neelakantan and Pranav Shyam and Girish Sastry and Amanda Askell and Sandhini Agarwal and Ariel Herbert. Language Models are Few-Shot Learners , booktitle =. 202...

  163. [171]

    Aho and Jeffrey D

    Alfred V. Aho and Jeffrey D. Ullman , title =. 1972

  164. [172]

    Publications Manual , year = "1983", publisher =

  165. [173]

    Chandra and Dexter C

    Ashok K. Chandra and Dexter C. Kozen and Larry J. Stockmeyer , year = "1981", title =. doi:10.1145/322234.322243

  166. [174]

    Zhiheng Xi and Wenxiang Chen and Xin Guo and Wei He and Yiwen Ding and Boyang Hong and Ming Zhang and Junzhe Wang and Senjie Jin and Enyu Zhou and Rui Zheng and Xiaoran Fan and Xiao Wang and Limao Xiong and Yuhao Zhou and Weiran Wang and Changhao Jiang and Yicheng Zou and Xian...

  167. [175]

    CoRR , volume =

    Yuxin Wang and Yiran Guo and Yining Zheng and Zhangyue Yin and Shuo Chen and Jie Yang and Jiajun Chen and Xuanjing Huang and Xipeng Qiu , title =. CoRR , volume =. 2025 , url =. doi:10.48550/ARXIV.2504.06766 , eprinttype =. 2504.06766 , timestamp =

  168. [176]

    CoRR , volume =

    John Schulman and Filip Wolski and Prafulla Dhariwal and Alec Radford and Oleg Klimov , title =. CoRR , volume =. 2017 , url =. 1707.06347 , timestamp =

  169. [177]

    CoRR , volume =

    Zhiheng Xi and Wenxiang Chen and Boyang Hong and Senjie Jin and Rui Zheng and Wei He and Yiwen Ding and Shichun Liu and Xin Guo and Junzhe Wang and Honglin Guo and Wei Shen and Xiaoran Fan and Yuhao Zhou and Shihan Dou and Xiao Wang and Xinbo Zhang and Peng Sun and Tao Gui and...

  170. [178]

    CoRR , volume =

    Mengsong Wu and Tong Zhu and Han Han and Chuanyuan Tan and Xiang Zhang and Wenliang Chen , title =. CoRR , volume =. 2024 , url =. doi:10.48550/ARXIV.2405.08355 , eprinttype =. 2405.08355 , timestamp =

  171. [179]

    CoRR , volume =

    Jinze Bai and Shuai Bai and Yunfei Chu and Zeyu Cui and Kai Dang and Xiaodong Deng and Yang Fan and Wenbin Ge and Yu Han and Fei Huang and Binyuan Hui and Luo Ji and Mei Li and Junyang Lin and Runji Lin and Dayiheng Liu and Gao Liu and Chengqiang Lu and Keming Lu and Jianxin M...

  172. [180]

    Self-Distillation Bridges Distribution Gap in Language Model Fine-Tuning , booktitle =

    Zhaorui Yang and Tianyu Pang and Haozhe Feng and Han Wang and Wei Chen and Minfeng Zhu and Qian Liu , editor =. Self-Distillation Bridges Distribution Gap in Language Model Fine-Tuning , booktitle =. 2024 , url =. doi:10.18653/V1/2024.ACL-LONG.58 , timestamp =

  173. [181]

    and Deepali Aneja and Zeyu Jin and Ramani Duraiswami and Dinesh Manocha , title =

    Sreyan Ghosh and Chandra Kiran Reddy Evuru and Sonal Kumar and Ramaneswaran S. and Deepali Aneja and Zeyu Jin and Ramani Duraiswami and Dinesh Manocha , title =. Forty-first International Conference on Machine Learning,. 2024 , url =

  174. [182]

    9th International Conference on Learning Representations,

    Dan Hendrycks and Collin Burns and Steven Basart and Andy Zou and Mantas Mazeika and Dawn Song and Jacob Steinhardt , title =. 9th International Conference on Learning Representations,. 2021 , url =

  175. [183]

    CoRR , volume =

    Karl Cobbe and Vineet Kosaraju and Mohammad Bavarian and Mark Chen and Heewoo Jun and Lukasz Kaiser and Matthias Plappert and Jerry Tworek and Jacob Hilton and Reiichiro Nakano and Christopher Hesse and John Schulman , title =. CoRR , volume =. 2021 , url =. 2110.14168 , timestamp =

  176. [184]

    Evaluating Large Language Models Trained on Code , journal =

    Mark Chen and Jerry Tworek and Heewoo Jun and Qiming Yuan and Henrique Pond. Evaluating Large Language Models Trained on Code , journal =. 2021 , url =. 2107.03374 , timestamp =

  177. [185]

    ToolSword: Unveiling Safety Issues of Large Language Models in Tool Learning Across Three Stages , booktitle =

    Junjie Ye and Sixian Li and Guanyu Li and Caishuang Huang and Songyang Gao and Yilong Wu and Qi Zhang and Tao Gui and Xuanjing Huang , editor =. ToolSword: Unveiling Safety Issues of Large Language Models in Tool Learning Across Three Stages , booktitle =. 2024 , url =. doi:10...

  178. [186]

    Knowledge Mechanisms in Large Language Models:

    Mengru Wang and Yunzhi Yao and Ziwen Xu and Shuofei Qiao and Shumin Deng and Peng Wang and Xiang Chen and Jia. Knowledge Mechanisms in Large Language Models:. CoRR , volume =. 2024 , url =. doi:10.48550/ARXIV.2407.15017 , eprinttype =. 2407.15017 , timestamp =

  179. [187]

    Learning or Self-aligning? Rethinking Instruction Fine-tuning , booktitle =

    Mengjie Ren and Boxi Cao and Hongyu Lin and Cao Liu and Xianpei Han and Ke Zeng and Guanglu Wan and Xunliang Cai and Le Sun , editor =. Learning or Self-aligning? Rethinking Instruction Fine-tuning , booktitle =. 2024 , url =. doi:10.18653/V1/2024.ACL-LONG.330 , timestamp =

  180. [188]

    Self-Refine: Iterative Refinement with Self-Feedback , booktitle =

    Aman Madaan and Niket Tandon and Prakhar Gupta and Skyler Hallinan and Luyu Gao and Sarah Wiegreffe and Uri Alon and Nouha Dziri and Shrimai Prabhumoye and Yiming Yang and Shashank Gupta and Bodhisattwa Prasad Majumder and Katherine Hermann and Sean Welleck and Amir Yazdanbakh...

  181. [189]

    Reflexion: language agents with verbal reinforcement learning , booktitle =

    Noah Shinn and Federico Cassano and Ashwin Gopinath and Karthik Narasimhan and Shunyu Yao , editor =. Reflexion: language agents with verbal reinforcement learning , booktitle =. 2023 , url =

  182. [190]

    CoRR , volume =

    Junjie Ye and Yuming Yang and Qi Zhang and Tao Gui and Xuanjing Huang and Peng Wang and Zhongchao Shi and Jianping Fan , title =. CoRR , volume =. 2024 , url =. doi:10.48550/ARXIV.2409.15825 , eprinttype =. 2409.15825 , timestamp =

  183. [191]

    7th International Conference on Learning Representations,

    Ilya Loshchilov and Frank Hutter , title =. 7th International Conference on Learning Representations,. 2019 , url =

  184. [192]

    CoRR , volume =

    An Yang and Baosong Yang and Binyuan Hui and Bo Zheng and Bowen Yu and Chang Zhou and Chengpeng Li and Chengyuan Li and Dayiheng Liu and Fei Huang and Guanting Dong and Haoran Wei and Huan Lin and Jialong Tang and Jialin Wang and Jian Yang and Jianhong Tu and Jianwei Zhang and...

  185. [193]

    CoRR , volume =

    Aohan Zeng and Bin Xu and Bowen Wang and Chenhui Zhang and Da Yin and Diego Rojas and Guanyu Feng and Hanlin Zhao and Hanyu Lai and Hao Yu and Hongning Wang and Jiadai Sun and Jiajie Zhang and Jiale Cheng and Jiayi Gui and Jie Tang and Jing Zhang and Juanzi Li and Lei Zhao and...

  186. [194]

    RoTBench:

    Junjie Ye and Yilong Wu and Songyang Gao and Caishuang Huang and Sixian Li and Guanyu Li and Xiaoran Fan and Qi Zhang and Tao Gui and Xuanjing Huang , editor =. RoTBench:. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing,. 2024 , url =

  187. [195]

    Rossi and Somdeb Sarkhel and Chao Zhang , title =

    Yuchen Zhuang and Xiang Chen and Tong Yu and Saayan Mitra and Victor Bursztyn and Ryan A. Rossi and Somdeb Sarkhel and Chao Zhang , title =. CoRR , volume =. 2023 , url =. doi:10.48550/ARXIV.2310.13227 , eprinttype =. 2310.13227 , timestamp =

  188. [196]

    CoRR , volume =

    Zhaoyang Liu and Zeqiang Lai and Zhangwei Gao and Erfei Cui and Zhiheng Li and Xizhou Zhu and Lewei Lu and Qifeng Chen and Yu Qiao and Jifeng Dai and Wenhai Wang , title =. CoRR , volume =. 2023 , url =. doi:10.48550/ARXIV.2310.17796 , eprinttype =. 2310.17796 , timestamp =

  189. [197]

    ToolEyes: Fine-Grained Evaluation for Tool Learning Capabilities of Large Language Models in Real-world Scenarios , booktitle =

    Junjie Ye and Guanyu Li and Songyang Gao and Caishuang Huang and Yilong Wu and Sixian Li and Xiaoran Fan and Shihan Dou and Tao Ji and Qi Zhang and Tao Gui and Xuanjing Huang , editor =. ToolEyes: Fine-Grained Evaluation for Tool Learning Capabilities of Large Language Models ...

  190. [198]

    Scalable training of

    Andrew, Galen and Gao, Jianfeng , booktitle=. Scalable training of

  191. [199]

    Dan Gusfield , title =. 1997

  192. [200]

    Tetreault , title =

    Mohammad Sadegh Rasooli and Joel R. Tetreault , title =. Computing Research Repository , volume =. 2015 , url =

  193. [201]

    A Framework for Learning Predictive Structures from Multiple Tasks and Unlabeled Data , Volume =

    Ando, Rie Kubota and Zhang, Tong , Issn =. A Framework for Learning Predictive Structures from Multiple Tasks and Unlabeled Data , Volume =. Journal of Machine Learning Research , Month = dec, Numpages =

  194. [202]

    ToolSandbox:

    Jiarui Lu and Thomas Holleis and Yizhe Zhang and Bernhard Aumayer and Feng Nan and Haoping Bai and Shuang Ma and Shen Ma and Mengyu Li and Guoli Yin and Zirui Wang and Ruoming Pang , editor =. ToolSandbox:. Findings of the Association for Computational Linguistics:. 2025 , url...

  195. [203]

    Mobile-Bench: An Evaluation Benchmark for LLM-based Mobile Agents , booktitle =

    Shihan Deng and Weikai Xu and Hongda Sun and Wei Liu and Tao Tan and Jianfeng Liu and Ang Li and Jian Luan and Bin Wang and Rui Yan and Shuo Shang , editor =. Mobile-Bench: An Evaluation Benchmark for LLM-based Mobile Agents , booktitle =. 2024 , url =. doi:10.18653/V1/2024.AC...

  196. [204]

    In Prospect and Retrospect: Reflective Memory Management for Long-term Personalized Dialogue Agents , booktitle =

    Zhen Tan and Jun Yan and I. In Prospect and Retrospect: Reflective Memory Management for Long-term Personalized Dialogue Agents , booktitle =. 2025 , url =. doi:10.18653/V1/2025.ACL-LONG.413 , timestamp =

  197. [205]

    Enabling Personalized Long-term Interactions in LLM-based Agents through Persistent Memory and User Profiles , journal =

    Rebecca Westh. Enabling Personalized Long-term Interactions in LLM-based Agents through Persistent Memory and User Profiles , journal =. 2025 , url =. doi:10.48550/ARXIV.2510.07925 , eprinttype =. 2510.07925 , timestamp =

  198. [206]

    Zollo and Andrew Wei Tung Siah and Naimeng Ye and Ang Li and Hongseok Namkoong , title =

    Thomas P. Zollo and Andrew Wei Tung Siah and Naimeng Ye and Ang Li and Hongseok Namkoong , title =. The Thirteenth International Conference on Learning Representations,. 2025 , url =

  199. [207]

    StableToolBench: Towards Stable Large-Scale Benchmarking on Tool Learning of Large Language Models , booktitle =

    Zhicheng Guo and Sijie Cheng and Hao Wang and Shihao Liang and Yujia Qin and Peng Li and Zhiyuan Liu and Maosong Sun and Yang Liu , editor =. StableToolBench: Towards Stable Large-Scale Benchmarking on Tool Learning of Large Language Models , booktitle =. 2024 , url =. doi:10....

  200. [208]

    Patil and Huanzhi Mao and Fanjia Yan and Charlie Cheng

    Shishir G. Patil and Huanzhi Mao and Fanjia Yan and Charlie Cheng. The Berkeley Function Calling Leaderboard. Forty-second International Conference on Machine Learning,. 2025 , url =

  201. [209]

    Jize Wang and Zerun Ma and Yining Li and Songyang Zhang and Cailian Chen and Kai Chen and Xinyi Le , editor =. Advances in Neural Information Processing Systems 37: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024, Vancouver, BC, Canada, December 1...

  202. [210]

    Cohen and Emine Yilmaz , editor =

    Zheng Zhao and Clara Vania and Subhradeep Kayal and Naila Khan and Shay B. Cohen and Emine Yilmaz , editor =. PersonaLens:. Findings of the Association for Computational Linguistics,. 2025 , url =. doi:10.18653/V1/2025.FINDINGS-ACL.927 , timestamp =

  203. [211]

    Mem-PAL: Towards Memory-based Personalized Dialogue Assistants for Long-term User-Agent Interaction , booktitle =

    Zhaopei Huang and Qifeng Dai and Guozheng Wu and Xiaopeng Wu and Xubin Li and Tiezheng Ge and Wenxuan Wang and Qin Jin , editor =. Mem-PAL: Towards Memory-based Personalized Dialogue Assistants for Long-term User-Agent Interaction , booktitle =. 2026 , url =. doi:10.1609/AAAI....

  204. [212]

    McAuley , title =

    Yuanzhe Hu and Yu Wang and Julian J. McAuley , title =. CoRR , volume =. 2025 , url =. doi:10.48550/ARXIV.2507.05257 , eprinttype =. 2507.05257 , timestamp =

  205. [213]

    Personalized Safety in LLMs:

    Yuchen Wu and Edward Sun and Kaijie Zhu and Jianxun Lian and Jos. Personalized Safety in LLMs:. CoRR , volume =. 2025 , url =. doi:10.48550/ARXIV.2505.18882 , eprinttype =. 2505.18882 , timestamp =

  206. [214]

    CoRR , volume =

    Jiazheng Sun and Te Yang and Jiayang Niu and Mingxuan Li and Yongyong Lu and Ruimeng Yang and Xin Peng , title =. CoRR , volume =. 2025 , url =. doi:10.48550/ARXIV.2509.20729 , eprinttype =. 2509.20729 , timestamp =

  207. [215]

    CoRR , volume =

    Zhenhailong Wang and Haiyang Xu and Junyang Wang and Xi Zhang and Ming Yan and Ji Zhang and Fei Huang and Heng Ji , title =. CoRR , volume =. 2025 , url =. doi:10.48550/ARXIV.2501.11733 , eprinttype =. 2501.11733 , timestamp =

  208. [216]

    Mobile-Agent-v2: Mobile Device Operation Assistant with Effective Navigation via Multi-Agent Collaboration , booktitle =

    Junyang Wang and Haiyang Xu and Haitao Jia and Xi Zhang and Ming Yan and Weizhou Shen and Ji Zhang and Fei Huang and Jitao Sang , editor =. Mobile-Agent-v2: Mobile Device Operation Assistant with Effective Navigation via Multi-Agent Collaboration , booktitle =. 2024 , url =

  209. [217]

    CoRR , volume =

    Yanchu Guan and Dong Wang and Zhixuan Chu and Shiyu Wang and Feiyue Ni and Ruihua Song and Longfei Li and Jinjie Gu and Chenyi Zhuang , title =. CoRR , volume =. 2023 , url =. doi:10.48550/ARXIV.2312.06677 , eprinttype =. 2312.06677 , timestamp =

  210. [218]

    AutoDroid: LLM-powered Task Automation in Android , booktitle =

    Hao Wen and Yuanchun Li and Guohong Liu and Shanhui Zhao and Tao Yu and Toby Jia. AutoDroid: LLM-powered Task Automation in Android , booktitle =. 2024 , url =. doi:10.1145/3636534.3649379 , timestamp =

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.