Pith. sign in

REVIEW 3 major objections 6 minor 66 references

This paper argues that evaluating personal AI assistants requires replaying the same temporal change event across different users' accumulated state—memories, skills, tool configurations, policies—and measuring how failures propagate across

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 14:57 UTC pith:NU2VFINO

load-bearing objection A transparent, well-scoped position paper whose C1-C4 requirement is a useful design lens; the audit's negative result is real but fragile to coding choices, so read it as a design argument, not a definitive gap. the 3 major comments →

arxiv 2607.21635 v1 pith:NU2VFINO submitted 2026-07-20 cs.LG

Toward User-Conditioned Evaluation of Personal LLM Agents under Temporal Interventions

classification cs.LG
keywords personal agentsLLM agent evaluationtemporal interventionuser-conditioned statebenchmark designadaptation metricsgap analysisfailure propagation
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper argues that evaluating personal AI assistants needs a new kind of benchmark test: take a single temporal change event—an API schema update, a policy tightening—and replay it across different users' accumulated state (stored memories, learned skills, tool configurations, safety policies), then measure how failures propagate across components. It formalizes this as four conditions C1–C4: an explicit temporal intervention, persistent state carried across it, induced cross-dimensional effects, and variation in user state. A focused audit of 15 public benchmark protocols finds that none satisfies all four under the paper's deliberately narrow coding rules, though several come close. The paper then proposes a minimal benchmark design and five reporting metrics so that future evaluations can measure user-conditioned adaptation directly rather than treating memory, tools, skills, and safety in isolation. The practical point: a change that is harmless for a light user can cascade into unrelated failures for a power user, and current benchmarks cannot see that.

Core claim

The central claim is that personal-agent evaluation should be user-conditioned: adaptation quality must be measured as Q(A, e_i | u_j)—agent A under change event e_i given persistent user state u_j—rather than a generic Q(A, e_i) over an average population. The paper operationalizes this as four conditions that a benchmark protocol must satisfy: C1 an explicit temporal change event perturbs one adaptation dimension; C2 the agent carries persistent state across that change; C3 the protocol measures induced effects on another dimension; C4 the protocol varies user-conditioned state or measures how the same change behaves under different persistent contexts. Auditing 15 public benchmark protoco

What carries the argument

The key machinery is the four-condition protocol C1–C4, applied to the agent tuple A=(M,T,K,S,C) (core model, external tools, user-conditioned memory, learned skills, safety/policy constraints). C4 is the load-bearing novelty: it forces the benchmark to hold the change event fixed while varying the persistent user state, making user state part of the system under test rather than just input. The accompanying metric family—α_L (adaptation latency), γ (graceful degradation), σ (safety preservation), κ (cross-dimensional coherence), and ρ (user-conditioned regression rate)—turns the protocol into a concrete reporting discipline with executable definitions and explicit denominators.

Load-bearing premise

The central negative result depends on the paper's deliberately narrow operationalization of the four conditions—especially C1's requirement of an explicit exogenous event injected after initial state construction and C4's requirement of explicit variation in user-conditioned state—applied to a 15-protocol denominator that excludes acknowledged close boundary cases; a broader operationalization or a larger protocol set could narrow or eliminate the identified gap.

What would settle it

Find a public benchmark protocol in the audited space that injects an explicit change event after initial user-state construction, carries persistent user state across it, measures an induced effect on a second adaptation dimension, and varies user-conditioned state—i.e., yields a 'Y' for all of C1, C2, C3, and C4 under the paper's codebook. Alternatively, demonstrating any one such protocol in the wider literature outside the 15-protocol denominator would directly narrow the claimed gap.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Future personal-agent benchmarks should include profile initialization, replayed fixed intervention events, event–dependency graphs, and per-user regression suites, so that the same update can be tested across light and power user states.
  • The five reporting metrics standardize how adaptation quality is reported: latency to sustained correct behavior, graceful failure acknowledgment, safety/policy preservation, stability of unaffected dimensions, and regression rate on previously solved user tasks.
  • Benchmarks that evaluate memory, tools, skills, or safety in isolation cannot detect user-conditioned failure propagation, so conclusions drawn from their aggregate scores may overstate real-world robustness of personal agents.
  • The same change event can produce different failure scopes for different user states—a power user with stale memories and learned skills can regress across unrelated workflows where a light user merely hits a missing-key error—so evaluation without user-state conditioning is insufficient for personal agents.
  • The concrete analytics/reporting assistant example shows how a single API schema change can be scored across direct recovery, graceful fallback, safety preservation, and regression in unaffected tasks, illustrating the benchmark card design.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If the C1–C4 gap is real, the most direct next step is a small-scale implementation study: build the proposed benchmark card for one domain, run several LLM agents under static, reactive, and versioned adaptation policies, and check whether the five metrics separate the policies as the paper's synthetic example suggests.
  • The negative result is bounded by the 15-protocol denominator and the narrow operationalization of C1 (requiring an explicit exogenous event after initial state construction). Including boundary cases like τ-bench or ST-WebAgentBench with added versioned event replay could close the gap, so the finding is better read as a call to extend existing protocols than as proof that no such benchmark can e
  • The metric family could generalize beyond LLM agents to any software system with persistent user state—e.g., regression testing of personalization features or recommender systems—but the paper does not claim this, and the definitions would need adaptation to non-agent settings.
  • The proposed metrics are demonstrated on simulated rollouts and a single oracle-compatibility trace, not on deployed agents; their practical behavior (e.g., judge reliability, sensitivity to the sustained-success threshold m) remains untested and is the natural target for empirical validation.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. This position paper argues that evaluating personal LLM agents requires 'user-conditioned evaluation': replaying the same temporal intervention (e.g., an API schema change or policy update) across different persistent user states and measuring how failures propagate across agent components. The authors formalize this as four conditions C1–C4 (explicit temporal intervention, persistent state across the intervention, induced cross-dimensional effects, variation in user-conditioned state), audit 15 public benchmark protocols against these conditions, and report that none satisfies all four. They then propose five reporting metrics (α_L, γ, σ, κ, ρ), a minimal benchmark-card design, and a synthetic trace-based computability check. The negative result is explicitly scoped to the audited set and operationalization, and the audit includes a codebook, inclusion criteria, and inter-coder reliability statistics.

Significance. If the C1–C4 formalization is accepted as a design requirement, the paper identifies a concrete and plausible blind spot in current personal-agent evaluation: existing benchmarks largely test single capabilities (tools, memory, skills, safety) in isolation, not the propagation of a fixed change across varied persistent user states. The paper's strengths are its transparency — a documented audit protocol, codebook, inclusion criteria, boundary-case analysis, and reported inter-coder agreement — and its careful scoping of the negative result as a focused gap analysis rather than an exhaustive claim. The proposed metrics are clearly labeled as candidate reporting tools, not as validated measures, and the synthetic example is a useful sanity check for computability. The main weaknesses are the selective denominator (close boundary cases are excluded from the 15-protocol set) and the relatively low inter-coder reliability on C1, which is the most consequential condition for the negative result.

major comments (3)
  1. [§4, Tables 5–6; Appendix A.1] The central negative result is conditional on a 15-protocol denominator that excludes three acknowledged close boundary cases: τ-bench, ST-WebAgentBench, and AgentEval (Table 6). The stated inclusion criteria in Appendix A.1 ('defines an LLM-agent evaluation task suite, provides reproducible metrics or an executable procedure, and targets at least one adaptation dimension') appear to apply to τ-bench and ST-WebAgentBench, yet Table 5 provides no rationale for their exclusion and τ-bench is absent from Table 5 entirely. Because the headline finding is 'no protocol in the audited set satisfies C1–C4,' the denominator must be principled and reproducible. I recommend coding these boundary cases with the same codebook, or explicitly stating which inclusion criterion each fails, and adding them to Table 5/7. If they fail C1–C4, the negative result is strengthened; if any pass under a reasonabl
  2. [Appendix A, Table 4] C1 has the lowest inter-coder reliability: agreement .800 and Gwet's AC1 .607, versus .885–1.000 for C2–C4. C1 is the first and most consequential condition, since it drives the exclusion of stateful protocols such as ToolSandbox, MemoryArena, WAREX, and ReliabilityBench. The paper reports that the final Inter. conclusion did not change after adjudication, but it does not report the number, direction, or resolution of C1 disagreements. Given that the boundary-case analysis hinges on what counts as an 'exogenous temporal event,' the audit should include a decision tree or explicit positive/negative examples for C1, and should report per-case disagreements. As written, C1 coding is not sufficient to bear the weight of the boundary-case exclusions.
  3. [§5, Table 2; Appendix B] The proposed C1–C4-satisfying benchmark is only a paper design. The synthetic rollouts in Table 8 are generated from assumed policy behavior (static, reactive, versioned), not from a real implementation of the benchmark card. To demonstrate that C1–C4 are not over-constrained, provide a minimal implemented pilot — even a small two-profile instance using a hosted LLM — that instantiates Table 2 and computes α_L, γ, σ, κ, ρ on actual agent traces. The current 'parser/oracle compatibility check' in Appendix B is a useful sanity check but does not show that the benchmark is administrable or that the metrics behave as intended on real trajectories.
minor comments (6)
  1. [Tables 4 and 7] The label 'Inter.' is used without definition in Table 4 and Table 7; define it in the caption or in the codebook (it appears to mean the C1–C4 interaction label).
  2. [Figure 2] The caption says 'white cells are complete gaps (0)' but the legend for 1–2 benchmarks is 'light cells.' Add a clear legend with exact cell counts or a supplementary CSV to improve readability.
  3. [§5] The text promises reporting of 'parameter sensitivity' for the sustained-success threshold m and the baseline threshold τ, but no sensitivity analysis is shown. Add a small example (e.g., α_L as a function of m on the synthetic trace) or explicitly state that sensitivity analysis is future work.
  4. [Appendix B] The 'hosted-LLM parser/oracle compatibility check' names 'DeepSeek deepseek-v4-flash' but does not provide the exact API version, date, prompt template, or oracle rubric. Without these, the check is not reproducible.
  5. [References] Several references are dated 2026 and may be preprint-only. Please verify that all arXiv identifiers and venue tags are accurate and add DOIs/URLs for preprint-only items.
  6. [§1 and Figure 1] The phrase 'the 3/15 near-miss pattern for temporal persistent state' is unclear relative to the 15-protocol denominator, especially because τ-bench appears in the boundary table but not in the denominator. Clarify which protocols are the three near-misses.

Circularity Check

0 steps flagged

No significant circularity: the gap finding is an explicitly scoped, empirically coded audit, and the proposed metrics are design proposals rather than fitted predictions.

full rationale

The paper's central negative result is an audit finding, not a derived prediction: a 15-protocol denominator is disclosed, C1–C4 are explicitly defined coding criteria, the codebook is described in Appendix A, and two external coders re-coded the headline labels with reported agreement. No parameter is fitted to a subset of data and then reported as a predicted value; the proposed metrics (α_L, γ, σ, κ, ρ) are labeled as candidate reporting tools, and the simulated rollouts in Appendix B are explicitly called engineering diagnostics for parser/oracle compatibility, not empirical benchmark results. No load-bearing self-citation appears: the authors do not invoke their own prior uniqueness theorem or ansatz, and no cited result is doing the work of forcing the conclusion. The abstract and Section 4 repeatedly scope the claim ('Under our explicitly narrow operationalization', 'best read as a narrow negative result'), and the boundary-case table (Table 6) discloses that τ-bench, ST-WebAgentBench, and AgentEval are outside the denominator and states what extension would be needed. The lower Gwet's AC1 for C1 (0.607) indicates coding ambiguity, which is a reliability/generalizability concern rather than a circular reduction of the result to its own inputs. The derivation chain is self-contained: the paper proposes a definitional requirement, applies it to a bounded coded set, and reports the outcome without disguising the criteria as an external discovery.

Axiom & Free-Parameter Ledger

2 free parameters · 4 axioms · 0 invented entities

The paper introduces no physical or conceptual entities beyond its proposed metric definitions and benchmark design. Its main free parameters are user-specified thresholds in the proposed metric definitions, not fitted values. The key load-bearing assumptions are the agent tuple model and the C1-C4 operationalization, both of which are stated explicitly and scoped.

free parameters (2)
  • m (sustained-success threshold in α_L) = 3 (illustrative in Table 8)
    Design choice in the proposed metric definition; not fitted to data. Example uses m=3, T=6.
  • τ (baseline threshold for κ denominator) = not specified
    Proposed metric requires a threshold to define valid baselines; value is left to benchmark designer.
axioms (4)
  • domain assumption The agent tuple A=(M,T,K,S,C) captures the relevant components of a personal agent.
    The entire framework of user-conditioned evaluation rests on this decomposition of personal agents into model, tools, memory, skills, and safety constraints. Introduced in Section 2.
  • ad hoc to paper The C1-C4 conditions are the correct formalization of user-conditioned temporal-intervention evaluation.
    This is the paper's central proposal, not an established result. The claim that no protocol satisfies all four is conditioned on these definitions. Section 4.
  • domain assumption The 15-protocol denominator is representative of public agent evaluation protocols.
    Inclusion criteria are explicit, but the negative result depends on this set. Boundary cases are excluded and discussed separately. Appendix A.
  • domain assumption Inter-rater coding (agreement .947, AC1 .898) is valid evidence for the coding decisions.
    The reliability check supports the coding matrix, but only two coders were used and the coding is inherently subjective. Appendix A.

pith-pipeline@v1.3.0-alltime-deepseek · 14800 in / 8839 out tokens · 99600 ms · 2026-08-01T14:57:36.064805+00:00 · methodology

0 comments
read the original abstract

Personal agents maintain memories, learned skills, tool configurations, and policy state that evolve with each user. Existing agent benchmarks often evaluate these capabilities in isolation: tool benchmarks test invocation under fixed APIs, memory benchmarks test recall or forgetting, and safety benchmarks test static policy compliance. We argue that personal-agent evaluation requires a different protocol: replaying the same temporal intervention across different persistent user-conditioned states and measuring how failures propagate across agent components. We formalize this requirement as four conditions: explicit temporal intervention, persistent state across the intervention, induced cross-dimensional effects, and variation in user-conditioned state. A focused audit of public benchmark protocols selected by explicit inclusion criteria identifies several close cases. Under our explicitly narrow operationalization, we did not find a protocol in that audited set satisfying all four conditions. This claim is scoped as a focused gap analysis with bounded literature coverage. This position paper proposes a minimal benchmark design and candidate reporting metrics for user-conditioned adaptation. The result is a concrete design requirement for future personal-agent evaluation, with metrics used as reporting tools for that requirement.

Figures

Figures reproduced from arXiv: 2607.21635 by Junxian You, Pin Qian, Qiaolin Yu, Su Wang, Xiaoyuan Wang, Yihang Chen, Zhicheng Wang, Zhitong Guo.

Figure 1
Figure 1. Figure 1: Broader landscape of 24 unique benchmarks/systems related to adaptive agent evaluation, organized by adaptation [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: The 5 × 4 evaluation matrix. Each cell shows how many of 15 audited benchmarks have at least partial cover￾age of both the row dimension and column aspect. Dark cells indicate adequate coverage (≥3); light cells indicate sparse coverage (1–2); white cells are complete gaps (0). Fixed￾intervention cross-dimensional adaptation is unobserved in the audited set under C1–C4. denominator spans tool use [9, 27, 3… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

66 extracted references · 3 canonical work pages · 2 internal anchors

  1. [1]

    Maksym Andriushchenko, Alexandra Souly, Mateusz Dziemian, Derek Duenas, Maxwell Lin, Justin Wang, Dan Hendrycks, Andy Zou, Zico Kolter, Matt Fredrik- son, et al. 2025. AgentHarm: A Benchmark for Measuring Harmfulness of LLM Qian et al. Agents. International Conference on Learning Representations

  2. [2]

    Ahmed Nusayer Ashik, Shaowei Wang, Tse-Hsun Chen, Muhammad Asaduzza- man, and Yuan Tian. 2026. When LLMs Lag Behind: Knowledge Conflicts from Evolving APIs in Code Generation. arXiv preprint arXiv:2604.09515

  3. [3]

    Zhiyuan Cheng, Longying Lai, and Yue Liu. 2026. Resolving the Robustness- Precision Trade-off in Financial RAG through Hybrid Document-Routed Re- trieval. https://doi.org/10.48550/arXiv.2603.26815 arXiv:2603.26815 [cs.CL]

  4. [4]

    Zhiyuan Cheng, Longying Lai, Yue Liu, and Yu Sun. 2026. Toward Sustainable On-Device Intelligence: A Survey on Energy-Efficient RAG Systems with Small Language Models.A vailable at SSRN 6698538(2026). https://ssrn.com/abstract= 6698538

  5. [5]

    Pengfei Du. 2026. Memory for Autonomous LLM Agents: Mechanisms, Evalua- tion, and Emerging Frontiers. arXiv preprint arXiv:2603.07670

  6. [6]

    K. M. Ferdous, Dipayan Banik, Kowshik Chowdhury, and Shazibul Islam Shamim

  7. [7]

    Huan-ang Gao, Jiayi Geng, Wenyue Hua, Mengkang Hu, Xinzhe Juan, Hongzhang Liu, Shilong Liu, Jiahao Qiu, Xuan Qi, Qihan Ren, Yiran Wu, Hongru Wang, Han Xiao, Yuhang Zhou, Shaokun Zhang, Jiayi Zhang, Jinyu Xiang, Yixiong Fang, Qiwen Zhao, Dongrui Liu, Cheng Qian, Zhenhailong Wang, Minda Hu, Huazheng Wang, Qingyun Wu, Heng Ji, and Mengdi Wang. 2026. A Survey...

  8. [8]

    Dongxin Guo, Jikun Wu, and Siu Ming Yiu. 2026. AgentEval: DAG-Structured Step-Level Evaluation for Agentic Workflows with Error Propagation Tracking. arXiv preprint arXiv:2604.23581

  9. [9]

    Zhicheng Guo, Sijie Cheng, Hao Wang, Shihao Liang, Yujia Qin, Peng Li, Zhiyuan Liu, Maosong Sun, and Yang Liu. 2024. StableToolBench: Towards Stable Large- Scale Benchmarking on Tool Learning of Large Language Models. Findings of the Association for Computational Linguistics: ACL 2024

  10. [10]

    Aayush Gupta. 2026. ReliabilityBench: Evaluating LLM Agent Reliability Under Production-Like Stress Conditions. arXiv preprint arXiv:2601.06112

  11. [11]

    Xiao Han, Yao Xiao, Chenyu Wu, and Tongchen Zhang. 2026. How Early Is Early Enough? Design-Dependent Observation-Window Sufficiency in Subscription Churn Prediction. https://doi.org/10.48550/arXiv.2607.00473 arXiv:2607.00473 [cs.LG]

  12. [12]

    Zexue He, Yu Wang, Churan Zhi, Yuanzhe Hu, Tzu-Ping Chen, Lang Yin, Ze Chen, Tong Arthur Wu, Siru Ouyang, Zihan Wang, Jiaxin Pei, Julian McAuley, Yejin Choi, and Alex Pentland. 2026. MemoryArena: Benchmarking Agent Memory in Interdependent Multi-Session Agentic Tasks. arXiv preprint arXiv:2602.16313

  13. [13]

    Yuanzhe Hu, Yu Wang, and Julian McAuley. 2026. Evaluating Memory in LLM Agents via Incremental Multi-Turn Interactions. International Conference on Learning Representations

  14. [14]

    Wenyue Hua, Xianjun Yang, Mingyu Jin, Zelong Li, Wei Cheng, Ruixiang Tang, and Yongfeng Zhang. 2024. TrustAgent: Towards Safe and Trustworthy LLM- based Agents. Conference on Empirical Methods in Natural Language Processing. https://doi.org/10.48550/arXiv.2402.01586 arXiv:2402.01586 [cs.CL]

  15. [15]

    Yuxuan Jiang and Francis Ferraro. 2026. SCRIBE: Structured Mid-Level Supervi- sion for Tool-Using Language Models.arXiv preprint arXiv:2601.03555(2026). https://doi.org/10.48550/arXiv.2601.03555 arXiv:2601.03555 [cs.AI]

  16. [16]

    Yanna Jiang, Delong Li, Haiyu Deng, Baihe Ma, Xu Wang, Qin Wang, and Guang- sheng Yu. 2026. SoK: Agentic Skills – Beyond Tool Use in LLM Agents. arXiv preprint arXiv:2602.20867

  17. [17]

    Wang-Cheng Kang and Julian McAuley. 2018. Self-Attentive Sequential Recom- mendation. IEEE International Conference on Data Mining

  18. [18]

    Su Kara, Fazle Faisal, and Suman Nath. 2026. WAREX: Web Agent Reliabil- ity Evaluation on Existing Benchmarks. Transactions on Machine Learning Research

  19. [19]

    Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, et al

    James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A. Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, et al. 2017. Overcoming Catastrophic Forgetting in Neural Networks.Proceedings of the National Academy of Sciences114, 13 (2017), 3521– 3526

  20. [20]

    Cottereau, Ziwei Liu, Tat-Seng Chua, and Wei Tsang Ooi

    Lingdong Kong, Xian Sun, Wei Chow, Linfeng Li, Kevin Qinghong Lin, Xuan Billy Zhang, Song Wang, Rong Li, Qing Wu, Wei Gao, Yingshuo Wang, Shaoyuan Xie, Jiachen Liu, Leigang Qu, Shijie Li, Lai Xing Ng, Benoit R. Cottereau, Ziwei Liu, Tat-Seng Chua, and Wei Tsang Ooi. 2026. AI for Auto-Research: Roadmap & User Guide. arXiv preprint arXiv:2605.18661. https:/...

  21. [21]

    Yehuda Koren. 2009. Collaborative Filtering with Temporal Dynamics. InProceed- ings of the 15th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. Association for Computing Machinery, New York, NY, USA, 447–456

  22. [22]

    Inan, Sahar Abdelnabi, Janardhan Kulkarni, Lukas Wutschitz, Reza Shokri, Christopher G

    Guangchen Lan, Huseyin A. Inan, Sahar Abdelnabi, Janardhan Kulkarni, Lukas Wutschitz, Reza Shokri, Christopher G. Brinton, and Robert Sim. 2025. Con- textual Integrity in LLMs via Reasoning and Reinforcement Learning. In The Thirty-Ninth Annual Conference on Neural Information Processing Systems. arXiv:2506.04245 [cs.AI]

  23. [23]

    Guangchen Lan, Lian Xiong, Xin Zhou, Hejie Cui, Yuwei Zhang, Mao Li, Zhenyu Shi, Besnik Fetahu, Lihong Li, and Xian Li. 2026. Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy. arXiv preprint arXiv:2603.15646(2026). https://doi.org/10.48550/arXiv.2603.15646 arXiv:2603.15646 [cs.LG]

  24. [24]

    Ido Levy, Ben Wiesel, Sami Marreed, Alon Oved, Avi Yaeli, Nir Mashkif, and Segev Shlomov. 2026. ST-WebAgentBench: A Benchmark for Evaluating Safety and Trustworthiness in Web Agents. arXiv:2410.06703 [cs.AI] International Conference on Learning Representations (ICLR)

  25. [25]

    Dawei Li, Bohan Jiang, Liangjie Huang, Alimohammad Beigi, Chengshuai Zhao, Zhen Tan, Amrita Bhattacharjee, Yuxuan Jiang, Canyu Chen, Tianhao Wu, Kai Shu, Lu Cheng, and Huan Liu. 2025. From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judge. InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. Associ...

  26. [26]

    Jialin Li, Zhenhao Chen, Hanjun Luo, and Hanan Salam. 2026. PrefIx: Understand and Adapt to User Preference in Human-Agent Interaction. https://doi.org/10. 48550/arXiv.2602.06714 arXiv:2602.06714 [cs.HC]

  27. [27]

    Minghao Li, Yingxiu Zhao, Bowen Yu, Feifan Song, Hangyu Li, Haiyang Yu, Zhoujun Li, Fei Huang, and Yongbin Li. 2023. API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. https://doi.org/10.18653/ v1/2023.emnlp-main.187 arXiv:2304.08244 [cs.CL]

  28. [28]

    Yanshu Li, Jiaqian Li, Kuai Yu, Xi Xiao, Dongfang Liu, Tianyang Wang, and Ruixiang Tang. 2026. Personalize Your Large Vision-Language Models With In-Context Prompt Tuning. arXiv preprint arXiv:2605.31513. https://doi.org/10. 48550/arXiv.2605.31513 arXiv:2605.31513 [cs.CV]

  29. [29]

    Luyun Lin, Lixing Lin, Zhen Zhang, Moxuan Zheng, and Yiqing Wang. 2026. A Volume-Price-Adjusted MACD Trading Strategy with Sensitivity Calibration for U.S. Equity Indices. arXiv preprint arXiv:2604.26063. https://doi.org/10.48550/ arXiv.2604.26063 arXiv:2604.26063 [q-fin.TR]

  30. [30]

    Lixing Lin, Juli You, Yue Li, Luyun Lin, Yiqing Wang, Zhen Zhang, and Moxuan Zheng. 2026. Reflect-Guard: Enhancing LLM Safeguards against Adversarial Prompts via Logical Self-Reflection.arXiv preprint arXiv:2605.24834(2026). https: //doi.org/10.48550/arXiv.2605.24834 arXiv:2605.24834 [cs.CR]

  31. [31]

    Jiayuan Liu, Tianqin Li, Shiyi Du, Xin Luo, Haoxuan Zeng, Emanuel Tewolde, Tai Sing Lee, Tonghan Wang, Carl Kingsford, and Vincent Conitzer. 2026. The Memory Curse: How Expanded Recall Erodes Cooperative Intent in LLM Agents. arXiv preprint arXiv:2605.08060. https://doi.org/10.48550/arXiv.2605.08060 arXiv:2605.08060 [cs.CL]

  32. [32]

    Xiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu, Xuanyu Lei, Hanyu Lai, Yu Gu, Hangliang Ding, Kaiwen Men, Kejuan Yang, et al. 2024. AgentBench: Evaluating LLMs as Agents. International Conference on Learning Representations

  33. [33]

    Yue Liu, Zhiyuan Cheng, and Longying Lai. 2026. Improving the Completeness and Comparability of Segment Disclosures: A Large Language Model Approach. https://doi.org/10.48550/arXiv.2605.23924 arXiv:2605.23924 [cs.CL]

  34. [34]

    Jiarui Lu, Thomas Holleis, Yizhe Zhang, Bernhard Aumayer, Feng Nan, Haoping Bai, Shuang Ma, Shen Ma, Mengyu Li, Guoli Yin, Zirui Wang, and Ruoming Pang. 2025. ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities. Findings of the Association for Computational Linguistics: NAACL 2025

  35. [35]

    Xing Han Lù, Zdeněk Kasner, and Siva Reddy. 2024. WebLINX: Real-World Website Navigation with Multi-Turn Dialogue. International Conference on Machine Learning

  36. [36]

    Bodhisattwa Prasad Majumder, Bhavana Dalvi Mishra, Peter Jansen, Oyvind Tafjord, Niket Tandon, Li Zhang, Chris Callison-Burch, and Peter Clark. 2024. CLIN: A Continually Learning Language Agent for Rapid Task Adaptation and Generalization. Conference on Language Modeling

  37. [37]

    Mahmoud Mohammadi, Yipeng Li, Jane Lo, and Wendy Yip. 2025. Evaluation and Benchmarking of LLM Agents: A Survey. Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining

  38. [38]

    Patil, Ion Stoica, and Joseph E

    Charles Packer, Sarah Wooders, Kevin Lin, Vivian Fang, Shishir G. Patil, Ion Stoica, and Joseph E. Gonzalez. 2024. MemGPT: Towards LLMs as Operating Systems. International Conference on Learning Representations

  39. [39]

    Patil, Huanzhi Mao, Fanjia Yan, Charlie Cheng-Jie Ji, Vishnu Suresh, Ion Stoica, and Joseph E

    Shishir G. Patil, Huanzhi Mao, Fanjia Yan, Charlie Cheng-Jie Ji, Vishnu Suresh, Ion Stoica, and Joseph E. Gonzalez. 2025. The Berkeley Function Calling Leader- board (BFCL): From Tool Use to Agentic Evaluation of Large Language Models. International Conference on Machine Learning

  40. [40]

    Yujia Qin, Shihao Liang, Yining Ye, Kunlun Zhu, Lan Yan, Yaxi Lu, Yankai Lin, Xin Cong, Xiangru Tang, Bill Qian, et al . 2024. ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs. International Conference on Learning Representations

  41. [41]

    Tran, Jonah Samost, Maciej Kula, Ed Chi, and Maheswaran Sathiamoorthy

    Shashank Rajput, Nikhil Mehta, Anima Singh, Raghunandan Hulikal Keshavan, Trung Vu, Lukasz Heldt, Lichan Hong, Yi Tay, Vinh Q. Tran, Jonah Samost, Maciej Kula, Ed Chi, and Maheswaran Sathiamoorthy. 2023. Recommender Systems with Generative Retrieval. Advances in Neural Information Processing Systems

  42. [42]

    Shaina Raza, Ranjan Sapkota, Manoj Karkee, and Christos Emmanouilidis. 2025. TRiSM for Agentic AI: A Review of Trust, Risk, and Security Management in LLM-based Agentic Multi-Agent Systems. arXiv preprint arXiv:2506.04133

  43. [43]

    Maddison, and Tatsunori Hashimoto

    Yangjun Ruan, Honghua Dong, Andrew Wang, Silviu Pitis, Yongchao Zhou, Jimmy Ba, Yann Dubois, Chris J. Maddison, and Tatsunori Hashimoto. 2024. Identifying the Risks of LM Agents with an LM-Emulated Sandbox. International Conference on Learning Representations. Toward User-Conditioned Evaluation of Personal LLM Agents under Temporal Interventions

  44. [44]

    Xian Sun, Wei Gao, Yingshuo Wang, Lingdong Kong, Yanhang Li, Zhichao Fan, Zexin Zhuang, Wenlong Dong, Zhiyuan Zheng, Hrishikesh Paranjape, Abhishek Mandal, and Johnny R. Zhang. 2026. Beyond Accuracy: Measuring Bias Acknowledgment in Chain-of-Thought Reasoning for Responsible AI Evaluation. https://doi.org/10.48550/arXiv.2606.15127 arXiv:2606.15127 [cs.LG]...

  45. [45]

    Haoran Tan, Zeyu Zhang, Chen Ma, Xu Chen, Quanyu Dai, and Zhenhua Dong. 2025. MemBench: Towards More Comprehensive Evaluation on the Memory of LLM-based Agents. InFindings of the Association for Computa- tional Linguistics: ACL 2025. Association for Computational Linguistics, Vi- enna, Austria, 19336–19352. https://doi.org/10.18653/v1/2025.findings-acl.98...

  46. [46]

    Yicheng Tao, Yiqun Wang, Xiangchen Song, Xin Luo, Kai Liu, and Jie Liu. 2026. GRASP: Plan-Guided Graph Retrieval with Adaptive Fusion and Reranking on Semi-Structured Knowledge Bases.arXiv preprint arXiv:2605.30237(2026). https://doi.org/10.48550/arXiv.2605.30237 arXiv:2605.30237 [cs.IR]

  47. [47]

    Guanzhi Wang, Yuqi Xie, Yunfan Jiang, Ajay Mandlekar, Chaowei Xiao, Yuke Zhu, Linxi Fan, and Anima Anandkumar. 2024. Voyager: An Open-Ended Em- bodied Agent with Large Language Models. Transactions on Machine Learning Research

  48. [48]

    Yingshuo Wang, Xian Sun, Lingdong Kong, Wei Gao, Yanhang Li, Zhichao Fan, and Zexin Zhuang. 2026. Do Time Series Foundation Model Benchmarks Hide Regime-Dependent Failures? Evidence from Traffic Speed Forecasting. https://doi.org/10.48550/arXiv.2606.18367 arXiv:2606.18367 [cs.LG] Accepted at the Workshop on Forecasting as a New Frontier of Intelligence, ICML 2026

  49. [49]

    Yingshuo Wang, Xian Sun, Yanhang Li, Zhichao Fan, and Zexin Zhuang. 2026. Em- bedding Foundation Model Predictions in Discrete-Choice Models with Structural Guarantees. https://doi.org/10.48550/arXiv.2606.26432 arXiv:2606.26432 [cs.LG]

  50. [50]

    Chenyu Wu. 2026. Class Weighting versus Amount Conditioning in Credit-Card Fraud Detection: A Dollar-Metric Study with a Temporal Explanation Audit. https://doi.org/10.48550/arXiv.2607.14686 arXiv:2607.14686 [cs.CE]

  51. [51]

    Rong Wu, Xiaoman Wang, Jianbiao Mei, Pinlong Cai, Daocheng Fu, Cheng Yang, Licheng Wen, Xuemeng Yang, Yufan Shen, Yuxin Wang, and Botian Shi. 2025. EvolveR: Self-Evolving LLM Agents through an Experience-Driven Lifecycle. arXiv preprint arXiv:2510.16079

  52. [52]

    Ningning Xu, Yuxuan Jiang, Shubhashis Roy Dipta, and Hengyuan Zhang. 2025. Learning How to Use Tools, Not Just When: Pattern-Aware Tool-Integrated Reasoning. NeurIPS 2025 MATH-AI Workshop. https://doi.org/10.48550/arXiv. 2509.23292 arXiv:2509.23292 [cs.AI]

  53. [53]

    Renjun Xu and Yang Yan. 2026. Agent Skills for Large Language Models: Architecture, Acquisition, Security, and the Path Forward. arXiv preprint arXiv:2602.12430

  54. [54]

    Yutao Yang, Junsong Li, Qianjun Pan, Bihao Zhan, Yuxuan Cai, Lin Du, Jie Zhou, Kai Chen, Qin Chen, Xin Li, Bo Zhang, and Liang He. 2026. AutoSkill: Experience-Driven Lifelong Learning via Skill Self-Evolution. arXiv preprint arXiv:2603.01145

  55. [55]

    Shunyu Yao, Noah Shinn, Pedram Razavi, and Karthik Narasimhan. 2025. 𝜏- bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains. International Conference on Learning Representations

  56. [56]

    Asaf Yehudai, Lilach Eden, Alan Li, Guy Uziel, Yilun Zhao, Roy Bar-Haim, Arman Cohan, and Michal Shmueli-Scheuer. 2025. Survey on Evaluation of LLM-based Agents. arXiv preprint arXiv:2503.16416

  57. [57]

    Shin Yoo and Mark Harman. 2012. Regression Testing Minimization, Selection and Prioritization: A Survey.Software Testing, Verification and Reliability22, 2 (2012), 67–120

  58. [58]

    Miao Yu, Fanci Meng, Xinyun Zhou, Shilong Wang, Junyuan Mao, Linsey Pang, Tianlong Chen, Kun Wang, Xinfeng Li, Yongfeng Zhang, Bo An, and Qingsong Wen. 2025. A Survey on Trustworthy LLM Agents: Threats and Countermeasures. arXiv preprint arXiv:2503.09648

  59. [59]

    Yike Zhang, Zuodong Xiang, and Hailu Xu. 2026. Performance-Efficiency Trade- Offs in Human Preference Prediction: A Comparative Study of Traditional Ma- chine Learning and Large Language Models. Proceedings of the 31st IEEE Symposium on Computers and Communications (ISCC)

  60. [60]

    Zhexin Zhang, Shiyao Cui, Yida Lu, Jingzhuo Zhou, Junxiao Yang, Hongning Wang, and Minlie Huang. 2024. Agent-SafetyBench: Evaluating the Safety of LLM Agents. arXiv preprint arXiv:2412.14470

  61. [61]

    Yujie Zhao, Boqin Yuan, Junbo Huang, Haocheng Yuan, Zhongming Yu, Haozhou Xu, Lanxiang Hu, Abhilash Shankarampeta, Zimeng Huang, Wentao Ni, Yuan- dong Tian, and Jishen Zhao. 2026. AMA-Bench: Evaluating Long-Horizon Mem- ory for Agentic Applications. arXiv preprint arXiv:2602.22769

  62. [62]

    Junhao Zheng, Chengming Shi, Xidi Cai, Qiuke Li, Duzhen Zhang, Chenxing Li, Dong Yu, and Qianli Ma. 2026. Lifelong Learning of Large Language Model based Agents: A Roadmap.IEEE Transactions on Pattern Analysis and Machine Intelligence48, 5 (May 2026), 5552–5571. https://doi.org/10.1109/TPAMI.2025. 3650546 arXiv:2501.07278 [cs.AI]

  63. [63]

    Lucen Zhong, Zhengxiao Du, Xiaohan Zhang, Haiyi Hu, and Jie Tang. 2025. ComplexFuncBench: Exploring Multi-Step and Constrained Function Calling under Long-Context Scenario. arXiv preprint arXiv:2501.10132

  64. [64]

    Shanshan Zhong, Yi Lu, Jingjie Ning, Yibing Wan, Lihan Feng, Yuyi Ao, Leonardo F. R. Ribeiro, Markus Dreyer, Sean Ammirati, and Chenyan Xiong. 2026. Skill- LearnBench: Benchmarking Continual Learning Methods for Agent Skill Gener- ation on Real-World Tasks. arXiv preprint arXiv:2604.20087

  65. [65]

    Xinxue Zhu, Jiacong Wu, Xiaoyu Zhang, Tianlin Li, Yanzhou Mu, Juan Zhai, Chao Shen, Chunrong Fang, and Yang Liu. 2026. An Empirical Study of Bugs in Modern LLM Agent Frameworks. https://doi.org/10.48550/arXiv.2602.21806 arXiv:2602.21806 [cs.SE]

  66. [2026]

    https://doi.org/10.48550/arXiv.2603.27524 arXiv:2603.27524 [cs.SE] Accepted at the 23rd International Conference on Mining Software Repositories (MSR)

    Safer Builders, Risky Maintainers: A Comparative Study of Breaking Changes in Human vs Agentic PRs. https://doi.org/10.48550/arXiv.2603.27524 arXiv:2603.27524 [cs.SE] Accepted at the 23rd International Conference on Mining Software Repositories (MSR)