{"total":10,"items":[{"citing_arxiv_id":"2607.07504","ref_index":18,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"Do LLM-Generated Skills Make Better AI Data Scientists? A Component Ablation Across Data-Science Workflows","primary_cat":"cs.AI","submitted_at":"2026-07-08T15:00:16+00:00","verdict":"ACCEPT","verdict_confidence":"HIGH","novelty_score":5.0,"formal_verification":"none","one_line_summary":"Across 56 tasks, 9 model configurations, and 10,584 runs, LLM-generated skill files provided no reliable performance improvement over task-only prompting for data-science workflows.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.01874","ref_index":2,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"SkillCoach: Self-Evolving Rubrics for Evaluating and Enhancing Agentic Skill-Use","primary_cat":"cs.AI","submitted_at":"2026-07-02T08:28:51+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"SkillCoach introduces self-evolving rubrics derived from rollouts to evaluate and supervise four process dimensions of agentic skill-use separately from outcome success.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.13317","ref_index":19,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"SkillCAT: Contrastive, Assessment-Augmented and Topology-AwareSkill Self-Evolution for LLM Agents","primary_cat":"cs.CL","submitted_at":"2026-06-11T13:12:10+00:00","verdict":"CONDITIONAL","verdict_confidence":"HIGH","novelty_score":6.0,"formal_verification":"none","one_line_summary":"Contrastive success/failure evidence, replay-based patch validation, and topology-aware routing improve training-free skill self-evolution for LLM agents.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.11543","ref_index":16,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"SkillJuror: Measuring How Agent Skill Organization Changes Runtime Behavior","primary_cat":"cs.AI","submitted_at":"2026-06-10T01:11:50+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"Empirical study finds Progressive Disclosure raises distinct resources touched (1.18 to 3.85) and uptake events (1.33 to 3.92) per trajectory, adds 17 passing trials out of 410 (+4.1%), with gains task-dependent.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.20659","ref_index":19,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"Skill Coverage: A Test Adequacy Metric for Agent Skills","primary_cat":"cs.AI","submitted_at":"2026-06-09T10:16:05+00:00","verdict":"CONDITIONAL","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"Skill coverage measures which natural-language skill constraints an LLM agent trajectory exercises and passes, revealing low coverage on SkillsBench and enabling a 16% recovery of failed tasks via targeted skill emphasis.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.09421","ref_index":33,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"What Should a Skill Remember? Quality--Cost Trade-offs in Cost-Aware Skill Rewriting for Language Model Agents","primary_cat":"cs.CL","submitted_at":"2026-06-08T12:36:51+00:00","verdict":"CONDITIONAL","verdict_confidence":"MODERATE","novelty_score":6.0,"formal_verification":"none","one_line_summary":"Cost-aware skill rewriting that preserves task-relevant operational anchors reduces LLM-agent total token cost by ~7% on a 20-task held-out panel and ~15% across agent stacks while maintaining verifier quality.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.08755","ref_index":46,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"Co-Evolving Skill Generation and Policy Optimization","primary_cat":"cs.CL","submitted_at":"2026-06-07T17:55:55+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"Framework estimates context-dependent marginal utility of candidate skills via reward gaps in matched base vs. skill-augmented rollouts to filter skills and co-train policy as generator.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.05661","ref_index":41,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"Continual Learning Bench: Evaluating Frontier AI Systems in Real-World Stateful Environments","primary_cat":"cs.AI","submitted_at":"2026-06-04T03:43:28+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":8.0,"formal_verification":"none","one_line_summary":"CL-Bench is the first expert-validated benchmark for continual learning in frontier LLMs across six real-world domains, showing limited gains and that naive in-context learning outperforms dedicated memory systems.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.27366","ref_index":39,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"MUSE-Autoskill: Self-Evolving Agents via Skill Creation, Memory, Management, and Evaluation","primary_cat":"cs.AI","submitted_at":"2026-05-26T17:59:19+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":5.5,"formal_verification":"none","one_line_summary":"A skill-lifecycle agent (create, memory, manage, evaluate, refine) beats Hermes, Codex, and Claude Code on SkillsBench/SkillLearnBench and transfers skills better.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.13716","ref_index":57,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"SkillOps: Managing LLM Agent Skill Libraries as Self-Maintaining Software Ecosystems","primary_cat":"cs.SE","submitted_at":"2026-05-13T16:02:25+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"SkillOps maintains LLM skill libraries via Skill Contracts and ecosystem graphs, raising ALFWorld task success to 79.5% as a standalone agent and improving retrieval baselines by up to 2.9 points with near-zero library-time LLM cost.","context_count":1,"top_context_role":"baseline","top_context_polarity":"baseline","context_text":"Results are reported for a library of 200 skills over three independent seeds. SR is reported as mean ± standard deviation across seeds, with Wilson 95% confidence intervals. Method SR mean±std Wilson 95% CI ReAct [Yao et al., 2023] 12.8%±1.90pp [10.3, 15.8] SkillWeaver [Zheng et al., 2025] 50.3%±1.43pp [46.1, 54.4] Hybrid_Retrieval 58.2%±0.83pp [54.1, 62.2] GoS_Style [Liu et al., 2026] 61.1%±0.94pp [57.0, 65.0] LLM_Skill_Planner 70.6% ±0.31pp [66.7, 74.3] SkillOps_Full (ours) 79.5%±0.00pp [75.9, 82.6] As shown in Table 1, we evaluate SkillOps as astandalone agenton ALFWorld with a 200-skill library. SkillOps achieves the best task success rate, reaching79.5%SR with zero standard deviation across three seeds. It outperforms the strongest baseline, LLM_Skill_Planner, by +8."}],"limit":50,"offset":0}