{"paper":{"title":"Skill-Pro: Learning Reusable Skills from Experience via Non-Parametric PPO for LLM Agents","license":"http://creativecommons.org/licenses/by/4.0/","headline":"Skill-Pro lets LLM agents learn reusable procedural skills from past experiences without updating any parameters.","cross_cats":[],"primary_cat":"cs.AI","authors_text":"Haifeng Zhang, Haoxuan Li, Jun Wang, Mengyue Yang, Qirui Mi, Yisen Wang, Zhijian Ma","submitted_at":"2026-02-02T09:43:12Z","abstract_excerpt":"LLM-driven agents excel at sequential decision-making but often rely on on-the-fly reasoning, re-deriving solutions even in recurring scenarios. This insufficient experience reuse leads to computational redundancy and instability. To bridge this gap, we propose Skill-Pro, a framework enabling agents to autonomously learn reusable procedural skills from interaction experiences without parameter updates. By formalizing a Skill-MDP, Skill-Pro transforms passive episodic narratives into executable Skills defined by activation, execution, and termination conditions to ensure executability. To achie"},"claims":{"count":4,"items":[{"kind":"strongest_claim","text":"Skill-Pro achieves superior reuse rates and significant performance gains with extreme memory compression by autonomously learning reusable procedural skills from interaction experiences without parameter updates, using Skill-MDP formalization and Non-Parametric PPO for reliable verification.","source":"verdict.strongest_claim","status":"machine_extracted","claim_id":"C1","attestation":"unclaimed"},{"kind":"weakest_assumption","text":"That semantic gradients combined with a PPO Gate can generate and verify high-quality skills that remain executable and non-degrading across recurring scenarios without any post-hoc exclusions or capability loss.","source":"verdict.weakest_assumption","status":"machine_extracted","claim_id":"C2","attestation":"unclaimed"},{"kind":"one_line_summary","text":"Skill-Pro enables LLM agents to autonomously learn, verify, and reuse compact procedural skills from experience via a Skill-MDP formalization and Non-Parametric PPO, yielding higher reuse rates, performance gains, and extreme memory compression across in-domain, cross-task, and cross-agent settings.","source":"verdict.one_line_summary","status":"machine_extracted","claim_id":"C3","attestation":"unclaimed"},{"kind":"headline","text":"Skill-Pro lets LLM agents learn reusable procedural skills from past experiences without updating any parameters.","source":"verdict.pith_extraction.headline","status":"machine_extracted","claim_id":"C4","attestation":"unclaimed"}],"snapshot_sha256":"cfdb92aac6ed8d7fea254fb87fc5088033ab64844870fb551b448d279863cb37"},"source":{"id":"2602.01869","kind":"arxiv","version":3},"verdict":{"id":"de1adf4e-d978-4769-bb94-7f76d07fb912","model_set":{"reader":"grok-4.3"},"created_at":"2026-05-16T08:31:09.441739Z","strongest_claim":"Skill-Pro achieves superior reuse rates and significant performance gains with extreme memory compression by autonomously learning reusable procedural skills from interaction experiences without parameter updates, using Skill-MDP formalization and Non-Parametric PPO for reliable verification.","one_line_summary":"Skill-Pro enables LLM agents to autonomously learn, verify, and reuse compact procedural skills from experience via a Skill-MDP formalization and Non-Parametric PPO, yielding higher reuse rates, performance gains, and extreme memory compression across in-domain, cross-task, and cross-agent settings.","pipeline_version":"pith-pipeline@v0.9.0","weakest_assumption":"That semantic gradients combined with a PPO Gate can generate and verify high-quality skills that remain executable and non-degrading across recurring scenarios without any post-hoc exclusions or capability loss.","pith_extraction_headline":"Skill-Pro lets LLM agents learn reusable procedural skills from past experiences without updating any parameters."},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2602.01869/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"}