{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2025:HODV7EKFSHTMAY5EHWFYM37TCF","short_pith_number":"pith:HODV7EKF","schema_version":"1.0","canonical_sha256":"3b875f914591e6c063a43d8b866ff311508b0e801de4107aed9682c46d47e8bd","source":{"kind":"arxiv","id":"2509.08755","version":1},"attestation_state":"computed","paper":{"title":"AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.AI","cs.CL"],"primary_cat":"cs.LG","authors_text":"Baodai Huang, Chenyang Liao, Guanyu Li, Honglin Guo, Jiaqi Liu, Jiazheng Zhang, Jiecao Chen, Jixuan Huang, Junjie Ye, Qi Zhang, Rui Zheng, Tao Gui, Wei He, Wenxiang Chen, Xuanjing Huang, Xuesong Yao, Yiwen Ding, Yufei Xu, Yu-Gang Jiang, Zehui Chen, Zhengyin Du, Zhiheng Xi, Zuxuan Wu","submitted_at":"2025-09-10T16:46:11Z","abstract_excerpt":"Developing autonomous LLM agents capable of making a series of intelligent decisions to solve complex, real-world tasks is a fast-evolving frontier. Like human cognitive development, agents are expected to acquire knowledge and skills through exploration and interaction with the environment. Despite advances, the community still lacks a unified, interactive reinforcement learning (RL) framework that can effectively train such agents from scratch -- without relying on supervised fine-tuning (SFT) -- across diverse and realistic environments. To bridge this gap, we introduce AgentGym-RL, a new f"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2509.08755","kind":"arxiv","version":1},"metadata":{"license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","primary_cat":"cs.LG","submitted_at":"2025-09-10T16:46:11Z","cross_cats_sorted":["cs.AI","cs.CL"],"title_canon_sha256":"adf6cd3636adabfcd4166c4d1e17e5d8d49ceecfa2201e666742a2cfa69940ff","abstract_canon_sha256":"4f4d9e8b34fd2ff27aba8278152b6c1c7747cbc7f6dc704c920e668d6488d793"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T12:08:35.640083Z","signature_b64":"dRLECOcgLdWdJxNto+JP4x4K/BQNUyGvaMrtDrNW8IQ8NldjwOY3txM2QgC/0n6lllKZBHI1VUL+GBY+AWPOCg==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"3b875f914591e6c063a43d8b866ff311508b0e801de4107aed9682c46d47e8bd","last_reissued_at":"2026-07-05T12:08:35.639501Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T12:08:35.639501Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.AI","cs.CL"],"primary_cat":"cs.LG","authors_text":"Baodai Huang, Chenyang Liao, Guanyu Li, Honglin Guo, Jiaqi Liu, Jiazheng Zhang, Jiecao Chen, Jixuan Huang, Junjie Ye, Qi Zhang, Rui Zheng, Tao Gui, Wei He, Wenxiang Chen, Xuanjing Huang, Xuesong Yao, Yiwen Ding, Yufei Xu, Yu-Gang Jiang, Zehui Chen, Zhengyin Du, Zhiheng Xi, Zuxuan Wu","submitted_at":"2025-09-10T16:46:11Z","abstract_excerpt":"Developing autonomous LLM agents capable of making a series of intelligent decisions to solve complex, real-world tasks is a fast-evolving frontier. Like human cognitive development, agents are expected to acquire knowledge and skills through exploration and interaction with the environment. Despite advances, the community still lacks a unified, interactive reinforcement learning (RL) framework that can effectively train such agents from scratch -- without relying on supervised fine-tuning (SFT) -- across diverse and realistic environments. To bridge this gap, we introduce AgentGym-RL, a new f"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2509.08755","kind":"arxiv","version":1},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2509.08755/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2509.08755","created_at":"2026-07-05T12:08:35.639564+00:00"},{"alias_kind":"arxiv_version","alias_value":"2509.08755v1","created_at":"2026-07-05T12:08:35.639564+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2509.08755","created_at":"2026-07-05T12:08:35.639564+00:00"},{"alias_kind":"pith_short_12","alias_value":"HODV7EKFSHTM","created_at":"2026-07-05T12:08:35.639564+00:00"},{"alias_kind":"pith_short_16","alias_value":"HODV7EKFSHTMAY5E","created_at":"2026-07-05T12:08:35.639564+00:00"},{"alias_kind":"pith_short_8","alias_value":"HODV7EKF","created_at":"2026-07-05T12:08:35.639564+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":19,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2605.27760","citing_title":"SkillGrad: Optimizing Agent Skills Like Gradient Descent","ref_index":1,"is_internal_anchor":false},{"citing_arxiv_id":"2510.13727","citing_title":"From Refusal to Recovery: A Control-Theoretic Approach to Generative AI Guardrails","ref_index":64,"is_internal_anchor":false},{"citing_arxiv_id":"2605.20061","citing_title":"Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents","ref_index":39,"is_internal_anchor":false},{"citing_arxiv_id":"2605.15224","citing_title":"ICRL: Learning to Internalize Self-Critique with Reinforcement Learning","ref_index":35,"is_internal_anchor":false},{"citing_arxiv_id":"2601.06794","citing_title":"No More Stale Feedback: Co-Evolving Critics for Open-World Agent Learning","ref_index":18,"is_internal_anchor":false},{"citing_arxiv_id":"2603.00977","citing_title":"HiMAC: Hierarchical Macro-Micro Learning for Long-Horizon LLM Agents","ref_index":56,"is_internal_anchor":false},{"citing_arxiv_id":"2605.11775","citing_title":"Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control","ref_index":84,"is_internal_anchor":false},{"citing_arxiv_id":"2605.08715","citing_title":"AgentForesight: Online Auditing for Early Failure Prediction in Multi-Agent Systems","ref_index":56,"is_internal_anchor":false},{"citing_arxiv_id":"2605.14558","citing_title":"Resolving Action Bottleneck: Agentic Reinforcement Learning Informed by Token-Level Energy","ref_index":33,"is_internal_anchor":false},{"citing_arxiv_id":"2605.11775","citing_title":"Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control","ref_index":43,"is_internal_anchor":false},{"citing_arxiv_id":"2605.11706","citing_title":"GRAFT: Graph-Tokenized LLMs for Tool Planning","ref_index":20,"is_internal_anchor":false},{"citing_arxiv_id":"2605.12289","citing_title":"PriorZero: Bridging Language Priors and World Models for Decision Making","ref_index":26,"is_internal_anchor":false},{"citing_arxiv_id":"2605.08715","citing_title":"AgentForesight: Online Auditing for Early Failure Prediction in Multi-Agent Systems","ref_index":56,"is_internal_anchor":false},{"citing_arxiv_id":"2605.06642","citing_title":"StraTA: Incentivizing Agentic Reinforcement Learning with Strategic Trajectory Abstraction","ref_index":13,"is_internal_anchor":false},{"citing_arxiv_id":"2604.18975","citing_title":"Gated Coordination for Efficient Multi-Agent Collaboration in Minecraft Game","ref_index":35,"is_internal_anchor":false},{"citing_arxiv_id":"2604.18133","citing_title":"Multi-Agent Systems: From Classical Paradigms to Large Foundation Model-Enabled Futures","ref_index":99,"is_internal_anchor":false},{"citing_arxiv_id":"2605.06761","citing_title":"Weblica: Scalable and Reproducible Training Environments for Visual Web Agents","ref_index":39,"is_internal_anchor":false},{"citing_arxiv_id":"2604.20572","citing_title":"Ask Only When Needed: Proactive Retrieval from Memory and Skills for Experience-Driven Lifelong Agents","ref_index":31,"is_internal_anchor":false},{"citing_arxiv_id":"2605.02572","citing_title":"On Training Large Language Models for Long-Horizon Tasks: An Empirical Study of Horizon Length","ref_index":38,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/HODV7EKFSHTMAY5EHWFYM37TCF","json":"https://pith.science/pith/HODV7EKFSHTMAY5EHWFYM37TCF.json","graph_json":"https://pith.science/api/pith-number/HODV7EKFSHTMAY5EHWFYM37TCF/graph.json","events_json":"https://pith.science/api/pith-number/HODV7EKFSHTMAY5EHWFYM37TCF/events.json","paper":"https://pith.science/paper/HODV7EKF"},"agent_actions":{"view_html":"https://pith.science/pith/HODV7EKFSHTMAY5EHWFYM37TCF","download_json":"https://pith.science/pith/HODV7EKFSHTMAY5EHWFYM37TCF.json","view_paper":"https://pith.science/paper/HODV7EKF","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2509.08755&json=true","fetch_graph":"https://pith.science/api/pith-number/HODV7EKFSHTMAY5EHWFYM37TCF/graph.json","fetch_events":"https://pith.science/api/pith-number/HODV7EKFSHTMAY5EHWFYM37TCF/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/HODV7EKFSHTMAY5EHWFYM37TCF/action/timestamp_anchor","attest_storage":"https://pith.science/pith/HODV7EKFSHTMAY5EHWFYM37TCF/action/storage_attestation","attest_author":"https://pith.science/pith/HODV7EKFSHTMAY5EHWFYM37TCF/action/author_attestation","sign_citation":"https://pith.science/pith/HODV7EKFSHTMAY5EHWFYM37TCF/action/citation_signature","submit_replication":"https://pith.science/pith/HODV7EKFSHTMAY5EHWFYM37TCF/action/replication_record"}},"created_at":"2026-07-05T12:08:35.639564+00:00","updated_at":"2026-07-05T12:08:35.639564+00:00"}