{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2025:TIXCOYNHE7YB7V6RCGZ4OGPWJA","short_pith_number":"pith:TIXCOYNH","schema_version":"1.0","canonical_sha256":"9a2e2761a727f01fd7d111b3c719f6483e60e2aecb02764a7bd84c0aa2ecd824","source":{"kind":"arxiv","id":"2509.09265","version":1},"attestation_state":"computed","paper":{"title":"Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.CL"],"primary_cat":"cs.LG","authors_text":"Jiacai Liu, Jiawei Wang, Ke Wang, Lin Zhang, Xintao Wang, Yang Wang, Yingru Li, Yuan Lin, Yuqian Fu, Yu Yue","submitted_at":"2025-09-11T08:50:01Z","abstract_excerpt":"In long-horizon tasks, recent agents based on Large Language Models (LLMs) face a significant challenge that sparse, outcome-based rewards make it difficult to assign credit to intermediate steps. Previous methods mainly focus on creating dense reward signals to guide learning, either through traditional reinforcement learning techniques like inverse reinforcement learning or by using Process Reward Models for step-by-step feedback. In this paper, we identify a fundamental problem in the learning dynamics of LLMs: the magnitude of policy gradients is inherently coupled with the entropy, which "},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2509.09265","kind":"arxiv","version":1},"metadata":{"license":"http://creativecommons.org/licenses/by/4.0/","primary_cat":"cs.LG","submitted_at":"2025-09-11T08:50:01Z","cross_cats_sorted":["cs.CL"],"title_canon_sha256":"3e797667fc2ee02d7b2d2f1329f8f7c5c553f2c41b3fe190600bf466c81659b5","abstract_canon_sha256":"8365a2c704d6f63c4ee5ba4203337c2540a69b986f85b326e1febeedfb1e5a64"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T12:09:32.755815Z","signature_b64":"mti6w7qeG9jK24GNVri95yFjmtyH4KFrMOd1eGgTPtaA0xqjnsONrQKUB7ZRq0huFQBdNS9kxyvz3lMTbLROBw==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"9a2e2761a727f01fd7d111b3c719f6483e60e2aecb02764a7bd84c0aa2ecd824","last_reissued_at":"2026-07-05T12:09:32.755249Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T12:09:32.755249Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.CL"],"primary_cat":"cs.LG","authors_text":"Jiacai Liu, Jiawei Wang, Ke Wang, Lin Zhang, Xintao Wang, Yang Wang, Yingru Li, Yuan Lin, Yuqian Fu, Yu Yue","submitted_at":"2025-09-11T08:50:01Z","abstract_excerpt":"In long-horizon tasks, recent agents based on Large Language Models (LLMs) face a significant challenge that sparse, outcome-based rewards make it difficult to assign credit to intermediate steps. Previous methods mainly focus on creating dense reward signals to guide learning, either through traditional reinforcement learning techniques like inverse reinforcement learning or by using Process Reward Models for step-by-step feedback. In this paper, we identify a fundamental problem in the learning dynamics of LLMs: the magnitude of policy gradients is inherently coupled with the entropy, which "},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2509.09265","kind":"arxiv","version":1},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2509.09265/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2509.09265","created_at":"2026-07-05T12:09:32.755320+00:00"},{"alias_kind":"arxiv_version","alias_value":"2509.09265v1","created_at":"2026-07-05T12:09:32.755320+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2509.09265","created_at":"2026-07-05T12:09:32.755320+00:00"},{"alias_kind":"pith_short_12","alias_value":"TIXCOYNHE7YB","created_at":"2026-07-05T12:09:32.755320+00:00"},{"alias_kind":"pith_short_16","alias_value":"TIXCOYNHE7YB7V6R","created_at":"2026-07-05T12:09:32.755320+00:00"},{"alias_kind":"pith_short_8","alias_value":"TIXCOYNH","created_at":"2026-07-05T12:09:32.755320+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":10,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.25852","citing_title":"Semantic Consistency Policy Optimization for Reinforcement Learning of LLM Agents","ref_index":33,"is_internal_anchor":false},{"citing_arxiv_id":"2605.06130","citing_title":"Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning","ref_index":89,"is_internal_anchor":false},{"citing_arxiv_id":"2605.06130","citing_title":"Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning","ref_index":89,"is_internal_anchor":false},{"citing_arxiv_id":"2605.00380","citing_title":"ResRL: Boosting LLM Reasoning via Negative Sample Projection Residual Reinforcement Learning","ref_index":23,"is_internal_anchor":false},{"citing_arxiv_id":"2605.00425","citing_title":"AEM: Adaptive Entropy Modulation for Multi-Turn Agentic Reinforcement Learning","ref_index":46,"is_internal_anchor":false},{"citing_arxiv_id":"2604.07774","citing_title":"RoboAgent: Chaining Basic Capabilities for Embodied Task Planning","ref_index":107,"is_internal_anchor":false},{"citing_arxiv_id":"2605.00425","citing_title":"AEM: Adaptive Entropy Modulation for Multi-Turn Agentic Reinforcement Learning","ref_index":46,"is_internal_anchor":false},{"citing_arxiv_id":"2605.06130","citing_title":"Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning","ref_index":89,"is_internal_anchor":false},{"citing_arxiv_id":"2605.00380","citing_title":"ResRL: Boosting LLM Reasoning via Negative Sample Projection Residual Reinforcement Learning","ref_index":23,"is_internal_anchor":false},{"citing_arxiv_id":"2605.02178","citing_title":"T$^2$PO: Uncertainty-Guided Exploration Control for Stable Multi-Turn Agentic Reinforcement Learning","ref_index":21,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/TIXCOYNHE7YB7V6RCGZ4OGPWJA","json":"https://pith.science/pith/TIXCOYNHE7YB7V6RCGZ4OGPWJA.json","graph_json":"https://pith.science/api/pith-number/TIXCOYNHE7YB7V6RCGZ4OGPWJA/graph.json","events_json":"https://pith.science/api/pith-number/TIXCOYNHE7YB7V6RCGZ4OGPWJA/events.json","paper":"https://pith.science/paper/TIXCOYNH"},"agent_actions":{"view_html":"https://pith.science/pith/TIXCOYNHE7YB7V6RCGZ4OGPWJA","download_json":"https://pith.science/pith/TIXCOYNHE7YB7V6RCGZ4OGPWJA.json","view_paper":"https://pith.science/paper/TIXCOYNH","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2509.09265&json=true","fetch_graph":"https://pith.science/api/pith-number/TIXCOYNHE7YB7V6RCGZ4OGPWJA/graph.json","fetch_events":"https://pith.science/api/pith-number/TIXCOYNHE7YB7V6RCGZ4OGPWJA/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/TIXCOYNHE7YB7V6RCGZ4OGPWJA/action/timestamp_anchor","attest_storage":"https://pith.science/pith/TIXCOYNHE7YB7V6RCGZ4OGPWJA/action/storage_attestation","attest_author":"https://pith.science/pith/TIXCOYNHE7YB7V6RCGZ4OGPWJA/action/author_attestation","sign_citation":"https://pith.science/pith/TIXCOYNHE7YB7V6RCGZ4OGPWJA/action/citation_signature","submit_replication":"https://pith.science/pith/TIXCOYNHE7YB7V6RCGZ4OGPWJA/action/replication_record"}},"created_at":"2026-07-05T12:09:32.755320+00:00","updated_at":"2026-07-05T12:09:32.755320+00:00"}