Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T19:29:01.480681Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 14 inbound Pith citation observations for arXiv:2509.09265.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T19:29:01.480681Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T09:48:09.620197Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T20:20:07.648363Z
55 of 55 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 8aed9b11-2178-4d64-b947-6ad892f692a1 · outbound
Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents The Unreasonable Effectiveness of Entropy Minimization in LLM Reasoning
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6144339-f32c-440d-a59d-ee4fd68eaee3 · outbound
Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Open Deep Search: Democratizing Search with Open-source Reasoning Agents
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69199cbf-c8ef-4641-92ed-6d7e0819c1b2 · outbound
Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Unifying count-based exploration and intrinsic motivation
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e325a43-c496-45a7-97ee-128721ef66f1 · outbound
Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a05d97b5-0110-4c94-ad1b-28ebe8aa8c2d · outbound
Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Reasoning with Exploration: An Entropy Perspective
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de68e66a-70e6-4029-bada-9fd30d5028ef · outbound
Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Mind2web: Towards a generalist agent for the web
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90f14950-d6d6-49b2-a74a-88977f7ab7ac · outbound
Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents From Novice to Expert: LLM Agent Policy Optimization via Step-wise Reinforcement Learning
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df389e07-afe4-4d95-b7c9-51e573fb0dbc · outbound
Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Group-in-Group Policy Optimization for LLM Agent Training
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 596ef3ac-f757-494c-88ad-d727917f3bec · outbound
Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents One-shot Entropy Minimization
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c180ad2-db6c-4ea3-aa5a-3719bcc5a038 · outbound
Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents PaSa: An LLM Agent for Comprehensive Academic Paper Search
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a1376b0-98e0-453f-9a7b-6f45d307f65a · outbound
Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a98c37c0-4a0b-49d9-aec4-9c9c63dc83ec · outbound
Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a2f0c0c-7b53-403d-a0a3-9166689d0131 · outbound
Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38dc535f-1696-4f01-833e-38b59d342411 · outbound
Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2fd13494-fcc6-4011-89cc-4a00c826e275 · outbound
Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Vineppo: Accurate credit assignment in rl for llm mathematical reasoning
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30f6dae9-c4ec-4c99-b9ca-9a868bbf0e88 · outbound
Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Empowerment: A universal agent-centric measure of control
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35a7abdd-7c0c-4d50-a6f7-1171953a310c · outbound
Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Natural questions: a benchmark for question answering research
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f934fa58-2f68-4fe4-967a-97c41b32e0d2 · outbound
Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Search-o1: Agentic Search-Enhanced Large Reasoning Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1fb21e88-50dd-49a8-8d0b-b92dd067e492 · outbound
Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Logit Dynamics in Softmax Policy Gradient Methods
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d76f24e3-2abd-46f0-a4c9-b1c56c5755a7 · outbound
Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Let’s verify step by step
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13723c2e-df2d-4094-92f3-2598b454f8e0 · outbound
Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f632c4f-3d97-43b8-84de-1aa3165ac244 · outbound
Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Policy invariance under reward transformations: Theory and application to reward shaping
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6dc47bb-497e-43a0-a5c0-29aa152bc314 · outbound
Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Curiosity-driven exploration by self-supervised prediction
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 338e5196-915e-4c4f-a895-fdc764ad3033 · outbound
Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents On measures of entropy and information
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 496537d7-1db6-4c64-8b0d-8c1941c8e06e · outbound
Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Proximal Policy Optimization Algorithms
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c13f59ad-c478-4952-8302-f56a62205fe9 · outbound
Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Unresolved cited work
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 578313fc-82c3-4f14-a902-e6c44baed511 · outbound
Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4cbfc7c3-90c2-4ebe-9eac-f4d5ecd9477b · outbound
Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents HybridFlow: A Flexible and Efficient RLHF Framework
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2dafed0-13fb-4382-8f5c-c771c102a2c6 · outbound
Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Alfworld: Aligning text and embodied environments for interactive learning
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6552af1-c61e-4232-b173-973651c25c94 · outbound
Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Entropy-guided sequence weighting for efficient exploration in RL-based LLM fine-tuning
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63c6d6a4-c0da-4806-bee3-5efbee141112 · outbound
Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e5ac1ad-1a34-4f53-a377-5f6ec8f6946c · outbound
Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Chain-of-thought prompting elicits reasoning in large language models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6515c5bd-58c0-4cf4-b7d8-601ecd8d0062 · outbound
Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c23d4d99-89c4-44d1-af99-5536df5706c6 · outbound
Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Webagent-r1: Training web agents via end-to-end multi-turn reinforcement learning
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9524d2bd-fded-4fd1-aa92-d028f2277706 · outbound
Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents WebWalker: Benchmarking LLMs in Web Traversal
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79d06aa4-e3b2-4652-8bf4-333ada1bd8ef · outbound
Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents GPT-4V in Wonderland: Large Multimodal Models for Zero-Shot Smartphone GUI Navigation
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6100cb95-1387-4902-ae6c-048d1603b7c8 · outbound
Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b0c2e6a-f183-4658-8568-718b3f7dc21e · outbound
Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Webshop: Towards scalable real-world web interaction with grounded language agents
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f561e9c-1096-4605-be83-fe43692f07a8 · outbound
Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents React: Synergizing reasoning and acting in language models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0a364ad-2e6c-41c5-acfa-bb7a56632527 · outbound
Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93c12cfa-257a-402a-8797-7fedceea0e0a · outbound
Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents CodeAgent: Enhancing Code Generation with Tool-Integrated Agent Systems for Real-World Repo-level Coding Challenges
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9cadc56b-68ff-4aa3-b874-4412ffc301f0 · outbound
Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 697a7a5f-2530-4662-8311-315ed08504e2 · outbound
Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68eb5094-2f5b-4c44-a765-9f568ccab9f4 · outbound
Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1da1519-3ae0-4a72-9b73-a586bbd6349b · outbound
Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Siren’s song in the ai ocean: A survey on hallucination in large language models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3fcd945e-bbd1-4d25-9ba4-34ab0beecf44 · outbound
Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Learning to Reason without External Rewards
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f14c444-b12b-40e3-bcc2-a2bb4ad25b0e · outbound
Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b62c2e21-5fd7-44e5-b1da-582779056cf2 · outbound
Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Maximum entropy inverse reinforcement learning
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 499563ec-ca50-4acf-9bad-8a1cdb473d1c · outbound
Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents TTRL: Test-Time Reinforcement Learning
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1e2fe69-8840-4377-82db-ba5776667af7 · outbound
Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Unresolved cited work
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13689775-68ca-4d31-bf74-5c608d409620 · outbound
Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents stably all-correct
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4b72e7b-3861-44d5-96cd-595c8510e0f8 · outbound
Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents assistant
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 544b829d-5b1a-42a3-99a3-714b0b53bbb8 · outbound
Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Unresolved cited work
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a32333b-95bc-4e91-b1fd-9a0f0b8f2796 · outbound
Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents For each step, the advantage is scaled by g(Ht) and augmented by the future clarity bonus ζ·g ′(Ht+1), yielding the modulated advantageA mod as defined in our main formula (Eq
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9322fcd5-3d85-4ad9-a19e-fc4273418d69 · outbound
Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents assistant
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f75f5912-6283-4211-b65b-87c1c27300ea · inbound
Attention Illuminates LLM Reasoning: The Preplan-and-Anchor Rhythm Enables Fine-Grained Policy Optimization Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35632d2d-07c4-414f-ab1f-debbc018ba5b · inbound
RoboAgent: Chaining Basic Capabilities for Embodied Task Planning Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents
Reference 107
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 19f3b51f-5f6d-4cc8-9d2b-7ab238992ac9 · inbound
ResRL: Boosting LLM Reasoning via Negative Sample Projection Residual Reinforcement Learning Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 502ff5c2-6737-426d-b756-4871a14cb9a7 · inbound
ResRL: Boosting LLM Reasoning via Negative Sample Projection Residual Reinforcement Learning Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e34849f4-fe1f-4457-88a8-dfb5b5a9d82b · inbound
AEM: Adaptive Entropy Modulation for Multi-Turn Agentic Reinforcement Learning Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e91e85ca-8c04-49b9-83dc-32d945d0a665 · inbound
AEM: Adaptive Entropy Modulation for Multi-Turn Agentic Reinforcement Learning Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a537d1f3-4839-4130-a9e5-90f74f09a3fd · inbound
T$^2$PO: Uncertainty-Guided Exploration Control for Stable Multi-Turn Agentic Reinforcement Learning Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b267031b-ef27-4861-8b4f-2e522b92afc7 · inbound
Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 88280954-5462-4afb-a651-9c0410b18b05 · inbound
Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9645b42f-d3a0-415d-847f-ce765024c3cd · inbound
Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation af7daae2-ef69-4176-a0e9-588baf808caa · inbound
Semantic Consistency Policy Optimization for Reinforcement Learning of LLM Agents Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1b4ac855-29fe-4c3f-bd0b-6a6e657fdfe0 · inbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42bbac1a-559f-4d9c-b285-f70dfc65cc5b · inbound
Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents
Reference 111
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff2c5ad0-5fd8-464e-8202-baa1d11ad2fd · inbound
Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.