Pith. sign in

Paper Citation Record · LEDGER

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents

As of 7 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 14 inbound Pith citation observations for arXiv:2509.09265.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.09265 v1

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T19:29:01.480681Z

measured 69 of 69 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T09:48:09.620197Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T20:20:07.648363Z

Reference resolution

55 of 55 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved55
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8aed9b11-2178-4d64-b947-6ad892f692a1 · outbound

This paper cites The Unreasonable Effectiveness of Entropy Minimization in LLM Reasoning.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents The Unreasonable Effectiveness of Entropy Minimization in LLM Reasoning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:56.091838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:56.091838Z digest=sha256:f4426e5e65aeea26959ad04eeb789feac2f7e2d3ecb744154dace69436733ba2

Observation e6144339-f32c-440d-a59d-ee4fd68eaee3 · outbound

This paper cites Open Deep Search: Democratizing Search with Open-source Reasoning Agents.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Open Deep Search: Democratizing Search with Open-source Reasoning Agents

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:56.149783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:56.149783Z digest=sha256:74a28c03e61c795a6b53910f631b919da70ed2168086c0e43a0acc1d62ba50a1

Observation 69199cbf-c8ef-4641-92ed-6d7e0819c1b2 · outbound

This paper cites Unifying count-based exploration and intrinsic motivation.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Unifying count-based exploration and intrinsic motivation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:56.219728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:56.219728Z digest=sha256:5500f9ce3b94981379b5e647a6112b66f2da3a8963b7e9ece31f66ab4ed0e6d1

Observation 2e325a43-c496-45a7-97ee-128721ef66f1 · outbound

This paper cites SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:56.258016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:56.258016Z digest=sha256:3f89ca00bd9328e497e39f310d2d19357d2ed802759ed043715a515fe6fca7b9

Observation a05d97b5-0110-4c94-ad1b-28ebe8aa8c2d · outbound

This paper cites Reasoning with Exploration: An Entropy Perspective.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Reasoning with Exploration: An Entropy Perspective

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:56.355996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:56.355996Z digest=sha256:85445a2d587ade9ce2f396d27113ed623a5c70336330f5ca5c88e37b6a85932d

Observation de68e66a-70e6-4029-bada-9fd30d5028ef · outbound

This paper cites Mind2web: Towards a generalist agent for the web.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Mind2web: Towards a generalist agent for the web

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:56.497613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:56.497613Z digest=sha256:c6e35c9ef58df1c1340d0e7af6af67e9a68aca039884acd38db2ba1a7c674b68

Observation 90f14950-d6d6-49b2-a74a-88977f7ab7ac · outbound

This paper cites From Novice to Expert: LLM Agent Policy Optimization via Step-wise Reinforcement Learning.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents From Novice to Expert: LLM Agent Policy Optimization via Step-wise Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:56.630716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:56.630716Z digest=sha256:68ff9aaa73d5ea0d9d4290bc50b58165d48e6fdb01103b66e91dcc0091f46ba6

Observation df389e07-afe4-4d95-b7c9-51e573fb0dbc · outbound

This paper cites Group-in-Group Policy Optimization for LLM Agent Training.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Group-in-Group Policy Optimization for LLM Agent Training

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:56.713427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:56.713427Z digest=sha256:a79790726f4ad7dfbe3469864bd617f2f1f4f2c08d8b3a8a7419af7d1d8d6628

Observation 596ef3ac-f757-494c-88ad-d727917f3bec · outbound

This paper cites One-shot Entropy Minimization.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents One-shot Entropy Minimization

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:56.841114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:56.841114Z digest=sha256:1e66b717b39c81ec997b75e42db0e4309e99cb5cb202211c533d5a1aa6000262

Observation 5c180ad2-db6c-4ea3-aa5a-3719bcc5a038 · outbound

This paper cites PaSa: An LLM Agent for Comprehensive Academic Paper Search.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents PaSa: An LLM Agent for Comprehensive Academic Paper Search

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:56.997117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:56.997117Z digest=sha256:2e1e6f6b8fd8165b2a18d1f97e97061df37ff34bb34d63fcb6f0a20f422e4fce

Observation 8a1376b0-98e0-453f-9a7b-6f45d307f65a · outbound

This paper cites Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:57.112469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:57.112469Z digest=sha256:46087d28d3439fa8d058385bb81e86099ad36f906a7edf09ffe65013d39441d6

Observation a98c37c0-4a0b-49d9-aec4-9c9c63dc83ec · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-World GitHub Issues?.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:57.248343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:57.248343Z digest=sha256:f1d7d065238780f2300fbc7b042f424b063183e3d6373128af44929f1607a0e1

Observation 7a2f0c0c-7b53-403d-a0a3-9166689d0131 · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:57.347328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:57.347328Z digest=sha256:6c64508c43fa277dfde4fe4e1045a45edb18cfe85f890fc8f58fbbadc2bc0e18

Observation 38dc535f-1696-4f01-833e-38b59d342411 · outbound

This paper cites TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:57.447645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:57.447645Z digest=sha256:1f8ed450617a20bf87d4e574558afb7ea0b8812f992752d093a3e035beb13549

Observation 2fd13494-fcc6-4011-89cc-4a00c826e275 · outbound

This paper cites Vineppo: Accurate credit assignment in rl for llm mathematical reasoning.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Vineppo: Accurate credit assignment in rl for llm mathematical reasoning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:57.566435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:57.566435Z digest=sha256:34e4712f647385fa554c9714afe5f089cd7b619844d44f2558f8e4b9473af4af

Observation 30f6dae9-c4ec-4c99-b9ca-9a868bbf0e88 · outbound

This paper cites Empowerment: A universal agent-centric measure of control.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Empowerment: A universal agent-centric measure of control

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:57.653021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:57.653021Z digest=sha256:2d10e3eb35d7b01421a42ef709d1c6414f5e216bbd41044cb01b7ed92603c7c0

Observation 35a7abdd-7c0c-4d50-a6f7-1171953a310c · outbound

This paper cites Natural questions: a benchmark for question answering research.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Natural questions: a benchmark for question answering research

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:57.767525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:57.767525Z digest=sha256:f85de81f5bf820471ac557d2df6e925e1787f9f7e93757356aaa39c7fcae8451

Observation f934fa58-2f68-4fe4-967a-97c41b32e0d2 · outbound

This paper cites Search-o1: Agentic Search-Enhanced Large Reasoning Models.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Search-o1: Agentic Search-Enhanced Large Reasoning Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:57.876860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:57.876860Z digest=sha256:4af5b6a679b84cc2b4115b449741f6b754aa9ab2661fb8bd731df3ee69b7aebb

Observation 1fb21e88-50dd-49a8-8d0b-b92dd067e492 · outbound

This paper cites Logit Dynamics in Softmax Policy Gradient Methods.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Logit Dynamics in Softmax Policy Gradient Methods

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:58.006560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:58.006560Z digest=sha256:126050e6256e9c0df9fc65ee2d4c9248b3e59ada6b6a2211e32f8a888932835c

Observation d76f24e3-2abd-46f0-a4c9-b1c56c5755a7 · outbound

This paper cites Let’s verify step by step.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Let’s verify step by step

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:58.159200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:58.159200Z digest=sha256:b9fadd6b856be9c97a866244167e8d80f4cda1e15fa483c3c63adf018e6e13b8

Observation 13723c2e-df2d-4094-92f3-2598b454f8e0 · outbound

This paper cites Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:58.286475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:58.286475Z digest=sha256:3627a2836b6903ca92f43f87e6234403aa7bfd208691be5d8273bc135d337729

Observation 9f632c4f-3d97-43b8-84de-1aa3165ac244 · outbound

This paper cites Policy invariance under reward transformations: Theory and application to reward shaping.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Policy invariance under reward transformations: Theory and application to reward shaping

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:58.398421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:58.398421Z digest=sha256:c29be835de1a4fafb2f5a19611d8e996412a7d9c64c1b347ddc63d2c314d64fe

Observation c6dc47bb-497e-43a0-a5c0-29aa152bc314 · outbound

This paper cites Curiosity-driven exploration by self-supervised prediction.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Curiosity-driven exploration by self-supervised prediction

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:58.536248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:58.536248Z digest=sha256:451cffef946f70d5c9876cf46e4e3a84c9c104fa4c7bed3eb6d436b9042354a8

Observation 338e5196-915e-4c4f-a895-fdc764ad3033 · outbound

This paper cites On measures of entropy and information.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents On measures of entropy and information

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:58.698098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:58.698098Z digest=sha256:6c1cf4151d9e1fc1055d833101140aea3174c50b7d498be5187d100a77b6aa52

Observation 496537d7-1db6-4c64-8b0d-8c1941c8e06e · outbound

This paper cites Proximal Policy Optimization Algorithms.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Proximal Policy Optimization Algorithms

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:58.804689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:58.804689Z digest=sha256:1385a5a73224471ccc16b1bc350ec7e36f276e5a9ee7f153a27af2ffe43f9a0f

Observation c13f59ad-c478-4952-8302-f56a62205fe9 · outbound

This paper cites an unresolved cited work.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:58.939607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:58.939607Z digest=sha256:793e1267f788f51c352f6a88ad1aab67a51834e70ef74b5dd399f9a3fad3bc62

Observation 578313fc-82c3-4f14-a902-e6c44baed511 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:59.039537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:59.039537Z digest=sha256:c5ef0af84b6289a7de2c948234ae5806715d2191501140c03cc3e2ecae643c14

Observation 4cbfc7c3-90c2-4ebe-9eac-f4d5ecd9477b · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents HybridFlow: A Flexible and Efficient RLHF Framework

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:59.160640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:59.160640Z digest=sha256:fc8b3f09938d1ab06f45a0133193dc5ee812a8e88ea6bd10b7d4a300a89ca5e9

Observation e2dafed0-13fb-4382-8f5c-c771c102a2c6 · outbound

This paper cites Alfworld: Aligning text and embodied environments for interactive learning.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Alfworld: Aligning text and embodied environments for interactive learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:59.252460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:59.252460Z digest=sha256:7374e4a1ed6dbb9ca5b9ebfb7c68f27f0cac8200851106353b97fab4830e2cdd

Observation f6552af1-c61e-4232-b173-973651c25c94 · outbound

This paper cites Entropy-guided sequence weighting for efficient exploration in RL-based LLM fine-tuning.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Entropy-guided sequence weighting for efficient exploration in RL-based LLM fine-tuning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:59.417616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:59.417616Z digest=sha256:a2b804acf6739817021c6467bc7d30acac0b83dc1752ecd040aa5633065a242a

Observation 63c6d6a4-c0da-4806-bee3-5efbee141112 · outbound

This paper cites Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:59.532301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:59.532301Z digest=sha256:42781f72b7c71e46d98572560c477b267911bc68d96fb95f88fdab1744b53f3e

Observation 6e5ac1ad-1a34-4f53-a377-5f6ec8f6946c · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Chain-of-thought prompting elicits reasoning in large language models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:59.677626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:59.677626Z digest=sha256:5a1070153cb020db338bdb7feebded2b9967b549513785cb3dd524b83f63101f

Observation 6515c5bd-58c0-4cf4-b7d8-601ecd8d0062 · outbound

This paper cites SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:59.797922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:59.797922Z digest=sha256:476fd32a5c6ad77432c2be34d100a70f800977b14670d81f48c4a67553697a49

Observation c23d4d99-89c4-44d1-af99-5536df5706c6 · outbound

This paper cites Webagent-r1: Training web agents via end-to-end multi-turn reinforcement learning.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Webagent-r1: Training web agents via end-to-end multi-turn reinforcement learning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:59.907127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:59.907127Z digest=sha256:fa36cb60e36f81aeb59493a4dbf0844f77251b474715ff84a405a6ee9fbb48b5

Observation 9524d2bd-fded-4fd1-aa92-d028f2277706 · outbound

This paper cites WebWalker: Benchmarking LLMs in Web Traversal.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents WebWalker: Benchmarking LLMs in Web Traversal

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T19:29:00.015869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:29:00.015869Z digest=sha256:80db941205281431f26ccef0f54e9abc57a4e1fb3f6237846c5543902e5fa276

Observation 79d06aa4-e3b2-4652-8bf4-333ada1bd8ef · outbound

This paper cites GPT-4V in Wonderland: Large Multimodal Models for Zero-Shot Smartphone GUI Navigation.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents GPT-4V in Wonderland: Large Multimodal Models for Zero-Shot Smartphone GUI Navigation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T19:29:00.066242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:29:00.066242Z digest=sha256:3878c1b70b176a039345455a502df4fdc30d18d6e04bd500da37c41cd24d5106

Observation 6100cb95-1387-4902-ae6c-048d1603b7c8 · outbound

This paper cites HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T19:29:00.156802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:29:00.156802Z digest=sha256:11935f5540a1c4415e290df5a77f92f3f6d865e313cfea309d9b8c183714ace6

Observation 7b0c2e6a-f183-4658-8568-718b3f7dc21e · outbound

This paper cites Webshop: Towards scalable real-world web interaction with grounded language agents.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Webshop: Towards scalable real-world web interaction with grounded language agents

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T19:29:00.216580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:29:00.216580Z digest=sha256:5c2a688560435493df1cb9feed3b814cb4c06040a6983e5ee734881af0ec8d9d

Observation 6f561e9c-1096-4605-be83-fe43692f07a8 · outbound

This paper cites React: Synergizing reasoning and acting in language models.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents React: Synergizing reasoning and acting in language models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T19:29:00.260395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:29:00.260395Z digest=sha256:2e23c531c46d0a35bfe9dceee78595184298761bb69e3a262b92d251ecc11038

Observation b0a364ad-2e6c-41c5-acfa-bb7a56632527 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T19:29:00.312931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:29:00.312931Z digest=sha256:3eb31df303a3443acc5b71c400d838c514b5673208017cf2f9766e29a49b4790

Observation 93c12cfa-257a-402a-8797-7fedceea0e0a · outbound

This paper cites CodeAgent: Enhancing Code Generation with Tool-Integrated Agent Systems for Real-World Repo-level Coding Challenges.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents CodeAgent: Enhancing Code Generation with Tool-Integrated Agent Systems for Real-World Repo-level Coding Challenges

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-04T19:29:00.384563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:29:00.384563Z digest=sha256:b0ba4fa2cd0f94e08e3ba8b195ada7b0c921fd4594182337d40e75c52a33c2ee

Observation 9cadc56b-68ff-4aa3-b874-4412ffc301f0 · outbound

This paper cites Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T19:29:00.494940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:29:00.494940Z digest=sha256:d09e2d90d75ea8edf4ec2e3f80c680a710a5c60353bf95a80012aaa4f5da81e7

Observation 697a7a5f-2530-4662-8311-315ed08504e2 · outbound

This paper cites Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T19:29:00.555925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:29:00.555925Z digest=sha256:1ebd1293f94460c198e667c6c7009b45cee62b320dca5a51bab00291a55e7bd6

Observation 68eb5094-2f5b-4c44-a765-9f568ccab9f4 · outbound

This paper cites EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T19:29:00.634858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:29:00.634858Z digest=sha256:b0a5331ef4ff03bf6865bd30d480809a49d9275a32a0bad61c65d88152149d6c

Observation d1da1519-3ae0-4a72-9b73-a586bbd6349b · outbound

This paper cites Siren’s song in the ai ocean: A survey on hallucination in large language models.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Siren’s song in the ai ocean: A survey on hallucination in large language models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T19:29:00.713218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:29:00.713218Z digest=sha256:eade4f4280484570c56c878bc22e1750cca74b04b182de220cdfa98b3009cc66

Observation 3fcd945e-bbd1-4d25-9ba4-34ab0beecf44 · outbound

This paper cites Learning to Reason without External Rewards.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Learning to Reason without External Rewards

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-04T19:29:00.812232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:29:00.812232Z digest=sha256:5c3d3e1f3886b4c3de6c215f6d0804ccfb18085a191a63a015835fca1adc4393

Observation 8f14c444-b12b-40e3-bcc2-a2bb4ad25b0e · outbound

This paper cites DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T19:29:00.879986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:29:00.879986Z digest=sha256:78bf5bb4b93e885e6c1c7ba9865321355075f4ea977fefc18663c5daece3b428

Observation b62c2e21-5fd7-44e5-b1da-582779056cf2 · outbound

This paper cites Maximum entropy inverse reinforcement learning.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Maximum entropy inverse reinforcement learning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-04T19:29:00.944566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:29:00.944566Z digest=sha256:733e4014a267b8be85cb1cc5717617980dfb0ece86b78c0d63260b0f9689742e

Observation 499563ec-ca50-4acf-9bad-8a1cdb473d1c · outbound

This paper cites TTRL: Test-Time Reinforcement Learning.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents TTRL: Test-Time Reinforcement Learning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-04T19:29:01.018560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:29:01.018560Z digest=sha256:788ff6102d90417170f68ec8e6ddecbb6c9c7a357e35bbb26a7299461cee5351

Observation a1e2fe69-8840-4377-82db-ba5776667af7 · outbound

This paper cites an unresolved cited work.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Unresolved cited work

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T19:29:01.084911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:29:01.084911Z digest=sha256:3c47b2971c0d33e9185c62c18ae99f8f19778901fd8af2c99fa9f1b60e536a45

Observation 13689775-68ca-4d31-bf74-5c608d409620 · outbound

This paper cites stably all-correct.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents stably all-correct

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-04T19:29:01.211245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:29:01.211245Z digest=sha256:ad27fcf804f68fec331e2df87969ccee3edb03cf3653f83c9bae130a0f459f6f

Observation d4b72e7b-3861-44d5-96cd-595c8510e0f8 · outbound

This paper cites assistant.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents assistant

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-04T19:29:01.262634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:29:01.262634Z digest=sha256:a47ce7d21d7b40c4fb2244d59d28bcdd031b1a0fc793709699e8f5baf793bb18

Observation 544b829d-5b1a-42a3-99a3-714b0b53bbb8 · outbound

This paper cites an unresolved cited work.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Unresolved cited work

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-04T19:29:01.353844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:29:01.353844Z digest=sha256:2b24950043d34e5ff8834611be9b56b575b347ee5bea32dc7e17e747d4cdc1f0

Observation 3a32333b-95bc-4e91-b1fd-9a0f0b8f2796 · outbound

This paper cites For each step, the advantage is scaled by g(Ht) and augmented by the future clarity bonus ζ·g ′(Ht+1), yielding the modulated advantageA mod as defined in our main formula (Eq.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents For each step, the advantage is scaled by g(Ht) and augmented by the future clarity bonus ζ·g ′(Ht+1), yielding the modulated advantageA mod as defined in our main formula (Eq

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-04T19:29:01.414578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:29:01.414578Z digest=sha256:d6f75a30696bbda97996b1e85391ab4b0453676d65a0e3310bd49343476564da

Observation 9322fcd5-3d85-4ad9-a19e-fc4273418d69 · outbound

This paper cites assistant.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents assistant

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-04T19:29:01.480681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:29:01.480681Z digest=sha256:e04de8a067266a7a93754bf9dbe84ed4ffc314f655df1e2f1b83a1da714a5e6b

Pith citing papers

Observation f75f5912-6283-4211-b65b-87c1c27300ea · inbound

Attention Illuminates LLM Reasoning: The Preplan-and-Anchor Rhythm Enables Fine-Grained Policy Optimization cites this paper.

Attention Illuminates LLM Reasoning: The Preplan-and-Anchor Rhythm Enables Fine-Grained Policy Optimization Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-04T09:48:09.620197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T09:48:09.620197Z digest=sha256:f38e896f18dc95de04ccc17cc46c4e252d34d257a97c8ee1ea1a719f01145f20

Observation 35632d2d-07c4-414f-ab1f-debbc018ba5b · inbound

RoboAgent: Chaining Basic Capabilities for Embodied Task Planning cites this paper.

RoboAgent: Chaining Basic Capabilities for Embodied Task Planning Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents

Reference 107

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:15:57.221330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T18:15:08.727921Z digest=sha256:0eafff804a57bb256b073661b48e899d7f17c77151f4b2fe4082d1202f80b47d

Observation 19f3b51f-5f6d-4cc8-9d2b-7ab238992ac9 · inbound

ResRL: Boosting LLM Reasoning via Negative Sample Projection Residual Reinforcement Learning cites this paper.

ResRL: Boosting LLM Reasoning via Negative Sample Projection Residual Reinforcement Learning Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:41:45.936142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-09T19:26:57.596581Z digest=sha256:82899a3aae57cde7ee61e180769e2661687361aee8da4571aa4ffecf9530de33

Observation 502ff5c2-6737-426d-b756-4871a14cb9a7 · inbound

ResRL: Boosting LLM Reasoning via Negative Sample Projection Residual Reinforcement Learning cites this paper.

ResRL: Boosting LLM Reasoning via Negative Sample Projection Residual Reinforcement Learning Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:55:56.612400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-11T02:07:21.806345Z digest=sha256:f9c4f30ce423cfa5351e2c4978a7e900ab7914dd601167da753c308574e0c864

Observation e34849f4-fe1f-4457-88a8-dfb5b5a9d82b · inbound

AEM: Adaptive Entropy Modulation for Multi-Turn Agentic Reinforcement Learning cites this paper.

AEM: Adaptive Entropy Modulation for Multi-Turn Agentic Reinforcement Learning Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:26:17.685450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-09T19:56:27.170349Z digest=sha256:6b3f33f0931286e8aefea02e30630b3b5bb19577c3a3da132df901bb7fb04e02

Observation e91e85ca-8c04-49b9-83dc-32d945d0a665 · inbound

AEM: Adaptive Entropy Modulation for Multi-Turn Agentic Reinforcement Learning cites this paper.

AEM: Adaptive Entropy Modulation for Multi-Turn Agentic Reinforcement Learning Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:55:59.145694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-11T00:58:25.685484Z digest=sha256:22c16fe124aeb61bdde5a75632a7e413ccdaf8e4a0f34501beafdae9f54be0e5

Observation a537d1f3-4839-4130-a9e5-90f74f09a3fd · inbound

T$^2$PO: Uncertainty-Guided Exploration Control for Stable Multi-Turn Agentic Reinforcement Learning cites this paper.

T$^2$PO: Uncertainty-Guided Exploration Control for Stable Multi-Turn Agentic Reinforcement Learning Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T05:45:22.203001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T19:36:00.359351Z digest=sha256:bb8a1e100b9a6c423d3e8207b2a3f9a40a3772cfaf932c7ba13ea705376d6c69

Observation b267031b-ef27-4861-8b4f-2e522b92afc7 · inbound

Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning cites this paper.

Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents

Reference 89

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:06:08.516567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T10:23:52.522238Z digest=sha256:cddea0caf977d2b232bfd12c0f0b69dc02ecafe787c5bcc102fba4ea2857f957

Observation 88280954-5462-4afb-a651-9c0410b18b05 · inbound

Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning cites this paper.

Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents

Reference 89

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:00:56.957537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-11T02:00:00.663355Z digest=sha256:9bc2b503367d82f3030c965bbea5ca49e88f62e781e97d19e1d70218cb1995f4

Observation 9645b42f-d3a0-415d-847f-ce765024c3cd · inbound

Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning cites this paper.

Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents

Reference 89

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:17:28.392379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-13T07:17:13.708752Z digest=sha256:8ee982c2890f008fd56125aa388b5d3132e6a1f8587d371991eae78dbdbc9892

Observation af7daae2-ef69-4176-a0e9-588baf808caa · inbound

Semantic Consistency Policy Optimization for Reinforcement Learning of LLM Agents cites this paper.

Semantic Consistency Policy Optimization for Reinforcement Learning of LLM Agents Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-04T20:20:07.650170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-25T20:16:39.347676Z digest=sha256:d2ede55ea042f6421a20b264dceac3a9a90b4bcdd36fcda2d046795583629f51

Observation 1b4ac855-29fe-4c3f-bd0b-6a6e657fdfe0 · inbound

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training cites this paper.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents

Reference 96

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:13c179157e3caed4b27a85399afb358839e87e28f5ca6c804162521e99b389eb

Observation 42bbac1a-559f-4d9c-b285-f70dfc65cc5b · inbound

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration cites this paper.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents

Reference 111

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:27.159465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:27.159465Z digest=sha256:cca1d7f6b2a1c0d89f18ca7114c927af8c415865c99d3db5d0ca3560dce41b2e

Observation ff2c5ad0-5fd8-464e-8202-baa1d11ad2fd · inbound

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization cites this paper.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:08.449103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:08.449103Z digest=sha256:4cc4fb4604efa2759127cc923cc20115c7a4ed7ab3fbe000e03b93b2b443c6d4