Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-10T03:00:10.404757Z
Paper Citation Record · LEDGER
As of 4 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 1 inbound Pith citation observation for arXiv:2604.19485.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-10T03:00:10.404757Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-25T21:04:09.237686Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-04T19:40:07.675000Z
41 of 41 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 070ed26d-0fa9-424f-9df1-bf4ba480d6f7 · outbound
EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ea8e235b-a331-4da2-a83a-26c8f447d281 · outbound
EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training Averaged-dqn: Variance reduction and stabi- lization for deep reinforcement learning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b6d58115-abd7-41f2-b3a2-22602d0f85a2 · outbound
EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training Gui-shepherd: Reliable process reward and verification for long-sequence gui tasks, September 2025
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 332ede39-4f3e-4274-a4c8-71f714fb5e3e · outbound
EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training Learning without critics? revisiting grpo in classical reinforcement learning environments
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 679eeca9-d734-4a81-b785-a0934aa19ab0 · outbound
EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training The frozen lake problem
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ff2ef797-b3f5-4974-b805-8c16a1d9172b · outbound
EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training Unresolved cited work
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0d80652e-ef3e-4c30-b284-ee1ac9ecff8f · outbound
EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training Octobench: Benchmarking scaffold-aware instruction following in repository-grounded agentic coding
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9eb9b35f-10ab-4888-a16d-22fd91e52edc · outbound
EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training Agentic reinforced policy optimization
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 87b51482-1e6f-432c-a924-af2097d3298d · outbound
EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training Cl-bench: A benchmark for context learning
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0a1771d8-2915-4fb5-aa69-ac94eba01aa5 · outbound
EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training Only Relevant Information Matters: Filtering Out Noisy Samples to Boost RL
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0fe7c71c-3124-4601-ad39-b82e6c7743c7 · outbound
EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training doi: 10.1038/s41586-025-09422-z
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 604f67c0-6e79-480a-80db-ac9a988ea258 · outbound
EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 261bab36-5b15-41f4-8a4a-8e3d40d46b55 · outbound
EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training Towards understanding the optimization landscape of grpo and its variants
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 73a1a290-0e49-44ed-a61a-904f345cb386 · outbound
EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training Sokoban: Enhancing general single-agent search methods using domain knowledge , journal =
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 50c061d6-43fa-463f-b808-620fd33aec06 · outbound
EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training A new approach to linear filtering and prediction problems
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b52747d3-c75b-4382-a88f-5cc3f1972bfd · outbound
EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training Unifying ppo, dpo, and grpo: A theoretical and empirical study on llm post-training, November 2025
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 649c0255-4eee-4e12-89dd-4635c3126c8f · outbound
EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training Mm-doc-r1: Training agents for long document visual question answering through multi-turn reinforcement learning, April 2026
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7396d07c-85f9-4598-a8ed-c75c020622f8 · outbound
EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training Proximal policy optimization with adaptive generalized advantage estimate: Critic- aware refinements
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 52df4864-6f2b-440c-af71-48e8a41e4311 · outbound
EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training OpenAI o1 System Card
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 538fc613-7358-4188-bcca-00317d9669f3 · outbound
EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training An Elementary Introduction to Kalman Filtering
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d364b24c-9873-4bd4-bda2-afe96b0aeef1 · outbound
EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training Qwen2.5 Technical Report
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 43ac8606-efac-4d49-8654-aebead90d3e9 · outbound
EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training Qwen3.5: Towards native multimodal agents, February 2026
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation cadac5ea-75f9-442a-93de-d8ae094f112d · outbound
EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training The nuts and bolts of deep rl research, December 2016
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 08c7aba5-b672-4e70-847d-eafbee323e6b · outbound
EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training Proximal Policy Optimization Algorithms
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b008bad2-a9b0-422f-a879-59dfd3f8a87f · outbound
EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training High-Dimensional Continuous Control Using Generalized Advantage Estimation
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 04d1690d-680f-4dac-8a2a-8c8b2bae3462 · outbound
EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2beded7a-e5e3-42dd-ade5-0726477dbc3b · outbound
EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training ALFWorld: Aligning Text and Embodied Environments for Interactive Learning
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 121a48e8-329f-4a80-986e-58ba0a5f12b9 · outbound
EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training Unresolved cited work
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ae39a11b-b775-4b53-9ea8-f83edfa691d0 · outbound
EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training OpenAI GPT-5 System Card
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9de7384c-f87e-4d37-b39a-f70f4276010e · outbound
EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training Policy gradient meth- ods for reinforcement learning with function approximation
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation dac46bdd-4f89-495d-b50c-492ca214988c · outbound
EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training Reft: Reasoning with reinforced fine-tuning
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation aeed9c43-44f7-4d04-ac8c-0b077318213c · outbound
EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training R e FT : Reasoning with reinforced fine-tuning
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 799d480a-dd61-466a-a561-bff8d4801995 · outbound
EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training Enhancing llm-based search agents via contribution weighted group relative policy optimization, April
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 62ffc7c2-1021-4394-bd72-5fc45a609449 · outbound
EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training Enhancing LLM-based Search Agents via Contribution Weighted Group Relative Policy Optimization
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f2bb9c72-46d1-47e0-917c-46b24afd5490 · outbound
EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 479d39b5-3732-47da-91e6-59dcba4a505f · outbound
EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training Williams
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ec79021b-ecc1-44a4-a376-adf6d1657c86 · outbound
EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training Qwen3 Technical Report
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 849d9d33-71e8-4e6e-8cfc-14f36a3d2ac5 · outbound
EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training Narasimhan
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4c688259-e04a-42df-b0b1-1e60db5871f8 · outbound
EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 337a1bbe-7e7d-4c1c-99fd-a31b223f6a84 · outbound
EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training Secrets of RLHF in Large Language Models Part I: PPO
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 178facb6-b1e2-4c5c-9dc9-07f81c80c4a5 · outbound
EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training Instruction-Following Evaluation for Large Language Models
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3836ffba-fb46-4ec7-9a06-ecfcd9361482 · inbound
BiPACE: Bisimulation-Guided Policy Optimization with Action Counterfactual Estimation for LLM Agents EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.