Pith. sign in

Paper Citation Record · LEDGER

EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training

As of 4 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 1 inbound Pith citation observation for arXiv:2604.19485.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.19485 v1

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T03:00:10.404757Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-25T21:04:09.237686Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-04T19:40:07.675000Z

Reference resolution

41 of 41 outbound references displayed

  • verified exact26
  • verified fuzzy13
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 070ed26d-0fa9-424f-9df1-bf4ba480d6f7 · outbound

This paper cites Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs.

EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:42:22.769824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T03:00:10.404757Z digest=sha256:758cb9e454eecc8fe334c5375749e1dc66168ed5d1a0bbe8128aecfce73761dd

Observation ea8e235b-a331-4da2-a83a-26c8f447d281 · outbound

This paper cites Averaged-dqn: Variance reduction and stabi- lization for deep reinforcement learning.

EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training Averaged-dqn: Variance reduction and stabi- lization for deep reinforcement learning

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T18:31:56.607300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T03:00:10.404757Z digest=sha256:cc62cd8e6519c29d71f0127a7f0b4ab741314f1d4961dcc9baf027001c305352

Observation b6d58115-abd7-41f2-b3a2-22602d0f85a2 · outbound

This paper cites Gui-shepherd: Reliable process reward and verification for long-sequence gui tasks, September 2025.

EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training Gui-shepherd: Reliable process reward and verification for long-sequence gui tasks, September 2025

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:46:12.588601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T03:00:10.404757Z digest=sha256:79e5d5f53909087f50a3fc457089378619fa677b6ea0791e69db86f7719f4548

Observation 332ede39-4f3e-4274-a4c8-71f714fb5e3e · outbound

This paper cites Learning without critics? revisiting grpo in classical reinforcement learning environments.

EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training Learning without critics? revisiting grpo in classical reinforcement learning environments

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:46:11.525899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T03:00:10.404757Z digest=sha256:28aa05d040ff273dae03836d63c71d570df112298b58abd57c5cc366abf2b633

Observation 679eeca9-d734-4a81-b785-a0934aa19ab0 · outbound

This paper cites The frozen lake problem.

EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training The frozen lake problem

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T18:31:56.600144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T03:00:10.404757Z digest=sha256:2f6d8613fd5fb860430d766ce18505b4cdf9da000c12c7858e68234c2c75067e

Observation ff2ef797-b3f5-4974-b805-8c16a1d9172b · outbound

This paper cites an unresolved cited work.

EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training Unresolved cited work

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:46:13.179286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T03:00:10.404757Z digest=sha256:849bc9dd954e2a01479c066b15a38fa5d6db09d6499ba569ea98d142603530f3

Observation 0d80652e-ef3e-4c30-b284-ee1ac9ecff8f · outbound

This paper cites Octobench: Benchmarking scaffold-aware instruction following in repository-grounded agentic coding.

EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training Octobench: Benchmarking scaffold-aware instruction following in repository-grounded agentic coding

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:46:11.316396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T03:00:10.404757Z digest=sha256:709e91f8c640e9054a015a7c8d793b0466b793cd8d586bdc75149e9ee441c8ba

Observation 9eb9b35f-10ab-4888-a16d-22fd91e52edc · outbound

This paper cites Agentic reinforced policy optimization.

EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training Agentic reinforced policy optimization

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T18:31:56.596152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T03:00:10.404757Z digest=sha256:74480e6472e90c0f979bd3c840bbadf155d02a13a4ffb978bdc2bc1f48371569

Observation 87b51482-1e6f-432c-a924-af2097d3298d · outbound

This paper cites Cl-bench: A benchmark for context learning.

EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training Cl-bench: A benchmark for context learning

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:46:11.131519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T03:00:10.404757Z digest=sha256:b124abe8b9e3a12be45d8282f48355ffb8ada3ae1157135c4fb9cdc2aceae9f4

Observation 0a1771d8-2915-4fb5-aa69-ac94eba01aa5 · outbound

This paper cites Only Relevant Information Matters: Filtering Out Noisy Samples to Boost RL.

EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training Only Relevant Information Matters: Filtering Out Noisy Samples to Boost RL

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:46:13.306057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T03:00:10.404757Z digest=sha256:e9ee4638cbfddb105b625f95b6cbff76d143e68a53d1a512b046ef8ad9bdfcf5

Observation 0fe7c71c-3124-4601-ad39-b82e6c7743c7 · outbound

This paper cites doi: 10.1038/s41586-025-09422-z.

EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training doi: 10.1038/s41586-025-09422-z

Reference 11

Resolution
verified exact
doi, observed 2026-05-10T03:03:36.750129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T03:00:10.404757Z digest=sha256:a68dd566c4673b564d54d3a56c6878137c2605e28b282a12a3f5e15acbb70043

Observation 604f67c0-6e79-480a-80db-ac9a988ea258 · outbound

This paper cites REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization.

EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:21:45.673967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T03:00:10.404757Z digest=sha256:24cfb6a8669d342a507201ee0cedb2f4bb1d479e23782feadc8164f2d55fc6c3

Observation 261bab36-5b15-41f4-8a4a-8e3d40d46b55 · outbound

This paper cites Towards understanding the optimization landscape of grpo and its variants.

EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training Towards understanding the optimization landscape of grpo and its variants

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T18:31:56.590129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T03:00:10.404757Z digest=sha256:5fe3aba15c5f071e34bb2ba4b37f7215c53abbd72277819906a2d6929ba4ec79

Observation 73a1a290-0e49-44ed-a61a-904f345cb386 · outbound

This paper cites Sokoban: Enhancing general single-agent search methods using domain knowledge , journal =.

EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training Sokoban: Enhancing general single-agent search methods using domain knowledge , journal =

Reference 14

Resolution
verified exact
doi, observed 2026-05-10T03:03:36.747580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T03:00:10.404757Z digest=sha256:fce220f5c20a37976c87a582676b095d356cc413285ccd01a1a312af193cf2af

Observation 50c061d6-43fa-463f-b808-620fd33aec06 · outbound

This paper cites A new approach to linear filtering and prediction problems.

EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training A new approach to linear filtering and prediction problems

Reference 15

Resolution
verified exact
doi, observed 2026-05-10T03:03:36.745652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T03:00:10.404757Z digest=sha256:9ea827e395ec2680914bcfa134c61938e0de6f8e7fae8087551e0813f9756634

Observation b52747d3-c75b-4382-a88f-5cc3f1972bfd · outbound

This paper cites Unifying ppo, dpo, and grpo: A theoretical and empirical study on llm post-training, November 2025.

EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training Unifying ppo, dpo, and grpo: A theoretical and empirical study on llm post-training, November 2025

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T18:31:56.584422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T03:00:10.404757Z digest=sha256:eb4f69efdce2934c250087090d4e95e1ee35e59066d28e15475a758f2547da32

Observation 649c0255-4eee-4e12-89dd-4635c3126c8f · outbound

This paper cites Mm-doc-r1: Training agents for long document visual question answering through multi-turn reinforcement learning, April 2026.

EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training Mm-doc-r1: Training agents for long document visual question answering through multi-turn reinforcement learning, April 2026

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T18:31:56.586406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T03:00:10.404757Z digest=sha256:94add4ee56df44c4d27bcb8c5ca9814150deb44b7fd18552e62b9543d1a65b01

Observation 7396d07c-85f9-4598-a8ed-c75c020622f8 · outbound

This paper cites Proximal policy optimization with adaptive generalized advantage estimate: Critic- aware refinements.

EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training Proximal policy optimization with adaptive generalized advantage estimate: Critic- aware refinements

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T18:31:56.588276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T03:00:10.404757Z digest=sha256:26ec80813c205151f81a2df73d15126ff42b367b5d856cb9b101917a69788306

Observation 52df4864-6f2b-440c-af71-48e8a41e4311 · outbound

This paper cites OpenAI o1 System Card.

EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training OpenAI o1 System Card

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-11T12:46:10.654980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T03:00:10.404757Z digest=sha256:893c6cb7aa4d3a05ba15e25704056b08be423b8778c65b17531368c419242558

Observation 538fc613-7358-4188-bcca-00317d9669f3 · outbound

This paper cites An Elementary Introduction to Kalman Filtering.

EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training An Elementary Introduction to Kalman Filtering

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:46:12.249347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T03:00:10.404757Z digest=sha256:956bb30f8418bf52491bb407504d31c6a77a289c4dd8d1bb0527d902a931e4c3

Observation d364b24c-9873-4bd4-bda2-afe96b0aeef1 · outbound

This paper cites Qwen2.5 Technical Report.

EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training Qwen2.5 Technical Report

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-11T12:46:12.105785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T03:00:10.404757Z digest=sha256:33b951690dbfa4e5cbd4458bd127a65adce818536fbfe53c6a4f6a271cc5b2e9

Observation 43ac8606-efac-4d49-8654-aebead90d3e9 · outbound

This paper cites Qwen3.5: Towards native multimodal agents, February 2026.

EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training Qwen3.5: Towards native multimodal agents, February 2026

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T18:31:56.592566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T03:00:10.404757Z digest=sha256:15ece0154428e9bfd62e4afe5376194ce8cdbc4865d986c50b7f1f2acc8819cd

Observation cadac5ea-75f9-442a-93de-d8ae094f112d · outbound

This paper cites The nuts and bolts of deep rl research, December 2016.

EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training The nuts and bolts of deep rl research, December 2016

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T18:31:56.582364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T03:00:10.404757Z digest=sha256:efb6efdcf7301145a6e2efa5159c172c3537009ae72c4e206c11c61dda9568a8

Observation 08c7aba5-b672-4e70-847d-eafbee323e6b · outbound

This paper cites Proximal Policy Optimization Algorithms.

EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training Proximal Policy Optimization Algorithms

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-11T12:46:12.712352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T03:00:10.404757Z digest=sha256:f66b756ac1da3044aa739b353a5ec03d20695705e890f2a9b0d7b1a214671258

Observation b008bad2-a9b0-422f-a879-59dfd3f8a87f · outbound

This paper cites High-Dimensional Continuous Control Using Generalized Advantage Estimation.

EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training High-Dimensional Continuous Control Using Generalized Advantage Estimation

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-11T12:46:11.006886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T03:00:10.404757Z digest=sha256:3bd053ab3b9c113cce3c96c830c8d448c415e90cdc514b6c6562bcc3051f4396

Observation 04d1690d-680f-4dac-8a2a-8c8b2bae3462 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-11T12:46:11.484943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T03:00:10.404757Z digest=sha256:e67d71d0b057125150d4166b674b42472a8c200415f09c226048856a25a0c695

Observation 2beded7a-e5e3-42dd-ade5-0726477dbc3b · outbound

This paper cites ALFWorld: Aligning Text and Embodied Environments for Interactive Learning.

EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training ALFWorld: Aligning Text and Embodied Environments for Interactive Learning

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:39:15.970362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T03:00:10.404757Z digest=sha256:53753f7f6b334c1a7a10320cc0fc68bd6bd1d7c01e65068bece58e2529b2643d

Observation 121a48e8-329f-4a80-986e-58ba0a5f12b9 · outbound

This paper cites an unresolved cited work.

EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-05-22T18:31:56.594358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T03:00:10.404757Z digest=sha256:a43f01cd9e0db7802941c78124e6db08292836bb8efadb96afb44abdc562ba7d

Observation ae39a11b-b775-4b53-9ea8-f83edfa691d0 · outbound

This paper cites OpenAI GPT-5 System Card.

EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training OpenAI GPT-5 System Card

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-11T12:46:12.427917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T03:00:10.404757Z digest=sha256:e3b1059b8cdb754354ff4e313c8b52877efd7442bd5ed16f4643bb9860cac006

Observation 9de7384c-f87e-4d37-b39a-f70f4276010e · outbound

This paper cites Policy gradient meth- ods for reinforcement learning with function approximation.

EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training Policy gradient meth- ods for reinforcement learning with function approximation

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T18:31:56.598210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T03:00:10.404757Z digest=sha256:f167fcef70830be5ef60f93912732a4b0075545938f0b6fb638cf8d59dfd6655

Observation dac46bdd-4f89-495d-b50c-492ca214988c · outbound

This paper cites Reft: Reasoning with reinforced fine-tuning.

EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training Reft: Reasoning with reinforced fine-tuning

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T18:31:56.605493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T03:00:10.404757Z digest=sha256:48007fd844be6a7dad570eec2e0400e5cfdbf4f421ee03db8791f6cf77787726

Observation aeed9c43-44f7-4d04-ac8c-0b077318213c · outbound

This paper cites R e FT : Reasoning with reinforced fine-tuning.

EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training R e FT : Reasoning with reinforced fine-tuning

Reference 32

Resolution
metadata mismatch
doi, observed 2026-05-10T03:03:36.754019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T03:00:10.404757Z digest=sha256:c535d690c65acda0eeab25c17469cffb3487249cdb79ec8a353dea74f1a4cec8

Observation 799d480a-dd61-466a-a561-bff8d4801995 · outbound

This paper cites Enhancing llm-based search agents via contribution weighted group relative policy optimization, April.

EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training Enhancing llm-based search agents via contribution weighted group relative policy optimization, April

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T18:31:56.602025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T03:00:10.404757Z digest=sha256:104dc00557a14f113b2525e4242b7c1f32d7118563ce0e193ea6493a36e33515

Observation 62ffc7c2-1021-4394-bd72-5fc45a609449 · outbound

This paper cites Enhancing LLM-based Search Agents via Contribution Weighted Group Relative Policy Optimization.

EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training Enhancing LLM-based Search Agents via Contribution Weighted Group Relative Policy Optimization

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-11T12:46:13.099355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T03:00:10.404757Z digest=sha256:71bd24162e9a3b165bfbcffd643d7f8733d1f66b3fc9bf8cee8c27d4b74b1442

Observation f2bb9c72-46d1-47e0-917c-46b24afd5490 · outbound

This paper cites RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning.

EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:13:34.687091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T03:00:10.404757Z digest=sha256:5e9e18ea454742b403d95ac0d477c2cb9caa1f25a8223364ea8d805324018b72

Observation 479d39b5-3732-47da-91e6-59dcba4a505f · outbound

This paper cites Williams.

EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training Williams

Reference 36

Resolution
verified exact
doi, observed 2026-05-10T03:03:36.752206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T03:00:10.404757Z digest=sha256:7a131c4ec4bd881cea7deba590a9683efad7a96f4da50aca066c2ce254ae26f1

Observation ec79021b-ecc1-44a4-a376-adf6d1657c86 · outbound

This paper cites Qwen3 Technical Report.

EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training Qwen3 Technical Report

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-11T12:46:13.749355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T03:00:10.404757Z digest=sha256:714f32999409688a94719cf383b797defa2853acedbc990cfcb13d542702457b

Observation 849d9d33-71e8-4e6e-8cfc-14f36a3d2ac5 · outbound

This paper cites Narasimhan.

EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training Narasimhan

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T18:31:56.603672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T03:00:10.404757Z digest=sha256:04e10454a43b4cabd776c400a29e3f971cd141ac9b7c854a3af1f350148d7c18

Observation 4c688259-e04a-42df-b0b1-1e60db5871f8 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-11T12:46:13.509807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T03:00:10.404757Z digest=sha256:f19942ead708dd85e8d1dfaf39485c0de66bc30c65af4f2ec097aecfbdbc73bb

Observation 337a1bbe-7e7d-4c1c-99fd-a31b223f6a84 · outbound

This paper cites Secrets of RLHF in Large Language Models Part I: PPO.

EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training Secrets of RLHF in Large Language Models Part I: PPO

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-17T13:17:43.137115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T03:00:10.404757Z digest=sha256:0af78f91ec0e231d99e3195a6a167d0d066bc07eef399d03b2011a3ae1d67d2d

Observation 178facb6-b1e2-4c5c-9dc9-07f81c80c4a5 · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training Instruction-Following Evaluation for Large Language Models

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-05-11T12:46:12.866351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T03:00:10.404757Z digest=sha256:cf0e2ca16395d82f9d6703c76bf3a8ac0ea3319bf3423731c2ce095e9175f4df

Pith citing papers

Observation 3836ffba-fb46-4ec7-9a06-ecfcd9361482 · inbound

BiPACE: Bisimulation-Guided Policy Optimization with Action Counterfactual Estimation for LLM Agents cites this paper.

BiPACE: Bisimulation-Guided Policy Optimization with Action Counterfactual Estimation for LLM Agents EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-04T19:40:07.676769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-25T21:04:09.237686Z digest=sha256:6f4993465eea2df81e73dea626e297309de6d03cc79227907514a603e8cf31ca