Pith. sign in

Paper Citation Record · LEDGER

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards

As of 22 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 0 inbound Pith citation observations for arXiv:2511.23310.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2511.23310 v4

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T06:47:16.927539Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

47 of 47 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved47
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0624fa21-0f17-4d6c-8d54-877a33b328a6 · outbound

This paper cites Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.715714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.715714Z digest=sha256:00b58931fb39ff921f3b2bcca5de3dcad5ce1002550b8a11243e3525ee6c879d

Observation cdc33d7b-b6a7-4c7d-ab3c-81ed1730e6ee · outbound

This paper cites an unresolved cited work.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.720985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.720985Z digest=sha256:6cf62b76753783f7904a6b6e85f553257b9b2360fdd150daafa535c66f673f17

Observation adf254ae-7ba7-41aa-a72d-801cfec47859 · outbound

This paper cites an unresolved cited work.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.725978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.725978Z digest=sha256:6b7daa0379f166069b58b68c2b53cf10f957c50c9e0cd2120d8fbd575836cc0f

Observation 0cbba526-11e8-4349-b46e-84ff765fce1e · outbound

This paper cites an unresolved cited work.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.735535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.735535Z digest=sha256:78ade2bac228b38ac1cca2f5d233f7ee9e5b138ec68255b1733e1034d8c4fa9c

Observation 42a0004a-111e-42bf-928b-975c11a0dcb5 · outbound

This paper cites an unresolved cited work.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.739630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.739630Z digest=sha256:ea357e51746ac7eedbad524885631c6723a4c5498993348588e62a0671b07b1a

Observation b9c5c8dc-3789-4c5f-b578-07b39a20653d · outbound

This paper cites an unresolved cited work.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.749076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.749076Z digest=sha256:99ae1ee755f332f3a971e2bea9cd531ab03e0f1998d8f2212ef7593b561e3668

Observation 63dbecb6-46c8-41eb-bb11-94868fa0d084 · outbound

This paper cites Deep reinforcement learning from human preferences.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Deep reinforcement learning from human preferences

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.753929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.753929Z digest=sha256:e956afc2013bc434584131b178279c0d563d9eeac0ff21a460724e2e14acb385

Observation 661c9ad4-255c-49cb-9015-8a6eba4c3e1a · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Training Verifiers to Solve Math Word Problems

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.758362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.758362Z digest=sha256:227edc495dbf5f9e7c3699cc8907e0b35b2295a49743eae675c9a7face220046

Observation 0e304497-bff5-4a3c-9e4b-d67f9b92166e · outbound

This paper cites Learning rate schedules for faster stochas- tic gradient search.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Learning rate schedules for faster stochas- tic gradient search

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.763412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.763412Z digest=sha256:fe7648ec13755c33b7c9a4241e1012b2bd08913b0c423f0f3e9d9de0468111e7

Observation 239f436c-ad32-4b16-93cc-578fe2bac4eb · outbound

This paper cites Note on learning rate schedules for stochastic optimiza- tion.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Note on learning rate schedules for stochastic optimiza- tion

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.768142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.768142Z digest=sha256:7523a50d934af09ac1607cdcdc33b84a940342c60263a88b829f1eb355ef0e70

Observation 1f9a8a28-c780-4993-8e5e-d5e39d93e110 · outbound

This paper cites an unresolved cited work.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.772206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.772206Z digest=sha256:472ef5c1965ee07c34f8742a27f5fdf5b7ef789eb85435fb537d90361b95a840

Observation c0b4841c-b26c-41d0-befa-183b775023e8 · outbound

This paper cites Variance Reduction Tech- niques for Gradient Estimates in Reinforce- ment Learning.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Variance Reduction Tech- niques for Gradient Estimates in Reinforce- ment Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.776040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.776040Z digest=sha256:d7a2c6bf709d4949796d8b026432016eece112f17b9c43117f15ea0955e55151

Observation a002a902-d8b4-45b9-a9e1-a0cc5d4d337e · outbound

This paper cites OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.781021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.781021Z digest=sha256:cf6fa24e4989ce08cc41970200806f8dfc171d58b0ec9a87fe63a401ab0ce81f

Observation 4587cee8-b653-4aba-8d44-95657ecce8f5 · outbound

This paper cites ∆LNormalization: Rethink Loss Aggregation in RL VR.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards ∆LNormalization: Rethink Loss Aggregation in RL VR

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.785322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.785322Z digest=sha256:26779c85b0101286ca502fea51f8240dea614da4546defc8c8f2766db3d88e58

Observation 578bb1d3-8a0f-467f-8119-7cb6e6938399 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Measuring Mathematical Problem Solving With the MATH Dataset

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.790052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.790052Z digest=sha256:9bfaa6700adf3838de477f5ab73ec4e4fa841d5716875c9d8eb343cd68529ad6

Observation 0e1c747e-fd15-4ee6-bb25-3c02c9213ba0 · outbound

This paper cites REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.794213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.794213Z digest=sha256:d9aa7ece5ee80350bb661d77b7bbab4299dddc68cea22b69935160ecf7ef1146

Observation d7b88388-7154-4c8e-a414-6fb6cdcf752d · outbound

This paper cites Buy 4 REINFORCE Samples, Get a Baseline for Free!.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Buy 4 REINFORCE Samples, Get a Baseline for Free!

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.798542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.798542Z digest=sha256:c88ed66e51b7ebcd42e008379af8e224028132fabe23767ffbb42475ad2b9e72

Observation 6f97b477-1bbf-47cc-bcec-5b7fc3c5ca07 · outbound

This paper cites an unresolved cited work.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.802389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.802389Z digest=sha256:4f95356915a17239317caca7ee08b6b869f2f3887aa0c59a5a4743e9190c3a28

Observation ed123b43-a3e8-4bbe-9566-2c76e2078605 · outbound

This paper cites An Exponential Learning Rate Schedule for Deep Learning.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards An Exponential Learning Rate Schedule for Deep Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.806756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.806756Z digest=sha256:8636e39ec82eafb64b16db4dcff1d0c6155f382e4a417e5a31a467b6a9c87de4

Observation f658452d-372c-446c-8f25-3536e859720a · outbound

This paper cites Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.811511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.811511Z digest=sha256:65e9d8f8efac3be15b2b33692eea9cb7e4c30f6b8a5b678b285261100f03863e

Observation cd5d6fde-b090-4875-99ab-5af041d96c5e · outbound

This paper cites ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.815569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.815569Z digest=sha256:bf9d16cd4c2cde9463f7fd94158b10c4c58964e5e4a665b0ca047b2af9133fea

Observation 2481f592-cb45-4efc-8f9d-7efa6fdf1458 · outbound

This paper cites Let's Verify Step by Step.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Let's Verify Step by Step

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.820370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.820370Z digest=sha256:2c8b7a393e30241110214a2d2ca42de9049174188cf43fdae6a1c8bed54dfd57

Observation 61bdd327-aa75-4922-91b9-d084dec88f66 · outbound

This paper cites Hugging Face, 2023.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Hugging Face, 2023

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.824626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.824626Z digest=sha256:a4c3d00d49e54407df57a40f5a0f23fb3b5ce8cd80719bfe9aba76b390910417

Observation 5ee1bef9-544c-4d3f-992d-ad87e0e32ad1 · outbound

This paper cites Training language models to follow instructions with human feedback.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Training language models to follow instructions with human feedback

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.829212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.829212Z digest=sha256:6126236d71b62ea687660a0fbbb9de48a3b238691ff6eedd9529736d9fc90481

Observation 6ca7a5e7-58a6-4d71-967f-36044d761b95 · outbound

This paper cites an unresolved cited work.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.833625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.833625Z digest=sha256:094be8bd5796027ee8a041bfa17b1c372ddba9ad0438b278454381c7c76c601d

Observation fdd799f0-cf68-4120-830f-62c763e1bfe9 · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.838194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.838194Z digest=sha256:b40ba87a7c3c6dc5f56408d8e0580b8cd20e15ae250a9a22cd08b3a97fc39848

Observation 53c43904-127b-41c8-a4d9-a27bff661e3c · outbound

This paper cites an unresolved cited work.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.842538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.842538Z digest=sha256:8aed1ce31a4cf45e0c177b47895b7100f5062c058d957b9e804922928713e215

Observation b2d0d4d4-e43e-4743-b662-f81eecaeff01 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Proximal Policy Optimization Algorithms

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.851666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.851666Z digest=sha256:93133489e0c6caa73e83112d0815615a036258e8a69d585a8a564ecd223d421a

Observation ce0a7930-a724-4ed5-a02c-e53dbd105095 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.856427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.856427Z digest=sha256:c9f5c76868ba1a93e31b6f1481d671be9c44045b8161352ab598faeb32e482f5

Observation 59837580-e41f-4789-bad2-050fb675b09e · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards HybridFlow: A Flexible and Efficient RLHF Framework

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.860531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.860531Z digest=sha256:70e2ec1e3ab24f2dc2105b081843c1c9f318f71714bdb832955545583eef5061

Observation 6f6fa42e-8205-40e2-a885-07d7c53b4bff · outbound

This paper cites Solving Inequality Proofs with Large Language Models.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Solving Inequality Proofs with Large Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.865402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.865402Z digest=sha256:522a3dc92deaf7597f2aa6a9a04e0c4061a65106b9a7d8ee78260e7fd612358b

Observation 03d3676c-ed06-4713-afe5-3a48a78a2df4 · outbound

This paper cites Learning to summarize from human feedback.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Learning to summarize from human feedback

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.869562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.869562Z digest=sha256:d0c41ac16b82eeea1ab319e360dc556d5097e8d2bb73824d8188944eb5bb372f

Observation d354e495-9260-4fa8-a6ce-32bd312f9be2 · outbound

This paper cites Trust Region Preference Approximation: A simple and stable reinforcement learning algorithm for LLM reasoning.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Trust Region Preference Approximation: A simple and stable reinforcement learning algorithm for LLM reasoning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.873682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.873682Z digest=sha256:309c196a14fcf90c97e7219d38784e85607f5f53e30f9b852afa18bfc42c3634

Observation 8d833df9-5d94-440f-816d-05e97aca4710 · outbound

This paper cites Pass@K Policy Optimization: Solving Harder Reinforcement Learning Problems.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Pass@K Policy Optimization: Solving Harder Reinforcement Learning Problems

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.878594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.878594Z digest=sha256:03b2cf21b737eeb6cad352cccac85dec173e6b4d8fd49a19ab5333e2709e8a45

Observation 08af61d4-331e-4101-b547-749b90dc217c · outbound

This paper cites Self-Instruct: Aligning Language Models with Self-Generated Instruc- tions.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Self-Instruct: Aligning Language Models with Self-Generated Instruc- tions

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.883512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.883512Z digest=sha256:b3d20946e84dd4a019a9f6b85db8c57b2d0406bb2404b825a59dcf282cd15ea7

Observation a8e46482-e302-406f-9ac5-dc8dc9079681 · outbound

This paper cites an unresolved cited work.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.887324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.887324Z digest=sha256:d7e0f14ecd48f07ba5322c27f2820d5dd6f3a9bbab5d57c3e5338f4763d8a1fa

Observation 346fc257-fbc5-4f41-afff-457ce96bdd5e · outbound

This paper cites Qwen3 Technical Report.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Qwen3 Technical Report

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.896488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.896488Z digest=sha256:0ee6d4dabac8a169a82db5a650e7b450ed99339fda54f18ca577f095e012ab82

Observation f0549afe-c503-48f6-b347-d34c7c8c723c · outbound

This paper cites Optimizing Chain-of-Thought Reasoners via Gradient Variance Minimization in Rejection Sampling and RL.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Optimizing Chain-of-Thought Reasoners via Gradient Variance Minimization in Rejection Sampling and RL

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.900557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.900557Z digest=sha256:61651591f79d3032ab213da51e746014650014725a0382776af217355767c750

Observation 5d203758-1e76-44d3-9140-7655cad53b77 · outbound

This paper cites OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.891598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.891598Z digest=sha256:4d96f04647a838e39dd318d22b9b7d035ef60c24258a09fc283bb83a78adadd7

Observation 7e36d264-33fa-4622-8aa9-c856b9a9c289 · outbound

This paper cites an unresolved cited work.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Unresolved cited work

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.909663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.909663Z digest=sha256:a1ed1f86e33713ec5bb6291885eeb7fcd4cb873b8e1e3c3815207695634deaec

Observation 860f6358-57cb-4341-b488-26f773d3646a · outbound

This paper cites an unresolved cited work.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Unresolved cited work

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.913719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.913719Z digest=sha256:175f4dd23b6c3cea4c8f1852f4fe065318f32ac5e6222cb98d9fa39585f58701

Observation 75231d1f-e143-44eb-8ebf-f0baf1334369 · outbound

This paper cites an unresolved cited work.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.905721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.905721Z digest=sha256:590e1605cb747f7057c39b8de45ae91fe6fcbde74d117e1a1c4ebf737ba63437

Observation 9a9f983f-0143-4fbe-99b7-51d84f4debc0 · outbound

This paper cites an unresolved cited work.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Unresolved cited work

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.918458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.918458Z digest=sha256:99283f8b8e35e1c8b0d89ad94d0bb89cff5a58666847c0bd717781869f76208a

Observation 20550033-098c-4543-b10d-e51c4554550c · outbound

This paper cites Formally, we consider min {Nt}T−1 t=0 ,{G t}T−1 t=0 E L(θ T ) s.t.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Formally, we consider min {Nt}T−1 t=0 ,{G t}T−1 t=0 E L(θ T ) s.t

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.927539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.927539Z digest=sha256:4d3d33c3ae79b38b289e6c6289e7e1df017d343e6b0a2f431c172819c57b49f5

Observation e08aa9e3-008c-44d0-b1f5-b4efd877269b · outbound

This paper cites Practical recommendations for gradient-based training of deep architectures.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Practical recommendations for gradient-based training of deep architectures

Reference 2012

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.730490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.730490Z digest=sha256:487508dfb72edb98dc4da9f2dda45ed664d6107bcb188ebefb56c6950ed579ab

Observation 0d2ee59d-f3c4-464d-8a00-1d3a097ef619 · outbound

This paper cites Fine-Tuning Language Models with Advantage-Induced Policy Alignment.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Fine-Tuning Language Models with Advantage-Induced Policy Alignment

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.922717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.922717Z digest=sha256:f491720a224a63817789dcec04177f7dd2c085fa3ae9825bfb763761a19e7692

Observation 62febf29-27b9-4adc-8355-6fec57b6f8b8 · outbound

This paper cites Accelerating RL for LLM Reasoning with Optimal Advantage Regression.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Accelerating RL for LLM Reasoning with Optimal Advantage Regression

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.744697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.744697Z digest=sha256:1599adb8013ca7727a9bdf5e377925b8955691ca7d3f466aef0ad972a3bfd378

Pith citing papers

No inbound Pith citation observations are available.