Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-18T11:54:29.955833Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 0 inbound Pith citation observations for arXiv:2509.25424.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-18T11:54:29.955833Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
49 of 49 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 0e4122ff-ca3b-4075-a6d5-c739314e406b · outbound
Polychromic Objectives for Reinforcement Learning Minimax Regret Bounds for Reinforcement Learning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1914c4a8-fa7c-414b-b0ec-14e3df6247d2 · outbound
Polychromic Objectives for Reinforcement Learning Sutton, Mohammad Ghavamzadeh, and Mark Lee
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b9a0af0a-d0ac-4be2-be05-b257b6ddbf6e · outbound
Polychromic Objectives for Reinforcement Learning Babyai: A platform to study the sample ef- ficiency of grounded language learning
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c0e503b3-e523-4b8f-a701-ff9a49998555 · outbound
Polychromic Objectives for Reinforcement Learning Minigrid & Miniworld: Modular & Customizable Reinforcement Learning Environments for Goal-Oriented Tasks
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a28d9f83-96a2-4709-a56a-238ee8093fd0 · outbound
Polychromic Objectives for Reinforcement Learning Inference-aware fine-tuning for best-of-n sampling in large language models
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation dd7549a0-eae4-444e-8696-951a6f767169 · outbound
Polychromic Objectives for Reinforcement Learning The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 85e8162f-36cd-4e2d-a5b7-0d20e902af26 · outbound
Polychromic Objectives for Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 234935e7-595c-47ca-aef3-af07a69a4801 · outbound
Polychromic Objectives for Reinforcement Learning Off-Policy Actor-Critic
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b04fe9f6-bf56-4d94-a267-a3711a02feae · outbound
Polychromic Objectives for Reinforcement Learning The Vendi Score: A Diversity Evaluation Metric for Machine Learning
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 28dd0253-8e81-4ca1-9bec-701910b0481a · outbound
Polychromic Objectives for Reinforcement Learning Reinforcement Learning with Deep Energy-Based Policies
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8b3fdffd-3dba-4940-8a86-87cb9cda0513 · outbound
Polychromic Objectives for Reinforcement Learning Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ba013230-dea9-4666-b986-e95ca2f46054 · outbound
Polychromic Objectives for Reinforcement Learning Rewarding the Unlikely: Lifting GRPO Beyond Distribution Sharpening
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b736e6f2-fc7e-4476-b9b8-fb85ef0ec722 · outbound
Polychromic Objectives for Reinforcement Learning Marginalized state distribution entropy regularization in policy optimization
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 891ae641-6a20-459c-b44b-17d9b4e287e5 · outbound
Polychromic Objectives for Reinforcement Learning A natural policy gradient
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c31ca28c-2017-47c3-94e0-9fed02aa8782 · outbound
Polychromic Objectives for Reinforcement Learning Unresolved cited work
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 15c00602-be23-4d54-a125-c55e43d1c737 · outbound
Polychromic Objectives for Reinforcement Learning Kakade and John Langford
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1fb923ae-a3f1-4e42-be94-8fce0d68c6c1 · outbound
Polychromic Objectives for Reinforcement Learning VinePPO: Refining Credit Assignment in RL Training of LLMs
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e3ec2f14-269f-44eb-8395-1e8070994e67 · outbound
Polychromic Objectives for Reinforcement Learning One Solution is Not All You Need: Few-Shot Extrapolation via Structured MaxEnt RL
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1bdbf20f-7ede-4e7f-b3ea-34698fa88b5a · outbound
Polychromic Objectives for Reinforcement Learning Diverse Preference Optimization
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d0ae5256-a762-42b8-908d-a7886cedb03b · outbound
Polychromic Objectives for Reinforcement Learning Jointly Reinforcing Diversity and Quality in Language Model Generations
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 60c55b26-fb85-4d1a-9bf1-31bd405cbe5c · outbound
Polychromic Objectives for Reinforcement Learning Lillicrap, Jonathan J
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a255ff8d-3236-48bd-933c-e4d4ee1114dc · outbound
Polychromic Objectives for Reinforcement Learning Continuous control with deep reinforcement learning
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 60d0afb0-2839-4ec2-86e3-03d148585d63 · outbound
Polychromic Objectives for Reinforcement Learning Understanding R1-Zero-Like Training: A Critical Perspective
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d6bc9b26-11ab-4028-9d08-77562a8bf11e · outbound
Polychromic Objectives for Reinforcement Learning Roll the dice & look before you leap: Going beyond the creative limits of next-token prediction
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation aa64cf93-90a0-4ba0-851a-a2a8f0997d06 · outbound
Polychromic Objectives for Reinforcement Learning OpenAI o1 System Card
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 76417b47-4424-4e6d-839e-4fd20980bfe9 · outbound
Polychromic Objectives for Reinforcement Learning Training language models to follow instructions with human feedback
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 630d9676-82fc-46e6-a947-a643463414ce · outbound
Polychromic Objectives for Reinforcement Learning Efros, and Trevor Darrell
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6039360e-be90-4733-b6b6-5d194c6b1fa8 · outbound
Polychromic Objectives for Reinforcement Learning Qwq-32b: Embracing the power of reinforcement learning, March 2025
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 37f7dbc8-c5db-4fbd-84f9-c811685013d1 · outbound
Polychromic Objectives for Reinforcement Learning Tapered Off-Policy REINFORCE: Stable and efficient reinforcement learning for LLMs
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8bb215ef-ad8b-4635-a89c-be7418a0fa5d · outbound
Polychromic Objectives for Reinforcement Learning Trust Region Policy Optimization
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 76d28e10-b94c-4c31-aaf5-dd2a4298d077 · outbound
Polychromic Objectives for Reinforcement Learning Proximal Policy Optimization Algorithms
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation af8158fc-415f-44d3-86bc-83f44bd71a03 · outbound
Polychromic Objectives for Reinforcement Learning High-Dimensional Continuous Control Using Generalized Advantage Estimation
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a692778f-f9a5-4280-9270-d63c135da3be · outbound
Polychromic Objectives for Reinforcement Learning State Entropy Maximization with Random Encoders for Efficient Exploration
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b7bb512f-69f9-4d5d-9625-cb1889f37d94 · outbound
Polychromic Objectives for Reinforcement Learning Deterministic policy gradient algorithms
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d4ccc6a4-9dbb-47c8-9534-019b9f48bcab · outbound
Polychromic Objectives for Reinforcement Learning Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 72fefe37-dd24-4dca-8a40-c7316b8ce585 · outbound
Polychromic Objectives for Reinforcement Learning Outcome-based exploration for llm reasoning
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 93fc4f7e-c621-4a05-9be5-d73b16ef952b · outbound
Polychromic Objectives for Reinforcement Learning Outcome-based Exploration for LLM Reasoning
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a0f56d34-ea0b-4dbf-8066-aee5c87193a0 · outbound
Polychromic Objectives for Reinforcement Learning Sutton and Andrew G
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cacd5bfc-388b-4140-be75-d63d9f2713f7 · outbound
Polychromic Objectives for Reinforcement Learning Sutton, David McAllester, Satinder Singh, and Yishay Mansour
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d88b1852-3a90-4dcb-bae8-f1db68cd280a · outbound
Polychromic Objectives for Reinforcement Learning Optimizing Language Models for Inference Time Objectives using Reinforcement Learning
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e13e429c-6794-42ae-bd0c-04190cecee3f · outbound
Polychromic Objectives for Reinforcement Learning Sample Efficient Actor-Critic with Experience Replay
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1c486f10-7be1-4d2e-b8c6-5bcf9cea97fd · outbound
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e066ec79-ab5f-434f-bbb9-981820444130 · outbound
Polychromic Objectives for Reinforcement Learning The invisible leash: Why rlvr may or may not escape its origin
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 426ee7b8-c965-4566-a4da-0495b0120ce6 · outbound
Polychromic Objectives for Reinforcement Learning Younis, Rodrigo Perez-Vicente, John U
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 00f8689d-e627-40f2-bc05-57a53958efb7 · outbound
Polychromic Objectives for Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 00fbc4c2-9c5b-40c1-aa1b-b0aefdd8721c · outbound
Polychromic Objectives for Reinforcement Learning Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 41e0bbe0-176c-4be0-b935-3b4d5c574488 · outbound
Polychromic Objectives for Reinforcement Learning NoveltyBench: Evaluating Language Models for Humanlike Diversity
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f5cb03da-797e-420f-9839-d10cf6717f17 · outbound
Polychromic Objectives for Reinforcement Learning Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2c8cc3e7-1bb4-45ac-b3e7-7c18835230f0 · outbound
Polychromic Objectives for Reinforcement Learning Group Sequence Policy Optimization
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
No inbound Pith citation observations are available.