Pith. sign in

Paper Citation Record · LEDGER

Polychromic Objectives for Reinforcement Learning

As of 5 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 0 inbound Pith citation observations for arXiv:2509.25424.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.25424 v6

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-18T11:54:29.955833Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

49 of 49 outbound references displayed

  • verified exact32
  • verified fuzzy12
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch4

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0e4122ff-ca3b-4075-a6d5-c739314e406b · outbound

This paper cites Minimax Regret Bounds for Reinforcement Learning.

Polychromic Objectives for Reinforcement Learning Minimax Regret Bounds for Reinforcement Learning

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-18T11:56:19.827701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T11:54:29.955833Z digest=sha256:9fce4eca2dd40c0779c580f97e6907f26a142250b1aae2b6da4783a52f07ee85

Observation 1914c4a8-fa7c-414b-b0ec-14e3df6247d2 · outbound

This paper cites Sutton, Mohammad Ghavamzadeh, and Mark Lee.

Polychromic Objectives for Reinforcement Learning Sutton, Mohammad Ghavamzadeh, and Mark Lee

Reference 2

Resolution
verified exact
doi, observed 2026-05-18T11:56:19.701442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T11:54:29.955833Z digest=sha256:1e03485ace0cf8fcc1f8b21fd2cc5e32a4b1d1f5123c24b3bf5b25cfe87b3eda

Observation b9a0af0a-d0ac-4be2-be05-b257b6ddbf6e · outbound

This paper cites Babyai: A platform to study the sample ef- ficiency of grounded language learning.

Polychromic Objectives for Reinforcement Learning Babyai: A platform to study the sample ef- ficiency of grounded language learning

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T11:56:20.199891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T11:54:29.955833Z digest=sha256:61695aab3c25c4c432c6fe4fb860ddc73f8b68e8b33b8ec6fd442f2524698172

Observation c0e503b3-e523-4b8f-a701-ff9a49998555 · outbound

This paper cites Minigrid & Miniworld: Modular & Customizable Reinforcement Learning Environments for Goal-Oriented Tasks.

Polychromic Objectives for Reinforcement Learning Minigrid & Miniworld: Modular & Customizable Reinforcement Learning Environments for Goal-Oriented Tasks

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-18T11:56:20.033957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T11:54:29.955833Z digest=sha256:aa0419931c9e1eb0ce38034da96e692b150ec865ec36e99e6291bf8032b4bc5b

Observation a28d9f83-96a2-4709-a56a-238ee8093fd0 · outbound

This paper cites Inference-aware fine-tuning for best-of-n sampling in large language models.

Polychromic Objectives for Reinforcement Learning Inference-aware fine-tuning for best-of-n sampling in large language models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-18T11:56:19.894737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T11:54:29.955833Z digest=sha256:8f874299d6d562e47c7aabcb15090fd8fce938610a57144d4027631f344390b3

Observation dd7549a0-eae4-444e-8696-951a6f767169 · outbound

This paper cites The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models.

Polychromic Objectives for Reinforcement Learning The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-18T11:56:19.994404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T11:54:29.955833Z digest=sha256:e3d4af74bd7917ce1c7af016b507244f94f7a249e0a96ebac7c5f33390827e46

Observation 85e8162f-36cd-4e2d-a5b7-0d20e902af26 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Polychromic Objectives for Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-18T11:56:19.833356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T11:54:29.955833Z digest=sha256:168c91ef985d5856858ff5b0421cf2c108390749bb166f8b067af21bdaa56111

Observation 234935e7-595c-47ca-aef3-af07a69a4801 · outbound

This paper cites Off-Policy Actor-Critic.

Polychromic Objectives for Reinforcement Learning Off-Policy Actor-Critic

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-18T11:56:19.983519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T11:54:29.955833Z digest=sha256:19c0f649a720c1b1a58389c89114bfe212958452804de67f95b2f5432bd44b4b

Observation b04fe9f6-bf56-4d94-a267-a3711a02feae · outbound

This paper cites The Vendi Score: A Diversity Evaluation Metric for Machine Learning.

Polychromic Objectives for Reinforcement Learning The Vendi Score: A Diversity Evaluation Metric for Machine Learning

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-18T11:56:20.028364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T11:54:29.955833Z digest=sha256:9ec42eddde0443dc267ab359c6966f5391d93eaff806183540cf0e205d10b933

Observation 28dd0253-8e81-4ca1-9bec-701910b0481a · outbound

This paper cites Reinforcement Learning with Deep Energy-Based Policies.

Polychromic Objectives for Reinforcement Learning Reinforcement Learning with Deep Energy-Based Policies

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-18T11:56:19.864851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T11:54:29.955833Z digest=sha256:74f9b2b4993fce876e944994c326fcccc34c1fe6081f2dd17e5142b2633d9765

Observation 8b3fdffd-3dba-4940-8a86-87cb9cda0513 · outbound

This paper cites Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor.

Polychromic Objectives for Reinforcement Learning Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-18T11:56:19.859035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T11:54:29.955833Z digest=sha256:0d9912e1286e856ec8866fe6c49a114c247acdd8da3efd04e68b5856b906d4b7

Observation ba013230-dea9-4666-b986-e95ca2f46054 · outbound

This paper cites Rewarding the Unlikely: Lifting GRPO Beyond Distribution Sharpening.

Polychromic Objectives for Reinforcement Learning Rewarding the Unlikely: Lifting GRPO Beyond Distribution Sharpening

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-18T11:56:19.854228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T11:54:29.955833Z digest=sha256:6252301e063a7ded5c56041d7cb18258d4e5b506ac6037d77dd8482d348bd6f4

Observation b736e6f2-fc7e-4476-b9b8-fb85ef0ec722 · outbound

This paper cites Marginalized state distribution entropy regularization in policy optimization.

Polychromic Objectives for Reinforcement Learning Marginalized state distribution entropy regularization in policy optimization

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T11:56:20.193050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T11:54:29.955833Z digest=sha256:72eaf9a8c975cd7a61dc9e73685a1b4db9123ccabaa2fbc3c57975f4c92b954f

Observation 891ae641-6a20-459c-b44b-17d9b4e287e5 · outbound

This paper cites A natural policy gradient.

Polychromic Objectives for Reinforcement Learning A natural policy gradient

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T11:56:20.196096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T11:54:29.955833Z digest=sha256:ca9889a1afdb65774a0136655e9ca7b0437e5794aedcb9fde5a1da0fd10df5ac

Observation c31ca28c-2017-47c3-94e0-9fed02aa8782 · outbound

This paper cites an unresolved cited work.

Polychromic Objectives for Reinforcement Learning Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-05-18T11:56:20.189783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T11:54:29.955833Z digest=sha256:3d17d67117c7bc60906e1939768462d3acbbc2825e26c92303d774a11400468d

Observation 15c00602-be23-4d54-a125-c55e43d1c737 · outbound

This paper cites Kakade and John Langford.

Polychromic Objectives for Reinforcement Learning Kakade and John Langford

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T11:56:20.177156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T11:54:29.955833Z digest=sha256:0f4bd2cb2e33010cbecacff37c26125c3e5c21e9e28433037bd5ecb11c80d415

Observation 1fb923ae-a3f1-4e42-be94-8fce0d68c6c1 · outbound

This paper cites VinePPO: Refining Credit Assignment in RL Training of LLMs.

Polychromic Objectives for Reinforcement Learning VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T11:56:19.879623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T11:54:29.955833Z digest=sha256:4bb1cf53befc71c9c3d318f6e0318017620bf2b7e3e9854d775f9ba0156c0658

Observation e3ec2f14-269f-44eb-8395-1e8070994e67 · outbound

This paper cites One Solution is Not All You Need: Few-Shot Extrapolation via Structured MaxEnt RL.

Polychromic Objectives for Reinforcement Learning One Solution is Not All You Need: Few-Shot Extrapolation via Structured MaxEnt RL

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-18T11:56:19.844565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T11:54:29.955833Z digest=sha256:ef281a632edd4175b4a33f6174ba32b27ea4dfc70f0d1d440cacb5a2d977ae2e

Observation 1bdbf20f-7ede-4e7f-b3ea-34698fa88b5a · outbound

This paper cites Diverse Preference Optimization.

Polychromic Objectives for Reinforcement Learning Diverse Preference Optimization

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T11:56:19.999530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T11:54:29.955833Z digest=sha256:b32cee38869ab7108f5c9d1272d868fdbe544b58b57302ad1914d76ca8d5bd1f

Observation d0ae5256-a762-42b8-908d-a7886cedb03b · outbound

This paper cites Jointly Reinforcing Diversity and Quality in Language Model Generations.

Polychromic Objectives for Reinforcement Learning Jointly Reinforcing Diversity and Quality in Language Model Generations

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-18T11:56:20.020157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T11:54:29.955833Z digest=sha256:16921bd8238d6178ba56ba526bace2bc2f9079fb0988abd516972dc383e05572

Observation 60c55b26-fb85-4d1a-9bf1-31bd405cbe5c · outbound

This paper cites Lillicrap, Jonathan J.

Polychromic Objectives for Reinforcement Learning Lillicrap, Jonathan J

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T11:56:20.183627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T11:54:29.955833Z digest=sha256:cae30cbeab852aa4be037fbe326b4174cd4487c5a13741c32b19d08454e95ee9

Observation a255ff8d-3236-48bd-933c-e4d4ee1114dc · outbound

This paper cites Continuous control with deep reinforcement learning.

Polychromic Objectives for Reinforcement Learning Continuous control with deep reinforcement learning

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-18T11:56:20.039145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T11:54:29.955833Z digest=sha256:12c966fc39f9a2a7618ac9425a861b8a88518396646a9ab66adb5789c252d0f9

Observation 60d0afb0-2839-4ec2-86e3-03d148585d63 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

Polychromic Objectives for Reinforcement Learning Understanding R1-Zero-Like Training: A Critical Perspective

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-18T11:56:19.822497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T11:54:29.955833Z digest=sha256:4f3936243cbe10a6b7044ff8439f787f8a3ff92e558b83756804568106dd770c

Observation d6bc9b26-11ab-4028-9d08-77562a8bf11e · outbound

This paper cites Roll the dice & look before you leap: Going beyond the creative limits of next-token prediction.

Polychromic Objectives for Reinforcement Learning Roll the dice & look before you leap: Going beyond the creative limits of next-token prediction

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-18T11:56:19.989292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T11:54:29.955833Z digest=sha256:9ac294d58a47ccf9493dc7e39ffeb96533ea89e787d916d4c7c1f1614e53c1f7

Observation aa64cf93-90a0-4ba0-851a-a2a8f0997d06 · outbound

This paper cites OpenAI o1 System Card.

Polychromic Objectives for Reinforcement Learning OpenAI o1 System Card

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-18T11:56:19.849526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T11:54:29.955833Z digest=sha256:adc06ac315a11dd6d861955ca7739e143362bd91025c36719ca50dd4c8f5ed56

Observation 76417b47-4424-4e6d-839e-4fd20980bfe9 · outbound

This paper cites Training language models to follow instructions with human feedback.

Polychromic Objectives for Reinforcement Learning Training language models to follow instructions with human feedback

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-18T11:56:20.010004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T11:54:29.955833Z digest=sha256:798cd0622f01acf4a8ae1f2abe45fa4e3d2646aac97917e247d95c9fe773b7e5

Observation 630d9676-82fc-46e6-a947-a643463414ce · outbound

This paper cites Efros, and Trevor Darrell.

Polychromic Objectives for Reinforcement Learning Efros, and Trevor Darrell

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T11:56:20.186627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T11:54:29.955833Z digest=sha256:179a96f5d4aaaf896ebddba7690fdfb6908c5063e625ae097cd99fa4c3a9a65c

Observation 6039360e-be90-4733-b6b6-5d194c6b1fa8 · outbound

This paper cites Qwq-32b: Embracing the power of reinforcement learning, March 2025.

Polychromic Objectives for Reinforcement Learning Qwq-32b: Embracing the power of reinforcement learning, March 2025

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T11:56:20.180321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T11:54:29.955833Z digest=sha256:e0756b67f660994d5d00aba4af37ae5e20ba525f017baa112a6f503b34dc0a44

Observation 37f7dbc8-c5db-4fbd-84f9-c811685013d1 · outbound

This paper cites Tapered Off-Policy REINFORCE: Stable and efficient reinforcement learning for LLMs.

Polychromic Objectives for Reinforcement Learning Tapered Off-Policy REINFORCE: Stable and efficient reinforcement learning for LLMs

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T11:56:19.889672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T11:54:29.955833Z digest=sha256:1841bb043ec48c435bc3452ead9d58d32f098384d78d2dfaff72d242347730b5

Observation 8bb215ef-ad8b-4635-a89c-be7418a0fa5d · outbound

This paper cites Trust Region Policy Optimization.

Polychromic Objectives for Reinforcement Learning Trust Region Policy Optimization

Reference 30

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T11:56:19.874820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T11:54:29.955833Z digest=sha256:2c0ce9cf7be0b61605bd39ae61cdb7968d8c5b4258c7aeeea725740d86b7b312

Observation 76d28e10-b94c-4c31-aaf5-dd2a4298d077 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Polychromic Objectives for Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-18T11:56:19.900579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T11:54:29.955833Z digest=sha256:6a1fe018ff0ac620411873f30dbae24bae682d6166871306c3b7e976f36cd10b

Observation af8158fc-415f-44d3-86bc-83f44bd71a03 · outbound

This paper cites High-Dimensional Continuous Control Using Generalized Advantage Estimation.

Polychromic Objectives for Reinforcement Learning High-Dimensional Continuous Control Using Generalized Advantage Estimation

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-18T11:56:19.965871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T11:54:29.955833Z digest=sha256:8594fbd55f3c6e7ed3aca939be705bb127a7880ba28beea2699ae6340304617b

Observation a692778f-f9a5-4280-9270-d63c135da3be · outbound

This paper cites State Entropy Maximization with Random Encoders for Efficient Exploration.

Polychromic Objectives for Reinforcement Learning State Entropy Maximization with Random Encoders for Efficient Exploration

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-18T11:56:19.978592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T11:54:29.955833Z digest=sha256:3d3f0f00ebd1e9ae36c3f2335fddbb073d8b0b7e1ba45367f977a58abbbe9b8f

Observation b7bb512f-69f9-4d5d-9625-cb1889f37d94 · outbound

This paper cites Deterministic policy gradient algorithms.

Polychromic Objectives for Reinforcement Learning Deterministic policy gradient algorithms

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T11:56:20.161695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T11:54:29.955833Z digest=sha256:866a573f1458ec8c332829753cdb352ddc8671b4d86e3b1bb1e1e78523173bdd

Observation d4ccc6a4-9dbb-47c8-9534-019b9f48bcab · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Polychromic Objectives for Reinforcement Learning Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-18T11:56:19.905594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T11:54:29.955833Z digest=sha256:68304b4ea5a7772d5f4f1ce1a3f607672c0d48c117a2596583ece88a6d1d0670

Observation 72fefe37-dd24-4dca-8a40-c7316b8ce585 · outbound

This paper cites Outcome-based exploration for llm reasoning.

Polychromic Objectives for Reinforcement Learning Outcome-based exploration for llm reasoning

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T11:56:20.168482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T11:54:29.955833Z digest=sha256:46a61583e0d510ae5c0103fd0d8c1647bb8280ff53c94e026ea963b74994df42

Observation 93fc4f7e-c621-4a05-9be5-d73b16ef952b · outbound

This paper cites Outcome-based Exploration for LLM Reasoning.

Polychromic Objectives for Reinforcement Learning Outcome-based Exploration for LLM Reasoning

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-18T11:56:20.005263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T11:54:29.955833Z digest=sha256:9291f67b176aebdbda5c3db21a281c88558a6075cd263cfedc36ee19c0e0b2e9

Observation a0f56d34-ea0b-4dbf-8066-aee5c87193a0 · outbound

This paper cites Sutton and Andrew G.

Polychromic Objectives for Reinforcement Learning Sutton and Andrew G

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T11:56:20.174400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T11:54:29.955833Z digest=sha256:52a41d277d5eb8c706d58d1f3375525d95ddc900912c1ef77c3be9f5a36cbca3

Observation cacd5bfc-388b-4140-be75-d63d9f2713f7 · outbound

This paper cites Sutton, David McAllester, Satinder Singh, and Yishay Mansour.

Polychromic Objectives for Reinforcement Learning Sutton, David McAllester, Satinder Singh, and Yishay Mansour

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T11:56:20.171628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T11:54:29.955833Z digest=sha256:b49c348e5e5a57b7ee591b352b52e63a3243de866f308225c21b0682c375bc59

Observation d88b1852-3a90-4dcb-bae8-f1db68cd280a · outbound

This paper cites Optimizing Language Models for Inference Time Objectives using Reinforcement Learning.

Polychromic Objectives for Reinforcement Learning Optimizing Language Models for Inference Time Objectives using Reinforcement Learning

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-18T11:56:19.949595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T11:54:29.955833Z digest=sha256:6bb3168b9f6c378f17c4fad2e7190348671e664d53b78c67df88dd36e7bd2ea9

Observation e13e429c-6794-42ae-bd0c-04190cecee3f · outbound

This paper cites Sample Efficient Actor-Critic with Experience Replay.

Polychromic Objectives for Reinforcement Learning Sample Efficient Actor-Critic with Experience Replay

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-05-18T11:56:19.973162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T11:54:29.955833Z digest=sha256:17a3f213cde4a6499faf64e624f449e3be134354efe6e9206565cb2f095727fe

Observation 1c486f10-7be1-4d2e-b8c6-5bcf9cea97fd · outbound

This paper cites Williams.

Polychromic Objectives for Reinforcement Learning Williams

Reference 42

Resolution
verified exact
doi, observed 2026-05-18T11:56:19.717406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T11:54:29.955833Z digest=sha256:f6ded0e31d07cca6a67af21613f8e9f73b40af556db3b4deae076cc7bc8a58e8

Observation e066ec79-ab5f-434f-bbb9-981820444130 · outbound

This paper cites The invisible leash: Why rlvr may or may not escape its origin.

Polychromic Objectives for Reinforcement Learning The invisible leash: Why rlvr may or may not escape its origin

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-18T11:56:19.884243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T11:54:29.955833Z digest=sha256:b63a221d8adbba18eed717a8b8c2e7821ab2da006a386a9c550e0a39f1da8f6a

Observation 426ee7b8-c965-4566-a4da-0495b0120ce6 · outbound

This paper cites Younis, Rodrigo Perez-Vicente, John U.

Polychromic Objectives for Reinforcement Learning Younis, Rodrigo Perez-Vicente, John U

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T11:56:20.164947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T11:54:29.955833Z digest=sha256:9a5e1bfca6a34e51defb292ab85b1ce907dfd3e432b0fcb0f8f811dea05e7e75

Observation 00f8689d-e627-40f2-bc05-57a53958efb7 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Polychromic Objectives for Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-05-18T11:56:19.839256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T11:54:29.955833Z digest=sha256:8d1520e8d952c3d4b3047090d9632f2c0b894243073981bad7dc0cc78b36e6c6

Observation 00fbc4c2-9c5b-40c1-aa1b-b0aefdd8721c · outbound

This paper cites Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?.

Polychromic Objectives for Reinforcement Learning Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-05-18T11:56:19.869625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T11:54:29.955833Z digest=sha256:32ba0b343e8b4ce1b42c6edfac2a1e922577dd9307b0bda5d06f7b0d7dacfeaf

Observation 41e0bbe0-176c-4be0-b935-3b4d5c574488 · outbound

This paper cites NoveltyBench: Evaluating Language Models for Humanlike Diversity.

Polychromic Objectives for Reinforcement Learning NoveltyBench: Evaluating Language Models for Humanlike Diversity

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-18T11:56:19.942452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T11:54:29.955833Z digest=sha256:f4403d186299ff144b822be93d1bb1ece373bb1b471634d924875a0b6c7984ce

Observation f5cb03da-797e-420f-9839-d10cf6717f17 · outbound

This paper cites Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining.

Polychromic Objectives for Reinforcement Learning Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-18T11:56:19.958923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T11:54:29.955833Z digest=sha256:bc2f7e564a15135c39120ebdcfb313973865e01a3afc9acb414406a197e5d51d

Observation 2c8cc3e7-1bb4-45ac-b3e7-7c18835230f0 · outbound

This paper cites Group Sequence Policy Optimization.

Polychromic Objectives for Reinforcement Learning Group Sequence Policy Optimization

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-05-18T11:56:20.014957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T11:54:29.955833Z digest=sha256:e7a2b94ef13c6981ce1928931913327744fd5cb71a1ef389d79fa230aa8c5547

Pith citing papers

No inbound Pith citation observations are available.