Pith. sign in

Paper Citation Record · LEDGER

RLSR: Reinforcement Learning from Self Reward

As of 22 August 2026, this Paper Citation Record lists 12 of 12 outbound references and 5 inbound Pith citation observations for arXiv:2505.08827.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.08827 v2

Coverage vector

measured 12 of 12 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T22:08:14.380006Z

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T11:40:35.592082Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T21:25:38.647436Z

Reference resolution

12 of 12 outbound references displayed

  • verified exact1
  • verified fuzzy2
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 150bbdbe-00ab-475c-a282-e1d4dc0649a6 · outbound

This paper cites Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision.

RLSR: Reinforcement Learning from Self Reward Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T22:08:14.303725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:08:14.303725Z digest=sha256:0bbbbd2c87843fec712776103b547bbed0657ad7c2ff748e989f97f05b737400

Observation 2421c312-9b83-4b1f-a0eb-d417690eeaeb · outbound

This paper cites Training language models to follow instructions with human feedback.

RLSR: Reinforcement Learning from Self Reward Training language models to follow instructions with human feedback

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T22:08:14.310242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:08:14.310242Z digest=sha256:0ab384736b82c2319d606bea9f0227af5f2c05dbc531528d3f8396a573b76ced

Observation 02b5632a-1dbe-4f0c-897c-b075c6246e53 · outbound

This paper cites Qwen2.5 Technical Report.

RLSR: Reinforcement Learning from Self Reward Qwen2.5 Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T22:08:14.317312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:08:14.317312Z digest=sha256:77440164d110fd83b3bcfe5ed3ecd572039df6a40e6d466b56323076c57ab376

Observation e65875a3-b079-44c9-9bb9-2471363a8f63 · outbound

This paper cites Brief analysis of DeepSeek R1 and its implications for Generative AI.

RLSR: Reinforcement Learning from Self Reward Brief analysis of DeepSeek R1 and its implications for Generative AI

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-15T22:08:14.551971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T22:08:14.324042Z digest=sha256:ffe78c2005f9f349fe894513d029c0d4934b8dc3a83349962be2ca7d2287a4b7

Observation 06b6859c-5575-4d61-aae9-3d662716bd14 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

RLSR: Reinforcement Learning from Self Reward DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T22:08:14.331680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:08:14.331680Z digest=sha256:80e03a228c8536ece3a692dec1f81eb76e782364f7cff16158c0a48e938a19b9

Observation acfa1bba-7b98-4951-8586-a49314443f75 · outbound

This paper cites Tinyzero.https://github.

RLSR: Reinforcement Learning from Self Reward Tinyzero.https://github

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:08:14.647151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T22:08:14.338286Z digest=sha256:c6d57756133f7c133cc642763dee191d8cae3f2125f8edac1e389ff3d674e465

Observation f7042a25-d2eb-4740-ba26-1c2e9ab0f32e · outbound

This paper cites Self-Rewarding Language Models.

RLSR: Reinforcement Learning from Self Reward Self-Rewarding Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T22:08:14.346804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:08:14.346804Z digest=sha256:1ab81f9e38524fa365e35e6b34bfc10cc71ddd24440a7b45c8faa642e39889b8

Observation b61533b0-0246-4603-bf2a-f7752b6e0c6b · outbound

This paper cites LADDER: Self-Improving LLMs Through Recursive Problem Decomposition.

RLSR: Reinforcement Learning from Self Reward LADDER: Self-Improving LLMs Through Recursive Problem Decomposition

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T22:08:14.353504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:08:14.353504Z digest=sha256:4bee193c3327cc33bea8dc0c2f9d31a7f4f7f3ebccc9cf1b23bfd02ffa08ab9b

Observation 3e66ccfd-3195-44b2-ac3e-ed8e964fee0a · outbound

This paper cites WEAK-TO-STRONG GENERALIZATION: ELICITING STRONG CAPABILITIES WITH WEAK SUPERVISION.

RLSR: Reinforcement Learning from Self Reward WEAK-TO-STRONG GENERALIZATION: ELICITING STRONG CAPABILITIES WITH WEAK SUPERVISION

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:08:14.629739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T22:08:14.359518Z digest=sha256:6255f42cb269c99454674ec13abe54f812d9b36ea655cf3b0fafaa66a9dfd162

Observation 32cfb0a2-bf45-47c6-8b78-a45fda698d6d · outbound

This paper cites Supervising strong learners by amplifying weak experts.

RLSR: Reinforcement Learning from Self Reward Supervising strong learners by amplifying weak experts

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T22:08:14.365141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:08:14.365141Z digest=sha256:cc591fc227ab305be3e347aeb5cbdf28357eba02f833dcebcbeca211e9f7e5b9

Observation b90f31ec-350d-4086-a406-bd6824b2a561 · outbound

This paper cites Risks from Learned Optimization in Advanced Machine Learning Systems.

RLSR: Reinforcement Learning from Self Reward Risks from Learned Optimization in Advanced Machine Learning Systems

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T22:08:14.371218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:08:14.371218Z digest=sha256:095e56138f06b98ab20f2c929f8a60b7f9575012359aad08086370da2a3fc040

Observation 7114c5e3-3c39-401c-9e7b-02e4d0bfa71d · outbound

This paper cites CodeRL: Mastering Code Generation through Pretrained Models and Deep Reinforcement Learning.

RLSR: Reinforcement Learning from Self Reward CodeRL: Mastering Code Generation through Pretrained Models and Deep Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T22:08:14.380006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:08:14.380006Z digest=sha256:ef22c6acb52383280cd17e39694b03ef52a5c84aa76043231fbb9dc900744717

Pith citing papers

Observation 00c1700a-e9b8-41b0-a2e4-0874cb7e3076 · inbound

Self-Rewarding Vision-Language Model via Reasoning Decomposition cites this paper.

Self-Rewarding Vision-Language Model via Reasoning Decomposition RLSR: Reinforcement Learning from Self Reward

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-18T21:06:50.865244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-18T21:03:31.606674Z digest=sha256:5485c6691c63a6c5421a55eec309cafd7527264a56c1dc96fae24f1aa74c18da

Observation b784f6f3-78d9-4dde-b611-bb52576ec7e8 · inbound

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration cites this paper.

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration RLSR: Reinforcement Learning from Self Reward

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T12:10:23.385138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:cc66873594f0eac402c37cb826c16a31708a259b35e650447724b40bbd3cb13b

Observation 1eec3f3a-b3b7-4bd3-adb4-59b061e12f4f · inbound

Primal Generation, Dual Judgment: Self-Training from Test-Time Scaling cites this paper.

Primal Generation, Dual Judgment: Self-Training from Test-Time Scaling RLSR: Reinforcement Learning from Self Reward

Reference 101

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T20:47:58.125632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-14T20:46:38.558668Z digest=sha256:2724b3fdd8622d4a02e8b01cee4102392e2971ee20a637858a8f5d8d3d5ae599

Observation 74981873-42cf-49ce-8db6-a0512b782060 · inbound

More Convincing, Not More Correct: Self-Play Reward Hacking of Reference-Free LLM Judges cites this paper.

More Convincing, Not More Correct: Self-Play Reward Hacking of Reference-Free LLM Judges RLSR: Reinforcement Learning from Self Reward

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T21:25:38.648851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-08T21:20:07.569587Z digest=sha256:9e8084e0c33f6b3ef14c78ec59c2f9420e3e020e454e34905375e7bf7458474c

Observation e9e496ea-6c54-46c8-a5e7-970221b930b3 · inbound

Rewarding Better Thinking for LLM Preference Alignment cites this paper.

Rewarding Better Thinking for LLM Preference Alignment RLSR: Reinforcement Learning from Self Reward

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:35.592082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:35.592082Z digest=sha256:c3740322ad33ade70dc0e52c86915f6183fe0c1081f850210f01099932cc8edb