Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T22:08:14.380006Z
Paper Citation Record · LEDGER
As of 22 August 2026, this Paper Citation Record lists 12 of 12 outbound references and 5 inbound Pith citation observations for arXiv:2505.08827.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T22:08:14.380006Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-01T11:40:35.592082Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-08T21:25:38.647436Z
12 of 12 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 150bbdbe-00ab-475c-a282-e1d4dc0649a6 · outbound
RLSR: Reinforcement Learning from Self Reward Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2421c312-9b83-4b1f-a0eb-d417690eeaeb · outbound
RLSR: Reinforcement Learning from Self Reward Training language models to follow instructions with human feedback
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02b5632a-1dbe-4f0c-897c-b075c6246e53 · outbound
RLSR: Reinforcement Learning from Self Reward Qwen2.5 Technical Report
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e65875a3-b079-44c9-9bb9-2471363a8f63 · outbound
RLSR: Reinforcement Learning from Self Reward Brief analysis of DeepSeek R1 and its implications for Generative AI
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 06b6859c-5575-4d61-aae9-3d662716bd14 · outbound
RLSR: Reinforcement Learning from Self Reward DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation acfa1bba-7b98-4951-8586-a49314443f75 · outbound
RLSR: Reinforcement Learning from Self Reward Tinyzero.https://github
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation f7042a25-d2eb-4740-ba26-1c2e9ab0f32e · outbound
RLSR: Reinforcement Learning from Self Reward Self-Rewarding Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b61533b0-0246-4603-bf2a-f7752b6e0c6b · outbound
RLSR: Reinforcement Learning from Self Reward LADDER: Self-Improving LLMs Through Recursive Problem Decomposition
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e66ccfd-3195-44b2-ac3e-ed8e964fee0a · outbound
RLSR: Reinforcement Learning from Self Reward WEAK-TO-STRONG GENERALIZATION: ELICITING STRONG CAPABILITIES WITH WEAK SUPERVISION
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 32cfb0a2-bf45-47c6-8b78-a45fda698d6d · outbound
RLSR: Reinforcement Learning from Self Reward Supervising strong learners by amplifying weak experts
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b90f31ec-350d-4086-a406-bd6824b2a561 · outbound
RLSR: Reinforcement Learning from Self Reward Risks from Learned Optimization in Advanced Machine Learning Systems
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7114c5e3-3c39-401c-9e7b-02e4d0bfa71d · outbound
RLSR: Reinforcement Learning from Self Reward CodeRL: Mastering Code Generation through Pretrained Models and Deep Reinforcement Learning
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00c1700a-e9b8-41b0-a2e4-0874cb7e3076 · inbound
Self-Rewarding Vision-Language Model via Reasoning Decomposition RLSR: Reinforcement Learning from Self Reward
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation b784f6f3-78d9-4dde-b611-bb52576ec7e8 · inbound
Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration RLSR: Reinforcement Learning from Self Reward
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 1eec3f3a-b3b7-4bd3-adb4-59b061e12f4f · inbound
Primal Generation, Dual Judgment: Self-Training from Test-Time Scaling RLSR: Reinforcement Learning from Self Reward
Reference 101
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 74981873-42cf-49ce-8db6-a0512b782060 · inbound
More Convincing, Not More Correct: Self-Play Reward Hacking of Reference-Free LLM Judges RLSR: Reinforcement Learning from Self Reward
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation e9e496ea-6c54-46c8-a5e7-970221b930b3 · inbound
Rewarding Better Thinking for LLM Preference Alignment RLSR: Reinforcement Learning from Self Reward
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.