Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T13:35:56.253150Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 16 inbound Pith citation observations for arXiv:2505.21493.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T13:35:56.253150Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:26:22.158083Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-01T20:56:13.505929Z
60 of 60 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation a93d6290-e684-40aa-84ad-0648f6d7895e · outbound
Reinforcing General Reasoning without Verifiers Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b8912b9-44d6-4993-8c75-365c99efb6f4 · outbound
Reinforcing General Reasoning without Verifiers Unresolved cited work
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation da40036f-4d63-460f-9951-a80959496b15 · outbound
Reinforcing General Reasoning without Verifiers Bootstrapping language models with dpo implicit rewards
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 52ebb6e6-cb33-47bf-9fd6-91affe7684c7 · outbound
Reinforcing General Reasoning without Verifiers Language Models are Hidden Reasoners: Unlocking Latent Reasoning Capabilities via Self-Rewarding
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81e3becc-9596-4c67-8111-196dba531996 · outbound
Reinforcing General Reasoning without Verifiers Training Verifiers to Solve Math Word Problems
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d669c5b-0d3a-4ba5-9a8e-f0cf67637d61 · outbound
Reinforcing General Reasoning without Verifiers SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a802705-31d5-4dfe-9f52-c59d5cddffa1 · outbound
Reinforcing General Reasoning without Verifiers Scaling laws for reward model overoptimization
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41cd47ac-0822-4c76-8982-68592cfe2bbc · outbound
Reinforcing General Reasoning without Verifiers RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5dade0db-61c1-4c40-8094-7951010b67e5 · outbound
Reinforcing General Reasoning without Verifiers The Llama 3 Herd of Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f51e8249-efdf-4f50-a80c-83e7ace68b80 · outbound
Reinforcing General Reasoning without Verifiers DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 532f803e-3d01-46a4-aeea-6d6fa3efeaf9 · outbound
Reinforcing General Reasoning without Verifiers Olympiadbench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems, 2024
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ab29f53-b30b-494d-88ee-a43a1f872e9f · outbound
Reinforcing General Reasoning without Verifiers Measuring Mathematical Problem Solving With the MATH Dataset
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3317d73c-bec0-4a93-9f00-2ebfaa75998e · outbound
Reinforcing General Reasoning without Verifiers Measuring mathematical problem solving with the math dataset
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2bbc0d8-6d7c-4e9a-8559-ef1948d36c91 · outbound
Reinforcing General Reasoning without Verifiers Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b479d45d-7c4b-4ae8-800e-17e7f52fb058 · outbound
Reinforcing General Reasoning without Verifiers Self-Improvement in Language Models: The Sharpening Mechanism
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81cf8b30-cfc3-4b32-a192-89c7731e1297 · outbound
Reinforcing General Reasoning without Verifiers Gonzalez, Hao Zhang, and Ion Stoica
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d2439f9-9c29-4fb7-b7ca-c8cb2be5b697 · outbound
Reinforcing General Reasoning without Verifiers Tulu 3: Pushing Frontiers in Open Language Model Post-Training
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59922537-1b14-47ce-a2d4-346a1010e5b7 · outbound
Reinforcing General Reasoning without Verifiers Solving quantitative reasoning problems with language models.Advances in Neural Information Processing Systems, 35:3843–3857, 2022
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cadd1958-83e1-4260-a2fe-7ebce56bb0ee · outbound
Reinforcing General Reasoning without Verifiers Numinamath: The largest public dataset in ai4maths with 860k pairs of competition math problems and solutions.Hugging Face repository, 13:9, 2024
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb322e58-6b54-4677-94c9-26d5f8f67426 · outbound
Reinforcing General Reasoning without Verifiers Let’s verify step by step
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23d49f70-44e6-486d-8234-6972ef41ec39 · outbound
Reinforcing General Reasoning without Verifiers X-Reasoner: Towards Generalizable Reasoning Across Modalities and Domains
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f616da6-ee5f-48fa-a183-eb4a4e5ecd88 · outbound
Reinforcing General Reasoning without Verifiers Oat: A research-friendly framework for llm online alignment.https://github.com/sail-sg/oat, 2024
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 99224617-7326-4c1a-85f0-2dd010ea96bc · outbound
Reinforcing General Reasoning without Verifiers There may not be aha moment in r1-zero-like training — a pilot study
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8fd9a995-3dc6-479b-9102-e66be0b39631 · outbound
Reinforcing General Reasoning without Verifiers Understanding R1-Zero-Like Training: A Critical Perspective
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00842862-6da4-4cb5-947f-450fb3d521b8 · outbound
Reinforcing General Reasoning without Verifiers Deepcoder: A fully open-source 14b coder at o3-mini level, 2025
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f99e1cdc-3791-47c3-88d4-7fe27993c55d · outbound
Reinforcing General Reasoning without Verifiers Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Li Erran Li, Raluca Ada Popa, and Ion Stoica
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8604899-443b-4c3e-8189-b8cef03b172b · outbound
Reinforcing General Reasoning without Verifiers General-Reasoner: Advancing LLM Reasoning Across All Domains
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a64d249a-1c61-4104-b764-d954622e262e · outbound
Reinforcing General Reasoning without Verifiers Ng, Daishi Harada, and Stuart J
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4a747535-adc6-4e76-9340-bbe4f3af9954 · outbound
Reinforcing General Reasoning without Verifiers Learning to reason with llms, 2024
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation fedd8f05-30b9-4aac-8644-585576047986 · outbound
Reinforcing General Reasoning without Verifiers Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6de7d42f-7f10-4ba6-967c-3f7b9cc6ec85 · outbound
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8008712-a0b1-452c-a993-0fa78b48d8e1 · outbound
Reinforcing General Reasoning without Verifiers Training chain-of- thought via latent-variable inference.Advances in Neural Information Processing Systems, 36: 72819–72841, 2023
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 80afe6a7-a471-41af-9798-d8c60a771a76 · outbound
Reinforcing General Reasoning without Verifiers Direct preference optimization: Your language model is secretly a reward model
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1dd8e9e-70b0-4602-9844-d1cab760b508 · outbound
Reinforcing General Reasoning without Verifiers Learning to drive a bicycle using reinforcement learning and shaping
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation bff29ac9-4452-474e-a5a8-662cfbb18d31 · outbound
Reinforcing General Reasoning without Verifiers Gpqa: A graduate-level google-proof q&a benchmark
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca43c1f2-a018-4109-b4ff-3389791b2fca · outbound
Reinforcing General Reasoning without Verifiers Proximal Policy Optimization Algorithms
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7849d0c-5308-4572-aae7-00b1eb99e966 · outbound
Reinforcing General Reasoning without Verifiers DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation facf22d4-454e-4613-b858-c10da0a74038 · outbound
Reinforcing General Reasoning without Verifiers Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13680724-420a-4471-9655-96475f45143a · outbound
Reinforcing General Reasoning without Verifiers Sutton and Andrew G
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5e30f52-8d0b-4008-a8f2-fc41b25e47ad · outbound
Reinforcing General Reasoning without Verifiers Beyond Verifiable Rewards: Scaling Reinforcement Learning for Language Models to Unverifiable Data
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23fbb93f-8030-4c13-9649-781b8f2dca25 · outbound
Reinforcing General Reasoning without Verifiers Qwen3, April 2025
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a661f906-bd43-41e7-96cd-89f67894319d · outbound
Reinforcing General Reasoning without Verifiers Mmlu-pro: A more robust and challenging multi-task language understanding benchmark
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ae8c311-f20e-43d8-ad0d-9983c0c8b323 · outbound
Reinforcing General Reasoning without Verifiers Qwen2.5 Technical Report
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ab5a380-d554-4d25-a2ea-5ba8c30cf2dd · outbound
Reinforcing General Reasoning without Verifiers Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dfbbdad7-3e5b-4f76-a2f8-96a7f26b2e10 · outbound
Reinforcing General Reasoning without Verifiers DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b46341a-8a36-4346-a992-69c3bac38ab0 · outbound
Reinforcing General Reasoning without Verifiers Self-rewarding language models.International Conference on Machine Learning,
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ea9dad09-2bb7-4512-9ae8-730ae885cb81 · outbound
Reinforcing General Reasoning without Verifiers Naturalreasoning: Reasoning in the wild with 2.8 m challenging questions.arXiv preprint arXiv:2502.13124, 2025
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb9e4194-f947-421d-8802-662ec50072e4 · outbound
Reinforcing General Reasoning without Verifiers Scaling Relationship on Learning Mathematical Reasoning with Large Language Models
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9cb71e8a-757a-459b-a8f0-6bced49c0046 · outbound
Reinforcing General Reasoning without Verifiers Mammoth2: Scaling instructions from the web.Advances in Neural Information Processing Systems, 37:90629–90660, 2024
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd3fc11e-4ffb-4e99-b255-fabad4fbd632 · outbound
Reinforcing General Reasoning without Verifiers 7b model and 8k examples: Emerging reasoning with reinforcement learning is both effective and efficient.https://hkust-nlp.notion.site/simplerl-reason, 2025
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation db5b202b-4c9f-4cf3-8395-4e90c75576e0 · outbound
Reinforcing General Reasoning without Verifiers Fine-Tuning Language Models from Human Preferences
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec4bb7de-61fa-4edc-990c-b6107dad8896 · outbound
Reinforcing General Reasoning without Verifiers TTRL: Test-Time Reinforcement Learning
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation acce323f-2ddd-4261-a551-78492765c929 · outbound
Reinforcing General Reasoning without Verifiers Unresolved cited work
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e6763064-39f5-48b3-b01b-87fda9d95806 · outbound
Reinforcing General Reasoning without Verifiers This relation is usually given in a form where a graph or a formula relates period to absolute magnitude (a measure of intrinsic brightness)
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 120c5616-51b9-4211-9664-27fd8f35cbdc · outbound
Reinforcing General Reasoning without Verifiers Unresolved cited work
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0edea18e-4c38-4d9f-974d-6f07835b4b5f · outbound
Reinforcing General Reasoning without Verifiers ""everyone else is doing it
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ecf44972-dfd3-463b-97e4-8ff3ab1de69d · outbound
Reinforcing General Reasoning without Verifiers Their reasoning is fear-based, and they view rules as set by authority figures
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ef0107df-7832-4415-b647-576f0b7fd712 · outbound
Reinforcing General Reasoning without Verifiers what’s in it for me?
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0e3fc1d9-feba-4c02-b9d5-adcd9d901f15 · outbound
Reinforcing General Reasoning without Verifiers Unresolved cited work
Reference 1999
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2c3364b5-f502-45b4-9dda-64bae6f72e75 · outbound
Reinforcing General Reasoning without Verifiers Self-Rewarding Language Models
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5011bd04-bd69-40bb-b9b5-63b1463ceb1d · inbound
Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe Systems Reinforcing General Reasoning without Verifiers
Reference 167
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 881d37f7-a003-4bc1-a8de-289095639b6c · inbound
Reinforcement Pre-Training Reinforcing General Reasoning without Verifiers
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09f1d75b-3d1f-411a-81d3-dc06e696881d · inbound
Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks Reinforcing General Reasoning without Verifiers
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d9cb06d2-5012-4061-985a-434b37d473ca · inbound
URPO: A Unified Reward & Policy Optimization Framework for Large Language Models Reinforcing General Reasoning without Verifiers
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85b3a33d-241e-4e2f-ab1e-ef958c464377 · inbound
Reverse-Engineered Reasoning for Open-Ended Generation Reinforcing General Reasoning without Verifiers
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8b2fe70-730b-4ddc-8020-c762a6a2adf3 · inbound
Compute as Teacher: Turning Inference Compute Into Reference-Free Supervision Reinforcing General Reasoning without Verifiers
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation adc3cf64-8802-4479-ba29-f2d3a969ed65 · inbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Reinforcing General Reasoning without Verifiers
Reference 257
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57b2310b-f1db-4c8c-8d96-305438cda973 · inbound
Learning to Pose Problems: Reasoning-Driven and Solver-Adaptive Data Synthesis Reinforcing General Reasoning without Verifiers
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b607286b-d83d-43df-9474-bd80a6ef4580 · inbound
VI-CuRL: Stabilizing Verifier-Independent RL Reasoning via Confidence-Guided Variance Reduction Reinforcing General Reasoning without Verifiers
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 52f3cb1b-099f-4c53-9e14-867cb07c84b5 · inbound
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Reinforcing General Reasoning without Verifiers
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9c6b2e4e-3b78-4ece-bb46-3f61bff623b2 · inbound
Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning Reinforcing General Reasoning without Verifiers
Reference 187
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e01d6ce9-375f-44d7-b0c2-6248af681050 · inbound
G-Zero: Self-Play for Open-Ended Generation from Zero Data Reinforcing General Reasoning without Verifiers
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8727f605-db0e-467b-8e17-1e63a94ffd2b · inbound
Likelihood scoring for continuations of mathematical text: a self-supervised benchmark with tests for shortcut vulnerabilities Reinforcing General Reasoning without Verifiers
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a24852dd-a7ee-4c3c-9469-f35df85cafb7 · inbound
CARE-RL: Capability-Aware Reinforcement Learning for Mitigating Cross-Domain Conflicts Reinforcing General Reasoning without Verifiers
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d86298e8-0466-4040-9e5a-1d394911d33f · inbound
Trust Region On-Policy Distillation Reinforcing General Reasoning without Verifiers
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1c040be7-4644-4a71-ad7a-08c77ae88e6e · inbound
Predictive Divergence Masks for LLM RL Reinforcing General Reasoning without Verifiers
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.