Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:09:46.121588Z
Paper Citation Record · LEDGER
As of 23 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 12 inbound Pith citation observations for arXiv:2506.08745.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:09:46.121588Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T19:29:00.494940Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T13:59:51.882252Z
65 of 65 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation b6ffe7a6-8cd9-4dea-9d5d-dd5cad7aa535 · outbound
Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Kakade, Jason D
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 33bb0019-910d-4abe-974c-3fdef402e725 · outbound
Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Open r1: Evaluating llms on uncontaminated math competitions, 2025
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 349a80c7-031e-4bf9-b36e-80b6ca054ddd · outbound
Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Program Synthesis with Large Language Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a028dc47-6b3c-461a-b440-d01b6c095bd8 · outbound
Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Constitutional AI: Harmlessness from AI Feedback
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 126ac31b-a149-4f20-a242-22e69feca703 · outbound
Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Evaluating Large Language Models Trained on Code
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7a66c95-c231-4ac5-800a-9fb4f645469d · outbound
Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning An Empirical Study on Eliciting and Improving R1-like Reasoning Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a96f279b-e9e2-4748-8db7-cbbe8ca149b3 · outbound
Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Training Verifiers to Solve Math Word Problems
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ace5b57-f3d8-47b4-854e-7baf906a026f · outbound
Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning The Llama 3 Herd of Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de5575b2-157c-4405-a3db-3e040724983f · outbound
Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6f8b083-dc55-41ed-aaa8-cce6433666c1 · outbound
Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 740110bb-d645-412c-a1e5-969cf9887c27 · outbound
Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Olympiadbench: A challenging benchmark for promoting AGI with olympiad-level bilingual multimodal scientific problems
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 47db5fb1-f488-460d-a6a4-2b79b25aeae2 · outbound
Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Measuring mathematical problem solving with the MATH dataset
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 8f58d985-37a9-48c1-be99-122ef30061de · outbound
Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1a9ab83-840e-422d-820b-e131a5396b53 · outbound
Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f59237a1-bee6-407f-a99b-ea33077367e3 · outbound
Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68bf5204-e622-4186-a76a-b84f075529da · outbound
Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning OpenAI o1 System Card
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d0e0c34-3cd6-4d97-ab62-faadd798ef38 · outbound
Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Buy 4 REINFORCE samples, get a baseline for free! In Deep Reinforcement Learning Meets Structured Prediction, ICLR 2019 Workshop, 2019
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation ba7da1c7-439b-41fe-a38b-b24b3222e832 · outbound
Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Efficient memory management for large language model serving with pagedattention
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 1d58092d-d4fc-4dee-97c8-cca1226af5c6 · outbound
Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Tulu 3: Pushing Frontiers in Open Language Model Post-Training
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d523d22-44c7-4159-8605-808ce22af111 · outbound
Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Numinamath, 2024
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c71c6a20-c3f8-4038-9297-01d8d6730a0b · outbound
Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Let’s verify step by step
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation e1c935f6-3a28-48d5-aa9a-b50a8efa91e9 · outbound
Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning A Survey of Direct Preference Optimization
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59a01924-d375-4698-afa9-7e01cd44e84c · outbound
Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac2164ab-115b-4381-b36e-a55838303da1 · outbound
Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85670879-67ff-43f4-9d9c-f41cc542d951 · outbound
Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning On the global convergence rates of softmax policy gradient methods
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation a7d0a55c-b43f-4d58-98ad-299082650a33 · outbound
Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Unresolved cited work
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation dfad0528-64af-4b13-bad9-d394c6b8df70 · outbound
Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Iterative reasoning preference optimization
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 0d617e2b-74d7-46c1-8a74-dda527e8d774 · outbound
Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Are NLP models really able to solve simple math word problems? InNorth American Chapter of the Association for Computational Linguistics, 2021
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation d553303b-f8d8-4160-b70d-91155988e8dc · outbound
Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Manning, Stefano Ermon, and Chelsea Finn
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 0fab6a5d-2b04-45ea-97a3-48a1e71374dd · outbound
Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Semantic cosine similarity
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation ddc33db8-3652-4bdf-8cbd-be304684d6a0 · outbound
Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning What makes a reward model a good teacher? an optimization perspective.arXiv preprint arXiv:2503.15477, 2025
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b4faffc-9300-4686-b4b0-3249f98dd90d · outbound
Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Sentence-bert: Sentence embeddings using siamese bert-networks
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d3ca9b9-f63c-42e7-bb4b-d15eb909844e · outbound
Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Gpqa: A graduate-level google-proof q&a benchmark
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation af519356-a0cd-44dd-98d6-0ae87e455b75 · outbound
Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Proximal Policy Optimization Algorithms
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bca96c70-9f78-4100-bbdf-d6bd0ef41128 · outbound
Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning A mathematical theory of communication.The Bell system technical journal, 1948
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 6ce762ba-95d2-4297-88e7-3e04074a7a03 · outbound
Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89f5da95-4b79-4e1e-ba65-f83acba760b6 · outbound
Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 255ae497-a8d3-4e96-987b-23795bdf76c8 · outbound
Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning MIT press Cambridge, 1998
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5a48ae3-e6ce-46a4-9b33-513fc885edc4 · outbound
Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Commonsenseqa: A question answering challenge targeting commonsense knowledge
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation ed5c02dc-80b9-4fe9-9eb3-558a69b6038d · outbound
Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Le, Ed H
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation baf741f9-462f-4292-aa7b-819939d4589a · outbound
Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Latent Space Chain-of-Embedding Enables Output-free LLM Self-Evaluation
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 706440a4-c0a9-47c8-96af-1477ac5189b2 · outbound
Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Wong, Zhuosheng Zhang, and Rui Wang
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation c10d9300-976e-4ae8-8d9f-0ba9d28f4d26 · outbound
Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Mmlu-pro: A more robust and challenging multi-task language understanding benchmark
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 84ddd584-f7dc-4d90-beb6-3439d62ea7a8 · outbound
Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning CREAM: Consistency Regularized Self-Rewarding Language Models
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89ff06b2-70d6-4386-b387-4bacf9fbf833 · outbound
Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Chi, Quoc V
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 1fef196a-87d2-4433-a6a7-37af59368ef7 · outbound
Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b844e0e9-3376-412c-89aa-c3ac0f50d2bd · outbound
Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d8ec851-bcbe-4610-825e-f5be039ea645 · outbound
Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Self-rewarding correction for mathematical reasoning
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea577258-48cf-4c2f-b79d-f26916e0f3e0 · outbound
Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Qwen2.5 Technical Report
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56f7a275-2495-411c-a8cd-15faa612de94 · outbound
Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62e43cf3-1427-4cee-a808-e6800a55c9ed · outbound
Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Self-rewarding language models
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 85a54cc2-bdcb-4eea-9f1e-d18d4b028f3a · outbound
Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e11b935-4602-4260-818a-7337de2eb813 · outbound
Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d2ffa49-c387-402e-bb6f-05bb86a0dde2 · outbound
Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f99f6ab8-b206-410d-9992-fa579703938d · outbound
Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6ca236c-8010-4364-b52b-f308d2ab6a58 · outbound
Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Sample efficient reinforcement learning with reinforce
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 871a9734-5061-49f7-98a0-123e3038f3d5 · outbound
Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Reasoning with Reinforced Functional Token Tuning
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d96c8c71-2b16-4466-af1d-22da45d8d444 · outbound
Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30fd01d9-4a5a-4067-8c7f-6140ea62cb82 · outbound
Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Process-based Self-Rewarding Language Models
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97b1f075-2d3f-4a02-bc76-f2266f53f3a7 · outbound
Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f1fcc14-1af2-4ead-8c0c-fe95686c89db · outbound
Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Calibrated self-rewarding vision language models
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation b4962ff6-9a42-4688-8290-a53355ebb185 · outbound
Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Landscape of thoughts: Visualizing the reasoning process of large language models.arXiv preprint arXiv:2503.22165, 2025
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ba2c7f2-b2cc-4807-8c01-8a8bf8181b62 · outbound
Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Texygen: A benchmarking platform for text generation models
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation e93e34ce-8ff9-4559-8d0c-652a94ae0a90 · outbound
Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2d13d38-dc1c-46a8-8c78-dbf3c121cc03 · outbound
Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning TTRL: Test-Time Reinforcement Learning
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29319d30-5e71-4196-a3bf-7d2e642e1d67 · inbound
A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning
Reference 180
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 9cadc56b-68ff-4aa3-b874-4412ffc301f0 · inbound
Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ecaf77b-8d39-4576-b78f-eca6fc09c686 · inbound
ConsistRM: Improving Generative Reward Models via Consistency-Aware Self-Training Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 24083e58-3a7d-4795-ad66-57b4b709d172 · inbound
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 10de83ea-bbcd-421c-9cd4-af79fe64ff8e · inbound
Uncovering the Representation Geometry of Minimal Cores in Overcomplete Reasoning Traces Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 2b93e0d7-a109-4f6d-b3a5-c4f04b6f73ab · inbound
When Self-Belief Misleads: Active Label Acquisition for Reinforcement Learning with Verifiable Rewards Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 87aaad09-9d25-4276-aa0a-da8b8d618830 · inbound
Trust Region On-Policy Distillation Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation cd01867c-1484-4581-9ce0-6109352f8bd9 · inbound
Smart Picks in the Dark: Towards Efficient RLVR for Reasoning via Tracing Metacognitive Pivots Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 3c3e53ac-6859-4942-a940-366d111b5910 · inbound
GeoMin: Data-Efficient Semi-Supervised RLVR via Geometric Distribution Modeling Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation ed2c3dae-401c-4cf9-9c43-135827c3186a · inbound
Mental-R1: Aligning LLM Reasoning for Mental Health Assessment Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation ddbafa54-e33e-4dc9-9f48-ee9e5c8f1759 · inbound
Reinforcement Learning without Ground-Truth Solutions can Improve LLMs Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 0c0481c5-6659-4e70-a0b2-28879ea86c06 · inbound
H-OPD: Confidence Aware Heterogeneous Multi-Teacher Multimodal On-policy Distillation Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.