Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 31 inbound Pith citation observations for arXiv:2504.05812.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:22:13.915142Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
0
pith, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation f872eae1-12ef-47fa-8932-3b10e4dffb34 · inbound
Reinforcement Learning for Reasoning in Large Language Models with One Training Example Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ccb8f4d4-1967-4723-8ba4-f22c4e2ddf11 · inbound
The Unreasonable Effectiveness of Entropy Minimization in LLM Reasoning Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation faa37307-72ef-439b-9160-a7171164c542 · inbound
Reinforcing Video Reasoning with Focused Thinking Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d96c8c71-2b16-4466-af1d-22da45d8d444 · inbound
Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3270981c-bc09-460b-9d43-a36dafd58e5a · inbound
No Free Lunch: Rethinking Internal Feedback for LLM Reasoning Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4ab2f1a-28d1-45fc-9b71-6c201380e3f5 · inbound
Maximizing Prefix-Confidence at Test-Time Efficiently Improves Mathematical Reasoning Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 747bd84f-aaff-4c49-8042-15dc9311dd0b · inbound
Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 697a7a5f-2530-4662-8311-315ed08504e2 · inbound
Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 798cff28-f776-4c9a-a670-6146378a044d · inbound
Self-Evolving Vision-Language Models for Image Quality Assessment via Voting and Ranking Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5aab0f4-904b-48b3-b249-9fee88d46222 · inbound
Plan Then Action:High-Level Planning Guidance Reinforcement Learning for LLM Reasoning Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b5ae84b-f730-47fb-b671-73d40ce6f37d · inbound
Breaking the Self-Confirming Loop: Diagnosing and Mitigating Systemic Reward Bias in Self-Rewarding RL Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c101a8eb-45cf-44a7-a4e0-633897d475cb · inbound
CPMobius: Iterative Coach-Player Reasoning for Data-Free Reinforcement Learning Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e49f4d05-92ec-42b3-b369-bebae749d2dc · inbound
VI-CuRL: Stabilizing Verifier-Independent RL Reasoning via Confidence-Guided Variance Reduction Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e7444054-0638-4bca-97b1-15dfa8bd05c1 · inbound
SARL: Label-Free Reinforcement Learning by Rewarding Reasoning Topology Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 1620394a-5db7-476b-878b-594992143284 · inbound
Can LLMs Learn to Reason Robustly under Noisy Supervision? Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ace6045b-4a2b-45b5-83d6-fa07d474bbb3 · inbound
ZeroCoder: Can LLMs Improve Code Generation Without Ground-Truth Supervision? Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation fdaba3ed-b7b3-49d7-aaf4-79dd35dc93b1 · inbound
Eliciting Medical Reasoning with Knowledge-enhanced Data Synthesis: A Semi-Supervised Reinforcement Learning Approach Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b03f0231-88d5-41b8-972d-ca08f6ba4319 · inbound
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation bf6d7a23-a01b-4f6a-8ca1-5d8763337f9c · inbound
StraTA: Incentivizing Agentic Reinforcement Learning with Strategic Trajectory Abstraction Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation de5143a8-d07a-4421-9ad5-12e4fd889d43 · inbound
OracleTSC: Oracle-Informed Reward Hurdle and Uncertainty Regularization for Traffic Signal Control Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 012ab47c-5b0a-4a1f-82c7-1240e448662f · inbound
Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation bcb43b4b-7ee9-4305-bc98-03103ae8f655 · inbound
Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation aa455587-5ff9-4913-819e-cf313bf74a6f · inbound
PluRule: A Benchmark for Moderating Pluralistic Communities on Social Media Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8e5a1d7e-9f9f-4e73-951d-06be1ae44ca3 · inbound
Why Semantic Entropy Fails: Geometry-Aware and Calibrated Uncertainty for Policy Optimization Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 76cf23bc-fa70-4e46-ba30-31baa4c5f92e · inbound
When Self-Belief Misleads: Active Label Acquisition for Reinforcement Learning with Verifiable Rewards Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 24bdac98-83ff-46a8-911f-ab6accbe959b · inbound
On the Generalization Gap in Self-Evolving Language Model Reasoning Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9bf97f62-f105-4923-a36e-4f2194fd0b89 · inbound
Trust Region On-Policy Distillation Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e7df529a-4b27-42d3-acfc-bf84f707eb8d · inbound
Back on Track: Aligning Rewards and States for Reasoning in Diffusion Large Language Models Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e65cf1e1-7936-4b03-9069-6500ea32705a · inbound
Continual Test-Time Adaptation in Computer Vision: Methods, Benchmarks, and Future Directions Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation fa1b5eb4-f59a-4f18-8e82-07f57f55f4f6 · inbound
Continual Test-Time Adaptation in Computer Vision: Methods, Benchmarks, and Future Directions Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9cfbb828-d21e-44c6-8a5b-d32ece739d6f · inbound
On-Policy Self-Distillation without Any Supervision Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.