Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T23:35:27.149715Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 14 inbound Pith citation observations for arXiv:2506.17219.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T23:35:27.149715Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-05T11:39:47.889405Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T23:49:02.982506Z
44 of 44 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation dfd4a66e-a357-4b88-a12e-b6346ec93a34 · outbound
No Free Lunch: Rethinking Internal Feedback for LLM Reasoning Agarwal, S
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 926b5f14-6525-4f7f-9785-37760aea34a6 · outbound
No Free Lunch: Rethinking Internal Feedback for LLM Reasoning Agarwal, Z
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f8e4d7f2-e6dc-4b2a-be7c-528477446a5e · outbound
No Free Lunch: Rethinking Internal Feedback for LLM Reasoning Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e16d01a8-ccf6-4941-a2b3-ab8c6000985f · outbound
No Free Lunch: Rethinking Internal Feedback for LLM Reasoning Balunović, J
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cfbe4d4c-bfc8-488e-a31a-12c7f52c6c82 · outbound
No Free Lunch: Rethinking Internal Feedback for LLM Reasoning Unresolved cited work
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 50404a3b-569f-4601-8971-c42fbf504402 · outbound
No Free Lunch: Rethinking Internal Feedback for LLM Reasoning Training Verifiers to Solve Math Word Problems
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6614c350-9651-4508-8e8c-e2f073e6c23d · outbound
No Free Lunch: Rethinking Internal Feedback for LLM Reasoning The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c754fc7-8981-4231-aa49-bfa42e2afe9a · outbound
No Free Lunch: Rethinking Internal Feedback for LLM Reasoning One-shot Entropy Minimization
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75df10c3-956d-4229-9976-9a275710b2cc · outbound
No Free Lunch: Rethinking Internal Feedback for LLM Reasoning ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1b215df-0ed7-4d1c-aaa1-5094f7cc55bc · outbound
No Free Lunch: Rethinking Internal Feedback for LLM Reasoning Measuring Mathematical Problem Solving With the MATH Dataset
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a980ae08-ece3-4ac1-b7a1-efbb67bcd784 · outbound
No Free Lunch: Rethinking Internal Feedback for LLM Reasoning REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27c4927d-50d6-4850-b166-6cbde3e4869d · outbound
No Free Lunch: Rethinking Internal Feedback for LLM Reasoning AM-Thinking-v1: Advancing the Frontier of Reasoning at 32B Scale
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05893edd-2dce-4562-9d29-ab19d8dc1d05 · outbound
No Free Lunch: Rethinking Internal Feedback for LLM Reasoning Unresolved cited work
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a7b0ad1c-dfe7-4e4a-adae-ca65f8b91856 · outbound
No Free Lunch: Rethinking Internal Feedback for LLM Reasoning Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ca78c30-ebc9-4430-9860-f99091beae3d · outbound
No Free Lunch: Rethinking Internal Feedback for LLM Reasoning Lightman, V
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ed05b42c-bfbc-49c5-bc6e-4eebdd2cc916 · outbound
No Free Lunch: Rethinking Internal Feedback for LLM Reasoning DeepSeek-V3 Technical Report
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c94fc518-c1c3-4364-abb8-6dc70395d769 · outbound
No Free Lunch: Rethinking Internal Feedback for LLM Reasoning Self-Reflection Makes Large Language Models Safer, Less Biased, and Ideologically Neutral
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0253faf-9e29-44f2-85d1-a91013852723 · outbound
No Free Lunch: Rethinking Internal Feedback for LLM Reasoning Unresolved cited work
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30db9513-16ac-4cdd-8966-76cd4157e181 · outbound
No Free Lunch: Rethinking Internal Feedback for LLM Reasoning Learning to reason with llms
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4447a8c4-0263-42a2-a4c4-cadde5a67f21 · outbound
No Free Lunch: Rethinking Internal Feedback for LLM Reasoning Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19899ab7-647d-4c13-a644-a14f963266dd · outbound
No Free Lunch: Rethinking Internal Feedback for LLM Reasoning Proximal Policy Optimization Algorithms
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f267deac-333a-4d76-ab52-c96df614623d · outbound
No Free Lunch: Rethinking Internal Feedback for LLM Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60b13ce0-a760-45f6-8f41-30fb582936b4 · outbound
No Free Lunch: Rethinking Internal Feedback for LLM Reasoning HybridFlow: A Flexible and Efficient RLHF Framework
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7a3c7f6-e97c-4c6d-9cbd-0016472e1504 · outbound
No Free Lunch: Rethinking Internal Feedback for LLM Reasoning The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1598053d-27c7-4ab3-96c9-bca81f96e6fc · outbound
No Free Lunch: Rethinking Internal Feedback for LLM Reasoning Trust Region Preference Approximation: A simple and stable reinforcement learning algorithm for LLM reasoning
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b968bd5-19e6-4957-b0c1-aa930627ec87 · outbound
No Free Lunch: Rethinking Internal Feedback for LLM Reasoning Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b8000a2-3eef-4215-8443-6c85f6547315 · outbound
No Free Lunch: Rethinking Internal Feedback for LLM Reasoning Unresolved cited work
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6312f44-007a-44c5-b259-63c11071d68b · outbound
No Free Lunch: Rethinking Internal Feedback for LLM Reasoning LLaMA: Open and Efficient Foundation Language Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6aab5aa5-f816-4f7f-a513-c48bca8135aa · outbound
No Free Lunch: Rethinking Internal Feedback for LLM Reasoning Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 139c3792-1314-42b8-8e2e-79b50d00ab57 · outbound
No Free Lunch: Rethinking Internal Feedback for LLM Reasoning Unresolved cited work
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ff0c164-70cc-429d-9d13-dd5503f55605 · outbound
No Free Lunch: Rethinking Internal Feedback for LLM Reasoning Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b65b78f8-554e-445d-926c-e8492dce2f6b · outbound
No Free Lunch: Rethinking Internal Feedback for LLM Reasoning Qwen2.5 Technical Report
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb89ba1d-4170-464e-bf8f-329194e24067 · outbound
No Free Lunch: Rethinking Internal Feedback for LLM Reasoning Qwen3 Technical Report
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 757d5ee4-ed98-4d43-938c-8a2dcd9f9f35 · outbound
No Free Lunch: Rethinking Internal Feedback for LLM Reasoning Model Merging in LLMs, MLLMs, and Beyond: Methods, Theories, Applications and Opportunities
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbef77fb-62a0-4faf-b2c8-2b08f1c9c8dd · outbound
No Free Lunch: Rethinking Internal Feedback for LLM Reasoning DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f154043-c19f-4642-9927-04ee83a6481d · outbound
No Free Lunch: Rethinking Internal Feedback for LLM Reasoning Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97bae04e-4157-4166-85dd-bbe2f8c258f4 · outbound
No Free Lunch: Rethinking Internal Feedback for LLM Reasoning VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3270981c-bc09-460b-9d43-a36dafd58e5a · outbound
No Free Lunch: Rethinking Internal Feedback for LLM Reasoning Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58aa6f31-1d74-4a44-a9a7-cbe5f14acbcc · outbound
No Free Lunch: Rethinking Internal Feedback for LLM Reasoning Learning to Reason without External Rewards
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8243797-5ce4-49e5-bb40-31d0b449f2e0 · outbound
No Free Lunch: Rethinking Internal Feedback for LLM Reasoning Learning to Reason without External Rewards
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5b7a70d-2c48-4245-b798-8a6f865638f8 · outbound
No Free Lunch: Rethinking Internal Feedback for LLM Reasoning TTRL: Test-Time Reinforcement Learning
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1244017-d90e-4068-a597-e1f1b03ddb7c · outbound
No Free Lunch: Rethinking Internal Feedback for LLM Reasoning @esa (Ref
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e31cead1-e0c6-4f75-97de-6f1cd5759809 · outbound
No Free Lunch: Rethinking Internal Feedback for LLM Reasoning Unresolved cited work
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf265a47-e7a5-4bfc-b8ac-f31d8d4ded7d · outbound
No Free Lunch: Rethinking Internal Feedback for LLM Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c31aded5-24be-4437-bcd2-3240d11e67b6 · inbound
Towards Agents That Know When They Don't Know: Uncertainty as a Control Signal for Structured Reasoning No Free Lunch: Rethinking Internal Feedback for LLM Reasoning
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a355802a-b5ca-4911-8a28-4adf17f38350 · inbound
Self-Aligned Reward: Towards Effective and Efficient Reasoners No Free Lunch: Rethinking Internal Feedback for LLM Reasoning
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7dbf4b0e-705b-45b0-8083-43c8ee33bce4 · inbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle No Free Lunch: Rethinking Internal Feedback for LLM Reasoning
Reference 245
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40c57f60-d892-462b-a512-4717ad7896b5 · inbound
Breaking the Self-Confirming Loop: Diagnosing and Mitigating Systemic Reward Bias in Self-Rewarding RL No Free Lunch: Rethinking Internal Feedback for LLM Reasoning
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2848b1e9-59ce-4b6b-b311-091eed1c6631 · inbound
VI-CuRL: Stabilizing Verifier-Independent RL Reasoning via Confidence-Guided Variance Reduction No Free Lunch: Rethinking Internal Feedback for LLM Reasoning
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0fb45e22-e0be-4aff-bd9a-39337a2a47f8 · inbound
CoAct: Co-Active LLM Preference Learning with Human-AI Synergy No Free Lunch: Rethinking Internal Feedback for LLM Reasoning
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 52c9872f-c47e-458d-8c82-67b508e515b3 · inbound
Free Energy-Driven Reinforcement Learning with Adaptive Advantage Shaping for Unsupervised Reasoning in LLMs No Free Lunch: Rethinking Internal Feedback for LLM Reasoning
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9672e61c-fb56-47ed-8f38-3c53997e1484 · inbound
Experience Sharing in Mutual Reinforcement Learning for Heterogeneous Language Models No Free Lunch: Rethinking Internal Feedback for LLM Reasoning
Reference 114
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation eb793416-eb9a-43c6-a989-8f187363629e · inbound
D$^2$Evo: Dual Difficulty-Aware Self-Evolution for Data-Efficient Reinforcement Learning No Free Lunch: Rethinking Internal Feedback for LLM Reasoning
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3b80e509-f985-4f97-8f1d-80f17f5e2a4d · inbound
Detecting and Mitigating the Correct-Answer Extinction Window in Test-Time Reinforcement Learning with Majority Voting No Free Lunch: Rethinking Internal Feedback for LLM Reasoning
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a47f0fb5-0084-4b81-8874-201806c4e10a · inbound
Trust Region On-Policy Distillation No Free Lunch: Rethinking Internal Feedback for LLM Reasoning
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 48d91114-c729-4d3c-8c56-81760641e3dd · inbound
GeoMin: Data-Efficient Semi-Supervised RLVR via Geometric Distribution Modeling No Free Lunch: Rethinking Internal Feedback for LLM Reasoning
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bcbf05a3-fd56-4b0d-8577-c58130cfd916 · inbound
Continual Self-Improvement with Lightweight Experiential Latent Memories No Free Lunch: Rethinking Internal Feedback for LLM Reasoning
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 16f175c8-4cf7-42f1-9910-af1ac9887390 · inbound
Be Your Own Teacher: Steering Protein Language Models via Unsupervised Reward Optimization No Free Lunch: Rethinking Internal Feedback for LLM Reasoning
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.