Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:00:04.018651Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 15 inbound Pith citation observations for arXiv:2505.20282.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:00:04.018651Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T23:35:23.466690Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T09:09:43.563954Z
28 of 28 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c684cd86-0491-4eb8-95fb-bb4a6ab2d7f6 · outbound
One-shot Entropy Minimization The unreasonable effectiveness of entropy minimization in llm reasoning, 2025
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4a3f193-0cf9-4bdc-ba4e-988585ab6de8 · outbound
One-shot Entropy Minimization Program synthesis with large language models, 2021
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a95a396f-dea4-4c46-8a22-0230835b0dfd · outbound
One-shot Entropy Minimization Michaud, Jacob Pfau, Dmitrii Krasheninnikov, Xin Chen, Lauro Langosco, Peter Hase, Erdem Bıyık, Anca Dragan, David Krueger, Dorsa Sadigh, and Dylan Hadfield-Menell
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3a07bc81-2a6e-4ab0-a5c0-8198cdafe87c · outbound
One-shot Entropy Minimization Process reinforcement through implicit rewards, 2025
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42dda982-470f-4c9b-bf9d-4992cc052fbe · outbound
One-shot Entropy Minimization Unresolved cited work
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b8dc0b44-98dc-4fbe-9f3a-6c8f2bb0741b · outbound
One-shot Entropy Minimization Interpretable contrastive monte carlo tree search reasoning, 2024
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e67606fe-d368-43d6-818f-14ece745f27c · outbound
One-shot Entropy Minimization Mixed preference optimization: Reinforcement learning with data selection and better reference model, 2025
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6fec6666-3639-4126-b301-2700ec07e3c8 · outbound
One-shot Entropy Minimization Accelerate: Training and inference at scale made simple, efficient and adaptable
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9766c71c-6604-4d11-b589-0736c0a59b0d · outbound
One-shot Entropy Minimization Olympiadbench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems, 2024
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b922226-020b-40e3-9e93-c9e6b7e53bb6 · outbound
One-shot Entropy Minimization Reinforce++: A simple and efficient approach for aligning large language models, 2025
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b765e25c-b272-4d62-9c67-623b504bcdb9 · outbound
One-shot Entropy Minimization Open-reasoner-zero: An open source approach to scaling up reinforcement learning on the base model, 2025
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b39f15ba-6895-4345-9b8d-2d6431c914e6 · outbound
One-shot Entropy Minimization Solving quantitative reasoning problems with language models, 2022
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f8531b7-91ad-4f83-b3c8-f6b0ee3fd4d2 · outbound
One-shot Entropy Minimization Numinamath
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2137447-de90-4943-8098-d810c46ec7d6 · outbound
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63073ceb-54a4-4927-a2a8-6e3ba697d05e · outbound
One-shot Entropy Minimization Understanding r1-zero-like training: A critical perspective, 2025
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0f32483-2e78-480e-a2f4-719cada52967 · outbound
One-shot Entropy Minimization Introducing openai o1
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b58c0e40-63af-46ee-b959-f8d146d415fd · outbound
One-shot Entropy Minimization Introducing openai o3 and o4-mini, April 2025
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d6b42aab-48a4-4a57-be82-b52df24a7b10 · outbound
One-shot Entropy Minimization Direct preference optimization: Your language model is secretly a reward model
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99dc841e-745d-43eb-acbc-cfe6b83a565e · outbound
One-shot Entropy Minimization Lee, and Sanjeev Arora
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a49c8fb-eee5-461b-8ace-0d3b6a0d57ff · outbound
One-shot Entropy Minimization Proximal policy optimization algorithms, 2017
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 380c4873-a384-4d53-9d23-33dd185fb90b · outbound
One-shot Entropy Minimization Unresolved cited work
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50cb3015-10f7-4dde-984a-962aa8ba4b5a · outbound
One-shot Entropy Minimization Unresolved cited work
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5cbfd56f-65c5-4494-90dc-7a97648f838f · outbound
One-shot Entropy Minimization Reinforcement learning for reasoning in large language models with one training example, 2025
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27853730-a5ce-4d70-8821-b5b49ea5b3d4 · outbound
One-shot Entropy Minimization On memorization of large language models in logical reasoning
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3956a355-01cf-4690-a333-cc63d9648b98 · outbound
One-shot Entropy Minimization Logic-rl: Unleashing llm reasoning with rule-based reinforcement learning, 2025
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07c8ef16-87cf-44af-9b98-0391ea31b4c0 · outbound
One-shot Entropy Minimization Towards large reasoning models: A survey of reinforced reasoning with large language models, 2025
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6dc3434-764f-4e1c-a924-052e782cbb49 · outbound
One-shot Entropy Minimization Redstar: Does scaling long-cot data unlock better slow-reasoning systems?, 2025
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6f37e15-a6d9-4694-9d98-8a9b0357e692 · outbound
One-shot Entropy Minimization Simplerl-zoo: Investigating and taming zero reinforcement learning for open base models in the wild, 2025
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4c754fc7-8981-4231-aa49-bfa42e2afe9a · inbound
No Free Lunch: Rethinking Internal Feedback for LLM Reasoning One-shot Entropy Minimization
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 881e88db-a251-40ce-9704-853b60289a18 · inbound
THIRDEYE: Cue-Aware Monocular Depth Estimation via Brain-Inspired Multi-Stage Fusion One-shot Entropy Minimization
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43d43ab2-3aa5-49da-8317-3a3b660ecc44 · inbound
EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity One-shot Entropy Minimization
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72352bc2-0201-46b3-9b37-71dd8e686483 · inbound
Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning One-shot Entropy Minimization
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 596ef3ac-f757-494c-88ad-d727917f3bec · inbound
Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents One-shot Entropy Minimization
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7bad3a1-dcca-4979-b003-a142e7a21adb · inbound
Compute as Teacher: Turning Inference Compute Into Reference-Free Supervision One-shot Entropy Minimization
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a46c5b96-f8c7-43b3-9750-91915ed0bcb5 · inbound
Relationship-Centered Care: Relatedness and Responsible Design for Human Connections in Mental-Health Care One-shot Entropy Minimization
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc6f50a1-32fc-443c-b4c0-78c4f578d803 · inbound
Token-Level Policy Optimization: Linking Group-Level Rewards to Token-Level Aggregation via Sequence-Level Likelihood One-shot Entropy Minimization
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 17991fc5-8074-4b24-a417-d2f1b8ae79ab · inbound
HEALing Entropy Collapse: Enhancing Exploration in Few-Shot RLVR via Hybrid-Domain Entropy Dynamics Alignment One-shot Entropy Minimization
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 57adcbe9-771a-4d1f-8176-6e58f1eeb267 · inbound
SELF-EMO: Emotional Self-Evolution from Recognition to Consistent Expression One-shot Entropy Minimization
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 491ffeb3-5399-4d0a-bdba-ec7b8e9fc6ed · inbound
TEMPO: Scaling Test-time Training for Large Reasoning Models One-shot Entropy Minimization
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9f6b2ece-7564-4e72-b83b-32db22f1f6b5 · inbound
Experience Sharing in Mutual Reinforcement Learning for Heterogeneous Language Models One-shot Entropy Minimization
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 51ef5662-03e5-464b-941a-6202a389d6f2 · inbound
Rethinking Entropy Minimization in Test-Time Adaptation for Autoregressive Models One-shot Entropy Minimization
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 435b3f76-3273-44d9-9e93-76bbb75d8a1d · inbound
Trust Region On-Policy Distillation One-shot Entropy Minimization
Reference 240
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2b142ae2-05d1-4a9e-bfb1-53ef1010a7c4 · inbound
What are Key Factors for Updates in RL for LLM Reasoning? One-shot Entropy Minimization
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.