Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T00:11:14.891530Z
Paper Citation Record · LEDGER
As of 15 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 16 inbound Pith citation observations for arXiv:2412.01951.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T00:11:14.891530Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-11T20:53:37.179431Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T14:08:21.732193Z
68 of 68 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 1fb15654-2f89-4084-b60d-78fbafbc5140 · outbound
Self-Improvement in Language Models: The Sharpening Mechanism Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14864a9c-3e66-4e79-9a7e-946e425a786e · outbound
Self-Improvement in Language Models: The Sharpening Mechanism Unresolved cited work
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 27f039b6-6ea1-4fda-9230-a402e54bcd7f · outbound
Self-Improvement in Language Models: The Sharpening Mechanism Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f531f67-c03e-433f-bd4b-19dcb247b78a · outbound
Self-Improvement in Language Models: The Sharpening Mechanism J.1.2 Proof of Theorem 4.2′ Proof of Theorem 4.2′
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation f9750400-8623-4dcb-8275-b44a276f1434 · outbound
Self-Improvement in Language Models: The Sharpening Mechanism Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1a19f5b-306a-4106-87a7-90cc46f728b3 · outbound
Self-Improvement in Language Models: The Sharpening Mechanism BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de0cfaa2-1a94-4681-9f56-ce9a6e749911 · outbound
Self-Improvement in Language Models: The Sharpening Mechanism Foundations of Reinforcement Learning and Interactive Decision Making
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d41769f-0976-44c9-8b0e-c04e47963bfe · outbound
Self-Improvement in Language Models: The Sharpening Mechanism The Statistical Complexity of Interactive Decision Making
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 334e2da3-8166-41b8-9891-3350675f0b0e · outbound
Self-Improvement in Language Models: The Sharpening Mechanism REBEL: Reinforcement Learning via Regressing Relative Rewards
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d599889-188e-48a1-b903-6776752e606f · outbound
Self-Improvement in Language Models: The Sharpening Mechanism Measuring Massive Multitask Language Understanding
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ecca62f2-d97e-4dc6-a1f0-813308f74f10 · outbound
Self-Improvement in Language Models: The Sharpening Mechanism Measuring Mathematical Problem Solving With the MATH Dataset
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f99de5e0-db70-491d-a30c-3b37a29e4a47 · outbound
Self-Improvement in Language Models: The Sharpening Mechanism Amortizing intractable inference in large language models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f835749-1798-421a-b54d-e5b277411b0a · outbound
Self-Improvement in Language Models: The Sharpening Mechanism Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee9f7080-a4e1-43f1-8178-68a2f29df2df · outbound
Self-Improvement in Language Models: The Sharpening Mechanism Large Language Models Can Self-Improve
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a225960-04a1-4a92-8e2f-a658629ba70f · outbound
Self-Improvement in Language Models: The Sharpening Mechanism Mistral 7B
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ecb5d6a-a9cb-44b6-a322-befa3abcb542 · outbound
Self-Improvement in Language Models: The Sharpening Mechanism Provably Mitigating Overoptimization in RLHF: Your SFT Loss is Implicitly an Adversarial Regularizer
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7cea87d-19b0-4b6c-8247-e08c0aace764 · outbound
Self-Improvement in Language Models: The Sharpening Mechanism Auto-Regressive Next-Token Predictors are Universal Learners
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9d616e4-1187-4a71-9e08-6a2ced45bf4b · outbound
Self-Improvement in Language Models: The Sharpening Mechanism If beam search is the answer, what was the question?
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de860048-80a8-4c98-928b-e37c2eb7d89a · outbound
Self-Improvement in Language Models: The Sharpening Mechanism Controlled Decoding from Language Models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a8d4734-9aa9-4cac-9a4e-e10413807cee · outbound
Self-Improvement in Language Models: The Sharpening Mechanism West-of-N: Synthetic Preferences for Self-Improving Reward Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0a80c5d-efc9-4ee1-ac42-e7fcbf1dbff1 · outbound
Self-Improvement in Language Models: The Sharpening Mechanism Language Model Self-improvement by Reinforcement Learning Contemplation
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4bfc5069-9fde-48bd-a408-74f5244ffb56 · outbound
Self-Improvement in Language Models: The Sharpening Mechanism Understanding the Gains from Repeated Self-Distillation
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87f34528-7e9e-47f8-b25f-7d3625855b21 · outbound
Self-Improvement in Language Models: The Sharpening Mechanism The Entropy Enigma: Success and Failure of Entropy Minimization
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c566fc9-7d28-41df-b585-eb9d12e9ea1b · outbound
Self-Improvement in Language Models: The Sharpening Mechanism Recursive Introspection: Teaching Language Model Agents How to Self-Improve
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5580a3f5-f4cf-46f6-872d-ae9f8d2da078 · outbound
Self-Improvement in Language Models: The Sharpening Mechanism BOND: Aligning LLMs with Best-of-N Distillation
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6402f5d-daae-4bb7-bf04-dab266799e55 · outbound
Self-Improvement in Language Models: The Sharpening Mechanism Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2b150d1-58cd-4487-aa4b-176f7418581f · outbound
Self-Improvement in Language Models: The Sharpening Mechanism The Importance of Online Data: Understanding Preference Fine-tuning via Coverage
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2cece19f-bf61-46fd-b1b8-ff72e0ce75c4 · outbound
Self-Improvement in Language Models: The Sharpening Mechanism Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e213e65d-7b99-4c4f-99fd-08adad6f00cb · outbound
Self-Improvement in Language Models: The Sharpening Mechanism Tent: Fully Test-time Adaptation by Entropy Minimization
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ecd20b6-2550-40e1-b732-f15ab39e5c62 · outbound
Self-Improvement in Language Models: The Sharpening Mechanism Self-Taught Evaluators
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23e2b478-f42c-4ddb-a92f-465bd58563a5 · outbound
Self-Improvement in Language Models: The Sharpening Mechanism Chain-of-Thought Reasoning Without Prompting
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f244995-23bd-4bb2-ab14-0cd4770b9397 · outbound
Self-Improvement in Language Models: The Sharpening Mechanism Self-Instruct: Aligning Language Models with Self-Generated Instructions
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33a550c8-070a-4b15-bf0b-f0ed9f0d6ce8 · outbound
Self-Improvement in Language Models: The Sharpening Mechanism Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea76dee6-6ef3-4b61-bfe5-e21682a2bc57 · outbound
Self-Improvement in Language Models: The Sharpening Mechanism Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 182c56cc-2c66-43db-b674-7b15889833d9 · outbound
Self-Improvement in Language Models: The Sharpening Mechanism Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5212328-04eb-4a7d-ac0d-ed1f8725245c · outbound
Self-Improvement in Language Models: The Sharpening Mechanism Asymptotics of Language Model Alignment
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e98b4f2e-9bb9-48cf-aa18-cc9c55bd8095 · outbound
Self-Improvement in Language Models: The Sharpening Mechanism Online Iterative Reinforcement Learning from Human Feedback with General Preference Model
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1fe62b64-63cb-4a12-91b6-bf7765fe99d4 · outbound
Self-Improvement in Language Models: The Sharpening Mechanism Self-Rewarding Language Models
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 474be5c1-9b8f-4b67-8682-abe60503e6c7 · outbound
Self-Improvement in Language Models: The Sharpening Mechanism Unresolved cited work
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation b0004a9f-e054-467e-9568-9e911add078d · outbound
Self-Improvement in Language Models: The Sharpening Mechanism LLM-as-a-Judge
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation bcae0b3a-9da3-4dd5-b997-43a78a11b2a8 · outbound
Self-Improvement in Language Models: The Sharpening Mechanism Unresolved cited work
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 485d398b-78f6-4fdd-9fab-b7a7f7f79349 · outbound
Self-Improvement in Language Models: The Sharpening Mechanism More sophisticated inference-time search strategies such tree search and MCTS (Yao et al., 2024; Wan et al., 2024; Mudgal et al., 2023; Zhao et al.,
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation d15a36d1-b38a-4be8-a229-e1d28059e7e3 · outbound
Self-Improvement in Language Models: The Sharpening Mechanism Perhaps most closely related to our work is Frei et al
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation cd1fac52-502e-41e7-9dad-ac4b99c7bc6d · outbound
Self-Improvement in Language Models: The Sharpening Mechanism Proof of Proposition C.1
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 7cdb21b2-5c02-465a-a7d3-4441a2eb07e4 · outbound
Self-Improvement in Language Models: The Sharpening Mechanism We quantify the quality of a sharpened model as follows
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 43c7ee13-b33b-4892-a93f-3f6453d3e602 · outbound
Self-Improvement in Language Models: The Sharpening Mechanism Now a⊤Γ−1 t−1a ≤ 1/λ with probability 1, where λ = λmin(Γ0)
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 02465288-f85a-49e2-b606-2e43efbb3d91 · outbound
Self-Improvement in Language Models: The Sharpening Mechanism Note that unlike the non-adaptive framework, the distribution overmi depends on the underlying instance I with which the algorithm interacts
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 4f13ea1f-2315-4d54-bf2b-dd97e64c56da · outbound
Self-Improvement in Language Models: The Sharpening Mechanism Unresolved cited work
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation f05ee8cb-53e6-40b7-b2b1-2996a591e203 · outbound
Self-Improvement in Language Models: The Sharpening Mechanism Initialize: π(1) ← πbase, D(0) ← ∅
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation b7cd25d5-9199-46fb-b847-93b6ad41f2f4 · outbound
Self-Improvement in Language Models: The Sharpening Mechanism Also suppose thatπ⋆ β ∈ Π where π⋆ β(y | x) ∝ π1+β−1 base (y | x)
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation b57a37d9-567b-42a2-b260-324885c83b69 · outbound
Self-Improvement in Language Models: The Sharpening Mechanism Chain of Thought Empowers Transformers to Solve Inherently Serial Problems
Reference 1973
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 100ba0b1-6cce-431d-ad6d-aa9ddcc9f180 · outbound
Self-Improvement in Language Models: The Sharpening Mechanism GPT-4 Technical Report
Reference 1983
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fbbf4766-f474-43ae-a5f0-3e499cdf04bf · outbound
Self-Improvement in Language Models: The Sharpening Mechanism Towards a theory of model distillation
Reference 1994
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1fa65628-c83b-4a35-a0b8-504f9e4fc9bb · outbound
Self-Improvement in Language Models: The Sharpening Mechanism Butterfly Effects of SGD Noise: Error Amplification in Behavior Cloning and Autoregression
Reference 1995
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f23a1978-c153-4140-af44-9127a7b0e72f · outbound
Self-Improvement in Language Models: The Sharpening Mechanism BoNBoN Alignment for Large Language Models and the Sweetness of Best-of-n Sampling
Reference 2004
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5642192-9f09-4e58-9c3c-65b637a8bb1f · outbound
Self-Improvement in Language Models: The Sharpening Mechanism Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models
Reference 2006
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9870b06-ed88-4bdf-8ac6-0216350d89f4 · outbound
Self-Improvement in Language Models: The Sharpening Mechanism In Defense of Pseudo-Labeling: An Uncertainty-Aware Pseudo-label Selection Framework for Semi-Supervised Learning
Reference 2011
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 645a2964-83fb-4947-b139-d1a25cb4b978 · outbound
Self-Improvement in Language Models: The Sharpening Mechanism PaLM 2 Technical Report
Reference 2014
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0308ba6c-f907-459d-a12b-5da8b4789a63 · outbound
Self-Improvement in Language Models: The Sharpening Mechanism LoRA: Low-Rank Adaptation of Large Language Models
Reference 2015
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76e57ec9-04a5-4467-86ee-0c576f8db4f9 · outbound
Self-Improvement in Language Models: The Sharpening Mechanism Proximal Policy Optimization Algorithms
Reference 2016
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a668f288-be40-42d7-9f6d-5f541cebedfa · outbound
Self-Improvement in Language Models: The Sharpening Mechanism Training Verifiers to Solve Math Word Problems
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f21e54c-60e5-48c2-8b3b-6e81dd482993 · outbound
Self-Improvement in Language Models: The Sharpening Mechanism Distillation $\approx$ Early Stopping? Harvesting Dark Knowledge Utilizing Anisotropic Information Retrieval For Overparameterized Neural Network
Reference 2018
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b84f714c-230d-40f4-acba-d4382f15c153 · outbound
Self-Improvement in Language Models: The Sharpening Mechanism The Llama 3 Herd of Models
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7948727-c316-488a-b7bc-9d176cc9fae7 · outbound
Self-Improvement in Language Models: The Sharpening Mechanism Variational Best-of-N Alignment
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9489302e-5150-489d-8ff0-5969331fd422 · outbound
Self-Improvement in Language Models: The Sharpening Mechanism Distilling the Knowledge in a Neural Network
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b33b96a4-3cf4-41af-84b2-d8c087597fa8 · outbound
Self-Improvement in Language Models: The Sharpening Mechanism Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8bbafa47-6fa0-4958-8e87-dfad73cd2737 · outbound
Self-Improvement in Language Models: The Sharpening Mechanism Retraining with Predicted Hard Labels Provably Increases Model Accuracy
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58aa67ff-9efd-4d3b-beea-8186a5fb3bec · outbound
Self-Improvement in Language Models: The Sharpening Mechanism Transferring Inductive Biases through Knowledge Distillation
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5053f7f-ac80-44c3-aa4a-e6ed349a16f9 · inbound
LIAR: Leveraging Inference Time Alignment (Best-of-N) to Jailbreak LLMs in Seconds Self-Improvement in Language Models: The Sharpening Mechanism
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 715ea1a9-cf4b-474f-b290-e9e02e7b081d · inbound
Self-Improving Transformers Overcome Easy-to-Hard and Length Generalization Challenges Self-Improvement in Language Models: The Sharpening Mechanism
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e68c7525-f7e4-48db-906e-ced8dded2667 · inbound
The Unreasonable Effectiveness of Entropy Minimization in LLM Reasoning Self-Improvement in Language Models: The Sharpening Mechanism
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation b479d45d-7c4b-4ae8-800e-17e7f52fb058 · inbound
Reinforcing General Reasoning without Verifiers Self-Improvement in Language Models: The Sharpening Mechanism
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 179cc3d2-eb8c-4a8f-a54d-9aa6771e01ee · inbound
Sample Complexity and Representation Ability of Test-time Scaling Paradigms Self-Improvement in Language Models: The Sharpening Mechanism
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2244849d-c21a-4ff2-82fb-0ec8ec7e670a · inbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Self-Improvement in Language Models: The Sharpening Mechanism
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c99e57b6-dc00-4d09-b4a1-c0c4d0d6ae38 · inbound
Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future Self-Improvement in Language Models: The Sharpening Mechanism
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c83422e7-447b-4537-a2bf-d80056733181 · inbound
Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation Self-Improvement in Language Models: The Sharpening Mechanism
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9e3b4c8-2b19-4b7e-8c88-eaf2aad9f39a · inbound
A Task-Centric Theory for Iterative Self-Improvement with Easy-to-Hard Curricula Self-Improvement in Language Models: The Sharpening Mechanism
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8022d91-06f9-4e48-80f0-26072928b072 · inbound
From Prior to Pro: Efficient Skill Mastery via Distribution Contractive RL Finetuning Self-Improvement in Language Models: The Sharpening Mechanism
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d7ed3f2-720a-485e-9dfe-8f0a24d7d7ac · inbound
The Role of Generator Access in Autoregressive Post-Training Self-Improvement in Language Models: The Sharpening Mechanism
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 0e98ced7-f560-41da-bdbd-eb0e9d7fb3c8 · inbound
Beyond Distribution Sharpening: The Importance of Task Rewards Self-Improvement in Language Models: The Sharpening Mechanism
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 73f72c26-913e-4e00-9097-f07e002fca28 · inbound
On the Generalization Gap in Self-Evolving Language Model Reasoning Self-Improvement in Language Models: The Sharpening Mechanism
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 51af3dc1-1f41-4b67-8962-6fdb0545fc90 · inbound
Trust Region On-Policy Distillation Self-Improvement in Language Models: The Sharpening Mechanism
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 70c40a63-4681-4129-b016-7a5f5524e943 · inbound
Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning Self-Improvement in Language Models: The Sharpening Mechanism
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation fc408979-3023-4995-b367-ff8699d39f6a · inbound
Select and Improve: Understanding the Mechanics of Post-Training for Reasoning Self-Improvement in Language Models: The Sharpening Mechanism
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.