Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T20:37:48.272553Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 1 inbound Pith citation observation for arXiv:2505.12457.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T20:37:48.272553Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-12T03:12:49.428954Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-12T03:16:19.329149Z
28 of 28 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation d95cc04d-e0aa-41e9-955b-192db53b1496 · outbound
UFO-RL: Uncertainty-Focused Optimization for Efficient Reinforcement Learning Data Selection Training Verifiers to Solve Math Word Problems
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5a7ad9e-50e4-4ca1-9b14-dcd9186fb452 · outbound
UFO-RL: Uncertainty-Focused Optimization for Efficient Reinforcement Learning Data Selection Open r1: A fully open reproduction of deepseek-r1, January 2025
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b6d4fa0-e60f-4c01-84bb-bba903eea854 · outbound
UFO-RL: Uncertainty-Focused Optimization for Efficient Reinforcement Learning Data Selection The Llama 3 Herd of Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02521e1c-2ff6-4669-8778-c8bc90c192cb · outbound
UFO-RL: Uncertainty-Focused Optimization for Efficient Reinforcement Learning Data Selection DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92924e94-15bd-458b-815c-a9de3c6b790b · outbound
UFO-RL: Uncertainty-Focused Optimization for Efficient Reinforcement Learning Data Selection Measuring massive multitask language understanding.Proceedings of the International Conference on Learning Representations (ICLR), 2021
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92b57636-be46-4f6d-9d6b-f498b408a9e1 · outbound
UFO-RL: Uncertainty-Focused Optimization for Efficient Reinforcement Learning Data Selection OpenAI o1 System Card
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f7742b3-60cd-4201-9607-2a7a28ce9821 · outbound
UFO-RL: Uncertainty-Focused Optimization for Efficient Reinforcement Learning Data Selection Mistral 7B
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a2f10d3-852e-4b8b-9e87-2f55cebfb27c · outbound
UFO-RL: Uncertainty-Focused Optimization for Efficient Reinforcement Learning Data Selection Efficient memory management for large language model serving with pagedattention
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6b93777-e0c5-4553-b9bc-13ac382145f1 · outbound
UFO-RL: Uncertainty-Focused Optimization for Efficient Reinforcement Learning Data Selection Reasoning Under 1 Billion: Memory-Augmented Reinforcement Learning for Large Language Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63cfa6f5-6848-4bea-8218-27a564635c80 · outbound
UFO-RL: Uncertainty-Focused Optimization for Efficient Reinforcement Learning Data Selection LIMR: Less is More for RL Scaling
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 927685ce-a55f-4d23-8b5a-65b4e8c71d1e · outbound
UFO-RL: Uncertainty-Focused Optimization for Efficient Reinforcement Learning Data Selection Let's Verify Step by Step
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a04e5741-1d71-4a98-ad63-2ea995e919f7 · outbound
UFO-RL: Uncertainty-Focused Optimization for Efficient Reinforcement Learning Data Selection WizardCoder: Empowering Code Large Language Models with Evol-Instruct
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f708a974-955d-4f41-acd2-61da25435205 · outbound
UFO-RL: Uncertainty-Focused Optimization for Efficient Reinforcement Learning Data Selection Tinyzero
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation ebf5aaa6-de94-492e-84a6-b8c3d422c460 · outbound
UFO-RL: Uncertainty-Focused Optimization for Efficient Reinforcement Learning Data Selection Proximal Policy Optimization Algorithms
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da097bbe-05df-4389-8c2a-4ae76b6d7e58 · outbound
UFO-RL: Uncertainty-Focused Optimization for Efficient Reinforcement Learning Data Selection Vygotsky’s zone of proximal development: Instructional implications and teachers’ professional development.English language teaching, 3(4):237–248, 2010
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 592f6d01-00d6-4e05-804e-8b1b5fa7809e · outbound
UFO-RL: Uncertainty-Focused Optimization for Efficient Reinforcement Learning Data Selection DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d12433b2-675b-4548-b109-4fd16cf88ee0 · outbound
UFO-RL: Uncertainty-Focused Optimization for Efficient Reinforcement Learning Data Selection Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90576817-ff7f-49ce-ad8d-3f8dd5d8a6dc · outbound
UFO-RL: Uncertainty-Focused Optimization for Efficient Reinforcement Learning Data Selection A Survey on Post-training of Large Language Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 987fb47c-4b71-49fe-a94a-b6cade2d47d0 · outbound
UFO-RL: Uncertainty-Focused Optimization for Efficient Reinforcement Learning Data Selection LESS: Selecting Influential Data for Targeted Instruction Tuning
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3262b2d-cab7-45bc-97a6-a12482c87147 · outbound
UFO-RL: Uncertainty-Focused Optimization for Efficient Reinforcement Learning Data Selection WizardLM: Empowering large pre-trained language models to follow complex instructions
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ebd5619-df64-4559-90ac-b53a8910d502 · outbound
UFO-RL: Uncertainty-Focused Optimization for Efficient Reinforcement Learning Data Selection Qwen2.5 Technical Report
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37ab6da3-c9ad-4bfc-b41f-65b3ac8190eb · outbound
UFO-RL: Uncertainty-Focused Optimization for Efficient Reinforcement Learning Data Selection LIMO: Less is More for Reasoning
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8afb3c14-f9c5-48f4-ad36-7d9664a3c2a8 · outbound
UFO-RL: Uncertainty-Focused Optimization for Efficient Reinforcement Learning Data Selection DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ee0bd85-2670-4060-b2d1-5cf54eab7afe · outbound
UFO-RL: Uncertainty-Focused Optimization for Efficient Reinforcement Learning Data Selection Supervised Fine-Tuning Achieve Rapid Task Adaption Via Alternating Attention Head Activation Patterns
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ff5a9fa-0fda-40b9-bc36-eb5cab39c37f · outbound
UFO-RL: Uncertainty-Focused Optimization for Efficient Reinforcement Learning Data Selection Deci- phering the impact of pretraining data on large language models through machine unlearning
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 9ee32d7c-cdd9-41a7-b324-4be260614525 · outbound
UFO-RL: Uncertainty-Focused Optimization for Efficient Reinforcement Learning Data Selection Beyond Similarity: A Gradient-based Graph Method for Instruction Tuning Data Selection
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 4f0bab86-5815-4175-9292-d5404ef1afc1 · outbound
UFO-RL: Uncertainty-Focused Optimization for Efficient Reinforcement Learning Data Selection Lima: Less is more for alignment.Advances in Neural Information Processing Systems, 36:55006–55021, 2023
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d2e64d9-4610-4995-b082-3ee594b24872 · outbound
UFO-RL: Uncertainty-Focused Optimization for Efficient Reinforcement Learning Data Selection Fine-Tuning Language Models from Human Preferences
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98825fa1-60da-4e9a-a865-5f7e05120346 · inbound
DARE: Difficulty-Adaptive Reinforcement Learning with Co-Evolved Difficulty Estimation UFO-RL: Uncertainty-Focused Optimization for Efficient Reinforcement Learning Data Selection
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.