Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-23T02:41:21.571824Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 89 of 89 outbound references and 1 inbound Pith citation observation for arXiv:2502.12272.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-23T02:41:21.571824Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-03T04:17:27.154486Z
A source-named dated measurement, never combined with another source.
Source: cited_works
89 of 89 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 4f7f7a4c-d5f1-4184-8a17-8f69d175f469 · outbound
Learning to Reason at the Frontier of Learnability DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c2cbddbf-206a-4b79-9e27-6848eb81e78e · outbound
Learning to Reason at the Frontier of Learnability Tulu 3: Pushing Frontiers in Open Language Model Post-Training
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation fdaf63a6-feec-45ff-9f96-01aa0bad26f3 · outbound
Learning to Reason at the Frontier of Learnability Learning to reason with llms
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b290130e-c07f-4672-9a7d-35b03ff5ed0b · outbound
Learning to Reason at the Frontier of Learnability Understanding R1-Zero-Like Training: A Critical Perspective
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7046d60f-b54a-4c5a-a42b-cca48dcecaad · outbound
Learning to Reason at the Frontier of Learnability Vineppo: Unlocking rl potential for llm reasoning through refined credit assignment
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 66fa54c1-421c-4938-8538-0c19b137cf03 · outbound
Learning to Reason at the Frontier of Learnability VinePPO: Refining Credit Assignment in RL Training of LLMs
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 79f22151-c393-4174-970b-f15eaa667fd7 · outbound
Learning to Reason at the Frontier of Learnability Group Robust Preference Optimization in Reward-free RLHF
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b2f396c2-77ad-4b05-8ab4-ceb0fcc0b99c · outbound
Learning to Reason at the Frontier of Learnability Proximal Policy Optimization Algorithms
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 137dd64d-fe7f-46dd-bbdb-6ecd5ba5c84e · outbound
Learning to Reason at the Frontier of Learnability Training Verifiers to Solve Math Word Problems
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 42f79281-0270-42b9-9207-899617751327 · outbound
Learning to Reason at the Frontier of Learnability Measuring Mathematical Problem Solving With the MATH Dataset
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0c878d87-271b-49f9-a748-107161b0410d · outbound
Learning to Reason at the Frontier of Learnability Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 95ee720f-f476-4f9e-b569-0f3e85e7c30c · outbound
Learning to Reason at the Frontier of Learnability Rho-1: Not All Tokens Are What You Need
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f955939b-2b2a-4736-afb8-5f09842eb4d9 · outbound
Learning to Reason at the Frontier of Learnability Qwen2.5 technical report
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6a7030aa-b42a-4730-9853-349c3df3b941 · outbound
Learning to Reason at the Frontier of Learnability MathScale: Scaling Instruction Tuning for Mathematical Reasoning
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8393f775-1f3a-43ca-937d-3c5d42a98506 · outbound
Learning to Reason at the Frontier of Learnability OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e66f894e-8ce2-45d8-b1a6-37b0f8284d24 · outbound
Learning to Reason at the Frontier of Learnability Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f494981d-ec13-4c91-babe-835993ab72ce · outbound
Learning to Reason at the Frontier of Learnability Proximal Curriculum for Reinforcement Learning Agents
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ebdc1395-d93e-4c1b-9d96-bd4315a448ef · outbound
Learning to Reason at the Frontier of Learnability Automatic goal generation for reinforcement learning agents
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1dce26a7-21f9-435d-a4e5-30ef10cb0337 · outbound
Learning to Reason at the Frontier of Learnability No Regrets: Investigating and Improving Regret Approximations for Curriculum Discovery
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6e35180a-fdf5-4874-95a9-d1a9af850592 · outbound
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation dca9b697-bd7c-4ec7-9771-f0cff09a2db2 · outbound
Learning to Reason at the Frontier of Learnability Asynchronous RLHF: Faster and More Efficient Off-Policy RL for Language Models
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation aaeb81a3-ee00-4047-b28f-a9f6e50632e0 · outbound
Learning to Reason at the Frontier of Learnability OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 10e0113a-ee17-49aa-9fcc-e8b7130c4ad7 · outbound
Learning to Reason at the Frontier of Learnability There may not be aha moment in r1-zero-like training — a pilot study
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ee5c9610-11d4-4c2e-856c-651230236188 · outbound
Learning to Reason at the Frontier of Learnability Numinamath
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d8c04222-da3e-444d-b7b9-697610f088e4 · outbound
Learning to Reason at the Frontier of Learnability Solving Quantitative Reasoning Problems with Language Models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 23390e2f-76d4-4c61-b3fb-2f047a4552c5 · outbound
Learning to Reason at the Frontier of Learnability Teaching Large Language Models to Reason with Reinforcement Learning
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e13af358-449a-4dc1-b110-c6013ee8c5fb · outbound
Learning to Reason at the Frontier of Learnability Prioritized Level Replay
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d59497fd-ba5c-400e-83c3-3288b371159f · outbound
Learning to Reason at the Frontier of Learnability Learning Montezuma's Revenge from a Single Demonstration
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b9767396-8989-4e33-86e3-0366a0ea8681 · outbound
Learning to Reason at the Frontier of Learnability Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5ba38af3-159d-467c-9c25-3b7f99c5bbf9 · outbound
Learning to Reason at the Frontier of Learnability DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a248f80a-a5e1-4951-93b9-13e1d36b7984 · outbound
Learning to Reason at the Frontier of Learnability Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b7293fb5-3be6-4a7c-9258-bd6653cd1ae0 · outbound
Learning to Reason at the Frontier of Learnability Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation fccbdeb1-b453-4849-a485-80f6c0eee41e · outbound
Learning to Reason at the Frontier of Learnability The llama 4 herd: The beginning of a new era of natively multimodal ai innovation
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 017d8ccd-e4a3-4d4e-a1d6-29eda5c39b62 · outbound
Learning to Reason at the Frontier of Learnability Emergent Complexity and Zero-shot Transfer via Unsupervised Environment Design
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 64e71307-ec78-4c6e-9ae0-e2a5656ffd7d · outbound
Learning to Reason at the Frontier of Learnability Minigrid & miniworld: Modular & customizable reinforcement learning environments for goal-oriented tasks
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2dbf633d-26d8-46a5-8cd3-57a13e2ed0ed · outbound
Learning to Reason at the Frontier of Learnability XLand-MiniGrid: Scalable Meta-Reinforcement Learning Environments in JAX
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8d6e0c89-2840-4822-a6f3-7c0dbb8d85cb · outbound
Learning to Reason at the Frontier of Learnability JaxMARL: Multi-Agent RL Environments and Algorithms in JAX
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 66e42759-9de4-40d8-8858-fdd92eea52f3 · outbound
Learning to Reason at the Frontier of Learnability OpenAI Gym
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6a88b35b-2f4c-4967-ab08-229c79b51f0a · outbound
Learning to Reason at the Frontier of Learnability JAX: composable transformations of Python+NumPy programs
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 75b9796e-604a-43ed-a377-36da011eea82 · outbound
Learning to Reason at the Frontier of Learnability Measuring short-form factuality in large language models
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b3d21010-bb44-459f-a500-7b07dff31324 · outbound
Learning to Reason at the Frontier of Learnability Evolving Curricula with Regret-Based Environment Design
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e35f324e-bd90-4fd6-bc13-56172bca0ec5 · outbound
Learning to Reason at the Frontier of Learnability Absolute Zero: Reinforced Self-play Reasoning with Zero Data
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ab6cbce9-9fc2-40f7-97e6-c4081d05ce6f · outbound
Learning to Reason at the Frontier of Learnability Unresolved cited work
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3de0a68b-6434-47f1-a568-07ee7e0a358c · outbound
Learning to Reason at the Frontier of Learnability Curriculum learning
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e4019767-2d6d-40b9-84de-c43f166aca00 · outbound
Learning to Reason at the Frontier of Learnability Learning and development in neural networks: The importance of starting small
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0d215df8-c324-44a5-adaa-71f935de9a1d · outbound
Learning to Reason at the Frontier of Learnability Online batch selection for faster training of neural networks
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation bda077f8-05ae-47ac-951f-036e3f34c4e6 · outbound
Learning to Reason at the Frontier of Learnability Online Batch Selection for Faster Training of Neural Networks
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 56f0f289-b948-4ea1-90a1-8512750e7c1d · outbound
Learning to Reason at the Frontier of Learnability Ordered SGD: A New Stochastic Optimization Framework for Empirical Risk Minimization
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 658cc2a6-ce97-4eb2-93cf-fe04e7e4dc14 · outbound
Learning to Reason at the Frontier of Learnability Accelerating Deep Learning by Focusing on the Biggest Losers
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8390b0f2-a71e-42e1-9618-87e0c85d49cc · outbound
Learning to Reason at the Frontier of Learnability Curriculum Learning by Transfer Learning: Theory and Experiments with Deep Networks
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation adf83da6-8a75-4fd5-8bb9-4405c3e7615a · outbound
Learning to Reason at the Frontier of Learnability Active learning literature survey
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation afc46c93-d264-4c35-9265-de39385a70cf · outbound
Learning to Reason at the Frontier of Learnability Confidence-based active learning
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0cd69aea-9ee3-423a-9512-d3105fcdcb0a · outbound
Learning to Reason at the Frontier of Learnability Selection via Proxy: Efficient Data Selection for Deep Learning
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b9113709-88e6-462b-9b33-c67eb55989bb · outbound
Learning to Reason at the Frontier of Learnability Prioritized Training on Points that are Learnable, Worth Learning, and Not Yet Learnt
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 898ef982-bd6f-4e4f-b954-3510565f7a36 · outbound
Learning to Reason at the Frontier of Learnability An Overview and a Benchmark of Active Learning for Outlier Detection with One-Class Classifiers
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f064bdf0-fd28-4c58-9430-a0768496b8c0 · outbound
Learning to Reason at the Frontier of Learnability Training deep models faster with robust, approximate importance sampling
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2a4d60cb-1c34-478e-bd5e-6b3a0e26c75d · outbound
Learning to Reason at the Frontier of Learnability Not all samples are created equal: Deep learning with importance sampling
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 341adf48-c1a3-4997-815f-3196b743385b · outbound
Learning to Reason at the Frontier of Learnability Self-paced learning for latent variable models
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ad35adb3-e42d-4c0e-92c1-9c32b935e279 · outbound
Learning to Reason at the Frontier of Learnability Automated Curriculum Learning for Neural Networks
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e480ceff-129f-4204-8390-55f1b6bf2d4f · outbound
Learning to Reason at the Frontier of Learnability Teacher-Student Curriculum Learning
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f7a01818-4359-47b6-a735-ded8010df04b · outbound
Learning to Reason at the Frontier of Learnability A survey of multi-task deep reinforcement learning
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4d81e9e8-2b97-4458-813d-3fd6076f4007 · outbound
Learning to Reason at the Frontier of Learnability Automatic curriculum learning through value disagreement
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c3e9b355-ff56-4a18-8a1b-b0961559ae68 · outbound
Learning to Reason at the Frontier of Learnability Automatic Curriculum Learning through Value Disagreement
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 64a38f54-5c17-4084-a444-728db0e96793 · outbound
Learning to Reason at the Frontier of Learnability Information-theoretic Task Selection for Meta-Reinforcement Learning
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 70dea3b6-cd80-46cf-8035-3d6f2862fc7a · outbound
Learning to Reason at the Frontier of Learnability Maximum Entropy Gain Exploration for Long Horizon Multi-goal Reinforcement Learning
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1f05f239-4449-48a3-b8f2-60e199228dce · outbound
Learning to Reason at the Frontier of Learnability Skew-Fit: State-Covering Self-Supervised Reinforcement Learning
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 727e61de-24b1-4793-b80a-fdf8899f22fe · outbound
Learning to Reason at the Frontier of Learnability CLIC: Curriculum Learning and Imitation for object Control in non-rewarding environments
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c34e21d6-43e9-4e82-bc52-0fda7635b0e5 · outbound
Learning to Reason at the Frontier of Learnability Goal-GAN: Multimodal Trajectory Prediction Based on Goal Position Estimation
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7124d511-8153-4b2c-9e55-ffa1792fb33e · outbound
Learning to Reason at the Frontier of Learnability Prioritized Experience Replay
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation cf3a376b-7f5e-449c-bb41-3270b1bba14a · outbound
Learning to Reason at the Frontier of Learnability In [48] the authors use the loss from a pre-trained model to estimate the difficulty of new samples for a freshly initialized network learning a new task
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e4942561-ca24-4a84-8d16-c660bc6881de · outbound
Learning to Reason at the Frontier of Learnability LILO can be seen as using return variance—or learnability—as an estimator of entropy or uncertainty
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e6b2d2c2-ed7d-4ac1-ad38-a25ab2db34a8 · outbound
Learning to Reason at the Frontier of Learnability This allows prioritizing samples that maximize the change in loss—i.e., the model’s learning progress [ 52, 11, 53, 49]
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 249db8c9-8b4f-4918-ba51-3441497bc233 · outbound
Learning to Reason at the Frontier of Learnability Self-paced learning [56] is an early approach that allows the model to determine the pace at which it incorporates harder examples with higher values of U
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c813b16f-bd7c-49aa-a48c-5d1e9b73ad32 · outbound
Learning to Reason at the Frontier of Learnability Each bullet point contains a claim and a hyperlink to the section of the paper that proves the claim
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5eb58b2e-1dfe-408b-a88e-b0a9712c4ceb · outbound
Learning to Reason at the Frontier of Learnability Section 7 also contains some limitations
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 78f48f7a-df59-48fa-9422-db9011825e5b · outbound
Learning to Reason at the Frontier of Learnability Guidelines: • The answer NA means that the paper does not include theoretical results
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3145bef9-d59f-475a-9721-c673f544bdd3 · outbound
Learning to Reason at the Frontier of Learnability The results in 6 were produced using open-source codebases [5] [4] and models, with some small additions of code by us
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a8ba68fb-4e0a-47d1-823d-1ead5b3a6ea7 · outbound
Learning to Reason at the Frontier of Learnability This is described in Section 5
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation dc7ecaf9-0b14-44d1-9a44-9487efbe03b1 · outbound
Learning to Reason at the Frontier of Learnability All the other hyperparameters for training are replicated directly from the VinePPO [ 5] and Oat [4] libraries, and the user is directed to these in Section 5
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d9bdbbaf-dec8-49ca-87ef-49ab1f1754ed · outbound
Learning to Reason at the Frontier of Learnability We have, however, provided training curves to aid the reader in interpreting the significance of the results
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d8e05a43-969a-4c9a-bbe6-3b08774999e8 · outbound
Learning to Reason at the Frontier of Learnability • The paper should indicate the type of compute workers CPU or GPU, internal cluster, or cloud provider, including relevant memory and storage
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4937820f-e32d-4279-bc46-91c07d4beddf · outbound
Learning to Reason at the Frontier of Learnability Guidelines: • The answer NA means that the authors have not reviewed the NeurIPS Code of Ethics
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a8a36f1f-ba68-4b49-a227-65ae6d3846c6 · outbound
Learning to Reason at the Frontier of Learnability Guidelines: • The answer NA means that there is no societal impact of the work performed
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7e20c96e-748b-44cd-8914-0ed7640fc2a7 · outbound
Learning to Reason at the Frontier of Learnability Guidelines: • The answer NA means that the paper poses no such risks
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6f8bbaa6-8709-40d1-8dab-d2b6437daf55 · outbound
Learning to Reason at the Frontier of Learnability The two libraries we used for training (VinePPO and Oat) are both fully open-source
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 990cf32f-c5b3-4c03-90c2-0d0db59eac4a · outbound
Learning to Reason at the Frontier of Learnability It very simple, and could be implemented from this paper alone
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 25126526-67c4-4dd8-9def-b2d074d6cd84 · outbound
Learning to Reason at the Frontier of Learnability Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8aa1cc51-a145-43a9-813f-4d0b0bb8c35d · outbound
Learning to Reason at the Frontier of Learnability Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a35994fa-622d-45e2-b5de-d565cf4ea1f2 · outbound
Learning to Reason at the Frontier of Learnability Answer: [NA] Justification: LLM usage was only used in a standard way for editing
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2e123e76-4eb5-44c5-8243-1d08c19a53f0 · inbound
Multi-Task GRPO: Reliable LLM Reasoning Across Tasks Learning to Reason at the Frontier of Learnability
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.