Pith. sign in

Paper Citation Record · LEDGER

Learning to Reason at the Frontier of Learnability

As of 5 August 2026, this Paper Citation Record lists 89 of 89 outbound references and 1 inbound Pith citation observation for arXiv:2502.12272.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.12272 v6

Coverage vector

measured 89 of 89 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-23T02:41:21.571824Z

measured 90 of 90 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T04:17:27.154486Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

89 of 89 outbound references displayed

  • verified exact46
  • verified fuzzy39
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4f7f7a4c-d5f1-4184-8a17-8f69d175f469 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Learning to Reason at the Frontier of Learnability DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-23T02:42:26.069190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:52c41e9538c9e015b450425d57b1f32e1b2719d65c8b2ba9b95e61e8130ac8c8

Observation c2cbddbf-206a-4b79-9e27-6848eb81e78e · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

Learning to Reason at the Frontier of Learnability Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-23T02:42:26.038058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:53f7c8e4a34e5ba12567eb85debfc50a644072b69567bee11d3f53949bbbbbcd

Observation fdaf63a6-feec-45ff-9f96-01aa0bad26f3 · outbound

This paper cites Learning to reason with llms.

Learning to Reason at the Frontier of Learnability Learning to reason with llms

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T02:47:27.556981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:bd14f969b49698fde02f732fc782301ab16817e4a9cc4d83434fda559bdd5ade

Observation b290130e-c07f-4672-9a7d-35b03ff5ed0b · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

Learning to Reason at the Frontier of Learnability Understanding R1-Zero-Like Training: A Critical Perspective

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-23T02:42:26.030208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:daea93bb086d265428af0e8f3b16330917cf0f3472200857694a52fd501bfd4e

Observation 7046d60f-b54a-4c5a-a42b-cca48dcecaad · outbound

This paper cites Vineppo: Unlocking rl potential for llm reasoning through refined credit assignment.

Learning to Reason at the Frontier of Learnability Vineppo: Unlocking rl potential for llm reasoning through refined credit assignment

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T02:47:27.553779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:7c81c76e2bbc953e92c63022c9052fa488fe07204c91708709c61d6c2b2a5545

Observation 66fa54c1-421c-4938-8538-0c19b137cf03 · outbound

This paper cites VinePPO: Refining Credit Assignment in RL Training of LLMs.

Learning to Reason at the Frontier of Learnability VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-23T02:42:26.025966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:46d2d0d718c1d9e533dc22a8df0a950c95ba0411b24f02b68ada4a1e5fc7b69c

Observation 79f22151-c393-4174-970b-f15eaa667fd7 · outbound

This paper cites Group Robust Preference Optimization in Reward-free RLHF.

Learning to Reason at the Frontier of Learnability Group Robust Preference Optimization in Reward-free RLHF

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-23T02:42:26.236094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:ef700741843e9d38f1c57dc2f445d6206674910c09fc9c116772c79805890f61

Observation b2f396c2-77ad-4b05-8ab4-ceb0fcc0b99c · outbound

This paper cites Proximal Policy Optimization Algorithms.

Learning to Reason at the Frontier of Learnability Proximal Policy Optimization Algorithms

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-23T02:42:26.231736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:f6902ad3c71e34aede9c34e6f09e05e8582dd374cc387f9eb2a6a2dfc632c928

Observation 137dd64d-fe7f-46dd-bbdb-6ecd5ba5c84e · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Learning to Reason at the Frontier of Learnability Training Verifiers to Solve Math Word Problems

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-23T02:42:26.227411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:1c6c9ace738768278a769a5ba05cde42e13e5286bf031f7faac36ee195e6e4d9

Observation 42f79281-0270-42b9-9207-899617751327 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Learning to Reason at the Frontier of Learnability Measuring Mathematical Problem Solving With the MATH Dataset

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-23T02:42:26.081240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:d6edced4c2f006e6c6320d28c9f910f8dcdd1a5d3d894b98d98b7db9160dc972

Observation 0c878d87-271b-49f9-a748-107161b0410d · outbound

This paper cites Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model.

Learning to Reason at the Frontier of Learnability Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-23T02:42:26.222795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:b392419eb0cc44cab5432c4ba8640ba54ab9d9c004cf7a15cf87bb7de5e946f3

Observation 95ee720f-f476-4f9e-b569-0f3e85e7c30c · outbound

This paper cites Rho-1: Not All Tokens Are What You Need.

Learning to Reason at the Frontier of Learnability Rho-1: Not All Tokens Are What You Need

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-23T02:42:26.218182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:a26ed2b91631d7838d00156773df8f98631988dbe03da50c9acabd2204f4c7e5

Observation f955939b-2b2a-4736-afb8-5f09842eb4d9 · outbound

This paper cites Qwen2.5 technical report.

Learning to Reason at the Frontier of Learnability Qwen2.5 technical report

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T02:47:27.549825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:fe0add013e978004d5b39c434669f39eb9a1f0f3665171fefcb978464b81392f

Observation 6a7030aa-b42a-4730-9853-349c3df3b941 · outbound

This paper cites MathScale: Scaling Instruction Tuning for Mathematical Reasoning.

Learning to Reason at the Frontier of Learnability MathScale: Scaling Instruction Tuning for Mathematical Reasoning

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-23T02:42:26.213553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:1a5b1b68bbd882716da7846bf90fb0be125161b9e58523ce82824cff13283da6

Observation 8393f775-1f3a-43ca-937d-3c5d42a98506 · outbound

This paper cites OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

Learning to Reason at the Frontier of Learnability OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-23T02:42:26.208213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:2e33f90457d60b5c5736ea7fc989db846734536a26e71527ef896507f918eae3

Observation e66f894e-8ce2-45d8-b1a6-37b0f8284d24 · outbound

This paper cites Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs.

Learning to Reason at the Frontier of Learnability Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-23T02:42:26.045213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:2c1a0850e85feef6da7c7198833e3c3adf2d133275242f201094bed04e761cbf

Observation f494981d-ec13-4c91-babe-835993ab72ce · outbound

This paper cites Proximal Curriculum for Reinforcement Learning Agents.

Learning to Reason at the Frontier of Learnability Proximal Curriculum for Reinforcement Learning Agents

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-23T02:42:26.073171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:ff2c64dd455917733e61fe8400da26e0ae0015d8680f933088faa50029e37fe2

Observation ebdc1395-d93e-4c1b-9d96-bd4315a448ef · outbound

This paper cites Automatic goal generation for reinforcement learning agents.

Learning to Reason at the Frontier of Learnability Automatic goal generation for reinforcement learning agents

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T02:47:27.523015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:ceaa9fec66784aaa3a663d9f9c7b3b1e31ae48298f1d3847a6a4e3d13d16c9e9

Observation 1dce26a7-21f9-435d-a4e5-30ef10cb0337 · outbound

This paper cites No Regrets: Investigating and Improving Regret Approximations for Curriculum Discovery.

Learning to Reason at the Frontier of Learnability No Regrets: Investigating and Improving Regret Approximations for Curriculum Discovery

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-23T02:42:26.141497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:8b4630277068f731414b2fc1dc8a3eb234ec4cd6caa925d7cadb692e8e037d1e

Observation 6e35180a-fdf5-4874-95a9-d1a9af850592 · outbound

This paper cites Williams.

Learning to Reason at the Frontier of Learnability Williams

Reference 20

Resolution
metadata mismatch
doi, observed 2026-05-23T02:42:25.514626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:69f8db147fd8a62791aac680a5c7bee9fcfad5e9aa658b7d6192a2cac0e2251a

Observation dca9b697-bd7c-4ec7-9771-f0cff09a2db2 · outbound

This paper cites Asynchronous RLHF: Faster and More Efficient Off-Policy RL for Language Models.

Learning to Reason at the Frontier of Learnability Asynchronous RLHF: Faster and More Efficient Off-Policy RL for Language Models

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-23T02:42:26.179007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:89a044acf41fe817e63cb1982ae52eaf46abb9522460d22661bc45083f760fe4

Observation aaeb81a3-ee00-4047-b28f-a9f6e50632e0 · outbound

This paper cites OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework.

Learning to Reason at the Frontier of Learnability OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-23T02:42:26.089482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:97c15bcbff39cc98719f539692eea141fb0ca88211497af94f3e415c4314cfd9

Observation 10e0113a-ee17-49aa-9fcc-e8b7130c4ad7 · outbound

This paper cites There may not be aha moment in r1-zero-like training — a pilot study.

Learning to Reason at the Frontier of Learnability There may not be aha moment in r1-zero-like training — a pilot study

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T02:47:27.519337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:692d13f4b73770f8606b4cefdc2afec0ba30a6611971a31e0e58d6a5fa137150

Observation ee5c9610-11d4-4c2e-856c-651230236188 · outbound

This paper cites Numinamath.

Learning to Reason at the Frontier of Learnability Numinamath

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T02:47:27.573705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:1f8f7ee387bb4754b26d4ab0dc201cbc57da0d6b85eade5b9fe1a0dab0e2e676

Observation d8c04222-da3e-444d-b7b9-697610f088e4 · outbound

This paper cites Solving Quantitative Reasoning Problems with Language Models.

Learning to Reason at the Frontier of Learnability Solving Quantitative Reasoning Problems with Language Models

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-23T02:42:26.053039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:3ddaac2ee98d22e580d10c9ec36e0ee8858c83c2c1d6b47c157e7901d1bcf974

Observation 23390e2f-76d4-4c61-b3fb-2f047a4552c5 · outbound

This paper cites Teaching Large Language Models to Reason with Reinforcement Learning.

Learning to Reason at the Frontier of Learnability Teaching Large Language Models to Reason with Reinforcement Learning

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-23T02:42:26.085841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:01dda4b686a64bf6003ab97c9c3d49b03b751ebdcd45eae0ae35265d998da9ad

Observation e13af358-449a-4dc1-b110-c6013ee8c5fb · outbound

This paper cites Prioritized Level Replay.

Learning to Reason at the Frontier of Learnability Prioritized Level Replay

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-23T02:42:26.065432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:2aa8ba5f55cfdd08819ef42f4c3111a890925a591abe9be11f6ced3f719021a1

Observation d59497fd-ba5c-400e-83c3-3288b371159f · outbound

This paper cites Learning Montezuma's Revenge from a Single Demonstration.

Learning to Reason at the Frontier of Learnability Learning Montezuma's Revenge from a Single Demonstration

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-23T02:42:26.057246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:178fde1f6c6cc7d548577e3def5507970e75af1737dd318659ec9c8f346b58e1

Observation b9767396-8989-4e33-86e3-0366a0ea8681 · outbound

This paper cites Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive.

Learning to Reason at the Frontier of Learnability Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-23T02:42:26.061849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:8d8ffba946a568d0fceab0669fd79c92a479edcb86ce1d5a03bc50c50950304a

Observation 5ba38af3-159d-467c-9c25-3b7f99c5bbf9 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Learning to Reason at the Frontier of Learnability DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-23T02:42:26.203304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:2437351c37ef581a568a62c5710bd0729f58530a04571b13ca95fb0253771b73

Observation a248f80a-a5e1-4951-93b9-13e1d36b7984 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

Learning to Reason at the Frontier of Learnability Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-23T02:42:26.188553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:d88148fbe7a057cff04e620ab881312cf0b5a1ef469c011aba592f7b9511531e

Observation b7293fb5-3be6-4a7c-9258-bd6653cd1ae0 · outbound

This paper cites Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning.

Learning to Reason at the Frontier of Learnability Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-23T02:42:26.183846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:cff614f204160d1d8008847644e2c0b6156c400efd6fce5e20e5a10e5e642b63

Observation fccbdeb1-b453-4849-a485-80f6c0eee41e · outbound

This paper cites The llama 4 herd: The beginning of a new era of natively multimodal ai innovation.

Learning to Reason at the Frontier of Learnability The llama 4 herd: The beginning of a new era of natively multimodal ai innovation

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T02:47:27.511125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:d28d95fb8b69cf435c0d0011f67562ee6889da1725aacd7e6ac23095c44ed42b

Observation 017d8ccd-e4a3-4d4e-a1d6-29eda5c39b62 · outbound

This paper cites Emergent Complexity and Zero-shot Transfer via Unsupervised Environment Design.

Learning to Reason at the Frontier of Learnability Emergent Complexity and Zero-shot Transfer via Unsupervised Environment Design

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-23T02:42:26.198274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:db5dccd5dbbc676e537ca3be0458c0c7b83f07ea813c1c1818431337835c3279

Observation 64e71307-ec78-4c6e-9ae0-e2a5656ffd7d · outbound

This paper cites Minigrid & miniworld: Modular & customizable reinforcement learning environments for goal-oriented tasks.

Learning to Reason at the Frontier of Learnability Minigrid & miniworld: Modular & customizable reinforcement learning environments for goal-oriented tasks

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T02:47:27.567335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:516be23fe00d3bf3b37e3f7ce388d28317bbb3065f761dbf5f15bdd1bff344a8

Observation 2dbf633d-26d8-46a5-8cd3-57a13e2ed0ed · outbound

This paper cites XLand-MiniGrid: Scalable Meta-Reinforcement Learning Environments in JAX.

Learning to Reason at the Frontier of Learnability XLand-MiniGrid: Scalable Meta-Reinforcement Learning Environments in JAX

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-23T02:42:26.048877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:bed95faacaf69c5216731f6f9e86539139d75bee56ca6d04dd244051f0b3d531

Observation 8d6e0c89-2840-4822-a6f3-7c0dbb8d85cb · outbound

This paper cites JaxMARL: Multi-Agent RL Environments and Algorithms in JAX.

Learning to Reason at the Frontier of Learnability JaxMARL: Multi-Agent RL Environments and Algorithms in JAX

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-07-07T02:17:02.673620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:6da9c5d958bc0207f348580303eb8faad4eb0641197d3c0b7421449869df80b6

Observation 66e42759-9de4-40d8-8858-fdd92eea52f3 · outbound

This paper cites OpenAI Gym.

Learning to Reason at the Frontier of Learnability OpenAI Gym

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-23T02:42:26.173933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:ce343a1d232623e6d580ac37f68e12853173a4093711ac6b1f12659a408f7781

Observation 6a88b35b-2f4c-4967-ab08-229c79b51f0a · outbound

This paper cites JAX: composable transformations of Python+NumPy programs.

Learning to Reason at the Frontier of Learnability JAX: composable transformations of Python+NumPy programs

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T02:47:27.507646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:f231f31c4f8fdb002b97d923b9f330484f0042b064eb404c78e10fdf49c97620

Observation 75b9796e-604a-43ed-a377-36da011eea82 · outbound

This paper cites Measuring short-form factuality in large language models.

Learning to Reason at the Frontier of Learnability Measuring short-form factuality in large language models

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-23T02:42:26.155841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:b1ed92ec54206198c214dc8908d968f8e4f2823b863db4ad1ba9755592922ed4

Observation b3d21010-bb44-459f-a500-7b07dff31324 · outbound

This paper cites Evolving Curricula with Regret-Based Environment Design.

Learning to Reason at the Frontier of Learnability Evolving Curricula with Regret-Based Environment Design

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-23T02:42:26.165070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:af2e2d4f3e33b1292e3053c06ba6512e7064579e3f7217779c9816e1a9c931ae

Observation e35f324e-bd90-4fd6-bc13-56172bca0ec5 · outbound

This paper cites Absolute Zero: Reinforced Self-play Reasoning with Zero Data.

Learning to Reason at the Frontier of Learnability Absolute Zero: Reinforced Self-play Reasoning with Zero Data

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-05-23T02:42:26.077279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:1d834f70ff541eaf484daa1e28941fbecfc872b33ea0feeed3cfec8688d57e47

Observation ab6cbce9-9fc2-40f7-97e6-c4081d05ce6f · outbound

This paper cites an unresolved cited work.

Learning to Reason at the Frontier of Learnability Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-05-23T02:47:27.563811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:e01f143e5111f34f5de52ee3b4ebe80800fe304ffa3131bf6f1967591d59837a

Observation 3de0a68b-6434-47f1-a568-07ee7e0a358c · outbound

This paper cites Curriculum learning.

Learning to Reason at the Frontier of Learnability Curriculum learning

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T02:47:27.503948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:4d675af1476a3279086ed3d864465690891c669bf0d8866e3c00aea56ba58341

Observation e4019767-2d6d-40b9-84de-c43f166aca00 · outbound

This paper cites Learning and development in neural networks: The importance of starting small.

Learning to Reason at the Frontier of Learnability Learning and development in neural networks: The importance of starting small

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T02:47:27.500421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:1798b50e031d840e499782b9fc3be154551de808210df26ce399632062d5ef46

Observation 0d215df8-c324-44a5-adaa-71f935de9a1d · outbound

This paper cites Online batch selection for faster training of neural networks.

Learning to Reason at the Frontier of Learnability Online batch selection for faster training of neural networks

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T02:47:27.496724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:c09d0450d3baa5d039f5b3e43460016b1e3e82c0b0804fea5301f55e700770c9

Observation bda077f8-05ae-47ac-951f-036e3f34c4e6 · outbound

This paper cites Online Batch Selection for Faster Training of Neural Networks.

Learning to Reason at the Frontier of Learnability Online Batch Selection for Faster Training of Neural Networks

Reference 47

Resolution
metadata mismatch
local_arxiv, observed 2026-05-23T02:42:26.116137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:6b189087214cc52aba4f686ce86b9cb8f2d28331741e2c787b194834ad4c40ed

Observation 56f0f289-b948-4ea1-90a1-8512750e7c1d · outbound

This paper cites Ordered SGD: A New Stochastic Optimization Framework for Empirical Risk Minimization.

Learning to Reason at the Frontier of Learnability Ordered SGD: A New Stochastic Optimization Framework for Empirical Risk Minimization

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-23T02:42:26.193403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:3b32055b19e8998620da5470227730da2a1da0b79b916cd32a8f550b14dab8fb

Observation 658cc2a6-ce97-4eb2-93cf-fe04e7e4dc14 · outbound

This paper cites Accelerating Deep Learning by Focusing on the Biggest Losers.

Learning to Reason at the Frontier of Learnability Accelerating Deep Learning by Focusing on the Biggest Losers

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-23T02:42:26.150995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:c2aab46306c16af58ad62cba8fcfed22b11a19378860b50ed730248db5cfe448

Observation 8390b0f2-a71e-42e1-9618-87e0c85d49cc · outbound

This paper cites Curriculum Learning by Transfer Learning: Theory and Experiments with Deep Networks.

Learning to Reason at the Frontier of Learnability Curriculum Learning by Transfer Learning: Theory and Experiments with Deep Networks

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-23T02:42:26.169796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:a968f3ae5fa3cda64a30e42792852fc01a02f9a96e8da560576301ebf7f2df89

Observation adf83da6-8a75-4fd5-8bb9-4405c3e7615a · outbound

This paper cites Active learning literature survey.

Learning to Reason at the Frontier of Learnability Active learning literature survey

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T02:47:27.492887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:220fd46a6c5d95d5d68251d774e95d0157943ecdd445fde75ba8c41bd63f501d

Observation afc46c93-d264-4c35-9265-de39385a70cf · outbound

This paper cites Confidence-based active learning.

Learning to Reason at the Frontier of Learnability Confidence-based active learning

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T02:47:27.489016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:e68d92e5d22d0d38e836bc5d5e78fa3c646407a1fcf35a3640b8df133812fdb9

Observation 0cd69aea-9ee3-423a-9512-d3105fcdcb0a · outbound

This paper cites Selection via Proxy: Efficient Data Selection for Deep Learning.

Learning to Reason at the Frontier of Learnability Selection via Proxy: Efficient Data Selection for Deep Learning

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-23T02:42:26.121575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:b5ed58f89e6b83884ac0f714f82fd617800bf9ffa42d7ef4862bbd9562f0d6c6

Observation b9113709-88e6-462b-9b33-c67eb55989bb · outbound

This paper cites Prioritized Training on Points that are Learnable, Worth Learning, and Not Yet Learnt.

Learning to Reason at the Frontier of Learnability Prioritized Training on Points that are Learnable, Worth Learning, and Not Yet Learnt

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-23T02:42:26.126116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:0589a9ecb6b9fbdb27f8854200f181d60c3bb2c10fcd4a11447b4ee1cf5d842e

Observation 898ef982-bd6f-4e4f-b954-3510565f7a36 · outbound

This paper cites An Overview and a Benchmark of Active Learning for Outlier Detection with One-Class Classifiers.

Learning to Reason at the Frontier of Learnability An Overview and a Benchmark of Active Learning for Outlier Detection with One-Class Classifiers

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-05-23T02:42:26.111402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:b7fbdcc562849ec29d56a9f2cb0588da7ee1004088b33903f4c9970310801be9

Observation f064bdf0-fd28-4c58-9430-a0768496b8c0 · outbound

This paper cites Training deep models faster with robust, approximate importance sampling.

Learning to Reason at the Frontier of Learnability Training deep models faster with robust, approximate importance sampling

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T02:47:27.622517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:e93202060fdcfd4975a0d75cd68ee2788c29446e2c4b6f311027b92156ffd571

Observation 2a4d60cb-1c34-478e-bd5e-6b3a0e26c75d · outbound

This paper cites Not all samples are created equal: Deep learning with importance sampling.

Learning to Reason at the Frontier of Learnability Not all samples are created equal: Deep learning with importance sampling

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T02:47:27.618558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:f8d80f165976930fe73975d1ea2eb0877c055a6f46bed16259fa8a03c5bf29c9

Observation 341adf48-c1a3-4997-815f-3196b743385b · outbound

This paper cites Self-paced learning for latent variable models.

Learning to Reason at the Frontier of Learnability Self-paced learning for latent variable models

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T02:47:27.634546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:47f77c06d488b37a3b728fa4f1d94877035e0a49a37ae3ddc4530b37d93d7ef1

Observation ad35adb3-e42d-4c0e-92c1-9c32b935e279 · outbound

This paper cites Automated Curriculum Learning for Neural Networks.

Learning to Reason at the Frontier of Learnability Automated Curriculum Learning for Neural Networks

Reference 59

Resolution
verified exact
local_arxiv, observed 2026-05-23T02:42:26.130839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:0bd8c4a43322ed68d37548bdc6fa5a5629b2df4687cd745763aa81f743e33f9b

Observation e480ceff-129f-4204-8390-55f1b6bf2d4f · outbound

This paper cites Teacher-Student Curriculum Learning.

Learning to Reason at the Frontier of Learnability Teacher-Student Curriculum Learning

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-05-23T02:42:26.097838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:ceb93574d8eb18623073c7da4e8947e2fc1df63248462478720af405dcefa554

Observation f7a01818-4359-47b6-a735-ded8010df04b · outbound

This paper cites A survey of multi-task deep reinforcement learning.

Learning to Reason at the Frontier of Learnability A survey of multi-task deep reinforcement learning

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T02:47:27.611853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:9e4db0741b38a93f393263cba4fb07cd054a6fbbbc8dfe24a66799367a5dbe7f

Observation 4d81e9e8-2b97-4458-813d-3fd6076f4007 · outbound

This paper cites Automatic curriculum learning through value disagreement.

Learning to Reason at the Frontier of Learnability Automatic curriculum learning through value disagreement

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T02:47:27.604553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:557c81950d76f81b429bd5ad6e8123f538418009dcd00d24ee9284a852d5aa28

Observation c3e9b355-ff56-4a18-8a1b-b0961559ae68 · outbound

This paper cites Automatic Curriculum Learning through Value Disagreement.

Learning to Reason at the Frontier of Learnability Automatic Curriculum Learning through Value Disagreement

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-23T02:42:26.106419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:58827b03980364972568e60244050a0dd64fbd43aed56fcf6fb93595bc270741

Observation 64a38f54-5c17-4084-a444-728db0e96793 · outbound

This paper cites Information-theoretic Task Selection for Meta-Reinforcement Learning.

Learning to Reason at the Frontier of Learnability Information-theoretic Task Selection for Meta-Reinforcement Learning

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-23T02:42:26.093984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:86c76e7cc2fed800637a8d86ce2d1e3ff70c62594fc9507de9caf0376383585d

Observation 70dea3b6-cd80-46cf-8035-3d6f2862fc7a · outbound

This paper cites Maximum Entropy Gain Exploration for Long Horizon Multi-goal Reinforcement Learning.

Learning to Reason at the Frontier of Learnability Maximum Entropy Gain Exploration for Long Horizon Multi-goal Reinforcement Learning

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-23T02:42:26.135156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:a6ddb4ab048e61a2749120d07d2e64943a0b60a868d95af988462e416b037aa4

Observation 1f05f239-4449-48a3-b8f2-60e199228dce · outbound

This paper cites Skew-Fit: State-Covering Self-Supervised Reinforcement Learning.

Learning to Reason at the Frontier of Learnability Skew-Fit: State-Covering Self-Supervised Reinforcement Learning

Reference 66

Resolution
metadata mismatch
arxiv_id, observed 2026-05-23T02:42:26.146328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:78bcfc180466a49353ef4f46b0317c9a54f6fe4669df9948ee534361716a3895

Observation 727e61de-24b1-4793-b80a-fdf8899f22fe · outbound

This paper cites CLIC: Curriculum Learning and Imitation for object Control in non-rewarding environments.

Learning to Reason at the Frontier of Learnability CLIC: Curriculum Learning and Imitation for object Control in non-rewarding environments

Reference 67

Resolution
verified exact
local_arxiv, observed 2026-05-23T02:42:26.101950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:7f55cb2eafde5e1723da3d24ee2fb20ac5215d1c8db45eb1b0e8d7726ff7eaa6

Observation c34e21d6-43e9-4e82-bc52-0fda7635b0e5 · outbound

This paper cites Goal-GAN: Multimodal Trajectory Prediction Based on Goal Position Estimation.

Learning to Reason at the Frontier of Learnability Goal-GAN: Multimodal Trajectory Prediction Based on Goal Position Estimation

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-23T02:42:26.034216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:82d4417b1a13bdbdcabf4b5290dfb2a64126fcddeede1b7d14329d049185c096

Observation 7124d511-8153-4b2c-9e55-ffa1792fb33e · outbound

This paper cites Prioritized Experience Replay.

Learning to Reason at the Frontier of Learnability Prioritized Experience Replay

Reference 69

Resolution
verified exact
local_arxiv, observed 2026-05-23T02:42:26.041639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:8a93c8c7c8374fbd24a85dcf336d81b5e147e9c11eb8bbc15c7eed6b0f3a9f12

Observation cf3a376b-7f5e-449c-bb41-3270b1bba14a · outbound

This paper cites In [48] the authors use the loss from a pre-trained model to estimate the difficulty of new samples for a freshly initialized network learning a new task.

Learning to Reason at the Frontier of Learnability In [48] the authors use the loss from a pre-trained model to estimate the difficulty of new samples for a freshly initialized network learning a new task

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T02:47:27.597336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:1f505d32c0a92565036c2d912b267ae9f88e2816996222c5cc467f8328631d1d

Observation e4942561-ca24-4a84-8d16-c660bc6881de · outbound

This paper cites LILO can be seen as using return variance—or learnability—as an estimator of entropy or uncertainty.

Learning to Reason at the Frontier of Learnability LILO can be seen as using return variance—or learnability—as an estimator of entropy or uncertainty

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T02:47:27.580531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:f1377d07f04b36770b0f54f3c06c8b9fa5901db71204e2be5b9e474ca110ae7b

Observation e6b2d2c2-ed7d-4ac1-ad38-a25ab2db34a8 · outbound

This paper cites This allows prioritizing samples that maximize the change in loss—i.e., the model’s learning progress [ 52, 11, 53, 49].

Learning to Reason at the Frontier of Learnability This allows prioritizing samples that maximize the change in loss—i.e., the model’s learning progress [ 52, 11, 53, 49]

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T02:47:27.585243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:0d60813c654686b06f989deb34563245cb695a0bebcead6b69628777bd4d31b0

Observation 249db8c9-8b4f-4918-ba51-3441497bc233 · outbound

This paper cites Self-paced learning [56] is an early approach that allows the model to determine the pace at which it incorporates harder examples with higher values of U.

Learning to Reason at the Frontier of Learnability Self-paced learning [56] is an early approach that allows the model to determine the pace at which it incorporates harder examples with higher values of U

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T02:47:27.608148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:3d5edb3ab7b5b8330f7f4e2ef9bdb8c931d600040cdb95dffef18877eead620c

Observation c813b16f-bd7c-49aa-a48c-5d1e9b73ad32 · outbound

This paper cites Each bullet point contains a claim and a hyperlink to the section of the paper that proves the claim.

Learning to Reason at the Frontier of Learnability Each bullet point contains a claim and a hyperlink to the section of the paper that proves the claim

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T02:47:27.534779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:981e5df43e1927537e2ed04fc994ebbbe74e803bf71bbabb3711527de37490d7

Observation 5eb58b2e-1dfe-408b-a88e-b0a9712c4ceb · outbound

This paper cites Section 7 also contains some limitations.

Learning to Reason at the Frontier of Learnability Section 7 also contains some limitations

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T02:47:27.527386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:f25430690617f16b31aa7503e5c2a43e1d458f108d351ade8df137c2b40efd9e

Observation 78f48f7a-df59-48fa-9422-db9011825e5b · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include theoretical results.

Learning to Reason at the Frontier of Learnability Guidelines: • The answer NA means that the paper does not include theoretical results

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T02:47:27.542651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:b609b6d13affd5063ef6329c4f683d929988efd32c5fea9a335b9808cff45df2

Observation 3145bef9-d59f-475a-9721-c673f544bdd3 · outbound

This paper cites The results in 6 were produced using open-source codebases [5] [4] and models, with some small additions of code by us.

Learning to Reason at the Frontier of Learnability The results in 6 were produced using open-source codebases [5] [4] and models, with some small additions of code by us

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T02:47:27.514929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:acaba03e8f426beae0cd584e015f6af6b2863a7d69bdfdc9c6347002a98948fd

Observation a8ba68fb-4e0a-47d1-823d-1ead5b3a6ea7 · outbound

This paper cites This is described in Section 5.

Learning to Reason at the Frontier of Learnability This is described in Section 5

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T02:47:27.546614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:56722316ccc8edbe71b83aea53df7fc8c5097e7c0e7db8e19b86a9e5e570d7c7

Observation dc7ecaf9-0b14-44d1-9a44-9487efbe03b1 · outbound

This paper cites All the other hyperparameters for training are replicated directly from the VinePPO [ 5] and Oat [4] libraries, and the user is directed to these in Section 5.

Learning to Reason at the Frontier of Learnability All the other hyperparameters for training are replicated directly from the VinePPO [ 5] and Oat [4] libraries, and the user is directed to these in Section 5

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T02:47:27.560504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:faae52557e05458a7e599d7d7515f745a1eedda5963467025ee4f2d2ab0817bc

Observation d9bdbbaf-dec8-49ca-87ef-49ab1f1754ed · outbound

This paper cites We have, however, provided training curves to aid the reader in interpreting the significance of the results.

Learning to Reason at the Frontier of Learnability We have, however, provided training curves to aid the reader in interpreting the significance of the results

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T02:47:27.614972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:15e9e47da5bb54c56f54792bdca59d477ae6dc2e55cc7746e69139794cd3308a

Observation d8e05a43-969a-4c9a-bbe6-3b08774999e8 · outbound

This paper cites • The paper should indicate the type of compute workers CPU or GPU, internal cluster, or cloud provider, including relevant memory and storage.

Learning to Reason at the Frontier of Learnability • The paper should indicate the type of compute workers CPU or GPU, internal cluster, or cloud provider, including relevant memory and storage

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T02:47:27.626385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:c33129529923eb72705661273065a9ee15c5865386810fa3a64ba4a6d72e32b9

Observation 4937820f-e32d-4279-bc46-91c07d4beddf · outbound

This paper cites Guidelines: • The answer NA means that the authors have not reviewed the NeurIPS Code of Ethics.

Learning to Reason at the Frontier of Learnability Guidelines: • The answer NA means that the authors have not reviewed the NeurIPS Code of Ethics

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T02:47:27.538684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:e5186050d8b9746f86d22d2285b28950a7eef79983e8ad443c42ad40c43383c3

Observation a8a36f1f-ba68-4b49-a227-65ae6d3846c6 · outbound

This paper cites Guidelines: • The answer NA means that there is no societal impact of the work performed.

Learning to Reason at the Frontier of Learnability Guidelines: • The answer NA means that there is no societal impact of the work performed

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T02:47:27.589234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:08bd8df05e91fdce676336f6f5201a7c96c8bacfec5f028aed53bb3ccd871bdd

Observation 7e20c96e-748b-44cd-8914-0ed7640fc2a7 · outbound

This paper cites Guidelines: • The answer NA means that the paper poses no such risks.

Learning to Reason at the Frontier of Learnability Guidelines: • The answer NA means that the paper poses no such risks

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T02:47:27.593649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:96aaee87e525bbbb40cc9a00f417fae3e3a2cf6d21435950d42bbc9781119da8

Observation 6f8bbaa6-8709-40d1-8dab-d2b6437daf55 · outbound

This paper cites The two libraries we used for training (VinePPO and Oat) are both fully open-source.

Learning to Reason at the Frontier of Learnability The two libraries we used for training (VinePPO and Oat) are both fully open-source

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T02:47:27.570800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:d976c0990281d2c1d944b194ee68d2bf48c6ee78b35fda51ec7c65097bf6c1c0

Observation 990cf32f-c5b3-4c03-90c2-0d0db59eac4a · outbound

This paper cites It very simple, and could be implemented from this paper alone.

Learning to Reason at the Frontier of Learnability It very simple, and could be implemented from this paper alone

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T02:47:27.576905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:456b7146faa299ff074d673e3db66afeda0dc5234509af4a483df6509d78d375

Observation 25126526-67c4-4dd8-9def-b2d074d6cd84 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects.

Learning to Reason at the Frontier of Learnability Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T02:47:27.600904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:c1468b2419dedb62ea762ac1c3908cbaf6a1c496733b5e686c8724b4a7a3c847

Observation 8aa1cc51-a145-43a9-813f-4d0b0bb8c35d · outbound

This paper cites Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects.

Learning to Reason at the Frontier of Learnability Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T02:47:27.630722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:65701fe19758e99db68c075a8c466ae1a353bc17420228578854ee6164b47b40

Observation a35994fa-622d-45e2-b5de-d565cf4ea1f2 · outbound

This paper cites Answer: [NA] Justification: LLM usage was only used in a standard way for editing.

Learning to Reason at the Frontier of Learnability Answer: [NA] Justification: LLM usage was only used in a standard way for editing

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T02:47:27.531069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:ebb2a0362215968c69f7725851b147d337a6400a3df7e54ec44ba165fb486677

Pith citing papers

Observation 2e123e76-4eb5-44c5-8243-1d08c19a53f0 · inbound

Multi-Task GRPO: Reliable LLM Reasoning Across Tasks cites this paper.

Multi-Task GRPO: Reliable LLM Reasoning Across Tasks Learning to Reason at the Frontier of Learnability

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T04:17:27.154486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:17:27.154486Z digest=sha256:c22c686b372221d4e150e6cc4f32ab1116897129eb98b3ec52898e0fa2b5d218