Pith. sign in

Paper Citation Record · LEDGER

Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 85 inbound Pith citation observations for arXiv:2312.06585.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.06585 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 85 of 85 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 85 of 85 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:29:18.697529Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T21:00:08.414065Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ab4126f8-e276-4acd-b497-4213e1f02631 · inbound

An Iterative Utility Judgment Framework Inspired by Philosophical Relevance via LLMs cites this paper.

An Iterative Utility Judgment Framework Inspired by Philosophical Relevance via LLMs Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-24T00:15:52.765918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-24T00:14:42.779546Z digest=sha256:45363d66c9a89ec6391938c3f0bcbee9f0287110b34d375f18481d5a429419e2

Observation 84baa3d2-b09e-4f5a-8843-044468b627a2 · inbound

Training Language Models to Self-Correct via Reinforcement Learning cites this paper.

Training Language Models to Self-Correct via Reinforcement Learning Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-17T12:04:10.514436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-17T12:04:10.210508Z digest=sha256:45bbbae2c5f5a972eaa1d8cb229cb7bc7d5428041bed517dc905bb7b6b88b016

Observation eb350d8c-a9ab-4afd-9829-fd1f3060fba9 · inbound

Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning cites this paper.

Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-21T01:42:19.127139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T01:42:19.004468Z digest=sha256:17cb3a06549b75d8eee8b2a61db2d73b116acb4297d7925631c73bde16d94d10

Observation 599203ec-1368-4c74-ac60-8e2e8612a6ba · inbound

Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision cites this paper.

Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T13:02:57.278896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:02:57.278896Z digest=sha256:8a6352c6c040e0265dd968bf191b9555773a682932ba946d39c6ad24f98db5b7

Observation 671dc132-9b2f-4109-9e30-31bbe50d3824 · inbound

Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models cites this paper.

Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T12:45:20.080762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:45:20.080762Z digest=sha256:a350e93bc81cf418407e9af6f2ef7295ca640b47e32e93626039f561e6f122e6

Observation bd27e811-fe89-4086-94dc-46179857da5f · inbound

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models cites this paper.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.701530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.701530Z digest=sha256:f5067f94ff6d06fa2526bc49d80624a6a508fbcfda72f7741457bd5ff92dad76

Observation 6725648c-4b04-4e3c-a9ad-83294051fed6 · inbound

Surveying the Effects of Quality, Diversity, and Complexity in Synthetic Data From Large Language Models cites this paper.

Surveying the Effects of Quality, Diversity, and Complexity in Synthetic Data From Large Language Models Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 179

Resolution
unresolved
no resolver link, observed 2026-08-11T22:57:01.993575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:57:01.993575Z digest=sha256:a407905bfae8e0d9ce8acbe97fae0752fe13f3f2fd8cd5d810b2111bb59e8b81

Observation 36080728-174d-4ffb-b838-7c44684f950c · inbound

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods cites this paper.

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 206

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:08:34.726256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T23:08:34.312466Z digest=sha256:3ae64767139f3857aa7fe6d89c74cd6aade9fee6a0c8a9601d8686a206d87fe5

Observation dfb035d9-6bd0-4e64-941e-10e63facb169 · inbound

From Multimodal LLMs to Generalist Embodied Agents: Methods and Lessons cites this paper.

From Multimodal LLMs to Generalist Embodied Agents: Methods and Lessons Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-11T17:55:13.556210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:55:13.556210Z digest=sha256:bd40cd6e0ec58abd4913d7644e5c8e8ce83e16ea955b9610b391ac79d192aa96

Observation 784b433b-36f8-4e23-9228-587c27d7f4a1 · inbound

HARP: A challenging human-annotated math reasoning benchmark cites this paper.

HARP: A challenging human-annotated math reasoning benchmark Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T17:34:19.630654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:34:19.630654Z digest=sha256:43ed94853c0b604024660ff0b9545c8e55e728fa512df53e5c616ef9a4d0d611

Observation 6db31b5e-4b51-4255-bb88-9c30d78efb5a · inbound

How to Synthesize Text Data without Model Collapse? cites this paper.

How to Synthesize Text Data without Model Collapse? Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T12:05:48.535997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:05:48.535997Z digest=sha256:a7356efd64b0fd788575814e38983cfd8321c5478672cc6b6225ffbd7191ca7f

Observation 265d8a8a-7cc2-41c6-a0ca-0744a93834d0 · inbound

Offline Reinforcement Learning for LLM Multi-Step Reasoning cites this paper.

Offline Reinforcement Learning for LLM Multi-Step Reasoning Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T10:50:19.089699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:50:19.089699Z digest=sha256:72710b4811fdd7972175c97a2b5d6bcf8659380b4e569136a3bed412c1a3a8b8

Observation 9f49c1a7-e212-4bc3-a9ba-bc786c6bdd6b · inbound

B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners cites this paper.

B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 2006

Resolution
unresolved
no resolver link, observed 2026-08-11T05:45:00.426531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:45:00.426531Z digest=sha256:640462ad085a313dbe48b9edaafa4ca1c702c2adb93d765e58af85523963501a

Observation dff38718-995e-490a-abe9-9cc6bbafe7eb · inbound

Diving into Self-Evolving Training for Multimodal Reasoning cites this paper.

Diving into Self-Evolving Training for Multimodal Reasoning Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:38.975103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:38.975103Z digest=sha256:309c40b5d4b3c3e698620b185fcc97953fa054315eb66df2d317f0f7d13d03fb

Observation a59cf79e-a084-41e7-bdb3-7e90ec44fad5 · inbound

Can Large Language Models Improve SE Active Learning via Warm-Starts? cites this paper.

Can Large Language Models Improve SE Active Learning via Warm-Starts? Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-10T23:05:01.531805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:05:01.531805Z digest=sha256:e4627e707b4cd87e85fdcc15a95f6c43267faa6c4b4befa5a666ac32d7793275

Observation 62c545dd-1679-48f5-8457-0d0f6884a179 · inbound

ReARTeR: Retrieval-Augmented Reasoning with Trustworthy Process Rewarding cites this paper.

ReARTeR: Retrieval-Augmented Reasoning with Trustworthy Process Rewarding Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:59.518256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:59.518256Z digest=sha256:9d5a6f2696066f9aa78e519f3f9e15db8b12f4e100eeca66b929d220ada7ef76

Observation 5e346fb3-ea88-4136-a0bd-213e91797571 · inbound

RedStar: Does Scaling Long-CoT Data Unlock Better Slow-Reasoning Systems? cites this paper.

RedStar: Does Scaling Long-CoT Data Unlock Better Slow-Reasoning Systems? Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T18:30:56.604968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:30:56.604968Z digest=sha256:58b93e4b7248622e0f8481fcc3d9d335afe8cf47351419a91ccf1f7e7a0cc072

Observation faeffd8c-00fd-4ec4-83ba-5d5f2b921b3e · inbound

BRiTE: Bootstrapping Reinforced Thinking Process to Enhance Language Model Reasoning cites this paper.

BRiTE: Bootstrapping Reinforced Thinking Process to Enhance Language Model Reasoning Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-09T22:20:13.811949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:20:13.811949Z digest=sha256:65e0ad106aa1da669e32d557304450a39662b990ba6938c365988b7e91f53cc9

Observation ff08f82e-cf01-4bb5-b43e-b7374402512b · inbound

To Code or not to Code? Adaptive Tool Integration for Math Language Models via Expectation-Maximization cites this paper.

To Code or not to Code? Adaptive Tool Integration for Math Language Models via Expectation-Maximization Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-09T18:07:52.191287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:07:52.191287Z digest=sha256:5b7008b1156757d4a851bed3f79e0bde0151ae16c231bbd76fc787d005df3201

Observation ac222a59-b313-4bcc-bae0-bc83508d9445 · inbound

Self-Improving Transformers Overcome Easy-to-Hard and Length Generalization Challenges cites this paper.

Self-Improving Transformers Overcome Easy-to-Hard and Length Generalization Challenges Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-09T14:54:29.266882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:54:29.266882Z digest=sha256:411a18f2d02fbf329d449db3d192f4f021df5bb6b0c2ace80b54cd1a889b6c97

Observation 9d71f355-0c72-46e5-90a0-db7e7604bc31 · inbound

QLASS: Boosting Language Agent Inference via Q-Guided Stepwise Search cites this paper.

QLASS: Boosting Language Agent Inference via Q-Guided Stepwise Search Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-09T11:47:31.305500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:47:31.305500Z digest=sha256:eb492b4088d159409203c31e417dc739b384a4bf35ff7fd7e87e0bf1a9bff1ee

Observation b96d62a6-a0d1-4148-bbf0-1a7fce3401f7 · inbound

Demystifying Long Chain-of-Thought Reasoning in LLMs cites this paper.

Demystifying Long Chain-of-Thought Reasoning in LLMs Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T01:29:59.416718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T01:29:59.353698Z digest=sha256:43af36feb7bd4d5953adf5df2fa9fa39745de623d2577e5a8cd018ca81f65724

Observation 2957bce5-2342-4b0d-b755-bf63fd216b50 · inbound

LLMs can be easily Confused by Instructional Distractions cites this paper.

LLMs can be easily Confused by Instructional Distractions Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-09T10:50:12.483673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:50:12.483673Z digest=sha256:7de18d07f43e5aa786a8fe5da34351b40d8f482d3f3d3d86cabe4688b3dccbc3

Observation 2c04f306-b46e-4265-ba4b-088a53fde1bc · inbound

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning cites this paper.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:51.152051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:51.152051Z digest=sha256:4696360f02c85d884c43c76802f1cf1603f454308573e6e6746e6da0ca39f89f

Observation 984f8bfb-1d13-4390-9e06-041949791329 · inbound

Policy Guided Tree Search for Enhanced LLM Reasoning cites this paper.

Policy Guided Tree Search for Enhanced LLM Reasoning Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-09T11:20:31.874537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:20:31.874537Z digest=sha256:e8d8ba36b8d9149a2355ec35858fa7699c3201fb206463885df98c20b190df6e

Observation f91fb950-3512-4d62-9f44-2f9d72aaa330 · inbound

From System 1 to System 2: A Survey of Reasoning Large Language Models cites this paper.

From System 1 to System 2: A Survey of Reasoning Large Language Models Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 192

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T01:36:24.273524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-13T01:36:23.845366Z digest=sha256:a2e2b2a9f0bdd589b689d704f470cf83461f6a983876eacd611a6e2bb83980f5

Observation e247596f-c4d8-45a6-9b26-2b9f341b0783 · inbound

OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles cites this paper.

OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-19T06:59:03.257668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T06:59:03.112252Z digest=sha256:6403799a822e6c7e462c099919fa8120a6bee0917f4085f2c7801e01a75c26b3

Observation 2bcb275e-3172-40ef-9dad-2bf894550a60 · inbound

Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe Systems cites this paper.

Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe Systems Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 159

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T21:42:10.791523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-22T21:39:49.832151Z digest=sha256:ce0313943d08db24253a9a982e0ee6e621636b4509fa795c0de26b63e3fb4743

Observation d7a45e69-b45a-45a8-98ac-eb70033b135c · inbound

Think, Prune, Train, Improve: Scaling Reasoning without Scaling Models cites this paper.

Think, Prune, Train, Improve: Scaling Reasoning without Scaling Models Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T10:29:18.697529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:29:18.697529Z digest=sha256:7398d9f28e81e0d7927f44d96ae4edb9da9734ddca5ef6e4c923ad6e082443fd

Observation c0da9749-2e12-439c-b471-e5f0f21f102a · inbound

Learning to Plan Before Answering: Self-Teaching LLMs to Learn Abstract Plans for Problem Solving cites this paper.

Learning to Plan Before Answering: Self-Teaching LLMs to Learn Abstract Plans for Problem Solving Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:51.778603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:51.778603Z digest=sha256:36f739ed7d25dd0c04e19c3059c6dd063fceaf3e6cc11558560739a5954a3457

Observation 0e8eca2d-c256-4575-8604-84bef9de3477 · inbound

Optimizing Chain-of-Thought Reasoners via Gradient Variance Minimization in Rejection Sampling and RL cites this paper.

Optimizing Chain-of-Thought Reasoners via Gradient Variance Minimization in Rejection Sampling and RL Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T00:59:21.839510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:59:21.839510Z digest=sha256:3b671b86344820b95644c147c97533af0486cc53872ba5c9bf85faaff352de79

Observation c867417d-8074-48fa-98a7-66b68caf33a9 · inbound

ToolACE-DEV: Self-Improving Tool Learning via Decomposition and EVolution cites this paper.

ToolACE-DEV: Self-Improving Tool Learning via Decomposition and EVolution Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T22:19:24.875235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:19:24.875235Z digest=sha256:220c5ace9f1edb4e5585c2c7b7e26deab2fbc4362ad55b3cc10aa5fc1afebd10

Observation 1b0d8153-130a-4b51-ad38-58be24bf47bb · inbound

Generalizing Large Language Model Usability Across Resource-Constrained cites this paper.

Generalizing Large Language Model Usability Across Resource-Constrained Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 116

Resolution
unresolved
no resolver link, observed 2026-08-15T22:08:55.035502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:08:55.035502Z digest=sha256:2631d618014fe34e5e48c41c1728ff5afb6400d8e96736a59af5072b64c9795c

Observation d1f3863e-7606-483e-a643-b3fc8dfbc7fe · inbound

SLearnLLM: A Self-Learning Framework for Efficient Domain-Specific Adaptation of Large Language Models cites this paper.

SLearnLLM: A Self-Learning Framework for Efficient Domain-Specific Adaptation of Large Language Models Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:50:05.227428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:50:05.227428Z digest=sha256:d927eae4d69263e8d2e6f3901a67c37f14b178078c4a6b99d499543ed602dbb4

Observation 2b550a2a-2d34-4a83-a7ac-7634953fc29d · inbound

Large Language Models for Planning: A Comprehensive and Systematic Survey cites this paper.

Large Language Models for Planning: A Comprehensive and Systematic Survey Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 215

Resolution
unresolved
no resolver link, observed 2026-08-07T14:12:05.364648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:12:05.364648Z digest=sha256:7040fa8f716577f2f9091a13b8f4c6cbbff7f9d89b1f4b10130b05b8d6ae1bf7

Observation e6128fac-ad80-4c3e-8161-404773e5b32b · inbound

Maximizing Confidence Alone Improves Reasoning cites this paper.

Maximizing Confidence Alone Improves Reasoning Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T13:07:50.094799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:07:50.094799Z digest=sha256:61c31be8df1bd15fb024a827c0ce385658dda12c1f721ae32b38c0b0b322a2b4

Observation 0d04703e-ce4e-42be-a82d-05ff6f566335 · inbound

HardTests: Synthesizing High-Quality Test Cases for LLM Coding cites this paper.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:57.636759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:57.636759Z digest=sha256:b45206d843aad02280eb257ce4fc351d999690274327116b7ce46c910501aff2

Observation e5e4dfcb-47a4-4f45-a5dd-578774dff898 · inbound

SPARQ: Synthetic Problem Generation for Reasoning via Quality-Diversity Algorithms cites this paper.

SPARQ: Synthetic Problem Generation for Reasoning via Quality-Diversity Algorithms Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T05:59:04.390203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:59:04.390203Z digest=sha256:9d6925736e234e4bb8fe51e142487ea2a51542bb861e0a6968ad78e74a83be2f

Observation 607c30cf-eaaf-4c8f-8c72-1a6015fcdd00 · inbound

CIIR@LiveRAG 2025: Optimizing Multi-Agent Retrieval Augmented Generation through Self-Training cites this paper.

CIIR@LiveRAG 2025: Optimizing Multi-Agent Retrieval Augmented Generation through Self-Training Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:18.850888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:18.850888Z digest=sha256:5520bff7bca8f2c087bdd74ea308dbbdbdae6d5ed69ae21aba232ce00cd86e03

Observation 500a8237-3561-48aa-9da6-c5369e86188f · inbound

Spectra 1.1: Scaling Laws and Efficient Inference for Ternary Language Models cites this paper.

Spectra 1.1: Scaling Laws and Efficient Inference for Ternary Language Models Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T21:58:35.473834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:58:35.473834Z digest=sha256:cdb1eeb04cbf475093712ccaa23edc3505f7d36231f35ba5050f601066a23ce7

Observation 60399d37-3763-4d34-9780-b5231395c56f · inbound

SyncLoop: A Multimodal Dual-Loop Framework for Self-Improving Mathematical Reasoning cites this paper.

SyncLoop: A Multimodal Dual-Loop Framework for Self-Improving Mathematical Reasoning Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T15:12:01.966932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:12:01.966932Z digest=sha256:973ff176b8706878a4ad292eec5fb742282e4a3cbe940b1da6e6d9c1dd2935e9

Observation 71059b91-6525-46b3-98d2-7379c09ed54a · inbound

Utilizing Training Data to Improve LLM Reasoning for Tabular Understanding cites this paper.

Utilizing Training Data to Improve LLM Reasoning for Tabular Understanding Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T16:22:50.862355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:22:50.862355Z digest=sha256:67bfa80c8d8f451bb5bc16866f21bc0aff70ee976a1cbccc42445bfffa2c306b

Observation 87b836da-7df4-457d-a007-3dbb733fc341 · inbound

ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding cites this paper.

ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T15:45:24.746306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:45:24.746306Z digest=sha256:7e06bb029208b4b2fafecced32198bcf74d33f6269706b495a6d396945747624

Observation 804b351d-d527-4650-a4b3-d6ec3000e6d0 · inbound

Learn from What We HAVE: History-Aware VErifier that Reasons about Past Interactions Online cites this paper.

Learn from What We HAVE: History-Aware VErifier that Reasons about Past Interactions Online Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T13:52:52.128802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:52:52.128802Z digest=sha256:2362ad5f090d4bbb6d0a7926df1716b6e1ad2ca983bd9c0de5f433d44f0f6827

Observation 1ffe2ae4-5226-4f86-b0fb-2daf5a0086f6 · inbound

Bridging the Capability Gap: Joint Alignment Tuning for Harmonizing LLM-based Multi-Agent Systems cites this paper.

Bridging the Capability Gap: Joint Alignment Tuning for Harmonizing LLM-based Multi-Agent Systems Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T18:50:10.554445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T18:50:10.554445Z digest=sha256:af678ef8bec6996f24495283bd3bc4f3e4569a24b4865731fdf88ae93278151d

Observation 890e5f64-f821-448e-8cc0-bf83393f31d0 · inbound

GPO: Learning from Critical Steps to Improve LLM Reasoning cites this paper.

GPO: Learning from Critical Steps to Improve LLM Reasoning Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T15:55:56.175915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:55:56.175915Z digest=sha256:99351000d2181f78ee8cf89b6abd1934ce45929d04cbd03048494c950f938b1c

Observation e5abc5f4-acb5-49b1-a537-180ae38e0bbb · inbound

rePIRL: Learn PRM with Inverse RL for LLM Reasoning cites this paper.

rePIRL: Learn PRM with Inverse RL for LLM Reasoning Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T13:14:10.981671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T13:13:13.293921Z digest=sha256:43db16aba9bcfc42ca9312955d667d3bdbb04bef6a5b2d8b058881cb62cd689f

Observation 790c996a-e57a-453a-a2cc-5e7228b8c615 · inbound

rePIRL: Learn PRM with Inverse RL for LLM Reasoning cites this paper.

rePIRL: Learn PRM with Inverse RL for LLM Reasoning Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-03T03:33:44.696113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:33:44.696113Z digest=sha256:accf7febdbdeeec149e39b7167c4fcfcf3fca27ac310c87fc9cd06045fafef45

Observation 5f340510-3d13-47fd-b388-cdd51ecdeb84 · inbound

A Task-Centric Theory for Iterative Self-Improvement with Easy-to-Hard Curricula cites this paper.

A Task-Centric Theory for Iterative Self-Improvement with Easy-to-Hard Curricula Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T02:43:36.662189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:43:36.662189Z digest=sha256:f1139005324b958492a181b8eb0574bdf90043d01a78f5da619569dffdc58305

Observation 8c37dfd8-83f4-4fe2-b809-033c678ab898 · inbound

Turbo Connection: Reasoning as Information Flow from Higher to Lower Layers cites this paper.

Turbo Connection: Reasoning as Information Flow from Higher to Lower Layers Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T22:08:42.714880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:08:42.714880Z digest=sha256:b1d08101816457dc2513610bf505ee804901e2cd8dffb5907d7f85e88d42e3d1

Observation 924abd59-3f5f-43c1-a9ff-3d048b97ebd4 · inbound

Space Syntax-guided Post-training for Residential Floor Plan Generation cites this paper.

Space Syntax-guided Post-training for Residential Floor Plan Generation Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T19:40:17.733380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T19:38:24.134463Z digest=sha256:5802f8c804cfdf2ea5dafde76e0726b9ddba6d5588faffd08bd4970a6ccdddcf

Observation c391cbcd-6896-4f3d-856f-c02b0e90f75a · inbound

PRISMA: Preference-Reinforced Self-Training Approach for Interpretable Emotionally Intelligent Negotiation Dialogues cites this paper.

PRISMA: Preference-Reinforced Self-Training Approach for Interpretable Emotionally Intelligent Negotiation Dialogues Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-10T10:14:10.818457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T05:00:26.238268Z digest=sha256:a5c31f6215daf8dea39c70718def6d665856189ba2003fa3b8d0db5f7b23a3ae

Observation 734bac73-6aee-4772-a492-e877d7328eae · inbound

$S^3$-R1: Learning to Retrieve and Answer Step-by-Step with Synthetic Data cites this paper.

$S^3$-R1: Learning to Retrieve and Answer Step-by-Step with Synthetic Data Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:46:06.930542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-09T15:08:53.731480Z digest=sha256:36083bf3fa85e7ad289a7e3e46a3361998709530add9d357d7ead7f94afa210d

Observation af2a7b0e-a535-4cb4-814f-64525ee14ee4 · inbound

$S^3$-R1: Learning to Retrieve and Answer Step-by-Step with Synthetic Data cites this paper.

$S^3$-R1: Learning to Retrieve and Answer Step-by-Step with Synthetic Data Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-01T00:55:12.106795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T00:48:54.797750Z digest=sha256:64a56b13a31f595f7cb6539c30a8ac36d820d7453a94ce556c3cd476a7b99cc7

Observation 83deb363-64c7-4970-958d-fc6bd40d1a98 · inbound

Segment-Aligned Policy Optimization for Multi-Modal Reasoning cites this paper.

Segment-Aligned Policy Optimization for Multi-Modal Reasoning Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T16:51:07.861678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-09T14:44:31.160543Z digest=sha256:8f2277348de8004e429259df8d858b97c55d8711bb239a8f3cf15ee8a2c298f0

Observation b8c432ad-a1c4-4228-8f84-5e9a210351c7 · inbound

Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients cites this paper.

Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:16:08.428351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-08T09:53:32.077464Z digest=sha256:0497f47e61011a01115aafbe19dc36066d9f939504b075d069a50086363ee5d1

Observation 4df2d3ef-ea52-4954-ac9b-091668ef48f3 · inbound

SOLAR: A Self-Optimizing Open-Ended Autonomous Agent for Lifelong Learning and Continual Adaptation cites this paper.

SOLAR: A Self-Optimizing Open-Ended Autonomous Agent for Lifelong Learning and Continual Adaptation Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-21T11:24:08.738134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T11:21:30.867480Z digest=sha256:be8e04492219d1d7fa5df79cd0ef795e9ef4e3bf29c3ef669ee3312d6648923b

Observation 2d8e60a0-23ff-491b-888b-6e971a037740 · inbound

RISE: Reliable Improvement in Self-Evolving Vision-Language Models cites this paper.

RISE: Reliable Improvement in Self-Evolving Vision-Language Models Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:39:40.598262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T05:38:26.590720Z digest=sha256:22b32791c261179e90d20f52be81e8af3a2ef3fe7a444df06c237b27a4455760

Observation 7fbeceff-a5d9-481c-975e-d297f43c2bed · inbound

RISE: Reliable Improvement in Self-Evolving Vision-Language Models cites this paper.

RISE: Reliable Improvement in Self-Evolving Vision-Language Models Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:24:57.549368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-30T17:19:09.596950Z digest=sha256:b515f9ba3125d546c27ecb4916c64f080de5be4bc678bfb978a24d5a7d0c28c5

Observation 66f7a843-37d7-49f4-bd12-4b296193d4ae · inbound

Self-Policy Distillation via Capability-Selective Subspace Projection cites this paper.

Self-Policy Distillation via Capability-Selective Subspace Projection Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-22T05:34:40.416173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-22T05:31:42.803465Z digest=sha256:2fb0870914c853e986246942837318638d07b4859b4491895b80f4a4805be5b3

Observation 5eb8abed-a77d-444c-800a-cad6063f909a · inbound

Peak-Then-Collapse and the Four Interface Channels of Knowledge-Graph Tool Use cites this paper.

Peak-Then-Collapse and the Four Interface Channels of Knowledge-Graph Tool Use Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T21:23:58.942626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T21:20:16.027196Z digest=sha256:52bae04cc2c8aa8a90d463f40622f9fdfc0be24c8d1c91964132b20779260023

Observation 162d4ffc-93b7-43b0-9ec3-d2daf9973525 · inbound

DRIFT: Decoupled Rollouts and Importance-Weighted Fine-Tuning for Efficient Multi-Turn Optimization cites this paper.

DRIFT: Decoupled Rollouts and Importance-Weighted Fine-Tuning for Efficient Multi-Turn Optimization Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T00:02:50.297157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T23:16:49.358792Z digest=sha256:a335fcd3317547479c3b759d14609cd4fa4ffbcbdedfdbdebd1b3bea6b34c18e

Observation b1ad4b1d-ebb2-4778-83a0-a03e9fe21758 · inbound

Robust Reasoning via Dynamic Token Selection for Distribution-Aligned Self-Distillation cites this paper.

Robust Reasoning via Dynamic Token Selection for Distribution-Aligned Self-Distillation Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:22:35.025991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-28T19:12:39.939329Z digest=sha256:1c58d579f3f445b5fab8493c965bfcb5d66e6ca983af79c0dec27e9242f73653

Observation 71a6711c-6d6e-4d83-8eba-186cf78d6b44 · inbound

Step-by-Step Optimization-like Reasoning in LLMs over Expanding Search Spaces cites this paper.

Step-by-Step Optimization-like Reasoning in LLMs over Expanding Search Spaces Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T08:46:48.965970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T05:46:26.938277Z digest=sha256:33f2a446a4b570a8914f52749313d9ea3b2a77092ca70ab496e73e6f282664bc

Observation 667f03c4-2391-4a3c-bcd0-3a76d4a86077 · inbound

Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients cites this paper.

Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 77

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T20:48:56.134444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-27T01:08:52.981296Z digest=sha256:1ada004c408bb145235f9011818c387dd19f1fb05a7a3ef4f57f2963b0c81aae

Observation e3ccd73e-cf5c-42c6-8855-488529460e81 · inbound

DRIFT: Refining Instruction Data via On-Policy Data Attribution cites this paper.

DRIFT: Refining Instruction Data via On-Policy Data Attribution Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-03T19:18:54.719160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-27T01:57:05.784589Z digest=sha256:23eb0ebc3d81ccef19bf93dbfc9b0d47023ab4209aaf608e2cf7db955b09115e

Observation f8137536-2793-4e73-828e-8f691d7d3b6d · inbound

Data Selection Through Iterative Self-Filtering for Vision-Language Settings cites this paper.

Data Selection Through Iterative Self-Filtering for Vision-Language Settings Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:49:44.969512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-26T09:22:47.537137Z digest=sha256:d9c98df63515a684e14448dd3784937ae471094a9ae3105c02e4a0943f1b73e1

Observation dcf81913-45a8-4a5f-bf89-88c9345346a9 · inbound

On-Policy Self-Distillation with Sampled Demonstrations Reduces Output Diversity cites this paper.

On-Policy Self-Distillation with Sampled Demonstrations Reduces Output Diversity Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-07-04T21:00:08.422731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-25T19:23:56.452083Z digest=sha256:cf69bfe11d32aabd4c036e7288e43baf23d761fdb2b50092aecf00c2b1d2ad75

Observation c88cb04e-be10-40ae-bda7-3e86ffdc71c1 · inbound

Paying More Attention to Visual Tokens in Self-Evolving Large Multimodal Models cites this paper.

Paying More Attention to Visual Tokens in Self-Evolving Large Multimodal Models Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T13:39:50.967732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-26T05:01:47.207296Z digest=sha256:3a4267ac8a8e677de3666562f65d6a733d8a2b1d3e0b3ddcb46174e125ca1ba8

Observation 6cb802be-2f28-43ef-a056-6c9296fdc956 · inbound

PHF: Privileged Hidden Flow for On-Policy Self-Distillation cites this paper.

PHF: Privileged Hidden Flow for On-Policy Self-Distillation Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-06-30T07:34:21.811633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-30T07:26:32.902086Z digest=sha256:7df88511b50c9997274e25a98cf860b6535986cb63a50da5e4121c37dcec4a06

Observation bd21824f-69a1-4582-8962-2589159d7811 · inbound

Active-GRPO: Adaptive Imitation and Self-Improving Reasoning for Molecular Optimization cites this paper.

Active-GRPO: Adaptive Imitation and Self-Improving Reasoning for Molecular Optimization Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-07-02T16:17:08.604206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-02T16:07:49.362916Z digest=sha256:940a756fac35f043181edd8e0feaf57dc56dc6150dd925d35d0ea459f7e178c0

Observation 90ab09a6-18ef-4af7-bb5b-b663d1979ec4 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 205

Resolution
unresolved
no resolver link, observed 2026-07-11T13:53:36.775836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T13:53:36.775836Z digest=sha256:3a021538c9e0ae613420397d5c2c71c73f7314396b09015ecb5dd8a0744bb20c

Observation a5c001a8-f646-4831-b290-bdd4ceadd008 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 206

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:56.405118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:56.405118Z digest=sha256:86d27cf17599c73423217cfe8a6f2a77be1d049cc617efc174b5626866637428

Observation 2d3c46f9-bc77-4557-bd9f-699103e384a7 · inbound

The Verifier is the Curriculum: Execution-Gated Self-Distillation for Cross-Family Game Generation cites this paper.

The Verifier is the Curriculum: Execution-Gated Self-Distillation for Cross-Family Game Generation Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-14T17:25:37.713857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T17:25:37.713857Z digest=sha256:4dea1c527d8729fc976c472d5f4938167b56c728f4da6772a6b054235abb8ebe

Observation 13dd479b-e987-4d72-adad-948b0c3a7dfb · inbound

UNIBROWSE: A Data-to-Agent Framework for Multimodal BrowseComp cites this paper.

UNIBROWSE: A Data-to-Agent Framework for Multimodal BrowseComp Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-14T10:51:16.019022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T10:51:16.019022Z digest=sha256:1707989de1c4d8d4f72ab70c01515a82a186d93e905b49c4ed3c37210e2aaaa4

Observation e0ff153d-3cab-4e56-aead-9b7e7bd48a79 · inbound

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration cites this paper.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 106

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:26.701252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:26.701252Z digest=sha256:da83a4df0cb308d88392f331196074d0efec526ac29c39c6f9925bea2c3cd4ce

Observation d1620c7f-8ad5-43ff-bc27-7b2489a9d2b6 · inbound

Answer-Conditioned Chains of Thought Degrade Verifiable-Reasoning Distillation in Large Language Models cites this paper.

Answer-Conditioned Chains of Thought Degrade Verifiable-Reasoning Distillation in Large Language Models Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T01:52:28.936526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:52:28.936526Z digest=sha256:44ef27596f97f734f5bb482e8370b497de023befeda77f4999efb650245c38d4

Observation f75ed8b3-e8bc-4b2e-9605-2afc67667824 · inbound

CoTu at EXACT 2026: Neuro-Symbolic Reasoning for Transparent Educational QA cites this paper.

CoTu at EXACT 2026: Neuro-Symbolic Reasoning for Transparent Educational QA Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T01:15:33.669821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:15:33.669821Z digest=sha256:fb333089eba9ff77e58e0029a94e66f2fa2f59cecf1c99dc3ac08f188a28bd71

Observation 8601a47d-6b23-4ca9-8ba9-1fc876d16795 · inbound

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information cites this paper.

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-01T12:54:34.843879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:54:34.843879Z digest=sha256:296527ea110d33fc1eb7460ae58efdc70a50c51aa296a7141c27562d971b6535

Observation 0be574e1-6fbf-4c6b-96a8-61d630984dcf · inbound

FormulaSPIN: Self-Play Fine-Tuning for Natural Language to Spreadsheet Formula Generation cites this paper.

FormulaSPIN: Self-Play Fine-Tuning for Natural Language to Spreadsheet Formula Generation Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T13:32:13.493804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T13:32:13.493804Z digest=sha256:72b6a62097b59ac95c4967bb30209757bb014db5e32a8a6af5307a5a21ab27fa

Observation 05e6d4e8-23ee-4cdc-b797-a7c0281ffebe · inbound

Bridging Compute- and Data-Optimal Pretraining cites this paper.

Bridging Compute- and Data-Optimal Pretraining Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 171

Resolution
unresolved
no resolver link, observed 2026-08-01T03:02:08.797292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:02:08.797292Z digest=sha256:56b2339610b4bf591c7c8dfd82d8b5d1233b8de4316e7df73fcee8618419bcc2

Observation 8f08c263-545e-4637-b701-934097b26025 · inbound

From Scoring to Acting: Outcome-Verified Comparative Self-Distillation for LLM Agents cites this paper.

From Scoring to Acting: Outcome-Verified Comparative Self-Distillation for LLM Agents Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-31T22:51:49.392233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T22:51:49.392233Z digest=sha256:256f21ce9b335b00b1f45d80e5a9074425843e5e7417f7c78d2a226c38155fec

Observation c2689870-d8ac-4522-a585-d2f1e16e5a9b · inbound

Recursive Synthesis for Long-Horizon Terminal Tasks cites this paper.

Recursive Synthesis for Long-Horizon Terminal Tasks Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T12:35:01.459919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:35:01.459919Z digest=sha256:2f6e44e944be9c61a238cfe5fd77a4b6e625c1674a30d6cbeefb4f79c6b16fa2

Observation 5e6c19ce-8f43-4cca-8b91-8acaebf75c24 · inbound

Recursive Synthesis for Long-Horizon Terminal Tasks cites this paper.

Recursive Synthesis for Long-Horizon Terminal Tasks Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T04:31:42.362344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:31:42.362344Z digest=sha256:d42e2dd3174d7735c7ae6821e39e0150040111b9fee42a3a61cf86795a7049b7

Observation c35085c4-e018-450c-96f8-9517ab8c39cb · inbound

ThinkRetrieve: Retrieval-Augmented Reasoning Traces for Test-Time Scaling cites this paper.

ThinkRetrieve: Retrieval-Augmented Reasoning Traces for Test-Time Scaling Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 147

Resolution
unresolved
no resolver link, observed 2026-08-12T14:10:45.595055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:10:45.595055Z digest=sha256:dad9bce5599eb4b71c191a00007c8edd6c8b6eaf351f2dfc81a3de0c6325474e