Pith. sign in

Paper Citation Record · LEDGER

Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 58 inbound Pith citation observations for arXiv:2312.06585.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.06585 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 58 of 58 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T14:27:51.152051Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T21:00:08.414065Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ab4126f8-e276-4acd-b497-4213e1f02631 · inbound

An Iterative Utility Judgment Framework Inspired by Philosophical Relevance via LLMs cites this paper.

An Iterative Utility Judgment Framework Inspired by Philosophical Relevance via LLMs Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-24T00:15:52.765918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-24T00:14:42.779546Z digest=sha256:f0d58b07129a23e54c507413dc8b7771b70a9a9e1cede1ca270501a040da11f4

Observation 84baa3d2-b09e-4f5a-8843-044468b627a2 · inbound

Training Language Models to Self-Correct via Reinforcement Learning cites this paper.

Training Language Models to Self-Correct via Reinforcement Learning Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-17T12:04:10.514436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-17T12:04:10.210508Z digest=sha256:82645c2997342a0f8fcf3f08535e282ab724c86a74c5c449b1d7d6eb30d9db96

Observation eb350d8c-a9ab-4afd-9829-fd1f3060fba9 · inbound

Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning cites this paper.

Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-21T01:42:19.127139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T01:42:19.004468Z digest=sha256:ba9445bcca6eb8c531d00979e45c6226b1e3ce35e1b2be69bf687c1d9369d6f4

Observation 36080728-174d-4ffb-b838-7c44684f950c · inbound

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods cites this paper.

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 206

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:08:34.726256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T23:08:34.312466Z digest=sha256:7fe7d8a21ce896bab7b266678c40ebb1c4471d45a825877b9e052c71b1b5c929

Observation b96d62a6-a0d1-4148-bbf0-1a7fce3401f7 · inbound

Demystifying Long Chain-of-Thought Reasoning in LLMs cites this paper.

Demystifying Long Chain-of-Thought Reasoning in LLMs Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T01:29:59.416718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T01:29:59.353698Z digest=sha256:c2f7a914dd6c576287f695067755509f6e784195820b047dd2f646160902b076

Observation 2c04f306-b46e-4265-ba4b-088a53fde1bc · inbound

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning cites this paper.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:51.152051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:51.152051Z digest=sha256:34c24bf5e1e1f7672f9a5c740603f0765bdc2a236a878c9e5f127c8c728901ee

Observation f91fb950-3512-4d62-9f44-2f9d72aaa330 · inbound

From System 1 to System 2: A Survey of Reasoning Large Language Models cites this paper.

From System 1 to System 2: A Survey of Reasoning Large Language Models Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 192

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T01:36:24.273524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T01:36:23.845366Z digest=sha256:18324d6fa7318b0379b9e3ca7ac87822f1a535a22ab2ab745c126404da565327

Observation e247596f-c4d8-45a6-9b26-2b9f341b0783 · inbound

OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles cites this paper.

OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-19T06:59:03.257668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T06:59:03.112252Z digest=sha256:9ef3f7a723b24b61a1973acfacfaf4611b6faa7550f2c761fec40cf539683a34

Observation 2bcb275e-3172-40ef-9dad-2bf894550a60 · inbound

Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe Systems cites this paper.

Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe Systems Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 159

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T21:42:10.791523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T21:39:49.832151Z digest=sha256:f9117eda05fe4fde743be81b0820f66f54ca9f33cdb7c845f84934049685823e

Observation d1f3863e-7606-483e-a643-b3fc8dfbc7fe · inbound

SLearnLLM: A Self-Learning Framework for Efficient Domain-Specific Adaptation of Large Language Models cites this paper.

SLearnLLM: A Self-Learning Framework for Efficient Domain-Specific Adaptation of Large Language Models Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:50:05.227428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:50:05.227428Z digest=sha256:ec8e49399f9c1a634cc8f36ad5cd33c80202e01f68eb9f3c7c2678b9db9be126

Observation 2b550a2a-2d34-4a83-a7ac-7634953fc29d · inbound

Large Language Models for Planning: A Comprehensive and Systematic Survey cites this paper.

Large Language Models for Planning: A Comprehensive and Systematic Survey Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 215

Resolution
unresolved
no resolver link, observed 2026-08-07T14:12:05.364648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:12:05.364648Z digest=sha256:fe98cd68c0aaabf5bd610d9d6cba066c67862aa2d1a5ae0db60c2585b4eaa931

Observation e6128fac-ad80-4c3e-8161-404773e5b32b · inbound

Maximizing Confidence Alone Improves Reasoning cites this paper.

Maximizing Confidence Alone Improves Reasoning Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T13:07:50.094799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:07:50.094799Z digest=sha256:1d42a340e0cc09c4391c806e097cbd094a1060452e4df00ffd9b7f43896fd2ea

Observation 0d04703e-ce4e-42be-a82d-05ff6f566335 · inbound

HardTests: Synthesizing High-Quality Test Cases for LLM Coding cites this paper.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:57.636759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:57.636759Z digest=sha256:e17d7b67dcf10188f4d942de8cb0d2e5cda21fbd20b4bd2d79e8ba4a718ad07d

Observation e5e4dfcb-47a4-4f45-a5dd-578774dff898 · inbound

SPARQ: Synthetic Problem Generation for Reasoning via Quality-Diversity Algorithms cites this paper.

SPARQ: Synthetic Problem Generation for Reasoning via Quality-Diversity Algorithms Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T05:59:04.390203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:59:04.390203Z digest=sha256:5d36dc0bf9c5fdc3c2b213d5547ad38ce238b38eca190317c6c8a8f4aaf3d3c4

Observation 607c30cf-eaaf-4c8f-8c72-1a6015fcdd00 · inbound

CIIR@LiveRAG 2025: Optimizing Multi-Agent Retrieval Augmented Generation through Self-Training cites this paper.

CIIR@LiveRAG 2025: Optimizing Multi-Agent Retrieval Augmented Generation through Self-Training Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T04:21:18.850888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:21:18.850888Z digest=sha256:dd41d7aaba4bc7b546ed17ac2f3fa5a81bbab80d4d7857961c30c365727c60f0

Observation 500a8237-3561-48aa-9da6-c5369e86188f · inbound

Spectra 1.1: Scaling Laws and Efficient Inference for Ternary Language Models cites this paper.

Spectra 1.1: Scaling Laws and Efficient Inference for Ternary Language Models Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T21:58:35.473834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:58:35.473834Z digest=sha256:3f9f91f09f96362798128801f4060afe078fb15ad66a3354e681b1cbff87fe67

Observation 60399d37-3763-4d34-9780-b5231395c56f · inbound

SyncLoop: A Multimodal Dual-Loop Framework for Self-Improving Mathematical Reasoning cites this paper.

SyncLoop: A Multimodal Dual-Loop Framework for Self-Improving Mathematical Reasoning Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T15:12:01.966932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:12:01.966932Z digest=sha256:8b6326a680e4ad98df75fda93dba1a7007fea87eea126a47e16bab353355f684

Observation 71059b91-6525-46b3-98d2-7379c09ed54a · inbound

Utilizing Training Data to Improve LLM Reasoning for Tabular Understanding cites this paper.

Utilizing Training Data to Improve LLM Reasoning for Tabular Understanding Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T16:22:50.862355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:22:50.862355Z digest=sha256:40c06757584a93b5c839bdf838584dbb8de2b12c710a63c446c011eb9d9b2a61

Observation 87b836da-7df4-457d-a007-3dbb733fc341 · inbound

ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding cites this paper.

ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T15:45:24.746306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:45:24.746306Z digest=sha256:9844d2e9030833ea1641a57e6e5bd71752738d0f4aaebb53e15bcfce2e561ebb

Observation 804b351d-d527-4650-a4b3-d6ec3000e6d0 · inbound

Learn from What We HAVE: History-Aware VErifier that Reasons about Past Interactions Online cites this paper.

Learn from What We HAVE: History-Aware VErifier that Reasons about Past Interactions Online Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T13:52:52.128802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:52:52.128802Z digest=sha256:3e28adcbd0f5976cf1bb163abcea12a48b954db11fbf019e2e37b55b4adcfa4a

Observation 1ffe2ae4-5226-4f86-b0fb-2daf5a0086f6 · inbound

Bridging the Capability Gap: Joint Alignment Tuning for Harmonizing LLM-based Multi-Agent Systems cites this paper.

Bridging the Capability Gap: Joint Alignment Tuning for Harmonizing LLM-based Multi-Agent Systems Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T18:50:10.554445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T18:50:10.554445Z digest=sha256:05ae218b639a04c26a5bf6c99d0122c8963098eb90790ddb6bae7312b43d080c

Observation e5abc5f4-acb5-49b1-a537-180ae38e0bbb · inbound

rePIRL: Learn PRM with Inverse RL for LLM Reasoning cites this paper.

rePIRL: Learn PRM with Inverse RL for LLM Reasoning Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T13:14:10.981671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T13:13:13.293921Z digest=sha256:ef45c79a21f10fdd3f2b785612e5e0519dda6e5dba3a3b5a2db7da17bb5fe477

Observation 790c996a-e57a-453a-a2cc-5e7228b8c615 · inbound

rePIRL: Learn PRM with Inverse RL for LLM Reasoning cites this paper.

rePIRL: Learn PRM with Inverse RL for LLM Reasoning Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-03T03:33:44.696113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:33:44.696113Z digest=sha256:8318e575858add4f7c9cc71e2b7896126880417c76e931589e32e0027901d51f

Observation 5f340510-3d13-47fd-b388-cdd51ecdeb84 · inbound

A Task-Centric Theory for Iterative Self-Improvement with Easy-to-Hard Curricula cites this paper.

A Task-Centric Theory for Iterative Self-Improvement with Easy-to-Hard Curricula Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T02:43:36.662189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:43:36.662189Z digest=sha256:a56c7c16c79ba408c1377220e0fa1f3ff675220ca0a2459022c94f655424daa1

Observation 8c37dfd8-83f4-4fe2-b809-033c678ab898 · inbound

Turbo Connection: Reasoning as Information Flow from Higher to Lower Layers cites this paper.

Turbo Connection: Reasoning as Information Flow from Higher to Lower Layers Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T22:08:42.714880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:08:42.714880Z digest=sha256:bba2a49f00d169638e0d063aa233151765451d210b043392f291b8883f1624cf

Observation 924abd59-3f5f-43c1-a9ff-3d048b97ebd4 · inbound

Space Syntax-guided Post-training for Residential Floor Plan Generation cites this paper.

Space Syntax-guided Post-training for Residential Floor Plan Generation Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T19:40:17.733380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T19:38:24.134463Z digest=sha256:987be43d18931b5408d3f1abfd3a8ed48dfce795cd07c947d59f5c2e82c4692a

Observation c391cbcd-6896-4f3d-856f-c02b0e90f75a · inbound

PRISMA: Preference-Reinforced Self-Training Approach for Interpretable Emotionally Intelligent Negotiation Dialogues cites this paper.

PRISMA: Preference-Reinforced Self-Training Approach for Interpretable Emotionally Intelligent Negotiation Dialogues Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-10T10:14:10.818457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T05:00:26.238268Z digest=sha256:a24cd27c5991b33f2a435ebd5b9f17d7bfc92186d654d295e83bb4db55977b17

Observation 734bac73-6aee-4772-a492-e877d7328eae · inbound

$S^3$-R1: Learning to Retrieve and Answer Step-by-Step with Synthetic Data cites this paper.

$S^3$-R1: Learning to Retrieve and Answer Step-by-Step with Synthetic Data Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:46:06.930542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T15:08:53.731480Z digest=sha256:df482be33bf6bc17af6972cb4225901d957379c6ab45719889343d9e88d19c96

Observation af2a7b0e-a535-4cb4-814f-64525ee14ee4 · inbound

$S^3$-R1: Learning to Retrieve and Answer Step-by-Step with Synthetic Data cites this paper.

$S^3$-R1: Learning to Retrieve and Answer Step-by-Step with Synthetic Data Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-01T00:55:12.106795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T00:48:54.797750Z digest=sha256:f84710459a889ade89536613e9ace0e3ca012e97009f67497a2e1144e7159819

Observation 83deb363-64c7-4970-958d-fc6bd40d1a98 · inbound

Segment-Aligned Policy Optimization for Multi-Modal Reasoning cites this paper.

Segment-Aligned Policy Optimization for Multi-Modal Reasoning Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T16:51:07.861678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T14:44:31.160543Z digest=sha256:8ddc8441f9cf2c372763157dde9822d114d9c127e048bfce82dfb16027466c0b

Observation b8c432ad-a1c4-4228-8f84-5e9a210351c7 · inbound

Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients cites this paper.

Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:16:08.428351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-08T09:53:32.077464Z digest=sha256:db9ac370048e93f16f4c47bb2f94115a8ceeed0a3108b769e3c0f5c7a69451f4

Observation 4df2d3ef-ea52-4954-ac9b-091668ef48f3 · inbound

SOLAR: A Self-Optimizing Open-Ended Autonomous Agent for Lifelong Learning and Continual Adaptation cites this paper.

SOLAR: A Self-Optimizing Open-Ended Autonomous Agent for Lifelong Learning and Continual Adaptation Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-21T11:24:08.738134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T11:21:30.867480Z digest=sha256:f77d9996fd713219053d9f19504601ff5b75225cb9e247e8294249c75bcffe4d

Observation 2d8e60a0-23ff-491b-888b-6e971a037740 · inbound

RISE: Reliable Improvement in Self-Evolving Vision-Language Models cites this paper.

RISE: Reliable Improvement in Self-Evolving Vision-Language Models Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:39:40.598262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T05:38:26.590720Z digest=sha256:40b769898a89c55e28e88239235a13bbb5e0b11b7a44d9cf741b5d039b8faddb

Observation 7fbeceff-a5d9-481c-975e-d297f43c2bed · inbound

RISE: Reliable Improvement in Self-Evolving Vision-Language Models cites this paper.

RISE: Reliable Improvement in Self-Evolving Vision-Language Models Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:24:57.549368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T17:19:09.596950Z digest=sha256:e8548fe7a996f2e778beb8675086cbe20bf8487bac6ca498b2d192d1c4be0eec

Observation 66f7a843-37d7-49f4-bd12-4b296193d4ae · inbound

Self-Policy Distillation via Capability-Selective Subspace Projection cites this paper.

Self-Policy Distillation via Capability-Selective Subspace Projection Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-22T05:34:40.416173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T05:31:42.803465Z digest=sha256:cfef224b3fc289a769b2947c6b4fc6035b3456a47f9b90bf974a245fd874edc6

Observation 5eb8abed-a77d-444c-800a-cad6063f909a · inbound

Peak-Then-Collapse and the Four Interface Channels of Knowledge-Graph Tool Use cites this paper.

Peak-Then-Collapse and the Four Interface Channels of Knowledge-Graph Tool Use Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T21:23:58.942626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T21:20:16.027196Z digest=sha256:618e98b312ebe1b70106e010bfeaa356aad7d28da250fa944c6a1060b3662a41

Observation 162d4ffc-93b7-43b0-9ec3-d2daf9973525 · inbound

DRIFT: Decoupled Rollouts and Importance-Weighted Fine-Tuning for Efficient Multi-Turn Optimization cites this paper.

DRIFT: Decoupled Rollouts and Importance-Weighted Fine-Tuning for Efficient Multi-Turn Optimization Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T00:02:50.297157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T23:16:49.358792Z digest=sha256:c06e910038d283bcc60e8b5dbc180adff4f75754c9cf7b5e853287b7e00cfb79

Observation b1ad4b1d-ebb2-4778-83a0-a03e9fe21758 · inbound

Robust Reasoning via Dynamic Token Selection for Distribution-Aligned Self-Distillation cites this paper.

Robust Reasoning via Dynamic Token Selection for Distribution-Aligned Self-Distillation Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:22:35.025991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T19:12:39.939329Z digest=sha256:b075fa5caa48ba29699bab2c77a45a905a90912f166d868b987d4ad4f1671a90

Observation 71a6711c-6d6e-4d83-8eba-186cf78d6b44 · inbound

Step-by-Step Optimization-like Reasoning in LLMs over Expanding Search Spaces cites this paper.

Step-by-Step Optimization-like Reasoning in LLMs over Expanding Search Spaces Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T08:46:48.965970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T05:46:26.938277Z digest=sha256:c76b8e25bd42fbf8c2b732123bbfb244d823db2531344517db808b245896d9a4

Observation 667f03c4-2391-4a3c-bcd0-3a76d4a86077 · inbound

Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients cites this paper.

Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 77

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T20:48:56.134444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T01:08:52.981296Z digest=sha256:25111686a6ef7dc678cb16f6eb98ed724d0d896802764f64ce9496ef39ce688d

Observation e3ccd73e-cf5c-42c6-8855-488529460e81 · inbound

DRIFT: Refining Instruction Data via On-Policy Data Attribution cites this paper.

DRIFT: Refining Instruction Data via On-Policy Data Attribution Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-03T19:18:54.719160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T01:57:05.784589Z digest=sha256:971aa6faeb0234740b0f0d28bf462069c445fdb65c0e4727d03b8d3f18bd867e

Observation f8137536-2793-4e73-828e-8f691d7d3b6d · inbound

Data Selection Through Iterative Self-Filtering for Vision-Language Settings cites this paper.

Data Selection Through Iterative Self-Filtering for Vision-Language Settings Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:49:44.969512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T09:22:47.537137Z digest=sha256:3d6b02d67399eda52d7e7602a95f86037102f4067011555ab6989bb4cf3d7916

Observation dcf81913-45a8-4a5f-bf89-88c9345346a9 · inbound

On-Policy Self-Distillation with Sampled Demonstrations Reduces Output Diversity cites this paper.

On-Policy Self-Distillation with Sampled Demonstrations Reduces Output Diversity Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-07-04T21:00:08.422731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-25T19:23:56.452083Z digest=sha256:6a341659e79a154cb6c62f73d57123144e58f652fe04386e0821a305fe681b71

Observation c88cb04e-be10-40ae-bda7-3e86ffdc71c1 · inbound

Paying More Attention to Visual Tokens in Self-Evolving Large Multimodal Models cites this paper.

Paying More Attention to Visual Tokens in Self-Evolving Large Multimodal Models Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T13:39:50.967732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T05:01:47.207296Z digest=sha256:a106a4e62e25e3ed5688f1c42f6115d9d9853177d21f730f9dd07414c0bec765

Observation 6cb802be-2f28-43ef-a056-6c9296fdc956 · inbound

PHF: Privileged Hidden Flow for On-Policy Self-Distillation cites this paper.

PHF: Privileged Hidden Flow for On-Policy Self-Distillation Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-06-30T07:34:21.811633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-30T07:26:32.902086Z digest=sha256:56ba7f643df9b980b4b8aad5835cbc40c06859780e19f4c38bc536b54491aeba

Observation bd21824f-69a1-4582-8962-2589159d7811 · inbound

Active-GRPO: Adaptive Imitation and Self-Improving Reasoning for Molecular Optimization cites this paper.

Active-GRPO: Adaptive Imitation and Self-Improving Reasoning for Molecular Optimization Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-07-02T16:17:08.604206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T16:07:49.362916Z digest=sha256:492479d8f86d2f4a50d5cebf65064dd2199d06d781557882dc9b8a985063f7b1

Observation 90ab09a6-18ef-4af7-bb5b-b663d1979ec4 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 205

Resolution
unresolved
no resolver link, observed 2026-07-11T13:53:36.775836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T13:53:36.775836Z digest=sha256:3971bfc0778461786b5da174a08a0c7385ef39d83328cc35b75f5261830ff99b

Observation a5c001a8-f646-4831-b290-bdd4ceadd008 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 206

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:56.405118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:56.405118Z digest=sha256:b57273ec5aff530464a4c5a41e7554d77e00f783df704ade4a8a6f66f02640b8

Observation 2d3c46f9-bc77-4557-bd9f-699103e384a7 · inbound

The Verifier is the Curriculum: Execution-Gated Self-Distillation for Cross-Family Game Generation cites this paper.

The Verifier is the Curriculum: Execution-Gated Self-Distillation for Cross-Family Game Generation Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-14T17:25:37.713857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T17:25:37.713857Z digest=sha256:5ea5a079eb52e8075f393ae907198f3624c17f6b3bf6f86bb05f2f41d6a1c2df

Observation 13dd479b-e987-4d72-adad-948b0c3a7dfb · inbound

UNIBROWSE: A Data-to-Agent Framework for Multimodal BrowseComp cites this paper.

UNIBROWSE: A Data-to-Agent Framework for Multimodal BrowseComp Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-14T10:51:16.019022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T10:51:16.019022Z digest=sha256:e5d8822d0dc4cbee6ee76487012655fd6e6f04ac045bebd2ad119f93c7f96871

Observation e0ff153d-3cab-4e56-aead-9b7e7bd48a79 · inbound

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration cites this paper.

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape CoT Calibration Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 106

Resolution
unresolved
no resolver link, observed 2026-08-02T03:56:26.701252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:56:26.701252Z digest=sha256:c282e554d89c33a4d7d80d49e384eac61af1ccde9e9f46464b868157f99f9caa

Observation d1620c7f-8ad5-43ff-bc27-7b2489a9d2b6 · inbound

Answer-Conditioned Chains of Thought Degrade Verifiable-Reasoning Distillation in Large Language Models cites this paper.

Answer-Conditioned Chains of Thought Degrade Verifiable-Reasoning Distillation in Large Language Models Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T01:52:28.936526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:52:28.936526Z digest=sha256:d6e7d7f47e9abf46a2206fcdb0a834e2b11e2167a2ff86419f4289625d55cc64

Observation f75ed8b3-e8bc-4b2e-9605-2afc67667824 · inbound

CoTu at EXACT 2026: Neuro-Symbolic Reasoning for Transparent Educational QA cites this paper.

CoTu at EXACT 2026: Neuro-Symbolic Reasoning for Transparent Educational QA Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T01:15:33.669821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:15:33.669821Z digest=sha256:39dad415796b8397affc002e024fb7eab83f77630c2312633632eacef3641ff5

Observation 8601a47d-6b23-4ca9-8ba9-1fc876d16795 · inbound

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information cites this paper.

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-01T12:54:34.843879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:54:34.843879Z digest=sha256:13312fa96f37ba56e42ba4efc72e5d4559086931de503b5beb04cce2a4b078de

Observation 0be574e1-6fbf-4c6b-96a8-61d630984dcf · inbound

FormulaSPIN: Self-Play Fine-Tuning for Natural Language to Spreadsheet Formula Generation cites this paper.

FormulaSPIN: Self-Play Fine-Tuning for Natural Language to Spreadsheet Formula Generation Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T13:32:13.493804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T13:32:13.493804Z digest=sha256:2c2202c5a0a01b3e2bf3f701924a5f450999ddeaa1fad438bf620e0ede90681e

Observation 05e6d4e8-23ee-4cdc-b797-a7c0281ffebe · inbound

Bridging Compute- and Data-Optimal Pretraining cites this paper.

Bridging Compute- and Data-Optimal Pretraining Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 171

Resolution
unresolved
no resolver link, observed 2026-08-01T03:02:08.797292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:02:08.797292Z digest=sha256:1f3155242b40c49f952455270504dc46161d5c037518d087104675e8f39b271a

Observation 8f08c263-545e-4637-b701-934097b26025 · inbound

From Scoring to Acting: Outcome-Verified Comparative Self-Distillation for LLM Agents cites this paper.

From Scoring to Acting: Outcome-Verified Comparative Self-Distillation for LLM Agents Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-31T22:51:49.392233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T22:51:49.392233Z digest=sha256:e885357fe18be15f8fee6ec7f86f95d19e14167fee587187f153e485d58a2df9

Observation c2689870-d8ac-4522-a585-d2f1e16e5a9b · inbound

Recursive Synthesis for Long-Horizon Terminal Tasks cites this paper.

Recursive Synthesis for Long-Horizon Terminal Tasks Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T12:35:01.459919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:35:01.459919Z digest=sha256:c7290e181cd4c13481ec8049966630e6a116017b55baafd7396fa2a67c87978b