{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2023:CGIBGO635F67SMMO7ZUYSHIKO5","short_pith_number":"pith:CGIBGO63","schema_version":"1.0","canonical_sha256":"1190133bdbe97df9318efe69891d0a7765158f2f07c750ba8010d727bcafe5c8","source":{"kind":"arxiv","id":"2312.06585","version":4},"attestation_state":"computed","paper":{"title":"Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models","license":"http://creativecommons.org/licenses/by-sa/4.0/","headline":"","cross_cats":[],"primary_cat":"cs.LG","authors_text":"Aaron Parisi, Abhishek Kumar, Alex Alemi, Alex Rizkowsky, Ankesh Anand, Avi Singh, Azade Nova, Behnam Neyshabur, Ben Adlam, Bernd Bohnet, Ethan Dyer, Gamaleldin Elsayed, Hanie Sedghi, Igor Mordatch, Isabelle Simpson, Izzeddin Gur, Jaehoon Lee, James Harrison, Jascha Sohl-Dickstein, Jasper Snoek, Jeffrey Pennington, Jiri Hron, John D. Co-Reyes, Kathleen Kenealy, Kelvin Xu, Kevin Swersky, Kshiteej Mahajan, Laura Culp, Lechao Xiao, Maxwell L. Bileschi, Noah Constant, Noah Fiedel, Peter J. Liu, Piyush Patil, Rishabh Agarwal, Roman Novak, Rosanne Liu, Tris Warkentin, Xavier Garcia, Yamini Bansal, Yundi Qian","submitted_at":"2023-12-11T18:17:43Z","abstract_excerpt":"Fine-tuning language models~(LMs) on human-generated data remains a prevalent practice. However, the performance of such models is often limited by the quantity and diversity of high-quality human data. In this paper, we explore whether we can go beyond human data on tasks where we have access to scalar feedback, for example, on math problems where one can verify correctness. To do so, we investigate a simple self-training method based on expectation-maximization, which we call ReST$^{EM}$, where we (1) generate samples from the model and filter them using binary feedback, (2) fine-tune the mo"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2312.06585","kind":"arxiv","version":4},"metadata":{"license":"http://creativecommons.org/licenses/by-sa/4.0/","primary_cat":"cs.LG","submitted_at":"2023-12-11T18:17:43Z","cross_cats_sorted":[],"title_canon_sha256":"e6f01419f3183ccef4ae6812e0bc345f062a4d13758792b5f93678a3dc639db4","abstract_canon_sha256":"b4ddb1d9e9a1501c3bbbf08491415a4f11b004e8a7be172142b48cca1fdcc66f"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T08:09:27.040334Z","signature_b64":"Iujldx8qKy2m5y1OB1wIRPf5/Zi04/bPm1DrFao03FofmWmioQsHcR3iuJwn+TuT3rQlBtVBk6G37n1vFQGPAg==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"1190133bdbe97df9318efe69891d0a7765158f2f07c750ba8010d727bcafe5c8","last_reissued_at":"2026-07-05T08:09:27.039855Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T08:09:27.039855Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models","license":"http://creativecommons.org/licenses/by-sa/4.0/","headline":"","cross_cats":[],"primary_cat":"cs.LG","authors_text":"Aaron Parisi, Abhishek Kumar, Alex Alemi, Alex Rizkowsky, Ankesh Anand, Avi Singh, Azade Nova, Behnam Neyshabur, Ben Adlam, Bernd Bohnet, Ethan Dyer, Gamaleldin Elsayed, Hanie Sedghi, Igor Mordatch, Isabelle Simpson, Izzeddin Gur, Jaehoon Lee, James Harrison, Jascha Sohl-Dickstein, Jasper Snoek, Jeffrey Pennington, Jiri Hron, John D. Co-Reyes, Kathleen Kenealy, Kelvin Xu, Kevin Swersky, Kshiteej Mahajan, Laura Culp, Lechao Xiao, Maxwell L. Bileschi, Noah Constant, Noah Fiedel, Peter J. Liu, Piyush Patil, Rishabh Agarwal, Roman Novak, Rosanne Liu, Tris Warkentin, Xavier Garcia, Yamini Bansal, Yundi Qian","submitted_at":"2023-12-11T18:17:43Z","abstract_excerpt":"Fine-tuning language models~(LMs) on human-generated data remains a prevalent practice. However, the performance of such models is often limited by the quantity and diversity of high-quality human data. In this paper, we explore whether we can go beyond human data on tasks where we have access to scalar feedback, for example, on math problems where one can verify correctness. To do so, we investigate a simple self-training method based on expectation-maximization, which we call ReST$^{EM}$, where we (1) generate samples from the model and filter them using binary feedback, (2) fine-tune the mo"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2312.06585","kind":"arxiv","version":4},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2312.06585/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2312.06585","created_at":"2026-07-05T08:09:27.039908+00:00"},{"alias_kind":"arxiv_version","alias_value":"2312.06585v4","created_at":"2026-07-05T08:09:27.039908+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2312.06585","created_at":"2026-07-05T08:09:27.039908+00:00"},{"alias_kind":"pith_short_12","alias_value":"CGIBGO635F67","created_at":"2026-07-05T08:09:27.039908+00:00"},{"alias_kind":"pith_short_16","alias_value":"CGIBGO635F67SMMO","created_at":"2026-07-05T08:09:27.039908+00:00"},{"alias_kind":"pith_short_8","alias_value":"CGIBGO63","created_at":"2026-07-05T08:09:27.039908+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":30,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.26091","citing_title":"On-Policy Self-Distillation with Sampled Demonstrations Reduces Output Diversity","ref_index":81,"is_internal_anchor":false},{"citing_arxiv_id":"2606.27373","citing_title":"Paying More Attention to Visual Tokens in Self-Evolving Large Multimodal Models","ref_index":27,"is_internal_anchor":false},{"citing_arxiv_id":"2606.23611","citing_title":"Data Selection Through Iterative Self-Filtering for Vision-Language Settings","ref_index":58,"is_internal_anchor":false},{"citing_arxiv_id":"2606.18216","citing_title":"Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients","ref_index":77,"is_internal_anchor":false},{"citing_arxiv_id":"2606.18307","citing_title":"DRIFT: Refining Instruction Data via On-Policy Data Attribution","ref_index":19,"is_internal_anchor":false},{"citing_arxiv_id":"2607.00531","citing_title":"Active-GRPO: Adaptive Imitation and Self-Improving Reasoning for Molecular Optimization","ref_index":49,"is_internal_anchor":false},{"citing_arxiv_id":"2606.05464","citing_title":"Step-by-Step Optimization-like Reasoning in LLMs over Expanding Search Spaces","ref_index":7,"is_internal_anchor":false},{"citing_arxiv_id":"2605.01248","citing_title":"$S^3$-R1: Learning to Retrieve and Answer Step-by-Step with Synthetic Data","ref_index":16,"is_internal_anchor":false},{"citing_arxiv_id":"2605.20914","citing_title":"RISE: Reliable Improvement in Self-Evolving Vision-Language Models","ref_index":35,"is_internal_anchor":false},{"citing_arxiv_id":"2606.29340","citing_title":"PHF: Privileged Hidden Flow for On-Policy Self-Distillation","ref_index":47,"is_internal_anchor":false},{"citing_arxiv_id":"2605.26037","citing_title":"Peak-Then-Collapse and the Four Interface Channels of Knowledge-Graph Tool Use","ref_index":2,"is_internal_anchor":false},{"citing_arxiv_id":"2605.31455","citing_title":"DRIFT: Decoupled Rollouts and Importance-Weighted Fine-Tuning for Efficient Multi-Turn Optimization","ref_index":22,"is_internal_anchor":false},{"citing_arxiv_id":"2606.00628","citing_title":"Robust Reasoning via Dynamic Token Selection for Distribution-Aligned Self-Distillation","ref_index":8,"is_internal_anchor":false},{"citing_arxiv_id":"2406.11290","citing_title":"An Iterative Utility Judgment Framework Inspired by Philosophical Relevance via LLMs","ref_index":5,"is_internal_anchor":false},{"citing_arxiv_id":"2504.01990","citing_title":"Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe Systems","ref_index":159,"is_internal_anchor":false},{"citing_arxiv_id":"2605.22675","citing_title":"Self-Policy Distillation via Capability-Selective Subspace Projection","ref_index":2,"is_internal_anchor":false},{"citing_arxiv_id":"2602.07832","citing_title":"rePIRL: Learn PRM with Inverse RL for LLM Reasoning","ref_index":25,"is_internal_anchor":false},{"citing_arxiv_id":"2605.20189","citing_title":"SOLAR: A Self-Optimizing Open-Ended Autonomous Agent for Lifelong Learning and Continual Adaptation","ref_index":66,"is_internal_anchor":false},{"citing_arxiv_id":"2605.20914","citing_title":"RISE: Reliable Improvement in Self-Evolving Vision-Language Models","ref_index":31,"is_internal_anchor":false},{"citing_arxiv_id":"2410.08146","citing_title":"Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning","ref_index":20,"is_internal_anchor":false},{"citing_arxiv_id":"2503.17352","citing_title":"OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles","ref_index":63,"is_internal_anchor":false},{"citing_arxiv_id":"2502.03373","citing_title":"Demystifying Long Chain-of-Thought Reasoning in LLMs","ref_index":4,"is_internal_anchor":false},{"citing_arxiv_id":"2409.12917","citing_title":"Training Language Models to Self-Correct via Reinforcement Learning","ref_index":62,"is_internal_anchor":false},{"citing_arxiv_id":"2602.22507","citing_title":"Space Syntax-guided Post-training for Residential Floor Plan Generation","ref_index":44,"is_internal_anchor":false},{"citing_arxiv_id":"2502.17419","citing_title":"From System 1 to System 2: A Survey of Reasoning Large Language Models","ref_index":192,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/CGIBGO635F67SMMO7ZUYSHIKO5","json":"https://pith.science/pith/CGIBGO635F67SMMO7ZUYSHIKO5.json","graph_json":"https://pith.science/api/pith-number/CGIBGO635F67SMMO7ZUYSHIKO5/graph.json","events_json":"https://pith.science/api/pith-number/CGIBGO635F67SMMO7ZUYSHIKO5/events.json","paper":"https://pith.science/paper/CGIBGO63"},"agent_actions":{"view_html":"https://pith.science/pith/CGIBGO635F67SMMO7ZUYSHIKO5","download_json":"https://pith.science/pith/CGIBGO635F67SMMO7ZUYSHIKO5.json","view_paper":"https://pith.science/paper/CGIBGO63","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2312.06585&json=true","fetch_graph":"https://pith.science/api/pith-number/CGIBGO635F67SMMO7ZUYSHIKO5/graph.json","fetch_events":"https://pith.science/api/pith-number/CGIBGO635F67SMMO7ZUYSHIKO5/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/CGIBGO635F67SMMO7ZUYSHIKO5/action/timestamp_anchor","attest_storage":"https://pith.science/pith/CGIBGO635F67SMMO7ZUYSHIKO5/action/storage_attestation","attest_author":"https://pith.science/pith/CGIBGO635F67SMMO7ZUYSHIKO5/action/author_attestation","sign_citation":"https://pith.science/pith/CGIBGO635F67SMMO7ZUYSHIKO5/action/citation_signature","submit_replication":"https://pith.science/pith/CGIBGO635F67SMMO7ZUYSHIKO5/action/replication_record"}},"created_at":"2026-07-05T08:09:27.039908+00:00","updated_at":"2026-07-05T08:09:27.039908+00:00"}