{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2025:XN7DDASU3QTZFGQHNMPKTPXZPU","short_pith_number":"pith:XN7DDASU","schema_version":"1.0","canonical_sha256":"bb7e318254dc27929a076b1ea9bef97d3f60bc607ff2cb9498c1944432032c59","source":{"kind":"arxiv","id":"2504.04022","version":1},"attestation_state":"computed","paper":{"title":"Rethinking Reflection in Pre-Training","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.AI"],"primary_cat":"cs.CL","authors_text":"Adarsh Chaluvaraju, Andrew Hojel, Andrew Ma, Anil Thomas, Anthony Polloreno, Ashish Tanwer, Ashish Vaswani, Burhan Drak Sibai, Divya Shivaprasad, Divya S Mansingka, Essential AI: Darsh J Shah, Ishaan Shah, Karl Stratos, Khoi Nguyen, Kurt Smith, Michael Callahan, Michael Pust, Mohit Parmar, Mrinal Iyer, Peter Rushton, Philip Monk, Platon Mazarakis, Ritvik Kapila, Saurabh Srivastava, Somanshu Singla, Tim Romanski, Yash Vanjani","submitted_at":"2025-04-05T02:24:07Z","abstract_excerpt":"A language model's ability to reflect on its own reasoning provides a key advantage for solving complex problems. While most recent research has focused on how this ability develops during reinforcement learning, we show that it actually begins to emerge much earlier - during the model's pre-training. To study this, we introduce deliberate errors into chains-of-thought and test whether the model can still arrive at the correct answer by recognizing and correcting these mistakes. By tracking performance across different stages of pre-training, we observe that this self-correcting ability appear"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2504.04022","kind":"arxiv","version":1},"metadata":{"license":"http://creativecommons.org/licenses/by/4.0/","primary_cat":"cs.CL","submitted_at":"2025-04-05T02:24:07Z","cross_cats_sorted":["cs.AI"],"title_canon_sha256":"ffa874d91bdb73c25b04b6e0cd70cdbc96d13f9b1d6e6cb2fd6930890349b3e8","abstract_canon_sha256":"890c7d4ec309404ab8912052524de2042b5cafc664dcff92282dd57d352f1d97"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T10:44:53.034913Z","signature_b64":"KpWYjarlBx60LlFruwSxtSNHp9eaGbWorKNOII8ZVUa2rdvmYM6UbIlYIALSe1fwuxwMsmnM/fXRGP3TFUL2DA==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"bb7e318254dc27929a076b1ea9bef97d3f60bc607ff2cb9498c1944432032c59","last_reissued_at":"2026-07-05T10:44:53.034367Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T10:44:53.034367Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Rethinking Reflection in Pre-Training","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.AI"],"primary_cat":"cs.CL","authors_text":"Adarsh Chaluvaraju, Andrew Hojel, Andrew Ma, Anil Thomas, Anthony Polloreno, Ashish Tanwer, Ashish Vaswani, Burhan Drak Sibai, Divya Shivaprasad, Divya S Mansingka, Essential AI: Darsh J Shah, Ishaan Shah, Karl Stratos, Khoi Nguyen, Kurt Smith, Michael Callahan, Michael Pust, Mohit Parmar, Mrinal Iyer, Peter Rushton, Philip Monk, Platon Mazarakis, Ritvik Kapila, Saurabh Srivastava, Somanshu Singla, Tim Romanski, Yash Vanjani","submitted_at":"2025-04-05T02:24:07Z","abstract_excerpt":"A language model's ability to reflect on its own reasoning provides a key advantage for solving complex problems. While most recent research has focused on how this ability develops during reinforcement learning, we show that it actually begins to emerge much earlier - during the model's pre-training. To study this, we introduce deliberate errors into chains-of-thought and test whether the model can still arrive at the correct answer by recognizing and correcting these mistakes. By tracking performance across different stages of pre-training, we observe that this self-correcting ability appear"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2504.04022","kind":"arxiv","version":1},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2504.04022/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2504.04022","created_at":"2026-07-05T10:44:53.034440+00:00"},{"alias_kind":"arxiv_version","alias_value":"2504.04022v1","created_at":"2026-07-05T10:44:53.034440+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2504.04022","created_at":"2026-07-05T10:44:53.034440+00:00"},{"alias_kind":"pith_short_12","alias_value":"XN7DDASU3QTZ","created_at":"2026-07-05T10:44:53.034440+00:00"},{"alias_kind":"pith_short_16","alias_value":"XN7DDASU3QTZFGQH","created_at":"2026-07-05T10:44:53.034440+00:00"},{"alias_kind":"pith_short_8","alias_value":"XN7DDASU","created_at":"2026-07-05T10:44:53.034440+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":6,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2605.23926","citing_title":"How Much Thinking is Enough? Quantifying and Understanding Redundancy in LLM Reasoning","ref_index":17,"is_internal_anchor":false},{"citing_arxiv_id":"2606.06096","citing_title":"OrderGrad: Optimizing Beyond the Mean with Order-Statistic Policy Gradient Estimation","ref_index":84,"is_internal_anchor":false},{"citing_arxiv_id":"2606.06080","citing_title":"On Advantage Estimates for Max@K Policy Gradients","ref_index":45,"is_internal_anchor":false},{"citing_arxiv_id":"2508.13755","citing_title":"Depth-Breadth Synergy in RLVR: Unlocking LLM Reasoning Gains with Adaptive Exploration","ref_index":13,"is_internal_anchor":false},{"citing_arxiv_id":"2509.21882","citing_title":"Position: The Hidden Costs and Measurement Gaps of Reinforcement Learning with Verifiable Rewards","ref_index":29,"is_internal_anchor":false},{"citing_arxiv_id":"2504.20571","citing_title":"Reinforcement Learning for Reasoning in Large Language Models with One Training Example","ref_index":21,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/XN7DDASU3QTZFGQHNMPKTPXZPU","json":"https://pith.science/pith/XN7DDASU3QTZFGQHNMPKTPXZPU.json","graph_json":"https://pith.science/api/pith-number/XN7DDASU3QTZFGQHNMPKTPXZPU/graph.json","events_json":"https://pith.science/api/pith-number/XN7DDASU3QTZFGQHNMPKTPXZPU/events.json","paper":"https://pith.science/paper/XN7DDASU"},"agent_actions":{"view_html":"https://pith.science/pith/XN7DDASU3QTZFGQHNMPKTPXZPU","download_json":"https://pith.science/pith/XN7DDASU3QTZFGQHNMPKTPXZPU.json","view_paper":"https://pith.science/paper/XN7DDASU","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2504.04022&json=true","fetch_graph":"https://pith.science/api/pith-number/XN7DDASU3QTZFGQHNMPKTPXZPU/graph.json","fetch_events":"https://pith.science/api/pith-number/XN7DDASU3QTZFGQHNMPKTPXZPU/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/XN7DDASU3QTZFGQHNMPKTPXZPU/action/timestamp_anchor","attest_storage":"https://pith.science/pith/XN7DDASU3QTZFGQHNMPKTPXZPU/action/storage_attestation","attest_author":"https://pith.science/pith/XN7DDASU3QTZFGQHNMPKTPXZPU/action/author_attestation","sign_citation":"https://pith.science/pith/XN7DDASU3QTZFGQHNMPKTPXZPU/action/citation_signature","submit_replication":"https://pith.science/pith/XN7DDASU3QTZFGQHNMPKTPXZPU/action/replication_record"}},"created_at":"2026-07-05T10:44:53.034440+00:00","updated_at":"2026-07-05T10:44:53.034440+00:00"}