{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2025:2JLTHYKDJIYCAF34KFMHT244XY","short_pith_number":"pith:2JLTHYKD","schema_version":"1.0","canonical_sha256":"d25733e1434a3020177c515879eb9cbe348307b389bed5d79765572fe0cf8f99","source":{"kind":"arxiv","id":"2506.09026","version":2},"attestation_state":"computed","paper":{"title":"e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.CL"],"primary_cat":"cs.LG","authors_text":"Amrith Setlur, Aviral Kumar, Charlie Snell, Ian Wu, Jeremy Greer, Matthew Y. R. Yang, Max Simchowitz, Virginia Smith","submitted_at":"2025-06-10T17:52:42Z","abstract_excerpt":"Test-time scaling offers a promising path to improve LLM reasoning by utilizing more compute at inference time; however, the true promise of this paradigm lies in extrapolation (i.e., improvement in performance on hard problems as LLMs keep \"thinking\" for longer, beyond the maximum token budget they were trained on). Surprisingly, we find that most existing reasoning models do not extrapolate well. We show that one way to enable extrapolation is by training the LLM to perform in-context exploration: training the LLM to effectively spend its test time budget by chaining operations (such as gene"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2506.09026","kind":"arxiv","version":2},"metadata":{"license":"http://creativecommons.org/licenses/by/4.0/","primary_cat":"cs.LG","submitted_at":"2025-06-10T17:52:42Z","cross_cats_sorted":["cs.CL"],"title_canon_sha256":"6a2ab98c24a0dcb28732f4de062978a6b193957834788d88ac217a18610cc29c","abstract_canon_sha256":"97d0810375bcf87d39326be6f6b6015e772d131daa96798130b962e08ee5b62a"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T11:20:56.447593Z","signature_b64":"S9aUuSB50NE2pGxKukose/PlMuXBZMdatAmYCq9knzAw8hCbrtkIUA/hFBvLYY/YXl5yPTFRjlVCLhv/LZFiDw==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"d25733e1434a3020177c515879eb9cbe348307b389bed5d79765572fe0cf8f99","last_reissued_at":"2026-07-05T11:20:56.447045Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T11:20:56.447045Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.CL"],"primary_cat":"cs.LG","authors_text":"Amrith Setlur, Aviral Kumar, Charlie Snell, Ian Wu, Jeremy Greer, Matthew Y. R. Yang, Max Simchowitz, Virginia Smith","submitted_at":"2025-06-10T17:52:42Z","abstract_excerpt":"Test-time scaling offers a promising path to improve LLM reasoning by utilizing more compute at inference time; however, the true promise of this paradigm lies in extrapolation (i.e., improvement in performance on hard problems as LLMs keep \"thinking\" for longer, beyond the maximum token budget they were trained on). Surprisingly, we find that most existing reasoning models do not extrapolate well. We show that one way to enable extrapolation is by training the LLM to perform in-context exploration: training the LLM to effectively spend its test time budget by chaining operations (such as gene"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2506.09026","kind":"arxiv","version":2},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2506.09026/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2506.09026","created_at":"2026-07-05T11:20:56.447106+00:00"},{"alias_kind":"arxiv_version","alias_value":"2506.09026v2","created_at":"2026-07-05T11:20:56.447106+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2506.09026","created_at":"2026-07-05T11:20:56.447106+00:00"},{"alias_kind":"pith_short_12","alias_value":"2JLTHYKDJIYC","created_at":"2026-07-05T11:20:56.447106+00:00"},{"alias_kind":"pith_short_16","alias_value":"2JLTHYKDJIYCAF34","created_at":"2026-07-05T11:20:56.447106+00:00"},{"alias_kind":"pith_short_8","alias_value":"2JLTHYKD","created_at":"2026-07-05T11:20:56.447106+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":7,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.06096","citing_title":"OrderGrad: Optimizing Beyond the Mean with Order-Statistic Policy Gradient Estimation","ref_index":83,"is_internal_anchor":false},{"citing_arxiv_id":"2606.06080","citing_title":"On Advantage Estimates for Max@K Policy Gradients","ref_index":44,"is_internal_anchor":false},{"citing_arxiv_id":"2605.03065","citing_title":"OGPO: Sample Efficient Full-Finetuning of Generative Control Policies","ref_index":42,"is_internal_anchor":false},{"citing_arxiv_id":"2601.20829","citing_title":"Training Reasoning Models on Saturated Problems via Failure-Prefix Conditioning","ref_index":17,"is_internal_anchor":false},{"citing_arxiv_id":"2603.04333","citing_title":"What Does Flow Matching Bring To TD Learning?","ref_index":55,"is_internal_anchor":false},{"citing_arxiv_id":"2604.18493","citing_title":"Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data","ref_index":23,"is_internal_anchor":false},{"citing_arxiv_id":"2605.03065","citing_title":"OGPO: Sample Efficient Full-Finetuning of Generative Control Policies","ref_index":175,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/2JLTHYKDJIYCAF34KFMHT244XY","json":"https://pith.science/pith/2JLTHYKDJIYCAF34KFMHT244XY.json","graph_json":"https://pith.science/api/pith-number/2JLTHYKDJIYCAF34KFMHT244XY/graph.json","events_json":"https://pith.science/api/pith-number/2JLTHYKDJIYCAF34KFMHT244XY/events.json","paper":"https://pith.science/paper/2JLTHYKD"},"agent_actions":{"view_html":"https://pith.science/pith/2JLTHYKDJIYCAF34KFMHT244XY","download_json":"https://pith.science/pith/2JLTHYKDJIYCAF34KFMHT244XY.json","view_paper":"https://pith.science/paper/2JLTHYKD","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2506.09026&json=true","fetch_graph":"https://pith.science/api/pith-number/2JLTHYKDJIYCAF34KFMHT244XY/graph.json","fetch_events":"https://pith.science/api/pith-number/2JLTHYKDJIYCAF34KFMHT244XY/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/2JLTHYKDJIYCAF34KFMHT244XY/action/timestamp_anchor","attest_storage":"https://pith.science/pith/2JLTHYKDJIYCAF34KFMHT244XY/action/storage_attestation","attest_author":"https://pith.science/pith/2JLTHYKDJIYCAF34KFMHT244XY/action/author_attestation","sign_citation":"https://pith.science/pith/2JLTHYKDJIYCAF34KFMHT244XY/action/citation_signature","submit_replication":"https://pith.science/pith/2JLTHYKDJIYCAF34KFMHT244XY/action/replication_record"}},"created_at":"2026-07-05T11:20:56.447106+00:00","updated_at":"2026-07-05T11:20:56.447106+00:00"}