{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2025:5DX7ITEOLFMRHPKEB76FSII3EI","short_pith_number":"pith:5DX7ITEO","schema_version":"1.0","canonical_sha256":"e8eff44c8e595913bd440ffc59211b2206f3e5a91ae6bf013e66f8b641a50d40","source":{"kind":"arxiv","id":"2507.09850","version":3},"attestation_state":"computed","paper":{"title":"The Challenge of Teaching Reasoning to LLMs Without RL or Distillation","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":[],"primary_cat":"cs.AI","authors_text":"Advaith Avadhanam, Alexan Ayrapetyan, Ashmit Dutta, Boris Ginsburg, Branislav Kisacanin, Dan Zhao, David Zhang, Dragan Masulovic, George Armstrong, Igor Gitman, Ivan Moshkov, Joonseok Kang, Leon Luo, Marius Stanean, Max Wang, Mihir Tandon, Sadegh Mahdavi, Shitij Govil, Shizhe Diao, Shubham Toshniwal, Sriram Ananthakrishnan, Sri Yanamandara, Titu Andreescu, Vedant Rathi, Wei Du","submitted_at":"2025-07-14T01:14:50Z","abstract_excerpt":"Reasoning-capable language models achieve state-of-the-art performance in diverse complex tasks by generating long, explicit Chain-of-Thought (CoT) traces. While recent works show that base models can acquire such reasoning traces via reinforcement learning or distillation from stronger models like DeepSeek-R1, previous works demonstrate that even short CoT prompting without fine-tuning is able to improve reasoning. We ask whether long CoT can be induced in a base model using only prompting or minimal tuning. Using just 20 long CoT examples from the reasoning model \\texttt{QwQ-32B-Preview}, we"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2507.09850","kind":"arxiv","version":3},"metadata":{"license":"http://creativecommons.org/licenses/by/4.0/","primary_cat":"cs.AI","submitted_at":"2025-07-14T01:14:50Z","cross_cats_sorted":[],"title_canon_sha256":"717b124fb9d55155cfb1168b7f7092789b151279bffa59ad9caaf38104310009","abstract_canon_sha256":"2c0bd347ac124c62a11e95fd7c262d65379577aff75f894010f09baa4f7f35b1"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T11:38:09.409054Z","signature_b64":"OjuqFfy3dxwNXQoDLeyeKaUq5S7ipoP8fQqBYMlQqqaAl2mXC3uQtI+3psVtMFuOtKsirybknDgu3c4H0px/AQ==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"e8eff44c8e595913bd440ffc59211b2206f3e5a91ae6bf013e66f8b641a50d40","last_reissued_at":"2026-07-05T11:38:09.408418Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T11:38:09.408418Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"The Challenge of Teaching Reasoning to LLMs Without RL or Distillation","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":[],"primary_cat":"cs.AI","authors_text":"Advaith Avadhanam, Alexan Ayrapetyan, Ashmit Dutta, Boris Ginsburg, Branislav Kisacanin, Dan Zhao, David Zhang, Dragan Masulovic, George Armstrong, Igor Gitman, Ivan Moshkov, Joonseok Kang, Leon Luo, Marius Stanean, Max Wang, Mihir Tandon, Sadegh Mahdavi, Shitij Govil, Shizhe Diao, Shubham Toshniwal, Sriram Ananthakrishnan, Sri Yanamandara, Titu Andreescu, Vedant Rathi, Wei Du","submitted_at":"2025-07-14T01:14:50Z","abstract_excerpt":"Reasoning-capable language models achieve state-of-the-art performance in diverse complex tasks by generating long, explicit Chain-of-Thought (CoT) traces. While recent works show that base models can acquire such reasoning traces via reinforcement learning or distillation from stronger models like DeepSeek-R1, previous works demonstrate that even short CoT prompting without fine-tuning is able to improve reasoning. We ask whether long CoT can be induced in a base model using only prompting or minimal tuning. Using just 20 long CoT examples from the reasoning model \\texttt{QwQ-32B-Preview}, we"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2507.09850","kind":"arxiv","version":3},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2507.09850/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2507.09850","created_at":"2026-07-05T11:38:09.408494+00:00"},{"alias_kind":"arxiv_version","alias_value":"2507.09850v3","created_at":"2026-07-05T11:38:09.408494+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2507.09850","created_at":"2026-07-05T11:38:09.408494+00:00"},{"alias_kind":"pith_short_12","alias_value":"5DX7ITEOLFMR","created_at":"2026-07-05T11:38:09.408494+00:00"},{"alias_kind":"pith_short_16","alias_value":"5DX7ITEOLFMRHPKE","created_at":"2026-07-05T11:38:09.408494+00:00"},{"alias_kind":"pith_short_8","alias_value":"5DX7ITEO","created_at":"2026-07-05T11:38:09.408494+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":0,"internal_anchor_count":0,"sample":[]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/5DX7ITEOLFMRHPKEB76FSII3EI","json":"https://pith.science/pith/5DX7ITEOLFMRHPKEB76FSII3EI.json","graph_json":"https://pith.science/api/pith-number/5DX7ITEOLFMRHPKEB76FSII3EI/graph.json","events_json":"https://pith.science/api/pith-number/5DX7ITEOLFMRHPKEB76FSII3EI/events.json","paper":"https://pith.science/paper/5DX7ITEO"},"agent_actions":{"view_html":"https://pith.science/pith/5DX7ITEOLFMRHPKEB76FSII3EI","download_json":"https://pith.science/pith/5DX7ITEOLFMRHPKEB76FSII3EI.json","view_paper":"https://pith.science/paper/5DX7ITEO","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2507.09850&json=true","fetch_graph":"https://pith.science/api/pith-number/5DX7ITEOLFMRHPKEB76FSII3EI/graph.json","fetch_events":"https://pith.science/api/pith-number/5DX7ITEOLFMRHPKEB76FSII3EI/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/5DX7ITEOLFMRHPKEB76FSII3EI/action/timestamp_anchor","attest_storage":"https://pith.science/pith/5DX7ITEOLFMRHPKEB76FSII3EI/action/storage_attestation","attest_author":"https://pith.science/pith/5DX7ITEOLFMRHPKEB76FSII3EI/action/author_attestation","sign_citation":"https://pith.science/pith/5DX7ITEOLFMRHPKEB76FSII3EI/action/citation_signature","submit_replication":"https://pith.science/pith/5DX7ITEOLFMRHPKEB76FSII3EI/action/replication_record"}},"created_at":"2026-07-05T11:38:09.408494+00:00","updated_at":"2026-07-05T11:38:09.408494+00:00"}