{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2024:XSXRTB7HLGWUVBNY5FEXQNIXW7","short_pith_number":"pith:XSXRTB7H","schema_version":"1.0","canonical_sha256":"bcaf1987e759ad4a85b8e949783517b7dac5c4eabf56f473a071121347fa27a3","source":{"kind":"arxiv","id":"2410.01792","version":2},"attestation_state":"computed","paper":{"title":"When a language model is optimized for reasoning, does it still show embers of autoregression? An analysis of OpenAI o1","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.AI"],"primary_cat":"cs.CL","authors_text":"Dan Friedman, Mathew D. Hardy, R. Thomas McCoy, Shunyu Yao, Thomas L. Griffiths","submitted_at":"2024-10-02T17:50:19Z","abstract_excerpt":"In \"Embers of Autoregression\" (McCoy et al., 2023), we showed that several large language models (LLMs) have some important limitations that are attributable to their origins in next-word prediction. Here we investigate whether these issues persist with o1, a new system from OpenAI that differs from previous LLMs in that it is optimized for reasoning. We find that o1 substantially outperforms previous LLMs in many cases, with particularly large improvements on rare variants of common tasks (e.g., forming acronyms from the second letter of each word in a list, rather than the first letter). Des"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2410.01792","kind":"arxiv","version":2},"metadata":{"license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","primary_cat":"cs.CL","submitted_at":"2024-10-02T17:50:19Z","cross_cats_sorted":["cs.AI"],"title_canon_sha256":"590f6d2f4152595f77374a11021b046f972ff6752b4a7d769060a545a81f558c","abstract_canon_sha256":"f2d793398db7c57aba58a752a23a62db0e7c5553ee299fef6b6098b8cfc5cd58"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T09:15:38.763266Z","signature_b64":"I6jJ4gpwZhVt1TDZHYP8lvgb9iAuNqxHL/IQI+LpyG3ssdPp1vOfHqs3JkLR/SrPEUxVE72fVS+toYhCpZB3Cw==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"bcaf1987e759ad4a85b8e949783517b7dac5c4eabf56f473a071121347fa27a3","last_reissued_at":"2026-07-05T09:15:38.762763Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T09:15:38.762763Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"When a language model is optimized for reasoning, does it still show embers of autoregression? An analysis of OpenAI o1","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.AI"],"primary_cat":"cs.CL","authors_text":"Dan Friedman, Mathew D. Hardy, R. Thomas McCoy, Shunyu Yao, Thomas L. Griffiths","submitted_at":"2024-10-02T17:50:19Z","abstract_excerpt":"In \"Embers of Autoregression\" (McCoy et al., 2023), we showed that several large language models (LLMs) have some important limitations that are attributable to their origins in next-word prediction. Here we investigate whether these issues persist with o1, a new system from OpenAI that differs from previous LLMs in that it is optimized for reasoning. We find that o1 substantially outperforms previous LLMs in many cases, with particularly large improvements on rare variants of common tasks (e.g., forming acronyms from the second letter of each word in a list, rather than the first letter). Des"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2410.01792","kind":"arxiv","version":2},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2410.01792/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2410.01792","created_at":"2026-07-05T09:15:38.762833+00:00"},{"alias_kind":"arxiv_version","alias_value":"2410.01792v2","created_at":"2026-07-05T09:15:38.762833+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2410.01792","created_at":"2026-07-05T09:15:38.762833+00:00"},{"alias_kind":"pith_short_12","alias_value":"XSXRTB7HLGWU","created_at":"2026-07-05T09:15:38.762833+00:00"},{"alias_kind":"pith_short_16","alias_value":"XSXRTB7HLGWUVBNY","created_at":"2026-07-05T09:15:38.762833+00:00"},{"alias_kind":"pith_short_8","alias_value":"XSXRTB7H","created_at":"2026-07-05T09:15:38.762833+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":4,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.19308","citing_title":"Enhancing Decision-Making with Large Language Models through Multi-Agent Fictitious Play","ref_index":34,"is_internal_anchor":false},{"citing_arxiv_id":"2501.09686","citing_title":"Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models","ref_index":95,"is_internal_anchor":false},{"citing_arxiv_id":"2604.01621","citing_title":"DWDP: Distributed Weight Data Parallelism for High-Performance LLM Inference on NVL72","ref_index":2,"is_internal_anchor":false},{"citing_arxiv_id":"2604.21632","citing_title":"To See the Unseen: on the Generalization Ability of Transformers in Symbolic Reasoning","ref_index":11,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/XSXRTB7HLGWUVBNY5FEXQNIXW7","json":"https://pith.science/pith/XSXRTB7HLGWUVBNY5FEXQNIXW7.json","graph_json":"https://pith.science/api/pith-number/XSXRTB7HLGWUVBNY5FEXQNIXW7/graph.json","events_json":"https://pith.science/api/pith-number/XSXRTB7HLGWUVBNY5FEXQNIXW7/events.json","paper":"https://pith.science/paper/XSXRTB7H"},"agent_actions":{"view_html":"https://pith.science/pith/XSXRTB7HLGWUVBNY5FEXQNIXW7","download_json":"https://pith.science/pith/XSXRTB7HLGWUVBNY5FEXQNIXW7.json","view_paper":"https://pith.science/paper/XSXRTB7H","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2410.01792&json=true","fetch_graph":"https://pith.science/api/pith-number/XSXRTB7HLGWUVBNY5FEXQNIXW7/graph.json","fetch_events":"https://pith.science/api/pith-number/XSXRTB7HLGWUVBNY5FEXQNIXW7/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/XSXRTB7HLGWUVBNY5FEXQNIXW7/action/timestamp_anchor","attest_storage":"https://pith.science/pith/XSXRTB7HLGWUVBNY5FEXQNIXW7/action/storage_attestation","attest_author":"https://pith.science/pith/XSXRTB7HLGWUVBNY5FEXQNIXW7/action/author_attestation","sign_citation":"https://pith.science/pith/XSXRTB7HLGWUVBNY5FEXQNIXW7/action/citation_signature","submit_replication":"https://pith.science/pith/XSXRTB7HLGWUVBNY5FEXQNIXW7/action/replication_record"}},"created_at":"2026-07-05T09:15:38.762833+00:00","updated_at":"2026-07-05T09:15:38.762833+00:00"}