{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2025:NDWUQGAFHHJVTKEFKQZDO26X3J","short_pith_number":"pith:NDWUQGAF","schema_version":"1.0","canonical_sha256":"68ed48180539d359a8855432376bd7da6b1979332e3cf594feeee07493718a16","source":{"kind":"arxiv","id":"2506.18183","version":3},"attestation_state":"computed","paper":{"title":"Reasoning about Uncertainty: Do Reasoning Models Know When They Don't Know?","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.CL"],"primary_cat":"cs.AI","authors_text":"Anirudha Majumdar, Christina Zhang, Justin Lidard, Ola Shorinwa, Tenny Yin, Zhiting Mei","submitted_at":"2025-06-22T21:46:42Z","abstract_excerpt":"Reasoning language models have set state-of-the-art (SOTA) records on many challenging benchmarks, enabled by multi-step reasoning induced using reinforcement learning. However, like previous language models, reasoning models are prone to generating confident, plausible responses that are incorrect (hallucinations). Knowing when and how much to trust these models is critical to the safe deployment of reasoning models in real-world applications. To this end, we explore uncertainty quantification of reasoning models in this work. Specifically, we ask three fundamental questions: First, are reaso"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2506.18183","kind":"arxiv","version":3},"metadata":{"license":"http://creativecommons.org/licenses/by/4.0/","primary_cat":"cs.AI","submitted_at":"2025-06-22T21:46:42Z","cross_cats_sorted":["cs.CL"],"title_canon_sha256":"654bb680806b6427c4a57e8604b12a7bd280031c3a97565892d078276100b630","abstract_canon_sha256":"b6140fd2a8039fcd79953955518bd89c68fc3b7afbd97735686a66b25a9c122a"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T11:39:12.495468Z","signature_b64":"gjTO5yBfTwFKiw0y1BjJBGae4SZbHhdSYxFzBgwwuzOPh014EVGVXoG7m10AB8/khIN+UDIvFEK0wVUdynRLCg==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"68ed48180539d359a8855432376bd7da6b1979332e3cf594feeee07493718a16","last_reissued_at":"2026-07-05T11:39:12.494946Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T11:39:12.494946Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Reasoning about Uncertainty: Do Reasoning Models Know When They Don't Know?","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.CL"],"primary_cat":"cs.AI","authors_text":"Anirudha Majumdar, Christina Zhang, Justin Lidard, Ola Shorinwa, Tenny Yin, Zhiting Mei","submitted_at":"2025-06-22T21:46:42Z","abstract_excerpt":"Reasoning language models have set state-of-the-art (SOTA) records on many challenging benchmarks, enabled by multi-step reasoning induced using reinforcement learning. However, like previous language models, reasoning models are prone to generating confident, plausible responses that are incorrect (hallucinations). Knowing when and how much to trust these models is critical to the safe deployment of reasoning models in real-world applications. To this end, we explore uncertainty quantification of reasoning models in this work. Specifically, we ask three fundamental questions: First, are reaso"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2506.18183","kind":"arxiv","version":3},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2506.18183/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2506.18183","created_at":"2026-07-05T11:39:12.494999+00:00"},{"alias_kind":"arxiv_version","alias_value":"2506.18183v3","created_at":"2026-07-05T11:39:12.494999+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2506.18183","created_at":"2026-07-05T11:39:12.494999+00:00"},{"alias_kind":"pith_short_12","alias_value":"NDWUQGAFHHJV","created_at":"2026-07-05T11:39:12.494999+00:00"},{"alias_kind":"pith_short_16","alias_value":"NDWUQGAFHHJVTKEF","created_at":"2026-07-05T11:39:12.494999+00:00"},{"alias_kind":"pith_short_8","alias_value":"NDWUQGAF","created_at":"2026-07-05T11:39:12.494999+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":8,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.13649","citing_title":"Operadic consistency: a label-free signal for compositional reasoning failures in LLMs","ref_index":32,"is_internal_anchor":false},{"citing_arxiv_id":"2607.01612","citing_title":"Scaling with Confidence: Calibrating Confidence of LLMs for Adaptive Test Time Scaling","ref_index":10,"is_internal_anchor":false},{"citing_arxiv_id":"2606.30814","citing_title":"When Calibration Rankings Reverse: Accuracy-Controlled Evaluation for Fair Comparison of LLMs","ref_index":144,"is_internal_anchor":false},{"citing_arxiv_id":"2603.17839","citing_title":"How do LLMs Compute Verbal Confidence","ref_index":12,"is_internal_anchor":false},{"citing_arxiv_id":"2509.21882","citing_title":"Position: The Hidden Costs and Measurement Gaps of Reinforcement Learning with Verifiable Rewards","ref_index":20,"is_internal_anchor":false},{"citing_arxiv_id":"2602.22474","citing_title":"When to Act, Ask, or Learn: Uncertainty-Aware Policy Steering","ref_index":18,"is_internal_anchor":false},{"citing_arxiv_id":"2604.22266","citing_title":"Large Language Models Decide Early and Explain Later","ref_index":9,"is_internal_anchor":false},{"citing_arxiv_id":"2604.19444","citing_title":"Unsupervised Confidence Calibration for Reasoning LLMs from a Single Generation","ref_index":180,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/NDWUQGAFHHJVTKEFKQZDO26X3J","json":"https://pith.science/pith/NDWUQGAFHHJVTKEFKQZDO26X3J.json","graph_json":"https://pith.science/api/pith-number/NDWUQGAFHHJVTKEFKQZDO26X3J/graph.json","events_json":"https://pith.science/api/pith-number/NDWUQGAFHHJVTKEFKQZDO26X3J/events.json","paper":"https://pith.science/paper/NDWUQGAF"},"agent_actions":{"view_html":"https://pith.science/pith/NDWUQGAFHHJVTKEFKQZDO26X3J","download_json":"https://pith.science/pith/NDWUQGAFHHJVTKEFKQZDO26X3J.json","view_paper":"https://pith.science/paper/NDWUQGAF","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2506.18183&json=true","fetch_graph":"https://pith.science/api/pith-number/NDWUQGAFHHJVTKEFKQZDO26X3J/graph.json","fetch_events":"https://pith.science/api/pith-number/NDWUQGAFHHJVTKEFKQZDO26X3J/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/NDWUQGAFHHJVTKEFKQZDO26X3J/action/timestamp_anchor","attest_storage":"https://pith.science/pith/NDWUQGAFHHJVTKEFKQZDO26X3J/action/storage_attestation","attest_author":"https://pith.science/pith/NDWUQGAFHHJVTKEFKQZDO26X3J/action/author_attestation","sign_citation":"https://pith.science/pith/NDWUQGAFHHJVTKEFKQZDO26X3J/action/citation_signature","submit_replication":"https://pith.science/pith/NDWUQGAFHHJVTKEFKQZDO26X3J/action/replication_record"}},"created_at":"2026-07-05T11:39:12.494999+00:00","updated_at":"2026-07-05T11:39:12.494999+00:00"}