{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2026:PJW27EJTZY2DMHX332FQKTCT7N","short_pith_number":"pith:PJW27EJT","schema_version":"1.0","canonical_sha256":"7a6daf9133ce34361efbde8b054c53fb699642678948de336cd0f514930a5658","source":{"kind":"arxiv","id":"2604.22167","version":2},"attestation_state":"computed","paper":{"title":"Estimating Tail Risks in Language Model Output Distributions","license":"http://creativecommons.org/licenses/by/4.0/","headline":"Creating unsafe versions of language models allows accurate estimation of rare harmful outputs with 10-20 times fewer samples than brute force.","cross_cats":["cs.AI"],"primary_cat":"cs.LG","authors_text":"He He, Kathleen McKeown, Raghav Singhal, Rajesh Ranganath, Rico Angell, Zachary Horvitz, Zhou Yu","submitted_at":"2026-04-24T02:30:46Z","abstract_excerpt":"Language models are increasingly capable and are being rapidly deployed on a population-level scale. As a result, the safety of these models is increasingly high-stakes. Fortunately, advances in alignment have significantly reduced the likelihood of harmful model outputs. However, when models are queried billions of times in a day, even rare worst-case behaviors will occur. Current safety evaluations focus on capturing the distribution of inputs that yield harmful outputs. These evaluations disregard the probabilistic nature of models and their tail output behavior. To measure this tail risk, "},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2604.22167","kind":"arxiv","version":2},"metadata":{"license":"http://creativecommons.org/licenses/by/4.0/","primary_cat":"cs.LG","submitted_at":"2026-04-24T02:30:46Z","cross_cats_sorted":["cs.AI"],"title_canon_sha256":"a813d455ed69faccf8b16ac08f27602d5bf9aa1c137fce7e24934da90601bf49","abstract_canon_sha256":"2821c22efcacd59ccb382c8380d4a04172564d3467d64dbac35cdd3708524874"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-06-11T01:09:36.050729Z","signature_b64":"RFbbNan+1HJmTISwvskAuerhZMZwrdPrrUOjET3a1BWE/LOt/dFqMEhHPTLG673ZGgw5M+jDhmaDjSejTJbpBQ==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"7a6daf9133ce34361efbde8b054c53fb699642678948de336cd0f514930a5658","last_reissued_at":"2026-06-11T01:09:36.049698Z","signature_status":"signed_v1","first_computed_at":"2026-06-11T01:09:36.049698Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Estimating Tail Risks in Language Model Output Distributions","license":"http://creativecommons.org/licenses/by/4.0/","headline":"Creating unsafe versions of language models allows accurate estimation of rare harmful outputs with 10-20 times fewer samples than brute force.","cross_cats":["cs.AI"],"primary_cat":"cs.LG","authors_text":"He He, Kathleen McKeown, Raghav Singhal, Rajesh Ranganath, Rico Angell, Zachary Horvitz, Zhou Yu","submitted_at":"2026-04-24T02:30:46Z","abstract_excerpt":"Language models are increasingly capable and are being rapidly deployed on a population-level scale. As a result, the safety of these models is increasingly high-stakes. Fortunately, advances in alignment have significantly reduced the likelihood of harmful model outputs. However, when models are queried billions of times in a day, even rare worst-case behaviors will occur. Current safety evaluations focus on capturing the distribution of inputs that yield harmful outputs. These evaluations disregard the probabilistic nature of models and their tail output behavior. To measure this tail risk, "},"claims":{"count":4,"items":[{"kind":"strongest_claim","text":"On benchmarks measuring misuse and misalignment, these estimates match brute-force Monte Carlo estimates using 10-20x fewer samples. For example, we can estimate probability of harmful outputs on the order of 10^-4 with just 500 samples.","source":"verdict.strongest_claim","status":"machine_extracted","claim_id":"C1","attestation":"unclaimed"},{"kind":"weakest_assumption","text":"That unsafe versions of the target model can be constructed such that importance sampling yields unbiased estimates of the original model's harmful output probabilities without introducing systematic bias from the modification process.","source":"verdict.weakest_assumption","status":"machine_extracted","claim_id":"C2","attestation":"unclaimed"},{"kind":"one_line_summary","text":"Importance sampling with unsafe model variants estimates tail probabilities of harmful language model outputs using 10-20x fewer samples than brute-force Monte Carlo.","source":"verdict.one_line_summary","status":"machine_extracted","claim_id":"C3","attestation":"unclaimed"},{"kind":"headline","text":"Creating unsafe versions of language models allows accurate estimation of rare harmful outputs with 10-20 times fewer samples than brute force.","source":"verdict.pith_extraction.headline","status":"machine_extracted","claim_id":"C4","attestation":"unclaimed"}],"snapshot_sha256":"295b40bce66dc5bd2db439f7b9a17bceb54c4d392973f2f9922275e66ab235f3"},"source":{"id":"2604.22167","kind":"arxiv","version":2},"verdict":{"id":"2c014260-a889-43ff-8305-060b2cc053f1","model_set":{"reader":"grok-4.3"},"created_at":"2026-05-08T12:23:01.623617Z","strongest_claim":"On benchmarks measuring misuse and misalignment, these estimates match brute-force Monte Carlo estimates using 10-20x fewer samples. For example, we can estimate probability of harmful outputs on the order of 10^-4 with just 500 samples.","one_line_summary":"Importance sampling with unsafe model variants estimates tail probabilities of harmful language model outputs using 10-20x fewer samples than brute-force Monte Carlo.","pipeline_version":"pith-pipeline@v0.9.0","weakest_assumption":"That unsafe versions of the target model can be constructed such that importance sampling yields unbiased estimates of the original model's harmful output probabilities without introducing systematic bias from the modification process.","pith_extraction_headline":"Creating unsafe versions of language models allows accurate estimation of rare harmful outputs with 10-20 times fewer samples than brute force."},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2604.22167/integrity.json","findings":[],"available":true,"detectors_run":[{"name":"ai_meta_artifact","ran_at":"2026-05-21T11:35:40.883991Z","status":"completed","version":"1.0.0","findings_count":0},{"name":"doi_compliance","ran_at":"2026-05-20T00:13:19.431935Z","status":"completed","version":"1.0.0","findings_count":0}],"snapshot_sha256":"3ce65ca15016b4bf4a37b88797955d29777ef48ea0ca74374281e7198110e323"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2604.22167","created_at":"2026-06-11T01:09:36.049832+00:00"},{"alias_kind":"arxiv_version","alias_value":"2604.22167v2","created_at":"2026-06-11T01:09:36.049832+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2604.22167","created_at":"2026-06-11T01:09:36.049832+00:00"},{"alias_kind":"pith_short_12","alias_value":"PJW27EJTZY2D","created_at":"2026-06-11T01:09:36.049832+00:00"},{"alias_kind":"pith_short_16","alias_value":"PJW27EJTZY2DMHX3","created_at":"2026-06-11T01:09:36.049832+00:00"},{"alias_kind":"pith_short_8","alias_value":"PJW27EJT","created_at":"2026-06-11T01:09:36.049832+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":1,"internal_anchor_count":1,"sample":[{"citing_arxiv_id":"2607.07184","citing_title":"Predicting LLM Safety Before Release by Simulating Deployment","ref_index":18,"is_internal_anchor":true}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/PJW27EJTZY2DMHX332FQKTCT7N","json":"https://pith.science/pith/PJW27EJTZY2DMHX332FQKTCT7N.json","graph_json":"https://pith.science/api/pith-number/PJW27EJTZY2DMHX332FQKTCT7N/graph.json","events_json":"https://pith.science/api/pith-number/PJW27EJTZY2DMHX332FQKTCT7N/events.json","paper":"https://pith.science/paper/PJW27EJT"},"agent_actions":{"view_html":"https://pith.science/pith/PJW27EJTZY2DMHX332FQKTCT7N","download_json":"https://pith.science/pith/PJW27EJTZY2DMHX332FQKTCT7N.json","view_paper":"https://pith.science/paper/PJW27EJT","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2604.22167&json=true","fetch_graph":"https://pith.science/api/pith-number/PJW27EJTZY2DMHX332FQKTCT7N/graph.json","fetch_events":"https://pith.science/api/pith-number/PJW27EJTZY2DMHX332FQKTCT7N/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/PJW27EJTZY2DMHX332FQKTCT7N/action/timestamp_anchor","attest_storage":"https://pith.science/pith/PJW27EJTZY2DMHX332FQKTCT7N/action/storage_attestation","attest_author":"https://pith.science/pith/PJW27EJTZY2DMHX332FQKTCT7N/action/author_attestation","sign_citation":"https://pith.science/pith/PJW27EJTZY2DMHX332FQKTCT7N/action/citation_signature","submit_replication":"https://pith.science/pith/PJW27EJTZY2DMHX332FQKTCT7N/action/replication_record"}},"created_at":"2026-06-11T01:09:36.049832+00:00","updated_at":"2026-06-11T01:09:36.049832+00:00"}