{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2024:5K2KI25L5F3KZVZXHAQCLN242L","short_pith_number":"pith:5K2KI25L","schema_version":"1.0","canonical_sha256":"eab4a46babe976acd737382025b75cd2ceb1de531e0572d29e8e23d3a3fc399d","source":{"kind":"arxiv","id":"2407.17436","version":2},"attestation_state":"computed","paper":{"title":"AIR-Bench 2024: A Safety Benchmark Based on Risk Categories from Regulations and Policies","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.AI"],"primary_cat":"cs.CY","authors_text":"Andy Zhou, Bo Li, Dawn Song, Jeffrey Ziwei Tan, Kevin Klyman, Minzhou Pan, Percy Liang, Ruoxi Jia, Yifan Mai, Yi Zeng, Yuheng Tu, Yu Yang","submitted_at":"2024-07-11T21:16:48Z","abstract_excerpt":"Foundation models (FMs) provide societal benefits but also amplify risks. Governments, companies, and researchers have proposed regulatory frameworks, acceptable use policies, and safety benchmarks in response. However, existing public benchmarks often define safety categories based on previous literature, intuitions, or common sense, leading to disjointed sets of categories for risks specified in recent regulations and policies, which makes it challenging to evaluate and compare FMs across these benchmarks. To bridge this gap, we introduce AIR-Bench 2024, the first AI safety benchmark aligned"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2407.17436","kind":"arxiv","version":2},"metadata":{"license":"http://creativecommons.org/licenses/by/4.0/","primary_cat":"cs.CY","submitted_at":"2024-07-11T21:16:48Z","cross_cats_sorted":["cs.AI"],"title_canon_sha256":"7d30aea6e567ba719055149cae11e11635fd67e538f29874b5d13357870253e9","abstract_canon_sha256":"d2d7fcc144b13058ea401db6dfb56128106491470af2e66a0d1309c5d407ab67"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T08:52:37.097680Z","signature_b64":"2obxrpSxjH5TIh5CRoAYeRTGKlNbdl2rvaX8DtJ5IRXt3yrOyW6ItBc9Tdwx+9hsgz7QtWJQarDFR+sdvD40Dg==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"eab4a46babe976acd737382025b75cd2ceb1de531e0572d29e8e23d3a3fc399d","last_reissued_at":"2026-07-05T08:52:37.097131Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T08:52:37.097131Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"AIR-Bench 2024: A Safety Benchmark Based on Risk Categories from Regulations and Policies","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.AI"],"primary_cat":"cs.CY","authors_text":"Andy Zhou, Bo Li, Dawn Song, Jeffrey Ziwei Tan, Kevin Klyman, Minzhou Pan, Percy Liang, Ruoxi Jia, Yifan Mai, Yi Zeng, Yuheng Tu, Yu Yang","submitted_at":"2024-07-11T21:16:48Z","abstract_excerpt":"Foundation models (FMs) provide societal benefits but also amplify risks. Governments, companies, and researchers have proposed regulatory frameworks, acceptable use policies, and safety benchmarks in response. However, existing public benchmarks often define safety categories based on previous literature, intuitions, or common sense, leading to disjointed sets of categories for risks specified in recent regulations and policies, which makes it challenging to evaluate and compare FMs across these benchmarks. To bridge this gap, we introduce AIR-Bench 2024, the first AI safety benchmark aligned"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2407.17436","kind":"arxiv","version":2},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2407.17436/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2407.17436","created_at":"2026-07-05T08:52:37.097197+00:00"},{"alias_kind":"arxiv_version","alias_value":"2407.17436v2","created_at":"2026-07-05T08:52:37.097197+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2407.17436","created_at":"2026-07-05T08:52:37.097197+00:00"},{"alias_kind":"pith_short_12","alias_value":"5K2KI25L5F3K","created_at":"2026-07-05T08:52:37.097197+00:00"},{"alias_kind":"pith_short_16","alias_value":"5K2KI25L5F3KZVZX","created_at":"2026-07-05T08:52:37.097197+00:00"},{"alias_kind":"pith_short_8","alias_value":"5K2KI25L","created_at":"2026-07-05T08:52:37.097197+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":14,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.19887","citing_title":"FinRED: An Expert-Guided Benchmark Generation and Evaluation Framework for Financial LLM Red-Teaming","ref_index":27,"is_internal_anchor":false},{"citing_arxiv_id":"2606.09178","citing_title":"Culturally-Adapted Red-Teaming Across East and Southeast Asian Contexts: A Methodological and Comparative Analysis","ref_index":17,"is_internal_anchor":false},{"citing_arxiv_id":"2606.08376","citing_title":"RiskNet: A large-scale dataset of AI risk incidents from news with alignment and multi-dimensional annotations","ref_index":14,"is_internal_anchor":false},{"citing_arxiv_id":"2606.03648","citing_title":"Safety Measurements for Fine-tuned LLMs Should be Grounded in Capability","ref_index":43,"is_internal_anchor":false},{"citing_arxiv_id":"2606.20626","citing_title":"Efficient Safety Benchmarking via Item Response Theory","ref_index":18,"is_internal_anchor":false},{"citing_arxiv_id":"2606.01481","citing_title":"SafeGen-Bench: Benchmarking Safety in Image-Conditioned Text-to-Video Generation","ref_index":62,"is_internal_anchor":false},{"citing_arxiv_id":"2606.04394","citing_title":"Beyond Single-Policy: Evaluating Composed Organization-Specific Policy Alignment in LLM Chatbots","ref_index":56,"is_internal_anchor":false},{"citing_arxiv_id":"2605.22643","citing_title":"Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety","ref_index":91,"is_internal_anchor":false},{"citing_arxiv_id":"2605.22643","citing_title":"Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety","ref_index":91,"is_internal_anchor":false},{"citing_arxiv_id":"2603.13933","citing_title":"OmniCompliance-100K: A Multi-Domain, Rule-Grounded, Real-World Safety Compliance Dataset","ref_index":17,"is_internal_anchor":false},{"citing_arxiv_id":"2603.20633","citing_title":"Seed1.8 Model Card: Towards Generalized Real-World Agency","ref_index":88,"is_internal_anchor":false},{"citing_arxiv_id":"2605.14152","citing_title":"ROK-FORTRESS: Measuring the Effect of Geopolitical Transcreation for National Security and Public Safety","ref_index":24,"is_internal_anchor":false},{"citing_arxiv_id":"2605.06213","citing_title":"Beyond Fixed Benchmarks and Worst-Case Attacks: Dynamic Boundary Evaluation for Language Models","ref_index":25,"is_internal_anchor":false},{"citing_arxiv_id":"2605.03179","citing_title":"A Validated Prompt Bank for Malicious Code Generation: Separating Executable Weapons from Security Knowledge in 1,554 Consensus-Labeled Prompts","ref_index":25,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/5K2KI25L5F3KZVZXHAQCLN242L","json":"https://pith.science/pith/5K2KI25L5F3KZVZXHAQCLN242L.json","graph_json":"https://pith.science/api/pith-number/5K2KI25L5F3KZVZXHAQCLN242L/graph.json","events_json":"https://pith.science/api/pith-number/5K2KI25L5F3KZVZXHAQCLN242L/events.json","paper":"https://pith.science/paper/5K2KI25L"},"agent_actions":{"view_html":"https://pith.science/pith/5K2KI25L5F3KZVZXHAQCLN242L","download_json":"https://pith.science/pith/5K2KI25L5F3KZVZXHAQCLN242L.json","view_paper":"https://pith.science/paper/5K2KI25L","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2407.17436&json=true","fetch_graph":"https://pith.science/api/pith-number/5K2KI25L5F3KZVZXHAQCLN242L/graph.json","fetch_events":"https://pith.science/api/pith-number/5K2KI25L5F3KZVZXHAQCLN242L/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/5K2KI25L5F3KZVZXHAQCLN242L/action/timestamp_anchor","attest_storage":"https://pith.science/pith/5K2KI25L5F3KZVZXHAQCLN242L/action/storage_attestation","attest_author":"https://pith.science/pith/5K2KI25L5F3KZVZXHAQCLN242L/action/author_attestation","sign_citation":"https://pith.science/pith/5K2KI25L5F3KZVZXHAQCLN242L/action/citation_signature","submit_replication":"https://pith.science/pith/5K2KI25L5F3KZVZXHAQCLN242L/action/replication_record"}},"created_at":"2026-07-05T08:52:37.097197+00:00","updated_at":"2026-07-05T08:52:37.097197+00:00"}