{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2024:SJ53GEXFSFIZRO5HJUT2DEVTXI","short_pith_number":"pith:SJ53GEXF","schema_version":"1.0","canonical_sha256":"927bb312e5915198bba74d27a192b3ba2ba2a3d488fafeed5a86531dbf0d70ea","source":{"kind":"arxiv","id":"2410.03769","version":2},"attestation_state":"computed","paper":{"title":"SciSafeEval: A Comprehensive Benchmark for Safety Alignment of Large Language Models in Scientific Tasks","license":"http://creativecommons.org/licenses/by-nc-nd/4.0/","headline":"","cross_cats":["cs.AI","cs.CR"],"primary_cat":"cs.CL","authors_text":"Bin Wu, Chuangxin Chu, Haotian Huang, Huajun Chen, Jingyu Lu, Kai Ma, Keyan Ding, Mei Li, Qiang Zhang, Tianhao Li, Tianyu Zeng, Xingkai Wang, Xuejing Yuan, Yujia Zheng, Zuoxian Liu","submitted_at":"2024-10-02T16:34:48Z","abstract_excerpt":"Large language models (LLMs) have a transformative impact on a variety of scientific tasks across disciplines including biology, chemistry, medicine, and physics. However, ensuring the safety alignment of these models in scientific research remains an underexplored area, with existing benchmarks primarily focusing on textual content and overlooking key scientific representations such as molecular, protein, and genomic languages. Moreover, the safety mechanisms of LLMs in scientific tasks are insufficiently studied. To address these limitations, we introduce SciSafeEval, a comprehensive benchma"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2410.03769","kind":"arxiv","version":2},"metadata":{"license":"http://creativecommons.org/licenses/by-nc-nd/4.0/","primary_cat":"cs.CL","submitted_at":"2024-10-02T16:34:48Z","cross_cats_sorted":["cs.AI","cs.CR"],"title_canon_sha256":"e1ace83eda243e9621aadcd60fc37c34a466e3d9fba3f48d31dfa5e8ccc15f3f","abstract_canon_sha256":"a984bbbb48d58e8c0b18823580091eaa102f3f39bca97c042a559d30698f2f7c"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T09:49:45.688463Z","signature_b64":"FKxagUq/e5g5PP1glZ7BJnYzWdW6qkZ7GhDiQ3Mn7307SY4mzFERsfKI0pNwHz61rd3u4BcJo65Cf7ldFhXCBQ==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"927bb312e5915198bba74d27a192b3ba2ba2a3d488fafeed5a86531dbf0d70ea","last_reissued_at":"2026-07-05T09:49:45.687944Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T09:49:45.687944Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"SciSafeEval: A Comprehensive Benchmark for Safety Alignment of Large Language Models in Scientific Tasks","license":"http://creativecommons.org/licenses/by-nc-nd/4.0/","headline":"","cross_cats":["cs.AI","cs.CR"],"primary_cat":"cs.CL","authors_text":"Bin Wu, Chuangxin Chu, Haotian Huang, Huajun Chen, Jingyu Lu, Kai Ma, Keyan Ding, Mei Li, Qiang Zhang, Tianhao Li, Tianyu Zeng, Xingkai Wang, Xuejing Yuan, Yujia Zheng, Zuoxian Liu","submitted_at":"2024-10-02T16:34:48Z","abstract_excerpt":"Large language models (LLMs) have a transformative impact on a variety of scientific tasks across disciplines including biology, chemistry, medicine, and physics. However, ensuring the safety alignment of these models in scientific research remains an underexplored area, with existing benchmarks primarily focusing on textual content and overlooking key scientific representations such as molecular, protein, and genomic languages. Moreover, the safety mechanisms of LLMs in scientific tasks are insufficiently studied. To address these limitations, we introduce SciSafeEval, a comprehensive benchma"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2410.03769","kind":"arxiv","version":2},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2410.03769/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2410.03769","created_at":"2026-07-05T09:49:45.688002+00:00"},{"alias_kind":"arxiv_version","alias_value":"2410.03769v2","created_at":"2026-07-05T09:49:45.688002+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2410.03769","created_at":"2026-07-05T09:49:45.688002+00:00"},{"alias_kind":"pith_short_12","alias_value":"SJ53GEXFSFIZ","created_at":"2026-07-05T09:49:45.688002+00:00"},{"alias_kind":"pith_short_16","alias_value":"SJ53GEXFSFIZRO5H","created_at":"2026-07-05T09:49:45.688002+00:00"},{"alias_kind":"pith_short_8","alias_value":"SJ53GEXF","created_at":"2026-07-05T09:49:45.688002+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":5,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2605.30693","citing_title":"Triaging Threats to Specialized Guardrails","ref_index":24,"is_internal_anchor":false},{"citing_arxiv_id":"2605.22643","citing_title":"Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety","ref_index":46,"is_internal_anchor":false},{"citing_arxiv_id":"2605.22643","citing_title":"Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety","ref_index":46,"is_internal_anchor":false},{"citing_arxiv_id":"2604.01444","citing_title":"Cooking Up Risks: Benchmarking and Reducing Food Safety Risks in Large Language Models","ref_index":14,"is_internal_anchor":false},{"citing_arxiv_id":"2605.04992","citing_title":"You Snooze, You Lose: Automatic Safety Alignment Restoration through Neural Weight Translation","ref_index":82,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/SJ53GEXFSFIZRO5HJUT2DEVTXI","json":"https://pith.science/pith/SJ53GEXFSFIZRO5HJUT2DEVTXI.json","graph_json":"https://pith.science/api/pith-number/SJ53GEXFSFIZRO5HJUT2DEVTXI/graph.json","events_json":"https://pith.science/api/pith-number/SJ53GEXFSFIZRO5HJUT2DEVTXI/events.json","paper":"https://pith.science/paper/SJ53GEXF"},"agent_actions":{"view_html":"https://pith.science/pith/SJ53GEXFSFIZRO5HJUT2DEVTXI","download_json":"https://pith.science/pith/SJ53GEXFSFIZRO5HJUT2DEVTXI.json","view_paper":"https://pith.science/paper/SJ53GEXF","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2410.03769&json=true","fetch_graph":"https://pith.science/api/pith-number/SJ53GEXFSFIZRO5HJUT2DEVTXI/graph.json","fetch_events":"https://pith.science/api/pith-number/SJ53GEXFSFIZRO5HJUT2DEVTXI/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/SJ53GEXFSFIZRO5HJUT2DEVTXI/action/timestamp_anchor","attest_storage":"https://pith.science/pith/SJ53GEXFSFIZRO5HJUT2DEVTXI/action/storage_attestation","attest_author":"https://pith.science/pith/SJ53GEXFSFIZRO5HJUT2DEVTXI/action/author_attestation","sign_citation":"https://pith.science/pith/SJ53GEXFSFIZRO5HJUT2DEVTXI/action/citation_signature","submit_replication":"https://pith.science/pith/SJ53GEXFSFIZRO5HJUT2DEVTXI/action/replication_record"}},"created_at":"2026-07-05T09:49:45.688002+00:00","updated_at":"2026-07-05T09:49:45.688002+00:00"}