{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2025:BUPIXQPXOQCTZJZJVXGWI6VDNW","short_pith_number":"pith:BUPIXQPX","schema_version":"1.0","canonical_sha256":"0d1e8bc1f774053ca729adcd647aa36db8cfdda65d3df4624ff67f5ac033ed32","source":{"kind":"arxiv","id":"2508.12733","version":2},"attestation_state":"computed","paper":{"title":"LinguaSafe: A Comprehensive Multilingual Safety Benchmark for Large Language Models","license":"http://creativecommons.org/licenses/by-nc-sa/4.0/","headline":"","cross_cats":["cs.AI"],"primary_cat":"cs.CL","authors_text":"Huacan Liu, Jiaxin Song, Jie Li, Lingyu Li, Meng Lingyu, Shixin Hong, Tianle Gu, Yan Teng, Yingchun Wang, Yixu Wang, Zhiyuan Ning","submitted_at":"2025-08-18T08:59:01Z","abstract_excerpt":"The widespread adoption and increasing prominence of large language models (LLMs) in global technologies necessitate a rigorous focus on ensuring their safety across a diverse range of linguistic and cultural contexts. The lack of a comprehensive evaluation and diverse data in existing multilingual safety evaluations for LLMs limits their effectiveness, hindering the development of robust multilingual safety alignment. To address this critical gap, we introduce LinguaSafe, a comprehensive multilingual safety benchmark crafted with meticulous attention to linguistic authenticity. The LinguaSafe"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2508.12733","kind":"arxiv","version":2},"metadata":{"license":"http://creativecommons.org/licenses/by-nc-sa/4.0/","primary_cat":"cs.CL","submitted_at":"2025-08-18T08:59:01Z","cross_cats_sorted":["cs.AI"],"title_canon_sha256":"04b6c8389959f5b351369fd5a75cab2dc2c218f3d2568efbee2e0ae55ea42ec4","abstract_canon_sha256":"33e6467a6fda8fbc6f1191751a87c391648949583bfd69272fd08e3ccd6cf15f"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T12:00:08.738675Z","signature_b64":"h0x5FpbUDe4Io6bItqrmgTpWyIC6LYMnui++Oonj4hTE2lGtMasGQ+GyNEJ4LOBbu24eZlNL5jzPhwITiDwwDA==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"0d1e8bc1f774053ca729adcd647aa36db8cfdda65d3df4624ff67f5ac033ed32","last_reissued_at":"2026-07-05T12:00:08.738193Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T12:00:08.738193Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"LinguaSafe: A Comprehensive Multilingual Safety Benchmark for Large Language Models","license":"http://creativecommons.org/licenses/by-nc-sa/4.0/","headline":"","cross_cats":["cs.AI"],"primary_cat":"cs.CL","authors_text":"Huacan Liu, Jiaxin Song, Jie Li, Lingyu Li, Meng Lingyu, Shixin Hong, Tianle Gu, Yan Teng, Yingchun Wang, Yixu Wang, Zhiyuan Ning","submitted_at":"2025-08-18T08:59:01Z","abstract_excerpt":"The widespread adoption and increasing prominence of large language models (LLMs) in global technologies necessitate a rigorous focus on ensuring their safety across a diverse range of linguistic and cultural contexts. The lack of a comprehensive evaluation and diverse data in existing multilingual safety evaluations for LLMs limits their effectiveness, hindering the development of robust multilingual safety alignment. To address this critical gap, we introduce LinguaSafe, a comprehensive multilingual safety benchmark crafted with meticulous attention to linguistic authenticity. The LinguaSafe"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2508.12733","kind":"arxiv","version":2},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2508.12733/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2508.12733","created_at":"2026-07-05T12:00:08.738248+00:00"},{"alias_kind":"arxiv_version","alias_value":"2508.12733v2","created_at":"2026-07-05T12:00:08.738248+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2508.12733","created_at":"2026-07-05T12:00:08.738248+00:00"},{"alias_kind":"pith_short_12","alias_value":"BUPIXQPXOQCT","created_at":"2026-07-05T12:00:08.738248+00:00"},{"alias_kind":"pith_short_16","alias_value":"BUPIXQPXOQCTZJZJ","created_at":"2026-07-05T12:00:08.738248+00:00"},{"alias_kind":"pith_short_8","alias_value":"BUPIXQPX","created_at":"2026-07-05T12:00:08.738248+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":9,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.07874","citing_title":"Safety is Contextual, LLM-Judges Are Not: Navigating the Rigid Priors of Evaluators","ref_index":59,"is_internal_anchor":false},{"citing_arxiv_id":"2606.11316","citing_title":"Sch\\\"utzen: Evaluating LLM Safety in Bulgarian and German Contexts","ref_index":13,"is_internal_anchor":false},{"citing_arxiv_id":"2606.19640","citing_title":"Creating Multilingual Mental Health Dialogue Datasets: Limits of Persona-Based Localization via Nationality and Language","ref_index":42,"is_internal_anchor":false},{"citing_arxiv_id":"2605.17173","citing_title":"Why Do Safety Guardrails Degrade Across Languages?","ref_index":16,"is_internal_anchor":false},{"citing_arxiv_id":"2605.14152","citing_title":"ROK-FORTRESS: Measuring the Effect of Geopolitical Transcreation for National Security and Public Safety","ref_index":13,"is_internal_anchor":false},{"citing_arxiv_id":"2605.14152","citing_title":"ROK-FORTRESS: Measuring the Effect of Geopolitical Transcreation for National Security and Public Safety","ref_index":14,"is_internal_anchor":false},{"citing_arxiv_id":"2605.05662","citing_title":"XL-SafetyBench: A Country-Grounded Cross-Cultural Benchmark for LLM Safety and Cultural Sensitivity","ref_index":30,"is_internal_anchor":false},{"citing_arxiv_id":"2605.00689","citing_title":"ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models","ref_index":13,"is_internal_anchor":false},{"citing_arxiv_id":"2605.06652","citing_title":"When No Benchmark Exists: Validating Comparative LLM Safety Scoring Without Ground-Truth Labels","ref_index":31,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/BUPIXQPXOQCTZJZJVXGWI6VDNW","json":"https://pith.science/pith/BUPIXQPXOQCTZJZJVXGWI6VDNW.json","graph_json":"https://pith.science/api/pith-number/BUPIXQPXOQCTZJZJVXGWI6VDNW/graph.json","events_json":"https://pith.science/api/pith-number/BUPIXQPXOQCTZJZJVXGWI6VDNW/events.json","paper":"https://pith.science/paper/BUPIXQPX"},"agent_actions":{"view_html":"https://pith.science/pith/BUPIXQPXOQCTZJZJVXGWI6VDNW","download_json":"https://pith.science/pith/BUPIXQPXOQCTZJZJVXGWI6VDNW.json","view_paper":"https://pith.science/paper/BUPIXQPX","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2508.12733&json=true","fetch_graph":"https://pith.science/api/pith-number/BUPIXQPXOQCTZJZJVXGWI6VDNW/graph.json","fetch_events":"https://pith.science/api/pith-number/BUPIXQPXOQCTZJZJVXGWI6VDNW/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/BUPIXQPXOQCTZJZJVXGWI6VDNW/action/timestamp_anchor","attest_storage":"https://pith.science/pith/BUPIXQPXOQCTZJZJVXGWI6VDNW/action/storage_attestation","attest_author":"https://pith.science/pith/BUPIXQPXOQCTZJZJVXGWI6VDNW/action/author_attestation","sign_citation":"https://pith.science/pith/BUPIXQPXOQCTZJZJVXGWI6VDNW/action/citation_signature","submit_replication":"https://pith.science/pith/BUPIXQPXOQCTZJZJVXGWI6VDNW/action/replication_record"}},"created_at":"2026-07-05T12:00:08.738248+00:00","updated_at":"2026-07-05T12:00:08.738248+00:00"}