{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2024:RT7RA3Y46W2UGSJ7JH75F3WBYC","short_pith_number":"pith:RT7RA3Y4","schema_version":"1.0","canonical_sha256":"8cff106f1cf5b543493f49ffd2eec1c0ac4efbd4e6a743b5c562b88a692a4b57","source":{"kind":"arxiv","id":"2408.12787","version":2},"attestation_state":"computed","paper":{"title":"LLM-PBE: Assessing Data Privacy in Large Language Models","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.AI"],"primary_cat":"cs.CR","authors_text":"Bingsheng He, Bo Li, Chulin Xie, Dan Hendrycks, Dawn Song, Jeffrey Tan, Junyi Hou, Junyuan Hong, Qinbin Li, Rachel Xin, Xavier Yin, Zhangyang Wang, Zhun Wang","submitted_at":"2024-08-23T01:37:29Z","abstract_excerpt":"Large Language Models (LLMs) have become integral to numerous domains, significantly advancing applications in data management, mining, and analysis. Their profound capabilities in processing and interpreting complex language data, however, bring to light pressing concerns regarding data privacy, especially the risk of unintentional training data leakage. Despite the critical nature of this issue, there has been no existing literature to offer a comprehensive assessment of data privacy risks in LLMs. Addressing this gap, our paper introduces LLM-PBE, a toolkit crafted specifically for the syst"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2408.12787","kind":"arxiv","version":2},"metadata":{"license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","primary_cat":"cs.CR","submitted_at":"2024-08-23T01:37:29Z","cross_cats_sorted":["cs.AI"],"title_canon_sha256":"b3c0a04a4e76ac56a14e97fea1f3f8ab81ef781c851cf2e0650140e1cc608bd3","abstract_canon_sha256":"78a6008709e1b2f59e1699277a095374d3edd63fc5d3d8b9d50374171d23c31a"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T09:03:55.597469Z","signature_b64":"X/lPjMoLS/L3QeLsNU7G3Fw8ae4EHLzmxSddBjHgctC5j0WyycH16u1DqDTrnr/mEHI1Mqs3+VDYvREJzwqwAQ==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"8cff106f1cf5b543493f49ffd2eec1c0ac4efbd4e6a743b5c562b88a692a4b57","last_reissued_at":"2026-07-05T09:03:55.596973Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T09:03:55.596973Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"LLM-PBE: Assessing Data Privacy in Large Language Models","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.AI"],"primary_cat":"cs.CR","authors_text":"Bingsheng He, Bo Li, Chulin Xie, Dan Hendrycks, Dawn Song, Jeffrey Tan, Junyi Hou, Junyuan Hong, Qinbin Li, Rachel Xin, Xavier Yin, Zhangyang Wang, Zhun Wang","submitted_at":"2024-08-23T01:37:29Z","abstract_excerpt":"Large Language Models (LLMs) have become integral to numerous domains, significantly advancing applications in data management, mining, and analysis. Their profound capabilities in processing and interpreting complex language data, however, bring to light pressing concerns regarding data privacy, especially the risk of unintentional training data leakage. Despite the critical nature of this issue, there has been no existing literature to offer a comprehensive assessment of data privacy risks in LLMs. Addressing this gap, our paper introduces LLM-PBE, a toolkit crafted specifically for the syst"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2408.12787","kind":"arxiv","version":2},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2408.12787/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2408.12787","created_at":"2026-07-05T09:03:55.597043+00:00"},{"alias_kind":"arxiv_version","alias_value":"2408.12787v2","created_at":"2026-07-05T09:03:55.597043+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2408.12787","created_at":"2026-07-05T09:03:55.597043+00:00"},{"alias_kind":"pith_short_12","alias_value":"RT7RA3Y46W2U","created_at":"2026-07-05T09:03:55.597043+00:00"},{"alias_kind":"pith_short_16","alias_value":"RT7RA3Y46W2UGSJ7","created_at":"2026-07-05T09:03:55.597043+00:00"},{"alias_kind":"pith_short_8","alias_value":"RT7RA3Y4","created_at":"2026-07-05T09:03:55.597043+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":5,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2607.01019","citing_title":"Toward a Unified Security and Privacy Framework for AI-Native 6G Networks","ref_index":58,"is_internal_anchor":false},{"citing_arxiv_id":"2606.00991","citing_title":"Large Language Models in Transportation Systems Management and Operations: From Text Reasoning to Multi-modal Decision Support","ref_index":114,"is_internal_anchor":false},{"citing_arxiv_id":"2604.02501","citing_title":"ECG Foundation Models and Medical LLMs for Agentic Cardiovascular Intelligence at the Edge: A Review and Outlook","ref_index":142,"is_internal_anchor":false},{"citing_arxiv_id":"2605.05340","citing_title":"How Far Are VLMs from Privacy Awareness in the Physical World? An Empirical Study","ref_index":17,"is_internal_anchor":false},{"citing_arxiv_id":"2605.05340","citing_title":"How Far Are VLMs from Privacy Awareness in the Physical World? An Empirical Study","ref_index":17,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/RT7RA3Y46W2UGSJ7JH75F3WBYC","json":"https://pith.science/pith/RT7RA3Y46W2UGSJ7JH75F3WBYC.json","graph_json":"https://pith.science/api/pith-number/RT7RA3Y46W2UGSJ7JH75F3WBYC/graph.json","events_json":"https://pith.science/api/pith-number/RT7RA3Y46W2UGSJ7JH75F3WBYC/events.json","paper":"https://pith.science/paper/RT7RA3Y4"},"agent_actions":{"view_html":"https://pith.science/pith/RT7RA3Y46W2UGSJ7JH75F3WBYC","download_json":"https://pith.science/pith/RT7RA3Y46W2UGSJ7JH75F3WBYC.json","view_paper":"https://pith.science/paper/RT7RA3Y4","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2408.12787&json=true","fetch_graph":"https://pith.science/api/pith-number/RT7RA3Y46W2UGSJ7JH75F3WBYC/graph.json","fetch_events":"https://pith.science/api/pith-number/RT7RA3Y46W2UGSJ7JH75F3WBYC/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/RT7RA3Y46W2UGSJ7JH75F3WBYC/action/timestamp_anchor","attest_storage":"https://pith.science/pith/RT7RA3Y46W2UGSJ7JH75F3WBYC/action/storage_attestation","attest_author":"https://pith.science/pith/RT7RA3Y46W2UGSJ7JH75F3WBYC/action/author_attestation","sign_citation":"https://pith.science/pith/RT7RA3Y46W2UGSJ7JH75F3WBYC/action/citation_signature","submit_replication":"https://pith.science/pith/RT7RA3Y46W2UGSJ7JH75F3WBYC/action/replication_record"}},"created_at":"2026-07-05T09:03:55.597043+00:00","updated_at":"2026-07-05T09:03:55.597043+00:00"}