{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2023:QLMKB2HH6QNGSX7KKF7B6DZVLZ","short_pith_number":"pith:QLMKB2HH","schema_version":"1.0","canonical_sha256":"82d8a0e8e7f41a695fea517e1f0f355e46327fd509acfae9303416c18a7840ca","source":{"kind":"arxiv","id":"2306.11698","version":5},"attestation_state":"computed","paper":{"title":"DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models","license":"http://creativecommons.org/licenses/by-sa/4.0/","headline":"","cross_cats":["cs.AI","cs.CR"],"primary_cat":"cs.CL","authors_text":"Bo Li, Boxin Wang, Chejian Xu, Chenhui Zhang, Chulin Xie, Dan Hendrycks, Dawn Song, Hengzhi Pei, Mantas Mazeika, Mintong Kang, Ritik Dutta, Rylan Schaeffer, Sang T. Truong, Sanmi Koyejo, Simran Arora, Weixin Chen, Yu Cheng, Zidi Xiong, Zinan Lin","submitted_at":"2023-06-20T17:24:23Z","abstract_excerpt":"Generative Pre-trained Transformer (GPT) models have exhibited exciting progress in their capabilities, capturing the interest of practitioners and the public alike. Yet, while the literature on the trustworthiness of GPT models remains limited, practitioners have proposed employing capable GPT models for sensitive applications such as healthcare and finance -- where mistakes can be costly. To this end, this work proposes a comprehensive trustworthiness evaluation for large language models with a focus on GPT-4 and GPT-3.5, considering diverse perspectives -- including toxicity, stereotype bia"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2306.11698","kind":"arxiv","version":5},"metadata":{"license":"http://creativecommons.org/licenses/by-sa/4.0/","primary_cat":"cs.CL","submitted_at":"2023-06-20T17:24:23Z","cross_cats_sorted":["cs.AI","cs.CR"],"title_canon_sha256":"83efe1b55f51fedc173bfc49a1ad7dc00baee5685f0735711b63e369b34cec65","abstract_canon_sha256":"0b21628c925d0292f81d4a17e2f1684bdb3c2cb00986fe92fa8faf51f0174935"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T07:49:32.155956Z","signature_b64":"mddOAv0VfyvkgiO6PRYhkLyCkAV1TbN8EOMCaSeGiXIW14ZPFoibi14F3Lp47dlCPVaVf7/AP0CLFZLbicRxDw==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"82d8a0e8e7f41a695fea517e1f0f355e46327fd509acfae9303416c18a7840ca","last_reissued_at":"2026-07-05T07:49:32.155546Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T07:49:32.155546Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models","license":"http://creativecommons.org/licenses/by-sa/4.0/","headline":"","cross_cats":["cs.AI","cs.CR"],"primary_cat":"cs.CL","authors_text":"Bo Li, Boxin Wang, Chejian Xu, Chenhui Zhang, Chulin Xie, Dan Hendrycks, Dawn Song, Hengzhi Pei, Mantas Mazeika, Mintong Kang, Ritik Dutta, Rylan Schaeffer, Sang T. Truong, Sanmi Koyejo, Simran Arora, Weixin Chen, Yu Cheng, Zidi Xiong, Zinan Lin","submitted_at":"2023-06-20T17:24:23Z","abstract_excerpt":"Generative Pre-trained Transformer (GPT) models have exhibited exciting progress in their capabilities, capturing the interest of practitioners and the public alike. Yet, while the literature on the trustworthiness of GPT models remains limited, practitioners have proposed employing capable GPT models for sensitive applications such as healthcare and finance -- where mistakes can be costly. To this end, this work proposes a comprehensive trustworthiness evaluation for large language models with a focus on GPT-4 and GPT-3.5, considering diverse perspectives -- including toxicity, stereotype bia"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2306.11698","kind":"arxiv","version":5},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2306.11698/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2306.11698","created_at":"2026-07-05T07:49:32.155604+00:00"},{"alias_kind":"arxiv_version","alias_value":"2306.11698v5","created_at":"2026-07-05T07:49:32.155604+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2306.11698","created_at":"2026-07-05T07:49:32.155604+00:00"},{"alias_kind":"pith_short_12","alias_value":"QLMKB2HH6QNG","created_at":"2026-07-05T07:49:32.155604+00:00"},{"alias_kind":"pith_short_16","alias_value":"QLMKB2HH6QNGSX7K","created_at":"2026-07-05T07:49:32.155604+00:00"},{"alias_kind":"pith_short_8","alias_value":"QLMKB2HH","created_at":"2026-07-05T07:49:32.155604+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":18,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.25782","citing_title":"Do Encoders Suffice? A Systematic Comparison of Encoder and Decoder Safety Judges for LLM Adversarial Evaluation","ref_index":11,"is_internal_anchor":false},{"citing_arxiv_id":"2606.09125","citing_title":"Unveiling Privacy Risks in Multi-modal Large Language Models: Task-specific Vulnerabilities and Mitigation Challenges","ref_index":20,"is_internal_anchor":false},{"citing_arxiv_id":"2605.22771","citing_title":"Reducing Political Manipulation with Consistency Training","ref_index":38,"is_internal_anchor":false},{"citing_arxiv_id":"2605.27763","citing_title":"A Paired Testing Protocol for Batch-Conditioned Refusal Robustness in LLM Serving","ref_index":22,"is_internal_anchor":false},{"citing_arxiv_id":"2409.18169","citing_title":"Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey","ref_index":154,"is_internal_anchor":false},{"citing_arxiv_id":"2605.22771","citing_title":"Reducing Political Manipulation with Consistency Training","ref_index":38,"is_internal_anchor":false},{"citing_arxiv_id":"2603.04459","citing_title":"Benchmark of Benchmarks: Unpacking Influence and Code Repository Quality in LLM Safety Benchmarks","ref_index":121,"is_internal_anchor":false},{"citing_arxiv_id":"2401.05561","citing_title":"TrustLLM: Trustworthiness in Large Language Models","ref_index":71,"is_internal_anchor":false},{"citing_arxiv_id":"2308.03825","citing_title":"\"Do Anything Now\": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models","ref_index":84,"is_internal_anchor":false},{"citing_arxiv_id":"2512.05439","citing_title":"BEAVER: An Efficient Deterministic LLM Verifier","ref_index":51,"is_internal_anchor":false},{"citing_arxiv_id":"2512.21110","citing_title":"Beyond Context: Large Language Models' Failure to Grasp Users' Intent","ref_index":2,"is_internal_anchor":false},{"citing_arxiv_id":"2309.10253","citing_title":"GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts","ref_index":60,"is_internal_anchor":false},{"citing_arxiv_id":"2402.17177","citing_title":"Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models","ref_index":110,"is_internal_anchor":false},{"citing_arxiv_id":"2605.04539","citing_title":"RLearner-LLM: Balancing Logical Grounding and Fluency in Large Language Models via Hybrid Direct Preference Optimization","ref_index":13,"is_internal_anchor":false},{"citing_arxiv_id":"2605.04539","citing_title":"RLearner-LLM: Balancing Logical Grounding and Fluency in Large Language Models via Hybrid Direct Preference Optimization","ref_index":13,"is_internal_anchor":false},{"citing_arxiv_id":"2605.05810","citing_title":"CXR-ContraBench: Benchmarking Negated-Option Attraction in Medical VLMs","ref_index":32,"is_internal_anchor":false},{"citing_arxiv_id":"2605.04539","citing_title":"RLearner-LLM: Balancing Logical Grounding and Fluency in Large Language Models via Hybrid Direct Preference Optimization","ref_index":13,"is_internal_anchor":false},{"citing_arxiv_id":"2604.15789","citing_title":"A Systematic Study of Training-Free Methods for Trustworthy Large Language Models","ref_index":49,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/QLMKB2HH6QNGSX7KKF7B6DZVLZ","json":"https://pith.science/pith/QLMKB2HH6QNGSX7KKF7B6DZVLZ.json","graph_json":"https://pith.science/api/pith-number/QLMKB2HH6QNGSX7KKF7B6DZVLZ/graph.json","events_json":"https://pith.science/api/pith-number/QLMKB2HH6QNGSX7KKF7B6DZVLZ/events.json","paper":"https://pith.science/paper/QLMKB2HH"},"agent_actions":{"view_html":"https://pith.science/pith/QLMKB2HH6QNGSX7KKF7B6DZVLZ","download_json":"https://pith.science/pith/QLMKB2HH6QNGSX7KKF7B6DZVLZ.json","view_paper":"https://pith.science/paper/QLMKB2HH","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2306.11698&json=true","fetch_graph":"https://pith.science/api/pith-number/QLMKB2HH6QNGSX7KKF7B6DZVLZ/graph.json","fetch_events":"https://pith.science/api/pith-number/QLMKB2HH6QNGSX7KKF7B6DZVLZ/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/QLMKB2HH6QNGSX7KKF7B6DZVLZ/action/timestamp_anchor","attest_storage":"https://pith.science/pith/QLMKB2HH6QNGSX7KKF7B6DZVLZ/action/storage_attestation","attest_author":"https://pith.science/pith/QLMKB2HH6QNGSX7KKF7B6DZVLZ/action/author_attestation","sign_citation":"https://pith.science/pith/QLMKB2HH6QNGSX7KKF7B6DZVLZ/action/citation_signature","submit_replication":"https://pith.science/pith/QLMKB2HH6QNGSX7KKF7B6DZVLZ/action/replication_record"}},"created_at":"2026-07-05T07:49:32.155604+00:00","updated_at":"2026-07-05T07:49:32.155604+00:00"}