Pith. sign in

REVIEW 7 cited by

Attestable Audits: Verifiable AI Safety Benchmarks Using Trusted Execution Environments

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2506.23706 v1 pith:L3NBXDUK submitted 2025-06-30 cs.AI cs.CLcs.CR

Attestable Audits: Verifiable AI Safety Benchmarks Using Trusted Execution Environments

classification cs.AI cs.CLcs.CR
keywords benchmarksmodelattestableauditsenvironmentsexecutionsafetytrusted
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Benchmarks are important measures to evaluate safety and compliance of AI models at scale. However, they typically do not offer verifiable results and lack confidentiality for model IP and benchmark datasets. We propose Attestable Audits, which run inside Trusted Execution Environments and enable users to verify interaction with a compliant AI model. Our work protects sensitive data even when model provider and auditor do not trust each other. This addresses verification challenges raised in recent AI governance frameworks. We build a prototype demonstrating feasibility on typical audit benchmarks against Llama-3.1.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. A TEE-Based Architecture for Confidential and Dependable Process Attestation in Authorship Verification

    cs.CR 2026-02 unverdicted novelty 7.0

    First TEE-based architecture for continuous process attestation with hardware tamper resistance, tiered assurance levels, Markov-chain dependability modeling, and resilient protocol achieving over 99.5% evidence chain...

  2. A TEE-Based Architecture for Confidential and Dependable Process Attestation in Authorship Verification

    cs.CR 2026-02 reject novelty 6.0

    Continuous evidence collection for authorship verification is moved into SGX enclaves with an availability model, but the trust-inversion security proof is conditional on a conjectural leakage bound that the paper's o...

  3. Position: Behavioural Assurance Cannot Verify the Safety Claims Governance Now Demands

    cs.LG 2026-05 unverdicted novelty 5.0

    Behavioral assurance is structurally unable to verify the latent safety properties demanded by AI governance frameworks enacted 2019-2026.

  4. From Specification to Deployment: Empirical Evidence from a W3C VC + DID Trust Infrastructure for Autonomous Agents

    cs.CR 2026-05 unverdicted novelty 5.0

    MolTrust deploys a W3C VC+DID trust infrastructure for AI agents with kernel-layer authorization, cross-protocol interoperability, and layered Sybil resistance, operational since March 2026 across eight verticals.

  5. Making AI-Assisted Grant Evaluation Auditable without Exposing the Model

    cs.CR 2026-04 unverdicted novelty 4.0

    A TEE-based remote attestation system creates signed evaluation bundles that link input hashes, model measurements, and outputs to make AI grant reviews verifiable without revealing proprietary components.

  6. LLM Harms: A Taxonomy and Discussion

    cs.CY 2025-12 reject novelty 3.0

    Proposes a five-bucket taxonomy of LLM harms and calls for dynamic auditing, but the systematic review behind it is not reproducible and contains mismatched citations.

  7. LLM Harms: A Taxonomy and Discussion

    cs.CY 2025-12 unverdicted novelty 3.0

    This paper proposes a taxonomy of LLM harms in five categories and suggests mitigation strategies plus a dynamic auditing system for responsible development.