Pith. sign in

REVIEW 2 cited by

Ascend-CC: Confidential Computing on Heterogeneous NPU for Emerging Generative AI Workloads

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.11888 v1 pith:47R76QVL submitted 2024-07-16 cs.CR

classification cs.CR
keywords ascend-ccdatahostmodeladdresscomputingconfidentialexisting
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Cloud workloads have dominated generative AI based on large language models (LLM). Specialized hardware accelerators, such as GPUs, NPUs, and TPUs, play a key role in AI adoption due to their superior performance over general-purpose CPUs. The AI models and the data are often highly sensitive and come from mutually distrusting parties. Existing CPU-based TEEs such as Intel SGX or AMD SEV do not provide sufficient protection. Device-centric TEEs like Nvidia-CC only address tightly coupled CPU-GPU systems with a proprietary solution requiring TEE on the host CPU side. On the other hand, existing academic proposals are tailored toward specific CPU-TEE platforms. To address this gap, we propose Ascend-CC, a confidential computing architecture based on discrete NPU devices that requires no trust in the host system. Ascend-CC provides strong security by ensuring data and model encryption that protects not only the data but also the model parameters and operator binaries. Ascend-CC uses delegation-based memory semantics to ensure isolation from the host software stack, and task attestation provides strong model integrity guarantees. Our Ascend-CC implementation and evaluation with state-of-the-art LLMs such as Llama2 and Llama3 shows that Ascend-CC introduces minimal overhead with no changes in the AI software stack.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. BOLT: Bandwidth-Optimized Lightning-Fast Oblivious Map powered by Secure HBM Accelerators

    cs.CR 2025-09 conditional novelty 8.0 of 10

    BOLT is an FPGA-based oblivious map accelerator that uses isolated HBM as an unobservable cache to cut the bandwidth blow-up of key-value lookups to O(1)+O(log log N).

  2. Verifying International Agreements on AI: Six Layers of Verification for Rules on Large-Scale AI Development and Deployment

    cs.CY 2025-07 conditional novelty 6.0 of 10

    Countries could verify compliance with international AI agreements through six redundant verification layers, provided the report's listed hardware and analysis challenges are solved.

Pith tools