Pith. sign in

REVIEW 5 cited by

East: Efficient and Accurate Secure Transformer Framework for Inference

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2308.09923 v1 pith:IVZUKBEG submitted 2023-08-19 cs.CR cs.AIcs.LG

classification cs.CRcs.AIcs.LG
keywords inferencesecuretransformertimeseastfunctionsprotocolsaccurate
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
abstract

Transformer has been successfully used in practical applications, such as ChatGPT, due to its powerful advantages. However, users' input is leaked to the model provider during the service. With people's attention to privacy, privacy-preserving Transformer inference is on the demand of such services. Secure protocols for non-linear functions are crucial in privacy-preserving Transformer inference, which are not well studied. Thus, designing practical secure protocols for non-linear functions is hard but significant to model performance. In this work, we propose a framework \emph{East} to enable efficient and accurate secure Transformer inference. Firstly, we propose a new oblivious piecewise polynomial evaluation algorithm and apply it to the activation functions, which reduces the runtime and communication of GELU by over 1.5$\times$ and 2.5$\times$, compared to prior arts. Secondly, the secure protocols for softmax and layer normalization are carefully designed to faithfully maintain the desired functionality. Thirdly, several optimizations are conducted in detail to enhance the overall efficiency. We applied \emph{East} to BERT and the results show that the inference accuracy remains consistent with the plaintext inference without fine-tuning. Compared to Iron, we achieve about 1.8$\times$ lower communication within 1.2$\times$ lower runtime.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Fast Plaintext-Ciphertext Matrix Multiplication from Additively Homomorphic Encryption

    cs.CR 2025-04 conditional novelty 6.0 of 10

    Cussen's compression-reconstruction algorithm, originally for plaintext matrix multiplication, is applied to plaintext-ciphertext matrix multiplication with additively homomorphic encryption, giving up to an order of ...

  2. CENTAUR: Bridging the Impossible Trinity of Privacy, Efficiency, and Performance in Privacy-Preserving Transformer Inference

    cs.LG 2024-12 conditional novelty 6.0 of 10

    CENTAUR speeds up privacy-preserving Transformer inference by permuting model weights and secret-sharing the input, at the cost of replacing provable privacy with empirical attack resistance.

  3. LLM Access Shield: Domain-Specific LLM Framework for Privacy Policy Compliance

    cs.CR 2025-05 conditional novelty 5.0 of 10

    An enterprise proxy that detects sensitive data in LLM prompts with a fine-tuned small model and replaces it with format-preserving encryption.

  4. Private Transformer Inference in MLaaS: A Survey

    cs.CR 2025-05 conditional novelty 4.0 of 10

    A structured survey of private transformer inference, comparing MPC- and HE-based methods and showing non-linear layers dominate overhead.

  5. A Survey on Private Transformer Inference

    cs.CR 2024-12 reject

    A literature survey on private transformer inference that is too incomplete to support its promised comparisons and evaluation guidelines.

Pith tools