Pith. sign in

REVIEW 1 cited by

Comet: A Communication-efficient and Performant Approximation for Private Transformer Inference

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.17485 v2 pith:PZ4W5FFO submitted 2024-05-24 cs.LG cs.AIcs.CR

classification cs.LGcs.AIcs.CR
keywords communicationapproximationcometinferencemodelsintroducemethodmodel
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

The prevalent use of Transformer-like models, exemplified by ChatGPT in modern language processing applications, underscores the critical need for enabling private inference essential for many cloud-based services reliant on such models. However, current privacy-preserving frameworks impose significant communication burden, especially for non-linear computation in Transformer model. In this paper, we introduce a novel plug-in method Comet to effectively reduce the communication cost without compromising the inference performance. We second introduce an efficient approximation method to eliminate the heavy communication in finding good initial approximation. We evaluate our Comet on Bert and RoBERTa models with GLUE benchmark datasets, showing up to 3.9$\times$ less communication and 3.5$\times$ speedups while keep competitive model performance compared to the prior art.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SETransformer: A Hybrid Attention-Based Architecture for Robust Human Activity Recognition

    cs.LG 2025-05 reject novelty 2.0 of 10

    SETransformer combines a Transformer encoder, channel attention, and attention pooling for WISDM activity recognition, but the architecture is permutation-invariant and the reported comparison omits the model itself.

Pith tools