REVIEW 4 major objections 6 minor 1 cited by
A Survey on Private Transformer Inference
T0 review · 4 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read Private transformer inference has two bottlenecks — large matrix multiplications and non-linear functions — and the surveyed systems differ mainly by cryptographic setup and which layer they optimize.
desk verdict A well-organized PTI survey that is currently too incomplete to use: the promised evaluation guidelines are missing and the comparison tables mislabel several systems. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying mechanism is a layer-wise decomposition of a transformer encoder into linear operations (matrix multiplications in attention and feed-forward layers) and non-linear operations (Softmax, GeLU, and LayerNorm), cross-classified by cryptographic setup (two-party, two-party with a trusted dealer, three-party). Within that grid, the load-bearing objects are: secret-sharing schemes with Beaver triples (precomputed shared randomness that turns secure multiplication into one communication round) or re-sharing (refreshing shares by exchanging noisy local results) for secure multiplication; homomorphic-encryption schemes (BFV, CKKS, RNS-CKKS) with ciphertext packing and SIMD operations for matrix multiplication; and approximation techniques for non-linear functions, including low-degree polynomials, Taylor/Maclaurin/Chebyshev/Fourier series, and look-up tables. These mechanisms let the survey compare systems along the same axes: each system's reported communication volume, runtime, and accuracy loss are tied to which mechanisms it uses and which layer it optimizes.
What would settle it
A reader can settle the central claim by cross-checking every row of Tables 3, 4, 8, 9, 10, and 12 against the cited papers' own reported numbers and reference lists. The draft already shows two misattributions (BOLT appears with marker [26] in Table 3 and [43] in Table 4; Curl appears with marker [17] instead of [50]), so a systematic verification of all rows would show whether the survey's comparisons are reliable.
Extended reading notes
Core claim
The paper's central claim is that the current state of private transformer inference is best understood not as a contest between homomorphic encryption and secure multi-party computation, but as a layered design problem: each transformer component imposes a different cryptographic cost, and each system can be described by which layer it optimizes and under which setup. The paper argues that in two-party setups, large matrix multiplications are a dominant bottleneck because secure multiplication requires extra privacy protection; in dealer-assisted and three-party setups, that bottleneck moves to non-linear layers, which now account for most of the runtime in both MPC and HE systems. It further claims that accuracy preservation is achieved mainly by replacing non-linear functions with crypto-friendly approximations (low-degree polynomials, Taylor, Maclaurin, Chebyshev, or Fourier series, and look-up tables) and then recovering accuracy through knowledge distillation. The paper concludes by proposing evaluation guidelines, arguing that fair comparison requires reporting communication volume, runtime, accuracy loss, and the security model together.
Load-bearing premise
The survey's usefulness depends on its tables faithfully representing the cited systems' security models, runtimes, and reference markers; if those entries are wrong or misattributed, the comparative conclusions drawn from the survey would be misleading.
Editorial extensions
If this is right
- If the survey's classification holds, future systems in the two-party setting should focus on jointly optimizing MatMul and non-linear layers, since neither alone determines end-to-end cost.
- If the per-layer breakdowns are accurate, optimizing Softmax, GeLU, and LayerNorm will produce larger end-to-end gains than further MatMul speedups for dealer and three-party systems.
- If the evaluation guidelines are adopted, reported numbers across studies become comparable, because current tables differ in network bandwidth, input size, and setup, making cross-paper comparison unreliable without normalization.
- If the approximation-and-distillation trend continues, accuracy preservation will carry an extra training cost that must be included in any resource comparison, not just inference runtime.
Reading between the lines
- Not stated in the paper, but a reader could infer that the setup choice encodes a trust assumption: two-party systems avoid extra trust but pay for matrix multiplication, while dealer and three-party systems push that cost onto an assumed-honest helper; the two-party direction is therefore the harder test for the field.
- Not stated in the paper, the per-layer tables imply a testable ordering under identical network conditions: an HE-only system will show near-zero communication but the longest runtime, a two-party hybrid will sit in the middle, and a three-party or dealer system will show the lowest runtime only if a helper is available.
- Not stated in the paper, the proposed evaluation guidelines could be turned into a community benchmark that normalizes communication per token, per-layer runtime, and GLUE accuracy loss, which would convert the survey's qualitative comparisons into reproducible numbers.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript is a survey of private transformer inference (PTI). It covers background on transformer architecture and cryptographic primitives (MPC, HE), reviews roughly thirty PTI systems from 2022–2024, categorizes them by setup (2PC, 2PC-Dealer, 3PC), and discusses protocols for linear layers (MatMul) and non-linear layers (Softmax, GeLU, LayerNorm). The abstract and introduction promise three contributions: a comprehensive review, a breakdown of challenges and typical solutions, and proposed evaluation guidelines for resource efficiency and privacy guarantees. The review portion is structured around comparison tables and per-layer discussions, but the manuscript is incomplete: Section 8 is empty, several cross-references appear as unresolved 'Section ??', and the promised evaluation guidelines do not materialize in the text.
Significance. If the survey were completed and made internally consistent, it would be a useful reference for researchers working on private transformer inference. The paper has several strengths: it collects recent work into a single narrative, reproduces the standard cryptographic background, gives explicit approximation formulas for Softmax, GeLU, and LayerNorm, and provides links to open-source implementations. The paper does not claim a novel cryptographic derivation; its value is survey-level. However, the current significance is substantially undercut by the missing conclusion, unresolved cross-references, and citation errors in the comparison tables, because a survey's primary value lies in the reliability of its organization and tables.
major comments (4)
- [Section 8 / Abstract] The paper advertises evaluation guidelines in the abstract and introduction, but no such guidelines appear anywhere in the manuscript. Section 8, titled 'CONCLUSION', is empty, and the future-directions section is referenced as 'Section ??' in the Introduction and in Section 5. This is a missing contribution, not a presentation issue: a reader cannot use the paper for one of its two advertised central claims.
- [Tables 3, 4, 8, 9] Several comparison tables mislabel cited systems, which undermines the survey's core comparative function. Table 3 lists BOLT as reference [26] rather than [43]; Table 4 lists SecFormer in both the 2PC group and the 2PC-Dealer group, and lists Curl as [17] even though reference [17] is SIGMA and the bibliography entry for Curl is [50]; Tables 8 and 9 label NEXUS as [43] rather than [64]. Because the tables are the main deliverable for comparing systems, these errors are load-bearing.
- [Section 7, Table 12] Table 12 mixes runtimes across different models (BERT-Base, GPT2-Base, LLaMA-7B, ViT-Base), different datasets, different input sizes, and different network settings (e.g., 5 Gbps with 1 ms latency, 3 Gbps with 0.8 ms, 100 Mbps with 80 ms) with no normalization, no stated methodology, and no hardware/software environment details. As presented, the table cannot support any cross-system ranking of resource efficiency, yet Section 7 claims to compare experimental results.
- [Sections 4.1, 5, and 6.2] The manuscript contains multiple unresolved cross-references and broken exposition: Section 4.1 says 'we first introduce a secure inference system setup in Section ??', Section 5 says 'Section ?? first provides a breakdown', and the text after the MatMul discussion refers to attention equations as '(??)-(??)'. In addition, Section 6.2 contains the incomplete sentence 'Tech Tips: The function GeLU(𝑥) . Besides, polynomials are still available...'. These are not isolated typos but indicate that parts of the draft are unfinished, making the survey difficult to follow.
minor comments (6)
- [Section 3 title] The section title 'PIVACY THREATS IN SECURE INFERENCE' should be 'PRIVACY THREATS IN SECURE INFERENCE'.
- [Sections 2.3.2 and 5.2] The phrase 'secrete sharing' appears multiple times and should be 'secret sharing'.
- [Section 6.1] In the Softmax Tech Tips, the phrase 'to server as F(x)' should be 'to serve as F(x)'.
- [Section 4.3] The paragraph on 'Stuides [1, 32, 59]' contains a typo: 'Stuides' should be 'Studies'. Additionally, the sentence immediately following 'THE-X [7]' is a dangling fragment with no accompanying claim, so the discussion of client computation is incomplete.
- [Reference [65]] The reference title contains a typo: 'latency efficiefnt' should be 'latency efficient'.
- [Front matter] The copyright line reads '© 2018 Copyright held by the owner/author(s)' while the manuscript is an arXiv 2024 submission; this date appears inconsistent and should be corrected.
Circularity Check
No circularity found: the survey's content is drawn from external work and contains no fitted parameters or self-referential derivation chain.
full rationale
This paper is a literature survey; it introduces no novel protocol, derives no result from an assumed conclusion, and fits no parameters. Its equations (e.g., the attention, GeLU, and LayerNorm definitions in Sections 2 and 6) are standard textbook definitions reproduced from the cited literature, not predictions derived from the survey's own inputs. The survey's comparisons rely on reported numbers from external systems, and any inaccuracies in those tables (such as the BOLT/NEXUS citation mismatches noted in the manuscript) are correctness and completeness concerns, not circularity. There is no self-citation chain that supports a load-bearing claim, no ansatz smuggled in via the authors' prior work, and no renamed known result presented as a new derivation. The proposed evaluation guidelines are underdeveloped, but absence of content is not circular reasoning. Therefore the circularity score is 0.
Assumptions & free parameters
Cite this review
Pith. "Pith review of A Survey on Private Transformer Inference." pith.science (2026). https://pith.science/paper/DCMPTJEH
@misc{pith2026241208145,
author = {Pith},
title = {Pith review of: A Survey on Private Transformer Inference},
year = {2026},
howpublished = {\url{https://pith.science/paper/DCMPTJEH}},
note = {Machine review of arXiv:2412.08145}
}
read the original abstract
Transformer models have revolutionized AI, enabling applications like content generation and sentiment analysis. However, their use in Machine Learning as a Service (MLaaS) raises significant privacy concerns, as centralized servers process sensitive user data. Private Transformer Inference (PTI) addresses these issues using cryptographic techniques such as Secure Multi-Party Computation (MPC) and Homomorphic Encryption (HE), enabling secure model inference without exposing inputs or models. This paper reviews recent advancements in PTI, analyzing state-of-the-art solutions, their challenges, and potential improvements. We also propose evaluation guidelines to assess resource efficiency and privacy guarantees, aiming to bridge the gap between high-performance inference and data privacy.
Figures
Forward citations
Cited by 1 Pith paper
-
Towards Efficient Privacy-Preserving Machine Learning: A Systematic Review from Protocol, Model, and System Perspectives
A structured survey of PPML efficiency optimizations, grouped into protocol, model, and system levels, with comparisons and future directions.
Reference graph
Works this paper leans on
-
[26]
Brian Knott, Shobha Venkataraman, Awni Hannun, Shubho Sengupta, Mark Ibrahim, and Laurens van der Maaten. 2021. Crypten: Secure multi-party computation meets machine learning. Advances in Neural Information Processing Systems 34 (2021), 4961–4973
work page 2021
-
[43]
Qi Pang, Jinhao Zhu, Helen Möllering, Wenting Zheng, and Thomas Schneider. 2024. BOLT: Privacy-Preserving, Accurate and Efficient Inference for Transformers. In 2024 IEEE Symposium on Security and Privacy (SP) . IEEE Computer Society, 130–130. Manuscript submitted to ACM A Survey on Private Transformer Inference 23
work page 2024
-
[17]
Kanav Gupta, Neha Jawalkar, Ananta Mukherjee, Nishanth Chandran, Divya Gupta, Ashish Panwar, and Rahul Sharma. 2023. SIGMA: secure GPT inference with function secret sharing. Cryptology ePrint Archive (2023)
work page 2023
-
[50]
Manuel B Santos, Dimitris Mouris, Mehmet Ugurbil, Stanislaw Jarecki, José Reis, Shubho Sengupta, and Miguel de Vega. 2024. Curl: Private LLMs through Wavelet-Encoded Look-Up Tables. Cryptology ePrint Archive (2024)
work page 2024
-
[64]
Jiawen Zhang, Jian Liu, Xinpeng Yang, Yinghao Wang, Kejia Chen, Xiaoyang Hou, Kui Ren, and Xiaohu Yang. 2024. Secure Transformer Inference Made Non-interactive. Cryptology ePrint Archive (2024)
work page 2024
-
[1]
Yoshimasa Akimoto, Kazuto Fukuchi, Youhei Akimoto, and Jun Sakuma. 2023. Privformer: Privacy-preserving transformer with mpc. In 2023 IEEE 8th European Symposium on Security and Privacy (EuroS&P) . IEEE, 392–410
work page 2023
-
[2]
Toshinori Araki, Jun Furukawa, Yehuda Lindell, Ariel Nof, and Kazuma Ohara. 2016. High-throughput semi-honest secure three-party computation with an honest majority. In Proceedings of the 2016 ACM SIGSAC Conference on Computer and Communications Security . 805–817
work page 2016
-
[3]
Donald Beaver. 1992. Efficient multiparty protocols using circuit randomization. In Advances in Cryptology—CRYPTO’91: Proceedings 11 . Springer, 420–432
work page 1992
Show all 69 references
-
[4]
Elette Boyle, Geoffroy Couteau, Niv Gilboa, and Yuval Ishai. 2018. Compressing vector OLE. In Proceedings of the 2018 ACM SIGSAC Conference on Computer and Communications Security . 896–912
2018
-
[5]
Nishanth Chandran, Divya Gupta, Aseem Rastogi, Rahul Sharma, and Shardul Tripathi. 2019. EzPC: Programmable and efficient secure two-party computation for machine learning. In 2019 IEEE European Symposium on Security and Privacy (EuroS&P) . IEEE, 496–511
2019
-
[6]
Dake Chen, Yuke Zhang, Souvik Kundu, Chenghao Li, and Peter A Beerel. 2023. RNA-ViT: Reduced-Dimension Approximate Normalized Attention Vision Transformers for Latency Efficient Private Inference. In 2023 IEEE/ACM International Conference on Computer Aided Design (ICCAD) . IEE...
2023
-
[7]
Tianyu Chen, Hangbo Bao, Shaohan Huang, Li Dong, Binxing Jiao, Daxin Jiang, Haoyi Zhou, Jianxin Li, and Furu Wei. 2022. The-x: Privacy-preserving transformer inference with homomorphic encryption. arXiv preprint arXiv:2206.00216 (2022)
2022 arXiv
-
[8]
Yuntian Chen, Xianjia Meng, Zhiying Shi, Zhiyuan Ning, and Jingzhi Lin. 2024. SecureTLM: Private inference for transformer-based large model with MPC. Information Sciences 667 (2024), 120429
2024
-
[9]
Edward Chou, Josh Beal, Daniel Levy, Serena Yeung, Albert Haque, and Li Fei-Fei. 2018. Faster cryptonets: Leveraging sparsity for real-world encrypted inference. arXiv preprint arXiv:1811.09953 (2018)
2018 arXiv
-
[10]
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805 (2018)
2018 arXiv
-
[11]
Yuanchao Ding, Hua Guo, Yewei Guan, Weixin Liu, Jiarong Huo, Zhenyu Guan, and Xiyong Zhang. 2023. East: Efficient and accurate secure transformer framework for inference. arXiv preprint arXiv:2308.09923 (2023)
2023 arXiv
-
[12]
Ye Dong, Wen-jie Lu, Yancheng Zheng, Haoqi Wu, Derun Zhao, Jin Tan, Zhicong Huang, Cheng Hong, Tao Wei, and Wenguang Cheng. 2023. Puma: Secure inference of llama-7b in five minutes. arXiv preprint arXiv:2307.12533 (2023)
2023
-
[13]
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. 2020. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint...
2020 arXiv
-
[14]
David Evans, Vladimir Kolesnikov, Mike Rosulek, et al. 2018. A pragmatic introduction to secure multi-party computation. Foundations and Trends® in Privacy and Security 2, 2-3 (2018), 70–246
2018
-
[15]
Craig Gentry. 2009. Fully homomorphic encryption using ideal lattices. In Proceedings of the forty-first annual ACM symposium on Theory of computing. 169–178
2009
-
[16]
Ian Goodfellow, Yoshua Bengio, and Aaron Courville. 2016. Deep learning. MIT press
2016
-
[18]
Meng Hao, Hongwei Li, Hanxiao Chen, Pengzhi Xing, Guowen Xu, and Tianwei Zhang. 2022. Iron: Private inference on transformers. Advances in neural information processing systems 35 (2022), 15718–15731
2022
-
[19]
Geoffrey E Hinton, Simon Osindero, and Yee-Whye Teh. 2006. A fast learning algorithm for deep belief nets. Neural computation 18, 7 (2006), 1527–1554
2006
-
[20]
Xiaoyang Hou, Jian Liu, Jingyu Li, Yuhan Li, Wen-jie Lu, Cheng Hong, and Kui Ren. 2023. Ciphergpt: Secure two-party gpt inference. Cryptology ePrint Archive (2023)
2023
-
[21]
Hai Huang and Yongjian Wang. 2024. SecBERT: Privacy-preserving pre-training based neural network inference system. Neural Networks 172 (2024), 106135
2024
-
[22]
Zhicong Huang, Wen-jie Lu, Cheng Hong, and Jiansheng Ding. 2022. Cheetah: Lean and fast secure{Two-Party} deep neural network inference. In 31st USENIX Security Symposium (USENIX Security 22) . 809–826
2022
-
[23]
Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu. 2019. Tinybert: Distilling bert for natural language understanding. arXiv preprint arXiv:1909.10351 (2019)
2019 arXiv
-
[24]
2018.{GAZELLE}: A low latency framework for secure neural network inference
Chiraag Juvekar, Vinod Vaikuntanathan, and Anantha Chandrakasan. 2018.{GAZELLE}: A low latency framework for secure neural network inference. In 27th USENIX security symposium (USENIX security 18) . 1651–1669
2018
-
[25]
Marcel Keller. 2020. MP-SPDZ: A versatile framework for multi-party computation. In Proceedings of the 2020 ACM SIGSAC conference on computer and communications security. 1575–1590
2020
-
[27]
Nishant Kumar, Mayank Rathee, Nishanth Chandran, Divya Gupta, Aseem Rastogi, and Rahul Sharma. 2020. Cryptflow: Secure tensorflow inference. In 2020 IEEE Symposium on Security and Privacy (SP) . IEEE, 336–353
2020
-
[28]
Dacheng Li, Hongyi Wang, Rulin Shao, Han Guo, Eric Xing, and Hao Zhang. 2022. MPCFORMER: FAST, PERFORMANT AND PRIVATE TRANS- FORMER INFERENCE WITH MPC. In The Eleventh International Conference on Learning Representations
2022
-
[29]
Shaohua Li, Kaiping Xue, Bin Zhu, Chenkai Ding, Xindi Gao, David Wei, and Tao Wan. 2020. Falcon: A fourier transform based approach for fast and secure convolutional neural network predictions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognitio...
2020
-
[30]
Xuanqi Liu and Zhuotao Liu. 2023. Llms can understand encrypted prompt: Towards privacy-computing friendly transformers. arXiv preprint arXiv:2305.18396 (2023)
2023 arXiv
-
[31]
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692 (2019)
2019 arXiv
-
[32]
Yanxin Liu and Qianqian Su. 2024. PPTIF: Privacy-Preserving Transformer Inference Framework for Language Translation. IEEE Access (2024)
2024
-
[33]
Natasha Lomas. 2023. Italy orders ChatGPT blocked citing data protection concerns. TechCrunch, March 31 (2023)
2023
-
[34]
Wen-jie Lu, Zhicong Huang, Zhen Gu, Jingyu Li, Jian Liu, Kui Ren, Cheng Hong, Tao Wei, and WenGuang Chen. 2023. Bumblebee: Secure two-party inference framework for large transformers. Cryptology ePrint Archive (2023)
2023
-
[35]
Brady D Lund and Ting Wang. 2023. Chatting about ChatGPT: how may AI and GPT impact academia and libraries? Library hi tech news 40, 3 (2023), 26–29
2023
-
[36]
Jinglong Luo, Yehong Zhang, Zhuo Zhang, Jiaqi Zhang, Xin Mu, Hui Wang, Yue Yu, and Zenglin Xu. 2024. SecFormer: Fast and Accurate Privacy-Preserving Inference for Transformer Models via SMPC. In Findings of the Association for Computational Linguistics ACL 2024 . 13333–13348
2024
-
[37]
2023.{SecretFlow- SPU}: A Performant and{User-Friendly} Framework for{Privacy-Preserving} Machine Learning
Junming Ma, Yancheng Zheng, Jun Feng, Derun Zhao, Haoqi Wu, Wenjing Fang, Jin Tan, Chaofan Yu, Benyu Zhang, and Lei Wang. 2023.{SecretFlow- SPU}: A Performant and{User-Friendly} Framework for{Privacy-Preserving} Machine Learning. In 2023 USENIX Annual Technical Conference (USE...
2023
-
[38]
Cecily Mauran. 2023. Whoops, Samsung workers accidentally leaked trade secrets via ChatGPT. Mashable [online]. Dostupné z: https://mashable. com/article/samsungchatgpt-leak-details (2023)
2023
-
[39]
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. 2016. Pointer sentinel mixture models. arXiv preprint arXiv:1609.07843 (2016)
2016 arXiv
-
[40]
Microsoft and OpenAI. 2023. Bing Chat. (2023). https://www.bing.com/search
2023
-
[41]
Jungho Moon, Dongwoo Yoo, Xiaoqian Jiang, and Miran Kim. 2024. THOR: Secure Transformer Inference with Homomorphic Encryption.Cryptology ePrint Archive (2024)
2024
-
[42]
OpenAI. 2022. ChatGPT. (2022). https://openai.com/blog/chatgpt
2022
-
[44]
Dongjin Park, Eunsang Lee, and Joon-Woo Lee. 2024. Powerformer: Efficient privacy-preserving transformer with batch rectifier-power max function and optimized homomorphic attention. Cryptology ePrint Archive (2024)
2024
-
[45]
Hongyuan Qu and Guangwu Xu. 2023. Improvements of Homomorphic Secure Evaluation of Inverse Square Root. In International Conference on Information and Communications Security . Springer, 110–127
2023
-
[46]
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019. Language models are unsupervised multitask learners. OpenAI blog 1, 8 (2019), 9
2019
-
[47]
Deevashwer Rathee, Dacheng Li, Ion Stoica, Hao Zhang, and Raluca Popa. 2024. MPC-Minimized Secure LLM Inference.arXiv preprint arXiv:2408.03561 (2024)
2024 arXiv
-
[48]
Deevashwer Rathee, Mayank Rathee, Rahul Kranti Kiran Goli, Divya Gupta, Rahul Sharma, Nishanth Chandran, and Aseem Rastogi. 2021. Sirnn: A math library for secure rnn inference. In 2021 IEEE Symposium on Security and Privacy (SP) . IEEE, 1003–1020
2021
-
[49]
Lorenzo Rovida and Alberto Leporati. 2024. Transformer-based language models and homomorphic encryption: An intersection with bert-tiny. In Proceedings of the 10th ACM International Workshop on Security and Privacy Analytics . 3–13
2024
-
[51]
Sijun Tan, Brian Knott, Yuan Tian, and David J Wu. 2021. CryptGPU: Fast privacy-preserving machine learning on the GPU. In2021 IEEE Symposium on Security and Privacy (SP) . IEEE, 1021–1038
2021
-
[52]
Ahmed Tlili, Boulus Shehata, Michael Agyemang Adarkwah, Aras Bozkurt, Daniel T Hickey, Ronghuai Huang, and Brighter Agyemang. 2023. What if the devil is my guardian angel: ChatGPT as a case study of using chatbots in education. Smart learning environments 10, 1 (2023), 15
2023
-
[53]
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971 (2023)
2023 arXiv
-
[54]
Iulia Turc, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. Well-read students learn better: On the importance of pre-training compact models. arXiv preprint arXiv:1908.08962 (2019)
2019 arXiv
-
[55]
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in neural information processing systems 30 (2017)
2017
-
[56]
Sameer Wagh, Shruti Tople, Fabrice Benhamouda, Eyal Kushilevitz, Prateek Mittal, and Tal Rabin. 2021. FALCON: Honest-Majority Maliciously Secure Framework for Private Deep Learning. Proceedings on Privacy Enhancing Technologies
2021
-
[57]
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman. 2018. GLUE: A multi-task benchmark and analysis platform for natural language understanding. arXiv preprint arXiv:1804.07461 (2018)
2018 arXiv
-
[58]
Weize Wang and Yi Kuang. 2024. CipherFormer: Efficient Transformer Private Inference with Low Round Complexity.arXiv preprint arXiv:2403.16860 (2024)
2024 arXiv
-
[59]
Yongqin Wang, G Edward Suh, Wenjie Xiong, Benjamin Lefaudeux, Brian Knott, Murali Annavaram, and Hsien-Hsin S Lee. 2022. Characterization of mpc-based private inference for transformer-based models. In 2022 IEEE International Symposium on Performance Analysis of Systems and So...
2022
-
[60]
Tianshi Xu, Lemeng Wu, Runsheng Wang, and Meng Li. 2024. PrivCirNet: Efficient Private Inference via Block Circulant Transformation. arXiv preprint arXiv:2405.14569 (2024)
2024 arXiv
-
[61]
Andrew C Yao. 1982. Protocols for secure computations. In 23rd annual symposium on foundations of computer science (sfcs 1982) . IEEE, 160–164
1982
-
[62]
Chenkai Zeng, Debiao He, Qi Feng, Xiaolin Yang, and Qingcai Luo. 2024. SecureGPT: A Framework for Multi-Party Privacy-Preserving Transformer Inference in GPT. IEEE Transactions on Information Forensics and Security (2024)
2024
-
[63]
Wenxuan Zeng, Meng Li, Wenjie Xiong, Tong Tong, Wen-jie Lu, Jin Tan, Runsheng Wang, and Ru Huang. 2023. Mpcvit: Searching for accurate and efficient mpc-friendly vision transformer with heterogeneous attention. In Proceedings of the IEEE/CVF International Conference on Compute...
2023
-
[65]
Yuke Zhang, Dake Chen, Souvik Kundu, Chenghao Li, and Peter A Beerel. 2023. Sal-vit: Towards latency efficiefnt private inference on vit using selective attention search with a learnable softmax approximation. In Proceedings of the IEEE/CVF International Conference on Computer...
2023
-
[66]
Chuan Zhao, Shengnan Zhao, Minghao Zhao, Zhenxiang Chen, Chong-Zhi Gao, Hongwei Li, and Yu-an Tan. 2019. Secure multi-party computation: theory, practice and applications. Information Sciences 476 (2019), 357–372
2019
-
[67]
Mengxin Zheng, Qian Lou, and Lei Jiang. 2023. Primer: Fast private transformer inference on encrypted data. In 2023 60th ACM/IEEE Design Automation Conference (DAC). IEEE, 1–6
2023
-
[68]
Itamar Zimerman, Allon Adir, Ehud Aharoni, Matan Avitan, Moran Baruch, Nir Drucker, Jenny Lerner, Ramy Masalha, Reut Meiri, and Omri Soceanu. 2024. Power-Softmax: Towards Secure LLM Inference over Encrypted Data. arXiv preprint arXiv:2410.09457 (2024)
2024 arXiv
-
[69]
Itamar Zimerman, Moran Baruch, Nir Drucker, Gilad Ezov, Omri Soceanu, and Lior Wolf. 2023. Converting transformers to polynomial form for secure inference over homomorphic encryption. arXiv preprint arXiv:2311.08610 (2023). Manuscript submitted to ACM 24 Yang et al. A MPC SETT...
2023 arXiv
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.