{"id":"af028ba9-492b-41fd-a46b-b96c5d2afdcd","arxiv_id":"2412.08145","paper_version":1,"verdict":"REJECT","confidence":"LOW","novelty_score":0.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A literature survey on private transformer inference that is too incomplete to support its promised comparisons and evaluation guidelines.","lead":"A survey of private transformer inference methods compares MPC- and HE-based systems for running models like BERT on encrypted inputs. The draft is incomplete, with missing sections, an empty conclusion, and several incorrect table citations.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The survey's promised evaluation guidelines are absent and its comparison tables are internally inconsistent, so the central comparative claim is unsupported.","rationale":"The reader's weakest assumption is that the tables faithfully represent the cited systems, and my check confirms that assumption is already violated by internal inconsistencies. I agree with the reader's REJECT verdict: the draft is not yet a usable survey. My emphasis is slightly broader: even if every citation typo were fixed, the promised evaluation guidelines do not exist in the draft, and Section 7's comparison table lacks any normalization methodology. Both problems are load-bearing because the abstract and introduction promise the guidelines, and the tables are the only evidence offered for the review's comparative claims. The concern is not about controversial judgments or technical errors in the underlying PTI papers; it is about the manuscript's own evidence layer being incomplete and unreliable. A repaired version could be valuable, but the current draft cannot support the central claim.","tokens_in":22149,"tokens_out":3591,"duration_ms":40646,"concrete_test":"In the final manuscript, locate the promised evaluation-guidelines contribution and reconstruct Table 4's setup columns from each cited paper's declared protocol. If no guidelines section exists, or if any paper must be moved between setup categories (e.g., SecFormer appears in both 2PC and 2PC-Dealer), the central comparative claim of the survey fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim of the paper has two parts: a comprehensive review and proposed evaluation guidelines. The review side fails where it must be most reliable: the comparison tables are internally inconsistent and reference the wrong systems. Table 4 lists SecFormer under both the 2PC group and the 2PC-Dealer group, which cannot both be correct; Table 4 gives Curl as [17] while the bibliography entry for Curl is [50]; Table 3 links BOLT to [26] instead of [43]; and Tables 8 and 9 label the NEXUS row as [43] rather than [64]. A reader cannot use these tables to compare PTI systems if the rows and citations do not reliably identify the systems being compared. The guidelines side is structurally absent: Section 8 is an empty 'CONCLUSION', and the future-directions section is referenced as 'Section ??'. Table 12 mixes runtimes across different models, datasets, input sizes, and network settings with no normalization or stated methodology, so any cross-system ranking is unsupported. This is a missing-support failure, not a disagreement with the field's consensus.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript is a survey of private transformer inference (PTI). It covers background on transformer architecture and cryptographic primitives (MPC, HE), reviews roughly thirty PTI systems from 2022–2024, categorizes them by setup (2PC, 2PC-Dealer, 3PC), and discusses protocols for linear layers (MatMul) and non-linear layers (Softmax, GeLU, LayerNorm). The abstract and introduction promise three contributions: a comprehensive review, a breakdown of challenges and typical solutions, and proposed evaluation guidelines for resource efficiency and privacy guarantees. The review portion is structured around comparison tables and per-layer discussions, but the manuscript is incomplete: Section 8 is empty, several cross-references appear as unresolved 'Section ??', and the promised evaluation guidelines do not materialize in the text.","tokens_in":22274,"tokens_out":4456,"duration_ms":45939,"significance":"If the survey were completed and made internally consistent, it would be a useful reference for researchers working on private transformer inference. The paper has several strengths: it collects recent work into a single narrative, reproduces the standard cryptographic background, gives explicit approximation formulas for Softmax, GeLU, and LayerNorm, and provides links to open-source implementations. The paper does not claim a novel cryptographic derivation; its value is survey-level. However, the current significance is substantially undercut by the missing conclusion, unresolved cross-references, and citation errors in the comparison tables, because a survey's primary value lies in the reliability of its organization and tables.","major_comments":[{"comment":"The paper advertises evaluation guidelines in the abstract and introduction, but no such guidelines appear anywhere in the manuscript. Section 8, titled 'CONCLUSION', is empty, and the future-directions section is referenced as 'Section ??' in the Introduction and in Section 5. This is a missing contribution, not a presentation issue: a reader cannot use the paper for one of its two advertised central claims.","section":"Section 8 / Abstract"},{"comment":"Several comparison tables mislabel cited systems, which undermines the survey's core comparative function. Table 3 lists BOLT as reference [26] rather than [43]; Table 4 lists SecFormer in both the 2PC group and the 2PC-Dealer group, and lists Curl as [17] even though reference [17] is SIGMA and the bibliography entry for Curl is [50]; Tables 8 and 9 label NEXUS as [43] rather than [64]. Because the tables are the main deliverable for comparing systems, these errors are load-bearing.","section":"Tables 3, 4, 8, 9"},{"comment":"Table 12 mixes runtimes across different models (BERT-Base, GPT2-Base, LLaMA-7B, ViT-Base), different datasets, different input sizes, and different network settings (e.g., 5 Gbps with 1 ms latency, 3 Gbps with 0.8 ms, 100 Mbps with 80 ms) with no normalization, no stated methodology, and no hardware/software environment details. As presented, the table cannot support any cross-system ranking of resource efficiency, yet Section 7 claims to compare experimental results.","section":"Section 7, Table 12"},{"comment":"The manuscript contains multiple unresolved cross-references and broken exposition: Section 4.1 says 'we first introduce a secure inference system setup in Section ??', Section 5 says 'Section ?? first provides a breakdown', and the text after the MatMul discussion refers to attention equations as '(??)-(??)'. In addition, Section 6.2 contains the incomplete sentence 'Tech Tips: The function GeLU(𝑥) . Besides, polynomials are still available...'. These are not isolated typos but indicate that parts of the draft are unfinished, making the survey difficult to follow.","section":"Sections 4.1, 5, and 6.2"}],"minor_comments":[{"comment":"The section title 'PIVACY THREATS IN SECURE INFERENCE' should be 'PRIVACY THREATS IN SECURE INFERENCE'.","section":"Section 3 title"},{"comment":"The phrase 'secrete sharing' appears multiple times and should be 'secret sharing'.","section":"Sections 2.3.2 and 5.2"},{"comment":"In the Softmax Tech Tips, the phrase 'to server as F(x)' should be 'to serve as F(x)'.","section":"Section 6.1"},{"comment":"The paragraph on 'Stuides [1, 32, 59]' contains a typo: 'Stuides' should be 'Studies'. Additionally, the sentence immediately following 'THE-X [7]' is a dangling fragment with no accompanying claim, so the discussion of client computation is incomplete.","section":"Section 4.3"},{"comment":"The reference title contains a typo: 'latency efficiefnt' should be 'latency efficient'.","section":"Reference [65]"},{"comment":"The copyright line reads '© 2018 Copyright held by the owner/author(s)' while the manuscript is an arXiv 2024 submission; this date appears inconsistent and should be corrected.","section":"Front matter"}],"recommendation":"major_revision","confidential_remarks":"This is clearly an incomplete draft rather than a finished survey. The citation mismatches and empty conclusion are fixable, but the advertised evaluation guidelines are entirely absent, so the revision will require substantial new content, not just editing. I would invite a resubmission after the missing sections are written and all tables are corrected, rather than recommending rejection outright."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Zhi, quick take on arXiv:2412.08145. The survey's organizing scheme is genuinely useful: classifying PTI work by setup (2PC, 2PC-Dealer, 3PC) and by layer type (MatMul, Softmax, GeLU, LayerNorm) gives readers a sensible map of the field, and the coverage of 2022-2024 work is broad. The detailed rundown of approximation techniques for nonlinearities is the strongest part, especially the equations comparing Softmax and GeLU approximations across systems. That alone could be a handy reference.\n\nBut the paper is not submission-ready. The abstract and introduction promise a set of evaluation guidelines; that contribution is absent, and Section 8 is an empty 'CONCLUSION' with the future-directions section still sitting as 'Section ??'. The comparison tables have citation errors that are not trivial: BOLT is linked to [26] (CrypTen) instead of [43]; Curl appears as [17] (SIGMA) instead of [50]; Tables 8 and 9 label NEXUS as [43] (BOLT) rather than [64]. Table 4 lists SecFormer twice, once under 2PC and once under 2PC-Dealer, which looks like a duplicate rather than intentional evaluation under both settings. Table 12 compares runtimes across different models, datasets, input sizes, and network settings with no normalization or stated methodology, so any cross-system ranking is unsupported. Some prose sentences also break off mid-thought (e.g., the Tech Tips on GeLU). These are mechanical problems, but they are load-bearing: a survey that misidentifies which paper a row belongs to cannot be used for comparison.\n\nWhat is here is largely a restatement of categories and numbers from the cited papers, which is normal for a survey, but the reader's critique is fair: the missing guidelines and broken references mean the current version is not a reliable reference. The authors know how to organize the literature; the work needs completion and a careful citation audit before it can serve its purpose.\n\nWho would benefit? A newcomer to PTI who wants a structured bibliography and a rundown of the main approximation strategies, after the errors are fixed. As it stands I would not send it to referees; I would return it to the authors with a do-not-submit verdict and invite a revised version. If they complete the missing sections and correct the tables, it becomes a moderate-value survey worth citing.","headline":"A well-organized PTI survey that is currently too incomplete to use: the promised evaluation guidelines are missing and the comparison tables mislabel several systems.","tokens_in":22817,"tokens_out":2678,"would_cite":false,"duration_ms":26797,"reading_group":"no","serious_thinker":"yes","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Private transformer inference has two bottlenecks — large matrix multiplications and non-linear functions — and the surveyed systems differ mainly by cryptographic setup and which layer they optimize.","keywords":["private transformer inference","secure multi-party computation","homomorphic encryption","Softmax approximation","GeLU approximation","secure matrix multiplication","privacy-preserving machine learning","transformer architecture"],"falsifier":"A reader can settle the central claim by cross-checking every row of Tables 3, 4, 8, 9, 10, and 12 against the cited papers' own reported numbers and reference lists. The draft already shows two misattributions (BOLT appears with marker [26] in Table 3 and [43] in Table 4; Curl appears with marker [17] instead of [50]), so a systematic verification of all rows would show whether the survey's comparisons are reliable.","tokens_in":21916,"feed_emoji":"🔐","tokens_out":9344,"duration_ms":85796,"temperature":0.7,"pith_summary":"This survey aims to establish that the design space of private transformer inference from 2022 to 2024 can be organized around two cryptographic bottlenecks: large matrix multiplications in the linear layers and complex non-linear functions (Softmax, GeLU, LayerNorm) in the attention and feed-forward blocks. It reviews the surveyed systems by their setup (two-party, two-party with a trusted dealer, and three-party), by the cryptographic tool they employ, and by which transformer component they optimize, and it proposes evaluation guidelines that report communication volume, runtime, and accuracy loss together. The survey's comparisons show a trade-off: homomorphic-encryption-based systems achieve low communication and non-interactivity but high computation, while MPC-hybrid systems trade communication and trust assumptions for speed. If the survey is accurate, a practitioner can use its tables to choose a system based on threat model and network environment rather than on isolated latency claims.","feed_headline":"Non-linear layers now dominate private transformer inference","feed_subtitle":"A 2022–2024 review shows the cost shifted from matrix multiplication to Softmax, GeLU, and LayerNorm.","key_machinery":"The carrying mechanism is a layer-wise decomposition of a transformer encoder into linear operations (matrix multiplications in attention and feed-forward layers) and non-linear operations (Softmax, GeLU, and LayerNorm), cross-classified by cryptographic setup (two-party, two-party with a trusted dealer, three-party). Within that grid, the load-bearing objects are: secret-sharing schemes with Beaver triples (precomputed shared randomness that turns secure multiplication into one communication round) or re-sharing (refreshing shares by exchanging noisy local results) for secure multiplication; homomorphic-encryption schemes (BFV, CKKS, RNS-CKKS) with ciphertext packing and SIMD operations for matrix multiplication; and approximation techniques for non-linear functions, including low-degree polynomials, Taylor/Maclaurin/Chebyshev/Fourier series, and look-up tables. These mechanisms let the survey compare systems along the same axes: each system's reported communication volume, runtime, and accuracy loss are tied to which mechanisms it uses and which layer it optimizes.","core_discovery":"The paper's central claim is that the current state of private transformer inference is best understood not as a contest between homomorphic encryption and secure multi-party computation, but as a layered design problem: each transformer component imposes a different cryptographic cost, and each system can be described by which layer it optimizes and under which setup. The paper argues that in two-party setups, large matrix multiplications are a dominant bottleneck because secure multiplication requires extra privacy protection; in dealer-assisted and three-party setups, that bottleneck moves to non-linear layers, which now account for most of the runtime in both MPC and HE systems. It further claims that accuracy preservation is achieved mainly by replacing non-linear functions with crypto-friendly approximations (low-degree polynomials, Taylor, Maclaurin, Chebyshev, or Fourier series, and look-up tables) and then recovering accuracy through knowledge distillation. The paper concludes by proposing evaluation guidelines, arguing that fair comparison requires reporting communication volume, runtime, accuracy loss, and the security model together.","pith_inferences":["Not stated in the paper, but a reader could infer that the setup choice encodes a trust assumption: two-party systems avoid extra trust but pay for matrix multiplication, while dealer and three-party systems push that cost onto an assumed-honest helper; the two-party direction is therefore the harder test for the field.","Not stated in the paper, the per-layer tables imply a testable ordering under identical network conditions: an HE-only system will show near-zero communication but the longest runtime, a two-party hybrid will sit in the middle, and a three-party or dealer system will show the lowest runtime only if a helper is available.","Not stated in the paper, the proposed evaluation guidelines could be turned into a community benchmark that normalizes communication per token, per-layer runtime, and GLUE accuracy loss, which would convert the survey's qualitative comparisons into reproducible numbers."],"forward_implications":["If the survey's classification holds, future systems in the two-party setting should focus on jointly optimizing MatMul and non-linear layers, since neither alone determines end-to-end cost.","If the per-layer breakdowns are accurate, optimizing Softmax, GeLU, and LayerNorm will produce larger end-to-end gains than further MatMul speedups for dealer and three-party systems.","If the evaluation guidelines are adopted, reported numbers across studies become comparable, because current tables differ in network bandwidth, input size, and setup, making cross-paper comparison unreliable without normalization.","If the approximation-and-distillation trend continues, accuracy preservation will carry an extra training cost that must be included in any resource comparison, not just inference runtime."],"supporting_citations":[{"why":"Supplies the early HE-only transformer inference system, establishing the pure-HE approach and its small-model limitation.","marker":"[7]"},{"why":"Introduces the hybrid ASS+HE MatMul protocol and LUT-based Softmax, a baseline for 2PC systems.","marker":"[18]"},{"why":"Supplies a leveled-HE MatMul protocol with column-wise packing and polynomial GeLU, a central 2PC hybrid system.","marker":"[43]"},{"why":"Gives the nearly communication-free RNS-CKKS system whose non-linear approximations anchor the HE-only comparisons.","marker":"[64]"},{"why":"Provides the 3PC replicated-secret-sharing system that demonstrates honest-majority setups shift the bottleneck to non-linear layers.","marker":"[12]"},{"why":"Shows a 2PC hybrid using preprocessing correlations and Taylor-approximated Softmax, used in the non-linear layer comparisons.","marker":"[20]"},{"why":"Introduces the IntrLeave packing and polynomial approximations used for large transformers, a key 2PC hybrid data point.","marker":"[34]"},{"why":"Establishes the 2PC-Dealer setup with Beaver triples and distillation-based accuracy recovery, central to the dealer discussion.","marker":"[28]"},{"why":"Demonstrates function-secret-sharing and reduced-bit-width LUTs for Softmax and GeLU, a key 2PC-Dealer system.","marker":"[17]"},{"why":"Provides a Fourier-series erf approximation for GeLU and a 2PC-Dealer comparison point for accuracy-preserving non-linear layers.","marker":"[36]"}],"fun_headline_variants":["Private transformers: non-linear layers are the new bottleneck","Where does private transformer cost go? Non-linear layers","Bottleneck shift: non-linear layers now dominate private transformer inference","In private transformers, non-linear layers now cost the most","Private transformer inference: non-linear layers take over"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The survey's usefulness depends on its tables faithfully representing the cited systems' security models, runtimes, and reference markers; if those entries are wrong or misattributed, the comparative conclusions drawn from the survey would be misleading.","fun_headline_variants_meta":{"raw":{"variants":["Private transformers: non-linear layers are the new bottleneck","Where does private transformer cost go? Non-linear layers","Bottleneck shift: non-linear layers now dominate private transformer inference","In private transformers, non-linear layers now cost the most","Private transformer inference: non-linear layers take over"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000417,"raw_usage":{"total_tokens":2094,"prompt_tokens":834,"completion_tokens":1260,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":450,"completion_tokens_details":{"reasoning_tokens":1181}},"tokens_in":450,"tokens_out":1260,"duration_ms":9222,"temperature":1.0,"reasoning_tokens":1181,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T18:08:49.070357+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A reader can settle the central claim by cross-checking every row of Tables 3, 4, 8, 9, 10, and 12 against the cited papers' own reported numbers and reference lists. The draft already shows two misattributions (BOLT appears with marker [26] in Table 3 and [43] in Table 4; Curl appears with marker [17] instead of [50]), so a systematic verification of all rows would show whether the survey's comparisons are reliable.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies a leveled-HE MatMul protocol with column-wise packing and polynomial GeLU, a central 2PC hybrid system."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Gives the nearly communication-free RNS-CKKS system whose non-linear approximations anchor the HE-only comparisons."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Shows a 2PC hybrid using preprocessing correlations and Taylor-approximated Softmax, used in the non-linear layer comparisons."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces the IntrLeave packing and polynomial approximations used for large transformers, a key 2PC hybrid data point."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes the 2PC-Dealer setup with Beaver triples and distillation-based accuracy recovery, central to the dealer discussion."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Demonstrates function-secret-sharing and reduced-bit-width LUTs for Softmax and GeLU, a key 2PC-Dealer system."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides a Fourier-series erf approximation for GeLU and a 2PC-Dealer comparison point for accuracy-preserving non-linear layers."}],"review_version":1}