{"id":"49b30e66-5333-45a7-a0a6-ebdf4e827606","arxiv_id":"2607.12504","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.5,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"PQC-TLS handshake exhaustion can keep a server at high CPU up to 88× longer and drops leading DL IDS recall/AU-ROC to near-random levels under mixed attack traffic.","lead":"Post-quantum TLS handshakes make servers far more vulnerable to handshake-flood DDoS and can blind modern deep-learning intrusion detectors. The work measures that amplification and releases a hybrid PQC attack dataset plus cloud scripts so others can reproduce and harden defenses.","discovery_kind":"extension","skeptic_critique":{"model":"grok-4.5","headline":"Abstract-only review cannot verify whether the 88× CPU-duration claim and IDS collapses rest on fair, representative measurement; the single-server testbed and detector setup remain the load-bearing uncheckable premise.","rationale":"The Reader already set UNVERDICTED / LOW confidence precisely because only the abstract is present and the representativeness of the small empirical testbed cannot be checked. That is the correct posture for an abstract-only empirical measurement paper whose strongest claims are quantitative amplification factors and detector-score collapses. No stronger internal inconsistency or circularity can be diagnosed without the body; manufacturing one would violate good-faith review. The concrete test above is the minimal experiment that would settle whether the load-bearing premise holds once artifacts appear. Until then the verdict stays UNVERDICTED.","tokens_in":2150,"tokens_out":503,"duration_ms":5433,"concrete_test":"Once the full paper and public dataset/scripts are available: re-run the exact AWS deployment scripts on a multi-core instance with TLS offload enabled (or at least L=2–4 cores) under the same PQC suites and attack intensity; recompute the high-CPU duration ratio versus classical TLS. If the ratio falls below ~10× (or IDS metrics recover above 0.8 AU-ROC / 80% recall under re-tuned hyperparameters), the headline amplification and blind-spot claims do not transfer and must be caveated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claims (up to 88× prolonged high-CPU under handshake exhaustion; exosphere recall ~50%; HyperVision AU-ROC ~0.49) are purely empirical. With only the abstract available, it is impossible to confirm that the single-server / ten-attacker AWS-style testbed, the chosen PQC suites, attack intensity, mixed legitimate traffic, multi-core/offload configuration, and IDS hyperparameters are free of confounds that would inflate the amplification factor or produce artificial blind spots. The reader correctly flags representativeness as the weakest assumption; that assumption is load-bearing because the paper offers no theoretical bound or multi-scale validation that would let the numbers transfer beyond the reported setup. Without body text, figures, or artifacts, neither the 88× figure nor the IDS degradation can be audited for fairness or scale.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The manuscript reports an empirical study of post-quantum cryptography (PQC) integrated into TLS 1.3, arguing that PQC handshake overhead amplifies handshake-exhaustion DDoS. On a testbed of one PQC-TLS server and ten attackers generating >16.5 GB of mixed legitimate and attack traffic, the authors claim PQC-TLS prolongs sustained high server CPU utilization by up to 88× relative to classical TLS. They further report that state-of-the-art deep-learning IDS degrade under PQC traffic (exosphere recall ~50%; HyperVision AU-ROC ~0.49). Contributions include root-cause analysis of IDS blind spots, a public PQC-DDoS hybrid dataset with attack timestamps and server resource traces, and open-sourced experimental code and AWS deployment scripts.","tokens_in":2300,"tokens_out":829,"duration_ms":10564,"significance":"If the measurements hold under fair, representative conditions, the work is timely and practically important: it quantifies a concrete security side-effect of the PQC transition in TLS and supplies a public hybrid dataset plus reproducible cloud scripts that the community currently lacks. The dual focus on attack amplification and IDS degradation is a useful framing for next-generation PQC-aware detectors. Significance is conditional on experimental fairness and transferability, which cannot be audited from the abstract alone.","major_comments":[{"comment":"Abstract claim of 'up to 88 times' prolonged high-CPU periods is load-bearing for the central amplification result, yet the abstract does not define the CPU threshold, the duration metric, the classical-TLS baseline configuration (cipher suites, key sizes, multi-core/offload), or error bars/replication. Without those definitions and a clear comparison protocol, the factor cannot be assessed for fairness or statistical reliability.","section":null},{"comment":"Abstract reports exosphere recall ~50% and HyperVision AU-ROC ~0.49 under PQC traffic. The load-bearing question is whether detectors were trained/tuned only on classical TLS and then evaluated zero-shot on PQC, or retrained with matched PQC features and hyperparameters. If the former, the 'blind spots' may reflect distribution shift rather than an intrinsic PQC failure mode; the abstract does not state the training/evaluation protocol.","section":null},{"comment":"The single-server, ten-attacker AWS-style testbed is the sole empirical basis for both the 88× and IDS claims. Representativeness (scale, concurrent legitimate load, hardware crypto offload, multi-core scheduling, network path) is not established in the abstract and is load-bearing for any claim that results transfer to production PQC-TLS deployments.","section":null}],"minor_comments":[{"comment":"Abstract should name the specific PQC KEMs/signatures and TLS library (e.g., OpenSSL/oqs-provider versions) used, so readers can judge suite choice and known performance characteristics.","section":null},{"comment":"Abstract should briefly state how 'mixed legitimate browsing' was generated and what fraction of traffic was attack vs. benign, to contextualize IDS metrics.","section":null},{"comment":"Dataset and code release claims are valuable; the abstract would be stronger if it named the license and archival location (e.g., DOI/Zenodo) rather than only asserting public release.","section":null}],"recommendation":"uncertain","confidential_remarks":"This is an abstract-only review; full text, figures, tables, and artifacts were not available. I cannot responsibly recommend accept/minor/major/reject without the methods, baselines, and statistical detail. If the full manuscript is supplied, the three major points above (CPU-duration definition and baseline, IDS train/eval protocol, testbed representativeness) should be the first items checked. Scope appears appropriate for cs.CR if the experiments are sound."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The one thing to know: this is an applied measurement paper claiming that PQC-TLS handshake exhaustion can keep a server at high CPU for up to 88× longer than classical TLS, and that two named DL IDS (exosphere, HyperVision) fall apart on the resulting hybrid traffic (recall ~50%, AU-ROC ~0.49). They also ship a 16.5 GB mixed dataset with attack timestamps and server resource traces, plus AWS scripts. That package is the real product.\n\nWhat is new is the quantification under a concrete PQC-TLS setup and the public hybrid corpus. Handshake DDoS and PQC size/CPU cost are not new ideas; tying them together with named detectors and releasing the traffic is useful engineering work for people rolling out PQC and for IDS researchers who need non-classical traffic. Circularity is not the issue—the claims are external measurements of CPU and third-party detectors.\n\nThe soft spot is exactly what the stress-test flags: we only have the abstract. Representativeness of the single-server / ten-attacker AWS testbed, cipher-suite choice, multi-core or offload configuration, attack intensity, mix of legitimate traffic, and IDS hyperparameter fairness are all uncheckable. The 88× figure and the near-random AU-ROC could be real or inflated by setup; without body, figures, or artifacts we cannot tell. That is a load-bearing limitation of this review, not a proven flaw in the paper.\n\nWho it is for: network-security and crypto-engineering people who care about PQC deployment and DDoS/IDS practice. Not a theory paper. It deserves a serious referee if the full text and artifacts match the abstract’s promises—especially the dataset and scripts. I would not desk-reject on the abstract alone; I would send it out and demand the experimental fairness checks. Bring it to reading group only after the PDF and data are in hand; until then treat the numbers as provisional.","headline":"Abstract-only empirical measurement of PQC-TLS handshake exhaustion (claimed 88× CPU-duration) and DL IDS collapse, plus a promised public hybrid dataset; numbers matter if the testbed is fair, but we cannot audit that yet.","tokens_in":2947,"tokens_out":518,"would_cite":false,"duration_ms":6077,"reading_group":"maybe","serious_thinker":"unclear","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"PQC-enabled TLS prolongs handshake-exhaustion high-CPU periods by up to 88 times and collapses deep-learning IDS detection to near-random levels.","keywords":["post-quantum cryptography","TLS 1.3","handshake exhaustion","DDoS","intrusion detection systems","PQC-TLS","deep learning IDS","hybrid traffic dataset"],"falsifier":"Re-run the identical handshake-exhaustion workload against a multi-core or hardware-accelerated PQC-TLS server and measure whether sustained high-CPU duration still multiplies by tens of times versus classical TLS; retrain the same IDS models on balanced PQC traffic and check whether AU-ROC remains near 0.5.","tokens_in":3012,"feed_emoji":"🔐","tokens_out":934,"duration_ms":18413,"temperature":0.7,"pith_summary":"This paper establishes that integrating post-quantum cryptography into TLS 1.3, while defending against quantum attacks, substantially amplifies handshake-exhaustion DDoS and creates detection blind spots for existing deep-learning intrusion systems. On an empirical testbed of one PQC-TLS server and ten attackers that produced more than 16.5 GB of mixed legitimate and attack traffic, PQC handshakes extend periods of sustained high server CPU utilization by as much as 88 times relative to classical TLS. State-of-the-art detectors then fail: exosphere recall falls to roughly 50 percent and HyperVision AU-ROC drops to 0.49. The authors quantify the causes of these failures, release a public hybrid traffic dataset with attack timestamps and resource traces, and open-source the full experimental stack so others can reproduce the findings and build PQC-aware defenses.","feed_headline":"PQC-TLS multiplies handshake DDoS CPU load by up to 88×","feed_subtitle":"Deep-learning IDS recall falls to ~50% and AU-ROC near random, exposing detection blind spots on PQC traffic.","key_machinery":"A controlled PQC-TLS testbed that mixes legitimate browsing with high-intensity handshake-exhaustion traffic, server-side CPU and resource monitoring, and side-by-side evaluation of deep-learning IDS (exosphere and HyperVision) on the resulting hybrid traces.","core_discovery":"When post-quantum cryptographic suites replace classical ones inside TLS 1.3, the extra handshake computation and communication cost multiplies the duration of sustained high server CPU under handshake-exhaustion attacks by up to 88 times and drives deep-learning IDS performance down to near-random levels (exosphere recall ~50 percent, HyperVision AU-ROC 0.49), exposing systematic blind spots that classical-trained detectors cannot see.","pith_inferences":["Even multi-core or offload-accelerated production servers may still experience amplified exhaustion if the per-handshake cost ratio between PQC and classical remains high.","Feature extractors that ignore PQC-specific size and timing signatures will systematically under-detect until they are retrained or redesigned.","Protocol designers may need client puzzles or rate-limiting mechanisms sized to the new PQC handshake costs rather than classical ones.","The same overhead that lengthens CPU spikes may also supply new side-channel signals that a PQC-aware detector could exploit if instrumented correctly."],"forward_implications":["Handshake-exhaustion attacks become far more effective against PQC-TLS servers than against classical TLS servers of equal capacity.","Deep-learning IDS trained or tuned on classical TLS traffic will miss a large fraction of the same attacks once PQC suites are enabled.","Defenders must redesign IDS features and models around the larger packet sizes and longer computational footprints of PQC handshakes.","The released hybrid dataset and AWS scripts give the community a reproducible baseline for building and comparing PQC-aware detectors."],"fun_headline_variants":["PQC-TLS extends high-CPU handshake attacks by 88×","Handshake exhaustion sustains server CPU 88× longer under PQC","PQC-TLS drops IDS recall to ~50% with AU-ROC near 0.49","PQC handshakes amplify DDoS load 88× and blind deep IDS","Classical-trained IDS miss PQC-TLS exhaustion at random levels"],"cache_read_input_tokens":128,"weakest_assumption_plain":"The single-server, ten-attacker cloud testbed with the chosen PQC suites, attack intensity, and mixed traffic is representative enough that the 88-times CPU-duration amplification and IDS blind spots transfer to production PQC-TLS deployments and fairly configured detectors.","fun_headline_variants_meta":{"raw":{"variants":["PQC-TLS extends high-CPU handshake attacks by 88×","Handshake exhaustion sustains server CPU 88× longer under PQC","PQC-TLS drops IDS recall to ~50% with AU-ROC near 0.49","PQC handshakes amplify DDoS load 88× and blind deep IDS","Classical-trained IDS miss PQC-TLS exhaustion at random levels"]},"model":"grok-4.5","effort":"low","cost_usd":0.004164,"raw_usage":{"total_tokens":1360,"prompt_tokens":903,"num_sources_used":0,"completion_tokens":86,"cost_in_usd_ticks":41640000,"prompt_tokens_details":{"text_tokens":903,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":371,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":903,"tokens_out":86,"duration_ms":3381,"temperature":1.0,"reasoning_tokens":371,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-15T05:34:19.830206+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Re-run the identical handshake-exhaustion workload against a multi-core or hardware-accelerated PQC-TLS server and measure whether sustained high-CPU duration still multiplies by tens of times versus classical TLS; retrain the same IDS models on balanced PQC traffic and check whether AU-ROC remains near 0.5.","supporting_citations":[],"review_version":1}