{"id":"1bcba6eb-e685-4414-bf93-069cf93916f5","arxiv_id":"2506.19943","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"In a local DNS testbed, quantum-resistant MLKEM and Falcon keep DNS latency near classical levels, while SPHINCS+ and HQC add large bandwidth and CPU overhead.","lead":"This paper tests quantum-resistant versions of three DNS security protocols in a containerized lab and compares their speed, bandwidth, CPU, and memory. It finds MLKEM and Falcon are practical for DNS, while SPHINCS+ and HQC add heavy overhead.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Loopback, single-resolver benchmark supports qualitative algorithm-size findings, but the 'practical latency' half of the central claim is not yet tied to real DNS resolution.","rationale":"The paper's central claim is mostly supported by its own data: message-size increases for SPHINCS+ and HQC are intrinsic algorithm properties, and the high-concurrency CPU and bandwidth results consistently separate lattice-based from hash/code-based schemes. The concern I identify is narrower than the reader's general testbed-representativeness worry: the benchmark measures only a stub-to-resolver exchange over loopback with the test domain hosted on the resolver, so the reported single-query latency is not end-to-end DNS resolution. This does not overturn the qualitative ranking, but it does mean the 'practical latency' assertion should be read as a local-benchmark result until tested on a realistic network path. The paper itself acknowledges hardware and network dependence in Section 5, which is a point in its favor. I therefore keep the reader's CONDITIONAL verdict unchanged: the analysis is useful and honestly scoped, and the proposed WAN test would settle whether the latency component of the central claim generalizes. I do not see an internal inconsistency that would justify rejection, nor a need to accept unconditionally without the additional validation.","tokens_in":26987,"tokens_out":11637,"duration_ms":117981,"concrete_test":"Deploy the same OQS-BIND setup on a public cloud VM with a separate authoritative server, run clients from a different host over the Internet (or with `tc netem` adding 20-50 ms RTT), and measure per-query latency distributions over at least 10 repeated trials for the Table 3 configurations. If MLKEM/Falcon combinations remain the fastest and SPHINCS+/HQC remain substantially slower in both median and tail latency, the qualitative central claim survives; if the gaps collapse or the ordering changes, the claim must be explicitly scoped to local benchmark conditions.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim's latency conclusion rests on Tables 3-8, which measure `dig +tls` and `+https` against a local OQS-BIND resolver that hosts the test domain (Section 4.1). The query never traverses root/TLD/authoritative servers, and loopback has zero network RTT; the reported 9-20 ms figures therefore include dig process startup and WSL2/Docker overhead, not end-to-end DNS resolution. Section 5 concedes that results are hardware-dependent and may not generalize to wide-area deployments. If a realistic WAN RTT (e.g., 20-100 ms) dominates, the absolute latency gaps seen in Table 3 (9.1 ms for MLKEM512+Falcon512 vs. 16.4 ms for SPHINCS+) become a small fraction of user-perceived latency, and the phrase 'significantly increase ... processing overhead' would need to be scoped to CPU and batch throughput rather than single-query latency. The bandwidth and high-concurrency CPU findings are less affected because they follow from fixed key/signature sizes and measured CPU saturation, but the 'practical latency' component of the central claim is under-supported as stated. No error bars or trial counts are reported, so sub-millisecond differences (e.g., 9.10 vs. 9.17 ms) are not statistically distinguishable.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents PQC-DNS, a benchmarking framework and experimental study of post-quantum cryptography applied to DNSSEC, DNS-over-TLS (DoT), and DNS-over-HTTPS (DoH). The authors integrate Open Quantum Safe (OQS) libraries, oqsprovider, and a forked BIND9 (OQS-BIND) into a containerized local testbed, then measure latency, bandwidth, CPU, and memory for legacy-only, PQC-only, and hybrid configurations across NIST security levels 1, 3, and 5. The central claim is that lattice-based primitives such as MLKEM and Falcon provide practical latency and resource profiles, while hash-based and code-based schemes such as SPHINCS+ and HQC impose substantially larger bandwidth and processing overhead, especially under concurrency. The paper also provides a security threat taxonomy for PQC-enabled DNS and discusses deployment, sustainability, and standardization barriers.","tokens_in":27190,"tokens_out":3775,"duration_ms":41436,"significance":"If the measurements are interpreted within their stated environment, the paper is a useful system-level comparison: it is, to my knowledge, one of the few studies that benchmarks DNSSEC, DoT, and DoH together under PQC configurations, and it makes a concrete artifact available on GitHub. The bandwidth results, which follow from fixed key and signature sizes, are qualitatively robust and consistently show SPHINCS+ and HQC consuming several times the bandwidth of legacy and lattice-based combinations across all tables. The multi-threaded CPU findings also provide a plausible scalability picture. The main value is therefore as an empirical reference for operators considering PQC migration, provided the latency claims are properly scoped to the local, loopback testbed. The paper is honest about this limitation in Section 5, but the abstract and Section 4.3 currently overstate the strength of the latency evidence.","major_comments":[{"comment":"The latency half of the central claim is not supported by the presented measurements. The benchmark uses a local OQS-BIND resolver hosting the test domain, with queries issued over loopback; no root, TLD, or authoritative traversal occurs and no WAN RTT is included. The reported 9-20 ms values therefore include dig process startup and WSL2/Docker overhead, not end-to-end DNS resolution. The paper itself concedes in Section 5 that results are hardware-dependent and may not generalize to wide-area deployments. With a realistic WAN RTT of 20-100 ms, the 9.1 ms versus 16.4 ms gap in Table 3 becomes a small fraction of user-perceived latency. Please either add an explicit WAN-RTT sensitivity analysis, or restrict the abstract and Section 4.3 claims to CPU-bound processing and batch throughput rather than single-query 'practical latency' in real deployments.","section":"§4.1, Tables 3-8, abstract"},{"comment":"No statistical characterization is reported for any metric. All tables give point estimates without standard deviations, confidence intervals, or the number of repeated sessions used to compute averages. The text highlights sub-millisecond differences such as 9.10 versus 9.17 ms in Table 3 and 8.87 versus 8.84 ms in Table 4, but without variance or trial counts these differences are not statistically distinguishable. Please report dispersion measures and repetition counts for each configuration, or explicitly mark small differences as not significant. This is load-bearing because the claim that MLKEM/Falcon 'match or outperform' legacy combinations rests on such differences.","section":"§4.2, Algorithm 1; Tables 3-15"},{"comment":"The scalability narrative is internally inconsistent with the data. The text preceding Table 12 states that latency remains unaffected as bandwidth scales, but Table 12 shows roughly 110 ms for 100 concurrent queries versus about 9 ms for single-query tests, Table 14 shows about 1,000 ms for 1,000 queries, and Table 15 shows about 11,000 ms for 10,000 queries. This is approximately linear degradation, not unaffected latency. The paragraph must be rewritten to describe the observed near-linear scaling, and the labels '1000 concurrent queries' and '10000 concurrent queries' should be clarified: the text describes sessions of 100 workers sending one query each, so the relationship between sessions, workers, and total query count needs to be stated precisely.","section":"§4.3, Tables 12-15"}],"minor_comments":[{"comment":"The phrase 'DNSSEC-enabled configurations ans security levels' contains a typo; it should read 'and security levels'.","section":"§4.3, paragraph before Table 3"},{"comment":"'Meanwile' should be 'Meanwhile'.","section":"§4.3, paragraph before Table 12"},{"comment":"Algorithm naming is inconsistent: Table 9 uses 'falconpadded512' and 'sphincssha2128fsimple', while Tables 10 and 11 use 'falcon512' and 'sphincssha2128f'. Please clarify whether these are intentionally different variants and use consistent names throughout.","section":"Tables 9-11"},{"comment":"The text states that the lowest numerical values in each subcategory are bolded, but the rendered tables do not show bolding. Please ensure the final PDF preserves the formatting or remove the claim.","section":"§4.3, table formatting"},{"comment":"The reference format line lists the date as 2018; this should be updated to the actual submission year.","section":"ACM Reference Format"},{"comment":"The normalization of /usr/bin/time 'Percent of CPU this job got' by 16 vCPUs is not standard and may obscure differences for single-threaded dig processes. Please document the raw values or justify the normalization with respect to how the tool reports CPU percentage.","section":"§4.2, CPU normalization"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope for a networking/security venue and the artifact availability is a genuine strength. The main risks are the gap between the abstract's latency claim and the loopback-only benchmark, the absence of statistical reporting, and the contradictory scalability discussion. These issues are fixable within the manuscript's scope, so I recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know this is a benchmark paper, not a research breakthrough. What it does well is put DNSSEC, DoT, and DoH on the same testbed under legacy, hybrid, and PQC-only configurations, which the cited prior work mostly treats separately. The headline qualitative result—MLKEM paired with MLDSA or Falcon is cheap, while SPHINCS+ and HQC are expensive in bandwidth and CPU—is well supported by the tables. Anyone who knows the algorithm sizes will not be surprised, but having it measured across all three DNS security mechanisms in one place is genuinely useful. Credit also goes to the OQS-BIND integration work and the multi-threaded tests at 100, 1000, and 10000 concurrent queries; that stress dimension is the most interesting part of the paper. The Limitations section is refreshingly honest about hardware dependence and generalizability, and the threat taxonomy, while qualitative, is a reasonable checklist. The benchmark suite is linked on GitHub, which is good, though there is no commit hash or artifact version to pin reproducibility.\n\nThe soft spot is exactly the one the stress-test note flags. The latency numbers come from dig +tls and +https against a local OQS-BIND resolver over loopback. The query never traverses a real DNS hierarchy, and the reported single-digit-to-20 ms values include WSL2/Docker and process overhead. On a wide-area path with 20–100 ms of RTT, the absolute differences between MLKEM and SPHINCS+ shrink to a small fraction of user-perceived latency. The bandwidth and server CPU findings are much more robust because they follow from fixed key and signature sizes plus measured CPU saturation, but the \"practical latency\" half of the abstract's central claim is under-supported as stated. Also, the paper reports no standard deviations, confidence intervals, or trial counts. The 100-query runs are averaged somewhere, but we never see variance, so sub-millisecond differences between configurations are not statistically distinguishable.\n\nThe formal model in Section 3 is just descriptive arithmetic, not a derivation, and the security analysis is a literature-based taxonomy. That is fine, but it means the contribution is empirical, not theoretical.\n\nThis paper is for DNS operators and protocol engineers who want a starting point for choosing PQC algorithms and for estimating extra bandwidth and CPU in their deployments. It is not for cryptographers. I would send it to a serious referee, but with clear conditions: report variance and trial counts, add at least one wide-area or emulated-RTT scenario, and publish the scripts at a tagged commit. With those changes it becomes a citable reference. As is, the qualitative findings are solid enough to warrant peer review rather than desk rejection.","headline":"A useful, well-scoped benchmark of PQC for DNSSEC/DoT/DoH; the qualitative bandwidth and CPU findings hold, but the latency claims are tied to a loopback testbed and need stronger evidence.","tokens_in":27754,"tokens_out":2617,"would_cite":true,"duration_ms":29848,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Post-quantum DNS is practical with lattice-based MLKEM and Falcon, while SPHINCS+ and HQC strain bandwidth and CPU.","keywords":["post-quantum cryptography","DNS","DNSSEC","DNS-over-TLS","DNS-over-HTTPS","ML-KEM","Falcon","SPHINCS+"],"falsifier":"Run the same post-quantum-enabled resolver on a public cloud instance and measure end-to-end DoT latency for MLKEM512 with Falcon512 versus X25519 with RSA2048 under real network round-trip time and background load; if the lattice pair no longer matches or beats the classical pair's latency, or if HQC and SPHINCS+ become competitive because transmission time dominates, the paper's central performance ranking fails. A second check is to push a resolver to 10,000 queries per second with SPHINCS+ signatures and observe whether server CPU exceeds 50 percent with severe latency growth, as the paper predicts.","tokens_in":26765,"feed_emoji":"🌐","tokens_out":7861,"duration_ms":77357,"temperature":0.7,"pith_summary":"This paper sets out to determine whether the Domain Name System can be made quantum-resistant without breaking its performance budget. It builds and benchmarks a full DNS stack—DNSSEC, DNS-over-TLS, and DNS-over-HTTPS—under classical, hybrid, and post-quantum-only cryptographic configurations. The central finding is that lattice-based key encapsulation and signatures, specifically MLKEM paired with MLDSA or Falcon, resolve DNS queries at latencies comparable to or better than today's RSA and elliptic-curve settings across all three standardized security levels. Hash-based SPHINCS+ and code-based HQC, by contrast, multiply bandwidth by 2 to 23 times, raise server CPU sharply, and degrade badly under high concurrency. The practical upshot, if these measurements hold in wider deployments, is that operators can migrate DNSSEC, DoT, and DoH to post-quantum cryptography by choosing lattice schemes and by planning mitigations for the bandwidth and denial-of-service risks that larger signatures create.","feed_headline":"MLKEM and Falcon make post-quantum DNS practical","feed_subtitle":"Controlled DNS benchmarks show lattice schemes can replace RSA and ECDSA without a latency penalty.","key_machinery":"The load-bearing object is the PQC-DNS benchmarking framework: a two-phase protocol model consisting of a TLS 1.3 handshake using a post-quantum key encapsulation mechanism and signature, followed by encrypted recursive DNS resolution with optional DNSSEC validation, together with a performance vector S_i(k,s) = (latency, bandwidth, client CPU, server CPU, memory) measured for each cryptographic configuration. The framework runs on a containerized testbed with a patched BIND9 resolver linked to OpenSSL and an open-source post-quantum provider, and it formalizes total resolution cost as a sum of handshake, transport, and DNSSEC terms. That structure lets the authors compare every (KEM, signature) pair under identical conditions across legacy-only, PQC-only, and hybrid deployment classes, which is what allows them to separate algorithm-family effects from protocol-layer effects.","core_discovery":"On the paper's own terms, the discovery is that the choice of post-quantum primitive, not the protocol, determines whether DNS can go quantum-resistant cheaply. Across DoT and DoH at security levels 1, 3, and 5, MLKEM512/768/1024 combined with MLDSA44/65/87 or Falcon512/1024 matched or beat every legacy FFDHE/ECDSA/RSA combination in end-to-end latency, often landing at 8.8 to 9.3 milliseconds versus 9.2 to 18.7 milliseconds for classical pairs, while raising bandwidth roughly 1.5 to 6 times. HQC-based key encapsulation raised latency to about 19 milliseconds at security level 1 and above 65 milliseconds at level 5, and SPHINCS+ signatures lifted single-query bandwidth to 38 to 87 kilobytes and server CPU to 3 to 3.5 percent. Under 10,000 concurrent queries, MLKEM with Falcon kept latency near 10.8 seconds with server CPU around 9 percent, while SPHINCS+ configurations reached 21 to 23 seconds, server CPU above 54 percent, and total transfer volumes over 380 megabytes per 10,000-query session. The paper also argues that DNSSEC validation itself stays computationally cheap: MLDSA44 and Falcon-512 raise bandwidth but leave resolution latency nearly unchanged relative to RSA or Ed25519, and it identifies downgrade, timing, fragmentation, and DDoS amplification as the main new risks introduced by PQC adoption.","pith_inferences":["A testable extension beyond the paper is to repeat the benchmarks over a real wide-area network; with network round-trip time in the tens of milliseconds, the sub-10-millisecond cryptographic differences would shrink relative to transmission, which would likely widen the usability gap between MLKEM/Falcon and HQC/SPHINCS+ rather than close it.","The single-host, loopback testbed likely understates the cost of large signatures on shared infrastructure; under multi-tenant CPU contention, the server-CPU dominance of SPHINCS+ and HQC could translate into packet loss or timeouts, making the preference for lattice schemes even stronger in production.","Energy and edge-device implications follow implicitly: because SPHINCS+ inflates both transfer size and CPU, resolvers with tight power budgets would face disproportionate costs, so protocol-level fragmentation and caching optimizations deserve priority in standardization alongside algorithm choice."],"forward_implications":["DoT and DoH resolvers can adopt MLKEM with MLDSA or Falcon without adding noticeable per-query latency, so the first wave of production post-quantum DNS can use existing protocol stacks.","SPHINCS+ and HQC should be avoided for latency-sensitive or high-throughput recursive resolvers, because their bandwidth and CPU costs grow with concurrency and security level.","DNSSEC zones can move to MLDSA44 or Falcon-512 with only modest bandwidth growth and no measured latency penalty, making signature-based validation feasible in the near term.","Hybrid classical-plus-PQC cipher suites preserve the option of pairing legacy KEMs with PQC signatures, or vice versa, without collapsing performance, easing the transition before PQC-only clients are universal.","Operational mitigations—rate limiting, client puzzles, application-layer fragmentation, and strict cipher-suite binding—become necessary for safe deployment of the larger-signature schemes."],"supporting_citations":[{"why":"Supplies the patched BIND9 resolver used as the measurement subject for DNSSEC, DoT, and DoH.","marker":"[12]"},{"why":"Defines the standardized post-quantum algorithm set whose security levels structure the benchmark comparisons.","marker":"[23]"},{"why":"Establishes the hybrid classical-plus-PQC key-exchange technique that the hybrid deployment class builds on.","marker":"[37]"},{"why":"Provides the application-layer fragmentation method the paper identifies as the mitigation for large post-quantum signatures.","marker":"[11]"},{"why":"Documents how larger post-quantum handshakes can be weaponized for denial-of-service, grounding the DDoS threat analysis.","marker":"[36]"},{"why":"Gives the prior SPHINCS+ DNSSEC evaluation that this paper extends to encrypted DNS and system-level measurement.","marker":"[42]"},{"why":"Demonstrates side-channel leakage in lattice-based KEM implementations, supporting the paper's timing-attack threat discussion.","marker":"[15]"},{"why":"Defines hybrid TLS 1.3 binding rules relevant to the downgrade-resistance claims and mitigations.","marker":"[38]"},{"why":"Provides prior ML-KEM TLS 1.3 performance results that the paper's lattice-scheme latency findings align with.","marker":"[43]"}],"fun_headline_variants":["Lattice algorithms make DNS quantum-safe without latency loss","Post-quantum DNS: MLKEM and Falcon win on speed","Quantum-safe DNS: Lattice beats hash-based in benchmarks","DNS post-quantum upgrade: Choose lattice, avoid SPHINCS+","Benchmarks show lattice PQC ready for DNS deployment"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The benchmark assumes that a local containerized testbed on a single desktop CPU, using loopback networking and no competing traffic, represents how real DNS resolvers behave; if real-world network latency, middleboxes, or multi-tenant CPU contention change the balance between cryptographic processing and transmission, the quantitative latency and CPU rankings may not generalize, though the qualitative bandwidth findings likely do.","fun_headline_variants_meta":{"raw":{"variants":["Lattice algorithms make DNS quantum-safe without latency loss","Post-quantum DNS: MLKEM and Falcon win on speed","Quantum-safe DNS: Lattice beats hash-based in benchmarks","DNS post-quantum upgrade: Choose lattice, avoid SPHINCS+","Benchmarks show lattice PQC ready for DNS deployment"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000262,"raw_usage":{"total_tokens":1694,"prompt_tokens":1141,"completion_tokens":553,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":757,"completion_tokens_details":{"reasoning_tokens":466}},"tokens_in":757,"tokens_out":553,"duration_ms":5742,"temperature":1.0,"reasoning_tokens":466,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T22:59:53.042459+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same post-quantum-enabled resolver on a public cloud instance and measure end-to-end DoT latency for MLKEM512 with Falcon512 versus X25519 with RSA2048 under real network round-trip time and background load; if the lattice pair no longer matches or beats the classical pair's latency, or if HQC and SPHINCS+ become competitive because transmission time dominates, the paper's central performance ranking fails. A second check is to push a resolver to 10,000 queries per second with SPHINCS+ signatures and observe whether server CPU exceeds 50 percent with severe latency growth, as the paper predicts.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the patched BIND9 resolver used as the measurement subject for DNSSEC, DoT, and DoH."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the standardized post-quantum algorithm set whose security levels structure the benchmark comparisons."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes the hybrid classical-plus-PQC key-exchange technique that the hybrid deployment class builds on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Documents how larger post-quantum handshakes can be weaponized for denial-of-service, grounding the DDoS threat analysis."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Gives the prior SPHINCS+ DNSSEC evaluation that this paper extends to encrypted DNS and system-level measurement."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Demonstrates side-channel leakage in lattice-based KEM implementations, supporting the paper's timing-attack threat discussion."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines hybrid TLS 1.3 binding rules relevant to the downgrade-resistance claims and mitigations."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides prior ML-KEM TLS 1.3 performance results that the paper's lattice-scheme latency findings align with."}],"review_version":1}