{"id":"0b5a19dc-7f37-4156-b919-4dd8ee958732","arxiv_id":"1908.07145","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The paper derives a closed-form correlation between template matching test items and shows that whitening the per-block counts produces statistically independent randomness tests.","lead":"This paper derives the joint distribution of p-values from two non-overlapping template matching randomness tests and proposes a linear transformation that makes the test items independent. The method could let the NIST SP800-22 test suite use a single rational significance level across its many template-based items.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The claimed exact independence after Eq. (40) rests on an unproved multivariate-CLT step and a neglected O(m/M) correlation remainder; without a direct residual-independence check the central claim is only asymptotic.","rationale":"The paper's analytic pairwise correlation formula and the large simulation in Figure 3 are genuine supporting evidence; the method is useful and the experiment is substantial. The weakest point is the gap between the asymptotic normal model used to construct LQ and the finite-M reality. Exact independence is a very strong property and cannot follow from the unproved multivariate CLT or from pairwise correlations alone. The O(m/M) remainder is small in the tested regime but is larger than stated in Eq. (31), so the headline claim should be read as approximate. This does not invalidate the method for SP800-22, where M=125000, but it should be stated and tested. The reader's conditional verdict is appropriate because the concern does not overturn the central claim on current evidence but does require explicit caveats and verification. The proposed empirical checks on residual correlations and normality would settle whether the approximation is benign in practice.","tokens_in":10792,"tokens_out":31506,"duration_ms":321443,"concrete_test":"Generate 10^6 null sequences (n=10^6, N=8) and compute C'_j from Eq. (40) for the 145 templates. Check (i) the empirical correlation matrix of C'_j against I_145: report the maximum off-diagonal |correlation| and a chi-square Q-Q plot of the eigenvalues; (ii) a multivariate normality test on C_j (e.g., Mardia skewness/kurtosis) or at least marginal Q-Q plots for several components. If the maximum off-diagonal correlation is below the Monte Carlo standard error (about 1.5 x 10^-4 for 10^6 blocks) and the normality checks do not flag, the practical independence claim survives; otherwise it is only approximate. Also compute the smallest eigenvalue of the 145 x 145 Sigma after the three removals to confirm it is safely positive and the inverse in Eq. (40) is well defined.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim---that replacing C_j by C'_j=(LQ)^{-1}C_j yields test items that are 'independent of each other' (Section 4.2)---requires C_j to be exactly multivariate normal with covariance Sigma whose off-diagonal entries are the asymptotic rho(T^(k),T^(l)) of Eq. (31). Two conditions are load-bearing and neither is established at the finite block length M=125000 used in the experiment. First, the multidimensional extension of Theorem 2.1 is asserted without proof; Cramer-Wold would supply it, but the paper never gives the argument, and exact multivariate normality is impossible for bounded count variables, so the transformed components are only approximately standard normal. Second, Eq. (31) drops a remainder that is actually O(m/M) in rho, not the stated O(m^2/(2^m M)): the numerator error is O(m/2^m), and dividing by the variance M/2^m gives O(m/M). At M=125000 this is about 7 x 10^-5, too small to matter in Figure 3, but it means the statement 'independent' is approximate, not exact. The paper should state the asymptotic nature of the claim explicitly and verify the residual correlations after whitening.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies the Non-overlapping Template Matching Test of NIST SP800-22 and addresses the fact that the 148 test items based on different 9-bit templates are not independent. It derives the joint distribution of two p-values under the null by modeling each block's standardized count as a bivariate normal variable, and it obtains an explicit joint cumulative distribution function in Eq. (35). The paper then proposes an orthogonalization transformation C'_j = (LQ)^{-1} C_j in Eq. (40), where L and Q diagonalize the covariance matrix built from the pairwise correlation coefficients in Eq. (31), and claims that the transformed test items are independent. The claim is supported by experiments using the Mersenne twister and AES-128.","tokens_in":10936,"tokens_out":18844,"duration_ms":173555,"significance":"The problem is practically important: using multiple template-matching tests with a joint significance level requires understanding and removing the dependence among test items. The paper's main contributions are a closed-form pairwise correlation coefficient, a detailed derivation of the joint CDF of two p-values, and a whitening procedure for the 148 test items. A notable strength is that the correlation coefficients are computed from the analytic formula rather than fitted, and the experimental validation uses two independent generators. If the asymptotic gaps are closed, the orthogonalization method could be useful for defining an explicit overall criterion in NIST SP800-22. However, the central claim that the transformed test items are 'independent of each other' is stronger than what is actually established, because the underlying multivariate normality and the covariance matrix are only asymptotic approximations.","major_comments":[{"comment":"The count of ordered pairs (k,l) in {1,...,M-m+1}^2 with |k-l| >= m is (M-2m+1)(M-2m+2), not (M-m+1)(M-3m+2) as written in Eqs. (19) and (28). The difference is m(m-1), which after division by 2^{2m} is O(m^2/2^{2m}) and therefore does not change the leading-order correlation coefficient. However, the remainder statement in Eq. (31) is not correct as stated. If one retains the exact factor M-m+1 in the near-diagonal sum, the finite-M correction to rho is of order 1/M with a coefficient depending on the template overlap indicators; if one instead replaces every M-m+1 by M before forming the ratio, the residual is O(m/M). In either case the bound O(m^2/(2^m M)) given in Eq. (31) is too small. The paper should replace this with a correct bound and explicitly state that the rho values used to build the covariance matrix are asymptotic quantities.","section":"Eqs. (19), (28), and (31)"},{"comment":"The multidimensional extension of Theorem 2.1 is asserted without proof, and exact multivariate normality is impossible for the bounded, discrete block counts. Consequently, Eq. (43) and the statement that the components of C'_j 'follow the standard normal distribution independent of each other' hold only asymptotically. The paper should either supply a rigorous multivariate CLT argument, for example via Cramér-Wold applied to the m-dependent sequence of indicators, or cite an appropriate theorem. It should also state consistently that the independence after the transformation is asymptotic. In addition, Section 4.2 should report a direct residual-dependence check, such as the maximum absolute empirical correlation among the transformed test items over the 10^6 sequences, rather than relying solely on the rejection-count histogram in Figure 3.","section":"Section 3.1 and Section 4.1"},{"comment":"The procedure for removing templates to make the covariance matrix nonsingular is incomplete. The text gives examples of dependent template groups (100000000 and 000000001; 011111111 and 111111110; and the four templates 001010101, 010101011, 101010100, 110101010) but does not prove that removing one element from each of these groups actually leaves a nonsingular 145x145 covariance matrix, nor does it explain how the complete list of dependent templates was obtained. Since the transformation in Eq. (40) requires det(Sigma) != 0, the paper should report the eigenvalues of Sigma before and after the removals, or otherwise verify that the remaining 145-template covariance matrix is invertible.","section":"Section 4.1, template removal"}],"minor_comments":[{"comment":"The symbol sigma is defined by Eq. (4) as M(1/2^m - (2m-1)/2^{2m}), which is the variance, not the standard deviation, of c_j. The standardization in Eq. (2) then incorrectly divides by the variance instead of the standard deviation. This notational inconsistency should be fixed, for example by writing sigma^2 in Eq. (4).","section":"Eq. (4)"},{"comment":"The agreement between the experimental and theoretical joint distributions in Figures 1 and 2 is assessed only visually. A quantitative discrepancy measure or a formal goodness-of-fit statistic for the two-dimensional histograms would make the validation more convincing.","section":"Section 3.2"},{"comment":"Figure 3 shows the rejection-count histogram before and after transformation but does not provide confidence bands or a goodness-of-fit test against the expected binomial distribution. Adding these would strengthen the claim that the transformed test items are independent.","section":"Section 4.2 and Figure 3"},{"comment":"The text says 'we need to remove either 100000000 or 000000001, and 011111111 or 111111110. Finally, we need to remove 001010101 or 010101011 or 101010100 or 110101010,' but the experiment removes 100000000, 111111110, and 001010101. The relation between the general removal rule and the specific templates excluded in the experiment should be stated explicitly.","section":"Section 4.1"},{"comment":"Reference [4] contains 'at el.' which should be 'et al.'; the equation display after Eq. (53) in the appendix has a typographical artifact 'd¯Xd¯Y' that should be cleaned up; and the caption of Figure 3 lacks axis labels.","section":"References and typography"}],"recommendation":"major_revision","confidential_remarks":"The paper addresses a relevant problem and the main idea is sound, but the rigor level needs improvement: the multivariate CLT step is a genuine gap, the remainder bound in Eq. (31) is incorrect, and the claimed exact independence should be stated as asymptotic. These issues are fixable within the paper's scope, so I would not reject the manuscript on these grounds."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a genuinely useful paper for anyone working with NIST SP800-22. It gives the first analytic pairwise correlation formula for the Non-overlapping Template Matching Test items, and the orthogonalization transformation is a sensible way to get independent test items. The core idea is sound and the experiment supports it. But the paper needs revision before I'd rely on it: the \"independent\" claim is overstated, and a few derivations have rough edges.\n\nWhat's new: Eq. (31) and the joint p-value CDF (35)-(36) are not in the earlier empirical literature. The whitening step is textbook, but the analytic covariance matrix is what makes it concrete. The 10^6-sequence simulations with two generators are a real check.\n\nSoft spots, in proportion:\n\n- The multivariate CLT for the vector of standardized counts is asserted in one sentence. A Cramér-Wold argument would work but is not given. Since counts are bounded, the transformed components are only approximately standard normal. Fine asymptotically, but the paper should own the approximation.\n\n- The exact count of disjoint window pairs in (19) and (28) is off: it should be (M-2m+1)(M-2m+2), not (M-m+1)(M-3m+2). The error is O(m^2/2^{2m}) in the raw covariance and translates to a remainder of order O(m^2/(2^m M)) after dividing by the variance—which is what the paper states, but only after a cancellation that isn't shown. The stress-test's claim of an O(m/M) remainder doesn't land on reading; the count mismatch gives a smaller term. Still, the derivation is sloppy and should be rewritten.\n\n- The post-transformation test statistic is never written down explicitly. It's clear from context that you sum squares of the transformed components over blocks, but it should be stated.\n\n- Figure 3 has no error bars or goodness-of-fit, but with 10^6 sequences the visible agreement is convincing enough.\n\n- The template removals based on eigenvector checks are asserted, not shown. Minor.\n\nNone of this breaks the central result. The whitening method will give approximately independent items in the large-M limit, and the experiment supports it. The reader's conditional verdict is right. Who this is for: anyone doing randomness testing, especially with NIST SP800-22; also people applying whitening to discrete count statistics. It deserves a serious referee; I'd send it for peer review with a request for revisions, not desk-reject.","headline":"Useful analytic dependence structure for NIST template tests with a sound whitening idea, but the paper overstates exact independence and needs a few corrections before I'd fully trust it.","tokens_in":11557,"tokens_out":5158,"would_cite":true,"duration_ms":48366,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F03","62H20","62E15","60F05"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that a linear transformation of the standardized template counts, $C'_j = (LQ)^{-1}C_j$, renders the Non-overlapping Template Matching Test items independent under the null hypothesis.","keywords":["randomness tests","NIST SP800-22","non-overlapping template matching","p-value independence","orthogonalization","multivariate normal approximation","chi-square test","correlation coefficient"],"falsifier":"Generate a large null sample with block length $M$ no larger than, say, $2^m$ for $m=9$, apply the orthogonalizing transform (Eq. (40)), and test the empirical covariance of the transformed per-block vectors or the joint distribution of the transformed p-values against the i.i.d. uniform prediction; any systematic off-diagonal correlation or nonuniformity at that $M$ would show that the infinite-$M$ normality and remainder-neglect assumptions are doing real work.","tokens_in":10475,"feed_emoji":"🎲","tokens_out":12085,"duration_ms":113265,"temperature":0.7,"pith_summary":"The paper targets a known flaw in the NIST SP800-22 randomness test suite: the 148 test items based on the Non-overlapping Template Matching Test are not independent, so no single significance level applies to the suite as a whole. It derives the joint distribution of two such test items' p-values from the multivariate normal approximation of the per-block template counts, and then proposes an orthogonalizing transformation that produces independent test items. After removing three templates whose counts are deterministically tied, the remaining 145 items are claimed to behave as independent tests. The experimental section, using one million $10^{6}$-bit sequences from two different generators, reports that the number of rejecting items per sequence matches the independence expectation. If the claim holds, randomness test users can fix a rational suite-wide criterion with an explicit significance level.","feed_headline":"A linear transform makes NIST's 145 template tests independent","feed_subtitle":"Independent p-values mean the whole suite can keep one explicit significance level.","key_machinery":"The machinery is the orthogonalization of the multivariate normal vector $C_j$. After computing the $R\\times R$ covariance matrix $\\Sigma$ from the pairwise correlation formula (Eq. (31)), the paper chooses an orthogonal matrix $L$ and a diagonal scaling $Q$ so that $Q^\\top L^\\top \\Sigma^{-1} LQ = I$, and transforms $C_j$ to $C'_j = (LQ)^{-1}C_j$; the proof is that the density of $C'_j$ becomes a product of standard normals. The distributional identity underlying the two-test-item dependency analysis is Eq. (35), the joint CDF $F_{N,\\rho}(X_N,Y_N)$, an infinite series in $\\rho^{2r}$ involving incomplete gamma functions, obtained by inverting the characteristic function of the correlated chi-square pair.","core_discovery":"The central discovery is that the dependency among Non-overlapping Template Matching Test items can be removed by a linear change of variables. For large block length $M$, the standardized per-block count vectors $C_j$ are treated as jointly normal with a covariance matrix $\\Sigma$ whose off-diagonal entries are the pairwise correlations $\\rho(T^{(k)}, T^{(\\ell)})$ given by Eq. (31). Writing $\\Sigma^{-1} = L\\Lambda L^\\top$ and rescaling by $Q$, the map $C'_j = (LQ)^{-1}C_j$ has identity covariance, so the components of $C'_j$ are standard normal and independent; hence the chi-square statistics built from each component are independent test items. The paper also derives the explicit joint cumulative distribution of two chi-square statistics (Eq. (35)) and of the two p-values, verified by comparing experimental and theoretical two-dimensional p-value distributions. Removing templates that create zero eigenvalues, the author reports that the number of test items rejecting a sequence follows the independent-items expectation for 145 templates.","pith_inferences":["At finite block length $M$ the neglected remainder $O(m^2/(2^m M))$ in Eq. (31) means the orthogonalized items are only asymptotically independent; a practical guide would need to state how large $M$ must be for the approximation to hold.","The same recipe might extend to other parametric NIST tests whose block statistics are approximately normal, potentially turning the whole SP800-22 suite into independent items, but the paper does not demonstrate this.","Independence under the null does not by itself say how the transformed items behave under a biased generator; power against alternatives is a separate question the paper does not address.","The derived joint CDF (35) could be used as a standalone copula to calibrate pairs of dependent p-values, which may be useful when only a small subset of templates is relevant."],"forward_implications":["Under the independence claim, the number of the 145 template-based items that reject a sequence at level $\\alpha$ is binomially distributed, so a suite-wide pass/fail rule can be set with known Type I error.","The explicit two-p-value joint CDF (Eq. (35)) lets a user assess or correct for dependence between any pair of templates without running the full orthogonalization.","A single combined p-value for the Non-overlapping Template Matching Test family can be obtained, e.g., by Fisher's method, since the transformed items are independent under the null.","The construction explains which templates must be excluded (e.g., 100000000 and 000000001) to avoid singular covariance, turning a numerical necessity into a principled test-set reduction."],"supporting_citations":[{"why":"Defines the NIST SP800-22 suite and the Non-overlapping Template Matching Test whose 148 default template items are the objects of the dependence analysis.","marker":"[4]"},{"why":"Provides the central limit theorem for dependent random variables used to justify the one-dimensional normality of per-block counts and, by extension, the multivariate normality assumption.","marker":"[14]"},{"why":"Supplies the Mersenne Twister generator used to produce one million 10^6-bit sequences for the experimental verification of the joint distributions and the orthogonalization.","marker":"[15]"},{"why":"Supplies AES-128 in counter mode as a second generator in the orthogonalization experiment, checking that the independence result is not specific to one generator.","marker":"[16]"}],"fun_headline_variants":["Independence achieved for NIST template tests via linear map","Breaking template test correlations with a single transform","How to make all 145 template tests truly independent","Linear transform decouples NIST's template p-values","One transform erases template test dependencies"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole construction rests on treating the per-block template counts, after standardizing them, as jointly normal already at the block sizes used in practice, and on ignoring the finite-block-size remainder in the correlation formula.","fun_headline_variants_meta":{"raw":{"variants":["Independence achieved for NIST template tests via linear map","Breaking template test correlations with a single transform","How to make all 145 template tests truly independent","Linear transform decouples NIST's template p-values","One transform erases template test dependencies"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000115,"raw_usage":{"total_tokens":1022,"prompt_tokens":844,"completion_tokens":178,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":460,"completion_tokens_details":{"reasoning_tokens":105}},"tokens_in":460,"tokens_out":178,"duration_ms":2503,"temperature":1.0,"reasoning_tokens":105,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T12:25:53.801760+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Generate a large null sample with block length $M$ no larger than, say, $2^m$ for $m=9$, apply the orthogonalizing transform (Eq. (40)), and test the empirical covariance of the transformed per-block vectors or the joint distribution of the transformed p-values against the i.i.d. uniform prediction; any systematic off-diagonal correlation or nonuniformity at that $M$ would show that the infinite-$M$ normality and remainder-neglect assumptions are doing real work.","supporting_citations":[{"cited_title":"A Statistical Test Suite for Random and Pseudorandom Number Gen- erators for Cryptographic Applications,","cited_arxiv_id":null,"evidence_quote":"Defines the NIST SP800-22 suite and the Non-overlapping Template Matching Test whose 148 default template items are the objects of the dependence analysis."},{"cited_title":"The central limit theorem for dependent random variables,","cited_arxiv_id":null,"evidence_quote":"Provides the central limit theorem for dependent random variables used to justify the one-dimensional normality of per-block counts and, by extension, the multivariate normality assumption."},{"cited_title":"Mersenne twister: a 623-dimensionally equidistributed uniform pseudo-random number generator,","cited_arxiv_id":null,"evidence_quote":"Supplies the Mersenne Twister generator used to produce one million 10^6-bit sequences for the experimental verification of the joint distributions and the orthogonalization."},{"cited_title":"Advanced encryption standard,","cited_arxiv_id":null,"evidence_quote":"Supplies AES-128 in counter mode as a second generator in the orthogonalization experiment, checking that the independence result is not specific to one generator."}],"review_version":1}