{"id":"eccaf304-bfb6-48e2-8a10-947baaf9b8dd","arxiv_id":"1908.06090","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"For paired-comparison experiments with second-order (three-attribute) interactions, D-optimal designs are uniform on one to three comparison depths, and the paper provides these depths and weights for many settings.","lead":"This paper derives optimal ways to design paired-comparison experiments, where people rate pairs of product profiles, when three attributes interact in their effect on preference. The results give benchmark designs for both full and partial profile experiments with any number of attribute levels.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Theorem 1's single-depth optimality is asserted without proof and its uniqueness fails (K=4, S=3, v=2 tie); the design characterization needs a tie-handling statement.","rationale":"The reader's conditional verdict is appropriate. The main structural claim, Theorem 3, appears salvageable: the variance function is a cubic with positive leading coefficient, and once the sign analysis is written out explicitly, the support set is forced to be contained in {S, d*, d*+1}. The variance formulas in Lemma 1 and Theorem 2 check out algebraically, and Corollary 1 is internally consistent. However, the weakest point is Theorem 1, which is stated as a theorem but justified only by numerical maximization of h3(d) for the cases in Table 1. The paper itself flags this by saying the table values were obtained by calculating h3(d) and determining the maximum. Moreover, uniqueness of the maximizing depth is not merely unproved; it is false for K=4, S=3, v=2, where h3(1)=h3(3). This tie does not destroy single-depth optimality, but it does mean the phrase 'the optimal comparison depth' and the later use of d* need a precise tie-breaking rule or a weaker existence statement. The full-vector designs in Table 2 are only numerically verified, and verification values are given only for full profiles, so those entries are also less secure. Still, none of these issues overturns the main support characterization, so the verdict stays conditional rather than being upgraded or rejected.","tokens_in":17503,"tokens_out":27004,"duration_ms":265044,"concrete_test":"Evaluate h3(d) from Lemma 1 for (K,S,v)=(4,3,2) and check whether h3(1)=h3(3). If the values coincide, Theorem 1's uniqueness assertion is false and a tie-handling rule is required before Table 2's d* can be called 'the' optimal comparison depth.","verdict_should_be":"UNCHANGED","load_bearing_attack":"After Theorem 1, the paper states that Table 1 was obtained 'by first calculating the values of h3(d) and determining the maximum'; no general proof of optimality is given, and Theorem 1 is worded as if the maximizing depth d* is unique. The uniqueness claim is false already inside the table: for K=4, S=3, v=2, Lemma 1 gives h3(d) proportional to d(4d^2-18d+20), so h3(1)=h3(3) and both depths are D-optimal for the second-order-interaction block. Thus Table 1's entry '1' is only one of two valid depths. The later construction in Table 2 (e.g. K=4, S=3, v=2 entry (1,0.900)) and the phrase 'the optimal comparison depth' depend on knowing a maximizer of h3 and on tie-breaking. For untabulated (K,S,v) the paper gives no proof that h3 has a single maximizer, and this example shows that non-adjacent ties occur. This does not by itself break Theorem 3's support bound, but it means the claimed benchmark characterization 'supported on S, d*, d*+1' is not fully specified unless ties are resolved and the existence/choice of d* is proved.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies approximate D-optimal designs for linear paired comparison experiments in which the alternatives are described by K attributes with a common number v of levels, under full or partial profiles of strength S, and the model includes main effects, first-order interactions, and second-order (three-attribute) interactions. The design region is partitioned into orbits of fixed comparison depth d, and the paper derives the information matrix of uniform designs on these orbits (Lemma 1), the variance function of invariant designs (Theorem 2), and a formula for the variance of single-depth designs (Corollary 1). The main structural claims are that a single depth d* optimizes the second-order interaction block (Theorem 1) and that the D-optimal design for the full parameter vector is supported on at most three depths, S, d*, and d*+1 (Theorem 3). Numerical tables list optimal depths and weights for selected values of K, S, and v, with Kiefer-Wolfowitz checks reported for full profiles in Table 3.","tokens_in":17734,"tokens_out":22416,"duration_ms":207134,"significance":"If the main characterization is correct, the paper provides benchmark D-optimal designs for paired comparison experiments with second-order interactions for general level numbers, extending earlier results for binary attributes and for first-order interaction models. The derivations are largely self-contained: Lemma 1 is proved by explicit combinatorial counts, Theorem 2 gives a closed-form variance formula, and the D-optimality of the tabulated full-profile designs is checked numerically by the Kiefer-Wolfowitz equivalence theorem. The formulas for h3(d) and lambda(d) are concrete and directly usable. However, the proof of the central support theorem is only sketched, and the statement of Theorem 1 is not correct as written; these issues are repairable but currently leave the main structural claims unsubstantiated.","major_comments":[{"comment":"The theorem asserts the existence of a single comparison depth d* that maximizes h3(d), but no proof is given and the uniqueness assertion is false as stated. For K=4, S=3, v=2, Lemma 1 gives h3(d) proportional to d(4d^2 - 18d + 20), so h3(1)=h3(3) > h3(2)=0; both depths 1 and 3 are D-optimal for the second-order interaction block, yet Table 1 reports only d*=1. Since the later construction in Table 2 and the phrase 'the optimal comparison depth' depend on a selected maximizer, the paper needs either a proof of uniqueness for all (K,S,v), which the example shows is impossible, or a consistent tie-handling rule and an explicit statement of which maximizer is used. The absence of such a statement also leaves the behavior for untabulated parameter values unsupported.","section":"Section 4, Theorem 1"},{"comment":"The proof of Theorem 3 is only a sketch and does not establish the claimed support pattern. From Theorem 2, V(d, xi) is a cubic polynomial in d with positive leading coefficient, but to conclude that equality V(d, xi*)=p can occur only at depths forming a set {S, d*, d*+1} one must analyze the sign of V-p on the integer depths; the manuscript merely says 'by the shape of the variance function' and omits the required sign-change argument. In addition, the symbol d* in Theorem 3 is not defined and does not coincide with the h3-maximizer of Theorem 1 in the paper's own examples: for K=5, S=4, v=2, Table 1 gives d*=4 for the interaction block, while the full-model design in Table 2 is supported on depths 2 and 4. A complete proof and a clear definition of d* are needed before the central characterization can be accepted.","section":"Section 4, Theorem 3"},{"comment":"The paper states that the D-optimality of the designs in Table 2 has been checked numerically via the Kiefer-Wolfowitz equivalence theorem, but Table 3 displays normalized variance values only for the full-profile case S=K. For the partial-profile rows (S<K), which form a substantial part of Table 2, no normalized variance values are shown, so the numerical verification is not reproducible from the manuscript and the optimality of those entries is not documented. Please provide the corresponding values for partial profiles or state explicitly where the verification can be found.","section":"Section 4, Tables 2 and 3"}],"minor_comments":[{"comment":"In the sentence following equation (12), 'f1(i1) ⊗ f1(i2) ⊗ f3(i3)' should read 'f1(i3)' instead of 'f3(i3)'.","section":"Section 3, Eq. (12)"},{"comment":"In the third displayed equation of the proof of Lemma 1, the term 'f1(jk) ⊗ f1(jl) ⊗ g(jm)' should be 'f1(jk) ⊗ f1(jl) ⊗ f1(jm)'.","section":"Appendix, proof of Lemma 1, third displayed equation"},{"comment":"The sentence 'For S >= 4 numerical computations indicate that at most two different comparison depths S and d* may be required' should be labeled as a conjecture or supported by a proof, since the paper currently gives no formal result covering this case.","section":"Section 4, after Theorem 3"},{"comment":"The caption 'boldface 1 corresponds to the optimal comparison depths d*' is confusing because the table entries are not visibly bold in the text; please clarify that entries equal to 1 are the normalized variance at support depths and that boldface marks the optimal depth.","section":"Section 4, Table 3 caption"},{"comment":"The first sentence of the abstract is grammatically awkward and should be rewritten for clarity.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":"The concrete tie in Table 1 for K=4, S=3, v=2 is a decisive counterexample to the uniqueness wording of Theorem 1 and should be addressed directly. The proof of Theorem 3 is a sketch that will need to be written out in full, including the sign analysis on the integer grid and a consistent definition of d*. The combinatorial derivations in Lemma 1 and Theorem 2 appear careful and are likely salvageable, so I see this as a major-revision rather than a rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nShort version: this paper gives the right next step from Graßhoff et al. (2003) by working out D-optimal invariant designs for paired comparisons with three-attribute interactions at general level counts, and the main formulas check out. But Theorem 1 is asserted rather than proved, and the uniqueness claim in it is actually false for a tabulated case. That needs fixing before the characterization is fully trustworthy.\n\nWhat is genuinely new: the h3(d) formula in Lemma 1, the variance formula in Theorem 2, and the support-size bound in Theorem 3. The derivations are explicit, combinatorial, and I had no trouble following the counting in the appendix. The h3(d) expression is closed-form and directly checkable; I verified the K=4, S=3, v=2 example. So the basic machinery is solid and re-implementable without code.\n\nThe soft spots are real but narrow. Theorem 1 states there exists a single comparison depth d* that is D-optimal for the interaction block, with no proof. The text says Table 1 came from calculating h3(d) and taking the maximum. That is a numerical claim, not a theorem. Worse, the uniqueness is false inside the table: for K=4, S=3, v=2, h3(d) is proportional to d(4d^2 - 18d + 20), which gives h3(1) = h3(3). So both d=1 and d=3 are optimal for that block, and Table 1 lists only one. The later Table 2 entries and the phrase \"the optimal comparison depth\" depend on tie-breaking, and for untested (K,S,v) there is no proof of a unique maximizer. This does not break Theorem 3's support bound, but it means the benchmark characterization is under-specified without a tie-handling statement.\n\nTheorem 3's proof is a sketch: the cubic variance function argument is correct in spirit, but the step where the three roots must be S, d*, d*+1 is not written out. Given the positive leading coefficient and the boundary conditions, it is repairable, but it is not fully rigorous as printed. The Kiefer-Wolfowitz checks in Table 3 are numerical; fine as evidence, but they don't replace a proof for all parameter values.\n\nThe citation pattern looks honest: the paper builds on Graßhoff et al. and on Nyarko-Schwabe, and says so. No fitting to data, no invented parameters.\n\nWho is this for? Anyone working on optimal design for conjoint analysis or partial-profile paired comparisons, especially as a benchmark for exact designs. It deserves a serious referee; the gaps are repairable, not fatal. My advice: send it out, and ask the author to prove or amend Theorem 1, address ties explicitly, and expand the sketch of Theorem 3.\n\nBest","headline":"Useful extension of paired-comparison optimal designs to second-order interactions, with a real gap in Theorem 1 that is repairable and a tie-failure that should be acknowledged.","tokens_in":18281,"tokens_out":720,"would_cite":true,"duration_ms":8719,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62K05","62J15","62K15"],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper proves that paired comparison experiments with second-order interactions have D-optimal designs supported on at most three comparison depths: S, d*, and d*+1.","keywords":["paired comparisons","D-optimality","second-order interactions","comparison depth","partial profiles","profile strength","invariant designs","conjoint analysis"],"falsifier":"Compute $h_3(d)$ from the explicit formula in Lemma 1 for a parameter triple not covered by Table 1, such as $K=11$, $S=10$, $v=9$ or $K=12$, $S=6$, $v=5$, and check whether the maximum over $d=1,\\ldots,S$ is attained at two distinct depths; if it is, Theorem 1 is false and the claimed three-depth support may fail. Alternatively, run a numerical search over invariant designs for the full parameter vector and look for a D-optimal design whose support contains four different comparison depths, which would directly contradict Theorem 3.","tokens_in":17269,"feed_emoji":"📊","tokens_out":8645,"duration_ms":74027,"temperature":0.7,"pith_summary":"The paper asks how to design paired comparison experiments, in which respondents rate or choose between two product profiles, when the statistical model includes three-attribute (second-order) interactions and the profiles may be partial instead of full. It proves that the D-optimal design never needs more than three distinct comparison depths (the number of attributes in which the two profiles differ): the full profile strength S and at most two adjacent intermediate depths $d^*$ and $d^*+1$. By restricting attention to invariant designs that are uniform on each depth, the information matrix becomes diagonal with explicitly known entries, so the optimal design is a simple weighted mixture of designs on these few depths. This reduces a combinatorial search over all possible choice sets to a small explicit calculation, yielding benchmark designs for real experiments.","feed_headline":"At most three profile depths give optimal paired comparisons","feed_subtitle":"The proof limits D-optimal choice-design search to a few simple mixtures, making benchmarks feasible.","key_machinery":"The central object is the comparison depth $d$, the number of attributes in which the two alternatives differ, and the uniform invariant design $\\bar{\\xi}_d$ that assigns equal weight to every pair of that depth. The load-bearing identity is the diagonal information matrix of Lemma 1: the matrix is block-diagonal with blocks proportional to $h_1(d)$, $h_2(d)$, and $h_3(d)$, where $h_3(d)$ is an explicit cubic in $d$ capturing the three-attribute interaction information. The Kiefer-Wolfowitz equivalence theorem converts D-optimality into a bound on the variance function $V(d,\\bar{\\xi})$, which is also a cubic in $d$; since a cubic can equal the parameter count $p$ at most three times, the support of an optimal design is limited to three depths, and the structure of the cubic identifies them as $S$, $d^*$, and $d^*+1$.","core_discovery":"In the second-order interactions model with $K$ attributes, $v$ common levels, and profile strength $S$, the paper proves that the uniform design on a single comparison depth $d^*$ is D-optimal for the second-order interaction block alone, where $d^*$ maximizes the explicit cubic function $h_3(d)$ (Theorem 1). For the full parameter vector, the D-optimal design is supported on at most three comparison depths, namely $S$, $d^*$, and $d^*+1$ (Theorem 3). The proof shows the variance function of any invariant design is a cubic polynomial in the depth $d$, so the equivalence theorem allows at most three depths to achieve the parameter-count bound; the shape of the cubic forces these depths to be exactly $S$, $d^*$, and $d^*+1$ when three are needed. Numerical tables list the optimal depths and weights for $K=4,\\ldots,10$ and $v=2,\\ldots,8$, with one, two, or three supporting depths depending on the parameters.","pith_inferences":["This suggests a general pattern: in paired comparison models with up to $r$-factor interactions, a D-optimal design for the full parameter vector may be supported on at most $r+1$ comparison depths; the paper's cubic argument proves the case $r=3$ but does not state the generalisation.","An experimenter could turn the benchmark design into a concrete exact design by rounding the optimal weights on the few supporting depths; the paper notes the benchmark use but does not detail the rounding step.","A proof that $h_3(d)$ is strictly unimodal, perhaps via its derivative, would close the only numerical gap and turn Table 1 into a theorem; that is a natural next step.","The invariance argument assumes all attributes have the same number of levels; extending to unequal level counts would require a new argument, since the diagonal information structure would no longer hold."],"forward_implications":["For every combination of attribute count $K$, level count $v$, and profile strength $S$, a D-optimal paired comparison design can be built by mixing uniform designs on at most the comparison depths $S$, $d^*$, and $d^*+1$.","The optimal depth $d^*$ for second-order interactions is the maximizer of the explicit cubic $h_3(d)$, so no numerical search over full designs is needed.","No single comparison depth can be D-optimal for main effects, first-order interactions, and second-order interactions simultaneously, since the three blocks have different preferred depths ($S$, an intermediate depth, and $d^*$).","For large numbers of levels $v$, the tables show the optimal intermediate depth $d^*$ drops to $S-2$ in the partial-profile case $S=K-1$, so optimal designs become less demanding of full profile differences.","These benchmark designs are starting points for constructing exact designs or fractions with a reasonable number of comparisons."],"supporting_citations":[{"why":"Supplies the formulas for h1(d) and h2(d) and the variance-function method that the paper extends to the second-order block.","marker":"Graßhoff et al. (2003)"},{"why":"Provides the equivalence theorem used to certify D-optimality of the constructed designs.","marker":"Kiefer and Wolfowitz (1960)"},{"why":"Establishes the binary-attribute case v=2 that this paper generalizes to arbitrary v.","marker":"Nyarko and Schwabe (2019)"},{"why":"Gives optimal full-profile designs for main effects and first-order interactions, the starting point of the model.","marker":"van Berkum (1987)"},{"why":"Introduces the approximate-design framework in which D-optimality is defined.","marker":"Kiefer (1959)"},{"why":"Provides the invariance argument that justifies restricting attention to uniform designs on comparison-depth orbits.","marker":"Schwabe (1996)"},{"why":"Determines optimal designs for main-effects models, used to identify the full-depth design as optimal for main effects.","marker":"Graßhoff et al. (2004)"}],"fun_headline_variants":["Cubic proof: only three comparison depths needed for D-optimality","D-optimal paired comparisons use at most three profile depths","Three depths enough: optimal paired comparisons proof","Pair comparisons: optimal design needs only depths S, d*, d*+1","Cut search: at most three depths for optimal paired comparisons"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The construction depends on the assumption that $h_3(d)$ has a unique maximizer $d^*$ among the depths $1,\\ldots,S$ for every combination of attribute count $K$, level count $v$, and profile strength $S$; the paper verifies this only numerically for the parameter values in its tables, offering no proof, so a parameter combination with two equal maxima would break the single-depth optimality of the interaction block and alter the support claimed in Theorem 3.","fun_headline_variants_meta":{"raw":{"variants":["Cubic proof: only three comparison depths needed for D-optimality","D-optimal paired comparisons use at most three profile depths","Three depths enough: optimal paired comparisons proof","Pair comparisons: optimal design needs only depths S, d*, d*+1","Cut search: at most three depths for optimal paired comparisons"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000197,"raw_usage":{"total_tokens":1280,"prompt_tokens":774,"completion_tokens":506,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":390,"completion_tokens_details":{"reasoning_tokens":420}},"tokens_in":390,"tokens_out":506,"duration_ms":5145,"temperature":1.0,"reasoning_tokens":420,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T12:57:36.424980+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compute $h_3(d)$ from the explicit formula in Lemma 1 for a parameter triple not covered by Table 1, such as $K=11$, $S=10$, $v=9$ or $K=12$, $S=6$, $v=5$, and check whether the maximum over $d=1,\\ldots,S$ is attained at two distinct depths; if it is, Theorem 1 is false and the claimed three-depth support may fail. Alternatively, run a numerical search over invariant designs for the full parameter vector and look for a D-optimal design whose support contains four different comparison depths, which would directly contradict Theorem 3.","supporting_citations":[{"cited_title":"and Wolfowitz, J","cited_arxiv_id":null,"evidence_quote":"Provides the equivalence theorem used to certify D-optimality of the constructed designs."},{"cited_title":"and Schwabe, R","cited_arxiv_id":null,"evidence_quote":"Establishes the binary-attribute case v=2 that this paper generalizes to arbitrary v."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Gives optimal full-profile designs for main effects and first-order interactions, the starting point of the model."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces the approximate-design framework in which D-optimality is defined."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the invariance argument that justifies restricting attention to uniform designs on comparison-depth orbits."}],"review_version":1}