{"id":"1cc67428-594c-44d4-9b38-c1e98896a5c2","arxiv_id":"2506.00889","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"For any two outcome probabilities, the complementary log ratio approximates the risk ratio more closely than the odds ratio does, with a proof using the Aranda-Ordaz family of link functions.","lead":"This paper proposes using complementary log-log models instead of logistic regression when researchers want to approximate risk ratios for binary outcomes. It proves, within a family of link functions, that the complementary log ratio is always at least as close to the risk ratio as the odds ratio.","discovery_kind":"first_principles","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Corollary 1 is mathematically sound; the load-bearing gap is the unsupported extrapolation from estimand-level comparison to finite-sample regression practice.","rationale":"The reader's conditional verdict is appropriate, and I agree with the assessment that the mathematical core is sound. My independent re-derivation of Theorem 1 and Corollary 1 found only a minor derivative typo in the eAppendix that does not affect the convexity argument. The main unresolved issue is the scope of the central claim: the theorem compares fixed estimands, whereas the paper's stated purpose is to recommend an estimation method for practitioners. That practical leap needs finite-sample evidence, which is promised in the abstract but absent from the supplied text. I also note that the statement 'λ=0 gives the smallest approximation error within the Aranda-Ordaz family' is true only on the paper's explicit domain 0 ≤ λ ≤ 1; extending the same formula to λ = −1 gives W_{−1}(θ)=θ and WR(−1)=RR exactly, so the domain restriction is doing real work. None of this overturns Corollary 1; it narrows the practical claim to what is actually proved.","tokens_in":6188,"tokens_out":12703,"duration_ms":134436,"concrete_test":"Run a finite-sample simulation on a grid of saturated models (e.g., p1 in {0.2, 0.5, 0.8}, p0 chosen so RR is in {0.5, 1.25, 2}, n in {100, 500, 2000}, ≥1000 replicates). Fit logistic and cloglog GLMs to each simulated binary sample, compute the estimated OR and exp(beta_cloglog), and compare both to the true RR using relative bias, MSE, and the proportion of replicates where the cloglog-based estimate is closer. If the cloglog estimates are not at least comparable on MSE, or if the simulation results are not reported, the abstract's simulation claim and the Section 4 recommendation require qualification.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Corollary 1 is a deterministic, estimand-level statement: for fixed 0 < p0 ≠ p1 < 1, the maximal relative discrepancy B(λ) is strictly increasing on 0 ≤ λ ≤ 1, so B(0) < B(1). The proof in eAppendix A.1 is essentially correct, although the displayed second derivative should be (λ+1)(1−θ)^{−λ−2}; convexity, and hence Lemma 1, still holds. The load-bearing gap is the transition from this inequality to the paper's practical recommendation (Section 4) that complementary log-log models should be preferred to logistic models for approximating risk ratios. In any fitted GLM, the estimated OR and the estimated CLR (exp of the cloglog coefficient) are random; a comparison of estimands does not control their sampling variability, standard errors, or mean squared error around the true RR. The abstract states that 'Simulation studies further reinforce our theoretical findings,' but no simulation section, table, or code appears in the supplied manuscript, so the finite-sample support for the practical claim cannot be checked. If the cloglog estimator has larger variance or bias at realistic sample sizes and prevalences, the recommendation may fail even though the estimand inequality is true.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper defines the complementary log ratio (CLR) as log(1−p1)/log(1−p0), the estimand of a complementary log-log regression, and compares it with the odds ratio (OR) as an approximation to the risk ratio RR = p1/p0. Using the Aranda-Ordaz family Wλ(θ) with λ=0 giving the cloglog transformation and λ=1 giving the logit, the paper proves Lemma 1 (WR(λ) always lies on the same side of RR as the inequality between p0 and p1), Theorem 1 (the maximum relative discrepancy B(λ) is strictly increasing on [0,1]), and Corollary 1 (CLR is closer to RR than OR under the maximum relative discrepancy measure). The paper then recommends complementary log-log models as a practical alternative to logistic models for risk ratio approximation.","tokens_in":6331,"tokens_out":7387,"duration_ms":69611,"significance":"The theoretical result is clean, elementary, and parameter-free: for any fixed p0 and p1, CLR is a better approximation to RR than OR in the maximum-relative-discrepancy sense, and λ=0 is the minimizer within the restricted Aranda-Ordaz family. This is a useful reference result for categorical data analysis and is proved with a short convexity and monotonicity argument. The strength of the manuscript lies in this deterministic, estimand-level ordering. Its limitation is that the practical recommendation in Section 4 and the abstract's claim of simulation support go beyond what the proof establishes, because nothing is said about sampling variability or the behavior of estimated CLR in finite samples.","major_comments":[{"comment":"The paper's practical recommendation—that practitioners should prefer complementary log-log models to logistic models when the goal is to approximate risk ratios—is based on Corollary 1, which is a deterministic statement about the true estimands CLR and OR for fixed p0 and p1. In any fitted GLM, the estimated OR and estimated CLR are random variables, and the proof says nothing about their sampling variability, standard errors, or mean squared error around the true RR. A lower-bias estimand can still yield a worse practical approximation if its estimator has larger variance, so the transition from Corollary 1 to the practical claim in Section 4 requires additional support, e.g., asymptotic variance calculations, a finite-sample simulation study, or a substantially qualified conclusion.","section":"Section 4 (Discussion) and Corollary 1"},{"comment":"The abstract states that 'Simulation studies further reinforce our theoretical findings,' but the supplied manuscript contains no simulation section, table, or code. Since the practical claim rests on finite-sample behavior, the absence of the promised simulations makes that support uncheckable. The authors should either add a simulation study that compares estimated CLR and estimated OR in fitted regressions, or remove and explicitly qualify the simulation claim.","section":"Abstract and full text"}],"minor_comments":[{"comment":"The displayed second derivative of Wλ(θ) is incorrect: it should be (λ+1)(1−θ)^{−λ−2}, not λ(1−θ)^{−λ}. The convexity conclusion is unaffected because (λ+1)>0, but the formula should be corrected.","section":"eAppendix A.1 (proof of Lemma 1)"},{"comment":"The proof invokes 'without loss of generality' that 0<p0<p1<1, but it does not explicitly justify why this is WLOG. The justification is that B(λ) is symmetric under exchanging p0 and p1, so adding one sentence about this symmetry would make the argument complete.","section":"eAppendix A.1 (proof of Theorem 1)"},{"comment":"The citation 'Zhang and Kai (1998)' is incorrect: the author is Kai F. Yu (JAMA 280(19), 1690–1691, 1998), so the in-text citations and the reference entry 'Zhang, J. and F. Y. Kai (1998)' should be corrected to 'Zhang, J. and Yu, K. F. (1998)'.","section":"References and in-text citations"},{"comment":"The displayed definition of B(λ) has a line-break artifact that makes it appear as max(RR WR(λ), WR(λ) RR); it should be typeset clearly as max(RR/WR(λ), WR(λ)/RR). The surrounding text makes the intent clear, but the equation should be reformatted.","section":"Section 3, definition of B(λ)"}],"recommendation":"major_revision","confidential_remarks":"The core mathematical result appears correct and is a genuine contribution as an estimand-level comparison. The main deficiency is the disconnect between that theoretical result and the practical recommendation, amplified by the abstract's unfulfilled promise of simulation studies. A revision that adds a simulation study comparing fitted cloglog and logistic regressions, or that narrows the claims to the deterministic approximation property, would make the paper publishable. Also note the citation error for Zhang and Yu, which should be fixed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The headline: the math is sound and the result is genuinely new — complementary log ratio is provably closer to risk ratio than odds ratio for any two outcome probabilities — but the paper's practical pitch outruns what the proof alone supports.\n\nWhat it does well: Corollary 1 is not stated in the cited literature. The proof is clean: convexity of the Aranda-Ordaz transform gives the inequality, and the monotonicity argument with h(x)=x e^x/(e^x−1) is solid. Presenting logit and cloglog as special cases of the Aranda-Ordaz family is a genuinely nice unifying move, and the paper is short and readable. The derivation is parameter-free, so there is no circularity.\n\nThe soft spots: first, the abstract promises 'simulation studies further reinforce our theoretical findings,' but no simulations appear in the supplied manuscript. If that is an artifact of the version, fine; as presented, the empirical support cannot be checked. Second, and more substantive, the theorem is about estimands, not estimators. For fixed p0 and p1, CLR is closer to RR than OR. But in a fitted GLM, both the estimated OR and the estimated CLR are random. The proof says nothing about their variances, standard errors, or mean squared error around the true RR. So the recommendation to prefer cloglog in routine regression practice is a step beyond the theorem. The estimand-level result is useful and may well carry over, but that needs simulation or at least an explicit caveat. There is a minor typo in the appendix: the second derivative should be (λ+1)(1−θ)^{−λ−2}, but convexity still holds, so the proof is unaffected.\n\nWho it's for: biostatisticians and epidemiologists who routinely report odds ratios for common outcomes and want a simple way to get closer to risk ratios. The theoretical result is a nice, citable fact. A serious referee should engage with it: the proof deserves checking and the practical claims need to be brought in line with the evidence. My recommendation: send it to peer review, and ask for either the simulations or a rewritten abstract and discussion that limit the claim to the estimand comparison.","headline":"Sound, genuinely new estimand-level result; the practical recommendation needs more support.","tokens_in":6925,"tokens_out":2169,"would_cite":true,"duration_ms":20418,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62J12"],"pacs":[],"model":"deepseek-v4-flash","headline":"For any two outcome probabilities, the complementary log ratio is always closer to the risk ratio than the odds ratio is.","keywords":["risk ratio approximation","complementary log-log model","odds ratio","Aranda-Ordaz family","binary regression","link function","complementary log ratio","maximum relative discrepancy"],"falsifier":"Fix any $0<p_0\\ne p_1<1$ and compute $RR=p_1/p_0$, $CLR=\\log(1-p_1)/\\log(1-p_0)$, $OR=\\frac{p_1/(1-p_1)}{p_0/(1-p_0)}$, and the two maxima $\\max(RR/CLR,CLR/RR)$ and $\\max(RR/OR,OR/RR)$; a single pair with the former not smaller than the latter would refute Corollary 1. Equivalently, check whether $\\frac{d}{d\\lambda}\\log WR(\\lambda)$ ever becomes negative for some $\\lambda\\in(0,1)$ and $p_0\\ne p_1$, which would overturn Theorem 1.","tokens_in":5914,"feed_emoji":"📈","tokens_out":12411,"duration_ms":100985,"temperature":0.7,"pith_summary":"The paper sets out to show that the effect measure of a complementary log-log model, which it names the complementary log ratio, always tracks the risk ratio more closely than the odds ratio does. The comparison is exact: for any two distinct outcome probabilities in the unit interval, the maximum relative discrepancy between the risk ratio and the complementary log ratio is strictly smaller than the same discrepancy for the odds ratio. The proof works inside the Aranda-Ordaz family of link transformations, a one-parameter family that contains the complementary log-log link at one end and the logit link at the other. Because the complementary log-log model is a standard GLM already available in R and SAS, the practical payoff is that practitioners can get closer risk-ratio approximations simply by changing the link function. The paper does not argue against direct risk-ratio estimation; it offers the complementary log-log model as an accessible alternative when direct methods are unstable.","feed_headline":"Complementary log-log beats logistic at risk ratios","feed_subtitle":"The complementary log ratio always lands closer to the risk ratio than the odds ratio, and you can switch in R or SAS.","key_machinery":"The central object is the Aranda-Ordaz family of transformations $W_\\lambda(\\theta)=\\{(1-\\theta)^{-\\lambda}-1\\}/\\lambda$ for $0<\\lambda\\le1$ and $W_0(\\theta)=-\\log(1-\\theta)$, which interpolates between the complementary log-log link ($\\lambda=0$) and the logit link ($\\lambda=1$). The proof works by writing the generalized ratio as $WR(\\lambda)=(e^{\\lambda b}-1)/(e^{\\lambda a}-1)$ with $a=-\\log(1-p_0)$ and $b=-\\log(1-p_1)$, then showing that the derivative of $\\log WR(\\lambda)$ is positive because the function $h(x)=x e^x/(e^x-1)$ is strictly increasing for $x>0$. Monotonicity of $WR(\\lambda)$ plus convexity of $W_\\lambda$ in $\\theta$ yields monotonicity of the discrepancy $B(\\lambda)$ and the direction of the over- or underestimation.","core_discovery":"The paper's central result is that, for every pair of distinct outcome probabilities $p_0,p_1\\in(0,1)$, the complementary log ratio $CLR=\\log(1-p_1)/\\log(1-p_0)$ is a better approximation to the risk ratio $RR=p_1/p_0$ than the odds ratio $OR$ is, in the sense of the maximum relative discrepancy $B=\\max(RR/\\text{measure},\\text{measure}/RR)$. This is Corollary 1, and it follows from the stronger Theorem 1: within the Aranda-Ordaz family $W_\\lambda(\\theta)$, the function $B(\\lambda)=\\max(RR/WR(\\lambda),WR(\\lambda)/RR)$ is strictly increasing in $\\lambda$ over $[0,1]$. Since $WR(1)=OR$ and $WR(0)=CLR$, the complementary log-log link minimizes approximation error inside the family while the logit link maximizes it. Lemma 1 supplies the directional part: $WR(\\lambda)$ always overestimates the risk ratio when $RR>1$ and underestimates it when $RR<1$, so the discrepancy never flips sign.","pith_inferences":["Inference: the proof is deterministic and says nothing about estimation uncertainty; in finite samples, the complementary log ratio estimator's advantage in bias could be outweighed by larger variance or worse confidence-interval coverage, which the paper does not address.","Inference: the monotonicity result suggests the Aranda-Ordaz parameter could be tuned, for example choosing a small positive $\\lambda$ to retain some logistic-like inferential behavior while keeping most of the risk-ratio bias reduction, although the paper only advocates the $\\lambda=0$ endpoint.","Inference: because $CLR=\\log(1-p_1)/\\log(1-p_0)$ is a ratio of log-survival probabilities, the result connects naturally to proportional-hazards thinking, hinting that complementary log-log regression may be especially appropriate in settings where a hazard-ratio-style interpretation is already intended.","Inference: a testable extension is to compare, in simulation, the relative bias and coverage of cloglog versus logistic links for estimating risk ratios under covariate adjustment; the paper's theory predicts smaller relative bias for cloglog, but the sampling behavior is an open question."],"forward_implications":["In any binary comparison with distinct outcome probabilities, no pair $(p_0,p_1)$ can make the odds ratio a better risk-ratio approximation than the complementary log ratio under the maximum relative discrepancy measure.","Inside the Aranda-Ordaz family, approximation error is ordered by the parameter: $\\lambda=0$ (complementary log-log) is best, $\\lambda=1$ (logit) is worst, and intermediate links fall between.","When $RR>1$, both the odds ratio and the complementary log ratio overestimate the risk ratio, and when $RR<1$ both underestimate it; the complementary log ratio's deviation is always the smaller one.","Because the complementary log-log model is a standard GLM available in R and SAS, switching from logistic regression to a complementary log-log model gives a closer risk-ratio approximation with essentially no extra implementation cost.","The paper's recommendation is scope-limited: direct risk-ratio estimation remains preferable when it is stable, and the complementary log-log model is proposed as a fallback, not a replacement."],"supporting_citations":[{"why":"Supplies the one-parameter family of transformations that unifies logit and complementary log-log links, the setting in which the paper proves monotonicity.","marker":"Aranda-Ordaz (1981)"},{"why":"Establishes that odds ratios overestimate risk ratios above one and underestimate them below one, the fact Lemma 1 generalizes and the motivation for the whole comparison.","marker":"Zhang and Kai (1998)"},{"why":"Provides the complementary log ratio expression and its rare-outcome approximation to the risk ratio, and frames the conversion problem the paper targets.","marker":"VanderWeele (2020)"},{"why":"Shows complementary log-log regression produces covariate-adjusted prevalence ratios, supporting the paper's practical claim that the model is a usable alternative.","marker":"Penman and Johnson (2009)"},{"why":"Documents the common 10% prevalence cutoff beyond which odds ratios no longer approximate risk ratios, the practical threshold that motivates the paper.","marker":"Hosmer Jr et al. (2013)"}],"fun_headline_variants":["CLR consistently closer to risk ratio than odds ratio","Complementary log-log link: best risk ratio approximation","For common outcomes, CLR trumps OR for risk ratio","Switch to complementary log-log for accurate risk ratios","Risk ratio? Complementary log-log beats logit"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The proof compares the complementary log ratio and the odds ratio as exact functions of the two true outcome probabilities, so it does not account for estimation error, convergence problems, or the standard errors of fitted models; if those sources of variability are large enough, the practical advantage of the complementary log-log model could shrink or disappear.","fun_headline_variants_meta":{"raw":{"variants":["CLR consistently closer to risk ratio than odds ratio","Complementary log-log link: best risk ratio approximation","For common outcomes, CLR trumps OR for risk ratio","Switch to complementary log-log for accurate risk ratios","Risk ratio? Complementary log-log beats logit"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000237,"raw_usage":{"total_tokens":1519,"prompt_tokens":971,"completion_tokens":548,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":587,"completion_tokens_details":{"reasoning_tokens":482}},"tokens_in":587,"tokens_out":548,"duration_ms":5354,"temperature":1.0,"reasoning_tokens":482,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T11:54:02.121320+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Fix any $0<p_0\\ne p_1<1$ and compute $RR=p_1/p_0$, $CLR=\\log(1-p_1)/\\log(1-p_0)$, $OR=\\frac{p_1/(1-p_1)}{p_0/(1-p_0)}$, and the two maxima $\\max(RR/CLR,CLR/RR)$ and $\\max(RR/OR,OR/RR)$; a single pair with the former not smaller than the latter would refute Corollary 1. Equivalently, check whether $\\frac{d}{d\\lambda}\\log WR(\\lambda)$ ever becomes negative for some $\\lambda\\in(0,1)$ and $p_0\\ne p_1$, which would overturn Theorem 1.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the one-parameter family of transformations that unifies logit and complementary log-log links, the setting in which the paper proves monotonicity."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes that odds ratios overestimate risk ratios above one and underestimate them below one, the fact Lemma 1 generalizes and the motivation for the whole comparison."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the complementary log ratio expression and its rare-outcome approximation to the risk ratio, and frames the conversion problem the paper targets."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Shows complementary log-log regression produces covariate-adjusted prevalence ratios, supporting the paper's practical claim that the model is a usable alternative."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Documents the common 10% prevalence cutoff beyond which odds ratios no longer approximate risk ratios, the practical threshold that motivates the paper."}],"review_version":1}