Pith. sign in

REVIEW 4 major objections 4 minor 26 references

This paper claims that in circular logistic regression, the choice of link function matters most when the circular predictor is broadly dispersed and the response is unbalanced; under high predictor concentration, symmetric links are more s

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

A simulation and two applications compare logit, probit, cloglog, Cauchit, and skew-logit links for circular logistic regression, but the headline claim about dispersed predictors is contradicted by the study's own table.

T0 review reviewed 2026-08-02 challenge →

load-bearing objection The model is a standard cos-sin GLM and the simulation contradicts the abstract's own headline; the response-balance condition in the abstract is never manipulated. the 4 major comments →

arxiv 2607.13264 v1 pith:LPRCLODS submitted 2026-07-14 stat.ME

Evaluation of Circular Logistic Regression Models with Asymmetric Link Functions

classification stat.ME MSC 62H1162J12
keywords circular logistic regressionasymmetric link functionsvon Mises distributiondirectional datageneralized linear modelsskew-logit linkAIC model comparisonquasi-complete separation
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper develops a circular logistic regression model that expresses the circular predictor through its cosine and sine, so any standard binary link function can be plugged in. It then asks when the choice of link actually matters. Using simulations with a von Mises predictor, it finds that when the predictor is broadly dispersed and the response is unbalanced, link choice has real impact; but when the predictor is highly concentrated, symmetric links (logit, probit) are stable and the asymmetric skew-logit link performs worse and fails more often due to quasi-complete separation. The practical payoff is guidance for applied researchers: use simple symmetric links unless the response is very lopsided and the predictor covers a wide angular range.

Core claim

On the paper's own terms, the central discovery is a concentration-dependent interaction between the circular predictor and the link function. With a dispersed circular predictor (κ=3), all five links perform comparably, and the skew-logit link actually achieves the lowest average deviance, reflecting its extra flexibility. With a concentrated predictor (κ=12), average AIC is nearly identical for logit, probit, cloglog, and Cauchit, while the skew-logit link is distinctly worse (27.67 versus about 23.5–23.8) and the overall failure rate from quasi-complete separation jumps to 10.6%. The paper interprets this as near-collinearity among cos(θ), sin(θ), and the intercept when the angles are tig

What carries the argument

The cosine–sine linear predictor η = β0 + β1 cos(θ) + β2 sin(θ) is the central device. It converts a circular covariate into two ordinary engineered covariates, so a standard generalized linear model can fit any monotone link function. The paper compares the symmetric logit and probit links against asymmetric cloglog, Cauchit, and a skew-logit power link (where the logistic cdf is raised to a power σ estimated with the coefficients). Von Mises–distributed predictors with two concentration levels (κ=3 and κ=12) provide the stress test; AIC and deviance are the comparison metrics.

Load-bearing premise

The load-bearing premise is that the average-AIC comparison is fair: replications that produced quasi-complete separation were excluded from each link's average, but the paper never reports how many failures each link contributed, so if the skew-logit link failed more often, its remaining replications are not the same subset used for the other links.

What would settle it

Count the number of estimation failures per link under κ=12. If the skew-logit link accounts for most of the 53 failures, then its average AIC of 27.67 is computed on an easier subset of replications, and the conclusion that skew-logit fits worse under high concentration would no longer follow; recomputing averages on the intersection of replications that converge for all links would settle it.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • If the claim holds, applied analysts should default to the logit or probit link for circular logistic regression unless the response is markedly unbalanced, because the symmetric links are more stable and parsimonious.
  • When a circular predictor is highly concentrated, researchers should expect a higher rate of quasi-complete separation and may need larger samples or wider observation windows before adding link flexibility.
  • The recovered amplitude A and phase angle θ0 from the fitted β1 and β2 give a directly interpretable summary of the directional effect; standard errors for these derived quantities (via delta method) would make the method more usable.
  • The same cosine–sine predictor unifies circular regression across the binomial and Poisson families, so the link-function comparison generalizes to a broader circular GLM framework.
  • Software that automates forming the cosine–sine covariates and fitting multiple links would lower the barrier to using these models.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Inference: The paper's conclusion that asymmetric links are 'prone to instability' under high concentration would be more credible if per-link failure counts were reported; inferring that the comparison may be biased if failures cluster in the skew-logit link.
  • Inference: An extension the paper leaves implicit: a likelihood-ratio test or cross-validated comparison on the same replications would avoid AIC's parameter penalty and directly test whether the skew-logit's extra parameter ever buys predictive accuracy.
  • Inference: A testable extension: generate unbalanced response configurations with dispersed predictors to directly measure when the asymmetric links outperform symmetric ones, since the current simulation uses a balanced-ish design and does not vary response imbalance.
  • Inference: Because the failure rate jumps with concentration, a diagnostic based on the circular spread (e.g., mean resultant length of the predictor) could warn practitioners before fitting a flexible link.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The paper proposes a circular logistic regression model in which a binary or binomial response is regressed on the cosine and sine of a circular predictor, and compares five link functions (logit, probit, cloglog, Cauchit, and a skew-logit power link). A Monte Carlo simulation with von Mises concentration κ=3 and κ=12 (n=100, B=500) compares average AIC and deviance across links, and two real-data examples (rainfall occurrence with wind direction; monthly earthquake counts) illustrate the model. The abstract concludes that link choice matters most when the circular predictor is broadly dispersed and the response is markedly unbalanced, and that symmetric links are preferred under high concentration.

Significance. The modeling device is straightforward and useful: expressing the circular predictor through cos(θ) and sin(θ) reduces the problem to a standard GLM with engineered covariates, so any binomial link can be used with existing software. This addresses a real gap in applied circular statistics. If the evaluative claims were well supported, the paper would give practitioners concrete guidance on link selection. However, the main abstract conclusion is contradicted by the paper's own simulation results, and a key condition (response balance) is never manipulated. No code or data are provided, so the simulation cannot be independently checked. The contribution is methodological/evaluative rather than theoretical.

major comments (4)
  1. [Abstract; Section 6, Table 2] The abstract's central claim — that link choice matters most when the circular predictor is broadly dispersed and the response is markedly unbalanced — is not supported by the reported design and is contradicted by Table 2. The simulation varies only κ (3 vs. 12); response balance is never varied. Table 2 shows average AIC ranges of 36.87–38.11 at κ=3 (spread 1.24) and 23.54–27.67 at κ=12 (spread 4.13). Thus the only material link-dependent difference occurs under the concentrated predictor, not the dispersed one. Section 6 itself states that 'under the dispersed predictor (κ=3), all five links produce broadly comparable average AIC values.' The 'markedly unbalanced' condition is absent from the entire simulation design.
  2. [Section 6 (Design); Section 9] The simulation excludes replications that produced quasi-complete separation 'under any link' from the corresponding average, but per-link failure counts are never reported. This is load-bearing for the Section 9 conclusion that 'the skew-logit link ... failed considerably more often due to quasi-complete separation.' If failures are concentrated in certain links, the average AIC values in Table 2 are computed over different subsets of replications, so direct comparison of those averages is biased. The authors should report the number of failures per link and per κ, and use a common set of replications across links or a method robust to separation (e.g., Firth's penalized likelihood).
  3. [Section 8 (Discussion); Section 6; Section 7.1] The practical guideline 'Match the link to the balance of the response' is presented as a consequence of this paper's simulation, but the simulation never varies response balance. The only empirical support offered is the Macomb rainfall data with a 44/59 split, which is explicitly described as 'moderately but not severely unbalanced.' The cited literature (Chen et al., Collett, Gómez-Déniz et al.) addresses non-circular regression and cannot fill the gap for the circular setting. To support the abstract's claim, the design must include conditions with markedly unbalanced responses (e.g., rare events) in addition to the two concentration regimes.
  4. [Section 6, Table 2] The simulation reports only average AIC over 500 replications, with no Monte Carlo standard errors, quantiles, or replication-level variation. Statements such as 'all five links produce broadly comparable average AIC values' and 'the skew-logit link performs distinctly worse' have no measure of uncertainty. Given that AIC is a random quantity, the authors should report standard errors or confidence intervals for the averages (or a paired comparison across links) so the reader can judge whether the differences are meaningful.
minor comments (4)
  1. [Section 5; Section 4; Section 3] There are several typographical errors: 'funcyion' in Section 5; 'acrophase, which the angle' in Section 4; 'under the unbalance response variable ratio' in Section 3; 'T erm' in Table 4. These should be corrected.
  2. [Section 6] The text says 'the corresponding deviances follow the same ordering, since all symmetric and the cloglog links use three estimated parameters; the skew-logit link uses four.' This is not automatic: the AIC penalty differs for the skew-logit link, and no deviance table is provided. Please report deviance values or clarify the claim.
  3. [Section 7.2, Table 5] The description of the earthquake count data is confusing: the table lists multiple values for some months and the observation window (Jan 2013–Jul 2014) has 19 months, yet the sample size is not stated. Please clarify the data structure and the effective number of observations.
  4. [Throughout] The figures (Figure 1, Figure 2, Figure 3) are referenced but not included in the manuscript text. Ensure the final version contains all figures.

Circularity Check

0 steps flagged

No significant circularity: the simulation is self-contained and the central comparison does not reduce to its inputs, though the abstract's scope claims are not supported by the design.

full rationale

The paper's central evaluation is a Monte Carlo simulation in which binary responses are generated from a stated logit-link model with a cosine-sine predictor, and the five competing link functions are each fitted to the same simulated data sets (Section 6). There is no fitted parameter relabeled as a prediction, and the AIC comparison is a direct empirical comparison, not a quantity defined in terms of the fitted values or of the authors' prior results. The cosine-sine linear predictor is standard (Jammalamadaka and SenGupta, 2001), and the skew-logit link is defined explicitly as the logistic cdf raised to a power, with the additional parameter estimated by maximum likelihood; this is an adopted model specification, not a result derived from itself. The paper does cite the author's own earlier work (Tasdan, 2026; Dawotola and Tasdan, 2025), but those citations are used for background, for a companion circular Poisson example, and for the bioassay comparison of links; they are not load-bearing for the simulation's conclusion. The main weaknesses are correctness/design issues, not circularity: the abstract's claim about 'markedly unbalanced' responses is untested because response balance is never manipulated, and the claim about 'broadly dispersed' predictors is contradicted by Table 2, where the largest link-dependent AIC spread occurs at κ=12 (concentrated), not κ=3 (dispersed). Likewise, the exclusion of replications with quasi-complete separation is a potential bias in the averaged AIC, but it is not a definitional or fitting-loop circularity. Because the central derivation is self-contained against an explicit data-generating process, the circularity score is low; the self-citations are minor and non-essential.

Axiom & Free-Parameter Ledger

4 free parameters · 3 axioms · 0 invented entities

No new particles, forces, or entities are introduced. The skew-logit power link is prior work (Dawotola and Tasdan, 2025; Smithson and Shou, 2017). The main unstated debts are the model-space assumption and the exclusion-rule assumption.

free parameters (4)
  • Simulation true coefficients (β0, β1, β2) = (1, 2, 3)
    Chosen by hand in Section 6; the amplitude A≈3.61 on the logit scale sets how unbalanced and how separable the simulated responses are, directly shaping the AIC gap and separation rates.
  • von Mises concentration κ regimes = κ = 3 and κ = 12
    Chosen to represent 'dispersed' and 'concentrated' predictors; the entire conclusion about link choice depending on concentration rests on these two values.
  • Sample size n and replications B = n=100, B=500
    Chosen by hand; the quasi-complete-separation rate (10.6% at κ=12) is a small-sample artifact, and a different n would change the headline instability result.
  • Skew-logit shape parameter σ = estimated by ML, value not reported in Table 4
    The skew-logit link's extra shape parameter is estimated jointly with β; its identifiability under κ=12 drives the headline 'instability' result, yet no estimates or standard errors are reported for σ.
axioms (3)
  • domain assumption The cosine–sine linear predictor, E[Y|θ]=β0+A cos(θ−θ0), is an adequate and complete representation of a circular covariate's effect on a binary response.
    Inherited from linear-circular regression (Jammalamadaka and SenGupta, 2001) and applied without justification to the probit/cloglog/Cauchit/skew-logit scales in Section 5.
  • domain assumption Quasi-complete separation failures can be excluded without biasing the comparison of average AIC across links.
    Section 6 excludes failed replications; the conclusion in Section 9 that skew-logit fails 'considerably more often' implicitly assumes the exclusions are link-independent, but no per-link counts are provided.
  • standard math The von Mises distribution is the appropriate null/generative model for the circular predictor in the simulation.
    Section 2 defines von Mises as the circular analogue of the normal and Section 6 uses it to generate θ; this is standard directional-statistics practice.

reviewed 2026-08-02 · how reviews work

0 comments
Cite this review

Pith. "Pith review of Evaluation of Circular Logistic Regression Models with Asymmetric Link Functions." pith.science (2026). https://pith.science/paper/LPRCLODS

@misc{pith2026260713264,
  author       = {Pith},
  title        = {Pith review of: Evaluation of Circular Logistic Regression Models with Asymmetric Link Functions},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/LPRCLODS}},
  note         = {Machine review of arXiv:2607.13264}
}
Share X Bluesky LinkedIn Reddit HN
abstract

Circular (directional) data arise whenever observations are measured as angles on the unit circle, such as wind direction, time of day, or calendar phase, and require statistical methods that respect the periodicity of the domain $[0; 2\pi)$. While circular-linear and linear-circular regression models are well established, regression models for a binary or binomial response observed jointly with a circular predictor remain largely undeveloped, with the sole closely related study restricted to the symmetric logit link. This paper develops and evaluates a circular logistic regression framework in which the linear predictor is expressed through the cosine and sine of the circular covariate, and compares the performance of symmetric link functions (logit, probit) against asymmetric alternatives (complementary log-log, Cauchit, and a skew-logit power link) under a generalized linear model formulation. A Monte Carlo simulation generates circular predictors from the von Mises distribution under two concentration regimes and evaluates model fit using the Akaike Information Criterion (AIC) and deviance. The methodology is illustrated with two real data sets: daily rainfall occurrence and wind direction recorded in Macomb, Illinois, and monthly earthquake counts in Western Anatolia, Turkiye, the latter used to connect the binary circular model to the related circular Poisson regression framework for count outcomes. Results indicate that the choice of link function matters most when the circular predictor is broadly dispersed and the response is markedly unbalanced; under high concentration of the predictor, symmetric links are preferred and asymmetric links are prone to instability. Practical guidelines and directions for future software development are discussed.

Figures

Figures reproduced from arXiv: 2607.13264 by Feridun Tasdan.

Figure 1
Figure 1. Figure 1: Average AIC by link function under dispersed ( [PITH_FULL_IMAGE:figures/full_fig_p006_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Fitted log-odds of rainfall as a function of wind direction, from the circular skew-logit [PITH_FULL_IMAGE:figures/full_fig_p008_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Circular plot of monthly earthquake counts by calendar month. [PITH_FULL_IMAGE:figures/full_fig_p009_3.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

26 extracted references

  1. [1]

    and Lund, U

    Agostinelli, C. and Lund, U. (2017). R package ``circular'': Circular statistics. CRAN

  2. [2]

    Agresti, A. (2002). Categorical Data Analysis . Wiley, 2nd edition

  3. [3]

    B., and Holm, A

    Breen, R., Karlson, K. B., and Holm, A. (2018). Interpreting and understanding logits, probits, and other nonlinear probability models. Annual Review of Sociology , 44(1):39--54

  4. [4]

    K., and Shao, Q.-M

    Chen, M.-H., Dey, D. K., and Shao, Q.-M. (1999). A new skewed link model for dichotomous quantal response data. Journal of the American Statistical Association , 94(448):1172--1186

  5. [5]

    Collett, D. (2003). Modelling Binary Data . Chapman & Hall/CRC, 2nd edition

  6. [6]

    and Santner, T

    Czado, C. and Santner, T. J. (1992). The effect of link misspecification on binary regression inference. Journal of Statistical Planning and Inference , 33:213--231

  7. [7]

    Dawotola, T. B. and Tasdan, F. (2025). A comparative analysis of some link functions for binomial regression models with applications to bioassay data. International Journal of Scientific Research and Modern Technology , 3(12):106--116

  8. [8]

    Fisher, N. I. (1993). Statistical Analysis of Circular Data . Cambridge University Press

  9. [9]

    Fisher, N. I. and Lee, A. J. (1992). Regression models for an angular response. Biometrics , 48(3):665--677

  10. [10]

    G\'omez-D\'eniz, E., Calder\'in-Ojeda, E., and G\'omez, H. W. (2022). Asymmetric versus symmetric binary regression: A new proposal with applications. Symmetry , 14(4):733

  11. [11]

    Hornik, K. (2014). R package ``circular'': Circular statistics (version 0.4-93)

  12. [12]

    Jammalamadaka, S. R. and SenGupta, A. (2001). Topics in Circular Statistics . World Scientific Publishing

  13. [13]

    and Khan, S

    Kadhem, A.-D. and Khan, S. (2017). Logistic regression for circular data. AIP Conference Proceedings , 1842(1):030022

  14. [14]

    D., and Malkemper, E

    Landler, L., Ruxton, G. D., and Malkemper, E. P. (2021). Advice on comparing two independent samples of circular data in biology. Scientific Reports , 11:20337

  15. [15]

    Li, J. (2014). Choosing the Proper Link Function for Binary Data . PhD thesis, The University of Texas at Austin

  16. [16]

    Mardia, K. V. and Jupp, P. E. (2009). Directional Statistics . Wiley

  17. [17]

    and Nelder, J

    McCullagh, P. and Nelder, J. A. (1989). Generalized Linear Models . Chapman & Hall/CRC, 2nd edition

  18. [18]

    Pewsey, A., Neuhauser, M., and Ruxton, G. D. (2013). Circular Statistics in R . Oxford University Press

  19. [19]

    Rivest, L.-P., Duchesne, T., Nicosia, A., and Fortin, D. (2015). A general angular regression model for the analysis of data on animal movement in ecology. Journal of the Royal Statistical Society: Series C , 64(3):445--463

  20. [20]

    Shim, H., Bonifay, W., and Wiedermann, W. (2023). Parsimonious asymmetric item response theory modeling with the complementary log-log link. Behavior Research Methods , 55(1):200--219

  21. [21]

    and Shou, Y

    Smithson, M. and Shou, Y. (2017). Cdf-quantile distributions for modelling random variables on the unit interval. British Journal of Mathematical and Statistical Psychology , 70(3):412--438

  22. [22]

    Tasdan, F. (2026). Modelling of count data in circular statistics. In Duman, O. and Erkus-Duman, E., editors, Approximation Theory and Special Functions , volume 503 of Springer Proceedings in Mathematics & Statistics , pages 593--602. Springer Nature

  23. [23]

    and Cetin, M

    Tasdan, F. and Cetin, M. (2013). A simulation study on the influence of ties on uniform scores test for circular data. Journal of Applied Statistics , 41(5):1137--1146

  24. [24]

    and Yeniay, O

    Tasdan, F. and Yeniay, O. (2014). Power study of circular anova test against nonparametric alternatives. Hacettepe Journal of Mathematics and Statistics , 43(1):97--115

  25. [25]

    and Yeniay, O

    Tasdan, F. and Yeniay, O. (2018). A comparative simulation of multiple testing procedures in circular data problems. Journal of Applied Statistics , 45(2):255--269

  26. [26]

    and Al Luhayb, A

    Zaidi, A. and Al Luhayb, A. S. M. (2023). Two statistical approaches to justify the use of the logistic function in binary logistic regression. Mathematical Problems in Engineering , 2023:5525675

This paper was first reviewed by deepseek-v4-flash on August 2, 2026.