{"id":"9ff61e9b-9745-4109-9d40-0bc65e9b0888","arxiv_id":"1908.06512","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A mixture cure Cox model, which separates recipients into an engaged group and a never-open group, predicts email time-to-open better than standard classifiers and regressors on a large marketing dataset.","lead":"This paper applies survival analysis, a method from medical statistics, to predict how quickly a recipient will open a marketing email. A mixture model that splits recipients into likely and unlikely openers gives the most accurate waiting-time estimates when most emails are never opened.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"MM's reported MRAD(O) advantage over linear regression at the 3-hour window is within the sampling noise documented in the paper itself, so the 'best accuracy' claim is not established.","rationale":"The paper's contribution is the empirical claim that the mixture model is the most accurate model for time-to-open. I looked for the weakest condition on which that claim depends. The most load-bearing issue is not the proportional-hazards assumption or the feature set; it is the statistical reliability of the single-number comparison that carries the headline. The paper's own Section 4.4 says confidence intervals are needed to order models by MRAD(O) and that MRAD(O) is sensitive to input variations. On Validation, LR beats MM at C=3h; on Test, after retraining, MM beats LR by 0.372, a value smaller than the bootstrap SD of MM's MRAD(O) at that setting. Table 4 shows the MM result is knife-edge in the percentile p, going from best at p=5 to the always-censor baseline at p=10. The cure assumption is the plausible reason for this fragility, since censored cases assigned to the non-prone state make the prediction essentially a two-point rule, but the decisive test is whether the Table 6 gap survives a bootstrap CI. The reader's conditional verdict already calls for confidence intervals; this stress-test explains why that requirement is load-bearing rather than cosmetic. I do not see an internal mathematical inconsistency in the MM construction, and I am not claiming any result is manufactured; the issue is that the central quantitative claim is not yet supported by the evidence presented.","tokens_in":15020,"tokens_out":16625,"duration_ms":189752,"concrete_test":"Bootstrap the combined Train+Validation recipients (e.g., 200 resamples), refit all models on each resample, and evaluate on the held-out Test set at C=3h and p=5; report the 95% confidence interval for the difference in MRAD(O) between MM and linear regression. If the interval includes zero or favors LR, the central claim is unsupported. This directly tests whether the 0.372-point gap in Table 6 is real or within the sampling variability the paper itself documents.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim rests on Table 6: at C=3 hours, MM has MRAD(O)=7.381 versus 7.753 for linear regression, a difference of 0.372. No uncertainty is reported for this out-of-time comparison. Section 4.4's bootstrap experiment on Validation gives MM a mean MRAD(O) of 9.277 with standard deviation 1.854 at the same C=3h and p=5, while LR has mean 8.215 and SD 0.036; on Validation, LR is actually better. Only after retraining on Train+Validation and evaluating on a single Test split does MM appear better, by an amount smaller than the bootstrap SD of MM. The paper itself states that such intervals are 'required if we were to reliably order the different models.' Table 4 compounds the fragility: at C=3h, MM's MRAD(O) jumps from 9.499 at p=5 to 26.641 at p=10, exactly the always-censor baseline value, meaning the claimed advantage exists only at one percentile. The likely mechanism is the cure assumption: censored cases are absorbed into the non-prone state, turning MM's time prediction into a two-point rule (C or the p=5 quantile). A single split at one percentile therefore does not support the headline claim that MM achieves the best accuracy.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper addresses prediction of whether and when a recipient opens a marketing email. It casts the problem in a survival-analysis framework with censoring windows of 3, 6, and 12 hours and compares (i) simple baselines, (ii) logistic/linear regression, (iii) Cox proportional hazards with a linear relative hazard, (iv) Cox with a gradient-boosted hazard, and (v) a mixture cure-style Cox model in which a logistic regression assigns a latent probability of being 'prone to open' and a Cox model governs open times for the prone group. Using a large proprietary email dataset split into Training, Validation, and Test sets, the paper reports AUC and mean relative absolute deviation, checks proportional-hazards assumptions with Schoenfeld residuals, examines the effect of the survival percentile p used for time prediction, runs a bootstrap stability experiment, and evaluates out-of-time on Test. The headline claim is that the mixture model achieves the best time-to-open accuracy, especially at the 3-hour censoring window.","tokens_in":15281,"tokens_out":5684,"duration_ms":55777,"significance":"If the empirical claim were established at the reported precision, the paper would make a useful applied contribution: it demonstrates a principled survival-analysis treatment of 'never-event' recipients in email marketing, where the open rate is low, and it checks model assumptions carefully. The evaluation is structured, with separate train/validation/test sets, bootstrap stability results, and an out-of-time holdout, which is a strength. The assumption checks (Schoenfeld residuals and group survival curves) are appropriate for the Cox framework. However, the central 'best accuracy' claim is not supported at the required precision: the reported Test advantage is smaller than the model's own bootstrap variability, and the advantage disappears at the next percentile. Because these concerns are about evidence rather than about the model's mathematical viability, they can be addressed by additional analysis.","major_comments":[{"comment":"The central claim that MM is the most accurate model for time-to-open rests on a single Test split. At C=3 hours, Table 6 reports MRAD(O)=7.381 for MM versus 7.753 for LR, a difference of 0.372. Table 5 shows that on the Validation set the bootstrap standard deviation of MM's MRAD(O) is 1.854 and its mean is 9.277, whereas LR has mean 8.215 with SD 0.036; on Validation, LR is actually better. The paper itself states in §4.4 that confidence intervals would be 'required if we were to reliably order the different models.' No such interval is supplied for the Test comparison, so the ordering MM<LR at C=3h is not established.","section":"§4.5, Table 6; §4.4, Table 5"},{"comment":"The reported time-prediction advantage of MM is confined to a single percentile choice. For C=3 hours, MM's MRAD(O) is 9.499 at p=5 but jumps to 26.641 at p=10, exactly the value of the constant censoring-window baseline B in Table 3. The same pattern holds at C=6 and 12 hours, where p>=25 (or p>=10) reproduces the baseline. Since §4.5 selects p as the value with lowest validation MRAD, the Test result depends on a one-point selection in the percentile grid. The paper should report Test MRAD(O) for all p values (at least p=5 and p=10) and account for the selection, or otherwise show that the comparison is not an artifact of this choice.","section":"§4.3, Table 4"},{"comment":"The mixture model's latent state is defined as 'prone to the event' but is estimated from data censored at 3, 6, or 12 hours, while opens are monitored for 10 days. With C=3 hours, roughly 57% of eventual opens are censored (§3.1). Under Eq. (9), censored individuals can be assigned to the non-prone state with survival probability 1, so L_i conflates 'will never open' with 'did not open within the window.' This makes the engagement interpretation and the predicted times depend on the chosen censoring horizon. The paper should assess sensitivity of the MM advantage to this conflation, for example by evaluating a model that allows late opens or by reporting the proportion of censored observations assigned to L=0.","section":"§2.2, Eq. (9); §3.1"}],"minor_comments":[{"comment":"The likelihood expression is not written in a standard form: the hazard term h(t_i | L_i, X_i) appears inside an exponent indexed by the latent state L_i, even though L_i is unobserved. Please clarify the exact EM objective used for estimation.","section":"§2.2, Eq. (10)"},{"comment":"Using the same symbol LR for both logistic and linear regression, with only a footnote to disambiguate, makes Tables 3, 5, and 6 hard to read; consider separate names or explicit column labels.","section":"§3.1.1"},{"comment":"The phrase '10 bootrapped samples' should read '10 bootstrapped samples'; with only 10 bootstrap replicates, the standard-deviation estimates themselves have high variability, so the reported SDs should be interpreted cautiously.","section":"§4.4"},{"comment":"Only CPH-L Schoenfeld residuals are shown, although the text states that the procedure can be applied to GBM and MM; either show the analogous plots or state explicitly that the displayed checks are for the linear Cox model.","section":"§3.2, Figure 3"},{"comment":"Table 4 does not include the LR baseline, which makes it harder to compare the p-sensitivity of the winning model with the linear-regression baseline; adding LR would strengthen the analysis.","section":"§4.3, Table 4"}],"recommendation":"major_revision","confidential_remarks":"The paper has real strengths, especially the structured evaluation and assumption checks, and the issues identified are about evidence quality rather than fundamental errors. I would ask for paired bootstrap or repeated split evaluation on Test, plus a report of the p-sensitivity of the Test results, before endorsing the 'best accuracy' statement. The current single-split, single-percentile comparison is too fragile for the headline claim."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: the paper applies a known mixture cure Cox model to a large marketing email dataset, and the evaluation is careful enough to be worth reading, but the headline claim that the mixture model is “best” does not survive close reading. On the held-out Test split at the 3-hour censoring window, the reported MRAD(O) advantage over linear regression (7.381 vs 7.753) is smaller than the bootstrap standard deviation the paper itself reports for MM on Validation (SD 1.854). On Validation, LR actually wins (8.215 vs 9.277). The paper openly says such intervals would be required to reliably order models; that is the right caveat, but it undercuts the abstract's claim.\n\nWhat is genuinely useful: the paper is a competent demonstration that a cure mixture can be fitted to engagement data with heavy censoring, and it does a solid job checking assumptions—Schoenfeld residuals for proportionality, a log-rank test for the two latent groups, and a bootstrap stability analysis. The train/validation/test protocol is clean, and the discussion of why median survival time is undefined under heavy censoring is clear. The observation that MM's MRAD(O) collapses to the constant-censor baseline once the percentile moves from 5 to 10 (Table 4) is reported honestly, but it reveals that the advantage is fragile, not robust. In fact, that collapse suggests the model's time predictions are essentially a two-point rule: censoring window for most users, p=5 quantile for a minority.\n\nThe main soft spots, in proportion: (1) The central claim rests on a single test split at one percentile, with no confidence intervals; the authors themselves flag the need for intervals. (2) AUC for MM is not better than the alternatives, so the case for joint modeling rests entirely on MRAD(O), where the edge is small and unstable. (3) The cure assumption treats censored cases as never-openers; with only 10 days of observation and 3-hour windows, late openers get absorbed into the non-prone group, which could bias the comparison. (4) Proprietary data and a limited feature description hamper reproducibility, though that is typical for industry papers.\n\nWho this is for: applied researchers in email marketing or engagement prediction who want to see a careful application of cure models in a high-censoring setting. It deserves a serious referee, but the authors should be pushed to qualify the “best accuracy” claim and provide interval estimates on the test comparison.","headline":"Careful application of a known cure-mixture Cox model to email open-time data, but the headline accuracy claim is fragile: the MM advantage over linear regression on the test split is within the sampling noise the paper itself documents.","tokens_in":15820,"tokens_out":1839,"would_cite":false,"duration_ms":18551,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that a mixture survival model predicts email open times more accurately than alternatives in a domain where most emails are never opened.","keywords":["email interaction data","survival analysis","time-to-event prediction","enterprise email marketing","Cox proportional hazards model","mixture model","censoring","latent engagement state"],"falsifier":"Re-run the comparison after extending the observation window to 30 days and re-estimating each model with a 12-hour censoring window; if the mixture model's MRAD(O) no longer beats linear regression and CoxPH, then the never-open assumption, not genuine timing accuracy, was producing the reported gain.","tokens_in":14803,"feed_emoji":"📧","tokens_out":15993,"duration_ms":140176,"temperature":0.7,"pith_summary":"This paper addresses a practical marketing question: for a mass email with time-sensitive content, when will a given recipient actually open it? Survival analysis models the time until an event and handles cases where the event has not been observed, but standard versions assume everyone eventually opens; because a large majority of marketing emails are never opened, the paper argues for a mixture model that assigns each recipient a latent state — prone to open or not. It combines a Cox proportional-hazards model for open times among the prone with a logistic regression for the probability of being prone, so one survivor curve yields both the open probability and the predicted open time. On a large real-world marketing email dataset, using observation cutoffs of 3, 6, and 12 hours after which unopened emails are treated as censored, the mixture model produced the lowest mean relative absolute deviation in predicted open times on the held-out Test set, for example 7.381 versus 7.753 for linear regression and 10.009 for standard CoxPH at the 3-hour cutoff. If the claim holds, senders of time-sensitive campaigns could size their reach and set offer windows more accurately than with separate classification and regression pipelines.","feed_headline":"Separating never-openers from openers gives best email open-time forecasts","feed_subtitle":"On marketing email, the latent-state mixture beats linear regression and standard time-to-event models.","key_machinery":"The machine that carries the argument is the mixture survivor function $$S_i(t|X_i)=\\pi(Z_i)\\,S(t|L_i=1,X_i)+(1-\\pi(Z_i)),$$ where $\\pi(Z_i)$ is a logistic-regression estimate of the probability that recipient $i$ belongs to the latent prone-to-open state, and $S(t|L_i=1,X_i)=S_0(t)^{\\exp(\\beta^T X_i)}$ is the Cox proportional-hazards survivor curve for that state. The $(1-\\pi)$ term anchors non-prone recipients at survival probability 1, so each individual's predicted time-to-open is read from a low percentile of this survivor curve, the paper uses the 5th percentile, rather than the median, which is undefined when the curve flattens above 0.5. Parameters are estimated by an expectation-maximization (EM) procedure because the state indicator $L_i$ is latent. This machinery separates will-eventually-open from never-opens and fits the timing model only on the former group.","core_discovery":"On email data where roughly 8% of recipients open a message within ten days, the paper claims that a cure-mixture version of the Cox model is the right way to predict time-to-open. Each recipient has a latent indicator: those who will not open the email have survival probability 1 at every time, while those who may open follow a proportional-hazards model with a baseline hazard estimated nonparametrically, and the mixing probability is predicted by logistic regression from engagement features. On the held-out Test set, the mixture model attains a mean relative absolute deviation over opened emails of 7.381, 9.340, and 11.653 for censoring windows of 3, 6, and 12 hours, lower than linear regression (7.753, 10.194, 13.409), CoxPH with a linear relative hazard (10.009, 23.227, 29.131), and gradient-boosted Cox (13.787, 18.339, 21.080), while matching those models in area under the ROC curve. The central discovery is that explicitly modeling a subpopulation that never opens is what makes time predictions accurate when censored rows dominate.","pith_inferences":["Editorial extension: the same mixture-Cox construction should transfer to other domains with large never-event subpopulations, such as customer churn, app re-engagement, or support-ticket resolution, where censoring rates are high and a latent long-term non-event state is plausible.","Editorial inference: because anyone who opens after the censoring window is assigned to the non-prone state, the latent states themselves depend on the chosen horizon; one could re-fit at a 12-hour window and evaluate at 3 hours to see whether the predicted early-window distribution changes.","Beyond the paper: the reported advantage is aggregate; a follow-up could examine whether the mixture model's gain concentrates in recipients with high historical engagement or in cold recipients, which would identify the segment driving the improvement."],"forward_implications":["For marketing campaigns with open rates near 8–10%, the paper's results imply that a mixture cure model should be used instead of plain Cox regression or linear regression for open-time prediction, because its held-out mean relative absolute deviation over opened emails (MRAD(O)) is lower for every censoring window tested.","A single survival model can provide both the probability that an email is opened and the predicted time of opening, so practitioners need not train a separate classifier and regressor, provided they accept the latent never-open state.","The censoring window is a modeling choice that changes both the training data and the ranking of models: at 3 hours over half of eventual opens look censored and the mixture model's edge over linear regression is modest, while at 6 and 12 hours the gap widens.","Because the median survival time is undefined in this regime, time-to-open should be predicted from a low percentile, the 5th, of the individual survivor curve; higher percentiles saturate at the censoring window and match the baseline's error.","Model selection should use a metric aligned with the task: tuning by AUC does not necessarily minimize MRAD, so deployments that care most about timing accuracy should validate on MRAD."],"supporting_citations":[{"why":"Introduces the Cox proportional-hazards partial likelihood that the paper extends into a mixture model.","marker":"[10]"},{"why":"Provides the mixture and cure-model formulation that supplies the latent prone-to-open state and its EM likelihood.","marker":"[4, 18]"},{"why":"Gives the nonparametric baseline-hazard estimator used to build individual survivor curves and percentile predictions.","marker":"[29]"},{"why":"Supplies the residual test used to check the proportional-hazards assumption before comparing models.","marker":"[19]"},{"why":"Nonparametric survivor-curve estimator used to visualize the data and compare the two engagement groups.","marker":"[22]"},{"why":"Elastic net regularizes the logistic and linear regression baselines and the CoxPH relative hazard.","marker":"[41]"},{"why":"Closest prior work on time-to-event modeling in email, providing the comparison point for this application.","marker":"[12]"},{"why":"Linear regression baselines that exclude censored rows; they are the timing baselines the mixture model must beat.","marker":"[5, 13]"},{"why":"Gradient-boosted Cox baseline (CPH-G) that the experiments compare against the mixture model.","marker":"[33]"}],"fun_headline_variants":["Cure model for never-openers boosts email open-time forecasts","Mixture model separates never-openers to predict email open time","For unopened-dominant emails, latent-state mixture wins on open time","Cure Cox model beats regression when many emails never get opened","Latent never-open state improves email open-time predictions"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that a recipient placed in the non-prone state has survival probability exactly 1 forever, while the data only observe ten days, so anyone who opens after the censoring window is treated as a never-opener; if those late openers behave differently, the model's advantage could be an artifact.","fun_headline_variants_meta":{"raw":{"variants":["Cure model for never-openers boosts email open-time forecasts","Mixture model separates never-openers to predict email open time","For unopened-dominant emails, latent-state mixture wins on open time","Cure Cox model beats regression when many emails never get opened","Latent never-open state improves email open-time predictions"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000442,"raw_usage":{"total_tokens":2272,"prompt_tokens":1012,"completion_tokens":1260,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":628,"completion_tokens_details":{"reasoning_tokens":1173}},"tokens_in":628,"tokens_out":1260,"duration_ms":9984,"temperature":1.0,"reasoning_tokens":1173,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T12:43:13.590823+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the comparison after extending the observation window to 30 days and re-estimating each model with a 12-hour censoring window; if the mixture model's MRAD(O) no longer beats linear regression and CoxPH, then the never-open assumption, not genuine timing accuracy, was producing the reported gain.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces the Cox proportional-hazards partial likelihood that the paper extends into a mixture model."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Gives the nonparametric baseline-hazard estimator used to build individual survivor curves and percentile predictions."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the residual test used to check the proportional-hazards assumption before comparing models."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Nonparametric survivor-curve estimator used to visualize the data and compare the two engagement groups."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Elastic net regularizes the logistic and linear regression baselines and the CoxPH relative hazard."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Closest prior work on time-to-event modeling in email, providing the comparison point for this application."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Gradient-boosted Cox baseline (CPH-G) that the experiments compare against the mixture model."}],"review_version":1}