{"id":"ff7a966c-6050-43a1-87b6-51e4317dd469","arxiv_id":"2412.00823","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A hierarchical Bayesian model of reported campus sexual assault counts estimates that the rise in reports from 2014 to 2018 reflects rising reporting rates, not rising incidence.","lead":"This paper builds a Bayesian model that separates how many sexual assaults actually happen on campus from how many get reported, using national crime survey data as the tiebreaker. Applied to U.S. college data from 2014 to 2019, it finds the reported increase was mostly a rise in reporting, not a rise in assaults.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The Section 5.2 reporting-rate trend is largely determined by the prior variance ratio between per-year noise in lambda and p, which Appendix E does not test.","rationale":"The reader's weakest_assumption focuses on NCVS prior levels, which is a real limitation for absolute incidence and reporting rate estimates. But the paper's headline claim is about the temporal trend, and the reader judged that trend robust across the Appendix E scenarios. My review identifies a separate, more direct threat to that trend: the model has no year-level fixed effects, so the observed aggregate increase must be absorbed by the per-year noise terms eta and delta. Since the likelihood identifies only lambda*p, the posterior cannot learn whether lambda or p moved; it can only reflect the prior. The prior makes p much more variable year-to-year (SD 0.5 on logit scale) than lambda (SD 0.1 on log scale), so the model 'prefers' to explain the 2014-2018 increase in reported counts as a rise in reporting probability. This is not a mathematical error; it is a substantive prior assumption. The authors may have intended this, but they do not defend the ratio, and Appendix E's sensitivity analysis shifts the prior means (e.g., NCVS levels) without varying the variance ratio. Therefore the single most load-bearing concern is that the central temporal claim is an implication of an untested prior variance choice. A swap/equal-variance refit is a concrete, feasible check that would settle it. If the trend survives, the paper's conclusion stands; if not, the conclusion needs major qualification. This supports maintaining the CONDITIONAL verdict, with the condition being this additional sensitivity analysis. I agree with the reader that the NCVS mapping is a separate limitation, hence partial agreement, but the variance ratio is more central to the strongest claim.","tokens_in":20635,"tokens_out":7273,"duration_ms":71031,"concrete_test":"Refit the model with two alternative per-year noise configurations, keeping all other priors, likelihood, and data exactly as in Section 4: (A) swap the scales: eta ~ N(0, 0.5), delta ~ N(0, 0.1); (B) equalize them: eta ~ N(0, 0.25), delta ~ N(0, 0.25). For each configuration, record the posterior median systemwide reporting rate per year (the quantity in Figure 11) and the posterior probability that the 2018 reporting rate exceeds the 2014 rate. If the increasing trend and high posterior probability survive both alternative configurations, the Section 5.2 claim is robust. If the trend flattens, reverses, or becomes highly uncertain, the claim is an artifact of the chosen eta/delta variance ratio and should be substantially qualified or withdrawn.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"Because the likelihood (Eqs. 1-2) identifies only the product lambda_ij * p_ij (Section 4.5), the posterior split of year-to-year changes in observed reports between incidence and reporting probability is governed by the priors on the per-year noise terms. Equation (3) puts year noise eta_ij ~ N(0, 0.1) on log(lambda_ij), while Equation (5) puts year noise delta_ij ~ N(0, 0.5) on logit(p_ij). On the log scale relevant to the observed count, this makes p roughly five times more mobile year-to-year than lambda. The aggregate rise in reported assaults shown in Figure 2 can be represented either by rising incidence or by rising reporting probability; the data alone cannot distinguish these, and the much smaller prior variance on eta strongly steers the posterior toward attributing the rise to p. Thus the central claim in Section 5.2--that the increase in reported assaults is more likely due to rising reporting rates--is driven by this prior variance ratio, not by an identified feature of the data. Appendix E varies the prior means (NCVS levels, scenarios a-e) but never the relative scales of eta and delta, so the reported sensitivity analysis does not probe this load-bearing assumption. The model could be right, but the authors supply no empirical justification for choosing SD 0.1 for eta and 0.5 for delta, and this choice is decisive for the headline temporal conclusion.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript proposes a hierarchical Bayesian model for Clery Act campus sexual assault counts, treating observed reports x_ij as binomial thinnings of latent true counts z_ij with reporting probability p_ij and Poisson incidence rate lambda_ij. After marginalizing out z_ij, the observed counts follow Poisson(lambda_ij p_ij), so the data identify only the product of incidence and reporting probability; informative priors based on NCVS statistics are used to separate the two components. The model is fitted with HMC/Stan to 1,973 institutions over 2014-2019, with held-out predictive checks, a comparison of pooling schemes, and posterior estimates of incidence and reporting rates. The headline result is that true incidence was roughly flat while reporting rates rose from about 17% in 2014 to about 24% in 2018, making rising reporting rates the more likely explanation for the observed increase in reported assaults.","tokens_in":20985,"tokens_out":6308,"duration_ms":76385,"significance":"The paper is a serious and mostly careful contribution to the underreported-count-data literature: the marginalization derivation in Appendix B is correct, the split predictive checks are methodologically appropriate, and Appendices C-E contain unusually thorough model comparison and sensitivity analyses. The use of external NCVS statistics to set informative priors is legitimate and does not involve circular reasoning. However, the central substantive claim about the temporal trend in reporting rates is not empirically identified; it depends on the relative prior variances of the year-level noise terms, and the reported sensitivity analysis does not vary that ratio. Consequently, the paper's headline conclusion is a conditional modeling outcome rather than an established empirical finding. The framework and software are valuable, but the claims need reframing or additional robustness analysis.","major_comments":[{"comment":"The central claim that the 2014-2019 increase in reported assaults is attributable to rising reporting rates is not identified by the data. Because the likelihood is Poisson(lambda_ij p_ij), and Section 4.5 correctly explains that posterior estimates of lambda and p approach their prior conditional on the identified product, the decomposition of year-to-year changes in observed reports between incidence and reporting probability is governed by the priors on the per-year noise terms. Equation (3) sets eta_ij ~ N(0, 0.1) on log(lambda_ij), while Equation (5) sets delta_ij ~ N(0, 0.5) on logit(p_ij); on the log scale relevant to the observed count, this makes p roughly five times more mobile year-to-year than lambda. The upward trend in Figure 2 can be represented either by rising incidence or by rising reporting probability, and the much smaller prior variance on eta strongly steers the posterior toward the latter. Appendix E varies prior means (scenarios a-e) but never the relative variances of eta and delta, so the sensitivity analysis does not probe the assumption that drives the headline result. The authors should provide external empirical justification for these variance scales or report the posterior trend under a range of eta/delta variance ratios; if the reporting-rate trend reverses under plausible ratios, Section 5.2 should be substantially softened.","section":"Section 5.2; Section 4.5; Eqs. (3) and (5)"},{"comment":"The post hoc data modifications are material and are not subjected to sensitivity analysis. Collapsing the 104 Nebraska-Lincoln reports from 2017 to a single report reduces that year's reported total by 103 (from 119 to 16); the Ohio State Strauss exclusion removes 30 reports in 2018 and 97 in 2019; Michigan State University is dropped entirely; and similar collapses are applied to Wells College and Genesee Community College. These choices can affect both systemwide and school-level estimates, yet no analysis reports results on the unmodified data or under less aggressive treatments (for example, excluding one affected school at a time or modeling the repeat-victim counts as a separate category). Given that Appendix D shows the model is sensitive to correlated reporting decisions, the treatment of these extreme records deserves the same scrutiny as the prior specifications.","section":"Appendix A"},{"comment":"The mapping from NCVS statistics to the Clery-reportable sexual assault construct is not established. Clery Act counts use specific offense definitions and require reports to campus authorities or local police, while NCVS measures victimization that may not be reported to any authority. Because the likelihood identifies only the product lambda p, the absolute posterior levels of incidence and reporting rate inherit the prior means. Appendix E varies the prior means over a plausible range of underreporting, but it does not address definitional mismatch or the possibility that the bias in NCVS self-reports varies over time. The paper should either provide an explicit argument that NCVS estimates are commensurate with Clery-reportable campus assaults or add a scenario in which the prior means shift over time; this is relevant to the absolute estimates in Figures 10-11 even if the variance-ratio issue in the previous comment is the more direct threat to the trend conclusion.","section":"Sections 4.1-4.2 and Appendix E"}],"minor_comments":[{"comment":"The sentence describing the priors for the reporting-probability model says \"priors on epsilon, eta, and intercepts beta_0 are chosen,\" which repeats the incidence-model description; it should refer to gamma, delta, and alpha_0.","section":"Section 4.2"},{"comment":"The predictive sampling draws eta_ij ~ N(0, 0.2) and, for new schools, gamma_i ~ N(0, 1) and epsilon_i ~ N(0, 0.5), which do not match the prior variances stated in Eqs. (3) and (5) (0.1, 1.25, and 0.75 respectively). Please clarify whether these are deliberate prior-predictive choices or typographical errors.","section":"Algorithm 2"},{"comment":"The phrase \"on the behalf of the behalf of the US Bureau of Justice Statistics\" contains a duplicated phrase and should be corrected.","section":"Section 3"},{"comment":"In the sentence beginning \"The proposed the power law relationship,\" the word \"the\" is repeated.","section":"Section 5.1"},{"comment":"The word \"distributioin\" is misspelled in the model-comparison subsection.","section":"Appendix C"},{"comment":"The sentence \"the true trend in incidence could conceivably be flat or even increasing\" appears to express uncertainty about whether incidence is flat or decreasing; please rephrase to say what the posterior intervals actually permit.","section":"Section 5.2"}],"recommendation":"major_revision","confidential_remarks":"The prior-variance ratio issue is the key obstacle: the paper's headline conclusion about reporting-rate trends is not identified by the data, and Appendix E does not vary the relevant hyperparameters. A revision that either justifies the eta/delta variance ratio with external evidence or re-runs the analysis across a grid of ratios and still finds the reporting-rate trend would resolve this. The Appendix A data modifications also need a robustness check. If those analyses are added, the paper could become acceptable; in its current form, the central claim is not supported."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick heads-up on Bradshaw and Blei's paper. The modeling is clean, the derivation is correct, and the authors are unusually upfront about identifiability, but the headline temporal result is not established by the data. The likelihood identifies only the product of incidence and reporting rate. The split between the two is governed by the prior variances on the year noise terms: log(lambda) has sd 0.1 and logit(p) has sd 0.5. Because p is five times more mobile on the log scale, the posterior attributes most of the observed rise in reports to p. The paper never varies that variance ratio, and Appendix E only changes the prior means. So the Section 5.2 claim that the increase is more likely due to rising reporting rates is doing the prior's work.\n\nWhat does the paper do well? It is the first school-level decomposition of reporting rate and true incidence for Clery data, which is a real contribution to applied social statistics. The Poisson thinning derivation is correct, the held-out predictive checks are appropriate, and the sensitivity analyses in Appendices D and E are better than what most applied papers report. The authors also clearly state in Section 4.5 that only the product lambda times p is identified. The data edits in Appendix A are disclosed and seem defensible, though they are unmodeled assumptions.\n\nThe soft spots: the NCVS priors carry the absolute levels, and the mapping from NCVS constructs to Clery-reportable events is not tight. The post hoc edits could shift results if the judgment calls are wrong. But the biggest issue is the untested variance ratio, which is load-bearing for the central temporal conclusion. This is a paper for applied statisticians, university administrators, and policy people who read Clery data. It deserves a serious referee, and the referee should require either a genuine sensitivity analysis over the eta and delta scales or a much more hedged conclusion. I would bring it to a reading group—it is a clean case study in non-identifiability and prior choice—and I would cite it if I worked on underreporting.","headline":"A careful, transparent application of underreporting models to campus Clery data, but the headline 'rising reporting rates' result is driven by an untested prior variance ratio, not by the data itself.","tokens_in":21538,"tokens_out":3454,"would_cite":true,"duration_ms":31473,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F15","62P25"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper separates true assault incidence from reporting rates at 1,973 colleges, finding that the 2014-2018 rise in reported campus assaults came mostly from more reporting, not more assaults.","keywords":["sexual assault","underreporting","Bayesian hierarchical model","count data","campus crime statistics","informative priors","Hamiltonian Monte Carlo","identifiability"],"falsifier":"Carry out an anonymous victimization survey at a representative sample of the same campuses, asking about incidents that match the campus crime-report definitions, and compare the implied reporting rate, campus reports divided by survey-estimated victimizations, with the model's posterior median reporting rate; if the survey-based rate is flat or falling from 2014 to 2018, or if true incidence rises while reporting holds steady, the paper's central attribution fails.","tokens_in":20388,"feed_emoji":"📈","tokens_out":14281,"duration_ms":119585,"temperature":0.7,"pith_summary":"This paper tries to separate two quantities hidden inside each college's annual reported sexual assault count: how many assaults actually happen and how likely a student is to report one. Because the reported number alone only fixes the product of these two quantities, the authors add national victimization-survey statistics as informative priors to break the tie, and fit a hierarchical Bayesian model to campus crime data from 2014-2019. Their central finding is that the nationwide rise in reported assaults over this period is better explained by rising reporting rates than by a rise in the true number of assaults. The model also produces school-by-school estimates, and those vary widely: some schools report under 10% of assaults, others around half or more. If the estimates are right, they give administrators a way to read year-over-year changes in their own campus crime numbers instead of taking reported counts at face value.","feed_headline":"Reported campus assaults rose on better reporting, not more assaults","feed_subtitle":"Bayesian model of campus crime data pegs 2014-2018 rise in reports to higher reporting rates, not more assaults.","key_machinery":"The engine is binomial thinning: the observed count $x_{ij}$ is drawn as $\\mathrm{Bin}(z_{ij}, p_{ij})$ from a latent Poisson count $z_{ij} \\sim \\mathrm{Poisson}(\\lambda_{ij})$, so marginally $x_{ij} \\mid \\lambda_{ij}, p_{ij} \\sim \\mathrm{Poisson}(\\lambda_{ij} p_{ij})$. This marginal identity lets the sampler integrate out the discrete latent $z$ and so use gradient-based Hamiltonian Monte Carlo; unreported counts $u_{ij} = z_{ij} - x_{ij}$ are then drawn from $\\mathrm{Poisson}(\\lambda_{ij}(1-p_{ij}))$ to reconstruct the latent total. The identifiability gap, data fix only $\\lambda_{ij} p_{ij}$, is closed by informative priors on $\\lambda$ and $p$ calibrated to national victimization survey estimates, with school-level intercepts and covariates providing partial pooling across the 1,973 schools. The paper shows that even this prior information cannot be fully overtaken by data for the absolute levels, only for covariate slopes, which is why its headline trend claim is more stable than its point estimates.","core_discovery":"The paper's central claim is that the reported increase in campus sexual assaults from 2014 to 2019 is more likely attributable to an increase in reporting rates than to an increase in the true number of assaults. Under the fitted model, the posterior median reporting rate for the college population rose from 17.3% in 2014 to 24.2% in 2018, while the posterior median incidence stayed roughly flat at 2.6 to 2.8 assaults per 1000 students. This separation comes from a hierarchical binomial-thinning model in which the true assault count at each school-year is Poisson and each assault is reported independently with a school-specific probability; observed counts identify only the product $\\lambda p$, so informative priors drawn from national victimization statistics act as the tiebreaker. The same model yields per-school reporting probabilities that range from very low to about 74%, implying that a one-size-fits-all reading of reported campus crime statistics is unreliable. It also associates lower reporting probabilities with junior colleges, religiously oriented institutions, and schools with more need-based aid recipients, and estimates that expected per-capita incidence is higher at smaller schools.","pith_inferences":["The same marginal-Poisson augmentation could be reused in any underreported-count setting where an external source anchors either the event rate or the reporting probability; the independence-of-reporting assumption is the main thing that would have to be checked.","If the upward reporting trend has continued since 2019, reported counts and true incidence could move in opposite directions in coming years, so administrators who use raw reports as a success metric will get the sign wrong.","Because the paper's sensitivity analysis keeps the rising-reporting conclusion even under a 75% downward shift in the prior reporting rate, the direction of the national trend is more trustworthy than the absolute posterior levels.","A direct validation would compare the model's school-level reporting rates with reporting rates measured from a redesigned national victimization survey once campus-specific estimates become available; mismatches would show exactly where the prior transfer breaks down."],"forward_implications":["The 2014-2018 rise in national reported totals is more plausibly a sign that victims are coming forward than a sign that more assaults are occurring.","For a school with a low estimated reporting rate, a year-over-year increase in reports can plausibly come from reporting-rate variation alone; for a high-reporting school, the same increase is harder to explain without an incidence change.","Schools with the lowest estimated reporting probabilities, such as junior colleges, religiously oriented institutions, and schools with many need-based aid recipients, have reported counts that understate their true incidence by the most.","Expected per-capita assault incidence is higher at smaller schools, so comparing raw reported counts across schools of different sizes is misleading."],"supporting_citations":[{"why":"Proves the core identifiability failure: observed Poisson counts identify only the product of the rate and the reporting probability, and uninformative priors cannot separate the two, motivating the paper's informative-prior strategy.","marker":"Moreno and Giron (1998)"},{"why":"Supplies the hierarchical underreporting framework with an informative prior on the reporting probability that the paper extends to covariates and school-level random effects.","marker":"Stoner, Economou and Drummond Marques da Silva (2019)"},{"why":"Models underreported count data with unit-specific incidence and reporting probabilities, the hierarchical structure this paper builds on.","marker":"Dvorzak and Wagner (2016)"},{"why":"Provides national college-age victimization and reporting figures used to motivate and calibrate the informative priors.","marker":"Sinozich and Langton (2014)"},{"why":"Provides the national victimization-survey incidence and police-reporting estimates by age group that anchor the prior distributions for incidence and reporting probability.","marker":"Morgan and Thompson (2020)"}],"fun_headline_variants":["Campus assault rise is a reporting surge, not crime surge","Bayesian model: better reporting, not more assaults, drives rise","Underreported assaults vary wildly by college, model finds","Rising campus assault reports reflect improved reporting, not incidence","Reporting rates up, assault rates flat: Bayesian look at campuses"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that the national victimization-survey statistics used to set the priors give an unbiased picture of true assault incidence and reporting behavior on college campuses; because the observed data pin down only the product of these two quantities, any bias in those priors shifts the absolute estimates.","fun_headline_variants_meta":{"raw":{"variants":["Campus assault rise is a reporting surge, not crime surge","Bayesian model: better reporting, not more assaults, drives rise","Underreported assaults vary wildly by college, model finds","Rising campus assault reports reflect improved reporting, not incidence","Reporting rates up, assault rates flat: Bayesian look at campuses"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00076,"raw_usage":{"total_tokens":3385,"prompt_tokens":967,"completion_tokens":2418,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":583,"completion_tokens_details":{"reasoning_tokens":2333}},"tokens_in":583,"tokens_out":2418,"duration_ms":14756,"temperature":1.0,"reasoning_tokens":2333,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T04:58:43.384931+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Carry out an anonymous victimization survey at a representative sample of the same campuses, asking about incidents that match the campus crime-report definitions, and compare the implied reporting rate, campus reports divided by survey-estimated victimizations, with the model's posterior median reporting rate; if the survey-based rate is flat or falling from 2014 to 2018, or if true incidence rises while reporting holds steady, the paper's central attribution fails.","supporting_citations":[{"cited_title":"and D RUMMOND MARQUES DA SILVA, G","cited_arxiv_id":null,"evidence_quote":"Supplies the hierarchical underreporting framework with an informative prior on the reporting probability that the paper extends to covariates and school-level random effects."}],"review_version":1}