{"id":"d7badcc5-ba0f-4920-adac-7f21f3d78a45","arxiv_id":"1908.08116","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A new design and inference pipeline lets researchers estimate how the probability a lawyer replies to an email changes with the strength of a name's racial signal, tested with an underpowered Florida pilot.","lead":"This paper builds a statistical framework for email audit experiments that test whether lawyers treat potential clients differently based on how strongly their names signal race. It uses names with graded racial associations and estimates the shape of the response curve, though the demonstration study in Florida was too small to detect any effect.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The fixed-ξ assumption is the load-bearing point: Eq. (6) treats MTurk-estimated race levels as known covariates; ignoring their measurement error biases MLE and overstates precision, which also undermines the k=6 design simulation.","rationale":"The reader's weakest assumption identifies the same load-bearing point, and the paper's own Section 6 limitation makes it in-scope. I checked the likelihood and Fisher-information algebra: apart from a sign typo in Appendix A.1 for ∂²ℓ/∂a∂b, Eq. (10) is consistent with the expected information, so the concern is not internal inconsistency. The threat is calibration: the method's validity and the design recommendation both depend on ξ being known or at least measured with negligible error. This does not require changing the reader's conditional verdict; it sharpens the condition. The empirical null result is honestly reported and the framework is a genuine methodological contribution, but the central inference claim stands only if the fixed-ξ assumption is either corrected or shown to be harmless.","tokens_in":15729,"tokens_out":12958,"duration_ms":130469,"concrete_test":"Compute the per-name sampling variance of the MTurk proportions used in Figure 5 (or use a conservative n=50 per name). Run the Appendix B power simulation and the Table 2 analysis with measurement error: for each Monte Carlo draw, generate true ξ_j from a Beta distribution with mean equal to the reported level and variance equal to the survey variance, simulate responses under the model, then estimate parameters while treating the reported ξ̂_j as fixed. Compare empirical coverage of the nominal 95% CI for β−α and the δ1/δ2 power curves to the paper's Figures 7-8. If coverage falls below 90% or power changes by more than 10 percentage points, the fixed-ξ assumption is load-bearing and the 'ready for use' claim must be conditioned on precise, target-population measurement of ξ.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 2 defines ξ as the probability that a receiver identifies a name as black, and Section 3's likelihood (6) and Fisher information (10) treat the ξ_i assigned to each lawyer as known constants. In the implemented design (Section 4.2), ξ is not set by the experimenter; it is estimated as the proportion of MTurk raters who call the name black. The estimation noise in ξ̂ is ignored in the MLE, in the expected information matrix, and in the D-optimality and power simulations of Section 4.1 and Appendix B. This is a classical errors-in-variables problem: even zero-mean noise in a nonlinear covariate biases estimates of α, β, and γ and makes the plug-in covariance matrix inconsistent, so the nominal 95% intervals for β−α and the power curves are not calibrated. More importantly, the design conclusion that k=6 is robust presumes the six selected names have the intended equally spaced ξ values; if the MTurk estimates are noisy or lawyers' perceptions differ systematically from MTurk respondents, the actual design is no longer the one simulated. The authors flag this in Section 6 ('we have assumed that the input level ξ is a constant and can be chosen without noise') but do not quantify the impact on their central claims.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper develops a design-of-experiments and maximum-likelihood framework for email audit studies that measure racial discrimination in access to lawyers using more than two race-signal levels. It defines a race level ξ∈(0,1) as the probability that a receiver identifies a sender's name as black, models the response probability π(ξ) with a parametric family (2)-(5) indexed by boundary probabilities α, β and shape parameter γ, and derives an iterative MLE, the asymptotic normal distribution of the estimator (Proposition 1), and delta-method covariance results for α, β and predictions (Corollaries 1-2). The design section uses D-optimality simulations over k levels to recommend six equally spaced treatment levels, describes name selection using Florida voter-file data and MTurk perception surveys, and presents a blocked randomization of 899 Florida lawyers. The empirical analysis finds no statistically significant treatment effect, and the authors are explicit that resource constraints prevented a definitive test. The paper closes by acknowledging that the treatment level is assumed known without noise.","tokens_in":16041,"tokens_out":11131,"duration_ms":105104,"significance":"If the statistical claims hold, the paper makes a useful methodological contribution to correspondence audit design: it replaces the common binary black/white name manipulation with a multi-level perceived-race treatment and provides a coherent likelihood-based estimation and power-analysis workflow. The derivations are standard but carefully executed, the design simulations are honest about the large sample sizes required, and the paper explicitly flags the fixed-ξ assumption and the pilot nature of the empirical application. The main value is as a template for planning multi-level audit experiments rather than as a source of substantive estimates, since the Florida pilot lacks power.","major_comments":[{"comment":"The inference and design results treat each ξ_i as a fixed, exactly known treatment level, but in the implemented experiment ξ_i is not controlled; it is a point estimate (an MTurk proportion or an objective Bayes probability) of the perception probability for a selected name. Measurement error in a nonlinear covariate generally biases the MLE of (α, β, γ) and makes the plug-in covariance matrix from Proposition 1 inconsistent, so the nominal 95% intervals and the power curves in Appendix B are not calibrated to the actual randomized treatment. The acknowledgment in Section 6 (\"we have assumed that the input level ξ is a constant and can be chosen without noise\") states the problem but does not quantify its effect on the central claims. I ask the authors to add a sensitivity analysis (e.g., simulation under plausible true ξ values and MTurk sample sizes) or an errors-in-variables formulation, and to state explicitly which ξ values (objective, subjective, or a combination) were used in the Section 5 analysis.","section":"Section 3 (Eqs. (6), (10)), Section 4.2, Section 6"},{"comment":"The text says the preliminary proportions showed a lower response at the lowest race level (very likely black) than at the highest (very likely white), and that the model was therefore fit under α<β; however, all three rows of Table 2 report α̂>β̂, which is the decreasing-branch case, and the reported intervals are centered at α̂−β̂ rather than β̂−α̂ (e.g., overall midpoint 0.1569 equals 0.3687−0.2118). The heading \"95% CI for ˆβ−ˆα\" is therefore inconsistent with the entries, and the description of which model branch was fitted is internally contradictory. This must be corrected and reconciled because it determines the sign convention for the estimated response curve and the interpretation of the confidence interval for the treatment effect.","section":"Section 5 and Table 2"}],"minor_comments":[{"comment":"The equation for updating b is said to be solved \"for a\"; it should be solved for b.","section":"Section 3, Algorithm Step 7"},{"comment":"In \"det(IN(θ) = ∑Ni=1 Ii(θ))\" a closing parenthesis is missing, and the cross-reference to \"Ii(θ) is given by (1)\" should point to Eq. (10), not Eq. (1).","section":"Section 4.1"},{"comment":"The y-axis label \"Estimated subjective probability\" does not say whether the probability is of identifying the name as black or as white; since ξ is defined as black-identification probability, this needs to be stated.","section":"Figure 5"},{"comment":"The clause \"when there is a large treatment effect (α−β = 0.6, the asymptotic intervals...\" is missing a closing parenthesis.","section":"Appendix B, Figure 7 caption"},{"comment":"The Wald (1943) entry (\"Transactions of the American Mathematical Society. Health Educ. Res. 54, 426–482\") and the Pager et al. (2006) entry (whose title appears to belong to a different paper) need to be corrected.","section":"References"},{"comment":"Minor typographical errors include \"adu it studies\" in Section 1 and \"nave Bayesian\" in Section 4.2.","section":"Sections 1 and 4.2"}],"recommendation":"major_revision","confidential_remarks":"The paper is honest about its limitations and the requested revisions are feasible within the manuscript's scope; I do not see grounds for rejection. The empirical sign inconsistencies suggest that the authors should re-check all reported estimates against the actual fitted model branch before resubmission."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear [Colleague],\n\nHere's the quick take on Libgober and Dasgupta: it's a real methods contribution to audit studies, giving a way to estimate a dose-response curve over racial signaling rather than the usual two-name black/white comparison. The statistical core is sound; the soft spots are known but not resolved.\n\nWhat's new: The paper adapts a probability weighting function to model response as a function of ξ, the probability a receiver reads a name as black, and works out the MLE, Fisher information, and delta-method standard errors. The design section goes beyond platitudes: simulations to choose the number of levels (k=6), a naive Bayes pipeline for name selection from the Florida voter file, and MTurk validation. The pilot is honestly reported as underpowered, and the power analysis in Appendix B is unusually frank about sample size demands. The math in Section 3 checks out as far as I can see.\n\nWhere it's soft. First and most important: the likelihood and the design treat each name's ξ as a known constant, but in the actual experiment ξ is a point estimate from MTurk. Measurement error in a nonlinear covariate doesn't just inflate standard errors; it biases the MLE of α, β, and γ, and it means the k=6 design simulations are run on levels that may not be the ones actually administered. The authors flag this in Section 6 but don't quantify or fix it. This isn't a reason to kill the paper, but a serious revision needs to confront it, even if only with a sensitivity analysis or a measurement-error model.\n\nSecond, Section 5 has an internal inconsistency: the text says the fitted model assumed α<β, but the reported estimates (Table 2) all have α>β. That has to be a typo or a leftover from an earlier draft, and it obscures what the pilot actually showed.\n\nThird, and more of a design concern: there is one name per level. Any idiosyncratic effect of \"Latasha\" versus \"Nicole\" is completely confounded with the race level. The method assumes names differ only through ξ. Given the whole point is to vary ξ continuously, this is a real limitation, though common in early dose-response work.\n\nBottom line: this deserves a serious referee. It's a useful template for statisticians designing audit studies, and the honest power analysis is a plus. I wouldn't cite the empirical null, but I'd consider citing the model and design approach, with a caveat about ξ. Send it to review, with instructions to push on measurement error and name-specific effects.","headline":"A genuine statistical template for graded audit studies, with a known-but-unaddressed measurement error problem in the race levels and a Section 5 inconsistency that should be fixed before publication.","tokens_in":16513,"tokens_out":5276,"would_cite":true,"duration_ms":53042,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62K05","62F12","62P25"],"pacs":[],"model":"deepseek-v4-flash","headline":"An email audit with six graded name levels can estimate the full response curve of lawyer bias, not just a black-white gap.","keywords":["access to justice","design of experiments","maximum likelihood","binary response","email audit studies","racial discrimination","race level","inverted-S response function"],"falsifier":"Run the six-level design with a total sample of at least 2000 units in a market where a prior binary-name audit already shows a large black-white response gap, and estimate $\\alpha$, $\\beta$, and $\\gamma$; if the fitted 95% confidence interval for $\\alpha-\\beta$ still contains zero, the design fails to recover a known effect. Separately, ask the lawyers themselves to rate the six sender names on the same black-white scale used to fix $\\xi$; if their average ratings deviate from the assigned levels by more than sampling error, the fixed-$\\xi$ assumption is violated.","tokens_in":15565,"feed_emoji":"⚖️","tokens_out":11530,"duration_ms":107777,"temperature":0.7,"pith_summary":"Standard email-audit experiments compare only very-black-sounding and very-white-sounding names, so they can detect discrimination but cannot say whether the effect is confined to extreme names or grows steadily with every hint of race. This paper establishes a statistical framework in which each sender name carries a graded \"race level\" $\\xi$ in $(0,1)$, and the goal is inference on the shape of the response probability $\\pi(\\xi)$ as a function of $\\xi$. The response function is parameterized by boundary rates $\\alpha$ and $\\beta$ and a single shape parameter $\\gamma$, with $\\gamma=1$ giving a linear relationship, $\\gamma>1$ a logistic curve, and $\\gamma<1$ an inverted-S curve; the paper derives maximum-likelihood estimation, asymptotic standard errors, and prediction variances for this model. On the design side, simulations of the Fisher information determinant indicate that six equally spaced race levels are near-optimal across a wide range of parameter values, and the paper demonstrates a complete workflow of name selection, blocking, randomization, and analysis on a pilot sample of lawyers. The pilot's estimate of the treatment gap is not statistically significant, but the paper's contribution is the inference-and-design machinery, which it argues is ready for markets where discrimination is already known to exist.","feed_headline":"Six name levels reveal the shape of lawyer race bias","feed_subtitle":"A single parameter now separates linear, logistic, and inverted-S response curves, and six levels are enough to estimate it.","key_machinery":"The load-bearing object is the probability-weighting function $g(\\gamma,\\xi)=\\xi^{\\gamma}/(\\xi^{\\gamma}+(1-\\xi)^{\\gamma})$, taken from the psychology literature on how people weight probabilities. It converts the shape of the race-response curve into a single scalar: $\\gamma=1$ recovers a straight line between the boundary rates, $\\gamma>1$ produces the steep-midrange curve the paper calls logistic, and $\\gamma<1$ produces the extreme-sensitive inverted-S curve. The function enters the likelihood through the reparameterization $\\pi(\\xi)=(a+g)/(a+b)$ or $(b-g)/(a+b)$, where $a$ and $b$ are deterministic functions of the boundary rates $\\alpha$ and $\\beta$, and it is the derivative $g'(\\gamma,\\xi)$ that appears in the score equations and the Fisher information matrix. All of the paper's inference results, including the MLE algorithm, Proposition 1, the delta-method corollaries, and the D-optimality simulations, flow from this single parametric family.","core_discovery":"On the paper's own terms, the core discovery is that a graded-treatment email audit can be designed and analyzed so that the qualitative question \"is lawyer discrimination linear or concentrated at extreme names?\" becomes a parameter-estimation problem. Using the odds-ratio function $g(\\gamma,\\xi)=\\xi^{\\gamma}/(\\xi^{\\gamma}+(1-\\xi)^{\\gamma})$ inside a likelihood for binary responses, the paper shows that the boundary response probabilities $\\alpha$ and $\\beta$ and the shape parameter $\\gamma$ are identifiable, gives their asymptotic normal approximation, and derives a delta-method covariance for $\\alpha$ and $\\beta$ and a prediction variance for $\\pi(\\xi)$ at any new level. Design simulations then show that six equally spaced levels of $\\xi$ keep the determinant of the Fisher information high over a wide grid of $(\\alpha,\\beta,\\gamma)$, making the design a near-optimal default. The pilot application, using 899 lawyers, six names per gender, and blocking on lawyer race, gender, and practice area, yields point estimates $\\hat{\\alpha}=0.3687$, $\\hat{\\beta}=0.2118$, and $\\hat{\\gamma}=0.688$, with a 95% confidence interval for $\\beta-\\alpha$ that contains zero, so the pilot cannot reject a null treatment effect but does demonstrate the intended analysis.","pith_inferences":["Editorial inference: the same shape parameter $\\gamma$ could be exported to other continuous signals, such as accent strength, neighborhood cues, or credit-score proxies, so that a single number characterizes whether any service market discriminates in a threshold or gradient fashion; this is a direct generalization the paper does not spell out.","Editorial inference: the fixed-$\\xi$ assumption could be tested by collecting race ratings of the actual sender names from a sample of the same profession being audited; if lawyer ratings differ systematically from the survey panel's proportions, the model would need a measurement-error layer for $\\xi$, which the authors flag as future work.","Editorial inference: an adaptive design, with a first wave to localize the response curve and a second wave concentrating treatment levels there, could recover shape information at lower total sample cost than the fixed equally spaced design, an extension suggested by the paper's own power constraints.","Editorial inference: the pilot's wide confidence interval, combined with the paper's power analysis, implies that conclusions of \"no discrimination\" from small correspondence audits should be treated cautiously; failure to reject could simply reflect sample size rather than the absence of bias."],"forward_implications":["The distinction between linear and nonlinear discrimination becomes a Wald test of $\\gamma=1$, so an adequately powered audit can say whether bias is a smooth gradient or switches on only for very distinctive names.","With six equally spaced levels and equal allocation, the design is near-D-optimal across a wide grid of $\\alpha$, $\\beta$, and $\\gamma$, meaning future audits can adopt it without running bespoke optimal-design calculations for each new market.","The iterative MLE and asymptotic standard errors apply to any binary-outcome correspondence audit with a continuous treatment, so the method generalizes beyond legal services to rental, employment, and other service-provision contexts.","Power simulations show that resolving shape requires large samples, with $\\gamma$ far from 1 and large $|\\beta-\\alpha|$ most favorable; the 899-lawyer pilot is underpowered for small effects, and its non-significant result should not be read as evidence against discrimination.","The design is presented as ready for use where discrimination has already been established; in undiagnosed markets, the sample sizes needed for shape inference are prohibitive."],"supporting_citations":[{"why":"Supplies the canonical correspondence-audit template and the extreme-name binary convention that this paper replaces with graded race levels.","marker":"Bertrand and Mullainathan (2004)"},{"why":"Provides evidence that perceived race of names is a continuous variable, motivating treatment levels beyond the binary shortcut.","marker":"Gaddis (2017a)"},{"why":"Contributes the naive-Bayes name-race probability calculation used to identify candidate names at intermediate levels.","marker":"Fryer and Levitt (2004)"},{"why":"Supplies the probability-weighting function $g(\\gamma,\\xi)$ whose parameter $\\gamma$ defines the family of response shapes.","marker":"Gonzalez and Wu (1997)"},{"why":"Gives the asymptotic maximum-likelihood theory behind the normal approximation and Fisher-information covariance in Proposition 1.","marker":"Boos and Stefanski (2013)"},{"why":"Introduces locally optimal designs, the D-optimality idea the paper adapts through simulation across parameter settings.","marker":"Chernoff (1953)"},{"why":"Provides the D-optimality criterion used to compare candidate numbers of treatment levels in the design simulations.","marker":"Atkinson et al. (2007)"}],"fun_headline_variants":["Statistical model maps shape of lawyer race bias","Six name levels estimate lawyer response curve","Pilot email audit shows how to detect lawyer bias","Graded names distinguish linear from extreme bias"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise, flagged by the authors themselves, is that the race level $\\xi$ assigned to each sender name, meaning the proportion of a survey panel who read the name as black, is an exact and noise-free constant describing how lawyers perceive that name; if lawyers' perceptions differ systematically, every estimated response curve, standard error, and design choice is miscalibrated.","fun_headline_variants_meta":{"raw":{"variants":["Statistical model maps shape of lawyer race bias","Six name levels estimate lawyer response curve","Pilot email audit shows how to detect lawyer bias","Graded names distinguish linear from extreme bias"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00029,"raw_usage":{"total_tokens":1702,"prompt_tokens":954,"completion_tokens":748,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":570,"completion_tokens_details":{"reasoning_tokens":692}},"tokens_in":570,"tokens_out":748,"duration_ms":7488,"temperature":1.0,"reasoning_tokens":692,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T11:49:48.288519+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the six-level design with a total sample of at least 2000 units in a market where a prior binary-name audit already shows a large black-white response gap, and estimate $\\alpha$, $\\beta$, and $\\gamma$; if the fitted 95% confidence interval for $\\alpha-\\beta$ still contains zero, the design fails to recover a known effect. Separately, ask the lawyers themselves to rate the six sender names on the same black-white scale used to fix $\\xi$; if their average ratings deviate from the assigned levels by more than sampling error, the fixed-$\\xi$ assumption is violated.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the canonical correspondence-audit template and the extreme-name binary convention that this paper replaces with graded race levels."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Contributes the naive-Bayes name-race probability calculation used to identify candidate names at intermediate levels."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the probability-weighting function $g(\\gamma,\\xi)$ whose parameter $\\gamma$ defines the family of response shapes."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Gives the asymptotic maximum-likelihood theory behind the normal approximation and Fisher-information covariance in Proposition 1."},{"cited_title":"Donev, and R","cited_arxiv_id":null,"evidence_quote":"Provides the D-optimality criterion used to compare candidate numbers of treatment levels in the design simulations."}],"review_version":1}