Pith. sign in

REVIEW 4 major objections 5 minor 36 references

The Effect of Perceived Race and Gender on Police Language Use: Experimental Evidence from VR Simulations

T0 review · 4 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read This paper claims that police officers speak measurably less deferentially to VR characters perceived as Black men, and that this per-exchange deficit accumulates over the course of a conversation.

desk verdict A useful experiment and an honest LLM-methods comparison, but the headline 'marginal ATE' is not what it claims because it conditions on a post-treatment mediator; the paper needs major revision, not rejection. read the letter →

arxiv 2608.05050 v1 pith:43FTM42T submitted 2026-08-05 cs.CY cs.AIcs.CL

classification cs.CYcs.AIcs.CL
keywords policelanguagedeferencevirtualrealitysimulationaveragetreatmenteffectcausalinferencelargemodelsracialbiasde-escalationtraining
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper claims that police officers speak measurably less deferentially to virtual reality characters they perceive as Black men, and that this deficit repeats with each conversational exchange rather than being a one-time reaction. Using a treatment defined by randomly assigned character demographics in VR simulations, the authors estimate a marginal average treatment effect on officer deference of $-0.107$ points on a $0$--$10$ scale for the next statement, statistically significant after multiple-testing correction. Across a typical scene of about sixteen exchanges, this per-turn effect accumulates to a drop of several points in deference, which the authors argue can push encounters toward breakdown. They also report that the effect is not uniform: Black male and Black female officers show larger negative effects, while White and biracial or multiracial female officers speak more deferentially to Black man characters. A secondary contribution is methodological: the paper recommends mixed effects models with inverse-propensity treatment weighting and LLM-created text features for estimating causal effects in multilevel text data.

What carries the argument

The load-bearing object is the causal graph in Figure 1 with treatment $D=1$ for a perceived Black man character and $D=0$ for all other character demographics, outcome $U_{t+1}$ the deference of the officer's next statement, and $C$ the historical context (prior dialogue plus average deference up to time $t$). The analysis estimates the propensity score $P[D=1 \mid S, C, M, T]$ with a mixed effects logistic model, then fits a mixed effects linear outcome model for $U_{t+1}$ on the treatment and covariates, and computes the ATE as the difference between inverse-probability-weighted treated and control predicted outcomes. The text context is featurized by taking the last-token embedding of Llama 3.1 8B for each prompt and reducing it with PCA before entering the models. This machinery is what turns the randomly assigned character demographics into a causal estimate of per-exchange deference change.

What would settle it

Re-estimate the ATE without conditioning on the historical context $C$ (or with $C$ modeled as a mediator rather than a confounder). If the total effect of perceiving a Black man character is zero or positive while the $C$-adjusted effect is negative, the per-exchange deference deficit is an artifact of conditioning on a mediator; if the total effect is at least as negative as $-0.107$, the accumulative deficit claim survives.

Watch

Extended reading notes

Core claim

In VR field experiments where 79 U.S. police officers interacted with randomly assigned characters whose skin tone and gender varied, officers' next statements were scored for deference by crowd workers on a $0$--$10$ scale. The paper's central claim is that perceiving a Black man character causes a per-exchange decrease in deference: the overall marginal ATE is $-0.107$ ($p < 0.001$, Benjamini-Hochberg $q = 0.014$), with larger negative effects among Black male officers ($-0.197$) and Black female officers ($-0.258$), and positive effects among White female officers ($+0.414$) and biracial/multiracial female officers ($+0.188$). These marginal effects are small per turn but accumulate: the authors estimate a bus-scene conversation of sixteen exchanges loses almost 4 points of deference from start to finish. The paper also claims its Pipeline I estimator--mixed effects models with inverse propensity treatment weighting and Llama-3.1-8B featurized dialogue context--is valid for detecting such effects, based on synthetic data with known ground-truth effects, and that a fine-tuned LLM regression approach (Pipeline II) is not yet reliable for field data.

Load-bearing premise

The causal estimates treat the conversation's history (prior dialogue and average deference) as a confounder to be adjusted away, but that history is itself shaped by the character's race; if this post-treatment adjustment is wrong, the reported per-exchange effect is not the total effect of the character on deference.

Editorial extensions

If this is right

  • A 16-exchange bus scene carries an estimated cumulative deference loss of roughly 4 points on the $0$--$10$ scale, which is the difference between a cordial and a tense exchange.
  • Because the effect is per-exchange, de-escalation training can target each turn: officers can be taught to maintain or rebuild deference within a single conversation.
  • The positive ATEs for White and biracial/multiracial female officers indicate that not all officers show the pattern, so training content may need to be tailored by officer demographics.
  • Officers with a BA or MA show significantly higher deference (0.34 and 0.54 points), suggesting education-level interventions or training-as-education can shift language use.
  • For multilevel text data, the paper's recommended pipeline (mixed effects + iptw + LLM featurization) should be used instead of fine-tuned LLM prediction models for ATE estimation.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The reported ATE conditions on the historical context $C$, which the paper's own causal graph places on the causal path from $D$ to $U$; the estimate is therefore a controlled direct effect, not the total effect of encountering a Black man character, and the accumulated-conversation claim may differ under a mediation analysis that lets $C$ absorb part of the effect.
  • The 'All Others' control group pools White men, White women, and Black women; the contrast cannot distinguish whether the deference deficit is driven by the character's race, gender, or the specific combination, and a fully crossed design would separate these.
  • The synthetic validation with inserted ground-truth effects demonstrates the pipeline can detect an effect of a given sign, but it does not validate the magnitude of the field ATE; a follow-up could compare the VR per-exchange slopes against body-camera transcripts scored with the same deference rubric.
  • The VR character is a proxy for a Black adult male, not an actual person; whether these estimates transfer to live encounters is a testable question, and the paper's own reliance on real-world body-camera results suggests a direct comparison is feasible.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper analyzes data from VR simulations in which police officers interact with virtual characters whose perceived race and gender are randomly assigned. The treatment is the character demographic Black Man versus a composite All Others group, and the outcome is crowd-worker-rated deference of the officer's next statement. The authors develop two pipelines for estimating an average treatment effect: Pipeline I uses mixed-effects models with inverse propensity weighting and LLM text features, and Pipeline II uses fine-tuned LLM prediction models with doubly robust estimation. The headline result is a statistically significant negative overall ATE of -0.107 on a 0-10 deference scale, with subgroup variation by officer demographics and scene. The paper also compares methods on partially synthetic data with inserted treatment effects and recommends Pipeline I for multilevel text data. The central claim is that police officers speak less deferentially to Black Man characters and that these per-turn effects accumulate over conversations.

Significance. If the causal claim holds, the finding is socially important and methodologically interesting: it uses a randomized VR field experiment with external crowd-worker outcome ratings, which is a substantial strength. The paper also contributes an annotated dataset and a systematic comparison of LLM-based and mixed-effects approaches to ATE estimation with multilevel text data. The synthetic-data validation is a genuine attempt to test the pipelines under known ground truth. However, the significance is qualified by a load-bearing identification problem: the headline estimate conditions on a post-treatment mediator, so the reported quantity is not the marginal ATE claimed in the abstract and conclusion.

major comments (4)
  1. [Our Causal Model; Eq. (1); Table 3] The target estimand is mislabeled. The causal graph in Fig. 1 and the text state that D causes C, the historical context including past dialogue and average deference. C is therefore a post-treatment mediator, not a pre-treatment confounder. Including C in both the propensity model P[D=1|S,C,M,T] and the outcome model in Eq. (1) conditions on this mediator, so the reported -0.107 estimates a controlled direct effect at fixed conversational history rather than the marginal ATE of being assigned a Black Man character. Moreover, because D is randomly assigned, adding post-treatment C to the propensity model can induce selection bias rather than remove confounding. The abstract and conclusion should be corrected: either re-estimate without C to obtain the marginal ATE, or explicitly relabel the results as conditional direct effects and drop the accumulation claim.
  2. [Results; Table 3; Conclusion] The accumulation argument is not a valid total-effect calculation. The text multiplies the per-turn ATE by the number of exchanges (e.g., -0.248 for the bus scene times 16 exchanges gives about 4 points) and states that marginal ATEs accumulate. If part of D's effect operates through the time-varying context C, then a direct effect conditional on C omits the indirect path, and the sum of such direct effects is not the total causal effect of the character assignment over a conversation. The claim of 'two to several points difference' therefore overstates what the analysis identifies. A g-computation or longitudinal marginal structural model would be needed to estimate the cumulative total effect.
  3. [Methodology; Table 3] The treatment contrast conflates race and gender. D=1 is Black Man, while D=0 is a composite of White Woman, Black Woman, and White Man. The reported effects therefore compare Black Man against a mixed control group, not against White Man or Black Woman separately. If the effect of perceived gender or race differs across the three control conditions, the composite ATE is not a meaningful counterfactual for the claim that 'police officers speak less deferentially to Black Man virtual characters versus all other characters.' The paper should report pairwise contrasts or clearly define the estimand as the average effect versus the pooled control, and adjust the language of the substantive conclusion accordingly.
  4. [Table 3; Model Validation Approach] The field ATEs are reported without standard errors or confidence intervals. The text states that bootstrap resampling with 1000 iterations was used, so these quantities should be reported, especially for subgroups with small N such as officer_latina_female (N=53) and officer_white_female (N=52). Additionally, the synthetic-data validation does not test the actual identification structure: the synthetic data are constructed from Black Man conversations only, with an inserted deference shift, and do not include the composite control group or the time-varying mediator structure of the field data. The statement that the model 'is valid' on the basis of this validation is therefore stronger than the evidence supports.
minor comments (5)
  1. [Model Validation Approach; Table 2] The text refers to 'geometric p=2 derived perturbations,' but the table and surrounding text use p=0.2, p=0.3, and p=1.0; this is either a typo or a mismatch between table and prose.
  2. [Eq. (1)] The notation in Eq. (1), such as 'c5 C5' and 't2 T2', is confusing; the coefficients and indexing should be defined consistently, preferably with standard notation such as beta and gamma vectors.
  3. [Table 6 caption] The caption contains the typo 'sythetic' for 'synthetic,' and the label 'Dr' for 'DR' (doubly robust) is inconsistent with the rest of the paper.
  4. [Results] The text says the overall, store, and bus scene subgroups have negative ATEs, but Table 3 also shows a negative and statistically significant house_scene ATE of -0.159; the sentence should be corrected to include it.
  5. [Related Work] The reference to Voigt et al. (2017) is cited as validating the finding in the 'real world,' but that is external correlational evidence, not validation of the present experiment; the wording should be softened to 'consistent with.'

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity: the treatment is randomly assigned and the outcome is externally rated; the mediator-conditioning concern is a methodological issue, not a circular derivation.

full rationale

The paper's central causal claim is not circular. The treatment D (Black Man VR character) is randomly assigned per the cited Doan et al. (2021) design; the outcome U is crowd-worker-rated deference, which is measured independently of the model; and the propensity and outcome models are fit to that external outcome. The synthetic-data validation inserts a known effect and checks whether the estimator recovers it; this is a test of the estimator, not a fit of the target ATE. There is author overlap in the Doan et al. (2021) citation, but that citation supplies the experimental data and randomization, not an unverified uniqueness theorem or a forbidden alternative, so it is not load-bearing circularity. The paper's own causal graph (Fig. 1) has D causing C, yet the modeling conditions on C in the propensity and outcome models; this makes the reported quantity a conditional/direct effect rather than a total marginal ATE. That is a causal-identification and interpretation problem, not a case of the derivation reducing to its inputs by construction. No equation in the paper is defined in terms of its target, and no fitted parameter is renamed as a prediction. Hence the circularity score is low.

Assumptions & free parameters 3 free parameters · 4 assumptions · 0 invented entities

The central claim depends mainly on experimental randomization and external crowd ratings rather than on fitted constants. The hand-chosen modeling inputs are the PCA dimensionality, IPTW weight trimming bounds, and the geometric distribution used for synthetic perturbations; the composite control group and the treatment of conversation history as a confounder are substantive assumptions.

free parameters (3)
  • PCA_components_50 = 50
    Number of principal components retained from Llama 3.1 8B embeddings of dialogue context, used as covariates in the outcome and propensity models; no sensitivity analysis reported.
  • IPTW_weight_trim_bounds = 0.01 to 0.99
    Trimming bounds for inverse propensity weights, chosen without sensitivity analysis; removes observations with extreme propensity scores.
  • synthetic_geometric_p_values = 0.2, 0.3, 1.0
    Parameters for the geometric distribution that generated synthetic deference perturbations; used for model validation, not for the field ATEs.
assumptions (4)
  • domain assumption Stable unit treatment value assumption: no interference between officer-scene units and consistency of potential outcomes.
    Required for ATE identification from the randomized assignment; not stated explicitly in the paper.
  • ad hoc to paper The composite control D=0 (White Woman, Black Woman, White Man) is a valid counterfactual for the Black Man treatment.
    Collapsing three demographics into one control group means the ATE mixes racial and gender differences; see Methodology, D definition.
  • ad hoc to paper Historical context C is a confounder rather than a mediator.
    Fig. 1 shows D causes C; including C in propensity and outcome models assumes it can be adjusted for without destroying the treatment effect.
  • domain assumption LLM embeddings plus PCA capture the conversational context needed for unbiased adjustment.
    Outcome and propensity models rely on 50 PCA components of Llama 3.1 8B embeddings; no robustness checks reported.

how reviews work

0 comments
Cite this review

Pith. "Pith review of The Effect of Perceived Race and Gender on Police Language Use: Experimental Evidence from VR Simulations." pith.science (2026). https://pith.science/paper/43FTM42T

@misc{pith2026260805050,
  author       = {Pith},
  title        = {Pith review of: The Effect of Perceived Race and Gender on Police Language Use: Experimental Evidence from VR Simulations},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/43FTM42T}},
  note         = {Machine review of arXiv:2608.05050}
}
read the original abstract

Against the backdrop of violence in police interactions with the U.S. public, we explore how deferentially police officers speak to virtual characters depicted as Black adult males in vir- tual reality (VR) simulations. We evaluate the effect of seeing and communicating with these characters through a causal in- ference lens, where the assignment of the Black man character to a police officer and simulation is the treatment variable. Our (marginal) average treatment effect AT E measures the social impact of the character on the deference of officer statements with each turn of the conversation. Soberingly, we find that most officers speak less deferentially to Black man characters, except for White, biracial, and multiracial female officers, es- pecially in settings where the VR character was known to be a suspect. Across a full conversation of a typical VR scene, these marginal AT Es can result in notable changes in def- erence of tone (two to several points difference on a scale of 0-10), above and beyond that due to the initial effect of per- ceiving a Black male character. Even more disconcerting is that this can contribute to conversation breakdowns that po- tentially result in violence or danger to both the public and the police. We also explored the capabilities of large language models (LLMs) for ATE estimation. From our methods com- parison analysis, including model validation against synthetic data, we provide unique scientific insights on LLM-assisted methodologies for ATE estimation. As such, for ATE esti- mation with multilevel data with text, we recommend mixed effects models with the inverse propensity treatment weighted (iptw) approach, which utilized an LLM for text feature cre- ation. While we also tested LLMs for finetuning prediction models ultimately for ATE estimation, we conclude they are an area for further development and refinement.

Figures

Figures reproduced from arXiv: 2608.05050 by the authors.

Figure 1
Figure 1. Causal Graph Here we depict the treatment effect of the VR character on the deference ratings (U) of the police of￾ficer’s next statements in a police scene simulation. Confounders also affecting U include C, the scene’s context including historical dialogue and deference ratings of officer statements, and T for the scene (house, bus stop, or convenience store) of the VR simula￾tion. Survey variables (S) proxy B, an… view at source ↗

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

36 extracted references · 24 canonical work pages

  1. [1]

    American journal of epidemiology , volume=

    Doubly robust estimation of causal effects , author=. American journal of epidemiology , volume=. 2011 , publisher=

  2. [2]

    CaM-Gen:Causally-aware Metric-guided Text Generation

    Cam-gen: Causally-aware metric-guided text generation , author=. arXiv preprint arXiv:2010.12795 , year=

  3. [3]

    Multivariate behavioral research , volume=

    An introduction to propensity score methods for reducing the effects of confounding in observational studies , author=. Multivariate behavioral research , volume=. 2011 , publisher=

  4. [4]

    What police chiefs and sheriffs need to know about collecting and analyzing use-of-force data , booktitle=

    Police Executive Research Forum , year=. What police chiefs and sheriffs need to know about collecting and analyzing use-of-force data , booktitle=

  5. [5]

    Annual Review of Criminology , volume=

    Exceptionally lethal: American police killings in a comparative perspective , author=. Annual Review of Criminology , volume=. 2023 , publisher=

  6. [6]

    2007 , publisher=

    Data analysis using regression and multilevel/hierarchical models , author=. 2007 , publisher=

  7. [7]

    Understanding

    Eric Tang and Bangding Yang and Xingyou Song , journal=. Understanding. 2025 , url=

  8. [8]

    ACM computing surveys (CSUR) , volume=

    Deep learning--based text classification: a comprehensive review , author=. ACM computing surveys (CSUR) , volume=. 2021 , publisher=

Show all 36 references
  1. [9]

    The Annals of Family Medicine , volume=

    Multilevel modeling and practice-based research , author=. The Annals of Family Medicine , volume=. 2005 , publisher=

  2. [10]

    arXiv preprint arXiv:2406.17169 , year=

    Multi-logieval: Towards evaluating multi-step logical reasoning ability of large language models , author=. arXiv preprint arXiv:2406.17169 , year=

  3. [11]

    R package version , volume=

    nlme: Linear and nonlinear mixed effects models , author=. R package version , volume=

  4. [12]

    Journal of statistical software , volume=

    Fitting linear mixed-effects models using lme4 , author=. Journal of statistical software , volume=

  5. [13]

    Proceedings of the national Academy of sciences , volume=

    Language from police body camera footage shows racial disparities in officer respect , author=. Proceedings of the national Academy of sciences , volume=. 2017 , publisher=

  6. [14]

    , title =

    Meta Platforms,Inc. , title =. 2024 , howpublished =

  7. [15]

    Department of Justice, Office of Justice Programs , institution =

    U.S. Department of Justice, Office of Justice Programs , institution =. T3 - Tact, Tactics, and Trust Training and Technical Assistance , url =. 2019 , publisher =

  8. [16]

    Motz and Jennifer Cherkauskas , institution =

    Robin Engel and Nicholas Corsaro and Ryan T. Motz and Jennifer Cherkauskas , institution =. Evaluation of Integrating Communications, Assessment, and Tactics (ICAT) Training with the Indianapolis Metropolitan Police , url =. 2025 , publisher =

  9. [17]

    Policing: An International Journal , volume=

    Moving the needle: can training alter officer perceptions and use of de-escalation? , author=. Policing: An International Journal , volume=. 2021 , publisher=

  10. [18]

    arXiv preprint arXiv:2403.06963 , year=

    The pitfalls of next-token prediction , author=. arXiv preprint arXiv:2403.06963 , year=

  11. [19]

    Policy Insights from the Behavioral and Brain Sciences , volume=

    The social psychology of racially biased policing: Evidence-based policy responses , author=. Policy Insights from the Behavioral and Brain Sciences , volume=. 2020 , publisher=

  12. [20]

    Proceedings of the National Academy of Sciences , volume=

    Escalated police stops of Black men are linguistically and psychologically distinct in their earliest moments , author=. Proceedings of the National Academy of Sciences , volume=. 2023 , publisher=

  13. [21]

    Clinical Simulation in Nursing , volume=

    Enhancing learning outcomes through AI-driven simulation in nursing education: A systematic review , author=. Clinical Simulation in Nursing , volume=. 2025 , publisher=

  14. [22]

    2000 , publisher=

    Mixed-effects models in S and S-PLUS , author=. 2000 , publisher=

  15. [23]

    The Thirteenth International Conference on Learning Representations , year=

    Better autoregressive regression with LLMs via regression-aware fine-tuning , author=. The Thirteenth International Conference on Learning Representations , year=

  16. [24]

    arXiv preprint arXiv:2503.04381 , year=

    TRACT: Regression-Aware Fine-tuning Meets Chain-of-Thought Reasoning for LLM-as-a-Judge , author=. arXiv preprint arXiv:2503.04381 , year=

  17. [25]

    World Journal of Gastroenterology , volume=

    Machine learning models and over-fitting considerations , author=. World Journal of Gastroenterology , volume=

  18. [26]

    International Conference on Human-Computer Interaction , pages=

    Evaluation of a virtual reality simulation tool for studying bias in police-civilian interactions , author=. International Conference on Human-Computer Interaction , pages=. 2021 , organization=

  19. [27]

    Group Processes & Intergroup Relations , volume=

    Policing at the crossroads: An intergroup communication accommodation perspective , author=. Group Processes & Intergroup Relations , volume=. 2024 , publisher=

  20. [28]

    Indian Journal of Thoracic and Cardiovascular Surgery , volume=

    The importance of randomization in clinical research , author=. Indian Journal of Thoracic and Cardiovascular Surgery , volume=. 2022 , publisher=

  21. [29]

    The Leadership Quarterly , volume=

    Causal inference with observational data: A tutorial on propensity score analysis , author=. The Leadership Quarterly , volume=. 2023 , publisher=

  22. [30]

    British Journal of Occupational Therapy , volume=

    Achieving inter-rater agreement and inter-rater reliability to assess fidelity of an occupation-based coaching (OBC) clinical trial intervention , author=. British Journal of Occupational Therapy , volume=. 2025 , publisher=

  23. [31]

    Biometrika , volume=

    The central role of the propensity score in observational studies for causal effects , author=. Biometrika , volume=. 1983 , publisher=

  24. [32]

    arXiv preprint arXiv:2403.08295 , year=

    Gemma: Open models based on gemini research and technology , author=. arXiv preprint arXiv:2403.08295 , year=

  25. [33]

    arXiv preprint arXiv:1907.11692 , year=

    Roberta: A robustly optimized bert pretraining approach , author=. arXiv preprint arXiv:1907.11692 , year=

  26. [34]

    Statistics in medicine , volume=

    Propensity score weighting with multilevel data , author=. Statistics in medicine , volume=. 2013 , publisher=

  27. [35]

    Multivariate Behavioral Research , volume=

    Causal inference with multilevel data: a comparison of different propensity score weighting approaches , author=. Multivariate Behavioral Research , volume=. 2022 , publisher=

  28. [36]

    Advances in neural information processing systems , volume=

    Attention is all you need , author=. Advances in neural information processing systems , volume=

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.