{"id":"17d4253d-a79d-4dc2-b19f-aeb630666663","arxiv_id":"2603.28675","paper_version":2,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":3.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Aggregate accuracy metrics obscure demographic disparities in false positive and false negative rates for law enforcement facial recognition systems.","lead":"The paper argues that overall accuracy numbers for facial recognition tools used by police hide large differences in mistakes made for different demographic groups. A smart generalist should read it because relying on a single average score can lead to unfair outcomes like wrongful stops or missed identifications in real security work.","discovery_kind":"extension","skeptic_critique":{"model":"grok-4.3","headline":"No significant objection identified","rationale":"The reader's weakest_assumption correctly flags the missing link from error-rate disparities to quantified real-world harm, yet this does not undermine the narrower claim that aggregate accuracy alone fails to reveal those disparities. Because the paper does not assert that subgroup rates are sufficient for final fairness judgments, the identified gap is real but not load-bearing for the headline argument. The UNVERDICTED status and low confidence therefore remain appropriate given the abstract-only review.","tokens_in":1692,"tokens_out":291,"duration_ms":38194,"concrete_test":"Locate the empirical observations section; extract any reported pairs of systems or models with aggregate accuracy within 2% of each other, then verify whether the reported subgroup FPR/FNR differences exceed 5 percentage points and whether base-rate or threshold sensitivity was tested.","verdict_should_be":"UNCHANGED","load_bearing_attack":"No significant objection identified. The central claim—that aggregate accuracy can obscure subgroup disparities in FPR and FNR—is mathematically direct: overall accuracy is a prevalence-weighted average, so equal aggregates are compatible with arbitrarily different per-group error rates. The abstract's reference to empirical observations of systems with similar aggregate accuracy but divergent fairness profiles is consistent with this fact and with prior documented cases in facial recognition. The extension to law enforcement operational risks follows without internal contradiction or unstated assumption that would invalidate the argument as stated.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript argues that aggregate accuracy is an insufficient metric for evaluating the fairness and reliability of facial recognition systems in law enforcement contexts. Through subgroup-level analysis of false positive rates (FPR) and false negative rates (FNR), it shows that systems with similar overall accuracy can exhibit substantially different fairness profiles across demographic groups. The paper discusses the operational risks of misclassification in high-stakes applications and advocates for fairness-aware evaluation frameworks and model-agnostic auditing strategies.","tokens_in":1790,"tokens_out":383,"duration_ms":46875,"significance":"If the empirical observations hold, this work usefully highlights a well-known but practically important limitation of aggregate metrics in algorithmic fairness. The logical point that overall accuracy is a prevalence-weighted average compatible with arbitrary subgroup disparities is direct and relevant to responsible deployment of facial recognition in law enforcement. The emphasis on post-deployment auditing adds a pragmatic angle to the existing literature on disaggregated performance evaluation.","major_comments":[{"comment":"Abstract: The central claim rests on 'empirical observations' that systems with similar aggregate accuracy exhibit substantially different fairness profiles. However, the manuscript provides no specific subgroup FPR/FNR values, sample sizes, demographic breakdowns, or references to the underlying datasets or studies, leaving the magnitude and statistical reliability of the reported disparities unquantified and difficult to evaluate.","section":null}],"minor_comments":[{"comment":"The abstract and introduction would benefit from a short, explicit statement of the demographic groups analyzed and the sources of the empirical observations to allow readers to assess generalizability immediately.","section":null},{"comment":"Consider adding a brief discussion of how FPR/FNR disparities translate to concrete operational risks (e.g., wrongful arrest rates or missed identifications) rather than leaving the link at a high level.","section":null}],"recommendation":"minor_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive feedback and the recommendation for minor revision. We address the single major comment below.","responses":[{"response":"We agree that the abstract would be strengthened by greater specificity. In the revised manuscript we will add concrete examples drawn from the existing literature on commercial facial recognition systems (e.g., NIST FRVT reports and academic benchmarks), including illustrative FPR and FNR values across demographic subgroups, approximate sample sizes, and explicit citations to the source datasets and studies. These additions will quantify the disparities and allow readers to assess their statistical reliability without altering the paper’s conceptual focus.","revision_made":"yes","referee_comment":"[—] Abstract: The central claim rests on 'empirical observations' that systems with similar aggregate accuracy exhibit substantially different fairness profiles. However, the manuscript provides no specific subgroup FPR/FNR values, sample sizes, demographic breakdowns, or references to the underlying datasets or studies, leaving the magnitude and statistical reliability of the reported disparities unquantified and difficult to evaluate."}],"tokens_in":1271,"tokens_out":233,"duration_ms":40075,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main point is that a single accuracy number for facial recognition systems can look fine while false positive and false negative rates differ sharply across demographic groups. In law enforcement that matters because errors can lead to wrongful stops or overlooked suspects. The paper makes this case by walking through how overall accuracy is just a weighted average that can stay high even when subgroups perform poorly.","headline":"Aggregate accuracy can hide demographic disparities in law enforcement facial recognition, a point the paper applies to operational risks but does not advance with new data or derivations.","tokens_in":2246,"tokens_out":147,"would_cite":false,"duration_ms":28286,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":{"model":"grok-4.3","evidence":[{"relation":"unclear","rs_module":"IndisputableMonolith/Cost/FunctionalEquation.lean","rs_theorem":"washburn_uniqueness_aczel","paper_passage":"aggregate accuracy is an insufficient metric... subgroup-level error distribution, including false positive rate (FPR) and false negative rate (FNR)"},{"relation":"unclear","rs_module":"IndisputableMonolith/Foundation/RealityFromDistinction.lean","rs_theorem":"reality_from_one_distinction","paper_passage":"systems with similar overall accuracy can exhibit substantially different fairness profiles"}],"headline":"Paper on aggregate accuracy vs. subgroup FPR/FNR disparities in facial recognition has no overlap with RS forcing chain or cost machinery","alignment":"orthogonal","rationale":"The paper's core argument (accuracy as prevalence-weighted average can mask arbitrary per-group FPR/FNR differences; empirical tables showing race/age variation despite fixed aggregate accuracy) is standard classification theory with no connection to recognition cost J(x), phi-ladder, 8-tick periodicity, or any theorem in the RS Lean corpus (e.g., reality_from_one_distinction, AbsoluteFloorClosure, AlexanderDuality, Cost.FunctionalEquation). Domain is applied AI ethics/auditing; RS has no opinion on fairness metrics or law-enforcement deployment.","tokens_in":43685,"confidence":"high","tokens_out":314,"duration_ms":8964,"cache_read_input_tokens":32896,"cache_creation_input_tokens":0},"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Aggregate accuracy metrics can mask substantial differences in error rates for different demographic groups in facial recognition systems.","keywords":["facial recognition","fairness","aggregate accuracy","law enforcement","false positive rate","false negative rate","demographic groups","evaluation metrics"],"falsifier":"Deployment data from a law enforcement facial recognition system that tracks the demographic breakdown of actual false positives and false negatives and checks whether the groups with higher rates experience corresponding increases in negative outcomes like wrongful arrests.","tokens_in":2595,"feed_emoji":"⚖️","tokens_out":612,"duration_ms":51226,"temperature":0.7,"pith_summary":"The paper establishes that a single overall accuracy figure for facial recognition systems does not reveal whether performance is consistent across demographic groups. By examining false positive and false negative rates separately for each subgroup, it shows that two systems with nearly identical aggregate accuracy can still produce very different patterns of mistakes. This matters in law enforcement because uneven error rates can lead to more frequent wrongful identifications or overlooked threats for particular populations. The argument calls for evaluation methods that prioritize subgroup fairness over headline accuracy numbers.","feed_headline":"Accuracy scores hide demographic disparities in facial recognition","feed_subtitle":"Systems can match on overall accuracy while differing sharply in false positive and false negative rates for demographic groups, creatingune","key_machinery":"Subgroup-level analysis of false positive rates (FPR) and false negative rates (FNR) across demographic groups, which exposes performance variations that aggregate accuracy conceals.","core_discovery":"Aggregate accuracy is an insufficient metric for evaluating the fairness and reliability of facial recognition systems in high-stakes environments. Through analysis of subgroup-level error distributions including false positive rate and false negative rate, the paper demonstrates how aggregate performance metrics can obscure critical disparities across demographic groups. Systems with similar overall accuracy can exhibit substantially different fairness profiles, and this has implications for operational risks such as wrongful suspicion or missed identification in law enforcement applications.","pith_inferences":["These findings imply that similar metric problems could affect other high-stakes AI applications beyond facial recognition.","Quantifying the translation from error rates to actual societal harm would strengthen the case for changing evaluation practices.","Testing these auditing methods on commercial facial recognition tools could provide concrete examples of the disparities."],"forward_implications":["Operational decisions in law enforcement may carry unequal risks for different demographic groups despite high reported accuracy.","Fairness-aware evaluation approaches become necessary to identify and address these hidden disparities.","Model-agnostic auditing strategies enable ongoing assessment of deployed systems.","Responsible AI deployment requires moving beyond accuracy as the primary evaluation metric."],"fun_headline_variants":["Aggregate accuracy masks subgroup disparities in facial recognition","Accuracy obscures demographic fairness gaps in law enforcement AI","Facial recognition accuracy fails to ensure fairness across groups","Overall accuracy masks fairness gaps in law enforcement facial recognition"],"cache_read_input_tokens":64,"weakest_assumption_plain":"That measuring and comparing false positive and false negative rates at the subgroup level supplies enough information to judge fairness without needing additional data on real-world consequences of those errors.","fun_headline_variants_meta":{"raw":{"variants":["Aggregate accuracy masks subgroup disparities in facial recognition","Accuracy obscures demographic fairness gaps in law enforcement AI","Facial recognition accuracy fails to ensure fairness across groups","Overall accuracy masks fairness gaps in law enforcement facial recognition"]},"model":"grok-4.3","cost_usd":0.011859,"raw_usage":{"total_tokens":5099,"prompt_tokens":657,"num_sources_used":0,"completion_tokens":60,"cost_in_usd_ticks":118590500,"prompt_tokens_details":{"text_tokens":657,"audio_tokens":0,"image_tokens":0,"cached_tokens":64},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":4382,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":657,"tokens_out":60,"duration_ms":60030,"temperature":1.0,"reasoning_tokens":4382,"cache_read_input_tokens":64,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-05-21T09:29:51.025130+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Deployment data from a law enforcement facial recognition system that tracks the demographic breakdown of actual false positives and false negatives and checks whether the groups with higher rates experience corresponding increases in negative outcomes like wrongful arrests.","supporting_citations":[],"review_version":1}