{"id":"e2a2a262-b0da-4d78-b058-b7071b85b597","arxiv_id":"2604.15038","paper_version":2,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Different fairness metrics frequently disagree on model bias levels, quantified via a new Fairness Disagreement Index that remains high across thresholds and configurations in face recognition experiments.","lead":"This paper shows that common fairness metrics in machine learning often produce conflicting results about whether a model exhibits demographic bias, demonstrated through experiments on face recognition systems. A smart generalist should read it because single-metric fairness reports may be unreliable for high-stakes decisions in areas like biometrics and healthcare.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"Generalizability of disagreement findings from face recognition to broader ML fairness practices","rationale":"The reader's weakest_assumption directly identifies the same representativeness gap. Full-text experiments appear confined to the face-recognition setting described in the abstract, so the UNVERDICTED verdict and low confidence are still warranted; no new cross-domain controls or sensitivity analyses are indicated in the provided summary.","tokens_in":1751,"tokens_out":301,"duration_ms":21189,"concrete_test":"Re-run the exact multi-metric pipeline and FDI computation on a tabular fairness benchmark (e.g., Adult or COMPAS dataset) using the same metric collection and group partitions; if average FDI drops below 0.3 or sign disagreements fall below 20 % of cases, the general claim about evaluation reliability weakens.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that fairness metrics produce contradictory conclusions about model bias and that single-metric reporting is therefore insufficient—rests on treating the face-recognition experiments (standard demographic partitions, fixed set of error-rate and performance metrics) as representative. Face recognition involves high-dimensional continuous inputs, specific group definitions (e.g., race/gender bins), and performance regimes that differ from tabular or low-dimensional settings common in healthcare or lending. Without evidence that FDI magnitudes and contradiction rates transfer, the leap from “in this controlled setting metrics disagree” to “current fairness evaluation practices have a critical limitation” remains unsupported.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper claims that fairness metrics in ML often produce contradictory assessments of demographic bias, as shown via systematic experiments on face recognition models across group partitions and a fixed set of error-rate and performance metrics. It introduces the Fairness Disagreement Index (FDI) to quantify inconsistency, reports that disagreement remains high across thresholds and configurations, and concludes that single-metric reporting is insufficient for reliable bias assessment.","tokens_in":1843,"tokens_out":459,"duration_ms":29302,"significance":"If the central empirical observation holds and generalizes, the work is significant for providing a quantitative tool (FDI) to diagnose a practical limitation in fairness evaluation and for using a controlled face-recognition setting to isolate metric disagreement. This could encourage multi-metric reporting standards in high-stakes applications.","major_comments":[{"comment":"§4.2 (Experimental Setup) and Table 1: The FDI values and disagreement rates are reported without sample sizes, standard errors, or any statistical test for whether observed disagreements exceed chance levels under the null of metric consistency; this is load-bearing for the claim that disagreement 'remains high' and leads to 'contradictory conclusions.'","section":"§4.2 and Table 1"},{"comment":"§5 (Discussion and Conclusions): The assertion that the results highlight 'a critical limitation in current fairness evaluation practices' and that single-metric reporting is insufficient rests on treating the face-recognition setting (high-dimensional inputs, race/gender bins, error-rate metrics) as representative, yet no transfer experiments or sensitivity analysis to tabular/low-dimensional domains (e.g., lending or healthcare) are provided.","section":"§5"}],"minor_comments":[{"comment":"Abstract: The acronym FDI is used before its expansion; expand on first use.","section":"Abstract"},{"comment":"§3.1: The FDI formula is introduced without an explicit statement of the aggregation function over metric pairs or the normalization range; add a short derivation or pseudocode for reproducibility.","section":"§3.1"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We are grateful to the referee for their insightful review. The comments have prompted us to strengthen the statistical rigor and clarify the scope of our findings. We respond to each major comment in turn.","responses":[{"response":"We concur that the absence of sample sizes, standard errors, and statistical tests weakens the evidential basis for the claim that disagreement remains high. We will revise §4.2 and Table 1 to include the number of trials or partitions, bootstrap-derived standard errors for the FDI and disagreement rates, and a statistical test (e.g., a binomial test or permutation test under the null that metrics are consistent) to evaluate if the observed rates exceed chance. These additions will be incorporated in the next version of the manuscript.","revision_made":"yes","referee_comment":"[§4.2 and Table 1] §4.2 (Experimental Setup) and Table 1: The FDI values and disagreement rates are reported without sample sizes, standard errors, or any statistical test for whether observed disagreements exceed chance levels under the null of metric consistency; this is load-bearing for the claim that disagreement 'remains high' and leads to 'contradictory conclusions.'"},{"response":"The face recognition setting was deliberately chosen to provide a controlled environment with high-dimensional data and explicit demographic partitions, allowing us to focus on metric disagreement without confounding factors from other domains. Nevertheless, we accept that demonstrating generalizability would bolster the broader conclusions. In the revised Discussion and Conclusions (§5), we will include additional text on the potential applicability to other domains such as tabular data in lending or healthcare, along with a qualitative sensitivity analysis based on the metric properties. We will also qualify the claims to note that while the limitation is illustrated in this setting, further empirical validation across domains is warranted. Full transfer experiments are not included in this revision due to scope, but we will highlight this as a direction for future work.","revision_made":"partial","referee_comment":"[§5] §5 (Discussion and Conclusions): The assertion that the results highlight 'a critical limitation in current fairness evaluation practices' and that single-metric reporting is insufficient rests on treating the face-recognition setting (high-dimensional inputs, race/gender bins, error-rate metrics) as representative, yet no transfer experiments or sensitivity analysis to tabular/low-dimensional domains (e.g., lending or healthcare) are provided."}],"tokens_in":1358,"tokens_out":514,"duration_ms":41898,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main takeaway is that this paper shows common fairness metrics can point in opposite directions on the same face recognition model, and the Fairness Disagreement Index is their way of putting a number on how often that happens. The experiments use standard demographic splits and a mix of error-rate and performance metrics, and the disagreement holds across thresholds and setups. That part is straightforward and worth noting for anyone who audits biometric systems.","headline":"Fairness metrics disagree in these face recognition runs and the FDI quantifies it, but the jump to a general flaw in ML fairness practices rests on narrow evidence.","tokens_in":2317,"tokens_out":160,"would_cite":false,"duration_ms":45297,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":{"model":"grok-4.3","evidence":[{"relation":"unclear","rs_module":"IndisputableMonolith/Cost/FunctionalEquation.lean","rs_theorem":"washburn_uniqueness_aczel","paper_passage":"We introduce the Fairness Disagreement Index (FDI)... pairwise metric disagreement... ranking disagreement... FDI = 1/N² Σ [α D_ij + (1-α) R_ij]"},{"relation":"unclear","rs_module":"IndisputableMonolith/Foundation/RealityFromDistinction.lean","rs_theorem":"reality_from_one_distinction","paper_passage":"Using face recognition as a controlled experimental setting... LFW proxy groups... FaceNet/ArcFace embeddings"}],"headline":"ML fairness-metric disagreement study (FDI) has no structural overlap with RS","alignment":"orthogonal","rationale":"The paper's central contribution is an empirical index (FDI) quantifying inconsistency among standard ML fairness metrics (FPR/FNR/accuracy disparities, Wasserstein) on face-verification proxies. This lives entirely in the domain of statistical evaluation practices for supervised models; RS contains no theorems about fairness definitions, metric trade-offs, or disagreement indices. No J-cost, φ-ladder, 8-tick periodicity, or parameter-free constant derivation appears. Hence orthogonal per rubric.","tokens_in":44987,"confidence":"high","tokens_out":308,"duration_ms":14092,"cache_read_input_tokens":38528,"cache_creation_input_tokens":0},"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Fairness assessments can vary significantly depending on the choice of metrics, leading to contradictory conclusions about model bias.","keywords":["fairness metrics","demographic fairness","bias assessment","machine learning","fairness disagreement index","face recognition","model bias"],"falsifier":"Conducting the same analysis on a healthcare or risk assessment model and finding that all fairness metrics agree on the presence or absence of bias would contradict the main claim.","tokens_in":2625,"feed_emoji":"⚖️","tokens_out":573,"duration_ms":64284,"temperature":0.7,"pith_summary":"The paper investigates the consistency of fairness evaluations in machine learning by applying multiple common fairness metrics to models in a face recognition setting. It finds that these metrics often produce conflicting results on whether demographic bias is present. To address this, the authors develop the Fairness Disagreement Index to measure how much the metrics disagree. This is important because high-stakes applications depend on these assessments to decide if models are safe to use, but inconsistent signals undermine confidence in the process.","feed_headline":"Fairness metrics disagree on bias in the same model","feed_subtitle":"Face recognition tests show metrics can lead to opposite conclusions, so relying on one is risky for fairness checks.","key_machinery":"The Fairness Disagreement Index (FDI) that captures the degree of inconsistency across fairness metrics applied to the same system.","core_discovery":"Using face recognition as a controlled experimental setting, we evaluate model performance across multiple group partitions under a range of commonly used fairness metrics. Our results demonstrate that fairness assessments can vary significantly depending on the choice of metrics, leading to contradictory conclusions regarding model bias. To quantify this phenomenon, we introduce the Fairness Disagreement Index (FDI), a measure designed to capture the degree of inconsistency across fairness metrics. We further show that disagreement remains high across thresholds and model configurations.","pith_inferences":["Teams deploying ML systems might benefit from always reporting a set of fairness metrics along with their disagreement score.","This inconsistency could be tested in other areas like credit lending or job screening to see if it is widespread.","Methods to choose or combine metrics when they disagree could be developed as a follow-up."],"forward_implications":["Single-metric reporting is insufficient for reliable bias assessment.","Model bias conclusions can reverse based on which fairness metric is selected.","Disagreement between metrics stays high no matter the threshold or model configuration used."],"fun_headline_variants":["Metrics conflict on bias in the same ML model","Different metrics yield opposing bias conclusions","Fairness assessments flip with metric selection","Multi-metric analysis reveals assessment inconsistencies"],"cache_read_input_tokens":64,"weakest_assumption_plain":"The patterns of metric disagreement found in face recognition tasks with standard demographic splits hold for fairness assessment in machine learning more broadly.","fun_headline_variants_meta":{"raw":{"variants":["Metrics conflict on bias in the same ML model","Different metrics yield opposing bias conclusions","Fairness assessments flip with metric selection","Multi-metric analysis reveals assessment inconsistencies"]},"model":"grok-4.3","cost_usd":0.0069,"raw_usage":{"total_tokens":3129,"prompt_tokens":684,"num_sources_used":0,"completion_tokens":50,"cost_in_usd_ticks":69003000,"prompt_tokens_details":{"text_tokens":684,"audio_tokens":0,"image_tokens":0,"cached_tokens":64},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":2395,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":684,"tokens_out":50,"duration_ms":28722,"temperature":1.0,"reasoning_tokens":2395,"cache_read_input_tokens":64,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-05-21T09:22:31.704038+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Conducting the same analysis on a healthcare or risk assessment model and finding that all fairness metrics agree on the presence or absence of bias would contradict the main claim.","supporting_citations":[],"review_version":2}