{"id":"99f6fb9b-57a2-4b1d-acac-8f6c73b82fda","arxiv_id":"2506.17490","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Open-source LLMs quote Black mortgage applicants higher interest rates than identical white applicants; a control-vector intervention reduces the gap by about a third on average, but the reduction is measured in-sample.","lead":"This paper tests whether five open-source AI language models give different mortgage terms to identical applicants who differ only by race. It finds systematic gaps and proposes a layer-level control-vector fix that shrinks the gaps by about a third on average, though the fix is tuned on the same data used to measure it.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"In-sample selection of the control-vector scale α* on the same 197 profiles used to report reductions means the headline up-to-70% mitigation is a training-score result, not an out-of-sample finding.","rationale":"The reader's weakest assumption identifies exactly the same load-bearing concern: α* is selected on the same 197 profiles used to report the mitigation reductions, so the headline reductions are in-sample. This is the most serious threat to the paper's central claim because the bias-detection results (counterfactual discrepancies across models and prompts) are robust and independent of the remediation, whereas the mitigation claim is the paper's main advertised contribution. The concern is concrete: the line search in Section 8 (Eq. 2) minimizes L(α, a, b) over the full set X, and Tables 7 and 8 report remediation MAE at that optimal α* on the same X. Because α = 0 is on the grid, the reported reductions are mechanically non-negative in-sample. The fix is straightforward—hold out profiles before selecting α* and report out-of-sample reductions—which is why the reader's CONDITIONAL verdict is appropriate. The secondary issue that 'without impairing overall model performance' is never measured is also real, but it is secondary because the disparity reduction itself would still be valuable even if the performance claim were dropped. I credit the paper for its reproducible, deterministic experimental design, open-source models, and clear counterfactual framework; these are real strengths. The concern does not invalidate the bias-detection findings, but it does mean the mitigation headline is not yet supported. A conditional acceptance requesting held-out validation and a performance metric is the right call, so no verdict change is needed.","tokens_in":22812,"tokens_out":4873,"duration_ms":53266,"concrete_test":"Split the 197 simulated profiles into a training set (e.g., 147) and a held-out set (50). Select α* by line search on the training set using the same grid [−0.2, 0.2] step 0.02, then compute the mean absolute error reduction between white and Black applicants on the held-out set for each model and prompt in Tables 7 and 8. If the held-out reduction is less than half of the reported in-sample reduction (or is negative for a majority of cells), the headline up-to-70% / 33%-average mitigation claim is not supported. Separately, report at least one performance metric for the remediated model, such as calibration of interest rates against a reference schedule or accuracy on a task with known labels, to substantiate the 'without impairing overall model performance' assertion.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central mitigation claim rests on α*, the control-vector scale chosen by line search in Section 8 (Eq. 2) over candidate values from −0.2 to 0.2 at 0.02 intervals. The objective L(α, a, b) is evaluated on X, the full set of 197 simulated applicant profiles, and the same set is used to report the 'Remediation MAE' in Tables 7 and 8. Because α = 0 is in the search grid, the minimized MAE is guaranteed to be no larger than the baseline MAE; any reported reduction is therefore an in-sample training improvement. The up-to-70% and 33%-average reductions are thus not independent estimates but artifacts of optimizing on the evaluation data. This circularity is confined to the scalar α* (the control vector itself is built from a separate contrasting-pair dataset), but α* is precisely the parameter that determines the magnitude of the headline effect. A related but distinct gap is the unsupported claim that remediation occurs 'without impairing overall model performance': no performance metric, ground-truth comparison, or output-quality check is reported anywhere in the paper, so even a validated disparity reduction would not establish that the intervention is cost-free. If α* does not generalize to new applicants or if the intervention degrades outputs in unmeasured ways, the abstract's central claims about effective and safe mitigation would not hold.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"Using deterministic counterfactual prompts on five open-source LLMs, the paper measures racial bias in simulated mortgage lending decisions, finding mean interest-rate gaps between Black- and white-labeled applicants that often exceed the 13.1 bp historical benchmark from Popick (2022). It then applies representation-engineering control vectors, constructed from contrasting-pair data, and reports reductions in the mean absolute discrepancy of up to 70% (around 33% on average). The authors also provide layer-wise analyses of where race is encoded and show that proxy indicators such as alma mater activate the same representation space as explicit race cues.","tokens_in":23195,"tokens_out":7230,"duration_ms":70102,"significance":"The bias-detection design is a genuine strength: counterfactual comparisons of deterministic models on identical hard inputs give a clean, reproducible measurement, and the layer-wise concept-intensity analysis offers a useful diagnostic tool. The control-vector object itself is constructed from an independent contrasting-pair dataset, so the detection and the steering direction are not circular. However, the paper's headline remediation claim is an in-sample statement: the scalar alpha* is tuned on the same 197 profiles used to report the reductions, and no metric of overall output quality is provided. These gaps are fixable with a holdout validation and performance checks, but they currently make the abstract's 'up to 70%' and 'without impairing performance' claims overstatements.","major_comments":[{"comment":"The optimal scale alpha* is selected by minimizing the objective L(alpha, a, b) over the same simulated set X of 197 profiles on which Baseline MAE and Remediation MAE are reported. Because the line search over [-0.2, 0.2] includes alpha = 0, the reported Remediation MAE is mechanically no larger than Baseline MAE, and the reductions (e.g., 14 to 3.9 bps for Mistral expanded-direct in Table 8) are in-sample fitted values, not predictions for new applicants. The up-to-70% and 33%-average reductions in the abstract are therefore not yet established out of sample. Please add a holdout split, cross-validation, or a separate calibration set, and report the reductions on applicant profiles not used to choose alpha*.","section":"Section 8, Eq. (2), Tables 7 and 8"},{"comment":"The claim that remediation occurs 'without impairing overall model performance' is unsupported. Tables 7 and 8 report only the race-gap MAE and discrepancy frequency; there is no measure of whether the steered model's interest-rate or approval outputs remain reasonable, calibrated, or accurate relative to any ground truth (e.g., consistency with the applicant's creditworthiness or with the unsteered model's overall rate distribution). Given the paper's own Appendix A.1 caution about steering-vector brittleness and side effects, a minimal performance audit (e.g., distribution of suggested rates, approval rates by credit score, or a task-accuracy check) is needed before making the cost-free claim.","section":"Section 8 and Abstract"},{"comment":"The headline 'reduces racial disparities' is based on the mean absolute discrepancy, but the frequency of discrepant pairs can move in the opposite direction. For Mistral v0.3 with the simple prompt and indirect race indicator, the mean discrepancy falls from 29.3 to 16.1 bps while the frequency of discrepant profiles rises from 58 to 72. The abstract's blanket claim should be qualified to mean-size reductions, and the paper should discuss disagreement between the two marginals.","section":"Section 8, Table 7"}],"minor_comments":[{"comment":"The right panel is labeled 'approval confidence' with values between -1 and 1, but the EBNF grammar restricts outputs to 'yes' or 'no'; the paper does not explain how confidence scores were computed. Please clarify.","section":"Section 5.1 and Figure 1"},{"comment":"The abstract's 'identical in all attributes except race' is only true for the direct-race experiments; in the proxy experiment the prompts also differ in university name. Please qualify.","section":"Abstract and Section 5.3"},{"comment":"The PDP-GP intervals for the indirect race indicator appear to overlap, and the text says the differences are not statistically significant, yet the section concludes with counterfactual evidence of disparity; please state the inferential status of the PDP-GP comparisons more precisely or add a formal test.","section":"Section 6, Figures 3 and 4"},{"comment":"Representation engineering is credited to Zhuo et al. (2023) in the text, but the cited Zhuo et al. paper is about ChatGPT ethics, not representation engineering; the technique as described matches Zou et al. (2023) (arXiv:2310.01405). Please correct the citation.","section":"Section 7 and References"},{"comment":"The notation in Eq. (1) is overloaded: G maps an input vector, but X_a is a set of profiles; please define what G(X_a) returns (e.g., a vector of outputs over the subset) and clarify the L1 norm over profiles.","section":"Section 8, Eq. (1)"},{"comment":"The text reports a maximum reduction of 72%, while the abstract says 'up to 70%'; please reconcile the numbers and state how the 33%-average figure is computed from Tables 7 and 8.","section":"Section 9 vs Abstract"}],"recommendation":"major_revision","confidential_remarks":"The paper's equations make the in-sample nature of alpha* selection transparent rather than hidden, and the authors are up-front about several limitations. The main concern is that the abstract overstates the remediation result. With a holdout validation and a performance audit, this could become a solid contribution; I recommend major revision rather than rejection. There is also a citation discrepancy (Zhuo vs Zou) that the editors may want checked."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nShort version: the bias measurement is real and worth taking seriously; the headline mitigation number is not, at least not as reported. The 70% reduction is a training score, not an out-of-sample result.\n\nWhat's new here is the application of representation engineering to credit decisions, plus the layer-wise evidence that a proxy like alma mater activates the same internal race representation as an explicit race label. The deterministic counterfactual design across five open-source models is clean: EBNF-constrained outputs, temperature zero, simulated profiles matched to the empirical mortgage distribution, and an economic benchmark from Popick (2022). The bias-detection part is solid and reproducible.\n\nThe soft spot is exactly where the stress-test lands. In Section 8, alpha* is chosen by line search over the same 197 profiles used to report the Remediation MAE in Tables 7 and 8. Because alpha = 0 is in the grid, the optimized MAE is guaranteed to be no larger than the baseline. So the up-to-70% and 33%-average reductions are fitted values, not predictions about new applicants. This does not invalidate the detection result, but it means the mitigation claim needs a held-out set or cross-validation before it can be believed.\n\nRelated gap: the abstract says 'without impairing overall model performance,' but no performance metric is reported anywhere. In a counterfactual experiment there's no ground-truth label, but you can check that the control vector doesn't distort other attributes of the output — e.g., that the recommended rate for white applicants remains sensible, or that the ordering of risk by credit score is preserved. That's missing.\n\nMinor issues: alpha* sits at the boundary of the search grid for several models (0.2 or -0.2), which suggests the optimum may lie outside the range; the simulated profiles come from a multivariate normal fitted to confidential data, and there's no sensitivity analysis on those parameters.\n\nBottom line: this paper is for people auditing LLM fairness in finance. The bias-measurement toolkit is a genuine contribution. The remediation claim needs a re-analysis with a proper train/test split, plus some output-quality check, before it can be taken at face value.\n\nI'd send it to peer review — a serious referee can require the out-of-sample evaluation, and the detection half deserves to be in the literature. I would not cite the 70% number, but I would cite the framework.\n\nBest.","headline":"Solid, reproducible bias measurement; the headline mitigation effect is an in-sample fitted value, and the 'no performance loss' claim is unmeasured.","tokens_in":23583,"tokens_out":3159,"would_cite":true,"duration_ms":32088,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Race-labelled loan applicants get worse LLM terms even with identical finances — and steering the model's insides removes about a third of the gap.","keywords":["large language models","mortgage lending","racial bias","proxy discrimination","representation engineering","control vectors","counterfactual testing","AI in finance"],"falsifier":"Hold out a fresh sample of simulated applicant profiles generated from the same empirical distribution, pick $\\alpha^*$ using only the original 197 profiles, then measure the between-group rate gap on the held-out profiles; the central claim is falsified if the average gap reduction falls near zero or reverses sign on that out-of-sample data. A second check would test whether the gap survives when the racial cue is given in a form never seen in the contrasting-pair dataset, such as a dialect or neighborhood name.","tokens_in":22584,"feed_emoji":"🏦","tokens_out":7175,"duration_ms":70860,"temperature":0.7,"pith_summary":"This paper tries to establish that open-source large language models, when asked to act as mortgage loan officers, give systematically worse terms to otherwise identical applicants described as Black, and that the gap is bigger than the comparable gap in human lending. The authors build a counterfactual test bench: the same simulated applicant profile is presented to the model with only the racial cue changed, so any difference in the quoted rate or approval decision must be caused by that cue. They find the disparity persists when race is not stated outright but is inferred from a proxy such as the applicant's university. They then show the bias is not just in the output: layer-by-layer activation readings locate a 'race direction' inside the model, and subtracting a scaled version of that direction from the hidden states reduces the average interest-rate gap by about a third, up to 70 percent in the strongest cases, without hurting overall model performance. The implied stakes are that lenders cannot audit or fix algorithmic discrimination by looking only at prompts and outputs.","feed_headline":"Control vector cuts LLM racial loan gap by a third","feed_subtitle":"Counterfactual audits on five open models find race-based rate gaps above human history; an inner steering fix removes a third.","key_machinery":"The central object is a control vector, built by representation engineering. For a set of contrasting sentence pairs that differ only in 'Black' versus 'white', the model's hidden activations are collected at every transformer block; the per-layer differences are stacked, and PCA's first principal component gives a single direction in activation space that tracks the race concept. Scaling that vector by a coefficient $\\alpha$ and adding it to the hidden states during generation steers the model away from racial coding, with the optimal $\\alpha^*$ chosen by line search over $[-0.2,0.2]$ to minimize the between-group discrepancy objective $L(\\alpha,a,b) = |G(X_a)-G'(X_b|\\alpha)|_1$. Concept-intensity scores from the same machinery locate where race is encoded and validate that proxy inputs such as alma mater trigger the same layers. This machinery does the double duty of diagnosing bias and remediating it, because it operates on internal representations rather than on prompts or output filters.","core_discovery":"Across five locally run open-source LLMs, counterfactual applicants identical in income, credit score, loan amount, LTV, DTI, and age receive different loan recommendations when one word in the prompt identifies them as Black rather than white; the typical quoted-rate gap exceeds the roughly 13 basis points observed in historical human mortgage data. The disparity is largest near credit-score thresholds and at low scores, shrinks but does not disappear when a full financial profile is supplied, and reappears when race is carried by a proxy such as the applicant's alma mater. Representation-vector heatmaps show the models encode the race cue in early layers for the interest-rate task and in final layers for approval, and that the proxy activates the same internal regions as the explicit label. Injecting a control vector — the first principal component of activation differences between matched Black and white sentence pairs, scaled by a tuned coefficient — into the model's hidden states reduces the mean absolute rate discrepancy to roughly a third of baseline, with the largest single reduction reaching about 70 percent, while leaving task performance effectively intact.","pith_inferences":["My inference: the same control-vector recipe should transfer to other protected attributes such as gender, age, or religion, but the paper tests only race in a credit task, so that transfer is unverified.","My inference: the 33 percent average reduction is probably optimistic until $\\alpha^*$ is validated on held-out profiles, because the same profiles both choose the tuning parameter and report the improvement.","My inference: the layer-localization heatmaps suggest a cheap monitoring tool — track concept-intensity scores on live prompts — but the authors do not propose a deployment monitor, so this is an extension, not their claim.","My inference: if the observed gap shrinks when more financial variables are supplied, part of the mechanism behaves like statistical discrimination, but distinguishing that from taste-based bias would require an experiment where credit risk varies independently of race; the paper does not attempt that."],"forward_implications":["Prompt-level instructions such as 'do not discriminate' do not remove the bias; it persists inside model representations, so output-only audits can miss it.","Removing explicit race fields is not enough, because proxy inputs like a university name activate the same internal race-coding layers.","A control vector can be constructed quickly from small contrast sets and applied without retraining or changing model weights, which makes representation-level remediation practical for local deployments.","If lenders deploy LLMs at scale, the route to fair-lending compliance runs through internal-activation checks rather than output filters alone."],"supporting_citations":[{"why":"This citation supplies the historical human mortgage-lending interest-rate gap the paper compares LLM discrepancies against.","marker":"Popick (2022)"},{"why":"This citation provides the other human-market racial disparity baseline used to judge the economic magnitude of the LLM gaps.","marker":"Hurtado and Sakong (2024)"},{"why":"This citation supplies the representation-engineering method for building control vectors from activation differences.","marker":"Zou et al. (2023)"},{"why":"This citation inspires the counterfactual correspondence-study design and the use of race proxies in applicant profiles.","marker":"Bertrand and Mullainathan (2004)"},{"why":"This citation anchors the in-group versus out-group loan-officer experiment comparing officer race and applicant outcomes.","marker":"Frame et al. (2024)"},{"why":"This citation supplies the deterministic counterfactual inference setup and the Gaussian-process partial-dependence approximation used for interpretation.","marker":"Cook et al. (2023)"},{"why":"This citation gives the fintech-era lending discrimination benchmark that motivates the question of algorithmic bias.","marker":"Bartlett et al. (2022)"},{"why":"This citation supports the claim that AI systems use benign features as proxies for protected attributes.","marker":"Prince and Schwarcz (2019)"}],"fun_headline_variants":["LLM loan bias exceeds human history; steering fix cuts it 70%","Inner vector halts AI racial bias in mortgage rates","Counterfactual audit exposes LLM race gap; control vector shrinks it","Race bias in AI lending tamed by 70% with control-vector","LLMs show racial loan gaps above humans; vector fix trims 70%"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the single control-vector scale chosen by searching over the same 197 simulated profiles on which the reductions are reported will keep working for new applicants; if that optimum is an in-sample fit, the headline 33 percent average reduction does not transfer.","fun_headline_variants_meta":{"raw":{"variants":["LLM loan bias exceeds human history; steering fix cuts it 70%","Inner vector halts AI racial bias in mortgage rates","Counterfactual audit exposes LLM race gap; control vector shrinks it","Race bias in AI lending tamed by 70% with control-vector","LLMs show racial loan gaps above humans; vector fix trims 70%"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000551,"raw_usage":{"total_tokens":2611,"prompt_tokens":907,"completion_tokens":1704,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":523,"completion_tokens_details":{"reasoning_tokens":1607}},"tokens_in":523,"tokens_out":1704,"duration_ms":12477,"temperature":1.0,"reasoning_tokens":1607,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T19:07:20.483345+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Hold out a fresh sample of simulated applicant profiles generated from the same empirical distribution, pick $\\alpha^*$ using only the original 197 profiles, then measure the between-group rate gap on the held-out profiles; the central claim is falsified if the average gap reduction falls near zero or reverses sign on that out-of-sample data. A second check would test whether the gap survives when the racial cue is given in a form never seen in the contrasting-pair dataset, such as a dialect or neighborhood name.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"This citation supplies the historical human mortgage-lending interest-rate gap the paper compares LLM discrepancies against."},{"cited_title":"and Sakong, J","cited_arxiv_id":null,"evidence_quote":"This citation provides the other human-market racial disparity baseline used to judge the economic magnitude of the LLM gaps."},{"cited_title":"and Mullainathan, S","cited_arxiv_id":null,"evidence_quote":"This citation inspires the counterfactual correspondence-study design and the use of race proxies in applicant profiles."},{"cited_title":"S., Huang, R., Jiang, E","cited_arxiv_id":null,"evidence_quote":"This citation anchors the in-group versus out-group loan-officer experiment comparing officer race and applicant outcomes."},{"cited_title":"R., Kazinnik, S., Hansen, A","cited_arxiv_id":null,"evidence_quote":"This citation supplies the deterministic counterfactual inference setup and the Gaussian-process partial-dependence approximation used for interpretation."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"This citation gives the fintech-era lending discrimination benchmark that motivates the question of algorithmic bias."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"This citation supports the claim that AI systems use benign features as proxies for protected attributes."}],"review_version":2}