{"id":"e8ac750a-d167-4d64-b1b2-42852332e24a","arxiv_id":"2506.16560","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"An external black-box audit of Meta's VRS finds it reduces delivery variance but lowers unique reach and raises cost per person reached, and equal budget splitting outperforms VRS on both.","lead":"This paper tests Meta's Variance Reduction System, which was created after a 2022 settlement with the US Department of Justice to reduce discrimination in housing ad delivery. It finds that VRS lowers demographic variance but also reduces how many unique people see opportunity ads and raises advertiser costs, and that a simple budget-splitting approach performs better on both.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Reach/cost findings rest on an unvalidated paired design: VRS and no-VRS ads run on disjoint audience partitions, so the 9.82% reach drop and 12.02% CPP rise may reflect audience or competition differences rather than VRS itself.","rationale":"The paper makes a valuable external audit contribution: black-box paired ads, public data release, and a concrete alternative strategy. The strongest claim, however, is that VRS reduces opportunity exposure and raises advertiser cost. That claim requires the reach/CPP differences to be attributable to VRS, not to the disjoint audience partitions used in the paired design. The paper itself flags this confound in §4.1.3 but does not resolve it with uncertainty quantification or a crossover design. The reader's weakest assumption was the self-estimated eligible ratio in §4.1.5; that is a legitimate concern for the variance-reduction claim, but the cost/reach finding is more central to the paper's policy conclusion and is independently vulnerable. I therefore partially agree with the reader: the eligible-ratio issue is real, yet the audience-partition confound is the more load-bearing problem for the strongest claim. Because the reader already returned CONDITIONAL, my read does not change the verdict; it strengthens the conditionality. A clean crossover or mixed-effects reanalysis of the released data would settle whether the concern lands.","tokens_in":23566,"tokens_out":6960,"duration_ms":75075,"concrete_test":"Re-analyze the released dataset using a paired mixed-effects model: reach or CPP ~ VRS + (1|ad_creative) + (1|audience_partition), and report bootstrap confidence intervals for the VRS coefficient; if the 95% CI for the reach difference includes zero, or for the CPP difference includes zero, the headline claim is not established. Additionally, run a crossover experiment for the hair-product and education ads: on each of several audience partitions, run no-VRS on partition A and VRS on partition B in week 1, then swap treatments in week 2; if the direction and size of the reach/CPP gap persist under swapped assignment, the confound is ruled out.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central empirical claim that VRS reduces utility for users and advertisers (§4.2.3) depends on comparing the same ad with and without VRS. But the paper deliberately runs the two versions on separate audience partitions to avoid self-competition (§4.1.3). Consequently, each point in Figure 5 is not a controlled pair: the no-VRS counterfactual is measured on a different set of users with different competition from other advertisers, different match rates, and different activity levels. The authors acknowledge this potential confound but do not quantify it. They report only mean changes (-9.82% reach, +12.02% CPP) across 36 experiments, with no standard errors, confidence intervals, or a mixed-effects model that would partition variance attributable to audience versus VRS. The same design carries into §5: the budget-splitting arm runs four separate campaigns, so the VRS-versus-splitting comparison is also not a same-audience controlled contrast. If the reach/cost gap is driven by the specific audience partitions assigned to VRS, the paper's policy conclusion that VRS 'passes the cost of decreasing variance to advertisers' would not be supported. The separate eligible-ratio estimation issue in §4.1.5 is real but secondary here: even with a correct baseline, the cost/reach result still needs a valid counterfactual.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper analyzes the 2022 Meta/DoJ settlement and Meta's Variance Reduction System (VRS) for housing ad delivery. It makes three contributions: (1) a policy analysis arguing that the settlement's impression-based variance metric, coverage rule, and lack of a non-leveling-down constraint permit implementations that do not increase individual access to opportunity; (2) a black-box experimental evaluation comparing otherwise identical ads run with and without VRS, reporting that VRS reduces demographic variance but also reduces mean reach by 9.82% and raises cost per 1,000 reach (CPP) by 12.02%, thereby passing compliance costs to advertisers; and (3) a comparison with a budget-splitting alternative that, in the authors' experiments, increases reach for all demographic groups and lowers CPP relative to VRS. The authors also examine VRS in employment and credit domains, finding weaker variance reduction than in housing. Data from the experiments are made public.","tokens_in":23851,"tokens_out":3683,"duration_ms":38805,"significance":"If the empirical findings are correct, this is the first fully independent, advertiser-accessible external audit of VRS, and it provides timely evidence for regulators and civil-rights stakeholders about how settlement metrics interact with real delivery outcomes. The paper's policy analysis of the settlement terms is thoughtful and largely independent of the experiments. The black-box methodology and public data release are valuable for reproducibility and for future audits of ad delivery systems. However, the central empirical claims about VRS's effect on reach and cost currently rest on a confounded paired design and on a circular baseline for the variance evaluation, and the reported aggregate effects lack confidence intervals or significance tests. These issues are load-bearing for the paper's main conclusions, so the manuscript needs substantial revision before the empirical claims can be considered established.","major_comments":[{"comment":"The paired design runs VRS-enabled and no-VRS ads on disjoint audience partitions to avoid self-competition, so each point in Figure 5 compares outcomes measured on different sets of users. Differences in match rates, auction competition, and user activity between partitions could produce the reported -9.82% reach and +12.02% CPP changes even if VRS had no effect. The paper acknowledges this confound but does not quantify it. Please provide per-pair differences with standard errors or confidence intervals, and ideally a mixed-effects model that includes audience partition as a random effect or a same-audience crossed design. Without such analysis, the central claim that VRS passes compliance costs to advertisers is not supported by the data as reported.","section":"§4.1.3, §4.2.3"},{"comment":"The eligible-ratio baseline used to evaluate VRS is estimated as the mean of the delivery ratios of the VRS-enabled ads themselves, justified by assuming that VRS works and by the law of large numbers. The subsequent claim that VRS reduces variance is therefore partly circular: the baseline is derived from the very system under evaluation. The post-matching audience-size data described in Appendix D provide a more independent basis for the eligible ratio, but the main text does not use them. Please re-estimate eligibility using the audience-match API fractions, or at minimum perform a sensitivity analysis over a plausible range of eligible ratios, and report whether the variance-reduction conclusion survives.","section":"§4.1.5, §4.2.1"},{"comment":"The experiments trigger VRS by declaring non-housing ads as housing ads, and the six creatives include hair products and golfing, which are not economic-opportunity ads. The global reach and CPP results in §4.2.3 are averaged over all 36 experiments, including these non-opportunity creatives, but the paper's policy conclusion is about 'fewer exposures to opportunities for individuals.' The effect of VRS on reach and cost may differ for education, insurance, and financial ads, which are the actual opportunity categories. Please report the reach and CPP results separately for the opportunity ads (EA, EB, IA, FA) and for the non-opportunity ads, and discuss whether the aggregate conclusion holds for opportunity ads alone.","section":"§4.1.1, §4.2.3"},{"comment":"The budget-splitting comparison in Figure 6 also uses separate campaigns on separate audience partitions for the VRS arm and the split arm, with no confidence intervals or significance tests. The claim that budget-splitting 'outperforms VRS' by increasing reach for all groups and reducing CPP is therefore subject to the same audience-confounding concern as the main paired comparison. Please provide per-pair comparisons with uncertainty quantification and, if possible, run the VRS and split arms on the same or crossed audience partitions.","section":"§5.2"}],"minor_comments":[{"comment":"The first sentence contains a typo: 'resulted is a first-of-its-kind change' should be 'resulted in a first-of-its-kind change.'","section":"Abstract"},{"comment":"The name 'Pesysakhovich et al.' should be 'Peysakhovich et al.' to match reference [59].","section":"§2.4"},{"comment":"The labels 'vrs-credit1', 'vrs-credit2', etc. on the y-axis of Figure 4a are not defined in the caption or text; please explain what the axis represents and define 'Declared as.'","section":"Figure 4"},{"comment":"The scatterplots show individual points but no indication of variability across the three audience replications; adding paired differences with error bars or a small-multiples per-creative breakdown would make the aggregate claims easier to interpret.","section":"Figures 5 and 6"},{"comment":"The decision to omit 'flipped' audiences to save costs weakens the location-based race proxy; prior work cited by the authors used flipped audiences as a control. At minimum, please state explicitly that DMA-specific confounds remain possible and discuss how this could affect the race-specific results.","section":"Appendix B"},{"comment":"The errata note indicates that the published FAccT version contained a miscalculation for the Male demographic attribute in Figures 5b and 5c. Please confirm that all numbers in the text, including the 9.82% and 12.02% aggregates, correspond to the corrected figures.","section":"§4.2.3 and Errata"}],"recommendation":"major_revision","confidential_remarks":"The paper is likely to be of interest to FAccT readers, and the policy analysis plus public dataset are genuine contributions. My main reservation is that the empirical core—the reach/cost tradeoff and the budget-splitting comparison—is presented without the statistical rigor needed to support the strong causal language ('VRS passes the cost...'). The circular eligible-ratio estimation compounds this. I would not reject the paper, because the issues are addressable with additional analysis and re-estimation, but the revision needs to be substantive rather than cosmetic. I also note some ethical/terms-of-service risk in deliberately mislabeling non-housing ads as housing ads; the authors acknowledge the practice, but the reviewer may want to consider whether this is acceptable to highlight as a method for other auditors."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's the short version: this paper is the first independent black-box evaluation of Meta's Variance Reduction System, and it's worth reading for the settlement-terms analysis alone. The empirical reach/cost findings are plausible and consistent across 36 runs, but the experimental design leaves a real confound unquantified, and the variance-reduction claim leans on a self-estimated baseline.\n\nWhat's actually new: the authors toggle VRS on and off by declaring non-housing ads as housing, which is a clever low-cost trick that lets external auditors isolate a deployed mitigation system. They are the first to measure VRS's effect on unique reach and advertiser cost in the wild. The settlement critique in Section 3 is sharp and stands on its own: impressions-versus-reach, the coverage slack that lets Meta exclude the largest campaigns, and the leveling-down risk are each argued clearly and quantified with Ad Library data. The budget-splitting comparison in Section 5 is also genuinely useful as an existence proof that the cost increase is not inevitable.\n\nThe soft spots are real but proportionate. The paired design runs VRS and no-VRS copies on disjoint audience partitions to avoid self-competition; that means each pair is not a controlled contrast, since audiences differ in competition, match rates, and activity. The authors acknowledge this in Section 4.1.3, but they do not quantify it, and there are no confidence intervals or significance tests on the -9.82% reach and +12.02% CPP figures. The same design issue carries into the budget-splitting comparison, which runs four separate campaigns. On top of that, the eligible-ratio baseline used for the variance-reduction claim is not measured from Meta's data; it is the mean delivery ratio of the VRS-enabled ads themselves. That creates a circularity problem for the specific claim that VRS reduces variance against the true eligible ratio. The authors are transparent about the limitation, but it is the weakest link in the empirical chain.\n\nNone of this sinks the paper. The settlement-terms analysis does not depend on the experiments at all, and the reach/cost direction is consistent enough across many independent runs that I would be surprised if it flipped entirely. The citation pattern is honest, with proper credit to prior paired-ad audits and to Gelauff et al. for budget splitting. No invented entities, no hidden parameters beyond the per-group eligible ratio estimates.\n\nWho is this for: regulators, platform auditors, and fairness researchers who want to know whether a landmark settlement's chosen metric actually helps people. It deserves a serious referee. I would send it to review, and I would ask for uncertainty quantification, a robustness check on the eligible-ratio baseline, and either a same-audience cohort or a formal model of the audience-partition confound before treating the cost/reach numbers as settled.","headline":"First independent black-box audit of Meta's VRS; settlement-terms critique is the strongest part, while the reach/cost findings are plausible but rest on an imperfectly controlled design.","tokens_in":24360,"tokens_out":1642,"would_cite":true,"duration_ms":18172,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that Meta's Variance Reduction System satisfies the settlement's impression-variance metric but reaches fewer unique users and raises advertiser cost per person reached, and that a simple budget-splitting alternative…","keywords":["algorithmic discrimination","ad delivery","variance reduction system","fairness metrics","external audit","settlement compliance","black-box methodology","fairness-utility trade-off"],"falsifier":"Run the same paired VRS/no-VRS campaigns but compute variance against an independently obtained eligible ratio, for example the demographic impression distribution across all advertisers in the same geographic areas from public ad-library data or a compliance report, and check whether VRS still reduces variance below the 10% threshold; if it does not, the paper's central empirical claim fails. A second decisive test would compare reach and cost under VRS with reach and cost under budget-splitting at larger budgets and longer durations, since the paper's 24-hour, $20 campaigns may not reflect steady-state delivery.","tokens_in":23395,"feed_emoji":"📉","tokens_out":6001,"duration_ms":57047,"temperature":0.7,"pith_summary":"This paper asks whether the 2022 settlement between Meta and the U.S. Department of Justice actually improves access to housing and other opportunity ads, and whether the Variance Reduction System (VRS) Meta built to comply is a good way to do it. The authors argue the settlement's chosen metrics are too weak: they count impressions rather than unique individuals, treat small and large ads equally in coverage targets, and permit leveling down, reducing everyone's exposure to hit variance targets. Using paired real-world ad campaigns run by ordinary advertisers, they find VRS does reduce variance by the settlement's own formula, but mean reach falls by about 10% and cost per 1,000 people reached rises by about 12%, so for a fixed budget fewer unique users see opportunity ads. They then test a transparent alternative, splitting a campaign's budget equally across demographic subgroups, and find it increases reach for all groups while lowering cost, concluding the utility loss of VRS is an implementation choice, not an inevitability.","feed_headline":"Meta's anti-bias ad system reaches fewer people, costs more","feed_subtitle":"Paired real-ad experiments show the settlement's metric passes costs to advertisers; an equal-budget split does better.","key_machinery":"The load-bearing object is the settlement's variance metric, defined as half the summed absolute difference between a demographic group's eligible ratio and its delivery ratio, and the VRS bid multiplier that Meta's machine-learning module applies to steer delivery toward the eligible ratio. The paper also relies on its black-box paired-campaign design: the same ad creative and budget run once with VRS enabled and once without, on disjoint balanced custom audiences, so any difference in delivery is attributable to VRS. Its eligible-ratio baseline is estimated externally as the mean delivery ratio across all VRS-enabled ads, justified by the law of large numbers under the assumption that VRS works.","core_discovery":"On the paper's own terms, the central discovery is that the settlement's compliance framework and Meta's implementation of it decouple 'fairness' as measured by impression variance from access to opportunities as experienced by individuals. The paper shows three structural gaps in the settlement: variance is defined over impressions rather than reach, coverage thresholds count ads rather than people and therefore allow the platform to leave the largest campaigns unregulated, and nothing forbids leveling down. Its field experiments with 36 paired ad campaigns show VRS-enabled ads reduce variance below the 10% threshold for race in all 18 race experiments, but mean reach drops by 9.82% and cost per 1,000 reached rises by 12.02%, with the leveling-down pattern appearing in several replications. The paper further claims its budget-splitting strategy dominated VRS in 36 paired experiments, increasing reach for Black, White, female, and male users while reducing cost, which it reads as evidence that the expense and reduced exposure of VRS are artifacts of Meta's implementation.","pith_inferences":["The eligible-ratio estimation is the main circularity risk: because the baseline is derived from VRS's own delivery outcomes, an independent baseline from platform-side aggregate impression data would be the cleanest confirmation that the measured variance reduction is real rather than an artifact.","Reframing compliance metrics in terms of unique reach would align the settlement with its stated goal of opportunity access, and would make VRS's reduced reach directly noncompliant rather than merely undesirable.","The budget-splitting result suggests a concrete, testable design for platforms: expose demographic budget-allocation controls to advertisers, letting fairness constraints be set explicitly instead of through an opaque bid multiplier.","A natural extension of the paper's findings is that future settlements should specify non-degradation constraints on reach and cost alongside variance targets, otherwise compliant systems can still reduce access to opportunities."],"forward_implications":["Under the settlement's own compliance metrics, VRS does reduce variance for housing-tagged ads, including below 5% in 15 of 18 race experiments, so the mechanism is not inert.","For a fixed advertiser budget, VRS reduces the number of unique users who see opportunity ads, with mean reach down 9.82% and cost per 1,000 reached up 12.02%, implying compliance costs are passed to advertisers and users.","Voluntary expansion of VRS to employment and credit ads does not achieve the same variance reduction as housing: variance stayed above 10%, comparable to no-VRS delivery, so voluntary extension should not be assumed equivalent.","Coverage targets based on the share of ads allow the platform to exclude the largest campaigns; using public political-ad data, excluding the largest 19% of ads would exempt about 78.9% of impressions, far more than the 19% a random exclusion would exempt.","Budget-splitting, with an equal budget per demographic subgroup and separate campaigns, increases reach for all groups and lowers cost relative to VRS, showing that higher cost and lower reach are not inherent to fairness interventions."],"supporting_citations":[{"why":"The settlement terms that mandate the Variance Reduction System and define the compliance framework the paper critiques.","marker":"[70]"},{"why":"Meta's technical report defining eligible ratio, delivery ratio, variance formula, and the VRS bid-multiplier mechanism.","marker":"[53]"},{"why":"The reinforcement-learning paper describing the module that adjusts bids and the differential privacy noise added to variance measurements.","marker":"[72]"},{"why":"The external reviewer's compliance metrics report supplying coverage targets and the data schema used to verify them.","marker":"[32]"},{"why":"The paired-audit methodology for isolating ad-delivery bias that the paper adapts to isolate VRS's effect.","marker":"[3]"},{"why":"Prior auditing of job-ad delivery that supplies the race-from-location proxy and evidence of delivery bias in opportunity ads.","marker":"[35]"},{"why":"Prior auditing of education ads whose creatives are reused and whose bias evidence motivates two of the paper's ad categories.","marker":"[37]"},{"why":"Simulation results on fairness interventions in ad delivery causing leveling down, which the paper tests with real ads.","marker":"[9]"},{"why":"The prior proposal of budget-splitting for demographically fair outcomes that the paper implements as its alternative.","marker":"[29]"},{"why":"Public data on advertiser budgets and reach used to quantify how selective VRS application could exempt large shares of impressions.","marker":"[75]"}],"fun_headline_variants":["Meta's anti-bias system cuts reach, raises costs","Ad bias fix reaches fewer, costs more – study","Meta's fairness fix reduces ads, not bias","VRS: less variance, less reach, higher cost","Settlement loopholes let Meta level down ads"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper's estimate of the eligible ratio, the baseline against which VRS's variance reduction is judged, is not measured from the platform's eligible-audience data; it is the mean of the delivery ratios of the VRS-enabled ads themselves, so if VRS's delivery is itself skewed, the measured reduction in variance may be an artifact of the baseline.","fun_headline_variants_meta":{"raw":{"variants":["Meta's anti-bias system cuts reach, raises costs","Ad bias fix reaches fewer, costs more – study","Meta's fairness fix reduces ads, not bias","VRS: less variance, less reach, higher cost","Settlement loopholes let Meta level down ads"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000197,"raw_usage":{"total_tokens":1420,"prompt_tokens":1055,"completion_tokens":365,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":671,"completion_tokens_details":{"reasoning_tokens":289}},"tokens_in":671,"tokens_out":365,"duration_ms":3628,"temperature":1.0,"reasoning_tokens":289,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T19:23:29.998654+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same paired VRS/no-VRS campaigns but compute variance against an independently obtained eligible ratio, for example the demographic impression distribution across all advertisers in the same geographic areas from public ad-library data or a compliance report, and check whether VRS still reduces variance below the 10% threshold; if it does not, the paper's central empirical claim fails. A second decisive test would compare reach and cost under VRS with reach and cost under budget-splitting at larger budgets and longer durations, since the paper's 24-hour, $20 campaigns may not reflect steady-state delivery.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The settlement terms that mandate the Variance Reduction System and define the compliance framework the paper critiques."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Meta's technical report defining eligible ratio, delivery ratio, variance formula, and the VRS bid-multiplier mechanism."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The reinforcement-learning paper describing the module that adjusts bids and the differential privacy noise added to variance measurements."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The external reviewer's compliance metrics report supplying coverage targets and the data schema used to verify them."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Public data on advertiser budgets and reach used to quantify how selective VRS application could exempt large shares of impressions."}],"review_version":2}