{"id":"c88dfafb-d12e-4f90-ac36-caa282e0496e","arxiv_id":"2502.06439","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"A 2024 black-box audit of an Italian car insurance comparator confirms that birthplace, age, city, education, and other personal attributes still influence quoted premiums, and that some companies vary quote availability by profile.","lead":"This paper repeats and extends a 2021 audit of Italian car insurance comparison websites, using thousands of fake driver profiles to test whether personal details change prices. It finds that birthplace, age, city, education, and other attributes still move quoted premiums in early 2024, and that some insurers give different numbers of quotes to different profiles.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Top-5 price averages mix price levels with quote availability: the RQ1 birthplace/city effects may partly reflect different company sets, not just premium differences.","rationale":"The reader's weakest assumption identifies exactly the load-bearing issue: the top-5 statistic is ill-defined when quote counts differ and, more generally, confounds the identity of the quoters with the prices they charge. I agree with this and extend it: even when five quotes are present, the set of company-products entering the average may differ across profiles, and Figure 1 shows this is not merely a theoretical possibility. The paper's use of randomized query order, control pairs, and an open repository are genuine strengths, and the sign tests are appropriate as an initial screen. The reason I keep the verdict UNCHANGED rather than moving to REJECT is that the concern is fixable by a matched reanalysis, and the direction of the reported effects is large enough that a composition effect is plausible but not established. I did not find a stronger objection: the external-validity threats are honestly stated, the lack of multiple-comparison correction is secondary given the number of significant tests, and the RQ2 evidence, while descriptive, is not the main load-bearing claim about pricing. The proposed concrete test would settle whether the central pricing claim survives the confound.","tokens_in":10350,"tokens_out":3467,"duration_ms":42248,"concrete_test":"Recompute the RQ1 top-5 and top-1 comparisons on the subset of company-product offers that appear for both members of each profile pair, i.e., a matched common-quote universe; report how profiles with fewer than five quotes are handled. Alternatively, compute within-company, same-product price differences for each attribute pair and average those, so company composition is held fixed. If the birthplace MA/MI and city NA/MI median differences remain large and significant under the matched analysis, the concern is settled; if they shrink toward zero or reverse, the headline must be reframed as an availability/access effect rather than a pure pricing effect.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central RQ1 evidence (Tables 3-4 and the abstract's claim that birthplace is the main discriminatory pricing factor) rests on comparing averages of the five cheapest quotes across profiles. But the comparator output is sparse and availability is attribute-dependent: Table 2 shows companies appearing at rates from 3% to 100%, and Figure 1 shows, e.g., C2 quotes fewer Moroccan profiles, C6 quotes only Milan residents, and C3/C6 appear selectively by risk class. The paper nowhere defines the top-5 average for profiles with fewer than five quotes, nor restricts the comparison to a common set of company/product offers. If, for a pair such as birthplace MA vs MI, the set of quoting companies differs by profile, the reported median differences (125-252 EUR in top-1/top-5) conflate the prices charged by companies with whether companies appear at all. Since RQ2 documents that availability depends on protected attributes, this composition effect is plausible and directly load-bearing for the strongest claim. The effect may still constitute discrimination in access, but the specific pricing claim would be weakened unless the analysis separates price level from quote availability.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports an updated and extended algorithmic audit of an Italian car insurance comparison website, replicating a 2021 study by Fabris et al. The authors collected 7,680 black-box queries in January 2024 using a double nested randomization design and control pairs. RQ1 analyzes whether protected attributes (gender, birthplace, age) and socio-demographic attributes (city, marital status, education, profession) directly influence quoted premiums, using paired sign tests on top-1 and top-5 price differences. RQ2 examines whether these attributes plus driving attributes (car, km driven, class) influence the number of quotes returned. The paper reports significant price differences for birthplace, age, city, education, and profession, and differences in quote availability by car type and risk class, with some company-specific patterns for birthplace and city. The authors conclude that demographic variables still affect pricing and that quote availability can deny equal opportunities.","tokens_in":10568,"tokens_out":5741,"duration_ms":46397,"significance":"If the findings withstand scrutiny, this is a valuable, time-resolved audit of a live commercial system. The study's strengths include a paired profile design, randomized query ordering, control pairs as a baseline for noise, a large query volume, and the public release of data and analysis scripts. The RQ2 finding that quote availability varies with driver characteristics is an important contribution in its own right. However, the load-bearing top-5 analysis conflates price levels with quote availability, and the abstract's 'main discriminatory factor' claim is not fully consistent with the paper's own descriptive statistics. These issues require substantive revision before the central claims can be accepted as stated.","major_comments":[{"comment":"Section 3.2 defines the top5 statistic as 'the averages of the five cheapest quotes for every profile' but does not state how profiles with fewer than five quotes are handled, nor does it restrict the averaging to a common set of company/product offers. Table 2 shows company appearance rates from 3% to 100%, and Figure 1 shows, for example, that C6 quotes only Milan residents and that C3/C6 appear only for risk class 18. Since RQ2 itself demonstrates that availability depends on protected and socio-demographic attributes, the top5 price differences in Tables 3-4 (e.g., Birthplace MA vs MI median 252€; City NA vs MI median 278€) may reflect a composition effect—different sets of quoting companies—rather than price-level differences alone. This is load-bearing for the abstract's claim that birthplace is the main discriminatory pricing factor. The authors should (a) define and justify the treatment of profiles with fewer than five quotes, and (b) recompute the top5 analysis on a common set of companies or per-company matched quotes as a robustness check, or explicitly reframe the result as a combined access-and-pricing effect.","section":"§3.2, Tables 3-4, Figure 1"},{"comment":"The abstract and conclusions state that 'birthplace remaining the main discriminatory factor,' but Section 4.1 and the Discussion in Section 5 report that City shows greater differences (Tables 3 and 4: City NA vs MI median 147€ top1 and 278€ top5 vs Birthplace MA vs MI 125€ top1 and 252€ top5; means 367€ vs 148€ and 657€ vs 371€). If 'main' is intended only among protected attributes (gender, birthplace, age), this should be stated explicitly in the abstract and conclusions; otherwise the claim is inconsistent with the paper's own data.","section":"Abstract; §4.1; §5; Tables 3-4"},{"comment":"The significance tests in Tables 3 and 4 are unadjusted repeated sign tests (11 profile pairs × 2 analyses), and the tie-handling rule is unspecified. The tables classify differences within ±5€ as 'Ties5,' but it is not clear whether the sign test discards zero differences and whether the Ties5 tolerance is applied before or after the sign test. All p-values are reported only as '<0.05' without exact values, which obscures the evidence for borderline pairs such as Birthplace RO vs MI (median 2-8€). For RQ2, Section 4.2 reports availability patterns from Figure 1 without any statistical test, yet the research question asks whether attributes 'influence' the number of quotes; the authors should add a formal comparison (e.g., chi-square or Fisher exact tests) or explicitly label those findings as descriptive.","section":"§3.2; Tables 3-4; §4.2"}],"minor_comments":[{"comment":"The manuscript uses 'italian comparison website' with a lowercase 'i'; please capitalize to 'Italian comparison website'.","section":"§3.3"},{"comment":"The y-axis label 'f(C2)' etc. is not defined in the caption; please state that it is the proportion of profiles for which the company provides at least one quote, and note explicitly that C1 and C5 are omitted because they appear for 100% of profiles.","section":"§4.2; Figure 1"},{"comment":"Replace the p-value column's '<0.05' entries with exact p-values (e.g., p=0.003) to allow readers to assess the strength of evidence, especially for pairs with small median differences such as Birthplace RO vs MI.","section":"Tables 3-4"},{"comment":"The threats-to-validity section omits the availability-composition threat identified in Major Comment 1 and the multiple-comparisons issue in Major Comment 3; please update Section 6 to address these internal validity threats explicitly.","section":"§6"}],"recommendation":"major_revision","confidential_remarks":"This is a solid empirical audit with a strong design, but the top-5 composition issue is central to the RQ1 claim and needs a robustness analysis before the paper can be accepted. The scope fits a software engineering venue focusing on fairness testing and empirical methods."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: this is a useful, honest replication audit. The new measurements are real – January 2024 data from a live comparator, new attributes (marital status, education, profession), and a documented quote-availability analysis. It does not claim a new method, which is fine. It does what a good audit should: paired profiles differing in one attribute, control pairs, randomization, sign tests, and a stated open repository for data and scripts.\n\nThe main claim – that pricing still differentiates on birthplace, city, age, and education – is supported for the top1 and top5 analyses. The effect sizes are large, and the control pairs behave as expected. That part holds up.\n\nThe soft spot is real and the paper does not address it. The top5 analysis compares averages of the five cheapest quotes per profile, but the paper never says what happens when a profile receives fewer than five quotes. Table 2 shows company presence varies from 3% to 100%, and Figure 1 shows company-specific gaps by birthplace, city, car type, and risk class. So the top5 average is a mix of which companies appear and what those companies charge. For the strongest comparisons (Morocco vs. Milan, Naples vs. Milan), the median differences of 125–252 € could partly reflect a different set of quoting insurers, not just higher premiums from the same insurers. That is load-bearing for the abstract's \"birthplace remains the main discriminatory factor\" claim. The fix is straightforward: restrict the top-5 average to a common set of company/products, or report price level and availability separately.\n\nTwo smaller issues. The paper runs many sign tests without multiple-comparison correction; with roughly 20 tests, a few p<0.05 are expected by chance. RQ2 is reported descriptively with no confidence intervals or significance tests, so the \"only two companies\" conclusion is softer than it looks. Both are fixable.\n\nThe replication design also changes the attribute values from the 2021 study (removing Rome and Romania, swapping Ghana/Laos for Morocco/China). That is explained, but it means the comparison to the original audit is not exact.\n\nWho is this for? Anyone working on algorithmic auditing, fairness in pricing, or software testing for non-discrimination. It gives a current, regulator-relevant data point. It deserves a serious referee; I would accept it with major revisions to fix the top-5 definition and add significance testing for RQ2. The central finding is likely to survive, but the strongest claims need the composition effect separated out.","headline":"A useful, honest replication audit with real new measurements, but the top-5 price analysis confounds price levels with quote availability and needs fixing before the strongest claims fully land.","tokens_in":11096,"tokens_out":2401,"would_cite":true,"duration_ms":26431,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Italian car insurance still prices drivers by birthplace and demographics, audit finds.","keywords":["algorithmic bias","non-discrimination testing","car insurance","black-box audit","price discrimination","quote availability","empirical software engineering"],"falsifier":"If a re-analysis restricted to profiles for which all five companies actually returned quotes, or an independent audit on a different comparator, found the birthplace and city price gaps shrink to near zero, the paper's central claim of direct price discrimination would be refuted.","tokens_in":10183,"feed_emoji":"🚗","tokens_out":6514,"duration_ms":50199,"temperature":0.7,"pith_summary":"This paper attempts to establish whether the pricing algorithms of a leading Italian car insurance comparison website still treat protected and socio-demographic attributes unequally, years after an earlier audit found they did. Using 7,680 black-box queries run in January 2024, the authors compare quotes for driver profiles that differ in exactly one attribute and count how many companies return offers for each profile. They report that birthplace remains the strongest price factor—drivers born in Morocco or China are quoted hundreds of euros more than identical Milan-born drivers—and that city of residence, age, education, and profession also shift premiums substantially. They further find that for some companies, car type, risk class, birthplace, and city determine whether any quote is offered. The study matters because if these results are correct, a regulated consumer market is still being priced in a way that restricts equal opportunity.","feed_headline":"Italian car insurance still prices by birthplace, audit finds","feed_subtitle":"A 7,680-query black-box audit shows foreign-born drivers pay hundreds more and some profiles get fewer quotes.","key_machinery":"The central mechanism is a paired-profile black-box audit: automated queries to the comparator with profiles that differ in exactly one attribute, plus control pairs (two identical queries) to measure background noise. Price differences are summarized as the median and percentiles of pairwise differences for the cheapest quote (top-1) and the average of the five cheapest quotes (top-5), with a sign test for significance; quote availability is measured as the fraction of profiles for which each company appears. This design lets the authors separate attribute-driven price effects from random fluctuation and from the set of companies that respond.","core_discovery":"The paper confirms and extends a 2021 algorithmic audit of the Italian car insurance market. As of January 2024, demographic variables still significantly affect quoted premiums, with birthplace remaining the main discriminatory factor: in the comparison of the five cheapest offers, drivers born in Morocco are charged on average €371 more than otherwise identical drivers born in Milan, and drivers born in China €200 more; the median top-1 gap for Morocco is €125. City of residence produces even larger median differences in the top-5 analysis (Naples vs Milan: €278), while a 25-year-old pays a median €211 more over five quotes than a 32-year-old. Profiles without a qualification pay a median €99 more over five quotes than Master's graduates, and jobseekers €22 more. In addition, the number of quotes a user receives is not equal: company C2 offers fewer quotes to profiles born in Morocco, C3 and C6 appear almost exclusively for the highest risk class, and C4 and C6 are absent for certain car types. Control pairs (identical queries) show essentially zero noise in 98% of cases, indicating the observed gaps are not artifacts of price fluctuation.","pith_inferences":["A re-analysis that matches companies across profiles—comparing only quotes from companies that appear for both members of a pair—could separate the role of quote availability from the role of pricing within each company, and might change the magnitude of the top-5 gaps.","Because the comparator itself may run A/B tests and customize offers, a future audit with repeated identical queries over longer periods could disentangle the comparator's own logic from the insurers' pricing rules.","The finding that some insurers target the highest risk class could mean the market segments by risk in unexpected ways; a welfare analysis would be needed to judge whether fewer quotes for safe drivers is a harm or a strategic choice.","The same paired-profile methodology could be ported to other regulated insurance markets or to consumer credit, where quote availability is also a channel of potential discrimination."],"forward_implications":["Regulators and consumer watchdogs can treat the results as evidence that automated pricing in Italian car insurance still violates the spirit of non-discrimination rules, prompting renewed oversight.","The audit protocol, including control pairs, can be rerun periodically to track whether price gaps by birthplace or city widen, shrink, or shift to other attributes.","If some profiles receive fewer quotes, users in those groups face a smaller choice set, which likely reduces their ability to find a competitive price, compounding the price penalty.","The fact that education and marital status affect premiums suggests that even attributes not directly regulated can act as proxies for age or risk; if they remain after controlling for those, they are themselves a fairness concern."],"supporting_citations":[{"why":"the prior audit this study replicates; supplies the original methodology and the gender/birthplace results it updates.","marker":"[9]"},{"why":"official road-accident statistics used to argue higher Naples prices do not match recorded accident frequency.","marker":"[18]"},{"why":"the 2011 European Court of Justice ruling that banned gender-based insurance pricing, which motivates the non-discrimination testing.","marker":"[7]"},{"why":"Article 21 of the EU Charter of Fundamental Rights, the source of the protected attributes under test.","marker":"[8]"},{"why":"the EU directive barring nationality-based premium surcharges, against which the birthplace findings are assessed.","marker":"[5]"}],"fun_headline_variants":["Birthplace still costs €371 more in Italian car insurance","Car insurance audit: birthplace, city, age all still skew prices","Foreign-born drivers overpay: Italian car insurance audit","Italian car insurers still discriminate by birthplace: €371 gap","Audit finds Italian car pricing still biased by birthplace and profile"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The top-5 price comparison assumes that the five cheapest quotes are comparable across profiles; the paper never states how profiles that receive fewer than five quotes are handled, so the reported averages may reflect differences in which companies appear rather than pure price differences.","fun_headline_variants_meta":{"raw":{"variants":["Birthplace still costs €371 more in Italian car insurance","Car insurance audit: birthplace, city, age all still skew prices","Foreign-born drivers overpay: Italian car insurance audit","Italian car insurers still discriminate by birthplace: €371 gap","Audit finds Italian car pricing still biased by birthplace and profile"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000823,"raw_usage":{"total_tokens":3630,"prompt_tokens":1006,"completion_tokens":2624,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":622,"completion_tokens_details":{"reasoning_tokens":2541}},"tokens_in":622,"tokens_out":2624,"duration_ms":17007,"temperature":1.0,"reasoning_tokens":2541,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-08T15:25:15.720075+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"If a re-analysis restricted to profiles for which all five companies actually returned quotes, or an independent audit on a different comparator, found the birthplace and city price gaps shrink to near zero, the paper's central claim of direct price discrimination would be refuted.","supporting_citations":[{"cited_title":"In: Proceedings of the 2021 AAAI/ACM Conference on AI, Ethics, and Society","cited_arxiv_id":null,"evidence_quote":"the prior audit this study replicates; supplies the original methodology and the gender/birthplace results it updates."},{"cited_title":"Anno 2022","cited_arxiv_id":null,"evidence_quote":"official road-accident statistics used to argue higher Naples prices do not match recorded accident frequency."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"the 2011 European Court of Justice ruling that banned gender-based insurance pricing, which motivates the non-discrimination testing."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Article 21 of the EU Charter of Fundamental Rights, the source of the protected attributes under test."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"the EU directive barring nationality-based premium surcharges, against which the birthplace findings are assessed."}],"review_version":1}