Pith. sign in

REVIEW 3 major objections 4 minor 29 references

Testing software for non-discrimination: an updated and extended audit in the Italian car insurance domain

T0 review · 3 major / 4 minor · reviewed 2026-08-08 · deepseek-v4-flash

Pith's one-line read Italian car insurance still prices drivers by birthplace and demographics, audit finds.

desk verdict A useful, honest replication audit with real new measurements, but the top-5 price analysis confounds price levels with quote availability and needs fixing before the strongest claims fully land. read the letter →

arxiv 2502.06439 v1 pith:6HOC3D3G submitted 2025-02-10 cs.SE cs.AIcs.HCcs.LG

classification cs.SEcs.AIcs.HCcs.LG
keywords algorithmicbiasnon-discriminationtestingcarinsuranceblack-boxauditpricediscriminationquoteavailabilityempiricalsoftwareengineering
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper attempts to establish whether the pricing algorithms of a leading Italian car insurance comparison website still treat protected and socio-demographic attributes unequally, years after an earlier audit found they did. Using 7,680 black-box queries run in January 2024, the authors compare quotes for driver profiles that differ in exactly one attribute and count how many companies return offers for each profile. They report that birthplace remains the strongest price factor—drivers born in Morocco or China are quoted hundreds of euros more than identical Milan-born drivers—and that city of residence, age, education, and profession also shift premiums substantially. They further find that for some companies, car type, risk class, birthplace, and city determine whether any quote is offered. The study matters because if these results are correct, a regulated consumer market is still being priced in a way that restricts equal opportunity.

What carries the argument

The central mechanism is a paired-profile black-box audit: automated queries to the comparator with profiles that differ in exactly one attribute, plus control pairs (two identical queries) to measure background noise. Price differences are summarized as the median and percentiles of pairwise differences for the cheapest quote (top-1) and the average of the five cheapest quotes (top-5), with a sign test for significance; quote availability is measured as the fraction of profiles for which each company appears. This design lets the authors separate attribute-driven price effects from random fluctuation and from the set of companies that respond.

What would settle it

If a re-analysis restricted to profiles for which all five companies actually returned quotes, or an independent audit on a different comparator, found the birthplace and city price gaps shrink to near zero, the paper's central claim of direct price discrimination would be refuted.

Watch

Extended reading notes

Core claim

The paper confirms and extends a 2021 algorithmic audit of the Italian car insurance market. As of January 2024, demographic variables still significantly affect quoted premiums, with birthplace remaining the main discriminatory factor: in the comparison of the five cheapest offers, drivers born in Morocco are charged on average €371 more than otherwise identical drivers born in Milan, and drivers born in China €200 more; the median top-1 gap for Morocco is €125. City of residence produces even larger median differences in the top-5 analysis (Naples vs Milan: €278), while a 25-year-old pays a median €211 more over five quotes than a 32-year-old. Profiles without a qualification pay a median €99 more over five quotes than Master's graduates, and jobseekers €22 more. In addition, the number of quotes a user receives is not equal: company C2 offers fewer quotes to profiles born in Morocco, C3 and C6 appear almost exclusively for the highest risk class, and C4 and C6 are absent for certain car types. Control pairs (identical queries) show essentially zero noise in 98% of cases, indicating the observed gaps are not artifacts of price fluctuation.

Load-bearing premise

The top-5 price comparison assumes that the five cheapest quotes are comparable across profiles; the paper never states how profiles that receive fewer than five quotes are handled, so the reported averages may reflect differences in which companies appear rather than pure price differences.

Editorial extensions

If this is right

  • Regulators and consumer watchdogs can treat the results as evidence that automated pricing in Italian car insurance still violates the spirit of non-discrimination rules, prompting renewed oversight.
  • The audit protocol, including control pairs, can be rerun periodically to track whether price gaps by birthplace or city widen, shrink, or shift to other attributes.
  • If some profiles receive fewer quotes, users in those groups face a smaller choice set, which likely reduces their ability to find a competitive price, compounding the price penalty.
  • The fact that education and marital status affect premiums suggests that even attributes not directly regulated can act as proxies for age or risk; if they remain after controlling for those, they are themselves a fairness concern.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A re-analysis that matches companies across profiles—comparing only quotes from companies that appear for both members of a pair—could separate the role of quote availability from the role of pricing within each company, and might change the magnitude of the top-5 gaps.
  • Because the comparator itself may run A/B tests and customize offers, a future audit with repeated identical queries over longer periods could disentangle the comparator's own logic from the insurers' pricing rules.
  • The finding that some insurers target the highest risk class could mean the market segments by risk in unexpected ways; a welfare analysis would be needed to judge whether fewer quotes for safe drivers is a harm or a strategic choice.
  • The same paired-profile methodology could be ported to other regulated insurance markets or to consumer credit, where quote availability is also a channel of potential discrimination.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. This paper reports an updated and extended algorithmic audit of an Italian car insurance comparison website, replicating a 2021 study by Fabris et al. The authors collected 7,680 black-box queries in January 2024 using a double nested randomization design and control pairs. RQ1 analyzes whether protected attributes (gender, birthplace, age) and socio-demographic attributes (city, marital status, education, profession) directly influence quoted premiums, using paired sign tests on top-1 and top-5 price differences. RQ2 examines whether these attributes plus driving attributes (car, km driven, class) influence the number of quotes returned. The paper reports significant price differences for birthplace, age, city, education, and profession, and differences in quote availability by car type and risk class, with some company-specific patterns for birthplace and city. The authors conclude that demographic variables still affect pricing and that quote availability can deny equal opportunities.

Significance. If the findings withstand scrutiny, this is a valuable, time-resolved audit of a live commercial system. The study's strengths include a paired profile design, randomized query ordering, control pairs as a baseline for noise, a large query volume, and the public release of data and analysis scripts. The RQ2 finding that quote availability varies with driver characteristics is an important contribution in its own right. However, the load-bearing top-5 analysis conflates price levels with quote availability, and the abstract's 'main discriminatory factor' claim is not fully consistent with the paper's own descriptive statistics. These issues require substantive revision before the central claims can be accepted as stated.

major comments (3)
  1. [§3.2, Tables 3-4, Figure 1] Section 3.2 defines the top5 statistic as 'the averages of the five cheapest quotes for every profile' but does not state how profiles with fewer than five quotes are handled, nor does it restrict the averaging to a common set of company/product offers. Table 2 shows company appearance rates from 3% to 100%, and Figure 1 shows, for example, that C6 quotes only Milan residents and that C3/C6 appear only for risk class 18. Since RQ2 itself demonstrates that availability depends on protected and socio-demographic attributes, the top5 price differences in Tables 3-4 (e.g., Birthplace MA vs MI median 252€; City NA vs MI median 278€) may reflect a composition effect—different sets of quoting companies—rather than price-level differences alone. This is load-bearing for the abstract's claim that birthplace is the main discriminatory pricing factor. The authors should (a) define and justify the treatment of profiles with fewer than five quotes, and (b) recompute the top5 analysis on a common set of companies or per-company matched quotes as a robustness check, or explicitly reframe the result as a combined access-and-pricing effect.
  2. [Abstract; §4.1; §5; Tables 3-4] The abstract and conclusions state that 'birthplace remaining the main discriminatory factor,' but Section 4.1 and the Discussion in Section 5 report that City shows greater differences (Tables 3 and 4: City NA vs MI median 147€ top1 and 278€ top5 vs Birthplace MA vs MI 125€ top1 and 252€ top5; means 367€ vs 148€ and 657€ vs 371€). If 'main' is intended only among protected attributes (gender, birthplace, age), this should be stated explicitly in the abstract and conclusions; otherwise the claim is inconsistent with the paper's own data.
  3. [§3.2; Tables 3-4; §4.2] The significance tests in Tables 3 and 4 are unadjusted repeated sign tests (11 profile pairs × 2 analyses), and the tie-handling rule is unspecified. The tables classify differences within ±5€ as 'Ties5,' but it is not clear whether the sign test discards zero differences and whether the Ties5 tolerance is applied before or after the sign test. All p-values are reported only as '<0.05' without exact values, which obscures the evidence for borderline pairs such as Birthplace RO vs MI (median 2-8€). For RQ2, Section 4.2 reports availability patterns from Figure 1 without any statistical test, yet the research question asks whether attributes 'influence' the number of quotes; the authors should add a formal comparison (e.g., chi-square or Fisher exact tests) or explicitly label those findings as descriptive.
minor comments (4)
  1. [§3.3] The manuscript uses 'italian comparison website' with a lowercase 'i'; please capitalize to 'Italian comparison website'.
  2. [§4.2; Figure 1] The y-axis label 'f(C2)' etc. is not defined in the caption; please state that it is the proportion of profiles for which the company provides at least one quote, and note explicitly that C1 and C5 are omitted because they appear for 100% of profiles.
  3. [Tables 3-4] Replace the p-value column's '<0.05' entries with exact p-values (e.g., p=0.003) to allow readers to assess the strength of evidence, especially for pairs with small median differences such as Birthplace RO vs MI.
  4. [§6] The threats-to-validity section omits the availability-composition threat identified in Major Comment 1 and the multiple-comparisons issue in Major Comment 3; please update Section 6 to address these internal validity threats explicitly.

Circularity Check

0 steps flagged · score 0.0 of 10

No circular step: the audit is an independent measurement against live system outputs, with controls; self-citations are design inheritance, not load-bearing evidence.

full rationale

The paper's results are measured quantities, not derived predictions. RQ1 is computed as differences between insurance quotes for pairs of profiles that vary in one attribute, and RQ2 as the frequency of company offers per attribute value. No parameter is fitted to a subset of data and then 'predicted' on a related subset; no equation defines the conclusion into the metric. The control pairs provide an empirical noise baseline rather than an assumed outcome. The central evidence (e.g., the Moroccan- versus Milan-born top-1/top-5 median differences of 125 EUR and 252 EUR) is direct descriptive statistics of newly scraped January 2024 comparator output. The self-citations to Fabris et al. (2021) are used to motivate replication and to reuse the double nested randomization procedure; they are procedural references, not uniqueness theorems or unverified premises whose acceptance forces the current results. The skeptic's concern about the top-5 average mixing price level with quote availability is a genuine measurement-validity threat, because the paper never defines how profiles with fewer than five offers are averaged, but it is not circularity: it questions what the statistic measures, not whether the statistic was constructed from the conclusion it supports.

Assumptions & free parameters 1 free parameters · 4 assumptions · 0 invented entities

The central empirical claim rests on the design assumptions listed above rather than on mathematical axioms. The only hand-set analysis threshold is the Ties5 tolerance; no model parameters are fitted to data, and no new entities are introduced.

free parameters (1)
  • Ties5 tolerance = ±5 €
    Hand-set threshold for classifying a price difference as a tie; directly determines Ties5 percentages in Tables 3-4 and the comparison to control pairs. No sensitivity analysis is reported.
assumptions (4)
  • domain assumption Quotes returned by the chosen comparison website faithfully represent the pricing algorithms under audit.
    The audit queries one intermediary; if the site rewrites or filters offers, measured prices may not reflect insurers' actual pricing. Location: Sections 2 and 3.3.
  • domain assumption Profiles differing in only one attribute isolate the effect of that attribute.
    Requires no hidden correlated factor such as session timing, A/B assignment, or quote availability to systematically differ between paired profiles. Location: Section 3.2.
  • domain assumption Control pairs capture all non-modelled noise.
    The comparison of test-pair differences to control-pair differences assumes identical queries experience the same background variability as attribute-varying queries. Location: Section 3.3.
  • ad hoc to paper Repeated sign tests without multiple-comparison correction are an acceptable significance procedure.
    The paper reports many p-values across attribute pairs and does not correct for multiplicity, yet interprets each at the 0.05 level. Location: Section 3.2 and Tables 3-4.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Testing software for non-discrimination: an updated and extended audit in the Italian car insurance domain." pith.science (2026). https://pith.science/paper/6HOC3D3G

@misc{pith2026250206439,
  author       = {Pith},
  title        = {Pith review of: Testing software for non-discrimination: an updated and extended audit in the Italian car insurance domain},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/6HOC3D3G}},
  note         = {Machine review of arXiv:2502.06439}
}
read the original abstract

Context. As software systems become more integrated into society's infrastructure, the responsibility of software professionals to ensure compliance with various non-functional requirements increases. These requirements include security, safety, privacy, and, increasingly, non-discrimination. Motivation. Fairness in pricing algorithms grants equitable access to basic services without discriminating on the basis of protected attributes. Method. We replicate a previous empirical study that used black box testing to audit pricing algorithms used by Italian car insurance companies, accessible through a popular online system. With respect to the previous study, we enlarged the number of tests and the number of demographic variables under analysis. Results. Our work confirms and extends previous findings, highlighting the problematic permanence of discrimination across time: demographic variables significantly impact pricing to this day, with birthplace remaining the main discriminatory factor against individuals not born in Italian cities. We also found that driver profiles can determine the number of quotes available to the user, denying equal opportunities to all. Conclusion. The study underscores the importance of testing for non-discrimination in software systems that affect people's everyday lives. Performing algorithmic audits over time makes it possible to evaluate the evolution of such algorithms. It also demonstrates the role that empirical software engineering can play in making software systems more accountable.

Figures

Figures reproduced from arXiv: 2502.06439 by the authors.

Figure 1
Figure 1. Presence of companies as a percentage of offers for each attribute. 4.2 Output Variability (RQ2) [PITH_FULL_IMAGE:figures/full_fig_p009_1.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

29 extracted references · 23 canonical work pages

  1. [1]

    Rondina et al

    ACM Code 2018 Task Force: ACM Code of Ethics and Professional Conduct (2018), https://www.acm.org/code-of-ethics 12 M. Rondina et al

  2. [2]

    International Journal of Artificial Intelligence in Education32(4), 1052–1092 (Dec 2022)

    Baker, R.S., Hawn, A.: Algorithmic Bias in Education. International Journal of Artificial Intelligence in Education32(4), 1052–1092 (Dec 2022). https://doi.org/ 10.1007/s40593-021-00285-9

  3. [3]

    Brun,Y.,Meliou,A.:Softwarefairness.In:Proceedingsofthe201826thACMJoint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering. pp. 754–759. ESEC/FSE 2018, Association for Computing Machinery, New York, NY, USA (Oct 2018). https://doi.org/10. 1145/3236024.3264838

  4. [4]

    Communications of the ACM66(1), 100 (Dec 2022)

    Conitzer, V., Hadfield, G.K., Vallor, S.: Technical Perspective: The Impact of Au- diting for Algorithmic Bias. Communications of the ACM66(1), 100 (Dec 2022). https://doi.org/10.1145/3571152

  5. [5]

    Council of the European Union: Council Directive 2000/43/EC of 29 June 2000 implementingtheprincipleofequaltreatmentbetweenpersonsirrespectiveofracial or ethnic origin (Jun 2000), http://data.europa.eu/eli/dir/2000/43/oj/eng

  6. [6]

    https: //doi.org/10.48550/arXiv.2403.20089

    Deck, L., Müller, J.L., Braun, C., Zipperling, D., Kühl, N.: Implications of the AI Act for Non-Discrimination Law and Algorithmic Fairness (Mar 2024). https: //doi.org/10.48550/arXiv.2403.20089

  7. [7]

    European Court of Justice: Association Belge des Consommateurs Test-Achats ASBL and Others v Conseil des ministres (Mar 2011), https://eur-lex.europa.eu/ legal-content/en/TXT/?uri=CELEX:62009CJ0236

  8. [8]

    European Union Agency For Fundamental Rights: EU Charter of Fundamental Rights - Title III: Quality - Article 21 - Non-discrimination (Apr 2015), http: //fra.europa.eu/en/eu-charter/article/21-non-discrimination

Show all 29 references
  1. [9]

    In: Proceedings of the 2021 AAAI/ACM Conference on AI, Ethics, and Society

    Fabris, A., Mishler, A., Gottardi, S., Carletti, M., Daicampi, M., Susto, G.A., Silvello, G.: Algorithmic Audit of Italian Car Insurance: Evidence of Unfairness in Access and Pricing. In: Proceedings of the 2021 AAAI/ACM Conference on AI, Ethics, and Society. pp. 458–468. AIES...

  2. [10]

    In: Proceedings of the 22nd International Confer- ence on Software Engineering and Knowledge Engineering

    Feldt, R., Magazinius, A.: Validity threats in empirical software engineering research-an initial survey. In: Proceedings of the 22nd International Confer- ence on Software Engineering and Knowledge Engineering. pp. 374–379. Knowl- edge Systems Institute Graduate School, Redwo...

  3. [11]

    In: Proceedings of the 2017 11th Joint Meeting on Foundations of Software Engineering

    Galhotra, S., Brun, Y., Meliou, A.: Fairness testing: Testing software for discrimi- nation. In: Proceedings of the 2017 11th Joint Meeting on Foundations of Software Engineering. pp. 498–510. ESEC/FSE 2017, Association for Computing Machinery, New York, NY, USA (Aug 2017). ht...

  4. [12]

    European Jour- nal of Law and Economics50(3), 405–435 (Dec 2020)

    Gautier, A., Ittoo, A., Van Cleynenbreugel, P.: AI algorithms, price discrimination and collusion: A technological, economic and legal perspective. European Jour- nal of Law and Economics50(3), 405–435 (Dec 2020). https://doi.org/10.1007/ s10657-020-09662-6

  5. [13]

    International Journal of Environmental Research and Public Health17(23), 9041 (Jan 2020)

    Gomes-Franco, K., Rivera-Izquierdo, M., Martín-delosReyes, L.M., Jiménez- Mejías, E., Martínez-Ruiz, V.: Explaining the Association between Driver’s Age and the Risk of Causing a Road Crash through Mediation Analysis. International Journal of Environmental Research and Public ...

  6. [14]

    Goodman, E.P., Trehu, J.: ALGORITHMIC AUDITING: CHASING AI AC- COUNTABILITY. Santa Clara High Technology Law Journal39(3), 289 (May 2023), https://digitalcommons.law.scu.edu/chtlj/vol39/iss3/1 An updated and extended audit in the Italian car insurance domain 13

  7. [15]

    https: //doi.org/10.48550/arXiv.2402.08101

    Groves, L., Metcalf, J., Kennedy, A., Vecchione, B., Strait, A.: Auditing Work: Exploring the New York City algorithmic bias audit regime (Feb 2024). https: //doi.org/10.48550/arXiv.2402.08101

  8. [16]

    Standard, International Organization for Standardization, Geneva, CH (2023), https://www.iso.org/standard/78177.html

    ISO: ISO/IEC 25019:2023 - Systems and software engineering — Systems and soft- ware Quality Requirements and Evaluation (SQuaRE) — Quality-in-use model. Standard, International Organization for Standardization, Geneva, CH (2023), https://www.iso.org/standard/78177.html

  9. [17]

    Standard, International Organization for Standardization, Geneva, CH (2023), https://www.iso.org/standard/80655.html

    ISO: ISO/IEC 25059:2023 - Software engineering — Systems and software Qual- ity Requirements and Evaluation (SQuaRE) — Quality model for AI systems. Standard, International Organization for Standardization, Geneva, CH (2023), https://www.iso.org/standard/80655.html

  10. [18]

    Anno 2022

    Istituto Nazionale di Statistica: Incidenti stradali in Italia. Anno 2022. Tech. rep., ISTAT (2022), https://www.istat.it/it/archivio/286933

  11. [19]

    Business Research13(3), 795–848 (Nov 2020)

    Köchling, A., Wehner, M.C.: Discriminated by an algorithm: A systematic review of discrimination and fairness by algorithmic decision-making in the context of HR recruitment and HR development. Business Research13(3), 795–848 (Nov 2020). https://doi.org/10.1007/s40685-020-00134-w

  12. [20]

    The Lancet Digital Health4(5), e384– e397 (May 2022)

    Liu, X., Glocker, B., McCradden, M.M., Ghassemi, M., Denniston, A.K., Oakden- Rayner, L.: The medical algorithmic audit. The Lancet Digital Health4(5), e384– e397 (May 2022). https://doi.org/10.1016/S2589-7500(22)00003-6

  13. [21]

    AI and Ethics2(1), 233–245 (Feb 2022)

    Malek, M.A.: Criminal courts’ artificial intelligence: The way it reinforces bias and discrimination. AI and Ethics2(1), 233–245 (Feb 2022). https://doi.org/10.1007/ s43681-022-00137-9

  14. [22]

    Science366(6464), 447–453 (Oct 2019)

    Obermeyer, Z., Powers, B., Vogeli, C., Mullainathan, S.: Dissecting racial bias in an algorithm used to manage the health of populations. Science366(6464), 447–453 (Oct 2019). https://doi.org/10.1126/science.aax2342

  15. [23]

    In: Proceed- ings of the 2019 AAAI/ACM Conference on AI, Ethics, and Society

    Raji, I.D., Buolamwini, J.: Actionable Auditing: Investigating the Impact of Pub- licly Naming Biased Performance Results of Commercial AI Products. In: Proceed- ings of the 2019 AAAI/ACM Conference on AI, Ethics, and Society. pp. 429–435. AIES ’19, Association for Computing M...

  16. [24]

    Communications of the ACM66(1), 101–108 (Dec 2022)

    Raji, I.D., Buolamwini, J.: Actionable Auditing Revisited: Investigating the Im- pact of Publicly Naming Biased Performance Results of Commercial AI Products. Communications of the ACM66(1), 101–108 (Dec 2022). https://doi.org/10.1145/ 3571151

  17. [25]

    Rini van Solingen, Basili, V., Caldiera, G., Rombach, H.D.: Goal Question Metric (GQM) Approach. In: J.J. Marciniak (ed.) Encyclopedia of Software Engineering. John Wiley & Sons, Ltd, USA (2002). https://doi.org/10.1002/0471028959.sof142

  18. [26]

    Proceedings of the ACM on Human-Computer Interaction 5(CSCW2), 433:1–433:29 (Oct 2021)

    Shen, H., DeVos, A., Eslami, M., Holstein, K.: Everyday Algorithm Auditing: Un- derstanding the Power of Everyday Users in Surfacing Harmful Algorithmic Be- haviors. Proceedings of the ACM on Human-Computer Interaction 5(CSCW2), 433:1–433:29 (Oct 2021). https://doi.org/10.1145/3479577

  19. [27]

    The New York City Council: A Local Law to amend the administrative code of the city of New York, in relation to automated employment decision tools (2021), https://legistar.council.nyc.gov/LegislationDetail.aspx?ID=4344524& GUID=B051915D-A9AC-451E-81F8-6596032FA3F9&Options=ID%...

  20. [28]

    In: Proceedings of the 1st ACM Confer- ence on Equity and Access in Algorithms, Mechanisms, and Optimization

    Vecchione, B., Levy, K., Barocas, S.: Algorithmic Auditing and Social Justice: Lessons from the History of Audit Studies. In: Proceedings of the 1st ACM Confer- ence on Equity and Access in Algorithms, Mechanisms, and Optimization. pp. 1–9. 14 M. Rondina et al. EAAMO ’21, Asso...

  21. [29]

    Government Infor- mation Quarterly 38(4), 101619 (Oct 2021)

    Vetrò, A., Torchiano, M., Mecati, M.: A data quality approach to the identification of discrimination risk in automated decision making systems. Government Infor- mation Quarterly 38(4), 101619 (Oct 2021). https://doi.org/10.1016/j.giq.2021. 101619

Pith tools

Reviewed August 8, 2026 · model on record in the stance chip above.