{"id":"732bc7e7-d9ae-4442-89dd-1c5d5057fa76","arxiv_id":"2509.00109","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A systematic review of 24 feedback-loop-aware bias mitigation studies yields a six-dimension taxonomy for recommender systems, revealing that only six studies report both fairness and performance.","lead":"This paper reviews 24 studies on bias mitigation in recommender systems that were tested over multiple retraining rounds, and builds a six-dimension taxonomy for them. It provides a practical checklist for choosing feedback-loop-aware mitigation methods and maps the field's open gaps, such as missing shared simulators.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Abstract's 'only six studies report both fairness and performance' conflicts with Table 2, which codes seven Evaluation Focus 'Both' entries (7/24 = 29%, matching §4.6).","rationale":"The reader's weakest assumption is search completeness, which is a legitimate risk already acknowledged in the paper's limitations. But the more immediate, load-bearing problem is internal inconsistency: the abstract's 'only six use both' is contradicted by the paper's own Table 2 and by the §4.6 percentage. This is not a matter of coverage or external databases; it is a checkable arithmetic/coding discrepancy in the central result. Because the review's value rests on a verified map of the 24 studies, a headline count that does not match the coding table undermines confidence in the taxonomy's reliability until resolved. The taxonomy itself and the 24-study set may well survive a corrected count, and the paper is transparent about search limitations, so I do not see grounds to reject or to move beyond the reader's conditional verdict. The concern strengthens the case for the requested corrections rather than changing the verdict category.","tokens_in":10941,"tokens_out":7495,"duration_ms":70680,"concrete_test":"Independently recount the Evaluation Focus column for all 24 rows in Table 2, cross-checking each 'Both' entry against the study's reported metrics in the online appendix. If exactly seven studies are 'Both,' correct the abstract to 'seven' and keep §4.6 at 29%; if exactly six, identify the miscoded Table 2 row and correct the table and §4.6. Separately, recompute the screening flow from raw query results to reconcile the 150/158/216 numbers with Figure 2, confirming that the final 24-study set is reproducible.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract foregrounds the finding that 'most studies report either fairness or performance; only six use both.' This is not supported by the paper's own coded data. In Table 2, the Evaluation Focus column is 'Both' for seven studies: [18], [11], [6], [13], [10], [23], and [33]. Section 4.6 says 29% of studies investigated both fairness and performance, and 7/24 = 29.2%; the abstract's 6/24 would be 25%. So either the abstract count is wrong, or one of the Table 2 codes is wrong, or the §4.6 percentage is wrong. This matters because the 'only six' figure is used as evidence of a research gap and is part of the headline contribution, not a minor detail. A systematic review's central descriptive claims must be internally consistent with its coding table. The same verification should also resolve the screening-count inconsistencies (150 human / 158 LLM / 216 full-text vs. the Figure 2 numbers), which currently prevent an independent check of the study-selection arithmetic.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript presents a systematic literature review (SLR) of bias mitigation techniques for recommender systems that explicitly account for AI feedback loops and are evaluated in multi-round simulations or live A/B tests. The authors screen 347 records from ACM Digital Library, IEEE Xplore, and arXiv, supplemented by papers from Pagan et al.'s feedback-loop classification, and retain 24 primary studies published between 2019 and 2025. Each study is coded on six dimensions: mitigation technique, biases addressed, dynamic testing set-up, evaluation focus, application domain, and ML task, resulting in a proposed taxonomy (Figure 5 and Table 2). The paper reports that in-processing methods dominate, that A/B tests tend to use performance-only metrics, and that a small number of studies evaluate both fairness and performance. The authors also discuss limitations, including the restricted database choice and the absence of widely used simulation benchmarks.","tokens_in":11268,"tokens_out":6110,"duration_ms":50035,"significance":"If the results are accepted after the inconsistencies below are resolved, the review would make a useful contribution by consolidating a small but growing literature on feedback-loop-aware bias mitigation. Its strengths include a transparent reporting of search strings and access dates, a two-stage screening procedure with a second evaluator, a detailed coding table (Table 2) that allows independent checking of the taxonomy mapping, and the use of an established taxonomy-development method (Nickerson et al.). The paper's descriptive findings, such as the dominance of in-processing interventions and the rarity of studies reporting both fairness and performance, are directly informative for both researchers and practitioners. The internal inconsistencies identified below, however, currently undermine the reliability of the headline numbers and the reproducibility of the selection process, so the contribution is best viewed as provisional until those points are corrected.","major_comments":[{"comment":"The abstract states that 'only six' of the 24 studies report both fairness and performance, but this contradicts the paper's own coding. In Table 2, the Evaluation Focus column is 'Both' for seven studies: [18], [11], [6], [13], [10], [23], and [33]. Section 4.6 reports that 29% of studies investigated both, and 7/24 = 29.2%, which matches Table 2 but not the abstract's 'six'. Since this figure is used as evidence of a research gap and appears in the headline contribution, the abstract must be corrected to 'seven' or the table and percentage must be revised to be consistent with a count of six.","section":"Abstract and §4.6 / Table 2"},{"comment":"The screening counts reported in the text are not consistent with the flowchart. The text says the first screening 'led to 150 papers', the LLM screening 'identified 158 papers' with '93 overlapping with our initial set', and 'This yielded 216 non-duplicate papers'. However, 150 + 158 - 93 = 215, not 216, and the numbers surviving screening stage 1 in Figure 2 are 202 (ACM), 8 (IEEE), 2 (arXiv), and 4 (Pagan et al.), which sum to 216. The relationship among the 150 human-screened papers, the 158 LLM-identified papers, and the 216 papers in Figure 2 is not explained. Please reconcile these numbers or provide a detailed breakdown so the study-selection arithmetic can be independently checked.","section":"§3 and Figure 2"},{"comment":"The inclusion criteria are applied inconsistently with respect to preprints. The text states that 'Only research papers from conferences, workshops, and journals were considered; extended abstracts, and posters were excluded', yet later an arXiv search is added 'to also include the most up-to-date research', Figure 2 explicitly labels 'Pre-Prints and studies included through other methods', and Table 2 includes arXiv preprints such as [23] and [33]. This contradiction affects the systematic nature of the review; please clarify whether preprints are within scope, and if so, adjust the stated inclusion criteria and discuss the implications for the review's conclusions.","section":"§3, Figure 2, Table 2"}],"minor_comments":[{"comment":"The caption says the table maps papers 'across our seven core dimensions', but the paper consistently describes a six-dimensional taxonomy; the Paper identifier column is not a taxonomy dimension. Please change 'seven' to 'six'.","section":"Table 2 caption"},{"comment":"The percentages (54%, 17%, 29%) correspond to 13, 4, and 7 studies, respectively. Reporting the raw counts alongside the percentages would prevent the ambiguity that led to the abstract's 'six' versus 'seven' discrepancy.","section":"§4.6"},{"comment":"The figure label 'liDimensions' appears to be a typographical artifact and should read 'Dimensions'.","section":"Figure 5"},{"comment":"The sentence 'This led to 150 papers' is ambiguous because it is not clear whether this number refers only to the human screening of the ACM database or to the combined screening of all sources; please clarify.","section":"§3"},{"comment":"The mapping in Figure 1 shows four bias types, but the text says the taxonomy focuses on four of Suresh and Guttag's classifications; the relationship between Figure 1 and the 'Depends on bias' category in Table 2 could be stated more explicitly.","section":"§2.1 and Figure 1"}],"recommendation":"major_revision","confidential_remarks":"The paper addresses a relevant and timely topic, and the basic structure of the SLR is sound. However, the abstract's headline number contradicts the paper's own coding table, the screening arithmetic is internally inconsistent, and the inclusion criteria regarding preprints are ambiguous. These are load-bearing issues for a systematic review, but they appear fixable with careful revision; I would encourage the editor to request a revised version that reconciles all reported counts and criteria."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is a solid, workmanlike systematic review. Its main contribution is a six-dimensional taxonomy for bias mitigation in recommender systems evaluated under feedback loops, with four new mitigation subclasses (Causal Inference, Learning Approach, Learning Problem, Add-On) that go beyond the standard pre/in/post schema. The aggregate observations—most studies use simulation over A/B tests, in-processing dominates, shared simulators are scarce—are genuinely useful for practitioners and researchers. The method is mostly transparent: search strings, access dates, a two-stage screening process, and a second evaluator for final selection. Credit where due: this is honest, useful progress in the recommender fairness subfield.\n\nThe soft spots are real but addressable. The most serious is an internal contradiction in a central descriptive claim. The abstract says 'only six use both' fairness and performance. Your own Table 2 codes seven studies with Evaluation Focus 'Both' ([18], [11], [6], [13], [10], [23], [33]), and Section 4.6 gives 29% (7/24). That is a direct inconsistency in the headline finding, not a minor typo. Either the abstract is wrong or the coding table is wrong; a systematic review needs its central numbers to line up.\n\nSecond, the screening counts do not reconcile. The text reports 150 human-screened papers, 158 from the LLM, 216 non-duplicate full texts, and 24 selected, but the Figure 2 flowchart appears to use different numbers (e.g., 202, 180). Independent verification of the selection arithmetic is currently hard. Third, the database restriction to ACM, IEEE, and arXiv is acknowledged as a limitation, but the LLM-assisted screening step is not validated—no measure of agreement with human screening beyond overlap. These are fixable weaknesses, not fatal ones.\n\nNet: the taxonomy is a useful organizing device and the review maps a young field accurately. I would send this to peer review, because the contribution is real and the flaws are correctable. A careful referee should require the count and flowchart to be reconciled and the LLM screening step to be reported with precision.\n\nFor a reading group, it is a good example of taxonomy-building in applied ML fairness, worth a slot. I would not cite the current version until the internal numbers are fixed.","headline":"A competent SLR with a genuinely useful taxonomy, but its headline claim about 'only six studies' is contradicted by its own Table 2's seven 'Both' entries—fix that before relying on the synthesis.","tokens_in":11620,"tokens_out":2769,"would_cite":false,"duration_ms":24245,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A systematic review maps 24 studies that test bias mitigation inside recommender feedback loops","keywords":["systematic literature review","bias mitigation","recommender systems","AI feedback loops","bias amplification","taxonomy","fairness evaluation","dynamic evaluation"],"falsifier":"Run the same inclusion criteria over additional bibliographic databases and include extended abstracts and posters; if this yields more than a handful of additional studies that test bias mitigation under multi-round retraining and report both fairness and performance, the claimed scarcity of 6 out of 24 would be shown to be a search artifact.","tokens_in":1562,"feed_emoji":"🔁","tokens_out":2202,"duration_ms":70467,"temperature":0.7,"pith_summary":"This paper tries to establish a reliable map of what is actually known about correcting recommender-system bias when the system keeps learning from its own outputs. It argues that most bias-mitigation research is tested on static data and therefore says little about long-term fairness, while 24 studies do test mitigation under repeated retraining in simulations or live A/B tests. If the map is right, practitioners gain a six-dimensional checklist for choosing mitigation methods and researchers gain a list of the field's most urgent gaps, most notably the scarcity of shared simulators and the fact that only six studies measure fairness and performance together.","feed_headline":"Only 24 studies test bias fixes inside AI feedback loops","feed_subtitle":"A review of 347 papers finds only 24 evaluate bias fixes under retraining, and only six measure fairness and performance together.","key_machinery":"The carrying object is the six-dimensional taxonomy, built with the iterative conceptual-to-empirical method of [30]. It starts from the pre-, in-, and post-processing pipeline stages of [7] and adds recommender-specific classes such as Causal Inference-based, Learning Approach, Learning Problem, and Add-On. The taxonomy's work is to make 24 heterogeneous studies comparable on a common grid and to expose where the field is thin. The central phenomenon it organizes is the ML Model feedback loop, in which a system retrains on the very instances its own predictions caused to be observed, such as only recommended items receiving user feedback.","core_discovery":"The central claim is a field-level map: screening 347 records yields 24 primary studies published between 2019 and 2025 that explicitly address ML Model feedback loops and evaluate bias mitigation in multi-round simulation or live A/B tests. The paper organizes these studies along six dimensions—mitigation technique, bias addressed, dynamic testing setup, evaluation focus, application domain, and ML task—and reports that 17 of 24 interventions are in-processing, that A/B tests mostly use performance metrics, and that only six studies report both fairness and performance. This is presented as evidence that the field lacks shared simulators, standardized metrics, and combined fairness-and-performance evaluation.","pith_inferences":["A natural extension beyond the paper would be to use the six studies that report both fairness and performance as seeds for a standardized evaluation protocol, fixing the number of simulation rounds and the metric set so future mitigation claims can be compared directly.","The fact that A/B tests almost always report performance metrics suggests that live industry settings monitor fairness informally or not at all; if that is right, the next bottleneck is measurement practice, not mitigation algorithms.","Since the paper notes that applying the 90% static-evaluation estimate from [25] to other biases would exclude most studies, the field's apparent shortfall could shrink as evaluation practice improves; the taxonomy would then need richer classes for non-simulation validation."],"forward_implications":["If the taxonomy holds, practitioners can classify any new feedback-loop-aware mitigation study quickly and compare it against the 24 existing studies on six dimensions.","The 24 studies become a baseline: claims that a bias mitigation method works under feedback loops should be tested in simulations or A/B tests with multiple retraining rounds, not on static splits.","The dominance of in-processing interventions (17 of 24) suggests that future mitigation work will likely concentrate on learning algorithms, loss functions, and reward design rather than on data or output fixes.","The absence of shared simulators, if real, means cross-study comparability is currently limited, and the taxonomy can serve as a specification for the shared benchmarking environment the field lacks.","Because only six studies report both fairness and performance, current dynamic evaluations do not yet show whether fairness gains come at an acceptable performance cost."],"supporting_citations":[{"why":"Defines ML Model feedback loops, the phenomenon the review uses as its central inclusion criterion.","marker":"[31]"},{"why":"Supplies the bias categories (representation, measurement, historical, evaluation) used to code the studies.","marker":"[35]"},{"why":"Provides the pre-, in-, and post-processing mitigation classes that the taxonomy extends.","marker":"[7]"},{"why":"Gives the iterative method used to build the taxonomy.","marker":"[30]"},{"why":"Is the prior popularity-bias survey used for comparison and for the estimate that around 90% of studies lack dynamic evaluation.","marker":"[25]"},{"why":"Offers the RS bias taxonomy that the review distinguishes its own scope from.","marker":"[9]"},{"why":"Demonstrates that mitigation can fail over many retraining rounds, motivating the dynamic-evaluation criterion.","marker":"[1]"},{"why":"Exemplifies a dynamic fairness framework with model updates, cited among recent shifts toward dynamic bias auditing.","marker":"[39]"},{"why":"Exemplifies a fairness-constrained dynamic recommender framework cited alongside FADE.","marker":"[16]"}],"fun_headline_variants":["Just 24 of 347 AI bias papers handle feedback loops","AI bias fixes rarely tested under retraining loops","Only 6 AI bias studies track fairness and performance","Bias mitigation in AI loops: a 24-study reality check","AI feedback loops expose bias fix gaps: 24 studies"],"cache_read_input_tokens":13824,"weakest_assumption_plain":"The map's validity depends on the literature search being complete: if relevant studies live in databases, formats, or venues the search did not cover, including other databases and excluded extended abstracts and posters, then the reported gaps could be artifacts of the search rather than features of the field.","fun_headline_variants_meta":{"raw":{"variants":["Just 24 of 347 AI bias papers handle feedback loops","AI bias fixes rarely tested under retraining loops","Only 6 AI bias studies track fairness and performance","Bias mitigation in AI loops: a 24-study reality check","AI feedback loops expose bias fix gaps: 24 studies"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00052,"raw_usage":{"total_tokens":2472,"prompt_tokens":853,"completion_tokens":1619,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":469,"completion_tokens_details":{"reasoning_tokens":1548}},"tokens_in":469,"tokens_out":1619,"duration_ms":10299,"temperature":1.0,"reasoning_tokens":1548,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T16:42:14.861759+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same inclusion criteria over additional bibliographic databases and include extended abstracts and posters; if this yields more than a handful of additional studies that test bias mitigation under multi-round retraining and report both fairness and performance, the claimed scarcity of 6 out of 24 would be shown to be a search artifact.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the bias categories (representation, measurement, historical, evaluation) used to code the studies."},{"cited_title":"Enhancing New-item Fairness in Dynamic Recommender Systems","cited_arxiv_id":"2504.21362","evidence_quote":"Exemplifies a fairness-constrained dynamic recommender framework cited alongside FADE."}],"review_version":2}