{"id":"21d4075a-036b-4a1f-bcad-ce47e610d11f","arxiv_id":"2508.21179","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"A new synthetic CV dataset, generated from donated real CVs, is proposed as a benchmark for fairness-aware algorithmic hiring research.","lead":"The authors built 1,730 fake résumés modeled on real CVs donated by volunteers, with markers like gender, age, LGBT status, minority status, and religion. The dataset is offered as a benchmark for testing whether hiring algorithms unfairly rank people by these attributes.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Validation checks only marginal distributions and human indistinguishability, not whether a fairness audit on synthetic CVs reproduces the group-level bias signal of real CVs.","rationale":"I agree with the reader's CONDITIONAL verdict and with several of the stated limitations, including circular validation and incomplete coverage of protected groups. However, the single most load-bearing concern is not primarily about the representativeness of the 1,211 donated CVs; it is about the validity of the evidence offered for the central utility claim. Even a perfectly representative reference pool would not guarantee that the synthetic corpus preserves the conditional relationships between protected attributes and CV content that determine algorithmic bias. The paper's validation establishes marginal distributional similarity and human indistinguishability, but not that a fairness audit on synthetic CVs would reproduce the bias patterns found in real CVs. That missing link is directly testable with a ranking-based fairness audit. Because the paper is explicitly positioned as a benchmarking resource, this gap supports keeping the verdict CONDITIONAL rather than accepting the claim at face value. I do not see grounds for rejection: the donation-driven, privacy-preserving design is a genuine contribution, and the authors acknowledge the coverage limitations. The concern is about the lack of evidence for the specific benchmarking use case, not about the integrity of the work.","tokens_in":21288,"tokens_out":5919,"duration_ms":67970,"concrete_test":"Compare a concrete fairness audit on both corpora. Using the 1,211 donated CVs and the 1,730 synthetic CVs, choose one protected attribute with enough data (e.g., gender) and one sector (e.g., ICT or business). Build a standard retrieval/ranking baseline (BM25 over parsed education+experience+skills text, or logistic regression on the parsed feature set) and compute a group fairness metric, e.g., the normalized difference in mean retrieval score between women and men for several job queries. Repeat exactly on the synthetic subset generated with the gender parameter. If the synthetic audit reverses the sign or changes the magnitude of the gender gap by more than 0.1 (normalized score scale) relative to the real-CV audit, the claim that synthetic CVs are 'as useful as real documents' for fairness benchmarking fails. This check targets the missing link between distributional realism and fai","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that the synthetic CVs are 'likely to be as useful as the real documents for studying bias in algorithmic hiring' and suitable as a fairness benchmark (Sec. 4.2.1). The support offered is (i) univariate Jensen-Shannon divergence between reference and synthetic distributions for sector, experience, gender, age, LGBTQ+, minority, and foreignness, and (ii) a subjective test in which crowdworkers identify real vs synthetic at 53% accuracy. Neither establishes the claimed utility. A fairness benchmark needs the joint, conditional relationship between protected attributes and the CV content rankers use—education, roles, seniority, skills—to match the real data; marginal similarity is partly circular because the generator was fit to the same reference set. The manuscript itself flags consequences of the 20-CV threshold: only six sectors remain, each synthetic CV carries one demographic attribute, disability and non-Christian religions are absent, and intersectional profiles do not exist (Secs. 3.3.2, 4.1, 5.1.2). These are acknowledged, but the unaddressed empirical question is whether an actual bias audit on the synthetic corpus would reach the same conclusions as on the donated corpus. If the generator's shuffling of institutions, clustering of items, and single-attribute conditioning attenuates or amplifies group differences in CV content, fairness metrics computed on the synthetic data will mislead rather than benchmark. The 53% Turing-style result and JS scores cannot detect this.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces a privacy-preserving pipeline for generating synthetic CVs from real donated CVs and presents a resulting dataset of 1,730 synthetic CVs across six job sectors, along with demographic attributes such as age, gender, religion, LGBTQ+ status, minority status, and perceived foreignness. The generator fits parametric distributions (e.g., Weibull) on the donated data, uses clustering to assemble coherent education/experience/skills sections, and applies manual verification. The authors claim the synthetic dataset can serve as a benchmarking standard for fairness-aware hiring research and is 'likely to be as useful as the real documents' for studying bias. Validation is based on univariate distribution comparisons using Jensen–Shannon divergence and a crowdsourced human-indistinguishability test. The paper also documents substantial limitations: only six sectors, a single demographic attribute per CV, absence of disability and non-Christian religious profiles, and replication of source-sample biases.","tokens_in":21652,"tokens_out":3508,"duration_ms":39700,"significance":"If the central claim were properly validated, this would be a valuable contribution: there is currently no publicly available CV corpus with protected attributes suitable for fairness benchmarking, and the EU AI Act explicitly encourages synthetic data for bias testing. The authors also make good-faith efforts on ethics and privacy, including ethical review, anonymization, separation of sensitive fields into intermediate tables, a minimum-cell rule for named entities, and a manual coherence review. The dataset is distributed under a license that aims to prevent re-identification. However, the validation provided does not establish that bias measurements on the synthetic CVs reproduce the bias signals of the real CVs, which is the load-bearing requirement for the paper's benchmarking claims.","major_comments":[{"comment":"The Jensen–Shannon divergence is computed between the synthetic and the reference distributions for the same variables (sector, experience, age, gender, etc.) that were used to fit the generator. This is a fidelity check against the training reference, not an independent measure of utility for fairness benchmarking. Moreover, fairness audits depend on the joint/conditional relationship between protected attributes and CV content (education, roles, seniority, skills), and univariate marginals do not capture such relationships. The claim 'likely to be as useful as the real documents' therefore goes beyond the evidence. Please add a downstream validation: run a fairness audit on both the reference and synthetic corpora and show that the group-level conclusions match.","section":"Section 4.2.1 / Table 9"},{"comment":"The subjective evaluation reports 53% overall accuracy, 0.52 probability of correctly classifying a real CV, and 0.55 probability that a synthetic CV is mistaken for real. These numbers are essentially at chance and, by themselves, indicate only that humans cannot easily distinguish the synthetic CVs; they do not show that bias-relevant signals are preserved. In addition, the statistic that 90% of synthetic CVs were labeled real by at least one of three evaluators is weak, since with three independent binary judgments a high proportion of items will receive at least one 'real' label by chance. A fairness benchmark needs to demonstrate that the synthetic data reproduce group-level differences in CV content, not merely that they look plausible.","section":"Section 4.2.2"},{"comment":"The generator's minimum of 20 matching real CVs per parameter combination restricts the synthetic dataset to six job sectors, prevents combinations of demographic attributes, and excludes disability and non-Christian religious profiles entirely. These limitations are acknowledged in Section 5.1.2, but they directly undermine the paper's framing of the dataset as a standard fairness benchmark for 'people from diverse backgrounds' (Section 1). Since the disclosed limitations mean the corpus cannot be used to study bias for several protected groups and cannot support intersectional analyses, the central claim and title-level framing should be revised, or the authors should provide evidence that the included groups and sectors are sufficient for the stated benchmarking goals.","section":"Sections 3.3.2 and 4.1"},{"comment":"The reference pool is 78% Spanish, overrepresents ICT/engineering/business/clerical sectors, and underrepresents manual-labor sectors, minority religions, and disabled people. The authors acknowledge that the synthetic corpus inherits and possibly amplifies these biases. This is not itself a flaw if the dataset is positioned as a case study, but the paper proposes the dataset as a benchmarking standard with 'representative' characteristics. The manuscript should either add evidence about how the source-sample composition affects fairness measurements, or substantially narrow the generality claims. A comparison of fairness metrics on the real donated CVs versus the synthetic CVs would also clarify how much of the original bias signal survives the generation process.","section":"Sections 3.1 and 5.1.2"}],"minor_comments":[{"comment":"Typo: 'Foreigness' should be 'Foreignness'. Also, the sentence 'The closer the score closer to 0' needs rewording, and the caption for Figure 3 should state how excluded reference sectors affect the displayed percentages and the JS score.","section":"Section 4.2.1 / Table 9"},{"comment":"The subjective test would benefit from reporting inter-annotator agreement (e.g., Fleiss' kappa) and from including attention checks beyond excluding extremely fast submissions. The current report of 90% of synthetic CVs labeled as real by at least one evaluator is not very informative, as noted above.","section":"Section 4.2.2"},{"comment":"The manual coherence review is described as performed by a single research assistant trained by the authors. Please clarify whether a second annotator was used, or whether a reliability study was conducted; otherwise the 4-out-of-5 cutoff is difficult to interpret.","section":"Section 3.2.2"}],"recommendation":"major_revision","confidential_remarks":"The paper addresses an important gap and the privacy-aware construction process is a strength. However, the validation does not yet support the central benchmarking claim; the authors should either add a direct comparison of fairness audit results on the reference and synthetic corpora, or reframe the contribution as a methodology demonstration with a more limited dataset. I do not see this as a rejection because the missing evidence is within the scope of a revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know about this paper. First, it is a genuine attempt to fill a real gap: a public fairness-oriented CV benchmark that includes sensitive attributes like LGBTQ+ status, minority status, and perceived foreignness. Second, the validation is not adequate for the paper's central claim that the synthetic CVs are 'likely to be as useful as the real documents' for studying bias. The stress-test note is right: the checks only cover univariate margins and human indistinguishability, not whether a fairness audit on synthetic data produces the same group-level conclusions as an audit on the donated CVs.\n\nThe donation campaign and privacy-preserving intermediate tables are the strongest part. Splitting the data into three unrelated tables and only feeding the generator with de-identified counts and shuffled entities is a thoughtful design. The manual coherence review is also a real effort, and the paper is unusually candid in Section 5.1.2 about the consequences of the 20-CV threshold: only six sectors survive, each CV carries a single demographic attribute, disability and non-Christian religions are absent, and intersectional profiles are impossible. These are acknowledged constraints, not hidden ones.\n\nThe soft spots are the circular JS validation and the 53% subjective accuracy. The JS scores are computed against the same reference distribution that was used to fit the Weibull parameters, so they are a measure of model fit, not an independent test of utility. And the Turing-style test telling you that crowdworkers can't tell real from synthetic does not tell you whether the synthetic CVs preserve the statistical relationship between protected attributes and the content that a ranking system actually uses (education, roles, seniority, skills). The paper would be much stronger with one simple experiment: run a well-known fairness audit (e.g., a resume-ranking model with a bias metric) on both the real and synthetic sets and compare the resulting bias estimates. Without something like that, the central claim is overreach.\n\nThere are also smaller issues: the dataset is not open (only under a restrictive license), and no code is released, which makes reproducibility hard. The citation pattern looks fine; the related work table is useful.\n\nBottom line: this deserves a serious referee. It is a resource paper with a plausible pipeline and a dataset that does fill a gap, but the authors should be asked to either strengthen the validation or temper the 'as useful as the real documents' statement. Worth a reading group discussion, but I wouldn't build on it yet.","headline":"A well-intentioned synthetic CV dataset with novel sensitive attributes, but the validation doesn't back the claim that it reproduces the bias signal real CVs would give you.","tokens_in":22094,"tokens_out":3091,"would_cite":false,"duration_ms":30073,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper presents a pipeline for generating synthetic CVs from anonymized donated resumes and a dataset of 1,730 documents that preserves demographic distributions for fairness benchmarking.","keywords":["synthetic data","algorithmic hiring","fairness","benchmarking","CV dataset","data donation","bias detection"],"falsifier":"Run the same resume-ranking algorithm on the original real donated CVs and on the synthetic CVs, then compare fairness metrics such as disparate impact by gender, age, or origin. If the synthetic data fails to reproduce the direction or magnitude of the real data's discriminatory outcomes, the 'as useful as real' claim would be falsified.","tokens_in":21215,"feed_emoji":"📄","tokens_out":4622,"duration_ms":49250,"temperature":0.7,"pith_summary":"The paper aims to fill a gap: there is no publicly available collection of CVs with rich sensitive attributes for studying algorithmic hiring discrimination. It provides a method to generate synthetic CVs from real CVs donated by consenting workers, and releases 1,730 such CVs annotated with demographic variables such as age, gender, LGBTQ+ status, minority status, and perceived foreignness. The central claim is that these synthetic CVs are likely to be as useful as real documents for studying bias in algorithmic hiring, because their distributional properties match the reference pool and people cannot reliably tell them apart. If true, this gives researchers and companies a privacy-preserving benchmark to test whether ranking algorithms discriminate by protected characteristics.","feed_headline":"1,730 synthetic CVs give researchers a testbed for hiring bias","feed_subtitle":"Built from anonymized donated CVs, the set preserves real demographic distributions while shielding donors.","key_machinery":"The core mechanism is a generator driven by a reference pool of parsed, anonymized CVs split into three tables: anonymized CVs, education-experience-skill combinations, and named entities. The generator samples section sizes from Weibull distributions fitted to real CVs with matching attributes, fills the sections by clustering education and experience items and drawing from shuffled institution lists, and enforces coherence rules (chronological order, duration caps, no duplicates). A minimum of 20 real CVs per parameter combination is required to protect donor anonymity, and a final manual coherence review keeps only CVs rated highly for plausibility.","core_discovery":"The authors claim that a dataset of 1,730 synthetic CVs, generated from 1,211 donated real CVs through a hybrid automatic-and-manual pipeline, is likely to be as useful as real CVs for studying bias in algorithmic hiring, while protecting donors' identities. They support this with distribution comparisons showing near alignment between synthetic and reference data, with Jensen-Shannon divergence scores ranging from 0.01 for gender to 0.25 for LGBTQ+ status, and with a crowdsourcing test in which evaluators could distinguish real from synthetic CVs only 53% of the time. They position the dataset as a candidate standard benchmark for detecting biased ranking functions in candidate screening.","pith_inferences":["Because each synthetic CV carries only one demographic attribute, intersectional analyses (e.g., minority women, older LGBTQ+ applicants) cannot be run directly on this dataset; future versions or pooled parameters would be needed.","The privacy threshold of 20 matching CVs per parameter combination silently excludes small or underrepresented groups, so the dataset's coverage of protected attributes should itself be audited before it is used as a benchmark.","The indistinguishability result suggests a testable extension: comparing human screening of synthetic versus real CVs could reveal whether perceived authenticity affects reviewer decisions, not just ranking algorithms.","The 'as useful as real' claim could be quantified downstream by comparing fairness metrics from the same ranking pipeline run on synthetic and real CVs."],"forward_implications":["Fairness-aware ranking algorithms can be evaluated before deployment against a common corpus without exposing real applicants.","Researchers at different institutions can reproduce and compare bias-detection results because the dataset is accessible under a license agreement.","The generation methodology can be reused with other donation pools to build workforce-specific benchmarking datasets.","The synthetic approach aligns with the EU AI Act's recommendation to prefer synthetic data for detecting and mitigating bias."],"supporting_citations":[{"why":"Supplies the reference dataset of donated real CVs and describes the data donation campaign.","marker":"[51]"},{"why":"Provides the univariate distribution comparison method used to validate synthetic data utility.","marker":"[15]"},{"why":"Motivates the subjective evaluation design for testing whether synthetic CVs are indistinguishable from real ones.","marker":"[28]"},{"why":"Provides the Weibull distribution used to model the number of items in each generated CV section.","marker":"[49]"},{"why":"Used to map between companies and roles and between degrees and institutions when filling synthetic CV content.","marker":"[48]"},{"why":"Provides the agglomerative clustering algorithm used to group education, experience, and skills items.","marker":"[54]"},{"why":"Cites the EU AI Act's article on synthetic data as the regulatory motivation for this approach.","marker":"[7]"}],"fun_headline_variants":["Synthetic CVs benchmark hiring-bias detection tools","1,730 fake CVs train fairer hiring algorithms","Donated real CVs spawn synthetic test set","Test hiring fairness with synthetic CVs","New synthetic CV dataset targets hiring bias"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The donated pool of 1,211 CVs—mostly Spanish and overrepresenting certain sectors while underrepresenting manual-labor professions and some demographic groups—is representative and large enough that the generator's 20-CV minimum per parameter combination still produces a dataset covering the diversity needed for fairness benchmarking.","fun_headline_variants_meta":{"raw":{"variants":["Synthetic CVs benchmark hiring-bias detection tools","1,730 fake CVs train fairer hiring algorithms","Donated real CVs spawn synthetic test set","Test hiring fairness with synthetic CVs","New synthetic CV dataset targets hiring bias"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000181,"raw_usage":{"total_tokens":1121,"prompt_tokens":700,"completion_tokens":421,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":444,"completion_tokens_details":{"reasoning_tokens":351}},"tokens_in":444,"tokens_out":421,"duration_ms":4651,"temperature":1.0,"reasoning_tokens":351,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T14:31:01.201158+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same resume-ranking algorithm on the original real donated CVs and on the synthetic CVs, then compare fairness metrics such as disparate impact by gender, age, or origin. If the synthetic data fails to reproduce the direction or magnitude of the real data's discriminatory outcomes, the 'as useful as real' claim would be falsified.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the reference dataset of donated real CVs and describes the data donation campaign."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the univariate distribution comparison method used to validate synthetic data utility."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Motivates the subjective evaluation design for testing whether synthetic CVs are indistinguishable from real ones."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the Weibull distribution used to model the number of items in each generated CV section."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the agglomerative clustering algorithm used to group education, experience, and skills items."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Cites the EU AI Act's article on synthetic data as the regulatory motivation for this approach."}],"review_version":1}