{"id":"4948c195-7887-4c9b-a587-aba5c58c5bf7","arxiv_id":"2508.00908","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"A 40-stakeholder interview study links job seekers', recruiters', companies', and job portal staff's fairness concerns to fairness metric categories for candidate recommendation.","lead":"This paper maps how job seekers, recruiters, companies, and job portal staff define unfairness in algorithmic hiring and how those definitions relate to fairness metrics for candidate recommendation. It gives hiring platforms and auditors a multi-stakeholder frame for choosing fairness measures.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central claim depends on a valid, reproducible mapping from 40 interview accounts to fairness metric categories; nothing readable supports that the mapping is not an interpretive construction.","rationale":"The reader's weakest assumption identifies the same load-bearing point: the paper's contribution is a stakeholder-derived mapping, and that mapping is only trustworthy if the coding and mapping steps are validated rather than being an artifact of the researchers' interpretive choices. My review cannot inspect the methods because the supplied full text is garbled, but the abstract alone is enough to locate the risk. The test I propose directly checks whether the mapping is reproducible from the qualitative data; a negative result would mean the central claim is unsupported, while a positive result would clear the concern. Since the reader already marked the paper UNVERDICTED on grounds of unreadability, my concern does not move the verdict; it sharpens the reason the verdict should stay UNVERDICTED pending a readable manuscript and evidence of coding validity. I find no evidence of fraud or bad faith, and I am not treating the paper's qualitative methodology as invalid merely because it is not quantitative; the concern is specifically about the chain from interview transcripts to formal metric categories.","tokens_in":15506,"tokens_out":1830,"duration_ms":22589,"concrete_test":"Obtain a readable version of the manuscript and locate the interview coding and mapping methods. Then take a random sample of 8-10 anonymized interview excerpts or the paper's own coded units, blind them, and have two independent coders apply the reported codebook to assign each excerpt to the paper's fairness metric categories; compute Cohen's kappa. Also check whether the paper reports saturation or respondent validation. If kappa is below 0.6, or the re-coded excerpts do not reproduce the published mapping, the central stakeholder-grounded mapping is not established. If no codebook or mapping table appears in the methods, that itself settles the concern.","verdict_should_be":"UNCHANGED","load_bearing_attack":"To make good on the paper's central claim, the authors must show that the mapping from stakeholder experiences to existing fairness metric categories is demonstrably grounded in the interview data rather than imposed by the researchers. The abstract says the interviews were used to \"co-design definitions of fairness as well as metrics\" and that the authors \"attempt to reconcile and map\" these perspectives to metric categories. The supplied full text is unreadable, so no coding scheme, inter-coder reliability check, saturation analysis, or member-checking step can be inspected. The load-bearing point is not that qualitative research must use a single fixed method; it is that this paper's contribution is precisely a mapping, and without evidence of coding validity or stakeholder confirmation the mapping could be an artifact of the authors' interpretive choices. Because this is the only substantive bridge between interview data and the concluding metric mapping, it is the weakest link in the argument.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper addresses multi-sided fairness in candidate recommendation for algorithmic hiring. It reports semi-structured interviews with 40 stakeholders—job seekers, companies, recruiters, and job portal employees—to explore lived experiences of unfairness, co-design fairness definitions and metrics, and map these to existing fairness metric categories. The abstract is readable, but the body text supplied for review is corrupted or otherwise unreadable, so almost none of the methodological detail, mapping tables, or analysis can be inspected. The claimed contribution is a stakeholder-grounded mapping from experienced unfairness to concrete metric categories, which is plausible but currently unsupported by the available text.","tokens_in":15658,"tokens_out":3243,"duration_ms":35014,"significance":"The topic is timely and important: algorithmic hiring is a high-risk AI application under the EU AI Act, and prior fairness analyses have focused on single-side fairness. A rigorous multi-stakeholder mapping grounded in 40 interviews would be a useful contribution to FATE research and to practitioners. The paper does not appear to ship machine-checked proofs or reproducible code; its strength would lie in the qualitative methodology and the defensibility of the interview-to-metric mapping. Credit is due for the multi-stakeholder design and the explicit attempt to reconcile conflicting perspectives. However, that value is conditional on evidence that is currently unreadable.","major_comments":[{"comment":"The body of the manuscript as provided is in an unreadable encoding; none of the interview protocol, participant demographics, coding scheme, inter-coder reliability, saturation analysis, or reconciliation procedure can be inspected. Because the paper's central claim is a mapping from stakeholder interviews to fairness metric categories, this is not a cosmetic issue: the load-bearing evidence is inaccessible.","section":"Full Text (entire body)"},{"comment":"The abstract claims that interviews were used to 'co-design definitions of fairness as well as metrics' and to 'reconcile and map' these to existing metric categories. The visible text provides no trace of the coding or mapping procedure, so the reader cannot determine whether the mapping is grounded in the data or imposed by the researchers. The paper needs to present, in readable form, the coding scheme, example quotes aligned to codes, and the explicit mapping table from codes to metric categories.","section":"Abstract"},{"comment":"The fragments that appear to be tables (e.g., the matrix-like blocks after the abstract) are not legible. Consequently the paper cannot be checked for whether 40 interviews, distributed over the four stakeholder groups, support the breadth of the claimed fairness concerns and metric coverage. The paper should include a readable table of stakeholder counts, code frequencies, and the mapping with confidence or agreement measures.","section":"Tables and figures (garbled)"},{"comment":"No readable limitations statement is visible; if one exists in the corrupted portion, it is inaccessible. The interpretive nature of qualitative-to-metric mapping needs explicit discussion, including saturation, researcher positionality, and the extent to which stakeholders confirmed the final mapping.","section":"Limitations (if present)"}],"minor_comments":[{"comment":"The abstract uses 'we attempt' to describe the reconciliation; the final manuscript should state the method and success criteria more precisely.","section":"Abstract"},{"comment":"The references section cannot be read; the manuscript should ensure all prior work on multi-stakeholder fairness and algorithmic hiring is cited correctly.","section":"References"},{"comment":"There are no visible figure or table numbers; once the encoding is fixed, all tables and figures need clear captions and in-text references.","section":"Full Text"}],"recommendation":"major_revision","confidential_remarks":"The supplied text appears to be corrupted by an encoding conversion, making substantive review impossible. I recommend the editors request a clean, readable version before further review. Based on the abstract alone, the topic fits the journal, and there is no sign of a fatal flaw, but the central mapping cannot be assessed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Know this before anything else: the abstract describes a plausibly useful piece of work, but the full text we were sent is unreadable. It is garbled throughout, so no one can evaluate the methods, the results, or the mapping tables. The paper cannot be reviewed in this state.\n\nWhat the abstract actually offers: a multi-stakeholder fairness analysis for candidate recommendation in algorithmic hiring, based on semi-structured interviews with 40 stakeholders (job seekers, recruiters, companies, job portal employees). The authors co-design fairness definitions and metrics with participants, then try to reconcile those views and map them to existing fairness metric categories. That is a real gap: most fairness work in hiring focuses on single-side fairness, and multi-stakeholder recommender fairness is usually studied outside hiring. The direction is right, and the EU AI Act hook gives it practical relevance.\n\nThe concern is the mapping itself. The whole contribution is the translation from lived experience to metric categories. The abstract says 'attempt to reconcile and map'—that is interpretive work. To trust it, you need to see the coding scheme, the interview protocol, how the co-design sessions were structured, whether there was any inter-coder agreement or member checking, and how conflicts between stakeholder definitions were resolved. None of that is visible here. The stress-test note is fair: the load-bearing step is exactly what we cannot inspect.\n\nThe abstract also claims this is the first interview-driven co-design mapping for candidate recommendation in hiring. That is a strong claim and needs careful positioning against the existing multi-stakeholder fairness literature. Again, not checkable.\n\nNone of this is fatal to the underlying idea. The sample size is decent, the stakeholder groups make sense, and making fairness metrics accountable to affected people is worth serious attention. If the body delivers on the abstract, this deserves proper peer review.\n\nMy recommendation: do not send this to referees as-is. They cannot read it. Ask the authors for a clean, complete manuscript and resubmission. If they provide one, I would be glad to see it refereed.","headline":"Potentially useful multi-stakeholder fairness mapping, but the supplied full text is unreadable, so the paper cannot be evaluated as submitted.","tokens_in":16127,"tokens_out":3250,"would_cite":false,"duration_ms":32274,"reading_group":"no","serious_thinker":"unclear","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Fairness in algorithmic hiring is multi-sided, and the paper maps 40 stakeholders' lived experiences of unfairness to concrete fairness metric categories for candidate recommendation.","keywords":["algorithmic hiring","candidate recommendation","multi-stakeholder fairness","fairness metrics","human-in-the-loop","semi-structured interviews","co-design","job portal"],"falsifier":"An independent research team would code the same 40 interview transcripts against the paper's metric-category scheme; if the two codings agree only weakly on which stakeholder statements map to which metric categories, the mapping is not a stable property of the stakeholders' expressed concerns.","tokens_in":15332,"feed_emoji":"⚖️","tokens_out":7767,"duration_ms":73992,"temperature":0.7,"pith_summary":"Algorithmic hiring systems that rank candidates for recruiters have mostly been audited for fairness from a single viewpoint, typically the job seeker's. This paper argues that fairness in such systems is multi-sided: job seekers, companies posting jobs, recruiters using the system, and the recruitment platform itself all have legitimate fairness concerns. Based on semi-structured interviews with 40 stakeholders across those four groups, it co-designs definitions of fairness and candidate metrics for each group, then maps these definitions onto existing categories of fairness metrics. If the mapping holds, fairness evaluation of a candidate recommender can be broadened from one-sided parity checks to a multi-sided set of measurable criteria that reflect the lived experiences of all parties.","feed_headline":"40 interviews yield a multi-sided fairness map for hiring AI","feed_subtitle":"Lived unfairness from job seekers, recruiters, employers, and platforms becomes checkable metric categories.","key_machinery":"The central object is the mapping from stakeholder fairness definitions to existing categories of fairness metrics, generated through semi-structured interviews with 40 stakeholders (job seekers, companies, recruiters, and job portal employees). The interviews are used to co-design fairness definitions and candidate metrics; the paper then reconciles and maps these definitions onto existing fairness metric categories suited to a candidate recommender system, defined here as a system that recommends relevant candidate CVs to human recruiters in a human-in-the-loop hiring scenario. The mapping itself is the load-bearing mechanism: it converts lived experiences of unfairness into concrete, testable metric categories, thereby turning multi-stakeholder fairness from a principle into an evaluation checklist.","core_discovery":"Past analyses of fairness in algorithmic hiring have been restricted to single-side fairness, typically checking whether a recommender treats job seekers equally across protected groups. This paper claims that candidate recommendation is a multi-stakeholder problem: job seekers, the companies posting jobs, the recruiters who use the system, and the recruitment agency or job portal itself all have fairness interests that a fairness evaluation should reflect. The authors conducted semi-structured interviews with 40 stakeholders from these four groups, used the interviews to explore lived experiences of unfairness, co-designed definitions of fairness and metrics that might capture those experiences, and then attempted to reconcile and map these different and sometimes conflicting perspectives to existing categories of fairness metrics relevant to a human-in-the-loop candidate recommender that shows candidate CVs to human recruiters. The central claim is that a stakeholder-grounded, multi-sided fairness mapping is feasible: the concerns of all four stakeholder groups can be expressed as a set of existing fairness metric categories, so fairness in algorithmic hiring becomes a multi-sided evaluation rather than a one-sided parity check.","pith_inferences":["A natural next step beyond the paper would be to implement the mapped metric categories and measure how much each one changes actual candidate rankings, since the paper stops at the mapping itself.","If the mapping is validated on fresh data, the same interview-to-metric pipeline could be transferred to other multi-stakeholder recommenders such as news ranking or marketplace matching, where fairness concerns also differ across sides.","The conflicting stakeholder definitions suggest that fairness here is best treated as a multi-objective problem, so system design could be framed as constrained optimization across the mapped metric categories rather than selection of a single global metric.","An independent research team re-coding the same 40 transcripts against the paper's metric categories would provide a reliability check that the mapping is not an artifact of the original researchers' interpretive choices."],"forward_implications":["Fairness audits of hiring recommenders can expand from job-seeker parity to a multi-sided checklist that also covers recruiter, organization, and platform concerns.","The mapping gives system designers a concrete starting set of metric categories to implement and monitor when building or evaluating human-in-the-loop candidate ranking.","Where stakeholder fairness definitions conflict, the mapping exposes the trade-offs explicitly, so choosing among metrics becomes a visible design decision rather than a hidden default.","In the EU AI Act context, the mapped metric categories offer one way to operationalize high-risk fairness requirements for candidate recommendation."],"supporting_citations":[],"fun_headline_variants":["40 interviews map fairness needs to hiring AI metrics","Multi-sided fairness map for hiring AI from 40 interviews","Hiring AI fairness: mapping 40 stakeholder voices to metrics","From 40 stakeholder interviews to a multi-sided hiring fairness map"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that fairness concerns voiced by 40 interviewed stakeholders can be reliably translated into quantitative metric categories, so that the resulting mapping reflects the stakeholders' own views rather than the researchers' interpretive construction.","fun_headline_variants_meta":{"raw":{"variants":["40 interviews map fairness needs to hiring AI metrics","Multi-sided fairness map for hiring AI from 40 interviews","Hiring AI fairness: mapping 40 stakeholder voices to metrics","From 40 stakeholder interviews to a multi-sided hiring fairness map"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001089,"raw_usage":{"total_tokens":4592,"prompt_tokens":1032,"completion_tokens":3560,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":648,"completion_tokens_details":{"reasoning_tokens":3493}},"tokens_in":648,"tokens_out":3560,"duration_ms":27918,"temperature":1.0,"reasoning_tokens":3493,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T12:26:48.864275+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"An independent research team would code the same 40 interview transcripts against the paper's metric-category scheme; if the two codings agree only weakly on which stakeholder statements map to which metric categories, the mapping is not a stable property of the stakeholders' expressed concerns.","supporting_citations":[],"review_version":1}