{"id":"878a5fc1-73c2-4969-9c46-d599d5df4dcd","arxiv_id":"2509.00658","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A new face benchmark with four visual domains and fairness-sensitive labels provides larger measured distribution shifts and lower baseline performance than existing fairness datasets.","lead":"The paper introduces FACE4FAIRSHIFTS, a 100,000-image face benchmark spanning photo, art, cartoon, and sketch domains with 42 human-annotated attributes. It aims to give fairness and domain-generalization researchers a testbed where distribution shifts and demographic correlations are both present.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central 'harder and more stable' claim depends on unvalidated majority-voted labels; without per-domain annotation agreement or noise rates, the gaps in Tables 4/8 could be label-noise artifacts.","rationale":"The paper has real supporting evidence: extensive baseline comparisons, three-run averages, standard deviations in Table 10, and use of standard repositories like DomainBed. The central issue is not lack of experiments but that every empirical comparison in Section 3 treats Face4FairShifts' human labels as ground truth, and those labels are exactly the component whose quality is asserted but not measured. The paper's own Appendix D admits subjectivity in attributes such as 'attractiveness' and potential bias from web-crawled sources. Without per-domain inter-annotator agreement or a noise model, the observed lower accuracy and compressed baseline variance could be explained by label noise rather than by harder domain shifts or fairer benchmarking. This is not a disagreement with external consensus; it is an internal gap between the claim and the evidence. The concrete test above would settle it: a unanimous-label subset removes noise as an explanation, and per-domain kappa localizes where noise lives. If gaps persist on unanimous labels, the concern is resolved. The data-release and checklist inconsistencies noted by the reader are real but secondary; they affect reproducibility and process, not the correctness of the 'harder' claim. For these reasons, the reader's CONDITIONAL verdict remains appropriate: the benchmark is promising, but the central empirical claim should not be fully trusted until annotation quality is demonstrated.","tokens_in":29002,"tokens_out":4582,"duration_ms":58614,"concrete_test":"Perform a second, independent annotation pass on a stratified sample (e.g., 500 images per domain) with the same 42-label protocol and compute per-domain and per-attribute inter-annotator agreement (Fleiss' kappa) and label-noise rates. Then re-run the key comparisons (Tables 4 and 8, age/gender) on the subset of images with unanimous first-pass labels. If per-domain agreement is substantially lower in Art/Sketch, and the Face4FairShifts-vs-FairFace/UTK-FairFace accuracy gaps and epsilon/+- stability gaps shrink or invert on the unanimous subset, the 'harder and more stable' conclusion is partially an artifact of annotation noise; if the gaps persist on unanimous labels, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 2.3 states 66 paid annotators, at least 3 labels per image, majority voting, and a 5-person QC random review, but reports no inter-annotator agreement (e.g., Fleiss' kappa), no disagreement rate, and no per-domain or per-attribute error analysis. Appendix D explicitly concedes that attributes such as 'Attractive' are subjective and that web-crawled sources may carry sampling bias. The paper's headline empirical claims — that Face4FairShifts is 'greater challenges' and 'more stable' (Tables 3, 4, 8) — are measured against age/gender/race labels. If crawled Art and Sketch domains have systematically noisier labels (stylized faces make Age/Race/Appearance harder to judge), two artifacts follow: (1) accuracy drops and (2) all baselines converge to a common noise ceiling, shrinking the epsilon and +/- stability indicators. Those are exactly the observations used to claim the dataset is harder and baselines are more stable. The LLM comparison in Fig. 2(right) cannot validate the human labels: it uses them as ground truth, is aggregated across domains, and only shows humans and LLMs disagree at some rate. Thus the most load-bearing assumption — label quality — is unquantified, and the core empirical conclusion is not currently falsifiable from the paper as written.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces Face4FairShifts, a large facial-image benchmark for fairness and robustness under visual domain shifts. It contains 100K images across four domains (Photo, Art, Cartoon, Sketch) and is claimed to carry 42 annotations across 15 attributes covering demographic and facial features. The authors describe data collection from existing datasets and web crawling, manual annotation by 66 paid annotators with majority voting, and a 5-person QC team. They then present extensive experiments across four research areas: fairness learning, OOD generalization, OOD detection, and fairness-aware OOD generalization, comparing Face4FairShifts with FairFace and UTK-FairFace on covariate-shift magnitude, task difficulty, and cross-baseline stability. The headline empirical claims are that Face4FairShifts poses greater challenges to current baselines while yielding more stable performance across methods.","tokens_in":29339,"tokens_out":4995,"duration_ms":61470,"significance":"If the label quality and statistical claims hold, this is a potentially valuable resource for the FairOG community: it provides a large, multi-domain facial dataset with rich attribute annotations, and the t-SNE/JS-divergence analysis suggests larger natural covariate shifts than the commonly used FairFace/UTK-FairFace partitions. The breadth of the experimental evaluation—covering fairness, OOD generalization/detection, and FairOG—is a strength, and the public release of code and dataset increases usability. However, the benchmark's validity and its headline 'harder and more stable' conclusions rest on the quality of human annotations, which are not quantified, and on statistical stability indicators that lack error bars. Several internal inconsistencies (e.g., 39 vs. 42 annotations and 14 vs. 15 attributes) and a broken Figure 4 (all DI cells show NaN) further reduce confidence in the paper as written.","major_comments":[{"comment":"The paper reports majority voting and a 5-person QC random review but gives no inter-annotator agreement (e.g., Fleiss' kappa), no disagreement rates, and no per-domain or per-attribute label-noise analysis. This is load-bearing because all downstream comparisons—accuracy, fairness metrics, and the 'harder/more stable' conclusions—use these labels as ground truth. In particular, systematically noisier labels in crawled Art/Sketch domains (which Appendix D itself concedes may be subjective) could produce exactly the observed pattern: lower accuracy and compressed cross-baseline variance due to a common noise ceiling. The LLM comparison in Figure 2(right) cannot validate human labels because it treats them as ground truth and is aggregated across domains. Please report agreement statistics and, ideally, robustness checks such as training on high-agreement subsets or estimating label noise","section":""},{"comment":"The lower heatmaps in Figure 4 display the string 'NaN' in every cell of every domain. If these are the actual DI values, the paper's claimed fairness disparities are not reported and the statement that 'Figure 4 highlights substantial correlation shifts' is unsupported. If this is a rendering artifact, the figure must be regenerated. Either way, the current figure cannot be used to verify the DI-based fairness characterization, and the formula's k/1/k transformation needs a clear definition of which group is unprivileged and how zero or undefined ratios are handled.","section":""},{"comment":"The number of annotations and attributes is inconsistent throughout the manuscript: the abstract says '39 annotations within 14 attributes', while the Introduction, Figure 1, and Figure 3 describe 42 annotations across 15 attributes; Table 2 and Table 9 list 39 annotations. Since the annotation count is a core dataset property, this factual inconsistency must be resolved in the final version and all occurrences checked.","section":""},{"comment":"The stability claim rests on the ϵ and ± indicators computed from per-baseline metric values. However, only Table 10 reports standard deviations; Tables 4–8 are averages over three runs without any variance or significance information. The differences in ϵ and ± across datasets may be within run-to-run noise, especially given the small number of runs. Please report per-run spread or confidence intervals for the stability indicators, or otherwise justify that the observed stability differences are statistically meaningful.","section":""}],"minor_comments":[{"comment":"The term 'sensory OOD detection' is used without definition. Clarify that it refers to covariate/sensory-level shifts, and how it differs from semantic OOD detection.","section":""},{"comment":"The 'minimum acceptable image size' is said to be manually determined but its value is never reported, and the manual Photoshop cropping step is described only qualitatively. For reproducibility, provide the exact size threshold and the number of images discarded at each filtering stage.","section":""},{"comment":"Figure 2(right) compares LLM and human annotations only via accuracy/F1 against human labels. This demonstrates disagreement but not human correctness; report human inter-annotator agreement instead, and consider domain-wise breakdowns rather than the current aggregate.","section":""},{"comment":"The checklist answers 'NA' for 'New assets' and 'Safeguards' despite the paper introducing a new, web-crawled dataset. Provide a data sheet, license information, terms of use for crawled images, and a broader-impacts discussion for this facial dataset.","section":""},{"comment":"Typographical issue: 'no-bread' should be 'no-beard'. Also, the text says 'all 14 attributes' while the figure and the rest of the paper use 15 attributes.","section":""}],"recommendation":"major_revision","confidential_remarks":"This is a dataset-contribution paper with a useful scope, but the current submission is not verifiable because the central label-quality premise is unquantified and one of the key figures (Figure 4) appears to be broken. The fixes—adding annotation agreement statistics, repairing the figure, and reconciling annotation counts—are within the scope of a revision. I would not reject on these grounds, but the claims of 'greater challenges' and 'more stable performance' should not be accepted without the requested evidence. Also note that the NeurIPS checklist answers are not fully consistent with the paper's own contribution (e.g., 'New assets: NA')."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short take: this is a real contribution—a purpose-built facial benchmark for FairOG with natural domain shifts—but the paper currently undercuts itself with inconsistent annotation counts, an unreleased dataset, and no inter-annotator agreement numbers. The core benchmark idea is sound; the empirical claims ride on label quality that is asserted rather than demonstrated.\n\nWhat's new and good: Existing FairOG work has been making do with artificial splits (FairFace by race; UTK-FairFace as two-source). Face4FairShifts gives you four visually distinct face domains (Photo, Art, Cartoon, Sketch) with 42 human-annotated labels across 15 attributes, covering demographic and facial features. That combination doesn't exist elsewhere. The data collection effort is serious: 66 paid annotators, at least 3 labels per image, majority voting, a 5-person QC team, and 12.6M annotation instances. The t-SNE and JS-divergence analysis gives reasonable evidence that the covariate shifts are larger than in prior FairOG datasets. And the evaluation touches four relevant areas (fairness learning, OOD generalization, OOD detection, FairOG) with sensible baselines.\n\nWhere it's soft: The abstract says 39 annotations within 14 attributes; the main text says 42 within 15; Table 2 lists 39. That is sloppy for a benchmark paper. Bigger issue: no inter-annotator agreement, no per-domain label noise rates, no disagreement analysis. The stress-test note is right that stylized faces in Art and Sketch are harder to label for age/race/appearance, so if the crawled domains are noisier, the 'greater challenges and more stable baselines' results could simply be a label-noise ceiling—all methods converging to the same poor performance. The LLM comparison in Fig. 2 validates the humans against themselves as ground truth; it doesn't quantify label noise. Also, the NeurIPS checklist has direct contradictions (Q13 says no new assets; Q14 says no crowdsourcing, despite the dataset and paid annotators). The main tables lack error bars; standard deviations appear only for fairness learning in the appendix.\n\nThe benchmark is still worth building and the paper is worth engaging with, but those claims need to be re-examined once the data and annotation metadata are released. This deserves a serious referee, mostly because the field needs a dataset like this and the authors have done the hard collection work. My recommendation: send it to review, but make data release, annotation agreement, and the inconsistency fixes conditions.","headline":"A potentially useful fairness benchmark that needs its label-quality evidence and internal consistency fixed before the 'harder and more stable' claims can be trusted.","tokens_in":29830,"tokens_out":2738,"would_cite":false,"duration_ms":35168,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper introduces Face4FairShifts, a 100,000-image face benchmark with natural domain shifts and sensitive-attribute correlations, and reports that it is harder and more stable than existing fairness benchmarks.","keywords":["fairness benchmark","facial attribute dataset","domain generalization","out-of-distribution detection","correlation shift","distribution shift","visual domains"],"falsifier":"Re-annotate a random sample of, say, 1,000 images per domain with an independent annotation team and measure per-attribute agreement against the released labels. If agreement is much lower in Art and Sketch, or near chance for subjective attributes like Attractive, then the reported disparate-impact heatmaps and the claim that the dataset is harder would reflect annotation noise, not domain shift.","tokens_in":28904,"feed_emoji":"🎭","tokens_out":7581,"duration_ms":86904,"temperature":0.7,"pith_summary":"Face4FairShifts is a proposed benchmark of 100,000 facial images in four visually distinct domains—Photo, Art, Cartoon, and Sketch—each labeled with 42 human annotations across 15 attributes. The authors aim to give fairness and domain-generalization research a dataset where domain shifts and correlations between sensitive attributes and labels arise naturally, not from artificially splitting one dataset. Across four research areas, they report that state-of-the-art baselines perform worse on Face4FairShifts than on existing benchmarks, while baseline rankings stay more consistent. If the central claim holds, the benchmark provides a common, harder testbed for deciding whether fair models survive real visual-style changes and for developing methods that do.","feed_headline":"100K face images test fairness across four visual domains","feed_subtitle":"Face4FairShifts shows fair models lose ground when faces move from photos to art, cartoons, and sketches.","key_machinery":"The load-bearing object is the benchmark itself: 100,000 face images across Photo, Art, Cartoon, and Sketch, each image carrying 42 binary labels in 15 attribute groups, majority-voted from at least three annotators. The mechanism is the domain construction: different rendering styles create natural covariate shifts, while the disparate-impact heatmaps show that correlations between sensitive attributes and class labels change across domains, so a model that is fair in one style can become unfair in another. This is what makes the benchmark a testbed for fairness-aware domain adaptation rather than just another attribute dataset.","core_discovery":"Face4FairShifts is claimed to be the first facial benchmark with both genuine covariate shifts—the same attributes rendered as photos, artwork, cartoons, and sketches—and measurable shifts in sensitive-attribute/label correlations from one domain to the next. Because the domains come from separate sources and web crawls rather than from splitting one dataset, the covariate shift is larger, quantified by Jensen-Shannon divergence in feature space. In fairness learning, OOD generalization, OOD detection, and fairness-aware OOD generalization, the paper reports that baselines score lower on Face4FairShifts while their relative ordering stays more stable, which it interprets as a harder and more","pith_inferences":["Beyond the paper: the benchmark could also serve as a general stress test for face attribute classifiers, not just fairness-specific methods, since it varies rendering style while keeping identity-related attributes roughly constant.","Beyond the paper: an independent re-annotation study on a sample of images would clarify how much of the reported hardness comes from annotation noise rather than domain shift, because the paper reports no inter-annotator agreement statistics.","Beyond the paper: stratifying model performance by age or race within each domain could reveal which demographic groups lose the most accuracy when style changes, a comparison the paper does not report."],"forward_implications":["Fairness-learning baselines (LFR, GSR, AD, CSAD, FNF) all show larger demographic-parity and equalized-odds gaps on Face4FairShifts than on CelebA, UTKFace, FairFace, or UTK-FairFace.","Domain-generalization baselines (ERM, IRM, GDRO, Mixup, MMD, MBDG) score lower accuracy and F1 on the new benchmark, with smaller variance across methods, so fair-generalization results on older datasets should be rechecked here.","OOD detectors separate the four visual domains well in inter-domain sensory detection, meaning the benchmark's domain shift is genuinely visible in feature space rather than being an artificial partition.","The dataset gives fairness-aware OOD generalization methods a common testbed where correlations between sensitive attributes and labels occur naturally and vary across domains."],"supporting_citations":[{"why":"Supplies the 30,000-image Photo domain and the annotation scheme the benchmark builds on.","marker":"[48]"},{"why":"Supplies museum artwork faces for the Art domain after filtering out black-and-white images.","marker":"[40]"},{"why":"Supplies crawled portrait faces in varied artistic styles for the Art domain.","marker":"[72]"},{"why":"Supplies watercolor and oil-painting images for the Art domain.","marker":"[37]"},{"why":"Supplies most of the Cartoon domain images.","marker":"[90]"},{"why":"Supplies the 1,194 hand-drawn sketches in the Sketch domain.","marker":"[86]"},{"why":"Supplies professional artist sketches for the Sketch domain.","marker":"[24]"},{"why":"The benchmark's main comparison dataset for fairness and domain generalization.","marker":"[38]"},{"why":"Provides the semi-synthetic UTK-FairFace comparison and the FCR baseline method.","marker":"[7]"},{"why":"Provides the FEDORA baseline and the convention of treating racial groups as domains that Face4FairShifts replaces with natural style domains.","marker":"[89]"}],"fun_headline_variants":["Fairness models fail when faces turn to art and cartoons","New 100K face benchmark exposes fairness gap across domains","Faces in art and cartoons break fairness models","Domain shifts in faces unmask fairness failures"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The benchmark's validity rests on majority-voted human annotations being accurate and equally reliable in all four domains, but no inter-annotator agreement, label-noise rate, or quality-control outcome is reported, and attributes like Attractive are admitted to be subjective.","fun_headline_variants_meta":{"raw":{"variants":["Fairness models fail when faces turn to art and cartoons","New 100K face benchmark exposes fairness gap across domains","Faces in art and cartoons break fairness models","Domain shifts in faces unmask fairness failures"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000175,"raw_usage":{"total_tokens":1074,"prompt_tokens":645,"completion_tokens":429,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":389,"completion_tokens_details":{"reasoning_tokens":367}},"tokens_in":389,"tokens_out":429,"duration_ms":6034,"temperature":1.0,"reasoning_tokens":367,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T13:20:31.958360+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-annotate a random sample of, say, 1,000 images per domain with an independent annotation team and measure per-attribute agreement against the released labels. If agreement is much lower in Art and Sketch, or near chance for subjective attributes like Attractive, then the reported disparate-impact heatmaps and the claim that the dataset is harder would reflect annotation noise, not domain shift.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the 30,000-image Photo domain and the annotation scheme the benchmark builds on."},{"cited_title":"Karras, M","cited_arxiv_id":null,"evidence_quote":"Supplies museum artwork faces for the Art domain after filtering out black-and-white images."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies crawled portrait faces in varied artistic styles for the Art domain."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies watercolor and oil-painting images for the Art domain."},{"cited_title":"Zheng, Y","cited_arxiv_id":null,"evidence_quote":"Supplies most of the Cartoon domain images."},{"cited_title":"Zhang, X","cited_arxiv_id":null,"evidence_quote":"Supplies the 1,194 hand-drawn sketches in the Sketch domain."},{"cited_title":"Karkkainen and J","cited_arxiv_id":null,"evidence_quote":"The benchmark's main comparison dataset for fairness and domain generalization."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the FEDORA baseline and the convention of treating racial groups as domains that Face4FairShifts replaces with natural style domains."}],"review_version":1}