Pith. sign in

REVIEW 2 cited by

FAIntbench: A Holistic and Precise Benchmark for Bias Evaluation in Text-to-Image Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.17814 v6 pith:CPSZCPZL submitted 2024-05-28 cs.CV cs.AI

classification cs.CVcs.AI
keywords biasesfaintbenchmodelsbiasbenchmarkevaluateevaluationholistic
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The rapid development and reduced barriers to entry for Text-to-Image (T2I) models have raised concerns about the biases in their outputs, but existing research lacks a holistic definition and evaluation framework of biases, limiting the enhancement of debiasing techniques. To address this issue, we introduce FAIntbench, a holistic and precise benchmark for biases in T2I models. In contrast to existing benchmarks that evaluate bias in limited aspects, FAIntbench evaluate biases from four dimensions: manifestation of bias, visibility of bias, acquired attributes, and protected attributes. We applied FAIntbench to evaluate seven recent large-scale T2I models and conducted human evaluation, whose results demonstrated the effectiveness of FAIntbench in identifying various biases. Our study also revealed new research questions about biases, including the side-effect of distillation. The findings presented here are preliminary, highlighting the potential of FAIntbench to advance future research aimed at mitigating the biases in T2I models. Our benchmark is publicly available to ensure the reproducibility.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Harm is not Universal: Community-Specific Toxicity Detection is Urgently Needed

    cs.CV 2026-07 conditional novelty 6.0 of 10

    SoTA T2I toxicity detectors miss ~35% of disability-community harms; zero-shot CTD fails below random, while ICL/VQA/LoRA improve but stay well below general TD performance.

  2. Auditing the Audit: Five Failure Modes in Benchmark-Validity Audits

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Five silent pipeline failure modes can manufacture conclusions in perturbation-based benchmark-validity audits; a six-point due-diligence gate left all ten cells of a two-model safety case study non-confirmatory.

Pith tools