Pith. sign in

REVIEW 6 cited by

BIGbench: A Unified Benchmark for Evaluating Multi-dimensional Social Biases in Text-to-Image Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.15240 v6 pith:3FBZODQS submitted 2024-07-21 cs.CV

BIGbench: A Unified Benchmark for Evaluating Multi-dimensional Social Biases in Text-to-Image Models

classification cs.CV
keywords biasesbigbenchmodelsbiasanalysisbenchmarkbigbench2024different
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Text-to-Image (T2I) generative models are becoming increasingly crucial due to their ability to generate high-quality images, but also raise concerns about social biases, particularly in human image generation. Sociological research has established systematic classifications of bias. Yet, existing studies on bias in T2I models largely conflate different types of bias, impeding methodological progress. In this paper, we introduce BIGbench, a unified benchmark for Biases of Image Generation, featuring a carefully designed dataset. Unlike existing benchmarks, BIGbench classifies and evaluates biases across four dimensions to enable a more granular evaluation and deeper analysis. Furthermore, BIGbench applies advanced multi-modal large language models to achieve fully automated and highly accurate evaluations. We apply BIGbench to evaluate eight representative T2I models and three debiasing methods. Our human evaluation results by trained evaluators from different races underscore BIGbench's effectiveness in aligning images and identifying various biases. Moreover, our study also reveals new research directions about biases with insightful analysis of our results. Our work is openly accessible at https://github.com/BIGbench2024/BIGbench2024/.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Gender Artifacts from Art History to Text-to-Image Generation

    cs.CV 2026-06 unverdicted novelty 7.0

    Introduces the StyleGender dataset and PixelSGA/MaskSGA metrics showing that text-to-image models amplify gender artifacts present in artistic styles beyond historical baselines.

  2. The Moving Target: A Longitudinal Audit of Trustworthiness Drift Across Twelve Checkpoints of Open-Source Chat LLMs

    cs.SE 2026-07 conditional novelty 6.5

    Across 12 checkpoints of Yi, Qwen, Mistral and Gemma, mean absolute adjacent-generation trust-score drift is 8.00 pp—3.6× an independence-based no-drift null—and remains elevated under leave-one-out and strict-scoring checks.

  3. Chains That See, Answers That Don't: A Multi-Aspect Evaluation Recipe for Forced Chain-of-Thought on Video-MME

    cs.CV 2026-06 conditional novelty 6.0

    Forced CoT produces video-dependent reasoning chains but does not improve MCQ accuracy on Qwen2.5-VL with Video-MME and causes a small drop on the 7B variant.

  4. Prototypicality Bias Reveals Blindspots in Multimodal Evaluation Metrics

    cs.CV 2026-01 conditional novelty 6.0

    Prototypicality bias: common text-to-image metrics systematically prefer plausible-but-wrong images over correct non-prototypical ones; PROTOSCORE mitigates but does not eliminate the failure.

  5. AInstein: Can LLMs Solve Research Problems From Parametric Memory Alone?

    cs.AI 2025-10 unverdicted novelty 6.0

    LLMs generate valid solutions to over 70% of AI research problems from parametric memory alone but rediscover the exact published approach less than 19% of the time, with performance limited by cross-domain analogical...

  6. Estimating Treatment Effects for Depression in Longitudinal Therapy Switching Settings

    stat.AP 2026-06 conditional novelty 5.0

    In switched MDD follow-up data, causal forest outperformed seven baselines at next-visit counterfactual prediction, and adjusted treatment effects were modest (about 0.3–0.9 HAMD-17 points).