Pith. sign in

REVIEW 3 major objections 5 minor 141 references

Multimodal LLMs reproduce instrument–gender stereotypes with strength ordered by modality: text > vision > audio.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 17:33 UTC pith:RRL6TQG6

load-bearing objection New parallel multimodal bias benchmark; the audio<vision<text ranking doesn't survive one of its two audio models' failure to recognize instruments. the 3 major comments →

arxiv 2607.26355 v1 pith:RRL6TQG6 submitted 2026-07-29 cs.CL

Symphony of Bias: Exploring Gender Associations with Musical Instruments in Multimodal LLMs

classification cs.CL
keywords gender biasmultimodal LLMsmusical instrumentsstereotype alignmentmodality comparisonnon-binary genderSymphony-Bias
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper tries to establish that multimodal LLMs reproduce long-documented cultural stereotypes linking musical instruments to gender, and that the strength of that reproduction depends on input modality. Using a new parallel dataset of over 50,000 text, image, and audio samples spanning 22 instruments, the authors find that about 92% of instrument-level model judgments align with prior social-science and survey-based stereotypes. Harp and drums are the most consistently gendered across all evaluated models and modalities. The paper further argues that stereotype alignment is weakest when models hear the instrument, stronger when they see it, and strongest when they read about it, suggesting that modality-specific training signals carry different amounts of gendered information.

Core claim

The central claim is that instrument–gender associations are encoded in LLMs in a modality-dependent hierarchy: audio < vision < text. The authors operationalize this with two measures. The Gender-Association-Score (GAS) quantifies how far an instrument's likelihood distribution over male, female, and non-binary deviates from an equal baseline; the Alignment-Bias-Score (ABS) subtracts the non-stereotypical from the stereotypical binary gender's score for each instrument. Across 220 model–instrument pairs, 92% align with documented stereotypes, with harp and drums showing the largest and most consistent ABS values. Text-only inputs produce the strongest ABS, vision inputs intermediate, and au

What carries the argument

The load-bearing machinery is a parallel multimodal dataset plus two derived scores. Each scenario is a neutral sentence about a person playing an instrument, and the same semantic template is presented as text alone, as text paired with an image, or as text paired with an audio clip of random notes rendered per instrument, so any difference in response is attributable to the input modality. GAS and ABS convert the model's Likert-style likelihood judgments into a signed measure of stereotype alignment: GAS measures deviation from uniform gender likelihood, and ABS measures the gap between the stereotypically expected and the alternative binary gender. A 47-participant human survey supplies t

Load-bearing premise

The audio-modality conclusion assumes that low or negative Alignment-Bias-Scores from audio models reflect genuinely weak stereotyped associations, not the models' failure to recognize the instrument; the paper itself reports that one audio model identifies instruments with under 10% accuracy.

What would settle it

Re-run the audio probes while filtering to trials where an independent closed-set instrument-recognition check confirms the model heard the correct instrument, then compare ABS with and without the filter. If filtered audio ABS rises toward the vision and text levels, the 'audio is least biased' claim collapses into a perception artifact; if it stays low, the claim survives.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • If the hierarchy holds, text-based LLM interfaces such as chatbots and search are the most likely to reinforce instrument–gender stereotypes, while voice- and sound-based interfaces may be comparatively safer.
  • Since 92% of model judgments align with documented stereotypes, downstream systems that recommend instruments or generate music-related content will likely propagate these associations unless actively countered.
  • Harp and drums are the strongest and most consistent carriers of the stereotype, making them natural test cases for monitoring and for debiasing interventions.
  • Scale and explicit reasoning reduce measured bias but do not remove it: after reasoning, 91% of outcomes still align with societal stereotypes, so incremental model improvements are unlikely to solve the problem.
  • Non-binary associations are unstable and often treated as an underspecified fallback category, particularly in vision models, meaning binary-only evaluations understate the bias landscape.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: the audio result is only as strong as the audio models' instrument recognition. One audio model identifies instruments with less than 10% accuracy (reported in the paper's Appendix F), so its near-zero or negative ABS may reflect failure to recognize the instrument rather than absence of stereotype encoding.
  • Editorial inference: because acoustic training data rarely labels the performer's gender, audio is a plausible lower-bias input channel; shifting consumer interactions from text to sound could measurably reduce stereotype reinforcement without model retraining, a claim worth testing in a user study.
  • Editorial inference: ABS excludes non-binary by construction, so 'alignment with social stereotypes' here means alignment with binary stereotypes; a future metric measuring stereotype strength against a neutral baseline across all three categories could give a fuller picture.
  • Editorial inference: the parallel-dataset design could be reused to ablate bias sources more finely, for example by removing the instrument name from the text while keeping the image or audio, isolating how much of the association comes from the linguistic label versus the perceptual signal.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. This paper introduces Symphony-Bias, a parallel text/image/audio dataset covering 22 musical instruments, and evaluates 10 multimodal LLMs using textual Likert-scale likelihood prompts to probe gender associations (male, female, non-binary). Two metrics are defined: Gender-Association Score (GAS) and Alignment-Bias Score (ABS). The paper claims that 92% of instrument-level outcomes align with prior social-science stereotypes, that harp and drums show the most consistent gendered associations, that alignment is weakest in audio, stronger in vision, and strongest in text, and that non-binary associations are less coherent. A 47-participant survey is used both to derive stereotype labels and as human ground truth for correlation analysis.

Significance. If the central claims were fully supported, this would be a useful contribution to multimodal bias evaluation: the domain (instrument–gender stereotypes) is genuinely underexplored, and the parallel text/image/audio dataset is a valuable resource. The paper also includes several good practices: prompt-order randomization, reasoning vs. no-reasoning comparisons, a log-probability proxy validation, bootstrap robustness checks, and an auxiliary audio-recognition sanity check. However, the headline quantitative claims are not currently supported by the reported analysis. The 92% alignment figure counts near-zero and non-significant positive ABS values, and the audio-modality conclusion rests heavily on one model that largely fails to recognize the stimuli. These are fixable with reanalysis and additional experiments, but they are load-bearing for the paper's main contributions.

major comments (3)
  1. [Section 6 (RQ1), Table 3] The headline '92% of model–instrument pairs align with social stereotypes' is computed from the sign of ABS, but many positive cells are effectively zero and many are not statistically significant. For example, in the Gemini 2.5 Flash row the ABS is 0.000 for acoustic guitar, bassoon, cello, flute, keyboard, oboe, etc., and numerous other entries are between 0.001 and 0.005 while marked non-significant (●). A near-zero, non-significant difference between male and female GAS is not evidence of alignment. The 92% figure should be recomputed using a minimum effect threshold and/or significance-aware counting, and the exact set of rows used should be stated: Table 3 displays 15 model/input rows, not the N=22×10=220 implied in the text.
  2. [Section 6 (RQ2), Table 3, Appendix F] The audio<vision<text modality ranking is not robust because one of the two audio-text models, Qwen2-Audio, identifies instruments with less than 10% accuracy (Appendix F), whereas Music-Flamingo achieves 71%. The paper acknowledges this in Section 6 but still uses Qwen2-Audio's near-zero or negative ABS values as evidence that audio models encode weaker stereotypes. For Qwen2-Audio, the audio input carries essentially no instrument identity, so its ABS reflects recognition failure, not modality-specific stereotype encoding. Removing Qwen2-Audio leaves a single audio model; moreover, Music-Flamingo's audio row is not uniformly low (e.g., clarinet 0.113, drums 0.110, saxophone 0.122). The modality-dependence claim therefore needs either additional audio models with adequate instrument recognition or an explicit reframing as an exploratory case study.
  3. [Section 4.1, Table 2, and Section 6 'Human Study' (Figure 2)] The same 47-participant survey is used both to construct the binary stereotype labels g*(I) in Table 2 (which feed the ABS) and to compute the human–model Pearson correlations in Figure 2. This makes the human correlation partly dependent on the same judgments that define the 'stereotyped' direction. The overlap is not complete, since Table 2 also draws on social-science literature, but the validation is not independent. The authors should either use held-out survey participants for the correlation, use literature-only labels for ABS, or explicitly acknowledge the dependence and present Figure 2 as descriptive rather than confirmatory of RQ1.
minor comments (5)
  1. [Section 6 (RQ2)] The first paragraph says text-only models 'generally exhibit lower bias scores than vision–text and audio–text models,' while the next paragraph says 'text-only inputs show the strongest bias.' These statements appear contradictory unless the comparison is carefully restricted to the same underlying models with different input modalities. Please clarify the intended contrast.
  2. [Table 3] The table contains 15 model/input rows (five text-only, six vision/text variants, four audio/text variants), yet the 92% calculation is described as N=22×10=220. Specify which rows are included in the headline percentage and make the row naming consistent with Section 4.2.
  3. [Appendix E] The prompt-sensitivity appendix is incomplete: 're-running the main . experiments' has a stray period, the results are said to be 'reported in table' with no table number, and 'comstrains' is a typo. The four prompt variants are listed, but the quantitative shift (±2 percentage points) is not shown in a table.
  4. [Section H and Appendix references] Section H ends with 'Tables Table 7–Section H report the corresponding results,' which is not a resolvable reference. Please cite the specific appendix tables.
  5. [Table 2 and Section 3.1] Table 2 lists only 16 instruments grouped into female and male categories, while the evaluation uses 22 instruments. Acoustic guitar, bassoon, glockenspiel, harmonica, horn, oboe, piccolo, and tuba are missing from the table; please provide the complete mapping and explain how these instruments were assigned in the ABS computation.

Circularity Check

1 steps flagged

The binary 'alignment with human perceptions' validation reuses the same 47-participant survey that defined the instrument stereotype labels (Table 2), making that specific validation partly tautological; the main 92% alignment and modality-ordering claims are not circular, though the audio claim has a separate confound.

specific steps
  1. self definitional [Section 4.1 / Table 2; Section 5 (ABS definition); Section 6 RQ1 (human study, Figure 2)]
    "we conducted a survey in which 47 participants rated the tendency of each instrument to be associated with each gender using a Likert scale ... By combining insights from both sources, we categorized instruments into the following groups as shown in Table 2. [Table 2 caption:] These associations are based on our conducted survey. ... [ABS:] Let g*(I) denote the stereotypically associated gender of instrument I as identified in prior social science literature ... [human study:] we conducted a controlled human study ... The study included 47 participants ... Following the approach of Farsi et al"

    The same 47-participant survey is used twice: once (Section 4.1, Table 2) to set the gendered labels g*(I) that define ABS = GAS_{g*} - GAS_{bar-g}, and again (Section 6, Figure 2) as the 'human responses' against which model outputs are correlated. A model's positive ABS is, by construction, agreement with the survey-derived label, so the reported binary 'alignment with human perceptions' is not an independent validation; it re-measures the input that defined the metric. The circularity is partial: instruments such as harp, drums, and flute have independent support in the cited social-science literature, so the 92% claim retains external content, but the human-study validation for the full 22-instrument set is tautological for binary dimensions.

full rationale

The derivation chain is: survey + literature -> Table 2 labels -> ABS scores -> 92% alignment; and survey -> human correlations. The first link is an empirical measurement, not circular, because model outputs are independent of the labels. The second link is the circular step described above: the human-study 'validation' is the same input that constructed the binary stereotype labels. The audio<vision<text ordering (RQ2) is not circular: it is an empirical comparison of model scores, though it is confounded by Qwen2-Audio's <10% instrument recognition (Appendix F), which is a validity risk rather than a definitional equivalence. The self-citation to Farsi et al. (2025) for Pearson correlation is not load-bearing; it is a standard statistical method. No ansatz or uniqueness theorem is smuggled via self-citation. Overall score 4: one self-referential validation exists, but the central 92% alignment claim and the modality ordering are not forced by construction and retain independent content.

Axiom & Free-Parameter Ledger

3 free parameters · 5 axioms · 0 invented entities

The central claims rest on hand-assigned stereotype labels, a small non-independent survey used for both construction and validation, a uniform baseline, and synthetic audio stimuli. No new physical or mathematical entities are introduced.

free parameters (3)
  • Instrument stereotype labels g*(I) = Table 2: Female = violin, flute, harp, clarinet, cello, ukulele; Male = piano, electric guitar, drums, bass guitar, trum
    These labels are chosen from literature plus the authors' 47-participant Likert survey. Every ABS value and the 92% alignment count depends on this hand-set categorization.
  • Likert mapping (very low..very high -> 1..5) = 1, 2, 3, 4, 5
    Ordinal textual likelihoods are mapped to integers and normalized. Different mappings could change GAS magnitudes and some sign decisions.
  • Uniform gender baseline 1/3 = 1/3
    GAS measures deviation from equal likelihood across three gender categories. This baseline defines what counts as 'biased' and is assumed rather than derived.
axioms (5)
  • domain assumption Textual Likert probabilities are a valid proxy for model gender associations across all modalities.
    Validated against log-probabilities only on three text models with ~87% sign agreement (Section 7), then extended to vision and audio models without direct validation.
  • domain assumption The 47-participant survey provides reliable ground-truth gender stereotypes for all 22 instruments.
    The survey is small and self-selected, yet it is used both to build Table 2 and as the human benchmark for correlations (Sections 4.1 and 6).
  • domain assumption Synthetic random-note Cubase audio preserves real-world instrument gender cues.
    The audio stimuli are generated from random notes rather than real performances, which removes genre and performance context that may contribute to human acoustic stereotypes (Section 3.1).
  • domain assumption Presenting 'a Male, a Female, a Non-binary' as a closed set prevents demographic-prior leakage.
    Prompt design choice in Section 4.3 intended to stop models from defaulting to population frequencies; no ablation isolates this effect.
  • standard math Normalizing likelihoods across the three gender categories makes scores comparable across scenarios.
    Used in the GAS definition in Section 5 to remove scale differences between scenarios.

pith-pipeline@v1.3.0-daily-deepseek · 34776 in / 11139 out tokens · 100577 ms · 2026-08-01T17:33:23.120652+00:00 · methodology

0 comments
read the original abstract

Large language models (LLMs) are increasingly embedded in everyday life and widely used for information seeking, raising concerns about their potential to perpetuate social biases and reinforce stereotypes. In this study, we investigate gender bias in LLMs through the lens of their associations with musical instruments. Building on social-science research on the cultural gender-typing of instruments, we introduce Symphony-Bias, a parallel multimodal dataset spanning text, vision, and audio. We evaluate ten multimodal models with diverse architectures and scales across 22 musical instruments, analyzing how they associate each instrument with three gender categories: {male, female, non-binary}, across three modalities: {text, vision, audio}. Our results show that 92\% of instrument-level outcomes align with prior social-science findings, with the harp and drums showing particularly consistent gendered associations across all evaluated models and modalities. We further find that alignment with social stereotypes is weakest in audio, stronger in vision, and strongest in text, suggesting that modality-specific representations can differentially amplify gendered associations with musical instruments.\footnote{The Symphony-Bias dataset will be publicly released upon acceptance of the paper.}

Figures

Figures reproduced from arXiv: 2607.26355 by Donya Rooein, Farhan Farsi, Mohammad Heydari Rad, Negar Heidary, Shayan Bali.

Figure 1
Figure 1. Figure 1: Examples from our multimodal dataset. The [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Pearson correlation heatmap between models [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Gender distribution of participants 31.9% 19.1% 23.4% 14.9% 10.7% 18-24 25-34 35-44 45-54 55+ [PITH_FULL_IMAGE:figures/full_fig_p013_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Age distribution of participants B.2 Survey Format We conducted our survey using Google Forms and invited approximately 430 individuals via email. Out of those, 40 participants accepted our invitation and took part in the survey. The design of the survey and a sample of questions are presented in Figures 5 and 6. C Instrument Pitch Ranges Each instrument has its own pitch range, which is displayed in [PIT… view at source ↗
Figure 5
Figure 5. Figure 5: Some entries of the survey form Model Name a-guitar b-guitar bassoon cello clarinet drums e-guitar flute glockenspiel harmonica harp horn keyboard oboe piano piccolo saxophone trombone trumpet tuba ukulele violin Text only models Qwen2.5-7B-Instruct 0.0164 ● 0.0294 ● -0.0175 ● 0.0979 ● 0.0253 ● 0.0838 ● 0.0359● 0.2294 ● -0.0069● -0.0095 ● 0.4478 ● 0.0958 ● 0.0187● 0.0249● 0.0698 ● 0.0174 ● -0.0599 ● 0.0498… view at source ↗
Figure 6
Figure 6. Figure 6: Instruction provided to participants in our survey. [PITH_FULL_IMAGE:figures/full_fig_p015_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Education level of participants 8.51% 78.72% 12.77% Homosexual Heterosexual Prefer not to say [PITH_FULL_IMAGE:figures/full_fig_p016_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: Sexual orientation of participants "Male": "likelihood" } {SCENARIO} {The instrument is shown in the image. (for image-text-to-text models evaluating based on image-text input modal￾ity)} {The instrument’s sound is given in the audio. (for audio-text-to-text models evaluating based on audio-text input modal￾ity)} What is the gender of the person? Model Gender Prediction Prompt – Variant 3 You are a profess… view at source ↗
Figure 9
Figure 9. Figure 9: Gender Association Score (GAS) across musical instruments for the Claude 3 Haiku model. Positive values [PITH_FULL_IMAGE:figures/full_fig_p018_9.png] view at source ↗
Figure 10
Figure 10. Figure 10: Gender Association Score (GAS) across musical instruments for the Gemini 2.5 Flash model. Positive [PITH_FULL_IMAGE:figures/full_fig_p019_10.png] view at source ↗
Figure 11
Figure 11. Figure 11: Gender Association Score (GAS) across musical instruments for the InternVL3.5 8B model with image. [PITH_FULL_IMAGE:figures/full_fig_p020_11.png] view at source ↗
Figure 12
Figure 12. Figure 12: Gender Association Score (GAS) across musical instruments for the InternVL3.5 8B model without [PITH_FULL_IMAGE:figures/full_fig_p021_12.png] view at source ↗
Figure 13
Figure 13. Figure 13: Gender Association Score (GAS) across musical instruments for the LLaVA-Next-8B model with image. [PITH_FULL_IMAGE:figures/full_fig_p022_13.png] view at source ↗
Figure 14
Figure 14. Figure 14: Gender Association Score (GAS) across musical instruments for the LLaVA-Next-8B model without [PITH_FULL_IMAGE:figures/full_fig_p023_14.png] view at source ↗
Figure 15
Figure 15. Figure 15: Gender Association Score (GAS) across musical instruments for the Mistral-7B-Instruct-v0.3 model. [PITH_FULL_IMAGE:figures/full_fig_p024_15.png] view at source ↗
Figure 16
Figure 16. Figure 16: Gender Association Score (GAS) across musical instruments for the Music-Flamingo model with audio. [PITH_FULL_IMAGE:figures/full_fig_p025_16.png] view at source ↗
Figure 17
Figure 17. Figure 17: Gender Association Score (GAS) across musical instruments for the Music-Flamingo model without [PITH_FULL_IMAGE:figures/full_fig_p026_17.png] view at source ↗
Figure 18
Figure 18. Figure 18: Gender Association Score (GAS) across musical instruments for the Qwen2-Audio-7B model with audio. [PITH_FULL_IMAGE:figures/full_fig_p027_18.png] view at source ↗
Figure 19
Figure 19. Figure 19: Gender Association Score (GAS) across musical instruments for the Qwen2-Audio-7B model without [PITH_FULL_IMAGE:figures/full_fig_p028_19.png] view at source ↗
Figure 20
Figure 20. Figure 20: Gender Association Score (GAS) across musical instruments for the Qwen2.5-7B-Instruct model. [PITH_FULL_IMAGE:figures/full_fig_p029_20.png] view at source ↗
Figure 21
Figure 21. Figure 21: Gender Association Score (GAS) across musical instruments for the Qwen2.5-14B-Instruct model. [PITH_FULL_IMAGE:figures/full_fig_p030_21.png] view at source ↗
Figure 22
Figure 22. Figure 22: Gender Association Score (GAS) across musical instruments for the Qwen2.5-VL-7B model with image. [PITH_FULL_IMAGE:figures/full_fig_p031_22.png] view at source ↗
Figure 23
Figure 23. Figure 23: Gender Association Score (GAS) across musical instruments for the Qwen2.5-VL-7B model without [PITH_FULL_IMAGE:figures/full_fig_p032_23.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

141 extracted references · 26 canonical work pages · 2 internal anchors

  1. [1]

    Harold F Abeles and Susan Yank Porter. 1978. http://www.jstor.org/stable/3344880 The sex-stereotyping of musical instruments . Journal of research in music education, 26(2):65--75

  2. [4]

    American Psychological Association . 2015. https://doi.org/10.1037/a0039906 Guidelines for Psychological Practice with Transgender and Gender Nonconforming People . American Psychological Association

  3. [5]

    Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, Humen Zhong, Yuanzhi Zhu, Mingkun Yang, Zhaohai Li, Jianqiang Wan, Pengfei Wang, Wei Ding, Zheren Fu, Yiheng Xu, and 8 others. 2025. https://arxiv.org/abs/2502.13923 Qwen2.5-vl technical report . Preprint, arXiv:2502.13923

  4. [6]

    Shayan Bali, Farhan Farsi, Mohammad Hosseini, Adel Khorramrouz, and Ehsaneddin Asgari. 2026. Detecting subtle biases: An ethical lens on underexplored areas in ai language models biases. In Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers), pages 7352--7379

  5. [9]

    Yupeng Chang, Yi Chang, and Yuan Wu. 2026. https://arxiv.org/abs/2408.04556 Ba-lora: Bias-alleviating low-rank adaptation to mitigate catastrophic inheritance in large language models . Preprint, arXiv:2408.04556

  6. [10]

    Yunfei Chu, Jin Xu, Qian Yang, Haojie Wei, Xipin Wei, Zhifang Guo, Yichong Leng, Yuanjun Lv, Jinzheng He, Junyang Lin, Chang Zhou, and Jingren Zhou. 2024. https://arxiv.org/abs/2407.10759 Qwen2-audio technical report . Preprint, arXiv:2407.10759

  7. [11]

    Gheorghe Comanici, Eric Bieber, and Mike Schaekermann. et al. 2025. https://arxiv.org/abs/2507.06261 Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities . Preprint, arXiv:2507.06261

  8. [16]

    Creswell and Vicki L

    John W. Creswell and Vicki L. Plano Clark. 2018. https://doi.org/10.1111/j.1753-6405.2007.00096.x Designing and Conducting Mixed Methods Research , 3 edition. SAGE

  9. [17]

    Judith K Delzell and David A Leppla. 1992. http://www.jstor.org/stable/3345559 Gender association of musical instruments and preferences of fourth-grade students for selected instruments . Journal of research in music education, 40(2):93--103

  10. [19]

    Dillman, Jolene D

    Don A. Dillman, Jolene D. Smyth, and Leah Melani Christian. 2014. https://doi.org/10.1002/9781394260645 Internet, Phone, Mail, and Mixed-Mode Surveys: The Tailored Design Method . Wiley

  11. [20]

    John Eros. 2008. Instrument selection and gender stereotypes: A review of recent literature. Update: Applications of Research in Music Education, 27(1):57--64

  12. [21]

    Payberah

    Farhan Farsi, Shayan Bali, Fatemeh Valeh, Parsa Ghofrani, Alireza Pakniat, Kian Kashfipour, and Amir H. Payberah. 2025. https://arxiv.org/abs/2510.19616 Pbbq: A persian bias benchmark dataset curated with human-ai collaboration for large language models . Preprint, arXiv:2510.19616

  13. [22]

    Patrick M Fortney, J David Boyle, and Nicholas J DeCarbo. 1993. http://www.jstor.org/stable/3345477 A study of middle school band students' instrument choices . Journal of Research in Music Education, 41(1):28--39

  14. [24]

    Sreyan Ghosh, Arushi Goel, Lasha Koroshinadze, Sang gil Lee, Zhifeng Kong, Joao Felipe Santos, Ramani Duraiswami, Dinesh Manocha, Wei Ping, Mohammad Shoeybi, and Bryan Catanzaro. 2025. https://arxiv.org/abs/2511.10289 Music flamingo: Scaling music understanding in audio language models . Preprint, arXiv:2511.10289

  15. [25]

    Lucy Green. 1997. https://doi.org/10.1017/CBO9780511585456 Music, Gender, Education . Cambridge University Press

  16. [26]

    Philip A Griswold and Denise A Chroback. 1981. http://www.jstor.org/stable/3344680 Sex-role associations of music instruments and occupations by gender and major . Journal of Research in Music Education, 29(1):57--62

  17. [29]

    Caner Hazirbas, Alicia Sun, Yonathan Efroni, and Mark Ibrahim. 2024. https://arxiv.org/abs/2402.07329 The bias of harmful label associations in vision-language models . Preprint, arXiv:2402.07329

  18. [32]

    Phillip Howard, Avinash Madasu, Tiep Le, Gustavo Lujan Moreno, Anahita Bhiwandiwalla, and Vasudev Lal. 2024. https://arxiv.org/abs/2312.00825 Socialcounterfactuals: Probing and mitigating intersectional social biases in vision-language models with counterfactual examples . Preprint, arXiv:2312.00825

  19. [33]

    Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed. 2023. https://arxiv.org/abs/2310.0...

  20. [35]

    Hadas Kotek, Rikker Dockum, and David Sun. 2023. https://doi.org/10.1145/3582269.3615599 Gender bias and stereotypes in large language models . In Proceedings of the ACM collective intelligence conference, pages 12--24

  21. [36]

    Krosnick and Stanley Presser

    Jon A. Krosnick and Stanley Presser. 2010. https://doi.org/10.1016/C2013-0-11411-0 Question and questionnaire design . In Handbook of Survey Research, 2 edition. Emerald

  22. [37]

    Bo Li, Kaichen Zhang, Hao Zhang, Dong Guo, Renrui Zhang, Feng Li, Yuanhan Zhang, Ziwei Liu, and Chunyuan Li. 2024. https://llava-vl.github.io/blog/2024-05-10-llava-next-stronger-llms/ Llava-next: Stronger llms supercharge multimodal capabilities in the wild

  23. [38]

    Kun Li, Lai Man Po, Hongzheng Yang, Xuyuan Xu, Kangcheng Liu, and Yuzhi Zhao. 2025. https://doi.org/10.48550/arXiv.2509.11620 Aesbiasbench: Evaluating bias and alignment in multimodal language models for personalized image aesthetic assessment . In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 7618--7631

  24. [39]

    Rensis Likert. 1932. https://psycnet.apa.org/record/1933-01885-001 A technique for the measurement of attitudes . Archives of Psychology, 22(140):1--55

  25. [40]

    Yi-Cheng Lin, Tzu-Quan Lin, Chih-Kai Yang, Ke-Han Lu, Wei-Chih Chen, Chun-Yi Kuan, and Hung-yi Lee. 2024. https://doi.org/10.48550/arXiv.2407.06957 Listen and speak fairly: a study on semantic gender bias in speech integrated large language models . In 2024 IEEE Spoken Language Technology Workshop (SLT), pages 439--446. IEEE

  26. [42]

    Geoff Norman. 2010. https://doi.org/10.1007/s10459-010-9222-y Likert scales, levels of measurement and the ``laws'' of statistics . Advances in Health Sciences Education, 15(5):625--632

  27. [43]

    Alicia Parrish, Angelica Chen, Nikita Nangia, Vishakh Padmakumar, Jason Phang, Jana Thompson, Phu Mon Htut, and Samuel R. Bowman. 2022. https://arxiv.org/abs/2110.08193 Bbq: A hand-built bias benchmark for question answering . Preprint, arXiv:2110.08193

  28. [45]

    Ehud Reiter and Robert Dale. 1997. https://doi.org/10.1017/S1351324997001502 Building applied natural language generation systems . Natural Language Engineering

  29. [46]

    Christina Richards, Walter Pierre Bouman, and Meg-John Barker. 2016. https://doi.org/10.1057/978-1-137-51053-2 Genderqueer and Non-Binary Genders . Palgrave Macmillan

  30. [48]

    Philip Sedgwick. 2012. Pearson’s correlation coefficient. Bmj, 345

  31. [50]

    Lisa M Stronsick, Samantha E Tuft, Sara Incera, and Conor T McLennan. 2018. https://doi.org/10.1177/0305735617734629 Masculine harps and feminine horns: Timbre and pitch level influence gender ratings of musical instruments . Psychology of Music, 46(6):896--912

  32. [51]

    Susan M Tarnowski. 1993. https://doi.org/10.1177/875512339301200103 Gender bias and musical instrument preference . Update: Applications of Research in Music Education, 12(1):14--21

  33. [52]

    Qwen Team. 2024. https://qwenlm.github.io/blog/qwen2.5/ Qwen2.5: A party of foundation models

  34. [53]

    Vishesh Thakur. 2023. https://doi.org/10.48550/arXiv.2307.09162 Unveiling gender bias in terms of profession across llms: Analyzing and addressing sociological implications . arXiv preprint arXiv:2307.09162

  35. [56]

    Qian Wang, Zhenheng Tang, and Bingsheng He. 2025 b . Can llm simulations truly reflect humanity? a deep dive. In The Fourth Blogpost Track at ICLR 2025

  36. [57]

    Weiyun Wang, Zhangwei Gao, Lixin Gu, Hengjun Pu, Long Cui, Xingguang Wei, Zhaoyang Liu, Linglin Jing, Shenglong Ye, Jie Shao, Zhaokai Wang, Zhe Chen, Hongjie Zhang, Ganlin Yang, Haomin Wang, Qi Wei, Jinhui Yin, Wenhao Li, Erfei Cui, and 56 others. 2025 c . https://arxiv.org/abs/2508.18265 Internvl3.5: Advancing open-source multimodal models in versatility...

  37. [60]

    Gina MF Wych. 2012. https://doi.org/10.1177/8755123312437049 Gender and instrument associations, stereotypes, and stratification: A literature review . Update: Applications of Research in Music Education, 30(2):22--31

  38. [63]

    Aho and Jeffrey D

    Alfred V. Aho and Jeffrey D. Ullman , title =. 1972

  39. [64]

    Publications Manual , year = "1983", publisher =

  40. [65]

    Chandra and Dexter C

    Ashok K. Chandra and Dexter C. Kozen and Larry J. Stockmeyer , year = "1981", title =. doi:10.1145/322234.322243

  41. [66]

    Scalable training of

    Andrew, Galen and Gao, Jianfeng , booktitle=. Scalable training of

  42. [67]

    Dan Gusfield , title =. 1997

  43. [68]

    Tetreault , title =

    Mohammad Sadegh Rasooli and Joel R. Tetreault , title =. Computing Research Repository , volume =. 2015 , url =

  44. [69]

    A Framework for Learning Predictive Structures from Multiple Tasks and Unlabeled Data , Volume =

    Ando, Rie Kubota and Zhang, Tong , Issn =. A Framework for Learning Predictive Structures from Multiple Tasks and Unlabeled Data , Volume =. Journal of Machine Learning Research , Month = dec, Numpages =

  45. [70]

    Are you listening , year=

    Hello , author=. Are you listening , year=

  46. [71]

    Psychology of Music , volume=

    Masculine harps and feminine horns: Timbre and pitch level influence gender ratings of musical instruments , author=. Psychology of Music , volume=. 2018 , publisher=

  47. [72]

    Journal of Research in Music Education , volume=

    A study of middle school band students' instrument choices , author=. Journal of Research in Music Education , volume=. 1993 , publisher=

  48. [73]

    Update: Applications of Research in Music Education , volume=

    Gender bias and musical instrument preference , author=. Update: Applications of Research in Music Education , volume=. 1993 , publisher=

  49. [74]

    Musical Instruments for Girls, Musical Instruments for Boys

    “Musical Instruments for Girls, Musical Instruments for Boys”: Italian Primary and Middle School Students’ Beliefs About Gender Appropriateness of Musical Instruments , author=. Education Sciences , volume=. 2025 , publisher=. doi:10.3390/educsci15040474 , URL =

  50. [75]

    Update: Applications of Research in Music Education , volume=

    Gender and instrument associations, stereotypes, and stratification: A literature review , author=. Update: Applications of Research in Music Education , volume=. 2012 , publisher=

  51. [76]

    Journal of Research in Music Education , volume=

    Sex-role associations of music instruments and occupations by gender and major , author=. Journal of Research in Music Education , volume=. 1981 , publisher=

  52. [77]

    Psychology of Music , volume=

    Children's gender-typed preferences for musical instruments: An intervention study , author=. Psychology of Music , volume=. 2000 , publisher=

  53. [78]

    Psychology of music , volume=

    Boys' and girls' preferences for musical instruments: A function of gender? , author=. Psychology of music , volume=. 1996 , publisher=

  54. [79]

    Journal of Research in Music Education , volume=

    Psychological sex type and preferences for musical instruments in fourth and fifth graders , author=. Journal of Research in Music Education , volume=. 1997 , publisher=

  55. [80]

    Journal of Research in Music Education , pages=

    Gender and musical instruments: Winds of change? , author=. Journal of Research in Music Education , pages=. 1994 , publisher=

  56. [81]

    Journal of research in music education , volume=

    The sex-stereotyping of musical instruments , author=. Journal of research in music education , volume=. 1978 , publisher=

  57. [82]

    , author=

    Investigating the cognitive structure of stereotypes: Generic beliefs about groups predict social judgments better than statistical beliefs. , author=. Journal of Experimental Psychology: General , volume=. 2017 , publisher=

  58. [83]

    Early childhood education journal , volume=

    Preschoolers’ perceptions of gender appropriate toys and their parents’ beliefs about genderized behaviors: Miscommunication, mixed messages, or hidden truths? , author=. Early childhood education journal , volume=. 2007 , publisher=

  59. [84]

    Journal of Language and Linguistic Studies , volume=

    The impact of gendered language on our communication and perception across contexts and domains , author=. Journal of Language and Linguistic Studies , volume=

  60. [85]

    , author=

    Workplace Gender Bias: Not Just Between Strangers. , author=. North American Journal of Psychology , volume=

  61. [86]

    Archives of Psychology , volume =

    Likert, Rensis , title =. Archives of Psychology , volume =. 1932 , url =

  62. [87]

    Advances in Health Sciences Education , volume =

    Norman, Geoff , title =. Advances in Health Sciences Education , volume =. 2010 , url =

  63. [88]

    and Smyth, Jolene D

    Dillman, Don A. and Smyth, Jolene D. and Christian, Leah Melani , title =. 2014 , doi =

  64. [89]

    and Presser, Stanley , title =

    Krosnick, Jon A. and Presser, Stanley , title =. Handbook of Survey Research , edition =. 2010 , url =

  65. [90]

    1997 , url=

    Green, Lucy , title =. 1997 , url=

  66. [91]

    International Journal of Music Education , volume =

    Hallam, Susan and Rogers, Lynne and Creech, Andrea , title =. International Journal of Music Education , volume =. 2016 , doi =

  67. [92]

    and Plano Clark, Vicki L

    Creswell, John W. and Plano Clark, Vicki L. , title =. 2018 , doi =

  68. [93]

    2015 , doi =

    Guidelines for Psychological Practice with Transgender and Gender Nonconforming People , publisher =. 2015 , doi =

  69. [94]

    2016 , url =

    Richards, Christina and Bouman, Walter Pierre and Barker, Meg-John , title =. 2016 , url =

  70. [95]

    Proceedings of Machine Learning Research , year =

    Ellipsoidal conformal inference for Multi-Target Regression , author =. Proceedings of Machine Learning Research , year =

  71. [96]

    Large Language Models Sensitivity to The Order of Options in Multiple-Choice Questions

    Pezeshkpour, Pouya and Hruschka, Estevam. Large Language Models Sensitivity to The Order of Options in Multiple-Choice Questions. Findings of the Association for Computational Linguistics: NAACL 2024. 2024. doi:10.18653/v1/2024.findings-naacl.130

  72. [97]

    2022 , eprint=

    BBQ: A Hand-Built Bias Benchmark for Question Answering , author=. 2022 , eprint=

  73. [98]

    2023 , publisher=

    LLMs and AI: Understanding its reach and impact , author=. 2023 , publisher=

  74. [99]

    and Rossi, Ryan A

    Gallegos, Isabel O. and Rossi, Ryan A. and Barrow, Joe and Tanjim, Md Mehrab and Kim, Sungchul and Dernoncourt, Franck and Yu, Tong and Zhang, Ruiyi and Ahmed, Nesreen K. Bias and Fairness in Large Language Models: A Survey. Computational Linguistics. 2024. doi:10.1162/coli_a_00524

  75. [100]

    arXiv preprint arXiv:2307.09162 , year=

    Unveiling gender bias in terms of profession across LLMs: Analyzing and addressing sociological implications , author=. arXiv preprint arXiv:2307.09162 , year=

  76. [101]

    Evaluating Gender Bias of LLM s in Making Morality Judgements

    Bajaj, Divij and Lei, Yuanyuan and Tong, Jonathan and Huang, Ruihong. Evaluating Gender Bias of LLM s in Making Morality Judgements. Findings of the Association for Computational Linguistics: EMNLP 2024. 2024. doi:10.18653/v1/2024.findings-emnlp.928

  77. [102]

    Sports and Women ' s Sports: Gender Bias in Text Generation with Olympic Data

    Biester, Laura. Sports and Women ' s Sports: Gender Bias in Text Generation with Olympic Data. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 2: Short Papers). 2025. doi:10.18653/v1/2025.naacl-short.17

  78. [103]

    Advanced Robotics , volume=

    A survey of multimodal deep generative models , author=. Advanced Robotics , volume=. 2022 , publisher=

  79. [104]

    When Audio and Text Disagree: Revealing Text Bias in Large Audio-Language Models

    Wang, Cheng and Deng, Gelei and Yang, Xianglin and Qiu, Han and Zhang, Tianwei. When Audio and Text Disagree: Revealing Text Bias in Large Audio-Language Models. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. doi:10.18653/v1/2025.emnlp-main.246

  80. [105]

    2025 IEEE 19th International Conference on Automatic Face and Gesture Recognition (FG) , pages=

    Investigating Social Biases in Multimodal LLMs , author=. 2025 IEEE 19th International Conference on Automatic Face and Gesture Recognition (FG) , pages=. 2025 , organization=

Showing first 80 references.