REVIEW 3 major objections 4 minor 21 references
All ten large language models tested showed consistent directional preferences on every workplace topic, and they rejected disfavored claims more strongly than they endorsed the opposite claims.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-03 11:14 UTC pith:HFE2GGZ6
load-bearing objection The framework is thoughtfully built but the manuscript reports no results at all: the abstract's empirical claims about ten models and the rejection/endorsement asymmetry are unverifiable, and the asymmetry itself is plausibly an artifact of symmetric Likert scoring. the 3 major comments →
BiasLab: A Multilingual Dual-Framing Framework for LLM Bias Measurement, Applied to Workplace and HR Contexts
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
BiasLab operationalizes bias as a systematic directional preference measured through mirrored probe pairs: an affirmative assertion favoring Target A and a reverse assertion favoring Target B, structurally identical except for target substitution. Across six workplace/HR topics and twelve languages, all ten evaluated models showed consistent directional preferences on every topic, and a recurrent asymmetric pattern emerged in which models rejected disfavored claims more strongly than they endorsed their opposites. This asymmetry is invisible to single-frame designs, which would only see the net direction, not the difference in intensity between rejection and endorsement.
What carries the argument
The load-bearing mechanism is the dual-framing prompt pair with deterministic target substitution, combined with randomized instruction wrappers, a fixed-choice Likert response format, LLM-based agreement labeling, and polarity-aligned scoring (reverse-frame scores are negated before aggregation). This machinery isolates directional preference from prompt artifacts and makes cross-model, cross-language comparisons possible.
Load-bearing premise
The polarity-aligned scoring assumes that a 'disagree' response is the exact intensity mirror of an 'agree' response; if models simply use the disagree anchors more emphatically than the agree anchors, the reported asymmetry would appear for any topic without indicating any target-specific bias.
What would settle it
Run BiasLab on a pair of genuinely interchangeable, value-neutral targets (e.g., two arbitrary colors) and check whether the rejection-vs-endorsement asymmetry still appears. If it does, the asymmetry is an anchor-intensity artifact of the scoring, not a target-specific bias; if it disappears, the asymmetry is tracking something about the topics themselves. A second check: measure the LLM judge's mapping of 'agree' versus 'disagree' intensities on controlled responses to identical statements.
If this is right
- Organizations can compare and vet LLMs on directional bias before adopting them for hiring or other workplace decisions.
- The asymmetry between rejection and endorsement provides a new indicator of model behavior that single-frame audits miss.
- Multilingual consistency can be inspected directly: a preference that appears in one language but not another signals language-dependent bias.
- Neutrality rates distinguish abstention or refusal from genuine balance, so a near-zero bias score is not automatically 'no bias'.
- Because probes are generated deterministically and all raw outputs are logged, the pipeline supports reproducible audits across time and model versions.
Where Pith is reading between the lines
- If the rejection-stonger-than-endorsement asymmetry is an artifact of anchor intensity—models or the judge using 'strongly disagree' more emphatically than 'strongly agree'—the reported asymmetry would appear even on neutral or arbitrary topics; the paper does not test this, but a control condition with two interchangeable targets would settle it.
- The dual-framing logic could be adapted to measure framing sensitivity in human respondents or to benchmark debiasing interventions, not just to audit existing models.
- The paper's focus on forced-choice Likert responses may undercount bias that lives in hedged, selective, or refusals-then-compliant outputs; a combined forced-choice plus open-ended protocol would likely reveal additional directional patterns.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript introduces BiasLab, a multilingual framework for measuring output-level directional bias in LLMs using mirrored affirmative/reverse prompt pairs, randomized instructional wrappers, forced-choice Likert responses, an LLM-based judge for normalization, and polarity-aligned scoring. The abstract makes strong empirical claims: ten LLMs were evaluated across six workplace/HR topics and twelve languages, yielding 43,200 responses, with consistent directional preferences across all models and topics and a recurring asymmetry in which rejection of disfavored claims is stronger than endorsement of favored ones. However, the body of the manuscript contains no Results section, no data tables, no model inventory, no topic-language matrix, and no statistical summaries. It is purely a framework/methods description. The limitations sections (3.4, 3.8) acknowledge measurement-error and descriptive-statistics concerns but do not address the core validity threat to the headline asymmetry.
Significance. If the framework works as intended, it could be a useful open-source instrument for comparing LLM output biases across languages and prompt perturbations. The dual-framing design, deterministic target substitution, explicit handling of neutrality, forced-choice response constraints, and availability of code and demo artifacts are genuine strengths. However, the paper's central contribution as stated in the abstract—the empirical finding of universal directional preference and rejection/endorsement asymmetry across ten models—is entirely unsupported by the submitted manuscript. Furthermore, the asymmetry claim is threatened by an uncalibrated scoring assumption. The potential methodological value is therefore conditional on substantial additional work and evidence.
major comments (3)
- [Abstract vs. full text] The abstract reports a complete empirical study—10 models, 6 topics, 12 languages, 30 iterations per framing, 43,200 responses—and states that 'all ten models showed consistent directional preferences across every topic' and that a 'recurring asymmetric pattern' emerged. The body of the manuscript contains none of this: there is no Results section, no table of per-model or per-topic scores, no list of the ten models, no description of the six topics, and no statistical summary such as mean bias scores, t-statistics, or effect sizes. The central empirical claims are therefore unverifiable and, as written, appear nowhere in the paper besides the abstract. This is a load-bearing omission, not a presentation issue.
- [Section 2.6 (Scoring and polarity alignment)] The main empirical novelty—that models reject disfavored claims more strongly than they endorse opposites—is an artifact-prone consequence of the scoring scheme. The paper maps Strongly agree to +2, Agree to +1, Disagree to −1, and Strongly disagree to −2, and then negates reverse-framing scores before aggregation. This assumes that a disagreement anchor is the exact intensity mirror of an agreement anchor. The paper further instructs the judge to assign 'Strongly' only when explicit intensifiers are present. If negative disagreement intensifiers are more frequent or more emphatic than positive agreement intensifiers in the test languages, the asymmetry would appear even for content-free or neutral probes. No no-bias control, per-anchor calibration, or cross-valence intensity validation is reported. Section 3.4 concedes that the judge may be a source of measurement error, but it does not
- [Sections 2.7 and 3.8] The statistical reporting uses one-sample t-tests and Cohen's d on the polarity-aligned ordinal scores, treating the Likert categories as an interval scale. Section 3.8 correctly states that these are descriptive indicators, not inferential population claims. However, the abstract's sweeping claim that 'all ten models showed consistent directional preferences across every topic' depends on these very summaries, and no summaries are reported. Moreover, without a baseline condition measuring the scoring scheme's behavior on symmetric or neutral content, a t-test against zero cannot distinguish target-specific bias from anchor-intensity imbalance. The manuscript needs either a no-bias control condition or an explicit calibration of the ordinal scale across valences before the asymmetry claim can be supported.
minor comments (4)
- [Section numbering] There are two subsections numbered 2.6; the second is 'Scoring and polarity alignment.' There is also no Section 2.2, jumping from 2.1 to 2.3. The numbering should be corrected.
- [Title and abstract consistency] The title and abstract emphasize workplace and HR contexts (gender in leadership, employment gaps, age in hiring, remote/office, four-day/five-day, AI-assisted hiring). The body does not mention any of these topics or any HR-specific application. The framework description is generic, which is fine methodologically, but the mismatch is confusing.
- [Section 2.1] The text says the reverse framing is an 'exact mirror' with 'all other lexical and syntactic structure identical,' but Section 2.3 states that language-specific prompt variants are generated to be 'fluent and idiomatic.' These two goals can conflict; the manuscript should clarify how synonym/idiom changes are controlled or audited, especially for the multilingual case.
- [Figures] Figures 1 and 2 are referenced but not described in sufficient detail; if they are included in the actual submission, the captions should explicitly show how the reverse prompt is a mirror of the affirmative prompt and how wrappers are sampled.
Circularity Check
No significant circularity: the headline asymmetry is an empirical aggregate, not a consequence of the scoring definition, and the only self-citation is not load-bearing.
full rationale
The derivation chain in BiasLab consists of constructing mirrored affirmative/reverse probes, normalizing raw outputs via an LLM-based judge, mapping labels to an ordinal scale, negating reverse-frame scores for polarity alignment, and aggregating to a directional mean. None of these steps defines the target result in terms of itself. The headline claim—'models rejected disfavored claims more strongly than they endorsed their opposites'—is presented as an observed pattern across 43,200 responses, not as an algebraic consequence of the scoring scheme. The polarity alignment convention (Sec. 2.6) sets sign and ordinal spacing, but it does not by construction determine whether 'Strongly disagree' is used more often than 'Strongly agree'; that is an empirical property of the collected labels, however confounded by judge and anchor effects. No fitted parameter is later relabeled as a prediction, and no uniqueness theorem or ansatz is imported from the authors' prior work. The only self-citation (ref. 11) appears in the bibliography and is not used to justify the framework's load-bearing choices. The limitations sections acknowledge real validity concerns—LLM-judge measurement error (Sec. 3.4), descriptive rather than inferential statistics (Sec. 3.8), forced-choice realism (Sec. 3.2)—but these are measurement/interpretation issues, not circular reductions. The supplied text omits a results section, making the empirical claims unverifiable here, yet that is an evidentiary gap rather than a circular derivation. Accordingly, no circular step can be exhibited with the specificity required by the review rules.
Axiom & Free-Parameter Ledger
free parameters (4)
- Likert score mapping =
+2/+1/0/-1/-2
- Strongly-label intensifier criterion =
explicit intensifiers only
- Wrapper iteration count N =
30 per framing
- Topic and language selection =
6 topics, 12 languages
axioms (5)
- domain assumption Polarity alignment via negation of reverse-frame scores yields a valid directional preference measure.
- domain assumption Mirrored prompts are semantically equivalent across languages and framings.
- domain assumption The LLM-based judge maps raw outputs to agreement labels with acceptable accuracy.
- domain assumption Fixed-choice Likert responses are comparable across models and languages.
- domain assumption t-tests and Cohen's d are meaningful descriptive indicators for model-output samples.
read the original abstract
Background: Large language models (LLMs) harbor systematic biases that are particularly consequential in workplace and HR contexts, where their outputs increasingly influence hiring, job design, and organizational decisions. Existing bias-evaluation approaches remain methodologically fragmented, limiting practitioners' ability to assess deployment risks. Objective: This study introduces BiasLab, a multilingual dual-framing framework to quantify and compare directional output-level bias in LLMs, demonstrated across six workplace and HR-relevant topics. Methods: BiasLab combines mirrored affirmative and reverse prompt pairs, randomized wrapper perturbations, fixed-choice response constraints, and polarity-aligned scoring. Ten LLMs were evaluated across six topics (gender in leadership, employment gap candidates, age in hiring, remote versus office work, four-day versus five-day work weeks, and AI-assisted versus human-only hiring), spanning 12 languages and 30 iterations per framing direction, yielding 43,200 responses. Results: All ten models showed consistent directional preferences across every topic. A recurring asymmetric pattern emerged in which models rejected disfavored claims more strongly than they endorsed their opposites, a distinction invisible to single-frame designs. Conclusions: BiasLab provides a standardized, reproducible instrument for measuring directional preferences across models. Whether a preference constitutes bias in a fairness sense is topic-dependent: for protected attributes such as gender and age it maps onto equal-employment standards, whereas elsewhere it is better described as systematic preference. The framework lets organizations compare and vet models before adopting them for hiring.
Reference graph
Works this paper leans on
-
[1]
Min, B., Ross, H., Sulem, E., Veyseh, A. P. B., Nguyen, T. H., Sainz, O., Agirre, E., Heintz, I., & Roth, D. (2023). Recent Advances in Natural Language Processing via Large Pre- trained Language Models: A Survey. ACM Comput. Surv., 56(2), 30:1-30:40. https://doi.org/10.1145/3605943
doi:10.1145/3605943 2023
-
[2]
Li, Y., Wang, S., Ding, H., & Chen, H. (2023). Large Language Models in Finance: A Survey. 4th ACM International Conference on AI in Finance, 374–382. https://doi.org/10.1145/3604237.3626869
arXiv 2023
-
[3]
Nazi, Z. A., & Peng, W. (2024). Large Language Models in Healthcare and Medical Domain: A Review. Informatics, 11(3), 57. https://doi.org/10.3390/informatics11030057
-
[4]
Gan, W., Qi, Z., Wu, J., & Lin, J. C.-W. (2023). Large Language Models in Education: Vision and Opportunities. 2023 IEEE International Conference on Big Data (BigData), 4776–
2023
-
[5]
Lu, J. G., Song, L. L., & Zhang, L. D. (2025). Cultural tendencies in generative AI. Nature Human Behaviour. https://doi.org/10.1038/s41562-025-02242-1
-
[6]
Guo, Y., Guo, M., Su, J., Yang, Z., Zhu, M., Li, H., Qiu, M., & Liu, S. S. (2024). Bias in Large Language Models: Origin, Evaluation, and Mitigation (arXiv:2411.10915). arXiv. https://doi.org/10.48550/arXiv.2411.10915
-
[7]
Arzaghi, M., Carichon, F., & Farnadi, G. (2025). Understanding Intrinsic Socioeconomic Biases in Large Language Models. In Proceedings of the 2024 AAAI/ACM Conference on AI, Ethics, and Society (pp. 49–60). AAAI Press
2025
-
[8]
Hu, T., Kyrychenko, Y., Rathje, S., Collier, N., Van Der Linden, S., & Roozenbeek, J. (2024). Generative language models exhibit social identity biases. Nature Computational Science, 5(1), 65–75. https://doi.org/10.1038/s43588-024-00741-1
-
[9]
Kaneko, M., Imankulova, A., Bollegala, D., & Okazaki, N. (2022). Gender Bias in Masked Language Models for Multiple Languages. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 2740–2750. https://doi.org/10.18653/v1/2022.naacl-main.197
-
[10]
Rozado, D. (2024). The Political Preferences of LLMs (arXiv:2402.01789). arXiv. https://doi.org/10.48550/arXiv.2402.01789
-
[11]
Guey, W., Bougault, P., Moura, V. D. de, Zhang, W., & Gomes, J. O. (2025). Mapping Geopolitical Bias in 11 Large Language Models: A Bilingual, Dual-Framing Analysis of U.S.-China Tensions (arXiv:2503.23688). arXiv. https://doi.org/10.48550/arXiv.2503.23688
-
[12]
Liu, Z. (2023). Cultural Bias in Large Language Models: A Comprehensive Analysis and Mitigation Strategies. Journal of Transcultural Communication, 3(2), 224–244. https://doi.org/10.1515/jtc-2023-0019
-
[13]
Bouguettaya, A., Stuart, E. M., & Aboujaoude, E. (2025). Racial bias in AI-mediated psychiatric diagnosis and treatment: A qualitative comparison of four large language models. Npj Digital Medicine, 8(1), 332. https://doi.org/10.1038/s41746-025-01746-4
-
[14]
Resnik, P. (2024). Large Language Models are Biased Because They Are Large Language Models. Computational Linguistics. https://doi.org/10.48550/ARXIV.2406.13138
work page internal anchor Pith review Pith/arXiv arXiv doi:10.48550/arxiv.2406.13138 2024
-
[15]
Jiao, T., Zhang, J., Xu, K., Li, R., Du, X., Wang, S., & Song, Z. (2024). Enhancing Fairness in LLM Evaluations: Unveiling and Mitigating Biases in Standard-Answer-Based Evaluations. Proceedings of the AAAI Symposium Series, 4(1), 56–59. Crossref. https://doi.org/10.1609/aaaiss.v4i1.31771
-
[16]
Navigli, R., Conia, S., & Ross, B. (2023). Biases in Large Language Models: Origins, Inventory, and Discussion. Journal of Data and Information Quality, 15(2), 1–21. Crossref. https://doi.org/10.1145/3597307
doi:10.1145/3597307 2023
-
[17]
Gallegos, I. O., Rossi, R. A., Barrow, J., Tanjim, M. M., Kim, S., Dernoncourt, F., Yu, T., Zhang, R., & Ahmed, N. K. (2024). Bias and Fairness in Large Language Models: A Survey. Computational Linguistics, 50(3), 1097–1179. https://doi.org/10.1162/coli_a_00524
-
[18]
Karinshak, E., Hu, A., Kong, K., Rao, V., Wang, J., Wang, J., & Zeng, Y. (2024). LLM- GLOBE: A Benchmark Evaluating the Cultural Values Embedded in LLM Output (arXiv:2411.06032). arXiv. https://doi.org/10.48550/arXiv.2411.06032
-
[19]
Doan, T. V., Wang, Z., Hoang, N. N. M., & Zhang, W. (2024). Fairness in Large Language Models in Three Hours. Proceedings of the 33rd ACM International Conference on Information and Knowledge Management, 5514–5517. https://doi.org/10.1145/3627673.3679090
arXiv 2024
-
[20]
Gilardi, F., Alizadeh, M., & Kubli, M. (2023). ChatGPT outperforms crowd workers for text- annotation tasks. Proceedings of the National Academy of Sciences, 120(30), e2305016120. https://doi.org/10.1073/pnas.2305016120
-
[4785]
https://doi.org/10.1109/BigData59044.2023.10386291
arXiv 2023
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.