Pith. sign in

REVIEW 3 major objections 4 minor 3 cited by

Language models, especially small open-source ones, reverse their ethical stance when instructions are phrased as prohibitions, and that instability should bar autonomous high-stakes deployment.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

Small open-weight LLMs endorse prohibited actions 24% of the time under affirmative framing but 77-100% under negated framings, a polarity swing that threatens high-stakes AI deployment.

T0 review reviewed 2026-08-03 challenge →

load-bearing objection Simple-negation effect is plausible; compound-negation headlines rest on a mislabeled framing predicate and the parsing pipeline is under-specified. the 3 major comments →

arxiv 2601.21433 v2 pith:KTILHR35 submitted 2026-01-29 cs.AI

Framing Instability in LLM Ethical Stance: Auditing Negation Sensitivity in Moral Dilemmas

classification cs.AI
keywords language model auditnegation sensitivityframing effectsethical decision-makingmoral dilemmasalgorithmic accountabilityprompt robustnessgovernance metric
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

At issue is a simple question: does a model that approves 'they should do X' disapprove 'they should not do X'? The paper audits 16 models on 14 moral dilemmas with polarity-paired framings and finds the answer is often no, most sharply for small open-weight models, which endorse the proposed action 24% of the time under affirmative framing but 77% under ordinary negation and 100% under a compound negation construction. The paper formalizes the swing as the Negation Sensitivity Index (NSI) and shows the effect is structural rather than random: deterministic decoding raises NSI by 16%, stances flip with high confidence, and financial, business, and military dilemmas are roughly twice as fragile as medical ones. Because audit logs, contestability, and governance frameworks assume instructions mean what they say, the paper argues that negation-sensitive models cannot be safely entrusted with autonomous high-stakes decisions and proposes tiered certification thresholds keyed to NSI.

Core claim

The central claim is that negation is not compositional in current LLM ethical reasoning: agreement with 'should X' and disagreement with 'should not X' — two expressions of the same endorsement — do not mirror each other. Instead, small open-weight models increase their endorsement of the action under negation, reaching 77% under 'should not' and 100% under compound negation, a 317% relative swing from the 24% affirmative baseline. Commercial models are more stable but still shift by 19-128%, and cross-model agreement drops from roughly 73-75% on affirmative framings to 62% on negated ones. The paper further reports that the instability is domain-dependent, that deterministic decoding expos

What carries the argument

The load-bearing object is the Negation Sensitivity Index (NSI), defined as the maximum minus minimum action-endorsement rate across four framings: F0 'should action', F1 'should not action', F2 'goal even if action', and F3 'not goal if action'. Endorsement is computed via logical polarity normalization: for affirmative framings, agreement counts as endorsement; for negated framings, disagreement counts as endorsement. NSI turns two paired agreement questions into a single 0-1 scale of stance stability, and the paper's governance proposal — tiered certification with domain-adjusted thresholds — is constructed directly from NSI values. The metric is mathematically equivalent to the Syntactic

Load-bearing premise

The compound-negation result depends on treating disagreement with 'should not {goal} if it means {action}' as endorsement of {action}; since F3 negates the goal rather than the action, this equivalence is the load-bearing premise, and if it fails the 100% ceiling and 317% swing are not measures of negation sensitivity.

What would settle it

Re-run the audit with action-scoped negation in the compound condition (e.g., 'They should NOT rob, even if it means not saving his daughter'); if the 100% endorsement under compound negation collapses, the reported ceiling is an artifact of goal-scope negation. A complementary check: have human annotators classify each model response to F3 as endorsing the action, endorsing the goal, or neither, and compare the human-coded action-endorsement rate with the rate assigned by the paper's normalization formula.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • A model with NSI ≥ 0.50 would fall in the proposed Tier C and require real-time human confirmation for every prohibition-type instruction in high-stakes contexts.
  • Audit logs cannot be trusted as evidence of what a negation-sensitive model was actually constrained to do, because a logged 'do not approve' can be executed as approval.
  • Domain-adjusted thresholds are needed: the paper proposes Tier A at NSI < 0.10 for financial, business, and military applications versus < 0.20 for medical, education, and science.
  • Worst-case testing should include deterministic (temperature 0) runs, since stochastic sampling masks rather than causes the instability.
  • Reasoning-enabled variants can roughly halve NSI but do not fix compound-negation failures, so explicit deliberation is a partial mitigation, not a substitute for auditing.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The F3 framing negates the goal, not the action: 'should NOT save his daughter if it means he must rob' can be rejected because saving the daughter is not endorsed, not because robbing is endorsed. The paper's normalization treats disagreement with F3 as endorsement of robbing, so the 100% compound-negation ceiling and the 317% swing may overstate action endorsement if goal-scope and action-scope
  • The claimed 2x domain gap between financial and medical fragility could be partly driven by scenario confounds (e.g., the financial scenarios involve more salient or emotionally loaded action keywords), a testable alternative to the 'clearer training signal' explanation.
  • If binary agree/disagree formats inflate apparent instability as the human-coding check suggests, NSI values computed from forced-choice responses may overstate fragility relative to open-ended or abstention-allowed elicitation; governance thresholds may need recalibration to a non-forced-choice baseline.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper audits 16 LLMs on 14 ethical dilemmas using four framing conditions (F0: 'should X', F1: 'should NOT X', F2: 'goal even if X', F3: 'NOT goal if X'). It normalizes binary agree/disagree responses into action endorsement via Eq. (1), defines the Negation Sensitivity Index (NSI) as the maximum endorsement swing across framings, and reports that open-source models endorse prohibited actions 77% of the time under simple negation and 100% under compound negation, with a 317% relative swing from the affirmative baseline. The paper further reports domain variation, robustness checks with deterministic decoding, and proposes a tiered certification framework with domain-adjusted NSI thresholds. The central claim is that instruction-following is not compositionally robust to negation, so accountability mechanisms that assume logged instructions govern behavior are unreliable for the affected models.

Significance. If the reported magnitudes are valid, the paper makes a significant contribution by connecting negation robustness to algorithmic accountability, proposing a measurable governance metric (NSI), and providing concrete deployment guidance. The study design has several strengths: polarity-paired prompts, repeated sampling, temperature ablation, FDR correction, a 16-model cross-regime sample, and a stated reproducibility package. The domain-level analysis (financial vs. medical fragility) is a useful empirical observation. However, the main headline results rest on a forced-choice response mapping whose parsing protocol is undocumented, and on a semantically invalid treatment of the compound framing F3. The abstract also promises human coding that is absent from the body and contradicted by §6. The empirical claim is therefore not yet supported to the standard the paper asserts, although the underlying question and the F0/F1 measurement approach are sound enough to warrant revision rather than rejection.

major comments (3)
  1. [Abstract; §3.5; §6] The abstract states that human coding of a response sample confirms the instability is genuine and shows that binary proxies overstate its magnitude, but no human coding is reported in the body. §3.5 describes only statistical tests and ablations; §6 explicitly states 'We did not collect human baselines.' This is a direct internal contradiction on the central validation path. Because Eq. (1) maps d=0 under F1 to action endorsement, any refusal, abstention, or 'cannot take a stance' response mechanically coded as 'disagree' inflates the reported 77% endorsement. The manuscript never specifies how free-form outputs were parsed into agree/disagree, what 'quality filtering' removed, or how abstentions were handled. Please provide the parsing protocol, abstention/exclusion rates, and either the promised human coding or a removal of the claim.
  2. [§3.2, Eq. (1); Table 2] The normalization treats F3 as a negated framing of the action (neg(F3)=1), so disagreeing with F3 is counted as endorsing the action. But F3 is 'They should NOT {goal} if it means they must {action}': the negation attaches to the goal, not to the action. Disagreeing with F3 may mean 'they should save the daughter even if they rob,' but it may also mean 'they should save the daughter without robbing' or 'the conditional is false'; it does not logically entail endorsing the prohibited action. This mapping is what drives the OSS 100% under F3 and the reported 317% swing. The F2/F3 pair is not a polarity pair for the action. The compound-negation results and any NSI component built on F3 should be removed or redefined with a logically valid action-level mapping; only the F0/F1 pair is a defensible basis for the simple-negation claim.
  3. [§4.2, Table 8] The certification thresholds (Tier A <0.20, Tier B 0.20–0.49, Tier C ≥0.50; domain multipliers 0.10/0.15) are justified as 'natural breakpoints in our distribution' and are calibrated on the same 16-model sample that is then classified. This is in-sample fitting; no procedure is given for setting thresholds independently, and no out-of-sample validation or stakeholder-based calibration is reported. Since the NSI governance framework is presented as a central contribution, the thresholds should be labeled exploratory and accompanied by a validation protocol rather than offered as ready-to-use regulatory standards.
minor comments (4)
  1. [§3.6; §3.7] Cross-reference errors: §3.6 says 'Table 1 presents our core finding,' but Table 1 lists model categories; the endorsement results are in Table 2. §3.7 says 'Table 2 summarizes the pattern' for domain results, but the domain table is Table 3.
  2. [References [11], [12], [13]] These references contain placeholder author names ('A. Firstauthor') and incomplete metadata. They must be completed or removed before publication.
  3. [§1.4; §6] The conclusion says the failures emerge from 'ordinary English sentences expressing ordinary negation,' while §6 acknowledges that the F2/F3 compound constructions are 'controlled constructions rather than naturalistic utterances.' Please reconcile these statements, since the 100% compound-negation result specifically depends on the unnatural F3 construction.
  4. [Figure 4] The n values shown for different models range from 9 to 14. If some model-scenario cells were excluded by 'quality filtering,' the exclusion criteria and counts should be reported; otherwise the per-cell denominator for Table 2 percentages is not transparent.

Circularity Check

0 steps flagged

No self-referential derivation in the core NSI measurement; the 77%/100% figures are descriptive statistics from observed responses. Minor post-hoc threshold calibration and non-load-bearing self-citations do not make the result circular.

full rationale

The core derivation chain is not circular. Eq. (1) is an explicit coding convention mapping binary decisions to action endorsement; Eq. (2) defines NSI as a range statistic over observed endorsement rates. The headline figures (24%→77%→100%) are computed from model responses, not derived from the definition of NSI or from any fitted parameter. No quantity is defined in terms of the conclusion it is used to support. The F3 concern raised by the skeptic is a semantic-validity problem: Eq. (1) treats F3 as a negated framing of the action even though F3 negates the goal; if that mapping is wrong, the F3-based numbers do not measure action endorsement, but this is a measurement artifact, not a circular derivation. Similarly, the abstract's promised human coding is absent from §3.5 and §6 states 'We did not collect human baselines'; that is missing evidence and an internal inconsistency, not circularity. The certification thresholds in §4.2 are calibrated to the same observed distribution ('natural breakpoints in our distribution') and then used to classify the same models; this is post-hoc fitting rather than independent prediction, and the paper itself calls for external stakeholder validation before adoption. The metric is explicitly disclosed as 'mathematically equivalent to the Syntactic Variation Index,' so renaming it NSI is a naming choice, not an undisclosed reduction. Self-citations [6,7,9] provide background context and are not load-bearing for the empirical measurement. Overall, the empirical result is self-contained against external benchmarks; only minor self-citation and threshold-calibration issues warrant a low non-zero score.

Axiom & Free-Parameter Ledger

2 free parameters · 5 axioms · 1 invented entities

The central empirical claim relies on the assumption that F3 is an action-negation and on binary forced-choice as a proxy for stance; both are questionable and the second is conceded to overstate magnitude. The governance thresholds are post-hoc calibrations from the same data. No physical entities are invented; NSI is a proposed metric with a clear computational definition but limited novelty.

free parameters (2)
  • NSI certification tier thresholds = 0.20 and 0.50 (general); domain-adjusted 0.10/0.15/0.35/0.40
    Chosen post hoc from the observed NSI distribution as 'natural breakpoints' in Section 4.2, not derived from an external standard or pre-registered.
  • Domain risk multipliers = Financial/business/military stricter than medical/education/science
    Based on the observed domain NSI means in Table 3; these are hand-set governance parameters, not independently validated thresholds.
axioms (5)
  • ad hoc to paper F3 'should not {goal} if it means {action}' is a negated framing of the action, so Eq (1) maps agreement onto action endorsement.
    Used in Sections 3.1-3.2 and Table 2. This is false for F3 because the negated element is the goal, not the action; it drives the '100% under compound negation' claim.
  • domain assumption Binary agree/disagree responses can be mapped to action endorsement and are a faithful proxy for model stance.
    Used throughout Section 3.2; the abstract concedes binary proxies overstate instability, yet NSI inherits the forced-choice format.
  • domain assumption Fourteen scenarios across seven domains are representative enough to generalize domain fragility.
    Section 3.4 states scenarios present 'genuine ethical tension'; no external validation of domain representativeness and no human baselines were collected.
  • domain assumption Temperature ablation on seven models supports structural rather than stochastic failure for all sixteen models.
    Section 3.8 and Appendix B test only 7 of 16 models at T=0.0; the generalization to all models is inferred, not measured.
  • domain assumption Models should be polarity-invariant for ethical judgments; any deviation is a failure rather than legitimate framing sensitivity.
    Sections 1.1 and 3.2 rely on this normative assumption; the paper acknowledges human framing effects exist and does not compare against a human baseline.
invented entities (1)
  • Negation Sensitivity Index (NSI) independent evidence
    purpose: Quantify maximum action-endorsement swing across polarity-paired framings for governance certification.
    Defined in Eq (2) and computable on any model outputs, so it has a falsifiable handle. However, the paper admits it is mathematically equivalent to the Syntactic Variation Index, so it is a relabeled metric rather than a new construct.

reviewed 2026-08-03 · how reviews work

0 comments
Cite this review

Pith. "Pith review of Framing Instability in LLM Ethical Stance: Auditing Negation Sensitivity in Moral Dilemmas." pith.science (2026). https://pith.science/paper/KTILHR35

@misc{pith2026260121433,
  author       = {Pith},
  title        = {Pith review of: Framing Instability in LLM Ethical Stance: Auditing Negation Sensitivity in Moral Dilemmas},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/KTILHR35}},
  note         = {Machine review of arXiv:2601.21433}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Language models are increasingly consulted on ethically consequential questions, yet the stance a model expresses may not survive a change in framing. We audit 16 models across 14 ethically fraught dilemmas using polarity-paired proposals ("They should X" / "They should not X"). A model's judgment of the underlying action should not reverse merely because the question is phrased as a prohibition rather than a prescription and yet, we find systematic deviations from this invariance including wholesale endorsement flips, indicating that ethical decisions are vulnerable to framing instability. Small open-weight models (1-4B parameters) endorse a proposed action 24% of the time under affirmative framing but up to 100% under negated framings, a swing of as much as 76 percentage points. Human coding of a response sample confirms the instability is genuine while showing that binary agree/disagree proxies over-state its magnitude, suggesting that an LLM judge cannot replace human coders because it silently collapses abstentions and mirrors the very forced-choice bias under study. Commercial models are for the most part more stable but still shift substantially, with cross-model agreement dropping from 73% on the bare affirmative framing to 59% under simple negation. We argue that because binary agree/disagree formats both inflate apparent endorsement and mask polarity-dependence, single-phrasing audits can misreport a model's ethical stance, and we propose the Negation Sensitivity Index (NSI) as a complement that measures stance stability directly. A model whose stance flips with phrasing cannot be relied upon in any high-stakes decision scenario.

Figures

Figures reproduced from arXiv: 2601.21433 by Jon Chun, Katherine Elkins.

Figure 1
Figure 1. Figure 1: Action endorsement rate (LPN-normalized) by [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Inter-model agreement by framing type. Affirma [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Negation sensitivity (NSI/SVI) by model and domain. Color scale: green (0, robust) to red (100, fragile). Models sorted [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Model rankings by Negation Sensitivity Index with [PITH_FULL_IMAGE:figures/full_fig_p008_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Confidence vs. negation sensitivity scatter plot by [PITH_FULL_IMAGE:figures/full_fig_p010_5.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Benchmarking LLM Competence on Logical Inference over Probability Operators

    cs.CL 2026-07 conditional novelty 6.0

    Only 9 of 29 tested LLMs beat a chance-level competence floor on valid/invalid inferences over probability operators like "probably" and "might"; most answer from a fixed yes/no bias.

  2. Benchmarking LLM Competence on Logical Inference over Probability Operators

    cs.CL 2026-07 conditional novelty 6.0

    Only 9 of 29 LLMs exceed chance on inferences over probability operators, and most models answer from a stable Yes/No bias rather than tracking the logic.

  3. Benchmarking LLM Competence on Logical Inference over Probability Operators

    cs.CL 2026-07 conditional novelty 6.0

    Most of 29 LLMs answer probability-operator inference questions from a fixed Yes/No bias, with only 9 clearing a 50% competence floor.

Reference graph

Works this paper leans on

38 extracted references · 1 canonical work pages · cited by 1 Pith paper

  1. [1]

    Yuntao Bai, Saurav Kadavath, Sharan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Catherine McKinnon, et al. 2022. Constitutional AI: Harmlessness from AI Feedback. arXiv preprint arXiv:2212.08073 (2022). https://arxiv.org/abs/2212.08073

  2. [2]

    Marcel Binz and Eric Schulz. 2023. Using Cognitive Psychology to Understand GPT-3. Proceedings of the National Academy of Sciences 120, 6 (2023), e2218523120. https://doi.org/10.1073/pnas.2218523120

  3. [3]

    Bowman and George E

    Samuel R. Bowman and George E. Dahl. 2021. What Will It Take to Fix Benchmark- ing in Natural Language Understanding?. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Hu- man Language Technologies (NAACL-HLT 2021) . Association for Computational Linguistics, Online, 4843–4855. https://d...

  4. [4]

    Joy Buolamwini and Timnit Gebru. 2018. Gender Shades: Intersectional Accu- racy Disparities in Commercial Gender Classification. In Conference on Fairness, Accountability and Transparency (FAT* 2018, Vol. 81) . PMLR, New York, NY, USA, 77–91. http://proceedings.mlr.press/v81/buolamwini18a.html

  5. [5]

    Alexandra Chouldechova. 2017. Fair Prediction with Disparate Impact: A Study of Bias in Recidivism Prediction Instruments. Big Data 5, 2 (2017), 153–163. https://doi.org/10.1089/big.2016.0047

  6. [6]

    Jon Chun, Christian Schroeder de Witt, and Katherine Elkins. 2024. Comparative Global AI Regulation: Policy Perspectives from the EU, China, and the US. arXiv preprint arXiv:2410.21279 (2024). https://arxiv.org/abs/2410.21279

  7. [7]

    Jon Chun and Katherine Elkins. 2024. Informed AI Regulation: Comparing the Ethical Frameworks of Leading LLM Chatbots Using an Ethics-Based Audit to Assess Moral Reasoning and Normative Values. arXiv preprint arXiv:2402.01651 (2024). https://arxiv.org/abs/2402.01651

  8. [8]

    Cynthia Dwork, Moritz Hardt, Toniann Pitassi, Omer Reingold, and Richard Zemel. 2012. Fairness Through Awareness. In Proceedings of the 3rd Innovations in Theoretical Computer Science Conference (ITCS ’12) . Association for Comput- ing Machinery, New York, NY, USA, 214–226. https://doi.org/10.1145/2090236. 2090255

  9. [9]

    Francisco Eiras, Aleksandar Petrov, Bertie Vidgen, Christian Schroeder, Fabio Pizzati, Katherine Elkins, et al. 2024. Risks and Opportunities of Open-Source Generative AI. arXiv preprint arXiv:2405.08597 (2024). https://arxiv.org/abs/2405. 08597

  10. [10]

    European Commission. 2021. Proposal for a Regulation Laying Down Harmonised Rules on Artificial Intelligence (Artificial Intelligence Act). COM(2021) 206 fi- nal. https://digital-strategy.ec.europa.eu/en/library/proposal-regulation-laying- down-harmonised-rules-artificial-intelligence

  11. [11]

    Firstauthor, B

    A. Firstauthor, B. Secondauthor, C. Thirdauthor, and D. Fourauthor. 2024. A Comprehensive Survey on AI Governance. arXiv preprint arXiv:2508.08789 (2024). https://arxiv.org/abs/2508.08789

  12. [12]

    Firstauthor, B

    A. Firstauthor, B. Secondauthor, C. Thirdauthor, and D. Fourauthor. 2024. MAQA: A Multimodal QA Benchmark for Negation. arXiv preprint arXiv:xxxx.xxxxx (2024). https://arxiv.org/

  13. [13]

    Firstauthor, B

    A. Firstauthor, B. Secondauthor, C. Thirdauthor, and D. Fourauthor. 2024. OR- Bench: An Over-Refusal Benchmark for Large Language Models. In OpenReview Preprint. https://openreview.net/forum?id=obYVdcMMIT

  14. [14]

    Deep Ganguli, Liane Lovitt, Jackson Kernion, Amanda Askell, Yuntao Bai, Anna Chen, Anna Goldie, Danny Hernandez, Neal DasSarma, Tom Henighan, et al. 2022. Red Teaming Language Models to Reduce Harms. arXiv preprint arXiv:2209.07858 (2022). https://arxiv.org/abs/2209.07858

  15. [15]

    Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daumé III, and Kate Crawford. 2021. Datasheets for Datasets. Commun. ACM 64, 12 (2021), 86–92. https://doi.org/10.1145/3458723

  16. [16]

    Anumanchipalli, and Julia Hockenmaier

    Federico Germani, Giovanni Spitale, Yong Chen, Gopala K. Anumanchipalli, and Julia Hockenmaier. 2025. Source Framing Triggers Systematic Bias in Large Language Models. Science Advances 11, 45 (2025), eadz2924. https://doi.org/10. 1126/sciadv.adz2924

  17. [17]

    Dan Hendrycks, Collin Burns, Steven Basart, Andrew Critch, Jerry Li, Dawn Song, and Jacob Steinhardt. 2021. Aligning AI with Shared Human Values. In Proceedings of the 9th International Conference on Learning Representations (ICLR 2021). OpenReview. https://openreview.net/forum?id=dNy_RKzJacY

  18. [18]

    Amirhossein Hosseini, Vera Demberg, and Barbara Plank. 2023. This Is Not a Dataset: A Large Negation Benchmark to Challenge Large Language Models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP 2023) . Association for Computational Linguistics, Singapore, to appear. https://aclanthology.org/

  19. [19]

    Anna Jobin, Marcello Ienca, and Effy Vayena. 2019. The Global Landscape of AI Ethics Guidelines. Nature Machine Intelligence 1, 9 (2019), 389–399. https: //doi.org/10.1038/s42256-019-0088-2

  20. [20]

    Jon Kleinberg, Sendhil Mullainathan, and Manish Raghavan. 2017. Inherent Trade- Offs in the Fair Determination of Risk Scores. InProceedings of the 8th Innovations in Theoretical Computer Science Conference (ITCS 2017). Schloss Dagstuhl–Leibniz- Zentrum für Informatik, Saarbrücken, Germany, 43:1–43:23. https://doi.org/10. 4230/LIPIcs.ITCS.2017.43

  21. [21]

    Dar, Laurel Orr, Scott Johnston, et al

    Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michi- hiro Yasunaga, Yian Zhang, Gagan B. Dar, Laurel Orr, Scott Johnston, et al. 2022. Holistic Evaluation of Language Models. arXiv preprint arXiv:2211.09110 (2022). https://arxiv.org/abs/2211.09110

  22. [22]

    Zhiwei Liu, Yupen Cao, et al. 2026. Same Claim, Different Judgment: Benchmark- ing Scenario-Induced Bias in Multilingual Financial Misinformation Detection (MFMD-Scen). arXiv preprint arXiv:2601.05403 (2026). https://arxiv.org/abs/2601. 05403

  23. [24]

    Müller, and Luciano Floridi

    Jakob Mökander, Jonas Schuett, Vincent C. Müller, and Luciano Floridi. 2023. Auditing Large Language Models: A Three-Layered Approach. AI and Ethics 3, 1 (2023), . https://doi.org/10.1007/s43681-023-00281-5

  24. [25]

    National Institute of Standards and Technology. 2023. Artificial Intelligence Risk Management Framework (AI RMF 1.0) . Technical Report NIST AI 100-1. NIST, Gaithersburg, MD, USA. https://www.nist.gov/itl/ai-risk-management- framework

  25. [26]

    Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schul- man, John Hilton, Luke Kelton, Fraser Miller, Maddie Simens, Jacob Acevedo, Kamal Ndousse Chu, et al. 2022. Training Language Models to Follow Instructions with Human Feedback. In Advances in Neural Inf...

  26. [27]

    Inioluwa Deborah Raji, Andrew Smart, Rebecca White, Margaret Mitchell, Timnit Gebru, Ben Hutchinson, Jamila Smith-Loud, Daniel Theron, and Parker Barnes

  27. [28]

    Cynthia Rudin. 2019. Stop Explaining Black Box Machine Learning Models for High Stakes Decisions and Use Interpretable Models Instead. Nature Machine Intelligence 1, 5 (2019), 206–215. https://doi.org/10.1038/s42256-019-0048-x

  28. [29]

    Smith, and Hannaneh Hajishirzi

    Michael Sclar, Rahul Palamuttam, Ari Holtzman, Noah A. Smith, and Hannaneh Hajishirzi. 2024. Quantifying Language Models’ Sensitivity to Spurious Features in Prompt Design. arXiv preprint arXiv:2310.11324 (2024). https://arxiv.org/abs/ 2310.11324

  29. [30]

    Selbst, danah Boyd, Sorelle A

    Andrew D. Selbst, danah Boyd, Sorelle A. Friedler, Suresh Venkatasubramanian, and Janet Vertesi. 2019. Fairness and Abstraction in Sociotechnical Systems. In Proceedings of the Conference on Fairness, Accountability, and Transparency (FAccT ’19). Association for Computing Machinery, New York, NY, USA, 59–68. https://doi.org/10.1145/3287560.3287598

  30. [31]

    Mukund Sharma, Dorottya Demszky, Alex Warstadt, Ethan Perez, and Samuel R. Bowman. 2024. Towards Understanding Sycophancy in Language Models. In Proceedings of the Twelfth International Conference on Learning Representations (ICLR 2024). OpenReview. https://openreview.net/forum?id=17593

  31. [32]

    Jeonghwan So, Jihyung Lee, Sunghyun Park, et al. 2025. Thunder-NUBench: A Benchmark for LLMs’ Sentence-Level Negation Understanding. arXiv preprint arXiv:2506.14397 (2025). https://arxiv.org/abs/2506.14397

  32. [33]

    Thinh Hung Truong, Timothy Baldwin, Karin Verspoor, and Trevor Cohn. 2023. Language Models Are Not Naysayers: An Analysis of Language Models on Negation Benchmarks. In Proceedings of the 12th Joint Conference on Lexical and Computational Semantics (*SEM 2023). Association for Computational Linguistics, Toronto, Canada, 101–117. https://arxiv.org/abs/2306.08189

  33. [34]

    Amos Tversky and Daniel Kahneman. 1981. The Framing of Decisions and the Psychology of Choice. Science 211, 4481 (1981), 453–458. https://doi.org/10.1126/ science.7455683

  34. [35]

    Sandra Wachter, Brent Mittelstadt, and Chris Russell. 2018. Counterfactual Explanations Without Opening the Black Box: Automated Decisions and the GDPR. Harvard Journal of Law & Technology 31, 2 (2018), 841–887. https: //jolt.law.harvard.edu/

  35. [36]

    Boxin Wang, Jiyue Wang, Rui Shao, Mislav Balunovic, Ce Zhang, Bo Li, and Michael Zhang. 2023. DecodingTrust: A Comprehensive Assessment of Trust- worthiness in GPT Models. In Advances in Neural Information Processing Sys- tems (NeurIPS 36). Neural Information Processing Systems Foundation. https: //arxiv.org/abs/2306.11698

  36. [37]

    Alexander Wei, Nika Haghtalab, and Jacob Steinhardt. 2023. Jailbroken: How Does LLM Safety Training Fail?. In Advances in Neural Information Processing Systems (NeurIPS 36) . https://proceedings.neurips.cc/paper_files/paper/2023/ hash/fd6613131889a4b656206c50a8bd7790-Abstract-Conference.html

  37. [38]

    Ziang Zhao, Xianghao Wei, Chunyang Xie, Zongxi Li, Shumin Xu, Shen Feng, Shizhu He, and Kang Liu. 2024. Social Bias Evaluation for Large Language Models Requires Prompt Variations. In Findings of the Association for Computational Lin- guistics: EMNLP 2024. Association for Computational Linguistics, Miami, Florida, USA, 13355–13380. https://aclanthology.or...

  38. [2020]

    In Proceedings of the 2020 Conference on Fair- ness, Accountability, and Transparency (FAccT ’20)

    Closing the AI Accountability Gap: Defining an End-to-End Framework for Internal Algorithmic Auditing. In Proceedings of the 2020 Conference on Fair- ness, Accountability, and Transparency (FAccT ’20). Association for Computing Machinery, New York, NY, USA, 33–44. https://doi.org/10.1145/3351095.3372873

This paper was first reviewed by deepseek-v4-flash on August 3, 2026.