Pith. sign in

REVIEW 3 major objections 4 minor 57 references

Quantization leaves short-form safety checks intact while open-ended generation volunteers stereotypes in roughly one of four answers across eight languages, a gap standard evaluation misses.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 08:32 UTC pith:LVCAFKGS

load-bearing objection Well-built measurement stack and honest hedging; the missing BF16 open-ended cell is exactly the cell the title's causal claim needs. the 3 major comments →

arxiv 2607.21063 v1 pith:LVCAFKGS submitted 2026-07-23 cs.CL cs.CYcs.HC

QuantiBias: Benchmarking Quantization-Induced Bias in LLMs

classification cs.CL cs.CYcs.HC
keywords quantizationLLM biasstereotype biasopen-ended generationsafety evaluationmodel compressionmultilingual biasbenchmark
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Almost every large model is deployed quantized: trained in full precision, then compressed. This paper argues that the principal side effect of that compression is a selective gap: the short-form behaviors that standard safety screens measure—refusal, over-refusal, and multiple-choice bias avoidance—stay flat across the precision ladder, while the same model, asked open-ended questions, volunteers stereotype-endorsing content in roughly one in four answers across all eight languages probed, under an independent judge. The paper introduces QuantiBias, a benchmark that detects this gap by pairing a generative multilingual stereotype probe with refusal and multiple-choice controls, measuring true bits per weight rather than nominal labels, contrasting reasoning on and off, and rating content severity. If the selective gap is real, then a quantized model that passes every standard safety check can still reach users measurably more biased—and a quantized build must be re-evaluated on open-ended generation, not short-form scores alone.

Core claim

Quantization preserves short-form safety behaviors—refusal, over-refusal control, multiple-choice bias avoidance—while open-ended generation grows more biased: an independent judge flags a stereotype in close to one in four open-ended answers across all eight languages. The selective gap is the core finding, present at every precision and under both independent and in-family judges; the compression slope is judge-dependent and treated as provisional. The mechanism: quantizer round-off acts as bounded noise of scale set by the measured bit width, flipping only decisions with narrow logit margins, and open-ended stereotype avoidance is such a narrow-margin behavior because it was never a direc

What carries the argument

The margin-crossing rate R(b) = integral rho(m) Phi(-m/sigma(b)) dm, with sigma(b) ∝ 2^{-b}, is the identity that explains the selectivity: behaviors with margins well above the noise scale stay flat, while near-boundary behaviors flip first. QuantiBias operationalizes this by scoring a generative multilingual probe against measured effective bits per weight, with reasoning on/off and severity ratings, alongside the short-form controls.

Load-bearing premise

The entire result depends on trusting the out-of-family judge's stereotype-endorsing labels as a valid measure of bias; if those labels are an artifact of the judge, the claimed one-in-four rate and the selective gap would not be established.

What would settle it

Re-score the full-sample open-ended generations, including the one-bit rung, with an independent ensemble of out-of-family judges and human annotators. If the one-in-four rate falls to the in-family judge's low single digits, the selective gap would be a labeling artifact rather than a property of quantized models.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • A quantized model can pass refusal, over-refusal, and multiple-choice bias checks while still volunteering stereotypes in roughly one in four open-ended answers; standard safety evaluation alone is therefore insufficient for quantized builds.
  • Quantized models need re-evaluation on open-ended generation; QuantiBias provides a measurement protocol that isolates that channel and rates content severity.
  • Reasoning before answering roughly halves the open-ended stereotype rate on one backbone but leaves it unchanged on another, so reasoning is a family-dependent safeguard, not a universal one.
  • An independent generative bias benchmark sharing no items with the main probe shows the same direction, indicating the effect is not an artifact of a single prompt set.
  • Nominal quantizer labels overstate compression by 11–40%; indexing to measured bits per weight is necessary to compare builds fairly.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If the margin-crossing account is right, the severity of endorsed stereotypes should rise with compression even where the endorsement rate looks flat; a graded severity measurement across the full ladder would test that prediction directly.
  • The same selective-gap logic likely applies to other post-training perturbations—pruning, activation-precision changes, or low-rank updates—because any bounded weight noise will flip whichever behaviors sit behind narrow margins first.
  • Until the full-sample cells are re-scored by an independent ensemble, the only well-supported claim is the selectivity, not the exact one-in-four rate.
  • Because the benchmark isolates open-ended generation, it could be adapted to monitor other safety-adjacent behaviors that are not direct training targets, such as sycophancy or hallucinated justifications, for the same selective degradation under quantization.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. QuantiBias introduces a benchmark for open-ended stereotype bias in quantized LLMs, pairing a multilingual generative probe (MBTP) with refusal, over-refusal, multiple-choice, and capability controls, and indexing results to measured effective bits per weight rather than nominal quantizer labels. Across a Qwen3.6-27B ladder, a Gemma-4-31B ladder, and a five-family screen, the paper reports that standard short-form safeguards (refusal, BBQ, over-refusal) remain flat under compression while open-ended stereotype endorsement is high (~24–27% under an independent judge), and that the compression slope is judge-dependent and explicitly labeled provisional. A margin-crossing model and a calibration-coverage argument are offered as the mechanism, with frequency and severity as separate harm channels.

Significance. If the central result holds, the paper identifies a practically important blind spot: quantized builds can pass standard short-form safety screens while exhibiting high open-ended stereotype bias. The strengths are substantial: results are indexed to measured bpw, decoding is held fixed, the harness is checkpointed and released, independent out-of-family judges are used alongside an in-family judge, the main claims are hedged with unusual precision, and an independent benchmark (CEB) replicates the within-ladder rise. However, the causal claim 'quantization-induced bias' is not yet supported because no full-precision open-ended baseline is reported, and the headline rate rests on a single independent judge that the limitations section says must be re-scored with an ensemble before reliance. These are load-bearing gaps that can be fixed within the manuscript's scope.

major comments (3)
  1. [§5, Table A3, Figure A1] The central claim 'quantization-induced bias' and the abstract's 'principal side effect is increased bias' require comparing quantized builds against the same model at full precision under the same judge and decoding. Table A3 reports MBTP rows only at 5.235, 2.789, and 1.128 bpw; the 16.00 bpw row is a dash. Figure A1 likewise starts at Q4, and the independent-judge rates (0.238/0.267/0.266) have no BF16 anchor. Without that baseline, the one-in-four open-ended rate may simply be the model's full-precision open-ended behavior, and the 'selective gap' would not be attributable to quantization. The margin model in §7, R(b) − R0 ∝ 2^(−2b), also requires an R0 anchor that is never measured. This should be fixed by running the same MBTP cells at 16.00 bpw with both judges and adding them to Table A3 and Figure A1.
  2. [§5, Limitations (1)] The headline 'roughly one in four' is produced by a single independent judge (Claude Sonnet-5) over the full ladder. The limitations section states that the full-sample diagonal cells 'must be re-scored with the independent ensemble before any single level is relied on,' and that the existing pilot covers only two middle rungs at n=160. Since the one-in-four figure is the core evidence for the selective gap, the paper should either perform the ensemble re-scoring for all reported levels or explicitly present the headline as conditional on a single judge. The reported human annotation instability (18.6% flips) reinforces that label validity is not yet established at the level the abstract's language implies.
  3. [§7, Appendix C.2] The calibration-coverage mechanism is supported by the statement 'our data is itself the evidence for that premise': the claim that perplexity and coding accuracy are preserved while open-ended bias degrades is the very phenomenon the paper aims to establish, and it inherits the missing-BF16-baseline problem from M1. This is a circularity concern, not a mere presentation issue. Please either provide independent evidence for low coverage of bias-relevant directions (for example, sensitivity or attribution analyses on calibration data) or explicitly label the empirical premise as an assumption rather than using the observed pattern as both premise and conclusion.
minor comments (4)
  1. [§7, Eq. (1)] The notation 'E∥Δb∥2 = c 2^(−2b)' followed by 'Σb := E[ΔbΔbᵀ] = c 2^(−2b) I' is dimensionally inconsistent: the scalar second moment and the covariance matrix cannot share the same constant c without clarification. State the per-coordinate variance explicitly and use separate constants.
  2. [Table 2 / Appendix D] Several Chinese and Japanese exhibit strings render as mojibake or broken ideographs (e.g., entries for Family structure and Appearance). The verbatim exhibits are central to the severity claim; please ensure the PDF/final rendering preserves the original scripts.
  3. [Author block] The author line contains a formatting artifact: 'ThomasLordDepartmentofComputerScience'. This should be corrected.
  4. [Figure A6] The caption already says the curve is not a fit, which is good, but the plotted curve may still be misread as fitted to the three points. Consider labeling it 'illustrative model curve' directly on the figure, not only in the caption.

Circularity Check

1 steps flagged

Minor circularity in the C.2 mechanism; central empirical benchmark is self-contained.

specific steps
  1. other [Appendix C.2, Proposition 2 discussion (Section 7 mechanism)]
    "The inequality is exact given the allocation model; the empirical premise is that safety and bias directions genuinely have low Ck. Our data is itself the evidence for that premise: perplexity and coding accuracy, which are high-coverage, are preserved, while open-ended bias, which is low-coverage, degrades, exactly the signature Equation2 predicts."

    Proposition 2 derives that low calibration coverage Ck forces large margin-shift variance and hence degradation of that behavior. The paper then uses the observed degradation pattern itself (perplexity/coding preserved, open-ended bias degraded) as the evidence that bias directions have low Ck. Since Ck is never independently measured, the mechanism does not independently predict the selective gap; it restates the observed outcome as the premise and then 'predicts' the same outcome. This is a supporting explanatory step, not the central measurement, so it is minor rather than fatal.

full rationale

QuantiBias is primarily an empirical benchmark paper. The headline selective-gap measurement — open-ended stereotype endorsement near 24–27% under an independent judge while refusal, over-refusal, and BBQ controls stay flat — is an external measurement on fixed prompts with fixed judges; it is not fitted to the benchmark's own outputs. The compression slope is explicitly labeled provisional and judge-dependent, and the margin-crossing model in Section 7 is explicitly labeled 'not a proof' with Figure A6 described as 'not a fit'; the severity slope is called a prediction. The self-citations (Ferrara 2024, 2026; Chand et al. 2026) are motivational and not load-bearing. The only circular step I can exhibit is in Appendix C.2, where the low-calibration-coverage premise is supported by the very data the mechanism is meant to explain ('Our data is itself the evidence for that premise'). This is a real but limited circularity in the explanatory mechanism, not in the central measurement. A separate, non-circular concern — the missing BF16 open-ended MBTP baseline — weakens the causal 'quantization-induced' framing but is not a derivation-to-input reduction.

Axiom & Free-Parameter Ledger

0 free parameters · 4 axioms · 0 invented entities

No fitted constants enter the headline benchmark results; the margin model uses an unspecified noise-scale constant and margin distribution, but these are not fitted to data and the R(b) curve is explicitly not a fit. The empirical rates are direct judge counts. The axioms listed are the contestable premises under the mechanistic explanation and the measurement claim.

axioms (4)
  • domain assumption Uniform round-off with diagonal covariance Σ_b = c 2^{-2b}I
    Used as the exact form (Appendix C, Eq. 1) for the dose-response and precision-floor results; assumes symmetric bounded error and uncorrelated weight perturbations.
  • domain assumption Linearized logit margins: ε_k = ⟨∇_W m_k, Δ_b⟩
    Justifies the Gaussian-tail flip probability R(b); a small-noise approximation that the paper itself says is not expected to hold at the 1-bit rung.
  • ad hoc to paper Margin distribution ρ(m) with wide margins for refusal/MCQ and narrow margins for open-ended stereotype refusal
    Load-bearing structural premise of the mechanistic explanation; no direct margin measurements exist, and support comes only from consistency with the observed pattern.
  • domain assumption Out-of-family LLM judge labels approximate human stereotype judgments
    Ground truth for the one-in-four claim; acknowledged to be judge-family-dependent and not yet validated by the full independent ensemble or human annotation.

pith-pipeline@v1.3.0-alltime-deepseek · 30856 in / 13450 out tokens · 145928 ms · 2026-08-01T08:32:57.422328+00:00 · methodology

0 comments
read the original abstract

Almost every large language model that reaches a broad audience is quantized: trained in full precision, then compressed for efficiency. This step is assumed harmless and its safety is rarely re-checked. We find its principal side effect is increased bias that standard safety evaluation misses. Holding the model, its training, and the prompts fixed, a quantized model still refuses harmful requests, still avoids over-refusing benign prompts, and still selects the unbiased multiple-choice answer. Yet asked an open-ended question, the same model volunteers stereotypes in all eight languages we probe, in roughly one in four open-ended answers under an independent judge (~24% to ~27% across the compression ladder): it passes every standard check and still reaches users measurably more biased. The selective gap is a robust finding; whether open-ended bias further increases with compression is less certain, sensitive to the judge that scores it. We address both with \textbf{QuantiBias}, a benchmark that pairs a generative, multilingual stereotype probe with the refusal and multiple-choice controls that isolate open-ended generation, contrasts each build with and without reasoning, and rates the content severity of what it generates. Across two backbone models (Qwen and Gemma), a five-family screen, and eight benchmarks, quantizers allocate their extra precision by capability data that carries no bias-prevention signal, and reasoning before answering roughly halves the effect on some families while doing nothing on others. A quantized build must be re-evaluated for open-ended bias, not only on the short-form safeguards it already passes.

Figures

Figures reproduced from arXiv: 2607.21063 by Emilio Ferrara.

Figure 1
Figure 1. Figure 1: The selective safety gap under compression. Quantization compresses an aligned model from 16 to about one bit per weight. The short-form checks a release is screened on, refusing harmful requests and avoiding biased multiple-choice answers, still pass. Yet in open-ended generation the same model volunteers stereotypes the checklist never registers, a gap at every precision. QuantiBias measures it: a genera… view at source ↗
Figure 2
Figure 2. Figure 2: The selective safety gap, measured. Safe￾guard, capability, and MBTP benchmarks across the Qwen3.6-27B ladder (more compression rightward). The upper panel holds: refusal, multiple-choice bias avoidance (BBQ), and capability (GSM8K, MMLU￾Redux, declining only at one bit) stay near full precision, so a standard audit reports it unchanged. The lower panel is what it misses: scoring the same reasoning-off gen… view at source ↗
Figure 4
Figure 4. Figure 4: The reasoning safeguard is family￾dependent (lean judge). Lean-judge MBTP endorse￾ment rate, reasoning off versus on, for both backbones at two rungs. Reasoning roughly halves the rate on Qwen but not on Gemma, where the manipulation is verified (2K–3K-character traces). The reasoning contrast, and its limit. Turning reasoning on roughly halves the rate at every rung and flattens the compression slope from… view at source ↗
Figure 5
Figure 5. Figure 5: An independent generative-bias bench￾mark. CEB biased-continuation rate across three rungs (reasoning off, n ≈ 960 per rung), sharing no items with MBTP, under the lean and independent judges. Both rise with compression, independent above lean as with MBTP; the one-bit rung is a separate ternary backbone, so the Q4-to-IQ2 rise on the shared ladder is the robust comparison. eration is thus a family-dependen… view at source ↗
Figure 6
Figure 6. Figure 6: Five-family screen, both judges. Q8 (open circle) and Q2 (filled) stereotype rate per family with 95% intervals: lean judge (Qwen3-8B, n ≤ 320) above, independent (Claude Sonnet-5, n = 60) below, on their own scales. The independent judge sees three to four times the lean rate; within every family the Q8 and Q2 intervals overlap. interval), so the screen names which families dif￾fer in level without settli… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

57 extracted references · 27 linked inside Pith

  1. [1]

    Machine Learning with Applications , volume =

    Ferrara, Emilio , title =. Machine Learning with Applications , volume =. 2024 , doi =

  2. [2]

    AI , volume =

    Chand, Shireen and Baca, Faith and Ferrara, Emilio , title =. AI , volume =. 2026 , doi =

  3. [3]

    and Macready, William G

    Wolpert, David H. and Macready, William G. , title =. IEEE Transactions on Evolutionary Computation , volume =. 1997 , doi =

  4. [4]

    Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics , year =

    Nadeem, Moin and Bethke, Anna and Reddy, Siva , title =. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics , year =

  5. [5]

    , title =

    Parrish, Alicia and Chen, Angelica and Nangia, Nikita and Padmakumar, Vishakh and Phang, Jason and Thompson, Jana and Htut, Phu Mon and Bowman, Samuel R. , title =. Findings of the Association for Computational Linguistics: ACL 2022 , year =

  6. [6]

    Advances in Neural Information Processing Systems (NeurIPS), Datasets and Benchmarks Track , year =

    Souly, Alexandra and Lu, Qingyuan and Bowen, Dillon and others , title =. Advances in Neural Information Processing Systems (NeurIPS), Datasets and Benchmarks Track , year =. 2402.10260 , archivePrefix =

  7. [7]

    Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics , year =

    R. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics , year =

  8. [8]

    arXiv preprint arXiv:2407.02408 , year =

    Wang, Song and Wang, Peng and Zhou, Tong and Dong, Yushun and Tan, Zhen and Li, Jundong , title =. arXiv preprint arXiv:2407.02408 , year =

  9. [9]

    Advances in Neural Information Processing Systems (NeurIPS) , year =

    Egashira, Kazuki and Vero, Mark and Staab, Robin and He, Jingxuan and Vechev, Martin , title =. Advances in Neural Information Processing Systems (NeurIPS) , year =

  10. [10]

    arXiv preprint arXiv:2502.15799 , year =

    Kharinaev, Artyom and Moskvoretskii, Viktor and Shvetsov, Egor and Studenikina, Kseniia and Bykov, Mikhail and Burnaev, Evgeny , title =. arXiv preprint arXiv:2502.15799 , year =

  11. [11]

    arXiv preprint arXiv:2601.12033 , year =

    Al Hakim, Muhammad Alif and Wicaksono, Alfan Farizki and Koto, Fajri , title =. arXiv preprint arXiv:2601.12033 , year =

  12. [12]

    and Rosenfeld, Elan and Kolter, J

    Cohen, Jeremy M. and Rosenfeld, Elan and Kolter, J. Zico , title =. International Conference on Machine Learning (ICML) , year =. 1902.02918 , archivePrefix =

  13. [13]

    International Conference on Machine Learning (ICML) , year =

    Nagel, Markus and Amjad, Rana Ali and van Baalen, Mart and Louizos, Christos and Blankevoort, Tijmen and Welling, Max , title =. International Conference on Machine Learning (ICML) , year =. 2004.10568 , archivePrefix =

  14. [14]

    International Conference on Learning Representations (ICLR) , year =

    Frantar, Elias and Ashkboos, Saleh and Hoefler, Torsten and Alistarh, Dan , title =. International Conference on Learning Representations (ICLR) , year =. 2210.17323 , archivePrefix =

  15. [15]

    and Keutzer, Kurt , title =

    Dong, Zhen and Yao, Zhewei and Gholami, Amir and Mahoney, Michael W. and Keutzer, Kurt , title =. IEEE/CVF International Conference on Computer Vision (ICCV) , year =

  16. [16]

    International Conference on Artificial Intelligence and Statistics (AISTATS) , year =

    Isik, Berivan and Weissman, Tsachy and No, Albert , title =. International Conference on Artificial Intelligence and Statistics (AISTATS) , year =. 2102.08329 , archivePrefix =

  17. [17]

    Beyond Perplexity: Multi-dimensional Safety Evaluation of

    Xu, Zhichao and Gupta, Ashim and Li, Tao and Bentham, Oliver and Srikumar, Vivek , journal=. Beyond Perplexity: Multi-dimensional Safety Evaluation of

  18. [18]

    Proc.\ EACL , year=

    Marcuzzi, Federico and Ning, Xuefei and Schwartz, Roy and Gurevych, Iryna , title=. Proc.\ EACL , year=

  19. [19]

    IEEE Cloud Summit , year=

    Rath, Plawan Kumar and Maliakkal, Rahul , title=. IEEE Cloud Summit , year=. 2605.15208 , archivePrefix=

  20. [20]

    and Lotfi, Sanae and Chen, Irene Y

    Hua, Stanley Z. and Lotfi, Sanae and Chen, Irene Y. , title=. arXiv preprint arXiv:2602.06181 , year=

  21. [21]

    arXiv preprint arXiv:2606.10154 , year=

    Quality Is Not a Safety Proxy Under Quantization , author=. arXiv preprint arXiv:2606.10154 , year=

  22. [22]

    Proc.\ MLSys , year=

    Lin, Ji and Tang, Jiaming and Tang, Haotian and Yang, Shang and Chen, Wei-Ming and Wang, Wei-Chen and Xiao, Guangxuan and Dang, Xingyu and Gan, Chuang and Han, Song , title=. Proc.\ MLSys , year=. 2306.00978 , archivePrefix=

  23. [23]

    Proc.\ ICML , year=

    Xiao, Guangxuan and Lin, Ji and Seznec, Mickael and Wu, Hao and Demouth, Julien and Han, Song , title=. Proc.\ ICML , year=. 2211.10438 , archivePrefix=

  24. [24]

    arXiv preprint arXiv:2402.17764 , year=

    Ma, Shuming and Wang, Hongyu and Ma, Lingxiao and Wang, Lei and Wang, Wenhui and Huang, Shaohan and Dong, Li and Wang, Ruiping and Xue, Jilong and Wei, Furu , title=. arXiv preprint arXiv:2402.17764 , year=

  25. [25]

    2023 , howpublished=

    Gerganov, Georgi and others , title=. 2023 , howpublished=

  26. [26]

    and Andrews, Pierre and Smith, Eric Michael and Hansanti, Prangthip and Ropers, Christophe and Kalbassi, Elahe and Gao, Cynthia and Licht, Daniel and Wood, Carleigh , title=

    Costa-juss\`a, Marta R. and Andrews, Pierre and Smith, Eric Michael and Hansanti, Prangthip and Ropers, Christophe and Kalbassi, Elahe and Gao, Cynthia and Licht, Daniel and Wood, Carleigh , title=. Proc.\ EMNLP , year=

  27. [27]

    arXiv preprint arXiv:2406.07243 , year=

    Neplenbroek, Vera and Bisazza, Arianna and Fern\'andez, Raquel , title=. arXiv preprint arXiv:2406.07243 , year=

  28. [28]

    Proc.\ NAACL , year=

    Mitchell, Margaret and Attanasio, Giuseppe and others , title=. Proc.\ NAACL , year=

  29. [29]

    and Zhang, Hao and Gonzalez, Joseph E

    Zheng, Lianmin and Chiang, Wei-Lin and Sheng, Ying and Zhuang, Siyuan and Wu, Zhanghao and Zhuang, Yonghao and Lin, Zi and Li, Zhuohan and Li, Dacheng and Xing, Eric P. and Zhang, Hao and Gonzalez, Joseph E. and Stoica, Ion , title=. Proc.\ NeurIPS Datasets and Benchmarks , year=. 2306.05685 , archivePrefix=

  30. [30]

    and Feng, Shi , title=

    Panickssery, Arjun and Bowman, Samuel R. and Feng, Shi , title=. Proc.\ NeurIPS , year=. 2404.13076 , archivePrefix=

  31. [31]

    , title=

    Shannon, Claude E. , title=. IRE National Convention Record , volume=

  32. [32]

    and Thomas, Joy A

    Cover, Thomas M. and Thomas, Joy A. , title=. 2006 , doi=

  33. [33]

    and Bialek, William , title=

    Tishby, Naftali and Pereira, Fernando C. and Bialek, William , title=. 37th Allerton Conference on Communication, Control, and Computing , pages=. 1999 , note=. physics/0004057 , archivePrefix=

  34. [34]

    IEEE Information Theory Workshop (ITW) , year=

    Tishby, Naftali and Zaslavsky, Noga , title=. IEEE Information Theory Workshop (ITW) , year=

  35. [35]

    and Solla, Sara A

    LeCun, Yann and Denker, John S. and Solla, Sara A. , title=. Advances in Neural Information Processing Systems (NeurIPS) , year=

  36. [36]

    , title=

    Hassibi, Babak and Stork, David G. , title=. Advances in Neural Information Processing Systems (NeurIPS) , year=

  37. [37]

    Advances in Neural Information Processing Systems (NeurIPS) , year=

    Frantar, Elias and Singh, Sidak Pal and Alistarh, Dan , title=. Advances in Neural Information Processing Systems (NeurIPS) , year=. 2208.11580 , archivePrefix=

  38. [38]

    2019 , note=

    Hooker, Sara and Courville, Aaron and Clark, Gregory and Dauphin, Yann and Frome, Andrea , title=. 2019 , note=. 1911.05248 , archivePrefix=

  39. [39]

    2020 , note=

    Hooker, Sara and Moorosi, Nyalleng and Clark, Gregory and Bengio, Samy and Denton, Emily , title=. 2020 , note=. 2010.03058 , archivePrefix=

  40. [40]

    2022 , note=

    Elhage, Nelson and Hume, Tristan and Olsson, Catherine and others , title=. 2022 , note=. 2209.10652 , archivePrefix=

  41. [41]

    Proc.\ EMNLP-IJCNLP , year=

    Sheng, Emily and Chang, Kai-Wei and Natarajan, Premkumar and Peng, Nanyun , title=. Proc.\ EMNLP-IJCNLP , year=

  42. [42]

    Proc.\ EACL , year=

    Liu, Yang , title=. Proc.\ EACL , year=

  43. [43]

    Proc.\ ACL , year=

    Hartvigsen, Thomas and Gabriel, Saadia and Palangi, Hamid and Sap, Maarten and Ray, Dipankar and Kamar, Ece , title=. Proc.\ ACL , year=

  44. [44]

    Proceedings of ACL , year=

    Kaneko, Masahiro and Bollegala, Danushka and Baldwin, Timothy , title=. Proceedings of ACL , year=

  45. [45]

    Proceedings of NAACL , year=

    Gema, Aryo Pradipta and Leang, Joshua Ong Jun and Hong, Giwon and others , title=. Proceedings of NAACL , year=

  46. [46]

    2021 , note=

    Cobbe, Karl and Kosaraju, Vineet and Bavarian, Mohammad and others , title=. 2021 , note=. 2110.14168 , archivePrefix=

  47. [47]

    Future Internet , volume =

    Ferrara, Emilio , title =. Future Internet , volume =. 2026 , doi =

  48. [48]

    Findings of the Association for Computational Linguistics: ACL 2025 , year=

    Jin, Jiho and Kang, Woosung and Myung, Junho and Oh, Alice , title=. Findings of the Association for Computational Linguistics: ACL 2025 , year=

  49. [49]

    arXiv preprint arXiv:2412.06134 , year=

    Liu, Zhao and Xie, Tian and Zhang, Xueru , title=. arXiv preprint arXiv:2412.06134 , year=

  50. [50]

    NeurIPS 2024 Workshop on Safe Generative AI , year=

    Wataoka, Koki and Takahashi, Tsubasa and Ri, Ryokan , title=. NeurIPS 2024 Workshop on Safe Generative AI , year=. 2410.21819 , archivePrefix=

  51. [51]

    Findings of the Association for Computational Linguistics: EMNLP 2025 , year=

    Wu, Xuyang and Nian, Jinming and Wei, Ting-Ruen and Tao, Zhiqiang and Wu, Hsin-Tai and Fang, Yi , title=. Findings of the Association for Computational Linguistics: EMNLP 2025 , year=

  52. [52]

    Guan, Melody Y. and Joglekar, Manas and Wallace, Eric and Jain, Saachi and Barak, Boaz and Helyar, Alec and Dias, Rachel and Vallone, Andrea and Ren, Hongyu and Wei, Jason and Chung, Hyung Won and Toyer, Sam and Heidecke, Johannes and Beutel, Alex and Glaese, Amelia , title=. arXiv preprint arXiv:2412.16339 , year=

  53. [53]

    Leak, Cheat, Repeat: Data Contamination and Evaluation Malpractices in Closed-Source

    Balloccu, Simone and Schmidtov\'a, Patr\'icia and Lango, Mateusz and Du. Leak, Cheat, Repeat: Data Contamination and Evaluation Malpractices in Closed-Source. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (EACL) , year=

  54. [54]

    Advances in Neural Information Processing Systems (NeurIPS) , year=

    Tran, Cuong and Fioretto, Ferdinando and Kim, Jung-Eun and Naidu, Rakshit , title=. Advances in Neural Information Processing Systems (NeurIPS) , year=. 2205.13574 , archivePrefix=

  55. [55]

    arXiv preprint arXiv:2306.06238 , year=

    Dam, Harvey and Joseph, Vinu and Bhaskara, Aditya and Gopalakrishnan, Ganesh and Muralidharan, Saurav and Garland, Michael , title=. arXiv preprint arXiv:2306.06238 , year=

  56. [56]

    , title=

    Yong, Zheng-Xin and Menghini, Cristina and Bach, Stephen H. , title=. NeurIPS Workshop on Socially Responsible Language Modelling Research (SoLaR) , year=. 2310.02446 , archivePrefix=

  57. [57]

    2026 , eprint =

    Ghafouri, Bijean and Choi, Eun Cheol and Dey, Priyanka and Ferrara, Emilio , booktitle =. 2026 , eprint =