Pith. sign in

REVIEW 4 major objections 5 minor 300 references

A different-family second model that reads a judge's reasoning trace corrects LLM bias better than any single fixed auditor.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-03 00:50 UTC pith:SLG2VHPZ

load-bearing objection Real insight about cross-family auditors and the limits of standalone resistance, but a protocol contradiction in the appendix and an overstrong claim need fixing. the 4 major comments →

arxiv 2607.28636 v1 pith:SLG2VHPZ submitted 2026-05-19 cs.CL cs.CY

Chain-of-Models: Cross-Model Auditing for Bias-Robust LLM Judges

classification cs.CL cs.CY
keywords LLM-as-judgecognitive biascross-model auditingreasoning traceauditor selectionfunctional diversitysycophancybias mitigation
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper tries to establish that the best way to correct a biased LLM judge is not to prompt it harder or swap in the most bias-resistant model, but to add a second model from a different training family that reads the judge's full reasoning trace and issues the final judgment. Across nine models from six families, four cognitive biases, and four factual datasets, it argues that this auditor must be chosen per bias type: the model most resistant to a bias when answering alone is often the worst at correcting another model's biased reasoning, and no single auditor wins on every bias. The paper operationalizes this as Chain-of-Models (CoM), a routing rule that scores candidate auditors by functional diversity, standalone bias resistance, and a calibrated audit-effectiveness estimate, and shows under a calibration/test split that the rule reaches 0.884 accuracy across four biased slices, beating the strongest fixed auditor (0.824) and the no-audit baseline (0.805). A sympathetic reader would care because this offers a deployable, weight-free way to improve LLM-as-judge pipelines in high-stakes domains where human evaluation does not scale.

Core claim

Auditor identity, not standalone quality, decides whether a second model corrects a biased judge. With generator fixed at Qwen2.5-72B-Instruct, standalone bias resistance does not predict audit effectiveness: Kimi-K2.5, most resistant standalone on three biases, is the weakest auditor on bandwagon (0.400) and authority (0.500). The best auditor is bias-specific: GPT-4o wins bandwagon, authority, and distraction; GLM-5 wins sycophancy, where GPT-4o auditing is worse than no audit. A per-bias rule combining functional diversity, standalone resistance, and calibrated audit effectiveness reaches 0.884 test accuracy versus 0.824 for the best fixed auditor and 0.805 for no audit.

What carries the argument

The load-bearing mechanism is the two-model Chain-of-Models pipeline: M1 produces a reasoning trace and answer, then M2 receives the same prompt plus M1's full trace and re-judges, with no bias-specific instructions. Auditor selection is driven by a weighted score for each candidate: functional diversity (cosine distance between the generator's and candidate's behavioral fingerprint vectors), per-bias standalone resistance (candidate's own biased accuracy), and calibrated audit effectiveness e(M1, M2, b), the chain accuracy estimated on a held-out calibration split. The e term is what rules out the highly standalone-resistant Kimi-K2.5; conditioning on bias type b is what routes sycophancy t

Load-bearing premise

The selection advantage depends on knowing each query's bias type at inference time and on the 25-example calibration estimate of audit effectiveness being accurate enough to rank candidate auditors reliably.

What would settle it

Withhold the bias label and infer it with a learned classifier, or bootstrap the 25 calibration examples per cell: if the chosen auditor frequently flips or the selector's margin over always-on GPT-4o disappears, the central claim fails.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • A deployed LLM judge can be made more accurate on biased inputs without retraining or prompt engineering, by routing flagged queries to the auditor calibrated for that bias.
  • Standalone bias-resistance leaderboards should not be used to choose auditors; a model's own robustness says little about how well it corrects another model's biased reasoning.
  • Same-model or same-family self-audit is insufficient; the auditor's training lineage should differ from the judge's to catch correlated blind spots.
  • Per-bias routing beats always-on auditing at the same per-query cost, with the gain concentrated on the bias where the fixed auditor is weakest.
  • The selection rule transfers to subjective preference judging, but audit-effectiveness estimates must be re-estimated on the target domain.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The resistance/effectiveness inversion suggests audit quality is about complementary failure modes between generator and auditor, not individual capability; a cheap trace-level disagreement score might predict it without costly calibration runs.
  • The rule assumes the bias type is known per query; a learned bias detector tested end-to-end would determine whether the method works on unlabeled in-the-wild prompts.
  • With only 25 calibration examples per bias–dataset cell, the effectiveness estimate is noisy; bootstrapping those cells would show how often the auditor ranking flips, and larger calibration sets could justify learning the scoring weights instead of fixing them.
  • The best auditor differing between factual and subjective sycophancy implies auditor rankings can be task-family-labile, so transferring a routing policy across domains without recalibration is risky.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper proposes Chain-of-Models (CoM), a sequential auditing pipeline in which a second LLM receives the first model's reasoning trace and produces a final judgment. Fixing the generator as Qwen2.5-72B-Instruct, the authors evaluate five cross-family auditors and one same-family ladder across four cognitive biases on MMLU-Pro and a subjective DPO sycophancy track. They report two main empirical findings: standalone bias resistance does not predict audit effectiveness (e.g., Kimi-K2.5 is most standalone-resistant but the weakest auditor on bandwagon and authority), and the best auditor is bias-specific (GPT-4o for bandwagon/authority/distraction; GLM-5 for factual sycophancy). They then propose a per-bias selector scoring auditors by functional DNA distance, standalone resistance, and calibration-split audit effectiveness, and claim 0.884 held-out accuracy across biased slices versus 0.824 for always-on GPT-4o and 0.805 for no audit.

Significance. If substantiated, the paper would make a useful contribution: it treats auditor identity as a design variable, provides a concrete selection rule, and releases code and configurations. The calibration/test split is a sound way to avoid selection-on-test, and the two headline findings are practically relevant. However, the protocol inconsistency in Appendix K.3 and the lack of uncertainty quantification currently prevent acceptance.

major comments (4)
  1. [§2.1 / §G.3 vs. Appendix K.3] The stated protocol and Appendix K.3 contradict each other. Section 2.1 says the auditor 'is not told that bias may be present' and G.3 says 'the original prompt—including any injected bias cue—is passed verbatim to Mi.' Appendix K.3, described as 'the key mechanism,' instead shows M2 receiving the clean prompt without the bandwagon cue and with the trace labeled 'Advisory: Another analyst's reasoning (may contain biases).' If K.3 reflects the actual experimental condition, the reported chain accuracies are confounded by cue removal and explicit bias flagging, and the central claim that a cross-family auditor resists the same cue M1 fell for is not supported. The authors must reconcile this: either K.3 is outside the evaluation protocol and must be relabeled/replaced with a verbatim G.3 trace, or G.3 is wrong and the experiments must be redone.
  2. [§3.3, Eq. (4), Table 12] The paper claims the selector combines functional diversity, standalone resistance, and empirical audit effectiveness, but the evidence shows the empirical-effectiveness term alone reproduces the result. Table 12's 'effectiveness-only' weighting (0,0,1) selects GPT-4o with the same 0.830 authority accuracy as the default (0.2,0.3,0.5); the diversity-only and resistance-only weightings collapse. No ablation tests whether removing d or r from the full score changes the selection. The three-axis claim is therefore not supported by the reported data. Please either demonstrate that d and r contribute beyond e (e.g., ablations that drop each term, or cases where e is unavailable) or reframe the selector as primarily e-based.
  3. [§3.3, Table 7; Appendix F] The calibration split uses only n=25 examples per bias–dataset cell to estimate e(M1, Mi, b), and all reported accuracies are point estimates without confidence intervals. With 25 calibration examples, the auditor ranking can easily flip under sampling noise; the overall 0.884 vs. 0.824 improvement could be inflated by a lucky calibration draw, especially since the entire gain is concentrated in the sycophancy slice. The paper acknowledges this in the limitations but does not quantify the risk. Please provide bootstrap confidence intervals, a calibration-size sensitivity analysis, or a larger calibration set.
  4. [§3.2] The statement that audit effectiveness 'cannot be predicted' from standalone bias resistance is stronger than the evidence supports. The paper demonstrates two counterexamples (Kimi-K2.5 and DeepSeek-V3) and shows that the best standalone model is not the best auditor, but this does not establish a general impossibility of any prediction from standalone resistance. A correlational analysis or a more tempered wording—e.g., 'standalone resistance is not a reliable predictor'—would match the data. This matters because the finding is one of the two headline claims motivating the selector.
minor comments (5)
  1. [Appendix C] Appendix C references 'the D-6 collapse on bandwagon (§3.2),' but Section 3.2 does not mention D-6 chains. Either add the D-6 results to the main text or fix the cross-reference.
  2. [Appendix F vs. Appendix I] Appendix F says 'An earlier trace-only detector pilot reported in §I had much lower recall on bandwagon and distraction,' but Appendix I does not describe such a pilot. The cross-reference appears to point to nonexistent content.
  3. [References] Several references are incomplete or placeholder-style, e.g., 'Jane Li and Others. 2024' and 'Nurit Cohen Inger and 1 others. 2026 ... Co-author list pending verification before camera-ready.' These need full author lists and verified details for a journal submission.
  4. [Appendix K.3] The K.3 trace uses GPT-4o-mini as M1, whereas the main experiments fix M1 as Qwen2.5-72B-Instruct. If this trace is intended as an illustrative example rather than a main-protocol run, it should be explicitly labeled as such.
  5. [§2.2] The text says 'over 500 experiments (∼100,000 API calls),' but the stated design (9 models × 4 biases × 4 datasets × 50 plus chains) corresponds to a different count. Please clarify the arithmetic or describe how the count was derived.

Circularity Check

0 steps flagged

No significant circularity: the selector is calibrated on a disjoint split and the headline result is an empirical evaluation, not a derivation from its own inputs.

full rationale

The paper's only predictive claim is the per-bias selector (Eq. 4), which combines d, r, and e. The empirical-effectiveness term e is estimated on a held-out calibration split (25 examples per bias-dataset cell) and the reported 0.884 accuracy is measured on disjoint test examples, so the headline result is not fitted to the target. The finding that standalone resistance does not predict audit effectiveness is an empirical comparison of measured r and e, not a definitional equivalence. The functional-diversity term comes from Wu et al. (2025), which shares two co-authors with this paper, but that citation is not load-bearing: the Appendix H ablation shows that dropping e (while retaining d and r) selects Kimi-K2.5 and collapses to 0.480, while the default selection is driven mainly by e. I also flag a serious internal inconsistency that is a validity concern rather than a circularity: Section 2.1 and Appendix G.3 state that the auditor receives the original biased prompt verbatim and is not told bias may be present, whereas Appendix K.3 describes the auditor as seeing the clean prompt without the bandwagon cue plus an 'Advisory: Another analyst's reasoning (may contain biases)' label and calls this 'the key mechanism.' If K.3 reflects the actual experimental protocol, the chain-accuracy measurements that populate e are confounded by cue removal and bias flagging, but that is a correctness/confound issue, not a derivation-equivalence or self-citation circularity.

Axiom & Free-Parameter Ledger

3 free parameters · 4 axioms · 0 invented entities

The central result relies on a hand-set routing rule with three terms, a known-bias assumption, and a calibration split of 25 samples. No new physical entities are introduced. The diversity metric comes from the authors' own prior work, adding a mild self-referential component.

free parameters (3)
  • Routing weights (alpha, beta, gamma) = (0.2, 0.3, 0.5)
    Hand-set to prioritize empirical effectiveness; swept in Appendix H but not learned from data. Selection rule collapses to 0.480 when e is omitted.
  • Calibration split size = 25 examples per (bias, dataset) cell
    Chosen split; e estimates are noisy at this n and no confidence intervals are reported.
  • Bias cue injection templates = e.g., '87%', '90%'
    Cue strengths are fixed and not varied; detector achieves 100% recall by construction on these templates.
axioms (4)
  • domain assumption Functional diversity d(Mi, Mj) from Wu et al. (2025) captures shared blind spots
    Eq. 4 includes d; the paper inherits the LLM-DNA representation without revalidating it on these models besides Figure 2 distances.
  • domain assumption Bias type b is known at inference time
    Section 3.3 states b comes from data-source labels or deployment context; if not, the selector cannot route.
  • domain assumption Auditors are not informed that bias may be present
    G.3 prompt template omits bias warnings; however Appendix K.3 shows a bias advisory, creating inconsistency.
  • standard math Cosine distance is a valid functional diversity metric
    Def. 1 uses cosine distance; standard, but the probes are only 5 prompts.

pith-pipeline@v1.3.0-alltime-deepseek · 22269 in / 11234 out tokens · 92649 ms · 2026-08-03T00:50:04.110270+00:00 · methodology

0 comments
read the original abstract

LLMs increasingly serve as automated judges, but their judgments remain vulnerable to cognitive biases. Existing mitigations mostly rely on prompt-driven debiasing, which is brittle across bias types, or human evaluation, which does not scale. We study \emph{Chain-of-Models} (CoM), an automated audit pipeline in which a second model inspects the first model's reasoning trace before producing the final judgment. The key design question is whether the auditor should be the same model, a same-family model, or a different-family model. Across 9 models from 6 families, 4 cognitive biases, and 4 factual datasets, we find that auditor identity matters in two ways. First, standalone bias resistance does not predict audit effectiveness: Kimi-K2.5 is the strongest standalone model on several biases, yet is a weak auditor for Qwen2.5-72B's biased traces. Second, the best auditor is bias-specific: GPT-4o is strongest on bandwagon, authority, and distraction, while GLM-5 is strongest on sycophancy. We operationalize these findings with a per-bias auditor selection rule that, given the bias type, scores candidates along functional diversity, per-bias standalone resistance, and calibrated audit effectiveness. Under a calibration/test split, the selector reaches the highest accuracy across the four biased slices ($0.884$ vs.\ $0.824$ for the strongest single fixed auditor and $0.805$ for the no-audit baseline). We release data, configurations, and an LLM-agent skill at https://anonymous.4open.science/r/chain-of-models-B585 .

Figures

Figures reproduced from arXiv: 2607.28636 by Bingsheng He, Nuo Chen, Qian Wang, Zhanzhi Lou, Zhenheng Tang.

Figure 1
Figure 1. Figure 1: Auditor identity matters. (A) Bias blind spot intuition: bias is easier to detect in another reasoner than in oneself, so the auditor should not be the same model as the judge. (B) M1 answers under a cognitive-bias cue: valid evidence in its trace supports answer B, but the cue pushes toward A and M1 pivots to the wrong answer. (C) A different-family auditor M2 receives M1’s trace and final answer, challen… view at source ↗
Figure 2
Figure 2. Figure 2: Pairwise functional DNA distances (co￾sine, computed over 5 probe prompts; cf. Wu et al., 2025). Bold cells: within-family pairs (Qwen→Qwen and GPT-mini→GPT-4o). Kimi-K2.5 is the largest func￾tional outlier (distances 0.103–0.145 to all other mod￾els); Qwen↔GPT distances (0.045–0.065) are at within￾family levels, indicating functional convergence despite distinct training organizations. family gain we repo… view at source ↗
Figure 3
Figure 3. Figure 3: Two bias-induced reasoning patterns visible in the trace. Left: acknowledge-but-defer (factual)— M1 identifies correct evidence then abandons it un￾der authority pressure. Right: fabricated justification (subjective)—M1 invents a rationale aligned with the bandwagon cue. Both signatures are exposed in the trace and detectable by a heterogeneous M2. 50 samples—enough to detect the headline effects but not s… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

300 extracted references · 1 canonical work pages

  1. [1]

    Personality and Social Psychology Bulletin , volume=

    The bias blind spot: Perceptions of bias in self versus others , author=. Personality and Social Psychology Bulletin , volume=. 2002 , publisher=

  2. [2]

    1957 , publisher=

    A Theory of Cognitive Dissonance , author=. 1957 , publisher=

  3. [3]

    Psychological Bulletin , volume=

    The case for motivated reasoning , author=. Psychological Bulletin , volume=. 1990 , publisher=

  4. [4]

    Essai sur l'application de l'analyse

    de Condorcet, Marquis , year=. Essai sur l'application de l'analyse

  5. [5]

    Wu, Zhaomin and Zhao, Haodong and Wang, Ziyang and Guo, Jizhou and Wang, Qian and He, Bingsheng , journal=

  6. [6]

    International Workshop on Multiple Classifier Systems , pages=

    Ensemble methods in machine learning , author=. International Workshop on Multiple Classifier Systems , pages=. 2000 , publisher=

  7. [7]

    International Conference on Learning Representations , year=

    Self-Consistency Improves Chain of Thought Reasoning in Language Models , author=. International Conference on Learning Representations , year=

  8. [8]

    International Conference on Machine Learning , year=

    Improving Factuality and Reasoning in Language Models through Multiagent Debate , author=. International Conference on Machine Learning , year=

  9. [9]

    Min, Sewon and Krishna, Kalpesh and Lyu, Xinxi and Lewis, Mike and Yih, Wen-tau and Koh, Pang Wei and Iyyer, Mohit and Zettlemoyer, Luke and Hajishirzi, Hannaneh , booktitle=

  10. [10]

    International Conference on Learning Representations , year=

    Teaching Large Language Models to Self-Debug , author=. International Conference on Learning Representations , year=

  11. [11]

    arXiv preprint arXiv:2305.20050 , year=

    Let's Verify Step by Step , author=. arXiv preprint arXiv:2305.20050 , year=

  12. [12]

    International Conference on Learning Representations , year=

    Mixture-of-Agents Enhances Large Language Model Capabilities , author=. International Conference on Learning Representations , year=

  13. [13]

    International Conference on Learning Representations , year=

    Towards Understanding Sycophancy in Language Models , author=. International Conference on Learning Representations , year=

  14. [14]

    arXiv preprint arXiv:2503.10814 , year=

    Thinking machines: A survey of llm based reasoning strategies , author=. arXiv preprint arXiv:2503.10814 , year=

  15. [15]

    arXiv preprint arXiv:2402.14016 , year=

    Is llm-as-a-judge robust? investigating universal adversarial attacks on zero-shot llm assessment , author=. arXiv preprint arXiv:2402.14016 , year=

  16. [16]

    Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) , pages=

    Deceptive Semantic Shortcuts on Reasoning Chains: How Far Can Models Go without Hallucination? , author=. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) , pages=

  17. [17]

    arXiv preprint arXiv:2410.05229 , year=

    Gsm-symbolic: Understanding the limitations of mathematical reasoning in large language models , author=. arXiv preprint arXiv:2410.05229 , year=

  18. [18]

    arXiv preprint arXiv:2502.16169 , year=

    Patterns over principles: The fragility of inductive reasoning in llms under noisy observations , author=. arXiv preprint arXiv:2502.16169 , year=

  19. [19]

    arXiv preprint arXiv:2506.04210 , year=

    Does Thinking More always Help? Understanding Test-Time Scaling in Reasoning Models , author=. arXiv preprint arXiv:2506.04210 , year=

  20. [20]

    Yuanbing Zhu and Zhenheng Tang and Xiang Liu and Ang Li and Bo Li and Xiaowen Chu and Bo Han , booktitle=. Oracle. 2025 , url=

  21. [21]

    2025 , eprint=

    One Token to Fool LLM-as-a-Judge , author=. 2025 , eprint=

  22. [22]

    2025 , eprint=

    Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety , author=. 2025 , eprint=

  23. [23]

    JailbreakLo

    Fanjunduo Wei and Zhenheng Tang and Rongfei Zeng and Tongliang Liu and Chengqi Zhang and Xiaowen Chu and Bo Han , booktitle=. JailbreakLo. 2025 , url=

  24. [24]

    ICML 2025 Workshop on Data in Generative Models - The Bad, the Ugly, and the Greats , year=

    Ghost in the Cloud: Your Geo-Distributed Large Language Models Training is Easily Manipulated , author=. ICML 2025 Workshop on Data in Generative Models - The Bad, the Ugly, and the Greats , year=

  25. [25]

    ICLR 2025 Workshop on Foundation Models in the Wild , year=

    AgentTaxo: Dissecting and Benchmarking Token Distribution of LLM Multi-Agent Systems , author=. ICLR 2025 Workshop on Foundation Models in the Wild , year=

  26. [26]

    The 63rd Annual Meeting of the Association for Computational Linguistics , year=

    MegaAgent: A Large-Scale Autonomous LLM-based Multi-Agent System Without Predefined SOPs , author=. The 63rd Annual Meeting of the Association for Computational Linguistics , year=

  27. [27]

    2025 , journal=

    Mediator: Memory-efficient LLM Merging with Less Parameter Conflicts and Uncertainty Based Routing , author=. 2025 , journal=

  28. [28]

    2025 , journal=

    Can LLMs Maintain Fundamental Abilities under KV Cache Compression? , author=. 2025 , journal=

  29. [29]

    2025 , journal=

    ChunkKV: Semantic-Preserving KV Cache Compression for Efficient Long-Context LLM Inference , author=. 2025 , journal=

  30. [30]

    Proceedings of the 42th International Conference on Machine Learning , series =

    Can Compressed LLMs Truly Act? An Empirical Evaluation of Agentic Capabilities in LLM Compression , author=. Proceedings of the 42th International Conference on Machine Learning , series =

  31. [31]

    Qian Wang and Zhenheng Tang and Bingsheng He , booktitle=. Can

  32. [32]

    The Lottery

    Zhenheng Tang and Xiang Liu and Qian Wang and Peijie Dong and Bingsheng He and Xiaowen Chu and Bo Li , booktitle=. The Lottery

  33. [33]

    2024 , journal=

    FusionLLM: A Decentralized LLM Training System on Geo-distributed GPUs with Adaptive Compression , author=. 2024 , journal=

  34. [34]

    arXiv preprint arXiv:2308.13387 , year=

    Do-not-answer: A dataset for evaluating safeguards in llms , author=. arXiv preprint arXiv:2308.13387 , year=

  35. [35]

    Journal of Experimental Social Psychology , volume=

    The effort heuristic , author=. Journal of Experimental Social Psychology , volume=. 2004 , publisher=

  36. [36]

    2025 , eprint=

    Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge , author=. 2025 , eprint=

  37. [37]

    2024 , eprint=

    A Survey on Data Selection for LLM Instruction Tuning , author=. 2024 , eprint=

  38. [38]

    arXiv preprint arXiv:2309.07045 , year=

    Safetybench: Evaluating the safety of large language models with multiple choice questions , author=. arXiv preprint arXiv:2309.07045 , year=

  39. [39]

    arXiv preprint arXiv:2407.17436 , year=

    Air-bench 2024: A safety benchmark based on risk categories from regulations and policies , author=. arXiv preprint arXiv:2407.17436 , year=

  40. [40]

    arXiv preprint arXiv:2406.14598 , year=

    Sorry-bench: Systematically evaluating large language model safety refusal behaviors , author=. arXiv preprint arXiv:2406.14598 , year=

  41. [41]

    arXiv preprint arXiv:2501.09686 , year=

    Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models , author=. arXiv preprint arXiv:2501.09686 , year=

  42. [42]

    arXiv preprint arXiv:2504.18333 , year=

    Adversarial Attacks on LLM-as-a-Judge Systems: Insights from Prompt Injections , author=. arXiv preprint arXiv:2504.18333 , year=

  43. [43]

    arXiv preprint arXiv:2408.01605 , year=

    Cyberseceval 3: Advancing the evaluation of cybersecurity risks and capabilities in large language models , author=. arXiv preprint arXiv:2408.01605 , year=

  44. [44]

    arXiv preprint arXiv:2402.05044 , year=

    Salad-bench: A hierarchical and comprehensive safety benchmark for large language models , author=. arXiv preprint arXiv:2402.05044 , year=

  45. [45]

    arXiv preprint arXiv:2404.13161 , year=

    Cyberseceval 2: A wide-ranging cybersecurity evaluation suite for large language models , author=. arXiv preprint arXiv:2404.13161 , year=

  46. [46]

    First Conference on Language Modeling , year=

    AutoDAN: interpretable gradient-based adversarial attacks on large language models , author=. First Conference on Language Modeling , year=

  47. [47]

    Advances in Neural Information Processing Systems , volume=

    Jailbroken: How does llm safety training fail? , author=. Advances in Neural Information Processing Systems , volume=

  48. [48]

    arXiv preprint arXiv:2406.18510 , year=

    WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language Models , author=. arXiv preprint arXiv:2406.18510 , year=

  49. [49]

    arXiv preprint arXiv:2410.09024 , year=

    Agentharm: A benchmark for measuring harmfulness of llm agents , author=. arXiv preprint arXiv:2410.09024 , year=

  50. [50]

    arXiv preprint arXiv:2410.02644 , year=

    Agent security bench (asb): Formalizing and benchmarking attacks and defenses in llm-based agents , author=. arXiv preprint arXiv:2410.02644 , year=

  51. [51]

    arXiv preprint arXiv:2312.14197 , year=

    Benchmarking and defending against indirect prompt injection attacks on large language models , author=. arXiv preprint arXiv:2312.14197 , year=

  52. [52]

    arXiv preprint arXiv:2403.02691 , year=

    Injecagent: Benchmarking indirect prompt injections in tool-integrated large language model agents , author=. arXiv preprint arXiv:2403.02691 , year=

  53. [53]

    arXiv preprint arXiv:2308.01263 , year=

    Xstest: A test suite for identifying exaggerated safety behaviours in large language models , author=. arXiv preprint arXiv:2308.01263 , year=

  54. [54]

    arXiv preprint arXiv:2307.15043 , year=

    Universal and transferable adversarial attacks on aligned language models , author=. arXiv preprint arXiv:2307.15043 , year=

  55. [55]

    arXiv preprint arXiv:2408.15221 , year=

    Llm defenses are not robust to multi-turn human jailbreaks yet , author=. arXiv preprint arXiv:2408.15221 , year=

  56. [56]

    arXiv preprint arXiv:2410.05295 , year=

    Autodan-turbo: A lifelong agent for strategy self-exploration to jailbreak llms , author=. arXiv preprint arXiv:2410.05295 , year=

  57. [57]

    arXiv preprint arXiv:2404.07921 , year=

    Amplegcg: Learning a universal and transferable generative model of adversarial suffixes for jailbreaking both open and closed llms , author=. arXiv preprint arXiv:2404.07921 , year=

  58. [58]

    arXiv preprint arXiv:2501.12948 , year=

    Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning , author=. arXiv preprint arXiv:2501.12948 , year=

  59. [59]

    arXiv preprint arXiv:2412.19437 , year=

    Deepseek-v3 technical report , author=. arXiv preprint arXiv:2412.19437 , year=

  60. [60]

    2025 , url =

    OpenAI , title =. 2025 , url =

  61. [61]

    2024 , eprint=

    CYBERSECEVAL 3: Advancing the Evaluation of Cybersecurity Risks and Capabilities in Large Language Models , author=. 2024 , eprint=

  62. [62]

    2024 , eprint=

    WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs , author=. 2024 , eprint=

  63. [63]

    arXiv preprint arXiv:2407.21783 , year=

    The llama 3 herd of models , author=. arXiv preprint arXiv:2407.21783 , year=

  64. [64]

    arXiv preprint arXiv:2406.12845 , year=

    Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts , author=. arXiv preprint arXiv:2406.12845 , year=

  65. [65]

    arXiv preprint arXiv:2409.10164 , year=

    Quantile regression for distributional reward models in rlhf , author=. arXiv preprint arXiv:2409.10164 , year=

  66. [66]

    arXiv preprint arXiv:2403.13787 , year=

    Rewardbench: Evaluating reward models for language modeling , author=. arXiv preprint arXiv:2403.13787 , year=

  67. [67]

    arXiv preprint arXiv:2410.21276 , year=

    Gpt-4o system card , author=. arXiv preprint arXiv:2410.21276 , year=

  68. [68]

    arXiv preprint arXiv:2310.03693 , year=

    Fine-tuning aligned language models compromises safety, even when users do not intend to! , author=. arXiv preprint arXiv:2310.03693 , year=

  69. [69]

    2017 , publisher =

    Expert Political Judgment: How Good Is It? How Can We Know? -- New Edition , author =. 2017 , publisher =

  70. [70]

    1993 , publisher=

    Ideals and illusions: On reconstruction and deconstruction in contemporary critical theory , author=. 1993 , publisher=

  71. [71]

    arXiv preprint arXiv:2501.18841 , year=

    Trading inference-time compute for adversarial robustness , author=. arXiv preprint arXiv:2501.18841 , year=

  72. [72]

    arXiv preprint arXiv:2411.01111 , year=

    Rule based rewards for language model safety , author=. arXiv preprint arXiv:2411.01111 , year=

  73. [73]

    2024 , author=

    Deliberative Alignment: Reasoning Enables Safer Language Models. 2024 , author=. URL https://arxiv. org/abs/2412.16339 , year=

  74. [74]

    arXiv preprint arXiv:2410.12784 , year=

    Judgebench: A benchmark for evaluating llm-based judges , author=. arXiv preprint arXiv:2410.12784 , year=

  75. [75]

    2020 , eprint=

    BERTScore: Evaluating Text Generation with BERT , author=. 2020 , eprint=

  76. [76]

    arXiv preprint arXiv:2404.06395 , year=

    Minicpm: Unveiling the potential of small language models with scalable training strategies , author=. arXiv preprint arXiv:2404.06395 , year=

  77. [77]

    arXiv preprint arXiv:2308.12261 , year=

    Prompt2model: Generating deployable models from natural language instructions , author=. arXiv preprint arXiv:2308.12261 , year=

  78. [78]

    Py-DPO Dataset , howpublished =

  79. [79]

    NSFW (not safe for work) content DPO dataset , howpublished =

  80. [80]

    2023 , author =

    LMSys Chat Platform , howpublished =. 2023 , author =

Showing first 80 references.