REVIEW 3 major objections 5 minor 8 references
A dataset of 951 expert-rated conceptual critiques shows that LLM judgment of arguments tracks general model capability, while vendor thinking modes do not systematically help.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-31 01:25 UTC pith:WVM5C6Y7
load-bearing objection Solid new expert-rated multi-axis dataset for conceptual critique judgment; capability tracking is real but partly confounded by style/source cues the authors already flag. the 3 major comments →
A dataset of rated conceptual arguments
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Contextualized arguments on conceptual questions can be rated along multiple expert dimensions more reliably than bottom-line conclusions. On a new dataset of 951 such critiques with 1,458 expert ratings, language-model judges’ agreement with those ratings tracks general capability rankings, while enabling thinking or reasoning modes does not systematically raise scores.
What carries the argument
A multi-axis expert rubric (centrality, strength, correctness, clarity, dead weight, single issue, overall) plus two scoring functions: a weighted pairwise ranking error over critiques of the same position, and a custom weighted absolute-error loss that scores the product strength×centrality rather than the two axes separately.
Load-bearing premise
The expert multi-axis ratings—especially overall and strength times centrality—are stable enough ground truth for ranking models, despite subjectivity, interpretation disputes, and style confounds between strong and weak critiques.
What would settle it
Collect a large new set of double-blind expert ratings on held-out critiques; if independent experts reverse the existing overall or strength×centrality orderings at scale, or if weaker models match frontier models once source and style cues are equalized, the claim that the scores measure conceptual judgment capability fails.
If this is right
- Models can be ranked as conceptual-argument judges without needing settled answers on philosophy or AI safety.
- The same multi-axis critique scores can serve as a target for training or selecting models that assist on conceptual questions.
- Debate-style oversight in domains without ground truth can reuse this critique-evaluation setup.
- Reasoning-mode training aimed at math and code does not automatically transfer to conceptual critique judgment.
- Human–human agreement still beats the best models, so the benchmark is not yet saturated.
Where Pith is reading between the lines
- The rating pipeline could supply preference data to train critique-writing or multi-agent debate models specifically for conceptual domains.
- Because model-written critiques are disproportionately weak and long author-written ones almost always strong, future releases will need style-matched and adversarially hard negatives to keep the signal clean.
- If strength×centrality carries most of the useful signal, simpler two-axis labeling interfaces might scale human annotation without losing ranking power.
- The weak transfer from thinking modes suggests conceptual judgment may need post-training objectives different from verifiable domains.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces a dataset of 951 expert-rated argumentative critiques of 442 position texts on conceptual questions (AI safety, decision theory, ethics, politics, etc.), with 1,458 multi-axis ratings (centrality, strength, correctness, clarity, dead weight, single issue, overall) from six expert raters. It motivates evaluating contextualized arguments rather than bottom-line conclusions when ground truth is unavailable, fully specifies a rubric (§F), analyzes scoring pitfalls (§A–B), and defines two metrics: a weighted pairwise ranking error rate and a custom multi-axis loss that uses strength×centrality. Baseline CoT benchmarking (Tables 5–7, §4) finds that model performance tracks general capability rankings while vendor “thinking”/reasoning modes do not systematically help. Inter-rater validation and a small multi-rater test set are reported in §C; limitations including style/source confounds are stated in §2.
Significance. If the ratings are sufficiently stable ground truth, this is a useful public resource for a capability class that standard verifiable benchmarks miss and that is directly relevant to AI-safety debate, cooperative AI, and philosophical reasoning. Strengths include an unusually detailed rubric, explicit treatment of strength–centrality ambiguity via the product, documented human–human gaps (§C: CO/AK/CN vs EC custom losses ~0.14–0.16 vs GPT-5 ~0.21), and a clear negative finding that current reasoning-mode training does not transfer. The work is primarily a dataset-and-benchmark contribution rather than a methodological breakthrough; its lasting value depends on whether later releases reduce confounds and broaden multi-rater coverage so that capability-tracking is not partly artifactual.
major comments (3)
- [§2 Limitations; Tables 5–7; §4] §2 Limitations and Tables 5–7 / §4: The central claim that judge performance tracks general capability (and that thinking modes do not help) is load-bearing, yet the paper itself states that model-written critiques are disproportionately weak, long author-written critiques almost always strong, and “it’s not that difficult to achieve strong performance… by picking up on these distributional differences,” with source often obvious. Table 5 aggregates 255 positions / 856 pairs scored almost entirely against one primary rater (EC: 946/1458). Without a same-source or style-controlled split (or an ablation that holds length/provenance fixed), ordinal wins may partly reflect surface-cue detection rather than conceptual judgment. A controlled re-evaluation or stratified reporting is needed before the capability-alignment and “thinking doesn’t transfer” claims can be treated as clean.
- [§C; Table 1] §C and dataset construction: Ceiling estimates rest on human–human custom losses against EC first ratings (~0.14–0.16) and a 52-critique multi-rater test set. Most of the leaderboard signal still comes from a single primary rater; double-rated coverage is limited (CO 211, AK 130, CN 103), and active-learning selection used LLM judges. For a dataset whose value is as expert ground truth, the manuscript should either expand multi-rater agreement on the full evaluation slice or report model rankings restricted to multiply-rated items, and clarify how post-discussion revisions and LLM-assisted double-checks affect the labels used in Tables 5–7.
- [§3; §B; §2] §3 and §B scoring functions: Comparison-based loss is limited by “relatively few positions with a good spread of good and bad critiques” and many single-critique or all-bad positions (§2). Custom loss weights (4 for strength×centrality, 2 for correctness/clarity, 1 for dead weight/single issue) and the clarity<0.5 gate are free parameters. Sensitivity of the reported model orderings to these choices and to the subset of multi-critique positions should be shown; otherwise it is unclear how stable the Tables 5–7 rankings are under reasonable metric variants the paper itself contemplates.
minor comments (5)
- [Table 4; §4] Table 4 / §4 “Example model failure”: The prose says the table shows Opus 4 on Critique 1 of Table 2, but the displayed trace is not clearly labeled with model identity in the table caption; make the model and prompt condition explicit.
- [§G.1] Prompt in §G.1 still uses outdated “argument” terminology for position texts; align with the body’s “position” language to avoid confusion for replicators.
- [Tables 5–7] Tables 5–7: State n (positions/pairs/dialogues) and the exact human reference (EC first vs revised) in every caption; CI construction is mentioned only lightly.
- [§5] Related work (§5): DebateBench / VivesDebate / IBM Project Debater / LSAT links are appropriate; a short explicit contrast table (length, expert vs crowd, multi-axis vs unidimensional, conceptual vs empirical domain) would help readers place the contribution.
- [§4; §B.2] Minor typos and wording: e.g., “have have worse judgment” (§4), “a omparison-based” (§B.2), “throws away’ a lot of data a lot of data” (§B.2).
Circularity Check
Empirical dataset/benchmark paper: model scores are compared to external human ratings; no derivation folds fitted inputs or self-cited uniqueness into a forced ‘prediction.’
full rationale
The paper’s load-bearing chain is (i) collect position texts and critiques from mixed sources, (ii) obtain multi-axis expert human ratings under a written rubric, (iii) define two losses that compare model ratings to those human labels (weighted pairwise ranking error within positions; custom multi-axis absolute loss), (iv) report that model losses track general capability ladders and that vendor ‘thinking’ modes do not systematically help. Human overall / strength×centrality scores are external labels, not quantities defined from the models being ranked. There is no equation in which a fitted parameter is renamed a prediction, no uniqueness theorem imported from the authors to forbid alternatives, and no ansatz smuggled in via self-citation. Minor self-reference exists in construction—authors rate critiques they or models wrote; LLM judges helped select hard critiques (active-learning-like filtering) and were used when double-checking large rater disagreements (§2; §C)—but those steps affect which items enter the set or how humans revise, not the algebraic identity of the reported leaderboard losses. Capability-tracking could still be partly artifactual due to style/source confounds the paper itself flags; that is a validity/confound concern, not circularity by construction. Honest finding: no significant circularity.
Axiom & Free-Parameter Ledger
free parameters (2)
- custom_loss_axis_weights =
4 / 2 / 1 weights; clarity threshold 0.5
- human_rating_scales_0_to_1
axioms (5)
- domain assumption Contextualized arguments on conceptual questions can be evaluated more reliably and with less subjectivity than bottom-line conclusions on those questions.
- domain assumption Evaluating arguments is useful for making progress on bottom-line conceptual conclusions.
- domain assumption Expert multi-axis ratings (especially overall and strength×centrality) are sufficiently consistent to serve as supervision/benchmark targets for model comparison.
- ad hoc to paper For ambiguous strength vs centrality allocation, only the product strength×centrality is a stable target; isolated axis errors on those two fields are not meaningful.
- standard math Standard definitions of ranking losses and absolute-error aggregates (Kendall-style pairwise error, weighted variants) are appropriate external metrics.
invented entities (3)
-
Conceptual questions (as a task class)
no independent evidence
-
Seven-axis critique rubric (centrality, strength, correctness, clarity, dead weight, single issue, overall)
no independent evidence
-
Weighted pairwise ranking error rate and custom weighted loss
no independent evidence
read the original abstract
Large language models have improved rapidly on tasks with verifiable answers, such as mathematics and programming. Much less is known about their ability to reason about what we call conceptual questions: questions for which no ground truth is realistically accessible and no widely accepted resolution methodology exists, but on which progress can still be made by debating arguments. Most philosophical questions are of this kind, as are central components of questions in AI safety, decision theory, and social choice. Our approach is based on the view that while bottom-line conclusions on such questions are hard to evaluate, individual contextualized arguments can be evaluated far more reliably. We therefore introduce a dataset of 951 argumentative critiques of 442 position texts, spanning topics from AI safety and decision theory to ethics and politics, with 1,458 ratings by six expert raters along dimensions including centrality, strength, correctness, and clarity. We propose two scoring functions and benchmark a range of models. Performance tracks general capability rankings.
Reference graph
Works this paper leans on
-
[1]
Utilitarianism, decision theory and eternity
Arntzenius, Frank (2014). “Utilitarianism, decision theory and eternity”. In:Philosophical Perspec- tives28, pp. 31–58. Askell, Amanda (2018). “Pareto Principles in Infinite Ethics”. PhD thesis. New York University. Barkhordar, Ehsan et al. (2024). “Why the Unexpected? Dissecting the Political and Economic Bias in Persian Small and Large Language Models”....
Pith/arXiv arXiv 2014
-
[8]
broader context
If the model rates the two critiques as equally good, i.e., ˆra,i ovr = ˆra,j ovr, then the loss is 1/2. We then average this loss over pairs of critiques of the same position, and then we average these averages over positions. Call the resulting number thepairwise ranking error rate. This is essentially equivalent to what is sometimes called the Kendall ...
1958
-
[37]
AI research considerations for human existential safety (ARCHES)
13, pp. 15359–15367. Critch, Andrew and David Krueger (2020). “AI research considerations for human existential safety (ARCHES)”. In:arXiv preprint arXiv:2006.04948. Dafoe, Allan et al. (2020). “Open problems in cooperative AI”. In:arXiv preprint arXiv:2012.08630. Durmus, Esin et al. (Apr. 9, 2024).Measuring the Persuasiveness of Language Models.url:https...
Pith/arXiv arXiv 2020
-
[39]
26, pp. 27410–27418. Lim, Gionnieve and Simon T Perrault (2024). “Evaluation of an llm in identifying logical fallacies: A call for rigor when adopting llms in hci research”. In:Companion Publication of the 2024 Conference on Computer-Supported Cooperative Work and Social Computing, pp. 303–308. McAleese, Nat et al. (2024). “Llm critics help catch llm bug...
Pith/arXiv arXiv 2024
-
[1093]
A large-scale dataset for argument quality ranking: Construction and analysis
url:https://aclanthology.org/P19-1093/. Gretz, Shai et al. (2020). “A large-scale dataset for argument quality ranking: Construction and analysis”. In:Proceedings of the AAAI Conference on Artificial Intelligence. Vol
2020
-
[2022]
Debating with More Persuasive LLMs Leads to More Truthful Answers
Ed. by Yoav Goldberg, Zornitsa Kozareva, and Yue Zhang. Abu Dhabi, United Arab Emirates: Association for Computational Linguistics, pp. 7180– 7198.doi:10.18653/v1/2022.findings-emnlp.532.url:https://aclanthology.org/2022. findings-emnlp.532/. Khan, Akbir et al. (2024). “Debating with More Persuasive LLMs Leads to More Truthful Answers”. In:International C...
Pith/arXiv arXiv 2022
-
[7160]
Whose opinions do language models reflect?
Santurkar, Shibani et al. (2023). “Whose opinions do language models reflect?” In:International Conference on Machine Learning. PMLR, pp. 29971–30004. Saunders, William et al. (2022). “Self-critiquing models for assisting human evaluators”. In:arXiv preprint arXiv:2206.05802. Schoenegger, Philipp et al. (2025). “Large Language Models Are More Persuasive T...
Pith/arXiv arXiv 2023
-
[7813]
Deepseek-R1: Incentivizing reasoning capability in LLMs via reinforcement learning
Guo, Daya et al. (2025). “Deepseek-R1: Incentivizing reasoning capability in LLMs via reinforcement learning”. In:arXiv preprint arXiv:2501.12948. Habernal, Ivan and Iryna Gurevych (Aug. 2016). “Which argument is more convincing? Analyzing and predicting convincingness of Web arguments using bidirectional LSTM”. In:Proceedings of the 54th Annual Meeting o...
Pith/arXiv arXiv 2025
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.