REVIEW 2 major objections 5 minor 24 references
Rules or Character? Scaling Laws for AI Safety Design
T0 review · 2 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read Character fragility, not deployment scale, sets the best AI safety mix.
desk verdict Useful stylized model with a clear priority result, but the headline 'structural' scale-monotonicity is actually a numerical artifact of an unvaried fragility-leakage channel. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the resource-allocation coefficient $\alpha \in [0,1]$ representing the fraction of safety resources devoted to character shaping as opposed to rule enforcement, together with the expected-harm objective $E_{\mathrm{harm}}(\alpha, T, M) = T\big[(1-q(M))L_{\mathrm{normal}}(\alpha, M) + q(M)L_{\mathrm{CMF}}(\alpha)\big]$. The mechanism that carries the argument is asymmetric scale dependence: deployment scale $T$ enters only through edge-case pressure $M=\rho_{\mathrm{edge}} T$, which degrades filter quality $\varepsilon(\alpha, M)$ and raises common-mode failure probability $q(M)$, while the character fragility rate $p_{\mathrm{frag}}(\alpha) = p^{(0)}_{\mathrm{frag}} \alpha^n$ and the shaped harm $g_\alpha$ stay $T$-independent. This asymmetry makes larger scale penalize rule-heavy designs through two channels and favors character shaping, while $p^{(0)}_{\mathrm{frag}}$ imposes a double penalty on high-$\alpha$ designs, which is why it dominates the optimum.
What would settle it
A measurement study that estimates $p_{\mathrm{frag}}$ for a safety-trained model on out-of-distribution inputs at two or more deployment scales (or population diversities) would settle the structural claim: if $p_{\mathrm{frag}}$ rises with scale, the model's monotonicity result is an artifact of its $T$-independent fragility assumption rather than a robust scaling law.
Extended reading notes
Core claim
The paper's central claim is that the optimal character weight $\alpha^*$ is almost never pure character shaping, is either interior or at the rules-only boundary, and weakly increases with deployment scale, but the single most influential determinant is the baseline character fragility rate $p^{(0)}_{\mathrm{frag}}$. Across three scenario parameterizations (optimistic, moderate, pessimistic), $\alpha^*$ shifts by at most $+0.21$ as $T$ goes from $10^2$ to $10^8$, whereas sweeping $p^{(0)}_{\mathrm{frag}}$ over its plausible range shifts $\alpha^*$ by $0.50$, from about $0.70$ to $0.20$. The paper also claims that CVaR-based and expected-harm-based optima converge at large $T$ because the Pareto context multiplier is $\alpha$-independent, so tail heaviness rescales harm but does not reorder designs.
Load-bearing premise
The model's results depend on the premise that deploying at larger scale weakens filters and raises common-mode failure risk but never increases the per-interaction rate at which trained character shaping fails; if wider deployment itself made shaped behavior more fragile, the predicted shift toward character shaping would no longer be guaranteed.
Editorial extensions
If this is right
- If character fragility is the dominant lever, then measuring fragility under distributional shift becomes a prerequisite for any principled safety architecture decision.
- Reducing $p_{\mathrm{frag}}$ from 10% to 1% shifts the optimal design by roughly $+0.30$, a larger effect than improving filter quality, which shifts it by only $+0.07$.
- In the low-fragility regime (below roughly 5%), the optimal mix is essentially flat across six orders of magnitude of deployment scale, so the scaling question becomes moot.
- The optimal policy is robust to the choice of risk criterion: expected-harm and CVaR optima converge for large $T$, and the optimum does not change with the Pareto tail exponent.
- Stronger character-shaping capability (larger $\Delta\mu$) lowers $\alpha^*$ because of diminishing returns, meaning a moderate character allocation plus continued filter investment can beat aggressive character reliance.
Reading between the lines
- Beyond the paper, the results suggest a concrete measurement agenda: benchmark suites that estimate $p^{(0)}_{\mathrm{frag}}$ on held-out distributional shifts could be used directly as input to architecture choice, even before deployment-scale forecasts are made.
- Beyond the paper, the model predicts that if deployment expansion itself raises fragility (so $p_{\mathrm{frag}}$ depends on $T$), the monotonic scaling law $\Delta\alpha^* \geq 0$ could reverse; an empirical comparison of fragility on original versus novel user populations would adjudicate the model's core structural assumption.
- Beyond the paper, the $\alpha$-separable Pareto structure implies that interventions capping worst-case context damage (for example, domain restrictions) would reduce catastrophic risk without changing the optimal mix, a corollary the paper leaves implicit.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript introduces a stylized comparative-statics model of AI safety design as a resource allocation α between character shaping and rule enforcement. It incorporates scale-dependent filter degradation, common-mode failure, and character fragility, derives closed-form expected harm under a Gaussian action model and a multiplicative Pareto damage model, and supplements this with CVaR estimates from count-level Monte Carlo simulation. The principal findings are that the optimal α is interior or at the rules-only endpoint, is weakly non-decreasing in deployment scale T across the explored scenarios, and is dominated by the baseline character fragility rate p_frag^0, with a claimed robustness of the optimal design to the choice of risk criterion and to the Pareto tail exponent.
Significance. If the results hold, the paper provides a useful formal vocabulary for reasoning about safety-design tradeoffs and a concrete research-prioritization message: measure and reduce character fragility. The derivations from Eqs. (1)-(14) are transparent, the Monte Carlo procedure is sensible, and the sensitivity and phase-diagram analyses are welcome additions. The paper is also unusually honest about several modeling simplifications, including the T-independence of p_frag and the lack of empirical calibration for key parameters. The main weaknesses are that the headline 'structural' claims are not fully supported by the model's own equations, and one robustness result is asserted without adequate theoretical justification.
major comments (2)
- [Discussion (The Scale Effect Is Real but Regime-Dependent); Eqs. (7), (12)] The assertion that 'T has no channel through which it degrades character shaping' is contradicted by the model's own fragility-leakage term. In Eq. (12), L_normal contains p_frag(α)·ε_frag·g_frag, and ε_frag = min(factor×ε(α,M), 1.0) with ε(α,M) increasing in M through Eq. (7). Since p_frag(α) is increasing in α, this term grows more strongly with M for high-α designs, a scale-driven channel that penalizes character shaping and can counteract the filter-degradation and CMF channels that favor high α. The claimed near-tautological monotonicity Δα* ≥ 0 is therefore not a structural consequence of the model; it is, at best, a numerical outcome of the explored grid. The phase diagrams in Table 2 vary only three parameter pairs, and none varies the ε_frag factor, so the reported zero-negative-cells result does not cover this channel. The authors should either soften the structural claim and present the monotonicity as a conditional numerical result, or extend the phase diagrams to include the ε_frag factor and demonstrate the sign of Δα* over that grid.
- [Tail Risk (CVaR), Table 4; Discussion (Robustness to Risk Criterion and Tail Severity)] The claimed invariance of the CVaR-optimal α* to the Pareto tail exponent α_PL is not explained by α-separability of the context multiplier. For sums of independent products S_i(α)·X_i, the tail of the total harm has index α_PL and a coefficient proportional to E[S_i(α)^{α_PL}]; the α that minimizes this tail quantity generally depends on α_PL. Multiplication by an independent Pareto random variable preserves the ordering of expected harm but not, in general, the ordering under CVaR. The numerical finding in Table 4 (α* = 0.50 for all α_PL) and the accompanying Discussion text require a proper asymptotic justification or a demonstration that any dependence is smaller than the grid resolution; as written, the theoretical explanation is insufficient.
minor comments (5)
- [Character Fragility (paragraph defining ε_frag)] The notation ε_frag is introduced in text and used in Eq. (12); please make the M-dependence explicit (e.g., ε_frag(α,M)) so that the scale dependence of the fragility branch is visible in the equations rather than being easy to overlook.
- [Simulation Results, Table 2] The phase-diagram summary would be more informative if it reported the range of Δα* for each grid (e.g., the Δμ×ε_ceiling grid has a much narrower range than Δμ×p_frag^0) and if it included a grid that varies the ε_frag factor, which is the channel most relevant to the monotonicity claim.
- [Proposition 1 (Comparative Statics)] The proof sketch considers only the filter term in L_normal and omits the fragility-leakage and CMF terms; since the proposition is used to draw conclusions about filter-technology improvements, the proof should either include these terms or be explicitly labelled as a numerical tendency verified in Figure 4.
- [Simulation Results, Figure 5] The CVaR-based α* at small T is reported with Monte Carlo noise of ±0.10; reporting bootstrap confidence intervals for α*_CVaR (as done for the CVaR magnitudes in Table 4) would strengthen the convergence claim.
- [Abstract and Conclusion] The phrase 'interior or at the rules-only boundary' is awkward; 'interior or at the rules-only endpoint' is clearer and avoids suggesting that α*=1 is a boundary of the same kind.
Circularity Check
No significant circularity: the headline results are computed consequences of the paper's stated equations, not fitted inputs or load-bearing self-citations.
full rationale
The paper's derivation chain is self-contained and transparent. Expected harm in Equations (10)-(14) follows from the stated Gaussian action model, the Pareto context multiplier, and the explicit functional forms for filter degradation, CMF probability, and character fragility. The optimal character weight alpha*(T), the sensitivity rankings, and the CVaR comparisons are analytic or Monte Carlo consequences of those equations rather than fitted parameters renamed as predictions. The scale-monotonicity result Delta-alpha* >= 0 is not an input assumption: it is a computed outcome over the explored grids, and the paper explicitly labels it a structural property of the current model while noting that a model with p_frag(alpha,T) could reverse it. Likewise, the CVaR invariance to the tail exponent is derived from the alpha-separability of the Pareto multiplier, and the paper states the conditional scope of that claim. There are no load-bearing self-citations and no uniqueness theorems imported from the authors' prior work; the cited external results (e.g., Constitutional Classifiers, heavy-tailed damage data) are used only as parameter anchors and motivation. The concern that epsilon_frag creates a scale-dependent penalty on character shaping is a substantive modeling or correctness question, but it is not circularity: the paper's equations are fully specified, and the monotonicity claim does not assume the conclusion it reports.
Assumptions & free parameters
free parameters (20)
- p_frag^(0) baseline character fragility rate =
0.01 (Optimistic), 0.03 (Moderate), 0.10 (Pessimistic)
- Delta_mu safety shift =
1.5 / 1.0 / 0.5
- r_sigma variance reduction =
0.6 / 0.7 / 0.9
- epsilon_min filter quality ceiling =
0.02 / 0.03 / 0.05
- epsilon_max,base =
0.10 / 0.15 / 0.25
- epsilon_ceiling =
0.25 / 0.30 / 0.50
- rho_edge =
0.01 / 0.05 / 0.10
- beta_d =
0.5 / 1.0 / 2.0
- d0 =
0.05 / 0.10 / 0.30
- beta_q =
0.3 / 0.5 / 1.0
- e0 =
0.2 / 0.3 / 0.5
- mu_frag =
0.0 / -0.5 / -1.0
- sigma_frag =
1.0 / 1.2 / 1.5
- epsilon_frag factor =
1.0 / 1.0 / 2.0
- alpha_PL tail exponent =
3.0 / 2.5 / 2.0
- n fragility exponent =
2 (default; 0.5 to 4 in sensitivity)
- k filter degradation exponent =
1.0
- M_ref reference scale =
1e6
- tau safety threshold =
-2.0
- mu0, sigma0 baseline distribution =
0.0, 1.0
assumptions (8)
- domain assumption Actions before intervention are drawn from P0=N(mu0, sigma0^2), and shaped behavior is N(mu(alpha), sigma(alpha)^2).
- domain assumption Harm is h(a)=(tau-a)_+ * X with X~Pareto(1, alpha_PL), independent of a and alpha.
- ad hoc to paper p_frag(alpha)=p_frag^(0)*alpha^n with n=2 and p_frag independent of T.
- ad hoc to paper T enters only through M=rho_edge*T, degrading filters and raising CMF probability, with no effect on character shaping.
- domain assumption During a common-mode failure all filters are disabled, while character shaping continues to reduce harm.
- ad hoc to paper Filter leakage epsilon(alpha,M) follows the exponential saturation form in Eq. (7), and CMF probability q(M) follows Eq. (8).
- standard math The Pareto mean E[X]=alpha_PL/(alpha_PL-1) for alpha_PL>1.
- standard math CVaR is estimated by count-level Monte Carlo with 10,000-50,000 replications and bootstrap confidence intervals.
Cite this review
Pith. "Pith review of Rules or Character? Scaling Laws for AI Safety Design." pith.science (2026). https://pith.science/paper/TFIQGI6N
@misc{pith2026260813345,
author = {Pith},
title = {Pith review of: Rules or Character? Scaling Laws for AI Safety Design},
year = {2026},
howpublished = {\url{https://pith.science/paper/TFIQGI6N}},
note = {Machine review of arXiv:2608.13345}
}
read the original abstract
Artificial Intelligence (AI) safety systems combine character shaping (e.g., Reinforcement Learning from Human Feedback [RLHF], Constitutional AI), which modifies behavioral distributions at training time, with rule enforcement (e.g., output filters, safety classifiers), which blocks harmful outputs at inference time, yet little formal analysis exists on how their optimal balance should change as deployment scales increase. We introduce a stylized comparative-statics model that parameterizes safety design as a resource allocation alpha in [0,1] between these two approaches, incorporating scale-dependent filter degradation, common-mode failures, and character fragility -- the risk that shaped behavior degrades or collapses under novel conditions. Under a multiplicative Pareto damage model, we derive closed-form expected harm and supplement it with tail-risk (CVaR) analysis via Monte Carlo simulation. Across three scenarios (optimistic, moderate, pessimistic), the optimal alpha* is interior or at the rules-only boundary and shifts weakly toward character shaping as deployment scale T grows, from negligible (Delta alpha* = +0.01) to pronounced (Delta alpha* = +0.21) depending on scenario. The dominant parameter is the baseline character fragility rate p^(0)_frag, which shifts alpha* by 0.50 across its range -- far exceeding the effect of tail severity, filter quality, or common-mode failure probability. CVaR and expected-harm optima converge at large T. These results suggest that safety architecture decisions depend less on deployment scale per se than on the reliability of character shaping under distributional shift.
Figures
Figures from the paper (2 more)
Reference graph
Works this paper leans on
-
[1]
Amodei, D.; Olah, C.; Steinhardt, J.; Christiano, P.; Schulman, J.; and Man \'e , D. 2016. Concrete Problems in AI Safety. arXiv preprint arXiv:1606.06565
arXiv 2016
-
[2]
Aristotle. 1999. Nicomachean Ethics. Indianapolis: Hackett Publishing Company, 2nd edition. Translated by T. Irwin
work page 1999
-
[3]
Askell, A. 2018. Pareto Principles in Infinite Ethics. Ph.D. thesis, New York University
work page 2018
-
[4]
Aven, T. 2016. Risk Assessment and Risk Management: Review of Recent Advances on Their Foundation. European Journal of Operational Research, 253(1): 1--13
work page 2016
-
[5]
Bai, Y.; Kadavath, S.; Kundu, S.; Askell, A.; Kernion, J.; Jones, A.; Chen, A.; Goldie, A.; Mirhoseini, A.; McKinnon, C.; et al. 2022. Constitutional AI : Harmlessness from AI feedback. arXiv preprint arXiv:2212.08073
arXiv 2022
-
[6]
Boehm, B. W. 1981. Software Engineering Economics. Prentice-Hall
work page 1981
-
[7]
F.; Leike, J.; Brown, T.; Martic, M.; Legg, S.; and Amodei, D
Christiano, P. F.; Leike, J.; Brown, T.; Martic, M.; Legg, S.; and Amodei, D. 2017. Deep reinforcement learning from human preferences. In Advances in Neural Information Processing Systems, volume 30
work page 2017
-
[8]
Edwards, B.; Hofmeyr, S.; and Forrest, S. 2016. Hype and Heavy Tails: A Closer Look at Data Breaches. Journal of Cybersecurity, 2(1): 3--14
work page 2016
Show all 24 references
-
[9]
I.; Lukošiūtė, K.; Chen, A.; Goldie, A.; Mirhoseini, A.; Olsson, C.; Hernandez, D.; et al
Ganguli, D.; Askell, A.; Schiefer, N.; Liao, T. I.; Lukošiūtė, K.; Chen, A.; Goldie, A.; Mirhoseini, A.; Olsson, C.; Hernandez, D.; et al. 2023. The capacity for moral self-correction in large language models. arXiv preprint arXiv:2302.07459
2023 arXiv
-
[10]
Hendrycks, D.; and Gimpel, K. 2017. A Baseline for Detecting Misclassified and Out-of-Distribution Examples in Neural Networks. In Proceedings of ICLR
2017
-
[11]
M.; Maxwell, T.; Cheng, N.; et al
Hubinger, E.; Denison, C.; Mu, J.; Lambert, M.; Tong, M.; MacDiarmid, M.; Lanham, T.; Ziegler, D. M.; Maxwell, T.; Cheng, N.; et al. 2024. Sleeper agents: Training deceptive LLMs that persist through safety training. arXiv preprint arXiv:2401.05566
2024 arXiv
-
[12]
Hursthouse, R. 1999. On Virtue Ethics. Oxford: Oxford University Press
1999
-
[13]
B.; Chess, B.; Child, R.; Gray, S.; Radford, A.; Wu, J.; and Amodei, D
Kaplan, J.; McCandlish, S.; Henighan, T.; Brown, T. B.; Chess, B.; Child, R.; Gray, S.; Radford, A.; Wu, J.; and Amodei, D. 2020. Scaling Laws for Neural Language Models. arXiv preprint arXiv:2001.08361
2020 arXiv
-
[14]
Kaplan, S.; and Garrick, B. J. 1981. On the Quantitative Definition of Risk. Risk Analysis, 1(1): 11--27
1981
-
[15]
Leveson, N. G. 2011. Engineering a Safer World: Systems Thinking Applied to Safety. MIT Press
2011
-
[16]
Maillart, T.; and Sornette, D. 2010. Heavy-Tailed Distribution of Cyber-Risks. The European Physical Journal B, 75(3): 357--364
2010
-
[17]
Noller, J. 2026. Artificial moral characters: constitutional AI and the challenge of alignment. AI and Ethics, 6(2)
2026
-
[18]
L.; Mishkin, P.; Zhang, C.; Agarwal, S.; Slama, K.; Ray, A.; et al
Ouyang, L.; Wu, J.; Jiang, X.; Almeida, D.; Wainwright, C. L.; Mishkin, P.; Zhang, C.; Agarwal, S.; Slama, K.; Ray, A.; et al. 2022. Training language models to follow instructions with human feedback. In Advances in Neural Information Processing Systems, volume 35, 27730--27744
2022
-
[19]
Perrow, C. 1984. Normal Accidents: Living with High-Risk Technologies. Basic Books
1984
-
[20]
Quiñonero-Candela, J.; Sugiyama, M.; Schwaighofer, A.; and Lawrence, N. D., eds. 2009. Dataset Shift in Machine Learning. MIT Press
2009
-
[21]
Rasmussen, J. 1997. Risk Management in a Dynamic Society: A Modelling Problem. Safety Science, 27(2--3): 183--213
1997
-
[22]
Reason, J. 1990. Human Error. Cambridge University Press
1990
-
[23]
T.; and Uryasev, S
Rockafellar, R. T.; and Uryasev, S. 2000. Optimization of Conditional Value-at-Risk. Journal of Risk, 2: 21--42
2000
-
[24]
Sharma, M.; Tong, M.; Mu, J.; Wei, J.; Kruthoff, J.; Goodfriend, S.; Ong, E.; Peng, A.; Agarwal, R.; Anil, C.; et al. 2025. Constitutional Classifiers: Defending against universal jailbreaks across thousands of hours of red teaming. arXiv preprint arXiv:2501.18837
2025 arXiv
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.