Pith. sign in

REVIEW 2 major objections 5 minor 24 references

Rules or Character? Scaling Laws for AI Safety Design

T0 review · 2 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read Character fragility, not deployment scale, sets the best AI safety mix.

desk verdict Useful stylized model with a clear priority result, but the headline 'structural' scale-monotonicity is actually a numerical artifact of an unvaried fragility-leakage channel. read the letter →

arxiv 2608.13345 v1 pith:TFIQGI6N submitted 2026-08-13 cs.AI

classification cs.AI
keywords AIsafetycharactershapingruleenforcementscalinglawfragilitycommon-modefailureCVaRParetodamage
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to establish that the optimal balance between training-time character shaping and inference-time rule enforcement is governed mainly by a single quantity: the rate at which shaped safe behavior fails under novel conditions. It models safety design as a resource split between the two approaches, derives closed-form expected harm, and finds that the best split moves only weakly toward character shaping as deployment scale grows, while the baseline fragility rate moves it by 0.50 across its range. The authors argue that if this is right, safety research should prioritize measuring and reducing character fragility over tuning for scale.

What carries the argument

The central object is the resource-allocation coefficient $\alpha \in [0,1]$ representing the fraction of safety resources devoted to character shaping as opposed to rule enforcement, together with the expected-harm objective $E_{\mathrm{harm}}(\alpha, T, M) = T\big[(1-q(M))L_{\mathrm{normal}}(\alpha, M) + q(M)L_{\mathrm{CMF}}(\alpha)\big]$. The mechanism that carries the argument is asymmetric scale dependence: deployment scale $T$ enters only through edge-case pressure $M=\rho_{\mathrm{edge}} T$, which degrades filter quality $\varepsilon(\alpha, M)$ and raises common-mode failure probability $q(M)$, while the character fragility rate $p_{\mathrm{frag}}(\alpha) = p^{(0)}_{\mathrm{frag}} \alpha^n$ and the shaped harm $g_\alpha$ stay $T$-independent. This asymmetry makes larger scale penalize rule-heavy designs through two channels and favors character shaping, while $p^{(0)}_{\mathrm{frag}}$ imposes a double penalty on high-$\alpha$ designs, which is why it dominates the optimum.

What would settle it

A measurement study that estimates $p_{\mathrm{frag}}$ for a safety-trained model on out-of-distribution inputs at two or more deployment scales (or population diversities) would settle the structural claim: if $p_{\mathrm{frag}}$ rises with scale, the model's monotonicity result is an artifact of its $T$-independent fragility assumption rather than a robust scaling law.

Watch

Extended reading notes

Core claim

The paper's central claim is that the optimal character weight $\alpha^*$ is almost never pure character shaping, is either interior or at the rules-only boundary, and weakly increases with deployment scale, but the single most influential determinant is the baseline character fragility rate $p^{(0)}_{\mathrm{frag}}$. Across three scenario parameterizations (optimistic, moderate, pessimistic), $\alpha^*$ shifts by at most $+0.21$ as $T$ goes from $10^2$ to $10^8$, whereas sweeping $p^{(0)}_{\mathrm{frag}}$ over its plausible range shifts $\alpha^*$ by $0.50$, from about $0.70$ to $0.20$. The paper also claims that CVaR-based and expected-harm-based optima converge at large $T$ because the Pareto context multiplier is $\alpha$-independent, so tail heaviness rescales harm but does not reorder designs.

Load-bearing premise

The model's results depend on the premise that deploying at larger scale weakens filters and raises common-mode failure risk but never increases the per-interaction rate at which trained character shaping fails; if wider deployment itself made shaped behavior more fragile, the predicted shift toward character shaping would no longer be guaranteed.

Editorial extensions

If this is right

  • If character fragility is the dominant lever, then measuring fragility under distributional shift becomes a prerequisite for any principled safety architecture decision.
  • Reducing $p_{\mathrm{frag}}$ from 10% to 1% shifts the optimal design by roughly $+0.30$, a larger effect than improving filter quality, which shifts it by only $+0.07$.
  • In the low-fragility regime (below roughly 5%), the optimal mix is essentially flat across six orders of magnitude of deployment scale, so the scaling question becomes moot.
  • The optimal policy is robust to the choice of risk criterion: expected-harm and CVaR optima converge for large $T$, and the optimum does not change with the Pareto tail exponent.
  • Stronger character-shaping capability (larger $\Delta\mu$) lowers $\alpha^*$ because of diminishing returns, meaning a moderate character allocation plus continued filter investment can beat aggressive character reliance.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Beyond the paper, the results suggest a concrete measurement agenda: benchmark suites that estimate $p^{(0)}_{\mathrm{frag}}$ on held-out distributional shifts could be used directly as input to architecture choice, even before deployment-scale forecasts are made.
  • Beyond the paper, the model predicts that if deployment expansion itself raises fragility (so $p_{\mathrm{frag}}$ depends on $T$), the monotonic scaling law $\Delta\alpha^* \geq 0$ could reverse; an empirical comparison of fragility on original versus novel user populations would adjudicate the model's core structural assumption.
  • Beyond the paper, the $\alpha$-separable Pareto structure implies that interventions capping worst-case context damage (for example, domain restrictions) would reduce catastrophic risk without changing the optimal mix, a corollary the paper leaves implicit.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. The manuscript introduces a stylized comparative-statics model of AI safety design as a resource allocation α between character shaping and rule enforcement. It incorporates scale-dependent filter degradation, common-mode failure, and character fragility, derives closed-form expected harm under a Gaussian action model and a multiplicative Pareto damage model, and supplements this with CVaR estimates from count-level Monte Carlo simulation. The principal findings are that the optimal α is interior or at the rules-only endpoint, is weakly non-decreasing in deployment scale T across the explored scenarios, and is dominated by the baseline character fragility rate p_frag^0, with a claimed robustness of the optimal design to the choice of risk criterion and to the Pareto tail exponent.

Significance. If the results hold, the paper provides a useful formal vocabulary for reasoning about safety-design tradeoffs and a concrete research-prioritization message: measure and reduce character fragility. The derivations from Eqs. (1)-(14) are transparent, the Monte Carlo procedure is sensible, and the sensitivity and phase-diagram analyses are welcome additions. The paper is also unusually honest about several modeling simplifications, including the T-independence of p_frag and the lack of empirical calibration for key parameters. The main weaknesses are that the headline 'structural' claims are not fully supported by the model's own equations, and one robustness result is asserted without adequate theoretical justification.

major comments (2)
  1. [Discussion (The Scale Effect Is Real but Regime-Dependent); Eqs. (7), (12)] The assertion that 'T has no channel through which it degrades character shaping' is contradicted by the model's own fragility-leakage term. In Eq. (12), L_normal contains p_frag(α)·ε_frag·g_frag, and ε_frag = min(factor×ε(α,M), 1.0) with ε(α,M) increasing in M through Eq. (7). Since p_frag(α) is increasing in α, this term grows more strongly with M for high-α designs, a scale-driven channel that penalizes character shaping and can counteract the filter-degradation and CMF channels that favor high α. The claimed near-tautological monotonicity Δα* ≥ 0 is therefore not a structural consequence of the model; it is, at best, a numerical outcome of the explored grid. The phase diagrams in Table 2 vary only three parameter pairs, and none varies the ε_frag factor, so the reported zero-negative-cells result does not cover this channel. The authors should either soften the structural claim and present the monotonicity as a conditional numerical result, or extend the phase diagrams to include the ε_frag factor and demonstrate the sign of Δα* over that grid.
  2. [Tail Risk (CVaR), Table 4; Discussion (Robustness to Risk Criterion and Tail Severity)] The claimed invariance of the CVaR-optimal α* to the Pareto tail exponent α_PL is not explained by α-separability of the context multiplier. For sums of independent products S_i(α)·X_i, the tail of the total harm has index α_PL and a coefficient proportional to E[S_i(α)^{α_PL}]; the α that minimizes this tail quantity generally depends on α_PL. Multiplication by an independent Pareto random variable preserves the ordering of expected harm but not, in general, the ordering under CVaR. The numerical finding in Table 4 (α* = 0.50 for all α_PL) and the accompanying Discussion text require a proper asymptotic justification or a demonstration that any dependence is smaller than the grid resolution; as written, the theoretical explanation is insufficient.
minor comments (5)
  1. [Character Fragility (paragraph defining ε_frag)] The notation ε_frag is introduced in text and used in Eq. (12); please make the M-dependence explicit (e.g., ε_frag(α,M)) so that the scale dependence of the fragility branch is visible in the equations rather than being easy to overlook.
  2. [Simulation Results, Table 2] The phase-diagram summary would be more informative if it reported the range of Δα* for each grid (e.g., the Δμ×ε_ceiling grid has a much narrower range than Δμ×p_frag^0) and if it included a grid that varies the ε_frag factor, which is the channel most relevant to the monotonicity claim.
  3. [Proposition 1 (Comparative Statics)] The proof sketch considers only the filter term in L_normal and omits the fragility-leakage and CMF terms; since the proposition is used to draw conclusions about filter-technology improvements, the proof should either include these terms or be explicitly labelled as a numerical tendency verified in Figure 4.
  4. [Simulation Results, Figure 5] The CVaR-based α* at small T is reported with Monte Carlo noise of ±0.10; reporting bootstrap confidence intervals for α*_CVaR (as done for the CVaR magnitudes in Table 4) would strengthen the convergence claim.
  5. [Abstract and Conclusion] The phrase 'interior or at the rules-only boundary' is awkward; 'interior or at the rules-only endpoint' is clearer and avoids suggesting that α*=1 is a boundary of the same kind.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the headline results are computed consequences of the paper's stated equations, not fitted inputs or load-bearing self-citations.

full rationale

The paper's derivation chain is self-contained and transparent. Expected harm in Equations (10)-(14) follows from the stated Gaussian action model, the Pareto context multiplier, and the explicit functional forms for filter degradation, CMF probability, and character fragility. The optimal character weight alpha*(T), the sensitivity rankings, and the CVaR comparisons are analytic or Monte Carlo consequences of those equations rather than fitted parameters renamed as predictions. The scale-monotonicity result Delta-alpha* >= 0 is not an input assumption: it is a computed outcome over the explored grids, and the paper explicitly labels it a structural property of the current model while noting that a model with p_frag(alpha,T) could reverse it. Likewise, the CVaR invariance to the tail exponent is derived from the alpha-separability of the Pareto multiplier, and the paper states the conditional scope of that claim. There are no load-bearing self-citations and no uniqueness theorems imported from the authors' prior work; the cited external results (e.g., Constitutional Classifiers, heavy-tailed damage data) are used only as parameter anchors and motivation. The concern that epsilon_frag creates a scale-dependent penalty on character shaping is a substantive modeling or correctness question, but it is not circularity: the paper's equations are fully specified, and the monotonicity claim does not assume the conclusion it reports.

Assumptions & free parameters 20 free parameters · 8 assumptions · 0 invented entities

The model introduces no new physical entities. The key free parameters are hand-chosen scenario anchors, with p_frag^(0) singled out in the paper as the dominant lever. The structural results (monotonicity, tail-invariance) rest on the ad hoc T-asymmetry and alpha-separability assumptions listed above.

free parameters (20)
  • p_frag^(0) baseline character fragility rate = 0.01 (Optimistic), 0.03 (Moderate), 0.10 (Pessimistic)
    Central sensitivity parameter; no empirical calibration; chosen as scenario anchor.
  • Delta_mu safety shift = 1.5 / 1.0 / 0.5
    Determines how much training shifts the mean safety score; varied across scenarios.
  • r_sigma variance reduction = 0.6 / 0.7 / 0.9
    Controls variance shrinkage from character shaping.
  • epsilon_min filter quality ceiling = 0.02 / 0.03 / 0.05
    Anchored to Constitutional Classifiers 4.4% jailbreak rate; mapped to per-interaction leakage.
  • epsilon_max,base = 0.10 / 0.15 / 0.25
    Filter leakage when no resources are given to filters (alpha=1).
  • epsilon_ceiling = 0.25 / 0.30 / 0.50
    Upper leakage under maximal scale-driven degradation.
  • rho_edge = 0.01 / 0.05 / 0.10
    Fraction of interactions that are edge cases.
  • beta_d = 0.5 / 1.0 / 2.0
    Sensitivity of filter pattern discovery to deployment scale.
  • d0 = 0.05 / 0.10 / 0.30
    Diffusion fraction of discovered vulnerability patterns.
  • beta_q = 0.3 / 0.5 / 1.0
    Sensitivity of common-mode failure discovery to scale.
  • e0 = 0.2 / 0.3 / 0.5
    Probability that a discovered systemic vulnerability is actually triggered.
  • mu_frag = 0.0 / -0.5 / -1.0
    Mean of fragility distribution; negative values mean worse than baseline behavior.
  • sigma_frag = 1.0 / 1.2 / 1.5
    Spread of the fragility distribution.
  • epsilon_frag factor = 1.0 / 1.0 / 2.0
    Filter effectiveness during fragility events; scenario-dependent.
  • alpha_PL tail exponent = 3.0 / 2.5 / 2.0
    Pareto tail index; not directly measured for AI incidents, anchored to cybersecurity and software cost analogies.
  • n fragility exponent = 2 (default; 0.5 to 4 in sensitivity)
    Shape of fragility growth with alpha; quantitative results depend on this functional form.
  • k filter degradation exponent = 1.0
    Linear degradation of filter quality with alpha; not varied.
  • M_ref reference scale = 1e6
    Normalizes edge-case pressure; arbitrary reference count.
  • tau safety threshold = -2.0
    Threshold in the Gaussian action space; normalization choice.
  • mu0, sigma0 baseline distribution = 0.0, 1.0
    Location and scale normalization of the baseline action distribution.
assumptions (8)
  • domain assumption Actions before intervention are drawn from P0=N(mu0, sigma0^2), and shaped behavior is N(mu(alpha), sigma(alpha)^2).
    Only the first two moments are modeled; the Gaussian tail makes g_alpha tractable but discards heavy-tailed action behavior, as the Limitations state.
  • domain assumption Harm is h(a)=(tau-a)_+ * X with X~Pareto(1, alpha_PL), independent of a and alpha.
    This alpha-separability makes the argmin over alpha independent of alpha_PL and underlies the CVaR invariance result.
  • ad hoc to paper p_frag(alpha)=p_frag^(0)*alpha^n with n=2 and p_frag independent of T.
    The power law is chosen, not measured; the T-independence drives the monotonicity result.
  • ad hoc to paper T enters only through M=rho_edge*T, degrading filters and raising CMF probability, with no effect on character shaping.
    Recognized in the paper as making Delta-alpha*>=0 near-tautological; if falsified, the central scaling direction could reverse.
  • domain assumption During a common-mode failure all filters are disabled, while character shaping continues to reduce harm.
    This creates the inner wall that favors character shaping at large scale.
  • ad hoc to paper Filter leakage epsilon(alpha,M) follows the exponential saturation form in Eq. (7), and CMF probability q(M) follows Eq. (8).
    Specific functional forms are chosen; the qualitative results depend on their monotone structure.
  • standard math The Pareto mean E[X]=alpha_PL/(alpha_PL-1) for alpha_PL>1.
    Standard calculation used in Eq. (3).
  • standard math CVaR is estimated by count-level Monte Carlo with 10,000-50,000 replications and bootstrap confidence intervals.
    Standard simulation practice; not independently verified without code.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Rules or Character? Scaling Laws for AI Safety Design." pith.science (2026). https://pith.science/paper/TFIQGI6N

@misc{pith2026260813345,
  author       = {Pith},
  title        = {Pith review of: Rules or Character? Scaling Laws for AI Safety Design},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/TFIQGI6N}},
  note         = {Machine review of arXiv:2608.13345}
}
read the original abstract

Artificial Intelligence (AI) safety systems combine character shaping (e.g., Reinforcement Learning from Human Feedback [RLHF], Constitutional AI), which modifies behavioral distributions at training time, with rule enforcement (e.g., output filters, safety classifiers), which blocks harmful outputs at inference time, yet little formal analysis exists on how their optimal balance should change as deployment scales increase. We introduce a stylized comparative-statics model that parameterizes safety design as a resource allocation alpha in [0,1] between these two approaches, incorporating scale-dependent filter degradation, common-mode failures, and character fragility -- the risk that shaped behavior degrades or collapses under novel conditions. Under a multiplicative Pareto damage model, we derive closed-form expected harm and supplement it with tail-risk (CVaR) analysis via Monte Carlo simulation. Across three scenarios (optimistic, moderate, pessimistic), the optimal alpha* is interior or at the rules-only boundary and shifts weakly toward character shaping as deployment scale T grows, from negligible (Delta alpha* = +0.01) to pronounced (Delta alpha* = +0.21) depending on scenario. The dominant parameter is the baseline character fragility rate p^(0)_frag, which shifts alpha* by 0.50 across its range -- far exceeding the effect of tail severity, filter quality, or common-mode failure probability. CVaR and expected-harm optima converge at large T. These results suggest that safety architecture decisions depend less on deployment scale per se than on the reliability of character shaping under distributional shift.

Figures

Figures reproduced from arXiv: 2608.13345 by the authors.

Figure 1
Figure 1. presents the central result: α ∗ (T) across the three scenarios. In all cases, pure character (α = 1) is never opti￾mal; the optimum is either interior or, in the Pessimistic sce￾nario at small T, at the rules-only boundary (α ∗ = 0). The optimal character weight increases weakly with deployment scale T, but the magnitude of this shift varies substantially across scenarios( [PITH_FULL_IMAGE:figures/full_fig_p006_1.png] view at source ↗
Figure 2
Figure 2. Phase diagram: scale-induced shift in optimal design over the ∆µ × p (0) frag parameter space. Each cell shows ∆α ∗ = α ∗ (108 ) − α ∗ (102 ). Red indicates that larger scale favors more character-reliant design; blue would in￾dicate the opposite. Across all 400 cells, ∆α ∗ ≥ 0. Col￾ored circles mark the three scenarios. This monotonicity is a structural property of the model (see the Discussion). This result is str… view at source ↗
Figure 3
Figure 3. shows this dominant parameter in detail: α ∗ de￾creases monotonically from 0.70 (at p (0) frag = 0.005) to 0.20 (at p (0) frag = 0.40) — a swing of 0.50. This result carries a clear practical implication: the most consequential input to safety architecture design is the es￾timated reliability of character shaping under distributional Parameter Swept range ∆α ∗ at T = 106 p (0) frag 0.005 → 0.40 −0.50 ∆µ 0.2 → 2.0 −0… view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: Numerical verification of Proposition 1: im￾proving filter technology lowers α ∗ . α ∗ (T = 106 ) as a function of the filter quality ceiling εmin across all three sce￾narios. In every case, α ∗ increases monotonically with εmin, confirming ∂α∗/∂εmin > 0. As filter tec…
Figure 6
Figure 6. Figure 6: decomposes Eharm into its normal-operation and CMF components at T = 106 . At low α, the CMF com￾ponent (purple) is visible as a non-negligible share of to￾tal harm: because gα ≈ g0 when character shaping is minimal, the system remains exposed when filters are dis￾able…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

24 extracted references · 18 canonical work pages

  1. [1]

    Amodei, D.; Olah, C.; Steinhardt, J.; Christiano, P.; Schulman, J.; and Man \'e , D. 2016. Concrete Problems in AI Safety. arXiv preprint arXiv:1606.06565

  2. [2]

    Aristotle. 1999. Nicomachean Ethics. Indianapolis: Hackett Publishing Company, 2nd edition. Translated by T. Irwin

  3. [3]

    Askell, A. 2018. Pareto Principles in Infinite Ethics. Ph.D. thesis, New York University

  4. [4]

    Aven, T. 2016. Risk Assessment and Risk Management: Review of Recent Advances on Their Foundation. European Journal of Operational Research, 253(1): 1--13

  5. [5]

    Bai, Y.; Kadavath, S.; Kundu, S.; Askell, A.; Kernion, J.; Jones, A.; Chen, A.; Goldie, A.; Mirhoseini, A.; McKinnon, C.; et al. 2022. Constitutional AI : Harmlessness from AI feedback. arXiv preprint arXiv:2212.08073

  6. [6]

    Boehm, B. W. 1981. Software Engineering Economics. Prentice-Hall

  7. [7]

    F.; Leike, J.; Brown, T.; Martic, M.; Legg, S.; and Amodei, D

    Christiano, P. F.; Leike, J.; Brown, T.; Martic, M.; Legg, S.; and Amodei, D. 2017. Deep reinforcement learning from human preferences. In Advances in Neural Information Processing Systems, volume 30

  8. [8]

    Edwards, B.; Hofmeyr, S.; and Forrest, S. 2016. Hype and Heavy Tails: A Closer Look at Data Breaches. Journal of Cybersecurity, 2(1): 3--14

Show all 24 references
  1. [9]

    I.; Lukošiūtė, K.; Chen, A.; Goldie, A.; Mirhoseini, A.; Olsson, C.; Hernandez, D.; et al

    Ganguli, D.; Askell, A.; Schiefer, N.; Liao, T. I.; Lukošiūtė, K.; Chen, A.; Goldie, A.; Mirhoseini, A.; Olsson, C.; Hernandez, D.; et al. 2023. The capacity for moral self-correction in large language models. arXiv preprint arXiv:2302.07459

  2. [10]

    Hendrycks, D.; and Gimpel, K. 2017. A Baseline for Detecting Misclassified and Out-of-Distribution Examples in Neural Networks. In Proceedings of ICLR

  3. [11]

    M.; Maxwell, T.; Cheng, N.; et al

    Hubinger, E.; Denison, C.; Mu, J.; Lambert, M.; Tong, M.; MacDiarmid, M.; Lanham, T.; Ziegler, D. M.; Maxwell, T.; Cheng, N.; et al. 2024. Sleeper agents: Training deceptive LLMs that persist through safety training. arXiv preprint arXiv:2401.05566

  4. [12]

    Hursthouse, R. 1999. On Virtue Ethics. Oxford: Oxford University Press

  5. [13]

    B.; Chess, B.; Child, R.; Gray, S.; Radford, A.; Wu, J.; and Amodei, D

    Kaplan, J.; McCandlish, S.; Henighan, T.; Brown, T. B.; Chess, B.; Child, R.; Gray, S.; Radford, A.; Wu, J.; and Amodei, D. 2020. Scaling Laws for Neural Language Models. arXiv preprint arXiv:2001.08361

  6. [14]

    Kaplan, S.; and Garrick, B. J. 1981. On the Quantitative Definition of Risk. Risk Analysis, 1(1): 11--27

  7. [15]

    Leveson, N. G. 2011. Engineering a Safer World: Systems Thinking Applied to Safety. MIT Press

  8. [16]

    Maillart, T.; and Sornette, D. 2010. Heavy-Tailed Distribution of Cyber-Risks. The European Physical Journal B, 75(3): 357--364

  9. [17]

    Noller, J. 2026. Artificial moral characters: constitutional AI and the challenge of alignment. AI and Ethics, 6(2)

  10. [18]

    L.; Mishkin, P.; Zhang, C.; Agarwal, S.; Slama, K.; Ray, A.; et al

    Ouyang, L.; Wu, J.; Jiang, X.; Almeida, D.; Wainwright, C. L.; Mishkin, P.; Zhang, C.; Agarwal, S.; Slama, K.; Ray, A.; et al. 2022. Training language models to follow instructions with human feedback. In Advances in Neural Information Processing Systems, volume 35, 27730--27744

  11. [19]

    Perrow, C. 1984. Normal Accidents: Living with High-Risk Technologies. Basic Books

  12. [20]

    Quiñonero-Candela, J.; Sugiyama, M.; Schwaighofer, A.; and Lawrence, N. D., eds. 2009. Dataset Shift in Machine Learning. MIT Press

  13. [21]

    Rasmussen, J. 1997. Risk Management in a Dynamic Society: A Modelling Problem. Safety Science, 27(2--3): 183--213

  14. [22]

    Reason, J. 1990. Human Error. Cambridge University Press

  15. [23]

    T.; and Uryasev, S

    Rockafellar, R. T.; and Uryasev, S. 2000. Optimization of Conditional Value-at-Risk. Journal of Risk, 2: 21--42

  16. [24]

    Sharma, M.; Tong, M.; Mu, J.; Wei, J.; Kruthoff, J.; Goodfriend, S.; Ong, E.; Peng, A.; Agarwal, R.; Anil, C.; et al. 2025. Constitutional Classifiers: Defending against universal jailbreaks across thousands of hours of red teaming. arXiv preprint arXiv:2501.18837

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.