Pith. sign in

h4rm3l: A language for composable jailbreak attack synthesis

5 Pith papers cite this work. Polarity classification is still indexing.

5 Pith papers citing it

citation-role summary

background 1 baseline 1

citation-polarity summary

years

2026 4 2025 1

representative citing papers

Characterizing Model-Native Skills

cs.AI · 2026-04-19 · conditional · novelty 6.0

Recovering an orthogonal basis from model activations yields a model-native skill characterization that improves reasoning Pass@1 by up to 41% via targeted data selection and supports inference steering, outperforming human-characterized alternatives.

A Red-Team Study of Anthropic Fable 5 & Opus 4.8 Models

cs.CR · 2026-06-16 · unverdicted · novelty 4.0

Red-teaming of Fable 5 and Opus 4.8 shows adaptive automated attacks succeed on 6-11% of harmful intents, producing 2322 panel-confirmed harmful outputs despite resistance to static obfuscation.

citing papers explorer

Showing 5 of 5 citing papers.