← back to paper
arxiv: 2506.14682 · 2 revisions
AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models