Pith. sign in

hub

InProceedings of the Inter- national Conference on Learning Representations (ICLR)

16 Pith papers cite this work. Polarity classification is still indexing.

16 Pith papers citing it

hub tools

citation-role summary

method 1

citation-polarity summary

roles

method 1

polarities

background 1

representative citing papers

AutoRAN: Automated Hijacking of Safety Reasoning in Large Reasoning Models

cs.LG · 2025-05-16 · unverdicted · novelty 8.0

AutoRAN automates hijacking of safety reasoning in large reasoning models by simulating execution with a weaker model and iteratively exploiting reasoning patterns from refusals, reaching near-100% success on AdvBench, HarmBench, and StrongReject.

Searching for Privacy Risks in LLM Agents via Simulation

cs.CR · 2025-08-14 · unverdicted · novelty 7.0

A search-based simulation method uses LLMs as optimizers to discover escalating privacy attacks like impersonation and corresponding defenses like identity-verification state machines in LLM agent dialogues.

A StrongREJECT for Empty Jailbreaks

cs.LG · 2024-02-15 · conditional · novelty 6.0

StrongREJECT provides a standardized benchmark and evaluator for jailbreak attacks that aligns better with human judgments than prior methods and reveals that successful jailbreaks often reduce model capabilities.

MESA: Improving MoE Safety Alignment via Decentralized Expertise

cs.LG · 2026-05-30 · unverdicted · novelty 5.0

MESA decentralizes safety duties in MoE LLMs via expert capacity reallocation and dynamic routing refinement based on optimal transport theory, yielding robust defense on harmful benchmarks while preserving helpfulness.

A Red-Team Study of Anthropic Fable 5 & Opus 4.8 Models

cs.CR · 2026-06-16 · unverdicted · novelty 4.0

Red-teaming of Fable 5 and Opus 4.8 shows adaptive automated attacks succeed on 6-11% of harmful intents, producing 2322 panel-confirmed harmful outputs despite resistance to static obfuscation.

citing papers explorer

Showing 16 of 16 citing papers.