Pith. sign in

REVIEW 2 cited by

Dynamic Jailbreaking Attack

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2510.02422 v4 pith:QVNZ3VC7 submitted 2025-10-02 cs.CR cs.AI

classification cs.CRcs.AI
keywords optimizationdynamicadversarialpromptsstrategytargetattackcandidate
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Existing gradient-based jailbreak attacks typically optimize a fixed-length adversarial suffix toward a predefined target response with a static optimization strategy. However, this fully static formulation undermines the effectiveness, efficiency and flexibility of gradient-based jailbreaking because (i) A predefined target usually lies in the low-probability region of a safety-aligned LLM's conditional output distribution, forcing the optimization to pursue an unlikely response pattern; (ii) Simple affirmative targets may even mislead LLMs to generate affirmative responses that are not highly relevant to the prompts; (iii) Fixed optimization strategy and suffix length treat all prompts equally, leading to limited attack capability for hard prompts and redundant capacity for easy ones. To address these limitations, we propose Dynamic Jailbreaking Attack (DJA), a parameter-free gradient-based jailbreak framework using dynamic candidate exploration, dynamic relevant targets and dynamic optimization strategy to craft adversarial prompts. In each optimization round, DJA samples multiple candidate target responses directly from the LLM's distribution conditioned on the current adversarial prompt. Among these candidates, DJA employs a multi-objective scorer to select an optimal target that satisfies multi-dimensional criteria such as harmfulness, relevance, and usefulness. Moreover, DJA introduces a parameter-free dynamic optimization strategy that allocates adversarial effort based on real-time feedback, adapting suffix length, candidate sampling capacity, and optimization iterations according to the difficulty of each harmful prompt. In an extensive evaluation of 40 safety-aligned LLMs (12 model families, scaling from 0.5B to 32B), DJA achieves a 100% ASR across all LLMs, requiring only 13.68 optimization rounds on average (10 iterations per round).

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ProbGuard: Calibrated Safety Risk Estimation from LLM Output Distributions

    cs.LG 2026-08 conditional novelty 6.0 of 10

    ProbGuard uses early output distributions and Monte Carlo sampling to estimate calibrated unsafe-generation probability, beating 13 baselines on Brier/ECE and limiting jailbreak success to at most 1%.

  2. REIN: Bridging the Gap between Reasoning and Reliability via Reflection and Abstention Alignment

    cs.AI 2026-08 conditional novelty 5.0 of 10

    REIN trains reasoning models with reflection-veracity and abstention rewards, letting them answer or abstain in a single pass, and reports large drops in a false-endorsement metric across four benchmarks.

Pith tools