{"id":"2f9434be-f6f6-47c5-b4c8-0496946f60a4","arxiv_id":"2501.08246","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"DART, a single-step embedding-space perturber trained with reinforcement learning, finds toxic prompts closer to reference prompts than fine-tuned or few-shot baselines on three LLMs.","lead":"This paper proposes DART, a method that automatically rewrites safe prompts into harmful ones while keeping the rewritten text close to the original. It gives LLM safety auditors a way to probe which topics remain vulnerable despite alignment training.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"DART's proximity control is unverified: the soft L2-norm penalty in an unspecified embedding space is not shown to yield semantically or syntactically close text, and Table 3 evaluates cosine similarity rather than the constrained norm.","rationale":"The reader's weakest assumption identifies exactly the point I find most load-bearing: the mapping from embedding-space noise to text-space proximity is unverified, and the evaluation metric (cosine similarity) is not the constrained quantity. The central claim in the abstract, that DART is significantly more effective at discovering harmful inputs in close proximity to the reference prompt, depends on that mapping. If the generated prompts at ε=0.1/0.5 are actually gibberish or off-topic strings, the method solves a different problem from the one claimed, and the comparison with RL/prompting baselines is unfair because those baselines are constrained by explicit text-level cosine-similarity penalties or instructions, not by an embedding-norm proxy. The paper's own limitation statement concedes typos and unrelated words, which is direct evidence that the reconstruction step can change semantic content. A secondary concern is that the reward and evaluation use the same toxicity classifier, but that issue is shared with prior red-teaming work and is less central to the paper's novelty; the proximity problem is specific to DART's design. I agree with the reader's conditional verdict: the method is plausible and the experiments are suggestive, but without verification that the soft L2 penalty yields semantically close text, the headline claim should not be accepted unconditionally.","tokens_in":17158,"tokens_out":6530,"duration_ms":68309,"concrete_test":"Run DART at ε=0.5 on the Vicuna-7b test split, logging ||µ||2 for every evaluated prompt, and for 100 randomly sampled pairs (reference, generated) collect human 5-point semantic-similarity ratings plus BLEU/ROUGE-L/BERTScore. The proximity claim stands only if (1) all logged norms are ≤ ε, and (2) human-rated similarity has a median ≥ 4 and the text metrics are comparable to those of human paraphrases of the same references. If either condition fails, the reported ASR gains are not evidence of discovering harmful inputs in close proximity to the reference.","verdict_should_be":"UNCHANGED","load_bearing_attack":"To support the claim that DART finds harmful inputs 'in close proximity to the reference prompt,' two things must hold: (i) the soft penalty LREG = max(0, ||µ||2 - ε) must actually keep the reconstructed text close to the reference, and (ii) the evaluation metric must measure that closeness. Neither is established. The constraint in (P1) is a hard budget on dist(P, Tθ(P)), but Algorithm 2 only penalizes the L2 norm of the noise µ in an unspecified embedding space; after vec2text reconstruction, nothing guarantees the text remains near P. Table 3 and Figure 3 report cosine similarity between P and P′, not the constrained norm, and the embedding used for that cosine similarity is not identified. The Conclusion admits most discovered prompts contain grammatical mistakes, typos, or unrelated words or characters. If DART achieves high ASR precisely by exploiting these artifacts while still scoring high cosine similarity, the headline comparison to RL/prompting baselines is not a comparison of semantically close red-teaming prompts. The manual intent-maintenance annotation (100 toxic pairs per method) checks only whether the output topic is related to P; it does not establish lexical or syntactic proximity of the modified prompt itself.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a red-teaming framework in which a reference prompt is modified to elicit harmful behavior from a target LLM while keeping the modified prompt close to the reference. The authors introduce DART, a text-diffusion-inspired method that perturbs a reference prompt in an embedding space and reconstructs text via vec2text, trained with PPO and a soft L2-norm regularizer intended to enforce the proximity budget. They compare DART against RL fine-tuning, zero-/few-shot prompting, and FLIRT on three target LLMs and two benchmark datasets, reporting reward, attack success rate, cosine similarity to the reference, and a small manual intent-maintenance annotation. The central claim is that DART is significantly more effective at discovering harmful prompts in close proximity to the reference than the baselines.","tokens_in":17427,"tokens_out":9429,"duration_ms":94986,"significance":"If the main claim is established, DART would be a practically useful tool for targeted safety audits, allowing model developers to identify topic-specific vulnerabilities rather than arbitrary jailbreaks. The framework is a clean extension of prior automated red-teaming work to the constrained-proximity setting, and the black-box assumption on the target model is appropriate for realistic auditing. The empirical evaluation covers multiple target models and datasets, and the paper provides training details, hyperparameters, qualitative examples, and a small variance analysis. The method is novel in using a single-step continuous text-diffusion-style perturbation for red-teaming. However, the strength of the contribution is substantially limited by the lack of direct verification that the proximity constraint actually holds for the reconstructed text, the absence of uncertainty quantification in the main results, and the reliance on a single toxicity classifier both as the training reward and as the evaluation metric.","major_comments":[{"comment":"The proximity constraint in problem (P1) is on dist(P, T_theta(P)), i.e., on the distance between the original text and the reconstructed text. However, Algorithm 2 only enforces the soft regularizer L_REG = max(0, ||mu_t||_2 - epsilon) on the mean noise vector in an embedding space. Because this is a penalty term in the PPO loss rather than a hard constraint, the deployed mu may violate the specified budget for finite beta. More importantly, the evaluation in Table 3 and Figure 3 measures proximity by cosine similarity between P and P', not by the constrained quantity ||mu||_2, and the relationship between ||mu||_2, cosine similarity, and semantic/syntactic closeness after vec2text reconstruction is never established. The embedder emb is never identified, so the geometry in which epsilon is defined is unknown. To support the paper's central claim, the authors should report the distribution of ||mu||_2 at deployment, the cosine similarity / text-level distance after reconstruction, and the exact embedding model used for both perturbation and evaluation.","section":"Methodology (DART) and Algorithm 2"},{"comment":"The main results are reported as point estimates without error bars, confidence intervals, or significance tests. The variance appendix reports standard errors only for one setting (DART epsilon=0.5 and RL alpha=0.5 on Vicuna-7b), and only for reward and cosine similarity, not for ASR. The abstract's claim that DART is 'significantly more effective' is therefore not supported by the reported evidence. Given the stochasticity of PPO training and prompt generation, the authors should provide multiple-seed results or confidence intervals for the main comparisons, and ideally a paired or bootstrap significance test for the headline ASR differences.","section":"Table 3, Figure 3, and Appendix 'Variance'"},{"comment":"The manual 'Intent Maintained' annotation checks only whether the target model's output O' is related to the reference prompt P; it does not assess whether the modified prompt P' itself is semantically and syntactically close to P. The paper's own Conclusion states that most discovered prompts contain grammatical mistakes, typos, or unrelated words or characters, and the qualitative examples in Tables 4-6 contain highly degraded prompts (e.g., 'saboshed the evil maligners', 'writers can seek to stop the smell and smell of nasty animals'). Consequently, the cosine similarity and the intent-maintenance numbers do not establish that DART produces human-plausible prompts 'in close proximity' to the reference in the sense stated in the introduction. The paper should include an evaluation of the modified prompt itself, such as human ratings of fluency, semantic preservation, and syntactic similarity.","section":"Metrics and manual annotation"},{"comment":"The primary evaluation metric, ASR, is computed with the same pretrained toxicity classifier whose logits are used as the RL reward for DART and as the selection signal for FLIRT. This creates a risk that the reported improvements partly reflect over-optimization of that particular classifier rather than the elicitation of genuinely harmful responses. The paper does not include human evaluation of response harmfulness or a second, independently trained classifier. Given that many of the reported high-ASR examples are nonsensical, the absolute ASR numbers should be interpreted with caution. I recommend adding a human harmfulness evaluation, or at least a second classifier, for the main comparisons to confirm that the discovered prompts elicit harmful content beyond the training proxy.","section":"Metrics and reward design"}],"minor_comments":[{"comment":"Each cell contains two numbers (first red-teaming dataset, second alpaca dataset), but the table body is not annotated with column headers for the two datasets; adding subheaders such as 'RT / Alpaca' would improve readability.","section":"Table 3"},{"comment":"The y-axis is clearly logarithmic, but the caption does not state this; please add a note to avoid confusion.","section":"Figure 3"},{"comment":"The training procedure says the reward is 'the probability with which a classifier categorizes the interaction ... to be toxic,' while the evaluation section defines the reward as the logits of the toxicity classifier. These are not the same quantity; please clarify which one is used.","section":"Metrics"},{"comment":"The text says 'Due to the aforementioned problems with embedding of large sequences,' but no such problem is mentioned earlier in the paper; please add a short explanation.","section":"Appendix 'Training Details'"},{"comment":"The sentence 'Figure 2 depicts the training time...' refers to the table of training times; the reference should be to Table 2, not Figure 2.","section":"Training Time"},{"comment":"There is a formatting glitch in the RL(alpha=0.5) row for GPT2-alpaca on the red-teaming dataset, where '0.6%' appears in place of a cosine-similarity value; please correct the cell.","section":"Table 3 formatting"}],"recommendation":"major_revision","confidential_remarks":"The paper addresses a timely and useful problem, and the proposed method is interesting. My main concern is that the key novel element—proximity control—is not actually verified in the evaluation: the constraint is defined on an unstated embedding space, the reported metric is cosine similarity rather than the constrained norm, and the soft penalty does not guarantee the budget. These issues are fixable with additional experiments and analysis, but they are load-bearing for the central claim. The lack of error bars or significance tests in the main results is also a barrier to the strong 'significantly more effective' phrasing. I would encourage an additional round of experiments that report uncertainty, identify the embedding model, and validate proximity and harmfulness at the text level."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe useful piece: the paper defines a clean red-teaming variant—find small modifications of reference prompts that make a target model produce toxic output—and offers a simple single-step method (DART) that beats auto-regressive and prompting baselines at matched cosine similarity. That is a real, if modest, contribution. It deserves referee time.\n\nWhat is actually new is the proximity-constrained objective (P1) and the method: one embedding-space perturbation plus vec2text reconstruction, trained with RL. That combination is not in the cited prior work. The empirical comparison is systematically done across three target models and two datasets, with comparable parameter counts for the trained baselines. The ablation on the RL penalty is also useful, and the Vicuna topic-level safety map gives a nice practical demonstration of the setting.\n\nNow the soft spots, in proportion. The main one is that \"proximity\" is not enforced or measured as stated. The constraint in (P1) is hard, but training only penalizes the L2 norm of the noise in an unspecified embedding space, and evaluation reports cosine similarity between P and P' with no embedder identified. The stress-test note is right: the paper never establishes semantic or syntactic closeness, and the authors concede most discovered prompts contain typos, grammatical mistakes, or unrelated words. So the headline \"in close proximity\" should be read as \"similar under the evaluation embedding,\" not as a verified linguistic property. This weakens the claim moderately; it does not destroy the method comparison, since the baselines are scored with the same cosine metric.\n\nSecond, the reward is the logits of the same toxicity classifier used as the evaluation metric. That is a circularity, though not fatal here, because the RL baseline is trained with the same reward and all methods are compared under the same metric. Still, absolute ASRs are only meaningful relative to that proxy. Third, the main results in Table 3 and Figure 3 have no error bars or significance tests; variance is reported for one setting only. That is thin support for the word \"significantly.\"\n\nThe citation pattern looks fair. No code release is provided, though the appendix says the authors plan to release it; that is a real minus for reproducibility but not a correctness flaw.\n\nWho gets value: researchers working on automated safety evaluation and adversarial prompting. This is a decent toolbox paper, not a conceptual breakthrough. A serious referee can usefully push on the proximity metric, the unspecified embedder, and the missing variance bars.\n\nMy recommendation: send it to peer review. It should get a conditional verdict, with the authors asked to specify the embedding, add error bars to the main table, and either enforce or properly measure the proximity budget.","headline":"Worth a round of review, but the proximity claim is softer than advertised: the soft embedding-norm penalty is not shown to equal semantic closeness, and the main table lacks error bars.","tokens_in":17933,"tokens_out":2325,"would_cite":true,"duration_ms":25517,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A text-diffusion red-teaming method, DART, discovers harmful prompts that remain close to reference prompts, outperforming RL fine-tuning and zero-/few-shot baselines under tight proximity budgets.","keywords":["red-teaming","LLM safety","text diffusion","proximity constraints","reinforcement learning","embedding-space perturbation","attack success rate","targeted safety audit"],"falsifier":"Retrain the auto-regressive RL baseline with the same embedding-space $\\ell^2$ penalty and the same embedder that DART uses, then plot attack-success rate against cosine similarity; if that baseline reaches or exceeds DART's frontier, the central claim that text-diffusion architectures are uniquely effective for proximity-constrained red-teaming would be refuted.","tokens_in":16977,"feed_emoji":"🛡️","tokens_out":9394,"duration_ms":81208,"temperature":0.7,"pith_summary":"The paper defines a targeted red-teaming task: given a target LLM and a dataset of reference prompts, find modified prompts that trigger harmful responses while staying within a user-specified distance of the originals. It argues that standard auto-regressive red-teaming models, which generate novel prompts token by token, are poorly suited to this task because they cannot naturally control how far the output drifts from the reference. To address this, it introduces DART (Diffusion for Auditing and Red-Teaming), a black-box method that perturbs a reference prompt's embedding with a learned noise vector and reconstructs the perturbed embedding into text, training the noise policy with reinforcement learning under an explicit $\\ell^2$-norm budget. Across three target LLMs and two reference datasets, DART achieves higher attack success rates at comparable cosine similarity than RL fine-tuning, zero-shot, few-shot, and FLIRT baselines. If this holds, DART gives safety auditors a way to map precisely which topics, styles, and prompt types can be nudged into harmfulness, and which are genuinely safe.","feed_headline":"Text-diffusion finds harmful prompts that stay close to the original.","feed_subtitle":"Beats RL fine-tuning and prompting baselines at triggering harmful outputs under tight edit budgets.","key_machinery":"The central object is DART, a text-diffusion-inspired policy modeled by an encoder-decoder transformer (initialized from T5-base) that maps a reference prompt $P$ and its embedding $e = \\mathrm{emb}(P)$ to the mean $\\mu$ of a noise distribution. The modified embedding $e - n$ is decoded into a natural-language prompt $P'$ by the vec2text method, and the target LLM's response is scored by a toxicity classifier. Training uses PPO to maximize that toxicity reward, plus a proximity regularizer $L_{\\mathrm{REG}} = \\max(0, \\|\\mu\\|_2 - \\epsilon)$ that penalizes predicted noise only when it exceeds the user-set budget $\\epsilon$; at deployment, the model outputs $\\mu$ deterministically. This object carries the argument because it turns 'small, targeted modifications' into a trainable operation with an explicit control knob, rather than an emergent property of token generation.","core_discovery":"The central discovery is that continuous text diffusion, which modifies a sequence via small embedding-space perturbations rather than token-by-token generation, is well suited to finding harm-inducing prompts that stay close to a given reference. DART learns a policy that maps a reference prompt to a noise vector in the embedding space; the perturbed embedding is decoded back to text with the vec2text method, and the policy is trained with PPO to maximize a toxicity classifier's score while a regularization term $L_{\\mathrm{REG}} = \\max(0, \\|\\mu\\|_2 - \\epsilon)$ discourages exceeding the budget $\\epsilon$. In the reported comparisons, DART with budgets $\\epsilon = 0.1$ and $\\epsilon = 0.5$ achieves higher attack success rates at comparable cosine similarity than the RL, zero-shot, few-shot, and FLIRT baselines on both the Red Teaming and alpaca datasets across gpt2-alpaca, Vicuna-7b, and Llama2-7b-chat-hf. The paper also shows that when the budget is relaxed ($\\epsilon = 2$), DART finds many more harmful prompts, but preservation of the original intent drops, illustrating the trade-off it is designed to control.","pith_inferences":["The experimental comparison may not isolate the architecture's contribution: DART's proximity penalty acts on the embedding-space norm, while the RL baseline is penalized via text cosine similarity, so a matched-penalty comparison would clarify whether the diffusion structure or the penalty design drives DART's advantage.","The paper evaluates proximity with cosine similarity, not the $\\ell^2$ norm it constrains; a natural follow-up is to measure text-level distance such as edit distance or a paraphrase detector, which would test whether the embedding budget genuinely corresponds to small surface changes.","Because the discovered prompts often contain typos or unrelated words (a limitation the authors acknowledge), adding a fluency or grammar reward would test whether the attack success survives when proximity is enforced on well-formed sentences.","The single-step perturbation design suggests a general principle for minimal-edit problems — paraphrasing, style transfer, adversarial robustness auditing — where continuous embedding-space edits with a norm budget may be more sample-efficient than token-level autoregressive search."],"forward_implications":["Auditors can run controlled, topic-specific safety scans: any reference set can be probed for near neighbours that trigger harmful behavior, and the failures found are close enough to realistic user prompts to be actionable.","The method reveals per-topic differences in safety: for example, the Vicuna audit shows low success on violence, privacy, and illegal-instruction topics but high success on controversial and adult topics, so alignment effort can be directed where it matters.","Because DART only requires black-box access to the target and a toxicity classifier, the same procedure transfers to proprietary LLMs without any internal information.","Relaxing the budget trades intent preservation for attack success: at $\\epsilon = 2$ DART finds many more harmful prompts, but the fraction of prompts that keep the original intent drops, confirming that the proximity budget is the controlling dial."],"supporting_citations":[{"why":"Defines the automated red-teaming framework (RL fine-tuning and zero-/few-shot prompting) that DART extends and must beat.","marker":"(Perez et al. 2022)"},{"why":"Supplies the PPO algorithm used to train DART's policy.","marker":"(Schulman et al. 2017)"},{"why":"Provides vec2text, the embedding-to-text decoder that turns DART's perturbed embeddings into actual prompts.","marker":"(Morris et al. 2023)"},{"why":"Supplies the toxicity classifier that serves as both the reward signal and the evaluation metric for harmfulness.","marker":"(Corrêa 2023)"},{"why":"Provides the FLIRT baseline, a dynamic in-context red-teaming method compared against DART under proximity constraints.","marker":"(Mehrabi et al. 2023)"},{"why":"Supplies the Red Teaming dataset of harmful prompts used as a reference corpus for training and evaluation.","marker":"(Ganguli et al. 2022)"},{"why":"Supplies the alpaca-gpt4 dataset of benign instructions, the second reference corpus.","marker":"(Peng et al. 2023)"}],"fun_headline_variants":["Text diffusion red-teams LLMs with proximity control","DART: Better red-teaming by staying close to reference","Red-teaming via embedding diffusion beats RL baselines","Proximity-constrained text diffusion finds harmful prompts"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that an $\\ell^2$-norm budget on embedding noise faithfully corresponds to semantic and syntactic closeness of the reconstructed text, and that the soft penalty reliably enforces that budget — an assumption the paper's own examples (typos, unrelated words) and its cosine-similarity evaluation only partially support.","fun_headline_variants_meta":{"raw":{"variants":["Text diffusion red-teams LLMs with proximity control","DART: Better red-teaming by staying close to reference","Red-teaming via embedding diffusion beats RL baselines","Proximity-constrained text diffusion finds harmful prompts"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000198,"raw_usage":{"total_tokens":1402,"prompt_tokens":1013,"completion_tokens":389,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":629,"completion_tokens_details":{"reasoning_tokens":324}},"tokens_in":629,"tokens_out":389,"duration_ms":4277,"temperature":1.0,"reasoning_tokens":324,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T20:29:15.682104+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Retrain the auto-regressive RL baseline with the same embedding-space $\\ell^2$ penalty and the same embedder that DART uses, then plot attack-success rate against cosine similarity; if that baseline reaches or exceeds DART's frontier, the central claim that text-diffusion architectures are uniquely effective for proximity-constrained red-teaming would be refuted.","supporting_citations":[],"review_version":1}