{"total":13,"items":[{"citing_arxiv_id":"2607.00572","ref_index":59,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment","primary_cat":"cs.AI","submitted_at":"2026-07-01T07:58:16+00:00","verdict":"ACCEPT","verdict_confidence":"HIGH","novelty_score":7.0,"formal_verification":"none","one_line_summary":"HARC couples harmfulness and refusal directions at prompt and response positions, yielding the best robustness-capability-usability trade-off among major safety methods.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.29441","ref_index":9,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Closing the Activation-Cone Blind Spot: Response-Time Probing and Unified Defense","primary_cat":"cs.CR","submitted_at":"2026-06-28T15:05:12+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"Response-time linear probing on first generated tokens detects prefilling attacks missed by prompt-time activation defenses, achieving 0/40 attack success and 0% false positives across seven models while composing orthogonally with AlphaSteer.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.24552","ref_index":37,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Ellipsoid Control: A White-list Jailbreak Defense via Benign Latent Modeling","primary_cat":"cs.CR","submitted_at":"2026-05-23T12:39:25+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"Ellipsoid Control is a white-list test-time jailbreak defense that fits an anisotropic ellipsoid from benign activations to constrain projected gradient descent updates, aiming to improve the safety-utility tradeoff over black-list RepE methods.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.24154","ref_index":59,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Palette: A Modular, Controllable, and Efficient Framework for On-demand Authorized Safety Alignment Relaxation in LLMs","primary_cat":"cs.AI","submitted_at":"2026-05-22T19:22:17+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":5.0,"formal_verification":"none","one_line_summary":"Palette identifies refusal directions via multi-objective search, internalizes them through lightweight adaptation, and supports on-demand multi-domain authorization via independent learning and parameter merging.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.18104","ref_index":61,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Safety Geometry Collapse in Multimodal LLMs and Adaptive Drift Correction","primary_cat":"cs.AI","submitted_at":"2026-05-18T09:16:55+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":5.0,"formal_verification":"none","one_line_summary":"Multimodal LLMs suffer Safety Geometry Collapse from modality-induced drift that reduces refusal separability; ReGap corrects drift at inference time using self-rectification signals to restore safety without retraining.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.15687","ref_index":17,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"ASRU: Activation Steering Meets Reinforcement Unlearning for Multimodal Large Language Models","primary_cat":"cs.CL","submitted_at":"2026-05-15T07:22:43+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"ASRU combines activation redirection and reward-optimized fine-tuning to unlearn cross-modal sensitive knowledge in MLLMs, reporting +24.6% better unlearning effectiveness and 5.8x higher generation quality on Qwen3-VL while preserving utility with limited retained data.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.12890","ref_index":47,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Steer-to-Detect: Probing Hidden Representations for Detection of LLM-Generated Texts","primary_cat":"stat.AP","submitted_at":"2026-05-13T02:14:21+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"Steer-to-Detect learns a steering vector injected into LLM hidden states to boost class separability and applies hypothesis testing with finite-sample Type I/II error guarantees for generated-text detection.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.08930","ref_index":19,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Internalizing Safety Understanding in Large Reasoning Models via Verification","primary_cat":"cs.AI","submitted_at":"2026-05-09T13:05:00+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"Training large reasoning models only on safety verification tasks internalizes safety understanding and boosts robustness to out-of-domain jailbreaks, providing a stronger base for reinforcement learning alignment than standard supervised fine-tuning.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.06342","ref_index":33,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Don't Lose Focus: Activation Steering via Key-Orthogonal Projections","primary_cat":"cs.CL","submitted_at":"2026-05-07T14:29:18+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"SKOP uses key-orthogonal projections to steer LLM activations while preserving attention patterns on focus tokens, cutting utility degradation by 5-7x and retaining over 95% of standard steering efficacy.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.01167","ref_index":11,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Minimizing Collateral Damage in Activation Steering","primary_cat":"cs.LG","submitted_at":"2026-05-01T23:52:54+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"Activation steering is cast as constrained optimization that minimizes collateral damage by weighting perturbations according to the empirical second-moment matrix of activations instead of assuming isotropy.","context_count":1,"top_context_role":"background","top_context_polarity":"background","context_text":"First, let's diagonalizeA=QDQ ⊤ so that we can solve orthogonal component-wise. Notice thatAis diagonalizable because it is a symmetric matrix. Lety=Q ⊤uandc=Q ⊤b. Equation (10) becomes: r2Dy−λy=−rc r2βiyi −λy i =−rc i, i= 1, . . . , p−2 yi = −rci r2βi −λ We enforce the constraint: ∥y∥2 =y ⊤y=u ⊤QQ⊤u=u ⊤u= 1 p−2X i=1 y2 i = p−2X i=1 r2c2 i (r2βi −λ) 2 = 1 ⇒  X i r2c2 i Y j̸=i (r2βj −λ) 2   − p−2Y k (r2βk −λ) 2 = 0(11) Equation (11) is a polynomial of λ with degree 2(p−2) . It has finite number of roots (at most 2(p−2) ). Any valid critical point of ϕ(x) must be generated by λ that solves (11). So does any critical point of J(x). Therefore the set of critical points of J is Finite. ⇒Ω , the set of accumulation points of x(t) is also finite. Since Ω is also connected, it must have only 1"},{"citing_arxiv_id":"2604.12616","ref_index":42,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Every Picture Tells a Dangerous Story: Memory-Augmented Multi-Agent Jailbreak Attacks on VLMs","primary_cat":"cs.AI","submitted_at":"2026-04-14T11:44:59+00:00","verdict":null,"verdict_confidence":null,"novelty_score":null,"formal_verification":null,"one_line_summary":null,"context_count":1,"top_context_role":"background","top_context_polarity":"background","context_text":"Exist- ing attacks use (a) textual manipulation, (b) visual perturba- tions, (c) typographic or (d) harmful images. (e) Ours Mem- Jack exploits original natural images via multi-agent visual- semantic camouflage with memory-augmented reflection. architectural convergence fundamentally alters and drastically ex- pands the adversarial attack surface [ 42]. While contemporary safety alignment techniques have been proven highly effective in unimodal environments, they frequently fail to generalize across the multimodal interface [28, 48]. The semantic gap between visual perception and text generation acts as an unconstrained conduit, allowing benign visual elements to be weaponized [17]. Therefore, because the information density of images is much higher than that"},{"citing_arxiv_id":"2604.12384","ref_index":40,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Preventing Safety Drift in Large Language Models via Coupled Weight and Activation Constraints","primary_cat":"cs.AI","submitted_at":"2026-04-14T07:17:55+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"Coupled constraints on weight updates in a safety subspace and regularization of SAE-identified safety features preserve LLM refusal behaviors during fine-tuning better than weight-only or activation-only methods.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2604.07727","ref_index":32,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"TrajGuard: Streaming Hidden-state Trajectory Detection for Decoding-time Jailbreak Defense","primary_cat":"cs.CR","submitted_at":"2026-04-09T02:22:44+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"TrajGuard detects jailbreaks by tracking how hidden-state trajectories move toward high-risk regions during decoding, achieving 95% defense rate with 5.2 ms/token latency across tested attacks.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null}],"limit":50,"offset":0}