Pith. sign in

hub

2025 IEEE Conference on Secure and Trustworthy Machine Learning (SaTML) , year =

16 Pith papers cite this work, alongside 94 external citations. Polarity classification is still indexing.

16 Pith papers citing it
94 external citations · external index

hub tools

citation-role summary

background 3

citation-polarity summary

years

2026 16

roles

background 3

polarities

background 3

representative citing papers

PRISM: Recovering Instruction Sets from Language Model Activations

cs.AI · 2026-06-08 · unverdicted · novelty 7.0

PRISM is a new activation-conditioned model that recovers full sets of simultaneous instructions from LLM hidden states via judge-guided GRPO training and outperforms prior activation-to-language methods on security-relevant tasks.

Adaptive Probe-based Steering for Robust LLM Jailbreaking

cs.CR · 2026-05-19 · unverdicted · novelty 5.0

Adaptive probe-based steering guided by model extraction and activation statistics improves LLM jailbreak success rates from 6% to 70% average harmfulness without extra contrastive prompts or manual tuning.

citing papers explorer

Showing 16 of 16 citing papers.