Pith. sign in

Designing and interpreting probes with control tasks.arXiv preprint arXiv:1909.03368

6 Pith papers cite this work, alongside 12 external citations. Polarity classification is still indexing.

6 Pith papers citing it
12 external citations · external index

citation-role summary

background 1

citation-polarity summary

years

2026 6

verdicts

UNVERDICTED 6

roles

background 1

polarities

unclear 1

representative citing papers

PRISM: Recovering Instruction Sets from Language Model Activations

cs.AI · 2026-06-08 · unverdicted · novelty 7.0

PRISM is a new activation-conditioned model that recovers full sets of simultaneous instructions from LLM hidden states via judge-guided GRPO training and outperforms prior activation-to-language methods on security-relevant tasks.

Contextual Linear Activation Steering of Language Models

cs.CL · 2026-04-27 · unverdicted · novelty 6.0

CLAS dynamically adapts linear activation steering strengths to context, outperforming fixed-strength steering and matching or exceeding ReFT and LoRA on eleven benchmarks across four model families with limited labeled data.

citing papers explorer

Showing 6 of 6 citing papers.