Pith. sign in

REVIEW

Sparse Activation Editing for Reliable Instruction Following in Narratives

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2505.16505 v1 pith:HBBZ2ASJ submitted 2025-05-22 cs.CL cs.AIcs.HC

classification cs.CLcs.AIcs.HC
keywords instructionfollowingcomplexconcise-saeeditinginstructionslanguagenarratives
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Complex narrative contexts often challenge language models' ability to follow instructions, and existing benchmarks fail to capture these difficulties. To address this, we propose Concise-SAE, a training-free framework that improves instruction following by identifying and editing instruction-relevant neurons using only natural language instructions, without requiring labelled data. To thoroughly evaluate our method, we introduce FreeInstruct, a diverse and realistic benchmark of 1,212 examples that highlights the challenges of instruction following in narrative-rich settings. While initially motivated by complex narratives, Concise-SAE demonstrates state-of-the-art instruction adherence across varied tasks without compromising generation quality.

Discussion (0). Sign in to comment.

Pith tools