Pith. sign in

REVIEW 2 cited by

Adversarial Representation Engineering: A General Model Editing Framework for Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.13752 v3 pith:EBPIF5GW submitted 2024-04-21 cs.LG cs.AIcs.CLcs.CRmath.OC

classification cs.LGcs.AIcs.CLcs.CRmath.OC
keywords editingrepresentationmodelengineeringadversarialframeworkgenerallanguage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Since the rapid development of Large Language Models (LLMs) has achieved remarkable success, understanding and rectifying their internal complex mechanisms has become an urgent issue. Recent research has attempted to interpret their behaviors through the lens of inner representation. However, developing practical and efficient methods for applying these representations for general and flexible model editing remains challenging. In this work, we explore how to leverage insights from representation engineering to guide the editing of LLMs by deploying a representation sensor as an editing oracle. We first identify the importance of a robust and reliable sensor during editing, then propose an Adversarial Representation Engineering (ARE) framework to provide a unified and interpretable approach for conceptual model editing without compromising baseline performance. Experiments on multiple tasks demonstrate the effectiveness of ARE in various model editing scenarios. Our code and data are available at https://github.com/Zhang-Yihao/Adversarial-Representation-Engineering.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Shaking to Reveal: Perturbation-Based Detection of LLM Hallucinations

    cs.AI 2025-06 conditional novelty 6.0 of 10

    SSP adds a learned, sample-specific noise prompt to an LLM input and scores hallucination by the cosine shift in intermediate representations, outperforming output-confidence baselines on QA benchmarks.

  2. Advancing LLM Safe Alignment with Safety Representation Ranking

    cs.CL 2025-05 reject novelty 4.0 of 10

    Safety Representation Ranking (SRR) trains a lightweight transformer on internal LLM hidden states to rank candidate responses by safety, reporting high pairwise accuracy on safety benchmarks.

Pith tools