Pith. sign in

InInternational Conference on Learning Representations

8 Pith papers cite this work. Polarity classification is still indexing.

8 Pith papers citing it

citation-role summary

baseline 1

citation-polarity summary

years

2026 8

roles

baseline 1

polarities

baseline 1

representative citing papers

CURE:Circuit-Aware Unlearning for LLM-based Recommendation

cs.IR · 2026-04-04 · unverdicted · novelty 7.0

CURE disentangles LLM recommendation circuits into forget-specific, retain-specific, and task-shared modules with tailored update rules to achieve more effective unlearning than weighted baselines.

Distributed Sparse Interventions in Language Models

cs.LG · 2026-07-08 · conditional · novelty 6.0

Sparse interventions on 8–64 neurons distributed across layers can activate task behavior in instruction-tuned LLMs, outperforming first-order linear steering approaches by modeling nonlinear neuron interactions.

Judge Circuits

cs.CL · 2026-05-15 · unverdicted · novelty 6.0 · 2 refs

A sparse, format-independent Latent Evaluator circuit in mid-to-late MLPs computes judgments across tasks; ablating it removes judgment while leaving world knowledge intact.

ADAG: Automatically Describing Attribution Graphs

cs.CL · 2026-04-08 · unverdicted · novelty 6.0

ADAG automates description of attribution graphs in language model interpretability by combining gradient-based attribution profiles, a new clustering algorithm, and an LLM explainer-simulator to recover interpretable circuits and identify steerable clusters for jailbreaks.

Fast & Faithful Function Vectors

cs.CL · 2026-06-03 · unverdicted · novelty 4.0

LRP-based attention head selection and distributed application improve the efficiency and accuracy of function vectors for steering LLMs compared to prior choices.

citing papers explorer

Showing 8 of 8 citing papers.

  • CURE:Circuit-Aware Unlearning for LLM-based Recommendation cs.IR · 2026-04-04 · unverdicted · none · ref 20

    CURE disentangles LLM recommendation circuits into forget-specific, retain-specific, and task-shared modules with tailored update rules to achieve more effective unlearning than weighted baselines.

  • From Attribution to Action: A Human-Centered Application of Activation Steering cs.AI · 2026-04-13 · conditional · none · ref 27

    Activation steering of SAE-attributed components lets practitioners move from correlational inspection to causal hypothesis testing on CLIP failures, with trust shifting to observed model responses (N=8 experts).

  • Distributed Sparse Interventions in Language Models cs.LG · 2026-07-08 · conditional · none · ref 24

    Sparse interventions on 8–64 neurons distributed across layers can activate task behavior in instruction-tuned LLMs, outperforming first-order linear steering approaches by modeling nonlinear neuron interactions.

  • When Attribution Patching Lies: Diagnosis and a Second-Order Correction cs.LG · 2026-06-05 · unverdicted · none · ref 9

    Dominant error in attribution patching arises from downstream non-linearities; a single HVP correction removes the leading error term and matches Integrated Gradients accuracy at lower cost across 124M-9B models.

  • Judge Circuits cs.CL · 2026-05-15 · unverdicted · none · ref 5 · 2 links

    A sparse, format-independent Latent Evaluator circuit in mid-to-late MLPs computes judgments across tasks; ablating it removes judgment while leaving world knowledge intact.

  • Seeing Through Circuits: Faithful Mechanistic Interpretability for Vision Transformers cs.AI · 2026-04-15 · unverdicted · none · ref 20

    Edge-based circuits in vision transformers can be automatically recovered to explain and steer model computations for classification and adversarial behaviors.

  • ADAG: Automatically Describing Attribution Graphs cs.CL · 2026-04-08 · unverdicted · none · ref 2

    ADAG automates description of attribution graphs in language model interpretability by combining gradient-based attribution profiles, a new clustering algorithm, and an LLM explainer-simulator to recover interpretable circuits and identify steerable clusters for jailbreaks.

  • Fast & Faithful Function Vectors cs.CL · 2026-06-03 · unverdicted · none · ref 37

    LRP-based attention head selection and distributed application improve the efficiency and accuracy of function vectors for steering LLMs compared to prior choices.