Pith. sign in

Monotonic Representation of Numeric Properties in Language Models

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

Language models (LMs) can express factual knowledge involving numeric properties such as Karl Popper was born in 1902. However, how this information is encoded in the model's internal representations is not understood well. Here, we introduce a simple method for finding and editing representations of numeric properties such as an entity's birth year. Empirically, we find low-dimensional subspaces that encode numeric properties monotonically, in an interpretable and editable fashion. When editing representations along directions in these subspaces, LM output changes accordingly. For example, by patching activations along a "birthyear" direction we can make the LM express an increasingly late birthyear: Karl Popper was born in 1929, Karl Popper was born in 1957, Karl Popper was born in 1968. Property-encoding directions exist across several numeric properties in all models under consideration, suggesting the possibility that monotonic representation of numeric properties consistently emerges during LM pretraining. Code: https://github.com/bheinzerling/numeric-property-repr

citation-role summary

background 1

citation-polarity summary

fields

cs.LG 1

years

2025 1

verdicts

CONDITIONAL 1

roles

background 1

polarities

unclear 1

representative citing papers

Harmonic Loss Trains Interpretable AI Models

cs.LG · 2025-02-03 · conditional · novelty 5.0

Harmonic loss, which scores logits by inverse Euclidean distance to class prototypes with a scale-invariant normalization, yields more interpretable class centers, reduced grokking, and faster convergence across MLPs, transformers, MNIST, and a small GPT-2.

citing papers explorer

Showing 1 of 1 citing paper.

  • Harmonic Loss Trains Interpretable AI Models cs.LG · 2025-02-03 · conditional · none · ref 15 · internal anchor

    Harmonic loss, which scores logits by inverse Euclidean distance to class prototypes with a scale-invariant normalization, yields more interpretable class centers, reduced grokking, and faster convergence across MLPs, transformers, MNIST, and a small GPT-2.