Pith. sign in

World-model interpretability is all we need.AI Alignment Forum, January 2023

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it

citation-role summary

background 1

citation-polarity summary

fields

cs.AI 1

years

2025 1

verdicts

REJECT 1

roles

background 1

polarities

background 1

representative citing papers

Linear Spatial World Models Emerge in Large Language Models

cs.AI · 2025-06-03 · reject · novelty 5.0

Spatial relation words in LLaMA and Qwen models form antipodal, roughly orthogonal directions in a low-dimensional subspace, and steering along these directions changes the model's output.

citing papers explorer

Showing 1 of 1 citing paper.

  • Linear Spatial World Models Emerge in Large Language Models cs.AI · 2025-06-03 · reject · none · ref 4

    Spatial relation words in LLaMA and Qwen models form antipodal, roughly orthogonal directions in a low-dimensional subspace, and steering along these directions changes the model's output.