REVIEW 3 major objections 4 minor 17 references
Truth in small language models is a single knowledge-gated direction that attention builds and the feed-forward network overwrites.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 20:04 UTC pith:GXFNXPXE
load-bearing objection A serious, self-correcting empirical study of truth directions in small LMs — worth refereeing, but the headline 'propagate vs. overwrite' law is conditional on a pre-frame control that leaves an admitted indirect-path confound. the 3 major comments →
The Anatomy of a Truth Direction: Knowledge-Dependent Dimensionality, a Relational Law, and a Convergent Category Geometry in Small Language Models
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central claim: a single SVD axis over true/false minimal-pair differences captures most of the truth signal on facts a model knows, and its failure on unknown or heterogeneous material is itself the signal—the dimensionality of truth is knowledge-dependent. Polarity's second dimension exists but is recoverable without polarity labels. Inside the residual stream, attention propagates truth frames it did not write, the feed-forward network opposes the current frame, and the post-peak erosion is causally carried by the SwiGLU value stream. Per-category truth axes form a semantically signed arrangement that converges across model families, knowledge-gated and quantified as classical attenuat
What carries the argument
The central object is the truth axis: the top right-singular vector of the SVD of hidden-state differences over true/false minimal pairs, read at the last token and identified without labels up to one global sign (Lemma 1). It costs two dot products per token and needs no training. Carrying the argument are three further instruments: an exact additive decomposition of residual states into attention and FFN contributions, gated by an identity check; a pre-frame versus post-frame control that removes the containment circularity and exposes the relational law; and an exact SwiGLU split into gate and value terms (plus a norm term for the sandwich-normalized architecture) that attributes post-pea
Load-bearing premise
The pre-frame control is assumed to remove all effects of the FFN contribution being measured, but indirect paths—accumulated earlier FFN writes that shape the pre frame—are not removed, and Empirical Law 1's attribution depends on those indirect paths being negligible.
What would settle it
Ablate the FFN in the layers before a frame block and re-fit the pre frame: if the axis moves materially, the measured 'opposition' of the FFN is an artifact of earlier writes rather than a genuine law; if the axis is stable, indirect contamination is negligible. The paper's own seven falsifications targeted the one-dimensionality claim, so a decisive test of the relational law should target the contamination path directly.
If this is right
- A training-free, two-dot-product axis can trace where truth is built (attention in middle layers), where it is degraded (FFN value stream after the peak), and why the single-axis reading fails (unknown or heterogeneous material).
- On facts a model knows, adding a second dimension to the truth axis adds no separation; the one-dimensional reading survives a seven-experiment falsification battery.
- The same axis has two readings: its sign classifies true vs false, and its separability measures how well the model knows the material—a geometric, label-free confidence estimate.
- Per-category truth axes form a semantically signed arrangement that converges across model families; the convergence is knowledge-gated, so the arrangement is a property of the knowledge domain, not the training recipe.
- A measured dictionary of per-category axes and FFN write directions is exported, with a four-criterion contract for any sparse autoencoder claiming to do better.
Where Pith is reading between the lines
- If the knowledge gate is real, the same mechanism predicts where hallucination is geometrically invisible: facts a model does not know produce overlapping true/false representations that no single direction can separate, suggesting detection must shift to per-atom expansion (e.g., learned dictionaries) in the low-knowledge regime.
- The propagate/overwrite law suggests a candidate microscopic writer for macroscopic rotational-dynamics observations in factual-constraint processing; connecting the two could turn a correlational anatomy into a circuit-level mechanism.
- The private geometry observed in one model family's dense bundle, if confirmed, implies that training diet (e.g., code and mathematics corpora) leaves a measurable residue in how truth is represented, beyond what shared knowledge explains; a fourth family with a known training pedigree would stress this.
- The consensus sign gauge may generalize beyond truth axes: any SVD-based analysis with per-component orientations identified only up to sign can use the leading eigenvector of the cosine matrix to stabilize cross-condition comparisons.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper extends Bürger et al. (2024) in three directions: the knowledge-dependence of truth-direction dimensionality, the architectural component that builds and degrades the truth direction, and the semantic mixture underlying it. Part I, on Qwen2.5-1.5B, constructs a training-free axis from the SVD of hidden-state minimal pairs, reports held-out AUC 0.938 on 45 curated pairs (permutation p≈0.005, full-probe ceiling 0.989), shows the axis-to-probe gap grows on less-known and heterogeneous material, recovers the supervised polarity direction t_P without polarity labels, and reports that seven falsification attempts fail to overturn one-dimensionality on known facts. An exact attention/FFN decomposition and causal ablations attribute signal construction to attention in middle layers and post-peak erosion to the FFN. Part II extends the framework to five models across three families, introducing truth frames (pre and post), a fixed-frame gap statistic, and a SwiGLU split. The headline claims are: Empirical Law 1 (attention propagates frames it did not write; the FFN opposes them), the causal attribution of post-peak erosion to the SwiGLU value stream, a semantically signed per-category geometry that converges across families under a spectral consensus sign gauge, a knowledge-gated interpretation of that convergence as classical attenuation, and a model-private Qwen geometry at high sample density. The paper includes extensive registered predictions, negative results, tooling
Significance. If the central claims hold, the paper makes a substantial contribution: a cheap, training-free probe competitive with trained probes on known facts; a mechanistic, cross-family law about attention and FFN contributions; and a quantitative, knowledge-gated account of when single-direction truth reading fails. The manuscript's methodological discipline is a genuine strength: a one-command-per-number reproducibility contract, held-out protocols with permutation nulls that price layer selection, pre-registered predictions with failures reported as results, exact decomposition identity checks, and a synthetic axis-provenance control. The paper is also unusually honest, explicitly flagging confounds (Secs. 12 and 26) and correcting its own earlier interpretations. However, the most novel claim — Empirical Law 1 — rests on a pre-frame control whose admitted residual contamination has not been ruled out, and the knowledge-dependent dimensionality claim of Part I remains partially confounded with dataset and template identity. These are correctness risks on load-bearing points, not merely presentation issues.
major comments (3)
- [Secs. 16.1, 17.3, 26; Table 9] Empirical Law 1 depends on the pre-frame control. The pre frame F_pre(b) is fit on h_b + a_b, which contains all FFN writes from earlier blocks. Section 26 admits: 'The pre frame removes the direct containment circularity but not indirect paths (accumulated earlier writes shape the pre frame).' If the pre-frame axis is substantially aligned with the subspace of these earlier FFN writes, then the measured negative d'(f_L; F_pre(b)) for L > b could be autocorrelation between successive FFN writes rather than a general 'propagate vs. overwrite' regularity. The synthetic axis-provenance control (Sec. 16.6) validated the pre-frame collapse on planted circular writers, but it did not model this specific indirect-path contamination. I suggest a direct test: regress the pre-frame axis onto the span of FFN writes from blocks < b, and re-measure the sub-diagonal of Table 9 after abating earlier FF
- [Sec. 5.2, Table 1, Sec. 12] The knowledge-dependent dimensionality claim compares 45 curated pairs (behaviorally verified, one template set) with 250 CounterFact pairs (know-rate 30%, many templates) and a CounterFact/TruthfulQA mix. The axis-to-probe gap differences (0.051, 0.062, 0.203) are therefore confounded with dataset identity, template syntax, and topic heterogeneity. The paper itself notes in Sec. 12 that the 'unknown-fact' reading is 'confounded with sentence-format effects' and that filtering by a behavioral label is future work. Since the statement 'when the model knows a fact, its truth content is concentrated on a single direction' is one of the paper's two central contributions, the claim needs a within-dataset, within-template comparison of behaviorally known vs. unknown facts. As it stands, the evidence supports a monotone relation between know-rate and axis gap at the dataset level, but not the s
- [Secs. 21, 22, 26; Table 13] The arrangement law is promoted to a law at the registered protocol, but the underlying category geometry has no bootstrap error bars (admitted in Sec. 26: 'no bootstrap error bars'), and the law is stated only after a sign-gauge repair. The sign instability documented in Sec. 22 is substantial: three of six Qwen seed re-samplings anti-correlate with the canonical matrix under the per-category gauge, and the dense cross-family comparison fails under that gauge. The consensus gauge repairs this, but it is selected post hoc from the same data. The authors argue convincingly that the consensus gauge is computed independently per model and that relative signs are gauge-invariant, but the report should include bootstrap confidence intervals on the Mantel statistics and on the per-category cosines, and a sensitivity analysis showing the law's conclusions are stable across reasonable alternativ
minor comments (4)
- [Throughout] The manuscript alternates between 'hidden level' and 'block' indexing (Part I peak layer 16 for Qwen2.5-1.5B; Table 6 peak p=15). This is consistent but easy to misread; a single explicit conversion sentence in Sec. 4 or 15 would help.
- [Sec. 4.2] 'Know-rate' is defined as a string-match lower bound on 200 prompts, and Sec. 22 later shows its per-relation version correlates only weakly with adherence (Spearman +0.21). This limitation is disclosed, but the narrative sometimes leans on the word 'knows' as if it were an established behavioral label; rephrasing to 'behavioral lower-bound know-rate' throughout would reduce over-reading.
- [Sec. 22, Eq. (8)] The spectral consensus gauge is introduced via the leading eigenvector of the cosine matrix. It would help to state explicitly that the eigenvector is defined only up to sign and that the final sign is fixed by majority, as is done in the text, but the equation itself could carry a comment to that effect.
- [Sec. 23] The private-geometry section is presented as a finding, though it rests on one dense bundle per family and its three predictions are registered but untested. The text does mark it as 'beyond the gate' and the mechanism as 'declared, not established,' which is appropriate; consider labeling it Tier C given the absence of a tested prediction.
Circularity Check
No significant circularity: the one admitted containment artifact is corrected, and the final laws rest on pre-frame controls plus independent held-out evidence.
full rationale
The paper's derived chain is self-contained rather than circular. The only place the word 'circularity' appears is in the admitted limitation that the pre-frame 'removes the direct containment circularity but not indirect paths' (Sec. 26), and in the correction of the earlier post-frame flip (Sec. 17.2), where the axis fit on h_{b+1} literally contained f_b and therefore dragged itself toward its own ingredient. That measurement is explicitly exposed as an artifact and is not load-bearing for the final claims: Empirical Law 1 is validated on the sub-diagonal of the pre-frame matrices, where neither a_L nor f_L is contained in the frame, and the erosion attribution uses frame-free ablations and an exact algebraic SwiGLU split (Eq. 5). The central dimensionality claim is evaluated held-out with permutation nulls against external datasets (CounterFact, TruthfulQA) and a full-probe ceiling, not fitted into the conclusion. The consensus gauge is computed independently per model, and the paper explicitly notes that two independent gauges cannot collude. The residual indirect-path contamination is a real, declared validity threat to the law's mechanistic reading, but it is a correctness risk, not a definitional reduction of the law to its inputs. No self-citation is load-bearing, and no fitted parameter is renamed as a prediction in the final derivation chain.
Axiom & Free-Parameter Ledger
free parameters (6)
- Peak layer p per model =
15, 16, 7, 9, 11 (Qwen-1.5B, Qwen-3B, Llama-1B, Llama-3B, Gemma)
- Consensus sign gauge s* =
leading eigenvector of C (Eq. 8)
- Knowledge-restriction threshold tau =
0.6 (within-category held-out AUC on both families)
- Stability floor and significance margins =
2 sigma and 0.03 AUC/d'
- Category protocol K and pairs per relation =
K=8 (top relations) with 60 pairs; K=33 with 60 and 888 pairs
- Registered magnitude windows for confirmations =
e.g., R1 cosine window [+0.15, +0.45]
axioms (7)
- domain assumption Minimal-pair hidden-state differences isolate truth from topic and surface form
- domain assumption Truth is linearly readable in the residual stream
- domain assumption Greedy-completion know-rate measures behavioral knowledge
- standard math Residual stream decomposes additively into attention and FFN contributions
- domain assumption The pre-frame excludes all effects of the measured FFN write
- ad hoc to paper Sign comparability of per-category axes across models
- ad hoc to paper One-factor latent arrangement for attenuation/reliability
invented entities (3)
-
Truth frame (pre/post)
no independent evidence
-
Knowledge gate / classical attenuation
independent evidence
-
Private geometry (Qwen dense bundle)
no independent evidence
read the original abstract
B\"urger et al. (2024) demonstrated that truth representations in large language models are universal across statement polarity but reside within a multidimensional subspace. We extend this framework along three questions: how the dimensionality of the subspace depends on the model's knowledge, which architectural component builds the truth direction, and what the direction is a mixture of. In Part I, a training-free directional probe derived from the SVD of hidden-state minimal pairs shows that the dimensionality of truth is knowledge-dependent: the signal concentrates on a single axis for known facts and diffuses as knowledge decreases. In Part II, a relational law emerges across multiple model families: attention propagates truth frames, the feed-forward network opposes the current block's frame, and post-peak decay is causally attributed to the SwiGLU value stream. Furthermore, per-category truth axes form a semantically signed arrangement that converges across families. Stress tests expose a sign instability in this orientation, which we repair with a spectral consensus gauge to sharpen the convergence into a knowledge-gated law. Finally, a replication campaign on Gemma-2-2b, extending our decomposition tools to accommodate its sandwich normalization, confirms these laws and attributions. We quantify the knowledge gate as classical attenuation and isolate a stable, model-specific private geometry.
Figures
Reference graph
Works this paper leans on
-
[1]
C. Burns, H. Ye, D. Klein, J. Steinhardt.Discovering Latent Knowledge in Language Models Without Supervision.arXiv:2212.03827, 2022 (ICLR 2023)
Pith/arXiv arXiv 2022
-
[2]
S. Marks, M. Tegmark.The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets.arXiv:2310.06824, 2023 (COLM 2024)
Pith/arXiv arXiv 2023
-
[3]
K. Li, O. Patel, F. Viégas, H. Pfister, M. Wattenberg.Inference-Time Intervention: Eliciting Truthful Answers from a Language Model.NeurIPS, 2023
2023
-
[4]
Zou et al.Representation Engineering: A Top-Down Approach to AI Transparency
A. Zou et al.Representation Engineering: A Top-Down Approach to AI Transparency. arXiv:2310.01405, 2023
Pith/arXiv arXiv 2023
-
[5]
L. Bürger, F. A. Hamprecht, B. Nadler.Truth is Universal: Robust Detection of Lies in LLMs. arXiv:2407.12831, 2024 (NeurIPS 2024)
Pith/arXiv arXiv 2024
-
[6]
H. Cunningham, A. Ewart, L. Riggs, R. Huben, L. Sharkey.Sparse Autoencoders Find Highly Interpretable Features in Language Models.arXiv:2309.08600, 2023
Pith/arXiv arXiv 2023
-
[7]
Bricken et al.Towards Monosemanticity: Decomposing Language Models With Dictionary Learning.Transformer Circuits Thread, Anthropic, 2023
T. Bricken et al.Towards Monosemanticity: Decomposing Language Models With Dictionary Learning.Transformer Circuits Thread, Anthropic, 2023
2023
-
[8]
N. Elhage et al.Toy Models of Superposition.Transformer Circuits Thread, Anthropic, 2022 (arXiv:2209.10652)
Pith/arXiv arXiv 2022
-
[9]
Elhage et al.A Mathematical Framework for Transformer Circuits.Transformer Circuits Thread, Anthropic, 2021
N. Elhage et al.A Mathematical Framework for Transformer Circuits.Transformer Circuits Thread, Anthropic, 2021
2021
-
[10]
G. Bar-Shalom, F. Frasca, Y. Galron, Y. Ziser, H. Maron.Beyond Token Probes: Hallucination Detection via Activation Tensors with ACT-ViT.arXiv:2510.00296, 2025
arXiv 2025
-
[11]
K. Meng, D. Bau, A. Andonian, Y. Belinkov.Locating and Editing Factual Associations in GPT.NeurIPS, 2022. Dataset mirror:NeelNanda/counterfact-tracing
2022
-
[12]
ICR Probe: Tracking Hidden State Dynamics for Reliable Hallucination Detection in LLMs. arXiv:2507.16488, 2025
Pith/arXiv arXiv 2025
-
[13]
LLM Hallucination Detection: A Fast Fourier Transform Method Based on Hidden Layer Temporal Signals (HSAD).arXiv:2509.13154, 2025
arXiv 2025
-
[14]
How Transformers Reject Wrong Answers: Rotational Dynamics of Factual Constraint Process- ing.Preprint, February 2026
2026
-
[15]
The HeadGenome Taxonomy: A Structural and Causal Map of Attention Head Specialization in Transformers.Preprint, OpenReview iG4i2pc4wg
-
[16]
M. Huh, B. Cheung, T. Wang, P. Isola.The Platonic Representation Hypothesis.ICML 2024. arXiv:2405.07987
Pith/arXiv arXiv 2024
-
[17]
S. Lin, J. Hilton, O. Evans.TruthfulQA: Measuring How Models Mimic Human Falsehoods. ACL, 2022. 41
2022
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.