REVIEW 2 major objections 4 minor
Falsehood and Impossibility Are Different Directions in an AI's Representation of Language
T0 review · 2 major / 4 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read A language model's internal geometry treats impossibility as a separate direction from falsehood.
desk verdict A careful small-model probe study showing truth and impossibility as near-orthogonal directions, but the 'impossible' category contains at least one genuinely true item and the below-chance AUC is misread as chance. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the linear probe readout on residual-stream states, together with the angular comparison of probe weight vectors. Each probe is a regularized logistic regression fitted on grouped, family-held-out folds; the impossibility probe contrasts impossible statements against true, false, and improbable ones, while the truth probe contrasts false against true and the anomaly probe contrasts anomalous against the three possible conditions. The double dissociation—the truth probe cannot order impossible against false, while the impossibility probe can—and the near-orthogonality of the fitted directions carry the argument. A pretrained sparse autoencoder at the impossibility peak (layer 15) supplies a feature-level confirmation.
What would settle it
Train the same impossibility probe on a new set of 'necessary falsehoods' built with completely different surface forms (for example, generated by conjoining each proposition with its negation in varied vocabularies) while holding topic vocabulary matched across conditions; if probe AUC on held-out families drops substantially, the direction is capturing lexical templates rather than impossibility. Alternatively, replace the necessary-falsehood items with pragmatically odd but logically possible sentences; if the probe still separates them from contingent falsehood at AUC 1.00, the direction encodes oddity, not modality.
Extended reading notes
Core claim
The paper's central claim is that in the residual stream of Gemma 3 4B IT, impossibility and falsehood are carried by nearly orthogonal linear directions. The truth direction, which separates true from false statements with balanced accuracy 0.93, fails to separate impossible from merely false statements (AUC 0.20); an impossibility probe fitted on held-out topic families does separate them at AUC 1.00. The absolute cosine between the two directions stays at or below 0.12 beyond depth 10. Sparse autoencoder features at layer 15 repeat this geometry: features selective for impossibility fire on anomalous sentences but rarely on contingent falsehoods. The author states the result as a representational observation, not a claim that impossible statements are intrinsically meaningless.
Load-bearing premise
The handwritten 'necessary falsehood' stimuli are assumed to be false under every admissible interpretation that preserves ordinary meanings and background constraints; if they are actually only pragmatically odd or lexically distinctive, the fitted impossibility direction captures a narrower property and the dissociation would not be about modality.
Editorial extensions
If this is right
- Verbally, the model collapses contingent falsehood into 'contradiction', so behavior alone would not reveal the internal dissociation; activation measurements do.
- Truth-probing results cannot be read as locating impossibility: a model can be accurate on factual truth while its truth direction is blind to modal status.
- The impossibility direction transfers across topic families and, partially, across the philosophical and modality stimulus sets, so it is not a lexical-template artifact.
- Necessary falsehood and semantic anomaly are related but separable in this geometry, so the model does not simply treat contradictions as meaningless.
- The study provides a template for asking whether larger models sharpen or dissolve the same separation, and whether the model ever uses this direction when it answers.
Reading between the lines
- If the near-orthogonality holds in larger models, then a model's tendency to call falsehoods 'contradictions' may be a shallow verbal policy layered on top of a more structured internal geometry; interventions aimed at truthfulness might be adjusting the wrong axis.
- The dissociation suggests that ordinary language itself contains enough distributional evidence for a learner to separate the contingent from the impossible without explicit logical supervision, since no labels of this kind were used in pretraining.
- A causal test the paper does not run: patching or steering along the impossibility direction should change answers about contradictions while leaving answers about factual lies unchanged; if it does, the direction is used, not merely decodable.
- A further extension would replace the handwritten necessary-falsehood stimuli with generated contradictions that share no lexical template, to test whether the direction tracks logical form rather than topic-specific wording.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper reports a probing study of Gemma 3 4B IT. Using 85 philosophical prompts and 75 topic-matched statements in five conditions (true, contingent false, improbable, anomalous, necessary false), the author trains linear probes on residual-stream activations. The main findings are that a truth probe does not separate impossible from false statements (AUC 0.20) while an impossibility probe does (AUC 1.00) with grouped held-out families, and that the truth and impossibility directions are nearly orthogonal, with impossibility closer to semantic anomaly. The SAE features at layer 15 are reported to repeat this geometry. The paper interprets this as evidence that necessary falsehoods are represented as distinct from contingent falsehoods and nearer to anomaly.
Significance. The design has real strengths: held-out-family cross-validation, a permutation test with Bonferroni correction, TF-IDF surface baselines, cross-dataset transfer, and public data and code. If the stimulus validation is repaired, the double dissociation would be a valuable empirical footnote connecting formal semantics and mechanistic interpretability. However, the central quantitative claims rest entirely on the validity of the hand-authored 'necessary falsehood' category, and the SAE feature-selection step is in-sample, so the mechanistic claim is weaker than the probe result.
major comments (2)
- [§4.1 (Stimuli), Table 4] The temporal-reversal item 'arrived before it departed' is not a necessary falsehood under ordinary English: a train can arrive at a station and then depart from it, making the sentence true. Since this item is one of the 15 impossibility stimuli, the impossibility probe's AUC 1.00 (Table 2) and the orthogonality cosine (Fig. 2c) are computed from a contaminated category. Please validate all 15 items with independent necessity judgments (or formal paraphrase), replace or remove invalid items, and rerun all analyses.
- [§4.4 (SAE analysis), Table 3] The SAE features are selected for selectivity using the same 160 prompts on which firing prevalence is then reported, so the claim that SAE features 'repeat the probe geometry' is partly guaranteed by construction. The manuscript acknowledges this in §4.4, but the Results section should either present a held-out feature-selection analysis (e.g., select on training families and report on held-out families) or explicitly frame Table 3 as descriptive in-sample data rather than confirmatory evidence.
minor comments (4)
- [§2.2, Table 2] The truth-probe AUC of 0.20 for impossible vs. false is reported without a confidence interval; with 15 examples per class, please add bootstrap CIs or a permutation test to confirm that this value is not a stable inversion of the expected ordering.
- [Figure 2c caption] The caption says 'cosine' while the text in §2.2 says 'absolute cosine'; please make the terminology consistent.
- [Table 1] The entry 'True141 0 0' is a formatting issue; add spaces between the counts for readability.
- [§4.4, Table 3] Reporting mean log activations in addition to firing prevalence would make the SAE effect size clearer and would strengthen the descriptive comparison.
Circularity Check
Central probe dissociation is held-out and externally checked; only a minor disclosed SAE same-sample selection loop is present.
-
other
[Methods 4.4, Sparse-autoencoder analysis; Table 3]
"Feature selection and description use the same 160 prompts. No independent corpus was used, and no inferential statistics are attached to individual features."
Features are ranked by the standardized difference of mean activation on the same 160 prompts, and Table 3 then reports firing prevalence by condition on those same prompts. The observed selectivity (impossible > false, anomalous intermediate) is therefore by construction a property of the selection criterion rather than an out-of-sample finding. The paper explicitly discloses this and attaches no inferential statistics, and the central probe dissociation does not depend on the SAE section. This is a minor, non-load-bearing same-sample selection loop, not a circular derivation of the main claim.
full rationale
The main derivation is self-contained and not circular. The impossibility, truth, and anomaly probes are evaluated with grouped cross-validation that holds out entire topic families, so the key results (impossibility vs. false AUC 1.00, truth probe at chance on impossible vs. false, cosine at most 0.12) are held-out transfer measurements rather than fits on the evaluated items. Cross-dataset transfer to the 85 philosophical prompts (AUC 0.72-0.79) and TF-IDF surface baselines (0.67-0.77) provide external checks that the separation is not a pure lexical artifact. No load-bearing self-citations appear; the cited truth-probing and SAE works are external tools or prior results. The one same-sample element is the SAE feature analysis in Methods 4.4, where features are selected and described on the same 160 prompts; the paper explicitly labels these as correlational candidates and attaches no inferential statistics, and the central double dissociation does not rely on that section. The possible inaccuracy of the 'arrived before it departed' stimulus is a construct-validity concern rather than a circularity, since the operational category definition is an input assumption rather than a derived conclusion. Overall the core claim has independent content, so circularity is minor.
Assumptions & free parameters
free parameters (3)
- Logistic regression regularization C =
0.1
- Peak layer for impossibility probe =
depth 16 (transformer layer 15)
- SAE checkpoint =
layer 15 width 16k l0 big
assumptions (4)
- domain assumption Residual stream states at the final prompt token carry information about the semantic categories used in the probes.
- ad hoc to paper The hand-written "necessary falsehood" stimuli are false under every admissible interpretation that preserves ordinary meanings and background constraints.
- domain assumption Topic-matched construction controls for topic vocabulary, so family-level holdout is a valid generalization test.
- domain assumption Fixed instruction and greedy generation do not materially alter the representational geometry.
Cite this review
Pith. "Pith review of Falsehood and Impossibility Are Different Directions in an AI's Representation of Language." pith.science (2026). https://pith.science/paper/L6UBQQ4P
@misc{pith2026260812852,
author = {Pith},
title = {Pith review of: Falsehood and Impossibility Are Different Directions in an AI's Representation of Language},
year = {2026},
howpublished = {\url{https://pith.science/paper/L6UBQQ4P}},
note = {Machine review of arXiv:2608.12852}
}
read the original abstract
Language can describe states of affairs that are false and states of affairs that could not be the case at all. Whether an AI model internally distinguishes these failures remains unclear. I report an exploratory activation study of the multimodal open-weight model Gemma 3 4B IT using 85 prompts from 17 philosophical families and a topic-matched modality set of 15 topics, each expressed as a truth, contingent falsehood, improbable claim, semantic anomaly, and necessary falsehood. In its answers, the model conflates contingent falsehood with contradiction, labeling 12 of 15 false statements "contradiction." Its activations show a different pattern. A linear truth probe separates impossible from true statements (AUC 0.93) but not impossible from false statements (AUC 0.20). An impossibility probe evaluated on held-out topic families separates necessary from contingent falsehood at AUC 1.00, peaking at layer 15 with balanced accuracy 0.97 (Bonferroni-adjusted P=0.018). The truth and impossibility directions are close to orthogonal, whereas the impossibility direction partially overlaps a semantic anomaly direction while remaining distinguishable from it. Sparse autoencoder features at the same layer repeat this geometry. Features selective for impossibility also fire on anomalous sentences but rarely on contingent falsehoods. In this model's activation space, necessary falsehoods are not extreme cases of contingent falsehood but lie closer to the experimentally defined category of semantic anomaly. This representational proximity does not imply that impossible statements are intrinsically meaningless. These correlational observations from one small model offer an empirical footnote to an old philosophical distinction.
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.