Pith. sign in

REVIEW 2 major objections 4 minor

Falsehood and Impossibility Are Different Directions in an AI's Representation of Language

T0 review · 2 major / 4 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read A language model's internal geometry treats impossibility as a separate direction from falsehood.

desk verdict A careful small-model probe study showing truth and impossibility as near-orthogonal directions, but the 'impossible' category contains at least one genuinely true item and the below-chance AUC is misread as chance. read the letter →

arxiv 2608.12852 v2 pith:L6UBQQ4P submitted 2026-08-13 cs.CL cs.AI

classification cs.CLcs.AI
keywords mechanisticinterpretabilitycontradictionimpossibilitytruthprobingsparseautoencodersphilosophyoflanguagesemanticanomaly
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks whether a large language model distinguishes statements that are merely false from statements that could not be true. Using 160 prompts built from 17 philosophical families and 15 topic-matched conditions, the author finds a split between the model's words and its internal states: the model verbally labels ordinary falsehoods as 'contradiction', but its activation geometry keeps the two apart. A linear probe trained to detect impossibility separates necessary from contingent falsehood at perfect held-out AUC, while the direction that separates true from false is at chance on that same contrast. The central claim is that necessary falsehoods are not represented as extreme contingent falsehoods; they sit closer to semantic anomaly while remaining distinguishable from it.

What carries the argument

The load-bearing object is the linear probe readout on residual-stream states, together with the angular comparison of probe weight vectors. Each probe is a regularized logistic regression fitted on grouped, family-held-out folds; the impossibility probe contrasts impossible statements against true, false, and improbable ones, while the truth probe contrasts false against true and the anomaly probe contrasts anomalous against the three possible conditions. The double dissociation—the truth probe cannot order impossible against false, while the impossibility probe can—and the near-orthogonality of the fitted directions carry the argument. A pretrained sparse autoencoder at the impossibility peak (layer 15) supplies a feature-level confirmation.

What would settle it

Train the same impossibility probe on a new set of 'necessary falsehoods' built with completely different surface forms (for example, generated by conjoining each proposition with its negation in varied vocabularies) while holding topic vocabulary matched across conditions; if probe AUC on held-out families drops substantially, the direction is capturing lexical templates rather than impossibility. Alternatively, replace the necessary-falsehood items with pragmatically odd but logically possible sentences; if the probe still separates them from contingent falsehood at AUC 1.00, the direction encodes oddity, not modality.

Watch

Extended reading notes

Core claim

The paper's central claim is that in the residual stream of Gemma 3 4B IT, impossibility and falsehood are carried by nearly orthogonal linear directions. The truth direction, which separates true from false statements with balanced accuracy 0.93, fails to separate impossible from merely false statements (AUC 0.20); an impossibility probe fitted on held-out topic families does separate them at AUC 1.00. The absolute cosine between the two directions stays at or below 0.12 beyond depth 10. Sparse autoencoder features at layer 15 repeat this geometry: features selective for impossibility fire on anomalous sentences but rarely on contingent falsehoods. The author states the result as a representational observation, not a claim that impossible statements are intrinsically meaningless.

Load-bearing premise

The handwritten 'necessary falsehood' stimuli are assumed to be false under every admissible interpretation that preserves ordinary meanings and background constraints; if they are actually only pragmatically odd or lexically distinctive, the fitted impossibility direction captures a narrower property and the dissociation would not be about modality.

Editorial extensions

If this is right

  • Verbally, the model collapses contingent falsehood into 'contradiction', so behavior alone would not reveal the internal dissociation; activation measurements do.
  • Truth-probing results cannot be read as locating impossibility: a model can be accurate on factual truth while its truth direction is blind to modal status.
  • The impossibility direction transfers across topic families and, partially, across the philosophical and modality stimulus sets, so it is not a lexical-template artifact.
  • Necessary falsehood and semantic anomaly are related but separable in this geometry, so the model does not simply treat contradictions as meaningless.
  • The study provides a template for asking whether larger models sharpen or dissolve the same separation, and whether the model ever uses this direction when it answers.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the near-orthogonality holds in larger models, then a model's tendency to call falsehoods 'contradictions' may be a shallow verbal policy layered on top of a more structured internal geometry; interventions aimed at truthfulness might be adjusting the wrong axis.
  • The dissociation suggests that ordinary language itself contains enough distributional evidence for a learner to separate the contingent from the impossible without explicit logical supervision, since no labels of this kind were used in pretraining.
  • A causal test the paper does not run: patching or steering along the impossibility direction should change answers about contradictions while leaving answers about factual lies unchanged; if it does, the direction is used, not merely decodable.
  • A further extension would replace the handwritten necessary-falsehood stimuli with generated contradictions that share no lexical template, to test whether the direction tracks logical form rather than topic-specific wording.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 4 minor

Summary. This paper reports a probing study of Gemma 3 4B IT. Using 85 philosophical prompts and 75 topic-matched statements in five conditions (true, contingent false, improbable, anomalous, necessary false), the author trains linear probes on residual-stream activations. The main findings are that a truth probe does not separate impossible from false statements (AUC 0.20) while an impossibility probe does (AUC 1.00) with grouped held-out families, and that the truth and impossibility directions are nearly orthogonal, with impossibility closer to semantic anomaly. The SAE features at layer 15 are reported to repeat this geometry. The paper interprets this as evidence that necessary falsehoods are represented as distinct from contingent falsehoods and nearer to anomaly.

Significance. The design has real strengths: held-out-family cross-validation, a permutation test with Bonferroni correction, TF-IDF surface baselines, cross-dataset transfer, and public data and code. If the stimulus validation is repaired, the double dissociation would be a valuable empirical footnote connecting formal semantics and mechanistic interpretability. However, the central quantitative claims rest entirely on the validity of the hand-authored 'necessary falsehood' category, and the SAE feature-selection step is in-sample, so the mechanistic claim is weaker than the probe result.

major comments (2)
  1. [§4.1 (Stimuli), Table 4] The temporal-reversal item 'arrived before it departed' is not a necessary falsehood under ordinary English: a train can arrive at a station and then depart from it, making the sentence true. Since this item is one of the 15 impossibility stimuli, the impossibility probe's AUC 1.00 (Table 2) and the orthogonality cosine (Fig. 2c) are computed from a contaminated category. Please validate all 15 items with independent necessity judgments (or formal paraphrase), replace or remove invalid items, and rerun all analyses.
  2. [§4.4 (SAE analysis), Table 3] The SAE features are selected for selectivity using the same 160 prompts on which firing prevalence is then reported, so the claim that SAE features 'repeat the probe geometry' is partly guaranteed by construction. The manuscript acknowledges this in §4.4, but the Results section should either present a held-out feature-selection analysis (e.g., select on training families and report on held-out families) or explicitly frame Table 3 as descriptive in-sample data rather than confirmatory evidence.
minor comments (4)
  1. [§2.2, Table 2] The truth-probe AUC of 0.20 for impossible vs. false is reported without a confidence interval; with 15 examples per class, please add bootstrap CIs or a permutation test to confirm that this value is not a stable inversion of the expected ordering.
  2. [Figure 2c caption] The caption says 'cosine' while the text in §2.2 says 'absolute cosine'; please make the terminology consistent.
  3. [Table 1] The entry 'True141 0 0' is a formatting issue; add spaces between the counts for readability.
  4. [§4.4, Table 3] Reporting mean log activations in addition to firing prevalence would make the SAE effect size clearer and would strengthen the descriptive comparison.

Circularity Check

1 steps flagged · score 1.0 of 10

Central probe dissociation is held-out and externally checked; only a minor disclosed SAE same-sample selection loop is present.

  1. other [Methods 4.4, Sparse-autoencoder analysis; Table 3]
    "Feature selection and description use the same 160 prompts. No independent corpus was used, and no inferential statistics are attached to individual features."

    Features are ranked by the standardized difference of mean activation on the same 160 prompts, and Table 3 then reports firing prevalence by condition on those same prompts. The observed selectivity (impossible > false, anomalous intermediate) is therefore by construction a property of the selection criterion rather than an out-of-sample finding. The paper explicitly discloses this and attaches no inferential statistics, and the central probe dissociation does not depend on the SAE section. This is a minor, non-load-bearing same-sample selection loop, not a circular derivation of the main claim.

full rationale

The main derivation is self-contained and not circular. The impossibility, truth, and anomaly probes are evaluated with grouped cross-validation that holds out entire topic families, so the key results (impossibility vs. false AUC 1.00, truth probe at chance on impossible vs. false, cosine at most 0.12) are held-out transfer measurements rather than fits on the evaluated items. Cross-dataset transfer to the 85 philosophical prompts (AUC 0.72-0.79) and TF-IDF surface baselines (0.67-0.77) provide external checks that the separation is not a pure lexical artifact. No load-bearing self-citations appear; the cited truth-probing and SAE works are external tools or prior results. The one same-sample element is the SAE feature analysis in Methods 4.4, where features are selected and described on the same 160 prompts; the paper explicitly labels these as correlational candidates and attaches no inferential statistics, and the central double dissociation does not rely on that section. The possible inaccuracy of the 'arrived before it departed' stimulus is a construct-validity concern rather than a circularity, since the operational category definition is an input assumption rather than a derived conclusion. Overall the core claim has independent content, so circularity is minor.

Assumptions & free parameters 3 free parameters · 4 assumptions · 0 invented entities

The main free choices are the probe regularization strength, the selected peak layer, and the SAE checkpoint, all of which are conventional or data-driven rather than fitted to a quantitative theory. The load-bearing assumptions are that final-token residual states are informative for these categories and that the hand-built "necessary falsehood" stimuli really are impossible. No new theoretical entities are postulated.

free parameters (3)
  • Logistic regression regularization C = 0.1
    Used for all probes (Methods 4.3); chosen by hand, not tuned per layer or contrast. It is a conventional value, but it is a free choice that affects decision values and the fitted directions.
  • Peak layer for impossibility probe = depth 16 (transformer layer 15)
    Selected as the maximum of the depth profile after looking at the data; Bonferroni correction for 35 depths is applied, but the choice remains data-driven.
  • SAE checkpoint = layer 15 width 16k l0 big
    Chosen because it is the impossibility probe peak; the sparser checkpoint showed no selective features. This is a post-hoc selection on the same data.
assumptions (4)
  • domain assumption Residual stream states at the final prompt token carry information about the semantic categories used in the probes.
    The entire probing methodology (Methods 4.2-4.3) presupposes that a linear readout of the final-token residual state is a meaningful window into how the model represents the sentence. Probes establish decodability, not causal use, as the paper itself notes.
  • ad hoc to paper The hand-written "necessary falsehood" stimuli are false under every admissible interpretation that preserves ordinary meanings and background constraints.
    Methods 4.1 defines necessary falsehood operationally; the label is a human judgment. If these stimuli are not actually impossible, the central dissociation is about something else.
  • domain assumption Topic-matched construction controls for topic vocabulary, so family-level holdout is a valid generalization test.
    The five conditions in each family share topic words, but the author acknowledges TF-IDF baselines of 0.67-0.77 capture part of the signal; the assumption is that the remaining signal is semantic.
  • domain assumption Fixed instruction and greedy generation do not materially alter the representational geometry.
    All prompts share one template (Methods 4.2); whether the finding extends to other prompts is unknown, and the paper limits its claim accordingly.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Falsehood and Impossibility Are Different Directions in an AI's Representation of Language." pith.science (2026). https://pith.science/paper/L6UBQQ4P

@misc{pith2026260812852,
  author       = {Pith},
  title        = {Pith review of: Falsehood and Impossibility Are Different Directions in an AI's Representation of Language},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/L6UBQQ4P}},
  note         = {Machine review of arXiv:2608.12852}
}
read the original abstract

Language can describe states of affairs that are false and states of affairs that could not be the case at all. Whether an AI model internally distinguishes these failures remains unclear. I report an exploratory activation study of the multimodal open-weight model Gemma 3 4B IT using 85 prompts from 17 philosophical families and a topic-matched modality set of 15 topics, each expressed as a truth, contingent falsehood, improbable claim, semantic anomaly, and necessary falsehood. In its answers, the model conflates contingent falsehood with contradiction, labeling 12 of 15 false statements "contradiction." Its activations show a different pattern. A linear truth probe separates impossible from true statements (AUC 0.93) but not impossible from false statements (AUC 0.20). An impossibility probe evaluated on held-out topic families separates necessary from contingent falsehood at AUC 1.00, peaking at layer 15 with balanced accuracy 0.97 (Bonferroni-adjusted P=0.018). The truth and impossibility directions are close to orthogonal, whereas the impossibility direction partially overlaps a semantic anomaly direction while remaining distinguishable from it. Sparse autoencoder features at the same layer repeat this geometry. Features selective for impossibility also fire on anomalous sentences but rarely on contingent falsehoods. In this model's activation space, necessary falsehoods are not extreme cases of contingent falsehood but lie closer to the experimentally defined category of semantic anomaly. This representational proximity does not imply that impossible statements are intrinsically meaningless. These correlational observations from one small model offer an empirical footnote to an old philosophical distinction.

Discussion (0). Continue with ORCID to comment.

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.