Pith. sign in

REVIEW 3 major objections 7 minor 23 references

The geometry a language model uses over concepts is set by the in-context rule, not by a fixed stored world-model.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-31 15:11 UTC pith:2LEOQJOM

load-bearing objection Declarative context really does rewrite concept geometry and topology in capable models; the causal claim is a bit looser than the abstract, but the core result still stands. the 3 major comments →

arxiv 2607.24425 v1 pith:2LEOQJOM submitted 2026-07-27 cs.LG

Context Is King: How In-Context Specification Shapes the Geometry of Concepts

classification cs.LG
keywords in-context specificationconcept geometryrepresentational similarity analysisactivation patchingtopology typescale gatinglarge language models
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Large language models are often read as storing fixed geometric world-models: weekdays on a circle, months on another, hierarchies as trees. This paper argues that what the model actually uses is whatever structure the prompt declares. A plain declarative rule can force the same tokens into a cycle or a branching tree, even when those tokens have no pretrained meaning, and when the rule conflicts with a strong prior the context-set geometry dominates in capable models. Entity-substitution patching shows the map is causal: overwrite one entity’s activation with another’s and the model answers with the donor’s successor under the imposed order. A rough map appears even in small models; clean dominance and the full causal crossover appear only at larger scale. Operationally, recovered concept geometry should be read as context-dependent, not intrinsic.

Core claim

In capable models from the Gemma and Qwen families, a declarative in-context specification sets both which relations and which topology type the model’s entity geometry encodes. On conflict with a pretrained prior, the same pre-generation activations align strongly with the imposed structure and near-zero with the prior; the map forms even on arbitrary tokens; and activation patching produces a double dissociation in which the answer follows whichever successor the context specifies.

What carries the argument

Entity-layout RSA plus entity-substitution activation patching: mean-centered pairwise dissimilarities among entity centroids in the last-token pre-generation residual are scored against a theory template (cycle, line, or depth), and overwriting one entity’s residual with another’s steers the answer to the donor’s successor under the context’s order.

Load-bearing premise

That the mean-centered layout of entity centroids in one late residual state is the relational map the next-token computation actually runs on, rather than a correlated byproduct of some other mechanism.

What would settle it

Find a capable model where, under a conflicting declarative order, last-token entity geometry still matches the pretrained ring (high natural RSA, low imposed RSA) and entity-substitution patches answer with the pretrained successor regardless of context.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • A recovered weekday circle or hierarchy should be treated as the geometry under that prompt, not as a fixed property of the weights.
  • Declarative rules can impose topology types (cycle vs tree) that no relabeling of a stored shape could produce, including on meaning-free tokens.
  • Clean override and causal use of the imposed map are scale-gated within a family, so small-model geometry findings need scale checks before transfer.
  • Mechanistic edits that swap entity activations steer multi-hop successor answers through the context-built map, not through a fixed lookup table.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If context routinely overwrites stored relational geometry, interpretability pipelines that mine ‘world models’ from default prompts may be cataloguing default contexts rather than stable internal knowledge.
  • Training or alignment methods that assume fixed concept manifolds may need to treat relational structure as a runtime computation conditioned on the prompt.
  • A natural next test is whether partial orders, DAGs, or native chain-of-thought traces show the same context-set topology switch and causal crossover.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 7 minor

Summary. The manuscript argues that the relational geometry LLMs use over concept sets (weekdays, months, clock hours, arbitrary tokens) is set by the in-context declarative specification rather than being a fixed structure stored in the weights. The main evidence: (i) under a conflicting redefined order, RSA of mean-centered entity centroids at a late pre-generation layer reaches ~0.6–0.9 to the imposed structure vs. near-zero to the pretrained prior, surviving relabeling, anisotropy-null, serial-position, and non-structural-query controls; (ii) the specification sets topology type — cycle vs. depth-stratified tree on the same tokens, confirmed by effective dimension and persistent homology; (iii) entity-substitution activation patching shows a context-dependent crossover: the patched answer follows the donor's successor under whichever order the context specifies (1.00/1.00 in Gemma-31B and Qwen-27B); (iv) forming a rough map is cheap (present in small and base models) but clean dominance and the full causal dissociation are scale-gated, emerging only at Gemma-12B/31B and Qwen-27B. The paper is unusually well-controlled for this subfield, ships code and data, and hedges several claims appropriately in §7.

Significance. If the results hold, this reframes a live debate: concept manifolds have been read as fixed, stored world-models (Engels et al. 2024; Park et al. 2024), and this manuscript shows that under declarative conflict the used geometry follows the context, across two model families, with a scale-gated causal signature. The topology-type result — cycle vs. tree on the same tokens, and structure built on prior-free tokens — is the sharpest evidence against a relabeled-store account and, to my knowledge, novel. Strengths that raise confidence: the claims are falsifiable and tested rather than asserted; the templates are theory-specified rather than fitted; code, cached data, and explorers are released for reproduction; and the authors disclose boundary conditions (capable-model-only dominance, Qwen wrap weakness, line case capability-dependent) rather than hiding them. The finding also carries a practical caution for the interpretability field — geometry recovered without controlling context is not intrinsic — and the scale-gating result (mechanism present at 27–31B, absent or reversed at 2–9B) is a concrete transfer warning for small-model circuit work.

major comments (3)
  1. [§6.1, Abstract] Abstract and contribution 2 state that activation patching 'shows the map is causally used, not a probe correlate.' The experiment patches the full residual at the entity position (h_Eorig ← h_Epatch, §6.1), which simultaneously carries entity identity, its binding to the declared list, and its location in the RSA-measured manifold (§5). The clean 1.00/1.00 crossover shows that context-conditioned donor information redirects the successor computation — but a context-conditioned symbolic successor function over an entity code predicts exactly the same crossover with no causal role for the pairwise geometry. The wrong-slot and random-vector controls exclude position and norm confounds but neither removes identity/binding while preserving geometry. §7 concedes this ('not the manifold axis as the causal carrier'), but the abstract, §1 contribution 2, and §6.1's 'a computation, not a lookup'
  2. [§6.2 / Appendix E, Table 4] Table 4 reports the Imp and Nat arms 'each at its strongest layer,' i.e., per-arm post-hoc layer maxima over the swept grid, on 2 scrambles × 42 pairs. Selecting the maximum per arm inflates both arms and the apparent double dissociation, especially at 2–9B where the 'clean locus' thresholds (0.9/0.2) are near-boundary (E4B Nat 0.48→'—', Qwen-9B 0.89/1.00). Figure 7 shows the full sweep only for Gemma-31B. The scale-onset claim ('clean dissociation appears only at 12B/31B/27B') is load-bearing for contribution 3; please report both arms at a common, a priori layer per model (the sweep exists), or show the full per-model sweep curves, and state sensitivity of the onset scale to the 0.9/0.2 thresholds. Two scrambles is also thin for a binary per-pair rate; CIs should accompany Table 4.
  3. [§3.1, §5.1 vs. Appendix C/D] The paper's central object is 'the geometry the model uses,' read pre-generation (§3.1). Yet in the direct regime, imposed-order accuracy collapses for k≥2 even in flagship models (Gemma-31B 0.57, Qwen-27B 0.44 overall; Fig. 10) while the pre-generation ring is clean (RSA ~0.8) and the RSA–behavior correlation in Appendix I uses only k=1. So the state that is claimed to be 'used' coexists with failure to use it beyond one hop. The 'built pre-generation but traversed a hop at a time' framing (§6.2) is plausible but post hoc: if traversal is generative hop-by-hop, the relevant 'used' geometry for k≥2 is re-built at each generated step, not the pre-generation map the probe reads. Please sharpen the definition of 'used' to specify which behaviors the pre-generation geometry actually licenses (adjacency reading, patch crossover), and qualify the §3.1/§5.1 claim accordingly.
minor comments (7)
  1. [§3.2, Appendix A] 'Arbitrary, structure-free tokens' (§3.2, Appendix A) is overstated: everyday nouns (Apple, Tiger, Bridge) carry semantic/category neighborhoods in any LLM, so 'no prior to inherit' holds only for the *imposed relation type*, not for the token geometry. The 12/12 sibling-placement control partially addresses this; the phrasing should be softened to match.
  2. [Appendix H] At N=7 the cyclic template takes only 3 distance values and the symmetry orbit gives ~15/2000 null draws at the observed RSA (Appendix H caveat). This is honestly handled, but the days column carries much of the headline evidence (Figs. 1, 3, 7); consider foregrounding the N=12 concepts (p<0.001) as primary and demoting days to corroborating.
  3. [Appendix F, Table 5] Table 5 (wrap vs. co-occurrence) uses 3 scrambles with no CIs, and Qwen-27B's dRSA stays negative even with the adjacent wrap (−0.08); the text says the adjacent wrap 'closes it fully' for the capable model — true only for Gemma. State the cross-family asymmetry at first mention (§5.2) rather than deferring it.
  4. [Appendix H, Control 1] Control 1 reports the imposed order 'ties for first' among 5040 orderings with 15 orderings sharing the maximum RSA; please list what the other 14 are (presumably the 14-element dihedral orbit of the cycle) to reassure the reader the ties are symmetry images, not competitors.
  5. [Appendix E] The identity→binding transition (Appendix E) rests on a wrong-slot control sampled at 'two live layers' with a full sweep 'pending.' This is an intriguing mechanistic claim that currently rests on incomplete data; either complete the sweep or move the claim to future work.
  6. [Figure 5] Figures 1, 5A, 6 show 2D PCA projections; Table 8 usefully shows 2D/3D/full RSA agree for days, but no analogous check is given for the depth-4 tree (31 nodes) whose 3D rendering is doing visual work in Fig. 5A. A full-space depth-RSA is reported (+0.73) — please state explicitly that Fig. 5A is illustration only.
  7. [Appendix J] The QWERTY grid result (Appendix J) rests on 6 scrambles and a weak, model-dependent native prior (Gemma 'weak'); the claim that formation-and-override extends to 2-D is fine as an existence proof but 'not tied to a single topology' is stronger than one additional topology supports.

Circularity Check

0 steps flagged

No significant circularity: RSA templates and successor targets are external theory objects, not quantities fitted from the same residuals and re-predicted.

full rationale

This is an empirical intervention paper, not a first-principles derivation that could collapse into its inputs. The load-bearing quantities—Spearman RSA of mean-centered entity-centroid RDMs against fixed cyclic-distance or depth templates, and behavioral/patched successors under a declared order—are scored against externally specified theory objects, not parameters estimated from the same activations and then re-labeled as predictions. Controls (label-shuffle null preserving the residual cone, relabeling crossover across 7! orderings, co-mention vs wrap, sibling shuffle, wrong-slot and random-vector patch controls) are designed to break trivial self-consistency rather than to smuggle the result in. Citations (Engels, Park, Gurnee, Li/Nanda Othello, Kriegeskorte RSA, etc.) are external prior art; there is no author-overlapping uniqueness theorem or ansatz citation that forces the central claim. The open gap noted by the skeptic—that full-residual entity substitution does not isolate the pairwise RSA manifold axis as the causal carrier—is a causal-identification / construct-validity issue, not circularity by construction. Operational claims remain independently falsifiable on held-out scrambles, models, and topology types.

Axiom & Free-Parameter Ledger

4 free parameters · 5 axioms · 2 invented entities

This is an empirical interpretability paper. It does not rest on novel physical entities or long formal derivations; it rests on standard residual-stream causal analysis plus several measurement conventions (centroid RSA, mean-centering, fixed fractional layer, declarative prompt scaffolds) and the modeling assumption that those measurements track the map used for next-token behavior. Free choices are analysis hyperparameters more than fitted scientific constants; invented constructs are operational (imposed specification, context-set map), not new ontological particles.

free parameters (4)
  • probe_layer_fraction = 0.75 n_layers
    Fixed at 0.75 n_layers for all models after a sweep showing a late plateau; not fit to maximize every claim, but still a hand-chosen readout depth the geometry scores depend on.
  • mean_centering_of_centroids = subtract mean centroid before cosine RDM
    Required to expose order under residual anisotropy; alternative normalizations (whitening/CKA) are discussed only as rationale, so results are conditional on this preprocessing choice.
  • scramble_and_pair_sample_sizes = 10 scrambles RSA; 2×42 causal pairs
    Geometry averaged over 10 scrambles (6 for hierarchy); causal arm uses 2 scrambles × 42 pairs. These budgets affect confidence intervals and claimed cleanliness of crossover.
  • patch_success_thresholds_for_clean_locus = Imp≥0.9, wrong-slot<0.2
    Clean locus defined with Imp≥0.9 and wrong-slot <0.2; near-threshold mid-ladder models make onset scale somewhat definition-dependent.
axioms (5)
  • domain assumption Representational similarity between mean-centered entity-centroid RDMs and a theory dissimilarity template tracks the relational geometry present in residual activations.
    Section 4 adopts neuroscience-style RSA as the primary geometry score; convergent PCA/behavior/PH checks support it but do not prove uniqueness.
  • domain assumption The last prompt token’s pre-generation residual is the integrated state consumed by next-token computation for these tasks.
    Stated in §3.1/§4; token-local day states can still show the pretrained order (Fig. 4), so this readout choice is load-bearing.
  • domain assumption Overwriting residual activations of entity tokens with another entity’s cached residual tests whether entity identity is causally read through the context-defined successor map.
    Standard activation-patching assumption (§6.1); wrong-slot/random-vector controls reduce but do not eliminate pathway ambiguity.
  • domain assumption Declarative natural-language rules with no exemplars constitute an in-context specification of relations among entities.
    Framework §3; surface form (list vs edges, wrap) is ablated, but still assumes the model parses the rule as intended relations.
  • standard math Spearman correlation, participation ratio, and related rank/topology summaries are valid under N≤12 centroids in high-dimensional residuals after mean-centering.
    Used throughout; small-N cyclic symmetry caveats are discussed in Appendix H.
invented entities (2)
  • context-set map / imposed geometry independent evidence
    purpose: Name the relational structure asserted to be constructed or selected by the prompt and used for behavior under conflict with pretrained priors.
    Operational object defined via RSA templates plus causal successor tests; not a new physical substrate beyond residual-stream computation.
  • identity→binding stage (mid-network) no independent evidence
    purpose: Localize where patched entity identity becomes bound into the relational map versus token-local identity.
    Inferred from wrong-slot vs entity-slot patch depth profiles in Appendix E; useful mechanistic sketch but coarsely resolved and pending fuller sweeps.

pith-pipeline@v1.2.0-grok45-kimik3 · 24852 in / 3963 out tokens · 88289 ms · 2026-07-31T15:11:19.747954+00:00 · methodology

0 comments
read the original abstract

Large language models place structured concepts on geometrically faithful manifolds: weekdays lie on a circle, months on another, usually taken to be a fixed world-model the network stores and looks up. We show that context is king: the structure a model actually uses is set by the in-context specification. A declarative rule fixes not only which relations the geometry encodes but its topology type: the same tokens form a cycle or a branching tree on command, built even on arbitrary, meaning-free tokens with no prior to inherit, which a relabeled stored shape cannot do. When the specification conflicts with a strong pretrained prior, the context-set geometry dominates it in capable models, read from the same activations (representational similarity 0.6--0.9 to the imposed structure versus near-zero to the prior), across the priors we test and both families we study (Gemma, Qwen). Activation patching shows the map is causally used, not a probe correlate: swapping one entity's activation for another's makes the model answer with the other entity's successor under the imposed order. A rough map forms readily, present even in small and base models; what scale gates is using it cleanly: clean dominance and the causal crossover emerge only in the larger models (up to Gemma-31B and Qwen-27B) and weaken or reverse below, so a mechanism present in a large model can be absent in a smaller one of the same family. Whether the model builds this geometry anew or reconfigures a stored one we leave open; operationally, the geometry it uses is the one the context specifies.

Figures

Figures reproduced from arXiv: 2607.24425 by Elad David, Max Fomin.

Figure 1
Figure 1. Figure 1: Entity centroids for three cyclic concepts under a conflicting in-context order (Gemma [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: What a specification is. The same order stated as a numbered list or as adjacency edges (left); adding a wrap edge closes the chain into a cycle (middle; the open, no-wrap case is a weaker baseline); a declarative hierarchy is a different structure type, read as depth layers (right). 3.1 DEFINITIONS In-context specification. A declarative rule, given with no exemplars, asserting the relations among a set o… view at source ↗
Figure 3
Figure 3. Figure 3: Imposed-order RSA (green) vs. residual natural-order RSA (gray) per model in the same [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Imposed- and natural-order RSA by readout position ( [PITH_FULL_IMAGE:figures/full_fig_p006_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Imposed depth-4 tree on arbitrary tokens under neutral queries (Gemma/Qwen). [PITH_FULL_IMAGE:figures/full_fig_p007_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Seven weekdays under three specifications (Gemma-31B; per-condition 2D PCA, one [PITH_FULL_IMAGE:figures/full_fig_p008_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: The context-set map is causally used (Gemma-31B, days; entity-substitution patch hEorig ← hEpatch , 42 ordered pairs × 2 scrambles). (a) Across depth, the patched answer follows the donor entity’s successor when the overwrite is applied at mid-network, under both the imposed and the natural order; the wrong-slot control (writing the entity vector at a non-entity slot) leaks only at the earliest layers and … view at source ↗
Figure 8
Figure 8. Figure 8: Full-layer sweep (days): imposed-order RSA (solid) and residual natural-order RSA (dot [PITH_FULL_IMAGE:figures/full_fig_p013_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: Imposed ring recovered from each hop value [PITH_FULL_IMAGE:figures/full_fig_p014_9.png] view at source ↗
Figure 10
Figure 10. Figure 10: Accuracy by hop count k=1–6 per model in the direct (no-reasoning) regime (days, list, bare ≤3-token answers, two scrambles): natural order (dashed, per-k ceiling) vs. imposed order (solid), with the 1/7 chance line. One-shot accuracy degrades with k even for the capable models; the reasoning regimes recover ceiling ( [PITH_FULL_IMAGE:figures/full_fig_p014_10.png] view at source ↗
Figure 11
Figure 11. Figure 11: Imposed-order accuracy by hop count under three prompt regimes (days; [PITH_FULL_IMAGE:figures/full_fig_p015_11.png] view at source ↗
Figure 12
Figure 12. Figure 12: Imposed depth-4 tree across 36 query types (Gemma-31B). [PITH_FULL_IMAGE:figures/full_fig_p018_12.png] view at source ↗
Figure 13
Figure 13. Figure 13: Left: form × wrap 2×2 for Gemma-31B (PC1–PC2 of entity centroids, colored by imposed position, wrap edge red); per cell, closure ≈1 / dRSA ≥ 0 when the wrap is stated, closure ≫1 / dRSA < 0 when open. Right: the wrap-pair closure ratio for three models, open vs. wrap￾stated. entity centroids share a dominant direction, so raw cosines are high and barely distinguish the order￾ings (imposed-adjacent pairs a… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

23 extracted references · 16 linked inside Pith

  1. [1]

    and Obeso, O

    Arditi, A. and Obeso, O. and Syed, A. and Paleka, D. and Panickssery, N. and Gurnee, W. and Nanda, N. , title =. arXiv preprint arXiv:2406.11717 , year =

  2. [2]

    and Goel, S

    Diaconis, P. and Goel, S. and Holmes, S. , title =. arXiv preprint arXiv:0811.1477 , year =

  3. [3]

    , title =

    Ethayarajh, K. , title =. arXiv preprint arXiv:1909.00512 , year =

  4. [4]

    and Cope, D

    Jorgensen, O. and Cope, D. and Schoots, N. and Shanahan, M. , title =. arXiv preprint arXiv:2312.03813 , year =

  5. [5]

    and Kriegeskorte, N

    Lin, B. and Kriegeskorte, N. , title =. Proceedings of the National Academy of Sciences , volume =. 2024 , note =

  6. [6]

    and others , title =

    Wurgaft, D. and others , title =. arXiv preprint arXiv:2605.05115 , year =

  7. [7]

    and Michaud, E

    Engels, J. and Michaud, E. J. and Liao, I. and Gurnee, W. and Tegmark, M. , title =. arXiv preprint arXiv:2405.14860 , year =

  8. [8]

    and Tegmark, M

    Gurnee, W. and Tegmark, M. , title =. arXiv preprint arXiv:2310.02207 , year =

  9. [9]

    and Mur, M

    Kriegeskorte, N. and Mur, M. and Bandettini, P. , title =. Frontiers in Systems Neuroscience , volume =

  10. [10]

    and Hopkins, A

    Li, K. and Hopkins, A. K. and Bau, D. and Vi. Emergent world representations: exploring a sequence model trained on a synthetic task , journal =

  11. [11]

    and Rahtz, M

    Lieberum, T. and Rahtz, M. and Kram. Does circuit analysis interpretability scale? Evidence from multiple choice capabilities in. arXiv preprint arXiv:2307.09458 , year =

  12. [12]

    Lepori, M. A. and Linzen, T. and Yuan, A. and Filippova, K. , title =. arXiv preprint arXiv:2602.04212 , year =

  13. [13]

    and Bau, D

    Meng, K. and Bau, D. and Andonian, A. and Belinkov, Y. , title =. arXiv preprint arXiv:2202.05262 , year =

  14. [14]

    and Lee, A

    Nanda, N. and Lee, A. and Wattenberg, M. , title =. arXiv preprint arXiv:2309.00941 , year =

  15. [15]

    Park, C. F. and Lee, A. and Lubana, E. S. and Yang, Y. and Okawa, M. and Nishi, K. and Wattenberg, M. and Tanaka, H. , title =. arXiv preprint arXiv:2501.00070 , year =

  16. [16]

    and Choe, Y

    Park, K. and Choe, Y. J. and Jiang, Y. and Veitch, V. , title =. arXiv preprint arXiv:2406.01506 , year =

  17. [17]

    and Miranda, B

    Schaeffer, R. and Miranda, B. and Koyejo, S. , title =. arXiv preprint arXiv:2304.15004 , year =

  18. [18]

    , title =

    Kim, C. , title =. arXiv preprint arXiv:2512.17325 , year =

  19. [19]

    and Tay, Y

    Wei, J. and Tay, Y. and Bommasani, R. and Raffel, C. and Zoph, B. and Borgeaud, S. and Yogatama, D. and Bosma, M. and Zhou, D. and Metzler, D. and Chi, E. H. and Hashimoto, T. and Vinyals, O. and Liang, P. and Dean, J. and Fedus, W. , title =. arXiv preprint arXiv:2206.07682 , year =

  20. [20]

    and Ji-An, L

    Xiong, H.-D. and Ji-An, L. and Wilson, R. C. and Lee, K. and Wei, X.-X. , title =. arXiv preprint arXiv:2605.28854 , year =

  21. [21]

    Hosseini, E. A. and Li, Y. and Bahri, Y. and Campbell, D. and Lampinen, A. K. , title =. arXiv preprint arXiv:2601.22364 , year =

  22. [22]

    , title =

    Carlsson, G. , title =. Bulletin of the American Mathematical Society , volume =

  23. [23]

    and Morozov, D

    de Silva, V. and Morozov, D. and Vejdemo-Johansson, M. , title =. Discrete & Computational Geometry , volume =