REVIEW 3 major objections 6 minor 8 cited by
The Geometry of Tokens in Internal Representations of Large Language Models
T0 review · 3 major / 6 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read Token geometry carries the signature of prediction difficulty: intrinsic dimension correlates with next-token loss.
desk verdict A solid, honest token-level study of how intrinsic dimension tracks next-token loss, but the ID estimator is unvalidated at this scale, so the headline correlation needs an independent sanity check before I'd fully trust it. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the empirical measure of the token cloud at each layer: the probability distribution that puts equal mass on every token's position in the residual stream. To probe that measure the paper relies on the GRIDE intrinsic-dimension estimator, a likelihood-based nearest-neighbor method whose range-scaling-2 special case is the TWO-NN estimator, which converts the ratios of second-to-first nearest-neighbor distances into a local dimension estimate. That estimate carries the argument: the intrinsic-dimension profile across layers, the higher peak for shuffled prompts, and the correlation between intrinsic dimension and loss are all computed with it. Neighborhood overlap between adjacent layers and cosine similarity serve as complementary probes of how coherent token neighborhoods are and how aligned the token vectors become, and a chain through logits and softmax entropy connects the geometric quantity to the information-theoretic loss.
What would settle it
Run the same GRIDE/TWO-NN estimator on synthetic point clouds of known intrinsic dimension, using 1024 points embedded in 4096 dimensions with realistic anisotropy, correlated directions, and nonuniform density; if the estimate systematically misses the true dimension in this regime, the quantitative claims lose their support. A second check is to compare prompts matched for cross-entropy loss but with different syntactic or semantic structure: if their token-level intrinsic-dimension profiles differ substantially despite equal loss, then the intrinsic-dimension-loss correlation is not a stable signature of prediction difficulty.
Extended reading notes
Core claim
On the paper's own account, the central discovery is stated in Section 5: the intrinsic dimension of token representations across hidden layers is correlated with the average cross-entropy loss of the next-token probability distribution for a given prompt. Across the three decoder-only models studied, the Pearson correlation between $\log(\mathrm{ID})$ and loss is positive and significant, especially around the early-to-middle-layer intrinsic-dimension peak. The paper also establishes a layer-by-layer chain: the last-layer token representation is linearly unembedded into logits, the intrinsic dimension of the logits tracks the intrinsic dimension of the last layer ($\rho = 0.96$), the logits' intrinsic dimension correlates with the contextual entropy of the softmax output ($\rho = 0.43$ for one of the models), and the contextual entropy, averaged over a long prompt, is nearly the cross-entropy loss. Toy calculations with logits supported on a $D_M$-dimensional unit box or probability simplex give $\langle S\rangle \sim \log D_M$, suggesting that softmax entropy grows logarithmically with the intrinsic dimension of the logit manifold.
Load-bearing premise
The load-bearing premise is that the GRIDE/TWO-NN intrinsic-dimension estimate, computed from 1024 tokens in a 4096-dimensional residual stream, genuinely measures the local manifold dimension of the token representations; the estimator assumes locally uniform density and independent nearest-neighbor ratios, and if those assumptions fail in this regime, the intrinsic-dimension peak, the shuffle contrasts, and the intrinsic-dimension-loss correlation could be artifacts of the estimator rather than properties of the representations.
Editorial extensions
If this is right
- If the central claim is right, intrinsic dimension can serve as an unsupervised metric for evaluating model performance: prompts with higher loss are represented in higher-dimensional manifolds, and the correlation holds across three different model families.
- The shuffle experiments imply that natural syntactic and semantic structure compresses token representations: disrupting that structure raises the intrinsic-dimension peak, increases cosine alignment among tokens, and lowers neighborhood overlap around the peak.
- Because last-layer intrinsic dimension correlates with logit intrinsic dimension ($\rho = 0.96$), the correlation can be read off the residual stream without evaluating the full softmax distribution.
- The toy-model relation $\langle S\rangle \sim \log D_M$ suggests the softmax entropy is bounded by the intrinsic dimension of the logit manifold, giving a concrete geometric mechanism for the loss correlation.
Reading between the lines
- A natural extension the paper leaves implicit is a per-token version of the correlation: if prompt-level loss is tracked by token-cloud dimensionality, then per-token difficulty maps could be read off hidden states alone, flagging high-surprisal tokens without computing the output distribution.
- The shuffling contrast suggests that the geometry of the empirical measure acts as a structure detector that should transfer to non-text token streams, such as protein or image-patch sequences, and could flag distribution shift or out-of-domain inputs.
- The $\langle S\rangle \sim \log D_M$ calculation hints that constraining the logit manifold to a lower-dimensional subspace would reduce predictive entropy, a testable prediction about confidence and calibration that the paper does not make.
- Because the correlation is strongest near the early-to-middle-layer intrinsic-dimension peak, the layer at which the peak occurs could serve as a diagnostic of where a model commits to its prediction, connecting this work to layerwise latent-prediction analyses.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript studies token-level geometry of internal representations in three decoder-only LLMs (Llama 3 8B, Mistral 7B, Pythia 6.9B) through the lens of empirical measures. It computes intrinsic dimension (GRIDE/TWO-NN), neighborhood overlap, and cosine similarity across layers for 2244 Pile-10K prompts of 1024 tokens, compares structured prompts with block-shuffled versions, and reports a layerwise Pearson correlation between log intrinsic dimension and average next-token cross-entropy loss. Section 5.1 proposes an explanatory chain from last-layer ID to logits ID to contextual entropy to loss, supported by a toy softmax model. The abstract and conclusions additionally suggest that ID could be a metric for evaluating model performance across models.
Significance. If the central correlation is real, the paper makes a meaningful empirical contribution: it extends prior prompt-level studies (e.g., Cheng et al., 2023, 2024) to token-level point clouds inside a prompt and shows a consistent layerwise association across three independently trained models, with p-values mostly below 0.01. The analysis is direct rather than circular: the main quantity is a measured correlation, not a fitted parameter, and the authors are careful to label the softmax-box and Dirichlet calculations as toy examples. The scale analysis in Appendix C and the comparison with ESS-based prompt-level correlations in Table 1 are useful consistency checks. The reproducibility statement gives a code repository. However, the result is currently gated by two issues: the ID estimator is not validated in the specific regime used, and the pooled correlation is not shown to be robust to the Pile-10K source-domain structure. These are the main reasons the central claim is not yet secured.
major comments (3)
- [Section 3, Eq. (2)-(3)] The load-bearing premise that GRIDE/TWO-NN estimates the true local manifold dimension in this regime is not validated. The distribution in Eq. (2) requires local uniform density and independence of the neighbor ratios across points, and the paper applies it to N=1024 points in d=4096 without any synthetic test at this sample size and ambient dimension. The footnote asserting that local homogeneity 'is generally true' is not evidence. Because Fig. 6 correlates the log of this ID estimate with loss, an estimator artifact—say, sensitivity to distance concentration or to non-uniform density at second-neighbor scale in high dimension—could produce the reported layerwise correlations even if the true geometric dimension is unrelated to loss. The shuffling contrasts do not resolve this, since shuffling changes the point-cloud distribution in ways that could alter the estimator's bias. Required: validate GRIDE on synthetic manifolds of known dimension embedded in 4096 dimensions with N=1024, including non-uniform and anisotropic densities, and/or reproduce the ID profiles and correlations with an independent estimator (e.g., ESS, correlation dimension, or local PCA).
- [Section 5.1, Eqs. (11)-(15)] The toy model connects softmax entropy to the number of active logits D_M, not to the GRIDE estimate of the intrinsic dimension of the logit point cloud. The identification between the estimated ID and the parameter D_M is not established; in the empirical analysis, ID is estimated from a point cloud of logits in a vocabulary-sized ambient space, whereas the toy model presumes an explicit coordinate box of dimension D_M. The reported rho=0.43 between log ID and contextual entropy is therefore consistent with mechanisms other than the toy model. A concrete test is to generate synthetic logit clouds with known D_M at the relevant ambient dimension and sample size, and verify that GRIDE recovers D_M; alternatively, the text should explicitly state that the toy model is only an analogy and not a derivation of the observed correlation.
- [Section 5, Fig. 6] The Pearson correlation is computed on 2244 prompts pooled over the 22 source domains of Pile-10K. If both the average loss and the ID estimates vary systematically by source, the pooled correlation can be inflated by a domain-level confound, and the statement that 'prompts with a higher cross-entropy loss have token representations lying in higher dimensional manifolds' would not follow within domains. Please report source-stratified correlations or a partial correlation controlling for the Pile source, and show that the correlation holds for at least the most frequent sources.
minor comments (6)
- [Section 5.1, item 1] The Pearson coefficient rho=0.96 between the log ID of the last layer and the log ID of the logits is stated without a scatter plot, confidence interval, or model breakdown; please provide the supporting figure or at least the per-model values.
- [Section B.2] The text says Pythia has a lower ID peak than the other models 'though the significance is low'; because this is an explicit statement of low significance, either provide a formal significance test (e.g., a permutation test over prompts) or refrain from treating the difference as a model-level property.
- [Figure 7 caption] The sentence 'analysis of the correlation between the logits ID at scaling = 2 ... and the contextual entropy to the average contextual entropy' is grammatically garbled; please state explicitly which quantity is plotted on each axis in each panel.
- [Section 4.2] The variables x_{i,k} and r_{i,k} are used in the 'Distribution of tokens at the ID peak' paragraph but defined only in the footnote of the following paragraph; move the definition to the first occurrence.
- [Reproducibility section] The repository URL 'https://github.com/RitAreaSciencePark/token geometry' contains a space and could not be resolved during review; please ensure the link is valid and includes a README with the exact model revisions, filtering code, and GRIDE implementation needed to reproduce all figures.
- [Section 6] The claim that ID 'could be an important metric for evaluating model performance across different models' goes beyond the within-model, across-prompt correlation reported in Fig. 6; either add a cross-model analysis (e.g., relating peak ID to per-model loss while controlling for model identity) or soften the conclusion to the within-model setting actually tested.
Circularity Check
No significant circularity: the central ID–loss correlation is a directly measured empirical relation, and the explanatory chain is supported by independent measurements and an explicitly labeled toy model.
full rationale
The paper's central claim (Section 5) is a measured Pearson correlation between the GRIDE/TWO-NN intrinsic dimension of token representations and average next-token cross-entropy loss across 2244 prompts. This quantity is computed directly from the residual-stream point clouds and from the model's next-token probabilities; no parameter is fitted to the loss and then reported as a prediction, and no equation defines ID in terms of loss or vice versa. The explanatory chain in Section 5.1 uses empirically measured correlations (last-layer ID vs. logits ID, ρ=0.96; logits ID vs. contextual entropy, ρ=0.43) and a clearly labeled toy model relating softmax entropy to the dimension of a unit box or simplex. The toy model is an independent mathematical example, not an inference from the data, and it is not used to manufacture the main correlation. The observation that higher ID corresponds to nearest-neighbor ratios closer to unity follows directly from Equation (3), but it is presented as an interpretation of the estimator, not as a load-bearing derivation. Self-citations in the related-work section (e.g., [18], [21], [22]) motivate the geometric approach but do not supply the central result, which is internally computed and benchmarked against external models and prior prompt-level findings. The concern that the TWO-NN estimator may be unreliable at N=1024 in a 4096-dimensional residual stream is a correctness or validity risk, not a circularity: even if the estimator were biased, the correlation would still be between two independently computed empirical quantities. Therefore the derivation chain is self-contained and no circular step can be exhibited.
Assumptions & free parameters
free parameters (1)
- GRIDE range scaling n2/n1 =
2 (variants 4, 8)
assumptions (4)
- domain assumption GRIDE/TWO-NN estimator assumptions: local uniform density and independent nearest-neighbor ratios
- domain assumption Mean-field interaction picture: token dynamics depend on the current token representation and empirical measure, not on token labels
- ad hoc to paper Linear unembedding approximately preserves intrinsic dimension between last-layer representations and logits
- ad hoc to paper Toy logits model: D_M active entries uniform in [0,1] and remaining entries at -infinity
Cite this review
Pith. "Pith review of The Geometry of Tokens in Internal Representations of Large Language Models." pith.science (2026). https://pith.science/paper/N42RQLVK
@misc{pith2026250110573,
author = {Pith},
title = {Pith review of: The Geometry of Tokens in Internal Representations of Large Language Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/N42RQLVK}},
note = {Machine review of arXiv:2501.10573}
}
read the original abstract
We investigate the relationship between the geometry of token embeddings and their role in the next token prediction within transformer models. An important aspect of this connection uses the notion of empirical measure, which encodes the distribution of token point clouds across transformer layers and drives the evolution of token representations in the mean-field interacting picture. We use metrics such as intrinsic dimension, neighborhood overlap, and cosine similarity to observationally probe these empirical measures across layers. To validate our approach, we compare these metrics to a dataset where the tokens are shuffled, which disrupts the syntactic and semantic structure. Our findings reveal a correlation between the geometric properties of token embeddings and the cross-entropy loss of next token predictions, implying that prompts with higher loss values have tokens represented in higher-dimensional spaces.
Figures
Figures from the paper (18 more)
Forward citations
Cited by 8 Pith papers
-
Metaphor Tracer: A Theory-Informed Analysis of Hidden States
Hidden-state aggregator and differentiator scores, frozen on one text, track within-text organization across models and align with engineered registers and psychoanalytic marks while dissociating from information and ...
-
Attention's forward pass and Frank-Wolfe
Hardmax self-attention is shown to be a Frank-Wolfe iteration; with positive-definite key-query it converges to Voronoi-cell vertices, and a Markov-chain version of soft attention is metastable there for exponential-i...
-
Geometric Configurations of Perturbed Jailbreak Prompts
In six small open-weight LLMs, jailbreak prompts are linearly separable in last-token embeddings by surface form, but not by refusal/compliance behavior.
-
An Analysis of Residual-Stream Geometry Across Transformer Depth
Across six instruction-tuned transformers, residual-stream layer transitions follow a model-specific, condition-stable depth curve: large early and late updates, a quiet middle, near-flat rotation, and a rising final ...
-
Geometric Metrics and LLMs: What They Measure and When They Work
The paper's abstract claims that Schatten Norm and MOM reflect output length and that geometric features add modest classifier accuracy over text statistics, but the body instead reports consistent generator rankings ...
-
What's in a prompt? Language models encode literary style in prompt embeddings
Deep-layer embeddings of short literary excerpts carry enough information to identify their source book and author, with same-author works more confused, indicating style is encoded in the prompt representation.
-
Physics- and geometry-aware spatio-spectral graph neural operator for time-independent and time-dependent PDEs
A submission whose abstract describes a new graph neural operator for PDEs but whose full text is a different paper, leaving the claimed method and results unverifiable.
-
Position: Foundation Models Need Digital Twin Representations
A position paper proposes replacing token-based representations in foundation models with outcome-driven digital twin representations that explicitly encode physical and semantic structure.
Reference graph
Works this paper leans on
-
[1]
A mathematical theory of attention,
J. Vuckovic, A. Baratin, and R. T. des Combes, “A mathematical theory of attention,” 2020. https://arxiv.org/abs/2007.02876
arXiv 2020
-
[2]
A mathematical perspective on transformers,
B. Geshkovski, C. Letrouit, Y . Polyanskiy, and P. Rigollet, “A mathematical perspective on transformers,” 2024. https://arxiv.org/abs/2312.10794
arXiv 2024
-
[3]
Geometric dynamics of signal propagation predict trainability of transformers,
A. Cowsik, T. Nebabu, X.-L. Qi, and S. Ganguli, “Geometric dynamics of signal propagation predict trainability of transformers,” 2024. https://arxiv.org/abs/2403.02579. 11 A PREPRINT - JANUARY 22, 2025
arXiv 2024
-
[4]
A. Agrachev and C. Letrouit, “Generic controllability of equivariant systems and applications to particle systems and neural networks,” 2024. https://arxiv.org/abs/2404.08289
work page Pith review arXiv 2024
-
[5]
The emergence of clusters in self-attention dynamics,
B. Geshkovski, C. Letrouit, Y . Polyanskiy, and P. Rigollet, “The emergence of clusters in self-attention dynamics,” in Advances in Neural Information Processing Systems, A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, eds., vol. 36, pp. 57026–57037. Curran Associates, Inc., 2023. https://proceedings.neurips.cc/paper_files/paper/2023/...
work page 2023
-
[6]
Signal propagation in transformers: Theoretical perspectives and the role of rank collapse,
S. Anagnostidis, L. Biggio, L. Noci, A. Orvieto, S. P. Singh, and A. Lucchi, “Signal propagation in transformers: Theoretical perspectives and the role of rank collapse,” in Advances in Neural Information Processing Systems, A. H. Oh, A. Agarwal, D. Belgrave, and K. Cho, eds. 2022. https://openreview.net/forum?id=FxVH7iToXS
work page 2022
-
[7]
Revisiting over-smoothing in BERT from the perspective of graph,
H. Shi, J. GAO, H. Xu, X. Liang, Z. Li, L. Kong, S. M. S. Lee, and J. Kwok, “Revisiting over-smoothing in BERT from the perspective of graph,” inInternational Conference on Learning Representations. 2022. https://openreview.net/forum?id=dUV91uaXm3
work page 2022
-
[8]
Demystifying oversmoothing in attention-based graph neural networks,
X. Wu, A. Ajorlou, Z. Wu, and A. Jadbabaie, “Demystifying oversmoothing in attention-based graph neural networks,” in Thirty-seventh Conference on Neural Information Processing Systems. 2023. https://openreview.net/forum?id=Kg65qieiuB
work page 2023
Show all 60 references
-
[9]
Deep transformers without shortcuts: Modifying self-attention for faithful signal propagation,
B. He, J. Martens, G. Zhang, A. Botev, A. Brock, S. L. Smith, and Y . W. Teh, “Deep transformers without shortcuts: Modifying self-attention for faithful signal propagation,” in The Eleventh International Conference on Learning Representations. 2023. https://openreview.net/for...
2023
-
[10]
On the role of attention masks and layernorm in transformers,
X. Wu, A. Ajorlou, Y . Wang, S. Jegelka, and A. Jadbabaie, “On the role of attention masks and layernorm in transformers,” in The Thirty-eighth Annual Conference on Neural Information Processing Systems. 2024. https://openreview.net/forum?id=lIH6oCdppg
2024
-
[11]
Eliciting latent predictions from transformers with the tuned lens,
N. Belrose, Z. Furman, L. Smith, D. Halawi, I. Ostrovsky, L. McKinney, S. Biderman, and J. Steinhardt, “Eliciting latent predictions from transformers with the tuned lens,” 2023. https://arxiv.org/abs/2303.08112
2023 arXiv
-
[12]
Residual connections encourage iterative inference,
S. Jastrzebski, D. Arpit, N. Ballas, V . Verma, T. Che, and Y . Bengio, “Residual connections encourage iterative inference,” in International Conference on Learning Representations. 2018. https://openreview.net/forum?id=SJa9iHgAZ
2018
-
[13]
interpreting gpt: the logit lens,
nostalgebraist, “interpreting gpt: the logit lens,” LessWrong (2020) . https://www.lesswrong.com/ posts/AcKRB8wDpdaN6v6ru/interpreting-gpt-the-logit-lens
2020
-
[14]
A mathematical framework for transformer circuits,
N. Elhage, N. Nanda, C. Olsson, T. Henighan, N. Joseph, B. Mann, A. Askell, Y . Bai, A. Chen, T. Conerly, N. DasSarma, D. Drain, D. Ganguli, Z. Hatfield-Dodds, D. Hernandez, A. Jones, J. Kernion, L. Lovitt, K. Ndousse, D. Amodei, T. Brown, J. Clark, J. Kaplan, S. McCandlish, a...
2021
-
[15]
Intrinsic dimension of data representations in deep neural networks,
A. Ansuini, A. Laio, J. H. Macke, and D. Zoccolan, “Intrinsic dimension of data representations in deep neural networks,” in Proceedings of the 33rd International Conference on Neural Information Processing Systems. Curran Associates Inc., Red Hook, NY , USA, 2019
2019
-
[16]
Hierarchical nucleation in deep neural networks,
D. Doimo, A. Glielmo, A. Ansuini, and A. Laio, “Hierarchical nucleation in deep neural networks,” in Advances in Neural Information Processing Systems, H. Larochelle, M. Ranzato, R. Hadsell, M. F. Balcan, and H. Lin, eds., vol. 33, pp. 7526–7536. Curran Associates, Inc., 2020
2020
-
[17]
The intrinsic dimension of images and its impact on learning,
P. Pope, C. Zhu, A. Abdelkader, M. Goldblum, and T. Goldstein, “The intrinsic dimension of images and its impact on learning,” in International Conference on Learning Representations. 2021. https://openreview.net/forum?id=XJk19XzGq2J
2021
-
[18]
The geometry of hidden representations of large transformer models,
L. Valeriani, D. Doimo, F. Cuturello, A. Laio, A. Ansuini, and A. Cazzaniga, “The geometry of hidden representations of large transformer models,” in Advances in Neural Information Processing Systems, A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, eds., v...
2023
-
[19]
Bridging information-theoretic and geometric compression in language models,
E. Cheng, C. Kervadec, and M. Baroni, “Bridging information-theoretic and geometric compression in language models,” in Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, H. Bouamor, J. Pino, and K. Bali, eds., pp. 12397–12420. Association ...
2023
-
[20]
Emergence of a high-dimensional abstraction phase in language transformers,
E. Cheng, D. Doimo, C. Kervadec, I. Macocco, J. Yu, A. Laio, and M. Baroni, “Emergence of a high-dimensional abstraction phase in language transformers,” 2024. https://arxiv.org/abs/2405.15471
2024 arXiv
-
[21]
The representation landscape of few-shot learning and fine-tuning in large language models,
D. Doimo, A. P. Serra, A. ansuini, and A. Cazzaniga, “The representation landscape of few-shot learning and fine-tuning in large language models,” in The Thirty-eighth Annual Conference on Neural Information Processing Systems. 2024. https://openreview.net/forum?id=nmUkwoOHFO
2024
-
[22]
Persistent topological features in large language models,
Y . Gardinazzi, G. Panerai, K. Viswanathan, A. Ansuini, A. Cazzaniga, and M. Biagetti, “Persistent topological features in large language models,” 2024. https://arxiv.org/abs/2410.11042
2024 arXiv
-
[23]
Levels of analysis for machine learning,
J. B. Hamrick and S. Mohamed, “Levels of analysis for machine learning,” CoRR abs/2004.05107 (2020) , 2004.05107. https://arxiv.org/abs/2004.05107
2020 arXiv
-
[24]
A primer in BERTology: What we know about how BERT works,
A. Rogers, O. Kovaleva, and A. Rumshisky, “A primer in BERTology: What we know about how BERT works,” Transactions of the Association for Computational Linguistics 8 (2020) 842–866. https://aclanthology.org/2020.tacl-1.54
2020
-
[25]
Probing classifiers: Promises, shortcomings, and advances,
Y . Belinkov, “Probing classifiers: Promises, shortcomings, and advances,”Computational Linguistics 48 no. 1, (Mar., 2022) 207–219. https://aclanthology.org/2022.cl-1.7
2022
-
[26]
Interpretability and analysis in neural NLP,
Y . Belinkov, S. Gehrmann, and E. Pavlick, “Interpretability and analysis in neural NLP,” inProceedings of the 58th Annual Meeting of the Association for Computational Linguistics: Tutorial Abstracts, A. Savary and Y . Zhang, eds., pp. 1–5. Association for Computational Lingui...
2020
-
[27]
Learning Overcomplete Representations,
M. S. Lewicki and T. J. Sejnowski, “Learning Overcomplete Representations,” Neural Computation 12 no. 2, (Feb., 2000) 337–365. https://doi.org/10.1162/089976600300015826. eprint: https://direct.mit.edu/neco/article-pdf/12/2/337/814391/089976600300015826.pdf
2000 doi
-
[28]
Efficient sparse coding algorithms,
H. Lee, A. Battle, R. Raina, and A. Ng, “Efficient sparse coding algorithms,” in Advances in Neural Information Processing Systems, B. Sch¨olkopf, J. Platt, and T. Hoffman, eds., vol. 19. MIT Press, 2006. https://proceedings.neurips.cc/paper_files/paper/2006/file/ 2d71b2ae158c...
2006
-
[29]
Sparse overcomplete word vector representations,
M. Faruqui, Y . Tsvetkov, D. Yogatama, C. Dyer, and N. A. Smith, “Sparse overcomplete word vector representations,” in Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Process...
2015
-
[30]
Transformer feed-forward layers build predictions by promoting concepts in the vocabulary space,
M. Geva, A. Caciularu, K. Wang, and Y . Goldberg, “Transformer feed-forward layers build predictions by promoting concepts in the vocabulary space,” in Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, Y . Goldberg, Z. Kozareva, and Y . Zh...
2022
-
[31]
Future lens: Anticipating subsequent tokens from a single hidden state,
K. Pal, J. Sun, A. Yuan, B. Wallace, and D. Bau, “Future lens: Anticipating subsequent tokens from a single hidden state,” in Proceedings of the 27th Conference on Computational Natural Language Learning (CoNLL), J. Jiang, D. Reitter, and S. Deng, eds., pp. 548–560. Associatio...
2023
-
[32]
Clustering in causal attention masking,
N. Karagodin, Y . Polyanskiy, and P. Rigollet, “Clustering in causal attention masking,” inThe Thirty-eighth Annual Conference on Neural Information Processing Systems. 2024. https://openreview.net/forum?id=OiVxYf9trg
2024
-
[33]
Measure-to-measure interpolation using transformers,
B. Geshkovski, P. Rigollet, and D. Ruiz-Balet, “Measure-to-measure interpolation using transformers,” 2024. https://arxiv.org/abs/2411.04551
2024 arXiv
-
[34]
Dynamic metastability in the self-attention model,
B. Geshkovski, H. Koubbi, Y . Polyanskiy, and P. Rigollet, “Dynamic metastability in the self-attention model,”
-
[35]
Goodfellow, Y
I. Goodfellow, Y . Bengio, A. Courville, and Y . Bengio,Deep learning, vol. 1. MIT Press, 2016
2016
-
[36]
Classification and geometry of general perceptual manifolds,
S. Chung, D. D. Lee, and H. Sompolinsky, “Classification and geometry of general perceptual manifolds,” Physical Review X 8 no. 3, (2018) 031003
2018
-
[37]
Separability and geometry of object manifolds in deep neural networks,
U. Cohen, S. Chung, D. D. Lee, and H. Sompolinsky, “Separability and geometry of object manifolds in deep neural networks,” Nature communications 11 no. 1, (2020) 746. 13 A PREPRINT - JANUARY 22, 2025
2020
-
[38]
Unsupervised detection of semantic correlations in big data,
S. Acevedo, A. Rodriguez, and A. Laio, “Unsupervised detection of semantic correlations in big data,” 2024. https://arxiv.org/abs/2411.02126
2024 arXiv
-
[39]
Implicit geometry of next-token prediction: From language sparsity patterns to model representations,
Y . Zhao, T. Behnia, V . Vakilian, and C. Thrampoulidis, “Implicit geometry of next-token prediction: From language sparsity patterns to model representations,” 2024. https://arxiv.org/abs/2408.15417
2024 arXiv
-
[40]
Intrinsic dimension estimation for robust detection of AI-generated texts,
E. Tulchinskii, K. Kuznetsov, K. Laida, D. Cherniavskii, S. Nikolenko, E. Burnaev, S. Barannikov, and I. Piontkovskaya, “Intrinsic dimension estimation for robust detection of AI-generated texts,” inThirty-seventh Conference on Neural Information Processing Systems. 2023. http...
2023
-
[41]
Estimating the intrinsic dimension of datasets by a minimal neighborhood information,
E. Facco, M. d’Errico, A. Rodriguez, and A. Laio, “Estimating the intrinsic dimension of datasets by a minimal neighborhood information,” Scientific Reports 7 no. 1, (Sep, 2017) 12140
2017
-
[42]
Distributional results for model-based intrinsic dimension estimators,
F. Denti, D. Doimo, A. Laio, and A. Mira, “Distributional results for model-based intrinsic dimension estimators,” 2021. https://arxiv.org/abs/2104.13832
2021 arXiv
-
[43]
Hierarchical nucleation in deep neural networks,
D. Doimo, A. Glielmo, A. Ansuini, and A. Laio, “Hierarchical nucleation in deep neural networks,” in Proceedings of the 34th International Conference on Neural Information Processing Systems, NIPS ’20. Curran Associates Inc., Red Hook, NY , USA, 2020
2020
-
[44]
Introducing meta llama 3: The most capable openly available llm to date,
Meta, “Introducing meta llama 3: The most capable openly available llm to date,” 2024. https://ai.meta.com/blog/meta-llama-3/
2024
-
[45]
Mistral 7b,
A. Q. Jiang, A. Sablayrolles, A. Mensch, C. Bamford, D. S. Chaplot, D. de las Casas, F. Bressand, G. Lengyel, G. Lample, L. Saulnier, L. R. Lavaud, M.-A. Lachaux, P. Stock, T. L. Scao, T. Lavril, T. Wang, T. Lacroix, and W. E. Sayed, “Mistral 7b,” 2023. https://arxiv.org/abs/2...
2023 arXiv
-
[46]
Pythia: a suite for analyzing large language models across training and scaling,
S. Biderman, H. Schoelkopf, Q. Anthony, H. Bradley, K. O’Brien, E. Hallahan, M. A. Khan, S. Purohit, U. S. Prashanth, E. Raff, et al., “Pythia: a suite for analyzing large language models across training and scaling,” in Proceedings of the 40th International Conference on Mach...
2023
-
[47]
The Pile: An 800GB dataset of diverse text for language modeling,
L. Gao, S. Biderman, S. Black, L. Golding, T. Hoppe, C. Foster, J. Phang, H. He, A. Thite, N. Nabeshima, et al., “The Pile: An 800GB dataset of diverse text for language modeling,” arXiv preprint arXiv:2101.00027 (2020)
2020 arXiv
-
[48]
Pile-10k dataset,
N. Nanda, “Pile-10k dataset,” 2022. https://huggingface.co/datasets/NeelNanda/pile-10k
2022
-
[49]
How contextual are contextualized word representations? Comparing the geometry of BERT, ELMo, and GPT-2 embeddings,
K. Ethayarajh, “How contextual are contextualized word representations? Comparing the geometry of BERT, ELMo, and GPT-2 embeddings,” in Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural ...
2019
-
[50]
Mind the gap: Understanding the modality gap in multi-modal contrastive representation learning,
V . W. Liang, Y . Zhang, Y . Kwon, S. Yeung, and J. Y . Zou, “Mind the gap: Understanding the modality gap in multi-modal contrastive representation learning,” Advances in Neural Information Processing Systems 35 (2022) 17612–17625
2022
-
[51]
Abid: Angle based intrinsic dimensionality,
E. Thordsen and E. Schubert, “Abid: Angle based intrinsic dimensionality,” in Similarity Search and Applications: 13th International Conference, SISAP 2020, Copenhagen, Denmark, September 30 – October 2, 2020, Proceedings, p. 218–232. Springer-Verlag, Berlin, Heidelberg, 2020....
2020 doi
-
[52]
Estimating functions of probability distributions from a finite set of samples,
D. H. Wolpert and D. R. Wolf, “Estimating functions of probability distributions from a finite set of samples,” Phys. Rev. E 52 (Dec, 1995) 6841–6854. https://link.aps.org/doi/10.1103/PhysRevE.52.6841
1995 doi
-
[53]
Entropy and inference, revisited,
I. Nemenman, F. Shafee, and W. Bialek, “Entropy and inference, revisited,” in Proceedings of the 15th International Conference on Neural Information Processing Systems: Natural and Synthetic, NIPS’01, p. 471–478. MIT Press, Cambridge, MA, USA, 2001
2001
-
[54]
On the harmonic number (hn) upper and lower
M. V . (https://math.stackexchange.com/users/218419/mark viola), “On the harmonic number (hn) upper and lower ”classical” bounds: which of those is closest to hn?” Mathematics stack exchange. https://math.stackexchange.com/q/2534095. URL:https://math.stackexchange.com/q/253409...
2017
-
[55]
The developmental landscape of in-context learning,
J. Hoogland, G. Wang, M. Farrugia-Roberts, L. Carroll, S. Wei, and D. Murfet, “The developmental landscape of in-context learning,” 2024. https://arxiv.org/abs/2402.02364. 14 A PREPRINT - JANUARY 22, 2025
2024 arXiv
-
[56]
The shape of learning: Anisotropy and intrinsic dimensions in transformer-based models,
A. Razzhigaev, M. Mikhalchuk, E. Goncharova, I. Oseledets, D. Dimitrov, and A. Kuznetsov, “The shape of learning: Anisotropy and intrinsic dimensions in transformer-based models,” in Findings of the Association for Computational Linguistics: EACL 2024, Y . Graham and M. Purver...
2024
-
[57]
Evidence from fmri supports a two-phase abstraction process in language models,
R. Antonello and E. Cheng, “Evidence from fmri supports a two-phase abstraction process in language models,” in UniReps: 2nd Edition of the Workshop on Unifying Representations in Neural Models
-
[58]
Lines of thought in large language models,
R. Sarfati, T. J. B. Liu, N. Boull ´e, and C. J. Earls, “Lines of thought in large language models,” 2024. https://arxiv.org/abs/2410.01545
2024 arXiv
-
[59]
Low bias local intrinsic dimension estimation from expected simplex skewness,
K. Johnsson, C. Soneson, and M. Fontes, “Low bias local intrinsic dimension estimation from expected simplex skewness,” IEEE transactions on pattern analysis and machine intelligence 37 no. 1, (Jan, 2015) 196–202. 15 A PREPRINT - JANUARY 22, 2025 A Consistency Checks for the S...
2015
-
[2024]
https://arxiv.org/abs/2410.06833
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.