Pith. sign in

REVIEW 3 cited by

Layers at Similar Depths Generate Similar Activations Across LLM Architectures

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2504.08775 v3 pith:V62H6WBL submitted 2025-04-03 cs.CL cs.AI

classification cs.CLcs.AI
keywords layershareddifferentlayersllmsmodelsnearestneighbor
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

How do the latent spaces used by independently-trained LLMs relate to one another? We study the nearest neighbor relationships induced by activations at different layers of 24 open-weight LLMs, and find that they 1) tend to vary from layer to layer within a model, and 2) are approximately shared between corresponding layers of different models. Claim 2 shows that these nearest neighbor relationships are not arbitrary, as they are shared across models, but Claim 1 shows that they are not "obvious" either, as there is no single set of nearest neighbor relationships that is universally shared. Together, these suggest that LLMs generate a progression of activation geometries from layer to layer, but that this entire progression is largely shared between models, stretched and squeezed to fit into different architectures.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Contravariance Theory: Strong Alignment for Minimal Solutions to Hard Tasks

    cs.LG 2026-07 conditional novelty 7.5 of 10

    For minimal networks solving hard tasks, linear (weak) equivalence across adjacent layers forces unit-level axis alignment, and terminal linear equivalence zippers upstream to make representation convergence mathemati...

  2. Large Language Models Show Signs of Alignment with Human Neurocognition During Abstract Reasoning

    q-bio.NC 2025-08 unverdicted novelty 6.0 of 10

    Only the largest tested LLMs (about 70 billion parameters) match human accuracy on an abstract reasoning task, and the internal geometry of their best layers correlates moderately with human frontal EEG activity.

  3. Correlated Errors in Large Language Models

    cs.CL 2025-06 conditional novelty 6.0 of 10

    Large language models from different providers and architectures often make the same errors, and more accurate models are especially likely to share mistakes.

Pith tools