{"id":"6bdd48ea-720b-4a61-8390-1f37cca28785","arxiv_id":"2606.19249","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":4.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Empirical analysis of a ViT-Small model shows increasing effective dimensionality and decreasing anisotropy in representations as training progresses, especially in the CLS token.","lead":"The paper introduces the TGO framework to track spectral metrics like effective rank and anisotropy in a Vision Transformer during training on ImageNet-100. A smart generalist might read it to see whether training spreads or concentrates information inside these models.","discovery_kind":"unclear","skeptic_critique":{"model":"grok-4.3","headline":"Single training run on ViT-Small/16 + ImageNet-100 leaves open whether spectral redistribution is architecture- or run-specific","rationale":"The reader's weakest_assumption already isolates the exact point of fragility. No internal inconsistency in metric definitions or contradictory self-comparisons is visible from the given material; the limitation is purely one of scope and replication.","tokens_in":1722,"tokens_out":307,"duration_ms":13039,"concrete_test":"Retrain the identical ViT-Small/16 architecture on ImageNet-100 with two fresh random seeds; recompute Effective Rank, Participation Ratio, Spectral Entropy and Spectral Anisotropy on the CLS token at epochs 0, 50, 100 and final. If the monotonic increase in effective rank and entropy fails to appear (or reverses) in either run, the headline redistribution claim does not generalize beyond the original trajectory.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim asserts a progressive increase in effective dimensionality, participation ratio, spectral entropy and flatness (with decreasing anisotropy) as a general property of ViT representational geometry during training. All reported trends derive from one ViT-Small/16 checkpoint sequence on ImageNet-100; no other seeds, ViT depths/widths, or datasets are shown. Because the metrics are sensitive to the precise covariance structure of the final-layer activations, an idiosyncratic optimization path or dataset bias could produce the observed flattening without the phenomenon being reproducible.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript introduces the Transformer Geometry Observatory (TGO-I) framework for investigating the spectral geometry of Vision Transformer representations. On a single ViT-Small/16 model trained on ImageNet-100, it computes Effective Rank, Stable Rank, Participation Ratio, Spectral Entropy, Spectral Flatness, Spectral Anisotropy, covariance structure, eigenspectra, and singular value spectra across training, reporting a progressive increase in effective dimensionality, participation ratio, spectral entropy and flatness together with decreasing anisotropy—particularly pronounced in the final CLS token—contrary to the intuition that training concentrates variance into few dominant directions.","tokens_in":1821,"tokens_out":413,"duration_ms":17217,"significance":"If the trends prove robust, the work supplies a useful empirical baseline and systematic pipeline for spectral analysis of ViT activations, directly computing standard quantities without fitted parameters. This could inform representational studies in computer vision, though the single-run design restricts claims of generality.","major_comments":[{"comment":"Abstract: the central claim that the observed redistribution of variance constitutes a general property of ViT representational geometry during training is load-bearing for the paper's contribution, yet all trends derive from one ViT-Small/16 checkpoint sequence on ImageNet-100 with no reported error bars, statistical tests, ablation on metrics, or controls for training stochasticity.","section":"Abstract"},{"comment":"Abstract: the assumption that the chosen spectral metrics on this specific architecture, dataset subset, and training run capture general properties (rather than run- or data-specific covariance structure) is not tested; no additional seeds, ViT depths/widths, or datasets are shown to support reproducibility of the flattening trend.","section":"Abstract"}],"minor_comments":[{"comment":"The abstract could more explicitly separate descriptive observations on this model from broader implications for Vision Transformers.","section":null}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the detailed review and for identifying the need to carefully bound the scope of our empirical observations. TGO-I is presented as an initial case study within a larger framework, and we will revise the manuscript to make this explicit while preserving the reported measurements on the single ViT-Small/16 run.","responses":[{"response":"The manuscript reports concrete observations obtained from a single, fully specified training trajectory rather than asserting that the redistribution is a general property of all ViT models. The abstract employs the phrasing “we observe” and “our results reveal” to reflect this empirical scope. Because the spectral quantities are deterministic functions of the saved activations, error bars and statistical tests across stochastic runs are not applicable to the current data release. We will add an explicit limitations paragraph stating that the trends are specific to this architecture–dataset pair and that multi-seed controls are planned for subsequent TGO installments.","revision_made":"partial","referee_comment":"[Abstract] Abstract: the central claim that the observed redistribution of variance constitutes a general property of ViT representational geometry during training is load-bearing for the paper's contribution, yet all trends derive from one ViT-Small/16 checkpoint sequence on ImageNet-100 with no reported error bars, statistical tests, ablation on metrics, or controls for training stochasticity."},{"response":"TGO-I deliberately restricts itself to a canonical ViT-Small/16 model on ImageNet-100 to establish a reproducible baseline pipeline. The paper does not claim that the flattening trend generalizes beyond the reported configuration; the contribution lies in the systematic computation of the listed spectral descriptors on this setup. We will revise the abstract and introduction to foreground the case-study nature of the work and will add a dedicated “Future Work” subsection outlining extensions to additional seeds, depths, and datasets.","revision_made":"yes","referee_comment":"[Abstract] Abstract: the assumption that the chosen spectral metrics on this specific architecture, dataset subset, and training run capture general properties (rather than run- or data-specific covariance structure) is not tested; no additional seeds, ViT depths/widths, or datasets are shown to support reproducibility of the flattening trend."}],"tokens_in":1344,"tokens_out":480,"duration_ms":16334,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main observation is that variance redistributes across more dimensions as training progresses instead of concentrating in a few dominant ones. They track this with standard tools—effective rank, participation ratio, spectral entropy, flatness, and anisotropy—on activations from a single ViT-Small/16 checkpoint sequence.\n\nThe paper does the measurement work carefully. Computing the full set of related spectral quantities on the same data and showing they move together gives a coherent picture. The CLS token ending up with the highest effective dimensionality is a concrete detail worth noting.\n\nThe limitation is the scope. All trends come from one model, one dataset subset, and one training trajectory. No multiple seeds, no other ViT widths or depths, no different datasets, and no error bars or variability estimates. Spectral metrics are sensitive to the precise covariance structure, so an idiosyncratic run could produce the same flattening without it being a general ViT property. The abstract frames the result as contrary to common intuition about training, but that framing rests on limited evidence.\n\nThis is useful for researchers who study representation geometry in transformers and want empirical baselines on how covariance evolves. It is not a new method or a broad claim that has been stress-tested.\n\nI would bring the paper to a reading group to discuss the plots and what controls would be needed next. I would not cite it in my own work until the finding is checked on more runs and architectures. It deserves peer review because the measurements are reproducible in principle and the observation is clear enough to be worth testing further.","headline":"This is a clean but narrow measurement study: one ViT-Small training run on ImageNet-100 shows spectral flattening and rising effective dimensionality, especially in the CLS token.","tokens_in":2289,"tokens_out":390,"would_cite":false,"duration_ms":17476,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Vision Transformer training redistributes variance across representational dimensions rather than concentrating it.","keywords":["Vision Transformers","representational geometry","spectral analysis","effective dimensionality","anisotropy","CLS token","eigenspectra"],"falsifier":"If a different Vision Transformer model or training on a larger dataset shows decreasing effective dimensionality and increasing anisotropy in the CLS token, the observed phenomenon would not hold generally.","tokens_in":2619,"feed_emoji":"📊","tokens_out":535,"duration_ms":18865,"temperature":0.7,"pith_summary":"The paper presents the Transformer Geometry Observatory framework to examine the spectral properties of Vision Transformer representations. It trains a ViT-Small/16 on ImageNet-100 and computes multiple spectral metrics including effective rank and spectral entropy at different stages. The key finding is that training increases dimensional utilization with flatter eigenspectra and reduced anisotropy, contradicting the expectation of information concentration in few directions. This effect is strongest in the CLS token output.","feed_headline":"ViT training spreads variance across more dimensions","feed_subtitle":"Spectral metrics on ImageNet-100 show rising effective rank and flatter spectra, most in CLS token","key_machinery":"Spectral geometry analysis using metrics such as Effective Rank, Stable Rank, Participation Ratio, Spectral Entropy, Spectral Flatness, Spectral Anisotropy, and examination of covariance structure, eigenspectra, and singular value spectra.","core_discovery":"Analysis of ViT representations shows a consistent increase in dimensional utilization, decreasing anisotropy, increasing spectral entropy, and progressively flatter eigenspectra during training, with the final CLS token representation exhibiting the highest effective dimensionality and lowest anisotropy.","pith_inferences":["This suggests that ViTs may learn more robust features by spreading information across dimensions.","Similar redistribution might be observable in other transformer-based models beyond vision.","Training dynamics could be monitored using these spectral metrics to detect convergence or issues."],"forward_implications":["Effective dimensionality of representations increases over the course of training.","Anisotropy in the representations decreases as training progresses.","The CLS token becomes the representation with the highest effective dimensionality within the network.","Spectral entropy and participation ratio increase, indicating more uniform use of dimensions."],"fun_headline_variants":["ViT training raises dimensional utilization","Flatter eigenspectra emerge in ViT training","ViT representations decrease in anisotropy","Highest dimensionality in trained ViT CLS token","Spectral entropy increases in ViT training"],"cache_read_input_tokens":64,"weakest_assumption_plain":"The spectral properties observed in this specific ViT-Small/16 model trained on ImageNet-100 reflect general behavior of Vision Transformers rather than being unique to this setup.","fun_headline_variants_meta":{"raw":{"variants":["ViT training raises dimensional utilization","Flatter eigenspectra emerge in ViT training","ViT representations decrease in anisotropy","Highest dimensionality in trained ViT CLS token","Spectral entropy increases in ViT training"]},"model":"grok-4.3","cost_usd":0.005711,"raw_usage":{"total_tokens":2698,"prompt_tokens":612,"num_sources_used":0,"completion_tokens":62,"cost_in_usd_ticks":57112000,"prompt_tokens_details":{"text_tokens":612,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":2024,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":612,"tokens_out":62,"duration_ms":15217,"temperature":1.0,"reasoning_tokens":2024,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-26T21:38:45.737397+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"If a different Vision Transformer model or training on a larger dataset shows decreasing effective dimensionality and increasing anisotropy in the CLS token, the observed phenomenon would not hold generally.","supporting_citations":[],"review_version":1}