Pith. sign in

REVIEW 5 cited by

Reliability of CKA as a Similarity Measure in Deep Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2210.16156 v2 pith:L3OG5MLG submitted 2022-10-28 cs.LG cs.AIcs.CV

classification cs.LGcs.AIcs.CV
keywords beenrepresentationssimilaritydifferentalignmentbehaviourdatafunctional
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Comparing learned neural representations in neural networks is a challenging but important problem, which has been approached in different ways. The Centered Kernel Alignment (CKA) similarity metric, particularly its linear variant, has recently become a popular approach and has been widely used to compare representations of a network's different layers, of architecturally similar networks trained differently, or of models with different architectures trained on the same data. A wide variety of conclusions about similarity and dissimilarity of these various representations have been made using CKA. In this work we present analysis that formally characterizes CKA sensitivity to a large class of simple transformations, which can naturally occur in the context of modern machine learning. This provides a concrete explanation of CKA sensitivity to outliers, which has been observed in past works, and to transformations that preserve the linear separability of the data, an important generalization attribute. We empirically investigate several weaknesses of the CKA similarity metric, demonstrating situations in which it gives unexpected or counter-intuitive results. Finally we study approaches for modifying representations to maintain functional behaviour while changing the CKA value. Our results illustrate that, in many cases, the CKA value can be easily manipulated without substantial changes to the functional behaviour of the models, and call for caution when leveraging activation alignment metrics.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LAEF: A Lead-Agnostic ECG Foundation Model Towards Point-of-Care Diagnostics

    cs.LG 2026-08 conditional novelty 7.0 of 10

    A 7M-parameter graph ECG foundation model pre-trained with random lead dropout matches 12-lead models on full input and beats zero-padded baselines on most datasets with 1-2 leads.

  2. SOTAlign: Semi-Supervised Alignment of Unimodal Vision and Language Models via Optimal Transport

    cs.LG 2026-02 conditional novelty 6.0 of 10

    SOTAlign aligns frozen vision and language encoders with 10k pairs plus up to 1M unpaired samples, beating supervised baselines by 5-10 points on COCO retrieval and ImageNet classification.

  3. Can Biologically Plausible Temporal Credit Assignment Rules Match BPTT for Neural Similarity? E-prop as an Example

    cs.NE 2025-06 conditional novelty 6.0 of 10

    At matched task accuracy, e-prop trained RNNs reach neural data similarity comparable to BPTT trained RNNs on Mante 2013 and Sussillo 2015 datasets, with initialization and architecture influencing similarity more tha...

  4. Grounding Functional Similarity by Invariance-Aware Model Stitching

    cs.LG 2025-05 conditional novelty 6.0 of 10

    FuLA, a task-agnostic stitching objective that aligns intermediate features through the frozen end network, is claimed to be a more reliable functional similarity metric than task-based stitching.

  5. Reproducing Recurrent Transformers: The CoTFormer

    cs.LG 2026-07 conditional novelty 5.0 of 10

    A reproduction of CoTFormer confirms its perplexity results, finds its adaptive-compute claims fragile, and shows looped computation benefits p-hop retrieval but not inductive counting.

Pith tools