REVIEW 5 cited by
Reliability of CKA as a Similarity Measure in Deep Learning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Comparing learned neural representations in neural networks is a challenging but important problem, which has been approached in different ways. The Centered Kernel Alignment (CKA) similarity metric, particularly its linear variant, has recently become a popular approach and has been widely used to compare representations of a network's different layers, of architecturally similar networks trained differently, or of models with different architectures trained on the same data. A wide variety of conclusions about similarity and dissimilarity of these various representations have been made using CKA. In this work we present analysis that formally characterizes CKA sensitivity to a large class of simple transformations, which can naturally occur in the context of modern machine learning. This provides a concrete explanation of CKA sensitivity to outliers, which has been observed in past works, and to transformations that preserve the linear separability of the data, an important generalization attribute. We empirically investigate several weaknesses of the CKA similarity metric, demonstrating situations in which it gives unexpected or counter-intuitive results. Finally we study approaches for modifying representations to maintain functional behaviour while changing the CKA value. Our results illustrate that, in many cases, the CKA value can be easily manipulated without substantial changes to the functional behaviour of the models, and call for caution when leveraging activation alignment metrics.
Forward citations
Cited by 5 Pith papers
-
LAEF: A Lead-Agnostic ECG Foundation Model Towards Point-of-Care Diagnostics
A 7M-parameter graph ECG foundation model pre-trained with random lead dropout matches 12-lead models on full input and beats zero-padded baselines on most datasets with 1-2 leads.
-
SOTAlign: Semi-Supervised Alignment of Unimodal Vision and Language Models via Optimal Transport
SOTAlign aligns frozen vision and language encoders with 10k pairs plus up to 1M unpaired samples, beating supervised baselines by 5-10 points on COCO retrieval and ImageNet classification.
-
Can Biologically Plausible Temporal Credit Assignment Rules Match BPTT for Neural Similarity? E-prop as an Example
At matched task accuracy, e-prop trained RNNs reach neural data similarity comparable to BPTT trained RNNs on Mante 2013 and Sussillo 2015 datasets, with initialization and architecture influencing similarity more tha...
-
Grounding Functional Similarity by Invariance-Aware Model Stitching
FuLA, a task-agnostic stitching objective that aligns intermediate features through the frozen end network, is claimed to be a more reliable functional similarity metric than task-based stitching.
-
Reproducing Recurrent Transformers: The CoTFormer
A reproduction of CoTFormer confirms its perplexity results, finds its adaptive-compute claims fragile, and shows looped computation benefits p-hop retrieval but not inductive counting.
Discussion (0). Sign in to comment.