Pith. sign in

REVIEW 2 cited by

Determinant Estimation under Memory Constraints and Neural Scaling Laws

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.04424 v2 pith:WOBAA5KP submitted 2025-03-06 stat.ML cs.LGcs.NAmath.NA

classification stat.MLcs.LGcs.NAmath.NA
keywords matricesneurallawslog-determinantsscalingaccuratelycomputationalderive
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Calculating or accurately estimating log-determinants of large positive definite matrices is of fundamental importance in many machine learning tasks. While its cubic computational complexity can already be prohibitive, in modern applications, even storing the matrices themselves can pose a memory bottleneck. To address this, we derive a novel hierarchical algorithm based on block-wise computation of the LDL decomposition for large-scale log-determinant calculation in memory-constrained settings. In extreme cases where matrices are highly ill-conditioned, accurately computing the full matrix itself may be infeasible. This is particularly relevant when considering kernel matrices at scale, including the empirical Neural Tangent Kernel (NTK) of neural networks trained on large datasets. Under the assumption of neural scaling laws in the test error, we show that the ratio of pseudo-determinants satisfies a power-law relationship, allowing us to derive corresponding scaling laws. This enables accurate estimation of NTK log-determinants from a tiny fraction of the full dataset; in our experiments, this results in a $\sim$100,000$\times$ speedup with improved accuracy over competing approximations. Using these techniques, we successfully estimate log-determinants for dense matrices of extreme sizes, which were previously deemed intractable and inaccessible due to their enormous scale and computational demands.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Spectral Estimation with Free Decompression

    stat.ML 2025-06 conditional novelty 6.0 of 10

    Free decompression evolves a small submatrix spectrum into an estimate of a large matrix spectrum using a PDE derived from free probability, the Nica-Speicher free compression theorem.

  2. Models of Heavy-Tailed Mechanistic Universality

    stat.ML 2025-06 conditional novelty 6.0 of 10

    A new random matrix model with one structure parameter explains heavy-tailed spectra in trained networks, and yields scaling laws, optimizer-tail behavior, and a description of the five-plus-one phases of training.

Pith tools