Pith. sign in

REVIEW 6 cited by

Torch.manual_seed(3407) is all you need: On the influence of random seeds in deep learning architectures for computer vision

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2109.08203 v2 pith:S2XXH5O4 submitted 2021-09-16 cs.CV

classification cs.CV
keywords largeseedsarchitecturescomputerdeepinvestigatelearningmuch
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

In this paper I investigate the effect of random seed selection on the accuracy when using popular deep learning architectures for computer vision. I scan a large amount of seeds (up to $10^4$) on CIFAR 10 and I also scan fewer seeds on Imagenet using pre-trained models to investigate large scale datasets. The conclusions are that even if the variance is not very large, it is surprisingly easy to find an outlier that performs much better or much worse than the average.

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Pretrain on Small Synthetic Data, Scale Large for Free: Symmetry-Aware Foundation Model for Logic Rule Induction

    cs.LO 2026-08 conditional novelty 6.0 of 10

    A canonical decoder that reads literal scores yields exactly equivariant discrete DNF rules under the full rule-induction symmetry group, letting a small-data pretrained inducer transfer to schemas of up to 1024 atoms.

  2. GLID: Gated Local Intrinsic Dimension Repairs the Blind Spots of Face-Forgery Detectors

    cs.CR 2026-07 conditional novelty 6.0 of 10

    A gated, training-free local-intrinsic-dimension profile from a frozen ViT repairs face-forgery detectors on unseen GAN and diffusion axes, lifting generation-family AUC by +0.084.

  3. Grokking Is Conditional and Fragile: A Fully-Tractable, Multi-Seed Study at 12K Parameters

    cs.LG 2026-07 accept novelty 6.0 of 10

    In a fully tractable 12K Llama-style model, grokking is a conditional fragile phase transition gated by coverage (tracking modulus more than structure), weight decay, and floating-point reduction order, so evidence mu...

  4. Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs

    cs.CV 2025-07 conditional novelty 6.0 of 10

    LUViT jointly pretrains a ViT with masked auto-encoding and LoRA adapters in a frozen LLM block, reporting +0.4% ImageNet-1K accuracy and up to +2.2% on ImageNet-A over its own MAE baseline.

  5. Algorithmic Tradeoffs, Applied NLP, and the State-of-the-Art Fallacy

    cs.CY 2025-09 conditional novelty 5.0 of 10

    On a large corpus of UC admissions essays, correlated topic models plus linear regression matched or beat word embeddings, BERTopic, XGBoost, CNNs, and LSTM-based models, while LLM-generated essays showed strong but h...

  6. Multi-Loco: Unifying Multi-Embodiment Legged Locomotion via Reinforcement Learning Augmented Diffusion

    cs.RO 2025-06 conditional novelty 5.0 of 10

    A single diffusion-plus-residual-RL policy, trained on four robot morphologies using zero-padded observations and actions, outperforms per-robot PPO baselines in simulation and transfers to real robots.

Pith tools