REVIEW 3 major objections 6 minor 182 references
Smaller, Faster, Cheaper: Architectural Designs for Efficient Machine Learning
T0 review · 3 major / 6 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read Architectural design, not raw scale, can make neural networks smaller, faster, and cheaper without sacrificing performance, the dissertation argues, and it offers three concrete interventions as evidence.
desk verdict A well-organized thesis compilation of three solid efficiency papers, but the Flowers-102 split gap undercuts the headline data-efficiency claim. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The argument is carried by three named mechanisms. First, the Compact Convolutional Transformer tokenizer: an overlapping convolutional patch embedding, with kernel size larger than stride followed by max pooling, that preserves boundary information, paired with SeqPool, a learned attention-inspired pooling that contracts the token sequence into a vector instead of slicing a single class token. Second, Hydra Neighborhood Attention: a restricted-attention mechanism in which each attention head has its own window size and dilation, so dense local and sparse global receptive fields intermix; it reduces to standard neighborhood attention when heads share one kernel and to full self-attention when the window covers the image. Third, compositional invertibility of normalizing flows: because a flow is a composition of diffeomorphic layers, knowledge can be transferred as latent distillation, intermediate-latent distillation, and synthesized distillation, with a combined loss weighting each channel.
What would settle it
Retrain CCT-14/7×2 on the standard Flowers-102 split using the dissertation's exact recipe; if it yields roughly 68.85% instead of 97.19%, the Flowers-102 evidence for escaping the big-data paradigm fails on that split. Also rerun the CIFAR-10 CCT-7/3×1 experiments with a validation split fixed before any hyperparameter search; if the reported 98% drops materially, part of the gain is selection on the test set.
Extended reading notes
Core claim
On the dissertation's own terms, the core claim is that three architectural interventions each improve accuracy or likelihood per unit of compute and data in their settings, and together they answer affirmatively whether neural architectures can be smaller, faster, and cheaper without sacrificing performance. Chapter 3 shows that a Compact Convolutional Transformer with an overlapping convolutional tokenizer and SeqPool reaches 98.00% on CIFAR-10 with 3.76M parameters when trained from scratch, outperforming much larger ViTs and ResNets, and reaches 81.34% top-1 on ImageNet with distillation against DeiT-S's 81.16%. Chapter 4 shows that variadic attention heads, giving each head an independent window size and dilation, let a StyleGAN-style generator score FID 2.05 on FFHQ-256 at 48.92M parameters and 59.90 images per second, beating StyleSwin's 2.81 and StyleGAN-XL's 2.19 at lower parameter cost and higher throughput. Chapter 5 formalizes flow distillation into latent, intermediate-latent, and synthesized channels, showing a GLOW student with about a quarter of the teacher's parameters retaining roughly 99% of its BSDS300 log-likelihood, and intermediate-latent distillation improving image-generation FID on both CIFAR-10 and CelebA.
Load-bearing premise
The load-bearing premise is that the reported benchmark numbers are unbiased held-out estimates of generalization, but hyperparameters were selected by validation-set search with best results reported, and the Flowers-102 97.19% headline came from the Kaggle split while the standard split yields 68.85%, so if the premise fails the small-data claim loses its headline evidence.
Editorial extensions
If this is right
- Vision transformers can be trained from scratch on small and medium image datasets and still beat comparable ViTs and ResNets, so web-scale pretraining is not a prerequisite for transformer adoption in vision.
- Allowing attention heads to use independent window sizes and dilations lets generative models mix local and global structure, which improves FID at comparable parameter count and throughput on FFHQ-256.
- Because normalizing flows are compositions of invertible layers, teacher knowledge can be transferred in three directions, so a flow student a fraction of the teacher's size can retain most of its density-estimation performance.
- Efficiency gains can come from data ingress and egress, core attention design, and structure-aware distillation, so scaling data and parameters is not the only route to progress.
Reading between the lines
- Editorial inference: the Hydra-NA idea should transfer to other generative backbones, especially diffusion models where local texture and global layout both matter; swapping a latent diffusion model's attention blocks for variadic heads would test this directly.
- Editorial inference: the Flowers-102 split sensitivity suggests small-dataset transformer results may be fragile to dataset construction, so standardizing evaluation splits would make the escape-from-big-data claim falsifiable across labs.
- Editorial inference: the three flow-distillation channels could in principle extend to continuous flow-matching models by aligning trajectories in time rather than composing layers, though the dissertation only gestures at this possibility.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This dissertation argues that careful architectural design can make vision and generative models smaller, faster, and cheaper without sacrificing performance, and it presents three supporting interventions. Chapter 3 introduces Compact Convolutional Transformers (CCTs), which combine an overlapping convolutional tokenizer with SeqPool, and reports strong results on small datasets, ImageNet, and Flowers-102. Chapter 4 extends Neighborhood Attention to allow attention heads to have independent kernel sizes and dilations (Hydra-NA), and demonstrates improved FID, throughput, and parameter counts for StyleGAN-style image generation on FFHQ and LSUN Church. Chapter 5 formalizes three categories of knowledge distillation for normalizing flows (latent, intermediate-latent, and synthesized) and shows improved density estimation and image generation over a non-distilled student. The dissertation concludes that these results positively answer the question of whether architectures can be designed to be smaller, faster, and cheaper without sacrificing performance.
Significance. If the reported results hold, the thesis makes a useful empirical case that targeted architectural modifications (overlapping tokenization, per-head receptive fields, and flow-structure-aware distillation) can shift the efficiency-accuracy Pareto frontier. The work has several strengths: the CCT and StyleNAT code and checkpoints are released; the author is transparent about the Flowers-102 split problem and about the non-convergence of some StyleNAT runs; and the attention-map analysis and metric critique in Chapter 4 are thoughtful. The main significance is dampened by the evaluation-protocol issues in Chapter 3, which is the chapter most directly tied to the dissertation's 'escape the big-data paradigm' claim.
major comments (3)
- [Section 3.3.6, Table 8] The headline Flowers-102 result of 97.19% for CCT-14/7x2 without pretraining is obtained on the Kaggle split, and the text immediately discloses that retraining on the standard torchvision split gives 68.85%. A 28-point swing due to the split choice is far larger than the efficiency margins the chapter claims, so the statements 'able to achieve an accuracy of over 97% without the use of any pretraining data' and 'state of the art results' (99.76% with pretraining) are not supported on the standard benchmark. The chapter should report the standard-split numbers as primary and either remove the Kaggle-split claims or clearly re-label them as non-standard; the Section 6.1 summary should be revised to match.
- [Section 3.3.3, Section 3.3.5, Table 6] The text states that hyperparameters were selected by directed search on a validation set and that 'the best results we achieved' are reported, while Table 6 notes that numbers are 'best out of 4 runs' with random initializations. Because no separate validation split and no measure of run-to-run variability are provided, the reported accuracies in Table 3 are selected, not typical, estimates. This matters for Chapter 3's data-efficiency claims, several of which rest on small margins (e.g., CCT-7/3x1 vs. ResNet164 on CIFAR-100). Please report mean and standard deviation over the repeated runs, or at minimum the full set of per-run numbers, and make explicit which split was used for model selection.
- [Section 5.3.1, Section 5.3.2, Table 17] The claim that the GLOW student 'achieves 98.94% the accuracy' of the teacher on BSDS300 is based only on test log-likelihood. In the image-generation experiments the same method leaves a much larger relative gap: the CelebA teacher FID is 37.460, the non-distilled student is 68.127, and the ILKD student is 54.480. The text should temper the 'passing nearly all its knowledge' conclusion and present the interpolation-based FID of Table 18 as evidence of the transfer, since the raw FID gap indicates significant remaining quality loss.
minor comments (6)
- [Section 3.3.6 and Table 8 caption] The first paragraph of Section 3.3.6 states 'we are able to achieve an accuracy of over 97%' without the caveat about the Kaggle split, which only appears later in the section and in the table note; the caveat should accompany the first mention.
- [Section 3.3.5, Figure 9] The subcaption text for Figures 9b and 9c appears swapped: the text says 'Fig. 9c with sinusoidal' and 'Fig. 9b with a learnable positional embedding,' but the panels are captioned the reverse. Please check and correct the cross-references.
- [Chapter 2 text] The document contains several typographical errors, including 'attnetion' (Section 2.3.1), 'explinations' (Acknowledgements), 'Labratory' (Curriculum Vitae), and 'CIF AR-10' spacing throughout; a copyedit pass is needed.
- [Section 5.3.1] The sentence 'our final student GLOW model is has ≈25.5 as many parameters as the teacher model' is missing the percentage sign: it should read '≈25.5% as many parameters.'
- [Section 5.2.1.4, Eq. (5.9)] The loss weights λ_i are introduced without a normalization convention; the experimental section says they are 'percentages of the whole loss,' but this is not reflected in Equation (5.9), so please clarify whether the weights sum to one or are scaled elsewhere.
- [Chapter 3, Table 6] The main-results Table 3 does not state whether those numbers are also 'best out of 4 runs' like Table 6; please add an explicit statement about run selection for all tables in Chapter 3.
Circularity Check
No significant circularity: the dissertation's claims are empirical and benchmark-driven; disclosed evaluation-protocol caveats are validity concerns, not derivation-circularity.
full rationale
This dissertation does not present a formal derivation chain whose conclusions reduce to its inputs; the three contributions are architectural modifications evaluated against external baselines. Chapter 3's CCT is defined by explicit tokenizer and pooling equations (3.1)-(3.2) and compared to ResNet/ViT/DeiT baselines; Chapter 4's variadic attention heads are defined by (4.3)-(4.4) and ablated against StyleSwin and StyleGAN variants; Chapter 5's flow-distillation framework is defined by (5.6)-(5.9) and tested against teacher/student GLOW and MAF models. The main evaluation-protocol weaknesses are explicitly acknowledged in the text: hyperparameters were selected by directed search and the best of 4 runs are reported (Sections 3.3.3 and Table 6 caption, and Section 6.2.3.1), and the Flowers-102 result depends on the Kaggle split, with the torchvision-split retraining giving 68.85% rather than 97.19% (Section 3.3.6). These are real threats to the trustworthiness of the reported generalization estimates, but they are selection-bias and benchmark-validity problems, not circularity: no reported number is equivalent by construction to a fitted input, and no load-bearing claim is justified solely by an unverified self-citation. Self-citations to the authors' prior NAT/CCT works serve as background or as the same works being reported, with code and experimental details included, so they do not make the argument circular.
Assumptions & free parameters
free parameters (4)
- CCT tokenizer hyperparameters (kernel, stride, number of convolutions) =
3x3 kernel, stride 1, 1 convolution for CCT-7/3x1; variants use 7x7 or 3x3 with strides 1 or 2
- StyleNAT attention head kernel and dilation schedule =
kernel 7 with dilations 1,2,4,...,128 depending on resolution (Table 10); 8 splits for LSUN Church
- Distillation loss weights lambda_i =
lambda0=0.9, lambda1=lambda2=0.1; for SKD lambda0=0.85 and the others 0.075
- StyleNAT training iterations and LR-decay start =
LR-decay at 740k, stopped at 940k for FFHQ-256; 500k and 900k for FFHQ-1024
assumptions (5)
- domain assumption Reported CIFAR-10/100 and Flowers-102 accuracies are unbiased held-out estimates, with no separate validation split used for model selection.
- domain assumption The student's latent representation is at least as large as the latent data manifold, making full distillation possible.
- domain assumption The latent manifold hypothesis: information needed to generate an image is smaller than the image's dimensionality.
- domain assumption Neighborhood attention with per-head dilation can emulate global attention when the dilated window covers the image.
- standard math Change-of-variables formula for bijective differentiable maps
Cite this review
Pith. "Pith review of Smaller, Faster, Cheaper: Architectural Designs for Efficient Machine Learning." pith.science (2026). https://pith.science/paper/M7EPWACO
@misc{pith2026250719795,
author = {Pith},
title = {Pith review of: Smaller, Faster, Cheaper: Architectural Designs for Efficient Machine Learning},
year = {2026},
howpublished = {\url{https://pith.science/paper/M7EPWACO}},
note = {Machine review of arXiv:2507.19795}
}
read the original abstract
Major advancements in the capabilities of computer vision models have been primarily fueled by rapid expansion of datasets, model parameters, and computational budgets, leading to ever-increasing demands on computational infrastructure. However, as these models are deployed in increasingly diverse and resource-constrained environments, there is a pressing need for architectures that can deliver high performance while requiring fewer computational resources. This dissertation focuses on architectural principles through which models can achieve increased performance while reducing their computational demands. We discuss strides towards this goal through three directions. First, we focus on data ingress and egress, investigating how information may be passed into and retrieved from our core neural processing units. This ensures that our models make the most of available data, allowing smaller architectures to become more performant. Second, we investigate modifications to the core neural architecture, applied to restricted attention in vision transformers. This section explores how removing uniform context windows in restricted attention increases the expressivity of the underlying neural architecture. Third, we explore the natural structures of Normalizing Flows and how we can leverage these properties to better distill model knowledge. These contributions demonstrate that careful design of neural architectures can increase the efficiency of machine learning algorithms, allowing them to become smaller, faster, and cheaper.
Figures
Figures from the paper (22 more)
Reference graph
Works this paper leans on
-
[1]
Semdedup: Data-efficient learning at web-scale through semantic deduplication
Amro Abbas, Kushal Tirumala, D´ aniel Simig, Surya Ganguli, and Ari S Morcos. Semdedup: Data-efficient learning at web-scale through semantic deduplication. arXiv preprint arXiv:2303.09540 , 2023
arXiv 2023
-
[2]
Dbpedia: A nucleus for a web of open data
S¨ oren Auer, Christian Bizer, Georgi Kobilarov, Jens Lehmann, Richard Cyganiak, and Zachary Ives. Dbpedia: A nucleus for a web of open data. In The semantic web, pages 722–735. Springer, 2007
2007
-
[3]
Neural machine translation by jointly learning to align and translate, 2016
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. Neural machine translation by jointly learning to align and translate, 2016
2016
-
[4]
Distilling the knowledge from conditional normalizing flows
Dmitry Baranchuk, Vladimir Aliev, and Artem Babenko. Distilling the knowledge from conditional normalizing flows. In ICML Workshop on Invertible Neural Networks, Normalizing Flows, and Explicit Likelihood Models , 2021
2021
-
[5]
Semantic photo manipulation with a generative image prior
David Bau, Hendrik Strobelt, William Peebles, Jonas Wulff, Bolei Zhou, Jun-Yan Zhu, and Antonio Torralba. Semantic photo manipulation with a generative image prior. ACM Transactions on Graphics , 38(4):1–11, 2019
2019
-
[6]
Irwan Bello, Barret Zoph, Ashish Vaswani, Jonathon Shlens, and Quoc V. Le. Attention augmented convolutional networks. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , 2019
2019
-
[7]
Longformer: The long-document transformer
Iz Beltagy, Matthew E Peters, and Arman Cohan. Longformer: The long-document transformer. arXiv preprint arXiv:2004.05150 , 2020
arXiv 2004
-
[8]
Rianne van den Berg, Leonard Hasenclever, Jakub M. Tomczak, and Max Welling. Sylvester Normalizing Flows for Variational Inference, 2019. arXiv:1803.05649 [cs, stat]. 124
arXiv 2019
Show all 182 references
-
[9]
Unleashing transformers: Parallel token prediction with discrete absorbing diffusion for fast high-resolution image generation from vector-quantized codes
Sam Bond-Taylor, Peter Hessey, Hiroshi Sasaki, Toby P Breckon, and Chris G Willcocks. Unleashing transformers: Parallel token prediction with discrete absorbing diffusion for fast high-resolution image generation from vector-quantized codes. In European Conference on Computer ...
2022
-
[10]
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffr...
2020
-
[11]
Model compression
Cristian Bucilu˘a, Rich Caruana, and Alexandru Niculescu-Mizil. Model compression. In Proceedings of the 12th ACM SIGKDD international conference on Knowledge discovery and data mining , pages 535–541, 2006
2006
-
[12]
Proxylessnas: Direct neural architecture search on target task and hardware
Han Cai, Ligeng Zhu, and Song Han. Proxylessnas: Direct neural architecture search on target task and hardware. In ICLR, 2018
2018
-
[13]
Cascade r-cnn: Delving into high quality object detection
Zhaowei Cai and Nuno Vasconcelos. Cascade r-cnn: Delving into high quality object detection. In CVPR, 2018
2018
-
[14]
Ricky T. Q. Chen and Yaron Lipman. Flow matching on general geometries, 2024
2024
-
[15]
Ricky T. Q. Chen, Jens Behrmann, David Duvenaud, and J¨ orn-Henrik Jacobsen. Residual Flows for Invertible Generative Modeling, 2020. arXiv:1906.02735 [cs, stat]. 125
2020 arXiv
-
[16]
On the relationship between self-attention and convolutional layers
Jean-Baptiste Cordonnier, Andreas Loukas, and Martin Jaggi. On the relationship between self-attention and convolutional layers. In International Conference on Learning Representations, 2020
2020
-
[17]
Scalable high-resolution pixel-space image synthesis with hourglass diffusion transformers
Katherine Crowson, Stefan Andreas Baumann, Alex Birch, Tanishq Mathew Abraham, Daniel Z Kaplan, and Enrico Shippole. Scalable high-resolution pixel-space image synthesis with hourglass diffusion transformers. In Forty-first International Conference on Machine Learning , 2024
2024
-
[18]
Approximation with artificial neural networks.Faculty of Sciences, Etvs Lornd University, Hungary , 24(48):7, 2001
Bal´ azs Csan´ ad Cs´ aji et al. Approximation with artificial neural networks.Faculty of Sciences, Etvs Lornd University, Hungary , 24(48):7, 2001
2001
-
[19]
Randaugment: Practical automated data augmentation with a reduced search space
Ekin D Cubuk, Barret Zoph, Jonathon Shlens, and Quoc V Le. Randaugment: Practical automated data augmentation with a reduced search space. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, pages 702–703, 2020
2020
-
[20]
Flow matching in latent space
Quan Dao, Hao Phung, Binh Nguyen, and Anh Tran. Flow matching in latent space. arXiv preprint arXiv:2307.08698 , 2023
2023 arXiv
-
[21]
Flashattention-2: Faster attention with better parallelism and work partitioning
Tri Dao. Flashattention-2: Faster attention with better parallelism and work partitioning. In ICLR, 2023
2023
-
[22]
Flashattention: Fast and memory-efficient exact attention with io-awareness
Tri Dao, Daniel Y Fu, Stefano Ermon, Atri Rudra, and Christopher R´ e. Flashattention: Fast and memory-efficient exact attention with io-awareness. In NeurIPS, 2022
2022
-
[23]
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In CVPR, 2009
2009
-
[24]
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. In 126 Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human l...
2019
-
[25]
Ensemble methods in machine learning
Thomas G Dietterich. Ensemble methods in machine learning. In International workshop on multiple classifier systems , pages 1–15. Springer, 2000
2000
-
[26]
Density estimation using real nvp, 2017
Laurent Dinh, Jascha Sohl-Dickstein, and Samy Bengio. Density estimation using real nvp, 2017
2017
-
[27]
Dolatabadi, Sarah Erfani, and Christopher Leckie
Hadi M. Dolatabadi, Sarah Erfani, and Christopher Leckie. Invertible Generative Modeling using Linear Rational Splines, 2020. arXiv:2001.05168 [cs, stat]
2020 arXiv
-
[28]
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition at...
2021
-
[29]
Augmented Neural ODEs
Emilien Dupont, Arnaud Doucet, and Yee Whye Teh. Augmented Neural ODEs. In Advances in Neural Information Processing Systems . Curran Associates, Inc., 2019
2019
-
[30]
Cubic- Spline Flows, 2019
Conor Durkan, Artur Bekasov, Iain Murray, and George Papamakarios. Cubic- Spline Flows, 2019. arXiv:1906.02145 [cs, stat]
2019 arXiv
-
[31]
Softmax linear units
Nelson Elhage, Tristan Hume, Catherine Olsson, Neel Nanda, Tom Henighan, Scott Johnston, Sheer ElShowk, Nicholas Joseph, Nova DasSarma, Ben Mann, Danny Hernandez, Amanda Askell, Kamal Ndousse, Andy Jones, Dawn Drain, Anna Chen, Yuntao Bai, Deep Ganguli, Liane Lovitt, Zac Hatfi...
-
[32]
Visualizing higher-layer features of a deep network, 2009
Dumitru Erhan, Yoshua Bengio, Aaron Courville, and Pascal Vincent. Visualizing higher-layer features of a deep network, 2009
2009
-
[33]
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Jonathan Frankle and Michael Carbin. The lottery ticket hypothesis: Finding sparse, trainable neural networks. In International Conference on Learning Representations, 2019
2019
-
[34]
Fuhr and Michael Kallay
Richard D. Fuhr and Michael Kallay. Monotone linear rational spline interpolation. Computer Aided Geometric Design , 9(4):313–319, 1992
1992
-
[35]
A new algorithm for data compression
Philip Gage. A new algorithm for data compression. C Users J. , 12(2):23–38, 1994
1994
-
[36]
Swagan: A style-based wavelet-driven generative model
Rinon Gal, Dana Cohen Hochberg, Amit Bermano, and Daniel Cohen-Or. Swagan: A style-based wavelet-driven generative model. ACM Trans. Graph., 40(4), 2021
2021
-
[37]
Masked diffusion transformer is a strong image synthesizer
Shanghua Gao, Pan Zhou, Ming-Ming Cheng, and Shuicheng Yan. Masked diffusion transformer is a strong image synthesizer. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , pages 23164–23173, 2023
2023
-
[38]
Mahoney, and Kurt Keutzer
Amir Gholami, Sehoon Kim, Zhen Dong, Zhewei Yao, Michael W. Mahoney, and Kurt Keutzer. A survey of quantization methods for efficient neural network inference, 2021
2021
-
[39]
Ganalyze: Toward visual definitions of cognitive image properties
Lore Goetschalckx, Alex Andonian, Aude Oliva, and Phillip Isola. Ganalyze: Toward visual definitions of cognitive image properties. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , 2019
2019
-
[40]
Deep Learning
Ian Goodfellow, Yoshua Bengio, and Aaron Courville. Deep Learning. MIT Press, 2016. 128
2016
-
[41]
Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde- Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio
Ian J. Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde- Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. Generative adversarial nets. In Advances in Neural Information Processing Systems . Curran Associates, Inc., 2014
2014
-
[42]
Will Grathwohl, Ricky T. Q. Chen, Jesse Bettencourt, and David Duvenaud. Scalable reversible generative models with free-form continuous dynamics. In International Conference on Learning Representations , 2019
2019
-
[43]
Neural turing machines, 2014
Alex Graves, Greg Wayne, and Ivo Danihelka. Neural turing machines, 2014
2014
-
[44]
Densely connected normalizing flows, 2021
Matej Grci´ c, Ivan Grubiˇ si´ c, and Siniˇ saˇSegvi´ c. Densely connected normalizing flows, 2021
2021
-
[45]
Pattern theory: from representation to inference
Ulf Grenander and Michael I Miller. Pattern theory: from representation to inference. OUP Oxford, 2006
2006
-
[46]
Mamba: Linear-time sequence modeling with selective state spaces, 2024
Albert Gu and Tri Dao. Mamba: Linear-time sequence modeling with selective state spaces, 2024
2024
-
[47]
Starflow: Scaling latent normalizing flows for high-resolution image synthesis, 2025
Jiatao Gu, Tianrong Chen, David Berthelot, Huangjie Zheng, Yuyang Wang, Ruixiang Zhang, Laurent Dinh, Miguel Angel Bautista, Josh Susskind, and Shuangfei Zhai. Starflow: Scaling latent normalizing flows for high-resolution image synthesis, 2025
2025
-
[48]
Variational Inference with Orthogonal Normalizing Flows
Leonard Hasenclever, Jakub M Tomczak, and Max Welling. Variational Inference with Orthogonal Normalizing Flows. In Bayesian Deep Learning, 2017
2017
-
[49]
Neighborhood attention: Dynamic restriction of self-attention, 2023
Ali Hassani. Neighborhood attention: Dynamic restriction of self-attention, 2023
2023
-
[50]
Dilated neighborhood attention transformer, 2023
Ali Hassani and Humphrey Shi. Dilated neighborhood attention transformer, 2023. 129
2023
-
[51]
Escaping the big data paradigm with compact transformers, 2022
Ali Hassani, Steven Walton, Nikhil Shah, Abulikemu Abuduweili, Jiachen Li, and Humphrey Shi. Escaping the big data paradigm with compact transformers, 2022
2022
-
[52]
Neighborhood attention transformer
Ali Hassani, Steven Walton, Jiachen Li, Shen Li, and Humphrey Shi. Neighborhood attention transformer. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages 6185–6194, 2023
2023
-
[53]
Faster neighborhood attention: Reducing the O(n2) cost of self attention at the threadblock level, 2024
Ali Hassani, Wen-Mei Hwu, and Humphrey Shi. Faster neighborhood attention: Reducing the O(n2) cost of self attention at the threadblock level, 2024
2024
-
[54]
Generalized neighborhood attention: Multi-dimensional sparse attention at the speed of light, 2025
Ali Hassani, Fengzhe Zhou, Aditya Kane, Jiannan Huang, Chieh-Yun Chen, Min Shi, Steven Walton, Markus Hoehnerbach, Vijay Thakkar, Michael Isaev, Qinsheng Zhang, Bing Xu, Haicheng Wu, Wen mei Hwu, Ming-Yu Liu, and Humphrey Shi. Generalized neighborhood attention: Multi-dimensio...
2025
-
[55]
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2016
2016
-
[56]
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR, 2016
2016
-
[57]
Identity mappings in deep residual networks
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Identity mappings in deep residual networks. In ECCV, 2016
2016
-
[58]
Mask r-cnn
Kaiming He, Georgia Gkioxari, Piotr Doll´ ar, and Ross Girshick. Mask r-cnn. In ICCV, 2017
2017
-
[59]
Masked autoencoders are scalable vision learners
Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Doll´ ar, and Ross Girshick. Masked autoencoders are scalable vision learners. In Proceedings of 130 the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages 16000–16009, 2022
2022
-
[60]
Gans trained by a two time-scale update rule converge to a local nash equilibrium
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. Gans trained by a two time-scale update rule converge to a local nash equilibrium. In Advances in Neural Information Processing Systems . Curran Associates, Inc., 2017
2017
-
[61]
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. Distilling the knowledge in a neural network. arXiv preprint arXiv:1503.02531 , 2015
2015 arXiv
-
[62]
Geoffrey F. Hinton. Shape representation in parallel systems. In Proceedings of the 7th International Joint Conference on Artificial Intelligence - Volume 2 , page 1088–1096, San Francisco, CA, USA, 1981. Morgan Kaufmann Publishers Inc
1981
-
[63]
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. In Advances in Neural Information Processing Systems , pages 6840–6851. Curran Associates, Inc., 2020
2020
-
[64]
The convolution exponential and generalized sylvester flows
Emiel Hoogeboom, Victor Garcia Satorras, Jakub Tomczak, and Max Welling. The convolution exponential and generalized sylvester flows. In Advances in Neural Information Processing Systems, pages 18249–18260. Curran Associates, Inc., 2020
2020
-
[65]
Simple diffusion: End-to- end diffusion for high resolution images, 2023
Emiel Hoogeboom, Jonathan Heek, and Tim Salimans. Simple diffusion: End-to- end diffusion for high resolution images, 2023
2023
-
[66]
On the limitations of compute thresholds as a governance strategy, 2024
Sara Hooker. On the limitations of compute thresholds as a governance strategy, 2024
2024
-
[67]
Multilayer feedforward networks are universal approximators
Kurt Hornik, Maxwell Stinchcombe, and Halbert White. Multilayer feedforward networks are universal approximators. Neural networks, 2(5):359–366, 1989. 131
1989
-
[68]
Local relation networks for image recognition
Han Hu, Zheng Zhang, Zhenda Xie, and Stephen Lin. Local relation networks for image recognition. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , 2019
2019
-
[69]
Squeeze-and-excitation networks
Jie Hu, Li Shen, and Gang Sun. Squeeze-and-excitation networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2018
2018
-
[70]
Arbitrary style transfer in real-time with adaptive instance normalization
Xun Huang and Serge Belongie. Arbitrary style transfer in real-time with adaptive instance normalization. In Proceedings of the IEEE International Conference on Computer Vision (ICCV) , 2017
2017
-
[71]
Lawrence Zitnick
Drew A Hudson and C. Lawrence Zitnick. Generative adversarial transformers. Proceedings of the 38th International Conference on Machine Learning, ICML 2021 , 2021
2021
-
[72]
Lawrence Zitnick
Drew A Hudson and C. Lawrence Zitnick. Compositional transformers for scene generation. Advances in Neural Information Processing Systems NeurIPS 2021 , 2021
2021
-
[73]
Estimation of Non-Normalized Statistical Models by Score Matching
Aapo Hyv¨ arinen. Estimation of Non-Normalized Statistical Models by Score Matching. Journal of Machine Learning Research , 6(24):695–709, 2005
2005
-
[74]
Averaging weights leads to wider optima and better generalization
Pavel Izmailov, Dmitrii Podoprikhin, Timur Garipov, Dmitry Vetrov, and Andrew Gordon Wilson. Averaging weights leads to wider optima and better generalization. arXiv preprint arXiv:1803.05407 , 2018
2018 arXiv
-
[75]
Oneformer: One transformer to rule universal image segmentation
Jitesh Jain, Jiachen Li, Mang Tik Chiu, Ali Hassani, Nikita Orlov, and Humphrey Shi. Oneformer: One transformer to rule universal image segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 2989–2998, 2023. 132
2023
-
[76]
Semask: Semantically masked transformers for semantic segmentation
Jitesh Jain, Anukriti Singh, Nikita Orlov, Zilong Huang, Jiachen Li, Steven Walton, and Humphrey Shi. Semask: Semantically masked transformers for semantic segmentation. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) Workshops , pages 752–761, 2023
2023
-
[77]
Face Perception, chapter 43
Nancy Kanwisher and Galit Yovel. Face Perception, chapter 43. John Wiley & Sons, Ltd, 2009
2009
-
[78]
Progressive growing of GANs for improved quality, stability, and variation
Tero Karras, Timo Aila, Samuli Laine, and Jaakko Lehtinen. Progressive growing of GANs for improved quality, stability, and variation. In International Conference on Learning Representations, 2018
2018
-
[79]
A style-based generator architecture for generative adversarial networks
Tero Karras, Samuli Laine, and Timo Aila. A style-based generator architecture for generative adversarial networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2019
2019
-
[80]
Training generative adversarial networks with limited data
Tero Karras, Miika Aittala, Janne Hellsten, Samuli Laine, Jaakko Lehtinen, and Timo Aila. Training generative adversarial networks with limited data. In Advances in Neural Information Processing Systems , pages 12104–12114. Curran Associates, Inc., 2020
2020
-
[81]
Analyzing and improving the image quality of stylegan
Tero Karras, Samuli Laine, Miika Aittala, Janne Hellsten, Jaakko Lehtinen, and Timo Aila. Analyzing and improving the image quality of stylegan. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020
2020
-
[82]
Alias-free generative adversarial networks
Tero Karras, Miika Aittala, Samuli Laine, Erik H¨ ark¨ onen, Janne Hellsten, Jaakko Lehtinen, and Timo Aila. Alias-free generative adversarial networks. In Advances in Neural Information Processing Systems , pages 852–863. Curran Associates, Inc., 2021. 133
2021
-
[83]
Design amortization for bayesian optimal experimental design
Noble Kennamer, Steven Walton, and Alexander Ihler. Design amortization for bayesian optimal experimental design. Proceedings of the AAAI Conference on Artificial Intelligence, 37(7):8220–8227, 2023
2023
-
[84]
Soft truncation: A universal training technique of score-based diffusion model for high precision score estimation
Dongjun Kim, Seungjae Shin, Kyungwoo Song, Wanmo Kang, and Il-Chul Moon. Soft truncation: A universal training technique of score-based diffusion model for high precision score estimation. In Proceedings of the 39th International Conference on Machine Learning , pages 11201–11...
2022
-
[85]
Kingma and Jimmy Ba
Diederik P. Kingma and Jimmy Ba. Adam: A method for stochastic optimization, 2017
2017
-
[86]
Glow: Generative Flow with Invertible 1x1 Convolutions
Durk P Kingma and Prafulla Dhariwal. Glow: Generative Flow with Invertible 1x1 Convolutions. In Advances in Neural Information Processing Systems . Curran Associates, Inc., 2018
2018
-
[87]
Kingma and Ruiqi Gao
Diederik P. Kingma and Ruiqi Gao. Understanding diffusion objectives as the elbo with simple data augmentation, 2023
2023
-
[88]
Kingma, Tim Salimans, Rafal Jozefowicz, Xi Chen, Ilya Sutskever, and Max Welling
Diederik P. Kingma, Tim Salimans, Rafal Jozefowicz, Xi Chen, Ilya Sutskever, and Max Welling. Improving variational inference with inverse autoregressive flow, 2017
2017
-
[89]
Big transfer (bit): General visual representation learning, 2020
Alexander Kolesnikov, Lucas Beyer, Xiaohua Zhai, Joan Puigcerver, Jessica Yung, Sylvain Gelly, and Neil Houlsby. Big transfer (bit): General visual representation learning, 2020
2020
-
[90]
Revealing the dark secrets of BERT
Olga Kovaleva, Alexey Romanov, Anna Rogers, and Anna Rumshisky. Revealing the dark secrets of BERT. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP...
2019
-
[91]
Learning multiple layers of features from tiny images.(2009), 2009
Alex Krizhevsky, Geoffrey Hinton, et al. Learning multiple layers of features from tiny images.(2009), 2009
2009
-
[92]
The role of imagenet classes in fr´ echet inception distance
Tuomas Kynk¨ a¨ anniemi, Tero Karras, Miika Aittala, Timo Aila, and Jaakko Lehtinen. The role of imagenet classes in fr´ echet inception distance. In The Eleventh International Conference on Learning Representations , 2023
2023
-
[93]
Howard, Wayne Hubbard, and Lawrence Jackel
Yann LeCun, Bernhard Boser, John Denker, Donnie Henderson, R. Howard, Wayne Hubbard, and Lawrence Jackel. Handwritten digit recognition with a back- propagation network. In Advances in Neural Information Processing Systems . Morgan-Kaufmann, 1989
1989
-
[94]
Gradient-based learning applied to document recognition
Yann LeCun, L´ eon Bottou, Yoshua Bengio, and Patrick Haffner. Gradient-based learning applied to document recognition. Proceedings of the IEEE, 86(11):2278– 2324, 2002
2002
-
[95]
Smooth manifolds
John M Lee and John M Lee. Smooth manifolds. Springer, 2003
2003
-
[96]
ViTGAN: Training GANs with vision transformers
Kwonjoon Lee, Huiwen Chang, Lu Jiang, Han Zhang, Zhuowen Tu, and Ce Liu. ViTGAN: Training GANs with vision transformers. In International Conference on Learning Representations, 2022
2022
-
[97]
SNIP: SINGLE- SHOT NETWORK PRUNING BASED ON CONNECTION SENSITIVITY
Namhoon Lee, Thalaiyasingam Ajanthan, and Philip Torr. SNIP: SINGLE- SHOT NETWORK PRUNING BASED ON CONNECTION SENSITIVITY. In International Conference on Learning Representations , 2019
2019
-
[98]
xformers: A modular and hackable transformer modelling library
Benjamin Lefaudeux, Francisco Massa, Diana Liskovich, Wenhan Xiong, Vittorio Caggiano, Sean Naren, Min Xu, Jieru Hu, Marta Tintore, Susan Zhang, Patrick Labatut, Daniel Haziza, Luca Wehrstedt, Jeremy Reizenstein, and Grigory Sizov. xformers: A modular and hackable transformer ...
2022
-
[99]
Demographic bias effects on face image synthesis
Roberto Leyva, Victor Sanchez, Gregory Epiphaniou, and Carsten Maple. Demographic bias effects on face image synthesis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops , pages 3818–3826, 2024
2024
-
[100]
Visualizing the loss landscape of neural nets
Hao Li, Zheng Xu, Gavin Taylor, Christoph Studer, and Tom Goldstein. Visualizing the loss landscape of neural nets. In Advances in Neural Information Processing Systems. Curran Associates, Inc., 2018
2018
-
[101]
Learning question classifiers
Xin Li and Dan Roth. Learning question classifiers. In COLING 2002: The 19th International Conference on Computational Linguistics , 2002
2002
-
[102]
M. Lichman. Uci machine learning repository, 2013
2013
-
[103]
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll´ ar, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In ECCV, 2014
2014
-
[104]
Yaron Lipman, Ricky T. Q. Chen, Heli Ben-Hamu, Maximilian Nickel, and Matthew Le. Flow matching for generative modeling. In The Eleventh International Conference on Learning Representations, 2023
2023
-
[105]
Yaron Lipman, Marton Havasi, Peter Holderrieth, Neta Shaul, Matt Le, Brian Karrer, Ricky T. Q. Chen, David Lopez-Paz, Heli Ben-Hamu, and Itai Gat. Flow matching guide and code, 2024
2024
-
[106]
Accelerate tarflow sampling with gs-jacobi iteration, 2025
Ben Liu and Zhen Qin. Accelerate tarflow sampling with gs-jacobi iteration, 2025
2025
-
[107]
Flow straight and fast: Learning to generate and transfer data with rectified flow
Xingchao Liu, Chengyue Gong, and Qiang Liu. Flow straight and fast: Learning to generate and transfer data with rectified flow. In The Eleventh International Conference on Learning Representations, 2023. 136
2023
-
[108]
Deep learning face attributes in the wild
Ziwei Liu, Ping Luo, Xiaogang Wang, and Xiaoou Tang. Deep learning face attributes in the wild. In Proceedings of International Conference on Computer Vision (ICCV) , 2015
2015
-
[109]
Swin transformer: Hierarchical vision transformer using shifted windows
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. In ICCV, pages 10012–10022, 2021
2021
-
[110]
Swin transformer v2: Scaling up capacity and resolution
Ze Liu, Han Hu, Yutong Lin, Zhuliang Yao, Zhenda Xie, Yixuan Wei, Jia Ning, Yue Cao, Zheng Zhang, Li Dong, Furu Wei, and Baining Guo. Swin transformer v2: Scaling up capacity and resolution. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (...
2022
-
[111]
A convnet for the 2020s
Zhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer, Trevor Darrell, and Saining Xie. A convnet for the 2020s. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2022
2022
-
[112]
Kan: Kolmogorov-arnold networks
Ziming Liu, Yixuan Wang, Sachin Vaidya, Fabian Ruehle, James Halverson, Marin Soljaˇ ci´ c, Thomas Y Hou, and Max Tegmark. Kan: Kolmogorov-arnold networks. arXiv preprint arXiv:2404.19756 , 2024
2024 arXiv
-
[113]
Decoupled weight decay regularization, 2019
Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization, 2019
2019
-
[114]
Minh-Thang Luong, Hieu Pham, and Christopher D. Manning. Effective approaches to attention-based neural machine translation, 2015
2015
-
[115]
Macow: Masked convolutional generative flow
Xuezhe Ma, Xiang Kong, Shanghang Zhang, and Eduard Hovy. Macow: Masked convolutional generative flow. In Advances in Neural Information Processing Systems. Curran Associates, Inc., 2019
2019
-
[116]
Learning word vectors for sentiment analysis
Andrew Maas, Raymond E Daly, Peter T Pham, Dan Huang, Andrew Y Ng, and Christopher Potts. Learning word vectors for sentiment analysis. In Proceedings 137 of the 49th annual meeting of the association for computational linguistics: Human language technologies, pages 142–150, 2011
2011
-
[117]
Martin, C
D. Martin, C. Fowlkes, D. Tal, and J. Malik. A database of human segmented natural images and its application to evaluating segmentation algorithms and measuring ecological statistics. In Proc. 8th Int’l Conf. Computer Vision , pages 416–423, 2001
2001
-
[118]
Which training methods for GANs do actually converge? In Proceedings of the 35th International Conference on Machine Learning , pages 3481–3490
Lars Mescheder, Andreas Geiger, and Sebastian Nowozin. Which training methods for GANs do actually converge? In Proceedings of the 35th International Conference on Machine Learning , pages 3481–3490. PMLR, 2018
2018
-
[119]
Are sixteen heads really better than one? In Advances in Neural Information Processing Systems
Paul Michel, Omer Levy, and Graham Neubig. Are sixteen heads really better than one? In Advances in Neural Information Processing Systems . Curran Associates, Inc., 2019
2019
-
[120]
Automated flower classification over a large number of classes
Maria-Elena Nilsback and Andrew Zisserman. Automated flower classification over a large number of classes. In 2008 Sixth Indian Conference on Computer Vision, Graphics and Image Processing , pages 722–729, 2008
2008
-
[121]
Masked autoregressive flow for density estimation, 2017
George Papamakarios, Theo Pavlakou, and Iain Murray. Masked autoregressive flow for density estimation, 2017
2017
-
[122]
On aliased resizing and surprising subtleties in gan evaluation
Gaurav Parmar, Richard Zhang, and Jun-Yan Zhu. On aliased resizing and surprising subtleties in gan evaluation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages 11410–11420, 2022
2022
-
[123]
Image transformer
Niki Parmar, Ashish Vaswani, Jakob Uszkoreit, Lukasz Kaiser, Noam Shazeer, Alexander Ku, and Dustin Tran. Image transformer. In International Conference on Machine Learning (ICML) , 2018. 138
2018
-
[124]
Pytorch: An imperative style, high-performance deep learning library, 2019
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas K¨ opf, Edward Yang, Zach DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu F...
2019
-
[125]
GloVe: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher Manning. GloVe: Global vectors for word representation. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP) , pages 1532–1543, Doha, Qatar,
2014
-
[126]
B. T. Polyak and A. B. Juditsky. Acceleration of stochastic approximation by averaging. SIAM Journal on Control and Optimization , 30(4):838–855, 1992
1992
-
[127]
Designing network design spaces
Ilija Radosavovic, Raj Prateek Kosaraju, Ross Girshick, Kaiming He, and Piotr Dollar. Designing network design spaces. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2020
2020
-
[128]
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21(140):1–67, 2020
2020
-
[129]
Stand-alone self-attention in vision models
Prajit Ramachandran, Niki Parmar, Ashish Vaswani, Irwan Bello, Anselm Levskaya, and Jon Shlens. Stand-alone self-attention in vision models. In Advances in Neural Information Processing Systems . Curran Associates, Inc., 2019
2019
-
[130]
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj¨ orn Ommer. High-resolution image synthesis with latent diffusion models. In 139 Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 10684–10695, 2022
2022
-
[131]
Projected gans converge faster
Axel Sauer, Kashyap Chitta, Jens M¨ uller, and Andreas Geiger. Projected gans converge faster. In Advances in Neural Information Processing Systems , pages 17480–17492. Curran Associates, Inc., 2021
2021
-
[132]
Stylegan-xl: Scaling stylegan to large diverse datasets
Axel Sauer, Katja Schwarz, and Andreas Geiger. Stylegan-xl: Scaling stylegan to large diverse datasets. In ACM SIGGRAPH 2022 conference proceedings, pages 1–10, 2022
2022
-
[133]
Are emergent abilities of large language models a mirage? In Advances in Neural Information Processing Systems, pages 55565–55581
Rylan Schaeffer, Brando Miranda, and Sanmi Koyejo. Are emergent abilities of large language models a mirage? In Advances in Neural Information Processing Systems, pages 55565–55581. Curran Associates, Inc., 2023
2023
-
[134]
Jermyn, Joe Benton, and Buck Shlegeris
Adam Scherlis, Kshitij Sachan, Adam S. Jermyn, Joe Benton, and Buck Shlegeris. Polysemanticity and capacity in neural networks. CoRR, abs/2210.01892, 2022
2022 arXiv
-
[135]
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch. Neural machine translation of rare words with subword units. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages 1715– 1725, Berlin, Germany, 2016. Associat...
2016
-
[136]
Flashattention-3: Fast and accurate attention with asynchrony and low- precision
Jay Shah, Ganesh Bikshandi, Ying Zhang, Vijay Thakkar, Pradeep Ramani, and Tri Dao. Flashattention-3: Fast and accurate attention with asynchrony and low- precision. arXiv preprint arXiv:2407.08608 , 2024
2024 arXiv
-
[137]
Deep inside convolutional networks: Visualising image classification models and saliency maps, 2014
Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman. Deep inside convolutional networks: Visualising image classification models and saliency maps, 2014. 140
2014
-
[138]
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D Manning, Andrew Y Ng, and Christopher Potts. Recursive deep models for semantic compositionality over a sentiment treebank. In Proceedings of the 2013 conference on empirical methods in natural language proce...
2013
-
[139]
Stanley, David B
Kenneth O. Stanley, David B. D’Ambrosio, and Jason Gauci. A hypercube-based encoding for evolving large-scale neural networks. Artificial Life, 15(2):185–212, 2009
2009
-
[140]
Cresswell, Rasa Hosseinzadeh, Yi Sui, Brendan Leigh Ross, Valentin Villecroze, Zhaoyan Liu, Anthony L
George Stein, Jesse C. Cresswell, Rasa Hosseinzadeh, Yi Sui, Brendan Leigh Ross, Valentin Villecroze, Zhaoyan Liu, Anthony L. Caterini, Eric Taylor, and Gabriel Loaiza-Ganem. Exposing flaws of generative model evaluation metrics and their unfair treatment of diffusion models. ...
2023
-
[141]
Revisiting unreasonable effectiveness of data in deep learning era
Chen Sun, Abhinav Shrivastava, Saurabh Singh, and Abhinav Gupta. Revisiting unreasonable effectiveness of data in deep learning era. In Proceedings of the IEEE International Conference on Computer Vision (ICCV) , 2017
2017
-
[142]
The bitter lesson, 2019
Richard Sutton. The bitter lesson, 2019
2019
-
[143]
Rethinking the inception architecture for computer vision
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna. Rethinking the inception architecture for computer vision. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2016
2016
-
[144]
SAN: Inducing metrizability of GAN with discriminative normalized linear layer
Yuhta Takida, Masaaki Imaizumi, Takashi Shibuya, Chieh-Hsin Lai, Toshimitsu Uesaka, Naoki Murata, and Yuki Mitsufuji. SAN: Inducing metrizability of GAN with discriminative normalized linear layer. In The Twelfth International Conference on Learning Representations, 2024. 141
2024
-
[145]
An analysis of attention mechanisms: The case of word sense disambiguation in neural machine translation
Gongbo Tang, Rico Sennrich, and Joakim Nivre. An analysis of attention mechanisms: The case of word sense disambiguation in neural machine translation. In Proceedings of the Third Conference on Machine Translation: Research Papers , pages 26–35, Brussels, Belgium, 2018. Associ...
2018
-
[146]
Relay diffusion: Unifying diffusion process across resolutions for image synthesis, 2023
Jiayan Teng, Wendi Zheng, Ming Ding, Wenyi Hong, Jianqiao Wangni, Zhuoyi Yang, and Jie Tang. Relay diffusion: Unifying diffusion process across resolutions for image synthesis, 2023
2023
-
[147]
Antonio Torralba, Rob Fergus, and William T. Freeman. 80 million tiny images: A large data set for nonparametric object and scene recognition. IEEE Transactions on Pattern Analysis and Machine Intelligence , 30(11):1958–1970, 2008
1958
-
[148]
Fixing the train- test resolution discrepancy
Hugo Touvron, Andrea Vedaldi, Matthijs Douze, and Herve Jegou. Fixing the train- test resolution discrepancy. In Advances in Neural Information Processing Systems . Curran Associates, Inc., 2019
2019
-
[149]
Training data-efficient image transformers & distillation through attention
Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Herve Jegou. Training data-efficient image transformers & distillation through attention. In Proceedings of the 38th International Conference on Machine Learning , pages 10347–10357. PMLR, 2021
2021
-
[150]
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, L ukasz Kaiser, and Illia Polosukhin. Attention is all you need. In NeurIPS, 2017
2017
-
[151]
Scaling local self-attention for parameter efficient visual backbones, 2021
Ashish Vaswani, Prajit Ramachandran, Aravind Srinivas, Niki Parmar, Blake Hechtman, and Jonathon Shlens. Scaling local self-attention for parameter efficient visual backbones, 2021. 142
2021
-
[152]
A 0-shot self-attention mechanism for accelerated diagonal attention
Mario Viti, Nadiya Shvai, Arcadi Llanza, and Amir Nakib. A 0-shot self-attention mechanism for accelerated diagonal attention. In Proceedings of the Winter Conference on Applications of Computer Vision (WACV) , pages 7308–7315, 2025
2025
-
[153]
Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned
Elena Voita, David Talbot, Fedor Moiseev, Rico Sennrich, and Ivan Titov. Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics , pages 5797...
2019
-
[154]
Isomorphism, normalizing flows, and density estimation: Preserving relationships between data, 2022
Steven Walton. Isomorphism, normalizing flows, and density estimation: Preserving relationships between data, 2022
2022
-
[155]
Training compact transformers from scratch in 30 minutes with pytorch
Steven Walton, Ali Hassani, Abulikemu Abuduweili, and Humphrey Shi. Training compact transformers from scratch in 30 minutes with pytorch. medium.com/pytorch, 2021
2021
-
[156]
Efficient image generation with variadic attention heads
Steven Walton, Ali Hassani, Xingqian Xu, Zhangyang Wang, and Humphrey Shi. Efficient image generation with variadic attention heads. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, 2025
2025
-
[157]
Distilling normalizing flows
Steven Walton, Valeriy Klyukin, Maksim Artemev, Denis Derkach, Nikita Orlov, and Humphrey Shi. Distilling normalizing flows. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops , 2025
2025
-
[158]
Diffusion-GAN: Training GANs with diffusion
Zhendong Wang, Huangjie Zheng, Pengcheng He, Weizhu Chen, and Mingyuan Zhou. Diffusion-GAN: Training GANs with diffusion. In The Eleventh International Conference on Learning Representations, 2023. 143
2023
-
[159]
Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus
Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, Ed H. Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus. Emergent abilities of large language models. T...
2022
-
[160]
Pytorch image models, 2019
Ross Wightman. Pytorch image models, 2019
2019
-
[161]
Resnet strikes back: An improved training procedure in timm, 2021
Ross Wightman, Hugo Touvron, and Herv´ e J´ egou. Resnet strikes back: An improved training procedure in timm, 2021
2021
-
[162]
Convnext v2: Co-designing and scaling convnets with masked autoencoders
Sanghyun Woo, Shoubhik Debnath, Ronghang Hu, Xinlei Chen, Zhuang Liu, In So Kweon, and Saining Xie. Convnext v2: Co-designing and scaling convnets with masked autoencoders. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages 1613...
2023
-
[163]
Efficient streaming language models with attention sinks, 2024
Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han, and Mike Lewis. Efficient streaming language models with attention sinks, 2024
2024
-
[164]
Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms, 2017
Han Xiao, Kashif Rasul, and Roland Vollgraf. Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms, 2017
2017
-
[165]
Data-efficient instance generation from instance discrimination
Ceyuan Yang, Yujun Shen, Yinghao Xu, and Bolei Zhou. Data-efficient instance generation from instance discrimination. In Advances in Neural Information Processing Systems, pages 9378–9390. Curran Associates, Inc., 2021
2021
-
[166]
Lsun: Construction of a large-scale image dataset using deep learning with humans in the loop
Fisher Yu, Yinda Zhang, Shuran Song, Ari Seff, and Jianxiong Xiao. Lsun: Construction of a large-scale image dataset using deep learning with humans in the loop. arXiv preprint arXiv:1506.03365 , 2015
2015 arXiv
-
[167]
Cutmix: Regularization strategy to train strong classifiers with 144 localizable features
Sangdoo Yun, Dongyoon Han, Seong Joon Oh, Sanghyuk Chun, Junsuk Choe, and Youngjoon Yoo. Cutmix: Regularization strategy to train strong classifiers with 144 localizable features. In Proceedings of the IEEE/CVF International Conference on Computer Vision , pages 6023–6032, 2019
2019
-
[168]
Normalizing flows are capable generative models, 2024
Shuangfei Zhai, Ruixiang Zhang, Preetum Nakkiran, David Berthelot, Jiatao Gu, Huangjie Zheng, Tianrong Chen, Miguel Angel Bautista, Navdeep Jaitly, and Josh Susskind. Normalizing flows are capable generative models, 2024
2024
-
[169]
Styleswin: Transformer-based gan for high-resolution image generation
Bowen Zhang, Shuyang Gu, Bo Zhang, Jianmin Bao, Dong Chen, Fang Wen, Yong Wang, and Baining Guo. Styleswin: Transformer-based gan for high-resolution image generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages 11304–11314, 2022
2022
-
[170]
mixup: Beyond empirical risk minimization
Hongyi Zhang, Moustapha Cisse, Yann N Dauphin, and David Lopez-Paz. mixup: Beyond empirical risk minimization. arXiv preprint arXiv:1710.09412 , 2017
2017 arXiv
-
[171]
Self-attention generative adversarial networks
Han Zhang, Ian Goodfellow, Dimitris Metaxas, and Augustus Odena. Self-attention generative adversarial networks. In Proceedings of the 36th International Conference on Machine Learning , pages 7354–7363. PMLR, 2019
2019
-
[172]
Character-level convolutional networks for text classification
Xiang Zhang, Junbo Zhao, and Yann LeCun. Character-level convolutional networks for text classification. arXiv preprint arXiv:1509.01626 , 2015
2015 arXiv
-
[173]
Nested hierarchical transformer: Towards accurate, data-efficient and interpretable visual understanding
Zizhao Zhang, Han Zhang, Long Zhao, Ting Chen, Sercan ¨O Arik, and Tomas Pfister. Nested hierarchical transformer: Towards accurate, data-efficient and interpretable visual understanding. In Proceedings of the AAAI Conference on Artificial Intelligence, pages 3417–3425, 2022
2022
-
[174]
Improved transformer for high-resolution gans
Long Zhao, Zizhao Zhang, Ting Chen, Dimitris Metaxas, and Han Zhang. Improved transformer for high-resolution gans. In Advances in Neural Information Processing Systems, pages 18367–18380. Curran Associates, Inc., 2021. 145
2021
-
[175]
Improved consistency regularization for gans
Zhengli Zhao, Sameer Singh, Honglak Lee, Zizhao Zhang, Augustus Odena, and Han Zhang. Improved consistency regularization for gans. In Proceedings of the AAAI conference on artificial intelligence , pages 11033–11041, 2021
2021
-
[176]
Fast training of diffusion models with masked transformers
Hongkai Zheng, Weili Nie, Arash Vahdat, and Anima Anandkumar. Fast training of diffusion models with masked transformers. In Transactions on Machine Learning Research (TMLR), 2024
2024
-
[177]
Zheng, Tong Xu, and Enhong Chen
Hui Zhong, Zaiyi Chen, Chuan Qin, Zai Huang, Vincent W. Zheng, Tong Xu, and Enhong Chen. Adam revisited: a weighted past gradients perspective. Frontiers of Computer Science, 14(5), 2020
2020
-
[178]
Random erasing data augmentation
Zhun Zhong, Liang Zheng, Guoliang Kang, Shaozi Li, and Yi Yang. Random erasing data augmentation. In Proceedings of the AAAI Conference on Artificial Intelligence, pages 13001–13008, 2020
2020
-
[179]
Scene parsing through ade20k dataset
Bolei Zhou, Hang Zhao, Xavier Puig, Sanja Fidler, Adela Barriuso, and Antonio Torralba. Scene parsing through ade20k dataset. In CVPR, 2017
2017
-
[180]
Unpaired image-to- image translation using cycle-consistent adversarial networks
Jun-Yan Zhu, Taesung Park, Phillip Isola, and Alexei A Efros. Unpaired image-to- image translation using cycle-consistent adversarial networks. In Computer Vision (ICCV), 2017 IEEE International Conference on , 2017. 146
2017
-
[2014]
Association for Computational Linguistics
-
[2022]
https://transformer-circuits.pub/2022/solu/index.html
2022
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.