REVIEW 3 major objections 4 minor 300 references
The Good, The Efficient and the Inductive Biases: Exploring Efficiency in Deep Learning Through the Use of Inductive Biases
T0 review · 3 major / 4 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read This dissertation argues that two inductive biases—continuous modeling and symmetry preservation—improve deep learning efficiency across compute, data, parameters, and design.
desk verdict A candid, well-organized dissertation of peer-reviewed work; the six-axis taxonomy is useful, but the 'design efficiency' claim is asserted, not measured. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central objects are the continuous kernel parameterization—an MLP or multiplicative filter network, such as a SIREN (a multilayer perceptron with sine activations) or a MAGNet (a multiplicative filter network built from anisotropic Gabor functions), that maps relative coordinates to kernel values—and the Gaussian mask mechanism in FlexConv that makes kernel size differentiable. On the symmetry side, the central objects are group convolutions, group-equivariant attention, and partial group convolutions, which encode translation, rotation, and scale symmetries through weight sharing. These mechanisms carry the efficiency claims: the continuous-kernel family removes the dependence of parameter count on context length, supports resolution transfer and irregular data, and turns architecture search into gradient-based optimization; the symmetry-preserving family reduces the data and parameters needed while preserving prediction consistency under input symmetries.
What would settle it
Run the proposed continuous and equivariant models on a shared, larger-scale benchmark suite, for instance ImageNet-scale image classification and long-context language modeling, with matched parameter, compute, and data budgets alongside the thesis's main baselines; if the claimed compute, data, parameter, or design-efficiency advantages disappear or reverse, the central claim is falsified.
Extended reading notes
Core claim
The dissertation's central claim is that two inductive biases—continuous modeling and symmetry preservation—are broadly efficiency-improving design principles for deep learning. Continuous modeling parameterizes operations, notably convolutional kernels, as functions of continuous coordinates, so a fixed parameter budget yields arbitrarily large, resolution-agnostic kernels and makes architectural components learnable by gradient descent. Symmetry preservation builds transformations that respect data symmetries, improving data and parameter efficiency through weight sharing, though it can raise computational cost. The thesis evaluates these claims through a series of contributed methods and concludes that the biases yield gains in compute, data, parameter, and design efficiency, with acknowledged trade-offs.
Load-bearing premise
The load-bearing premise is that the efficiency gains reported on small benchmarks and against the specific baselines chosen in the author's own papers are representative of how these methods would behave in real-world, large-scale use, and that design efficiency is a well-defined quantity even though the thesis gives no metric for it.
Editorial extensions
If this is right
- One network architecture can be trained at low resolution and deployed at higher resolutions, or on irregularly sampled data, without redesigning the model.
- Convolutional layers can model entire sequences with global kernels under a fixed parameter budget, matching or beating recurrent and attention models on benchmark sequential tasks.
- Kernel sizes, layer widths, downsampling locations, and network depth become learnable by backpropagation, removing part of the manual architecture-design burden.
- Symmetry-preserving models achieve higher accuracy per training example and per parameter, but pay a computational overhead; partial equivariance can soften this trade-off by letting the model decide how much symmetry to enforce.
- Point-cloud pipelines can map irregular data to compact grids and then use standard grid convolutions, improving scalability while preserving competitive accuracy.
Reading between the lines
- A direct testable extension would combine continuous kernel parameterization with symmetry preservation, for instance equivariant continuous kernels for point clouds, and measure whether the two efficiency gains compound or partly cancel.
- The thesis does not define a quantitative metric for design efficiency; a reader who wants to verify that claim would need to operationalize it, for example as human hours or compute needed to reach a target accuracy on a new dataset.
- The small-benchmark evidence suggests a high-risk prediction: if scaled to large models and datasets, the fixed-parameter continuous kernel advantage may shrink relative to learned sparse or recurrent alternatives, because storing full-resolution kernel responses grows with input length.
- The symmetry part implies that in domains where symmetries are only approximate, fully equivariant models may underperform learnable partial equivariance; this suggests a broader principle of treating inductive bias strength as a tunable hyperparameter rather than a binary choice.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript is a PhD dissertation, posted on arXiv, which argues that two inductive biases—continuous modeling and symmetry preservation—can substantially improve the efficiency of deep learning along compute, data, parameter, and design axes. The first part (Chapters 2–6) presents continuous-kernel and continuous-architecture methods (CKConv, CCNN, gridification, FlexConv, DNArch), while the second part (Chapters 7–11) presents equivariant and partially equivariant architectures. The dissertation contributes an efficiency taxonomy (Sec. 1.2.1), a summary table of per-chapter efficiency contributions (Table 1.1), and a concluding chapter with explicit limitations. Because each chapter is based on previously published papers by the author, the central claim is a synthesis-level claim about efficiency rather than a report of new experiments. The review copy provided to me is truncated after Sec. 5.4.1, so Part II is assessed here only through its abstract, chapter summaries, and Table 1.1.
Significance. If the synthesis claim were adequately supported, the dissertation would provide a useful organizing perspective on how continuous modeling and symmetry preservation translate into resource savings. Several individual contributions have already survived peer review (ICLR, ICML, NeurIPS, TMLR), and the thesis contains genuinely useful analytic results, including the resolution-change formula in Eq. (2.5) and the alias-frequency bound in Eqs. (D.12)–(D.16). The thesis is also unusually candid: Ch. 12.1 explicitly acknowledges the computational costs of global convolutions, the input-dependence of long convolutional models, and the need for symmetry pre-specification. What is missing is a quantitative, unified efficiency protocol: the central claim is stated in the abstract and Table 1.1, but no chapter measures the resources it claims to save, and no independent verification of the efficiency attributions is provided. The value of the dissertation as a structured compendium of the author's contributions is clear; the value of its efficiency framing as a falsifiable research claim is not yet established.
major comments (3)
- [Sec. 1.2.1, Table 1.1, Chs. 2, 5, 6] Design efficiency is defined in Sec. 1.2.1 as 'the resources in terms of compute, human hours, memory, experimentation, etc. needed to design a high-performing architecture,' and Table 1.1 assigns design-efficiency checkmarks to Chapters 2, 3, 5, 6, and 11. However, no chapter reports a measurement of any of these quantities. For example, Ch. 5 claims that FlexConv relieves users from pre-specifying kernel sizes, and Ch. 6 claims that DNArch learns kernel sizes, widths, depths, and downsampling positions by backpropagation, but neither measures the human effort, number of configurations explored, or compute required to reach a target accuracy relative to manual design or a discrete NAS baseline. The abstract's claim of 'substantial benefits' for design efficiency is therefore non-falsifiable from the evidence in this manuscript. The dissertation's own limitation statement in Ch. 12.1 is qualitative and does not repair this gap, and the acknowledged sensitivity of CKConv to the hyperparameter ω0 (Sec. 2.6) is an example of a design cost that is never included in the accounting.
- [Sec. 2.5, Table 2.3] The compute-efficiency claim for Chapter 2 relies in part on comparisons that do not separate model capacity from training budget. In Table 2.3, the neural-ODE baselines on SC raw are reported with accuracy ≈10.0, and Sec. 2.5 states that this result is obtained after a single training epoch under a computational budget matched to the CKCNN. The text further states that 'CKCNNs trained on SC raw are able to outperform several Neural ODE models trained on the preprocessed data (SC).' This second comparison uses different input representations, and the first comparison gives the baselines no opportunity to show standard accuracy-vs-epoch behavior. To support the chapter's compute-efficiency checkmark, the thesis should provide accuracy-versus-compute curves or matched-budget comparisons in which the baselines are also trained to convergence or clearly shown to be unable to reach it within a much larger budget.
- [Sec. 1.4, Table 1.1] The synthesis-level efficiency claim is supported almost entirely by the author's own previously published papers. Each chapter in Sec. 1.4 is based on a paper authored or co-authored by the dissertation author, and the efficiency attributions in Table 1.1 are categorical checks rather than quantitative results. This is not a logical circularity in any derivation, but it does mean that the thesis's central claim is a summary of the author's published claims rather than an independent evaluation of them. If the dissertation is intended as a research synthesis rather than a compendium, it needs a dedicated quantitative meta-analysis over the constituent papers: a common set of efficiency metrics, matched training budgets, and effect sizes that would let a reader see whether 'substantial benefits' is supported across chapters and modalities.
minor comments (4)
- [Sec. 1.2.1] The sentence 'Architectures that required lower lower overall inversions are more financially efficient' contains a duplicated 'lower' and the likely intended word is 'investments' rather than 'inversions'.
- [Table 1.1 footnote] In the table's footnote, 'financial efficiently' should be 'financial efficiency'.
- [Sec. 4.6] The concluding paragraph says gridification allows 'performing neural operations in there'; this should be 'in the grid' or 'on the grid' for clarity.
- [Sec. 2.6] The admission that CKCNNs are very susceptible to the selection of ω0, and that finding a good value induces an important cost in hyperparameter search, is honest and useful; this cost should be reported quantitatively because it bears directly on the design-efficiency claims made for the method.
Circularity Check
No circular derivation found; the efficiency claims are supported by chapter-level external benchmarks, with self-referential synthesis and unmeasured 'design efficiency' as completeness risks rather than circularity.
full rationale
I walked the claimed derivation chain chapter by chapter. The abstract's conclusion that continuous modeling and symmetry preservation improve efficiency is a synthesis of the author's own previously published papers, but each constituent chapter reports experiments against external benchmarks (sMNIST, pMNIST, sCIFAR10, CIFAR-10, ModelNet40, PhysioNet, LRA, etc.) and independent complexity analyses, so the evidence is not merely a self-citation loop. The places where circularity would most plausibly hide do not actually reduce to inputs: in Ch. 2, Eq. (2.5), the resolution-change relation, is a derived consequence of sampling a continuous kernel, and the parameter-efficiency numbers are direct arithmetic comparisons against discrete global kernels; in Ch. 5, the alias-free frequency bound (Eqs. D.12-D.16) is derived analytically from the MAGNet/MFN basis representation and then used as a regularizer, with the subsequent resolution-generalization results being measurements rather than restatements of the bound; in Ch. 3, the 'necessary and sufficient' claim is an architectural design argument rather than an equation whose output is its input. The self-referential aspects are real but not logical circularity: Table 1.1 assigns 'design efficiency' checkmarks without a unified measurement protocol, and the abstract's 'substantial benefits' for design efficiency is therefore not quantitatively pinned down; the one-epoch neural-ODE baselines in Sec. 2.5 are an unequal comparison, not a fitted parameter renamed as a prediction. I found no specific step where a derived quantity is equal, by construction or by a self-citation chain, to the quantity it is supposed to predict.
Assumptions & free parameters
free parameters (3)
- omega0 (SIREN frequency prior) in CKConv kernels =
not stated as single value; varies in [1,100]
- Grid resolution in gridification =
chosen per dataset so grid points approx equal point cloud size (e.g., 10x10x10 for N=1000)
- Gaussian mask parameters (mu, sigma) in FlexConv =
learned during training
assumptions (3)
- domain assumption The six efficiency axes in Sec. 1.2.1 are disjoint and exhaustive; environmental and financial efficiency are derivable from the other axes.
- domain assumption Kernel generator networks (MLPs, SIRENs, MAGNets) can approximate the convolutional kernels required for the tasks considered.
- domain assumption The benchmarks used (MNIST, CIFAR, ModelNet40, LRA, etc.) are representative of the efficiency benefits claimed for real-world deep learning.
invented entities (1)
-
Design efficiency as a distinct efficiency axis
Cite this review
Pith. "Pith review of The Good, The Efficient and the Inductive Biases: Exploring Efficiency in Deep Learning Through the Use of Inductive Biases." pith.science (2026). https://pith.science/paper/2SXYPQJK
@misc{pith2026241109827,
author = {Pith},
title = {Pith review of: The Good, The Efficient and the Inductive Biases: Exploring Efficiency in Deep Learning Through the Use of Inductive Biases},
year = {2026},
howpublished = {\url{https://pith.science/paper/2SXYPQJK}},
note = {Machine review of arXiv:2411.09827}
}
read the original abstract
The emergence of Deep Learning has marked a profound shift in machine learning, driven by numerous breakthroughs achieved in recent years. However, as Deep Learning becomes increasingly present in everyday tools and applications, there is a growing need to address unresolved challenges related to its efficiency and sustainability. This dissertation delves into the role of inductive biases -- particularly, continuous modeling and symmetry preservation -- as strategies to enhance the efficiency of Deep Learning. It is structured in two main parts. The first part investigates continuous modeling as a tool to improve the efficiency of Deep Learning algorithms. Continuous modeling involves the idea of parameterizing neural operations in a continuous space. The research presented here demonstrates substantial benefits for the (i) computational efficiency -- in time and memory, (ii) the parameter efficiency, and (iii) design efficiency -- the complexity of designing neural architectures for new datasets and tasks. The second focuses on the role of symmetry preservation on Deep Learning efficiency. Symmetry preservation involves designing neural operations that align with the inherent symmetries of data. The research presented in this part highlights significant gains both in data and parameter efficiency through the use of symmetry preservation. However, it also acknowledges a resulting trade-off of increased computational costs. The dissertation concludes with a critical evaluation of these findings, openly discussing their limitations and proposing strategies to address them, informed by literature and the author insights. It ends by identifying promising future research avenues in the exploration of inductive biases for efficiency, and their wider implications for Deep Learning.
Figures
Figures from the paper (69 more)
Reference graph
Works this paper leans on
-
[1]
Convolutional neural networks for speech recognition
Ossama Abdel-Hamid, Abdel-rahman Mohamed, Hui Jiang, Li Deng, Gerald Penn, and Dong Yu. Convolutional neural networks for speech recognition. IEEE/ACM Transactions on audio, speech, and language processing , 22(10):1533–1545, 2014
2014
-
[2]
End-to-end en- vironmental sound classification using a 1d convolutional neural network
Sajjad Abdoli, Patrick Cardinal, and Alessandro Lameiras Koerich. End-to-end en- vironmental sound classification using a 1d convolutional neural network. Expert Systems with Applications, 136:252–263, 2019
2019
-
[3]
Applications of the generalized fourier transform in numerical linear algebra
Krister ˚Ahlander and Hans Munthe-Kaas. Applications of the generalized fourier transform in numerical linear algebra. BIT Numerical Mathematics, 45(4):819–850, 2005
2005
-
[4]
Noether networks: meta-learning useful conserved quantities
Ferran Alet, Dylan Doblar, Allan Zhou, Josh Tenenbaum, Kenji Kawaguchi, and Chelsea Finn. Noether networks: meta-learning useful conserved quantities. Ad- vances in Neural Information Processing Systems, 34, 2021
2021
-
[5]
Statistical applications for equivariant matrices
SH Alkarni. Statistical applications for equivariant matrices. International Journal of Mathematics and Mathematical Sciences, 25(1):53–61, 2001
2001
-
[6]
Deep scattering spectrum
Joakim And ´en and St´ephane Mallat. Deep scattering spectrum. IEEE Transactions on Signal Processing, 62(16):4114–4128, 2014
2014
-
[7]
Unitary evolution recurrent neural networks
Martin Arjovsky, Amar Shah, and Yoshua Bengio. Unitary evolution recurrent neural networks. In International Conference on Machine Learning, pages 1120–1128, 2016
2016
-
[8]
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton. Layer normalization. arXiv preprint arXiv:1607.06450, 2016
arXiv 2016
Show all 300 references
-
[9]
The uea multivariate time series classification archive, 2018
Anthony Bagnall, Hoang Anh Dau, Jason Lines, Michael Flynn, James Large, Aaron Bostrom, Paul Southam, and Eamonn Keogh. The uea multivariate time series classification archive, 2018. arXiv preprint arXiv:1811.00075, 2018
2018 arXiv
-
[10]
Neural machine trans- lation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. Neural machine trans- lation by jointly learning to align and translate. In Yoshua Bengio and Yann Le- Cun, editors, 3rd International Conference on Learning Representations, ICLR 2015, 191 192 BIBLIOGRAPHY San Diego, CA, USA...
2015 arXiv
-
[11]
An empirical evaluation of generic convolutional and recurrent networks for sequence modeling
Shaojie Bai, J Zico Kolter, and Vladlen Koltun. An empirical evaluation of generic convolutional and recurrent networks for sequence modeling. arXiv preprint arXiv:1803.01271, 2018
2018 arXiv
-
[12]
Trellis networks for sequence mod- eling
Shaojie Bai, J Zico Kolter, and Vladlen Koltun. Trellis networks for sequence mod- eling. arXiv preprint arXiv:1810.06682, 2018
2018 arXiv
-
[13]
Mad max: Affine spline insights into deep learning
Randall Balestriero and Richard Baraniuk. Mad max: Affine spline insights into deep learning. arXiv preprint arXiv:1805.06576, 2018
2018 arXiv
-
[14]
The quickhull algo- rithm for convex hulls
C Bradford Barber, David P Dobkin, and Hannu Huhdanpaa. The quickhull algo- rithm for convex hulls. ACM Transactions on Mathematical Software (TOMS), 22(4): 469–483, 1996
1996
-
[15]
Ai in healthcare: Ethical and privacy challenges
Ivana Bartoletti. Ai in healthcare: Ethical and privacy challenges. In Artificial Intelligence in Medicine: 17th Conference on Artificial Intelligence in Medicine, AIME 2019, Poznan, Poland, June 26–29, 2019, Proceedings 17, pages 7–10. Springer, 2019
2019
-
[16]
Diagnostic assessment of deep learning algorithms for detection of lymph node metastases in women with breast cancer
Babak Ehteshami Bejnordi, Mitko Veta, Paul Johannes Van Diest, Bram Van Gin- neken, Nico Karssemeijer, Geert Litjens, Jeroen AWM Van Der Laak, Meyke Hermsen, Quirine F Manson, Maschenka Balkenhol, et al. Diagnostic assessment of deep learning algorithms for detection of lymph ...
2017
-
[17]
B-spline {cnn}s on lie groups
Erik J Bekkers. B-spline {cnn}s on lie groups. In International Conference on Learning Representations , 2020. URL https://openreview.net/forum?id= H1gBhkBFDH
2020
-
[18]
Roto-translation covariant convolutional networks for medical image analysis
Erik J Bekkers, Maxime W Lafarge, Mitko Veta, Koen AJ Eppenhof, Josien PW Pluim, and Remco Duits. Roto-translation covariant convolutional networks for medical image analysis. In International Conference on Medical Image Computing and Computer-Assisted Intervention, pages 440–...
2018
-
[19]
Fast, expressive se (n) equivariant networks through weight- sharing in position-orientation space
Erik J Bekkers, Sharvaree Vadgama, Rob D Hesselink, Putri A van der Linden, and David W Romero. Fast, expressive se (n) equivariant networks through weight- sharing in position-orientation space. arXiv preprint arXiv:2310.02970, 2023
-
[20]
At- tention augmented convolutional networks
Irwan Bello, Barret Zoph, Ashish Vaswani, Jonathon Shlens, and Quoc V Le. At- tention augmented convolutional networks. arXiv preprint arXiv:1904.09925, 2019
1904 arXiv
-
[21]
Understanding and simplifying one-shot architecture search
Gabriel Bender, Pieter-Jan Kindermans, Barret Zoph, Vijay Vasudevan, and Quoc Le. Understanding and simplifying one-shot architecture search. In International conference on machine learning, pages 550–559. PMLR, 2018
2018
-
[22]
Learning long-term depen- BIBLIOGRAPHY 193 dencies with gradient descent is difficult
Yoshua Bengio, Patrice Simard, and Paolo Frasconi. Learning long-term depen- BIBLIOGRAPHY 193 dencies with gradient descent is difficult. IEEE transactions on neural networks , 5 (2):157–166, 1994
1994
-
[23]
Estimating or propa- gating gradients through stochastic neurons for conditional computation
Yoshua Bengio, Nicholas L ´eonard, and Aaron Courville. Estimating or propa- gating gradients through stochastic neurons for conditional computation. arXiv preprint arXiv:1308.3432, 2013
2013 arXiv
-
[24]
A comprehensive survey on hardware-aware neural architecture search
Hadjer Benmeziane, Kaoutar El Maghraoui, Hamza Ouarnoughi, Smail Niar, Mar- tin Wistuba, and Naigang Wang. A comprehensive survey on hardware-aware neural architecture search. arXiv preprint arXiv:2101.09336, 2021
2021 arXiv
-
[25]
Learning invariances in neural networks from training data
Gregory Benton, Marc Finzi, Pavel Izmailov, and Andrew G Wilson. Learning invariances in neural networks from training data. Advances in Neural Information Processing Systems, 33:17605–17616, 2020
2020
-
[26]
Stabi- lizing darts with amended gradient estimation on architectural parameters
Kaifeng Bi, Changping Hu, Lingxi Xie, Xin Chen, Longhui Wei, and Qi Tian. Stabi- lizing darts with amended gradient estimation on architectural parameters. arXiv preprint arXiv:1910.11831, 2019
1910 arXiv
-
[27]
Spectrotemporal resolution tradeoff in auditory processing as revealed by human auditory brainstem re- sponses and psychophysical indices
Gavin M Bidelman and Ameenuddin Syed Khaja. Spectrotemporal resolution tradeoff in auditory processing as revealed by human auditory brainstem re- sponses and psychophysical indices. Neuroscience letters, 572:53–57, 2014
2014
-
[28]
Recognition-by-components: a theory of human image under- standing
Irving Biederman. Recognition-by-components: a theory of human image under- standing. Psychological review, 94(2):115, 1987
1987
-
[29]
Experiment tracking with weights and biases, 2020
Lukas Biewald. Experiment tracking with weights and biases, 2020. URL https: //www.wandb.com/. Software available from wandb.com
2020
-
[30]
The role of temporal structure in human vision
Randolph Blake and Sang-Hun Lee. The role of temporal structure in human vision. Behavioral and cognitive neuroscience reviews, 4(1):21–42, 2005
2005
-
[31]
Lorentz group equivariant neural network for particle physics
Alexander Bogatskiy, Brandon Anderson, Jan Offermann, Marwah Roussi, David Miller, and Risi Kondor. Lorentz group equivariant neural network for particle physics. In International Conference on Machine Learning , pages 992–1002. PMLR, 2020
2020
-
[32]
Smash: one-shot model architecture search through hypernetworks
Andrew Brock, Theodore Lim, James M Ritchie, and Nick Weston. Smash: one-shot model architecture search through hypernetworks. arXiv preprint arXiv:1708.05344, 2017
2017 arXiv
-
[33]
Recognition by children of inverted photos of faces
Richard M Brooks and Alvin G Goldstein. Recognition by children of inverted photos of faces. Child Development, 1963
1963
-
[34]
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. Advances in neural informa- tion processing systems, 33:1877–1901, 2020. 194 ...
1901
-
[35]
Recognizing objects and faces
Vicki Bruce and Glyn W Humphreys. Recognizing objects and faces. Visual cogni- tion, 1(2-3):141–180, 1994
1994
-
[36]
Invariant scattering convolution networks
Joan Bruna and St ´ephane Mallat. Invariant scattering convolution networks. IEEE transactions on pattern analysis and machine intelligence, 35(8):1872–1886, 2013
2013
-
[37]
Proxylessnas: Direct neural architecture search on target task and hardware
Han Cai, Ligeng Zhu, and Song Han. Proxylessnas: Direct neural architecture search on target task and hardware. arXiv preprint arXiv:1812.00332, 2018
2018 arXiv
-
[38]
Once-for- all: Train one network and specialize it for efficient deployment
Han Cai, Chuang Gan, Tianzhe Wang, Zhekai Zhang, and Song Han. Once-for- all: Train one network and specialize it for efficient deployment. arXiv preprint arXiv:1908.09791, 2019
1908 arXiv
-
[39]
Gcnet: Non- local networks meet squeeze-excitation networks and beyond
Yue Cao, Jiarui Xu, Stephen Lin, Fangyun Wei, and Han Hu. Gcnet: Non- local networks meet squeeze-excitation networks and beyond. arXiv preprint arXiv:1904.11492, 2019
1904 arXiv
-
[40]
The concept of group and the theory of perception
Ernst Cassirer. The concept of group and the theory of perception. Philosophy and phenomenological research, 5(1):1–36, 1944
1944
-
[41]
A program to build e (n)- equivariant steerable cnns
Gabriele Cesa, Leon Lang, and Maurice Weiler. A program to build e (n)- equivariant steerable cnns. In International Conference on Learning Representations , 2021
2021
-
[42]
Antisymmetri- crnn: A dynamical system view on recurrent neural networks
Bo Chang, Minmin Chen, Eldad Haber, and Ed H Chi. Antisymmetri- crnn: A dynamical system view on recurrent neural networks. arXiv preprint arXiv:1902.09689, 2019
1902 arXiv
-
[43]
Principled weight initialization for hypernetworks
Oscar Chang, Lampros Flokas, and Hod Lipson. Principled weight initialization for hypernetworks. In International Conference on Learning Representations , 2020. URL https://openreview.net/forum?id=H1lma24tPB
2020
-
[44]
Di- lated recurrent neural networks
Shiyu Chang, Yang Zhang, Wei Han, Mo Yu, Xiaoxiao Guo, Wei Tan, Xiaodong Cui, Michael Witbrock, Mark A Hasegawa-Johnson, and Thomas S Huang. Di- lated recurrent neural networks. In Advances in neural information processing sys- tems, pages 77–87, 2017
2017
-
[45]
Learning augmentation distributions using transformed risk mini- mization
Evangelos Chatzipantazis, Stefanos Pertigkiozoglou, Edgar Dobriban, and Kostas Daniilidis. Learning augmentation distributions using transformed risk mini- mization. arXiv preprint arXiv:2111.08190, 2021
2021 arXiv
-
[46]
Recurrent neural networks for multivariate time series with missing values
Zhengping Che, Sanjay Purushotham, Kyunghyun Cho, David Sontag, and Yan Liu. Recurrent neural networks for multivariate time series with missing values. Scientific reports, 8(1):1–12, 2018
2018
-
[47]
Linear system theory and design
Chi-Tsong Chen. Linear system theory and design . Saunders college publishing, 1984
1984
-
[48]
A group-theoretic framework for data augmentation
Shuxiao Chen, Edgar Dobriban, and Jane H Lee. A group-theoretic framework for data augmentation. Journal of Machine Learning Research, 21(245):1–71, 2020. BIBLIOGRAPHY 195
2020
-
[49]
Stabilizing differentiable architecture search via perturbation-based regularization
Xiangning Chen and Cho-Jui Hsieh. Stabilizing differentiable architecture search via perturbation-based regularization. In International conference on machine learn- ing, pages 1554–1565. PMLR, 2020
2020
-
[50]
Progressive darts: Bridging the op- timization gap for nas in the wild
Xin Chen, Lingxi Xie, Jun Wu, and Qi Tian. Progressive darts: Bridging the op- timization gap for nas in the wild. International Journal of Computer Vision , 129: 638–655, 2021
2021
-
[51]
Graph-based global reasoning networks
Yunpeng Chen, Marcus Rohrbach, Zhicheng Yan, Yan Shuicheng, Jiashi Feng, and Yannis Kalantidis. Graph-based global reasoning networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 433–442, 2019
2019
-
[52]
Long short-term memory-networks for machine reading
Jianpeng Cheng, Li Dong, and Mirella Lapata. Long short-term memory-networks for machine reading. arXiv preprint arXiv:1601.06733, 2016
2016 arXiv
-
[53]
Rotdcf: Decomposition of convolutional filters for rotation-equivariant deep networks
Xiuyuan Cheng, Qiang Qiu, Robert Calderbank, and Guillermo Sapiro. Rotdcf: Decomposition of convolutional filters for rotation-equivariant deep networks. arXiv preprint arXiv:1805.06846, 2018
2018 arXiv
-
[54]
Some experiments on the recognition of speech, with one and with two ears
E Colin Cherry. Some experiments on the recognition of speech, with one and with two ears. The Journal of the acoustical society of America, 25(5):975–979, 1953
1953
-
[55]
Parallelizing legendre memory unit train- ing
Narsimha Chilkuri and Chris Eliasmith. Parallelizing legendre memory unit train- ing. arXiv preprint arXiv:2102.11417, 2021
2021 arXiv
-
[56]
Learning phrase representa- tions using rnn encoder-decoder for statistical machine translation
Kyunghyun Cho, Bart Van Merri ¨enboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio. Learning phrase representa- tions using rnn encoder-decoder for statistical machine translation. arXiv preprint arXiv:1406.1078, 2014
2014 arXiv
-
[57]
Automatic tagging using deep convolutional neural networks
Keunwoo Choi, George Fazekas, and Mark Sandler. Automatic tagging using deep convolutional neural networks. arXiv preprint arXiv:1606.00298, 2016
2016 arXiv
-
[58]
Xception: Deep learning with depthwise separable convolu- tions
Franc ¸ois Chollet. Xception: Deep learning with depthwise separable convolu- tions. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1251–1258, 2017
2017
-
[59]
Rethinking attention with performers
Krzysztof Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamas Sarlos, Peter Hawkins, Jared Davis, Afroz Mohiuddin, Lukasz Kaiser, et al. Rethinking attention with performers. arXiv preprint arXiv:2009.14794, 2020
2009 arXiv
-
[60]
Palm: Scaling language modeling with pathways
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebas- tian Gehrmann, et al. Palm: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311, 2022
2022 arXiv
-
[61]
A downsampled variant 196 BIBLIOGRAPHY of imagenet as an alternative to the CIFAR datasets
Patryk Chrabaszcz, Ilya Loshchilov, and Frank Hutter. A downsampled variant 196 BIBLIOGRAPHY of imagenet as an alternative to the CIFAR datasets. CoRR, abs/1707.08819, 2017. URL http://arxiv.org/abs/1707.08819
2017 arXiv
-
[62]
Darts-: robustly stepping out of performance collapse without indicators
Xiangxiang Chu, Xiaoxing Wang, Bo Zhang, Shun Lu, Xiaolin Wei, and Junchi Yan. Darts-: robustly stepping out of performance collapse without indicators. arXiv preprint arXiv:2009.01027, 2020
2009 arXiv
-
[63]
Empir- ical evaluation of gated recurrent neural networks on sequence modeling
Junyoung Chung, Caglar Gulcehre, KyungHyun Cho, and Yoshua Bengio. Empir- ical evaluation of gated recurrent neural networks on sequence modeling. arXiv preprint arXiv:1412.3555, 2014
2014 arXiv
-
[64]
An analysis of single-layer net- works in unsupervised feature learning
Adam Coates, Andrew Ng, and Honglak Lee. An analysis of single-layer net- works in unsupervised feature learning. InProceedings of the fourteenth international conference on artificial intelligence and statistics, pages 215–223. JMLR Workshop and Conference Proceedings, 2011
2011
-
[65]
Group equivariant convolutional networks
Taco Cohen and Max Welling. Group equivariant convolutional networks. In International conference on machine learning, pages 2990–2999. PMLR, 2016
2016
-
[66]
Steerable cnns
Taco S Cohen and Max Welling. Steerable cnns. arXiv preprint arXiv:1612.08498, 2016
2016 arXiv
-
[67]
Spherical cnns
Taco S Cohen, Mario Geiger, Jonas K ¨ohler, and Max Welling. Spherical cnns. In International Conference on Learning Representations, 2018
2018
-
[68]
A general theory of equivariant cnns on homogeneous spaces
Taco S Cohen, Mario Geiger, and Maurice Weiler. A general theory of equivariant cnns on homogeneous spaces. InAdvances in Neural Information Processing Systems, pages 9142–9153, 2019
2019
-
[69]
Gauge equivariant convolutional networks and the icosahedral cnn
Taco S Cohen, Maurice Weiler, Berkay Kicanaoglu, and Max Welling. Gauge equivariant convolutional networks and the icosahedral cnn. arXiv preprint arXiv:1902.04615, 2019
1902 arXiv
-
[70]
Very deep con- volutional networks for text classification
Alexis Conneau, Holger Schwenk, Lo ¨ıc Barrault, and Yann Lecun. Very deep con- volutional networks for text classification. arXiv preprint arXiv:1606.01781, 2016
2016 arXiv
-
[71]
Fast construction of k-nearest neighbor graphs for point clouds
Michael Connor and Piyush Kumar. Fast construction of k-nearest neighbor graphs for point clouds. IEEE transactions on visualization and computer graphics , 16(4):599–608, 2010
2010
-
[72]
On the relation- ship between self-attention and convolutional layers
Jean-Baptiste Cordonnier, Andreas Loukas, and Martin Jaggi. On the relation- ship between self-attention and convolutional layers. In International Conference on Learning Representations , 2020. URL https://openreview.net/forum?id= HJlnC1rKPB
2020
-
[73]
A volumetric method for building complex models from range images
Brian Curless and Marc Levoy. A volumetric method for building complex models from range images. In Proceedings of the 23rd annual conference on Computer graphics and interactive techniques, pages 303–312, 1996. BIBLIOGRAPHY 197
1996
-
[74]
Deformable convolutional networks
Jifeng Dai, Haozhi Qi, Yuwen Xiong, Yi Li, Guodong Zhang, Han Hu, and Yichen Wei. Deformable convolutional networks. In Proceedings of the IEEE international conference on computer vision, pages 764–773, 2017
2017
-
[75]
Very deep convo- lutional neural networks for raw waveforms
Wei Dai, Chia Dai, Shuhui Qu, Juncheng Li, and Samarjit Das. Very deep convo- lutional neural networks for raw waveforms. In 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 421–425. IEEE, 2017
2017
-
[76]
Transformer-xl: Attentive language models beyond a fixed-length context
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc V Le, and Ruslan Salakhutdinov. Transformer-xl: Attentive language models beyond a fixed-length context. arXiv preprint arXiv:1901.02860, 2019
1901 arXiv
-
[77]
Picture memory experiments
Kent Dallett, Sandra G Wilcox, and Lester D’andrea. Picture memory experiments. Journal of Experimental Psychology, 76(2p1):312, 1968
1968
-
[78]
Fundamental papers in wavelet theory
Ingrid Daubechies. Fundamental papers in wavelet theory . Princeton University Press, 2006
2006
-
[79]
J.G. Daugman. Complete discrete 2-d gabor transforms by neural networks for image analysis and compression. IEEE Transactions on Acoustics, Speech, and Signal Processing, 36(7):1169–1179, 1988. doi: 10.1109/29.1644
1988 doi
-
[80]
Language mod- eling with gated convolutional networks
Yann N Dauphin, Angela Fan, Michael Auli, and David Grangier. Language mod- eling with gated convolutional networks. In International conference on machine learning, pages 933–941, 2017
2017
-
[81]
Gru-ode-bayes: Continuous modeling of sporadically-observed time series
Edward De Brouwer, Jaak Simm, Adam Arany, and Yves Moreau. Gru-ode-bayes: Continuous modeling of sporadically-observed time series. In Advances in Neural Information Processing Systems, pages 7379–7390, 2019
2019
-
[82]
Deepsphere: a graph-based spherical cnn
Micha ¨el Defferrard, Martino Milani, Fr ´ed´erick Gusset, and Nathana ¨el Perraudin. Deepsphere: a graph-based spherical cnn. arXiv preprint arXiv:2012.15000, 2020
2012 arXiv
-
[83]
Au- tomatic symmetry discovery with lie algebra convolutional network
Nima Dehmamy, Robin Walters, Yanchen Liu, Dashun Wang, and Rose Yu. Au- tomatic symmetry discovery with lie algebra convolutional network. Advances in Neural Information Processing Systems, 34, 2021
2021
-
[84]
Insect cyborgs: Bio-mimetic feature gen- erators improve ml accuracy on limited data
Charles B Delahunt and J Nathan Kutz. Insect cyborgs: Bio-mimetic feature gen- erators improve ml accuracy on limited data. 2019
2019
-
[85]
Diffusion models beat gans on image synthesis
Prafulla Dhariwal and Alexander Nichol. Diffusion models beat gans on image synthesis. Advances in Neural Information Processing Systems, 34:8780–8794, 2021
2021
-
[86]
Affine self convolution
Nichita Diaconu and Daniel E Worrall. Affine self convolution. arXiv preprint arXiv:1911.07704, 2019
1911 arXiv
-
[87]
Learning to convolve: A generalized weight-tying approach
Nichita Diaconu and Daniel E Worrall. Learning to convolve: A generalized weight-tying approach. 2019. 198 BIBLIOGRAPHY
2019
-
[88]
End-to-end learning for music au- dio
Sander Dieleman and Benjamin Schrauwen. End-to-end learning for music au- dio. In 2014 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 6964–6968. IEEE, 2014
2014
-
[89]
Exploiting cyclic symmetry in convolutional neural networks
Sander Dieleman, Jeffrey De Fauw, and Koray Kavukcuoglu. Exploiting cyclic symmetry in convolutional neural networks. In International conference on machine learning, pages 1889–1898. PMLR, 2016
2016
-
[90]
Searching for a robust neural architecture in four gpu hours
Xuanyi Dong and Yi Yang. Searching for a robust neural architecture in four gpu hours. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1761–1770, 2019
2019
-
[91]
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xi- aohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arX...
2010 arXiv
-
[92]
Brp-nas: Prediction-based nas using gcns
Lukasz Dudziak, Thomas Chau, Mohamed Abdelfattah, Royson Lee, Hyeji Kim, and Nicholas Lane. Brp-nas: Prediction-based nas using gcns. Advances in Neural Information Processing Systems, 33:10480–10490, 2020
2020
-
[93]
Abstract algebra, volume 3
David Steven Dummit and Richard M Foote. Abstract algebra, volume 3. Wiley Hoboken, 2004
2004
-
[94]
Generative models as dis- tributions of functions
Emilien Dupont, Yee Whye Teh, and Arnaud Doucet. Generative models as dis- tributions of functions. arXiv preprint arXiv:2102.04776, 2021
2021 arXiv
-
[95]
Efficient multi- objective neural architecture search via lamarckian evolution
Thomas Elsken, Jan Hendrik Metzen, and Frank Hutter. Efficient multi- objective neural architecture search via lamarckian evolution. arXiv preprint arXiv:1804.09081, 2018
2018 arXiv
-
[96]
Neural architecture search: A survey
Thomas Elsken, Jan Hendrik Metzen, and Frank Hutter. Neural architecture search: A survey. The Journal of Machine Learning Research, 20(1):1997–2017, 2019
1997
-
[97]
Lipschitz recurrent neural networks
N Benjamin Erichson, Omri Azencot, Alejandro Queiruga, Liam Hodgkinson, and Michael W Mahoney. Lipschitz recurrent neural networks. arXiv preprint arXiv:2006.12070, 2020
2006 arXiv
-
[98]
Cross-domain 3d equivariant image embeddings
Carlos Esteves, Avneesh Sud, Zhengyi Luo, Kostas Daniilidis, and Ameesh Maka- dia. Cross-domain 3d equivariant image embeddings. In International Conference on Machine Learning, pages 1812–1822. PMLR, 2019
2019
-
[99]
Equivariant multi-view networks
Carlos Esteves, Yinshuang Xu, Christine Allen-Blanchette, and Kostas Daniilidis. Equivariant multi-view networks. InProceedings of the IEEE International Conference on Computer Vision, pages 1568–1577, 2019
2019
-
[100]
Spin-weighted spherical cnns
Carlos Esteves, Ameesh Makadia, and Kostas Daniilidis. Spin-weighted spherical cnns. Advances in Neural Information Processing Systems, 33, 2020. BIBLIOGRAPHY 199
2020
-
[101]
Pytorch lightning
William Falcon et al. Pytorch lightning. GitHub. Note: https://github.com/PyTorchLightning/pytorch-lightning, 3, 2019
2019
-
[102]
Reparameteriz- ing distributions on lie groups
Luca Falorsi, Pim de Haan, Tim R Davidson, and Patrick Forr ´e. Reparameteriz- ing distributions on lie groups. In The 22nd International Conference on Artificial Intelligence and Statistics, pages 3244–3253. PMLR, 2019
2019
-
[103]
Densely connected search space for more flexible neural architecture search
Jiemin Fang, Yuzhu Sun, Qian Zhang, Yuan Li, Wenyu Liu, and Xinggang Wang. Densely connected search space for more flexible neural architecture search. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages 10628–10637, 2020
2020
-
[104]
Multiplicative filter networks
Rizal Fathony, Anit Kumar Sahu, Devin Willmott, and J Zico Kolter. Multiplicative filter networks. In International Conference on Learning Representations , 2021. URL https://openreview.net/forum?id=OmtmcPkkhT
2021
-
[105]
Gener- alizing convolutional neural networks for equivariance to lie groups on arbitrary continuous data
Marc Finzi, Samuel Stanton, Pavel Izmailov, and Andrew Gordon Wilson. Gener- alizing convolutional neural networks for equivariance to lie groups on arbitrary continuous data. arXiv preprint arXiv:2002.12880, 2020
2002 arXiv
-
[106]
Residual pathway priors for soft equivariance constraints
Marc Finzi, Gregory Benton, and Andrew G Wilson. Residual pathway priors for soft equivariance constraints. Advances in Neural Information Processing Systems , 34, 2021
2021
-
[107]
A practical method for constructing equivariant multilayer perceptrons for arbitrary matrix groups
Marc Finzi, Max Welling, and Andrew Gordon Wilson. A practical method for constructing equivariant multilayer perceptrons for arbitrary matrix groups. In International Conference on Machine Learning, pages 3318–3328. PMLR, 2021
2021
-
[108]
J Fourier. M ´emoire sur la propagation de la chaleur dans les corps solides, pr´esent´e le 21 d ´ecembre 1807 `a l’institut national—nouveau bulletin des sciences par la soci´et´e philomatique de paris. i. In Paris: First European Conference on Signal Anal- ysis and Predictio...
-
[109]
The Lottery Ticket Hypothesis: On Sparse, Trainable Neural Net- works
Jonathan Frankle. The Lottery Ticket Hypothesis: On Sparse, Trainable Neural Net- works. PhD thesis, Massachusetts Institute of Technology, 2023
2023
-
[110]
The face-inversion effect as a deficit in the encoding of configural information: Direct evidence
Alejo Freire, Kang Lee, and Lawrence A Symons. The face-inversion effect as a deficit in the encoding of configural information: Direct evidence. Perception, 29 (2):159–170, 2000
2000
-
[111]
Se (3)- transformers: 3d roto-translation equivariant attention networks
Fabian B Fuchs, Daniel E Worrall, Volker Fischer, and Max Welling. Se (3)- transformers: 3d roto-translation equivariant attention networks. arXiv preprint arXiv:2006.10503, 2020
2006 arXiv
-
[112]
Neocognitron: A self-organizing neural network model for a mechanism of visual pattern recognition
Kunihiko Fukushima and Sei Miyake. Neocognitron: A self-organizing neural network model for a mechanism of visual pattern recognition. In Competition and cooperation in neural nets, pages 267–285. Springer, 1982. 200 BIBLIOGRAPHY
1982
-
[113]
Speaker-independent isolated word recognition based on empha- sized spectral dynamics
Sadaoki Furui. Speaker-independent isolated word recognition based on empha- sized spectral dynamics. In ICASSP’86. IEEE International Conference on Acoustics, Speech, and Signal Processing, volume 11, pages 1991–1994. IEEE, 1986
1991
-
[114]
Theory of communication
Dennis Gabor. Theory of communication. part 1: The analysis of information. Journal of the Institution of Electrical Engineers-Part III: Radio and Communication En- gineering, 93(26):429–441, 1946
1946
-
[115]
Deep symmetry networks
Robert Gens and Pedro M Domingos. Deep symmetry networks. In Advances in neural information processing systems, pages 2537–2545, 2014
2014
-
[116]
Frame- exit: Conditional early exiting for efficient video recognition
Amir Ghodrati, Babak Ehteshami Bejnordi, and Amirhossein Habibian. Frame- exit: Conditional early exiting for efficient video recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 15608– 15618, 2021
2021
-
[117]
Neural message passing for quantum chemistry
Justin Gilmer, Samuel S Schoenholz, Patrick F Riley, Oriol Vinyals, and George E Dahl. Neural message passing for quantum chemistry. In International conference on machine learning, pages 1263–1272. PMLR, 2017
2017
-
[118]
Rich feature hier- archies for accurate object detection and semantic segmentation
Ross Girshick, Jeff Donahue, Trevor Darrell, and Jitendra Malik. Rich feature hier- archies for accurate object detection and semantic segmentation. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 580–587, 2014
2014
-
[119]
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio. Understanding the difficulty of training deep feedforward neural networks. In Proceedings of the thirteenth international conference on artificial intelligence and statistics , pages 249–256. JMLR Workshop and Confer- ence Proceedings, 2010
2010
-
[120]
It’s raw! audio gen- eration with state-space models
Karan Goel, Albert Gu, Chris Donahue, and Christopher R ´e. It’s raw! audio gen- eration with state-space models. In International Conference on Machine Learning , pages 7616–7633. PMLR, 2022
2022
-
[121]
Physiobank, physiotoolkit, and physionet: components of a new research resource for complex physiologic signals
Ary L Goldberger, Luis AN Amaral, Leon Glass, Jeffrey M Hausdorff, Plamen Ch Ivanov, Roger G Mark, Joseph E Mietus, George B Moody, Chung-Kang Peng, and H Eugene Stanley. Physiobank, physiotoolkit, and physionet: components of a new research resource for complex physiologic si...
2000
-
[122]
Exploiting and coping with sparsity to accelerate DNNs on CPUs
Zhangxiaowen Gong. Exploiting and coping with sparsity to accelerate DNNs on CPUs. PhD thesis, 2021
2021
-
[123]
Dense steerable filter cnns for exploiting rotational symmetry in histology images
Simon Graham, David Epstein, and Nasir Rajpoot. Dense steerable filter cnns for exploiting rotational symmetry in histology images. arXiv preprint arXiv:2004.03037, 2020
2004 arXiv
-
[124]
Speech recogni- tion with deep recurrent neural networks
Alex Graves, Abdel-rahman Mohamed, and Geoffrey Hinton. Speech recogni- tion with deep recurrent neural networks. In 2013 IEEE international conference on acoustics, speech and signal processing, pages 6645–6649. Ieee, 2013. BIBLIOGRAPHY 201
2013
-
[125]
Transforms associated to square inte- grable group representations
Alex Grossmann, Jean Morlet, and T Paul. Transforms associated to square inte- grable group representations. i. general results. Journal of Mathematical Physics, 26 (10):2473–2479, 1985
1985
-
[126]
Hippo: Recurrent memory with optimal polynomial projections
Albert Gu, Tri Dao, Stefano Ermon, Atri Rudra, and Christopher R ´e. Hippo: Recurrent memory with optimal polynomial projections. arXiv preprint arXiv:2008.07669, 2020
2008 arXiv
-
[127]
Improving the gating mechanism of recurrent neural networks
Albert Gu, Caglar Gulcehre, Thomas Paine, Matt Hoffman, and Razvan Pascanu. Improving the gating mechanism of recurrent neural networks. In International Conference on Machine Learning, pages 3800–3809. PMLR, 2020
2020
-
[128]
Combining recurrent, convolutional, and continuous-time mod- els with linear state space layers.Advances in Neural Information Processing Systems, 34, 2021
Albert Gu, Isys Johnson, Karan Goel, Khaled Saab, Tri Dao, Atri Rudra, and Christopher R´e. Combining recurrent, convolutional, and continuous-time mod- els with linear state space layers.Advances in Neural Information Processing Systems, 34, 2021
2021
-
[129]
Efficiently modeling long sequences with structured state spaces
Albert Gu, Karan Goel, and Christopher Re. Efficiently modeling long sequences with structured state spaces. In International Conference on Learning Representations,
-
[130]
On the param- eterization and initialization of diagonal state space models
Albert Gu, Ankit Gupta, Karan Goel, and Christopher R ´e. On the param- eterization and initialization of diagonal state space models. arXiv preprint arXiv:2206.11893, 2022
2022 arXiv
-
[131]
Bag of baselines for multi-objective joint neural architecture search and hyperparameter optimization
Julia Guerrero-Viu, Sven Hauns, Sergio Izquierdo, Guilherme Miotto, Simon Schrodi, Andre Biedenkapp, Thomas Elsken, Difan Deng, Marius Lindauer, and Frank Hutter. Bag of baselines for multi-objective joint neural architecture search and hyperparameter optimization. arXiv prepr...
2021 arXiv
-
[132]
Stacnas: Towards stable and consistent optimization for differentiable neural architecture search
Li Guilin, Zhang Xing, Wang Zitong, Li Zhenguo, and Zhang Tong. Stacnas: Towards stable and consistent optimization for differentiable neural architecture search. 2019
2019
-
[133]
Multi-time-scale convolu- tion for emotion recognition from speech audio signals
Eric Guizzo, Tillman Weyde, and Jack Barnett Leveson. Multi-time-scale convolu- tion for emotion recognition from speech audio signals. InICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pages 6489–6493. IEEE, 2020
2020
-
[134]
Dai, and Quoc V
David Ha, Andrew M. Dai, and Quoc V . Le. Hypernetworks. In International Conference on Learning Representations, 2017. URL https://openreview.net/ forum?id=rkpACe1lx
2017
-
[135]
De moivre’s normal approximation to the binomial, 1733, and its generalization
Anders Hald. De moivre’s normal approximation to the binomial, 1733, and its generalization. A History of Parametric Statistical Inference from Bernoulli to Fisher, 1713–1935, pages 17–24, 2007
1935
-
[136]
Optimizing filter size in convolutional neural networks for facial 202 BIBLIOGRAPHY action unit recognition
Shizhong Han, Zibo Meng, Zhiyuan Li, James O’Reilly, Jie Cai, Xiaofeng Wang, and Yan Tong. Optimizing filter size in convolutional neural networks for facial 202 BIBLIOGRAPHY action unit recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognit...
2018
-
[137]
Learning both weights and connections for efficient neural network
Song Han, Jeff Pool, John Tran, and William Dally. Learning both weights and connections for efficient neural network. Advances in neural information processing systems, 28, 2015
2015
-
[138]
Complexity of linear regions in deep networks
Boris Hanin and David Rolnick. Complexity of linear regions in deep networks. arXiv preprint arXiv:1901.09021, 2019
1901 arXiv
-
[139]
Dilated convolution with learnable spacings
Ismail Khalfaoui Hassani, Thomas Pellegrini, and Timoth ´ee Masquelier. Dilated convolution with learnable spacings. In The Eleventh International Conference on Learning Representations , 2023. URL https://openreview.net/forum?id= Q3-1vRh3HOA
2023
-
[140]
Faster autoaugment: Learning augmentation strategies using backpropagation
Ryuichiro Hataya, Jan Zdenek, Kazuki Yoshizoe, and Hideki Nakayama. Faster autoaugment: Learning augmentation strategies using backpropagation. In Euro- pean Conference on Computer Vision, pages 1–16. Springer, 2020
2020
-
[141]
Delving deep into recti- fiers: Surpassing human-level performance on imagenet classification
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Delving deep into recti- fiers: Surpassing human-level performance on imagenet classification. In Proceed- ings of the IEEE international conference on computer vision, pages 1026–1034, 2015
2015
-
[142]
Spatial pyramid pool- ing in deep convolutional networks for visual recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Spatial pyramid pool- ing in deep convolutional networks for visual recognition. IEEE transactions on pattern analysis and machine intelligence, 37(9):1904–1916, 2015
1904
-
[143]
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016
2016
-
[144]
Gaussian error linear units (gelus)
Dan Hendrycks and Kevin Gimpel. Gaussian error linear units (gelus). arXiv preprint arXiv:1606.08415, 2016
2016 arXiv
-
[145]
An experimental investigation of past experience as a determinant of visual form perception
Mary Henle. An experimental investigation of past experience as a determinant of visual form perception. Journal of Experimental Psychology, 30(1):1, 1942
1942
-
[146]
A computer model of amplitude-modulation sensitivity of single units in the inferior colliculus
Michael J Hewitt and Ray Meddis. A computer model of amplitude-modulation sensitivity of single units in the inferior colliculus. The Journal of the Acoustical Society of America, 95(4):2145–2159, 1994
1994
-
[147]
Untersuchungen zu dynamischen neuronalen netzen
Sepp Hochreiter. Untersuchungen zu dynamischen neuronalen netzen. Diploma, Technische Universit¨ at M¨ unchen, 91(1), 1991
1991
-
[148]
Long short-term memory
Sepp Hochreiter and J ¨urgen Schmidhuber. Long short-term memory. Neural com- putation, 9(8):1735–1780, 1997
1997
-
[149]
Sparsity in deep learning: Pruning and growth for efficient inference and training BIBLIOGRAPHY 203 in neural networks
Torsten Hoefler, Dan Alistarh, Tal Ben-Nun, Nikoli Dryden, and Alexandra Peste. Sparsity in deep learning: Pruning and growth for efficient inference and training BIBLIOGRAPHY 203 in neural networks. The Journal of Machine Learning Research , 22(1):10882–11005, 2021
2021
-
[150]
Peters, Taco S
Emiel Hoogeboom, Jorn W.T. Peters, Taco S. Cohen, and Max Welling. Hexa- conv. In International Conference on Learning Representations , 2018. URL https: //openreview.net/forum?id=r1vuQG-CW
2018
-
[151]
Integer discrete flows and lossless compression
Emiel Hoogeboom, Jorn Peters, Rianne Van Den Berg, and Max Welling. Integer discrete flows and lossless compression. Advances in Neural Information Processing Systems, 32, 2019
2019
-
[152]
Argmax flows and multinomial diffusion: Towards non-autoregressive language models
Emiel Hoogeboom, Didrik Nielsen, Priyank Jaini, Patrick Forr ´e, and Max Welling. Argmax flows and multinomial diffusion: Towards non-autoregressive language models. arXiv preprint arXiv:2102.05379, 2021
2021 arXiv
-
[153]
Equivariant diffusion for molecule generation in 3d
Emiel Hoogeboom, Vıctor Garcia Satorras, Cl ´ement Vignac, and Max Welling. Equivariant diffusion for molecule generation in 3d. In International Conference on Machine Learning, pages 8867–8887. PMLR, 2022
2022
-
[154]
Waarom helpt kunstmatige intelligentie de arts en pati ¨ent nog zo weinig? 2022
Mark Hoogendoorn. Waarom helpt kunstmatige intelligentie de arts en pati ¨ent nog zo weinig? 2022
2022
-
[155]
Local relation networks for image recognition
Han Hu, Zheng Zhang, Zhenda Xie, and Stephen Lin. Local relation networks for image recognition. In Proceedings of the IEEE International Conference on Computer Vision, pages 3464–3473, 2019
2019
-
[156]
Efficient forward architecture search
Hanzhang Hu, John Langford, Rich Caruana, Saurajit Mukherjee, Eric J Horvitz, and Debadeepta Dey. Efficient forward architecture search. Advances in Neural Information Processing Systems, 32, 2019
2019
-
[157]
Squeeze-and-excitation networks
Jie Hu, Li Shen, and Gang Sun. Squeeze-and-excitation networks. In Proceedings of the IEEE conference on computer vision and pattern recognition , pages 7132–7141, 2018
2018
-
[158]
Randla-net: Efficient semantic segmentation of large-scale point clouds
Qingyong Hu, Bo Yang, Linhai Xie, Stefano Rosa, Yulan Guo, Zhihua Wang, Niki Trigoni, and Andrew Markham. Randla-net: Efficient semantic segmentation of large-scale point clouds. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 11108–...
2020
-
[159]
Dsnas: Direct neural architecture search without parameter re- training
Shoukang Hu, Sirui Xie, Hehui Zheng, Chunxiao Liu, Jianping Shi, Xunying Liu, and Dahua Lin. Dsnas: Direct neural architecture search without parameter re- training. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12084–12092, 2020
2020
-
[160]
Pointwise convolutional neu- ral networks
Binh-Son Hua, Minh-Khoi Tran, and Sai-Kit Yeung. Pointwise convolutional neu- ral networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 984–993, 2018. 204 BIBLIOGRAPHY
2018
-
[161]
Densely connected convolutional networks
Gao Huang, Zhuang Liu, Laurens Van Der Maaten, and Kilian Q Weinberger. Densely connected convolutional networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4700–4708, 2017
2017
-
[162]
sharpdarts: Faster and more accurate differentiable architecture search
Andrew Hundt, Varun Jain, and Gregory D Hager. sharpdarts: Faster and more accurate differentiable architecture search. arXiv preprint arXiv:1903.09900, 2019
1903 arXiv
-
[163]
Lietransformer: equivariant self-attention for lie groups
Michael J Hutchinson, Charline Le Lan, Sheheryar Zaidi, Emilien Dupont, Yee Whye Teh, and Hyunjik Kim. Lietransformer: equivariant self-attention for lie groups. In International Conference on Machine Learning , pages 4533–4543. PMLR, 2021
2021
-
[164]
Attention-based deep mul- tiple instance learning
Maximilian Ilse, Jakub M Tomczak, and Max Welling. Attention-based deep mul- tiple instance learning. ICML, 2018
2018
-
[165]
Invariance learning in deep neural networks with differen- tiable laplace approximations
Alexander Immer, Tycho van der Ouderaa, Gunnar R ¨atsch, Vincent Fortuin, and Mark van der Wilk. Invariance learning in deep neural networks with differen- tiable laplace approximations. Advances in Neural Information Processing Systems , 35:12449–12463, 2022
2022
-
[166]
Batch normalization: Accelerating deep net- work training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy. Batch normalization: Accelerating deep net- work training by reducing internal covariate shift. In International conference on machine learning, pages 448–456. PMLR, 2015
2015
-
[167]
Structured receptive fields in cnns
Jorn-Henrik Jacobsen, Jan Van Gemert, Zhongyu Lou, and Arnold WM Smeul- ders. Structured receptive fields in cnns. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 2610–2619, 2016
2016
-
[168]
Perceiver: General perception with iterative attention
Andrew Jaegle, Felix Gimeno, Andy Brock, Oriol Vinyals, Andrew Zisserman, and Joao Carreira. Perceiver: General perception with iterative attention. In Interna- tional conference on machine learning, pages 4651–4664. PMLR, 2021
2021
-
[169]
Categorical reparameterization with gumbel-softmax
Eric Jang, Shixiang Gu, and Ben Poole. Categorical reparameterization with gumbel-softmax. arXiv preprint arXiv:1611.01144, 2016
2016 arXiv
-
[170]
Active convolution: Learning the shape of convolu- tion for image classification
Yunho Jeon and Junmo Kim. Active convolution: Learning the shape of convolu- tion for image classification. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4201–4209, 2017
2017
-
[171]
Dynamic filter networks
Xu Jia, Bert De Brabandere, Tinne Tuytelaars, and Luc V Gool. Dynamic filter networks. In Advances in neural information processing systems, pages 667–675, 2016
2016
-
[172]
Cotr: Correspondence transformer for matching across images
Wei Jiang, Eduard Trulls, Jan Hosang, Andrea Tagliasacchi, and Kwang Moo Yi. Cotr: Correspondence transformer for matching across images. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 6207–6217, 2021
2021
-
[173]
The pre-response stimulus ensemble of neurons in the cochlear nucleus
PLM Johannesma. The pre-response stimulus ensemble of neurons in the cochlear nucleus. In Symposium on Hearing Theory, 1972. IPO, 1972. BIBLIOGRAPHY 205
1972
-
[174]
Mimic-iii, a freely accessible critical care database
Alistair EW Johnson, Tom J Pollard, Lu Shen, Li-wei H Lehman, Mengling Feng, Mohammad Ghassemi, Benjamin Moody, Peter Szolovits, Leo Anthony Celi, and Roger G Mark. Mimic-iii, a freely accessible critical care database. Scientific data, 3 (1):1–9, 2016
2016
-
[175]
Machine learning: Trends, perspectives, and prospects
Michael I Jordan and Tom M Mitchell. Machine learning: Trends, perspectives, and prospects. Science, 349(6245):255–260, 2015
2015
-
[176]
Spinalnet: Deep neural network with gradual input
HM Kabir, Moloud Abdar, Seyed Mohammad Jafar Jalali, Abbas Khosravi, Amir F Atiya, Saeid Nahavandi, and Dipti Srinivasan. Spinalnet: Deep neural network with gradual input. arXiv preprint arXiv:2007.03347, 2020
2007 arXiv
-
[177]
Recurrent continuous translation models
Nal Kalchbrenner and Phil Blunsom. Recurrent continuous translation models. In Proceedings of the 2013 Conference on Empirical Methods in Natural Language Process- ing, pages 1700–1709, 2013
2013
-
[178]
Neural architecture search with bayesian optimisation and optimal transport
Kirthevasan Kandasamy, Willie Neiswanger, Jeff Schneider, Barnabas Poczos, and Eric P Xing. Neural architecture search with bayesian optimisation and optimal transport. Advances in neural information processing systems, 31, 2018
2018
-
[179]
Construction of 3d orthogonal cover of a digital object
Nilanjana Karmakar, Arindam Biswas, Partha Bhowmick, and Bhargab B Bhat- tacharya. Construction of 3d orthogonal cover of a digital object. In Combinatorial Image Analysis: 14th International Workshop, IWCIA 2011, Madrid, Spain, May 23-25,
2011
-
[180]
Alias-free generative adversarial networks
Tero Karras, Miika Aittala, Samuli Laine, Erik H ¨ark¨onen, Janne Hellsten, Jaakko Lehtinen, and Timo Aila. Alias-free generative adversarial networks. arXiv preprint arXiv:2106.12423, 2021
2021 arXiv
-
[181]
Transformers are rnns: Fast autoregressive transformers with linear attention, 2020
Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and Franc ¸ois Fleuret. Transformers are rnns: Fast autoregressive transformers with linear attention, 2020
2020
-
[182]
van Gemert
Osman Semih Kayhan and Jan C. van Gemert. On translation invariance in cnns: Convolutional layers can exploit absolute spatial location. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2020
2020
-
[183]
Neural controlled differential equations for irregular time series
Patrick Kidger, James Morrill, James Foster, and Terry Lyons. Neural controlled differential equations for irregular time series. arXiv preprint arXiv:2005.08926 , 2020
2005 arXiv
-
[184]
Smpconv: Self-moving point representa- tions for continuous convolution
Sanghyeon Kim and Eunbyung Park. Smpconv: Self-moving point representa- tions for continuous convolution. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10289–10299, 2023
2023
-
[185]
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014. 206 BIBLIOGRAPHY
2014 arXiv
-
[186]
Auto-encoding variational bayes
Diederik P Kingma and Max Welling. Auto-encoding variational bayes. arXiv preprint arXiv:1312.6114, 2013
2013 arXiv
-
[187]
Semi-supervised classification with graph con- volutional networks
Thomas N Kipf and Max Welling. Semi-supervised classification with graph con- volutional networks. arXiv preprint arXiv:1609.02907, 2016
2016 arXiv
-
[188]
Convolutional networks with oriented 1d kernels
Alexandre Kirchmeyer and Jia Deng. Convolutional networks with oriented 1d kernels. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 6222–6232, 2023
2023
-
[189]
Reformer: The efficient trans- former, 2020
Nikita Kitaev, Łukasz Kaiser, and Anselm Levskaya. Reformer: The efficient trans- former, 2020
2020
-
[190]
Exploiting redundancy: Separable group convolutional networks on lie groups
David M Knigge, David W Romero, and Erik J Bekkers. Exploiting redundancy: Separable group convolutional networks on lie groups. In International Conference on Machine Learning, pages 11359–11386. PMLR, 2022
2022
-
[191]
Romero, Albert Gu, Efstratios Gavves, Erik J Bekkers, Jakub Mikolaj Tomczak, Mark Hoogendoorn, and Jan jakob Sonke
David M Knigge, David W. Romero, Albert Gu, Efstratios Gavves, Erik J Bekkers, Jakub Mikolaj Tomczak, Mark Hoogendoorn, and Jan jakob Sonke. Modelling long range dependencies in $n$d: From task-specific to a general purpose CNN. In The Eleventh International Conference on Lear...
2023
-
[192]
Openai’s ceo says the age of giant ai models is already over
Will Knight. Openai’s ceo says the age of giant ai models is already over. Wired, April, 17:2023, 2023
2023
-
[193]
Kodak dataset, 1991
Kodak. Kodak dataset, 1991. URL http://r0k.us/graphics/kodak/
1991
-
[194]
On the generalization of equivariance and convolution in neural networks to the action of compact groups
Risi Kondor and Shubhendu Trivedi. On the generalization of equivariance and convolution in neural networks to the action of compact groups. arXiv preprint arXiv:1802.03690, 2018
2018 arXiv
-
[195]
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al. Learning multiple layers of features from tiny images. Technical report, Citeseer, 2009
2009
-
[196]
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. Imagenet classification with deep convolutional neural networks. Advances in neural information processing systems, 25, 2012
2012
-
[197]
A simple weight decay can improve generaliza- tion
Anders Krogh and John Hertz. A simple weight decay can improve generaliza- tion. Advances in neural information processing systems, 4, 1991
1991
-
[198]
Regular se (3) group convolutions for volu- metric medical image analysis
Thijs P Kuipers and Erik J Bekkers. Regular se (3) group convolutions for volu- metric medical image analysis. arXiv preprint arXiv:2306.13960, 2023
2023 arXiv
-
[199]
Lafarge, Erik J
Maxime W. Lafarge, Erik J. Bekkers, Josien P . W. Pluim, Remco Duits, and Mitko Veta. Roto-translation equivariant convolutional networks: Application to histopathology image analysis. arXiv preprint arXiv:2002.08725, 2020
2002 arXiv
-
[200]
Temporal ensembling for semi-supervised learning
Samuli Laine and Timo Aila. Temporal ensembling for semi-supervised learning. arXiv preprint arXiv:1610.02242, 2016. BIBLIOGRAPHY 207
2016 arXiv
-
[201]
An empirical evaluation of deep architectures on problems with many factors of variation
Hugo Larochelle, Dumitru Erhan, Aaron Courville, James Bergstra, and Yoshua Bengio. An empirical evaluation of deep architectures on problems with many factors of variation. In Proceedings of the 24th international conference on Machine learning, pages 473–480. ACM, 2007
2007
-
[202]
Eval- uation of algorithms using games: The case of music tagging
Edith Law, Kris West, Michael I Mandel, Mert Bay, and J Stephen Downie. Eval- uation of algorithms using games: The case of music tagging. In ISMIR, pages 387–392, 2009
2009
-
[203]
A simple way to initialize recurrent networks of rectified linear units
Quoc V Le, Navdeep Jaitly, and Geoffrey E Hinton. A simple way to initialize recurrent networks of rectified linear units. arXiv preprint arXiv:1504.00941, 2015
2015 arXiv
-
[204]
Pointgrid: A deep network for 3d shape understanding
Truc Le and Ye Duan. Pointgrid: A deep network for 3d shape understanding. In Proceedings of the IEEE conference on computer vision and pattern recognition , pages 9204–9214, 2018
2018
-
[205]
MNIST handwritten digit database
Yann LeCun and Corinna Cortes. MNIST handwritten digit database. 2010. URL http://yann.lecun.com/exdb/mnist/
2010
-
[206]
Backpropagation applied to handwritten zip code recognition
Yann LeCun, Bernhard Boser, John S Denker, Donnie Henderson, Richard E Howard, Wayne Hubbard, and Lawrence D Jackel. Backpropagation applied to handwritten zip code recognition. Neural computation, 1(4):541–551, 1989
1989
-
[207]
Optimal brain damage
Yann LeCun, John Denker, and Sara Solla. Optimal brain damage. Advances in neural information processing systems, 2, 1989
1989
-
[208]
Gradient-based learning applied to document recognition
Yann LeCun, L ´eon Bottou, Yoshua Bengio, and Patrick Haffner. Gradient-based learning applied to document recognition. Proceedings of the IEEE , 86(11):2278– 2324, 1998
1998
-
[209]
Deep learning
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton. Deep learning. nature, 521 (7553):436–444, 2015
2015
-
[210]
Raw waveform- based audio classification using sample-level cnn architectures
Jongpil Lee, Taejun Kim, Jiyoung Park, and Juhan Nam. Raw waveform- based audio classification using sample-level cnn architectures. arXiv preprint arXiv:1712.00866, 2017
2017 arXiv
-
[211]
Sample- level deep convolutional neural networks for music auto-tagging using raw wave- forms
Jongpil Lee, Jiyoung Park, Keunhyoung Luke Kim, and Juhan Nam. Sample- level deep convolutional neural networks for music auto-tagging using raw wave- forms. arXiv preprint arXiv:1703.01789, 2017
2017 arXiv
-
[212]
Toward efficient deep learning with sparse neural networks
Namhoon Lee. Toward efficient deep learning with sparse neural networks. PhD thesis, University of Oxford, 2020
2020
-
[213]
Fnet: Mix- ing tokens with fourier transforms
James Lee-Thorp, Joshua Ainslie, Ilya Eckstein, and Santiago Ontanon. Fnet: Mix- ing tokens with fourier transforms. arXiv preprint arXiv:2105.03824, 2021
2021 arXiv
-
[214]
Exploiting learned symmetries in group equivariant convolutions
Attila Lengyel and Jan van Gemert. Exploiting learned symmetries in group equivariant convolutions. In 2021 IEEE International Conference on Image Processing (ICIP), pages 759–763. IEEE, 2021. 208 BIBLIOGRAPHY
2021
-
[215]
Group equivariant cap- sule networks
Jan Eric Lenssen, Matthias Fey, and Pascal Libuschewski. Group equivariant cap- sule networks. In Advances in Neural Information Processing Systems , pages 8844– 8853, 2018
2018
-
[216]
Deep rotation equivariant network
Junying Li, Zichen Yang, Haifeng Liu, and Deng Cai. Deep rotation equivariant network. Neurocomputing, 290:26–33, 2018
2018
-
[217]
Geometry- aware gradient algorithms for neural architecture search
Liam Li, Mikhail Khodak, Maria-Florina Balcan, and Ameet Talwalkar. Geometry- aware gradient algorithms for neural architecture search. arXiv preprint arXiv:2004.07802, 2020
2004 arXiv
-
[218]
Independently recur- rent neural network (indrnn): Building a longer and deeper rnn
Shuai Li, Wanqing Li, Chris Cook, Ce Zhu, and Yanbo Gao. Independently recur- rent neural network (indrnn): Building a longer and deeper rnn. In Proceedings of the IEEE conference on computer vision and pattern recognition , pages 5457–5466, 2018
2018
-
[219]
Hospedales, Neil Martin Robertson, and Yongxin Yang
Yonggang Li, Guosheng Hu, Yongtao Wang, Timothy M. Hospedales, Neil Martin Robertson, and Yongxin Yang. DADA: differentiable automatic data augmenta- tion. 2020
2020
-
[220]
Fourier neural opera- tor for parametric partial differential equations
Zongyi Li, Nikola Kovachki, Kamyar Azizzadenesheli, Burigede Liu, Kaushik Bhattacharya, Andrew Stuart, and Anima Anandkumar. Fourier neural opera- tor for parametric partial differential equations. arXiv preprint arXiv:2010.08895 , 2020
2010 arXiv
-
[221]
Darts+: Improved differentiable architecture search with early stopping
Hanwen Liang, Shifeng Zhang, Jiacheng Sun, Xingqiu He, Weiran Huang, Kechen Zhuang, and Zhenguo Li. Darts+: Improved differentiable architecture search with early stopping. arXiv preprint arXiv:1909.06035, 2019
1909 arXiv
-
[222]
Fast autoaugment
Sungbin Lim, Ildoo Kim, Taesup Kim, Chiheon Kim, and Sungwoong Kim. Fast autoaugment. Advances in Neural Information Processing Systems, 32, 2019
2019
-
[223]
Network in network
Min Lin, Qiang Chen, and Shuicheng Yan. Network in network. arXiv preprint arXiv:1312.4400, 2013
2013 arXiv
-
[224]
Context-gated convolution, 2019
Xudong Lin, Lin Ma, Wei Liu, and Shih-Fu Chang. Context-gated convolution, 2019
2019
-
[225]
Scale-covariant and scale-invariant gaussian derivative net- works
Tony Lindeberg. Scale-covariant and scale-invariant gaussian derivative net- works. In Scale Space and Variational Methods in Computer Vision :, volume 12679 of Springer Lecture Notes in Computer Science, pages 3–14. Springer Nature, 2021. ISBN 978-3-030-75548-5. doi: 10.1007/...
2021 arXiv
-
[226]
Idealized computational models for auditory receptive fields
Tony Lindeberg and Anders Friberg. Idealized computational models for auditory receptive fields. PLoS one, 10(3), 2015
2015
-
[227]
Scale-space theory for auditory signals
Tony Lindeberg and Anders Friberg. Scale-space theory for auditory signals. In BIBLIOGRAPHY 209 International Conference on Scale Space and Variational Methods in Computer Vision , pages 3–15. Springer, 2015
2015
-
[228]
Learning long-range spatial dependencies with horizontal gated recurrent units
Drew Linsley, Junkyung Kim, Vijay Veerabadran, Charles Windolf, and Thomas Serre. Learning long-range spatial dependencies with horizontal gated recurrent units. Advances in neural information processing systems, 31, 2018
2018
-
[229]
Auto-deeplab: Hierarchical neural architecture search for semantic image segmentation
Chenxi Liu, Liang-Chieh Chen, Florian Schroff, Hartwig Adam, Wei Hua, Alan L Yuille, and Li Fei-Fei. Auto-deeplab: Hierarchical neural architecture search for semantic image segmentation. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, page...
2019
-
[230]
Hierarchical representations for efficient architecture search
Hanxiao Liu, Karen Simonyan, Oriol Vinyals, Chrisantha Fernando, and Koray Kavukcuoglu. Hierarchical representations for efficient architecture search. In 6th International Conference on Learning Representations, ICLR 2018, Vancouver, BC, Canada, April 30 - May 3, 2018, Confer...
2018
-
[231]
Darts: Differentiable architec- ture search
Hanxiao Liu, Karen Simonyan, and Yiming Yang. Darts: Differentiable architec- ture search. arXiv preprint arXiv:1806.09055, 2018
2018 arXiv
-
[232]
Applying topological persis- tence in convolutional neural network for music audio signals
Jen-Yu Liu, Shyh-Kang Jeng, and Yi-Hsuan Yang. Applying topological persis- tence in convolutional neural network for music audio signals. arXiv preprint arXiv:1608.07373, 2016
2016 arXiv
-
[233]
Under- standing the difficulty of training transformers
Liyuan Liu, Xiaodong Liu, Jianfeng Gao, Weizhu Chen, and Jiawei Han. Under- standing the difficulty of training transformers. arXiv preprint arXiv:2004.08249 , 2020
2004 arXiv
-
[234]
More convnets in the 2020s: Scaling up kernels beyond 51x51 using sparsity
Shiwei Liu, Tianlong Chen, Xiaohan Chen, Xuxi Chen, Qiao Xiao, Boqian Wu, Tommi K¨arkk¨ainen, Mykola Pechenizkiy, Decebal Constantin Mocanu, and Zhangyang Wang. More convnets in the 2020s: Scaling up kernels beyond 51x51 using sparsity. In The Eleventh International Conference...
-
[235]
Point-voxel cnn for efficient 3d deep learning
Zhijian Liu, Haotian Tang, Yujun Lin, and Song Han. Point-voxel cnn for efficient 3d deep learning. Advances in Neural Information Processing Systems, 32, 2019
2019
-
[236]
A convnet for the 2020s
Zhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer, Trevor Darrell, and Saining Xie. A convnet for the 2020s. arXiv preprint arXiv:2201.03545, 2022
2022 arXiv
-
[237]
Supervised scale-regularized linear convolu- tionary filters
Marco Loog and Francois Lauze. Supervised scale-regularized linear convolu- tionary filters. In Gabriel Brostow Tae-Kyun Kim, Stefanos Zafeiriou and Krystian Mikolajczyk, editors, Proceedings of the British Machine Vision Conference (BMVC) , pages 162.1–162.11. BMVA Press, Sep...
2017 doi
-
[238]
SGDR: Stochastic gradient descent with warm 210 BIBLIOGRAPHY restarts
Ilya Loshchilov and Frank Hutter. SGDR: Stochastic gradient descent with warm 210 BIBLIOGRAPHY restarts. In International Conference on Learning Representations, 2017. URL https: //openreview.net/forum?id=Skq89Scxx
2017
-
[239]
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. In International Conference on Learning Representations , 2019. URL https:// openreview.net/forum?id=Bkg6RiCqY7
2019
-
[240]
Deep pro- gressive multi-scale attention for acoustic event classification
Xugang Lu, Peng Shen, Sheng Li, Yu Tsao, and Hisashi Kawai. Deep pro- gressive multi-scale attention for acoustic event classification. arXiv preprint arXiv:1912.12011, 2019
1912 arXiv
-
[241]
Extended batch nor- malization
Chunjie Luo, Jianfeng Zhan, Lei Wang, and Wanling Gao. Extended batch nor- malization. arXiv preprint arXiv:2003.05569, 2020
2003 arXiv
-
[242]
Effective approaches to attention-based neural machine translation
Minh-Thang Luong, Hieu Pham, and Christopher D Manning. Effective approaches to attention-based neural machine translation. arXiv preprint arXiv:1508.04025, 2015
2015 arXiv
-
[243]
Luna: Linear unified nested attention
Xuezhe Ma, Xiang Kong, Sinong Wang, Chunting Zhou, Jonathan May, Hao Ma, and Luke Zettlemoyer. Luna: Linear unified nested attention. Advances in Neural Information Processing Systems, 34:2441–2453, 2021
2021
-
[244]
Mega: moving average equipped gated attention
Xuezhe Ma, Chunting Zhou, Xiang Kong, Junxian He, Liangke Gui, Graham Neu- big, Jonathan May, and Luke Zettlemoyer. Mega: moving average equipped gated attention. arXiv preprint arXiv:2209.10655, 2022
2022 arXiv
-
[245]
Learning word vectors for sentiment analysis
Andrew Maas, Raymond E Daly, Peter T Pham, Dan Huang, Andrew Y Ng, and Christopher Potts. Learning word vectors for sentiment analysis. In Proceedings of the 49th annual meeting of the association for computational linguistics: Human language technologies, pages 142–150, 2011
2011
-
[246]
The concrete distri- bution: A continuous relaxation of discrete random variables
Chris J Maddison, Andriy Mnih, and Yee Whye Teh. The concrete distri- bution: A continuous relaxation of discrete random variables. arXiv preprint arXiv:1611.00712, 2016
2016 arXiv
-
[247]
Equivariance-aware ar- chitectural optimization of neural networks
Kaitlin Maile, Dennis George Wilson, and Patrick Forr ´e. Equivariance-aware ar- chitectural optimization of neural networks. In The Eleventh International Con- ference on Learning Representations , 2023. URL https://openreview.net/ forum?id=a6rCdfABJXg
2023
-
[248]
A wavelet tour of signal processing
St ´ephane Mallat. A wavelet tour of signal processing. Elsevier, 1999
1999
-
[249]
Group invariant scattering
St ´ephane Mallat. Group invariant scattering. Communications on Pure and Applied Mathematics, 65(10):1331–1398, 2012
2012
-
[250]
Application of artificial intelligence in health- care: chances and challenges
Ravi Manne and Sneha C Kantheti. Application of artificial intelligence in health- care: chances and challenges. Current Journal of Applied Science and Technology, 40 (6):78–89, 2021. BIBLIOGRAPHY 211
2021
-
[251]
Building a large annotated corpus of english: The penn treebank
Mary Ann Marcinkiewicz. Building a large annotated corpus of english: The penn treebank. Using Large Corpora, page 273, 1994
1994
-
[252]
Rotation equiv- ariant vector field networks
Diego Marcos, Michele Volpi, Nikos Komodakis, and Devis Tuia. Rotation equiv- ariant vector field networks. In Proceedings of the IEEE International Conference on Computer Vision, pages 5048–5057, 2017
2017
-
[253]
Scale equiv- ariance in cnns with vector fields
Diego Marcos, Benjamin Kellenberger, Sylvain Lobry, and Devis Tuia. Scale equiv- ariance in cnns with vector fields. arXiv preprint arXiv:1807.11783, 2018
2018 arXiv
-
[254]
Invariant and equivariant graph networks
Haggai Maron, Heli Ben-Hamu, Nadav Shamir, and Yaron Lipman. Invariant and equivariant graph networks. arXiv preprint arXiv:1812.09902, 2018
2018 arXiv
-
[255]
On learning sets of symmetric elements
Haggai Maron, Or Litany, Gal Chechik, and Ethan Fetaya. On learning sets of symmetric elements. arXiv preprint arXiv:2002.08599, 2020
2002 arXiv
-
[256]
Laughing hyena distillery: Extracting compact recurrences from convolutions
Stefano Massaroli, Michael Poli, Dan Fu, Hermann Kumbong, David W Romero, Rom Parnichkun, Aman Timalsina, Quinn McIntyre, Beidi Chen, Atri Rudra, Ce Zhang, Christopher R ´e, Stefano Ermon, and Yoshua Bengio. Laughing hyena distillery: Extracting compact recurrences from convol...
2023
-
[257]
Voxnet: A 3d convolutional neural net- work for real-time object recognition
Daniel Maturana and Sebastian Scherer. Voxnet: A 3d convolutional neural net- work for real-time object recognition. In 2015 IEEE/RSJ international conference on intelligent robots and systems (IROS), pages 922–928. IEEE, 2015
2015
-
[258]
Efficient-capsnet: Capsule network with self-attention routing.arXiv preprint arXiv:2101.12491, 2021
Vittorio Mazzia, Francesco Salvetti, and Marcello Chiaberge. Efficient-capsnet: Capsule network with self-attention routing.arXiv preprint arXiv:2101.12491, 2021
2021 arXiv
-
[259]
A logical calculus of the ideas immanent in nervous activity
Warren S McCulloch and Walter Pitts. A logical calculus of the ideas immanent in nervous activity. The bulletin of mathematical biophysics, 5:115–133, 1943
1943
-
[260]
Mogrifier lstm
G ´abor Melis, Tom ´aˇs Ko ˇcisk`y, and Phil Blunsom. Mogrifier lstm. arXiv preprint arXiv:1909.01792, 2019
1909 arXiv
-
[261]
Occupancy networks: Learning 3d reconstruction in function space
Lars Mescheder, Michael Oechsle, Michael Niemeyer, Sebastian Nowozin, and Andreas Geiger. Occupancy networks: Learning 3d reconstruction in function space. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recogni- tion, pages 4460–4470, 2019
2019
-
[262]
Nerf: Representing scenes as neural radiance fields for view synthesis
Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view synthesis. arXiv preprint arXiv:2003.08934, 2020
2003 arXiv
-
[263]
Designing neural net- works using genetic algorithms
Geoffrey F Miller, Peter M Todd, and Shailesh U Hegde. Designing neural net- works using genetic algorithms. In ICGA, volume 89, pages 379–384, 1989
1989
-
[264]
A simple neural attentive meta-learner
Nikhil Mishra, Mostafa Rohaninejad, Xi Chen, and Pieter Abbeel. A simple neural attentive meta-learner. arXiv preprint arXiv:1707.03141, 2017. 212 BIBLIOGRAPHY
2017 arXiv
-
[265]
On the number of linear regions of deep neural networks
Guido F Montufar, Razvan Pascanu, Kyunghyun Cho, and Yoshua Bengio. On the number of linear regions of deep neural networks. InAdvances in neural information processing systems, pages 2924–2932, 2014
2014
-
[266]
Properties of auditory stream formation
Brian CJ Moore and Hedwig E Gockel. Properties of auditory stream formation. Philosophical Transactions of the Royal Society B: Biological Sciences , 367(1591):919– 931, 2012
2012
-
[267]
Instant neu- ral graphics primitives with a multiresolution hash encoding
Thomas M ¨uller, Alex Evans, Christoph Schied, and Alexander Keller. Instant neu- ral graphics primitives with a multiresolution hash encoding. ACM Transactions on Graphics (ToG), 41(4):1–15, 2022
2022
-
[268]
Rectified linear units improve restricted boltz- mann machines
Vinod Nair and Geoffrey E Hinton. Rectified linear units improve restricted boltz- mann machines. In Icml, 2010
2010
-
[269]
Listops: A diagnostic dataset for latent tree learning
Nikita Nangia and Samuel R Bowman. Listops: A diagnostic dataset for latent tree learning. arXiv preprint arXiv:1804.06028, 2018
2018 arXiv
-
[270]
Deeparchitect: Automatically designing and training deep architectures
Renato Negrinho and Geoff Gordon. Deeparchitect: Automatically designing and training deep architectures. arXiv preprint arXiv:1704.08792, 2017
2017 arXiv
-
[271]
Robust deep learning for computer vision to counteract data scarcity and label noise
Duc Nguyen. Robust deep learning for computer vision to counteract data scarcity and label noise. PhD thesis, 01 2020
2020
-
[272]
S4nd: Modeling images and videos as multidimensional signals using state spaces
Eric Nguyen, Karan Goel, Albert Gu, Gordon W Downs, Preey Shah, Tri Dao, Stephen A Baccus, and Christopher R ´e. S4nd: Modeling images and videos as multidimensional signals using state spaces. arXiv preprint arXiv:2210.06583, 2022
-
[273]
Hyenadna: Long-range genomic sequence modeling at single nucleotide resolution
Eric Nguyen, Michael Poli, Marjan Faizi, Armin Thomas, Callum Birch-Sykes, Michael Wornow, Aman Patel, Clayton Rabideau, Stefano Massaroli, Yoshua Ben- gio, et al. Hyenadna: Long-range genomic sequence modeling at single nucleotide resolution. arXiv preprint arXiv:2306.15794, 2023
2023 arXiv
-
[274]
Optimal trans- port kernels for sequential and parallel neural architecture search
Vu Nguyen, Tam Le, Makoto Yamada, and Michael A Osborne. Optimal trans- port kernels for sequential and parallel neural architecture search. In International Conference on Machine Learning, pages 8084–8095. PMLR, 2021
2021
-
[275]
schyena: Foun- dation model for full-length single-cell rna-seq analysis in brain
Gyutaek Oh, Baekgyu Choi, Inkyung Jung, and Jong Chul Ye. schyena: Foun- dation model for full-length single-cell rna-seq analysis in brain. arXiv preprint arXiv:2310.02713, 2023
2023 arXiv
-
[276]
The role of context in object recognition
Aude Oliva and Antonio Torralba. The role of context in object recognition. Trends in cognitive sciences, 11(12):520–527, 2007
2007
-
[277]
Wavenet: A generative model for raw audio.arXiv preprint arXiv:1609.03499, 2016
Aaron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu. Wavenet: A generative model for raw audio.arXiv preprint arXiv:1609.03499, 2016
2016 arXiv
-
[278]
Deepsdf: Learning continuous signed distance functions for shape BIBLIOGRAPHY 213 representation
Jeong Joon Park, Peter Florence, Julian Straub, Richard Newcombe, and Steven Lovegrove. Deepsdf: Learning continuous signed distance functions for shape BIBLIOGRAPHY 213 representation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 165...
2019
-
[279]
Bam: Bottle- neck attention module
Jongchan Park, Sanghyun Woo, Joon-Young Lee, and In So Kweon. Bam: Bottle- neck attention module. arXiv preprint arXiv:1807.06514, 2018
2018 arXiv
-
[280]
Hyper- nerf: A higher-dimensional representation for topologically varying neural radi- ance fields
Keunhong Park, Utkarsh Sinha, Peter Hedman, Jonathan T Barron, Sofien Bouaziz, Dan B Goldman, Ricardo Martin-Brualla, and Steven M Seitz. Hyper- nerf: A higher-dimensional representation for topologically varying neural radi- ance fields. arXiv preprint arXiv:2106.13228, 2021
2021 arXiv
-
[281]
How to construct deep recurrent neural networks
Razvan Pascanu, Caglar Gulcehre, Kyunghyun Cho, and Yoshua Bengio. How to construct deep recurrent neural networks. arXiv preprint arXiv:1312.6026, 2013
2013 arXiv
-
[282]
On the difficulty of training recurrent neural networks
Razvan Pascanu, Tomas Mikolov, and Yoshua Bengio. On the difficulty of training recurrent neural networks. In International conference on machine learning , pages 1310–1318, 2013
2013
-
[283]
Attention
Harold Pashler. Attention. Psychology Press, 2016
2016
-
[284]
Causality: Models, Reasoning, and Inference
Judea Pearl. Causality: Models, Reasoning, and Inference . Cambridge University Press, Cambridge, UK, 2009. ISBN 978-0521895606
2009
-
[285]
Deep scattering spectrum with deep neu- ral networks
Vijayaditya Peddinti, TaraN Sainath, Shay Maymon, Bhuvana Ramabhadran, David Nahamoo, and Vaibhava Goel. Deep scattering spectrum with deep neu- ral networks. In 2014 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 210–214. IEEE, 2014
2014
-
[286]
Large ker- nel matters–improve semantic segmentation by global convolutional network
Chao Peng, Xiangyu Zhang, Gang Yu, Guiming Luo, and Jian Sun. Large ker- nel matters–improve semantic segmentation by global convolutional network. In Proceedings of the IEEE conference on computer vision and pattern recognition , pages 4353–4361, 2017
2017
-
[287]
Film: Visual reasoning with a general conditioning layer
Ethan Perez, Florian Strub, Harm De Vries, Vincent Dumoulin, and Aaron Courville. Film: Visual reasoning with a general conditioning layer. InProceedings of the AAAI conference on artificial intelligence, volume 32, 2018
2018
-
[288]
Elements of causal infer- ence: foundations and learning algorithms
Jonas Peters, Dominik Janzing, and Bernhard Sch ¨olkopf. Elements of causal infer- ence: foundations and learning algorithms. The MIT Press, 2017
2017
-
[289]
Efficient neural architecture search via parameters sharing
Hieu Pham, Melody Guan, Barret Zoph, Quoc Le, and Jeff Dean. Efficient neural architecture search via parameters sharing. In International conference on machine learning, pages 4095–4104. PMLR, 2018
2018
-
[290]
Environmental sound classification with convolutional neural net- works
Karol J Piczak. Environmental sound classification with convolutional neural net- works. In 2015 IEEE 25th International Workshop on Machine Learning for Signal Processing (MLSP), pages 1–6. IEEE, 2015. 214 BIBLIOGRAPHY
2015
-
[291]
Resolution learning in deep convolutional networks using scale-space theory
Silvia L Pintea, Nergis Tomen, Stanley F Goes, Marco Loog, and Jan C van Gemert. Resolution learning in deep convolutional networks using scale-space theory. arXiv preprint arXiv:2106.03412, 2021
2021 arXiv
-
[292]
Hyena hierarchy: Towards larger convolutional language models
Michael Poli, Stefano Massaroli, Eric Nguyen, Daniel Y Fu, Tri Dao, Stephen Baccus, Yoshua Bengio, Stefano Ermon, and Christopher R ´e. Hyena hierarchy: Towards larger convolutional language models. arXiv preprint arXiv:2302.10866 , 2023
2023 arXiv
-
[293]
Randomly weighted cnns for (music) audio classifi- cation
Jordi Pons and Xavier Serra. Randomly weighted cnns for (music) audio classifi- cation. In ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 336–340. IEEE, 2019
2019
-
[294]
End-to-end learning for music audio tagging at scale
Jordi Pons, Oriol Nieto, Matthew Prockup, Erik Schmidt, Andreas Ehmann, and Xavier Serra. End-to-end learning for music audio tagging at scale. arXiv preprint arXiv:1711.02520, 2017
2017 arXiv
-
[295]
Timbre analysis of music audio signals with convolutional neural networks
Jordi Pons, Olga Slizovskaia, Rong Gong, Emilia G´omez, and Xavier Serra. Timbre analysis of music audio signals with convolutional neural networks. In 2017 25th European Signal Processing Conference (EUSIPCO), pages 2744–2748. IEEE, 2017
2017
-
[296]
Volumetric and multi-view cnns for object classification on 3d data
Charles R Qi, Hao Su, Matthias Nießner, Angela Dai, Mengyuan Yan, and Leonidas J Guibas. Volumetric and multi-view cnns for object classification on 3d data. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 5648–5656, 2016
2016
-
[297]
Pointnet: Deep learn- ing on point sets for 3d classification and segmentation
Charles R Qi, Hao Su, Kaichun Mo, and Leonidas J Guibas. Pointnet: Deep learn- ing on point sets for 3d classification and segmentation. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 652–660, 2017
2017
-
[2011]
Springer, 2011
Proceedings 14, pages 70–83. Springer, 2011
2011
-
[2022]
URL https://openreview.net/forum?id=uYLFoz1vlAC
-
[2023]
URL https://openreview.net/forum?id=bXNl-myZkJl
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.