Mesa-optimization arises when learned models act as optimizers with objectives that can differ from their training loss, creating alignment risks in advanced machine learning.
Towards practical verification of machine learning: The case of computer vision systems.arXiv preprint arXiv:1712.01785
3 Pith papers cite this work. Polarity classification is still indexing.
years
2019 3representative citing papers
DriveFI, a Bayesian ML-based fault injection engine, identifies 561 safety-critical faults in AV systems in under 4 hours on NVIDIA and Baidu stacks, while random injection over weeks found none.
Invariance-inducing regularization using worst-case transformations reduces relative error by 20% on CIFAR10 transformed examples, improves standard accuracy on SVHN, outperforms equivariant networks, and proves no accuracy-robustness trade-off in the infinite data limit.
citing papers explorer
-
Risks from Learned Optimization in Advanced Machine Learning Systems
Mesa-optimization arises when learned models act as optimizers with objectives that can differ from their training loss, creating alignment risks in advanced machine learning.
-
ML-based Fault Injection for Autonomous Vehicles: A Case for Bayesian Fault Injection
DriveFI, a Bayesian ML-based fault injection engine, identifies 561 safety-critical faults in AV systems in under 4 hours on NVIDIA and Baidu stacks, while random injection over weeks found none.
-
Invariance-inducing regularization using worst-case transformations suffices to boost accuracy and spatial robustness
Invariance-inducing regularization using worst-case transformations reduces relative error by 20% on CIFAR10 transformed examples, improves standard accuracy on SVHN, outperforms equivariant networks, and proves no accuracy-robustness trade-off in the infinite data limit.