Closed-loop LLM search with AST-generated examples discovers non-standard channel widths that improve vision model performance over initial architectures on CIFAR-100.
From brute force to semantic insight: Performance-guided data transformation design with llms.arXiv preprint, arXiv:2601.03808, 2026
5 Pith papers cite this work. Polarity classification is still indexing.
years
2026 5representative citing papers
LEMUR 2 releases a multi-generator, multi-task neural-architecture corpus with real-device latency metadata intended as fuel for LLM-driven AutoML.
Automated search of 4463 heterogeneous 4-expert MoE models found enumeration bias anchoring the space to AirNet and ranked ShuffleNet/MobileNetV3 as top performers.
Across 3,938 CIFAR-10 runs, CosineAnnealingWarmRestarts leads on mean accuracy while scheduler preference is strongly architecture-dependent; CyclicLR helps some mobile/conv nets but not overall.
Empirical grid search over 18 loss-optimizer pairs on 33 LEMUR architectures shows cross-entropy with Adam/AdamW is most robust while NGL and SGD-based pairings vary sharply by model family.
citing papers explorer
-
Closed-Loop LLM Discovery of Non-Standard Channel Priors in Vision Models
Closed-loop LLM search with AST-generated examples discovers non-standard channel widths that improve vision model performance over initial architectures on CIFAR-100.
-
LEMUR 2: Unlocking Neural Network Diversity for AI
LEMUR 2 releases a multi-generator, multi-task neural-architecture corpus with real-device latency metadata intended as fuel for LLM-driven AutoML.
-
Systematic Exploration of 4-Expert Heterogeneous Mixture-of-Experts via Automated Pipeline Search
Automated search of 4463 heterogeneous 4-expert MoE models found enumeration bias anchoring the space to AirNet and ranked ShuffleNet/MobileNetV3 as top performers.
-
Systematic Evaluation of Learning Rate Scheduling Strategies Across Heterogeneous Architectures
Across 3,938 CIFAR-10 runs, CosineAnnealingWarmRestarts leads on mean accuracy while scheduler preference is strongly architecture-dependent; CyclicLR helps some mobile/conv nets but not overall.
-
Towards Robust Training in NNGPT AutoML Pipeline: A Loss-Optimizer Pairing Selection Study
Empirical grid search over 18 loss-optimizer pairs on 33 LEMUR architectures shows cross-entropy with Adam/AdamW is most robust while NGL and SGD-based pairings vary sharply by model family.