Re-evaluating four LLM code-efficiency benchmarks with 30-run statistical testing shows 93.89% of 'performant' implementations are indistinguishable from baselines; a multi-agent test-generation framework reveals hidden significant improvements in ~24% of previously non-significant tasks.
Journal of Machine learning research , volume=
3 Pith papers cite this work. Polarity classification is still indexing.
years
2026 3representative citing papers
Di-COT is an unsupervised contrastive method that stochastically partitions time-series windows into overlapping sub-blocks to learn representations without augmentation, reporting SOTA results on classification and transfer tasks across multiple benchmarks while cutting training time.
Sensitivity analysis of tactical wireless network design via Tabu Search reveals scale-dependent transitions where some parameters reshape topology while others mainly scale performance magnitude.
citing papers explorer
-
Rethinking Code Performance Benchmarks for LLMs
Re-evaluating four LLM code-efficiency benchmarks with 30-run statistical testing shows 93.89% of 'performant' implementations are indistinguishable from baselines; a multi-agent test-generation framework reveals hidden significant improvements in ~24% of previously non-significant tasks.
-
Divide and Contrast: Learning Robust Temporal Features without Augmentation
Di-COT is an unsupervised contrastive method that stochastically partitions time-series windows into overlapping sub-blocks to learn representations without augmentation, reporting SOTA results on classification and transfer tasks across multiple benchmarks while cutting training time.
-
Sensitivity Analysis of Tactical Wireless Network Design Under Realistic Operational Constraints
Sensitivity analysis of tactical wireless network design via Tabu Search reveals scale-dependent transitions where some parameters reshape topology while others mainly scale performance magnitude.