FrontierOR benchmark shows frontier LLMs outperform Gurobi on solution quality and efficiency in only 31% of one-shot cases and 50% with test-time evolution on hard large-scale optimization tasks.
arXiv preprint arXiv:2601.09635 (2026)
7 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
years
2026 7representative citing papers
OptSkills clusters optimization problems by archetypes, distills workflow skills from successful trajectories, and achieves 68.27% micro-averaged accuracy on diverse benchmarks while outperforming DeepSeek-V3.2-Thinking by 4.53% on MIPLIB-NL.
LLM agent translates user prompts into model patches and selects primal-aware re-optimization techniques for large-scale dynamic problems, shown on supply-chain and exam-scheduling cases.
AutoREM augments LLMs with a structured memory of failed reformulation trajectories to improve accuracy and efficiency on robust optimization tasks without parameter updates or expert knowledge.
InvEvolve evolves inventory policies using LLMs with RL and provides statistical safety guarantees, outperforming classical and DL methods on synthetic and real data.
A survey compiling roles, applications, benchmarks, challenges, and future directions for large language models in operations research.
citing papers explorer
-
FrontierOR: Benchmarking LLMs' Capacity for Efficient Algorithm Design in Large-Scale Optimization
FrontierOR benchmark shows frontier LLMs outperform Gurobi on solution quality and efficiency in only 31% of one-shot cases and 50% with test-time evolution on hard large-scale optimization tasks.
-
OptSkills: Learning Generalizable Optimization Skills from Problem Archetypes via Cluster-Based Distillation
OptSkills clusters optimization problems by archetypes, distills workflow skills from successful trajectories, and achieves 68.27% micro-averaged accuracy on diverse benchmarks while outperforming DeepSeek-V3.2-Thinking by 4.53% on MIPLIB-NL.
-
Democratizing Large-Scale Re-Optimization with LLM-Guided Model Patches
LLM agent translates user prompts into model patches and selects primal-aware re-optimization techniques for large-scale dynamic problems, shown on supply-chain and exam-scheduling cases.
-
Automated Reformulation of Robust Optimization via Memory-Augmented Large Language Models
AutoREM augments LLMs with a structured memory of failed reformulation trajectories to improve accuracy and efficiency on robust optimization tasks without parameter updates or expert knowledge.
-
InvEvolve: Evolving White-Box Inventory Policies via Large Language Models with Performance Guarantees
InvEvolve evolves inventory policies using LLMs with RL and provides statistical safety guarantees, outperforming classical and DL methods on synthetic and real data.
-
Large Language Models for Operations Research: A Comprehensive Survey
A survey compiling roles, applications, benchmarks, challenges, and future directions for large language models in operations research.
- ReLoop: Structured Modeling and Behavioral Verification for Reliable LLM-Based Optimization