OR-Space is a benchmark for LLM agents performing full-lifecycle optimization tasks across Build, Revise, and Explain modes in executable multi-artifact workspaces.
Mamo: A mathematical modeling benchmark with solvers
7 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
representative citing papers
DeInfer reduces parallel inference communication cost for decomposed LLMs by up to 78% by moving collective operations into the low-rank latent space and redesigning KV-cache reconstruction for static graph compatibility.
ReLoop closes the feasibility-correctness gap in LLM optimization code via structured generation and behavioral verification with parameter perturbations, reaching 100% executability and accuracy gains on benchmarks while releasing RetailOpt-190.
LLMs prompted with domain knowledge can generate runnable, numerically valid code for stiff and non-stiff ODEs on new diagnostic and 1000-task benchmarks.
PARM adapts reward models to multi-stage LLM pipelines via pipeline data and direct preference optimization, improving execution rate and solving accuracy on optimization benchmarks and showing transfer to GSM8K.
AutoOR uses synthetic data generation and RL post-training with solver feedback to enable 8B LLMs to autoformalize linear, mixed-integer, and non-linear OR problems, matching larger models on benchmarks.
Opt-Verifier adds structure-side and solution-side verification to LLM-generated optimization models and reports over 20% accuracy gains on standard benchmarks.
citing papers explorer
-
OR-Space: A Full-Lifecycle Workspace Benchmark for Industrial Optimization Agents
OR-Space is a benchmark for LLM agents performing full-lifecycle optimization tasks across Build, Revise, and Explain modes in executable multi-artifact workspaces.
-
Co-evolving Agent Architectures and Interpretable Reasoning for Automated Optimization
DeInfer reduces parallel inference communication cost for decomposed LLMs by up to 78% by moving collective operations into the low-rank latent space and redesigning KV-cache reconstruction for static graph compatibility.
-
ReLoop: Structured Modeling and Behavioral Verification for Reliable LLM-Based Optimization
ReLoop closes the feasibility-correctness gap in LLM optimization code via structured generation and behavioral verification with parameter perturbations, reaching 100% executability and accuracy gains on benchmarks while releasing RetailOpt-190.
-
SciML Agents: Write the Solver, Not the Solution
LLMs prompted with domain knowledge can generate runnable, numerically valid code for stiff and non-stiff ODEs on new diagnostic and 1000-task benchmarks.
-
PARM: Pipeline-Adapted Reward Model
PARM adapts reward models to multi-stage LLM pipelines via pipeline data and direct preference optimization, improving execution rate and solving accuracy on optimization benchmarks and showing transfer to GSM8K.
-
AutoOR: Scalably Post-training LLMs to Autoformalize Operations Research Problems
AutoOR uses synthetic data generation and RL post-training with solver feedback to enable 8B LLMs to autoformalize linear, mixed-integer, and non-linear OR problems, matching larger models on benchmarks.
-
Opt-Verifier: Unleashing the Power of LLMs for Optimization Modeling via Dual-Side Verification
Opt-Verifier adds structure-side and solution-side verification to LLM-generated optimization models and reports over 20% accuracy gains on standard benchmarks.