Pith. sign in

REVIEW 3 major objections 5 minor 96 references

SCOPE claims black-box combinatorial search improves by evolving synthetic objectives that guide a fixed search engine, and shows gains across 15+ problems under a 64-query budget.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 04:19 UTC pith:RTR55I3Y

load-bearing objection Genuinely new framing plus a carefully run study, but the 'consistently improves' claim outruns the five-run point estimates in the headline tables. the 3 major comments →

arxiv 2607.27630 v1 pith:RTR55I3Y submitted 2026-07-30 cs.AI

SCOPE: Synthetic Conditional Objectives for Policy Evolution in Black-Box Combinatorial Optimization

classification cs.AI
keywords black-box combinatorial optimizationsynthetic conditional objectivesLLM-driven program searchpolicy evolutionobjective designportfolio selectionThompson samplinglimited evaluation budget
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper tries to establish that in black-box combinatorial optimization, where the true objective can only be queried a few hundred times, the most useful thing to learn is not a surrogate model of the true objective or a new search operator, but a set of synthetic scoring functions that guide a fixed search engine. These synthetic conditional objectives are generated by a language model from the accumulated query history, judged by the hidden-objective quality of the policies they induce, and frozen into a small complementary portfolio; at deployment, query allocation among the portfolio's workers is decided by Thompson sampling. If this is right, it matters because scarce black-box evaluations get converted into dense, free guidance, and objective design becomes an evolvable artifact rather than a hand-tuned component. Across more than fifteen combinatorial problems at sizes 100 and 200, with local-search, route-search, and constructive backends and two language models, the method consistently ranks best among learned heuristic- and objective-design methods at 64 hidden-objective queries, and in several settings beats a white-box reference that spends 1,968 true-objective evaluations.

Core claim

The paper's central claim is that an objective for black-box combinatorial search need not approximate the hidden function; it only needs to induce a policy that finds good solutions. SCOPE operationalizes this by asking a language model to write executable, deterministic scoring functions objective(solution, instance) that respect the public problem contract and are conditioned on the accumulated black-box observations. Each function parameterizes a copy of a fixed search engine, and its worth is measured by the median normalized rank of the best oracle values its induced policy produces on training instances. A program graph then evolves these objectives: parent selection balances downstre

What carries the argument

The central object is the synthetic conditional objective: an executable, deterministic function generated by an LLM from the accumulated history O_t and restricted to the public solution/instance fields. Its role is to reshape the search trajectories of a fixed backend A rather than to reconstruct the hidden f. The framework's central identity is that downstream policy utility—measured by the oracle quality of solutions produced when A is run under the objective—is the correct criterion for judging an objective. Around this, the machinery includes a program graph with quality/novelty/exploration parent selection, contrastive evidence from pairwise ordering failures, behavioral signatures ba

Load-bearing premise

The load-bearing premise is that, for each benchmark, there exist executable deterministic scoring functions expressible through the public solution/instance fields whose induced search trajectories on training instances transfer to larger test instances under the fixed search engine.

What would settle it

Train SCOPE on n=50/75 instances and test on n=1,000 instances from the same generator; if its gap to a same-budget white-box reference worsens steadily with n (or becomes positive), the small-to-large transfer assumption fails.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Under a 64-query budget, learned synthetic objectives can beat a white-box reference that spends 1,968 true-objective evaluations, implying that search guidance can substitute for direct oracle access.
  • The advantage persists across local-search, route-set-search, and constructive backends and across two language models, suggesting the mechanism is not tied to one search implementation.
  • A portfolio of a few complementary objectives outperforms the single best validated objective, so maintaining several search landscapes is part of the gain.
  • Ablation studies show each component—quality-guided selection, contrastive evidence, archived mechanism retrieval, behavioral novelty—contributes; removing any one degrades results by roughly 2.5–7.5 percent.
  • Gray-box hints help but do not fully close the gap on stochastic or long-range-interaction objectives, indicating the method's dependence on the disclosed public contract.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If the central claim holds, the practical bottleneck in costly combinatorial search may be the evaluation criterion rather than the proposal operator; a natural test is applying the same objective-evolution loop to continuous or mixed-integer black-box problems with richer public contracts.
  • The framework implies discovered objectives are coupled to the fixed search engine and public representation; a concrete extension is to measure whether objectives found with one backend transfer to another backend without re-discovery.
  • The Thompson-sampling portfolio suggests viewing search landscapes as bandit arms; one could extend it to dynamically grow or shrink the portfolio as the budget or instance distribution changes.
  • Because the paper's generalization claim rests on small-to-large transfer within one generator, a stress test on much larger instances or shifted generators would clarify how far the discovered landscapes generalize.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper proposes SCOPE, a framework for black-box combinatorial optimization in which the hidden objective is never reconstructed. Instead, an LLM synthesizes executable synthetic objectives conditioned on accumulated oracle history; each objective parameterizes a copy of a fixed search engine, and a program graph evolves these objectives through parent selection, contrastive evidence, behavioral novelty filtering, and downstream-quality scoring. After validation, a small portfolio of objectives is frozen and deployed on unseen instances, with Thompson sampling allocating oracle queries among persistent workers. Experiments span more than fifteen problems at n=100 and n=200 with GLS, ALNS, and constructive backends, comparing SCOPE against heuristic-design and objective-design baselines under a 64-query budget. The paper claims that SCOPE consistently improves black-box search performance and generalizes across combinatorial structures, sometimes beating a White-Box reference that uses many more true-objective evaluations.

Significance. If the empirical claims hold, SCOPE is a useful contribution to LLM-based objective and heuristic design. The experimental protocol is unusually careful in several respects: strict train/validation/test separation, frozen programs before deployment, shared instances and seeds across methods, explicit LLM-call accounting, fixed downstream backends, and ablations of the main components. This discipline makes the reported comparisons substantially more trustworthy than is common in this area. However, the central 'consistently improves' claim is currently supported only by five-run point estimates without standard deviations, confidence intervals, or paired significance tests (Tables 1–3). Several decisive margins are under 1.5 percentage points, and the number of effective independent replicates is five. The evidence is strong enough to warrant a major revision, but not yet strong enough to establish the headline claim as stated.

major comments (3)
  1. [§5.1, Tables 1–3] The headline claim of 'consistent improvement' rests on five-run point estimates. Tables 1 and 2 report only gaps to White-Box, and Table 3 reports only mean percentage gaps; none reports standard deviations, confidence intervals, or significance tests. Because results are aggregated at the run level (Appendix D.6), the effective sample size is five. With five paired runs, even a perfect win record cannot yield p<0.05 on a two-sided Wilcoxon signed-rank test, and several reported margins are small (e.g., Table 1: RiskTour n=100 with gpt-4o-mini, SCOPE −0.33% vs. FunSearch 0.17%; LoadFlow n=100, SCOPE 1.85% vs. Eureka 3.30%). The authors should report per-run values or paired bootstrap/permutation confidence intervals, or substantially increase the number of independent discovery runs. Without this, the word 'consistently' is not supported by the evidence presented.
  2. [§5.3, Table 4] Table 4 is the only main table with mean±standard deviation, and it illustrates the problem: on Stochastic Machine Assignment at n=200, SCOPE has the best mean (3.779) but its standard deviation (0.091) is more than ten times that of EoH (0.018) and overlaps heavily with every baseline. The text states that 'SCOPE obtains the best learned-method result' at both sizes; statistically, this is at best an unverified ranking of noisy means. The same uncertainty quantification should be applied to the central comparisons in Tables 1–3. Since the deployment process is stochastic at the run level, reporting only endpoint means hides exactly the variability that determines whether the framework's advantage is real.
  3. [Appendix E.8; §6 Conclusion] The paper's own limitations section explicitly acknowledges that 'finite evidence is not semantic identification,' that portfolio selection is a finite-sample decision with unstable validation ranks, and that attribution under limited budgets is joint rather than causal. These are honest and appropriate caveats, but they stand in tension with the unqualified Conclusion that SCOPE 'consistently improves search across diverse combinatorial problems.' The authors should either soften the central claim to match the demonstrated evidence or provide the missing uncertainty analysis that would justify the stronger wording. A useful intermediate step is to report the fraction of paired runs/conditions where SCOPE outperforms each baseline, rather than only average gaps.
minor comments (5)
  1. [Title page] Typo in affiliation: 'Techonology' should be 'Technology.'
  2. [Appendix D.7] The source code is stated to be released only after publication and is currently unavailable. Given that the paper's main contribution is an empirical framework, an anonymized artifact or code supplement would materially aid review and reproducibility.
  3. [Figures 2 and 3] Figure 2 is very dense and the text is difficult to read at normal size; Figure 3 would benefit from error bars or confidence bands, especially since the time-series curves are based on five runs.
  4. [§5.2, Table 3] The 'MeanGap' row appears to be a family-level average of paired gaps, but the table does not state which problems are included in each family, nor whether the mean is weighted equally across problem sizes. This should be clarified in the table caption.
  5. [Appendix D.3] The validation candidate limit is described as 'At most 12 objectives,' while Appendix C.6 says 'Every valid candidate' is reevaluated. The relationship between the number of valid programs and the 12-program cap should be made explicit.

Circularity Check

0 steps flagged

No significant circularity: objectives are trained and validated on disjoint data, and test metrics use frozen programs.

full rationale

I traced the derivation chain from Eq. (4)-(6) through the program graph, validation-based portfolio selection, and frozen deployment. Synthetic objectives are generated from public solution/instance contracts and rank-only contrastive evidence (Eq. 8), using oracle observations only from disjoint training instances. Their utility is measured by the downstream oracle performance of the policies they induce on validation instances (Eq. 7, 12-13), and all selected programs are frozen before test instances are opened. Reported test gaps (Eq. 48-49) compare frozen portfolios on unseen instances against a finite-budget White-Box reference. No equation defines objective quality in terms of the test objective, and no parameter fitted to test data is renamed a prediction. The statement that a synthetic objective is valued by its induced downstream policy is an explicit evaluation choice, not a hidden reuse of the target. Self-citations (MOTIF, B2B) appear only in related-work positioning and do not carry a load-bearing premise or uniqueness theorem. The absence of error bars and the small margins in Table 1 are statistical power/correctness concerns, not circularity.

Axiom & Free-Parameter Ledger

8 free parameters · 6 axioms · 0 invented entities

The central claim rests on hand-set algorithmic hyperparameters and on assumptions about LLM capability, public-contract expressiveness, behavioral novelty as a diversity signal, and training-to-test transfer. No new physical entities or fitted physical constants are introduced; the 'synthetic conditional objective' is an algorithmic artifact rather than a postulated entity.

free parameters (8)
  • Quality-novelty-exploration weights λ_q, λ_n, λ_e
    Hand-set coefficients balancing downstream quality, behavioral novelty, and exploration in Eq. (11); values are not reported, and no sensitivity analysis is given.
  • Behavioral-distance mixture weight η
    Hand-set weight between rank distance and structural distance in Eq. (10); value not reported.
  • Contrast-evidence thresholds ε_u, ε_f = 0.15, 0.25
    Thresholds in Eq. (8) defining preference-reversal and under-separated pairs; chosen by hand and stated in Appendix C.3.
  • Novelty admission thresholds = 0.02, 0.08
    Behavioral-distance threshold and archive threshold used to reject redundant children before downstream evaluation; stated in Appendix D.3.
  • Active-set quality fraction = 0.67
    Fraction bounding novelty competition inside the quality-admissible region; stated in Appendix D.3.
  • Portfolio size M = 3
    Number of frozen objectives selected during validation; larger or smaller portfolios are not tested.
  • Discovery schedule and search workloads = 9 rounds; 3 children per phase; active graph 6; probe bank 5×12; k=5
    Hand-set protocol constants in Appendix D.3 that jointly determine the LLM-call budget and downstream evaluation effort.
  • Deployment protocol constants = population 16; 30 local steps per query; Beta(1,1) prior
    Hand-set deployment settings in Appendix D.5 and Algorithm 2; they affect how quickly each worker produces candidates.
axioms (6)
  • domain assumption The hidden objective f is available only through a finite value oracle, and each query costs a unit of budget.
    This is the problem definition in Section 3.1; SCOPE's entire query-accounting and comparison protocol depends on it.
  • domain assumption Useful search guidance can be expressed as deterministic executable objectives over the public solution/instance contract.
    Introduced in Eq. (4) and enforced by the static/dynamic contract checks in Appendix C.2. If hidden structure is not representable in public fields, the method cannot work.
  • domain assumption Contrastive preference-reversal and under-separated pair evidence is informative to the LLM for generating better objectives.
    The revision loop in Eq. (9) and the RANK_FEEDBACK template in Appendix C.4 rely on this empirical assumption about LLM behavior.
  • domain assumption Behavioral novelty measured by rank vectors on public probes and structural signatures approximates search-behavior diversity without leaking oracle outcomes.
    Eq. (10) and Appendix C.5 define novelty in this way; the entire graph-exploration mechanism depends on this proxy being informative.
  • domain assumption Downstream performance on training instances under the fixed search engine A transfers to the validation and test instance distributions at larger sizes.
    The discovery and portfolio-selection pipeline in Algorithm 1 assumes that objectives that induce good search on n=50 training instances remain useful at n=100 and n=200.
  • standard math Beta-Bernoulli Thompson sampling is a reasonable allocation rule for dividing query budget among frozen workers.
    Algorithm 2 and Eq. (14) use a standard Bayesian bandit procedure; no special guarantees beyond standard regret behavior are claimed.

pith-pipeline@v1.3.0-daily-deepseek · 37841 in / 14670 out tokens · 146436 ms · 2026-08-01T04:19:28.494746+00:00 · methodology

0 comments
read the original abstract

Black-box combinatorial optimization requires systematically identifying high-quality solutions under a limited evaluation budget, yet the unknown objective function provides little guidance for deciding where the search should explore next. We introduce SCOPE, a general framework for Synthetic Conditional Objectives for Policy Evolution in Black-Box Combinatorial Optimization. Rather than directly optimizing the inaccessible objective, SCOPE learns a set of synthetic objectives conditioned on the accumulated search history, where each objective is designed to expose a distinct and potentially useful preference over candidate solutions. These objectives are then used to evolve search policies that generate diverse candidates, whose true quality is subsequently assessed through black-box evaluations. The outer loop adaptively updates and selects synthetic objectives according to how effectively their induced policies discover promising regions. In contrast, the inner loop returns a portfolio of top-performing policies to reduce the risk of relying on a single surrogate preference. This formulation reframes objective design as a mechanism for guiding policy exploration, enabling the search process to exploit observed evidence while maintaining structured diversity across discrete solution spaces. Extensive experiments across multiple benchmark problems demonstrate that SCOPE consistently improves black-box search performance under limited evaluation budgets and generalizes well across diverse combinatorial structures.

Figures

Figures reproduced from arXiv: 2607.27630 by Huynh Thi Thanh Binh, Le Cong Bang, Nguyen Huu Duc, Nguyen Viet Tuan Kiet, Tran Cong Dao.

Figure 1
Figure 1. Figure 1: Comparison of conventional black-box optimiza [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Overview of SCOPE. During discovery, (1) each synthetic objective parameterizes a persistent worker of the search [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Multi-task performance of SCOPE. The left column [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

96 extracted references · 27 linked inside Pith

  1. [1]

    Communication, Simulation, and Intelligent Agents: Implications of Personal Intelligent Machines for Medical Education

    Clancey, William J. Communication, Simulation, and Intelligent Agents: Implications of Personal Intelligent Machines for Medical Education. Proceedings of the Eighth International Joint Conference on Artificial Intelligence (IJCAI-83)

  2. [2]

    Classification Problem Solving

    Clancey, William J. Classification Problem Solving. Proceedings of the Fourth National Conference on Artificial Intelligence

  3. [3]

    , title =

    Robinson, Arthur L. , title =. 1980 , doi =. https://science.sciencemag.org/content/208/4447/1019.full.pdf , journal =

  4. [4]

    New Ways to Make Microcircuits Smaller---Duplicate Entry

    Robinson, Arthur L. New Ways to Make Microcircuits Smaller---Duplicate Entry. Science

  5. [5]

    Clancey and Glenn Rennels , abstract =

    Diane Warner Hasling and William J. Clancey and Glenn Rennels , abstract =. Strategic explanations for a diagnostic consultation system , journal =. 1984 , issn =. doi:https://doi.org/10.1016/S0020-7373(84)80003-6 , url =

  6. [6]

    and Rennels, Glenn R

    Hasling, Diane Warner and Clancey, William J. and Rennels, Glenn R. and Test, Thomas. Strategic Explanations in Consultation---Duplicate. The International Journal of Man-Machine Studies

  7. [7]

    Poligon: A System for Parallel Problem Solving

    Rice, James. Poligon: A System for Parallel Problem Solving

  8. [8]

    Transfer of Rule-Based Expertise through a Tutorial Dialogue

    Clancey, William J. Transfer of Rule-Based Expertise through a Tutorial Dialogue

  9. [9]

    The Engineering of Qualitative Models

    Clancey, William J. The Engineering of Qualitative Models

  10. [10]

    2023 , eprint=

    Attention Is All You Need , author=. 2023 , eprint=

  11. [11]

    Pluto: The 'Other' Red Planet

    NASA. Pluto: The 'Other' Red Planet

  12. [12]

    Proceedings of the 35th International Conference on Machine Learning , series =

    Baptista, Ricardo and Poloczek, Matthias , title =. Proceedings of the 35th International Conference on Machine Learning , series =. 2018 , publisher =

  13. [13]

    Pawan and Dupont, Emilien and Ruiz, Francisco J

    Romera-Paredes, Bernardino and Barekatain, Mohammadamin and Novikov, Alexander and Balog, Matej and Kumar, M. Pawan and Dupont, Emilien and Ruiz, Francisco J. R. and Ellenberg, Jordan S. and Wang, Pengming and Fawzi, Omar and Kohli, Pushmeet and Fawzi, Alhussein , title =. Nature , volume =. 2024 , doi =

  14. [14]

    Proceedings of the 41st International Conference on Machine Learning , series =

    Liu, Fei and Xialiang, Tong and Yuan, Mingxuan and Lin, Xi and Luo, Fu and Wang, Zhenkun and Lu, Zhichao and Zhang, Qingfu , title =. Proceedings of the 41st International Conference on Machine Learning , series =. 2024 , publisher =

  15. [15]

    Advances in Neural Information Processing Systems , volume =

    Ye, Haoran and Wang, Jiarui and Cao, Zhiguang and Berto, Federico and Hua, Chuanbo and Kim, Haeyeon and Park, Jinkyoo and Song, Guojie , title =. Advances in Neural Information Processing Systems , volume =. 2024 , doi =

  16. [16]

    International Conference on Learning Representations , year =

    Ma, Yecheng Jason and Liang, William and Wang, Guanzhi and Huang, De-An and Bastani, Osbert and Jayaraman, Dinesh and Zhu, Yuke and Fan, Linxi and Anandkumar, Anima , title =. International Conference on Learning Representations , year =

  17. [17]

    and Schonlau, Matthias and Welch, William J

    Jones, Donald R. and Schonlau, Matthias and Welch, William J. , title =. Journal of Global Optimization , volume =. 1998 , doi =

  18. [18]

    International Conference on Machine Learning , pages=

    Scaling laws for reward model overoptimization , author=. International Conference on Machine Learning , pages=. 2023 , organization=

  19. [19]

    Proceedings of the AAAI Conference on Artificial Intelligence , volume=

    Motif: Multi-strategy optimization via turn-based interactive framework , author=. Proceedings of the AAAI Conference on Artificial Intelligence , volume=

  20. [20]

    International conference on machine learning , pages=

    Bayesian optimization of combinatorial structures , author=. International conference on machine learning , pages=. 2018 , organization=

  21. [21]

    Nature , volume=

    Mathematical discoveries from program search with large language models , author=. Nature , volume=. 2024 , publisher=

  22. [22]

    arXiv preprint arXiv:2401.02051 , year=

    Evolution of heuristics: Towards efficient automatic algorithm design using large language model , author=. arXiv preprint arXiv:2401.02051 , year=

  23. [23]

    Advances in neural information processing systems , volume=

    Reevo: Large language models as hyper-heuristics with reflective evolution , author=. Advances in neural information processing systems , volume=

  24. [24]

    arXiv preprint arXiv:2310.12931 , year=

    Eureka: Human-level reward design via coding large language models , author=. arXiv preprint arXiv:2310.12931 , year=

  25. [25]

    Proceedings of the 42nd International Conference on Machine Learning , series =

    Zheng, Zhi and Xie, Zhuoliang and Wang, Zhenkun and Hooi, Bryan , title =. Proceedings of the 42nd International Conference on Machine Learning , series =. 2025 , publisher =

  26. [26]

    International Conference on Learning Representations , year =

    Chen, Chentong and Zhong, Mengyuan and Sun, Jianyong and Fan, Ye and Shi, Jialong , title =. International Conference on Learning Representations , year =

  27. [27]

    Voudouris, Christos and Tsang, Edward P. K. , title =. European Journal of Operational Research , volume =. 1999 , doi =

  28. [28]

    Transportation Science , volume =

    Ropke, Stefan and Pisinger, David , title =. Transportation Science , volume =. 2006 , doi =

  29. [29]

    Proceedings of the AAAI Conference on Artificial Intelligence , volume =

    Kiet, Nguyen Viet Tuan and Dao, Tung and Tran, Cong Dao and Binh, Huynh Thi Thanh , title =. Proceedings of the AAAI Conference on Artificial Intelligence , volume =

  30. [30]

    arXiv preprint arXiv:2002.08155 , year =

    Feng, Zhangyin and Guo, Daya and Tang, Duyu and Duan, Nan and Feng, Xiaocheng and Gong, Ming and Shou, Linjun and Qin, Bing and Liu, Ting and Jiang, Daxin and Zhou, Ming , title =. arXiv preprint arXiv:2002.08155 , year =

  31. [31]

    arXiv preprint arXiv:2009.08366 , year =

    Guo, Daya and Ren, Shuo and Lu, Shuai and Feng, Zhangyin and Tang, Duyu and Liu, Shujie and Zhou, Long and Duan, Nan and Svyatkovskiy, Alexey and Fu, Shengyu and others , title =. arXiv preprint arXiv:2009.08366 , year =

  32. [32]

    arXiv preprint arXiv:2107.03374 , year =

    Chen, Mark and Tworek, Jerry and Jun, Heewoo and Yuan, Qiming and Pinto, Henrique Ponde de Oliveira and Kaplan, Jared and Edwards, Harri and Burda, Yuri and Joseph, Nicholas and Brockman, Greg and others , title =. arXiv preprint arXiv:2107.03374 , year =

  33. [33]

    arXiv preprint arXiv:2203.13474 , year =

    Nijkamp, Erik and Pang, Bo and Hayashi, Hiroaki and Tu, Lifu and Wang, Huan and Zhou, Yingbo and Savarese, Silvio and Xiong, Caiming , title =. arXiv preprint arXiv:2203.13474 , year =

  34. [34]

    Competition-Level Code Generation with

    Li, Yujia and Choi, David and Chung, Junyoung and Kushman, Nate and Schrittwieser, Julian and Leblond, R. Competition-Level Code Generation with. Science , volume =. 2022 , doi =

  35. [35]

    arXiv preprint arXiv:2204.05999 , year =

    Fried, Daniel and Aghajanyan, Armen and Lin, Jessy and Wang, Sida and Wallace, Eric and Shi, Freda and Zhong, Ruiqi and Yih, Wen-tau and Zettlemoyer, Luke and Lewis, Mike , title =. arXiv preprint arXiv:2204.05999 , year =

  36. [36]

    arXiv preprint arXiv:2305.06161 , year =

    Li, Raymond and Ben Allal, Loubna and Zi, Yangtian and Muennighoff, Niklas and Kocetkov, Denis and Mou, Chenghao and Marone, Marc and Akiki, Christopher and Li, Jia and Chim, Jenny and others , title =. arXiv preprint arXiv:2305.06161 , year =

  37. [37]

    arXiv preprint arXiv:2308.12950 , year =

    Rozi. arXiv preprint arXiv:2308.12950 , year =

  38. [38]

    and Li, Y

    Guo, Daya and Zhu, Qihao and Yang, Dejian and Xie, Zhenda and Dong, Kai and Zhang, Wentao and Chen, Guanting and Bi, Xiao and Wu, Y. and Li, Y. K. and Luo, Fuli and Xiong, Yingfei and Liang, Wenfeng , title =. arXiv preprint arXiv:2401.14196 , year =

  39. [39]

    arXiv preprint arXiv:2108.07732 , year =

    Austin, Jacob and Odena, Augustus and Nye, Maxwell and Bosma, Maarten and Michalewski, Henryk and Dohan, David and Jiang, Ellen and Cai, Carrie and Terry, Michael and Le, Quoc and Sutton, Charles , title =. arXiv preprint arXiv:2108.07732 , year =

  40. [40]

    arXiv preprint arXiv:2105.09938 , year =

    Hendrycks, Dan and Basart, Steven and Kadavath, Saurav and Mazeika, Mantas and Arora, Akul and Guo, Ethan and Burns, Collin and Puranik, Samir and He, Horace and Song, Dawn and Steinhardt, Jacob , title =. arXiv preprint arXiv:2105.09938 , year =

  41. [41]

    and Guha, Arjun and Greenberg, Michael and Jangda, Abhinav , title =

    Cassano, Federico and Gouwar, John and Nguyen, Daniel and Nguyen, Sydney and Phipps-Costin, Luna and Pinckney, Donald and Yee, Ming-Ho and Zi, Yangtian and Anderson, Carolyn Jane and Feldman, Molly Q. and Guha, Arjun and Greenberg, Michael and Jangda, Abhinav , title =. arXiv preprint arXiv:2208.08227 , year =

  42. [42]

    arXiv preprint arXiv:2306.03091 , year =

    Liu, Tianyang and Xu, Canwen and McAuley, Julian , title =. arXiv preprint arXiv:2306.03091 , year =

  43. [43]

    and Yang, John and Wettig, Alexander and Yao, Shunyu and Pei, Kexin and Press, Ofir and Narasimhan, Karthik , title =

    Jimenez, Carlos E. and Yang, John and Wettig, Alexander and Yao, Shunyu and Pei, Kexin and Press, Ofir and Narasimhan, Karthik , title =. arXiv preprint arXiv:2310.06770 , year =

  44. [44]

    arXiv preprint arXiv:2207.10397 , year =

    Chen, Bei and Zhang, Fengji and Nguyen, Anh and Zan, Daoguang and Lin, Zeqi and Lou, Jian-Guang and Chen, Weizhu , title =. arXiv preprint arXiv:2207.10397 , year =

  45. [45]

    and Lin, Xi Victoria , title =

    Ni, Ansong and Iyer, Srini and Radev, Dragomir and Stoyanov, Ves and Yih, Wen-tau and Wang, Sida I. and Lin, Xi Victoria , title =. arXiv preprint arXiv:2302.08468 , year =

  46. [46]

    Teaching Large Language Models to Self-Debug , journal =

    Chen, Xinyun and Lin, Maxwell and Sch. Teaching Large Language Models to Self-Debug , journal =

  47. [47]

    arXiv preprint arXiv:2303.11366 , year =

    Shinn, Noah and Cassano, Federico and Berman, Edward and Gopinath, Ashwin and Narasimhan, Karthik and Yao, Shunyu , title =. arXiv preprint arXiv:2303.11366 , year =

  48. [48]

    arXiv preprint arXiv:2401.08500 , year =

    Ridnik, Tal and Kredo, Dedy and Friedman, Itamar , title =. arXiv preprint arXiv:2401.08500 , year =

  49. [49]

    and Wettig, Alexander and Lieret, Kilian and Yao, Shunyu and Narasimhan, Karthik and Press, Ofir , title =

    Yang, John and Jimenez, Carlos E. and Wettig, Alexander and Lieret, Kilian and Yao, Shunyu and Narasimhan, Karthik and Press, Ofir , title =. arXiv preprint arXiv:2405.15793 , year =

  50. [50]

    arXiv preprint arXiv:2404.05427 , year =

    Zhang, Yuntong and Ruan, Haifeng and Fan, Zhiyu and Roychoudhury, Abhik , title =. arXiv preprint arXiv:2404.05427 , year =

  51. [51]

    and Gendreau, Michel and Hyde, Matthew and Kendall, Graham and Ochoa, Gabriela and

    Burke, Edmund K. and Gendreau, Michel and Hyde, Matthew and Kendall, Graham and Ochoa, Gabriela and. Hyper-Heuristics: A Survey of the State of the Art , journal =. 2013 , doi =

  52. [52]

    arXiv preprint arXiv:2303.06532 , year =

    Zhao, Qi and Duan, Qiqi and Yan, Bai and Cheng, Shi and Shi, Yuhui , title =. arXiv preprint arXiv:2303.06532 , year =

  53. [53]

    and Leyton-Brown, Kevin and St

    Hutter, Frank and Hoos, Holger H. and Leyton-Brown, Kevin and St. Journal of Artificial Intelligence Research , volume =. 2009 , doi =

  54. [54]

    and Leyton-Brown, Kevin , title =

    Hutter, Frank and Hoos, Holger H. and Leyton-Brown, Kevin , title =. Learning and Intelligent Optimization , series =. 2011 , publisher =

  55. [55]

    L. The. Operations Research Perspectives , volume =. 2016 , doi =

  56. [56]

    arXiv preprint arXiv:2405.20132 , year =

  57. [57]

    arXiv preprint arXiv:2412.14995 , year =

    Dat, Pham Vu Tuan and Doan, Long and Binh, Huynh Thi Thanh , title =. arXiv preprint arXiv:2412.14995 , year =

  58. [58]

    arXiv preprint arXiv:2605.06123 , year =

    Kiet, Nguyen Viet Tuan and Pham, Bui Dinh and Tung, Dao Van and Dao, Tran Cong and Binh, Huynh Thi Thanh , title =. arXiv preprint arXiv:2605.06123 , year =

  59. [59]

    and de Freitas, Nando , title =

    Shahriari, Bobak and Swersky, Kevin and Wang, Ziyu and Adams, Ryan P. and de Freitas, Nando , title =. Proceedings of the IEEE , volume =. 2016 , doi =

  60. [60]

    and Gavves, Efstratios and Welling, Max , title =

    Oh, ChangYong and Tomczak, Jakub M. and Gavves, Efstratios and Welling, Max , title =. Advances in Neural Information Processing Systems , volume =

  61. [61]

    and Nguyen, Vu and Osborne, Michael A

    Ru, Binxin and Alvi, Ahsan S. and Nguyen, Vu and Osborne, Michael A. and Roberts, Stephen J. , title =. Proceedings of the 37th International Conference on Machine Learning , series =. 2020 , publisher =

  62. [62]

    Evolutionary Computation: Comments on the History and Current State , journal =

    B. Evolutionary Computation: Comments on the History and Current State , journal =. 1997 , doi =

  63. [63]

    IEEE Transactions on Evolutionary Computation , volume =

    Krasnogor, Natalio and Smith, Jim , title =. IEEE Transactions on Evolutionary Computation , volume =. 2005 , doi =

  64. [64]

    Journal of Global Optimization , volume =

    Storn, Rainer and Price, Kenneth , title =. Journal of Global Optimization , volume =. 1997 , doi =

  65. [65]

    From Recombination of Genes to the Estimation of Distributions I

    M. From Recombination of Genes to the Estimation of Distributions I. Binary Parameters , booktitle =. 1996 , publisher =

  66. [66]

    , title =

    Rubinstein, Reuven Y. , title =. Methodology and Computing in Applied Probability , volume =. 1999 , doi =

  67. [67]

    Principles and Practice of Constraint Programming---CP98 , series =

    Shaw, Paul , title =. Principles and Practice of Constraint Programming---CP98 , series =. 1998 , publisher =

  68. [68]

    and Harada, Daishi and Russell, Stuart J

    Ng, Andrew Y. and Harada, Daishi and Russell, Stuart J. , title =. Proceedings of the 16th International Conference on Machine Learning , pages =. 1999 , publisher =

  69. [69]

    and Russell, Stuart J

    Ng, Andrew Y. and Russell, Stuart J. , title =. Proceedings of the 17th International Conference on Machine Learning , pages =. 2000 , publisher =

  70. [70]

    and Maas, Andrew L

    Ziebart, Brian D. and Maas, Andrew L. and Bagnell, J. Andrew and Dey, Anind K. , title =. Proceedings of the 23rd AAAI Conference on Artificial Intelligence , pages =

  71. [71]

    and Leike, Jan and Brown, Tom B

    Christiano, Paul F. and Leike, Jan and Brown, Tom B. and Martic, Miljan and Legg, Shane and Amodei, Dario , title =. Advances in Neural Information Processing Systems , volume =

  72. [72]

    and Mishkin, Pamela and Zhang, Chong and Agarwal, Sandhini and Slama, Katarina and Ray, Alex and others , title =

    Ouyang, Long and Wu, Jeffrey and Jiang, Xu and Almeida, Diogo and Wainwright, Carroll L. and Mishkin, Pamela and Zhang, Chong and Agarwal, Sandhini and Slama, Katarina and Ray, Alex and others , title =. Advances in Neural Information Processing Systems , volume =

  73. [73]

    arXiv preprint arXiv:2212.08073 , year =

    Bai, Yuntao and Kadavath, Saurav and Kundu, Sandipan and Askell, Amanda and Kernion, Jackson and Jones, Andy and Chen, Anna and Goldie, Anna and Mirhoseini, Azalia and McKinnon, Cameron and others , title =. arXiv preprint arXiv:2212.08073 , year =

  74. [74]

    and Darrell, Trevor , title =

    Pathak, Deepak and Agrawal, Pulkit and Efros, Alexei A. and Darrell, Trevor , title =. Proceedings of the 34th International Conference on Machine Learning , series =. 2017 , publisher =

  75. [75]

    and Schaul, Tom and Leibo, Joel Z

    Jaderberg, Max and Mnih, Volodymyr and Czarnecki, Wojciech M. and Schaul, Tom and Leibo, Joel Z. and Silver, David and Kavukcuoglu, Koray , title =. International Conference on Learning Representations , year =

  76. [76]

    International Conference on Learning Representations , year =

    Xie, Tianbao and Zhao, Siheng and Wu, Chen Henry and Liu, Yitao and Luo, Qian and Zhong, Victor and Yang, Yanchao and Yu, Tao , title =. International Conference on Learning Representations , year =

  77. [77]

    Robotics: Science and Systems , year =

    Ma, Yecheng Jason and Liang, William and Wang, Hung-Ju and Wang, Sam and Zhu, Yuke and Fan, Linxi and Bastani, Osbert and Jayaraman, Dinesh , title =. Robotics: Science and Systems , year =

  78. [78]

    International Conference on Learning Representations , year =

    Klissarov, Martin and D'Oro, Pierluca and Sodhani, Shagun and Raileanu, Roberta and Bacon, Pierre-Luc and Vincent, Pascal and Zhang, Amy and Henaff, Mikael , title =. International Conference on Learning Representations , year =

  79. [79]

    and D'Oro, Pierluca , title =

    Klissarov, Martin and Henaff, Mikael and Raileanu, Roberta and Sodhani, Shagun and Vincent, Pascal and Zhang, Amy and Bacon, Pierre-Luc and Precup, Doina and Machado, Marlos C. and D'Oro, Pierluca , title =. International Conference on Learning Representations , year =

  80. [80]

    and Coleman, Russell and Srinivasan, Ravi and Niekum, Scott , title =

    Brown, Daniel S. and Coleman, Russell and Srinivasan, Ravi and Niekum, Scott , title =. Proceedings of the 37th International Conference on Machine Learning , series =. 2020 , publisher =

Showing first 80 references.