{"id":"d7f0996c-480d-448c-8986-df9613619405","arxiv_id":"2607.16863","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Treasure Search Optimization — explorers that keep searching plus a hunter that only teleports to improvements — has a well-posed conditional McKean-Vlasov limit and an O(1/α) stationary guarantee near the global minimizer.","lead":"Treasure Search Optimization is a new swarm-style global optimizer that separates exploration (a cloud of 'explorer' particles) from exploitation (a single 'treasure hunter' that only jumps to genuinely better consensus points). The paper proves the mean-field equations are well-posed and that, in a variance-matched regime, the hunter's steady state lands within O(1/α) of the global minimum — a read worth having if you build derivative-free optimizers for inverse problems.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Theorem 3.1 proves existence of an equilibrium near xmin, not convergence to it; the abstract's 'settles' is unsupported.","rationale":"The Reader's verdict (CONDITIONAL) is appropriate, but the most load-bearing concern is not the variance-matching mismatch per se — it is that Theorem 3.1 is an existence result for a stationary state, not a convergence result. The abstract and conclusions overstate the theorem by using 'settles'. The paper itself admits convergence to equilibrium is open. The variance-matching mismatch is a related but distinct issue: it means the numerics do not even satisfy the theorem's hypotheses. Both concerns justify a conditional verdict, but the missing convergence proof is the deeper threat to the central claim because it affects the theoretical regime as well. A numerical long-time simulation in the variance-matched regime would directly test whether the dynamic system approaches the predicted equilibrium; if it does not, the central claim fails. I therefore keep the Reader's CONDITIONAL verdict unchanged.","tokens_in":37034,"tokens_out":9659,"duration_ms":82442,"concrete_test":"Simulate the particle TSO system (or a discretized mean-field version) in the variance-matched regime: choose d=2, f as a smooth double-well with known unique global minimizer xmin, and set σ, η (with η2=0), σJ such that σ²/(2η)=σJ². Run for long horizons T = 100, 200, 500, 1000 with α large, and record the hunter's distance to xmin over time. If the distance does not decay to O(1/α) and stay there, the 'settles' claim is unsupported. Also compare the explorer empirical covariance to the predicted σJ² I_d; lack of convergence to this Gaussian equilibrium would show the dynamic system does not select the proven steady state.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim as advertised — 'the hunter settles near the global minimum with error of order 1/α' (abstract) — is stronger than what Theorem 3.1 establishes. The theorem shows that, under the variance-matching condition (3.1) and Assumption 1, there exists a self-consistent steady-state consensus point mα satisfying mα = Tα(mα) with |mα − xmin| ≤ C/α. It does not show that the time-dependent conditional McKean-Vlasov dynamics (2.10)-(2.11) converge to this stationary state, nor that the particle system approaches it as t → ∞. The paper explicitly acknowledges this gap in Section 7: 'A quantitative conditional propagation of chaos result ... together with convergence to the equilibrium, is the subject of ongoing work.' Thus the value proposition — that TSO provably locates the global minimum — rests on an existence result for an equilibrium, not on a guarantee that the algorithm reaches it. This concern is independent of the variance-matching mismatch flagged in the Reader's verdict: even in the variance-matched regime, no convergence to equilibrium is proven. Without a Lyapunov or ergodicity result, the dynamics could in principle never approach the stationary state, making the O(1/α) guarantee vacuous for actual TSO runs.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces Treasure Search Optimization (TSO), a two-species interacting particle method for global optimization. A swarm of explorers performs exploration via jump-diffusions, while a single treasure hunter performs exploitation by drifting toward an objective-weighted consensus and teleporting when this improves the objective. The mean-field limit is formulated as a conditional McKean–Vlasov jump-diffusion SDE with common noise. The paper proves well-posedness for a smoothed version of the teleportation rule (Theorem 2.1), characterizes a Gaussian stationary state under a variance-matching condition (Section 3.1), and proves existence of a self-consistent steady-state consensus point within O(1/α) of the global minimizer (Theorem 3.1). It also gives a formal free-energy gradient interpretation of the consensus drift, proposes a post-processing Kalman calibration for uncertainty quantification, and reports numerical experiments on ODE-constrained problems and a Bayesian inverse problem.","tokens_in":37297,"tokens_out":9011,"duration_ms":87609,"significance":"If the main theorem and its assumptions are taken at face value, the paper contributes a derivative-free swarm optimizer with a provable proximity guarantee for an equilibrium of the mean-field dynamics, and it develops a quantitative self-consistent Laplace method that is more subtle than a direct application of the Laplace principle. The well-posedness result for conditional McKean–Vlasov jump-diffusions with common noise, built via pathwise Leray–Schauder and measurable selection, is also of independent interest. However, the advertised headline — that 'the hunter settles near the global minimum with error of order 1/α' — is not established: the theorem proves existence of a fixed point of the stationary-map Tα, not convergence of the dynamics to that fixed point. Moreover, the numerical experiments run in a parameter regime where the variance-matching condition (3.1) is violated, so Theorem 3.1 does not apply to the demonstrated algorithm. These gaps substantially weaken the contribution as presented, though they are potentially repairable within the manuscript's scope.","major_comments":[{"comment":"The abstract and concluding remarks claim the hunter 'settles near the global minimum with error of order 1/α'. Theorem 3.1 only proves the existence of a self-consistent fixed point mα = Tα(mα) with |mα − xmin| ≤ C/α. It does not show that the time-dependent conditional McKean–Vlasov dynamics (2.10)–(2.11), or the finite-particle system, converge to this equilibrium. Section 7 explicitly states that 'convergence to the equilibrium, is the subject of ongoing work.' Thus the central value proposition — a provable guarantee that TSO locates the global minimum — is not supported by the theorem. The claim should be reframed as an equilibrium existence result, or a Lyapunov/ergodicity analysis should be added.","section":"Abstract; Section 7; Theorem 3.1 (Eqs. (3.14)–(3.16))"},{"comment":"Theorem 3.1 depends crucially on the variance-matching condition σ²/(2(η1+η2)) = σJ² (Eq. (3.1)), which makes the stationary explorer law Gaussian with variance σJ². The numerical demonstrations do not satisfy this condition. In §6.2.3, η=1, σ=0.25, so σ²/(2η)=0.03125, while Eq. (6.15) gives σJ² ≈ 4.55. In §6.3, σ=0.8, σJ=2, so σ²/(2η)=0.32 but σJ²=4. In both cases the steady state is non-Gaussian and governed by the covariance formula (3.12), for which no O(1/α) proximity theorem is proven. The theory and the numerical evidence therefore concern two different regimes. Either the experiments should be rerun in the variance-matched regime, or the theory should be extended to the unmatched case, or the mismatch should be explicitly acknowledged.","section":"Section 3.1 and Section 6 (numerical experiments)"},{"comment":"Theorem 2.1 establishes well-posedness only for the smoothed jump-size Gε(m,y) = (m−y)Ψε(m,y) with Ψε satisfying (2.22). The algorithm actually implemented uses the hard teleportation rule (6.3), where the indicator 1{f(mk+1)<f(Ŷk+1)} is discontinuous. No existence or uniqueness result is provided for this hard-indicator mean-field SDE. Since contribution (i) claims well-posedness of the TSO system, the current result covers only an auxiliary smoothed variant. Please clarify whether the well-posedness extends to the hard rule, or restrict the claim accordingly.","section":"Theorem 2.1 and Section 6.1 (implementation)"}],"minor_comments":[{"comment":"The horizontal axis labels appear as '4 3 2 1 0 1 2 3 4' with no minus signs; likely a rendering issue. Please verify the final PDF displays negative values correctly.","section":"Figure 1"},{"comment":"The formula for σJ contains a square root that may be negative for some parameter choices; the text does not discuss feasibility constraints. A brief comment on when the matching condition admits a real solution would improve the reproducibility.","section":"Eq. (6.15)"},{"comment":"The Laplace calibration map TLap in Eq. (5.7) assumes ΣLap^{1/2} commutes implicitly with the identity; since both are symmetric and share eigenvectors, this is fine, but the notation (σ⋆²Id)^{−1/2} could be simplified. Also, the caption says σ⋆=1 while the text says the steady-state covariance is σ⋆²I; please keep the notation consistent.","section":"Section 5 and Figure 2"},{"comment":"A few typos and minor notational inconsistencies remain (e.g., 'Mα(E_N_t, Y_t)' vs. 'Mk' in Section 6.1; 'λJ' vs. 'λ_Y' in some places). A careful proofreading pass is recommended.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The core fixed-point theorem appears sound under the stated assumptions, but the manuscript's advertising overreaches in two ways: the dynamics-to-equilibrium gap is acknowledged in Section 7, and the numerics are run outside the variance-matching regime. Neither issue is necessarily fatal, but both are load-bearing and require substantive revision. The well-posedness result for the smoothed system is a genuine contribution, though the gap to the hard rule should be addressed. I would not recommend rejection because the mathematical core is defensible and the method shows empirical promise; the revision should align claims with results and either adjust experiments or theory."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth a look, but read it with the claims and the theorems separated. The architecture is genuinely new: a single hunter with Poisson teleport shared as common noise by an explorer swarm is not in the cited CBO/CBS/EKI/jump-CBO literature, and the conditional McKean-Vlasov jump-diffusion object is a real extension. Credit where it is due: Theorem 2.1 is honestly scoped to the smoothed indicator, and the pathwise Leray–Schauder plus measurable-selection construction looks sound. I traced Theorem 3.1 as well; the inward-pointing estimate, Brouwer step, and Laplace bounds are internally consistent. That is real work, and the paper deserves a serious referee.\n\nNow the soft spots. The reader and the stress-test note are both right, and these are not manufactured flaws. First, the abstract says the hunter 'settles near the global minimum with error of order 1/α'. What is proven is the existence of a self-consistent steady-state consensus point satisfying |mα − xmin| ≤ C/α. There is no convergence-to-equilibrium theorem, no Lyapunov function, no ergodicity result for the time-dependent dynamics. The paper itself acknowledges this in Section 7. So the advertised value proposition is materially stronger than the proven statement. That has to be fixed by proof or by reframing.\n\nSecond, the variance-matching condition (3.1) is the hinge of both headline theorems, and every reported numerical configuration runs far outside it. Section 6.2.3 gives σ²/(2η) = 0.03125 against σJ² ≈ 4.55; Section 6.3 gives 0.32 against σJ² = 4. When the condition fails, no Gaussian steady state and no O(1/α) statement is proven. So the theory and the demonstrations are about two different regimes. That is a scoping/reporting issue, not a detected error in the main proofs, but it is a serious one.\n\nSmaller issues: the claim that the O(1/α) rate depends on self-consistency is inaccurate — the Laplace bound is uniform over centers, so the rate is essentially direct. The numerical section has no error bars, no code, and no ablation against the author's own jump-CBO (KST23), which is the most natural baseline. Also, the well-posedness theorem only covers the smoothed indicator, while the implemented algorithm uses the hard teleport; the gap should be stated explicitly.\n\nBottom line: send it to peer review. It is a real contribution with a new object, a valid theorem (for the smoothed system), and promising numerics. But the authors need to either extend the theory to the operating regime, test explicitly in the variance-matched regime, or reframe the claims; add a convergence-to-equilibrium statement or clearly say it is open; and ship code and error bars.","headline":"Genuinely new two-agent swarm architecture with a real mean-field theorem, but the abstract overclaims: Theorem 3.1 proves an equilibrium exists near xmin, not that the dynamics converge to it, and the numerics run outside the variance-matched regime the theory assumes.","tokens_in":37910,"tokens_out":1438,"would_cite":false,"duration_ms":15397,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["90C26","60H10","65K10","65C30","65C35","60J76"],"pacs":[],"model":"deepseek-v4-flash","headline":"Treasure Search Optimization splits a swarm into explorers and a hunter and proves the hunter's steady state reaches the global minimum within O(1/α), under a variance-matching condition.","keywords":["Treasure Search Optimization","interacting particle systems","global optimization","conditional McKean-Vlasov SDE","common noise","jump-diffusion","Laplace approximation","uncertainty quantification"],"falsifier":"Run the TSO particle system on a one-dimensional double-well objective f with parameters satisfying the variance-matching condition, and separately with the same parameters except σJ changed so the condition fails by a factor of ten; for several values of α, measure the long-run mean of |hunter − xmin|. If the O(1/α) decay appears only in the matching case, the theorem's domain is exactly as stated; if it appears in both, the theory is narrower than the method; if it appears in neither after finite-N and discretization effects are controlled, the fixed-point Laplace mechanism is not the operat","tokens_in":36732,"feed_emoji":"🎯","tokens_out":6804,"duration_ms":62630,"temperature":0.7,"pith_summary":"The paper tries to establish that a derivative-free, two-agent swarm method—a diffuse cloud of explorers plus a single treasure hunter—has a provable steady-state guarantee: as the weighting parameter α grows, the hunter's position lands within O(1/α) of the global minimizer. This matters because most swarm methods balance exploration and exploitation inside one population, usually by annealing or degenerating noise, which forces a trade-off between premature collapse and slow escape from local minima. TSO removes that trade-off by design and supplies a theorem rather than only a heuristic. The proof passes through a conditional McKean-Vlasov mean-field limit with common Poissonian noise, an explicit Gaussian steady state under a variance-matching condition, and a self-consistent Laplace evaluation of the consensus map. A secondary contribution links the swarm's drift to a smoothed free energy, explaining why the swarm ignores small spurious traps and enabling a post-processing Kalman step for Bayesian uncertainty quantification.","feed_headline":"Swarm with a single hunter lands within 1/α of global minimum","feed_subtitle":"A theorem shows the hunter-explorer feedback loop reaches the global minimizer at rate 1/α, with no gradients needed.","key_machinery":"The load-bearing object is the self-consistent consensus map Tα(m) = (∫ x e^{−αf(x)} e^{−c0|x−m|²} dx) / (∫ e^{−αf(x)} e^{−c0|x−m|²} dx), with c0 = 1/(2σ⋆²). It encodes the explorers' Gibbs-weighted average once the swarm is treated as a Gaussian cloud centered at m with variance σ⋆². The variance-matching condition σ²/(2(η1+η2)) = σJ² makes that Gaussian ansatz an actual stationary solution of the jump-diffusion mean-field equation; when the condition fails, the steady-state covariance is no longer Gaussian and is instead governed by the non-Gaussian formula (3.12). The proof pairs the inward-pointing estimate (3.23) with Brouwer's fixed-point theorem to obtain mα = Tα(mα), then uses Laplac","core_discovery":"The central claim is Theorem 3.1: under the variance-matching condition σ²/(2(η1+η2)) = σJ² and Assumption 1 (f is C⁴ near a unique global minimizer with positive-definite Hessian and quadratic growth), the mean-field TSO system has a stationary regime in which the explorer law is Gaussian N(m⋆, σ⋆²I), the hunter sits at m⋆, and any self-consistent consensus point mα solving mα = Tα(mα) satisfies |mα − xmin| ≤ C/α for all sufficiently large α. This is the result that turns TSO from a plausible algorithm into one with a proximity guarantee: no gradient information and no annealing schedule, only a tunable weight. The proof is a fixed-point Laplace argument: an inward-pointing estimate gives a","pith_inferences":["The numerical sections choose parameters that violate the variance-matching condition, so on the paper's own definitions those experiments do not directly test Theorem 3.1; a direct numerical check of the C/α scaling under condition (3.1) would complete the loop.","The self-consistent Laplace argument suggests the 1/α rate is controlled by the Gaussian smoothing scale σ⋆; varying σ⋆ with α rather than fixing it might yield a different, possibly dimension-dependent, rate, though the paper does not claim this.","The hunter's monotone-improvement jumps give it a global memory, so the full system behaves like gradient flow on a time-dependent smoothed potential; replacing the deterministic accept rule with a probabilistic one is a testable variant.","The paper lists quantitative conditional propagation of chaos as future work; a finite-N convergence bound would translate the mean-field O(1/α) guarantee into a statement about the actual particle system, which the theorems so far do not provide."],"forward_implications":["If Theorem 3.1 is correct, TSO offers a derivative-free optimization method with a tunable steady-state proximity guarantee to the global minimum, without gradient evaluations or noise annealing.","The well-posedness result for the smoothed conditional McKean-Vlasov jump-diffusion mean-field limit gives the finite-particle algorithm a firm mathematical base.","The free-energy and Stein-kernel analysis indicates that the swarm's macroscopic center descends a smoothed landscape, explaining why the collective cloud can be insensitive to small spurious local traps.","In inverse problems, the equilibrium explorer cloud has an explicit Gaussian form with known algorithmic covariance, so an affine or Kalman post-processing step can convert it into a geometry-aware uncertainty estimate while the optimization itself remains derivative-free.","The common-noise conditional McKean-Vlasov framework with finite-activity Poisson jumps is a new setting in which existence and uniqueness are obtained by a pathwise fixed-point plus measurable-selection argument."],"fun_headline_variants":["Single hunter, explorer swarm: global minimum at 1/α rate","Treasure Search Optimization: one hunter beats local traps","Swarm with a teleporting hunter: proof of convergence to global min","TSO: no gradients, no annealing, just a hunter and explorers","How a hunter and explorers crack global optimization"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"Everything rests on the variance-matching equality σ²/(2(η1+η2)) = σJ², which forces the explorer cloud's steady state to be Gaussian; the paper itself calls the condition restrictive, and its numerical demonstrations run far outside this regime, so if that equality is not a natural operating point, the proven O(1/α) guarantee and the tested algorithm are not about the same parameter settings.","fun_headline_variants_meta":{"raw":{"variants":["Single hunter, explorer swarm: global minimum at 1/α rate","Treasure Search Optimization: one hunter beats local traps","Swarm with a teleporting hunter: proof of convergence to global min","TSO: no gradients, no annealing, just a hunter and explorers","How a hunter and explorers crack global optimization"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000884,"raw_usage":{"total_tokens":3696,"prompt_tokens":824,"completion_tokens":2872,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":568,"completion_tokens_details":{"reasoning_tokens":2785}},"tokens_in":568,"tokens_out":2872,"duration_ms":19138,"temperature":1.0,"reasoning_tokens":2785,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T19:46:52.762047+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the TSO particle system on a one-dimensional double-well objective f with parameters satisfying the variance-matching condition, and separately with the same parameters except σJ changed so the condition fails by a factor of ten; for several values of α, measure the long-run mean of |hunter − xmin|. If the O(1/α) decay appears only in the matching case, the theorem's domain is exactly as stated; if it appears in both, the theory is narrower than the method; if it appears in neither after finite-N and discretization effects are controlled, the fixed-point Laplace mechanism is not the operat","supporting_citations":[],"review_version":1}