{"id":"3a7f41b3-5370-40ff-ac63-0bc5552769d1","arxiv_id":"2507.07159","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":8,"one_line_summary":"VNA produces Sharpe-ratio-competitive portfolios on indices up to 2,008 assets, but the claimed speed advantage and universal finite-size scaling are not robustly supported.","lead":"This paper applies variational neural annealing, a physics-inspired method, to optimize stock portfolios with up to 2,008 assets, and compares its speed and quality with the commercial solver Mosek. It also reports a universal scaling law for annealing time, but the evidence for that law is weak.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The DFSS universal-scaling claim is not established: three overlapping indices are not three realizations of one ensemble, and a two-parameter spline collapse can overfit just three curves; an out-of-sample prediction test is needed.","rationale":"The reader and I identify the same load-bearing weakness: the DFSS universal-scaling result, a headline claim, is built on treating three overlapping real-world indices as three system sizes of one random ensemble, with only five trading months as disorder samples. The paper itself acknowledges the small number of disorder realizations and the relaxed constraint space, but it does not address the more fundamental issue that the indices are not random draws from a common distribution; their overlapping constituents, shared factor exposures, and different constraint regimes make the finite-size scaling comparison uncontrolled. With three curves and two free exponents plus a fitted spline, the data-collapse procedure can succeed even when no universal scaling exists. The bootstrap error bars in Appendix D are computed from within-index variability and therefore miss this cross-index bias. The claim of polynomial annealing time scaling and the extrapolation to larger future indices are thus not supported by the evidence as presented. I would not reject the entire VNA-for-portfolio-optimization program: the S&P 500 e_rel scaling in Fig. 2(a), the annealing-versus-classical-optimization ablation, and the Sharpe-ratio comparisons provide partial evidence that VNA can be useful. But the advertised universal finite-size collapse and the associated universal scaling exponents are not established, and the absence of released code or data prevents independent verification. A revised manuscript with a genuine out-of-sample prediction test and at least one independent system size could change this assessment; in its current form, REJECT is the appropriate verdict.","tokens_in":18580,"tokens_out":5214,"duration_ms":66339,"concrete_test":"Perform a leave-one-out prediction test for the DFSS collapse: fit (κ, μ) and the spline g using only two indices (e.g., S&P 500 and Russell 1000) and predict the Russell 3000 curve e_rel N^κ versus τ_a N^μ without refitting. If the observed Russell 3000 points fall outside the bootstrap confidence band of the prediction, or if the residuals show a systematic index-dependent bias, then the Fig. 8 collapse is an overfit artifact rather than universal scaling. As a stronger supplement, add an independent system size not used in the original fit, such as a randomly drawn 1,300-asset subset of the Russell 3000 universe, and require that the exponents and spline from the original three indices predict its curve within noise. Use at least 20 trading months rather than 5 for the disorder average.","verdict_should_be":"REJECT","load_bearing_attack":"The central DFSS claim (Section IV E, Eq. 15, Fig. 8) is not supported because the S&P 500, Russell 1000, and Russell 3000 are not three independent realizations of the same optimization ensemble. These universes are nested or heavily overlapping, so their return and covariance matrices share many assets and differ systematically in factor composition and constraint activity; treating five trading months as disorder realizations does not remove index-specific biases. With only three sizes, the two-parameter collapse e_rel = N^{-κ} g(τ_a N^μ) is flexible enough to make any three smooth decreasing curves appear to collapse, and the reported bootstrap errors reflect only within-index month-to-month noise, not the dominant error from using non-identical ensembles. The paper also performs the collapse on a reduced objective (Eq. 14) with only the normalization constraint, explicitly relaxing the constraint space in Section IV E, so even a valid collapse would describe a simpler problem than the full MINLP used in the Mosek benchmark. No out-of-sample or leave-one-index-out validation is provided, and no code or data artifacts are released. Because the universal polynomial-time-scaling statement is one of the two headline claims, this weakness is load-bearing for the paper's central message.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes Variational Neural Annealing (VNA) with recurrent neural networks as a classical solver for a mixed-integer nonlinear portfolio-optimization problem. The objective is written as an Ising-like Hamiltonian with penalty terms for normalization, transaction costs, volatility, and turnover. The authors report that VNA scales to more than 2,000 assets, compares favorably with Mosek in Sharpe ratio, and exhibits power-law decay of relative energy with annealing steps. They further perform a dynamical finite-size scaling analysis over the S&P 500, Russell 1000, and Russell 3000, reporting a data collapse with exponents mu = -0.58 +/- 0.08 and kappa = 0.16 +/- 0.09, and a polynomial annealing-time scaling. The paper concludes that VNA is a practical large-scale MINLP solver and that the scaling behavior is universal.","tokens_in":18905,"tokens_out":4446,"duration_ms":49880,"significance":"If the central claims were fully supported, the work would be a useful contribution connecting statistical-physics-inspired methods to practical large-scale portfolio optimization. The paper has genuine strengths: the Sharpe-ratio comparison against Mosek in Figures 6-7 is an external benchmark, the annealing-vs-direct-optimization ablation in Figures 3-4 is internally consistent, the use of real index data is appropriate for the application, and the authors are candid about limitations such as low feasibility rates in Appendix C. However, the headline universal-scaling claim is not established by the evidence presented: the three indices are not independent realizations of one optimization ensemble, the collapse uses only three system sizes, and the reduced Hamiltonian of Eq. (14) differs from the full MINLP benchmarked earlier. Because the abstract's two central claims are performance comparability and universal scaling, the scaling weakness is load-bearing.","major_comments":[{"comment":"The DFSS universal-scaling claim is not supported by the presented evidence. The three 'system sizes' are the S&P 500, Russell 1000, and Russell 3000, which are nested or heavily overlapping universes with different composition, factor structure, and constraint activity, not independent realizations of a single optimization ensemble. With only three system sizes, the two-parameter collapse e_rel = N^{-kappa} g(tau_a N^mu) is flexible enough to fit three smooth curves, and the bootstrap errors in Appendix D reflect only within-index month-to-month noise, not the dominant uncertainty from using non-identical ensembles. An out-of-sample test, such as leave-one-index-out prediction or the use of genuinely independent asset universes, is needed before universal behavior and polynomial annealing-time scaling can be claimed.","section":"Section IV E, Eq. (15), Fig. 8, Appendix D"},{"comment":"The relative energy e_rel is normalized by E_ref, which is the lowest objective value found by VNA itself at tau_a = 1000, or in Fig. 2(b) by the lowest value over N_units. Thus e_rel measures closeness to VNA's own best solution, not closeness to an external or true optimum. The exponents extracted in Fig. 8 may therefore characterize VNA's internal convergence rather than solution quality relative to the global optimum. For at least one system size, E_ref should be cross-checked against an independent solver such as Mosek on the same reduced objective, or the interpretation of the scaling exponents should be restricted accordingly.","section":"Section IV B, Eq. (13)"},{"comment":"The DFSS analysis is performed on a reduced Hamiltonian that includes only the normalization constraint, explicitly relaxing the full constraint set used in the Mosek benchmark. The abstract and conclusions use this analysis to support VNA's scalability on the full MINLP, but the connection is not established. Either the DFSS analysis should be repeated on the full constraint set, or the scalability claim should be explicitly limited to the reduced problem defined by Eq. (14).","section":"Section IV E, Eq. (14) vs Section IV D, Table I"},{"comment":"The time-to-solution comparison in Table I is not apples-to-apples. VNA times are fixed wall-clock runtimes on four NVIDIA A100 GPUs, whereas Mosek times are times to reach a relative MIP gap of 0.01%, or 1% for Russell 3000, on 64 CPU threads; the hardware, stopping rules, and post-hoc sample filtering differ. The text calls the comparison 'qualitative,' but the abstract states that VNA exhibits 'faster convergence on hard instances' without presenting controlled solution-quality-versus-time curves. Please provide per-method quality-versus-time comparisons or a same-budget comparison, or soften the claimed speed advantage.","section":"Section IV D, Table I"},{"comment":"The claim that VNA can identify near-optimal solutions for portfolios of more than 2,000 assets is qualified by the fact that valid solutions were not obtained for all trading months in Fig. 7, and Appendix C reports Russell 3000 feasibility as low as 0.129%. The fraction of months without valid solutions and the sensitivity to penalty coefficients should be stated in the main text, since they directly affect the headline practical claim.","section":"Section IV D and Appendix C"}],"minor_comments":[{"comment":"Several section headings contain unintended spaces, such as 'POR TFOLIO OPTIMIZA TION FORMULA TION', 'T ransaction Costs', and 'V ersus'; these should be corrected.","section":"Throughout"},{"comment":"The definition of E_ref differs between Fig. 2(a) and Fig. 2(b), but the text does not clearly distinguish the two definitions; please state the reference value used in each panel.","section":"Section IV B, Eq. (13)"},{"comment":"The statement that the annealing time scales as tau_a ~ N^{-mu} is confusing because mu is negative; writing the positive exponent tau_a ~ N^{0.58} would make the polynomial growth explicit.","section":"Section IV E, Eq. (15)"},{"comment":"The fit range for the power-law exponent -1.43(1) is described as 'over the last number of annealing steps'; please specify the exact fit range and the number of points used.","section":"Appendix A, Fig. 9"},{"comment":"The caption says each data point is averaged over five independent trading months, but it is not stated whether the same five months are used for all three indices; if they differ, the collapse is even less controlled and the caption should say so.","section":"Section IV E, Fig. 8 caption"}],"recommendation":"major_revision","confidential_remarks":"The universal-scaling claim is the main obstacle: with three overlapping indices and a two-parameter collapse, the current evidence does not support the abstract's statement of universal behavior and polynomial scaling. The Mosek Sharpe-ratio benchmark is the strongest part of the paper, and the DFSS section could be reframed as a case study with explicit caveats or strengthened with out-of-sample validation. I would not publish the scaling claim in its present form, but the requested fixes are within the scope of a major revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Fair and direct take: the paper's core contribution is a new application of variational neural annealing to a realistic MINLP portfolio problem, and the Sharpe-ratio comparison against Mosek is a genuine external benchmark. The ablation showing annealing beats direct optimization is internally consistent, and the authors are transparent about feasibility failures and the need to relax constraints in the scaling section. That is real, useful work.\n\nThe soft spots are in the two headline claims. The time-to-solution comparison (Table I) is not apples-to-apples: VNA numbers are fixed runtimes on four A100 GPUs, Mosek runs on CPUs with different stopping criteria, and for Russell 3000 the gap was loosened from 0.01% to 1%. The paper itself cautions to read the times qualitatively, but the abstract still says 'faster convergence on hard instances.' As presented, that claim is not supported.\n\nThe DFSS analysis (Sec. IV E, Fig. 8) is the weakest part. S&P 500, Russell 1000, and Russell 3000 are not three independent realizations of one ensemble—they are nested universes, so the extracted μ and κ are fit to three overlapping curves with five months of 'disorder.' With a two-parameter collapse plus a spline, three smooth curves will collapse fairly easily; the reported bootstrap errors only reflect within-index month-to-month noise. More importantly, the collapse uses the reduced objective (Eq. 14) with only the normalization constraint, a different problem from the one benchmarked against Mosek. There is no out-of-sample check, no leave-one-index-out prediction. So the 'universal behavior and polynomial annealing time scaling' statement, which is in the abstract, is not established. The authors do flag limited disorder averages, but that is not the main problem; the index-composition bias is.\n\nMinor: no code or data artifacts, and the penalty coefficients are many free parameters, so feasibility at scale remains a practical worry.\n\nWho is this for? People working on quantum-inspired or ML-based portfolio optimization will want the Mosek solution-quality comparison and the constraint formulation. People who work on scaling collapses will find the DFSS section a cautionary example. It deserves a serious referee rather than a desk reject, but the revision needs to reframe the scaling claim as speculative, make the timing comparison fair (or drop it), and ideally release the data or at least the fitted curves. I would accept it in principle at a venue that allows major revisions.","headline":"A legitimate new application of VNA to portfolio optimization, with a solid Mosek benchmark on solution quality; but the advertised speed advantage and universal-scaling claims don't survive close reading.","tokens_in":19394,"tokens_out":2376,"would_cite":true,"duration_ms":25479,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that variational neural annealing solves mixed-integer nonlinear portfolio optimization on universes of more than 2,000 assets, with solution quality comparable to a leading commercial solver and faster convergence on…","keywords":["portfolio optimization","variational neural annealing","mixed-integer nonlinear programming","recurrent neural networks","Ising model","finite-size scaling","Markowitz mean-variance","transaction costs"],"falsifier":"A decisive check would be to extend the same dynamical finite-size scaling to a fourth index or to synthetic portfolios of intermediate size with many more trading months; if the rescaled data do not collapse onto one curve, or if the fitted exponents drift with the chosen set of sizes, the claimed universality and $N^{0.58}$ annealing-time scaling are refuted. In parallel, letting Mosek run to its standard 0.01% gap on the Russell 3000 instance and comparing Sharpe ratios month by month would test the competitiveness claim directly.","tokens_in":18380,"feed_emoji":"📈","tokens_out":12437,"duration_ms":119072,"temperature":0.7,"pith_summary":"Portfolio rebalancing under real-world rules — transaction costs, turnover limits, and volatility caps — becomes a mixed-integer nonlinear program that commercial solvers struggle with beyond a few hundred assets. The paper claims this problem can be mapped onto a classical Ising-like Hamiltonian and solved with variational neural annealing (VNA), which trains an autoregressive recurrent neural network to sample from the Boltzmann distribution while the temperature is lowered. On real data from the S&P 500, Russell 1000, and Russell 3000 — up to 2,008 assets — VNA finds near-optimal portfolios with Sharpe ratios comparable to Mosek's and reaches them faster on hard instances. The paper also reports a dynamical finite-size scaling collapse of the residual energy across the three indices, with exponents implying that the annealing time needed for near-adiabatic performance grows only polynomially with the number of assets. If these results hold, VNA would be a practical large-scale solver for a routine but NP-hard financial operation and a new setting for universal nonequilibrium scaling.","feed_headline":"Neural annealing matches Mosek on 2,000-asset portfolios","feed_subtitle":"A physics-style annealing network matches a leading solver on real equity portfolios and shows one universal scaling law.","key_machinery":"The load-bearing object is the variational neural annealing algorithm. It uses a recurrent neural network as an autoregressive ansatz, $P_\\theta(x) = \\prod_i P_\\theta(x_i | x_{i-1}, \\ldots, x_1)$, and trains it at each temperature to minimize the variational free energy $F_\\theta = \\langle H\\rangle_\\theta + T \\sum_x P_\\theta(x) \\ln P_\\theta(x)$, so that $P_\\theta$ approaches the Boltzmann distribution $e^{-H/T}$. The temperature follows a geometric schedule from large $T$ to zero, and at the end the network generates uncorrelated candidate portfolios that are filtered for constraint satisfaction. The RNN gives $O(N)$ cost per gradient step and is what lets the method scale past 2,000 assets; for the finite-size scaling study the normalization constraint is enforced with both a penalty and a Lagrange-multiplier term.","core_discovery":"The central claim is that VNA, in its classical recurrent-neural-network form, solves the mixed-integer nonlinear portfolio optimization problem at real-world index scale. The authors encode the Markowitz objective with normalization, transaction-cost, turnover, and volatility constraints as a Hamiltonian $H(x)$ with penalty and Lagrange-multiplier terms, turning each asset's weight into a discrete decision variable. VNA minimizes the variational free energy $F_\\theta = \\langle H\\rangle_\\theta + T \\sum_x P_\\theta(x) \\ln P_\\theta(x)$ along a geometric annealing schedule, using an autoregressive RNN whose per-gradient-step cost is $O(N)$. On the S&P 500 the residual energy decays as a power law, $e_{\\mathrm{rel}} \\propto \\tau_a^{-1.43(1)}$; on the Russell 3000, VNA reaches Sharpe ratios close to Mosek's, with a reported time-to-solution of 2,580 seconds versus 6,709 seconds for Mosek (which had its gap tolerance relaxed to 1% because the standard 0.01% threshold was not reached in days). Treating the three indices as increasing system sizes and five trading months as disorder realizations, the authors obtain a data collapse of $e_{\\mathrm{rel}} N^{\\kappa}$ against $\\tau_a N^{\\mu}$ with $\\mu = -0.58 \\pm 0.08$ and $\\kappa = 0.16 \\pm 0.09$, which they interpret as universal critical dynamics and polynomial annealing-time scaling.","pith_inferences":["The Ising-style encoding is asset-agnostic, so the same pipeline should transfer to bond portfolios, multi-asset funds, or crypto baskets with discrete position sizes; the paper tests equities only.","The scaling collapse implies an out-of-sample forecast: a hypothetical 4,000-asset index should land on the same collapsed curve, so its residual energy could be predicted from the spline without solving the optimization; running that forecast would be a sharp test of universality.","The validity failures on Russell 3000 suggest that penalty tuning, not the neural sampler, is the main barrier to production use; an adaptive augmented-Lagrangian update scheme, which the authors list as future work, would be the natural remedy."],"forward_implications":["Portfolio optimization with transaction costs, turnover limits, and volatility caps can be treated as a generic Hamiltonian minimization, so the same VNA pipeline applies without requiring convexity or differentiability.","Near-optimal portfolios for universes of more than 2,000 assets are within reach of variational methods, and on hard instances VNA can reach comparable Sharpe ratios in less wall-clock time than a commercial solver.","The extracted exponents imply that the annealing steps needed for near-adiabatic performance grow roughly as $N^{0.58}$, a polynomial cost in the number of assets.","Physics-inspired diagnostics such as free-energy variance and Von Neumann entropy can serve as proxies for solution quality when the optimum is unknown.","Constraint feasibility, not energy minimization, becomes the practical bottleneck: on Russell 3000 the fraction of valid samples can fall to 0.129% for some months, so penalty and Lagrange-multiplier tuning is decisive."],"supporting_citations":[{"why":"Defines the Markowitz mean-variance objective that the paper's Hamiltonian encodes.","marker":"[11]"},{"why":"Supplies the Ising formulation of binary decision problems and the NP-hardness argument for the portfolio mapping.","marker":"[13]"},{"why":"Characterizes the mixed-integer nonlinear program class that makes large-scale portfolio optimization hard.","marker":"[14]"},{"why":"Earlier quantum-inspired tensor-network portfolio optimization at hundreds of assets that this work scales beyond.","marker":"[26]"},{"why":"Introduces the variational neural annealing algorithm and the variational free-energy training objective used throughout.","marker":"[57]"},{"why":"The commercial solver used as the benchmark for solution quality and time-to-solution.","marker":"[68]"},{"why":"Supplies the residual-energy scaling methodology and the Kibble-Zurek interpretation applied to the annealing-time dependence.","marker":"[75]"},{"why":"Provides the dynamical finite-size scaling ansatz that gives the collapse curve $e_{\\mathrm{rel}} N^{\\kappa} = g(\\tau_a N^{\\mu})$.","marker":"[79]"},{"why":"Provides the data-collapse fitting and parametric bootstrap used to estimate the exponents and their errors.","marker":"[86]"}],"fun_headline_variants":["Neural annealer matches Mosek on 2,000-asset portfolios","VNA rivals Mosek and shows universal scaling","Physics-inspired annealer solves 2,000-asset portfolios","Neural annealing hits Mosek performance, faster on hard cases","Portfolio optimization: neural annealer scales to 2,000 assets"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the S&P 500, Russell 1000, and Russell 3000 behave like three sizes of the same optimization problem, and that five trading months are enough trading periods to stand in for many random instances; if that premise fails, the universal scaling curve and its exponents are a fitting artifact.","fun_headline_variants_meta":{"raw":{"variants":["Neural annealer matches Mosek on 2,000-asset portfolios","VNA rivals Mosek and shows universal scaling","Physics-inspired annealer solves 2,000-asset portfolios","Neural annealing hits Mosek performance, faster on hard cases","Portfolio optimization: neural annealer scales to 2,000 assets"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000444,"raw_usage":{"total_tokens":2284,"prompt_tokens":1019,"completion_tokens":1265,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":635,"completion_tokens_details":{"reasoning_tokens":1174}},"tokens_in":635,"tokens_out":1265,"duration_ms":12903,"temperature":1.0,"reasoning_tokens":1174,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T18:48:27.025633+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A decisive check would be to extend the same dynamical finite-size scaling to a fourth index or to synthetic portfolios of intermediate size with many more trading months; if the rescaled data do not collapse onto one curve, or if the fitted exponents drift with the chosen set of sizes, the claimed universality and $N^{0.58}$ annealing-time scaling are refuted. In parallel, letting Mosek run to its standard 0.01% gap on the Russell 3000 instance and comparing Sharpe ratios month by month would test the competitiveness claim directly.","supporting_citations":[{"cited_title":"Portfolio selection,","cited_arxiv_id":null,"evidence_quote":"Defines the Markowitz mean-variance objective that the paper's Hamiltonian encodes."},{"cited_title":"Ising formulations of many NP prob- lems,","cited_arxiv_id":null,"evidence_quote":"Supplies the Ising formulation of binary decision problems and the NP-hardness argument for the portfolio mapping."},{"cited_title":"Mixed- integer nonlinear optimization,","cited_arxiv_id":null,"evidence_quote":"Characterizes the mixed-integer nonlinear program class that makes large-scale portfolio optimization hard."},{"cited_title":"Dynamic portfolio optimiza- tion with real datasets using quantum processors and quantum-inspired tensor networks,","cited_arxiv_id":null,"evidence_quote":"Earlier quantum-inspired tensor-network portfolio optimization at hundreds of assets that this work scales beyond."},{"cited_title":"End- to-end risk budgeting portfolio optimization with neural networks,","cited_arxiv_id":null,"evidence_quote":"Introduces the variational neural annealing algorithm and the variational free-energy training objective used throughout."},{"cited_title":"Gurobi optimizer,","cited_arxiv_id":null,"evidence_quote":"Provides the dynamical finite-size scaling ansatz that gives the collapse curve $e_{\\mathrm{rel}} N^{\\kappa} = g(\\tau_a N^{\\mu})$."}],"review_version":1}