{"id":"70a17faa-a05b-46b1-8a6f-186b625af569","arxiv_id":"2606.20903","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Reformulates risk-sensitive benchmarked asset allocation as an LQG stochastic differential game via free energy-entropy duality and develops a continuous-time q-learning actor-critic algorithm that learns optimal policies with high accuracy in a proof-of-concept.","lead":"The paper reformulates a continuous-time risk-sensitive asset allocation problem using free energy-entropy duality to turn it into a linear-quadratic-Gaussian stochastic game, then builds a q-learning actor-critic method whose actors learn portfolio and adversarial controls. A generalist reader might care because the approach gives an economic interpretation of learned policies and shows practical learning on equity data.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"Validity of exact free energy-entropy duality reformulation for benchmarked problem with uncontrolled state but controlled Itô term in reward","rationale":"The reader's weakest_assumption directly identifies the reformulation step as load-bearing, which matches the structure of the argument. The proof-of-concept calibration to U.S. equity data and accuracy claim presuppose that the explicit saddle-point solutions are available from this step. No machine-checked proof or independent verification is mentioned, so the concern remains the central one even after abstract-only review.","tokens_in":1727,"tokens_out":334,"duration_ms":10817,"concrete_test":"Starting from the problem statement in §2 (state SDE and terminal reward with controlled integral), re-derive the free energy-entropy duality transformation step-by-step and check whether the resulting dynamics and objective are exactly those of an LQG game under an equivalent measure, with no leftover terms depending on the original control.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The strongest claim requires that the benchmarked allocation problem (uncontrolled state dynamics, terminal reward containing a controlled Itô integral) admits an exact free energy-entropy duality mapping to an equivalent LQG stochastic differential game. This mapping is asserted to produce explicit finite- and infinite-horizon saddle-point solutions that directly motivate the quadratic critic and affine deterministic actors in the continuous-time q-learning method. If the duality step introduces approximation, residual control dependence, or fails to preserve equivalence under the changed measure, the explicit solutions, the actor-critic architecture, and the subsequent claim that actors recover the optimal policy with high accuracy would not hold.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper develops a reinforcement-learning method for continuous-time risk-sensitive benchmarked asset allocation. It uses free energy-entropy duality to recast the problem (uncontrolled state dynamics, terminal reward with controlled Itô integral) as an equivalent linear-quadratic-Gaussian stochastic differential game under a changed measure. Explicit finite- and infinite-horizon saddle-point solutions are derived; these motivate a continuous-time q-learning actor-critic scheme whose critic is quadratic and whose actors are deterministic affine maps for portfolio and adversarial controls. The learned allocation is interpreted via fractional Kelly decompositions. A proof-of-concept calibration to U.S. equity data reports that the actors recover the optimal policy with high accuracy.","tokens_in":1895,"tokens_out":611,"duration_ms":17785,"significance":"If the duality mapping is exact and the equivalence is preserved, the work supplies a mathematically grounded route from a non-standard risk-sensitive control problem to an actor-critic architecture whose functional forms are dictated by the saddle-point solution rather than chosen heuristically. The explicit solutions, the economic Kelly interpretation, and the reproducible numerical demonstration on real data constitute clear strengths. The approach could influence the design of model-based RL methods for other finance problems whose state-reward structure deviates from the classical Markovian template.","major_comments":[{"comment":"The central claim rests on an exact free energy-entropy duality that converts the benchmarked problem into an LQG game. The abstract states that the benchmarked problem 'does not directly fit the standard Markovian stochastic-control template' yet 'admits' the reformulation; the manuscript must supply the explicit change-of-measure construction and verify that the controlled Itô term in the terminal reward remains equivalent (i.e., does not introduce residual control dependence) after the duality is applied. Without this step-by-step verification, the subsequent derivation of the quadratic critic and affine actors cannot be regarded as load-bearing.","section":"Abstract and duality-reformulation section"},{"comment":"The numerical claim that 'the actors learn the optimal policy with high accuracy' is presented without reported error bounds, sensitivity to the discretization of the continuous-time q-learning update, or ablation on the data-exclusion window used for calibration. Because the architecture is justified by the duality-derived saddle point, any discrepancy between the learned policy and the analytic saddle point must be quantified to support the accuracy statement.","section":"Numerical implementation and results section"}],"minor_comments":[{"comment":"Notation for the free-energy functional and the entropy term should be introduced with explicit definitions before the duality statement is invoked.","section":"Introduction"},{"comment":"The paper should state whether the continuous-time q-learning algorithm is derived from first principles or obtained by taking a formal limit of a discrete-time counterpart; the derivation path affects reproducibility.","section":"Actor-critic method"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive and detailed comments. We will undertake a major revision to address the two points raised, adding the requested explicit derivations and quantitative numerical assessments.","responses":[{"response":"We agree that the change-of-measure construction and verification of Itô-term equivalence require explicit, step-by-step presentation. The submitted manuscript sketches the duality but does not contain the full Girsanov construction or the direct check that the controlled integral remains equivalent under the new measure. In the revision we will insert a dedicated subsection that (i) states the Girsanov kernel, (ii) derives the Radon–Nikodym derivative, and (iii) verifies that the terminal reward’s controlled Itô term acquires no additional control dependence after the measure change. This will make the subsequent saddle-point derivations load-bearing.","revision_made":"yes","referee_comment":"[Abstract and duality-reformulation section] The central claim rests on an exact free energy-entropy duality that converts the benchmarked problem into an LQG game. The abstract states that the benchmarked problem 'does not directly fit the standard Markovian stochastic-control template' yet 'admits' the reformulation; the manuscript must supply the explicit change-of-measure construction and verify that the controlled Itô term in the terminal reward remains equivalent (i.e., does not introduce residual control dependence) after the duality is applied. Without this step-by-step verification, the subsequent derivation of the quadratic critic and affine actors cannot be regarded as load-bearing."},{"response":"We accept that the accuracy claim must be supported by quantitative diagnostics. The original proof-of-concept omitted error metrics, discretization sensitivity, and ablation studies. In the revised numerical section we will report (i) L² and sup-norm discrepancies between the learned and analytic saddle-point controls, (ii) results for a range of discretization step sizes in the continuous-time q-learning update, and (iii) an ablation over the data-exclusion window used for calibration. These additions will quantify any discrepancy and substantiate the accuracy statement.","revision_made":"yes","referee_comment":"[Numerical implementation and results section] The numerical claim that 'the actors learn the optimal policy with high accuracy' is presented without reported error bounds, sensitivity to the discretization of the continuous-time q-learning update, or ablation on the data-exclusion window used for calibration. Because the architecture is justified by the duality-derived saddle point, any discrepancy between the learned policy and the analytic saddle point must be quantified to support the accuracy statement."}],"tokens_in":1474,"tokens_out":539,"duration_ms":25555,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The core move is the duality reformulation that turns the awkward benchmarked problem (uncontrolled state, controlled Ito term in the reward) into an equivalent LQG stochastic differential game. From there they extract explicit finite- and infinite-horizon saddle-point solutions, use the quadratic value function for the critic, and let the affine controls shape deterministic actors for allocation and the adversarial process. The fractional-Kelly reading of the learned policy is a useful byproduct, and the proof-of-concept on U.S. equity data indicates the portfolio actor receives a cleaner signal than the auxiliary actor.\n\nThe duality step carries the weight. If the mapping is exact and preserves equivalence under the changed measure with no residual control dependence, the explicit solutions and the actor-critic architecture follow directly. The abstract states that it does, and the implementation claims high accuracy, but any gap in the derivation would propagate to the learning claims. The paper does not appear to rely on post-hoc fitting that would make the central result circular.\n\nThis is for people working on continuous-time stochastic control or RL methods inside quantitative portfolio management. It is a targeted technical fix rather than a general framework, but the structure is clean enough that a referee familiar with LQG games and continuous-time RL could check the duality and the calibration details in one pass. I would send it to peer review.","headline":"The paper maps a benchmarked risk-sensitive allocation problem to an LQG game via free energy-entropy duality so the resulting saddle-point controls can directly motivate an actor-critic method, and the equity-data run shows the portfolio actor learning cleanly.","tokens_in":2397,"tokens_out":362,"would_cite":false,"duration_ms":9518,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Free energy-entropy duality reformulates risk-sensitive benchmarked asset allocation as an explicit linear-quadratic-Gaussian game that a continuous-time actor-critic method solves.","keywords":["reinforcement learning","risk-sensitive control","free energy-entropy duality","benchmarked allocation","actor-critic","linear-quadratic-Gaussian","continuous-time control","portfolio management"],"falsifier":"Run the actor-critic algorithm on the same calibrated U.S. equity parameters and compare the learned controls against the known closed-form saddle-point solution of the equivalent LQG game; large persistent deviation would falsify the claim that the method learns the optimal policy with high accuracy.","tokens_in":2623,"feed_emoji":"","tokens_out":773,"duration_ms":14024,"temperature":0.7,"pith_summary":"The paper establishes a reinforcement-learning method for continuous-time risk-sensitive benchmarked asset allocation in a partly model-based setting. The original problem does not fit standard Markovian control because the state is uncontrolled while the terminal reward includes a controlled Itô integral. Free energy-entropy duality converts the problem into an equivalent linear-quadratic-Gaussian stochastic differential game under a changed measure, which admits explicit finite- and infinite-horizon saddle-point solutions. These solutions directly motivate the structure of a q-learning actor-critic algorithm whose critic uses a quadratic value function and whose actors use affine controls for portfolio and adversarial decisions. Numerical calibration to U.S. equity data shows the learned policies recover the optimal allocation with high accuracy.","feed_headline":"Duality converts benchmarked allocation into explicit LQG game","feed_subtitle":"Reformulation supplies closed-form solutions that train a q-learning actor-critic method to high accuracy on U.S. equity data.","key_machinery":"Free energy-entropy duality, which converts the original benchmarked control problem into an equivalent linear-quadratic-Gaussian stochastic differential game under a changed probability measure and supplies the explicit saddle-point solutions that shape the actor-critic architecture.","core_discovery":"Free energy-entropy duality reformulates the benchmarked risk-sensitive allocation problem as a linear-quadratic-Gaussian stochastic differential game under an equivalent probability measure. This game possesses explicit saddle-point solutions for both finite and infinite horizons. The resulting quadratic value function and affine controls then determine the architecture of a continuous-time q-learning actor-critic method, with the portfolio allocation and adversarial control treated as deterministic actors. Implementation on U.S. equity data confirms that the actors recover the optimal policy to high accuracy and exhibits an asymmetry in which the portfolio actor receives a cleaner learning","pith_inferences":["The same duality step could be tested on other risk-sensitive problems whose state process is exogenous but whose payoff includes controlled stochastic integrals.","The observed asymmetry in learning signals suggests that training schedules could allocate more updates to the main allocation actor than to the adversary.","Fractional Kelly interpretations of the learned policy could be checked against standard mean-variance or log-optimal benchmarks on the same data set."],"forward_implications":["Explicit finite- and infinite-horizon saddle-point solutions exist for the reformulated LQG game.","The quadratic value function motivates the critic and the affine controls motivate deterministic actors in the q-learning algorithm.","The learned allocation policy admits an economic interpretation through fractional Kelly decompositions.","The portfolio actor receives a cleaner learning signal than the auxiliary adversarial actor."],"fun_headline_variants":["Duality recasts benchmarked allocation as LQG stochastic game","Explicit LQG saddle points train q-learning for asset allocation","Free energy duality supplies quadratic value for actor-critic method","Actor learns allocation from affine controls via duality reformulation"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The benchmarked allocation problem, whose state is uncontrolled while the terminal reward contains a controlled Itô integral, admits an exact free energy-entropy duality reformulation into an equivalent LQG game.","fun_headline_variants_meta":{"raw":{"variants":["Duality recasts benchmarked allocation as LQG stochastic game","Explicit LQG saddle points train q-learning for asset allocation","Free energy duality supplies quadratic value for actor-critic method","Actor learns allocation from affine controls via duality reformulation"]},"model":"grok-4.3","cost_usd":0.004577,"raw_usage":{"total_tokens":2269,"prompt_tokens":661,"num_sources_used":0,"completion_tokens":65,"cost_in_usd_ticks":45774500,"prompt_tokens_details":{"text_tokens":661,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1543,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":661,"tokens_out":65,"duration_ms":11440,"temperature":1.0,"reasoning_tokens":1543,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-26T14:24:23.781847+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Run the actor-critic algorithm on the same calibrated U.S. equity parameters and compare the learned controls against the known closed-form saddle-point solution of the equivalent LQG game; large persistent deviation would falsify the claim that the method learns the optimal policy with high accuracy.","supporting_citations":[],"review_version":1}