{"id":"812e3183-67b2-4a78-8a35-1c82e1dc469c","arxiv_id":"2506.05799","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":2.0,"correctness_risk":"high","formal_verification":"none","parameter_count":4,"one_line_summary":"On CSI 300 index options, gradient boosting ensembles (LGBM, XGBoost, NGBoost) achieve the lowest RMSE in most experiments, but the training set includes data from after the test period, invalidating the temporal realism claim.","lead":"An ensemble learning benchmark for CSI 300 index option pricing finds that LGBM, XGBoost, and NGBoost usually beat neural network and Black-Scholes baselines in raw pricing error. The paper's headline is tempered by a look-ahead bias in the data split that undercuts the claim of realistic financial simulation.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central model comparison is compromised by training on 2021 data to evaluate 2020: reported rankings may reflect temporal leakage, not model quality.","rationale":"The reader's weakest assumption correctly identifies the most load-bearing flaw: the training set includes September–December 2021 data while the test set is September–December 2020. This temporal leakage undermines the empirical basis for the central claim because all model rankings, score rates, and noise-robustness conclusions are computed under a non-causal data split. The paper's justification—that future data supplies 'denoised information'—is not a methodological argument; no denoising operation is defined, and the term merely relabels the leakage. The absence of error bars and statistical tests further weakens the word 'consistently,' but the leakage alone is sufficient to invalidate the reported comparison as evidence of real-world predictive performance. A causal re-run would settle the issue; if the ensemble advantage persists, the core finding might survive in revised form, but as presented the evidence does not support the claim.","tokens_in":10629,"tokens_out":4519,"duration_ms":42346,"concrete_test":"Retrain all nine models with the same hyperparameter settings (Table 1: In1 for the input/sliding-window experiments, ALL for the moneyness/noise experiments) using only January–August 2020 data, evaluate on September–December 2020, and recompute Tables 2, 4, and the score rates. If LGBM, NGBoost, and XGBoost are no longer the top three score-rate models, or if their margin over GA, LSTM, and MLP shrinks materially, the reported 'consistently outperform' conclusion is an artifact of look-ahead leakage rather than model quality.","verdict_should_be":"REJECT","load_bearing_attack":"The paper's central claim—that LGBM, NGBoost, and XGBoost 'consistently outperform other models'—rests on RMSE comparisons in the input and moneyness experiments (Tables 2 and 4). Those experiments use a training set containing September–December 2021 data while the test set is September–December 2020 (Section 4.1). This violates temporal causality: information from after the pricing date is available to every model during training. The paper calls this future data 'denoised information,' but no denoising algorithm or preprocessing step is described; it is simply future data added to the training set. Because boosted tree ensembles are highly flexible, they are precisely the models that can overfit to and exploit such leakage, so the observed ordering among models is not a reliable measure of predictive performance in a realistic out-of-sample setting. The noise experiment (Table 7) is especially affected: its 'denoised dataset' is the same future-augmented training set, so the comparison measures sensitivity to training-set composition rather than any genuine denoising procedure. If the ensemble advantage disappears under a causal temporal split, the central claim is unsupported.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript compares ensemble learning methods (LGBM, XGBoost, NGBoost, CatBoost, DeepForest, RF) with classical ML models (MLP, LSTM, GA) and Black-Scholes/BSM baselines for pricing CSI 300 index options. It reports four experiments: input-feature sets, moneyness subsets, sliding-window on/off, and a noise-robustness comparison. The authors claim that gradient boosting ensembles (LGBM, NGBoost, XGBoost) consistently outperform the other models, and they propose a 'parameter transfer' experimental strategy plus a scoring/weighted-evaluation mechanism intended to embed financial theory into model assessment.","tokens_in":10836,"tokens_out":9241,"duration_ms":81480,"significance":"If the empirical results were valid, the paper would provide a useful benchmark comparison of tree-based ensembles against neural networks and genetic algorithms on real Chinese index options, and the proposed theory-weighted evaluation mechanism might be a reproducible evaluation template. The study is transparent in reporting full RMSE/MSE tables and the data source, and it addresses practically relevant dimensions such as moneyness and local-feature extraction. However, the significance is substantially limited by a temporal leakage in the experimental design and by internal inconsistencies in the proposed scoring formulas; these issues affect the main empirical claims rather than only their presentation.","major_comments":[{"comment":"The training set is stated to include September-December 2021 data while the test set is September-December 2020, and the 2021 data is described as 'denoised information' without any denoising procedure being defined. This is a look-ahead bias: models trained on observations after the test period can use information that would not be available in a real out-of-sample pricing exercise. Because this split underlies Tables 2, 4, 6, and 7, the reported RMSE rankings and the noise-robustness comparison do not constitute a valid evaluation of predictive performance. The noise experiment in Table 7 compares the baseline training set against the future-augmented set, so it measures training-set composition rather than any denoising method. The experiments need to be rerun with a causal chronological split, or the authors must provide a detailed, leakage-free justification for why future data may be used to train a model evaluated on an earlier period.","section":"Section 4.1, Data and data processing method"},{"comment":"The definitions of the Score Rate are internally inconsistent. The text says e is 'the smallest numerical error observed within that sub-experiment,' but Tables 3 and 5 report different score rates for every model, which is only possible if e is the evaluated model's own error. The later phrase 'unless the numerical error of the evaluated model is greater than that of the BS model' also implies e is model-specific. As written, Eqs. (4)-(5) yield a single scalar per sub-experiment and cannot reproduce the tables. Please rewrite the definitions with explicit notation such as e_i for model i, state which quantity is the denominator in each formula, and give the weighted aggregation formula used across sub-experiments.","section":"Section 3.2, Eqs. (4)-(5)"},{"comment":"The claim that a lower Score Rate when the BS model is included 'implies that the BS model contributes positively' is not established. For a fixed model error e, Score_BS < Score_NP is equivalent to E_BS < E_NP, where E_BS and E_NP are the denominators in Eqs. (4) and (5). Thus the comparison only reflects the relative size of the largest errors used as denominators; it does not measure the BS model's causal contribution to the experiment. The corresponding conclusion in Section 6 should be removed or supported by a formal argument that rules out this alternative explanation.","section":"Section 5, Discussion (BS-model inference) and Section 6, Conclusion"}],"minor_comments":[{"comment":"The sentence 'the parameter settings of the remaining experiments in Table 6 are directly inherited from In1' appears to reference the wrong table; the input experiment sub-experiments are in Table 2, so the intended reference is likely Table 2 or a general statement about all later experiments.","section":"Section 3.1"},{"comment":"There is a typo in 'feild' which should be 'field.'","section":"Section 4.2"},{"comment":"Reference [15] contains an apparent typo in the author name 'Ivas,cu'; the comma should be removed or replaced with the proper diacritic.","section":"References"},{"comment":"The RMSE magnitudes in Table 6 are orders of magnitude smaller than those in Tables 2 and 4; please clarify whether the output variable is scaled differently in the sliding-window experiments or whether a different subset of data is used, so readers can compare results across tables.","section":"Table 6 versus Tables 2 and 4"},{"comment":"In Eq. (1), the notation f(x_i) under the derivative is ambiguous; the subscript f_{m-1}(x) in the evaluation point suggests the derivative should be written with an explicit model index, e.g., ∂L(y_i, f(x_i))/∂f(x_i) evaluated at f = f_{m-1}(x).","section":"Section 2.2, Eq. (1)"}],"recommendation":"major_revision","confidential_remarks":"The paper's main contribution is an empirical comparison, but the temporal leakage in the training/test split is a serious flaw that will require a full re-execution of the experiments. If the authors can rerun with a causal split and correct the scoring-formula definitions, the paper may become acceptable; if they cannot, the central claims should not stand. The 'first' claims in the Introduction and Conclusion (first to use ensemble learning, first to propose the evaluation mechanism) are also stronger than the literature review supports and should be moderated in revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: the main empirical claim—that LGBM, NGBoost, and XGBoost beat the other models—cannot be trusted as reported, because the training set includes September–December 2021 data while the test set is September–December 2020. That is a causal leak, not a 'denoising' step: no denoising procedure is described, and future observations are simply added to training. Boosted trees are flexible enough to exploit exactly this kind of leakage, so the observed ranking may reflect data snooping rather than model quality. The stress-test note lands.\n\nWhat is genuinely useful here is narrower. The paper extends the existing ensemble-learning-for-option-pricing line to CSI 300 options, with six input specifications and moneyness splits, and it is careful to test sliding windows and noise. The discussion of flexibility versus noise robustness is appropriately cautious: it reports a possible trade-off without overclaiming. The parameter-transfer heuristic (tune once, reuse) and the weighted score mechanism are small design choices, not major contributions. The claim to be 'first' to use ensemble learning as the central framework is contradicted by the paper's own citation of Ivas,cu (2021), and the score-rate formulas (4)–(5) are ambiguous: E, E_BS, and e are defined loosely, and the inference that a lower score rate 'with the BS model' proves BS contributes positively does not follow.\n\nThe other soft spots are real but secondary: no error bars, no statistical tests, no code or data release, and hyperparameters are not reported. Table 6 shows tiny RMSE differences in the sliding-window experiment, so the qualitative claims there are fragile. The noise experiment inherits the same temporal leak, so Table 7 measures training-set composition, not denoising.\n\nWho is this for? A reader wanting a quick benchmark of tree ensembles on Chinese index options might skim the tables, but the leak undermines the headline comparison. With a proper causal split and code and data, the benchmark could be useful; as it stands, the central result is unsupported. I would not put this in front of a referee. Desk reject, with an invitation to resubmit after a real temporal split.","headline":"The paper's central model comparison is invalid as reported because the training set contains future data relative to the test window, so the ensemble advantage may be leakage rather than genuine predictive skill.","tokens_in":94,"tokens_out":2263,"would_cite":false,"duration_ms":59802,"reading_group":"no","serious_thinker":"yes","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["91G20","68T05"],"pacs":[],"model":"deepseek-v4-flash","headline":"Ensemble learning—specifically LGBM, NGBoost, and XGBoost—consistently outperforms other tested models in pricing CSI 300 index options.","keywords":["Option pricing","Ensemble learning","Gradient boosting","LGBM","NGBoost","XGBoost","Black-Scholes model","Noise robustness"],"falsifier":"Re-run the input and moneyness experiments with a strictly chronological split—train on January–August 2020 only, test on September–December 2020—and compare RMSE rankings. If LGBM, NGBoost, or XGBoost lose their lead, or if the denoised-training-set effect in the noise experiment reverses, the central claim of ensemble superiority in these setups would be contradicted.","tokens_in":10389,"feed_emoji":"📈","tokens_out":6194,"duration_ms":54813,"temperature":0.7,"pith_summary":"This paper tries to establish that ensemble learning models—gradient-boosted tree methods LGBM, NGBoost, and XGBoost in particular—price options more accurately than classical machine learning models and the Black–Scholes baseline. Using CSI 300 index option data, the authors run four experiments covering input configurations, moneyness (the spot-to-strike price ratio), sliding-window local features, and noise robustness. In the two main accuracy experiments, the boosting methods consistently post the lowest root-mean-squared error, outperforming MLP, LSTM, genetic algorithms, random forest, DeepForest, CatBoost, Black–Scholes, and Black–Scholes–Merton. The paper also introduces a hybrid tuning strategy that transfers hyperparameters across experiments and a scoring mechanism that uses the Black–Scholes model as a theoretical anchor, arguing this makes the evaluation more financially grounded. If correct, the results suggest gradient boosting is a practical substitute for neural-network approaches in nonparametric option pricing.","feed_headline":"LGBM, NGBoost, XGBoost outperform rivals on CSI 300 options","feed_subtitle":"On RMSE, boosted-tree ensembles beat deep nets, genetic algorithms, and the Black-Scholes baseline in two main experiments.","key_machinery":"The load-bearing machinery is the family of gradient-boosted tree ensembles—XGBoost, LightGBM, and NGBoost—combined with two methodological innovations. The hybrid tuning strategy restricts hyperparameter search to one sub-experiment and transfers the chosen settings to all other experiments, which is meant to test generalization rather than per-experiment optimization. The evaluation mechanism defines a 'score rate' that compares each model's error to both the best error and the Black–Scholes error, together with a weighted scheme that up-weights sub-experiments using Black–Scholes inputs; higher scores indicate better performance, and the mechanism is designed to let financial theory participate in model evaluation. The experiments also use a training set that mixes 2020 and 2021 data to create what the paper calls a noise-controlled training set.","core_discovery":"In this study, the central discovery is that ensemble methods, especially LGBM, NGBoost, and XGBoost, consistently outperform other models in both major experiments, as measured by RMSE on CSI 300 index options. The paper argues that the structural flexibility and regularization of gradient-boosted trees give them an advantage over MLP and LSTM neural networks, genetic algorithms, and the classical Black–Scholes model. It further claims that a novel experimental strategy—tuning hyperparameters only in an initial sub-experiment and inheriting them elsewhere—improves robustness and mirrors how financial practitioners reuse working models, and that its scoring and weighted evaluation mechanism, which anchors comparisons to Black–Scholes, shows the theoretical model contributes positively to the experiment. The noise and sliding-window experiments are presented as evidence that flexibility and noise robustness are not simply opposed, but the conclusions there are explicitly left open.","pith_inferences":["Because the training set includes data from September–December 2021 while the test set is September–December 2020, a standard chronological split would need to be run to confirm that the reported rankings are not inflated by look-ahead information; this is an extension the paper does not make.","The ensemble advantage may transfer to other option markets (e.g., S&P 500 or FTSE 100) and other derivatives, but the paper only examines CSI 300 index options, so the breadth of the claim is untested.","The proposed score-rate formulas could be adapted to other asset-pricing tasks where a closed-form benchmark exists.","The paper's tentative flexibility/noise trade-off suggests a testable hypothesis: models that capture local information well should degrade more under noise; a dedicated study varying window length and noise level jointly could sharpen this."],"forward_implications":["Practitioners pricing index options can expect gradient-boosted tree ensembles (LGBM, NGBoost, XGBoost) to give lower RMSE than MLP, LSTM, GA, and Black–Scholes on similar data.","The score-rate mechanism provides a template for evaluations in which a theoretical finance model is used as a reference point, not just an accuracy benchmark.","The transfer of hyperparameters across experiments, if accepted, reduces the computational cost of tuning while testing generalization.","The sliding-window experiments imply that the benefit of local-feature extraction depends on moneyness: it helps OTM options more than ITM or ATM ones.","The noise experiments single out MLP as robust to the paper's denoised training set, while LGBM and XGBoost show consistent performance."],"supporting_citations":[{"why":"Establishes the Black–Scholes baseline used for comparison and for the score-rate formula.","marker":"[3]"},{"why":"Defines XGBoost, one of the three ensemble models central to the claim.","marker":"[7]"},{"why":"Defines NGBoost and its natural-gradient probabilistic formulation.","marker":"[9]"},{"why":"Provides the machine-learning option-pricing framework the paper extends and the genetic-algorithm comparison.","marker":"[15]"},{"why":"Defines LightGBM, the third core ensemble model.","marker":"[16]"},{"why":"Supplies the data-processing strategy and the introduction of the BSM dividend term used in the experiments.","marker":"[17]"},{"why":"Supplies the Black–Scholes–Merton extension used as a second theoretical baseline.","marker":"[22]"},{"why":"Defines DeepForest, the cascade-forest ensemble included in the comparisons.","marker":"[30]"}],"fun_headline_variants":["Boosted trees beat deep nets and Black-Scholes on options","Gradient boosting bests neural nets for option pricing","LGBM, NGBoost, XGBoost lead option pricing study","Ensemble methods win for CSI 300 option RMSE","Study: boosted trees outperform in option pricing"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The ranking of models rests on treating a training set that includes 2021 data as a valid noise-controlled basis for predicting a 2020 test period; if that temporal mixing introduces look-ahead information, the reported robustness and accuracy results would need to be reassessed.","fun_headline_variants_meta":{"raw":{"variants":["Boosted trees beat deep nets and Black-Scholes on options","Gradient boosting bests neural nets for option pricing","LGBM, NGBoost, XGBoost lead option pricing study","Ensemble methods win for CSI 300 option RMSE","Study: boosted trees outperform in option pricing"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000343,"raw_usage":{"total_tokens":1856,"prompt_tokens":886,"completion_tokens":970,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":502,"completion_tokens_details":{"reasoning_tokens":886}},"tokens_in":502,"tokens_out":970,"duration_ms":9707,"temperature":1.0,"reasoning_tokens":886,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T10:12:57.348805+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the input and moneyness experiments with a strictly chronological split—train on January–August 2020 only, test on September–December 2020—and compare RMSE rankings. If LGBM, NGBoost, or XGBoost lose their lead, or if the denoised-training-set effect in the noise experiment reverses, the central claim of ensemble superiority in these setups would be contradicted.","supporting_citations":[{"cited_title":"Xgboost: A scalable tree boosting system, in: Proceedings of the 22nd acm sigkdd international conference on knowledge discovery and data mining, pp","cited_arxiv_id":null,"evidence_quote":"Defines XGBoost, one of the three ensemble models central to the claim."},{"cited_title":"Ngboost: Natural gradient boosting for probabilistic pre- diction, in: International conference on machine learning, PMLR","cited_arxiv_id":null,"evidence_quote":"Defines NGBoost and its natural-gradient probabilistic formulation."},{"cited_title":"Option pricing using machine learning","cited_arxiv_id":null,"evidence_quote":"Provides the machine-learning option-pricing framework the paper extends and the genetic-algorithm comparison."},{"cited_title":"Lightgbm: A highly efficient gradient boosting decision tree","cited_arxiv_id":null,"evidence_quote":"Defines LightGBM, the third core ensemble model."},{"cited_title":"Option Pricing with Convolutional Kolmogorov-Arnold Networks","cited_arxiv_id":"2412.01224","evidence_quote":"Supplies the data-processing strategy and the introduction of the BSM dividend term used in the experiments."},{"cited_title":"Theory of rational option pricing","cited_arxiv_id":null,"evidence_quote":"Supplies the Black–Scholes–Merton extension used as a second theoretical baseline."},{"cited_title":"Deep forest","cited_arxiv_id":null,"evidence_quote":"Defines DeepForest, the cascade-forest ensemble included in the comparisons."}],"review_version":1}