{"id":"41682695-a20d-49ad-9165-c7c81c0d7de0","arxiv_id":"2508.15922","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"The abstract proposes QRS and probabilistic stacking for quantile forecasts of cryptocurrency realized variance, claiming QRS with linear models performs best, but the supplied full text is unrelated.","lead":"This preprint's abstract describes probabilistic forecasts of Bitcoin volatility using quantile methods from multiple base models, but the uploaded full text is a different paper entirely. Read the abstract with caution because the supporting body is missing.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Manuscript body is an unrelated paper; the QRS/volatility forecast claims have no supporting text or results.","rationale":"The reader's verdict of UNVERDICTED is appropriate. The weakest assumption identified by the reader—QRS residual resampling yields correctly calibrated quantiles—is indeed a critical unverified aspect, but the more immediate and severe problem is that the manuscript body does not even address QRS or cryptocurrency volatility. The reader correctly noted insufficient information; our stress test confirms this with the additional observation that the full text is a different article entirely. Because we cannot access the actual cryptocurrency paper content (if it exists elsewhere), we cannot verify or refute the empirical claim. Thus the verdict should remain UNVERDICTED, pending retrieval of the correct manuscript or clarification from the authors. We do not recommend REJECT because the mismatch could be a clerical error in the arXiv posting; however, as submitted, the claim is wholly unsupported. Our concrete test would settle whether the correct content is available and whether the results support the abstract.","tokens_in":1313,"tokens_out":2995,"duration_ms":32827,"concrete_test":"Retrieve the full source/PDF for arXiv:2508.15922 from arXiv or the publisher. Search the full text for the keywords 'QRS', 'realized variance', 'Bitcoin', 'cryptocurrency', 'quantile', and 'residual simulation'. If none of these terms appear in the body, the abstract and manuscript are mismatched, confirming the concern. If the terms do appear (e.g., in a corrected version), then check that the empirical comparisons in the results section actually support the abstract's claim that QRS consistently outperforms—including evaluation metrics, statistical tests, and robustness checks.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that QRS applied to linear log-volatility models outperforms sophisticated alternatives for cryptocurrency realized variance quantile forecasting—cannot be evaluated because the submitted full text is a completely different preprint on thermodynamically consistent neural yield functions for anisotropic plasticity. There is no mention of QRS, realized variance, Bitcoin, quantile regression, or any cryptocurrency data or experiments. The abstract and full text are internally inconsistent. For the abstract's claim to hold, the paper would need to contain the described empirical methodology, results, and comparisons; none of these are present. This is not a matter of subtle methodological weakness or an arguable assumption; the evidential basis for the claim is entirely missing, making the claim unsupported and unverifiable from the provided manuscript. This is the load-bearing concern: the document does not contain the study that the abstract purports to summarize.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The abstract of this submission claims a novel probabilistic forecasting framework for cryptocurrency realized variance, introducing the 'Quantile Estimation through Residual Simulation' (QRS) method, and reports that QRS applied to linear base models on log-transformed realized volatility outperforms more sophisticated alternatives for Bitcoin. The supplied full text, however, is a completely different preprint titled 'Thermodynamically Consistent Hybrid and Permutation-Invariant Neural Yield Functions for Anisotropic Plasticity' (arXiv:2508.15923v1). This body contains no mention of QRS, realized variance, Bitcoin, quantile forecasting, HAR/GARCH/ARFIMA, LASSO/SVR/MLP/LSTM, or any cryptocurrency data. The empirical results advertised in the abstract have no supporting derivation, experimental protocol, or evaluation in the manuscript text.","tokens_in":1531,"tokens_out":1754,"duration_ms":21012,"significance":"If the abstract's claims were supported, the paper would potentially offer a simple and practical baseline for probabilistic cryptocurrency volatility forecasting, with QRS as a lightweight alternative to fully conditional quantile models. The claimed robustness of probabilistic stacking would also be a useful practical contribution. However, because the manuscript body is an unrelated paper on plasticity, the claims cannot be assessed at all. There is no reproducible code, no machine-checked proof, and no falsifiable evidence in the submitted document. The significance of the result is therefore entirely unevaluated.","major_comments":[{"comment":"The manuscript body is an unrelated preprint on neural yield functions for anisotropic plasticity (arXiv:2508.15923v1). It contains no discussion of QRS, realized variance, cryptocurrency volatility, Bitcoin, or any forecasting experiment. Consequently, the central claim of the abstract—that QRS applied to linear base models on log-transformed realized volatility outperforms alternatives—has no supporting evidence in the submitted document. This is not a subtle methodological issue; the evidential basis for the paper's stated contribution is entirely absent.","section":"Full text (entire body after abstract)"},{"comment":"The abstract describes a 'first study' proposing probabilistic forecasting of cryptocurrency variance, but the supplied full text does not present any such study. There are no equations, no data description, no evaluation metrics (e.g., quantile loss, coverage, Winkler score), and no comparison tables. The phrase 'consistently outperforms' is therefore an unsupported assertion. As a reviewer, I cannot verify any of the empirical claims, including the assumed validity of residual-simulation quantile estimation.","section":"Abstract vs Full text"},{"comment":"The document's running arXiv identifier (2508.15923) is different from the submission identifier (2508.15922), and the title, authors, and subject matter are entirely different. This mismatch further confirms that the abstract and body are not part of the same work. If this is an accidental file swap, the correct manuscript should be provided; as submitted, the paper cannot be evaluated for publication.","section":"Manuscript metadata"}],"minor_comments":[{"comment":"No page numbers, section numbers, or references are provided in the body, making it impossible to navigate or verify any specific claim. This is a significant presentation deficiency, though secondary to the substantive mismatch.","section":"General formatting"},{"comment":"The title and abstract do not match the body. Even if the correct manuscript were uploaded, the authors should ensure that the abstract accurately reflects the content and that all arXiv identifiers are consistent.","section":"Title and abstract"}],"recommendation":"reject","confidential_remarks":"This submission appears to be an accidental inclusion of an unrelated paper. The arXiv ID in the footer (2508.15923) is not the submission's ID (2508.15922), and the subject matter is entirely different. If this is a pipeline or file-submission error, the editor may wish to invite a corrected submission. However, on the current record, the manuscript is not evaluable and cannot be accepted."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nYou need to know one thing up front: the submitted full text has nothing to do with the abstract. The abstract promises a systematic comparison of probabilistic volatility forecasts for Bitcoin using methods like QRS, HAR, GARCH, and LSTM. The body is a preprint on thermodynamically consistent neural yield functions for anisotropic plasticity. Same arXiv number? No. It looks like a file mix-up, but as it stands, the manuscript is internally incoherent.\n\nWhat the abstract describes is actually a decent idea. A clean benchmark of quantile forecasting methods for crypto realized variance, with a simple residual-simulation approach against fancier ML alternatives, would be a useful contribution to crypto risk management. The \"first systematic evaluation\" claim is checkable and plausible, though not assured. The QRS method is essentially a residual bootstrap on point forecasts, which is standard but can be effective. If the real paper contains what the abstract says, it would be worth a careful read.\n\nBut none of that is in front of us. There is no derivation, no experimental protocol, no data, no results, no code. The soft spot is not a subtle assumption about residual resampling; it is the complete absence of any evidence for the claims. The stress-test note is correct: the document does not contain the study it purports to summarize. I cannot verify novelty, soundness, calibration, or even whether the authors actually ran the experiments. The reader's UNVERDICTED verdict is the only honest one.\n\nI would not send this to peer review as is. A serious editor would desk reject or, more charitably, return it to the authors to correct the file. The abstract alone does not carry enough weight to justify referee time when the supporting text is missing. If a corrected manuscript arrives with the actual crypto content, then yes, I would send it out—the question is relevant and the proposed comparison is sensible.\n\nFor now: do not cite, do not discuss at reading group. Flag it as a submission error.\n\nBest,\n[You]","headline":"The abstract describes a useful crypto-volatility forecasting benchmark, but the manuscript body is an unrelated plasticity paper, so there is nothing to review.","tokens_in":1909,"tokens_out":1533,"would_cite":false,"duration_ms":18918,"reading_group":"no","serious_thinker":"no","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that a simple residual-resampling method, QRS, applied to linear models of log-transformed realized volatility, produces the most accurate probabilistic forecasts of Bitcoin volatility among a wide field of statistical and","keywords":["probabilistic forecasting","realized volatility","cryptocurrency","Bitcoin","quantile regression","residual simulation","model stacking","volatility forecasting"],"falsifier":"An out-of-sample test on a different cryptocurrency or a longer Bitcoin sample where QRS quantile coverage is systematically below the nominal level (e.g., the 5% quantile is exceeded more than 5% of the time) while a parametric GARCH quantile matches coverage would falsify the claim.","tokens_in":1283,"feed_emoji":"📉","tokens_out":5004,"duration_ms":49961,"temperature":0.7,"pith_summary":"This paper tries to establish that probabilistic forecasts of cryptocurrency realized variance can be built cheaply from ordinary point forecasts, and that the simplest route—simulating quantiles by resampling residuals from linear models of log-transformed realized volatility—is consistently more accurate than heavier statistical and machine-learning alternatives. The authors evaluate a wide range of base models, including HAR, GARCH, ARFIMA, LASSO, SVR, MLP, Random Forest, and LSTM, and wrap them in a probabilistic-stacking framework for Bitcoin. If the claim holds, traders and risk managers can obtain well-calibrated volatility quantiles without bespoke deep-learning architectures, and the bar for beating the linear-simulation baseline becomes clearer.","feed_headline":"Residual resampling wins on Bitcoin volatility quantile forecasts","feed_subtitle":"A simple residual-simulation method gives reliable volatility quantiles, beating fancier ML for Bitcoin.","key_machinery":"QRS (Quantile Estimation through Residual Simulation): a residual-resampling procedure that turns a point forecast of log realized volatility into a set of simulated trajectories by adding draws from the model's in-sample residual distribution to the point forecast. Exponentiating back to realized-variance space and taking quantiles yields probabilistic forecasts. The method's success is tied to linear base models and the log transformation, which stabilizes the variance of the target series.","core_discovery":"The central claim is that the Quantile Estimation through Residual Simulation (QRS) method, applied to point forecasts from linear base models on log-transformed realized volatility, consistently outperforms more sophisticated alternatives for probabilistic forecasting of Bitcoin's realized variance. The QRS method converts a point forecast into a full conditional distribution by resampling the residuals of the point-forecast model and adding them to the forecast, thereby generating simulated paths of log realized volatility whose quantiles approximate those of the future realized variance. The paper reports this as the first systematic evaluation of such variance-quantile forecasts in crypt","pith_inferences":["A natural extension is to stress-test the QRS calibration on quantile coverage tests; the paper's reported superiority may be sensitive to the evaluation horizon and rolling-window length.","The same residual-simulation idea could be adapted to forecast tail-risk measures such as Value-at-Risk and Expected Shortfall for crypto portfolios, where extreme quantiles matter most.","Because the method is distribution-free, it may degrade less than parametric alternatives during volatility regime shifts, a hypothesis the paper's described experiments do not directly isolate."],"forward_implications":["If QRS with linear bases is consistently best, then probabilistic volatility forecasting for Bitcoin can be done with transparent, reproducible models rather than opaque ML ensembles.","Residual simulation provides a distribution-free route to volatility quantiles, avoiding parametric GARCH distributional assumptions.","Probabilistic stacking of point-forecast-derived distributions inherits the strengths of the best base models, so gains come from input design (log realized volatility) rather than from the nonlinear learner.","The methodology can be applied directly to other cryptocurrencies or assets without retraining complex architectures."],"supporting_citations":[],"fun_headline_variants":["Residual simulation beats complex ML for Bitcoin volatility quantiles","QRS with linear models wins on Bitcoin variance quantiles","Bitcoin volatility quantiles: residual resampling outperforms","Simple residual method tops intricate ML for Bitcoin forecasts"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The method assumes that drawing residuals from the base model's historical errors reproduces the shape of the future conditional distribution of log realized volatility—that the error distribution is stable enough to resample.","fun_headline_variants_meta":{"raw":{"variants":["Residual simulation beats complex ML for Bitcoin volatility quantiles","QRS with linear models wins on Bitcoin variance quantiles","Bitcoin volatility quantiles: residual resampling outperforms","Simple residual method tops intricate ML for Bitcoin forecasts"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000442,"raw_usage":{"total_tokens":2065,"prompt_tokens":720,"completion_tokens":1345,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":464,"completion_tokens_details":{"reasoning_tokens":1280}},"tokens_in":464,"tokens_out":1345,"duration_ms":11436,"temperature":1.0,"reasoning_tokens":1280,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T17:38:26.734479+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"An out-of-sample test on a different cryptocurrency or a longer Bitcoin sample where QRS quantile coverage is systematically below the nominal level (e.g., the 5% quantile is exceeded more than 5% of the time) while a parametric GARCH quantile matches coverage would falsify the claim.","supporting_citations":[],"review_version":1}