{"id":"55b8bac8-f716-4520-8cd0-304778c11ed3","arxiv_id":"2608.12583","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A structured survey of diffusion-family generative models in finance, organized by financial data type, with an open-source reference repository.","lead":"This paper reviews and organizes the growing literature on diffusion-based generative models applied to financial data, grouping papers by data type: time series, limit order books, tabular records, and structured objects such as volatility surfaces. It provides a shared taxonomy, background on diffusion and flow models, and a list of open research directions, plus a GitHub repository for further detail.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'first survey' priority claim and corpus representativeness rest on an undocumented selection process; a systematic literature search can settle whether the central claim holds.","rationale":"I agree with the reader's weakest assumption: the central claim is a coverage and priority claim, and the absence of a methodology section makes it unfalsifiable from the manuscript alone. In good faith, I note real support: the repository, the table's explicit caveat, and the statement 'to the best of our knowledge' show that the authors are not hiding uncertainty. The concern is not that the authors are sloppy; it is that the survey's contribution cannot be checked without an external search. That is exactly what a conditional verdict should require. I would not escalate to reject because the paper is competent and the repository provides partial transparency; I would not accept outright until the search protocol or a verified list exists. The one concrete test above would move this from conditional to accept if the search finds no prior survey and no major omitted cluster. The background math appears sound, and I found no internal inconsistency that would be load-bearing for the survey's central claim.","tokens_in":14494,"tokens_out":4088,"duration_ms":42895,"concrete_test":"Run a reproducible search in arXiv, Scopus, and SSRN with an explicit date cutoff (e.g., before 12 Aug 2026) using title/abstract queries such as ('diffusion model' OR 'denoising diffusion' OR 'score-based' OR 'flow matching') AND ('finance' OR 'financial' OR 'stock' OR 'option' OR 'limit order book'), plus a separate search for 'survey' AND 'diffusion' AND 'finance'. Compare the resulting paper set against Table 1 and the GitHub repository, and check every returned title for a dedicated survey on diffusion-family models in finance. If a dedicated survey predates the submission date, the 'first survey' claim is falsified. If the search surfaces a core application area absent from the taxonomy, the representativeness assumption weakens.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract and Section 1 assert that this is the first survey dedicated specifically to diffusion-family models for financial data, and the survey's value is its coverage and taxonomy. For that claim to hold, two conditions must be true: no earlier dedicated survey exists, and the papers in Table 1 are representative of the relevant universe. Neither condition is backed by a documented method. Section 1 mentions adjacent surveys (synthetic data in finance [65], generic time-series diffusion [93], tabular diffusion [50], LLM agents [49, 70]) but does not say how the authors searched, what databases or date ranges were used, or what inclusion/exclusion criteria were applied. Table 1's caption explicitly calls it 'illustrative rather than exhaustive,' which is honest but also concedes that the corpus is a selection. The linked repository is a useful transparency artifact, but a list of papers is not a search protocol. A prior dedicated survey or a missing major cluster (for example, option-pricing or market-microstructure work covered only sparsely) would falsify or weaken the central claim. The hedge 'to the best of our knowledge' does not repair the underlying verifiability gap. This is the load-bearing weak point because the paper introduces no new methods or results; its contribution is precisely the completeness and priority of the review. The paper itself, in Section 7, calls for benchmark and evaluation discipline, yet it does not document the corpus-construction protocol for its own survey.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper surveys diffusion-family generative models applied to financial data. The survey is organized primarily by financial data modality: time series, limit order books, tabular data, and other structured financial objects such as correlation matrices, implied-volatility surfaces, yield curves, and derivative paths. It provides a compact technical background covering score-based diffusion, the score-SDE formulation, flow matching, guidance, sampling accelerators, and common architectural backbones. The bulk of the paper reviews roughly forty works drawn from the authors' open-source repository, and it closes with three open directions: finance-specific evaluation benchmarks, scaling laws for financial diffusion models, and moving from generation to decision-making. The paper claims in the abstract and introduction to be the first survey dedicated specifically to diffusion-family models for financial data.","tokens_in":14827,"tokens_out":4666,"duration_ms":49715,"significance":"If the corpus is representative and the priority claim holds, this survey would serve as a useful entry point for researchers entering the intersection of diffusion-based generative modeling and finance. The technical background is largely correct and the data-type-driven taxonomy is sensible and reasonably well executed. The open-problems section (§7) is constructive and identifies real weaknesses in the current literature, and the open-source repository is a useful transparency artifact. However, the survey introduces no new methods or results; its value is entirely in coverage and synthesis. That value is contingent on two claims—representativeness of the selected papers and priority of the 'first survey' assertion—neither of which is currently supported by a documented methodology.","major_comments":[{"comment":"The survey's central claim of comprehensiveness rests on an undocumented selection process. The paper states that it 'reviews the growing literature' and claims to be the first dedicated survey, but it provides no search protocol, database list, search strings, date range, or inclusion/exclusion criteria. Table 1 is explicitly captioned as 'illustrative rather than exhaustive,' which concedes that the corpus is a selection. Because the paper contributes no new methods or results, the completeness and representativeness of this selection are load-bearing; without a methodology section, the reader cannot verify the taxonomy or the priority claim. I recommend adding a 'Scope and Method' subsection describing how papers were identified, screened, and categorized, and either expanding the coverage to be exhaustive within a clearly stated scope or explicitly reframing the contribution as a structured overview of a representative subset.","section":"Section 1 and Table 1"},{"comment":"The claim 'this is the first survey dedicated specifically to diffusion-family models for financial data' is asserted only with the hedge 'to the best of our knowledge' and is not supported by a systematic literature check. The paper mentions adjacent surveys on synthetic data in finance [65], generic time-series diffusion [93], tabular diffusion [50], and LLM agents in finance [49, 70], but it does not demonstrate that no earlier dedicated survey of diffusion models for financial data exists. The linked repository is a list of papers, not a search protocol, so it does not repair this verifiability gap. This is particularly important because the priority claim is part of the paper's stated contribution; I ask the authors to either document a reproducible search that supports the claim or weaken the claim to 'to the best of our knowledge, no prior survey has focused on this specific combination,' with the search evidence reported in an appendix.","section":"Abstract and Section 1"},{"comment":"The paper calls for benchmark and evaluation discipline in the field, yet it does not apply such discipline to its own corpus construction. Section 7 argues that the field suffers from proprietary data, incompatible horizons, and weak baselines, but the survey itself does not report any quality screening (e.g., peer-reviewed versus preprint status), any verification of the contributions of the included papers, or any criteria for why some borderline works (such as ByteGen [47], which the paper itself says is 'not a standard diffusion model') are included. The open-problems list, however reasonable, is derived from a selection that the authors acknowledge is illustrative; the list might look different had the corpus been assembled with a documented protocol. Please add a description of the inclusion criteria and a per-paper quality/venue classification, at least in an appendix.","section":"Section 7"}],"minor_comments":[{"comment":"The phrase 'stable likelihood-based training' is imprecise: diffusion models are typically trained with a variational lower bound or score-matching objective rather than exact likelihood, and the standard DDPM objective is a denoising objective, not a likelihood. Consider rephrasing to 'stable training based on denoising objectives' or similar.","section":"Abstract and Section 1"},{"comment":"The inclusion of ByteGen [47] as an order-flow generation method, while the paper explicitly states it is not a standard diffusion model, is potentially confusing in a survey of diffusion-family models. Please clarify whether ByteGen is included as a baseline/context or whether the survey explicitly covers adjacent generative models; if the latter, state the inclusion criterion.","section":"Section 4.1"},{"comment":"The distinction between 'conditional generation' and 'prediction and trading' is based on the stated purpose of each paper rather than on architecture or objective, which is a reasonable organizational principle, but it leads to the same paper being discussed in one subsection while some of its generated outputs are also relevant to the other. A short sentence at the beginning of Section 3.3 noting that the division is functional rather than architectural would reduce ambiguity.","section":"Section 3.2 and Section 3.3"},{"comment":"Many entries in Table 1 are arXiv preprints with 2025 or 2026 dates, and the survey does not indicate which have been peer-reviewed. Since the survey aims to be a reference, please add a column or footnote indicating the publication status (peer-reviewed conference/journal, arXiv preprint, or forthcoming) for each entry.","section":"Table 1 and Repository"},{"comment":"There are several minor grammar issues, for example: Section 3.2 'Guo et al. [24] uses' should be 'Guo et al. [24] use'; Section 3.2 'Zarifis et al. [95] use multivariate energy price time series and generates conditional scenario paths' should be 'generate conditional scenario paths.'","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The paper is honest about the illustrative nature of Table 1 and the 'to the best of our knowledge' hedge, and the self-citations in [87], [88], and [89] are transparent. The main concern is methodological: a survey whose contribution is coverage and priority must document its corpus construction. I would encourage the editor to consider whether a quick literature check (e.g., by a second referee) can confirm that no earlier dedicated survey exists, since the authors' lack of a documented search makes the priority claim currently unverified. The manuscript fits a computational-finance venue well, and the open-problems section is a genuine strength."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The useful news: this is a clearly written, well-structured survey that does something I haven't seen elsewhere—it maps diffusion-family models onto the data objects finance actually cares about: time series, limit order books, tabular data, and structured objects like correlation matrices and yield curves. The background section is compact and accurate; the score-SDE and flow-matching material is standard but correct, and the authors are honest that Table 1 is illustrative, not exhaustive. The open repository is a practical addition. For someone entering this space, this is a genuinely helpful map.\n\nThe main soft spot is exactly what the stress-test flags: the 'first dedicated survey' claim is asserted, not demonstrated. There's no search protocol, no inclusion/exclusion criteria, and no discussion of how the corpus was assembled. That matters for a survey whose value depends on coverage, and it should be fixed in revision. But it's not fatal. The authors hedge with 'to the best of our knowledge,' they do cite the adjacent surveys, and the real contribution—the data-type taxonomy and curated organization—stands even if an earlier dedicated survey exists. A quick literature search can settle the priority question; the structural framing does not depend on it.\n\nA smaller limitation: the paper catalogs work but doesn't critically assess it. You won't learn which methods actually perform well or how they compare. That's normal for a first survey, but it keeps the value organizational rather than evaluative.\n\nOverall, this is a solid, honest piece of work. It deserves a serious referee, and I'd cite it. The missing search protocol is a legitimate but fixable gap; the core content is useful as is.","headline":"A useful, well-organized survey of diffusion models in finance; the 'first survey' claim is under-supported but the data-type taxonomy stands on its own.","tokens_in":15248,"tokens_out":2178,"would_cite":true,"duration_ms":24418,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This survey claims to be the first dedicated overview of diffusion-family generative models in finance, organizing the literature by the type of financial data being generated rather than by model family.","keywords":["Diffusion models","Flow models","Finance","Survey","Financial time series","Limit order books","Tabular data","Generative models"],"falsifier":"A reproducible literature search across preprint servers and bibliographic databases, using explicit queries for diffusion, score-based, and flow-matching models combined with finance terms and explicit date ranges and inclusion criteria, that surfaces a peer-reviewed survey dedicated to diffusion models for financial data published before this one would falsify the priority claim; finding a substantial body of diffusion finance work in a data modality the taxonomy omits would falsify its completeness claim.","tokens_in":14250,"feed_emoji":"📈","tokens_out":5826,"duration_ms":54199,"temperature":0.7,"pith_summary":"This paper claims that diffusion-family generative models, including denoising diffusion probabilistic models, score-based models, flow matching, and stochastic interpolants, have formed a distinct research area in finance and that this area is now mature enough to warrant a dedicated survey organized by financial data type. It reviews work on financial time series, limit order books, tabular data, and structured objects such as correlation matrices, implied-volatility surfaces, yield curves, and derivative paths, and it identifies evaluation, scaling, and decision-making as the open directions. The stated reason to care is that diffusion models' stochastic differential equation formulation, stable training, flexible conditioning, and mode coverage match the needs of financial data generation, stress testing, and portfolio and trading decisions.","feed_headline":"First dedicated survey maps diffusion models in finance","feed_subtitle":"A data-first taxonomy sorts a dispersed literature into time series, order books, tabular data, and structured objects.","key_machinery":"The survey's organizing device is a taxonomy that sorts papers by financial data type, separating time series, limit order books, tabular records, and other structured financial objects, and within each type by modeling goal such as unconditional generation, conditional generation, prediction and trading, synthesis and augmentation, and privacy and trustworthiness. The technical ground it uses to unify the field is the score-SDE formulation of diffusion models, in which a forward noising stochastic differential equation is reversed by learning a neural score or noise predictor, together with flow matching's alternative view of generation as learning a vector field that transports noise into data. This shared formalism is what lets the survey treat very different applications as instances of one design space.","core_discovery":"The central claim is that existing surveys cover synthetic data in finance, diffusion models for generic time series, diffusion for tabular data, and large language models in finance, but none is dedicated to diffusion-family models across financial applications. The paper therefore asserts, to the best of its knowledge, that it is the first survey dedicated specifically to diffusion-family models for financial data, and it defends a data-first reading: the literature is best organized not by model family but by the financial object being generated, with the validity of a generator defined by the constraints of that object, such as tick sizes and event coherence for order books, mixed-type and privacy constraints for tabular records, positive semidefiniteness for correlation matrices, and martingale conditions for pricing paths.","pith_inferences":["The priority claim is the most brittle part of the paper: a documented systematic search that surfaces any earlier dedicated survey of diffusion models for financial data would overturn that claim without necessarily damaging the taxonomic contribution.","The data-type split is not the only plausible axis; organizing the same literature by decision task, such as forecasting, simulation, privacy, or pricing, might reveal different clusters, and future surveys could test which axis better predicts methodological choices.","A testable extension of the data-first thesis is a transfer experiment: training a financial diffusion generator on one asset class and sampling another should degrade unless the model is conditioned appropriately, which would support the paper's claim that validity is tied to the financial object.","The paper's benchmark argument points to a concrete next step: a shared leaderboard covering the four data types with common metrics for fidelity, downstream utility, privacy, and constraint satisfaction."],"forward_implications":["Practitioners can use the data-type taxonomy as a directory to locate relevant diffusion work for their object of interest and to see which modeling goals have already been tried.","The open-problem list implies that the next phase should be finance-specific benchmark suites, because current comparisons are fragmented across proprietary data, incompatible horizons, preprocessing choices, and weak baselines.","It implies that scaling behavior of financial diffusion models should be studied directly, relating compute, data, and conditioning to generation quality rather than assuming scaling laws transfer from language or image models.","It implies a shift in research target from realistic samples toward decision value: portfolio construction, hedging, execution, and market making framed as diffusion-generated actions rather than merely market paths."],"supporting_citations":[{"why":"Existing survey of synthetic data applications in finance that the paper distinguishes itself from.","marker":"[65]"},{"why":"Existing survey of diffusion models for time series and spatio-temporal data, not finance-specific.","marker":"[93]"},{"why":"Existing survey of diffusion models for tabular data, not finance-specific.","marker":"[50]"},{"why":"Existing survey of large language models in finance, cited to show no dedicated diffusion-finance survey exists.","marker":"[49]"},{"why":"Existing survey of large language model agents for investment management, completing the comparison of adjacent finance surveys.","marker":"[70]"}],"fun_headline_variants":["First dedicated survey: diffusion models in finance","Diffusion models meet finance: a first survey","First finance survey of diffusion models","Data-first survey of diffusion models in finance","Survey maps diffusion models by financial data type"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the papers selected for review, chosen without any stated search or inclusion/exclusion protocol, represent the full universe of diffusion-family finance research; if significant work was missed, the taxonomy and the 'first survey' claim would not stand.","fun_headline_variants_meta":{"raw":{"variants":["First dedicated survey: diffusion models in finance","Diffusion models meet finance: a first survey","First finance survey of diffusion models","Data-first survey of diffusion models in finance","Survey maps diffusion models by financial data type"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000719,"raw_usage":{"total_tokens":3171,"prompt_tokens":830,"completion_tokens":2341,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":446,"completion_tokens_details":{"reasoning_tokens":2276}},"tokens_in":446,"tokens_out":2341,"duration_ms":16026,"temperature":1.0,"reasoning_tokens":2276,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T00:04:23.796149+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A reproducible literature search across preprint servers and bibliographic databases, using explicit queries for diffusion, score-based, and flow-matching models combined with finance terms and explicit date ranges and inclusion criteria, that surfaces a peer-reviewed survey dedicated to diffusion models for financial data published before this one would falsify the priority claim; finding a substantial body of diffusion finance work in a data modality the taxonomy omits would falsify its completeness claim.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Existing survey of large language model agents for investment management, completing the comparison of adjacent finance surveys."}],"review_version":1}