{"id":"a1bec366-696c-49ab-b8f0-5ca6efb25a4c","arxiv_id":"2412.15270","paper_version":2,"verdict":"REJECT","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":3,"one_line_summary":"A finance-tuned LLM reportedly beats strong baselines on Chinese financial exams, but the training data includes exam questions and no contamination check is reported.","lead":"Baichuan4-Finance is a finance-specialized LLM that the authors claim outperforms GPT-4o on Chinese financial certification exams. The report describes a data pipeline and a 'domain self-constraint' training strategy, but the evaluation may be inflated because the training data includes exam questions.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Pre-training corpus explicitly includes 'examination questions' (Sec. 3.3) while headline results come from public certification-exam benchmarks; no contamination check is reported, so the claimed SOTA margins may reflect memorization rather than capability.","rationale":"The reader's weakest-assumption analysis identifies benchmark contamination as the load-bearing threat, and the manuscript text supports that concern directly: Section 3.3 lists 'examination questions' as an explicit category among the 37 financial data sources used for continual pre-training, while the headline evaluations use public financial certification exams (FinanceIQ and FLAME-Cer). A model trained on 100B tokens of financial data that includes exam questions could plausibly memorize a substantial fraction of these public multiple-choice items, especially after annealing on high-quality financial sources. The paper provides no contamination audit, no overlap statistics, no evaluation on a chronologically held-out exam set, and no release of the model or data that would allow independent verification. These omissions make the reported double-digit margins over GPT-4o and Qwen2.5-72B-Instruct unreliable as evidence of genuine financial reasoning capability. The concern is concrete and internal to the paper's own description of its data pipeline, not an external disagreement about benchmark choices. One decisive test is to evaluate on a freshly created, post-cutoff set of certification questions; if accuracy drops substantially, memorization rather than capability is the likely driver. Because the reader's REJECT verdict already rests on this same weakness, no verdict change is needed.","tokens_in":12309,"tokens_out":2874,"duration_ms":28748,"concrete_test":"Assemble a held-out set of financial certification questions from the same exam boards (AFP, CPA, SPQ, etc.) that were first published after the model's training-data cutoff and were not publicly available before that date. Evaluate Baichuan4-Finance zero-shot on this set using the same prompting and scoring protocol as FinanceIQ and FLAME-Cer. If the held-out accuracy is materially lower than the published benchmark numbers (e.g., more than 5 points), contamination of the 100B-token financial corpus is the likely explanation and the headline comparison is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that Baichuan4-Finance outperforms GPT-4o by large margins on FinanceIQ and FLAME-Cer depends on those public benchmark questions being absent from the 100B-token financial pre-training corpus. Section 3.3 states the corpus comprises 37 financial data sources 'including research reports, academic papers, examination questions, and more.' FinanceIQ (7,173 multiple-choice questions from 10 Chinese financial certifications) and FLAME-Cer (14 financial certifications) are exactly the kind of exam content such a corpus would plausibly absorb from the web. The paper reports no n-gram or semantic overlap analysis between the pre-training corpus and these benchmarks, no deduplication against evaluation data, and no held-out exam set. The annealing phase (Sec. 3.5) additionally up-weights high-quality financial sources, which could increase exposure to certification-exam material. Without evidence that benchmark items are absent, the reported FLAME-Cer average of 93.62 versus 81.17 for Qwen2.5-72B-Instruct could be inflated by training-data memorization. This is a correctness risk to the abstract's 'significant margins' claim, not a stylistic disagreement.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports the development of Baichuan4-Finance, a financial LLM series built by continual pre-training of Baichuan4-Turbo on a mixed general/financial corpus, using a proposed 'domain self-constraint' training objective intended to preserve general knowledge. The resulting base and chat models are evaluated on general benchmarks (C-Eval, MMLU, GSM8K, etc.) and on two Chinese financial certification benchmarks, FinanceIQ and FLAME. The central claim is that Baichuan4-Finance-Base surpasses competitive baselines on financial tasks by significant margins without losing general capability, and that the chat model Baichuan4-Finance achieves state-of-the-art accuracy on FinanceIQ and FLAME-Cer, outperforming GPT-4o and strong open-source models.","tokens_in":12618,"tokens_out":3288,"duration_ms":29317,"significance":"If the reported results are trustworthy, the paper offers a useful engineering recipe for financial LLM continual pre-training, including a data-quality pipeline, a scaling-law-based mixture-ratio selection, and an alignment pipeline with human and AI feedback. The domain self-constraint objective is a plausible mechanism for avoiding catastrophic forgetting. However, the central empirical claims are currently not credible because the pre-training corpus explicitly includes examination questions (Section 3.3) while the headline benchmarks are certification exams (Section 5.1), and no contamination analysis is reported. The paper also relies on externally quoted baseline numbers for FLAME without verification. These issues undermine the validity of the claimed SOTA margins.","major_comments":[{"comment":"The pre-training data composition described in Section 3.3 includes 'examination questions' among the 37 financial data sources, while the main benchmarks FinanceIQ and FLAME-Cer consist of Chinese financial certification exam questions (Section 5.1). The paper reports no n-gram or semantic overlap analysis, no deduplication against evaluation data, and no held-out exam set. Consequently, the large margins reported in Table 1 (e.g., FLAME-Cer average 93.62 vs. 81.17 for Qwen2.5-72B-Instruct) could be inflated by memorization rather than genuine capability. This is a load-bearing issue for the abstract's claim of 'significant margins' and must be addressed with a rigorous contamination study before the results can be interpreted.","section":"Section 3.3, Section 5.1, Table 1"},{"comment":"The baselines for the FLAME benchmark are not evaluated by the authors; Section 5.3.1 states that the paper 'presents the evaluation results of the official institute' for GPT-4o, ERNIE-4.0-Turbo-128K, GLM-4-PLUS, Qwen2.5-72B-Instruct, and XuanYuan3-70B. No details are given about the evaluation protocol, prompting format, or how these numbers were obtained, making them unverifiable and potentially inconsistent with the zero-shot protocol used for Baichuan4-Finance. The authors should either run these baselines themselves under identical conditions or provide the exact evaluation setup and release the raw outputs for verification.","section":"Section 5.3.1, Tables 1 and 2"},{"comment":"All evaluation results are reported as single accuracy numbers without error bars, confidence intervals, or significance tests. Some differences are small (e.g., Table 3, CAA: Baichuan4-Finance and GPT-4o both score 37.50; Table 2, 'Analysis and Research': Baichuan4-Finance 45.45 vs. GPT-4o 45.45), so the sweeping claim of 'significant margins' is not statistically supported. The authors should provide multiple evaluation runs (e.g., different seeds or sample-based bootstrap) and report significance or at least variance, especially for the certification subcategories with small sample sizes.","section":"Section 5.2, Tables 1-3"},{"comment":"The scaling-law model in Eq. (2) has eight free parameters per source ({E_i, A_i, B_i, C_i, alpha_i, beta_i, gamma_i, eta_i}) and is applied to n=37 sources, yielding 296 parameters. The paper does not report how many small-model runs were used to fit these parameters, nor does it validate the fitted law on held-out mixture ratios or at the actual target model scale. Without such validation, the predicted optimal mixture ratio is at risk of overfitting, and the claim that the law 'predicts the downstream performance of arbitrary data mixture ratios' is not established.","section":"Section 3.3, Eq. (2)"}],"minor_comments":[{"comment":"The row labeled 'FindPQ' should be 'FundPQ' (Fund Practitioner Qualification) to be consistent with Section 5.1 and Table 1.","section":"Table 3"},{"comment":"'Ecomonist' is a misspelling of 'Economist' in the last certification row.","section":"Table 1"},{"comment":"The sentence 'we sample 200 probabilities of the reference model for each token' is unclear; it likely means sampling 200 tokens or using a Monte Carlo estimate of the KL divergence, but the wording should be clarified.","section":"Section 3.4"},{"comment":"The pass@5 evaluation for multiple-choice questions is ambiguous: if the model generates a single answer per run, then pass@5 over five independent runs is a different metric than typical pass@k for code generation. Please define how multiple answers are aggregated and whether the benchmark protocol supports this.","section":"Section 5.2.2"},{"comment":"The radar charts and bar figures lack axis labels and legends in the text; this makes it difficult to interpret the ablation results, especially the quantitative scale of the differences in Figure 5.","section":"Figures 2-5"},{"comment":"The citation for RoPE (Su et al., 2023) is incomplete and appears as a mixture of DOI and title; please provide the full reference.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The contamination issue is the most serious concern. Given that the pre-training corpus explicitly includes examination questions and the benchmarks are public certification exams, the authors must provide a thorough contamination analysis (e.g., n-gram overlap with deduplication) and, ideally, a hold-out set of newly written exam questions. If they cannot supply such evidence, the SOTA claims should be withdrawn or heavily qualified. The paper is otherwise a useful technical report, but it currently reads more like a system card than a refereed technical contribution; the evaluation gaps (unverified baselines, no error bars) will need to be closed for publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a competent industrial write-up of continual pretraining for a finance LLM, but the headline numbers on FinanceIQ and FLAME cannot be taken at face value because the pretraining corpus explicitly includes exam questions and the paper reports no contamination check. The domain self-constraint trick is neat and the ablations are informative, so the evidence for \"significant margins\" over GPT-4o is not clean.\n\nWhat is actually new: the objective in Eq. 3-4—KL to the reference model on general data, interpolated KL plus LM loss on financial data—is a clean formulation of the standard anti-forgetting idea, and the 1B-model ablation shows it preserves general capability better than plain continual pretraining. The PPO mechanism analysis (pass@1 vs pass@5) is a nice diagnostic. The data pipeline description is detailed and plausible.\n\nThe soft spots, in proportion: the load-bearing issue is contamination. Section 3.3 lists \"examination questions\" among the 37 financial sources, and FinanceIQ and FLAME-Cer are exactly certification exams. No n-gram overlap analysis, no dedup against eval data, and the annealing phase up-weights financial sources, raising exposure further. The 12-point gap on FLAME-Cer could be memorization. This is a correctness concern, not a stylistic one.\n\nSmaller issues: FLAME baselines are quoted from an \"official institute\" without the authors' own runs; there are no error bars or significance tests; and the scaling-law fit uses eight parameters per source (296 total) with no comparison to D-CPT Law or any alternative. Individually these are minor, but together they add to the sense that the empirical claims are under-supported.\n\nWho this is for: practitioners building finance-specialized models will get a useful recipe. Academic readers should treat the performance claims as unverified.\n\nRecommendation: I would desk reject for a top venue but might send it to a workshop with a request for a contamination analysis. If the authors add a held-out exam set or overlap check, the paper becomes substantially more interesting. As is, the central empirical claim does not support peer review.","headline":"A competent engineering report whose headline financial benchmark numbers are undermined by an unaddressed contamination risk, since the pretraining corpus explicitly includes exam questions and the benchmarks are the same kind of certification exams.","tokens_in":655,"tokens_out":2324,"would_cite":false,"duration_ms":45889,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A finance-specialized LLM beats GPT-4o on Chinese certification exams by using a KL-constrained training objective that keeps its general knowledge intact.","keywords":["financial large language model","continual pre-training","domain self-constraint","scaling law data mixture","RLHF","FinanceIQ","FLAME","finance certification benchmarks"],"falsifier":"Conduct a near-duplicate search (for example, 13-gram overlap after normalization) between the FinanceIQ and FLAME-Cer question sets and the financial pre-training corpus. If a non-trivial fraction of benchmark questions appear verbatim or near-verbatim in the corpus, the accuracy numbers are invalid as a comparison of reasoning ability. A second check: compare accuracy on questions first published after the corpus was collected with accuracy on older questions; a large gap would indicate memorization of training data rather than learned financial competence.","tokens_in":12154,"feed_emoji":"💹","tokens_out":7725,"duration_ms":58820,"temperature":0.7,"pith_summary":"This paper reports the development of Baichuan4-Finance, a pair of financial large language models built by continuing pre-training a general-purpose base model on a 500-billion-token mixture of general and financial data. The central technical claim is a 'domain self-constraint' training objective: on general documents the model is trained to keep its next-token distribution close to the original base model, while on financial documents it combines that constraint with standard language-modeling loss. The authors claim this lets the base model absorb financial knowledge without losing general ability, and that the instruction-tuned chat model sets new state-of-the-art accuracy on two Chinese financial certification benchmarks, FinanceIQ and FLAME-Cer, beating GPT-4o and strong open-source models. The result matters because it offers a general recipe for injecting specialized knowledge into an LLM while preserving its general competence, which is exactly the trade-off that usually makes domain adaptation costly.","feed_headline":"Finance LLM tops certification exams, keeps general skills","feed_subtitle":"A KL-constrained continual pre-training objective adds financial knowledge without wiping out what the base model already knew.","key_machinery":"The central object is the domain self-constraint continual pre-training objective. For a training document $x$, if the document is general data, the loss is the per-token KL divergence between the new model's next-token distribution $P_{\\theta_{\\mathrm{fin}}}(x_t \\mid x_{<t})$ and the reference base model's distribution $P_{\\theta_{\\mathrm{ref}}}(x_t \\mid x_{<t})$; if the document is financial data, the loss is $\\alpha \\cdot \\mathrm{KL} + \\mathcal{L}_{\\mathrm{lm}}$, where $\\mathcal{L}_{\\mathrm{lm}}$ is the standard negative log-likelihood. This lets the model memorize and reason about financial material while actively resisting drift of the general-language distribution. A second piece of machinery is the two-stage scaling-law procedure (D-CPT Law plus a loss-to-accuracy mapping) used to pick the 37-source financial data mixture ratio at limited training cost.","core_discovery":"On the paper's own terms, the discovery is that continual pre-training for a specialized domain does not have to trade away general knowledge. The proposed domain self-constraint objective applies a KL-divergence penalty between the fine-tuned model and the frozen reference base model on general-data documents, and applies the same penalty together with the ordinary log-likelihood loss on financial documents. Trained this way, Baichuan4-Finance-Base outperforms nearly all baselines on the FinanceIQ and FLAME financial benchmarks while scoring comparably to its backbone on C-Eval, CMMLU, MMLU, GSM8k-ZH, and HumanEval. After supervised fine-tuning and a reward-robust PPO alignment stage, the chat model Baichuan4-Finance reaches 93.62 average accuracy on FLAME-Cer and 79.23 average on FinanceIQ, both above GPT-4o and the open-source models evaluated.","pith_inferences":["If the KL-constrained objective works as described, it should transfer to other high-stakes domains like medicine or law, where preserving general language competence during domain adaptation is equally important; a testable extension would be applying the same objective to an English legal or medical corpus and measuring general benchmarks.","The reported accuracy gap between Baichuan4-Finance and GPT-4o is on Chinese-language certification benchmarks; on English or multilingual financial benchmarks the ordering may differ, since the training corpus is heavily Chinese.","The contamination concern cuts both ways: a near-duplicate scan between the financial corpus and the exam benchmarks would settle whether the gains reflect genuine financial reasoning or memorization, and the paper does not report such a scan.","The scaling-law mixture selection was tuned for validation loss and benchmark accuracy; it might be extended to predict other properties such as calibration or robustness, which matter for real financial deployment."],"forward_implications":["Domain-specific continual pre-training can be done with a KL constraint on general data and a mixed objective on domain data, reducing the catastrophic-forgetting penalty that normally comes with adding specialized knowledge.","The reported results on FinanceIQ and FLAME-Cer indicate that a chat model trained this way can outperform general-purpose proprietary and open-source models on Chinese financial certification exams.","The scaling-law-based data mixture selection provides a cost-effective way to set the training data ratio for many domain sources, rather than relying on heuristic weights.","The PPO ablation suggests that reinforcement learning moves the model's per-inference hit rate toward what the model can already achieve with multiple samples, rather than adding entirely new knowledge.","The FLAME-Sce results suggest the same recipe carries over to real-world financial applications such as compliance, document generation, and risk control, not just exam questions."],"supporting_citations":[{"why":"Provides the D-CPT Law that the paper uses to predict validation loss for candidate data mixture ratios and model sizes.","marker":"(Que et al., 2024)"},{"why":"Supplies the PPO algorithm that both inspires the domain self-constraint objective and is used in the RLHF alignment phase.","marker":"(Schulman et al., 2017)"},{"why":"Supplies the reward-robust RLHF framework used to stabilize PPO training for the chat model.","marker":"(Yan et al., 2024)"},{"why":"Provides the data de-duplication and annealing practices the continual pre-training pipeline follows.","marker":"(Dubey et al., 2024)"},{"why":"Provides the preference-data construction and reward-model loss used in the alignment stage.","marker":"(Ouyang et al., 2022)"},{"why":"Supplies the AI-feedback generation and verification method used to build preference pairs for uniquely-answerable finance tasks.","marker":"(Cui et al., 2024; Li et al., 2024b)"},{"why":"Supplies MMLU and CMMLU as the general-knowledge benchmarks used to check that domain training does not regress general abilities.","marker":"(Hendrycks et al., 2020)"},{"why":"Supplies GSM8k, the math benchmark translated to Chinese and used as a general capability check.","marker":"(Cobbe et al., 2021)"}],"fun_headline_variants":["Finance LLM passes exams, keeps general skills","KL trick adds finance knowledge without forgetting","Baichuan4-Finance: top marks, no skill loss","Domain self-constraint: finance edge, general stable","New training preserves LLM smarts while adding finance"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The reported accuracy comparisons are only meaningful if the FinanceIQ and FLAME-Cer exam questions were not part of the 100-billion-token financial pre-training corpus, even though the corpus is described as containing examination questions.","fun_headline_variants_meta":{"raw":{"variants":["Finance LLM passes exams, keeps general skills","KL trick adds finance knowledge without forgetting","Baichuan4-Finance: top marks, no skill loss","Domain self-constraint: finance edge, general stable","New training preserves LLM smarts while adding finance"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000247,"raw_usage":{"total_tokens":1570,"prompt_tokens":998,"completion_tokens":572,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":614,"completion_tokens_details":{"reasoning_tokens":495}},"tokens_in":614,"tokens_out":572,"duration_ms":5819,"temperature":1.0,"reasoning_tokens":495,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T13:51:47.003258+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Conduct a near-duplicate search (for example, 13-gram overlap after normalization) between the FinanceIQ and FLAME-Cer question sets and the financial pre-training corpus. If a non-trivial fraction of benchmark questions appear verbatim or near-verbatim in the corpus, the accuracy numbers are invalid as a comparison of reasoning ability. A second check: compare accuracy on questions first published after the corpus was collected with accuracy on older questions; a large gap would indicate memorization of training data rather than learned financial competence.","supporting_citations":[],"review_version":1}