{"id":"9d42cdb3-e35a-4d68-ba40-9cac12e97037","arxiv_id":"2412.19638","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A 1.2B language model trained with tensor-program hyperparameter transfer and WSD decay with SFT data mixing; the headline SOTA claim is not supported by the paper's own benchmark tables.","lead":"Xmodel-2 is a 1.2-billion-parameter language model trained with hyperparameter transfer, a WSD learning-rate schedule, and a tuned mix of reasoning data in the final decay phase. The report claims state-of-the-art small-model performance, but its own tables show other models score higher on complex reasoning and commonsense benchmarks.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper's central SOTA claim is contradicted by its own Table 3: Qwen2.5-1.5B scores 39.98 vs Xmodel-2's 39.62 on complex reasoning; Table 2 also shows lower commonsense performance.","rationale":"The reader's stated weakest assumption concerns transfer of wind-tunnel hyperparameters to the 1.2B model. That is a legitimate secondary issue, but the more load-bearing problem is internal: the paper's own Table 3 gives the complex-reasoning average to Qwen2.5-1.5B (39.98) rather than Xmodel-2 (39.62), and Table 2 shows several models above Xmodel-2 on commonsense. The SOTA statement in the abstract and Section 3 is therefore contradicted by the reported evidence, independent of any transfer assumptions. The agent-based result (Table 4) does support SOTA in that narrower category, and the release of checkpoints and code is credible, but the paper's primary claim, as worded, is false. My recommendation matches the reader's REJECT verdict, but for a different and more direct reason.","tokens_in":15072,"tokens_out":5865,"duration_ms":42624,"concrete_test":"Recompute the equal-weighted averages from Table 3's per-task numbers, and run the released Xmodel-2 and Qwen2.5-1.5B checkpoints through the same evaluation harness. If Qwen2.5-1.5B's mean remains at or above Xmodel-2's 39.62, the 'state-of-the-art in complex reasoning' claim is refuted.","verdict_should_be":"REJECT","load_bearing_attack":"The central claim (Abstract, Section 3) is that Xmodel-2 achieves state-of-the-art performance in complex reasoning and agent-based tasks among 1-2B models. Table 3 shows Xmodel-2's average across GSM8K, MATH, BBH, MMLU, HumanEval, and MBPP is 39.62, whereas Qwen2.5-1.5B scores 39.98, and on individual tasks Qwen2.5 leads on GSM8K (62.40 vs 55.88), MATH (28.28 vs 25.50), MMLU (59.72 vs 48.87), and MBPP (40.00 vs 29.20). Xmodel-2 only leads on BBH and HumanEval. In Table 2, Xmodel-2's commonsense average of 61.79 is below Phi-1.5-1.3B (65.68), DCLM-1B (63.81), Qwen2.5-1.5B (63.14), and SmolLM-1.7B (61.92). Thus the SOTA claim for complex reasoning is unsupported and, under the paper's own measurements, false. The claimed 29.31% improvement over baseline (Section 2.4) is not defined with a baseline, so it cannot be verified independently. The transfer-from-wind-tunnel assumption would matter only if the headline result were credible; here the headline result itself fails.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This technical report introduces Xmodel-2, a 1.2B-parameter decoder-only language model trained on about 1.5 trillion tokens with a Warmup-Stable-Decay learning-rate scheduler, a deep-and-thin Transformer architecture with GQA, a custom Unigram tokenizer, and embedding sharing. The authors describe wind-tunnel experiments on 6M and 54M models to select hyperparameters and SFT-data ratios for the decay phase, then evaluate the 1.2B model on commonsense reasoning, complex reasoning, and agent-interaction benchmarks. They also fit a test-time loss scaling curve on Wikitext-2. The headline claim is that Xmodel-2 achieves state-of-the-art performance in complex reasoning and agent-based tasks among 1–2B models; checkpoints and code are released publicly.","tokens_in":15417,"tokens_out":7087,"duration_ms":60836,"significance":"If the claims were accurate, the paper would be a useful contribution: it presents an open 1.2B model, a wind-tunnel hyperparameter-search methodology that could reduce HPO cost, a focused data-ratio study for the WSD decay phase, and an agent-task evaluation that is comparatively rare in small-model reports. The agent results in Table 4 are the strongest part of the empirical contribution. However, the central 'state-of-the-art' claim for complex reasoning and commonsense reasoning is directly contradicted by the paper's own Tables 2 and 3, and the key quantitative improvement claim (29.31%) is not defined with a baseline. The methodological transfer claim from 6M/54M to 1.2B is asserted rather than demonstrated. These issues are load-bearing and prevent the paper, as written, from supporting its advertised conclusions.","major_comments":[{"comment":"The claim of state-of-the-art performance in complex reasoning is contradicted by the paper's own results. In Table 3, Xmodel-2's average over GSM8K, MATH, BBH, MMLU, HumanEval, and MBPP is 39.62, while Qwen2.5-1.5B scores 39.98; Qwen2.5 also leads on GSM8K (62.40 vs 55.88), MATH (28.28 vs 25.50), MMLU (59.72 vs 48.87), and MBPP (40.00 vs 29.20). Section 3 additionally states that Xmodel-2 is SOTA 'especially in commonsense reasoning,' yet Table 2 shows Xmodel-2 at 61.79, below Phi-1.5 (65.68), DCLM-1B (63.81), Qwen2.5-1.5B (63.14), and SmolLM-1.7B (61.92). Please either revise the abstract and Section 3 to reflect the actual comparative results or provide a precisely defined alternative notion of 'state-of-the-art' with justification.","section":"Abstract, §3, Tables 2–3"},{"comment":"The sentence 'These adjustments, combined with optimized data mixing and processing, improved complex reasoning performance by 29.31% compared to our baseline' is not verifiable because the baseline configuration, the evaluation benchmarks, the prompting setup, and the measurement procedure are not specified. Without a defined baseline and without variance or error information, this number cannot be used as evidence for the efficacy of the data-ratio and mixing recipe. Please specify the exact baseline and the evaluation protocol.","section":"§2.4"},{"comment":"The training configuration of the final 1.2B model is justified solely by 'wind tunnel' experiments on 6M and 54M models. The abstract asserts 'seamless transfer of optimal configurations to larger models,' but Section 6 only states that the small-scale experiments 'confirmed the strategy's suitability'; no intermediate-scale check (e.g., 200M or 300M) or any quantitative comparison of optimal hyperparameters/data ratios across scales is provided. As written, the transfer is an assumption rather than a demonstrated result. Please provide additional transfer evidence or explicitly characterize the transfer as an assumption and soften the abstract's wording.","section":"§6 and Abstract"},{"comment":"The 'Post-training Scaling Law' is presented as a general law, but it is a curve fit to a single model on a single dataset (Wikitext-2), with no error bars, no held-out validation, no alternative functional form comparison, and no evidence that the fitted parameters extend to other models or datasets. The name overstates the finding. Please reframe this as an empirical observation about Xmodel-2 or provide evidence of generality across models and datasets.","section":"§4.2"}],"minor_comments":[{"comment":"The decay length is given as 'T is set to 5000 steps (20 billion tokens),' but with a batch size of 3.93 million tokens, 5000 steps corresponds to approximately 19.65 billion tokens; please correct the token count or explicitly state that it is approximate.","section":"§2.2"},{"comment":"In Figure 2, the legend order does not match the order in the pie charts, and the listed percentages do not sum to 100 for either panel; please clarify the data-mixture display so the composition is unambiguous.","section":"§2.4, Figure 2"},{"comment":"The fitted parameter values are written as 'a ∼ -0.575, b∼ 1.772, c∼ 32.840'; the tilde is nonstandard and ambiguous. Use '=' or report confidence intervals, and make clear that these are fitted values, not exact constants.","section":"§4.2, Eq. (1)"},{"comment":"The text says 'consistent decrease in perplexity,' but the y-axis of Figure 5 is labeled 'loss'; please make the terminology consistent.","section":"§4.2, Figure 5"},{"comment":"The HumanEval pass@1 score reported for Qwen2.5-1.5B is 5.49, which is substantially lower than Qwen2-1.5B's 20.73 and looks suspicious relative to Qwen2.5's other reasoning scores; please verify this value, as it directly affects the average comparison.","section":"Table 3"},{"comment":"The author block is formatted without clear separators between names; add commas or another unambiguous delimiter.","section":"Title page"}],"recommendation":"major_revision","confidential_remarks":"The paper is better viewed as a technical report than as a full journal article. The central SOTA claims are contradicted by the paper's own tables, which is a serious problem, but the errors are correctable by revising the claims and adding the missing baseline definitions and transfer evidence. The agent-task evaluation and open release are positive features. I would support resubmission after the authors address the major comments, especially the comparison claims and the definition of the 29.31% improvement."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know this paper overclaims its headline result. Its own Table 3 shows Qwen2.5-1.5B with a higher complex-reasoning average (39.98 vs 39.62) and Table 2 shows several models above it on commonsense. The agent-table claim is actually supported (Xmodel-2 tops Table 4), but the abstract and Section 3 claim SOTA across the board. That mismatch is the first thing anyone will notice.\n\nWhat is actually new: the paper reports a complete training recipe for a 1.2B model combining Tensor Programs (µP), the WSD scheduler, and a data-ratio search during decay. The empirical finding that an SFT mixing ratio around 60–69% in the decay phase improves complex reasoning is new and could be useful to people training small models. The model, code, and checkpoints are public, which is real value. The agent evaluation across HotpotQA, FEVER, AlfWorld, and WebShop is a useful addition; few small-model reports include that.\n\nThe soft spots are concentrated in the framing. The 29.31% improvement over baseline is never defined. The post-training scaling law is a fitted curve of loss vs token position on Wikitext-2, with no error bars, no predictive check, and a formula that is not obviously a law. The transfer from 6M/54M wind-tunnel experiments to the 1.2B model is asserted rather than demonstrated; there is no intermediate-scale validation. The data-ratio search is described only as \"over 400 trials,\" with no ablation table or variance. These are not minor omissions, but they are fixable.\n\nThe citation pattern is clean: prior work on µP, MiniCPM, and data mixing is credited appropriately, and I see no self-citation issue. The core problem is not the experiments; it is the distance between what the tables show and what the prose claims.\n\nFor peer review: I would send it out, because the released artifacts and the decay-phase data-ratio finding deserve scrutiny, and a referee could force the necessary corrections. But I would not accept it in this form. It needs a major revision that narrows the claims, defines the baseline, and presents the scaling fit as a descriptive curve. If you are choosing a small model to deploy, the checkpoint may be worth a look, but ignore the marketing.\n\nRecommendation: invite revision, not acceptance.","headline":"The paper's central SOTA claim is contradicted by its own benchmark tables, but the released model and the decay-phase data-ratio recipe have real value if the claims are rewritten.","tokens_in":15939,"tokens_out":3242,"would_cite":false,"duration_ms":29252,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Xmodel-2 reports that a 1.2-billion-parameter model can lead the 1-2B class on complex-reasoning and agent benchmarks by transferring hyperparameters from tiny wind-tunnel models and tuning the SFT data ratio in the WSD decay phase.","keywords":["small language models","1.2B parameter model","maximal update parametrization","hyperparameter transfer","WSD learning rate scheduler","data ratio optimization","complex reasoning","agent-based tasks"],"falsifier":"Train a mid-size model, for example 300M parameters, using the wind-tunnel-optimal settings and compare the 64 percent SFT decay mix against, say, a 50 percent mix; if the 64 percent mix does not win or the optimal ratio falls outside the 60-69 percent band, the transfer assumption that carries the paper's central claim would be falsified.","tokens_in":14880,"feed_emoji":"🧠","tokens_out":7398,"duration_ms":59009,"temperature":0.7,"pith_summary":"Xmodel-2 is a 1.2-billion-parameter language model built for reasoning tasks, and the paper's central claim is that such a small model can beat other 1-2B models on complex reasoning and agent benchmarks. The route to that result is a training recipe: use the maximal update parametrization so that optimal hyperparameters found on 6M- and 54M-parameter wind-tunnel models transfer unchanged to the 1.2B model, then anneal the learning rate with the Warmup-Stable-Decay scheduler while mixing 64 percent supervised fine-tuning data into the decay phase. The paper reports that this combination improves complex-reasoning performance by 29.31 percent over its baseline and gives top agent-task scores in the 1-2B size class. If the recipe is right, expensive large-scale hyperparameter sweeps can be replaced by cheap small-model experiments, lowering the cost of building capable small reasoners.","feed_headline":"A 1.2B model leads 1-2B models on reasoning and agent tasks","feed_subtitle":"Tiny 6M and 54M experiments set the hyperparameters; a 64 percent SFT mix in the decay phase drives the gains.","key_machinery":"The load-bearing mechanism is the maximal update parametrization (muP), a scaling scheme from the Tensor Programs line of work that fixes how weights are initialized and how learning rates and output scales change with width and depth, so that the same hyperparameters stay near-optimal across model sizes. It is what makes the 6M and 54M wind-tunnel experiments trustworthy for the 1.2B model. The second mechanism is the Warmup-Stable-Decay (WSD) learning-rate schedule, whose final decay phase is repurposed as a supervised fine-tuning stage: the paper anneals the learning rate exponentially while mixing SFT data into the pretraining stream, with the SFT share searched and set to 64 percent.","core_discovery":"On its own terms, the paper establishes that a 1.2B decoder-only model pretrained on 1.5 trillion tokens can reach state-of-the-art complex reasoning and agent performance among 1-2B parameter models. The discovery is not a new architecture but a transferable training configuration: maximal update parametrization makes the optimal learning rate, embedding scaling, and depth scaling stable across model sizes, so a 300-configuration Bayesian search on the 6M model, as opposed to a 570,000-configuration grid, yields settings that carry to the 1.2B model. The second half of the recipe is the decay phase of the WSD scheduler, where the paper mixes supervised fine-tuning data with pretraining data and searches the SFT ratio, finding the optimum between 60 and 69 percent and choosing 64 percent. This data-ratio search, combined with rule-based prompt diversification and SimHash deduplication, accounts for the reported 29.31 percent gain over the baseline in complex reasoning. The authors further claim that the same model is state of the art on ReAct-based agent tasks, including HotpotQA, FEVER, AlfWorld, and WebShop, among 1-2B models.","pith_inferences":["The reported 29.31 percent gain bundles several interventions, including the SFT ratio, chain-of-thought inclusion, prompt diversification, and SimHash deduplication, so an ablation isolating the SFT ratio alone would reveal which ingredient actually drives the improvement.","If the wind-tunnel transfer holds beyond this model family, the same 6M and 54M search could be used to tune decay-phase data mixes for other small-model releases, making the recipe a cheap template rather than a one-off result.","The fitted post-training power law on Wikitext-2 implies that extending context at test time lowers loss with diminishing returns; whether that curve has the same shape on math, code, or agent trajectories is an untested extension."],"forward_implications":["Hyperparameter and data-ratio search can be done mostly on 6M and 54M models, cutting the search cost from hundreds of thousands of grid configurations to a few hundred Bayesian trials.","The optimal SFT share in the WSD decay phase is a narrow band around 60-69 percent, so practitioners can fix the overall ratio and search only the internal category mix.","Chain-of-thought data placed under logic helps complex reasoning, and instruction-formatted math and code data beats pretraining-format data in the decay phase.","A 1.2B model with these settings reaches scores competitive with 1.5-1.8B baselines, suggesting strong reasoning ability does not require a large parameter count."],"supporting_citations":[{"why":"Supplies the maximal update parametrization and zero-shot hyperparameter transfer that the wind-tunnel experiments rely on.","marker":"[Yang et al., 2022]"},{"why":"Extends Tensor Programs to feature learning in deep networks, backing the claim that hyperparameters stay stable across depth and width.","marker":"[Yang et al., 2023]"},{"why":"Provides the WSD learning-rate scheduler and decay-stage SFT mixing that form the second half of the training recipe.","marker":"[Hu et al., 2024]"},{"why":"Prior evidence that small-model data-ratio experiments transfer to larger models, which the paper extends to WSD decay.","marker":"[Ye et al., 2024]"},{"why":"Defines the Llama-2-style architecture that Xmodel-2 modifies with grouped-query attention and embedding sharing.","marker":"[Touvron et al., 2023]"}],"fun_headline_variants":["1.2B model beats larger rivals on reasoning and agent tasks","Tiny 6M model's hyperparameter search sets 1.2B to SOTA","64% SFT mix in decay phase drives 1.2B to SOTA reasoning","Hyperparameter transfer from 6M to 1.2B yields SOTA on reasoning","1.2B model, 1.5T tokens, WSD scheduler: recipe for SOTA agents"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire recipe rests on the assumption that hyperparameters and data ratios found optimal on 6M- and 54M-parameter wind-tunnel models remain optimal for the 1.2B model without any validation at an intermediate scale.","fun_headline_variants_meta":{"raw":{"variants":["1.2B model beats larger rivals on reasoning and agent tasks","Tiny 6M model's hyperparameter search sets 1.2B to SOTA","64% SFT mix in decay phase drives 1.2B to SOTA reasoning","Hyperparameter transfer from 6M to 1.2B yields SOTA on reasoning","1.2B model, 1.5T tokens, WSD scheduler: recipe for SOTA agents"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000582,"raw_usage":{"total_tokens":2732,"prompt_tokens":931,"completion_tokens":1801,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":547,"completion_tokens_details":{"reasoning_tokens":1685}},"tokens_in":547,"tokens_out":1801,"duration_ms":12574,"temperature":1.0,"reasoning_tokens":1685,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T00:07:29.229750+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train a mid-size model, for example 300M parameters, using the wind-tunnel-optimal settings and compare the 64 percent SFT decay mix against, say, a 50 percent mix; if the 64 percent mix does not win or the optimal ratio falls outside the 60-69 percent band, the transfer assumption that carries the paper's central claim would be falsified.","supporting_citations":[],"review_version":1}