{"id":"be5bdce3-4c1d-48dc-a75a-6b65d9cec112","arxiv_id":"2505.10182","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"Continual pretraining on texts augmented with LLM-generated \"hidden thoughts\" improves Gemma2-9B's MMLU accuracy more than standard continual pretraining, with the largest gains on harder questions and across domains.","lead":"The paper trains Gemma2-9B on synthetic \"hidden thoughts\" added to STEM and law texts, then compares this \"Reasoning CPT\" against standard continual pretraining on MMLU. It reports consistent gains, especially on hard questions, and suggests reasoning skills learned in one domain transfer to others.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"MMLU gains may reflect test-time prompt alignment: the 2-shot evaluation prompt itself contains hidden thoughts, so Reasoning CPT is measured in its native format while standard CPT is not.","rationale":"The paper is a plausible, internally consistent empirical study, and the reader's conditional acceptance is reasonable. My stress-test singles out the evaluation-prompt confound as the most load-bearing issue: the central comparison uses a test-time format that the treatment specifically trains, while the control does not. The paper does provide useful controls—Figure 4 matches cumulative token count, and Appendix B shows loss trends—but neither controls for the fact that the 2-shot MMLU prompt itself demonstrates hidden thoughts. I agree with the reader that hidden-thought faithfulness is a key assumption, but I see the evaluation-prompt alignment as the sharper threat to the central quantitative claim because it offers a direct mechanism for the across-the-board gains, the cross-domain transfer, and the difficulty-dependent length pattern. The proposed concrete test is inexpensive and decisive: if the advantage persists under a neutral prompt, the prompt-alignment concern is rebutted; if it collapses, the paper's main claim needs to be weakened or reformulated. Secondary concerns (single base model, no seeds, no direct comparison to Ruan et al.) are real but less central. The paper deserves a conditional venue with the additional experiment requested, so I do not change the reader's verdict.","tokens_in":25218,"tokens_out":11158,"duration_ms":121154,"concrete_test":"Re-run the MMLU evaluations in Tables 1 and 2 with the same checkpoints but with a standard 2-shot prompt that omits hidden thoughts from the exemplars (e.g., keep only 'Problem ... Options ...' and the correct answer, per Appendix E.2 but without thought tags), keeping all sampling settings identical. If the Reasoning CPT minus standard CPT gap on overall MMLU, or on the Very Hard bin, shrinks materially or loses its monotonic difficulty trend, the headline gains are attributable to test-time prompt alignment rather than to continual pretraining on reconstructed thoughts.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline result (Tables 1 and 2) compares models under a 2-shot MMLU prompt whose exemplars contain <start_of_thought> hidden thoughts (Appendix E.2). This is the same thought-tag inference format that Reasoning CPT was trained on via Eq. (2), whereas standard CPT was trained only on raw text. The measured advantage may therefore reflect fluency with the thought-tag format rather than the reasoning content of the synthetic thoughts. This alternative can explain all three headline findings: universal gains (the format is domain-general), cross-domain transfer (the same style transfers across subjects), and difficulty-adaptive length (the training corpus has a positive original-text-length to thought-length correlation, Figure 6, which can teach a length heuristic rather than true difficulty calibration). Figure 4's token-count comparison rules out token budget but not format familiarity, because it uses the same hidden-thought evaluation prompt. Consistent with this concern, Appendix A (Table 3) shows that the prompt format alone shifts base Gemma2-9B's GSM8k accuracy by about 7 points, indicating the 2-shot format is not a neutral evaluation condition. A standard-prompt evaluation or a filler-thought training control is needed to separate the training-data effect from the test-time format effect.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces Reasoning CPT, a continual-pretraining method in which each original text is prepended with LLM-generated 'hidden thoughts' produced by Gemma2-9B-it from the source text. The authors compare Reasoning CPT with standard CPT on the same 150k-example STEM and Law corpora, using Gemma2-9B + LoRA, and evaluate on MMLU plus GSM8k diversity. They report consistent MMLU gains over standard CPT, larger gains on difficult questions, cross-domain transfer, adaptive thought length, and improved Pass@k. The central claim is that learning synthetic reconstructions of latent author reasoning is more effective than learning the texts alone.","tokens_in":25479,"tokens_out":5359,"duration_ms":53138,"significance":"If the central claim survives, the paper is a useful empirical contribution because it proposes a reward-free way to create reasoning-oriented training data from abundant text. The presentation is transparent: exact prompts, data sizes, hyperparameters, loss curves, token-matched comparisons, and concrete examples are included. The matched-token comparison in Figure 4 is a genuine attempt to control for the extra-token confound, and the Pass@k analysis in Section 5 is a meaningful check on output diversity. However, all headline results rest on a single base model, a single benchmark, and a single evaluation prompt, and the main comparison is vulnerable to a format-familiarity confound. The contribution is therefore promising but not yet established at the level claimed.","major_comments":[{"comment":"The skeptical concern about evaluation-prompt format is real. Section 3.1 states that the few-shot prompts follow Ruan et al. (2025) and include hidden thoughts enclosed by thought tags, and Appendix E.2 confirms that the MMLU evaluation prompt contains <start_of_thought> and <end_of_thought> exemplars. Equation (2) trains Reasoning CPT on exactly this tag structure, whereas standard CPT is trained on raw text without tags. Consequently, the 1.4-8 point gaps in Tables 1 and 2 may measure fluency with the hidden-thought prompt format rather than the reasoning content of the synthetic thoughts. This is not a neutral evaluation condition: Appendix A, Table 3 shows that on GSM8k the base model improves from 58.3 to 65.4 when the hidden-thought style is used. The authors should evaluate all models with a standard CoT prompt, or train a control that inserts syntactically similar but content-free filler between the same thought tags, before the central claim can be attributed to hidden-thought content.","section":"§3.1, Appendix E.2, Table 3"},{"comment":"All accuracy numbers are single-run point estimates from one base model and one training configuration. The claimed advantages are mostly 1-3 points overall and about 8 points on Very Hard questions; without variance estimates or multiple seeds it is impossible to tell whether the smaller gaps, such as 68.1 vs 66.7 for Law overall, are reliable. The authors should report standard deviations across at least three seeds, or otherwise bound the noise, before claiming that Reasoning CPT 'consistently outperforms' standard CPT.","section":"Tables 1-2, Figure 4"},{"comment":"The quality of the generated hidden thoughts is asserted rather than measured. The thoughts are produced by Gemma2-9B-it, the instruction-tuned sibling of the base model, with no human evaluation, no validation against the source text's actual reasoning, and no ablation against an alternative augmentation such as Ruan et al.'s background-knowledge style. Because the whole method hinges on the reconstructed thoughts being faithful and useful, the paper needs either a validation study or a control condition that keeps the thought-tag format constant while varying only the content.","section":"§2.2"}],"minor_comments":[{"comment":"The abstract says 'gains of up to 8 points on the most challenging problems,' which is the gap versus standard CPT, but the table also shows 10.5-11.2 point gains versus the base model; the basis for the '8 points' should be stated explicitly to avoid ambiguity.","section":"Abstract and Table 2"},{"comment":"The paper cites Gemma2-9B and Gemma2-9B-it as [Shao et al., 2024], but that reference is DeepSeekMath; the Gemma model card or the correct Google model citation should be used.","section":"§3.1 and references"},{"comment":"The MMLU prompt in Appendix E.2 appears to have missing or misplaced <start_of_thought> tags in the first two exemplars, which makes the exact evaluation condition ambiguous; the prompt should be reproduced exactly as used.","section":"Appendix E.2"},{"comment":"The difficulty labels come from GPT-4o with no validation or inter-annotator agreement; reporting a secondary labeling or at least a sample of human-checked labels would strengthen the difficulty-ordering claim.","section":"§3.2, Table 2"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is worth a revision rather than rejection. The central confound is fixable with a standard-prompt evaluation and a filler-thought training control, and the remaining concerns are about statistical robustness and validation of the generated thoughts. The paper would also be strengthened by releasing code and data, though that is not a requirement for the recommendation."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The headline result is likely inflated by an evaluation-prompt confound. The MMLU 2-shot prompt in Appendix E.2 includes hidden thoughts in the exemplars, which is exactly the format Reasoning CPT was trained on. Standard CPT was not. That gives Reasoning CPT a home-field advantage, and the stress-test alternative—format fluency rather than reasoning content—explains the universal gains, cross-domain transfer, and the length heuristic pattern. The paper's own Appendix A shows the format alone moves base Gemma2-9B's GSM8k accuracy by about 7 points (5-shot CoT 58.3 vs 1-shot hidden thoughts 65.4), so the evaluation condition is not neutral.\n\nWhat's actually new and good: the paper extends latent-thought CPT to a law corpus, compares against standard CPT at matched token counts in Figure 4, and adds difficulty-level, reasoning-length, and Pass@k analyses. Section 4's correlation analysis is honest, even if it suggests the \"adaptive reasoning depth\" is a learned length heuristic. The experimental design is otherwise clear, and the authors cite the closest prior work (Ruan et al. 2025).\n\nSoft spots beyond the eval prompt: no seeds or error bars, a single base model (Gemma2-9B), a single main benchmark, and no direct comparison to Ruan et al.'s method. These are secondary. The primary issue is the test-time prompt, not the training data. The matched-token-count result doesn't rule out format familiarity, because both curves use the hidden-thought prompt.\n\nVerdict: the central claim is plausible but not established. I'd send this to peer review, but only with a request for control conditions: a standard-prompt evaluation, a filler-thought training control (thought tags with unstructured text), multiple seeds, and ideally a direct run of Ruan et al. If the advantage survives a standard-prompt eval, it's a real result. Until then, I wouldn't cite it as evidence for latent-thought training.\n\nWorth taking to reading group as a clean case study of evaluation confounds, though. If I were the area chair, this would be a borderline-major-revision paper.","headline":"The MMLU evaluation prompt gives Reasoning CPT a home-field advantage, so the headline gains over standard CPT are not yet trustworthy.","tokens_in":25992,"tokens_out":3511,"would_cite":false,"duration_ms":33096,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Prepending LLM-generated hidden thoughts to expert texts during continual pretraining improves reasoning on MMLU, with the largest gains — about 8 points over standard CPT — on the hardest questions.","keywords":["continual pretraining","hidden thoughts","synthetic data","LLM reasoning","MMLU","cross-domain transfer","chain-of-thought","reasoning efficiency"],"falsifier":"A control that replaces the hidden-thought segment with same-length, same-format filler thoughts (or shuffled thoughts from other texts) would settle the claim: if it matches Reasoning CPT's gains, the reasoning content is not what drives the improvement. A second check would have domain experts judge whether the generated thoughts recover actual omitted reasoning steps from the source texts; low fidelity would undermine the proposed mechanism.","tokens_in":25026,"feed_emoji":"💭","tokens_out":5787,"duration_ms":50451,"temperature":0.7,"pith_summary":"This paper tries to show that a language model can learn to reason better by continuing to pretrain on synthetic text in which an LLM-generated 'hidden thought' is placed before an expert passage, on the premise that every text is the residue of its author's thinking. On MMLU, this Reasoning CPT beats standard continual pretraining on the same passages by about 1.4–1.8 points overall and by roughly 8 points on the hardest questions, with up to 3.3 points over the base model. The gains appear in domains never seen during training, and the trained models spontaneously spend fewer tokens on easy questions and more on hard ones. If the effect is real, it offers a way to build reasoning data from ordinary expert text without task-specific rewards.","feed_headline":"Synthetic hidden thoughts lift MMLU by up to 8 points","feed_subtitle":"Reasoning-style synthetic data beats text-only continual pretraining, with the biggest edge on the hardest questions.","key_machinery":"The central object is the training sequence $X = \\langle\\text{start\\_of\\_thought}\\rangle \\oplus H \\oplus \\langle\\text{end\\_of\\_thought}\\rangle \\oplus S$, where $H$ is an LLM-generated hidden-thought segment and $S$ is the original expert text, trained with the standard autoregressive next-token loss. Hidden thoughts are produced by Gemma2-9B-it under a prompt that elicits goal setting, background-knowledge recall, decision-making, and self-verification, so ordinary STEM and legal passages become explicit reasoning traces. The mechanism also includes a corpus-level correlation between original-text length and hidden-thought length (Spearman $\\rho=0.348$ for STEM, $\\rho=0.486$ for Law), which the paper identifies as the plausible driver of the models' difficulty-adaptive thinking length.","core_discovery":"Continual pretraining on synthetic sequences that prepend an LLM-generated hidden thought to an expert text — called Reasoning CPT — improves MMLU accuracy more than standard continual pretraining on the same texts, and the gap widens with problem difficulty. Reasoning CPT trained on STEM text reaches 69.1% overall versus 67.3% for standard CPT; trained on legal text it reaches 68.1% versus 66.7%. On Very Hard questions the advantage over standard CPT is about 8 points in both domains, with gains of 10.5–11.2 points over the base model. The skill transfers across domains: a model trained on legal hidden thoughts improves MMLU-STEM by 4.3 points, and models trained with hidden thoughts generate shorter thinking on easy questions and longer thinking on hard ones, matching the positive correlation between source-text length and thought length in the training corpus.","pith_inferences":["The adaptive-length behaviour may be a corpus-level heuristic: the training data's text-length/thought-length correlation could teach 'think until confident' rather than genuine difficulty awareness; a controlled corpus with decorrelated lengths would test this.","A confound remains: the few-shot evaluation prompts themselves include hidden-thought examples, so part of the gain could come from format familiarity rather than reasoning content; an evaluation with plain CoT prompts would separate the two.","If the mechanism generalizes, the same recipe could turn fiction, history, or scientific prose into reasoning data, extending the approach beyond reward-rich domains.","The results also suggest a cheaper alternative to RL for building reasoning: mine thoughts once with a strong generator, then continual-pretrain a base model, preserving output diversity that instruction tuning tends to narrow."],"forward_implications":["Reasoning CPT on either STEM or legal text raises MMLU across all four subject groups, not just the training domain.","The advantage over standard CPT grows with problem difficulty, reaching about 8 points on Very Hard questions in both training domains.","Models trained this way adapt reasoning length: fewer thinking tokens than CPT on easy questions and more on hard ones, with no accuracy loss on easy items.","Because the method needs no correctness labels or verifiable rewards, it can be applied to any high-quality text corpus.","The trained model retains diverse reasoning paths, improving GSM8k Pass@5 to 91.7% versus 81.2% Pass@1 for the instruction-tuned baseline."],"supporting_citations":[{"why":"Cited for the Gemma2-9B base model and Gemma2-9B-it generator; both the training target and the hidden-thought generator come from this model family.","marker":"[Shao et al., 2024]"},{"why":"OpenWebMath supplies the 150,000 STEM training texts (36.8M tokens) that both standard CPT and Reasoning CPT use.","marker":"[Paster et al., 2024]"},{"why":"The FreeLaw subset of The Pile supplies the 150,000 legal opinion texts (28.3M tokens) used as the Law corpus.","marker":"[Gao et al., 2021]"},{"why":"The MMLU benchmark is the evaluation instrument; all accuracy comparisons and difficulty-level splits are measured on it.","marker":"[Hendrycks et al., 2021]"},{"why":"The predecessor that introduced training on latent thoughts and whose few-shot prompt format with thought tags the evaluation reuses.","marker":"[Ruan et al., 2025]"},{"why":"LoRA defines the parameter-efficient training configuration (rank 64) used for all continual pretraining runs.","marker":"[Hu et al., 2022]"},{"why":"GSM8k is the math-word-problem test used to measure whether the models generate diverse reasoning paths.","marker":"[Cobbe et al., 2021]"},{"why":"Defines the Pass@k metric used to quantify diversity of solutions under repeated sampling.","marker":"[Chen et al., 2021]"}],"fun_headline_variants":["Hidden-thought CPT gains 8 points on hardest MMLU items","Synthetic thoughts transfer reasoning skills across domains","LLMs trained on hidden thoughts adapt reasoning depth","Up to 8-point MMLU edge: synthetic thoughts beat text CPT"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the LLM-generated hidden thoughts faithfully reconstruct the reasoning behind the source texts, so the measured improvements come from learning that reasoning content rather than from simply seeing more tokens in a thought-tagged format.","fun_headline_variants_meta":{"raw":{"variants":["Hidden-thought CPT gains 8 points on hardest MMLU items","Synthetic thoughts transfer reasoning skills across domains","LLMs trained on hidden thoughts adapt reasoning depth","Up to 8-point MMLU edge: synthetic thoughts beat text CPT"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00031,"raw_usage":{"total_tokens":1775,"prompt_tokens":962,"completion_tokens":813,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":578,"completion_tokens_details":{"reasoning_tokens":744}},"tokens_in":578,"tokens_out":813,"duration_ms":8177,"temperature":1.0,"reasoning_tokens":744,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T21:14:12.221826+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A control that replaces the hidden-thought segment with same-length, same-format filler thoughts (or shuffled thoughts from other texts) would settle the claim: if it matches Reasoning CPT's gains, the reasoning content is not what drives the improvement. A second check would have domain experts judge whether the generated thoughts recover actual omitted reasoning steps from the source texts; low fidelity would undermine the proposed mechanism.","supporting_citations":[{"cited_title":"Openwebmath: An open dataset of high-quality mathematical web text","cited_arxiv_id":null,"evidence_quote":"OpenWebMath supplies the 150,000 STEM training texts (36.8M tokens) that both standard CPT and Reasoning CPT use."}],"review_version":1}