{"id":"f2ac95a2-fbc8-4aff-934e-5d9a87bd6cf1","arxiv_id":"2607.18228","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Learned soft prefixes reliably flip correct syllogistic judgments in LLMs, transferring across unseen forms and interfaces and behaving mainly as a broad answer preference rather than a transferable logical operation.","lead":"By prepending small trainable 'soft prefix' vectors to prompts, this paper shows that large language models can be pushed to change correct answers to syllogistic logic problems—even on logical forms and wordings the prefixes never saw during training. The work offers a stress-test method for measuring how stable a model's logical judgments are under learned context.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Generality of 'unseen forms' claim rests on only 24 minority forms total; pooled intervals across all folds may place lower bound near chance.","rationale":"The reader identified the same load-bearing concern: the central transfer claim rests on only six targeted logical forms per condition, with wide bootstrap intervals and an explicit paper caveat about population precision. My read agrees; the repeated-split experiment partially mitigates the original six-form dependence by rotating all 24 minority forms, but those 24 forms remain the entire minority-class population, and the effect size varies substantially across folds. This means the quantitative magnitude of the 'unseen forms' effect is uncertain, even though the qualitative comparison to random controls is robust in all 16 splits. The paper is transparent about this limitation, and the verdict CONDITIONAL is appropriate: the claim likely has merit, but a precise population estimate requires either a larger independent form space or a pooled confidence interval that accounts for form-level clustering. No internal inconsistency or methodological fraud is apparent; the concern is statistical precision and external validity. Therefore, the reader's verdict remains unchanged, with this concern reinforcing the conditional acceptance.","tokens_in":27500,"tokens_out":10919,"duration_ms":95519,"concrete_test":"Pool the four repeated-split folds to obtain a 95% cluster-bootstrap confidence interval over all 24 minority logical forms for each model–direction condition (e.g., resample complete forms from the union of test folds). If the lower bound for any headline condition falls below 50% — the chance level for binary flip — the 'redirect many correct answers' claim on unseen forms is not established at the population level. As a complementary check, evaluate the pretrained prefixes on a fresh set of syllogistic forms from a different benchmark or a larger enumeration (e.g., adding a fourth term) to test whether transfer extends beyond the 24 minority forms.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract claims learned prefixes 'remain effective across unseen forms,' but the headline evidence uses only six targeted logical forms per condition (Section 4.4, Table 1). The repeated-split experiment (Section 4.6, Table 4) rotates all 24 minority forms through test, which strengthens the comparison to random controls (learned > random in all 16 cases), but these 24 forms are the entire minority-class population of the 256-form benchmark, not a sample from a larger space. Flip rates vary widely across folds — e.g., Qwen3.6 valid→invalid ranges from 56.4% to 93.6% (Table 4) — and the paper itself concedes 'six targeted forms are too few for a precise population estimate' (Section 4.4). A pooled estimate over all 24 forms would have a confidence interval whose lower bound could approach chance for the weaker conditions, so the quantitative '72–90%' flip range in the abstract is not a stable population estimate. The qualitative robustness to random controls is solid, but the claim that the effect generalizes across 'unseen forms' is only demonstrated for a small, fixed set of minority forms.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces a diagnostic framework that prepends learned soft prefixes to an exactly labeled syllogistic reasoning benchmark, with model weights frozen, and characterizes the resulting changes through controls (randomized A/B mappings, norm-matched random prefixes, best-of-1000 random search, neutral-text prefixes, rephrased prompts, reverse directions, cross-task transfer, score-model fits, and activation patching). The central claim is that learned prefixes can redirect many previously correct judgments on held-out logical forms and across wording/prompt changes in Qwen3.6 MoE, Qwen3-8B, and Gemma 4 31B, and that the dominant mechanism is a broad preference for one answer meaning rather than fixed-symbol forcing or a transferable logical operation. The paper further claims that simple score models predict Gemma's prefixed margins better than Qwen's, while often predicting Qwen's binary flips without predicting their final margins.","tokens_in":27744,"tokens_out":6072,"duration_ms":59703,"significance":"If the claims hold, this is a valuable controlled methodology for measuring the stability of formal reasoning under optimized continuous context. The paper's strengths are substantial: exact labels computed from explicit semantics, held-out logical forms, four repeated split rotations, paired random controls, a best-of-1000 random control, per-form bootstrap intervals, score-model analyses fitted only on development data, and unusually candid limitation statements. The qualitative result—learned prefixes strongly outperform matched random prefixes—appears well supported. The main weakness is that the quantitative generalization claims rest on a small, fixed set of minority forms, and the paper's own tables show high split-dependence. The manuscript is methodologically careful but overstates the precision of its 'unseen forms' transfer estimate.","major_comments":[{"comment":"The central 'unseen forms' claim is quantified from only six targeted logical forms per condition. The paper acknowledges this in §4.4, but the abstract still presents '72–90%' and 'remain effective across unseen forms' as a model-level result. The form-level bootstrap for Qwen3.6 valid→invalid rephrased is 72.0% [56.1, 83.6], and Table 4 shows learned flip varying from 56.4% to 93.6% across splits. The 16/16 comparison against random controls in §4.6 is solid evidence of a qualitative effect, but it does not support a stable population estimate or a general quantitative range. Please report split-level ranges and/or a pooled form-bootstrap interval over all 24 minority forms, and state the conclusion as an effect in the tested minority forms rather than a population-level transfer rate.","section":"§4.4, Table 1; §4.6, Table 4; Abstract"},{"comment":"The 'new wording, random strings' condition is presented as an interface change, but the benchmark always contains random-string rendering as one of its four wording styles (Appendix A), and prefix training uses four wording styles (§D.2, Table 18). If random-string renderings are in the training distribution, this condition tests unseen logical forms in a familiar surface style, not an unseen wording interface. Please state explicitly whether the tested prefixes were trained on random-string renderings. If they were, revise 'new wording' and 'interface changes' in the abstract and §4.4; if separate prefixes were trained per wording style, this should be described in the main text. The rephrased-prompt conditions remain the only clear novel-interface transfer.","section":"§4.4, Tables 1–2 vs. Appendix A, §D.2"}],"minor_comments":[{"comment":"The abstract says Qwen3.6 MoE flip rates 'remain between 72% and 90%', but Table 1's lower bound is 72.0% and Table 4's repeated-split range goes down to 56.4%. Please attach a scope qualifier such as 'in the fixed six-form conditions' to avoid implying a stable cross-split estimate.","section":"Abstract, §4.4"},{"comment":"Panels A and B use a symmetric logarithmic scale, which cannot display non-positive R² values in the usual way. Please state explicitly how negative R² values are represented or clipped; several reported values are negative.","section":"Figure 2"},{"comment":"The neutral-text control is reported for Qwen3.6 only. The main text says 'the tested neutral phrases' without noting this scope limitation; please add the model restriction in §4.5 or in the table caption.","section":"Appendix K, §4.5"},{"comment":"Target-answer rates are computed over different denominators ('eligible rows' vs. 'all rows'). The main text and captions distinguish these, but the distinction is easy to miss. Consider labeling the columns more explicitly, e.g., 'target-answer rate, eligible rows' and 'target-answer rate, all rows' in both tables.","section":"Tables 2 and 14"}],"recommendation":"major_revision","confidential_remarks":"This is a careful empirical study with unusually strong controls, and I do not see a destructive flaw in the core comparison. My recommendation is driven by the gap between the abstract's quantitative generalization claims and the small number of tested logical forms, plus the ambiguity about whether random-string wording is a genuinely unseen interface. Both are fixable with reframing and more precise reporting; the underlying methodology is sound."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a solid empirical paper, worth a serious referee. The core claim—optimized soft prefixes can flip correct syllogistic answers on held-out logical forms, far beyond matched random controls, and the effect is dominated by a broad answer preference rather than fixed-symbol forcing—is supported by a genuinely careful control design. The paper is also unusually explicit about its own limitations, which is rare and welcome.\n\nWhat is actually new: not the method (soft-prefix optimization and universal triggers are known), but the diagnostic use in an exactly labeled formal domain, with splits by logical form, randomized A/B mappings, norm-matched random prefixes, a best-of-1000 random search, reverse directions, and activation patching. The cross-model difference in score predictability—Gemma's answer margins are well approximated by simple score models, Qwen's are not—is a real finding, even if the interpretation is behavioral rather than mechanistic.\n\nWhat it does well: the controls are the strong part. A/B randomization rules out fixed-letter forcing. Random prefix controls, including the selected best-of-1000, rule out generic sensitivity. The repeated-split experiment rotating all 24 minority forms through test is a meaningful robustness check, and the learned-over-random advantage holds in all 16 comparisons. The paper also flags the small-form-count problem directly: six targeted forms are too few for a precise population estimate, and the bootstrap intervals show it.\n\nSoft spots, in proportion: the stress-test note is right that the 'unseen forms' generality is quantitatively brittle. The 24 minority forms are the entire minority class in the benchmark, not a sample from a larger space, and flip rates vary widely across folds—Qwen3.6 valid-to-invalid ranges from 56.4% to 93.6%. A pooled estimate could easily put lower bounds near chance for the weaker conditions. So the abstract's 72–90% range should not be read as a stable population quantity. That said, the qualitative claim—learned prefixes beat matched random prefixes consistently—does not depend on the precise point estimate, so the central argument survives. The other real caveat is that the reproducibility artifacts are promised but not yet available (Appendix P). That is addressable, not fatal.\n\nWho it is for: people working on LLM robustness, safety evaluation, and reasoning benchmarks. It deserves peer review, not desk rejection. My recommendation: send it, but ask the authors to add pooled estimates over all 24 forms with explicit lower bounds, and to release the artifacts before acceptance.","headline":"A careful, unusually honest empirical study showing learned soft prefixes override correct syllogistic judgments, with the main caveat that the quantitative 'unseen forms' range is less stable than the abstract suggests.","tokens_in":706,"tokens_out":921,"would_cite":true,"duration_ms":24267,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A short trainable embedding sequence prepended to a prompt can flip a large majority of a frozen LLM's correct syllogistic judgments, and the effect survives unseen forms, new wordings, and prompt changes.","keywords":["soft prefix","syllogistic reasoning","logical robustness","answer bias","prefix tuning","activation patching","large language models","transfer across forms"],"falsifier":"Evaluate the same training setup with a test split drawn from all 256 logical forms (e.g., a 50-form random holdout rather than the fixed six-form group). If flip rates on the broad holdout fall to within the random-prefix control range, the paper's claim that steering survives unseen forms is falsified for that broader space.","tokens_in":27378,"feed_emoji":"🧠","tokens_out":5770,"duration_ms":50487,"temperature":0.7,"pith_summary":"The paper asks whether optimized continuous context can override formal reasoning that a model already does correctly. It finds that a short trainable soft prefix, added while the model's weights stay frozen, redirects most targeted correct answers on a syllogistic benchmark, in three different model families, and keeps working on logical forms never seen during prefix training as well as under new wordings and prompt phrasings. The effect is dominated by a broad preference for one answer meaning (e.g., 'satisfiable' or 'invalid') rather than a fixed output symbol or a transferable logical operation, since randomized letter mappings are overcome and cross-task transfer is weak. The paper further shows that similar aggregate flip rates can hide model-specific behavior: simple score-based transformations approximate Gemma's answer margins well but predict which Qwen answers flip without predicting how far the margins move.","feed_headline":"Learned context flips most correct logic answers","feed_subtitle":"Tiny trainable prefixes steer frozen models on unseen syllogism forms and wording changes","key_machinery":"The central object is the soft prefix: a trainable sequence of embedding vectors prepended to the token embeddings of a frozen model. The diagnostic machinery: (1) an exactly labeled syllogistic benchmark with 256 logical forms rendered in four wordings, split by form so test forms are unseen; (2) randomized A/B answer mappings, which separate fixed-symbol forcing from answer-meaning bias; (3) score models (one-shift, two-shift, global/class-conditioned affine, isotonic) fitted to the target-minus-source margin to test whether the prefix acts as a uniform shift or an example-dependent change; and (4) activation patching, which restores subsets of unprefixed states to locate where the answer","core_discovery":"Across three model families, a short trainable sequence of embedding vectors prepended to an unmodified prompt flips a large majority of correctly answered syllogistic judgments on logical forms never seen during prefix training, and the effect persists across new wordings and rephrased prompts. Learned prefixes beat matched random controls in all 16 model–direction–split comparisons (37–99 percentage points); Gemma validity prefixes flip 54–56% versus under 1% for random. Randomized letter-to-meaning mappings show the effect follows answer meanings, not fixed symbols, and target-answer rates of 93–99% identify a broad answer preference. Weak cross-task transfer rules out a shared logical op","pith_inferences":["Editorial inference: the same controlled design — exact labels, held-out forms, matched random controls — could be applied to other formally defined tasks (arithmetic, temporal reasoning, graph query answering) to test whether answer-bias dominance is a general property of soft prefixes or specific to syllogistic choice tasks.","Editorial inference: because the prompts explicitly instruct the model to ignore text outside the syllogism block and to treat prefixes as untrusted, the prefixes' success suggests that such semantic-scope instructions do not provide a reliable firewall; this may be relevant to defenses against injected or adversarial context.","Editorial inference: the shuffling results (order matters little in Qwen3-8B, more in Qwen3.6, most in Gemma) suggest model-specific reliance on positional versus content information; this could be tested by varying prefix length, position, or positional-encoding schemes.","Editorial inference: since selected random prefixes transfer poorly even when the best of 1000 are chosen, gradient-based optimization appears to find structure that random search at the same norm does not; a stronger test would compare against random search with the same number of optimization-equivalent evaluations."],"forward_implications":["Benchmark accuracy on a fixed prompt does not measure stability: a correct formal judgment can be reversed by an opaque, optimized context while the formal problem and correct answer stay unchanged.","Steering from learned context transfers across logical forms and interface changes in all three models, so the effect is not memorization of particular items or prompt templates.","The dominant mechanism is a broad answer-meaning preference rather than a fixed output symbol or a shared logical operation; high flip with low damage on minority-to-majority directions therefore characterizes intervention strength, not selective logical editing.","The same aggregate flip rate can conceal different response patterns: Gemma's margins are well approximated by simple score transformations, while Qwen's flips are predictable in direction but not magnitude.","Selected prefixes also change generated answers, so the effect is not an artifact of forced-choice continuation scoring."],"fun_headline_variants":["Tiny prefixes flip logic answers on unseen forms","Learned vectors override correct logic judgments","Soft prefixes steer logic answers broadly","Prefix training flips most correct syllogisms","Answer bias drives logic flips, not symbol tricks"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The transfer claims are estimated from test splits containing only six targeted logical forms per condition, so the whole 'unseen forms' generalization rests on those six forms being representative of a much larger space of syllogistic forms.","fun_headline_variants_meta":{"raw":{"variants":["Tiny prefixes flip logic answers on unseen forms","Learned vectors override correct logic judgments","Soft prefixes steer logic answers broadly","Prefix training flips most correct syllogisms","Answer bias drives logic flips, not symbol tricks"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000271,"raw_usage":{"total_tokens":1508,"prompt_tokens":830,"completion_tokens":678,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":574,"completion_tokens_details":{"reasoning_tokens":621}},"tokens_in":574,"tokens_out":678,"duration_ms":6720,"temperature":1.0,"reasoning_tokens":621,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T15:36:12.718772+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Evaluate the same training setup with a test split drawn from all 256 logical forms (e.g., a 50-form random holdout rather than the fixed six-form group). If flip rates on the broad holdout fall to within the random-prefix control range, the paper's claim that steering survives unseen forms is falsified for that broader space.","supporting_citations":[],"review_version":1}