{"id":"bca7052a-e632-452b-9a29-4b49ee01026a","arxiv_id":"2607.14147","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Refusal under a prefill jailbreak is a shallow response-site computation: the harm representation stays intact, the failure lives in an early response window, and the dominant mechanism is passive autoregressive conditioning.","lead":"A one-line prefill (\"Sure, here is\") makes aligned language models comply with harmful requests even though their internal harm detector still reads the request as harmful. The new finding is that the refusal decision is a shallow, early response-site computation driven by generic autoregressive conditioning, which explains why prompt-side monitors work against this attack.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Base-model discriminator confounded by chat-template vs raw-completion format; without a templated base control the passive-mechanism claim is not settled.","rationale":"The paper is unusually careful: the dose-matched position control, held-out state-transfer reversal, preregistered attention knockout, and controls ledger all support the early-window localization and the representational-intactness half of the claim. The step where the argument reaches beyond its controls is the base-model discriminator. Because the conclusion that the safety gate is not actively suppressing but passively conditioned is exactly what the base model is supposed to show, and because the base model is evaluated in a different format (raw completion vs chat template), the inference is not yet clean. This is a standard counterfactual-identification issue, not an internal inconsistency or a dispute with consensus. The proposed test directly compares the base model under the same chat template; if it reproduces the knockout's prefill-specific drop, the passive reading is supported. If not, the paper should weaken the 'dominant mechanism is passive' claim to something like 'prefill dependence is also observed in an untuned base model under raw completion,' which would not distinguish passive conditioning from format-dependent effects. Either way, the early-window, representational-intactness, and response-site-localization results remain intact, so the reader's conditional verdict is appropriate.","tokens_in":31377,"tokens_out":5282,"duration_ms":58594,"concrete_test":"Run the §4.3 early prefill-edge attention knockout on Qwen2.5-1.5B base using the identical chat-template input as the instruct run (same user/assistant formatting, same forced 'Sure, here is' assistant prefix), rather than raw completion. Measure harmful-content rate under early knockout vs mass-matched control. If the prefill-specific drop (raw-completion 64%→25%, +39pp) is reproduced, the passive conclusion survives the format confound. If the drop attenuates toward the mass-matched control, the passive claim is not established and the verdict should be conditional on a matched-format control.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The active-vs-passive resolution is the paper's central open question, and it is settled by a base-model discriminator (§4.3, App. E). The instruct model is evaluated under the chat template; the base model is explicitly evaluated 'with a raw-completion format (no chat template)' (App. E). The counterfactual therefore differs in safety tuning and in tokenization/conditioning context simultaneously. A prefill-specific early-attention knockout collapse in a raw-completion base model may reflect generic autoregressive dependence on a recent forced prefix in an un-templated continuation, not the absence of a safety-specific gate in the instruct model. This is not an internal inconsistency, but it is a counterfactual-identification gap: the paper uses this control to conclude 'the prefill's grip is generic autoregressive conditioning, not safety-specific suppression' (§4.3) and to demote the MLP 'instruct excess' to decodability (§7). A templated base run is needed before the passive reading is established.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies the \"Sure, here is\" prefill jailbreak in open-weight instruction-tuned models (1.5–3.8B, with checks at 7B and 14B). It reports a dissociation: a linear probe trained on plain harmful/benign prompts reads the last-prompt-token hidden states of attacked prompts as harmful (0.91–0.98) while behavioral refusal collapses to chance. Because prefill tokens are appended to the response, this intactness is partly structural, and the paper is explicit about that. The paper then localizes the behavioral failure causally to an early response window using activation patching, state transfer, and a preregistered attention-knockout, and uses a base-model control to argue that the prefill's grip is generic autoregressive conditioning rather than safety-specific suppression. It concludes that a prompt-side monitor is immune to response-site attacks by construction, and that refusal is distributed in representation space but positionally fragile.","tokens_in":31506,"tokens_out":6178,"duration_ms":65852,"significance":"If the main claims hold, the paper makes a strong contribution: it moves the account of prefill jailbreaks from behavioral depth-shallowness to a causal, response-site mechanism, and it gives a structural boundary for representation-based monitoring. The paper is unusually careful methodologically: preregistered knockout, mass-matched and dose-matched controls, held-out transfers, judge-free coherence reads, a claims-and-evidence ledger, and explicit retraction/demotion of earlier interpretations. These strengths make the core dissociation and the early-window localization credible. The main unresolved risk is the identification of the passive mechanism, which depends on a base-model control that is confounded by formatting; this is fixable but currently load-bearing.","major_comments":[{"comment":"The passive-vs-active resolution rests on a confounded counterfactual. The instruct models are evaluated under the chat template; App. E states that the Qwen2.5-1.5B base model is run \"with a raw-completion format (no chat template)\". The base-vs-instruct comparison therefore changes safety tuning and conditioning context simultaneously. A prefill-specific attention-knockout collapse in a raw-completion base model may reflect generic dependence on a recent forced prefix in un-templated continuation, not the absence of a safety-specific gate. Because the abstract and §7 state the central question as resolved toward passive conditioning, this gap is load-bearing. A templated base control, or a base model run with the chat template (or an instruct model run in raw format), is required before the generic claim is established; the same confound colors the MLP path-patch and the logit-trace co","section":"§4.3, App. E"},{"comment":"The mass-matched knockout control matches total baseline attention mass but not serial position/recency. The prefill keys are the final prompt-position keys; if early-window attention to any recent key is what sustains compliance, a control removing equal mass from the most recent non-prefill keys might also collapse the harmful continuation. The current \"non-prefill prompt keys\" selection is not described as recency-matched. Please add a position/recency-matched control, or a key-position permutation, to separate \"prefill-specific\" from \"recent-position-specific.\" Without this, the statement that the grip operates through attention to the forced prefix specifically is not fully settled.","section":"§4.3, Table 3; App. E"}],"minor_comments":[{"comment":"When the base-model discriminator is first introduced, the main text should explicitly note that the base model is run in raw-completion format without the chat template; currently that fact appears only in App. E. This is essential for readers to assess the counterfactual.","section":"§4.3"},{"comment":"The 24-seed random-control distribution is right-skewed (range 0–60%, median 0.12). Reporting the median alongside the mean in Table 2 would be more informative than the mean alone, since the harm−random delta is driven partly by a few high outliers.","section":"Table 2 / App. D"},{"comment":"The PCA rank analysis at n=70 is appropriately hedged in the appendix, but the main text's phrase \"the signal is spread across directions\" could be read as a stronger rank claim than the data support. A one-sentence reminder that the high-rank reading is exploratory at n=70 would help.","section":"§5.1"}],"recommendation":"major_revision","confidential_remarks":"The reader's skeptic concern about the base-model discriminator is valid and is my main reason for major revision. The paper is otherwise exceptionally well controlled; I would prioritize a templated-base run and a recency-matched knockout control. If those confirm the passive reading, the paper would be close to acceptable; if not, the passive claim should be weakened to 'generic at least in part'."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The core finding is solid and worth knowing: under a prefill jailbreak, the harm representation on the prompt side stays bit-for-bit intact while behavioral refusal collapses, and refusal is localized to an early response window. The causal work here is genuinely new — the dose-matched position control, the held-out state-transfer reversal with a random control at 0%, the preregistered attention knockout with a mass-matched control, and the base-model discriminator. The controls ledger is unusually honest: several flashy effects get demoted or retracted, and the claims table is transparent about what passed. This is the most careful mechanistic account of prefill jailbreaks I have seen.\n\nThe soft spot is the one the stress-test flagged: the active-vs-passive conclusion rests on a base-model comparison that conflates safety tuning with formatting. The instruct models run under the chat template; the base model runs in raw-completion format, no template. That is a real confound. A prefill-specific collapse in a raw base model could reflect generic dependence on a forced recent prefix in an untemplated continuation, not the absence of a safety-specific gate. The paper never flags this format difference as a limitation, and it is load-bearing for the claim that the prefill's grip is generic autoregressive conditioning. A templated base-model run is needed before that reading is settled. The other caveats — withheld completions, small models at the mechanism level, AdvBench-calibrated probe — are disclosed and proportionate. The representational-intactness half is structural, as the paper itself says; the new content is the causal localization and the passive-vs-active account, and that account currently rests on a counterfactual that is not clean.\n\nWho is this for? Anyone working on jailbreak mechanisms, safety classifier placement, or interpretability of refusal. It deserves a serious referee. The central localization claim and the state-transfer result hold up; the passive-mechanism conclusion needs a templated base control before publication. I would recommend acceptance with that revision, or at minimum a request for the extra experiment.","headline":"Careful, well-controlled mechanistic study; the passive-conditioning conclusion needs a templated base-model control before it is settled.","tokens_in":32084,"tokens_out":1350,"would_cite":true,"duration_ms":17984,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A one-line prefill jailbreak can flip an aligned model to compliance while its internal harm representation stays fully intact — refusal, the paper argues, is a shallow response-site computation.","keywords":["prefill jailbreak","refusal mechanism","representation probing","activation patching","attention knockout","safety alignment","autoregressive conditioning","mechanistic interpretability"],"falsifier":"Find a checkpoint pair that differ only by safety tuning (same base, same chat template, with and without RLHF/DPO) and run the early prefill-attention knockout on both. If the base model does not show the prefill-specific collapse of harmful content (the 64%→25% pattern), the passive-reading claim is false. Alternatively, train a model with a different chat template but identical safety tuning: if the knockout effect disappears, the confound, not the mechanism, was responsible.","tokens_in":31128,"feed_emoji":"🛡️","tokens_out":3595,"duration_ms":34610,"temperature":0.7,"pith_summary":"The paper claims that the one-line prefill jailbreak ('Sure, here is') does not erase the model's representation of harm: a probe reads the complied-with prompts as harmful as the refused ones (0.91–0.98), while refusal behavior drops to chance. Refusal, it argues, is a shallow, response-site computation: the attack wins in an early window of the response, through attention to the forced prefix, and the dominant mechanism is generic autoregressive conditioning rather than safety-specific suppression. The causal evidence includes a dose-matched position control, state-transfer reversal, and attention knockout with a mass-matched control. If correct, a monitor reading the prompt-side representation is immune to this class of attack by construction, and refusal is positionally fragile even though it is distributed in representation space.","feed_headline":"Attack fools behavior, not the model's harm detection","feed_subtitle":"A probe reads prefill-flipped prompts as harmful as refused ones, so monitoring the prompt side can catch the jailbreak by construction.","key_machinery":"The prefill attack itself — an affirmative prefix appended to the model's own response — is the central object; because it never touches the prompt tokens, the prompt-side representation is invariant by construction. The paper's causal instruments are the early-window position control (dose-matched bands of the response), state transfer (injecting the plain-condition minus attack-condition residual), and the attention knockout that severs early response-to-prefill attention edges, with a mass-matched control. The base-model discriminator — running the same knockout on a non-safety-tuned base model — is what separates generic autoregressive conditioning from safety-specific suppression.","core_discovery":"Under the prefill attack, the model acts on the request as if it were benign while its internal read of the prompt is as harmful as ever: a linear probe scores the flipped prompts 0.91–0.98, level with the ones the model refuses, and behavioral refusal falls to chance. The refusal decision is computed at the response site, not at the prompt: restoring the harm direction over the first half of the response re-engages refusal as much as the whole response, while the second half is nearly inert, and transplanting the model's own refuse-state re-induces refusal in 74% of held-out cases. Knocking out early response attention to the prefill, but not an equal attention mass elsewhere, selectively s","pith_inferences":["If the passive account is right, defenses that focus on the response window will keep chasing the attack across positions; the durable placement is prompt-side, and this extends to any response-site jailbreak that occupies the early output window.","The harm-direction's partial write-handle suggests a possible 'representation repair' at the response site that could be combined with a monitor, but the paper's own steering nulls indicate this is not a clean intervention.","A direct test of the weakest assumption: use a checkpoint pair that differ only by safety tuning (same base, same chat template, with and without RLHF/DPO) and re-run the knockout — if the base model fails to show the prefill-specific collapse, the passive reading is confounded.","Prompt-modifying jailbreaks should behave differently from prefill: they attenuate or displace the prompt-side representation, so a representation monitor would be strictly weaker against them; this is a testable prediction the paper leaves implicit."],"forward_implications":["A prompt-side monitor that reads the last-prompt-token representation catches this attack at 100% with 0% false positives on in-distribution negatives, because the attack cannot alter that representation.","Refusal restoration is a model-dependent fallback: cutting the prefill's early-window attention re-engages refusal in instruction-tuned models but only degrades continuations in base models.","No single direction, head, or layer is a clean refusal handle; the decision is decodable but distributed, so single-direction steering fails while a full linear probe reads it.","At larger scale (7B, 14B) the dissociation holds and the passive mechanism persists in the content channel, though the behavioral refusal signature becomes model-dependent."],"fun_headline_variants":["Prefill jailbreak: refusal is a shallow, response-site computation","Probe sees harm, behavior complies: prefill attack hits response site","Refusal is a response-site trick; first half of reply is all the attack needs","Knockout reveals prefill attack is generic conditioning, not safety override","Monitor prompt-side representations: they still read harm under prefill attack"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The passive mechanism conclusion rests on treating a non-safety-tuned base model as a clean counterfactual for the instruction-tuned model; if base and instruct models differ in more than safety tuning (chat template, instruction following, pretraining distribution), the attribution of the knockout effect to generic autoregressive conditioning is weakened.","fun_headline_variants_meta":{"raw":{"variants":["Prefill jailbreak: refusal is a shallow, response-site computation","Probe sees harm, behavior complies: prefill attack hits response site","Refusal is a response-site trick; first half of reply is all the attack needs","Knockout reveals prefill attack is generic conditioning, not safety override","Monitor prompt-side representations: they still read harm under prefill attack"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000581,"raw_usage":{"total_tokens":2660,"prompt_tokens":922,"completion_tokens":1738,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":666,"completion_tokens_details":{"reasoning_tokens":1641}},"tokens_in":666,"tokens_out":1738,"duration_ms":14038,"temperature":1.0,"reasoning_tokens":1641,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T06:27:05.853128+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Find a checkpoint pair that differ only by safety tuning (same base, same chat template, with and without RLHF/DPO) and run the early prefill-attention knockout on both. If the base model does not show the prefill-specific collapse of harmful content (the 64%→25% pattern), the passive-reading claim is false. Alternatively, train a model with a different chat template but identical safety tuning: if the knockout effect disappears, the confound, not the mechanism, was responsible.","supporting_citations":[],"review_version":1}