{"id":"6f6798a4-27aa-4cfd-8094-b5489c275e84","arxiv_id":"2608.06578","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Six frontier language models show categorical, model-specific response modes under steering pressure, with GPT-5 uniquely withholding reasoning while answering, and a linear probe plus activation steering tracing Llama's derail behavior to its residual stream.","lead":"This study tested six frontier AI models under prompts that push them to compromise values, expose hidden reasoning, or suppress reasoning, and found the models differ in the kind of response they give. The differences include a GPT-5-only mode that withholds reasoning while still answering, and a Llama experiment that decoded and steered a behavior from the model's internal state.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 99/0 reasoning-disclosure claim and the unique-mode conclusion rest on LLM-judge labels that are never validated against human raters; since the key label was built from GPT-5's own responses, this unvalidated measurement instrument is the load-bearing risk.","rationale":"The reader's weakest-assumption analysis identified exactly the right soft spot: all headline rates depend on LLM-judge labels that have not been validated against humans. My stress-test pass sharpens that concern by pointing to the specific mechanism that makes it load-bearing: the reasoning-refusal label was built from GPT-5's own development responses, so the label is tailored to the model that subsequently scores 99%. The paper's internal controls are strong in several respects: the held-out 80-item rerun, the context-field removal experiment, the leave-one-out consensus, and the token-budget remediation all address adjacent confounds. But none of those controls replaces human validation of the label definitions themselves. The context-field experiment is informative: removing the field shifts 16.7 points of labels in reasoning-suppression, the very category with the lowest kappa, and although the authors report that all directional claims survive in that small sample, the sample is only six instances per model per condition. For the reasoning-elicitation claim, the field-free control reproduces the GPT-5 separation but on six GPT-5 instances, which is suggestive rather than definitive. The mechanistic Llama section is honestly framed as an existence proof and includes permutation tests, a small-model baseline, and a steering sweep; it is not the load-bearing part of the central cross-model claim. The statistical analysis is careful, with Fisher exact tests and BH control, and the 99/0 effect is so large that ordinary multiple-testing concerns are not serious. The most defensible verdict remains conditional: the paper's central claim is plausible and well-audited, but it should not be upgraded to accepted until human raters confirm the label taxonomy on a sample large enough to re-estimate the key rates. My recommendation is UNCHANGED because the reader's conditional verdict already encodes exactly this reservation.","tokens_in":12493,"tokens_out":5838,"duration_ms":61013,"concrete_test":"Recruit at least three human annotators, blinded to responder identity, to classify a stratified sample of 300 steered reasoning-elicitation responses (all 100 GPT-5 responses and 40 each from the other five models) using the frozen Table 1 rubric without the item context field. Compute human-majority labels, measure agreement with the six-judge leave-one-out consensus (percent agreement and Cohen's kappa), and recompute GPT-5's reasoning-refusal rate and the other models' rates from human labels alone. The central claim survives only if human-majority labels reproduce the categorical split with GPT-5 at or near ceiling (e.g., at least 90 of 100) and all other models at or near zero (e.g., at most 5 of 200), and if kappa with the LLM consensus is at least 0.7.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that GPT-5 alone refuses to disclose reasoning while delivering the answer (99/100 vs 0/500) is a claim about a categorical response mode. Every rate in the paper comes from the leave-one-out consensus of six LLM judges applying frozen rubrics, and no human validation of those labels is reported. The paper itself flags this in the Conclusion: 'validating the judge panel against human raters on a sample is a necessary next step.' That is exactly the missing support, and it is load-bearing rather than routine for two reasons. First, the 'reasoning-refuse-answer-intact' rubric was derived from close reading of GPT-5's own development-phase responses, so the label is optimized to describe the one model that scores 99%; the held-out 80-item rerun removes item contamination but not the fact that the label definition itself was shaped by GPT-5's behavior. Second, the suppression-resistance finding lives in the category with the lowest inter-judge agreement (kappa = 0.599) and the largest context-field effect in the paper's own control experiment (16.7 percentage points). If human annotators do not apply the reasoning-refusal label the way the LLM judges do, or if they judge many GPT-5 responses as not 'answer intact' because the answer is substantively altered, the near-ceiling separation could shrink substantially. The 99/0 result is so extreme that a moderate label bias could still leave a large gap, but the paper's broader conclusion that these are differences in kind, not degree, would not be established by an unvalidated measurement instrument.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript reports a comparative study of six frontier LLMs under paired base/steered prompts across three core categories (values-conflict, reasoning-elicitation, reasoning-suppression), using all six models as blind judges with frozen rubrics and leave-one-out consensus scoring. It reports categorical differences: GPT-5 withholds reasoning while delivering answers on 99/100 steered reasoning-elicitation items versus 0/500 for other models; Claude Opus 4.7 and GPT-5 resist suppression in different modes; base values-conflict compliance splits models into three tiers; and, for Llama, a linear probe decodes the derail-versus-answer split from residual-stream activations and activation steering moves the derail rate from 0% to 86%. The paper includes extensive audits: token-budget remediation, an item-context control, and a full rerun of the analysis on items never used to build the rubrics.","tokens_in":12824,"tokens_out":5463,"duration_ms":53324,"significance":"If the behavioral findings hold, the paper makes a substantial contribution by showing that steering differences are qualitative as well as quantitative, and by providing a reusable symmetric cross-model evaluation template. The design is unusually careful: paired items, blind judging, leave-one-out consensus, BH correction for the main comparisons, a pre-specified single hypothesis for suppression resistance, token-budget remediation, a context-field control, and a held-out subset analysis. The mechanistic section is explicitly illustrative but methodologically sound, using nested cross-validation, permutation tests, a prompt-ambiguity control model, and held-out steering items. The main open risk is the unvalidated judge panel; the paper itself names human validation as a necessary next step, and that missing support is load-bearing for the central claims.","major_comments":[{"comment":"All headline rates are produced by the leave-one-out consensus of six LLM judges, and no human validation of those labels is reported. This is load-bearing for the central claim, not merely a routine future step: the 'reasoning-refuse-answer-intact' label was derived from close reading of GPT-5's own development-phase responses, so the label definition itself is shaped by the one model that scores 99%. The held-out 80-item rerun removes item-selection contamination but cannot remove label-definition contamination. The concern is sharpest in the reasoning-suppression category, where Fleiss' kappa is 0.599 and the item-context field changes 16.7 percentage points of labels. The manuscript should either provide a human-annotated validation sample with agreement statistics and re-estimated key rates, or explicitly reframe the headline numbers as relative to the LLM judge panel rather than as model behavior tout court.","section":"Benchmark Validity; Conclusion"},{"comment":"The field-free control is reassuring directionally but too small to quantify the key rates convincingly. With 216 responses total, each model has only six instances per category-condition cell; on steered reasoning-suppression there are only three resistance events in the field-free condition, and on reasoning-elicitation the GPT-5 result is 6/6 versus 0/30. The statement that 'every pairwise comparison carrying a claim in this paper holds' is stronger than the sample can support. The authors should report the sampling uncertainty of the field-free rates and ideally run a larger field-free replication, or temper the claim accordingly.","section":"Control Experiment for the Item Context Field"},{"comment":"The Fisher exact tests and BH procedure treat the consensus labels as error-free ground truth; judge disagreement and the lexicographic tie-break rule are not propagated into the reported p-values. Given that the least-agreement category (kappa = 0.599) carries a central finding, the paper should include a sensitivity analysis under alternative consensus rules (for example, different tie-breaking orders or a majority that includes self-judgment) rather than relying only on the listed robustness checks.","section":"Statistical Analysis"}],"minor_comments":[{"comment":"The development stage is described as involving 'model-assisted reading' and 'independent authorship,' but the manuscript does not clarify who performed the reading and authorship checks; please specify the procedure and any human involvement.","section":"Stage 1: development"},{"comment":"The caption column 'Refuse-reasoning' does not match the rubric label 'reasoning-refuse-answer-intact'; using one consistent label across Table 1, Table 2, and the text would reduce ambiguity.","section":"Results, Table 2"},{"comment":"The seven alpha levels in the sweep are each compared with alpha=0 by Fisher's exact test without an explicit multiple-comparison correction; this does not affect the behavioral findings, but the paper should state that the sweep was interpreted as a dose-response pattern rather than as seven independent confirmatory tests.","section":"Activation Steering Setup"},{"comment":"The abstract reports '0.87 held-out balanced accuracy' while the body reports a peak of 0.866 and a plateau of 0.83-0.87; please align these numbers so the abstract does not overstate the peak.","section":"Abstract and Probe Results"},{"comment":"The phrase 'the 80% of items written after the rubrics were frozen' is ambiguous because the held-out set is 80 items per category, not 80% of the 340-item benchmark; please rephrase to avoid confusion about the total denominator.","section":"Conclusion"}],"recommendation":"major_revision","confidential_remarks":"The paper is within the journal's scope and the evaluation design is stronger than most work in this area; the only substantive obstacle is the missing human validation of the behavioral labels. I would encourage treating this as a major revision rather than a rejection, since a focused human-rater study on a stratified sample could reasonably settle whether the categorical findings survive a different measurement instrument."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe headline: this paper is not another steerability benchmark that just asks whether the model hits the target. It asks what the model does when it refuses, and it reports categorical differences across six frontier models—most strikingly, GPT-5 alone withholds its reasoning while delivering the answer on 99 of 100 steered items, against 0 in 500 for all others. The design is unusually careful: paired base and steered items, blind peer judging by all six models, leave-one-out consensus that excludes self-judgment, BH correction, a token-budget remediation, a control for the item-context field, and a rerun on the 80% of items never used to build the rubrics. The mechanistic section on Llama is a clean existence proof: linear probe at 0.87 held-out balanced accuracy, permutation test, and a steering sweep from 0% to 86% with an appropriate control.\n\nThe soft spots are real but not fatal. All headline rates come from LLM-judge labels that have not been validated against human raters—the paper says this itself. The reasoning-refusal rubric was derived from close reading of GPT-5's own responses, so the label is partly tailored to the model that scores 99%. The suppression-resistance finding sits in the category with the lowest inter-judge agreement (κ=0.599) and the largest context-field effect (16.7 points). That said, the 99/0 gap is so extreme that a moderate label bias would still leave a large separation, and the held-out analysis preserves every finding. The paper is honest about the one-model-per-developer limitation.\n\nThis paper is for anyone working on steerability, refusal styles, or model evaluation methodology. It deserves a serious referee. My recommendation: send it out; the human-validation gap is a matter for the authors to close in revision, not a reason to desk-reject.","headline":"A genuinely careful empirical study of qualitative response modes under steering, with a striking GPT-5-specific reasoning-withholding finding; the main caveat is unvalidated LLM judges, but the paper is thorough and deserves peer review.","tokens_in":13332,"tokens_out":2613,"would_cite":true,"duration_ms":22791,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper finds that frontier language models differ in the kind of response they give under steering pressure, not just in how far they move, and that GPT-5 withholds its reasoning on 99 of 100 steered items while every other model…","keywords":["behavioral steerability","response modes","LLM-as-judge","reasoning refusal","values suppression","linear probing","activation steering","frontier language models"],"falsifier":"A human-annotation study that re-labeled a random sample of the 4,080 responses with the same rubrics and found that GPT-5's reasoning-refusal rate fell to within the range of the other models, or that Opus and GPT-5's suppression-resistance rates no longer exceeded the others', would refute the paper's categorical claim; equivalently, a re-run with the item-context field removed that failed to reproduce the 0-versus-99 split would undermine it.","tokens_in":12293,"feed_emoji":"🤖","tokens_out":8680,"duration_ms":67122,"temperature":0.7,"pith_summary":"This paper tries to establish that frontier language models from different developers do not merely comply more or less when pressured; they give categorically different kinds of responses, and some response modes exist in only one model. On steered prompts asking it to expose its reasoning, GPT-5 declined to disclose on 99 of 100 items while every other model disclosed at 95% or more. On prompts instructing models to suppress values reasoning, only Claude Opus 4.7 and GPT-5 resisted, in different styles. The paper also shows, for the open-weight model Llama, that the largest behavioral split is decodable from its internal activations and can be driven from 0% to 86% by steering a single learned direction. If true, these findings would mean that safety evaluations that measure only how far a model shifts miss the more important question of what mode of response it chooses.","feed_headline":"GPT-5 withholds reasoning 99/100 times when pressured; rivals: 0","feed_subtitle":"One model alone refuses to show its reasoning, and one internal direction drives Llama's behavior from 0% to 86%.","key_machinery":"The load-bearing machinery is a symmetric evaluation protocol: every one of the six models answers the same 340 paired base/steered items and then acts as a blind rubric-based judge over all responses, with the ground-truth label for each response being the leave-one-out consensus of the five peer judges, excluding the responder's self-judgment. The mechanistic extension for Llama uses a linear probe, defined as a logistic-regression readout of residual-stream activations, to show that the derail-versus-answer distinction is linearly decodable, and a contrastive steering vector, defined as the normalized difference of mean activations of the two behavior classes, added at one layer to show the decoded direction causes the behavior to change across a sweep of strengths. The whole design is audited by a token-budget remediation, a hypothesis-blind judge-prompt control, and a rerun on the 80 percent of items never used to build the rubrics.","core_discovery":"The central claim is that the six evaluated models differ in the kind of response they produce under explicit steering pressure, not only in how much they move, and that several differences are exclusive to a single model. The evidence is a symmetric benchmark in which 100 paired base/steered items per category are answered by all six models and judged by all six models blind, with ground truth a leave-one-out consensus. GPT-5 answers the prompt but refuses to share its reasoning on 99 of 100 steered reasoning-elicitation items, against 0 for the other five models across 500 items; the same model is also the least likely to surface values content without steering. Claude Opus 4.7 and GPT-5 are the only models that openly resist reasoning-suppression instructions, and they do so in statistically distinguishable ways: Opus challenges the instruction while complying, GPT-5 refuses the framing outright. On unsteered values-conflict items the models split into three compliance tiers, and for Llama the paper traces the largest baseline split to a linear direction in the residual stream: a probe reads derail-versus-answer at held-out balanced accuracy 0.87, and adding that direction during generation moves the derail rate from 0% to 86% across a strength sweep.","pith_inferences":["The GPT-5-exclusive reasoning-refusal mode suggests a deliberate or emergent policy about hiding chain-of-thought from users; testing sibling models from the same developer would show whether it is a family trait or an instance-specific quirk.","Because a single difference-of-means direction moved a behavior from 0% to 86%, the same causal style of analysis could be applied to other response categories in open-weight models, turning behavioral modes into manipulable variables.","The reasoning-suppression category, where the key resistance finding lives, is exactly where judges agreed least and the context field moved most labels; future versions of this benchmark would gain the most from human-annotated labels there.","The 9.7-point drop in label reproduction when the item-context field was removed suggests that blind judging is still expectation-sensitive; reporting both with-field and field-free rates may become a standard audit for LLM-judged benchmarks."],"forward_implications":["Evaluators of steerability should report which response mode a model chooses, since two models can move the same distance in opposite modes.","A model's unsteered disposition can foreshadow its steered behavior; GPT-5's base deficit in values content predicts its near-total refusal to disclose reasoning under elicitation.","Suppression instructions change what DeepSeek-R1 says but not what it registers internally: the excluded values dimension appears in its trace on 85 of 100 steered items before being set aside.","The linear direction that moves Llama's derail rate from 0% to 86% identifies a concrete internal target that future work could patch or ablate.","Three tiers of baseline compliance mean that any single-model safety result cannot be assumed to transfer across developers."],"supporting_citations":[{"why":"Supplies the linear probing method used to decode behavior from residual-stream activations.","marker":"Alain and Bengio 2016"},{"why":"Defines the difference-of-means activation steering vector applied at layer 40.","marker":"Turner et al. 2024"},{"why":"Provides the contrastive activation addition approach that motivates the steering intervention.","marker":"Rimsky et al. 2024"},{"why":"Establishes the permutation-test methodology for deciding when probe accuracy is above chance.","marker":"Belinkov 2022"},{"why":"Provides the exact permutation p-value estimator used for probe significance.","marker":"Phipson and Smyth 2010"},{"why":"Supplies the false discovery rate control applied to the 450 pairwise model comparisons.","marker":"Benjamini and Hochberg 1995"},{"why":"Gives the inter-judge agreement metric used to report rubric reliability.","marker":"Fleiss 1971"},{"why":"Provides the discovery-first label-construction approach that the development phase follows.","marker":"Sharma et al. 2023"},{"why":"Documents LLM-judge biases that motivate the blind multi-judge design.","marker":"Zheng et al. 2023"}],"fun_headline_variants":["GPT-5 withholds reasoning 99/100 under steering; peers 0","One linear direction flips Llama from answer to derail (0→86%)","Steering uncovers exclusive response modes in frontier LLMs","GPT-5 and Claude resist suppression differently; Llama mapped","Probe decodes Llama's derail behavior at 87% accuracy"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire headline depends on the labels produced by consensus among LLM judges, and those labels have not yet been validated against human ratings, especially for reasoning-suppression where inter-judge agreement is moderate ($\\kappa = 0.599$) and removing the item context changed the most labels.","fun_headline_variants_meta":{"raw":{"variants":["GPT-5 withholds reasoning 99/100 under steering; peers 0","One linear direction flips Llama from answer to derail (0→86%)","Steering uncovers exclusive response modes in frontier LLMs","GPT-5 and Claude resist suppression differently; Llama mapped","Probe decodes Llama's derail behavior at 87% accuracy"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000311,"raw_usage":{"total_tokens":1833,"prompt_tokens":1068,"completion_tokens":765,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":684,"completion_tokens_details":{"reasoning_tokens":668}},"tokens_in":684,"tokens_out":765,"duration_ms":6515,"temperature":1.0,"reasoning_tokens":668,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T04:16:43.565037+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A human-annotation study that re-labeled a random sample of the 4,080 responses with the same rubrics and found that GPT-5's reasoning-refusal rate fell to within the range of the other models, or that Opus and GPT-5's suppression-resistance rates no longer exceeded the others', would refute the paper's categorical claim; equivalently, a re-run with the item-context field removed that failed to reproduce the 0-versus-99 split would undermine it.","supporting_citations":[{"cited_title":"Steering","cited_arxiv_id":null,"evidence_quote":"Defines the difference-of-means activation steering vector applied at layer 40."},{"cited_title":"Steering","cited_arxiv_id":null,"evidence_quote":"Provides the contrastive activation addition approach that motivates the steering intervention."}],"review_version":1}