{"id":"62c463ae-6e43-446b-b356-ae8ab9e5c3a1","arxiv_id":"2608.13329","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"The contrastive prompt used to build an evaluation-awareness probe determines the reported score and even the sign of its trend with model size, so the statistic is a property of the prompt rather than of the model.","lead":"This paper shows that a popular method for measuring whether AI models know they are being tested is controlled by the researcher's choice of prompt wording, not by the model. Switching the wording flips the conclusion about how the signal scales with model size, so published comparisons rest on an unstated choice.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The quantitative prescription (Eρ2 = 0.018, k ≈ 9) rests on one four-model Qwen2.5 ladder and a hand-written 6×6 wrapper crossing, with only three degrees of freedom for the model facet; Section 11 flags this, but the abstract's unqualified 'about ten prompts' remains under-supported.","rationale":"I read the paper as making a strong, well-supported qualitative claim: holding task text, scoring set, layer convention, and models fixed, the probe statistic and its scaling trend vary with the contrastive prompt, so a single-prompt study cannot support cross-model comparison. This claim is supported by the 36-wrapper distribution on Qwen2.5, the reproduction of both published correlations inside that range, the direction-geometry results (within-wrapper cosine 0.9995 vs across-wrapper 0.3749), and the surface-form separability control. The reader's conditional verdict is appropriate because the quantitative reliability coefficients and the prompt-count prescription are estimated from a narrow base: four models of one family and a designed wrapper set, with the paper itself flagging the three-degrees-of-freedom limitation. I agree with the reader that this is the weakest assumption, and I do not see a more load-bearing weakness. The additional ladders (Qwen3, OLMo-2) confirm the qualitative pattern, but the magnitudes differ substantially (Eρ2 from 0.012 to 0.263), so the abstract's unqualified 'small share' and 'about ten' overstate what is currently known. The Section L gate change is a separate integrity concern, but I do not treat it as the primary load-bearing issue because the score-variance argument does not depend on the slope-reproducibility result; nevertheless, it warrants the reader's attention and would benefit from independent verification of the 0.874 threshold's pre-specification. A concrete, feasible check is to re-estimate the variance components with a larger, more representative wrapper sample and additional families, which would settle whether the k≈9 prescription is stable. Until that is done, CONDITIONAL remains the right verdict.","tokens_in":18994,"tokens_out":8975,"duration_ms":89979,"concrete_test":"Generate a large, externally plausible wrapper sample: have an instruction-tuned LLM produce 100 evaluation/deployment wrapper pairs from a distribution intended to match real evaluation practice, then run the same crossed variance decomposition on the Qwen2.5 ladder and on at least two other model families with overlapping parameter ranges, recomputing Eρ2, the model share, and k. If k leaves the 5–20 range or the model share rises to the level of the wrapper×model share, the headline prescription is not stable.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative deliverable—that a defensible comparison needs about nine to ten prompts, with Eρ2 = 0.018 on the Qwen2.5 ladder—depends on variance components estimated from four models of one family (three degrees of freedom) and thirty-six wrappers that are a designed crossing, not a random sample of the prompts researchers would actually write. Section 11 acknowledges this directly. The magnitudes are ladder-specific: on OLMo-2 the model share is 13.1% and Eρ2 = 0.263, an order of magnitude larger than the headline 0.018, even though the qualitative ordering (wrapper×model exceeds item×model) persists. If the hand-written wrapper set is more diverse than the population of field practice, the wrapper×model component is overestimated, making k conservative; if it is narrower, k is underestimated and the prescription could be anti-conservative. The claim that a single-prompt design cannot support cross-model comparison does not depend on these magnitudes, but the paper's stated deliverable—'the number of prompts a defensible comparison requires'—does. A separate concern is the post-hoc gate change disclosed in Section L, which affects the evidence that the scaling sign is a property of the wrapper rather than unestimable noise; that issue is secondary because the variance-component argument for the score's prompt-dependence does not rely on the disputed gate.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies the contrastive-prompt probe statistic used in evaluation-awareness research. It argues that the statistic has an undisclosed free parameter: the choice of prompt wrapper. Holding task text, scoring set, layer convention, and models fixed, the paper shows that the reported score and the sign of its correlation with model scale change with wrapper choice, and that two published studies with opposite scaling signs are both reproducible within a single wrapper distribution. Using generalizability-theory variance components on a 6x6 wrapper crossing over four Qwen2.5 models, it estimates the model share at 0.8%, E-rho-squared at 0.018, and the wrapper-by-model interaction at 13.4 times the item-by-model interaction, concluding that more items cannot repair the measurement while varying wrappers can. It also shows that the scoring split is surface-form separable, that content-free directions recover much of the published scores, and it reanalyzes Chaudhary et al. (2025) and Manek (2026).","tokens_in":19257,"tokens_out":9954,"duration_ms":97778,"significance":"The qualitative result, that prompt choice changes the reported score and the scaling sign, is important and strongly supported. The paper follows best practices: released code, data manifests, pinned checkpoints, fixed seeds, byte-identical task-text assertions, pre-registered predictions, and explicit disclosure of a withdrawn layer prescription and a post-hoc gate change. The reanalysis is possible only because both prior groups released artifacts, and the paper credits them appropriately. If the prompt-dependence claim holds, single-prompt probe comparisons cannot support cross-model claims, which is a substantial correction to the literature. The quantitative prompt-count prescription is less secure and needs qualification before the abstract can state it unqualified.","major_comments":[{"comment":"The headline 'about ten prompts' is presented without the conditions the paper itself identifies. Section 11 states that E-rho-squared = 0.018 and k = 9 rest on four Qwen2.5 models (three degrees of freedom for the model facet) and on a hand-written 6x6 wrapper crossing that is a designed set rather than a random sample of prompts a researcher would write. The OLMo-2 ladder yields a model share of 13.1% and E-rho-squared = 0.263, an order of magnitude larger, and Figure 3 shows the measured Qwen2.5 curve asymptoting at 0.541, never reaching 0.80; the k = 9 figure is therefore an assumed-variance calculation, not the measured Qwen2.5 result. Please qualify the prompt-count claim as ladder- and wrapper-sample-specific, or provide a sensitivity analysis over plausible wrapper populations, and adjust the abstract accordingly.","section":"Abstract; Sections 8 and 11"},{"comment":"The variance-component estimates are presented as exact point estimates, but the model facet has only three degrees of freedom and the interaction ratios have wide intervals; Section B gives an F-based interval of [1.03, 52.8] for the related pooled interaction ratio. The main text should state that the 0.8% model share, E-rho-squared = 0.018, and the 13.4x ratio are point estimates from a single variance-component solve on one ladder, with the uncertainty and ladder-dependence described in the main text rather than only in the limitations section.","section":"Section 6; Figure 2; Section 11"}],"minor_comments":[{"comment":"The sentence 'so the slope is a property of the wrapper and determines is the word the measurement supports' is garbled and should be rewritten.","section":"Section 4"},{"comment":"The post-hoc change of the sign-agreement gate from pooled agreement 0.700 to top-quarter agreement 0.90 should be flagged in the main text wherever Section 4 cites sign agreement, so that readers know the conditional figures are post-hoc descriptive statistics; the 0.874 Spearman-Brown slope correlation should be presented as the gate-independent decision criterion.","section":"Section L"},{"comment":"The phrase 'seven Qwen2.5, gemma-2 and Llama-3.2 models' is ambiguous; it should read 'seven models from the Qwen2.5, gemma-2, and Llama-3.2 families.'","section":"Section 3"},{"comment":"The abstract's 'leaves one claim standing and one not' is stronger than the body's careful statement that only the layer-selection correction could be applied to Chaudhary et al.; please soften the abstract or add the caveat that the direction control remains open.","section":"Section 10"},{"comment":"When reporting that a content-free direction reaches 70-116% of each published value, please state explicitly whether the comparison is against the label-permuted floor or the AR(1) floor; Figure 4 plots the label-permuted floor, but the text's 'content-free direction' refers to the AR(1) control.","section":"Section 9"}],"recommendation":"major_revision","confidential_remarks":"For the editor: I see no grounds for rejection; the qualitative finding is robust and the reporting is unusually transparent. The main revision should focus on the quantitative prescription and on making the Section L gate change visible in the main text. The paper is well within scope for cs.LG."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"If you read one paper this month on probing methodology, make it this one. The core finding is genuinely new in its demonstration: holding task text, scoring set, layer convention, and models fixed, the choice of contrastive wrapper moves the reported statistic enough to flip the sign of the correlation with model size. Both published scaling values fall inside the range produced by one 6x6 crossing, which is a strong and clean argument that the published sign disagreement is a wrapper effect, not a scientific dispute. The qualitative claim holds on three model ladders, and the paper's honesty is exceptional: pinned checkpoints, byte-identical task text, disclosed withdrawn layer prescription, and a frank appendix describing a post-hoc gate change that flipped their verdict. That last item is a red flag, but they mitigated it by showing the slope reliability (Spearman-Brown corrected correlation around 0.87–0.90) clears the threshold they originally wrote, and the variance-component argument does not depend on the disputed gate. The reanalysis is also careful: the Chaudhary scaling result survives floor correction, while the Manek peak-depth claim is honestly reported as unidentified at this estimator's precision. The soft spot is exactly where the stress-test lands. The quantitative deliverables, Eρ² = 0.018 and k ≈ 9–10 prompts, come from one four-model Qwen2.5 ladder (three degrees of freedom for the model facet) and a hand-written wrapper set that is a designed crossing, not a random sample of field practice. The paper itself flags this in Section 11, and the OLMo-2 ladder gives Eρ² = 0.263, an order of magnitude higher, so the magnitudes are clearly ladder-specific. That does not undermine the qualitative conclusion that single-prompt comparisons are indefensible, but it does mean the headline 'about ten prompts' is over-specified for the evidence. Also, while the disclosed gate change is a strength in transparency, it is still an analytic-flexibility episode that referees should probe. Who this is for: anyone doing contrastive probing, interpretability, or evaluation-awareness research, plus methods folks who care about measurement reliability. It deserves a serious referee: the artifact discipline alone justifies peer review, and the claims are important enough that the community needs to engage. My recommendation: send it to review, but ask the authors to present the prompt-count prescription as a sensitivity-dependent estimate, with uncertainty intervals on the variance components and more ladders if possible.","headline":"A careful, artifact-backed measurement critique showing the contrastive wrapper decides probe scores and scaling sign; trust the qualitative result, but the 'about ten prompts' prescription rests on thinner evidence than the abstract suggests.","tokens_in":19822,"tokens_out":1626,"would_cite":true,"duration_ms":19037,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A probe direction records the phrasing of the prompt that built it, not a stable property of the model, so a single-prompt design cannot compare models.","keywords":["evaluation awareness","activation probing","prompt sensitivity","contrastive prompts","generalizability theory","variance components","scaling laws","LLM evaluation"],"falsifier":"Collect a fresh sample of prompts that researchers actually use in evaluation-awareness studies, cross them in the same factorial design on the same four models, and recompute the variance components; if the model share is large enough that $E\\rho^2$ reaches 0.80 with one wrapper per model, or if no wrapper choice moves the scaling correlation from positive to negative, the central claim would be overturned. A cheaper check is to hold a published wrapper fixed and search a large corpus of deployment phrasings to see whether the sign of the scaling trend is ever immovable.","tokens_in":18706,"feed_emoji":"🧠","tokens_out":9504,"duration_ms":82804,"temperature":0.7,"pith_summary":"The paper tries to establish that the score produced by a standard activation-probe instrument—contrasting a model's internal states on \"you are being evaluated\" prompts against deployment prompts—records the wording of the prompt more than the model. Holding task text, scoring set, layer convention, and models fixed, the paper varies only the six-by-six crossing of evaluation and deployment framings and finds the reported statistic, and even the sign of its trend with model size, follows the wrapper; both published signs of that trend are reproduced inside one design. A variance decomposition attributes $0.8\\%$ of the variance in the reported number to the model, with most of the remainder in the model-by-prompt interaction, so collecting more evaluation items cannot repair the measurement. The paper concludes that a single-prompt design cannot support comparison between models, and gives about ten prompts (fifty-six evaluation-arm framings when both arms are crossed) as the number a defensible comparison requires.","feed_headline":"A probe direction reports the prompt, not the model","feed_subtitle":"Swapping one wording flips the sign of model-size trends, so one prompt cannot compare models.","key_machinery":"The load-bearing object is the $6\\times 6$ factorial crossing of prompt wrappers: six phrasings that frame a prompt as an evaluation and six that frame it as deployment, wrapped around byte-identical task text, producing 36 contrastive directions per model. The statistic is the field's folded maximum $\\max_\\ell |\\mathrm{AUROC}_\\ell - 0.5|$. The argument is carried by a generalizability-theory variance decomposition of that statistic over facets (model, wrapper, item sample, implementation), which yields $E\\rho^2 = 0.018$ for the current single-prompt design and shows the model-by-wrapper interaction is 13.4 times the model-by-item component, so reliability grows with prompt count and not with item count.","core_discovery":"The central claim is that a direction built by subtracting mean activations across contrastive prompts is principally a property of the contrastive prompt, not of the model it is extracted from. Concretely, with 36 wrappers formed by crossing 6 evaluation framings with 6 deployment framings around byte-identical task text, the correlation between the maximum-layer AUROC statistic and log parameters ranges from $-0.988$ to $+0.939$, and both the positive scaling claim of one published study and the opposite-sign claim of another fall inside this single design's range. At matched interior layers, the direction reproduces at cosine $0.9995$ when the item sample is split but only $0.3749$ when the wrapper changes. A label-permuted control shows a content-free direction attains 70\\% to 116\\% of each published score, and a bag-of-words classifier separates the scoring split at AUROC $0.993$ from surface form alone. The paper's conclusion is that reported scores are not comparable across models unless prompts are treated as a random facet and many of them are averaged over.","pith_inferences":["The same contrastive difference-of-means construction is used for refusal, truthfulness, and steering directions, so the wrapper-dependence shown here likely applies to those reported scores too; a direct test would cross evaluation-style framings in each domain and compare variance components.","If the score follows the prompt, any result correlating a probe score with model scale without randomizing over prompt wording—including safety-relevant monitoring claims—should be read as evidence about a phrasing, not about the model family.","The surface-form separability result suggests scoring sets themselves may leak evaluation status through markers such as 'choice', 'would', and 'say'; future detection datasets should be constructed adversarially to balance surface form across labels.","A practical standardization follows: benchmark comparisons should pre-register a sample of prompt wrappers drawn from published practice and report the full score distribution, turning the wrapper facet from a hidden choice into a measured component."],"forward_implications":["The sign of a scaling relationship between evaluation-awareness score and model size is not identifiable from any single wrapper; both published signs are inside the range one factorial design produces on every ladder tested.","A single-prompt comparison between models is unsupported: the generalizability coefficient for current practice is $E\\rho^2 = 0.018$, and about $k\\approx 9$--$10$ prompts are needed for a defensible rank comparison, rising to 56 evaluation-arm framings when both arms are crossed.","Averaging over more evaluation items cannot fix the measurement—the ceiling with infinite items is $E\\rho^2 \\approx 0.021$—whereas averaging over wrappers approaches $0.541$.","The scoring split is largely separable from surface form (bag-of-words AUROC $0.993$), so a direction carrying no evaluation information reproduces 70--116\\% of each published score; models must be judged against their own label-permuted floor.","Of the two reanalyzed published claims, the positive scaling relationship survives floor correction ($r = +0.600$ to $+0.561$), while the peak-layer-depth claim is not identified: resampling items moves the argmax across 26\\% of the layer stack at the median."],"supporting_citations":[{"why":"Supplies the generalizability-theory facet decomposition that turns prompt choice into a measured variance component and yields the $E\\rho^2$ and $k\\approx 9$ prescriptions.","marker":"Cronbach et al., 1972"},{"why":"Supplies the control-task discipline (label-permuted floor) adapted to an untrained zero-shot direction in Section 9.","marker":"Hewitt & Liang, 2019"},{"why":"Provides the published positive-scaling claim and released artifacts whose layer-selection floor correction tests whether the scaling result survives.","marker":"Chaudhary et al., 2025"},{"why":"Provides the opposite-sign/peak-depth claim, released direction vectors, contrastive set, and notebook used for the floor and identifiability reanalyses.","marker":"Manek, 2026"},{"why":"Supplies the 400-item stages-oversight scoring split used as the evaluation-versus-deployment test set.","marker":"Laine et al., 2024"},{"why":"Supplies the refusal split used as a second concept to show the floor results generalize beyond evaluation awareness.","marker":"Arditi et al., 2024"}],"fun_headline_variants":["Probe direction is a property of its prompt","One prompt flips the sign of scaling trend","Prompt choice, not model, drives probe scores","Swap the wording, flip the trend","Probe direction: more about prompt than model"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The quantitative claims ($E\\rho^2 = 0.018$, $k\\approx 9$) rest on the assumption that the 36 hand-written wrappers and the four-model Qwen2.5 ladder are representative enough of real researcher prompt choices and model families to give stable variance estimates; the paper states this directly.","fun_headline_variants_meta":{"raw":{"variants":["Probe direction is a property of its prompt","One prompt flips the sign of scaling trend","Prompt choice, not model, drives probe scores","Swap the wording, flip the trend","Probe direction: more about prompt than model"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00021,"raw_usage":{"total_tokens":1465,"prompt_tokens":1052,"completion_tokens":413,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":668,"completion_tokens_details":{"reasoning_tokens":344}},"tokens_in":668,"tokens_out":413,"duration_ms":3934,"temperature":1.0,"reasoning_tokens":344,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T13:32:50.584537+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Collect a fresh sample of prompts that researchers actually use in evaluation-awareness studies, cross them in the same factorial design on the same four models, and recompute the variance components; if the model share is large enough that $E\\rho^2$ reaches 0.80 with one wrapper per model, or if no wrapper choice moves the scaling correlation from positive to negative, the central claim would be overturned. A cheaper check is to hold a published wrapper fixed and search a large corpus of deployment phrasings to see whether the sign of the scaling trend is ever immovable.","supporting_citations":[{"cited_title":"Representational Depth of Evaluation Awareness Shifts With Scale in Open-Weight Language Models","cited_arxiv_id":"2606.29196","evidence_quote":"Provides the opposite-sign/peak-depth claim, released direction vectors, contrastive set, and notebook used for the floor and identifiability reanalyses."}],"review_version":1}