{"id":"b458079f-ae3f-4a53-b026-fef3c9a3ffc2","arxiv_id":"2607.03598","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Models linearly represent Gricean communicative intent (recognize vs evaluate) from pretraining, but default readout often discards it; a late-layer discriminative direction recovers honoring as well as an explicit instruction.","lead":"Language models encode a sender's communicative intent in their hidden states even when default replies ignore it. Steering that internal direction recovers the intended behavior without a prompt, showing the failure is readout, not understanding.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.5","headline":"Causal handle is a searched late logistic direction, not shown to be the peak-probe representation itself; the untested factorial leaves the 'routes a represented intent' link under-specified.","rationale":"The reader's weakest_assumption correctly isolates the softest link in the causal half of the strongest claim. Representation evidence (surface-matched leave-one-phrasing-out, bag-of-words at chance, request-matched, valence transfer, base checkpoints, inferred intent, multi-family probe ceilings) is multi-controlled and solid. Behavioral discard/recovery is real and non-overlapping on the three models where the gap is open, with specificity controls (random/shuffled directions, near-orthogonality to feedback axis, opener-biasing refutations) that largely hold. The paper already flags the untested factorial and scopes secondary axes (support/help steering fails specificity) and model-specificity honestly. That is precisely why CONDITIONAL (not ACCEPT) is appropriate; the concern does not introduce a new internal inconsistency or require a harsher verdict. No stronger load-bearing flaw (e.g., lexical leakage surviving controls, or the measure inventing the discard) survives the reported checks. Thus the reader's verdict and confidence stand.","tokens_in":19549,"tokens_out":626,"duration_ms":19560,"concrete_test":"On Qwen2.5-3B (n=60 recognize set, same dose/sanity protocol as Section 6), extract and steer (1) the logistic-weight direction at layer 24 and (2) the difference-of-means direction at layer 30. Report bootstrap honoring CIs for both. If neither recovers (or only the original logistic@30 combination does), the handle is not simply the early representation moved later; if logistic@24 recovers cleanly, the layer search was unnecessary and the representational tie is tighter than currently shown.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The represent-then-lagging-readout claim needs the steering vector to be the same (or tightly corresponding) feature the early probe decodes. Section 6 states difference-of-means at the peak-probe layer (24) fails the sanity gate and can worsen discard, while recovery requires the logistic-weight direction at a later searched layer (30). The logistic-at-24 / DoM-at-30 factorial is explicitly untested, so it is unknown whether the effective handle is the early-decodable intent feature or a later, possibly distinct computation that merely correlates with the labels. Section 10 localization (probe saturates before steering onset) is consistent with lag but also with a different late feature. Without geometric linkage (cosine between early and late directions) or the missing cells, 'closely tied to the representation' and 'routes a represented intent rather than a generic feedback knob' rest on correlation plus successful late steering, not isolation of the early feature as the causal object.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper treats a sender's communicative intent (primarily recognize vs evaluate) as a first-class linear feature of language-model residual streams. A linear probe decodes that intent surface-independently from default-pass last-token activations across six models and four families, including base (pre-instruct) checkpoints and cases where intent must be pragmatically inferred; bag-of-words under leave-one-phrasing-out CV is at chance. Default behavior discards the intent on three of six models (unsolicited feedback on recognize shares). Where the gap is open, the logistic-weight direction at a per-model searched late layer is a causal handle: steering recovers honoring with a clean dose-response, matches an explicit intent prompt, is near-orthogonal to a feedback-behavior axis, and beats matched-norm random and difference-of-means controls. Depth sweeps show the probe saturates several layers before steering recovers. A second, lexically clean inferred axis (support vs help) is represented but does not yield a specific causal handle. Controls, human validation of the honoring measure, and nulls (including an inconclusive pre-registered geometry test) are reported throughout.","tokens_in":19838,"tokens_out":1436,"duration_ms":18262,"significance":"If the result holds, it reframes a common deployment failure (models answering surface content rather than what the sender was doing) as a readout problem on top of a robust, pretraining-learned representation, and it introduces a sender's goal—distinct from character ToM beliefs or user attributes—as a linear, steerable interpretability object. Strengths that raise the contribution above a standard probe-and-steer paper include: surface-matched identical suffixes with leave-one-phrasing-out and object-held-out CV; request-matched and valence-disentanglement construct controls; base-checkpoint results; human rater validation (Fleiss κ=0.76) of the behavioral measure; specificity against random, shuffled-label, difference-of-means, and opener-unembedding directions; near-orthogonality to the feedback axis; honest scoping of the support/help null and the inconclusive pre-registered geometry test; and released code/stimuli with a CPU-reproducible core chain. The represent-versus-readout decomposition with depth localization is a useful template for other pragmatic features.","major_comments":[{"comment":"Section 6 (and the localization in Section 10 / Table 3): the central claim that steering 'routes a represented intent' rests on a direction that is only 'closely tied' to the early probe feature. Difference-of-means at the peak-probe layer (24) fails the sanity gate and can worsen discard; recovery requires the logistic-weight direction at a later searched layer (30 on Qwen-3B). The paper correctly flags that the logistic-at-24 / DoM-at-30 factorial is untested, so it is unknown whether the effective handle is the early-decodable intent feature or a later, possibly distinct computation that merely correlates with the labels. Without either the missing factorial cells or a geometric link (e.g., cosine between the peak-probe direction and the validated steer direction, or transfer of the early direction to the late layer), 'closely tied to the representation' and 'routes a represented int","section":"Section 6; Section 10 / Table 3"},{"comment":"Section 7 / Table 2: for four of the six models the steer layer and coefficient are selected in-sample on the evaluation items; only Qwen-3B and Llama-8B receive a nested object-split validation. The paper flags this, but the non-overlapping recovery claims for Qwen-7B (and the ceiling claims for the other three) are therefore weaker than for the two split models. Either run the nested split for the remaining discard model(s) or demote those rows more clearly to exploratory status so the across-model stratification is not over-read.","section":"Section 7 / Table 2"}],"minor_comments":[{"comment":"Section 10: the paper notes that probe-before-steer is common even for acted-on features and that a matched non-discarded control is missing. Consider elevating this caveat one notch in the main text (not only the localization section), since the depth gap is used as part of the represent-then-lag narrative.","section":"Section 10"},{"comment":"Section 9 / Appendix F: on Qwen-3B the body-reorientation test is inconclusive at the 100-token budget because separation dilutes; the refutation there rests on near-orthogonality alone. State this more prominently when claiming the handle is not opener-token biasing across all three discard models.","section":"Section 9; Appendix F"},{"comment":"Section 8 / Appendix J: support-vs-help is correctly scoped to representation only after the specificity control fails. A one-sentence pointer in the abstract or contributions that the causal half of the story is established only on recognize/evaluate (and vent/solve) would prevent over-reading the third axis.","section":"Section 8; Abstract"},{"comment":"Figure 1 and Table 1: bag-of-words is reported as 0.48 in the figure caption and main text but the per-model table (Appendix K) shows 0.46–0.48; keep a single consistent value or note the range.","section":"Figure 1; Table 1; Appendix K"},{"comment":"Reproducibility: the Modal scripts and GitHub link are welcome; ensure the pre-registration document for the inconclusive geometry test (Appendix E) is in the release so the decision rule can be audited.","section":"Appendix E; Appendix M"}],"recommendation":"minor_revision","confidential_remarks":"The paper is unusually careful about nulls and self-limits for this literature; the main risk is that the abstract's 'routes a represented intent' phrasing slightly outruns the untested factorial in Section 6. I would not reject on that basis—the authors already hedge in the body—but I would require the geometric link or the factorial before accept. Fit for a strong empirical interpretability / CL venue is good; novelty of the object (sender goal) is real even if the steering machinery is standard."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The punchline is simple: models encode a sender’s recognize-vs-evaluate intent cleanly and surface-independently, including from base checkpoints and when the intent is only inferred, yet some of them still default to unsolicited feedback. Where that gap is open, a late discriminative direction recovers honoring about as well as an explicit prompt, and it is near-orthogonal to a pure feedback axis.\n\nWhat is actually new is the object, not the machinery. Probing and steering are standard; treating a sender’s communicative goal (not a character’s belief or a user’s attribute) as a linear, causal residual-stream feature, with a surface-matched leave-one-phrasing-out design and a represent-versus-readout split, is not. The controls are better than average for this genre: bag-of-words at chance, permutation ceilings, request-matched and valence-transfer checks, human rater validation of the honoring measure (Fleiss κ ≈ 0.76), dose-response with CIs, random-direction and opener-biasing nulls, and explicit reporting of the support/help steering failure and the inconclusive pre-registered geometry test. Code and stimuli are released. That honesty is load-bearing.\n\nThe soft spot that matters is the one the stress-test flags. Difference-of-means at the peak-probe layer fails; recovery needs the logistic-weight direction at a searched later layer, and the paper never runs the factorial that would show which change carries the effect. So “routes a represented intent” is a bit looser than the abstract sells—it is a direction closely tied to the labels at a late layer, not proven to be the early decodable feature itself. Depth localization is consistent with lag but also with a later, correlated computation. Secondary axes are weaker; discard is only 3/6 models; naturalistic n is small and author-written. None of that sinks the central empirical claim for recognize/evaluate on the discard models, but it does keep the causal story one notch short of airtight.\n\nThis is for people who care about pragmatic failures in assistants and about whether “the model doesn’t get it” is representation or readout. It deserves a serious referee. I would bring it to reading group and cite the represent-vs-readout framing and the surface-matched protocol. Send it out.","headline":"Careful probe-and-steer work that makes sender intent a real object; the early-representation to late-handle link is under-specified but the paper is honest about it.","tokens_in":20416,"tokens_out":580,"would_cite":true,"duration_ms":9660,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Language models represent what you meant; they often fail to act on it.","keywords":["communicative intent","linear probing","activation steering","readout lag","pragmatics","Gricean intent","residual stream","language models"],"falsifier":"On the three discard models, show that no direction extracted from the peak-probe layer (or a controlled factorial of logistic versus difference-of-means at peak versus late layers) can recover recognize-honoring without collapsing coherence or failing the specificity and near-orthogonality controls; or show that default honoring is already at ceiling once stimuli fully block request-detection and valence confounds.","tokens_in":20402,"feed_emoji":"🎯","tokens_out":709,"duration_ms":5321,"temperature":0.7,"pith_summary":"When you share something with a language model, it often answers the surface of the message and misses what you were doing by sending it: a finished project gets unsolicited critique, a late-night vent gets a risk assessment. This paper treats the sender's communicative intent as a first-class object of study and shows the failure is usually not ignorance. Across six open models in four families, a linear probe cleanly recovers the intent (recognize versus evaluate) from the model's default hidden states, even when the intent is never stated and must be inferred from context, and even in base pretraining checkpoints. Within a model the intent is decodable several layers before it drives the output; across models, whether the default reply honors it is model-specific. Where the gap is open, steering a direction closely tied to that representation recovers the intended behavior as well as an explicit instruction does, with no prompt, and the direction is near-orthogonal to a generic feedback knob. The practical upshot is a reframe: the model often already has what you meant; the fragile part is whether the readout routes it.","feed_headline":"Models know what you meant; they often ignore it","feed_subtitle":"Intent is linearly readable from pretraining; only some models act on it by default","key_machinery":"The represent-then-lagging-readout chain on a surface-matched recognize-versus-evaluate contrast: leave-one-phrasing-out linear probes on default-pass hidden states, depth localization of decodability versus steerability, and causal recovery by adding the discriminative (logistic-weight) intent direction at a per-model late layer.","core_discovery":"A sender's communicative intent is robustly represented in the residual stream of language models and can be read out linearly, surface-independently, from pretraining onward, including when the intent must be pragmatically inferred; the common failure is a lagging or model-specific readout that does not act on that representation. Where default behavior discards the intent, a discriminative direction at a searched late layer is a causal handle that recovers honoring.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Models encode intent robustly but often fail to act on it","Intent is linearly readable; behavior lags or is model-specific","LMs represent what you meant more reliably than they use it","Hidden states hold intent; default output frequently skips it","Models know the sender's goal yet only some honor it by default"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"That the steering direction is closely enough tied to the early-decoded representation that recovery counts as routing what was already represented, even though difference-of-means at the probe peak fails and the paper never isolates whether the later layer or the logistic direction is what makes the handle work.","fun_headline_variants_meta":{"raw":{"variants":["Models encode intent robustly but often fail to act on it","Intent is linearly readable; behavior lags or is model-specific","LMs represent what you meant more reliably than they use it","Hidden states hold intent; default output frequently skips it","Models know the sender's goal yet only some honor it by default"]},"model":"grok-4.5","effort":"low","cost_usd":0.003992,"raw_usage":{"total_tokens":1340,"prompt_tokens":916,"num_sources_used":0,"completion_tokens":68,"cost_in_usd_ticks":39920000,"prompt_tokens_details":{"text_tokens":916,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":356,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":916,"tokens_out":68,"duration_ms":3227,"temperature":1.0,"reasoning_tokens":356,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-12T01:16:58.939648+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"On the three discard models, show that no direction extracted from the peak-probe layer (or a controlled factorial of logistic versus difference-of-means at peak versus late layers) can recover recognize-honoring without collapsing coherence or failing the specificity and near-orthogonality controls; or show that default honoring is already at ceiling once stimuli fully block request-detection and valence confounds.","supporting_citations":[],"review_version":1}