{"id":"75fdb464-2191-4c28-9011-7ace5e06f830","arxiv_id":"2507.21107","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Concern-shifted prompts alter residual-stream curvature and salience in two small LLMs, but the evidence for a meaningful, concern-specific signal is weak and partly tautological.","lead":"This paper measures how the internal activation path of a language model bends when a prompt word changes to carry more emotional or moral weight. It proposes curvature and salience as interpretability metrics, but the reported effects largely reduce to the expected consequence that different inputs produce different activations.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No synonym-swap or non-concern substitution baseline is reported, so the central claim that curvature tracks 'concern intensity' is confounded with generic token-edit sensitivity; a null-control experiment is the decisive check.","rationale":"The reader's weakest assumption is exactly the one I find load-bearing: concern shifts are single-word edits, and no non-concern edit baseline is provided. I agree with the reader's identification. The centrality follows because all headline quantitative results (Table 5, Figures 7-8) are deltas relative to neutral scaffolds; they establish only that edits change geometry, not that 'concern' as a semantic variable does so in a way distinct from arbitrary token substitution. The residual-stream site selection is also post hoc (Section 4), and the salience-curvature anticorrelation is plausibly a scaling artifact since kappa scales as ||v||^{-2} while S scales as ||v||, but these are secondary to the missing null-control condition. The paper's limitations section is honest about the omission, which strengthens confidence that the authors would run the control; but as written the central claim is not secured. I recommend no change to the reader's REJECT verdict. The concrete test above would settle whether the concern lands.","tokens_in":16139,"tokens_out":6075,"duration_ms":72981,"concrete_test":"Run the identical pipeline of Sections 3.3-4.3 on 20 matched synonym-swap controls, replacing each concern-shift word with a frequency-matched neutral alternative (e.g., 'desperately' -> 'routinely', 'repeatedly' -> 'frequently') while preserving scaffold length and part of speech, plus a random-token substitution condition. Compute per-prompt delta-kappa and delta-S deltas and the positive-polarity Str>Mod test exactly as in Table 5. If the median |delta-kappa| for synonym swaps overlaps or exceeds the concern-shift median, or if the p = 0.006 scaling result reproduces under the control permutations, the concern-specific interpretation fails; if the concern deltas lie clearly outside the control distribution, the central claim survives this challenge.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3.2 defines a concern shift as replacing one word in a neutral scaffold (e.g., 'practice your delivery repeatedly' -> 'practice your delivery desperately'), and Section 4.3 attributes the resulting delta-kappa and delta-S values to concern intensity. The only comparison class is the neutral control; there is no non-concern lexical-substitution control such as synonym swaps or structure-preserving random replacements. Any token substitution will generically move residual activations, changing both S(t) = ||x_{t+1} - x_t||_G and kappa_i (the finite-difference formula in Section 3.4.2 is nonlinear in x). Thus the observed LLaMA scaling (Table 5: positive curvature p = 0.006, salience p = 0.016) could equally arise from generic lexical sensitivity. The paper itself concedes this in Section 6: 'No null baselining... no scrambled, synonym-swapped, and random-weight prompts were tested. Without these additional baselines, curvature magnitudes remain relatively interpreted.' This is not a peripheral limitation; it is the construct-validity condition for the central claim. Without a null distribution of delta-kappa under matched non-concern edits, 'concern-sensitive geometry' is indistinguishable from 'any word edit bends the residual stream.'","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes 'Curved Inference,' a geometric-interpretability framework that measures curvature and salience of the residual-stream trajectory of transformer LLMs under a pullback metric G = U^T U induced by the unembedding matrix. Using 20 hand-curated matched prompt pairs across seven semantic domains, the authors compare neutral controls with 'concern-shifted' variants in Gemma3-1b and LLaMA3.2-3b. They report that concern shifts alter residual-stream curvature and salience in both models, with LLaMA showing statistically significant scaling between moderate and strong positive-polarity variants. The paper also presents correlation analyses between metrics, a 'Semantic Lens' interpretation of attention/MLP layers, and a discussion of limitations including the absence of null controls.","tokens_in":16392,"tokens_out":3665,"duration_ms":46904,"significance":"The idea of treating the residual stream as a curve and measuring semantic reorientation via curvature is potentially interesting and complements existing interpretability tools. The authors include useful details on finite-difference curvature estimation, acknowledge the need for native-space (rather than projected) analysis, and make their code and prompts available. However, the headline claim—that curvature and salience scale with 'concern intensity'—is not currently established, because the empirical design lacks the decisive control conditions and statistical corrections needed to rule out generic token-edit sensitivity. If the authors can supply those controls and temper the claims accordingly, the framework could become a useful diagnostic; as it stands, the evidence is suggestive but not convincing.","major_comments":[{"comment":"The central claim that curvature tracks 'concern intensity' is confounded with generic lexical-substitution effects. The concern-shifted prompts differ from controls by replacing a single word (e.g., 'repeatedly' → 'desperately'), but no synonym-swap, structure-preserving random substitution, or other non-concern lexical baseline is reported. Any token substitution generically moves residual activations and changes both salience S(t) and the finite-difference curvature κ_i, so the observed LLaMA scaling (Table 5: positive curvature p = 0.006, salience p = 0.016) could arise from non-semantic sensitivity to any word edit. The paper itself concedes in §6: 'No null baselining... no scrambled, synonym-swapped, and random-weight prompts were tested.' This is not a peripheral limitation; it is the construct-validity condition for 'concern-sensitive geometry.' Without a null distribution of Δκ and ΔS under matched non-concern edits, the central claim is indistinguishable from 'any word edit bends the residual stream.'","section":"§3.2, §4.3, Table 5, §6"},{"comment":"The statistical support for the scaling claim is limited to one polarity in one model. Of the eight model–metric–polarity combinations in Table 5, only LLaMA positive-polarity curvature (p = 0.006) and salience (p = 0.016) reach conventional significance; all Gemma rows and both LLaMA negative-polarity rows are non-significant (p ≥ 0.127). The Abstract's statement that 'LLaMA exhibiting consistent, statistically significant scaling' overstates this pattern, and no multiple-comparison correction is applied across the eight tests. Moreover, the residual-stream site and the G = U^T U metric were selected after examining the data (as described in §4: 'This was not assumed a priori—it emerged through a comparative analysis'), so the reported p-values do not account for this selection. The authors should report corrected p-values, disclose the selection procedure, and soften the headline claim accordingly.","section":"§4.3, Table 5, Abstract"},{"comment":"The 'core result' in Appendix C is tautological rather than a substantive geometric criterion. Stating that a shift-induced step Δv not colinear with the control step v produces positive curvature is just the definition of curvature for a polygonal path bent into a two-dimensional plane; it does not establish that such bending reflects 'semantic divergence' or 'concern'. Similarly, the converse statement (colinear shifts produce zero curvature) is a geometric identity. The appendix is therefore not evidence for the empirical claims, and invoking it as support for κ_i legitimizes a circular argument. The authors should reframe this appendix as a sanity check of the discretization, not as a derivation that concern-induced semantic divergence produces curvature.","section":"Appendix C"},{"comment":"Cross-model curvature magnitudes are not comparable as reported because the pullback metric G = U^T U scales with the norm of U, which differs across models. Table 3 shows mean curvatures of ~3.7e-5 for Gemma versus ~0.074 for LLaMA—a factor of roughly 2000—yet the paper interprets this as LLaMA exhibiting 'early, strong, sustained semantic curvature' while Gemma shows 'weak, shallow curvature.' Without normalizing by a model-specific scale (e.g., a per-model null distribution of κ under random or neutral inputs, or reporting standardized effect sizes), these cross-model differences may simply reflect differences in unembedding norms or residual-stream step sizes. The 'representational budget' interpretation of the salience–curvature anticorrelation (r = −0.89 for LLaMA) is likewise unfalsifiable as stated, since no alternative mechanism is specified or tested.","section":"§4.2, Table 3, Table 4"}],"minor_comments":[{"comment":"There is a typo in the sentence 'sharp spikes in response to certain concernt-shifted prompts'—'concernt' should be 'concern'.","section":"§5.2"},{"comment":"The paper has several missing spaces and formatting artifacts, e.g., 'We proposeCurved Inference' and 'Curved Inferencereveals' in the Abstract; these should be corrected throughout.","section":"Abstract and §1"},{"comment":"The table would be clearer if it reported the number of prompts per polarity and the exact test used (e.g., Wilcoxon signed-rank), as well as effect sizes, rather than only p-values and ratios.","section":"Table 5"},{"comment":"The Limitations section is admirably candid, but several of its concessions (e.g., 'curvature magnitudes remain relatively interpreted') directly contradict the strength of the claims in the Abstract and Conclusions; the authors should align the framing of the main text with these limitations.","section":"§6"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a single-author preprint with a somewhat exploratory character. The core geometric framework is plausible, but the empirical foundation needs substantial additional work—especially null controls and corrected statistics—before it can support the stated conclusions. I would encourage the editor to treat the current version as an early-stage contribution rather than a definitive study. The self-citation to the author's FRESH model is fine, but the paper would benefit from a more neutral positioning relative to that framework. There is also a question of scope fit: as written, the paper is closer to a methods proposal with a limited pilot than to a fully validated interpretability result."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a competent, reproducible pilot that is over-interpreted. The core observation—changing one word changes residual-stream geometry—is trivially true; the meaningful question is whether the change tracks 'concern' specifically rather than any token edit. That question is left open because the only control is the neutral scaffold. No synonym-swap or structure-preserving non-concern substitutions were run, and Section 6 concedes this. Until that null baseline exists, 'concern-sensitive geometry' is indistinguishable from 'any word edit bends the residual stream.'\n\nWhat's genuinely new: the specific empirical pattern—discrete curvature and salience under the pullback metric G=U^TU scaling with hand-coded concern intensity in LLaMA positive polarity (p=0.006, p=0.016)—isn't in the cited literature. The framing itself is standard differential geometry; the contribution is the matched-pair extension and the fact that code and data are available. The UMAP audit is a nice piece of negative method work.\n\nSoft spots, in order of severity. First, the missing null control is load-bearing, not peripheral. Second, residual stream was selected as the only interesting site after looking at three sites, then significance was reported without multiple-comparison correction—so the p-values are optimistic. Third, the salience–curvature anticorrelation (r=-0.89 for LLaMA) is largely a definitional artifact: curvature divides by ||v||^3, salience is ||v||, so speed and turniness trade off mechanically. The 'representational budget' story doesn't address that. Fourth, Appendix C's 'core result'—a non-colinear shift induces curvature—is just the definition of curvature, dressed up as a criterion. It doesn't add information.\n\nThe paper is honest about its limitations, which I respect. The math is elementary but correct, and the code/data availability is real evidence. For a reader in interpretability, this is a useful cautionary example of how a legitimate geometric measurement can be over-attributed to a semantic variable. It is not a demonstrated result.\n\nRecommendation: I would not accept this as-is, but I would not desk-reject it either. The missing control is a fixable flaw, not a conceptual one. Send it to review with a clear demand for synonym-swap and random-substitution baselines, correction for multiple comparisons, and a re-derivation of the salience–curvature relationship. If the authors do that, the paper could become a credible pilot. As it stands, it's a well-executed first draft of an experiment.","headline":"A clean, reproducible geometric probe whose central claim is unconfirmed for lack of a synonym-swap control—worth referee time if the fix is made.","tokens_in":16899,"tokens_out":3995,"would_cite":false,"duration_ms":44767,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Concern-shifted wording measurably bends the residual-stream trajectories of transformer language models, and in LLaMA the effect scales with concern intensity.","keywords":["residual stream","curvature","salience","semantic metric","pullback metric","large language model interpretability","prompt engineering","concern-shifted prompts"],"falsifier":"Replace each concern word in the 20 scaffolds with a frequency- and length-matched neutral synonym; if curvature and salience deltas of the same size appear, the concern-specific claim is false. Alternatively, redo the pullback metric with randomly permuted unembedding rows: if the LLaMA scaling survives, the signal is not semantic.","tokens_in":15906,"feed_emoji":"📐","tokens_out":10885,"duration_ms":105968,"temperature":0.7,"pith_summary":"Curved Inference claims that the way a large language model's internal state bends as it reads a prompt is a measurable, semantics-sensitive quantity. Across 20 matched prompt scaffolds, the paper finds that shifting one word toward emotional, moral, or logical concern changes the residual-stream trajectory in both models, and in LLaMA3.2-3b the change in curvature and salience grows systematically and significantly as concern intensity increases. If correct, this gives a geometric, model-grounded signal for when and how strongly a model is reorienting its representation, with practical uses in alignment monitoring and architectural comparison.","feed_headline":"Curvature of LLMs' inner paths grows with prompt concern","feed_subtitle":"Semantic one-word shifts change where and how hard an LLM's residual stream bends, most clearly in LLaMA.","key_machinery":"The carrier of the argument is the pullback semantic metric $G = U^{\\top} U$, built from the unembedding matrix $U$ that maps residual vectors to logits; it redefines distances and angles in residual space so they are aligned with the model's output vocabulary. Under this metric, salience is the first-order step norm $S(t) = \\|x_{t+1} - x_t\\|_G$, and curvature is the second-order reorientation $\\kappa_i = \\sqrt{\\|a_i\\|_G^2 \\|v_i\\|_G^2 - \\langle a_i,v_i\\rangle_G^2} / \\|v_i\\|_G^3$, estimated by 3-point finite differences of residual vectors across layers. This metric makes curvature coordinate-invariant and ties bending to token-level semantic change; the discrete curvature formula plus matched prompt scaffolds (control versus one-word concern-shifted variants) is what lets the paper attribute curvature deltas to semantic concern.","core_discovery":"Concern-shifted prompts do not merely change model output; they deform the internal trajectory. Measuring per-token, per-layer curvature $\\kappa_i$ and salience $S(t)$ in native residual space under the pullback metric $G = U^{\\top} U$, the paper reports that LLaMA3.2-3b shows early, statistically significant scaling of both metrics with concern strength, while Gemma3-1b responds to concern but differentiates moderate from strong variants only weakly. The paper interprets this as evidence for two interacting geometric layers: a static latent conceptual structure in the embedding and unembedding matrices, and a contextual trajectory in the residual stream that bends according to prompt-specific inference. It also treats attention and MLP blocks as semantic lenses whose integrated effect is visible as curvature.","pith_inferences":["A direct extension would add null baselines such as synonym swaps, scrambled word order, and random-weight forward passes; if curvature deltas survive these controls unchanged, the concern-specific reading would need revision.","The two-layer geometry suggests tracking tokens that are singular in embedding space through varied prompts to see which contexts amplify or neutralize their latent irregularity.","Early-layer curvature differences between the two models could partly reflect tokenization or positional-encoding differences, so testing under a shared tokenizer would separate semantic effect from preprocessing.","If the effect is genuinely semantic rather than token-level, translated versions of the same prompts should show similar curvature shapes, extending the paper's cross-lingual speculation."],"forward_implications":["Concern-induced bends appear at the same token-layer positions whether the wording is moderate or strong, but the bends grow larger with stronger concern, so curvature tracks graded semantic pressure rather than being a generic response to lexical change.","LLaMA3.2-3b and Gemma3-1b show different geometric styles: LLaMA bends early, strongly, and persistently, while Gemma bends shallowly near layers 5-6, giving a quantitative handle on cross-model comparison.","Because salience and curvature are strongly anticorrelated in LLaMA ($r=-0.89$), the paper argues for a representational trade-off: high reorientation tends to come with shorter total movement, suggesting a kind of internal representational budget.","Low-dimensional projections distort the curvature signal below the paper's fidelity threshold, so native-space residual metrics are required to observe the effect.","Curvature spikes on morally or emotionally charged prompts could serve as real-time alignment telemetry, flagging risky internal states before they surface in output."],"supporting_citations":[{"why":"This reference defines the residual stream as the trajectory whose curvature is measured.","marker":"[1]"},{"why":"This paper supplies an earlier geometric interpretation of transformer representations that grounds the trajectory view.","marker":"[10]"},{"why":"This work models residual-stream belief-state geometry, which the paper contrasts with matched-trajectory curvature divergence.","marker":"[11]"},{"why":"This model card supplies the Gemma3-1b checkpoint and architecture analysed.","marker":"[14]"},{"why":"This model card supplies the LLaMA3.2-3b checkpoint and architecture analysed.","marker":"[15]"},{"why":"This work shows that token embeddings violate the manifold hypothesis, motivating the latent-versus-contextual geometry split.","marker":"[16]"},{"why":"This work shows cross-model embedding alignment, supporting a shared latent geometry across architectures.","marker":"[17]"},{"why":"This work establishes a privileged basis for the residual stream, supporting the metric-based semantic geometry.","marker":"[19]"}],"fun_headline_variants":["LLM residual paths bend more with prompt concern","Concern shifts warp LLM internal geometry","Residual stream curvature tracks semantic concern","LLaMA's inner bends scale with prompt concern","Geometry of LLM thought curves with concern"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that swapping a single word in a matched scaffold changes only the intended semantic concern, and that the output-projection geometry built from the unembedding matrix turns residual-stream bending into a measure of meaning-change rather than a generic reaction to any token substitution.","fun_headline_variants_meta":{"raw":{"variants":["LLM residual paths bend more with prompt concern","Concern shifts warp LLM internal geometry","Residual stream curvature tracks semantic concern","LLaMA's inner bends scale with prompt concern","Geometry of LLM thought curves with concern"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000475,"raw_usage":{"total_tokens":2360,"prompt_tokens":948,"completion_tokens":1412,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":564,"completion_tokens_details":{"reasoning_tokens":1344}},"tokens_in":564,"tokens_out":1412,"duration_ms":11378,"temperature":1.0,"reasoning_tokens":1344,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T19:03:33.999765+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Replace each concern word in the 20 scaffolds with a frequency- and length-matched neutral synonym; if curvature and salience deltas of the same size appear, the concern-specific claim is false. Alternatively, redo the pullback metric with randomly permuted unembedding rows: if the LLaMA scaling survives, the signal is not semantic.","supporting_citations":[{"cited_title":"Exploring the Residual Stream of Transformers","cited_arxiv_id":null,"evidence_quote":"This reference defines the residual stream as the trajectory whose curvature is measured."}],"review_version":1}