REVIEW 3 major objections 4 minor 21 references
Positional influence in causal residual Transformers follows an exact replicator equation along gradient flow, so primacy, recency, and lost-in-the-middle are conditional, checkable outcomes rather than universal architectural laws.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 17:14 UTC pith:BR3AVNOI
load-bearing objection Careful, honest theory paper: real analytic results for positional influence, with the fragile regularity assumption disclosed rather than buried. the 3 major comments →
An Adjoint-Sensitivity Framework for Lost-in-the-Middle Phenomena in Causal Residual Transformers
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper's core claim is that positional influence can be captured by the normalized adjoint-energy density m_s^GD(p) (Definition 2.12), and that along full-batch gradient flow this density obeys the exact replicator equation ∂_s m = m(r̄ − r_p). From this starting point, the paper proves an exact generator-term decomposition of the input adjoint into residual transmission, nonlocal Volterra-cone, and local channels, retaining all covariance cross terms; and it states Theorem 3.7, which gives sufficient channel-energy and correlation bounds under which the regional middle average is smaller than both boundary averages — i.e., the lost-in-the-middle index (74) is positive. The same framework
What carries the argument
The central object is the normalized adjoint-energy influence density m_s^GD(p), the Radon–Nikodym derivative of E∥P_0(p)∥², the expected squared input-adjoint energy at position p. The load-bearing identities are the exact gradient-flow evolution ∂_s m = m(r̄ − r_p) (Section 3.1) and the exact regional decomposition (136) of the terminal adjoint into residual, cone, and local terms with cross-covariances. This decomposition, together with the Volterra adjoint formula (88) for the causal-cone transport and the Duhamel identity (106) for residual persistence, is what turns primacy, recency, and lost-in-the-middle into quantities that can be bounded by independently estimable channel energies.
Load-bearing premise
The entire gradient-flow influence dynamics and the proposed regularizers depend on the lifted attention-plus-feed-forward vector field being twice continuously Fréchet differentiable in states and parameters, with uniformly bounded first and second derivatives along the trajectory (Proposition 3.1's 'additional sensitivity hypothesis'); the paper does not prove this from the architectural primitives, it excludes non-smooth activations like ReLU, and it may fail for heavy-tai
What would settle it
Run a small causal residual transformer with smooth activations under full-batch gradient descent, compute the normalized adjoint-energy density m_s at several training times, and check whether ∂_s m equals m(r̄ − r_p) pointwise to within numerical precision; a systematic violation would falsify the exact replicator claim. For non-smooth activations (e.g., ReLU), testing whether the density dynamics deviate measurably would delimit the assumption's scope.
If this is right
- If the replicator equation holds, positional bias is not a static property of attention weights but a dynamical quantity: relative adjoint-energy decay rates decide which positions gain influence during training.
- The exact Volterra adjoint transport gives a quantitative primacy channel: early positions receive contributions from all later positions in their causal cone, so the cone mass C(q) is a measurable proxy for left-boundary influence.
- Residual identity paths transmit terminal adjoint energy to depth zero, so a right-biased readout or task protocol (e.g., last-token evaluation) produces recency, but only when the terminal right-middle contrast exceeds the interaction error quantified in Proposition 3.5.
- Lost-in-the-middle is equivalent to positivity of the index LIM_δ (74), and Theorem 3.7 converts this into sufficient conditions on channel lower bounds, middle upper bounds, and correlation constants that can be estimated from finite-token Jacobians.
- The three proposed regularizers (influence balancing, positional reweighting, observability balancing) all act on well-defined surrogate measures with stated computational costs, but the paper shows that the outer-loop variants need not monotonically reduce the influence-based diagnostic.
Where Pith is reading between the lines
- If the replicator dynamics are robust, an early-training diagnostic based on r_p(p) — the positionwise adjoint-energy decay rate — could flag emerging position bias before it is visible in downstream retrieval scores.
- The sufficient-channel-energy criterion suggests a practical model audit: estimating the correlation constants ρ_{αβ,B} from data, rather than taking the pessimistic Cauchy–Schwarz value 1, could make the lost-in-the-middle certificate much sharper than the paper's worst-case bounds.
- The framework invites testing on non-smooth activations: since the replicator equation is proved under twice-differentiability, a subdifferential or smoothing version would determine whether the exact dynamics survive in standard ReLU transformers.
- Read together with the paper's own conditional cautions, the analysis implies that empirical U-shaped benchmark curves should be treated as protocol-dependent outcomes (initialization, data law, readout placement) rather than inevitable transformer behavior.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper develops a continuous-depth model of a causal residual Transformer and uses adjoint sensitivity to define a positional influence density m_s^GD(p). The main unconditional results are a residual-to-ODE convergence bound (Theorem 2.8), finite-token-to-Volterra attention approximation (Theorem 2.7), and a discrete-to-continuum adjoint consistency estimate (Theorem 2.9). Under an explicitly stated twice-Fréchet differentiability hypothesis (Prop. 3.1), the paper derives the normalized density's exact gradient-flow evolution (69) and rewrites it as replicator dynamics, then splits the input adjoint into residual/cone/local channels with exact regional energy identities (136). Primacy, recency, and lost-in-the-middle are characterized as conditional sufficient conditions (Prop 3.4, Prop 3.5, Theorem 3.7). Three interventions—influence balancing, positional reweighting, observability regularization—are proposed with complexity and bias caveats. Small algebraic simulations demonstrate control of direct surrogates.
Significance. If the assumptions are met, this would be a useful analytic framework: it separates unconditional approximation theory from mechanism-specific claims, gives exact identities that can be checked empirically, and avoids overclaiming universality. The paper's explicit limitation statements and reproducible algebraic experiments are strengths. The main weakness is that the central gradient-flow evolution is conditional on a twice-Fréchet differentiability hypothesis that is not proved for any concrete architecture and excludes ReLU; because this hypothesis underpins Eq. (69), the replicator dynamics and all regularizers that rely on it are not yet established for standard Transformers.
major comments (3)
- [Sec. 3.1, Prop. 3.1, Eq. (69)] The exact influence-density dynamics and subsequent regularizers depend on the assumption that G^ϑ(t,X)=ζ(t)F_{θ^ϑ_t}(X) is twice continuously Fréchet differentiable in (X,ϑ) with uniform derivative bounds on the trajectory set. This is not proved from architectural primitives; Assumption 2.1 only requires bounded first derivatives of σ, and Prop 2.4 establishes only Lipschitz regularity. ReLU activations violate the hypothesis. The authors flag this as an 'additional sensitivity hypothesis', but it remains load-bearing: without it, ϑ↦I^ϑ(p) need not be differentiable and (69) is not classically well-defined. The manuscript should either prove the hypothesis for a concrete smooth architecture (e.g., GeLU/softplus with regularized LN) or state an explicit example class satisfying it; otherwise the central replicator equation is a conditional statement with no verified instantiations.
- [Sec. 3.5, Theorem 3.7] The theorem is called 'Independently checkable channel-energy criterion', but the manuscript does not provide a finite-sample procedure or sample-complexity bound for estimating the channel upper/lower bounds u_{α,B}, ℓ_L, ℓ_R and correlation constants ρ_{αβ,B} from finite-token data. As written, the theorem certifies only that if such constants are known, the regional inequality follows. Since Proposition 3.4's κ_L lower bound requires joint coercivity over all (t,u,p,r), the practical verification of these bounds appears as hard as the original question. Please state what 'checkable' means and provide at least a direct finite-token estimator with consistency conditions for each constant.
- [Sec. 3.1, Eq. (60) and definition of R(ϑ)] The paper writes the full-batch gradient flow (60), but Fréchet differentiability of the population risk R(ϑ) = E[ℓ(Y,H^ϑ(X^ϑ_T))] with respect to the finite-dimensional parameter state is not established, nor are domination conditions for differentiating under the data expectation. Proposition 3.1 proves differentiability of I^ϑ, not of R. Since the replicator equation (69) uses ∇R(ϑ_s) explicitly, a missing proof or explicit assumption that R is C^1 with bounded gradient on U leaves a gap in the chain-rule derivation along the trajectory.
minor comments (4)
- [Sec. 2.2, Theorem 2.7] Typo: 'C_{R,η} < ∞ 1' should be 'C_{R,η} < ∞'. Also, the double subscript notation is unusual; suggest writing C_{R,η} consistently.
- [Sec. 2.4 and Sec. 3.2] The symbol E_L is used both for the cellwise density embedding (46) and for the left region E_L=[0,δ) in Sec. 3.2. This overloaded notation is confusing; rename the embedding, for example Emb_L.
- [References] References [3], [9], [14], [15], and [17] appear in the bibliography but are not cited in the text. Please cite them or remove them.
- [Sec. 4.5] The text refers to 'the warning in Step 7 of Algorithm 3', but Algorithm 3 has only steps 1–6. The step number should be corrected.
Circularity Check
Framework is self-contained; only the LIM-index equivalence is a labeled definitional restatement, and self-citations are not load-bearing.
specific steps
-
self definitional
[Theorem 1.1(v); Section 3.2, Eq. (74)]
"Lost-in-the-middle at scale δ is equivalent to positivity of the lost-in-the-middle index(74)."
The index (74) is defined as LIMδ(s) = Δδ(s)/(min{m_L,m_R}+ε0) with Δδ(s) = min{m_L,m_R} − m_M. Therefore LIMδ>0 is exactly the condition m_L>m_M and m_R>m_M, which is the paper's own definition of Lost-in-the-Middle ('Primacy means m_L>m_M, recency means m_R>m_M, and Lost-in-the-Middle means both'). The theorem restates the definition rather than deriving an independent fact; the paper itself labels such statements 'exact algebraic characterization' in Corollary 3.6.
full rationale
The derivation is self-contained. The influence density (58) is defined directly from the adjoint of the continuous-depth model, and the gradient-flow identity (69) is obtained by differentiating the normalized quotient m=I/Z along (60); the paper explicitly says 'The identity (69) is exact but not closed,' and the decay rate r_s(p) is defined from the same gradient, so the replicator form is an algebraic consistency identity, not a fitted prediction. Primacy, recency, and Lost-in-the-Middle conclusions are conditional on independently checkable channel-energy and correlation bounds (Proposition 3.4, Theorem 3.7), not on fitted constants. The self-citations [12,13] occur only in notation and positioning ('scalar depth profile as in [12]', 'explicit-Euler notation of [12]') and are not used to prove the main theorems; no uniqueness theorem is imported from the authors' prior work. The simulations are explicitly algebraic toy checks and disclaim any verification on shared-parameter Transformers. The sole definitional reduction is Theorem 1.1(v), where the LIM index is defined to have the same sign as the regional gap, making the stated equivalence a transparent restatement. Aside from this labeled algebraic equivalence, no load-bearing circularity was found.
Axiom & Free-Parameter Ledger
axioms (6)
- domain assumption Admissible parameter set Θ_ad is closed/bounded; regularized layer norm and activation have bounded first derivatives on bounded sets (Assumption 2.1).
- domain assumption Depth profile ζ is nonnegative, bounded, Lipschitz; depth-control path θ_t is Lipschitz/piecewise Lipschitz with bounded values (Assumption 2.2).
- ad hoc to paper Twice continuous Fréchet differentiability of G^ϑ(t,X) in (X,ϑ) with uniformly bounded first/second derivatives on the bounded trajectory set (Prop 3.1).
- ad hoc to paper L4 differentiability of P_0^ϑ into L^4(D⊗Leb) with uniform bound and Z(ϑ) ≥ z_0 > 0 (Prop 3.2, eq. (65)–(66)).
- domain assumption Data law has finite fourth moments E[‖X0‖^4 + ‖Y‖^4] < ∞ (eq. (62)).
- standard math Standard analytical tools: Hardy's inequality, Gronwall's lemma, Fubini, Cauchy–Schwarz, implicit function theorem.
read the original abstract
We develop an adjoint-sensitivity framework for positional influence in causal residual Transformers and separate unconditional analytic results from conditional boundary-shape conclusions. The principal unconditional theorem is the residual-to-depth-flow estimate for layer controls converging in $L^1$, complemented by a finite-token-to-Volterra attention estimate that explicitly controls the first cells near the causal endpoint. We define a normalized adjoint-energy influence density and derive its exact evolution along full-batch gradient flow. The adjoint admits an exact generator-term decomposition into residual transmission, nonlocal Volterra, and local channels, including all covariance cross terms. Causal masking can amplify early-position sensitivity and residual identity paths can transmit a right-localized terminal bias, but neither mechanism alone forces a U-shaped profile. We therefore state boundary advantages under independently checkable energy, correlation, and local-channel bounds; these conditions are sufficient rather than necessary. Finite-token influence balancing, positional reweighting, and task-aligned observability are presented as diagnostics or regularizers with explicit differentiation requirements, computational costs, and limitations. Controlled simulations illustrate that each intervention controls its designated surrogate, while observability balance or outer-loop reweighting need not monotonically reduce the influence-based Lost-in-the-Middle diagnostic.
Figures
Reference graph
Works this paper leans on
-
[1]
Quantifying attention flow in Transformers
Samira Abnar and Willem Zuidema. Quantifying attention flow in Transformers. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4190–4197, 2020.doi:10.18653/v1/2020.acl-main.385
-
[2]
Zico Kolter
Shaojie Bai, Vladlen Koltun, and J. Zico Kolter. Stabilizing equilibrium models by Jacobian regularization. InProceedings of the 38th International Conference on Machine Learning, volume 139 of PMLR, pages 554–565, 2021
2021
-
[3]
Alain Bensoussan, Tak Kwong Wong, Sheung Chi Phillip Yam, and Hongwei Yuan. A theory of first order mean field type control problems and their equations.Journal of the European Mathematical Society, published online first, 2026.doi:10.4171/JEMS/1781
-
[4]
Edoardo Calvello, Nikola B. Kovachki, Matthew E. Levine, and Andrew M. Stuart. Contin- uum attention for neural operators. arXiv:2406.06486, 2024.doi:10.48550/arXiv.2406. 06486
-
[5]
Ricky T. Q. Chen, Yulia Rubanova, Jesse Bettencourt, and David Duvenaud. Neural ordinary differential equations. InAdvances in Neural Information Processing Systems,
-
[6]
Borun D. Chowdhury. Lost in the middle at birth: An exact theory of Transformer position bias. arXiv:2603.10123, 2026.doi:10.48550/arXiv.2603.10123
-
[7]
A mean-field optimal control formulation of deep learning.Research in the Mathematical Sciences, 6:10, 2019
Weinan E, Jiequn Han, and Qianxiao Li. A mean-field optimal control formulation of deep learning.Research in the Mathematical Sciences, 6:10, 2019. doi:10.1007/ s40687-018-0172-y
2019
-
[8]
G. H. Golub and C. F. Van Loan.Matrix Computations. 4th ed., Johns Hopkins University Press, 2013. 40
2013
-
[9]
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 770–778, 2016.doi:10.1109/CVPR.2016.90
-
[10]
Judy Hoffman, Daniel A. Roberts, and Sho Yaida. Robust learning with Jacobian regular- ization. arXiv:1908.02729, 2019.doi:10.48550/arXiv.1908.02729
-
[11]
Found in the middle: Calibrating positional attention bias improves long context utilization
Cheng-Yu Hsieh, Yung-Sung Chuang, Chun-Liang Li, Zifeng Wang, Long Le, Abhishek Kumar, James Glass, Alexander Ratner, Chen-Yu Lee, Ranjay Krishna, and Tomas Pfister. Found in the middle: Calibrating positional attention bias improves long context utilization. InFindings of the Association for Computational Linguistics: ACL 2024, pages 14982–14995, 2024.do...
-
[12]
A First-Order Mean Field Control Analysis of Transformer Layers under Cross-Entropy Training
Cheng Huan and Hongwei Yuan. A first-order mean field control analysis of Transformer layers under cross-entropy training. arXiv:2606.23235, 2026.doi:10.48550/arXiv.2606. 23235
work page internal anchor Pith review Pith/arXiv arXiv doi:10.48550/arxiv.2606.23235 2026
-
[13]
A mean-field analysis of multi-head self-attention under cross-entropy training
Cheng Huan and Hongwei Yuan. A mean-field analysis of multi-head self-attention under cross-entropy training. arXiv:2606.10469, 2026.doi:10.48550/arXiv.2606.10469
-
[14]
On uniform-in-time diffusion approximation for stochastic gradient descent
Lei Li and Yuliang Wang. On uniform-in-time diffusion approximation for stochastic gradient descent. arXiv:2207.04922, 2022.doi:10.48550/arXiv.2207.04922
-
[15]
Qianxiao Li, Cheng Tai, and Weinan E. Stochastic modified equations and dynamics of stochastic gradient algorithms I: Mathematical foundations.Journal of Machine Learning Research, 20(40):1–47, 2019. arXiv:1811.01558.doi:10.48550/arXiv.1811.01558
-
[16]
Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang
Nelson F. Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang. Lost in the middle: How language models use long contexts. Transactions of the Association for Computational Linguistics, 12:157–173, 2024.doi: 10.1162/tacl_a_00638
-
[17]
RoFormer: Enhanced Transformer with rotary position embedding
Jianlin Su, Yu Lu, Shengfeng Pan, Ahmed Murtadha, Bo Wen, and Yunfeng Liu. RoFormer: Enhanced Transformer with rotary position embedding. arXiv:2104.09864, 2021. doi: 10.48550/arXiv.2104.09864
-
[18]
Axiomatic attribution for deep networks
Mukund Sundararajan, Ankur Taly, and Qiqi Yan. Axiomatic attribution for deep networks. InProceedings of the 34th International Conference on Machine Learning, volume 70 of PMLR, pages 3319–3328, 2017
2017
-
[19]
Gomez, Lukasz Kaiser, and Illia Polosukhin
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. InAdvances in Neural Information Processing Systems, 2017. arXiv:1706.03762.doi:10.48550/arXiv. 1706.03762
-
[20]
Efficient streaming language models with attention sinks
Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han, and Mike Lewis. Efficient streaming language models with attention sinks. InInternational Conference on Learning Representations, 2024. arXiv:2309.17453.doi:10.48550/arXiv.2309.17453. 41
-
[2018]
arXiv:1806.07366.doi:10.48550/arXiv.1806.07366
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.