Pith. sign in

REVIEW 1 major objections 7 minor 191 references

Depth can replace missing numerical precision only relative to a declared low-bit library, horizon, execution arithmetic, and routing model—and a structural floor no amount of depth can cross.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-30 23:29 UTC pith:NOXJIYPF

load-bearing objection Conditionally sound resource theory that cleanly separates structural floor, pure-depth synthesis, arithmetic phase, and pre-training certificates; the tube hypothesis is the real bridge to practice, not a hidden proof gap. the 1 major comments →

arxiv 2607.23390 v1 pith:NOXJIYPF submitted 2026-07-25 cs.LG math.OC

When Can Depth Replace Precision? A Resource Theory of Quantized Neural Computation

classification cs.LG math.OC
keywords quantized neural networksdepth–precision tradeoffsresidual networksrelaxed controlserror feedbackstructural floormixture-of-experts routingreachability
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper asks when stacking more low-bit residual steps can stand in for the numerical precision a network is missing, for a fixed input–output map. It treats a quantized residual network as a pure schedule that picks fields from a declared low-bit operation library over a fixed horizon, and characterizes the infinite-depth limit by relaxed controls. The distance from the target map to the closed relaxed reachable set is an exact structural floor: no optimizer and no extra depth can remove it for that library. Pure schedules approach the floor at a first-order rate under ordinary time regularity, but execution arithmetic can reverse the story—full-state write-back can freeze residual updates—while increment error feedback keeps a bounded carry and an exact lattice conservation law. For coherent high-precision comparators with first-order error, matching accuracy forces student depth to scale linearly with teacher depth. Primal and dual certificates can mark a design feasible, impossible, or unresolved before training.

Core claim

For a fixed target map and a declared low-bit residual library, the exact asymptotic limit of infinite low-bit depth is the distance from the target to the closed relaxed reachable set generated by that library. That distance is a structural floor no pure schedule can cross. Finite pure depth approaches the floor at rate O(1/D) under bounded-variation time dependence (and a Hölder-adjusted rate otherwise), but only when residual increments remain numerically visible; full-state write-back can add a growing penalty and freeze updates, while increment error feedback replaces that growth by a bounded carry. When a coherent high-precision comparator also has first-order error and the floor is ze

What carries the argument

The structural floor E_Ω,∞(F★): the distance from the target full map to the closed relaxed reachable set of the declared low-bit dictionary family. It separates what the library can express from what finite pure depth, metadata, arithmetic, and routing cost; pure schedules approach it by balanced switching / online simplex rounding, and verified primal–dual bounds turn the floor plus finite-resource radii into feasible / impossible / unresolved decisions.

Load-bearing premise

Everything ideal and implemented must stay inside one verified tube where the residual fields stay bounded and Lipschitz, so the theory does not cover attention or normalization that blow up, or arithmetic that overflows that tube.

What would settle it

Find a fixed target and declared low-bit residual library whose structural floor is zero, with coherent first-order high-precision error, yet whose best pure low-bit schedules either stay bounded away from the floor as depth grows or match the high-precision accuracy with depth growing much slower or much faster than linear in the comparator depth—under the paper’s execution and tube assumptions.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Before training, a dual lower bound above the tolerance certifies that no depth or optimizer can hit the target with that library.
  • Full-state activation write-back can make deeper low-bit nets worse; preserving residual increments (e.g. error-feedback carry) is required for depth to help.
  • Accuracy matching against a coherent first-order high-precision teacher forces low-bit depth on the order of teacher depth when the floor is zero.
  • Learned codebooks must be charged as metadata separate from schedule depth; logarithmic metadata bits can keep codebook error commensurate with first-order synthesis.
  • Hard routing only keeps a first-order depth law under isolated transversal events and a small-gain route–state loop, not under a frozen positive margin alone.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Quantization toolchains could add a pre-training screen that brackets the structural floor and rejects libraries whose dual bound already exceeds the tolerance.
  • Hardware paths that quantize the full residual state each microstep are in a different phase from increment-carry designs; kernel choice may matter as much as nominal bit width.
  • The same floor-plus-radius logic could grade looped or unrolled blocks: only refinements that stay coherent with a shared residual horizon earn a depth–precision exchange.
  • Unresolved certificates become a research queue of their own—pointing at which bound (floor, arithmetic, or route) must tighten next rather than treating failed training as non-representability.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

1 major / 7 minor

Summary. The paper develops a target-specific resource theory for when low-bit residual depth can replace numerical precision. A depth-D student is modeled as a pure schedule over a declared low-bit residual-field dictionary on a fixed horizon, with the state lifted to the full input-indexed map. The distance from the target to the closed relaxed reachable set is identified as the exact structural floor; pure schedules approach it at O(D^{-1}) under bounded-variation time dependence and O(D^{-ϑ}+D^{-1}) under ϑ-Hölder dependence (Theorems 3, 5). Execution arithmetic is shown to change the phase: full-state write-back contributes a Dρ_z certificate term with an exact scalar freeze result (Proposition 7), while increment error feedback telescopes the carry (Theorem 8) and admits a bit-exact common-lattice realization with explicit register widths (Proposition 9). A fixed, D-independent binary teacher has a closed-form optimal error Θ(D^{-1}) (Theorem 10), lifted to residual-ReLU and nonuniform two-token attention realizations (Proposition 12, Theorem 13), yielding D_match = Θ(L) for coherent first-order comparators (Corollary 15). Learned codebooks add a metadata resource with upper, packing, and allocation laws (Theorems 18, 19, 57); state-dependent routing is treated by a transversal-event small-gain theorem (Theorem 23). A primal–dual stack (HJB, support, affine, occupation-measure/SOS) yields feasible/impossible/unresolved decisions (Corollary 28). Companion software (QReplace) and

Significance. If the results hold, the paper supplies a unifying conditional framework that cleanly separates library geometry, synthesis depth, metadata, execution arithmetic, and routing — resources that the literature often collapses into a nominal bit width. Several strengths deserve explicit credit: (i) an exact closed-form fixed-teacher optimum with a nonasymptotic envelope (Theorem 10), making the first-order depth price sharp for one fixed target rather than only minimax; (ii) a Lean 4 artifact with hash-locked build logs and claim-level axiom audits kernel-checking twelve discrete-core statements, plus exhaustive executable verifiers for the attention converse and common-lattice arithmetic; (iii) a certified nonlinear matrix-valued accuracy-matching depth (D_match = 8 = 2L) proved by rational piecewise-affine bounds, with a prospective falsifiable prediction (calibrated D=9 vs certified 8) that was checked; (iv) an explicit evidence hierarchy and trust-boundary discussion that is more disciplined than typical for this area. The component tools (relaxed controls, sigma–delta feedback, occupation measures, hybrid transversality) are mature, and the paper says so; the contribution is the t

major comments (1)
  1. [§3.2 Assumption 1; §9.2–9.3; Theorems 6, 8, 29] The common synthesis tube is the load-bearing bridge for the master law (1) and for every QReplace verdict, and it carries a bootstrap risk the manuscript should address more directly. Assumption 1 simultaneously asserts forward invariance of K for all measurable relaxed controls, all mixed/pure Euler states and interpolation segments, and — via Theorems 6, 8, and 29 — all implemented finite-arithmetic prefixes, together with uniform B and L_z on the enlarged tube K_ρ. For attention blocks, the QK-product Lipschitz constant is state-range dependent (§9.2.2 bounds scores through B_Q, B_K), so the constants that define the tube are valid only on a tube whose existence is part of the hypothesis. The paper resolves this constructively for the contractive soft-threshold class (§9.1, Eq. (197) gives an explicit invariant ball), but the Transformer specialization provides only componentwise err
minor comments (7)
  1. [§8.5, Theorem 27] The main text flags that primal–dual equality requires a 'closed-image qualification detailed in Appendix A; that qualification is not automatic.' Please clarify in the main text that this qualification affects only the no-duality-gap statement, not the validity of dual lower certificates: Theorem 25 is proved directly by monotonicity along trajectories, so Corollary 28's 'certified impossible' verdicts do not depend on strong duality. As written, a reader could over-discount the decision rule or, conversely, over-credit the SOS hierarchy.
  2. [§10.8] The phrase 'causal validation of the theory's central distinction' is stronger than the design supports. The coherent-target arm (depth-D Euler refinement converging to the depth-32 refinement of the same field) is essentially standard Euler convergence and is expected a priori; the informative arm is the direct-target divergence. The 4-bit study uses three QAT seeds, so the Student-t intervals have two degrees of freedom, and the layer-1 hidden-map interval crosses zero (reported, but only mid-paragraph). Please temper the causal language, state what outcome would have falsified the mechanism, and note that the fitted slopes (−1.16, −0.86) come from five depth points.
  3. [Appendix A vs. main text] Notation drift: the appendices use C_{fh,b} and C^{unif}_{fh} where the main text uses C^{end}_{syn} and C^{unif}_{syn} (e.g., Theorem 30 vs. Theorem 3; Theorem 53 vs. Theorem 18). Please harmonize or add the correspondence to Table 3. Similarly, Φ_L(T) is defined twice (Eq. (14) and after Eq. (270)), and Eq. (65) uses u = D^{-1} in the main text but x = 1/D in Appendix A.2.
  4. [§1.4 vs. §8.7] QReplace is described in §1.4 as returning five outcomes (certified feasible, certified impossible, conditionally feasible, diagnostically promising, unresolved), while Corollary 28 defines a three-way decision. Please state the mapping between the two lists and which outcomes are proof-backed versus heuristic.
  5. [§4.2, Figure 2(a)] The write-back term Dρ_z is a worst-case certificate envelope; the only exact freeze result is scalar (Proposition 7). The text acknowledges this, but the 'phase diagram' framing of Figure 2(a) may be read as asserting realized U-shaped behavior in general architectures. A sentence clarifying that no multidimensional lower bound realizing the Dρ_z growth is known would calibrate Claim 3.
  6. [Eq. (36)] The α_k√q/2 term assumes a coordinatewise uniform activation grid in a q-dimensional Euclidean state; please state this explicitly, since the surrounding development is in the lifted Banach space Z.
  7. [References] A substantial fraction of the citations are 2025–2026 arXiv preprints (e.g., Chakrabarti et al. 2026; Park et al. 2026b; Zhao et al. 2026). Please indicate which are peer-reviewed and pin versions, since several novelty-boundary comparisons depend on them.

Circularity Check

0 steps flagged

No significant circularity: structural floor, synthesis rates, and converses are independently derived, not fitted or self-referential.

full rationale

The paper's load-bearing chain is definitional-plus-theorem, not circular. The structural floor E_Ω,∞(F*) is defined as dist(F*, R_Ω,rel); Theorems 3/5 then prove pure schedules approach that closed set at O(D^{-1}) or O(D^{-ϑ}+D^{-1}), so the distance is the asymptotic floor by Hausdorff convergence rather than by renaming the target. The fixed-teacher exact error (Theorem 10), residual-ReLU and attention embeddings, common-lattice conservation (Proposition 9), and packing/metadata laws are closed-form or constructive arguments under stated assumptions; they do not fit parameters from the quantities they claim to predict. Empirical sections are explicitly tiered as diagnostic and separated from theorem claims. Lean 4 checks and rational certificates are independent verification, not self-citation load-bearing. Assumption 1 (common tube) is a strong applicability hypothesis, not a circular step. No self-definitional loop, fitted-as-prediction, or uniqueness-via-author-citation pattern is present.

Axiom & Free-Parameter Ledger

3 free parameters · 8 axioms · 3 invented entities

The theory is conditional on declared modeling resources: a finite low-bit field library, shared residual horizon, lifted full-map state, verified execution tubes, and stated arithmetic/routing semantics. Standard analysis and hybrid-control tools are imported; the paper does not claim unconditional replacement of precision by depth for arbitrary trained blocks.

free parameters (3)
  • Library- and tube-dependent constants (B, Lz, Lt/Vt or Ht, ϑ, Csyn, LΩ, ho z/ ho D, ηG, route u r,Δr,χ)
    These are problem-declared bounds used in rates and certificates; not universal constants. Empirical sections estimate some (e.g. stability budgets) for diagnostics only.
  • Metadata metric dimension m and entropy prefactor CΩ
    Enter the 2^{-s/m} codebook term when a compact learned dictionary family is assumed; chosen from the declared parameter geometry.
  • Per-microstep resource charge ℓσ and total budget B
    Accounting units for joint depth-metadata allocation; must be declared consistently (bits, latency, energy).
axioms (8)
  • standard math Banach-space Carathéodory existence/uniqueness for bounded, strongly measurable, uniformly state-Lipschitz relaxed fields on a forward-invariant tube (Assumption 1 / Lemma 60).
    Standard ODE-in-Banach hypotheses; paper states them explicitly because endpoint sets require well-posedness.
  • standard math Online simplex/prefix discrepancy rounding with bound < J per coordinate (Lemma 2), used to get first-order pure-to-relaxed rates.
    Belongs to mixed-integer control / combinatorial integral approximation lineage the paper cites.
  • standard math Superposition principle for continuity equations / occupation measures (Ambrosio et al.) to equate modal measure programs with mixtures of relaxed trajectories (Theorem 27).
    Used for exactness of the infinite-dimensional primal; duality needs extra closed-image qualification the paper flags.
  • domain assumption Common verified tube containing ideal and implemented prefixes; Lipschitz and defect bounds only claimed on that tube (Assumption 1; Theorem 6).
    Load-bearing for all synthesis and hardware-transfer rates; not automatic for attention/LayerNorm globally.
  • domain assumption Coherent fixed-horizon residual refinement family shared by high-precision comparator and low-bit student when stating D=Θ(L) matching (Sections 1.1, 5, Corollary 15).
    Without coherence, subdividing unrelated blocks need not approach the original map; DistilBERT control illustrates failure.
  • domain assumption Isolated transversal top-k route events with small-gain χ<1 and isolation radii (Assumption 22 / Theorem 23); excludes simultaneous/grazing/chattering regimes.
    Hybrid necessity for first-order routed laws; frozen positive-margin arguments are explicitly insufficient for actual swaps.
  • domain assumption Increment quantizer residual radius and non-overflow of carry/state registers on declared ranges for error-feedback and common-lattice exactness (Theorem 8, Proposition 9).
    Arithmetic conclusions are semantics-relative; wrapping without invariant is excluded.
  • ad hoc to paper Deterministic global metadata + input-independent pure schedules for packing lower bound (Theorem 19); no uncharged continuous side information.
    Accounting convention matching the upper metadata theorem; other programming models need different resource ledgers.
invented entities (3)
  • Structural floor E_Ω,∞(F*) = dist(F*, closed relaxed reachable set of declared dictionary family) independent evidence
    purpose: Exact asymptotic invariant separating library expressivity from finite depth/optimization
    Defined from standard reachable-set geometry but elevated as the central neural resource invariant for one target and library.
  • Schedulewise arithmetic radius A_D and write-back vs error-feedback phase diagram independent evidence
    purpose: Quantify when execution preserves or reverses depth benefit
    Operational aggregate of transition defects; falsifiable via freeze diagnostic and register audits.
  • QReplace feasible/impossible/unresolved decision rule combining [L,U] floor bracket with finite-resource radius R_{D,s} independent evidence
    purpose: Pre-training certification interface
    Workflow/software object packaging the theorems; correctness depends on sound certificates fed in.

pith-pipeline@v1.2.0-grok45-kimik3 · 62537 in / 3941 out tokens · 76901 ms · 2026-07-30T23:29:41.500477+00:00 · methodology

0 comments
read the original abstract

When can additional low-bit residual computation replace missing numerical precision for a fixed input-output map? We model a quantized residual system over a fixed horizon as a pure schedule selecting fields from a declared low-bit operation library, and use relaxed controls to characterize its infinite-depth limit. The distance from the target to the closed relaxed reachable set is the exact structural floor: no increase in depth can remove it for that library. Pure schedules approach the relaxed class at rate $O(D^{-1})$ under bounded-variation time dependence and $O(D^{-\vartheta}+D^{-1})$ under Holder dependence of exponent $\vartheta$. Execution arithmetic can reverse this conclusion: full-state write-back introduces a $D\rho_z$ penalty and can freeze residual updates, whereas increment error feedback replaces this growth by a bounded carry term and obeys an exact common-lattice conservation law. A fixed-teacher converse makes this rate sharp: for coherent depth-$L$ first-order high-precision comparators, accuracy matching requires $D=\Theta(L)$. Learned codebooks add a metadata resource, while state-dependent routing introduces hybrid event conditions. Verified primal and dual bounds yield feasible, impossible, or unresolved decisions before training. Companion software implements the workflow, and Lean 4 machine-checks the exact discrete core. Depth replaces precision only relative to a declared library, horizon, execution semantics, and routing model.

Figures

Figures reproduced from arXiv: 2607.23390 by Mojtaba Soltanalian.

Figure 1
Figure 1. Figure 1: The resource-theoretic view. The low-bit library determines the relaxed geometry and therefore the structural floor. Depth synthesizes that geometry with a pure schedule, metadata selects or describes the library, and the execution semantics determine whether the microsteps remain numerically visible. Hard routing couples state and event-time errors. Primal and dual certificates compare the resulting achie… view at source ↗
Figure 2
Figure 2. Figure 2: Execution arithmetic creates a genuine phase diagram. (a) Ideal synthesis ap￾proaches the structural floor, whereas fixed-grid state write-back has a U-shaped envelope and a finite optimal depth. Increment error feedback preserves the first￾order law when its physical residual unit scales with the microstep. (b) In an exact scalar diagnostic, a fixed activation grid eventually erases every residual update;… view at source ↗
Figure 3
Figure 3. Figure 3: Reachability is a geometric property of the declared operation library. Pure depth￾D schedules form a discrete cloud. Balanced switching drives that cloud toward the relaxed reachable set. A compatible target has zero structural floor and only a finite-depth remainder; an incompatible target remains separated even as D → ∞. write-back, saturating integer arithmetic, two’s-complement wrapping, or an auxilia… view at source ↗
Figure 4
Figure 4. Figure 4: A fixed target can require a first-order depth price. (a) The exact optimum in Theorem 10 tracks its a/(2eD) asymptote and remains inside the proved nonasymptotic envelope. (b) Multiplication by D exposes the nonzero limiting constant a/(2e); the lower law is not generated by changing the target with D. All curves are analytic consequences of the theorem, not fitted slopes. Proof. Apply ℓ to the pure Euler… view at source ↗
Figure 5
Figure 5. Figure 5: Depth and global codebook metadata are distinct resources. (a) The excess-error surface Csyn/D + Cmeta2 −s/m for a representative m-dimensional family. The white-backed curve balances the two terms and is not an empirical fit. (b) Under the declared unit-cost budget in Section 6.3, logarithmic metadata keeps its error commensurate with the synthesis term and preserves the first-order budget law. For total … view at source ↗
Figure 6
Figure 6. Figure 6: A route-changing diagnostic consistent with Theorem 23. Left: a direct score perturbation makes the implemented route switch one grid point early, but transver￾sality confines the mismatch to the shaded certified event window. Right: the resulting state error is visible and remains below the small-gain envelope. The plotted curves are a deterministic theorem illustration, not a claim about a partic￾ular Mo… view at source ↗
Figure 7
Figure 7. Figure 7: The decision interface. A primal upper bound and a verified dual lower bound define an interval for the implemented target error after adding the finite-resource radius. The interval can certify feasibility, finite-depth impossibility, or asymptotic impossibility; overlap with the required tolerance is reported as unresolved rather than converted into a post-hoc verdict. 9 Architecture Specializations The … view at source ↗
Figure 8
Figure 8. Figure 8: An exact fixed-target attention converse. (a) The two-token state decomposes into a mean mode, which carries the binary scheduling obstruction, and a disagreement mode, which experiences a nonuniform attention contraction with eigenvalue µ = tanh(1). (b) The closed-form global optimum in (81) agrees with exhaustive schedule enumeration through depth 12 and approaches its predicted first-order asymptote. To… view at source ↗
Figure 9
Figure 9. Figure 9: Exact structural split in a soft-threshold residual network. Under one finite dictio￾nary, a compatible nonlinear target converges toward zero, while an incompatible target converges to the exact infinite-depth floor αs/6. The dashed floor is analytic and the finite-depth points are globally solved. 101 102 low-bit microdepth D 10−3 10−2 max fixed-witness terminal error Reachable target compatible target D… view at source ↗
Figure 10
Figure 10. Figure 10: A state-dependent attention structural split. The compatible target follows a first-order decay, whereas a target with a component in a missing output subspace remains at an exact invariant floor. The right panel displays a nonuniform, input-dependent attention matrix from the experiment. 54 [PITH_FULL_IMAGE:figures/full_fig_p054_10.png] view at source ↗
Figure 11
Figure 11. Figure 11: verifies the minimax D−1 law in the signed-ray soft-threshold family. Globally solved finite programs coincide with the analytic edge-gap formula at every tested depth. At D = 128, the exact values of DED are 0.2371, 0.1185, 0.0790, and 0.0339 for binary, ternary, uniform 2-bit, and uniform 3-bit gain alphabets. 101 102 depth D 10−3 10−2 minimax error Global optimum matches the edge formula 101 102 depth … view at source ↗
Figure 12
Figure 12. Figure 12: First-order synthesis around a nonlinear attention field. Compatible gain al￾phabets all produce D−1 terminal error, while DED approaches a dictionary￾dependent constant. The experiment separates the universal exponent from the nonuniversal depth price. Inside the feasible region, exact finite-depth count synthesis produces the matching depths in [PITH_FULL_IMAGE:figures/full_fig_p056_12.png] view at source ↗
Figure 13
Figure 13. Figure 13: Certified soft-threshold feasibility boundary and precision sweet spot. Orange cells have certified infinite-depth floors above the requested tolerance and are asymptotically infeasible. The plot does not infer finite-depth impossibility from the floor alone; isolated shallow exceptions are checked by the accompanying exact search. Blue cells report globally solved matching depths. After codebook and sche… view at source ↗
Figure 14
Figure 14. Figure 14: Attention-output codebook geometry determines the infinite-depth floor. The target is compared with the convex hull generated by each finite codebook. Binary and uniform 2-bit codebooks leave a visible geometric gap; the tested 3-bit polygon contains the target. This feasibility calculation is completed before any finite-depth search. 2 4 8 16 high-resolution depth L (tolerance 1/L) Binary Ternary 2-Bit 3… view at source ↗
Figure 15
Figure 15. Figure 15: Attention-codebook phase boundary and matched-resource optimum. The left panel separates positive-floor asymptotic infeasibility from finite matching depths. The right panel shows that the resource-optimal precision can be interior and can change with the target tolerance. affine propagation over inscribed and circumscribed 16-gons proves inf σ∈{0,1} 7 sup y∈B2 2 (1) [PITH_FULL_IMAGE:figures/full_fig_p05… view at source ↗
Figure 16
Figure 16. Figure 16: Finite-horizon amplification controls the depth multiplier. Left: the observed large-depth constant DED grows with the empirical stability budget and with more aggressive normalization/residual scaling. Right: the first tested depth attaining error 2.5 × 10−3 . Stability changes the constant, not the first-order mechanism. 25 50 75 100 125 low-bit depth D 0.1 0.2 0.3 minimum hard-routing margin A routing-… view at source ↗
Figure 17
Figure 17. Figure 17: Hard routing has a margin-controlled phase transition. The dashed boundary is the quantized-router perturbation radius. Below it, route changes create a persistent error; above it, routes agree and the D−1 synthesis regime reappears. sup y∈B2 2 (1) [PITH_FULL_IMAGE:figures/full_fig_p059_17.png] view at source ↗
Figure 18
Figure 18. Figure 18: Certified minimum accuracy-matching depth in a nonlinear matrix system. The large left panel excludes every D ≤ 7 and certifies that the selected depth-8 student is more accurate relative to the common depth-32 reference than the depth-4 high-precision comparator. The two right panels show the comparator and low-bit error fields relative to the depth-32 reference on a common color scale. The proof uses ex… view at source ↗
Figure 19
Figure 19. Figure 19: Pretrained-model causal control for depth coherence. The same DistilBERT feed-forward residual fields are evaluated under two targets. (a) Subdividing the original one-step block moves away from its map even without quantization. (b) Refining a common depth-32 residual-flow target reduces error rapidly at both tested layers. The contrast isolates target coherence before any low-bit optimization is introdu… view at source ↗
Figure 20
Figure 20. Figure 20: Four-bit learned dictionaries preserve the coherent depth benefit. Means and Student-t 95% intervals are computed across three independent QAT seeds for locally hardened pure schedules. (a) Hidden-map error decreases modestly at the near-saturated first layer and more clearly at layer 3. (b) Logit RMSE improves at both layers, with a seed-robust depth effect at layer 3. (c) Prediction agreement with the c… view at source ↗
Figure 21
Figure 21. Figure 21: Medium-scale attention validation. In an eight-token, width-16, two-head attention–MLP system, a compatible target preserves the D−1 mechanism, while a component in a missing output direction remains at its exposed floor. Curves show held-out errors over an input ensemble fixed before the depth sweep. The right panel omits the final reference-integration-limited point when estimating the first-order plate… view at source ↗
Figure 22
Figure 22. Figure 22: Attention scale sweep. Left: the representative large-depth constant DED for compatible targets. Right: the exposed missing-direction floor. The horizontal labels report (n, d, H): token count, model width, and number of heads. 8 10 12 14 16 18 20 depth D 10−5 10−4 10−3 10−2 finite-grid error Representability is not trainability 8 10 12 14 16 18 20 depth D 100 101 102 optimization gap ratio Heuristics can… view at source ↗
Figure 23
Figure 23. Figure 23: Optimization gap for recurrent soft-threshold atoms. Globally exhaustive schedule search reveals representations that local coordinate descent misses by orders of magnitude. This experiment demonstrates that failed local schedule search is not a certificate of positive-floor asymptotic infeasibility; more generally, no failed optimizer establishes nonrepresentability by itself. A broader eight-teacher mat… view at source ↗
Figure 24
Figure 24. Figure 24: Order and optimization matter for noncommuting attention atoms. Exhaustive global optima, balanced schedules, and local-search distributions are reported separately. The theorem supplies a constructive representability law; it does not guarantee that an arbitrary training heuristic reaches the frontier. numerical result are listed in [PITH_FULL_IMAGE:figures/full_fig_p066_24.png] view at source ↗
Figure 25
Figure 25. Figure 25: Relaxed-to-pure compilation gap in the pretrained experiment. Bars show (Ehard/Erelaxed − 1) × 100% for locally refined pure schedules. The small gaps demonstrate that the principal positive result is realized by pure programs rather than only by convex mixtures. The compact release contains the complete seed-level trajectories, candidate metrics, confidence intervals, depth contrasts, per-example arrays,… view at source ↗
Figure 26
Figure 26. Figure 26: Supplementary breadth evidence. (a) The eight-teacher matrix campaign contains exact matches, certified feasible upper bounds, positive-floor impossibilities, and unresolved cases. (b) Expanding the complete dictionary from two to four atoms improves the globally optimized error by a median factor of roughly 1.7 and by more than 5 in the strongest cases. Dictionary geometry, not nominal bit count alone, c… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

191 extracted references · 18 canonical work pages · 3 internal anchors

  1. [1]

    Communications on Pure and Applied Mathematics , volume =

    Ingrid Daubechies and Michel Defrise and Christine De Mol , title =. Communications on Pure and Applied Mathematics , volume =. 2004 , doi =

  2. [2]

    SIAM Journal on Imaging Sciences , volume =

    Amir Beck and Marc Teboulle , title =. SIAM Journal on Imaging Sciences , volume =. 2009 , doi =

  3. [3]

    Proceedings of the 27th International Conference on Machine Learning , pages =

    Karol Gregor and Yann LeCun , title =. Proceedings of the 27th International Conference on Machine Learning , pages =

  4. [4]

    Advances in Neural Information Processing Systems , volume =

    Xiaohan Chen and Jialin Liu and Zhangyang Wang and Wotao Yin , title =. Advances in Neural Information Processing Systems , volume =

  5. [5]

    Eldar , title =

    Vishal Monga and Yuelong Li and Yonina C. Eldar , title =. IEEE Signal Processing Magazine , volume =. 2021 , doi =

  6. [6]

    Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , pages =

    Kaiming He and Xiangyu Zhang and Shaoqing Ren and Jian Sun , title =. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , pages =. 2016 , doi =

  7. [7]

    Ricky T. Q. Chen and Yulia Rubanova and Jesse Bettencourt and David Duvenaud , title =. Advances in Neural Information Processing Systems , volume =

  8. [8]

    Inverse Problems , volume =

    Eldad Haber and Lars Ruthotto , title =. Inverse Problems , volume =. 2018 , doi =

  9. [9]

    Sander and Pierre Ablin and Gabriel Peyr

    Michael E. Sander and Pierre Ablin and Gabriel Peyr. Do Residual Neural Networks Discretize Neural Ordinary Differential Equations? , booktitle =. 2022 , doi =

  10. [10]

    Jack Warga , title =

  11. [11]

    Young , title =

    Laurence C. Young , title =

  12. [12]

    Advances in Neural Information Processing Systems , volume =

    Matthieu Courbariaux and Yoshua Bengio and Jean-Pierre David , title =. Advances in Neural Information Processing Systems , volume =

  13. [13]

    Journal of Machine Learning Research , volume =

    Itay Hubara and Matthieu Courbariaux and Daniel Soudry and Ran El-Yaniv and Yoshua Bengio , title =. Journal of Machine Learning Research , volume =

  14. [14]

    European Conference on Computer Vision , pages =

    Mohammad Rastegari and Vicente Ordonez and Joseph Redmon and Ali Farhadi , title =. European Conference on Computer Vision , pages =. 2016 , doi =

  15. [15]

    International Conference on Learning Representations , year =

    Yukun Ding and Jinglan Liu and Jinjun Xiong and Yiyu Shi , title =. International Conference on Learning Representations , year =

  16. [16]

    Advances in Neural Information Processing Systems , volume =

    Yaniv Blumenfeld and Dar Gilboa and Daniel Soudry , title =. Advances in Neural Information Processing Systems , volume =

  17. [18]

    Journal of Machine Learning Research , volume =

    Hongyu Wang and Shuming Ma and Lingxiao Ma and Lei Wang and Wenhui Wang and Li Dong and Shaohan Huang and Huaijie Wang and Jilong Xue and Ruiping Wang and Yi Wu and Furu Wei , title =. Journal of Machine Learning Research , volume =

  18. [22]

    International Conference on Learning Representations , year =

    Peter O'Connor and Max Welling , title =. International Conference on Learning Representations , year =

  19. [23]

    Annals of Mathematics , volume =

    Ingrid Daubechies and Ronald DeVore , title =. Annals of Mathematics , volume =. 2003 , doi =

  20. [24]

    C. Sinan G. One-Bit Sigma--Delta Quantization with Exponential Accuracy , journal =. 2003 , doi =

  21. [25]

    IEEE Transactions on Information Theory , volume =

    Felix Krahmer and Rayan Saab and Rachel Ward , title =. IEEE Transactions on Information Theory , volume =. 2012 , doi =

  22. [26]

    Advances in Neural Information Processing Systems , volume =

    Zechun Liu and Changsheng Zhao and Hanxian Huang and Sijia Chen and Jing Zhang and Jiawei Zhao and Scott Roy and Lisa Jin and Yunyang Xiong and Yangyang Shi and Lin Xiao and Yuandong Tian and Bilge Soran and Raghuraman Krishnamoorthi and Tijmen Blankevoort and Vikas Chandra , title =. Advances in Neural Information Processing Systems , volume =

  23. [28]

    Advances in Neural Information Processing Systems , volume =

    Adrian Bulat and Yassine Ouali and Georgios Tzimiropoulos , title =. Advances in Neural Information Processing Systems , volume =. 2024 , doi =

  24. [29]

    Advances in Neural Information Processing Systems , volume =

    Yamato Arai and Yuma Ichikawa , title =. Advances in Neural Information Processing Systems , volume =

  25. [30]

    Advances in Neural Information Processing Systems , volume =

    Banseok Lee and Dongkyu Kim and Youngcheon You and Youngmin Kim , title =. Advances in Neural Information Processing Systems , volume =

  26. [31]

    Proceedings of the 43rd International Conference on Machine Learning , year =

    Shigeng Wang and Chao Li and Yangyuxuan Kang and Jiawei Fan and Anbang Yao , title =. Proceedings of the 43rd International Conference on Machine Learning , year =

  27. [32]

    Gomez and Lukasz Kaiser and Illia Polosukhin , title =

    Ashish Vaswani and Noam Shazeer and Niki Parmar and Jakob Uszkoreit and Llion Jones and Aidan N. Gomez and Lukasz Kaiser and Illia Polosukhin , title =. Advances in Neural Information Processing Systems , volume =

  28. [34]

    Advances in Neural Information Processing Systems , volume =

    Biao Zhang and Rico Sennrich , title =. Advances in Neural Information Processing Systems , volume =

  29. [35]

    Advances in Neural Information Processing Systems , volume =

    Subhabrata Dutta and Tanya Gautam and Soumen Chakrabarti and Tanmoy Chakraborty , title =. Advances in Neural Information Processing Systems , volume =

  30. [36]

    Bulletin of the American Mathematical Society , volume =

    Borjan Geshkovski and Cyril Letrouit and Yury Polyanskiy and Philippe Rigollet , title =. Bulletin of the American Mathematical Society , volume =. 2025 , doi =

  31. [37]

    Proceedings of the 38th International Conference on Machine Learning , series =

    Hyunjik Kim and George Papamakarios and Andriy Mnih , title =. Proceedings of the 38th International Conference on Machine Learning , series =

  32. [38]

    Proceedings of the 38th International Conference on Machine Learning , series =

    George Dasoulas and Kevin Scaman and Aladin Virmaux , title =. Proceedings of the 38th International Conference on Machine Learning , series =

  33. [39]

    How Smooth Is Attention? , booktitle =

    Val. How Smooth Is Attention? , booktitle =

  34. [40]

    Proceedings of the 38th International Conference on Machine Learning , series =

    Yihe Dong and Jean-Baptiste Cordonnier and Andreas Loukas , title =. Proceedings of the 38th International Conference on Machine Learning , series =

  35. [41]

    Advances in Neural Information Processing Systems , volume =

    Xinyi Wu and Amir Ajorlou and Yifei Wang and Stefanie Jegelka and Ali Jadbabaie , title =. Advances in Neural Information Processing Systems , volume =. 2024 , doi =

  36. [42]

    Consensus Is All You Get: The Role of Attention in Transformers , booktitle =

  37. [43]

    Mahoney and Kurt Keutzer , title =

    Sehoon Kim and Amir Gholami and Zhewei Yao and Michael W. Mahoney and Kurt Keutzer , title =. Proceedings of the 38th International Conference on Machine Learning , series =

  38. [44]

    Proceedings of the 40th International Conference on Machine Learning , series =

    Guangxuan Xiao and Ji Lin and Mickael Seznec and Hao Wu and Julien Demouth and Song Han , title =. Proceedings of the 40th International Conference on Machine Learning , series =

  39. [45]

    International Conference on Learning Representations , year =

    Elias Frantar and Saleh Ashkboos and Torsten Hoefler and Dan Alistarh , title =. International Conference on Learning Representations , year =

  40. [46]

    Proceedings of Machine Learning and Systems , volume =

    Ji Lin and Jiaming Tang and Haotian Tang and Shang Yang and Wei-Ming Chen and Wei-Chen Wang and Guangxuan Xiao and Xingyu Dang and Chuang Gan and Song Han , title =. Proceedings of Machine Learning and Systems , volume =

  41. [47]

    Proceedings of the 41st International Conference on Machine Learning , series =

    Albert Tseng and Jerry Chee and Qingyao Sun and Volodymyr Kuleshov and Christopher De Sa , title =. Proceedings of the 41st International Conference on Machine Learning , series =

  42. [48]

    Proceedings of the 41st International Conference on Machine Learning , series =

    Wei Huang and Yangdong Liu and Haotong Qin and Ying Li and Shiming Zhang and Xianglong Liu and Michele Magno and Xiaojuan Qi , title =. Proceedings of the 41st International Conference on Machine Learning , series =

  43. [49]

    Proceedings of the 41st International Conference on Machine Learning , series =

    Shiyao Li and Xuefei Ning and Luning Wang and Tengxuan Liu and Xiangsheng Shi and Shengen Yan and Guohao Dai and Huazhong Yang and Yu Wang , title =. Proceedings of the 41st International Conference on Machine Learning , series =

  44. [50]

    Lan and Wanzin Yazar and Tristan Webb and Sayeh Sharify and Xin Wang , title =

    Zifei Xu and Alexander Y. Lan and Wanzin Yazar and Tristan Webb and Sayeh Sharify and Xin Wang , title =. Proceedings of the 4th NeurIPS Efficient Natural Language and Speech Processing Workshop , series =

  45. [51]

    Proceedings of the 41st International Conference on Machine Learning , series =

    Harshavardhan Adepu and Zhanpeng Zeng and Li Zhang and Vikas Singh , title =. Proceedings of the 41st International Conference on Machine Learning , series =

  46. [52]

    International Conference on Learning Representations , year =

    Noam Shazeer and Azalia Mirhoseini and Krzysztof Maziarz and Andy Davis and Quoc Le and Geoffrey Hinton and Jeff Dean , title =. International Conference on Learning Representations , year =

  47. [53]

    Journal of Machine Learning Research , volume =

    William Fedus and Barret Zoph and Noam Shazeer , title =. Journal of Machine Learning Research , volume =

  48. [54]

    Zhao and Andrew M

    Yanqi Zhou and Tao Lei and Hanxiao Liu and Nan Du and Yanping Huang and Vincent Y. Zhao and Andrew M. Dai and Zhifeng Chen and Quoc V. Le and James Laudon , title =. Advances in Neural Information Processing Systems , volume =

  49. [55]

    Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages =

    Damai Dai and Li Dong and Shuming Ma and Bo Zheng and Zhifang Sui and Baobao Chang and Furu Wei , title =. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages =. 2022 , doi =

  50. [56]

    International Conference on Learning Representations , year =

    Joan Puigcerver and Carlos Riquelme Ruiz and Basil Mustafa and Neil Houlsby , title =. International Conference on Learning Representations , year =

  51. [57]

    Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages =

    Chaodong Xiao and Zhengqiang Zhang and Lei Zhang , title =. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages =

  52. [58]

    Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages =

    Hyunha Hwang and Xuan Truong Nguyen and Hyuk-Jae Lee , title =. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages =

  53. [59]

    Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision , pages =

    Jiahe Qian and Peisong Wang and Zhengyang Zhuge and Qinghao Hu and Jian Cheng , title =. Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision , pages =. 2026 , doi =

  54. [60]

    Mathematical Programming , volume =

    Sebastian Sager and Gerhard Reinelt and Hans Georg Bock , title =. Mathematical Programming , volume =. 2009 , doi =

  55. [61]

    Mathematical Programming , volume =

    Sebastian Sager and Hans Georg Bock and Moritz Diehl , title =. Mathematical Programming , volume =. 2012 , doi =

  56. [62]

    Mathematical Methods of Operations Research , volume =

    Sebastian Sager and Michael Jung and Christian Kirches , title =. Mathematical Methods of Operations Research , volume =. 2011 , doi =

  57. [63]

    SIAM Journal on Control and Optimization , volume =

    Christian Kirches and Felix Lenders and Paul Manns , title =. SIAM Journal on Control and Optimization , volume =. 2020 , doi =

  58. [64]

    Shankar Sastry , title =

    Ramanarayan Vasudevan and Humberto Gonzalez and Ruzena Bajcsy and S. Shankar Sastry , title =. SIAM Journal on Control and Optimization , volume =. 2013 , doi =

  59. [66]

    Bowman , title =

    Alex Wang and Amanpreet Singh and Julian Michael and Felix Hill and Omer Levy and Samuel R. Bowman , title =. International Conference on Learning Representations , year =

  60. [67]

    Manning and Andrew Y

    Richard Socher and Alex Perelygin and Jean Wu and Jason Chuang and Christopher D. Manning and Andrew Y. Ng and Christopher Potts , title =. Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing , pages =. 2013 , doi =

  61. [68]

    Transformers: State-of-the-Art Natural Language Processing , booktitle =

    Thomas Wolf and Lysandre Debut and Victor Sanh and Julien Chaumond and Clement Delangue and Anthony Moi and Pierric Cistac and Tim Rault and R. Transformers: State-of-the-Art Natural Language Processing , booktitle =. 2020 , doi =

  62. [69]

    distilbert-base-uncased-finetuned-sst-2-english , year =

  63. [70]

    Geometric Path Enumeration for Equivalence Verification of Neural Networks , booktitle =

    Samuel Teuber and Marko Kleine B. Geometric Path Enumeration for Equivalence Verification of Neural Networks , booktitle =. 2021 , doi =

  64. [71]

    Mahoney and Kurt Keutzer , title =

    Zhewei Yao and Zhen Dong and Zhangcheng Zheng and Amir Gholami and Jiali Yu and Eric Tan and Leyuan Wang and Qijing Huang and Yida Wang and Michael W. Mahoney and Kurt Keutzer , title =. Proceedings of the 38th International Conference on Machine Learning , series =

  65. [73]

    Automatica , volume =

    Mathieu Claeys and Jamal Daafouz and Didier Henrion , title =. Automatica , volume =. 2016 , doi =

  66. [74]

    Burden and S

    Samuel A. Burden and S. Shankar Sastry and Daniel E. Koditschek and Shai Revzen , title =. SIAM Journal on Applied Dynamical Systems , volume =. 2016 , doi =

  67. [75]

    Kong and J

    Nathan J. Kong and J. Joe Payne and James Zhu and Aaron M. Johnson , title =. Proceedings of the IEEE , volume =. 2024 , doi =

  68. [76]

    Iris A. M. Huijben and Matthijs Douze and Matthew J. Muckley and Ruud J. G. van Sloun and Jakob Verbeek , title =. Proceedings of the 41st International Conference on Machine Learning , series =

  69. [80]

    Lasserre and Didier Henrion and Christophe Prieur and Emmanuel Tr

    Jean B. Lasserre and Didier Henrion and Christophe Prieur and Emmanuel Tr. Nonlinear Optimal Control via Occupation Measures and. SIAM Journal on Control and Optimization , volume =. 2008 , doi =

  70. [81]

    Transactions on Machine Learning Research , year =

    Ian Colbert and Giuseppe Franco and Fabian Grob and Jinjie Zhang and Rayan Saab , title =. Transactions on Machine Learning Research , year =

  71. [85]

    Proceedings of the AAAI Conference on Artificial Intelligence , volume =

    Pei Huang and Haoze Wu and Yuting Yang and Ieva Daukantas and Min Wu and Yedi Zhang and Clark Barrett , title =. Proceedings of the AAAI Conference on Artificial Intelligence , volume =. 2024 , doi =

  72. [86]

    Gradient Flows in Metric Spaces and in the Space of Probability Measures , edition =

    Luigi Ambrosio and Nicola Gigli and Giuseppe Savar. Gradient Flows in Metric Spaces and in the Space of Probability Measures , edition =. 2008 , doi =

  73. [87]

    International Conference on Learning Representations , year =

    Shihao Zhang and Haoyu Zhang and Ian Colbert and Rayan Saab , title =. International Conference on Learning Representations , year =

  74. [89]

    Burr and Liu Liu and Meng Wang , title =

    Mohammed Nowaz Rabbani Chowdhury and Kaoutar El Maghraoui and Hsinyu Tsai and Naigang Wang and Geoffrey W. Burr and Liu Liu and Meng Wang , title =. International Conference on Learning Representations , year =

  75. [90]

    SIAM Journal on Mathematics of Data Science , volume =

    Jinjie Zhang and Yixuan Zhou and Rayan Saab , title =. SIAM Journal on Mathematics of Data Science , volume =. 2023 , doi =

  76. [91]

    Proceedings of the 37th International Conference on Machine Learning , series =

    Markus Nagel and Rana Ali Amjad and Mart van Baalen and Christos Louizos and Tijmen Blankevoort , title =. Proceedings of the 37th International Conference on Machine Learning , series =

  77. [96]

    SIAM Journal on Mathematics of Data Science , year =

    Jinjie Zhang and Yixuan Zhou and Rayan Saab , title =. SIAM Journal on Mathematics of Data Science , year =

  78. [97]

    Proceedings of the IEEE/CVF International Conference on Computer Vision , pages =

    Ian Colbert and Alessandro Pappalardo and Jakoba Petri-Koenig , title =. Proceedings of the IEEE/CVF International Conference on Computer Vision , pages =

  79. [98]

    Proceedings of the 41st International Conference on Machine Learning , series =

    Ian Colbert and Alessandro Pappalardo and Jakoba Petri-Koenig and Yaman Umuroglu , title =. Proceedings of the 41st International Conference on Machine Learning , series =. 2024 , publisher =

  80. [101]

    The Expressive Power of Low Precision Softmax Transformers with (Summarized) Chain-of-Thought , journal =

    Moritz Br. The Expressive Power of Low Precision Softmax Transformers with (Summarized) Chain-of-Thought , journal =. 2026 , note =

Showing first 80 references.