{"id":"c5a400ff-e8c6-47a6-a139-f5f7f67ed0d1","arxiv_id":"2506.20699","paper_version":2,"verdict":"REJECT","confidence":"MODERATE","novelty_score":3.0,"correctness_risk":"high","formal_verification":"none","parameter_count":3,"one_line_summary":"Despite its abstract, the manuscript derives no new quantitative results: it relabels standard variational inference and known learning mechanisms as a 'Context-Content Uncertainty Principle' and skips the key equivalence proof.","lead":"The abstract of this submission promises a 'Structural Decoupling' theory with results about width and VC dimension, but the body text is a different manuscript proposing a 'Context-Content Uncertainty Principle' that recasts familiar inference ideas as entropy alignment. A generalist should read it as a cautionary example: the headline results never appear in the text, and the remaining theorems mostly assume or restate their conclusions.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The convergence and composition theorems are not proved as stated: Theorem 5 assumes its own conclusion, and the Appendix proofs of Theorems 6–8 swap Φ and Ψ, so the central derivation from CCUP is unsupported.","rationale":"The reader's REJECT verdict is supported. My stress test identifies the same family of weaknesses but places the load-bearing failure slightly differently: while the reader's weakest_assumption highlights Theorem 5's unproved contraction and the unmeasured H(Φ)≪H(Ψ) premise, the more decisive problem is that the appendix proofs of the convergence and composition theorems prove different statements (systematic Φ/Ψ swapping) and the key equivalence is skipped. This is not a matter of missing empirical calibration; the formal chain from CCUP to the claimed operational principles is internally incomplete. The submission-level abstract mismatch (StrLT/width/contractive-similarity operator) is an additional, independent rejection reason that I did not make my primary concern. No adjustment to the verdict is needed; if anything, the concrete test would sharpen the rejection into a falsifiable check of Theorem 8.","tokens_in":26249,"tokens_out":5367,"duration_ms":58328,"concrete_test":"Independently re-derive Theorem 8 using the theorem's own variables and update rule: assume DKL(p(Ψ|Φ_ℓ)∥p(Ψ|Φ_{ℓ-1}))≤ε and Φ_ℓ=f_ℓ(Φ_{ℓ-1}^{(1)},...,Φ_{ℓ-1}^{(n)}), and attempt to prove H(Ψ|Φ_ℓ)<H(Ψ|Φ_{ℓ-1}) without Jensen's inequality on conditional entropy. If the proof requires the swapped inequality H(Φ_{ℓ-1}|Ψ_ℓ)<H(Φ_{ℓ-1}|Ψ_{ℓ-1}^{(i)}) or fails for a simple two-level hierarchy (e.g., Φ_1 is a noisy aggregate of Φ_0 components with independent noise), then Theorem 8 is unproved and the hierarchical composition claim is not established.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim requires that the CCUP entropy asymmetry actually entails bootstrapped convergence (Theorems 5–7) and hierarchical composition (Theorem 8). This chain fails at its load-bearing joints. First, Theorem 5 (Section 4.1) assumes that the update operator F is contractive in KL divergence with γ<1 and that the entropy gap decreases monotonically; these are precisely the convergence conclusions, and they are never derived for the specific update rules in Eqs. (1)–(2). Second, the appendix proofs of Theorems 6–8 prove different statements from the theorems. Theorem 6's main-text update is Φ(t+1)=argmin_Φ[H(Ψ(t)|Φ)+λDKL(Φ∥Φ(t))], but Appendix K begins with Ψ(t+1)=argmin_Ψ[H(Φ(t)|Ψ)+λDKL(Ψ∥Ψ(t))] and proves convergence of Ψ(t), not Φ(t). Appendix L similarly proves convergence of Ψ_l with KL(p(Φ_l|Ψ_l)∥p(Φ_l|Ψ_l^{t-1})), while Theorem 7 asserts convergence of Φ_l with KL(p(Ψ_l|Φ_l)∥p(Ψ_l|Φ_l^{t-1})). Theorem 8's main statement is H(Ψ_{ℓ-1}|Φ_ℓ)<H(Ψ_{ℓ-1}|Φ_{ℓ-1}); Appendix M proves H(Φ_{ℓ-1}|Ψ_ℓ)<H(Φ_{ℓ-1}|Ψ_{ℓ-1}^{(i)}), relying on Jensen's inequality for conditional entropy in the conditioning variable, which fails in general. The proof of Theorem 1 also contains a sign error: from H(Ψ)≫H(Φ), mutual-information identities give H(Ψ|Φ)>H(Φ|Ψ), not the claimed H(Φ|Ψ)>H(Ψ|Φ). Proposition 2, the central mutual-reducibility claim, is explicitly 'proof skipped.' Thus the central assertion that all Layer 1–4 principles reduce to one entropy law is not actually derived.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces the Context-Content Uncertainty Principle (CCUP), the claim that inference under uncertainty is governed by an entropy asymmetry between high-entropy context and low-entropy structured content. From this principle it derives a four-layer hierarchy of operational principles: core inference constraints (structure-before-specificity, asymmetric inference flow, cycle-consistent bootstrapping, conditional compression), resource allocation mechanisms (precision-weighted attention, asymmetric learning rates, low-entropy memory attractors), temporal bootstrapping dynamics, and spatial hierarchical composition. The advertised contribution is that these principles reduce to one entropy-alignment law, with formal theorems on equivalence, convergence, and hierarchical uncertainty reduction, plus simulations in the extended version. The manuscript is organized around a dependency lattice and includes an appendix with proofs of the main results.","tokens_in":26676,"tokens_out":3305,"duration_ms":35589,"significance":"If the central derivation were correct, the paper would offer an impressively broad unification of information-theoretic, variational, and cognitive-science perspectives on inference, memory, attention, and hierarchical representation. The explicit attempt to reduce multiple design principles to a single entropy asymmetry is ambitious, and the layered presentation and dependency lattice are conceptually clear. However, the paper's scientific value depends entirely on the claimed theorems, and those theorems are not established: the appendix proofs of several central results prove different statements from the ones stated, the convergence theorems assume the mechanisms that they purport to prove, and the key equivalence proposition is explicitly unproved. Because these failures occur at the load-bearing joints of the derivation, the paper does not currently provide a reliable foundation for its claims, even as a theoretical framework.","major_comments":[{"comment":"The proof of Theorem 1 contains a sign error that reverses the central implication. The proof states that H(Ψ) ≫ H(Φ) implies H(Φ|Ψ) > H(Ψ|Φ). From the mutual-information identities H(Φ|Ψ) = H(Φ) − I(Φ;Ψ) and H(Ψ|Φ) = H(Ψ) − I(Φ;Ψ), the correct implication is H(Ψ|Φ) − H(Φ|Ψ) = H(Ψ) − H(Φ) > 0, hence H(Ψ|Φ) > H(Φ|Ψ). The claimed inequality is the opposite, and the remainder of the proof uses this reversed ordering to justify structure-before-specificity. This is not a typographical slip: the direction of the entropy ordering is the central premise of the paper, so the proof of Theorem 1 as written does not support its conclusion.","section":"Appendix D, Theorem 1"},{"comment":"Proposition 2 is the central mutual-reducibility claim that all four Layer 1 principles are equivalent under CCUP, but its proof is explicitly skipped ('proof skipped'). Appendix H provides only a qualitative paragraph describing how each principle could be viewed from the variational objective; it does not prove equivalence, transitivity, or reparameterization. Since the paper's abstract and Section 6 rely on the mutual reducibility of SbS, DIF, BB, and CC, this missing proof is load-bearing. The subsequent 'dependency lattice' and the derivation of Layers 2–4 inherit this gap.","section":"Section 2.3, Proposition 2"},{"comment":"The convergence theorems assume the conclusion. Theorem 5 assumes that the update operator F is contractive in KL divergence with 0 < γ < 1 and that the entropy gap H(Ψ(t)|Φ(t)) is monotonically decreasing; these are exactly the mechanisms that force convergence to a fixed point, and they are never derived for the specific updates in Eqs. (1)–(2). Appendix J is only a proof sketch that restates these assumptions as steps. The same circularity appears in Theorems 6 and 7, whose proofs in Appendices K and L again assume convexity, boundedness, and, in Theorem 6's proof, a 'contraction of posterior distance' that is asserted rather than shown. Without independent derivation of contractivity, the theorems do not establish that CCUP-aligned bootstrapping converges.","section":"Section 4.1, Theorems 5–7"},{"comment":"The appendix proofs of Theorems 6, 7, and 8 prove statements about different variables from the theorems. Theorem 6's main text updates Φ(t+1) = arg min_Φ [H(Ψ(t)|Φ) + λD_KL(Φ‖Φ(t))], but Appendix K begins with Ψ(t+1) = arg min_Ψ [H(Φ(t)|Ψ) + λD_KL(Ψ‖Ψ(t))] and proves convergence of Ψ(t). Theorem 7 similarly states convergence of Φ_l with KL terms p(Ψ_l|Φ_l), but Appendix L proves convergence of Ψ_l with KL terms p(Φ_l|Ψ_l). Theorem 8's main statement is H(Ψ_{ℓ−1}|Φ_ℓ) < H(Ψ_{ℓ−1}|Φ_{ℓ−1}), whereas Appendix M proves H(Φ_{ℓ−1}|Ψ_ℓ) < H(Φ_{ℓ−1}|Ψ_{ℓ−1}^{(i)}). These are not notational variants: they are different conditional entropies over different variables, and the stated theorems are not proved by the supplied arguments.","section":"Appendices K, L, M, Theorems 6–8"},{"comment":"The proof of Theorem 8 relies on the claim that conditional entropy is convex in the conditioning variable, using Jensen's inequality to conclude H(Φ_{ℓ−1}|Ψ_ℓ) ≤ Σ α_i H(Φ_{ℓ−1}|Ψ^{(i)}_{ℓ−1}). Conditional entropy is convex in the conditional distribution, not generally in the conditioning random variable or its parameterization. The appendix gives no condition under which the required convexity holds, so the inequality 'H(Φ_{ℓ−1}|Ψ_ℓ) < H(Φ_{ℓ−1}|Ψ^{(i)}_{ℓ−1})' in the proof is unsupported. This is the main step in the claim that hierarchical composition reduces conditional uncertainty.","section":"Appendix M, Theorem 8"},{"comment":"The formulas for the entropy-modulated control policy are declared rather than derived. The proof in Appendix I differentiates the objective L_t(r_t) with respect to α_t, η_t, and C_t, but the resulting stationarity conditions involve only the log penalties, not the dependences stated in Theorem 4. The theorem's displayed formulas α_t ∝ |∇_{Φ_t} H(Ψ_t|Φ_t)|, η_t ∝ H(Ψ_t|Φ_t)/(H(Φ_t)+ε), and C_t ∝ I(Ψ_t;Φ_t) do not follow from the given derivatives or from the stated objective; the proof introduces each formula as an interpretation rather than deriving it from the optimization. Since Theorem 4 is the formal content of Layer 2, this is a load-bearing gap in the derivation of the resource-allocation corollaries.","section":"Section 3, Theorem 4"}],"minor_comments":[{"comment":"In the definition of L_multi, the first term is written H(Φ^{(t)}_ℓ | Φ^{(t)}_ℓ), which is identically zero for a fixed value of the conditioning variable; presumably the intended term is H(Ψ^{(t)}_ℓ | Φ^{(t)}_ℓ) or an analogous cross-level entropy. This makes the displayed objective undefined as written.","section":"Section 4.2, H2 variational objective"},{"comment":"The text contains the typos 'stabibility' and 'plasciticity' for 'stability' and 'plasticity'; please correct these throughout.","section":"Section 4.2, Step 5"},{"comment":"The appendix labels the proof of Theorem 5 as 'Proof Sketch' and states that convergence follows because KL divergence 'induces a proper topology and the space is complete.' This is not a proof of contractive convergence for the specific update F; at minimum, the domain and metric completeness need to be specified precisely.","section":"Appendix J, Theorem 5"},{"comment":"The proof of Theorem 3 uses the inequality D_KL(q‖p) ≥ H(q) − H(p), which is not generally valid as written; the correct relation involves cross-entropy and holds with appropriate signs. While the final entropy bound may be repairable, the displayed inequality should be corrected or replaced by a proper standard identity.","section":"Section 2.4, Theorem 3"},{"comment":"The assertion that H(Φ) ≪ H(Ψ) is 'a hallmark of natural cognitive and learning systems' is supported only by a citation to the author's own preprint [12]. This premise is central to the entire framework, and the paper would be strengthened by independent empirical or theoretical support for this ordering, or by a clear statement that the results are conditional on it.","section":"Section 2.1, Lemma 2 and Remark"}],"recommendation":"reject","confidential_remarks":"The manuscript has the structure of a broad theoretical synthesis, but the formal core is not currently sound. The recurring pattern of appendix proofs proving statements about Ψ while main theorems state statements about Φ suggests that the proofs were not checked against the theorems. The skipped proof of Proposition 2 and the circular assumptions in Theorems 5–7 would require substantial new derivations rather than local corrections. In its present form the paper is likely to mislead readers about what has been established, and I cannot recommend publication. If the author wishes to resubmit, the scope would need to be reduced to the lemmas that are actually proved, with the unproved equivalence and convergence claims either proved under explicit conditions or removed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know two things. First, the arXiv abstract advertises Structural Learning Theory, width, a contractive-similarity operator, structural decoupling, and safety failures; none of those concepts appears anywhere in the body, which is a separate manuscript about a \"Context-Content Uncertainty Principle.\" Second, the body's formal core is not sound: the appendix proofs for Theorems 6–8 systematically prove statements about the wrong variable (Psi instead of Phi, and vice versa), Theorem 5 assumes the contraction and monotone entropy decrease that it then \"proves,\" Proposition 2 is explicitly \"proof skipped,\" and Theorem 1's proof contains a sign error.\n\nWhat the paper does well: it assembles a genuinely useful conceptual map. The layered taxonomy—structure-before-specificity, asymmetric inference flow, cycle-consistent bootstrapping, conditional compression, precision-weighted attention, slow/fast learning rates, attractor memory, curriculum order—is a coherent way to organize existing design intuitions. The examples (phantom limb, ventral-dorsal streams, narrative understanding) show broad reading across cognitive science and machine learning, and the paper honestly cites the standard results it reduces to (Kingma & Welling, Yu & Dayan, Bengio et al., Hopfield). That last point cuts against the novelty claim, but it is honest about provenance.\n\nThe soft spots are load-bearing, not cosmetic. The central equivalence class E1 is not proved; the variational formulation reduces to the standard ELBO; the entropy asymmetry H(Phi) << H(Psi) is asserted and supported only by the author's own preprints. Theorem 4's allocation formulas (alpha proportional to the entropy gradient, eta proportional to the entropy ratio, C proportional to mutual information) are declared in the proof, not derived from the stated Lagrangian. Theorem 8 uses Jensen's inequality on conditional entropy in the conditioning variable, which is generally false. The body's abstract promises computational simulations that do not exist in the text. And the abstract/body mismatch is not a minor style issue: it means the submission as packaged does not describe the work it contains.\n\nMy verdict: desk reject. The underlying idea—that many architectural inductive biases can be viewed through one entropy-asymmetry lens—is worth a careful essay or hypothesis paper, and I would be happy to read that version. But as submitted, the formal theorems do not hold, the abstract describes a different project, and there are no experiments. A serious referee would be doing the author's proofreading and rewriting work for them. I would not spend referee time on it as it stands.","headline":"Abstract and body are two different papers, and the body's formal claims don't survive contact with their own appendices.","tokens_in":27263,"tokens_out":3225,"would_cite":false,"duration_ms":36623,"reading_group":"no","serious_thinker":"no","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that inference under uncertainty reduces to a single entropy-asymmetry law: low-entropy structured content must be established before high-entropy context is interpreted.","keywords":["context-content uncertainty principle","entropy asymmetry","structure-before-specificity","variational inference","cycle-consistent bootstrapping","conditional compression","hierarchical composition","continual learning"],"falsifier":"Measure the marginal entropies of a trained model's representations versus its raw inputs on a standard dataset: if the representation entropy is not substantially below the input entropy, the CCUP asymmetry fails. Separately, run the paper's bootstrapped update Φ(t+1) = arg min_Φ [H(Ψ|Φ) + λD_KL(Φ ∥ Φ(t))] on a benchmark task and record D_KL(p(Ψ|Φ(t+1)) ∥ p(Ψ|Φ(t))) at each step; if the ratio to the previous step is not bounded by a constant below 1, Theorem 5's convergence claim does not hold for that update rule.","tokens_in":25838,"feed_emoji":"🧠","tokens_out":7790,"duration_ms":75136,"temperature":0.7,"pith_summary":"The paper tries to show that effective learning under uncertainty is not a collection of separate tricks but one informational law: high-entropy, variable context (sensory input, examples, linguistic form) must be interpreted through low-entropy, structured content (priors, schemas, goals). It states this as the Context-Content Uncertainty Principle (CCUP), which says optimal inference minimizes joint entropy by first establishing structure and then binding specific details. From that single asymmetry the paper derives a lattice of results: four core inference constraints are mutually reducible; attention, learning-rate separation, and attractor memory follow as corollaries; recursive bootstrapping converges to a fixed-point schema; and hierarchical composition reduces conditional uncertainty at every level. If correct, this would give one unified theoretical foundation for how brains and machines organize perception, memory, planning, and even safety failures such as hallucination, which the paper interprets as scaffold-resolution failures rather than output errors.","feed_headline":"One entropy law unifies learning, memory, and attention","feed_subtitle":"Low-entropy structure before high-entropy detail: a single principle behind generalization, bootstrapping, and safety failures.","key_machinery":"The load-bearing object is the entropy decomposition under broken symmetry: H(Φ, Ψ) = H(Φ) + H(Ψ|Φ) with the assumption H(Φ) ≪ H(Ψ), which makes the conditional term dominate and gives inference a preferred direction from content to context. The paper operationalizes this through the variational free energy F[q] = E_{q(Z|Ψ)}[−log p(Ψ|Z)] + D_KL(q(Z|Ψ) ∥ p(Z|Φ)), reading the KL term as a variational preconditioner that confines the posterior to a low-entropy submanifold and bounds its entropy. The temporal claims ride on the recursive update Φ(t+1) = F(Ψ(t), Φ(t)) together with two structural assumptions: the update operator contracts KL divergence with rate 0 < γ < 1, and the entropy gap H(Ψ(t)|Φ(t)) decreases monotonically; these are what force convergence to a fixed-point schema in Theorems 5–7.","core_discovery":"The central claim is the Context-Content Uncertainty Principle (CCUP): because structured content Φ has far lower entropy than contextual input Ψ (H(Φ) ≪ H(Ψ)), optimal inference proceeds by minimizing the conditional entropy H(Ψ|Φ) rather than treating Ψ and Φ symmetrically. Decomposing joint entropy as H(Φ, Ψ) = H(Φ) + H(Ψ|Φ) turns this asymmetry into a directional policy—structure-before-specificity: establish Φ first, then use it to constrain Ψ. From this decomposition, the paper derives that structure-before-specificity, asymmetric inference flow, cycle-consistent bootstrapping, and conditional compression form one reducible equivalence class; that precision-weighted attention, asymmetric learning rates, and memory attractors are the resulting control laws; that bootstrapped updates converge to fixed-point schemas when the update operator contracts KL divergence; and that hierarchical composition reduces conditional uncertainty level by level.","pith_inferences":["The CCUP law suggests a design rule the paper only sketches: any system that can separate its scaffold (slow, low-entropy structure) from its flow (fast, high-entropy specifics) should be more stable under distribution shift; this is directly testable by comparing single-model versus decoupled training on a non-stationary benchmark.","Because the four core constraints are claimed mutually reducible, the framework predicts that a system built around any one—say, conditional compression alone—should spontaneously exhibit the others, such as attention-like allocation and bootstrapped refinement; the paper's simulations all include the full CCUP objective, so a minimal-contrast experiment would isolate whether one constraint suffic","The same entropy-alignment cycle could be applied to problems the paper does not target, such as continual task discovery or open-world agents: a new context that cannot be aligned to existing low-entropy structure would signal the need for a new scaffold, which operationalizes discovering which contexts exist."],"forward_implications":["If CCUP is right, training should separate structural learning from specificity learning: the gradients that update the scaffold (content) and the gradients that optimize within-context flow should be distinct, because mixing them violates the entropy alignment that makes generalization efficient.","Attention, learning rates, and memory capacity become derived quantities rather than hyperparameters, set by entropy gradients: attend to inputs whose interpretation most depends on structure, learn content slowly and specifics fast, and store traces that maximize mutual information with structure.","A learning system built on cycle-consistent bootstrapping converges to stable fixed-point schemas, which would make continual learning and memory consolidation the same process as entropy minimization over time.","Hierarchical composition strictly reduces conditional uncertainty at each abstraction level, giving a principled reason for deep, compositional architectures over flat ones.","Safety failures—hallucination, reward-model boundary errors, deceptive alignment—are reinterpreted as scaffold-resolution or scaffold-preservation failures, shifting where interventions should target."],"supporting_citations":[{"why":"Supplies the free-energy objective that the variational formulation builds on.","marker":"[1]"},{"why":"Supports the claim that structure is a learned invariant representation, grounding the SbS principle.","marker":"[4]"},{"why":"Author's earlier preprint asserting the broken symmetry H(Φ) ≪ H(Ψ) that CCUP assumes.","marker":"[12]"},{"why":"Provides the chain rule and entropy inequalities used in the core decomposition.","marker":"[15]"},{"why":"Author's earlier work on inverted inference and recursive bootstrapping that the BB principle extends.","marker":"[35]"},{"why":"Shannon's source coding theorem underlying the conditional compression principle.","marker":"[36]"},{"why":"Preconditioning literature the paper invokes to interpret SbS as a variational preconditioner.","marker":"[37]"},{"why":"Curriculum learning work cited for the temporal staging principle in Layer 3.","marker":"[45]"}],"fun_headline_variants":["Learning by structure first: a new entropy law","Structural decoupling: how AI learns context and content","Entropy asymmetry predicts learning phase transitions","Safety failures as scaffold failures, not prediction errors"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"Everything downstream rests on the claim that in real systems structured content genuinely has far less entropy than incoming context, and that repeated structure updates shrink prediction differences by a fixed factor rather than merely oscillating.","fun_headline_variants_meta":{"raw":{"variants":["Learning by structure first: a new entropy law","Structural decoupling: how AI learns context and content","Entropy asymmetry predicts learning phase transitions","Safety failures as scaffold failures, not prediction errors"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000217,"raw_usage":{"total_tokens":1469,"prompt_tokens":1010,"completion_tokens":459,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":626,"completion_tokens_details":{"reasoning_tokens":401}},"tokens_in":626,"tokens_out":459,"duration_ms":5384,"temperature":1.0,"reasoning_tokens":401,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T22:47:18.427854+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Measure the marginal entropies of a trained model's representations versus its raw inputs on a standard dataset: if the representation entropy is not substantially below the input entropy, the CCUP asymmetry fails. Separately, run the paper's bootstrapped update Φ(t+1) = arg min_Φ [H(Ψ|Φ) + λD_KL(Φ ∥ Φ(t))] on a benchmark task and record D_KL(p(Ψ|Φ(t+1)) ∥ p(Ψ|Φ(t))) at each step; if the ratio to the previous step is not bounded by a constant below 1, Theorem 5's convergence claim does not hold for that update rule.","supporting_citations":[{"cited_title":"The free-energy principle: a unified brain theory?,","cited_arxiv_id":null,"evidence_quote":"Supplies the free-energy objective that the variational formulation builds on."},{"cited_title":"Representation learning: A review and new per- spectives,","cited_arxiv_id":null,"evidence_quote":"Supports the claim that structure is a learned invariant representation, grounding the SbS principle."},{"cited_title":"On Broken Symmetry in Cognition","cited_arxiv_id":"2303.06047","evidence_quote":"Author's earlier preprint asserting the broken symmetry H(Φ) ≪ H(Ψ) that CCUP assumes."},{"cited_title":"Inverted Inference and Recursive Bootstrapping: A Primal-Dual Theory of Structured Cognition","cited_arxiv_id":"2404.01183","evidence_quote":"Author's earlier work on inverted inference and recursive bootstrapping that the BB principle extends."},{"cited_title":"Preconditioning techniques for large linear systems: a survey,","cited_arxiv_id":null,"evidence_quote":"Preconditioning literature the paper invokes to interpret SbS as a variational preconditioner."}],"review_version":1}