{"id":"687c01c5-91bb-4500-9cfb-a50a1e30d827","arxiv_id":"2608.12791","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"A typed accounting separates record correlation from operational capital value in finite learning devices, with separation, capitalization-efficiency, and value-retention theorems.","lead":"This paper introduces a four-part thermodynamic accounting that separates what a finite learning device has memorized from what its memory is worth on future tasks. It proves that memorization and value can separate, and it identifies conditions under which update efficiency and value retention obey clean bounds.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Separation theorem depends on open Assumption-U: LB's zero capital gain is proved only for the restricted protocol class, so 'memorization without value' is not unconditionally established for all finite devices.","rationale":"The paper is unusually careful: it identifies the U-restriction as open, attaches an asterisk to interpretation, and restricts its theorems. The question is whether the central separation claim, as advertised ('for every n there is a device family...'), is robust to that open restriction. It is not, as far as the present text shows. The LB witness's zero Delta V is computed via Lemma 3, which relies on an imported chain bound proven for the restricted class. No argument in the paper establishes that the informed optimal value is unchanged when clause (g) is dropped; the paper explicitly says it may deviate. Since the companion [9] is unpublished, the reader cannot check either the U-reduction or the chain bound. Thus the most load-bearing assumption is exactly the one the reader identified. I found no internal algebraic error in the separation proof, the ledger identities, or the alignment witnesses; the numerical descriptions are detailed but the scripts are not shipped, which is a reproducibility gap rather than a separate logical flaw. The recommended action matches the reader's CONDITIONAL verdict: accept conditionally on supplying the companion artifacts, proving the U-reduction (or showing the LB value is U-independent), and shipping the verification scripts.","tokens_in":51078,"tokens_out":12269,"duration_ms":133145,"concrete_test":"Determine whether dropping clause (g) changes the LB informed supremum: for the n=3 LB instance, enumerate protocols in the class that satisfies (a)–(f) but not (g) (finite states, K=3, all-read access, degeneracy E=0) and compute the supremum of E[W_ext|M=m,tau_1]; if any protocol exceeds kT ln 2, the 'exactly zero' capital gain is an artifact of the U-restriction and Main Theorem I(ii) requires restatement. If the enumeration is infeasible, settle it by proving or disproving the U-reduction for the flat* class used in Lemma 3.","verdict_should_be":"UNCHANGED","load_bearing_attack":"All protocol classes carry the imported Assumption-U clause (g) (Sec. II A), and Sec. II E states explicitly that whether the restriction is without loss of generality is open; V is therefore a restricted-class-relative quantity. Main Theorem I's separation is exhibited on the LB device by pinning Gap values through the flat* extraction identity (Lemma 3), whose informed-branch converse is imported from the companion via Lemma 9 and applies to the same restricted class. If protocols violating clause (g) can exploit M–Y correlation that the U-restriction forbids, the informed supremum on LB could exceed kT ln 2, making Delta V positive along the record-copy sequence. The paper is transparent about this asterisk, but the consequence is load-bearing: the headline claim that record and world correlation grow while capital gain is exactly zero is a statement about U-conforming protocols, not about all physically possible finite learning devices. Because the companion [9] is unpublished and the U-reduction is not proved here, the separation theorem's unrestricted validity is unresolved.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper develops a typed accounting framework for finite-state learning devices, separating training-side fit, record-correlation stock, the update-side search ledger, and an operational capital value V(M;T,b) defined as the informed-minus-blind work gap under a deletion counterfactual. Main Theorem I establishes basic properties of V, including affinity, nonnegativity, recoding invariance, and a single-task reduction to a companion value, and constructs for every n a device family LB on which record correlation I(M;D) and world correlation I(M;W) grow by n ln2 along an admissible update sequence while the capital gain is exactly zero. Main Theorem II introduces the search ledger sigma_M, proves an exact flat* extraction identity, a universal ledger identity, and bounds the capitalization efficiency eta_cap = Delta V/(kT sigma_M) by 1 under flat* supports, (F5')-stability, and condition (f), with equality conditions and a regime map of identified failures. Main Theorem III gives a value-retention alignment schema with a two-layer positive domain, a boundary one-time-pad witness, four one-coordinate intervention witnesses, and a subject-difference identity for the retention gap. The manuscript is explicit that all results are exact finite-device statements inside a fixed operational framework and are not a theory of statistical generalization.","tokens_in":51331,"tokens_out":9767,"duration_ms":103025,"significance":"If the results are accepted, the paper makes a useful conceptual and technical contribution: it gives a precise, operationally anchored distinction between what a finite device has recorded and what that record is worth on future tasks, with clean device witnesses for separation, for the capitalization bound, and for the failure modes of that bound. The manuscript is unusually careful about scope: it ships machine-checked numerical verification scripts, attributes imported identities explicitly, and includes a detailed limitations section. The main caveat is that the central value quantity is restricted-class-relative and several load-bearing lemmas are imported from unpublished companion work, so the headline claims currently run ahead of what the manuscript proves unconditionally.","major_comments":[{"comment":"The separation theorem is stated without a qualifier in the abstract and in the theorem, but all protocol classes carry the imported Assumption-U restriction (clause (g), Sec. II A), and Sec. II E states that whether this restriction is without loss of generality is an open question. Because V is defined relative to the U-conforming protocol class, the claim that record and world correlation grow by n ln2 while capital gain is exactly zero is not established for all physically possible finite learning devices; a protocol violating clause (g) could in principle exploit the M-Y correlation that the U-restriction forbids and make Delta V positive along the record-copy sequence. Please either prove the U-reduction or state the separation theorem and all interpretation-level claims with the restricted-class qualifier in the abstract and in the statements of Main Theorem I and Definition 9.","section":"Sec. II E; Main Theorem I(ii)-(iii)"},{"comment":"The load-bearing lemmas are not proved in this manuscript: Lemma 9 (the conditional chain bound) is restated, but the text explains that the weight-bearing path depends on this statement rather than on the interior of an imported proof, and its proof is in the unpublished companion [9]; similarly, the gate bounds B1-B3 (Theorems 2-4 of Appendix D) are given as fixed statements whose proofs are in the companion. Because the separation theorem, the flat* extraction identity, the gate threshold transition, and the budget/access witnesses all rely on these imported inequalities, the paper is not self-contained and the results cannot be fully verified by a referee. Please include complete proofs of these statements or make the dependence explicit and condition acceptance on the companion's availability.","section":"Appendix B.2; Appendix D"}],"minor_comments":[{"comment":"Clause (g) names Assumption-U but does not state the conditional-independence property formally; please give the precise condition or a complete reference to the companion.","section":"Sec. II A"},{"comment":"Refs. [9], [10], and [16] are cited as unpublished; please add arXiv identifiers or stable versioning so readers can verify the imported statements.","section":"Sec. IX"},{"comment":"The notation table would benefit from an explicit entry distinguishing flat tasks (Definition 18) from flat* tasks (Definition 13); the asterisk convention is easy to miss.","section":"Appendix E"},{"comment":"The subjects g_a, g_b, and g_c are defined in Lemma 6; repeating their one-line definitions at Eq. (27) would improve readability.","section":"Eq. (27)"}],"recommendation":"major_revision","confidential_remarks":"The manuscript relies heavily on the author's own unpublished companions (Refs. [9], [10], [16]) for the protocol class, the deletion counterfactual, the conditional chain bound, and the gate bounds. Given the novelty claim in Sec. IX, I would ask the editor to require those companions or self-contained proofs before acceptance. The U-restriction open question should also be resolved or the central claims reframed so that the published statements match what is actually proved."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The short version: this is a careful, self-aware paper that does real work, but its headline result is explicitly hostage to an open assumption imported from an unpublished companion. I'd send it to a strong referee, but I wouldn't cite it until the companion lands and the U-reduction question is addressed.\n\nWhat's new: the four-component typed accounting with V as a deletion-based informed–blind work gap, the separation theorem (record correlation grows while V stays fixed) on the LB device family, the capitalization regime map with eta_cap <= 1 under flat* + (f), the equality conditions, and the four-coordinate alignment collapse. The paper is honest that the core inequality is a composition of Sagawa–Ueda and Goldt–Seifert steps and that the ledger identity is imported; the claimed contribution is the operational packaging, the regime boundaries, and the device witnesses. The explicit witnesses and the machine-checked verification of the arithmetic are real evidence, though the scripts are not shipped.\n\nThe soft spots are exactly where the reader puts them. All protocol classes carry Assumption-U; the paper itself says the reduction is open, and V is restricted-class-relative. The separation theorem—memorization with zero capital gain—is proved only within that class. If the U-restriction fails without loss of generality, the headline statement doesn't extend to all physically possible finite devices. That's not a hidden flaw; it is in Sec. II E and Sec. X. But it is load-bearing, and it is the difference between a conditional and an unconditional result. Also, several lemmas (conditional chain bound, gate bounds B1–B3) come from the unpublished companion [9], and the gate bound's bookkeeping constant C_G has no proven numerical upper bound. The paper is upfront about all of this, which earns it credit on honesty, but it means the external verification burden is not met.\n\nDo I think it holds up? The internal math looks coherent; I didn't find a contradiction. The flat* extraction identity and the regime map are the strongest parts. The gate threshold theorem is honest about its dependence on unproven C_G. The paper's own scoping—not a theory of statistical generalization—is appropriate.\n\nWho is this for? People working on stochastic thermodynamics of learning, information engines, and value-of-information. It deserves a serious referee: the framework is important if the companion checks out. I'd accept for review, with the expectation of conditional acceptance after the companion is available or the imports are proven.\n\nRecommendation: send it to peer review. Tell the referees to focus on the U-reduction and the companion dependencies; the rest is solid.","headline":"A careful, self-aware paper with real contributions, but its headline separation result is explicitly restricted to an open assumption class from an unpublished companion, so it deserves peer review but not citation yet.","tokens_in":51777,"tokens_out":1957,"would_cite":false,"duration_ms":18679,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":["05.70.-a","89.70.+c"],"model":"deepseek-v4-flash","headline":"The paper proves that a finite learning device can accumulate arbitrary record correlation and world correlation while gaining exactly zero operational capital value, and it bounds how much of the search ledger can be capitalized into…","keywords":["thermodynamics of learning","capital value","record correlation","search ledger","capitalization efficiency","value retention","fit–value alignment"],"falsifier":"On a flat$^{*}$ task with condition (f) satisfied, construct an admissible (F5$'$)-stable memory-local update whose measured capital gain satisfies $\\Delta V > kT\\,\\sigma_M$; Main Theorem II(iii) predicts $\\eta_{\\mathrm{cap}} \\le 1$, so any such device would falsify the capitalization bound.","tokens_in":50875,"feed_emoji":"🧠","tokens_out":8001,"duration_ms":71325,"temperature":0.7,"pith_summary":"What a finite learning device has recorded and what will hold value for it on future tasks are not the same quantity: this paper develops a typed four-component accounting — acquisition cost, physical dissipation, transport action, and task value — in which memorization is a correlation stock while future value is the informed–blind work gap on a stated task distribution, access structure, and per-run budget. Its first theorem proves separation: for every $n$ there is a device family whose record correlation $I(M;D)$ and world correlation $I(M;W)$ both grow by $n\\ln 2$ along admissible memory-local updates, yet the capital gain $\\Delta V$ is exactly zero, so memorizing environmental noise creates no value follows as a theorem rather than a definition. Its second theorem introduces the search ledger $\\sigma_M$ and proves a capitalization bound $\\eta_{\\mathrm{cap}} = \\Delta V/(kT\\sigma_M) \\le 1$ in the flat$^{*}$ regime with a no-discarded-record-condition, with exact equality conditions and an explicit regime map of how the bound fails outside it. Its third theorem gives a two-layer alignment domain for value retention under task-distribution shift — exact exchange with the side-information-adjusted record fit, and with the raw record stock under joint neutrality — plus a four-coordinate intervention theorem in which shift, budget, access, or record content alone reverses or restores the fit–value ranking.","feed_headline":"Proven: a device can memorize n bits while gaining zero future value","feed_subtitle":"A typed accounting proves record correlation and operational value can diverge, with exact bounds and an alignment map.","key_machinery":"The central object is the capital value $V(M;T,b)$, defined by a deletion counterfactual: the optimal expected work extractable over a future task distribution $T$ when the memory-read port is available, minus the optimum of a blind agent from whom every read port on $M$ has been deleted and who re-optimizes from scratch under the same tasks, access structure, and per-run budget $b$; learning is then an admissible memory-local update with $\\Delta V > 0$, making the learning predicate relative to $T$, access, and budget. Two further instruments carry the theorems: the search ledger $\\sigma_M = \\Delta I(M;D) + \\Sigma_{\\mathrm{total}}$, a memory-side subsystem account that charges the effective entropy change of the memory registers plus the heat sent to the bath, including the cost of blank pages; and the flat$^{*}$ task class — degenerate energies, no gates, a single manipulable register, static all-read access, and flat-conditional attainability — on which the extraction identity $\\mathrm{Gap}(M;\\tau,b) = kT\\,I(M;X_\\tau|Y_\\tau)$ holds exactly and budget-independently.","core_discovery":"The central discovery is a type separation: a finite learning device's record-correlation stock $J_D = I(M;D)$ and its operational capital value $V(M;T,b)$, defined as the informed–blind work gap on a future task distribution under a stated budget and access structure, are not the same quantity and can move independently. Main Theorem I exhibits, for every $n$, a device family on which both $I(M;D)$ and the world correlation $I(M;W)$ grow by $n\\ln 2$ along admissible memory-local updates while $\\Delta V = 0$. Main Theorem II introduces the search ledger $\\sigma_M = \\Delta I(M;D) + \\Sigma_{\\mathrm{total}}$ and proves the capitalization bound $\\eta_{\\mathrm{cap}} \\le 1$ in the flat$^{*}$ plus (f) regime, with necessary and sufficient equality conditions and a three-route regime map of failures. Main Theorem III gives a two-layer alignment domain: value equals $kT\\,I(M';D|Y)$ exactly without any independence assumption, equals $kT\\,I(M';D)$ under joint side-information neutrality $(M,D)\\perp Y$, and the boundary is exhibited by an explicit one-time-pad witness; a four-coordinate intervention theorem shows that shift, budget, access route, and record content each admit a paired setting where changing that coordinate alone reverses or restores the fit–value ranking.","pith_inferences":["If the companion's Assumption-U restriction is not without loss of generality, then all value statements in this paper, including the separation and alignment theorems, describe only the restricted protocol class; the unrestricted operational value could deviate, so the asterisk on 'capital' is more than a formality.","The four-coordinate intervention theorem suggests a diagnostic recipe for physical learning systems inside the flat$^{*}$ class: by deliberately moving shift, budget, access route, or record content and observing whether fit ranking and value ranking separate, one can identify which operational coordinate is binding for a given device.","The gate-budget regime where $\\eta_{\\mathrm{cap}} \\to \\infty$ at fixed information content implies that under a finite per-run budget, specification information can function as a key that work cannot buy; a testable extension would probe whether such enablement persists in physical implementations with noisy gates.","The separation of record correlation from value suggests that a thermodynamics-aware training objective should target $\\Delta V$ rather than $\\Delta I(M;D)$; the paper itself stops at the accounting level and does not propose such an algorithm, leaving that as a natural next step."],"forward_implications":["Record-correlation increase is not learning: on the LB device, copying environment noise into memory grows $I(M;D)$ and $I(M;W)$ by $n\\ln 2$ with zero capital gain, so memorizing noise creates no value is a theorem, not a naming choice.","Within the flat$^{*}$ plus (f) regime, value cannot outpace the search ledger: $\\eta_{\\mathrm{cap}} \\le 1$, and full capitalization means simultaneously no forgetting, no waste, no $Y$-contamination, and a reversible implementation; outside the regime the bound fails in identified routes that the regime map organizes.","Fit–value alignment is conditional: on flat tasks value equals $kT\\,I(M';D|Y)$ exactly without an independence assumption, and equals the raw record stock $kT\\,I(M';D)$ only under joint side-information neutrality $(M,D)\\perp Y$.","A blank-start cumulative bound $\\eta_{\\mathrm{cum}} \\le 1$ survives even without condition (f), so the books of a multi-step history are honest even when a per-update recycler shows $\\eta_{\\mathrm{cap}} = 2$.","Overfitting does not imply low efficiency and low efficiency does not imply overfitting: $\\eta_{\\mathrm{cap}}$ and $\\rho_{\\mathrm{gen}}$ admit no functional or monotone relation, and the full rectangle $(0,1]\\times[0,1]$ is realized by explicit devices and shifts.","Under task-distribution shift, the retention gap has a subject-difference identity: the per-update change in $L_{\\mathrm{gen}}$ equals exactly the increase of the three breakage subjects (forgetting, waste, $Y$-contamination) on the shifted book, provided the support is flat and draw exogeneity holds over the shifted support."],"supporting_citations":[{"why":"Supplies the entire protocol framework: budgeted protocol classes, the deletion counterfactual, Assumption-U, and the informed–blind gap used to define V.","marker":"[9]"},{"why":"Provides the Sagawa–Ueda-type bound that value cannot exceed $kT$ times information, a step composing the capitalization inequality.","marker":"[1]"},{"why":"Provides the memory-side subsystem ledger $\\Delta S(\\omega)+\\Delta Q$, the direct ancestor of the search ledger $\\sigma_M$, and the information–ledger bound used in Main Theorem II.","marker":"[2]"},{"why":"Supplies the bipartite information-flow entropy balances from which the universal ledger identity (Lemma 5) is imported.","marker":"[4]"},{"why":"Supplies the companion bipartite fluctuation-theoretic framework that, together with [4], underlies the universal ledger identity.","marker":"[5]"},{"why":"Establishes the thermodynamic equivalence of maximum-work and maximum-likelihood training whose structural features Proposition 2 isolates in the present vocabulary.","marker":"[8]"},{"why":"Supplies the gate-family concepts of degeneracy, transfer, and effective key length $k_{\\mathrm{eff}}$ used in the threshold retention theorem.","marker":"[16]"},{"why":"Provides the bits-layer inferential efficiency with $\\eta\\le 1$ and two correlation subjects, the closest existing analogue of the capitalization-efficiency package.","marker":"[17]"}],"fun_headline_variants":["Memorize n bits, gain zero future value: a proven separation","Record correlation can grow n ln 2 while capital gain stays exactly zero","Memory and value can diverge completely: proof from finite-state thermodynamics","Zero capital gain despite n-bit memory growth: exact separation theorem","Operational value can stay zero while memory stock grows n ln 2"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that all protocol classes carry the imported Assumption-U restriction, whose validity without loss of generality is left open in the companion framework; if that reduction fails, V is only a restricted-class-relative quantity, and the separation and alignment theorems would not describe all physically possible devices.","fun_headline_variants_meta":{"raw":{"variants":["Memorize n bits, gain zero future value: a proven separation","Record correlation can grow n ln 2 while capital gain stays exactly zero","Memory and value can diverge completely: proof from finite-state thermodynamics","Zero capital gain despite n-bit memory growth: exact separation theorem","Operational value can stay zero while memory stock grows n ln 2"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000651,"raw_usage":{"total_tokens":3130,"prompt_tokens":1237,"completion_tokens":1893,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":853,"completion_tokens_details":{"reasoning_tokens":1800}},"tokens_in":853,"tokens_out":1893,"duration_ms":14460,"temperature":1.0,"reasoning_tokens":1800,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T23:13:50.185115+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"On a flat$^{*}$ task with condition (f) satisfied, construct an admissible (F5$'$)-stable memory-local update whose measured capital gain satisfies $\\Delta V > kT\\,\\sigma_M$; Main Theorem II(iii) predicts $\\eta_{\\mathrm{cap}} \\le 1$, so any such device would falsify the capitalization bound.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the entire protocol framework: budgeted protocol classes, the deletion counterfactual, Assumption-U, and the informed–blind gap used to define V."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the Sagawa–Ueda-type bound that value cannot exceed $kT$ times information, a step composing the capitalization inequality."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the memory-side subsystem ledger $\\Delta S(\\omega)+\\Delta Q$, the direct ancestor of the search ledger $\\sigma_M$, and the information–ledger bound used in Main Theorem II."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the bipartite information-flow entropy balances from which the universal ledger identity (Lemma 5) is imported."},{"cited_title":"HenceE[W ext] of ev- eryP∈Prot b(Aτ−M,p) is a functional of the non-M marginals alone, and the blind branch does not depend on the coupling structureπ M","cited_arxiv_id":null,"evidence_quote":"Supplies the companion bipartite fluctuation-theoretic framework that, together with [4], underlies the universal ledger identity."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes the thermodynamic equivalence of maximum-work and maximum-likelihood training whose structural features Proposition 2 isolates in the present vocabulary."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the gate-family concepts of degeneracy, transfer, and effective key length $k_{\\mathrm{eff}}$ used in the threshold retention theorem."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the bits-layer inferential efficiency with $\\eta\\le 1$ and two correlation subjects, the closest existing analogue of the capitalization-efficiency package."}],"review_version":1}