{"id":"7370b58d-daa9-4282-b3d9-efab0b780256","arxiv_id":"2502.04088","paper_version":3,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Achieved information gain, the ideal-minus-remaining relative entropy, quantifies how much of an update to a belief state is actually correct and can be negative for misleading updates.","lead":"The paper introduces 'achieved information gain' (AIG), a three-state information measure that subtracts the remaining distance to an ideal belief from the ideal update's gain, so wrong updates can score negative. It offers axiomatic support, analytic examples, and suggests cognitive fidelity and efficiency ratios to guide resource-aware inference.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The uniqueness theorem is entirely driven by the Locality axiom; without independent justification for locality, AIG is just one of many proper gains that satisfy the other four axioms.","rationale":"The paper's central contribution is the definition and axiomatic derivation of AIG as the right quantitative measure of imperfect cognition, so the axiomatic derivation is the main formal support for that claim. I checked the derivation: it is internally consistent, and AIG does satisfy all five axioms. The single most load-bearing premise is Locality, because without it Eq. 38 is not forced and many proper scoring rules generate three-state gains satisfying the other axioms. The reader identifies this same weakest assumption. Replacing Locality with the quadratic scoring rule gives a concrete counterexample to uniqueness, so the central claim is conditional on a premise the paper motivates only by analogy and convenience. The reader's other flags, such as the false 'apparent information gain is always larger' statement in Section 5.8 and the omitted derivation in Section 5.10, are secondary and do not affect the definition or the axiomatic derivation itself. Since the verdict CONDITIONAL already reflects the appropriate level of confidence, no change to the reader's verdict is needed.","tokens_in":27847,"tokens_out":16983,"duration_ms":174707,"concrete_test":"Verify by direct differentiation that G_Q(IA,IB,I0) = <S(IB,s) - S(I0,s)>_{s|IA}, with S(P,s) = -sum_k (delta_{s,k} - P_k)^2, satisfies Additivity, Analyticity, Properness, and Calibration. Then evaluate G_Q for the two-outcome Bernoulli example of Section 5.1, e.g. p_A = 0.64, p_B = 0.6, p_0 = 0.5, and compare with Eq. 45. Since G_Q is not proportional to the AIG of Eq. 6, the uniqueness theorem in Section 4 fails if Locality is dropped, demonstrating that the conclusion is fully driven by Locality.","verdict_should_be":"UNCHANGED","load_bearing_attack":"In Section 4.1, the derivation uses Additivity, Locality, Analyticity, Properness, and Calibration. The key step is Eq. 36: Locality restricts the gain for a perfectly informed Alice, G(I_{s=s'}, IB, I0), to depend only on P(s'|IB). Together with Properness this yields the differential equation in Eq. 37, the logarithmic g(q) = lambda ln q + c in Eq. 38, and hence AIG in Eq. 41. The remaining axioms do not force a logarithm by themselves. A three-state gain built from the strictly proper, polynomial quadratic scoring rule S(P,s) = -sum_k (delta_{s,k} - P_k)^2 via G_Q(IA,IB,I0) = <S(IB,s) - S(I0,s)>_{s|IA} satisfies Additivity (linearity in P_A), Analyticity, Properness, and Calibration, but is not proportional to AIG. The paper offers no independent argument that a cognitive gain must ignore probabilities assigned to wrong states; the analogy to known scoring-rule axioms is stated, not derived. Thus the central 'right measure' claim is conditional on a single, under-justified axiom, making Locality the load-bearing point in the axiomatic uniqueness argument.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes the achieved information gain (AIG), defined as DS(IA,IB,I0) := DS(IA,I0) − DS(IA,IB) = ⟨ln(P(s|IB)/P(s|I0))⟩_{s|IA}, as a quantitative measure of the information obtained in an imperfect cognitive update from I0 to IB relative to an ideal state IA. The manuscript derives AIG axiomatically in Section 4 from additivity, locality, analyticity, properness, and calibration; relates it to KL divergence, mutual information, Rényi divergence, and scoring rules; works out closed-form AIG expressions for Bernoulli, binomial, Poisson, Beta, and Gaussian updates; analyzes mean-field and incomplete-data approximations; gives a Monte Carlo estimator for intractable posteriors; and defines cognitive fidelity and cognitive efficiency, applying the latter to sustainable computing decisions.","tokens_in":28129,"tokens_out":19026,"duration_ms":188321,"significance":"If the axiomatic claim is accepted, AIG is an attractive three-state information measure with a clear operational reading: the reduction in Bob's expected surprise from Alice's perspective. The paper's concrete contributions are valuable: the Bernoulli, binomial, Poisson, Beta, and Gaussian formulas check out; the mean-field calculation and the incomplete-data Wiener-filter example are correct and instructive; and the estimator in Section 5.9, together with the data-averaged identity in Eq. (94), is a practical tool for benchmarking approximate posteriors. The cognitive-efficiency discussion, especially the comparison in Eq. (100), gives a useful decision criterion. The axiomatic derivation is parameter-free up to the unit scale λ, and the relation to established scoring rules is clearly laid out. The main weakness is that the uniqueness theorem is conditional on the Locality axiom, which is asserted rather than independently justified.","major_comments":[{"comment":"The uniqueness theorem is driven entirely by the Locality axiom. A three-state gain G_Q(IA,IB,I0) = ⟨S(IB,s) − S(I0,s)⟩_{s|IA} with the strictly proper quadratic scoring rule S(P,s) = −Σ_k (δ_{s,k} − P_k)^2 satisfies Additivity (linearity in P(s|IA)), Analyticity, Properness (it is maximal at IB = IA), and Calibration (it vanishes for IB = I0), but it is not proportional to AIG because for atomic IA it depends on P(s''|IB) for s'' ≠ s', violating Locality. The paper asserts Locality without an independent normative argument that a cognitive gain should ignore probabilities assigned to wrong states; the observation in Section 3.4 that the paper's preferred scoring rules are local is not a derivation of that requirement. The central claim should therefore be qualified as characterizing AIG among local gains, or an independent justification of Locality should be supplied.","section":"Sec. 4.1, Locality axiom and Eqs. (35)-(41)"},{"comment":"The transition from the stationarity condition to the global solution g(s,q,I0) = λ ln q + c(s,I0) is under-specified. Equation (37) is evaluated only at q~ = P(s'|IA), which is a single value for each s' for a fixed IA. To obtain an ordinary differential equation in q~, one must explicitly quantify Properness over all Alice beliefs and observe that every q~ in (0,1) can be realized as P(s'|IA) for some IA. The text instead invokes the Analytical axiom, but infinite differentiability (C^∞) is not real analyticity, and a pointwise stationarity condition does not extend to an open set merely by smoothness. The author should either supply the missing quantification over IA or state real analyticity and show how it applies to the local identity.","section":"Sec. 4.1, Eqs. (37)-(38)"}],"minor_comments":[{"comment":"The sentence 'out of which the AIG is build as DS(IA,IB,I0) = DS(IA,IB) − DS(IA,I0)' has the sign reversed; it contradicts Eq. (6), Eq. (15), and the immediately following expression ⟨ln(P(s|IB)/P(s|I0))⟩_{s|IA}.","section":"Sec. 3.1, Eq. (14)"},{"comment":"The Lagrange-multiplier term λ(1 − ∫ P(s''|IB) ds'') is non-local in P(s''|IB), so the literal statement that the gain 'should only depend on P(s'|IB), but not on any probability he assigns to other cases' is not consistent with the displayed functional. Please clarify that Locality is required only on the simplex of normalized distributions.","section":"Sec. 4.1, Eq. (36)"},{"comment":"The quantity fB = day/decade is a dimensionless fraction, approximately 2.7 × 10^-4, not '0.27 h'. The sentence 'fB = day/decade ≈ 0.27 h' mixes units, and the subsequent Euro estimate should be re-expressed with a correctly unitized fraction of facility time.","section":"Sec. 6.3"},{"comment":"The expression '⟨DS(IA(d),IB(d),I0⟩' has mismatched angle brackets; it should read '⟨DS(IA(d),IB(d),I0)⟩_{d|I0}'.","section":"Sec. 5.9, Eq. (102)"},{"comment":"The text says the gain should be 'infinitely differentiable' and then concludes 'this means it must be analytical.' Infinite differentiability does not imply real analyticity; the wording should be corrected to avoid a false mathematical implication.","section":"Sec. 4.1, Analytical axiom"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a solid and useful contribution: the example calculations are correct and the practical estimator in Section 5.9 is a genuine asset. The main issue is that the axiomatic uniqueness claim overreaches relative to the axioms actually used. I recommend major revision rather than rejection because the gap can be fixed within the manuscript's scope: either justify Locality from more basic principles, or explicitly present the result as a characterization of AIG among local gains and discuss non-local alternatives such as the quadratic scoring rule. The sign error in Section 3.1 and the unit error in Section 6.3 should also be corrected."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The definition AIG = D(IA,I0) − D(IA,IB) is simple, sign-sensitive, and reduces to known quantities in the right limits; the worked examples from Bernoulli to Gaussian and mean-field are correct and genuinely instructive. The paper is a solid subfield-level contribution: it gives the information-theory community a useful three-state measure of achieved information and connects it to scoring rules and cognitive fidelity/efficiency.\n\nThe main weakness is the axiomatic derivation in Section 4. The uniqueness claim rests almost entirely on the Locality axiom, which says the gain for a perfectly informed Alice depends only on Bob's probability for the true state. The paper offers no independent justification for that constraint, and it is not innocent: a gain built from the quadratic scoring rule satisfies Additivity, Analyticity, Properness and Calibration but is not AIG. So the phrase 'right quantitative measure' is too strong; AIG is the unique local gain under those axioms, not the unique gain. The definition itself is still well motivated.\n\nTwo smaller issues. Section 5.8 says the apparent information gain is always larger than the achieved one. That is false in general; their own Bernoulli example (Fig. 2, pB=0.6, low p0) shows AIG exceeding the apparent gain. The claim holds in the particular incomplete-data illustrations, but the wording overreaches. Section 5.10 introduces an attention-weighted AIG without the promised derivation; the text says it is omitted for brevity. That is minor since it is an extension, but the paper should either provide the derivation or clearly label it as conjectural.\n\nNone of this undermines the central object. The formulas check out, the examples are reproducible from the text, and the limitations are mostly stated. I would send this to a competent referee; a revision that softens the uniqueness claim and fixes the 'always' sentence will make it a dependable contribution. The paper is most valuable for people evaluating approximate Bayesian inference or designing resource-aware algorithms; it will get cited for the AIG definition and the Gaussian formulas.","headline":"AIG is a clean and useful three-state information measure; the uniqueness claim leans too hard on the Locality axiom, but the definition and examples survive a fair reading.","tokens_in":28588,"tokens_out":6455,"would_cite":true,"duration_ms":58265,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["94A15","94A17","62B10","62F15"],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper proposes achieved information gain (AIG), defined as the ideal information gain minus the remaining gain after an imperfect update, and argues it is the correct quantitative measure of how much information a cognitive operation…","keywords":["achieved information gain","relative entropy","belief updating","cognitive fidelity","cognitive efficiency","scoring rules","Bayesian inference","sustainable computing"],"falsifier":"Take a Bernoulli situation with true rate $p_A=0.6$, initial belief $p_0=0.5$, and update to $p_B=0.9$: apparent gain is positive while AIG is negative, so observe whether decision-makers acting on $p_B$ do better or worse than those acting on $p_0$; if the wrong update reliably improves downstream decisions, the claim that AIG measures achieved information is undermined. Alternatively, replace the locality axiom by a non-local scoring-rule condition and check whether a different gain satisfying the remaining axioms can be constructed.","tokens_in":27629,"feed_emoji":"🧠","tokens_out":6938,"duration_ms":70362,"temperature":0.7,"pith_summary":"Relative entropy (the Kullback-Leibler divergence) rewards any update that makes a belief more definite, even if the new certainty is wrong. The paper proposes achieved information gain (AIG), which measures the information obtained in an imperfect cognitive update by subtracting what would still be needed to reach the ideal state from what an ideal update would have provided. Concretely, AIG is the expected log ratio of Bob's updated belief to his initial belief, averaged under Alice's ideal distribution. It is zero for no update, maximal for a perfect update, and negative when the update moves belief away from the truth. The paper derives AIG from axioms and shows how ratios of AIG define cognitive fidelity and cognitive efficiency for communication, inference, and memorization.","feed_headline":"A new measure tells when a belief update goes the wrong way","feed_subtitle":"It scores updates against the ideal belief, so confident but wrong updates count as negative information instead of positive.","key_machinery":"The central object is the three-state information gain $D_S(I_A,I_B,I_0) = \\langle \\ln[P(s|I_B)/P(s|I_0)] \\rangle_{s|I_A}$: an expected log-likelihood ratio of the updated belief to the initial belief, averaged over the ideal belief. It carries the argument because it makes direction matter: approaching the ideal belief yields positive gain, standing still yields zero, and moving away yields negative gain. The axiomatic derivation in Section 4 forces this logarithmic form by combining additivity over Alice's atomic beliefs with locality, which says that when Alice knows the exact state the gain depends only on Bob's probability for that state; properness then fixes the sign and calibration fixes the reference point.","core_discovery":"The central claim is that a belief update should be scored by where it ends relative to both where it started and where it should have ended, not by how much confidence it added. AIG, defined as $D_S(I_A,I_B,I_0) := D_S(I_A,I_0) - D_S(I_A,I_B) = \\langle \\ln[P(s|I_B)/P(s|I_0)] \\rangle_{s|I_A}$, is that score. Under axioms of additivity, locality, analyticity, properness, and calibration, the paper shows that any such gain must be proportional to AIG, so the logarithmic form is not a choice but a consequence. AIG inherits useful structural properties: it is anti-symmetric in the two Bob states, path-additive, reduces to ordinary relative entropy for perfect updates, and separates across variables even when the ideal distribution is correlated. The paper works through Bernoulli, binomial, Poisson, Beta, and Gaussian updates, mean-field approximations, incomplete data usage, and sampling-based estimates for intractable posteriors, and uses AIG to define cognitive fidelity (the ratio of AIG to ideal gain) and cognitive efficiency (the ratio of AIG to cost).","pith_inferences":["The uniqueness result is only as strong as the locality axiom; a reader who prefers scoring rules that also reward correct probabilities assigned to non-truth states will get a different measure, so testing the appeal of locality is the natural way to probe the paper's central claim.","AIG gives an immediate testable prediction for human or machine learning agents: updates that agents report as informative should correlate positively with AIG computed against a trusted posterior, and confidently wrong updates should correlate negatively.","In practice, AIG measures information relative to whoever supplies the ideal state; without an agreed-upon reference belief, its numerical value is a statement about agreement with that reference, not about absolute truth.","Because AIG is path-additive, it could serve as an online monitoring signal: accumulate gains along a chain of updates and flag episodes where the cumulative gain drops below zero."],"forward_implications":["Apparent information gain can be arbitrarily misleading; replacing it with AIG makes wrong-direction updates report negative information instead of positive confidence.","Cognitive fidelity and cognitive efficiency give a quantitative basis for choosing between cheap approximate methods and expensive accurate ones, including the trade-off between computation and data acquisition costs.","In repeated measurements AIG typically grows only logarithmically with data set size, so a low-fidelity method may need far more data to match the gain of a high-fidelity one.","For intractable posteriors, AIG can be estimated by sample averages when samples from the ideal posterior are available.","AIG is path-additive and anti-symmetric in the two belief states of the learner, so it behaves like a directed distance rather than a divergence."],"supporting_citations":[{"why":"defines relative entropy, the classical measure that AIG extends and contrasts with.","marker":"[1]"},{"why":"supplies the theory of strictly proper scoring rules against which AIG's locality and properness are positioned.","marker":"[8]"},{"why":"establishes the scoring-rule perspective on relative entropy that AIG is built from.","marker":"[10]"},{"why":"supplies the axiomatic style and the attention-entropy framework that the derivation extends.","marker":"[13]"},{"why":"provides the geometric variational inference whose Newton-type updates target increased AIG.","marker":"[22]"},{"why":"motivates memorization as information loss that AIG quantifies in the compressed-memory example.","marker":"[37]"},{"why":"motivates the sustainability and resource-allocation application of cognitive efficiency.","marker":"[50]"},{"why":"provides the Gaussian posterior update formula used in the incomplete-data example.","marker":"[60]"}],"fun_headline_variants":["New measure scores belief updates, can go negative when wrong","Axioms force the only sensible measure of information gain","Belief updates scored against ideal, wrong ones get negative marks","Measure of imperfect cognition: update info is gain minus ideal gap","Why confident but wrong updates are worse than no update at all"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire uniqueness argument hinges on the assumption that, when the ideal belief is certain a particular state is true, the score of an update should depend only on how much probability the updated belief gives to that state and not on what it says about other states.","fun_headline_variants_meta":{"raw":{"variants":["New measure scores belief updates, can go negative when wrong","Axioms force the only sensible measure of information gain","Belief updates scored against ideal, wrong ones get negative marks","Measure of imperfect cognition: update info is gain minus ideal gap","Why confident but wrong updates are worse than no update at all"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000705,"raw_usage":{"total_tokens":3218,"prompt_tokens":1022,"completion_tokens":2196,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":638,"completion_tokens_details":{"reasoning_tokens":2112}},"tokens_in":638,"tokens_out":2196,"duration_ms":16730,"temperature":1.0,"reasoning_tokens":2112,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-08T23:35:24.290471+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a Bernoulli situation with true rate $p_A=0.6$, initial belief $p_0=0.5$, and update to $p_B=0.9$: apparent gain is positive while AIG is negative, so observe whether decision-makers acting on $p_B$ do better or worse than those acting on $p_0$; if the wrong update reliably improves downstream decisions, the claim that AIG measures achieved information is undermined. Alternatively, replace the locality axiom by a non-local scoring-rule condition and check whether a different gain satisfying the remaining axioms can be constructed.","supporting_citations":[{"cited_title":"Strictly proper scoring rules, prediction, and estimation","cited_arxiv_id":null,"evidence_quote":"supplies the theory of strictly proper scoring rules against which AIG's locality and properness are positioned."},{"cited_title":"Optimal Belief Approximation","cited_arxiv_id":"1610.09018","evidence_quote":"establishes the scoring-rule perspective on relative entropy that AIG is built from."},{"cited_title":"Attention to Entropic Communication","cited_arxiv_id":"2307.11423","evidence_quote":"supplies the axiomatic style and the attention-entropy framework that the derivation extends."},{"cited_title":"Probability of error for optimal codes in a Gaussian channel","cited_arxiv_id":null,"evidence_quote":"motivates memorization as information loss that AIG quantifies in the compressed-memory example."}],"review_version":1}