REVIEW 2 major objections 6 minor 43 references
This paper argues that curiosity is a weighted choice among four values—immediate uncertainty reduction, cost, delayed return, and keeping questions open—and that the weights drift with experience and ecology, extending to multi-agent disco
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-02 08:16 UTC pith:YLE2VQJJ
load-bearing objection Useful conceptual synthesis, but the key worked example in Section 3 drops the O_i term and that can reverse the conclusion; still deserves a serious referee. the 2 major comments →
A framework for single and multi-agent human-AI curiosity ecosystems
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The core claim is that an agent's inquiry policy is not fixed: four weights—ψ for immediate uncertainty reduction, λ for cost, τ for long-horizon return, and φ for the value of keeping a question open—define the curiosity policy, and these weights change over time through a drift operator that depends on the agent's recent experience and the surrounding ecology. The paper shows, through an explicit inequality, that the same history of cheap answers can push an agent toward shallow inquiry in one regime (habituation) but not in another (abundance), because the cost-weight drift flips sign. It then extends the framework to populations, introducing a shared knowledge stock that grows only throu
What carries the argument
The central object is the value function Vi(q,t) = ψi E[Ii] − λi E[Ci] + τi E[Li] − φi E[Oi], whose four weights define the agent's curiosity policy. The drift operator θi(t+1) = θi(t) + Δθi(Xi(t), Mt) makes the weights move with a smoothed experience state Xi and ecology Mt; a linear-with-decay specification is given as one implementation. For multi-agent systems, the key mechanism is the shared knowledge stock St, which evolves by crediting reusable, non-redundant contributions once even if several agents pursue the same question, and the population-performance index RSt = A S Q^aQ (εD + D)^aD (εF + F)^aF. The workhorse inequality is Eq. 5, which determines when drift strengthens the pull
Load-bearing premise
The load-bearing premise is that the weights in the value function actually drift according to the posited sign patterns—specifically, that repeated cheap, low-generativity answers reduce the weight on long-horizon return and increase cost sensitivity; the paper itself flags in Section 8 that it is an open empirical question whether curiosity weights drift in the manner the framework captures.
What would settle it
Run a controlled study where participants first receive a block of cheap, quickly answered, low-generativity questions (fast resolution, no follow-up), then choose between more such questions and a difficult, high-generativity frontier question. The habituation regime predicts Δλ>0 and Δτ<0, making participants avoid the frontier question. If instead participants become more likely to pursue the frontier question, or if their long-horizon weight does not drop (measured through choice patterns or reported value), the central drift scenario fails. A second arm varies the ecology—abundant cheap q
If this is right
- If the framework holds, a person's or AI agent's curiosity policy is experience-dependent and ecology-dependent, so the same environment can produce opposite drift directions depending on which channel dominates.
- The inequality in Eq. 5 gives a quantitative criterion: cheap inquiry degrades curiosity only when the immediacy channel outweighs the cost channel; in abundance regimes the cost channel works against shallow inquiry.
- Multi-agent systems do not grow knowledge by adding agents alone; they need division of labor, a shared inspectable state, and reusable outputs, and the framework's parameters η and ζ presuppose that structure.
- The population-performance index RSt predicts that increases in inquiry volume only help if they are not outweighed by losses in topic diversity or frontier-directed effort.
- The framework sets up testable predictions about adaptive versus maladaptive curiosity, and about when peer coupling coordinates versus homogenizes a population.
Where Pith is reading between the lines
- The drift operator suggests an intervention: deliberately structuring an agent's inquiry history (e.g., interspersing generative, open-ended questions among quick ones) could steer weights toward long-horizon inquiry, which may inform educational curricula and content recommendation design.
- The abundance-regime prediction—that abundant cheap answers can lower cost-sensitivity and keep frontier questions attractive—could be tested experimentally by varying the base rate of cheap versus frontier questions and measuring choice behavior, not just the history itself.
- The framework's shared knowledge stock, with its once-credited reusable contributions, offers a natural way to quantify 'model collapse' in multi-agent AI: parallel sampling from similar priors without functional diversity produces high volume but near-zero growth in St.
- The distinction between empirical and desirable drift operators opens a normative research program: rather than asking only how weights do drift, we can ask how they should drift to maximize useful knowledge, which could guide reward design for discovery-oriented AI.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper proposes a conceptual framework for modeling curiosity as an ecosystem. For a single agent, the value of pursuing a question is decomposed into immediate uncertainty reduction, cost, long-horizon return, and the value of keeping the question open (Eq. 1). The weights on these terms define a 'curiosity policy' that can drift with experience and ecology (Eqs. 2–4). The central example concerns a choice between a cheap and a frontier question: under a 'habituation' regime, repeated cheap answers are argued to strengthen the pull toward the cheap question, while under an 'abundance' regime the cost channel acts in the opposite direction, with the net effect summarized by Eq. (5). The framework is then extended to multiple agents sharing a knowledge stock (Eqs. 6–7) and to population-level metrics of volume, diversity, and frontier share (Eqs. 8–10). The paper explicitly positions itself as a framework for generating future theories and for multi-agent AI design, rather than as an empirically validated model.
Significance. If the framework is accepted as a conceptual scaffold, it provides a useful common language for single-agent curiosity, preference drift, and multi-agent discovery. Its formal decomposition of inquiry value and its treatment of drifting weights are clear and internally consistent except for the issue noted below. The multi-agent extension is a valuable step toward thinking about curiosity at the ecosystem level, and the paper is candid about its limitations: Section 8 explicitly flags that the central drift mechanism is an open empirical question and distinguishes empirical from desirable drift. The paper does not overclaim empirical support and draws on prior work as motivation rather than as proof. Strengths include the explicit formal structure, the transparent treatment of free parameters, and the acknowledgment of competing effects. However, one load-bearing derivation currently omits a term that can reverse the central example, so the paper needs revision before the framework's signature claim is established.
major comments (2)
- [Section 3, Eq. (5)] The derivation of Eq. (5) omits the O_i term from Eq. (1). Under the paper's own habituation sign pattern (Delta_phi < 0) and the natural assumption that a frontier question has a larger option value of keeping the question open (Delta_O = O_c - O_f < 0), the omitted contribution -Delta_O * Delta_phi is negative, opposing the three terms retained in Eq. (5). Thus the claim that 'the inequality holds whenever the drifts take the signs above' is not an implication of the framework as stated. The full inequality is Delta_I Delta_psi + Delta_L Delta_tau > Delta_C Delta_lambda + Delta_O Delta_phi, and the extra term can make the habituation-regime conclusion fail. The paper says 'I ignore the O_i term for simplicity,' but no argument is given that Delta_O is negligible. This needs to be addressed explicitly, either by including the O-term and analyzing its effect or by providing a substantive
- [Section 3, Eqs. (3)-(4)] The sign pattern of the drift operator is posited, not derived. The central claim that the same history of cheap answers can push an agent toward shallow inquiry in one ecology but not another depends on the specific signs of Delta_psi, Delta_lambda, Delta_tau, Delta_phi in each regime. The paper does acknowledge in Section 8 that whether curiosity weights drift as captured by the framework is an empirical question. That acknowledgment is welcome, but the presentation in Section 3 sometimes reads as a stronger claim. I recommend making the conditional nature of the example explicit in the main text (e.g., 'under the assumed sign patterns') and stating more clearly that the framework currently provides an existence proof of a possible mechanism, not a prediction about actual agents.
minor comments (6)
- [Section 2, Eq. (1)] The text says 'All four are written as expectations at time t (E_{i,t}[·])', but the weights themselves (psi, lambda, tau, phi) are not expectations; only the terms I, C, L, O are. Please rephrase to avoid confusion.
- [Section 3, Eq. (4)] The linear update with decay is specified as Eq. (4), but then the text says weights are clipped at zero after each update. This means the actual update is not exactly Eq. (4) when clipping binds. Please clarify whether Eq. (4) describes the unclipped update and whether the subsequent sign analysis assumes clipping is not active.
- [Section 4, Eq. (7)] The shared-knowledge stock evolution depends on U(q,K_t) and d(q,K_t), but these are only described verbally, not defined operationally. For a conceptual framework this may be acceptable, but a remark about how these might be computed in future implementations would help.
- [Section 5, Eq. (10)] The population-performance index R_t^S is introduced with scale constant A_S and exponents a_Q, a_D, a_F and floors epsilon_D, epsilon_F. The text says 'small proportional changes satisfy Delta log R = ...' but this holds exactly only when D_t and F_t are positive; the special cases D_t=0 or F_t=0 are handled by convention. Please state this explicitly near the displayed equation.
- [Section 7] The sentence 'Differentially, in the model in the current manuscript, generativity g_i is not a uniform property of interchangeable items but varies from question to question' reads awkwardly. Consider rewording to 'In contrast, in the present model...'
- [Section 8] Typo: 'in a a graph form' should be 'in a graph form'. Also, Section 7 has 'it be can be adaptive' which should be 'it can be adaptive'.
Circularity Check
No significant circularity: the framework is conditional, non-fitted, and self-citations are motivational background.
full rationale
The manuscript is an explicitly conceptual framework rather than an empirical derivation. Its central formal step, Eq. 5 in Section 3, follows algebraically from Eq. 1 once the habituation/abundance drift signs are assumed; the paper does not claim those signs are derived from data, and Section 8 explicitly states: "An important empirical question is whether curiosity weights drift in a manner that is captured by the framework." No fitted parameter is renamed as a prediction: the drift operator, regime signs, and pursuit probability are proposed modeling choices, and the paper even disclaims prediction ("The model does not predict that cheaper access to answers always degrades curiosity"). The self-citations (Monosov 2024; Bromberg-Martin and Monosov 2020; Jezzini et al. 2021) are motivational or neurobiological background and are not used as load-bearing uniqueness constraints or to force any conclusion. The one analytical gap noted by a reader—ignoring the O_i term in Eq. 5—is an internal completeness/correctness issue, not circularity, because the conclusion is explicitly conditional on that simplification rather than obtained by assuming the conclusion. The population-performance index Eq. 10 is a defined measure, and the diversity/frontier statements are immediate implications of the Cobb-Douglas form, not disguised empirical claims. Therefore no circular step can be exhibited, and the score is 0.
Axiom & Free-Parameter Ledger
free parameters (12)
- ℓ0
- ρ_i
- ρ^soc_i
- κ_i
- η_i
- ζ_i
- ω_d, ω_p
- γ_X
- ε, W, w_M
- A_S, a_Q, a_D, a_F, ε_D, ε_F
- δ
- pursuit-rule parameters (threshold, temperature, outside-option value)
axioms (7)
- domain assumption The value of a question decomposes linearly into immediate uncertainty reduction, cost, delayed return, and open-question value with time-varying weights (Eq. 1).
- domain assumption Preferences depend on past consumption through a smoothed experience state X_i(t) with persistence γ_X (Eq. 2).
- domain assumption Weights drift according to θ_i(t+1)=θ_i(t)+Δθ_i(X_i(t),M_t), with clipping at zero (Eq. 3).
- domain assumption Generativity g_i(q,t) counts gross new knowledge gaps opened, and long-horizon own return is linear in g_i with a baseline ℓ0.
- domain assumption Agents' pursuit and success events are independent conditional on the current state (Eq. 7).
- ad hoc to paper The population-performance index R^S_t takes a Cobb–Douglas form with floors (Eq. 10).
- ad hoc to paper The example drift rule (Eq. 4) is a linear update with decay; other update rules can be substituted.
invented entities (4)
-
Curiosity policy θ_i(t)
no independent evidence
-
Experience state X_i(t)
no independent evidence
-
Shared knowledge stock S_t
no independent evidence
-
Reusable value U(q,K_t)
no independent evidence
read the original abstract
This paper offers a framework for considering curiosity as an ecosystem. First, it suggests that a single agent's inquiry policy (how, when, and why an agent asks a question) depends on how the agent values immediate uncertainty reduction, costs, delayed return, and the value of keeping the question open. A key concept in the framework is that the weights on these decision-related terms can change with experience. For example, a period of cheap, quickly answered questions may change the cost of inquiry on a short timescale and change which kinds of questions the agent is drawn to answer over a longer timescale. Second, these ideas are extended to many agents exploring a shared knowledge landscape, and there the framework tracks inquiry volume, topic diversity, frontier-directed inquiry, redundancy, and reusable knowledge. The result is a conceptual framework for studying curiosity ecology and for future efforts towards designing multi-agent AI systems for discovery.
Reference graph
Works this paper leans on
-
[1]
, title =
Arrow, Kenneth J. , title =. The Rate and Direction of Inventive Activity , editor =. 1962 , doi =
1962
-
[2]
Econometrica , volume =
Aghion, Philippe and Howitt, Peter , title =. Econometrica , volume =. 1992 , doi =
1992
-
[3]
, title =
Averbeck, Bruno B. , title =. PLoS Computational Biology , volume =. 2015 , doi =
2015
-
[4]
and Murphy, Kevin M
Becker, Gary S. and Murphy, Kevin M. , title =. Journal of Political Economy , volume =. 1988 , doi =
1988
-
[5]
and Sharot, Tali , title =
Bromberg-Martin, Ethan S. and Sharot, Tali , title =. Neuron , volume =. 2020 , doi =
2020
-
[6]
and Monosov, Ilya E
Bromberg-Martin, Ethan S. and Monosov, Ilya E. , title =. Current Opinion in Behavioral Sciences , volume =. 2020 , doi =
2020
-
[7]
and Feng, Yang-Yang and Ogasawara, Takaya and White, J
Bromberg-Martin, Ethan S. and Feng, Yang-Yang and Ogasawara, Takaya and White, J. Kael and Zhang, Kaining and Monosov, Ilya E. , title =. Nature Neuroscience , volume =. 2024 , doi =
2024
-
[8]
and Merel, Josh and Monosov, Ilya E
Bromberg-Martin, Ethan S. and Merel, Josh and Monosov, Ilya E. , title =. bioRxiv , year =
-
[9]
and Bromberg-Martin, Ethan S
Charpentier, Caroline J. and Bromberg-Martin, Ethan S. and Sharot, Tali , title =. Proceedings of the National Academy of Sciences USA , volume =. 2018 , doi =
2018
-
[10]
and Douglas, Paul H
Cobb, Charles W. and Douglas, Paul H. , title =. American Economic Review , volume =
-
[11]
and Averbeck, Bruno B
Costa, Vincent D. and Averbeck, Bruno B. , title =. Journal of Neuroscience , volume =. 2020 , doi =
2020
-
[12]
and Pindyck, Robert S
Dixit, Avinash K. and Pindyck, Robert S. , title =. 1994 , doi =
1994
-
[13]
and Hauser, Oliver P
Doshi, Anil R. and Hauser, Oliver P. , title =. Science Advances , volume =. 2024 , doi =
2024
-
[14]
and Zin, Stanley E
Epstein, Larry G. and Zin, Stanley E. , title =. Econometrica , volume =. 1989 , doi =
1989
-
[15]
and Rzhetsky, Andrey and Evans, James A
Foster, Jacob G. and Rzhetsky, Andrey and Evans, James A. , title =. American Sociological Review , volume =. 2015 , doi =
2015
-
[16]
Psychological Review , volume =
Gigerenzer, Gerd and Garcia-Retamero, Rocio , title =. Psychological Review , volume =. 2017 , doi =
2017
-
[17]
Journal of Economic Literature , volume =
Golman, Russell and Hagmann, David and Loewenstein, George , title =. Journal of Economic Literature , volume =. 2017 , doi =
2017
-
[18]
Trends in Cognitive Sciences , volume =
Gottlieb, Jacqueline and Oudeyer, Pierre-Yves and Lopes, Manuel and Baranes, Adrien , title =. Trends in Cognitive Sciences , volume =. 2013 , doi =
2013
-
[19]
and Trambaiolli, Lucas R
Jezzini, Ahmad and Bromberg-Martin, Ethan S. and Trambaiolli, Lucas R. and Haber, Suzanne N. and Monosov, Ilya E. , title =. Neuron , volume =. 2021 , doi =
2021
-
[20]
and Porteus, Evan L
Kreps, David M. and Porteus, Evan L. , title =. Econometrica , volume =. 1978 , doi =
1978
-
[21]
Psychological Bulletin , volume =
Loewenstein, George , title =. Psychological Bulletin , volume =. 1994 , doi =
1994
-
[22]
, title =
March, James G. , title =. Organization Science , volume =. 1991 , doi =
1991
-
[23]
Quarterly Journal of Economics , volume =
McDonald, Robert and Siegel, Daniel , title =. Quarterly Journal of Economics , volume =. 1986 , doi =
1986
-
[24]
, title =
Monosov, Ilya E. , title =. Nature Reviews Neuroscience , volume =. 2024 , doi =
2024
-
[25]
Psychological Review , volume =
Murayama, Kou , title =. Psychological Review , volume =. 2022 , doi =
2022
-
[26]
and Winter, Sidney G
Nelson, Richard R. and Winter, Sidney G. , title =
-
[27]
International Conference on Learning Representations (ICLR) , year =
Padmakumar, Vishakh and He, He , title =. International Conference on Learning Representations (ICLR) , year =. 2309.05196 , archivePrefix =
-
[28]
, title =
Romer, Paul M. , title =. Journal of Political Economy , volume =. 1990 , doi =
1990
-
[29]
and Heal, Geoffrey M
Ryder, Harl E. and Heal, Geoffrey M. , title =. Review of Economic Studies , volume =. 1973 , doi =
1973
-
[30]
Read , title =
Schultz, Wolfram and Dayan, Peter and Montague, P. Read , title =. Science , volume =. 1997 , doi =
1997
-
[31]
, title =
Shannon, Claude E. , title =. Bell System Technical Journal , volume =. 1948 , doi =
1948
-
[32]
, title =
Sharot, Tali and Sunstein, Cass R. , title =. Nature Human Behaviour , volume =. 2020 , doi =
2020
-
[33]
Nature , volume =
Shumailov, Ilia and Shumaylov, Zakhar and Zhao, Yiren and Papernot, Nicolas and Anderson, Ross and Gal, Yarin , title =. Nature , volume =. 2024 , doi =
2024
-
[34]
and Becker, Gary S
Stigler, George J. and Becker, Gary S. , title =. American Economic Review , volume =
-
[35]
and Barto, Andrew G
Sutton, Richard S. and Barto, Andrew G. , title =. 2018 , url =
2018
-
[36]
Management Science , volume =
Tversky, Amos and Simonson, Itamar , title =. Management Science , volume =. 1993 , doi =
1993
-
[37]
Science , volume =
Uzzi, Brian and Mukherjee, Satyam and Stringer, Michael and Jones, Ben , title =. Science , volume =. 2013 , doi =
2013
-
[38]
, title =
Weitzman, Martin L. , title =. Quarterly Journal of Economics , volume =. 1998 , doi =
1998
-
[39]
and Geana, Andra and White, John M
Wilson, Robert C. and Geana, Andra and White, John M. and Ludvig, Elliot A. and Cohen, Jonathan D. , title =. Journal of Experimental Psychology: General , volume =. 2014 , doi =
2014
-
[40]
, title =
Pratt, John W. , title =. Econometrica , volume =. 1964 , doi =
1964
-
[41]
Econometrica , volume =
Kahneman, Daniel and Tversky, Amos , title =. Econometrica , volume =. 1979 , doi =
1979
-
[42]
Annual Review of Economics , volume =
Gabaix, Xavier , title =. Annual Review of Economics , volume =. 2009 , doi =
2009
-
[43]
Nature Human Behaviour , year =
Wiradhany, Wisnu and Parry, Douglas and Aru, Jaan , title =. Nature Human Behaviour , year =
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.