REVIEW 3 major objections 4 minor 62 references
Optimal ambition in business, politics and life
T0 review · 3 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read The paper proves that in a finite search, the optimal satisfaction threshold is finite and strictly above the mean reward, and that overshooting that threshold costs more than undershooting it.
desk verdict A readable but overreaching formalization of folk ambition: the i.i.d. core is sound enough, yet the smooth-landscape extension rests on a wrong variance formula and the proof of $T^*>\mu$ has an unproven unimodality step. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the satisfaction threshold T, expressed as a number of standard deviations above or below the mean reward μ, which controls an AR(1) explore-exploit process: while rewards sit below T the searcher samples a new reward, and once a reward reaches T the searcher keeps it. The key identity is equation (2), which expresses the expected cumulative reward as a sum over the exploration length t_x of (μ_explore(T) t_x + μ_exploit(T)(t_max - t_x))(1 - Φ(T)) Φ(T)^{t_x}, with μ_explore(T) = -φ(T)/Φ(T) and μ_exploit(T) = φ(T)/(1-Φ(T)), where φ(·) and Φ(·) are the standard normal density and CDF. This identity makes the expected reward a unimodal hump-shaped function of T, and the proof that each summand is maximized at a positive threshold is what carries the claim that T* > μ.
What would settle it
Simulate equation (1) with φ close to 1 and a Gaussian reward distribution, compute the expected cumulative reward over a fixed horizon for a fine grid of thresholds T, and check whether the argmax exceeds the mean μ; if the optimum reaches or falls below μ for any φ>0, the general claim fails.
Extended reading notes
Core claim
The authors claim to prove that in a finite-horizon search, the optimal satisfaction threshold T* satisfies μ < T* < ∞, where μ is the mean reward across strategies. The proof is worked out for the maximally rugged Gaussian case, where successive rewards are independent, by writing the expected reward as a sum over the possible length of the exploration phase and showing that every summand peaks at a threshold above zero. The authors then argue, from numerical simulation and by scaling the threshold with the variance of the autoregressive process, that the same conclusion holds on smooth autocorrelated landscapes and under constant search costs, provided the costs do not eliminate the incentive to search altogether. They also derive the asymmetry that overambition is costlier than caution, and they connect the threshold logic to empirical patterns in online dating and college applications.
Load-bearing premise
The analytic proof that the optimal threshold exceeds the mean is carried out only for the maximally rugged Gaussian case (φ=0), and the extension to smooth, autocorrelated landscapes rests on a variance scaling that is not derived from the model's own recurrence equation.
Editorial extensions
If this is right
- A finite search horizon has a sweet spot: the always-settle and never-settle strategies both earn the mean on average, while a threshold strictly above the mean earns more.
- Longer searches justify higher ambition; as the horizon grows, the optimal threshold rises, and only an infinite horizon permits unbounded ambition.
- Reward landscapes matter: rugged, weakly autocorrelated landscapes and left-skewed reward distributions call for higher thresholds, while right-skewed distributions call for thresholds closer to the mean.
- Search costs lower the optimal threshold and the value of search, but the above-mean result holds as long as searching is profitable at all.
- Upward social comparison is doubly harmful: it raises the perceived mean and lowers the threshold, reducing both satisfaction and earned reward.
Reading between the lines
- For practical settings, the model suggests a measurable rule: set aspiration levels from the median or mean of the observable outcome distribution, then shift only modestly upward, rather than anchoring on the top performers.
- The difference between ambition and risk-taking means that policy advice should separate the two margins: left-skewed environments call for cautious actions but ambitious targets relative to the mean.
- A laboratory or natural experiment that lengthens the decision horizon should produce higher observed aspiration thresholds; this prediction is directly testable.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper formalizes a finite-horizon search model in which an agent repeatedly samples rewards from an autoregressive process and stops searching once the current reward reaches a satisfaction threshold T. The central claims are that the optimal T is finite and strictly above the mean reward, that overshooting T is costlier than undershooting it, that longer time horizons and left-skewed or rugged reward landscapes raise optimal ambition, that upward social comparison harms performance, and that search costs do not overturn the main result unless they make all search unprofitable. An analytic expression for expected reward is derived for the Gaussian, maximally rugged (φ=0) case, and the other results are obtained by analytic scaling or simulation. The paper closes with qualitative applications to dating, college admissions, economic policy, wealth, and elections.
Significance. If the central theorem is established, the paper provides a simple and intuitive formalization of 'finite but above-average ambition' with a clean asymmetry result and several falsifiable qualitative predictions. The analytic expression in Eq. (2), the openly available simulation code, and the transparent numerical experiments are strengths. The paper is also careful to present the empirical examples as illustrations rather than as fitted validations. However, the significance is currently limited by two load-bearing gaps: the proof for the Gaussian rugged case relies on an unproven unimodality assertion, and the extension to smooth (φ>0) landscapes is based on an internally inconsistent variance formula. These gaps affect the abstract's general claim that the optimal threshold is strictly larger than the mean reward across reward landscapes. The manuscript is therefore promising but requires substantive revision before the central claims can be accepted.
major comments (3)
- [Section C (page 5) and Fig. 3a] The statement that the AR(1) process has variance var[X_t]=(1+φ)/(1−φ) is incorrect and self-contradictory. For Eq. (1), X_t=φX_{t−1}+(1−φ)ε_t with ε_t i.i.d. N(0,1), the stationary variance is (1−φ)^2/(1−φ^2)=(1−φ)/(1+φ). The very next sentence says that 'autocorrelation reduces the subsequent variance on smooth landscapes,' which is consistent with the correct formula and not with the printed one. Because the analytic smooth-landscape curves in Fig. 3a are generated by scaling the threshold by 1/sqrt(var[X_t]) using this erroneous variance, the reported dependence of optimal T on φ is not supported. The manuscript should either correct the variance calculation and redo the scaling, or preferably simulate Eq. (1) directly for φ>0 and report whether T*>μ continues to hold.
- [Appendix A.4.d] The proof that the optimal threshold is strictly greater than the mean for φ=0 is incomplete. The key step asserts that every summand R(T,t_x)P(T,t_x) is unimodal, but this is not proven. The product of a reward component that grows roughly linearly in T and a probability component that is unimodal need not be unimodal, and the fact that the derivative of each summand at T=0 is positive does not by itself establish that the maximum of the total expected reward lies at T>0. The manuscript needs either a rigorous verification of unimodality (with explicit conditions) or a numerical proof/argument covering the relevant range of t_max. Without this, the main theorem for the rugged Gaussian case is not fully established.
- [Appendix A.2 and abstract] The paper claims that the optimal threshold increases with the search time t_max, but Appendix A.2 provides only a verbal intuition rather than a proof. Since this monotonicity is presented as one of the main results and is used in the interpretation of Fig. 2a, a formal argument or a direct numerical demonstration should be supplied. In addition, the abstract states that the optimal threshold is 'strictly larger than the mean of available rewards' as a general theorem, yet the proof is restricted to the φ=0 Gaussian case and the smooth-landscape extension is invalidated by the variance error in Section C. The claims should be restricted to the cases actually proven, or the missing cases should be established by correct analysis or simulation.
minor comments (4)
- [Section C, paragraph beginning 'Rugged landscapes create...'] There is a typo: 'aurocorrelation' should be 'autocorrelation.'
- [Appendix C, after Fig. 11] The text appears to contain a long corrupted string of '/uni...' tokens. If this is present in the source file rather than an artifact of extraction, it should be removed or replaced with the intended figure caption or text.
- [Reference [44]] Reference [44] is incomplete: it gives authors, title, and year but no journal, volume, or preprint identifier. Please provide full bibliographic information.
- [Equation (2)] The derivation of Eq. (2) assumes that the reward distribution is standard normal with μ=0 and σ=1; this should be stated explicitly before the equation, along with the affine rescaling that recovers general μ and σ.
Circularity Check
No significant circularity; the optimal-threshold result is derived from the model's own equations, with post-hoc examples and peripheral self-citations that are not load-bearing.
full rationale
The paper's central claim, that the optimal satisfaction threshold T* is finite and strictly above the mean reward, is derived within the model from the truncated-normal Mills-ratio structure of equation (2) and the Appendix A argument for the maximally rugged Gaussian case (phi = 0). The threshold is defined relative to the reward distribution's mean and standard deviation, but the conclusion T* > mu is not equivalent to that definition; it is a nontrivial optimization result. The smooth-landscape extension in Section C is a scaling argument that uses the AR(1) variance from the model's own dynamics, and the social-comparison and search-cost results are computed from the same equations or from simulations of equation (1). The empirical applications (dating, college applications, GDP growth, wealth, elections) are post-hoc illustrations and do not enter the derivation or supply fitted parameters. The only author-overlapping references, such as ref. [7] on fisheries and ref. [44] on left-skewed economic growth, support peripheral illustrative claims rather than the formal results. The apparent inversion of the AR(1) stationary variance in Section C, and the resulting gap in the generality of T* > mu for phi > 0, is a correctness and internal-consistency concern, not a circularity: no step reduces a prediction to a fitted input or to a self-citation chain. The derivation is self-contained against its own stated assumptions.
Assumptions & free parameters
assumptions (6)
- domain assumption Rewards follow the threshold-stopped AR(1) process in Eq. (1): X_t = φ X_{t-1} + (1-φ) ε_t if X_{t-1} < T, else X_t = X_{t-1}.
- domain assumption Agents maximize the undiscounted expected sum of rewards over t_max periods and have no risk aversion.
- domain assumption For the main proof, agents know the true reward distribution (mean μ and variance σ²).
- standard math In the Gaussian φ=0 case, the expected exploration reward is the inverse Mills ratio μ_explore(T) = -φ(T)/Φ(T) and the expected exploitation reward is μ_exploit(T) = φ(T)/(1-Φ(T)).
- ad hoc to paper The expected reward function E[reward](T) is unimodal, or at least its maximum lies at T > 0 when t_max > 1, and each summand R(T,t_x)P(T,t_x) can be treated as unimodal in the proof.
- ad hoc to paper For φ > 0, the expected reward can be obtained by rescaling the threshold by 1/sqrt(var[X_t]) with var[X_t] = (1+φ)/(1-φ).
Cite this review
Pith. "Pith review of Optimal ambition in business, politics and life." pith.science (2026). https://pith.science/paper/3NMDS46O
@misc{pith2026250210500,
author = {Pith},
title = {Pith review of: Optimal ambition in business, politics and life},
year = {2026},
howpublished = {\url{https://pith.science/paper/3NMDS46O}},
note = {Machine review of arXiv:2502.10500}
}
read the original abstract
In business, politics and life, folk wisdom encourages people to aim for above-average results, but to not let the perfect be the enemy of the good. Here, we mathematically formalize and extend this folk wisdom. We model a time-limited search for strategies having uncertain rewards. At each time step, the searcher either is satisfied with their current reward or continues searching. We prove that the optimal satisfaction threshold is both finite and strictly larger than the mean of available rewards -- matching the folk wisdom. This result is robust to search costs, unless they are high enough to prohibit all search. We show that being too ambitious has a higher expected cost than being too cautious. We show that the optimal satisfaction threshold increases if the search time is longer, or if the reward distribution is rugged (i.e., has low autocorrelation) or left-skewed. The skewness result reveals counterintuitive contrasts between optimal ambition and optimal risk taking. We show that using upward social comparison to assess the reward landscape substantially harms expected performance. We show how these insights can be applied qualitatively to real-world settings, using examples from entrepreneurship, economic policy, political campaigns, online dating and college admissions. We discuss implications of several possible extensions of our model, including intelligent search, reward landscape uncertainty and risk aversion.
Figures
Figures from the paper (7 more)
Reference graph
Works this paper leans on
-
[1]
Limiting behavior ast max → ∞ Recall that—in the special case of a rugged landscape with Gaussian rewards—we can write down the expected cumulative reward for a given agent by summing over all the possible lengths of the exploration phase (equation (2)). By substituting the expressions for the inverse Mills ratiosµ explore andµ exploit, we can simplify th...
-
[2]
The optimal satisfaction threshold is finite and increases as the total time increases For any threshold,T, the probability of finding a strat- egy that satisfies the threshold increases with the length of the search,t max. The amount of time one gets to ex- ploit a satisfactory strategy also increases int max, all else equal. Both of these patterns shift...
-
[3]
The optimal satisfaction threshold is strictly greater than the mean reward The reward in the ‘exploit’ phase (µexploit(T)) is always positive and increases inT. The expected number of time steps it takes to satisfy thresholdTis 1 1−Φ(T) : atT= 0, the expected length of the exploration phase is two time steps, and it increases exponentially as the thresho...
-
[4]
M. Jay,The defining decade: why your twenties matter– and how to make the most of them now(Twelve, 2012)
work page 2012
-
[5]
C. E. Lindblom, The science of “muddling through”, Public administration review , 79 (1959)
work page 1959
-
[6]
Grant,Originals: How non-conformists move the world(Penguin, 2017)
A. Grant,Originals: How non-conformists move the world(Penguin, 2017)
work page 2017
-
[7]
Y. Guan, M. B. Arthur, S. N. Khapova, R. J. Hall, and R. G. Lord, Career boundarylessness and career success: A review, integration and guide to future research, Jour- nal of Vocational Behavior110, 390 (2019)
work page 2019
-
[8]
M. T. Hayes,The limits of policy change: Incremental- ism, worldview, and the rule of law(Georgetown Univer- sity Press, 2002)
work page 2002
Show all 62 references
-
[9]
George, S
D. George, S. Luo, J. Webb, J. Pugh, A. Martinez, and J. Foulston, Couple similarity on stimulus characteristics and marital satisfaction, Personality and Individual Dif- ferences86, 126 (2015)
2015
-
[10]
M. G. Burgess, E. Carrella, M. Drexler, R. L. Axtell, R. M. Bailey, J. R. Watson, R. B. Cabral, M. Clemence, C. Costello, C. Dorsett,et al., Opportunities for agent- based modelling in human dimensions of fisheries, Fish and Fisheries21, 570 (2020)
2020
-
[11]
Skellett, B
B. Skellett, B. Cairns, N. Geard, B. Tonkes, and J. Wiles, Maximally rugged nk landscapes contain the highest peaks, inProceedings of the 7th annual conference on Ge- netic and evolutionary computation(2005) pp. 579–584
2005
-
[12]
Verel, P
S. Verel, P. Collard, and M. Clergue, Where are bottle- necks in nk fitness landscapes?, inThe 2003 Congress on Evolutionary Computation, 2003. CEC’03., Vol. 1 (IEEE, 2003) pp. 273–280
2003
-
[13]
Dimova, J
B. Dimova, J. W. Barnes, and E. Popova, Arbitrary ele- mentary landscapes & AR (1) processes, Applied Math- ematics Letters18, 287 (2005)
2005
-
[14]
MacQueen and R
J. MacQueen and R. G. Miller Jr, Optimal persistence policies, Operations Research8, 362 (1960)
1960
-
[15]
van Moerbeke, Optimal stopping and free boundary problems, The Rocky Mountain Journal of Mathematics 4, 539 (1974)
P. van Moerbeke, Optimal stopping and free boundary problems, The Rocky Mountain Journal of Mathematics 4, 539 (1974)
1974
-
[16]
Freeman, The secretary problem and its extensions: A review, International Statistical Review/Revue Interna- tionale de Statistique , 189 (1983)
P. Freeman, The secretary problem and its extensions: A review, International Statistical Review/Revue Interna- tionale de Statistique , 189 (1983)
1983
-
[17]
W. B. Powell, A unified framework for stochastic opti- mization, European journal of operational research275, 795 (2019)
2019
-
[18]
M. G. Kohn and S. Shavell, The theory of search, Journal of Economic Theory9, 93 (1974)
1974
-
[19]
D. F. Ciocan and V. V. Miˇ si´ c, Interpretable optimal stop- ping, Management Science68, 1616 (2022)
2022
-
[20]
Liu and Y
Z. Liu and Y. Mu, Optimal stopping methods for invest- ment decisions: A literature review, International Jour- nal of Financial Studies10, 96 (2022)
2022
-
[21]
G. H. Pyke, Optimal foraging theory: a critical review, Annual review of ecology and systematics15, 523 (1984)
1984
-
[22]
E. L. Charnov, Optimal foraging, the marginal value the- orem, Theoretical population biology9, 129 (1976)
1976
-
[23]
D. L. Kramer, The behavioral ecology of air breathing by aquatic animals, Canadian Journal of Zoology66, 89 (1988)
1988
-
[24]
R. J. Cowie, Optimal foraging in great tits (parus major), Nature268, 137 (1977)
1977
-
[25]
Schmid-Hempel, A
P. Schmid-Hempel, A. Kacelnik, and A. I. Houston, Hon- eybees maximize efficiency by not filling their crop, Be- havioral Ecology and Sociobiology17, 61 (1985)
1985
-
[26]
G. A. Parker and J. M. Smith, Optimality theory in evo- lutionary biology, Nature348, 27 (1990)
1990
-
[27]
Becker, P
S. Becker, P. Cheridito, and A. Jentzen, Deep optimal stopping, Journal of Machine Learning Research20, 1 (2019)
2019
-
[28]
A. A. Novikov, On distributions of first passage times and optimal stopping of AR (1) sequences, Theory of Probability & Its Applications53, 419 (2009)
2009
-
[29]
Christensen, Phase-type distributions and optimal stopping for autoregressive processes, Journal of Applied Probability49, 22 (2012)
S. Christensen, Phase-type distributions and optimal stopping for autoregressive processes, Journal of Applied Probability49, 22 (2012)
2012
-
[30]
Christensen, A
S. Christensen, A. Irle, and A. Novikov, An elementary approach to optimal stopping problems for AR (1) se- quences, Sequential Analysis30, 79 (2011)
2011
-
[31]
M. Guan, M. Lee, and A. Silva, Threshold models of human decision making on optimal stopping problems in different environments, inProceedings of the annual meeting of the cognitive science society, Vol. 36 (2014)
2014
-
[32]
J. P. Gerber, L. Wheeler, and J. Suls, A social compar- ison theory meta-analysis 60+ years on., Psychological bulletin144, 177 (2018)
2018
-
[33]
Muller and M.-P
D. Muller and M.-P. Fayant, On being exposed to su- perior others: Consequences of self-threatening upward social comparisons, Social and Personality Psychology Compass4, 621 (2010)
2010
-
[34]
E. E. Bruch and M. E. Newman, Aspirational pursuit of mates in online dating markets, Science Advances4, eaap9815 (2018)
2018
-
[35]
C. M. Hoxby and C. Avery,The missing “one-offs’: The hidden supply of high-achieving, low income stu- dents, Tech. Rep. (National Bureau of Economic Re- search, 2012)
2012
-
[36]
W. H. Greene,Econometric analysis(Pearson Education India, 2003)
2003
-
[37]
Benuzzi and M
M. Benuzzi and M. Ploner, Skewness-seeking behavior and financial investments, Annals of Finance20, 129 (2024)
2024
-
[38]
G. R. Goethals and J. M. Darley, Social comparison theory: An attributional approach, Social comparison processes: Theoretical and empirical perspectives , 259 (1977)
1977
-
[39]
J. Suls, R. Martin, and L. Wheeler, Social comparison: Why, with whom, and with what effect?, Current direc- tions in psychological science11, 159 (2002)
2002
-
[40]
H. M. Schulz, Reference group influence in consumer role rehearsal narratives, Qualitative market research: An in- ternational journal18, 210 (2015)
2015
-
[41]
L´ evy-Garboua and C
L. L´ evy-Garboua and C. Montmarquette, Reported job satisfaction: what does it mean?, The Journal of Socio- Economics33, 135 (2004)
2004
-
[42]
Bygren, Pay reference standards and pay satisfaction: what do workers evaluate their pay against?, Social Sci- ence Research33, 206 (2004)
M. Bygren, Pay reference standards and pay satisfaction: what do workers evaluate their pay against?, Social Sci- ence Research33, 206 (2004)
2004
-
[43]
Our World in Data, Annual growth of GDP per capita,https://ourworldindata.org/grapher/ gdp-per-capita-growth(2024)
2024
-
[44]
Dhiman, Top companies,https://www.kaggle.com/ code/shiivvvaam/top-global-companies/notebook 14 (2024)
S. Dhiman, Top companies,https://www.kaggle.com/ code/shiivvvaam/top-global-companies/notebook 14 (2024)
2024
-
[45]
Elgiriyewithana, Billionaires statis- tics dataset (2023),https://www
N. Elgiriyewithana, Billionaires statis- tics dataset (2023),https://www. kaggle.com/datasets/nelgiriyewithana/ billionaires-statistics-dataset(2023)
2023
-
[46]
Five Thirty Eight, 2020 election forecast, https://projects.fivethirtyeight.com/ 2020-election-forecast/(2020)
2020
-
[47]
M. G. Burgess, R. E. Langendorf, T. Ippolito, and R. Pielke Jr, Optimistically biased economic growth fore- casts and negatively skewed annual variation, (2020)
2020
-
[48]
Scheffer, B
M. Scheffer, B. Van Bavel, I. A. van de Leemput, and E. H. van Nes, Inequality in nature and society, Pro- ceedings of the National Academy of Sciences114, 13154 (2017)
2017
-
[49]
Lavie, U
D. Lavie, U. Stettner, and M. L. Tushman, Explo- ration and exploitation within and across organizations, Academy of Management annals4, 109 (2010)
2010
-
[50]
H. R. Greve, Exploration and exploitation in product in- novation, Industrial and corporate change16, 945 (2007)
2007
-
[51]
Richard, C
G. Richard, C. Guinet, J. Bonnel, N. Gasco, and P. Tix- ier, Do commercial fisheries display optimal foraging? the case of longline fishers in competition with odontocetes, Canadian Journal of Fisheries and Aquatic Sciences75, 964 (2018)
2018
-
[52]
R. H. MacArthur and E. R. Pianka, On optimal use of a patchy environment, The American Naturalist100, 603 (1966)
1966
-
[53]
H. A. Simon, A behavioral model of rational choice, The quarterly journal of economics , 99 (1955)
1955
-
[54]
Layard, Happiness and public policy: A challenge to the profession, The economic journal116, C24 (2006)
R. Layard, Happiness and public policy: A challenge to the profession, The economic journal116, C24 (2006)
2006
-
[55]
Layard,Happiness: Lessons from a new science(Pen- guin UK, 2011)
R. Layard,Happiness: Lessons from a new science(Pen- guin UK, 2011)
2011
-
[56]
K. J. Arrow and P. S. Dasgupta, Conspicuous consump- tion, inconspicuous leisure, The Economic Journal119, F497 (2009)
2009
-
[57]
Surowiecki,The wisdom of crowds(Vintage, 2005)
J. Surowiecki,The wisdom of crowds(Vintage, 2005)
2005
-
[58]
I. L. Janis, Groupthink, Psychology Today , 84 (1971)
1971
-
[59]
Berdahl, C
A. Berdahl, C. J. Torney, C. C. Ioannou, J. J. Faria, and I. D. Couzin, Emergent sensing of complex environments by mobile animal groups, Science339, 574 (2013)
2013
-
[60]
Dussutour, S
A. Dussutour, S. J. Simpson, E. Despland, and N. Cola- surdo, When the group denies individual nutritional wis- dom, Animal Behaviour74, 931 (2007)
2007
-
[61]
Kahneman, J
D. Kahneman, J. L. Knetsch, R. H. Thaler,et al., The en- dowment effect, loss aversion, and status quo bias, Jour- nal of Economic perspectives5, 193 (1991)
1991
-
[62]
C. G. Small,Expansions and asymptotics for statistics (Chapman and Hall/CRC, 2010)
2010
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.