{"id":"e6ea1ec3-7828-4e93-82e2-4f892eab1a2a","arxiv_id":"2507.02327","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":4,"one_line_summary":"In a mean-field productivity model, dedicated teachers are optimal only above a critical population size, peak at intermediate sizes, and never exceed half the population.","lead":"What if a group should sometimes keep some people out of production entirely, just to teach? A new model finds that dedicated teaching only pays once a population passes a minimum size, the optimal teacher share peaks at intermediate sizes, and it never exceeds half the population.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The Discussion's claim that 'nobody should be dedicated to teaching when the group size is very small' is not a robust consequence of Eq. (7): nc can be below 1, so a two-individual group can optimally have one teacher.","rationale":"The reader's weakest assumption was equal sharing of rewards, which the authors explicitly acknowledge as a limitation; it is a structural boundary of the model, not a hidden error. The more immediate problem is that a headline qualitative consequence—small groups should have zero dedicated teachers—is false for a non-empty parameter region of the paper's own model. This matters because it changes how the central claim can be advertised: the existence of a positive critical population size is not the same as a meaningful group-size threshold. The counterexample uses only positive parameters and the paper's own formulas, so it is internal. It also coheres with Section F's intuition that in the expertise-limited regime the optimal teacher fraction approaches 1/2 at intermediate n; at n=2, 1/2 is the only positive integer fraction. Therefore the paper should be accepted only conditional on rewording the abstract and Discussion, and possibly adding the constraint under which 'small groups' means n<nc. The formal equations, the 50% rule, and the existence of a threshold remain intact.","tokens_in":22316,"tokens_out":14707,"duration_ms":178954,"concrete_test":"Directly evaluate the model for h0=1, h1=10, f=1, μ=10, λ=0.001 at n=2: compute Eq. (4) at τ=0 and τ=1/2, or run the discrete Markov chain of Section C with nT=0 versus nT=1. If F(τ=1/2)>F(0), the unqualified 'small groups' claim in Section IV is falsified. For a general check, scan integer n≥2 and positive (η,ν,r) for cases with n>nc and τ+>0; the existence of one such case suffices to require rewording.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The formal threshold result (Eqs. 6-7) is derived correctly, but the paper over-translates it into a universal small-group prediction. Section IV states: 'nobody should be dedicated to teaching when the group size is very small, even when the skill cannot be self-learned.' This does not follow, because Eq. (7) only guarantees nc>0; it imposes no lower bound above 1. Since all parameters in Table I are merely positive, ν(r-1) can be large enough that nc<2. Concretely, with h0=1, h1=10, f=1, μ=10, λ=0.001 (so η=0.001, ν=10, r=10), Eq. (7) gives nc≈0.011. For n=2, Eq. (6) gives τ+≈0.167, and the discrete-state model of Section C, which allows only integer teacher counts, yields nT=1 (τ=1/2) as optimal: F(τ=1/2)≈4.59 versus F(0)≈1.01. So a 'very small' group of two should have 50% teachers, not zero. The problem is not the algebra but the interpretation: the existence proof in Section V.B establishes existence of a positive real nc, not that this threshold is meaningful for integer population sizes. The paper's own expertise-limited approximation in Section F even predicts τ≈1/2 when η≪ντn≪1, which is exactly the regime producing this counterexample. The abstract and Discussion should state that teaching is optimal only for n>nc, without asserting that nc is large or that small groups in general have no teachers.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper develops a mean-field ODE model of a fixed-size population divided into amateur hunters, expert hunters, and dedicated teachers, and asks what teacher fraction maximizes equilibrium per-capita productivity. The authors derive a critical population size nc below which the continuous optimum teacher fraction is zero (Eq. 7), show that the optimum rises and then falls with population size, and prove a uniform upper bound of 1/2 (Eq. 9) plus an asymptotic 25% rule near r = 1. They then extend the model to part-time teaching, finite teaching capacity, and three levels of expertise, and compare the results with empirical frequencies of teaching roles. The main advertised conclusions are that dedicated teaching requires a minimum group size, peaks at intermediate population sizes, never exceeds half the population, and that the three-stage model exhibits a two-tier transition in teacher allocation.","tokens_in":22713,"tokens_out":6449,"duration_ms":74647,"significance":"If the results are stated precisely, the paper provides a clean, analytically tractable optimization model of a teacher class, with closed-form expressions for the threshold, the optimal teacher fraction, and its upper bounds. The derivations in Eqs. (3)-(9) are transparent and self-contained, and the inequalities hold for all parameter values rather than being fitted to data. The model also generates falsifiable comparative statics, such as nc decreasing in education efficiency ν and in the expertise advantage r, which are useful for future empirical work. The main limitations are interpretive: the real-valued threshold nc does not by itself guarantee that small integer-sized groups should have no teachers, and the three-stage conclusions rely on a continuity assumption that the authors acknowledge may fail. These issues are fixable but currently affect the paper's headline claims.","major_comments":[{"comment":"The Discussion claim that \"nobody should be dedicated to teaching when the group size is very small, even when the skill cannot be self-learned\" is not a consequence of Eq. (7). Since all parameters are only constrained to be positive, nc can be arbitrarily small: for h0=1, h1=10, f=1, μ=10, λ=0.001, Eq. (7) gives nc≈0.011, so a group of n=2 is already above the threshold. For this example the continuous model gives τ+≈0.167, and the discrete model of Section C gives nT=1 (τ=1/2) as optimal because F(1/2)≈4.59 versus F(0)≈1.01. The correct statement is that teaching is optimal only for n>nc, without asserting that nc is large or that small groups in general lack teachers. This wording should be corrected in the abstract and Discussion, since it is the paper's central interpretive claim.","section":"Section IV and Eq. (7)"},{"comment":"The proof of the existence of nc contains a swapped case analysis. The text states that when ∂τF(0;n)>0, \"F decreases monotonically from F(0;n)>0 to F(1;n)=0, implying the absence of a peak,\" and that when ∂τF(0;n)<0, the intermediate value theorem gives a peak. The opposite is true: a positive derivative at τ=0, combined with ∂τF(1;n)<0 and concavity (∂²τF<0), implies a unique interior maximum, whereas a negative derivative at τ=0 implies monotone decrease. The same reversal appears in Section III.D. The corrected argument is straightforward and the formulas are otherwise correct, but this is the main formal proof of the threshold result and must be fixed.","section":"Section V.B and Section III.D"},{"comment":"The two-tier transition in the three-stage model is derived under the ad hoc assumption that the optimal τ emerges continuously from 0 and the optimal α emerges continuously from 1. The paper itself notes that simulations suggest discontinuous transitions \"in some parameter regime\" and therefore that the method does not cover those cases. As written, however, the main text presents the two critical population sizes n*c and n*α and the sequential transition as general findings. The claims should either be explicitly restricted to the parameter regime where the continuity assumption holds, or a proof should be supplied; otherwise the size-complexity conclusion in the Discussion is stronger than the analysis supports.","section":"Section III.I and Supporting Information III"},{"comment":"The caption of Table III states that all six extended models exhibit \"a minimum population size required before teacher allocation is beneficial,\" and Section III.H repeats this for the first four models. This is contradicted by the finite-capacity model in the same section: when K<nc, the model gives τopt=0 for all n, so no positive threshold exists. The text should distinguish the K>nc case from the K<nc case and should not assert the threshold phenomenon for all six models.","section":"Table III and Section III.H"}],"minor_comments":[{"comment":"The sentence claiming that the numbers in the set {τopt(n)n} are not integers is an overstatement: τopt(n)n may be integer for isolated parameter values or population sizes. What is true is that the continuous optimum need not coincide with an integer teacher count, which is the point of the discrete model.","section":"Section III.C"},{"comment":"The phrase \"In the first regime\" is confusing because two regimes are defined immediately before. Rewording to \"If ∂τF(0;n)<0\" and \"If ∂τF(0;n)>0\" would remove the ambiguity and also fix the case reversal noted above.","section":"Section III.D"},{"comment":"The conjectured upper bound M/(M-1) for the M-stage model should be flagged as a conjecture in the Results section as well as in the text where it is introduced, so that it is not mistaken for a proved theorem.","section":"Section III.I"}],"recommendation":"major_revision","confidential_remarks":"The central optimization results are sound and the paper is within scope for a mathematical biology journal. The revision should focus on the small-group interpretation of nc, the swapped case analysis in the proof, and the qualified status of the three-stage results. If these are corrected, the paper would be suitable for publication. I would also suggest that the authors consider stating explicitly in the abstract that the threshold applies to the real-valued optimization problem and may lie below the smallest integer group sizes of interest."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The core math is right. The threshold, the non-monotonic peak, the 50% bound, and the 25% rule all check out, and they are genuinely new as an application to dedicated teaching. Tannenbaum's model is structurally different, so the novelty claim is fair. The paper is also refreshingly honest about its main limitation: the equal-sharing, non-evolutionary assumption is stated up front and revisited in the Discussion. I give credit for that.\n\nThe soft spots are real but contained. The stress-test concern lands. Eq. (7) proves nc > 0, not nc > 1. With h0=1, h1=10, f=1, mu=10, lambda=0.001, you get nc about 0.011, and a two-person group optimally has one teacher (F jumps from roughly 1.0 to 4.6). So the Discussion's claim that 'nobody should be dedicated to teaching when the group size is very small' is an over-translation. The abstract's 'must exceed a critical size' is fine if read as n > nc, but the Discussion pushes it further than the math supports. Fixing that is a minor rewrite, not a structural change.\n\nMethods Section B has a swapped case analysis. It says a positive derivative at tau=0 implies no peak; it should say that implies a peak, given concavity and the negative derivative at tau=1. The conclusion is still correct, but the written proof is wrong as printed. That needs correction.\n\nThe three-stage results are heuristic. The paper itself calls the approach ad hoc and conjectures the 2/3 bound. The continuity assumption for continuous emergence is not proved, and the authors acknowledge a discontinuous regime may exist. This is fine as an extension, but the referee should push for a clearer statement of what is proved versus conjectured.\n\nNo code or data are shipped, but the model is simple enough to reimplement in an afternoon, and the Monte Carlo step is described sufficiently. Reproducibility is a non-issue.\n\nOverall: this is a clean, useful theory paper for people working on cultural evolution, division of labor, and the size-complexity hypothesis. It deserves a serious referee. I would accept it for peer review, with the expectation of minor-to-moderate revisions: fix the small-group overstatement, correct the swapped cases in Methods, and tighten the three-stage conjectures.","headline":"The equations hold up; the small-group gloss does not—nc can be below 1, and the Methods proof swaps two cases.","tokens_in":23197,"tokens_out":2698,"would_cite":true,"duration_ms":32549,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["92D15"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that a population must exceed a critical size before dedicating anyone to teaching raises per-capita productivity, and that the optimal teaching share peaks at intermediate group sizes and never exceeds one half.","keywords":["dedicated teaching","division of labor","population threshold","optimal teacher allocation","per-capita productivity","cultural transmission","group size","size-complexity hypothesis"],"falsifier":"Measure per-capita productivity in real or simulated groups of several sizes below and above the predicted $n_c$, with teacher fraction varied in controlled steps. The model requires productivity to be maximized at zero teachers whenever $n<n_c$ and at a unique positive fraction when $n>n_c$; a single group below $n_c$ whose measured productivity peaks at a positive teacher fraction would falsify the threshold claim.","tokens_in":22128,"feed_emoji":"🎓","tokens_out":9473,"duration_ms":112410,"temperature":0.7,"pith_summary":"Why a group would set aside people purely to teach others is the puzzle this paper addresses. The authors model a well-mixed population that splits into amateur hunters, expert hunters, and dedicated teachers who never hunt, with all hunting gains shared equally. Their central claim is that the per-capita productivity-maximizing teacher share is zero below a critical group size $n_c=(1+r\\eta)(1+\\eta)/(\\nu(r-1))$, rises to a peak at an intermediate population size, and can never exceed one half. If the model is right, small groups should never create a dedicated teaching class, and no productivity-optimizing society should ever put a majority of its members into teaching, while more complex multi-level skills generate successive teaching tiers as the group grows.","feed_headline":"No dedicated teachers below a critical group size","feed_subtitle":"A new model finds the optimal teaching share peaks at intermediate group size and never exceeds half the population.","key_machinery":"The engine of the argument is the productivity function $F(\\tau;n)=(1-\\tau)W(\\tau;n)$, the product of the workforce fraction and a weighted-average hunting rate $W$ that rises with teacher number. $W$ is a concave function of $\\tau$, so $F$ is concave in the teacher fraction; its marginal value at zero teachers, $\\partial_\\tau F(0;n)$, increases linearly with $n$. This structure forces two regimes: for small $n$ productivity decreases monotonically with $\\tau$, and for large $n$ the intermediate-value theorem plus concavity give a unique interior optimum, producing closed-form expressions for $n_c$, $n^*$, and the bound $\\tau_{\\rm opt}^*<1/2$.","core_discovery":"The paper's central discovery is a population threshold for dedicated teaching. At equilibrium of the mass-action dynamics, per-capita productivity is $F(\\tau;n)=\\frac{(1-\\tau)(h_0+h_1(\\eta+\\nu\\tau n))}{1+\\eta+\\nu\\tau n}$; optimizing over $\\tau$ gives $\\tau_{\\rm opt}=0$ for $n<n_c$ and $\\tau_{\\rm opt}=\\tau_+$ for $n>n_c$, where $n_c=\\frac{(1+r\\eta)(1+\\eta)}{\\nu(r-1)}$. A critical population size of this type appears across all the model variants considered; in the finite-capacity variant teaching is worthwhile only when the half-saturation size $K$ exceeds $n_c$, and the optimal teacher share then saturates at a positive value instead of decaying to zero. The paper also proves that the maximal optimal teacher fraction is always below $1/2$, scales as $\\frac{1}{4(1+\\eta)}(r-1)$ when the advantage of expertise is small, and shows that in a three-level model two sequential transitions occur: teachers appear at one critical size and begin instructing intermediate students only above a second, larger critical size.","pith_inferences":["If the group-level sharing assumption is replaced by individual- or kin-level selection, the threshold would become a function of relatedness and of who captures the returns to skill; small kin groups might then evolve teaching even below $n_c$.","The 50% cap is a sharp diagnostic: a real society or organization with more than half its members in non-producing instructional roles cannot be explained by this productivity-sharing mechanism alone and needs a different incentive story.","The sequential thresholds in the three-stage model suggest a testable account of institutional growth: beginner-level instruction should appear before advanced-level instruction as a school, firm, or laboratory grows; data on training budgets by organization size could test this ordering.","The finite-capacity model predicts that measured per-teacher contact rates, not just the teacher ratio, should determine whether teaching pays; comparing groups of different sizes with the same apparent teacher share would discriminate between bounded and unbounded teaching capacity."],"forward_implications":["A group with population size below $n_c$ maximizes per-capita productivity with zero dedicated teachers, even when expertise cannot be acquired by self-learning.","The optimal teacher share peaks at an intermediate population size and then declines roughly as $1/n$ for large populations in the baseline model.","No choice of parameters in the baseline model makes it optimal for more than half the population to teach; the uniform upper bound is one half.","With finite teaching capacity, a dedicated teacher class pays only when the capacity half-saturation size $K$ exceeds $n_c$, and the optimal teacher share then tends to a positive constant rather than to zero.","With three levels of expertise, teaching turns on at a first critical size and starts targeting intermediate students only above a second, larger critical size."],"supporting_citations":[{"why":"Supplies the stochastic-process formalism used to construct the discrete Markov chain and validate the continuous optimal profile.","marker":"[18]"},{"why":"Provides the prior division-of-labor threshold result with which the paper's critical population size is compared.","marker":"[21]"},{"why":"Supplies the size-complexity hypothesis that the multi-stage teacher allocation is said to agree with.","marker":"[37]"},{"why":"Defines vertical and oblique teaching and anchors the discussion of when the shared-incentive assumption can fail.","marker":"[46]"},{"why":"Provides the U.S. workforce statistic used to compare predicted teacher fractions with empirical ones.","marker":"[29]"},{"why":"Frames the grandmother hypothesis as empirical motivation that dedicated teaching occurs in natural populations.","marker":"[5]"}],"fun_headline_variants":["Teaching only pays above a critical group size","Optimal teacher share never exceeds half the population","Dedicated teachers emerge past a population threshold","When is a teacher worth it? Above a critical group size","The math of teaching: a population size minimum required"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the group maximizes per-capita productivity and that everyone shares equally in the gains from acquired skill; if individual incentives or kinship decide who teaches, the threshold and the fifty-percent cap are not necessarily what evolves.","fun_headline_variants_meta":{"raw":{"variants":["Teaching only pays above a critical group size","Optimal teacher share never exceeds half the population","Dedicated teachers emerge past a population threshold","When is a teacher worth it? Above a critical group size","The math of teaching: a population size minimum required"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00021,"raw_usage":{"total_tokens":1446,"prompt_tokens":1018,"completion_tokens":428,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":634,"completion_tokens_details":{"reasoning_tokens":354}},"tokens_in":634,"tokens_out":428,"duration_ms":4535,"temperature":1.0,"reasoning_tokens":354,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T20:33:02.019101+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Measure per-capita productivity in real or simulated groups of several sizes below and above the predicted $n_c$, with teacher fraction varied in controlled steps. The model requires productivity to be maximized at zero teachers whenever $n<n_c$ and at a unique positive fraction when $n>n_c$; a single group below $n_c$ whose measured productivity peaks at a positive teacher fraction would falsify the threshold claim.","supporting_citations":[{"cited_title":"Hawkes, Nature 428, 128 (2004)","cited_arxiv_id":null,"evidence_quote":"Supplies the stochastic-process formalism used to construct the discrete Markov chain and validate the continuous optimal profile."},{"cited_title":"Lahdenper¨ a, A","cited_arxiv_id":null,"evidence_quote":"Provides the prior division-of-labor threshold result with which the paper's critical population size is compared."},{"cited_title":"Shemesh, J","cited_arxiv_id":null,"evidence_quote":"Supplies the size-complexity hypothesis that the multi-stage teacher allocation is said to agree with."},{"cited_title":"Ellis, D","cited_arxiv_id":null,"evidence_quote":"Defines vertical and oblique teaching and anchors the discussion of when the shared-incentive assumption can fail."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the U.S. workforce statistic used to compare predicted teacher fractions with empirical ones."},{"cited_title":"As this expression shows, there is a critical population size nc required before it is beneﬁcial to dedicate anyone to teaching, given by nc = (1 + rη)(1 + η) ν(r 2 1)","cited_arxiv_id":null,"evidence_quote":"Frames the grandmother hypothesis as empirical motivation that dedicated teaching occurs in natural populations."}],"review_version":1}