{"id":"28afdaa1-8c79-40a8-8cfd-a30b51588ab6","arxiv_id":"1908.03482","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":8,"one_line_summary":"Cohort-level CD62L time-courses, when analyzed with an adapted Bellman-Harris model, reproduce the clonal-data conclusion that memory precursors arise before effectors during the expansion phase, though with acknowledged modeling caveats.","lead":"This paper adapts a statistical model of immune cell fates, originally built for single-cell lineage data, so it can be fitted to ordinary population-level time-courses. Applied to four published datasets, it concludes that, within the model's assumptions, memory T cells are generated before effectors, but it also raises the possibility that memory forms only after the expansion phase.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Memory First conclusion depends on curtailing data to an estimated expansion phase that excludes the late CD62L+ rise that Kinjo et al. interpret as post-expansion memory; the framework cannot see the alternative it is meant to rule out.","rationale":"The reader identifies the load-bearing premise as the assumption that memory precursors are generated within the strictly exponential expansion phase and that the fitted expansion endpoint is correct. My reading agrees, but the concern can be sharpened: the expansion-phase curtailment is not just a technical detail, it is the step that excludes the late CD62L+ rise that Kinjo et al. report on day 7 and interpret as memory appearing after the expansion phase. In the simplified two-path framework, Effector First predicts CD62L+ memory appearing after CD62L- effectors, so the late rise is exactly the signal that would favour Effector First; full-data fitting in Fig. 6A does favour Effector First in the relevant cases. The paper is transparent about this limitation, and its conclusion is explicitly 'within this modeling framework,' so the correct verdict remains CONDITIONAL rather than REJECT. No internal inconsistency is evident in the fitting procedure itself; the issue is model misspecification and endpoint sensitivity. A sensitivity analysis with alternative cutoff days would settle whether the Memory First ranking is robust or an artifact of the chosen boundary.","tokens_in":15113,"tokens_out":5475,"duration_ms":60967,"concrete_test":"Run the sensitivity analysis described above and compare the ranking for cutoffs days 3–7; also report whether the Effector First fit assigns near-zero probability to the memory-producing transition, which would indicate that the model is using the naive CD62L+ population to fit the early decline and is not actually testing memory-timing at all.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim in Section 5 is that, once the day of peak response is taken into account, the Memory First model gives consistent best fits across [2, 21, 15]. This conclusion is produced by a specific curtailment step: data are restricted to an estimated expansion phase (Fig. 6B/C), with the endpoint for [15] inferred from a log-linear regression in Fig. 5A (r2=0.81) as day 4–5. The concern is that this curtailment removes exactly the late-time signal that would distinguish Effector First. In the simplified two-path models (Fig. 3B), Effector First generates CD62L+ memory only after CD62L- effectors, so the late rise in CD62L+ proportion in Kinjo et al. on day 7 is the key evidence for that path. Fitting the full data up to day 8 (Fig. 6A) favoured Effector First in the cases where that late rise is present. Thus the 'consistent deduction' is conditional on the choice of boundary: if the true expansion phase ends later than the Fig. 5A estimate, or if memory is generated after the strictly exponential expansion phase, then the model class cannot represent the biologically relevant alternative. The authors acknowledge this in the abstract and Discussion, saying the alternative 'is a deduction not possible from the mathematical methods provided,' but the headline robustness claim still depends on the curtailment being correct. Because the late memory signal is excluded before model comparison, the comparison is not a test between the two differentiation orders for the question the paper is addressing.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper adapts the multi-type Bellman-Harris branching process model of Buchholz et al. (2013) to infer the order of memory/effector differentiation from cohort-level (population) time-course data rather than clonal data. The authors first validate the adaptation by fitting six representative differentiation networks to the non-clonal data reported in Buchholz et al. and recover the original 'Memory First' conclusion with similar parameter values (Fig. 4). They then reduce the model to two competing structures, Linear Memory First and Linear Effector First, and apply it to data from Badovinac et al. (2007), Schlub et al. (2010), and Kinjo et al. (2015). When all data up to day 8 are fitted, the results are contradictory; however, after curtailing the data to an estimated expansion phase (based on a log-linear regression relating adoptive transfer number to day of peak response) and imposing an upper bound on cell lifetimes, the Memory First model provides the better fit across all spleen/blood datasets (Fig. 6B-C). The paper openly acknowledges that the model cannot represent the possibility that memory precursors are generated after the strictly exponential expansion phase, an alternative supported by data from Kinjo et al.","tokens_in":15569,"tokens_out":2890,"duration_ms":33354,"significance":"If the central claim holds, the paper provides a practical method for evaluating differentiation order from commonly available population-level time-course data, which would substantially increase the utility of the Buchholz et al. modeling framework. The internal consistency check on the Buchholz cohort data is a genuine positive result: fitting six models to cohort proportions yields the same best model as fitting to clonal summary statistics, with five of six parameters being quantitatively similar (Fig. 4B). The paper is also commendably transparent about its limitations, explicitly flagging the post-expansion memory alternative and the model's inability to represent it. However, the cross-dataset conclusion depends on several external estimates (peak day, average family size, expansion-phase endpoint) that are not derived from the data being fit, and no sensitivity analysis is provided to show that the Memory First conclusion is robust to these estimates. Because the load-bearing assumptions are acknowledged but not stress-tested, the significance of the cross-dataset claim is currently conditional.","major_comments":[{"comment":"The central cross-dataset claim that 'Memory First provides the best fit' is obtained only after curtailing each dataset to an estimated expansion phase, and the curtailment endpoint is the single most consequential modeling choice. For the Kinjo et al. data, the endpoint (day 4-5) is inferred from a log-linear regression (Fig. 5A, r2=0.81) that extrapolates to 10^6 transferred cells. The manuscript itself notes in Section 4 that fitting the full day 0-8 data supports Effector First in several cases (Fig. 6A), and that Kinjo et al. report a late rise in CD62L+ proportion by day 7. Thus the comparison between Memory First and Effector First is conditional on excluding exactly the late-time data that would distinguish the two paths. A sensitivity analysis is needed: the regression has uncertainty, and the authors should show how the model ranking changes as the expansion-phase endpoint varies over a plausible range (e.g., day 4 to day 7 for the Kinjo data). Without such an analysis, the 'consistent deduction' claim in Section 5 is not yet supportable.","section":"Section 4, Fig. 6B/C; Section 5"},{"comment":"The estimate of average family size at peak for the Kinjo et al. experiment (23 cells per family) is derived from a log-linear regression of peak OT-1 proportion on transfer number plus several auxiliary assumptions (total lymphocyte count at peak is constant across experiments; the Buchholz average family size is representative). The authors state that a 'sanity check established there was no significant impact (data not shown)', but no details are given. This quantity is not merely a nuisance parameter: it sets the scale of the entire population trajectory and directly enters the objective function d3 through the squared relative error term (y1 - y1(theta))^2 / y1^2. A change in the family size estimate could plausibly change which model better matches the observed proportions, especially for small family sizes. The authors should report the sanity check and/or demonstrate robustness of the model ranking to a range of family-size estimates.","section":"Section 4, Fig. 5B/C; Eq. (d3)"},{"comment":"The objective function is changed between the Buchholz cohort fit (d2, a chi-square-like sum using estimated standard errors) and the other datasets (d3, a weighted mean squared error that scales only the family-size term by its observed value squared). The choice of d3 is justified as a way to balance scales, but no justification is given for why only the family-size term is normalized and why the proportional terms enter with unit weight. This weighting is not derived from a noise model, and it could affect the relative ranking of Memory First versus Effector First. The authors should either provide a principled derivation of d3 (e.g., as a likelihood under a stated error model) or show that the qualitative conclusions are unchanged under alternative reasonable weightings (e.g., normalizing every term by the observed value squared, or using relative errors for all terms).","section":"Section 3, objective function d2; Section 4, objective function d3"},{"comment":"The paper excludes the MLN data from Kinjo et al. because of 'high levels of migration', yet that dataset, when fitted with the same constrained procedure, returns Effector First as the better model (Section 4, Fig. 6D). Since the stated goal is to test the generality of the Memory First deduction across data from different papers and organs, the exclusion criterion should be made more specific and objective. The authors should either give a quantitative justification for excluding MLN data (e.g., a migration model or a goodness-of-fit threshold) or report the MLN result transparently as a limitation rather than as an exception that is set aside. As written, the exclusion risks making the claim of cross-dataset consistency circular.","section":"Section 4, Fig. 6D and MLN analysis"}],"minor_comments":[{"comment":"The abstract states that the memory-precursors-after-expansion possibility is 'a deduction not possible from the mathematical methods provided'; this is a welcome and important caveat, but the wording in the main text (Section 4) is stronger: the authors write that the method 'indicates' differences that 'could come from other [15] results'. For clarity, use consistent language throughout: clearly distinguish between 'the model prefers Memory First given its assumption of expansion-phase-only memory generation' and 'memory generation after the expansion phase is occurring in the data'.","section":"Abstract and Section 4"},{"comment":"The claim of 'remarkably similar' parameterization between clonal and cohort fits is based on five of six parameters having similar values, but the figure shows the naive-cell lifetime lambda_N differs substantially (0.34 vs 1.24 days). The text gives a plausible explanation (naive and TCMp are conflated), but this should be quantified: report the fitted value and its uncertainty for lambda_N in both cases, and state whether the difference is within a plausible range given the conflation.","section":"Section 3, Fig. 4B"},{"comment":"In Figure 2B, the top row shows 'all data up to day 8' while the bottom row shows 'data curtailed to expansion phase' for the same datasets. The curtailment endpoints are not marked on the figure, which makes it difficult for the reader to see exactly which time points are included in each fit. Provide vertical dashed lines or numeric labels indicating the last day retained for each dataset.","section":"Section 2, Fig. 2B"},{"comment":"The notation for the objective function d3 is imprecise: the subscript ranges in the sum are written as i=2 to k+1, but the text above defines y1 as the family size and y2,...,y_{k+1} as proportions. This is understandable, but the definition of k is not stated. Please define k explicitly as the number of time points at which proportions are observed.","section":"Section 4, Eq. (d3)"},{"comment":"The paper manually extracts data from graphs in the cited papers. To ensure reproducibility, the authors should provide the extracted numerical data as a supplementary table or a publicly accessible repository, along with the fitting code or a clear description of the optimization procedure (e.g., which solver, number of restarts, convergence criteria). This is particularly important because the paper repeatedly notes that data are limited and parameters are at boundary values.","section":"References and data availability"}],"recommendation":"major_revision","confidential_remarks":"The paper is honest about its limitations, which is a plus, but the honest statements also reveal that the headline cross-dataset conclusion rests on procedurally fragile estimates (peak-day regression, family-size regression) and on excluding the late-time signal that the competing model would need. The authors should be pushed to provide the sensitivity analyses they mention only in passing. The internal consistency result on the Buchholz cohort data (Fig. 4) is solid and could be the strongest part of the paper if the cross-dataset claims are appropriately softened or made conditional on the expansion-phase assumption."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know: this is not a rehash of Buchholz et al. The authors build two new objective functions (d2, d3), a transfer-number-based pipeline for estimating peak day and family size, and run the first systematic cohort-level test of the 304-model framework across four published datasets. The cleanest result is Fig. 4: fitting cohort data from Buchholz's own paper recovers the same best model and nearly the same parameters as the clonal fit. That is a real consistency check, and it passes. Credit is also due for candor—questionable parameterizations, graph-extracted data, log-linear assumptions, and the model's blind spot for post-expansion memory generation are all flagged in the abstract and discussion.\n\nThe soft spot is the load-bearing claim in Section 5. The 'once day of peak was taken into account, Memory First provides the best fit' conclusion depends on curtailing each dataset to an estimated expansion phase whose endpoint comes from a log-linear regression (r2=0.81) in Fig. 5A. For Kinjo et al. that endpoint is day 4–5, and the curtailment removes the day-7 rise in CD62L+ that is the main empirical signature of late memory generation. Fitting all data up to day 8 (Fig. 6A) actually favored Effector First in several cases. So the consistent Memory First verdict is partly a property of the chosen boundary, not a robust test of the biological timing question. The authors acknowledge this in prose, but the headline claim still wears the curtailment. Also notable: the 'biologically implausible' fits often hit the imposed five-day lifetime bound, so the model is straining in several datasets, and the MLN data from Kinjo goes the other way, dismissed rather quickly.\n\nWho gets value: immunologists and quantitative biologists who want to test differentiation-order hypotheses from routine population time-courses, and modelers building on the Buchholz framework. The method is a stepping stone, not a settled tool.\n\nI would send this to a serious referee. The methodological adaptation deserves publication even if the cross-dataset biological conclusion is weaker than the abstract implies. A good referee should push for a clearer statement that the curtailment boundary is part of the hypothesis being tested, not a preprocessing detail.","headline":"A genuinely useful adaptation of Buchholz's branching-process framework to cohort data, with an honest internal consistency check, but the multi-dataset Memory First conclusion rests on a curtailment choice that excludes the late-time signal most relevant to the biological question.","tokens_in":16031,"tokens_out":1768,"would_cite":true,"duration_ms":18401,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["60J80","92C37"],"pacs":[],"model":"deepseek-v4-flash","headline":"Population-level time courses can identify when memory T cells are made.","keywords":["adaptive immune response","CD8+ T cell differentiation","memory cell generation","multi-type Bellman-Harris process","cohort data","clonal data","CD62L phenotype","expansion phase"],"falsifier":"Use a cell-cycle reporter, as in the real-time tracking experiments cited here, to record the first day on which cells with a memory phenotype appear among non-cycling cells after the expansion peak; if memory-phenotype cells first emerge only after the exponential expansion has ended, the framework's Memory First deduction cannot be right.","tokens_in":14913,"feed_emoji":"🧬","tokens_out":6396,"duration_ms":63361,"temperature":0.7,"pith_summary":"This paper asks a practical question: can the timing of memory-cell generation in an immune response be inferred from ordinary population-level time-course data, rather than from the laborious single-cell 'clonal' tracking used in the 2013 study that introduced the question? The authors adapt that study's stochastic branching-process model so it can be fitted to cohort proportions of CD62L+ cells, first checking the adaptation on the original study's own non-clonal data and then applying it to time-course data from three other studies. Once the day of peak response is accounted for, the adapted model consistently ranks the Linear Memory First pathway ahead of the Linear Effector First pathway across those datasets. The paper also reports two important caveats: fitting data that extend beyond the expansion phase can flip the ranking toward Effector First, and the real-time cell-cycle data used in one study suggest memory may be generated after the expansion phase, a possibility the framework cannot represent.","feed_headline":"Memory-first T-cell order survives population-level data","feed_subtitle":"Adapting a 2013 clonal-tracking model to cohort time courses still picks memory before effectors across four datasets.","key_machinery":"The central object is a multi-type Bellman-Harris branching process: each cell has an exponentially distributed lifetime, then either divides into two cells of its own type or differentiates into another type, with probabilities fixed by a small parameter set. The original version tracked clonal family sizes, variances and correlations at one time point; the adaptation replaces those statistics with cohort-level proportions over time plus an average family size at peak, using a weighted mean-squared-error objective that keeps the exponentially growing population statistic from dominating the proportions. The second load-bearing piece is the log-linear regression linking the number of adoptively transferred cells to the day of peak response, which determines how much of each time-course belongs to the expansion phase.","core_discovery":"The central claim is that the modeling framework introduced in the 2013 clonal study is robust enough to answer the same biological question when fitted to far more widely available cohort data. Concretely, after replacing the clonal summary statistics with a weighted objective function based on population proportions plus average family size, and after curtailing each dataset to the estimated expansion phase, the Linear Memory First model — memory-precursor cells appear and proliferate before effector cells — is the best fit for all datasets examined, including the original study's own cohort data. The paper is explicit that this conclusion is conditional: it holds within a model that only allows differentiation during strictly exponential expansion with no cell death, and it requires an independent estimate of the day of peak response, which the authors derive by log-linear regression on adoptive-transfer number. Without that peak-day adjustment, the same method gives contradictory rankings, sometimes favouring Effector First.","pith_inferences":["If memory generation after the expansion phase is confirmed, the Memory First ranking should be reinterpreted as a constraint of the model, not a fact about biology; fitting a model that allows post-expansion memory formation is a direct test.","The log-linear peak-day estimator is a testable prediction: time-course experiments at intermediate transfer numbers should confirm or reject the inferred peak days before the method is trusted across datasets.","The same cohort-fitting adaptation could in principle be transferred to B-cell responses or other differentiation questions whenever a two-phenotype time course and a peak-day estimate are available."],"forward_implications":["Differentiation-order questions can be screened quickly from cohort time courses alone, without clonal barcoding.","The original clonal-based conclusion is reinforced: within the expansion-phase assumption, memory precursors precede effector cells.","Precursor transfer number must be controlled: high-transfer experiments shorten the expansion phase, and fitting beyond it reverses the model ranking.","The framework's parameter estimates are not reliable in detail; most best fits sit on boundary values and imply implausible naive-cell lifetimes.","Data from lymph nodes may not be analysable by this method because migration breaks the model's assumptions."],"supporting_citations":[{"why":"Supplies the branching-process model class, the clonal data, and the original Memory First deduction that this paper adapts to cohort data.","marker":"[3]"},{"why":"Supplies cohort time-course data at multiple precursor transfer numbers and the peak-day/expansion scaling relationship.","marker":"[2]"},{"why":"Supplies additional cohort data and the method for estimating average divisions and family size at peak from transfer number.","marker":"[21]"},{"why":"Supplies the high-transfer cohort dataset and the cell-cycle reporter evidence that memory may appear after the expansion phase.","marker":"[15]"},{"why":"Provides the mathematical theory of multi-type Bellman-Harris branching processes on which the model is built.","marker":"[9]"},{"why":"The companion clonal-barcoding study whose single-cell data motivate the paper's discussion of familial heterogeneity and model caveats.","marker":"[8]"}],"fun_headline_variants":["Memory-first T-cell order survives broader data, with peak-day caveat","Adapted model still backs memory-first, but peak timing is essential","Cohort data match memory-first, only after peak-day calibration","Memory-first inference robust across datasets when peak-day adjusted","Differentiation order: memory-first holds, conditional on peak-day fit"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The argument collapses if memory precursors are created after the strictly exponential expansion phase has ended, or if the fitted day of peak response derived from transfer number is wrong; the paper itself flags both as live concerns.","fun_headline_variants_meta":{"raw":{"variants":["Memory-first T-cell order survives broader data, with peak-day caveat","Adapted model still backs memory-first, but peak timing is essential","Cohort data match memory-first, only after peak-day calibration","Memory-first inference robust across datasets when peak-day adjusted","Differentiation order: memory-first holds, conditional on peak-day fit"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000241,"raw_usage":{"total_tokens":1597,"prompt_tokens":1093,"completion_tokens":504,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":709,"completion_tokens_details":{"reasoning_tokens":415}},"tokens_in":709,"tokens_out":504,"duration_ms":5565,"temperature":1.0,"reasoning_tokens":415,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T14:12:37.921324+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Use a cell-cycle reporter, as in the real-time tracking experiments cited here, to record the first day on which cells with a memory phenotype appear among non-cycling cells after the expansion peak; if memory-phenotype cells first emerge only after the exponential expansion has ended, the framework's Memory First deduction cannot be right.","supporting_citations":[{"cited_title":"Buchholz, Michael Flossdorf, Inge Hensel, Lorenz Kretschmer, Bianca Weissbrich, Patricia Gr¨af, Admar Verschoor, Matthias Schiemann, Thomas H¨ofer, and Dirk H","cited_arxiv_id":null,"evidence_quote":"Supplies the branching-process model class, the clonal data, and the original Memory First deduction that this paper adapts to cohort data."},{"cited_title":"Badovinac, Jodie S","cited_arxiv_id":null,"evidence_quote":"Supplies cohort time-course data at multiple precursor transfer numbers and the peak-day/expansion scaling relationship."},{"cited_title":"Schlub, Vladimir P","cited_arxiv_id":null,"evidence_quote":"Supplies additional cohort data and the method for estimating average divisions and family size at peak from transfer number."},{"cited_title":"Wellard, Paulus Mrass, William Ritchie, Atsushi Doi, Lois L","cited_arxiv_id":null,"evidence_quote":"Supplies the high-transfer cohort dataset and the cell-cycle reporter evidence that memory may appear after the expansion phase."},{"cited_title":"The Theory of Branching Processes","cited_arxiv_id":null,"evidence_quote":"Provides the mathematical theory of multi-type Bellman-Harris branching processes on which the model is built."},{"cited_title":"Rohr, Le ¨ıla Peri´e, Nienke van Rooij, Jeroen W","cited_arxiv_id":null,"evidence_quote":"The companion clonal-barcoding study whose single-cell data motivate the paper's discussion of familial heterogeneity and model caveats."}],"review_version":1}