{"id":"be650d14-35a5-4388-a1da-e03f19d4fc1b","arxiv_id":"1908.09081","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"Topological feature vectors (crockers) outperform hand-designed order parameters for classifying dynamics and recovering parameters in simulations of the D'Orsogna collective-motion model.","lead":"The paper tests whether topological summaries of simulated swarm data, called crockers, help machine learning recover the parameters and behavior types of a classic collective-motion model. Crockers beat traditional order parameters in both clustering and supervised classification, suggesting a problem-independent alternative for analyzing group motion.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"PCA is described as reducing features before the 5-fold CV split, so the 99.9% crocker accuracy may be inflated by test-set leakage; a nested-CV rerun is needed to verify the topological advantage.","rationale":"I read the paper as a careful, well-scoped methods comparison: same simulations, same classifiers, and an explicit attempt to align feature dimensionality. The confusion matrices and the controlled pipeline are genuine strengths, and the topological descriptors are clearly informative for collective-motion data. However, the single most load-bearing assumption for the supervised headline claim is not the initial-condition independence but the data-handling protocol around PCA and normalization. The paper does not say whether PCA is fit within cross-validation folds, and the global normalization computed across all simulations raises the same concern. If PCA and normalization see the test set before classification, the 99.9% accuracy could be an artifact. This is a concrete, fixable methodological question, not an accusation of fraud. The reader's weakest_assumption identifies initial-condition independence, which is real but less decisive: even if the model is multistable, the parameter labels are still known from the simulation setup, so the parameter-recovery task is not inherently ill-posed; the main risk would be poor generalization to new initial-condition distributions. I therefore partially agree with the reader. The conditional verdict remains appropriate: the paper should be accepted only if the authors confirm that PCA and normalization are nested inside the CV procedure, or re-run the analysis correctly. My concern does not move the verdict to reject, because the evidence may well survive the rerun, but it does mean the central quantitative claim is not yet fully secured.","tokens_in":14986,"tokens_out":3934,"duration_ms":45230,"concrete_test":"Rerun the supervised parameter-recovery pipeline with PCA (and the global position normalization constant) computed only on the training fold within each of the 5 CV folds, then apply the fitted transformation to the held-out test fold. Compare the out-of-sample accuracy for time-delayed b0 crockers reduced to 87 dimensions against the DNN(t) baseline under the same protocol, and report fold-wise accuracies with standard errors. If the crocker accuracy drops materially (for example, below the 91.1% DNN baseline), the headline claim of topological outperformance is unsupported as stated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that topological crockers outperform order parameters for supervised parameter recovery after matching feature dimension via PCA. The strongest reported number is 99.9% for time-delayed b0 crockers reduced to 87 dimensions, versus 91.1% for DNN(t) at 87 dimensions (Table III, Section III B). The paper states that PCA is used to reduce the dimensionality of input feature vectors, but it does not state that PCA is fit inside each training fold of the 5-fold cross-validation described in Section II E. If PCA is computed on all 2500 simulations before the train/test split, then each held-out simulation contributes to the principal axes used to project its own features. This is a form of information leakage that can inflate out-of-sample accuracy, and it is especially consequential for a 17,200-dimensional crocker reduced to 87 dimensions, where the projection is highly flexible. The comparison is also asymmetric: the DNN(t) baseline is reported at its native 87 dimensions without PCA, so it has no comparable leakage. A second, related leakage path is the global normalization constant in Section II C, computed across all simulations to scale position data; if this constant is estimated on the full dataset, test-set scale information enters training features. Neither issue necessarily invalidates the conclusion, but both are unstated protocol details that directly affect the headline comparison. The reader's initial-condition concern is legitimate but less load-bearing for this claim, because the ground-truth parameter labels are known from the simulation setup even if the dynamics are not independent of initial conditions; the more immediate threat to the reported accuracies is the PCA/CV protocol.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript presents a methodology for recovering model parameters and classifying collective motion phenotypes in the D'Orsogna self-propelled particle model by applying machine learning to two classes of features: traditional order parameters (polarization, angular momentum, absolute angular momentum, and mean nearest-neighbor distance) and topological summaries (persistent homology crockers, including time-delayed versions). The authors simulate 2500 trajectories across 25 parameter combinations, compute features, and apply k-medoids clustering and linear SVMs with 5-fold cross-validation, using PCA to match feature dimensionality. The central claim is that topological crockers outperform order parameters in both unsupervised parameter recovery (76.6% vs 49.9% for concatenated features) and supervised parameter recovery (99.9% vs 91.1% after PCA to 87 dimensions).","tokens_in":15279,"tokens_out":4806,"duration_ms":48325,"significance":"If the claims hold, the paper makes a useful methodological contribution by demonstrating that problem-agnostic topological summaries can outperform hand-crafted order parameters for inverse problems in collective motion, and it provides a careful comparison protocol (dimensionality matching via PCA, multiple phenotypes, and both unsupervised and supervised settings). The paper's strengths are its clean simulation library, the direct comparison on a canonical model, and the explicit use of PCA to equalize feature dimension. However, the headline comparison depends on a leakage-free cross-validation protocol; as written, the PCA and global normalization steps are ambiguous and may inflate the topological accuracy. The empirical claims are also point estimates without uncertainty quantification, and the initial-condition independence is asserted rather than quantified.","major_comments":[{"comment":"The PCA dimensionality reduction appears to be applied before the 5-fold cross-validation split: Section III B says 'we use principal component analysis (PCA) to reduce the dimensionality of our input feature vectors' but does not state that PCA is fit on the training folds only. If the PCA transformation is computed on all 2500 simulations, the held-out test simulations contribute to the principal axes used to project their own features, which is test-set leakage and can inflate the reported 99.9% accuracy for time-delayed b1 crockers in Table III. Because the central claim is that crockers outperform order parameters after matching dimensions, this protocol detail is load-bearing. The authors should either clarify that PCA was fit inside each training fold (for example, via a pipeline or nested cross-validation) or rerun the comparison with a leakage-free protocol; as written, the 87-dimensional crocker results in Table III are not directly comparable to the 87-dimensional DNN(t) results.","section":"Section II E / Section III B"},{"comment":"The global normalization of position data is described as performed 'across all simulations,' meaning the scaling constant is estimated from the full dataset including the test simulations. If this constant is not computed from training data only, the scale information from the test set enters the training features, constituting a second, independent path for information leakage. This could bias the comparison, especially for escape phenotypes where distances are clamped at 10 before normalization. Please state whether the normalization constant and the escape clamping threshold are determined from training data alone or justify that they are fixed physical constants independent of the dataset.","section":"Section II C"},{"comment":"The statement that 'the D'Orsogna model is typically independent of initial conditions' is unquantified. Because all 100 simulations per parameter pair are labeled with the same (C,l) values, and Table I assigns phenotypes deterministically to parameter pairs, any parameter values for which initial conditions affect the asymptotic state would make the labels ambiguous and would overstate both parameter-recovery and phenotype-classification accuracies. Please quantify the sensitivity to initial conditions, for example by reporting the distribution of terminal phenotypes over the 100 realizations per parameter pair, or by excluding or separately analyzing parameter sets with mixed phenotypes.","section":"Section II A"},{"comment":"All reported accuracies are point estimates without error bars or confidence intervals, and the 5-fold cross-validation is described only as a single procedure rather than a repeated or stratified resampling with variance estimates. For the central comparison (99.9% vs 91.1%) the conclusion is likely robust, but smaller differences, such as 97.0% vs 96.2% for b0 crockers before and after PCA, need uncertainty quantification to support the claim that the topological advantage is systematic rather than noise.","section":"Tables II and III"}],"minor_comments":[{"comment":"The sentence reporting 'a 0.8% drop in accuracy for position-only information and a 0.1% increase for time-delayed topological information' is ambiguous because the preceding sentence highlights the b1 time-delayed PCA result at 99.9%, while the cited numbers correspond to the b0 rows (97.0% to 96.2% and 99.6% to 99.7%). Please clarify which feature rows are being compared.","section":"Section III B"},{"comment":"The epsilon grid for persistent homology (200 log-spaced values between 10^-4 and 1), the escape clamping threshold (||xi||∞ = 10), the order-parameter downsampling factor (23), and the time delay (5Δt) are hand-chosen. A brief robustness check or a statement that the qualitative conclusions are insensitive to these choices would strengthen the paper.","section":"Section II C"},{"comment":"In the unsupervised setting, the number of clusters k is set to 25 because there are 25 parameter combinations, which uses label information in an otherwise unsupervised procedure. This is stated, but the paper should also discuss the extent to which this choice helps the k-medoids results and whether the comparison with a 'problem-independent' approach is compromised by this prior knowledge.","section":"Section II D"},{"comment":"No statement is provided about code or data availability. Given the methodological nature of the paper and the explicit reproducibility goals of machine learning studies, a repository with simulation and analysis scripts would be valuable for verifying the protocol, especially the cross-validation and PCA steps.","section":"General"},{"comment":"In the caption of Fig. 8, the phrase 'the 10 × 10 bin in the top left corner represents 100 simulations' is slightly confusing because the bins represent parameter pairs on a 5×5 grid; please check whether the bin dimensions should refer to the simulated realizations within each parameter pair rather than the parameter grid.","section":"Appendix B / Fig. 8"}],"recommendation":"major_revision","confidential_remarks":"The central empirical claim is plausible but the PCA leakage ambiguity in Sections II E and III B is the main risk; if the results survive a rerun with PCA fit inside the cross-validation folds, the paper is a solid contribution to applied TDA and machine learning for dynamical systems. The comparison is an independent test on a model by other authors, and the use of the authors' own crocker tools (Refs. 29 and 30) is appropriately disclosed. I would like to see the authors address the initial-condition sensitivity and provide uncertainty estimates for the accuracies before publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The thing to know: this is the first paper to use crockers as machine-learning features for an inverse problem, and the comparison is set up fairly—same simulations, same classifiers, only the feature representation changes, with PCA used to match dimensions. If the result survives a protocol fix, it is a useful recipe: on the D'Orsogna model, topology-based features beat traditional order parameters in both unsupervised clustering (76.6% vs 49.9% parameter recovery) and supervised SVM classification (99.6% vs 91.1% at native dimensions, or 99.9% vs 91.1% after PCA).\n\nWhat is genuinely new: earlier crocker papers did exploratory analysis and model selection; this is the first use for parameter recovery and phenotype classification. The authors also do not oversell—they note computational cost, the simulation-only demonstration, and the limitation to discrete grid labels.\n\nThe soft spots are real. The stress-test concern about PCA leakage is the most important. Section II.E describes 5-fold CV but does not say that PCA is fit inside each training fold. If PCA is computed on all 2500 simulations before the split, the held-out simulations contribute to the projection axes, and the 99.9% delayed-b1 number is inflated. The comparison is asymmetric: DNN(t) is reported at native 87 dimensions with no PCA, so it has no chance of leakage. The global normalization constant in Section II.C, computed across all simulations, is a second, smaller leakage path. These are unstated protocol details, not necessarily fatal—the topological advantage is large even without PCA—but the headline margin cannot be trusted until a nested-CV rerun is done.\n\nThe reader's initial-condition concern is legitimate but secondary. The task is parameter recovery, and the labels are the parameters used to generate each simulation, so the supervised problem is well-posed even if dynamics are not unique. Non-uniqueness would hurt accuracy rather than invalidate the setup, though the unquantified 'typically' should be checked.\n\nMinor issues: accuracies are point estimates without error bars; the epsilon grid, clamping threshold, and downsampling factor are hand-chosen. All addressable.\n\nWho this is for: TDA folks, collective-motion modelers, and anyone building ML pipelines on trajectory data. It deserves a serious referee; I'd send it to review with a request for nested CV, error bars, and a statement of where the normalization constant comes from.","headline":"First real attempt to turn crockers into machine-learning features for parameter recovery, with a mostly fair comparison—but the headline accuracy numbers need a nested-CV rerun before the topological advantage is trusted.","tokens_in":15864,"tokens_out":2732,"would_cite":true,"duration_ms":27965,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["55N31","37M05","68T05"],"pacs":["02.40.Re","05.45.-a","87.23.Cc"],"model":"deepseek-v4-flash","headline":"Topological shape summaries, not hand-designed order parameters, best recover the parameters behind collective motion.","keywords":["collective motion","topological data analysis","persistent homology","crockers","D'Orsogna model","machine learning","parameter recovery","swarm behavior"],"falsifier":"Re-run the 100-simulation protocol for a parameter pair that produces a collective swarm, such as $(C,\\ell)=(0.1,0.5)$, but draw initial positions and velocities from several different distributions; if the phenotype varies across initial-condition sets, the ground-truth labels are ambiguous and the reported accuracies overstate method performance.","tokens_in":14789,"feed_emoji":"🐟","tokens_out":11258,"duration_ms":104692,"temperature":0.7,"pith_summary":"This paper asks whether the shape of a moving group, summarized with topological data analysis, can reveal more about how that group was generated than the standard statistics of collective-motion research. The authors simulate 2,500 runs of the D'Orsogna model of self-propelled agents and summarize each run in two ways: with four traditional order parameters (polarization, angular momentum, absolute angular momentum, and nearest-neighbor distance) and with crockers, contour plots of the number of connected components and holes in the swarm's point cloud over time and length scale. Both summaries feed an unsupervised clusterer and a supervised classifier that try to recover which of 25 parameter pairs produced a simulation and which of the model's phenotypes it exhibits. The topological inputs win in both settings, and the best supervised result recovers parameters with 99.9% accuracy after the crockers are compressed to the same dimension as the best order parameter, which reaches 91.1%.","feed_headline":"Shape-based summaries decode swarm parameters at 99.9%","feed_subtitle":"Topology-based crockers beat hand-designed statistics at classifying flocking, milling, and escape.","key_machinery":"The central object is the crocker, Contour Realization Of Computed k-dimensional hole Evolution in the Rips complex, a contour plot of the Betti number $b_k$ as a function of time $t$ and persistence scale $\\epsilon$. It is computed by building Vietoris-Rips complexes on agent positions, sometimes embedded as time-delayed 4-dimensional points $(x_i(t_j), x_i(t_{j-5\\Delta t}))$, and tracking how many connected components ($b_0$) and loops ($b_1$) persist across scales. Vectorizing the crocker produces the feature vector that feeds the machine learning algorithms, providing a shape-based summary that needs no knowledge of which patterns are expected.","core_discovery":"The paper's discovery is that crocker feature vectors carry more discriminative information for classification and parameter recovery than order parameters, without being designed around the expected patterns. For unsupervised k-medoids clustering into 25 groups, concatenated $b_0$ and $b_1$ crockers from position data recover parameters with 76.6% accuracy and phenotypes with 99.8% accuracy, compared with 49.9% and 97.4% for concatenated order parameters. For supervised linear support vector machines, the best crocker input, a time-delayed $b_1$ crocker reduced by PCA to 87 dimensions, recovers parameters with 99.9% accuracy, while the best order parameter $D_{NN}(t)$ achieves 91.1% at the same dimension; a 200-fold compression of crockers costs less than one percentage point of accuracy. The authors conclude that topology-based summaries are a problem-independent alternative to hand-designed order parameters for analyzing collective motion.","pith_inferences":["This suggests crocker PCA coordinates could serve as empirical coordinates for a phase diagram of the model, letting transitions between phenotypes be detected automatically instead of by inspecting individual simulations.","A natural stress test is to add measurement noise to positions and watch how classification accuracy degrades; the stability of crockers to such noise would determine whether the method transfers to field observations of real groups.","Running the same comparison on the three-dimensional version of the model, where phenotype transitions differ, would test whether the advantage of crockers over order parameters is robust beyond the two-dimensional setting studied here."],"forward_implications":["Supervised parameter recovery from position-derived trajectory summaries can be nearly perfect, so a trained classifier offers a way to read model parameters off observed trajectories without fitting a differential-equation model by hand.","Crockers survive aggressive dimension reduction: compressing them roughly 200-fold to 87 dimensions costs less than one percentage point of accuracy, whereas the same compression cuts the concatenated order-parameter accuracy by about 20 points.","Because crockers require no prior expectation of the phenotypes, the same feature construction can be carried to other agent-based models or to experimental position data without redesign.","$b_0$ information alone nearly suffices for parameter recovery, and adding $b_1$ mostly lengthens the feature vector, so simple connectivity summaries may be enough for many classification tasks."],"supporting_citations":[{"why":"Supplies the model under study, its parameter ranges, and the phenotypes used as classification targets.","marker":"[9]"},{"why":"Defines crockers and applies topological data analysis to biological aggregation, providing the feature construction used here.","marker":"[29]"},{"why":"Source for the absolute angular momentum order parameter and the state-transition context behind the baseline summaries.","marker":"[31]"},{"why":"Provides the persistent homology framework and computational pipeline the crockers rely on.","marker":"[35]"},{"why":"The software used to compute persistent homology on the simulation point clouds.","marker":"[37]"},{"why":"The linear support vector machine method used for supervised parameter recovery.","marker":"[43]"},{"why":"Principal component analysis used to reduce crockers and order-parameter features to comparable dimensions.","marker":"[44]"}],"fun_headline_variants":["Crocker summaries beat order parameters for swarm behavior","Topology outperforms hand-crafted stats in collective motion","Shape-based features decode flocking and milling better","Crocker vectors classify swarm phases with higher accuracy"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"Every simulation is labeled with the parameter pair that supposedly generated it, which assumes the model's long-term behavior does not depend on initial conditions; the paper says this is 'typically' true but does not quantify the exceptions.","fun_headline_variants_meta":{"raw":{"variants":["Crocker summaries beat order parameters for swarm behavior","Topology outperforms hand-crafted stats in collective motion","Shape-based features decode flocking and milling better","Crocker vectors classify swarm phases with higher accuracy"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000293,"raw_usage":{"total_tokens":1680,"prompt_tokens":888,"completion_tokens":792,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":504,"completion_tokens_details":{"reasoning_tokens":731}},"tokens_in":504,"tokens_out":792,"duration_ms":6676,"temperature":1.0,"reasoning_tokens":731,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T11:21:58.146355+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the 100-simulation protocol for a parameter pair that produces a collective swarm, such as $(C,\\ell)=(0.1,0.5)$, but draw initial positions and velocities from several different distributions; if the phenotype varies across initial-condition sets, the ground-truth labels are ambiguous and the reported accuracies overstate method performance.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the model under study, its parameter ranges, and the phenotypes used as classification targets."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines crockers and applies topological data analysis to biological aggregation, providing the feature construction used here."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Source for the absolute angular momentum order parameter and the state-transition context behind the baseline summaries."},{"cited_title":"Otter , author M","cited_arxiv_id":null,"evidence_quote":"Provides the persistent homology framework and computational pipeline the crockers rely on."},{"cited_title":"Tralie , author N","cited_arxiv_id":null,"evidence_quote":"The software used to compute persistent homology on the simulation point clouds."},{"cited_title":"Jolliffe ,\\ @noop title Principal Component Analysis \\ ( publisher Springer Verlag ,\\ year 1986 ) NoStop","cited_arxiv_id":null,"evidence_quote":"Principal component analysis used to reduce crockers and order-parameter features to comparable dimensions."}],"review_version":1}