{"id":"6cce5e4f-0990-4714-8b75-c9c1f1e8dd81","arxiv_id":"2501.00714","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A two-phase sparse-regression method recovers interaction kernels of stochastic many-particle systems by fitting the residual of the mean-field PDE to kernel-density-estimated densities.","lead":"Scientists watch thousands of particles move and want to reverse-engineer the invisible force rules between them. This paper offers a two-stage computation that first selects the likely force terms and then tunes their strengths, testing it on four simulated particle systems.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The power-law validation in Table 1 is not a reconstruction: the dictionary r^p with integer p cannot represent r^{-1.5}, yet the 'True' column lists -1 and the learned -0.815 is presented as success, leaving the singular-kernel claim unsupported.","rationale":"The reader correctly identified the power-law dictionary mismatch as a structural assumption, but their weakest_assumption emphasized KDE accuracy and differentiability. I agree that KDE bandwidth and the Eq. (10) derivation slip are real concerns; however, the most load-bearing issue for the paper's stated conclusion is the explicit internal inconsistency in the power-law validation. The method may still be sound when the dictionary contains the true exponent, and the cubic, discontinuous, and 2D examples provide some positive support, but the singular power-law case is presented as a success despite the reported coefficient -0.815 versus the listed true -1. Because the paper's Section 5 claim explicitly rests on this example, the strength of the central claim is reduced until the experiment is redone with a dictionary capable of representing r^{-1.5}. The reader's CONDITIONAL verdict remains appropriate: the concern is concrete and testable, but it does not by itself invalidate the method for cases where the dictionary is correctly specified.","tokens_in":10732,"tokens_out":7833,"duration_ms":85999,"concrete_test":"Augment the dictionary in Section 3.1.1 with the basis function r^{-1.5} and rerun Case B of Section 4.2 using the same N, dt, cutoff radius rc, and algorithmic settings. If Phase II returns a coefficient close to -1 for r^{-1.5} (within about 5%) and suppresses the spurious r^{-1} term, then the original Table 1 reflects only dictionary misspecification and the sparse-regression engine is sound. If the pipeline fails to select r^{-1.5} or produces a coefficient far from -1, the central claim of accurate singular-kernel reconstruction is falsified. The authors should also report the KDE bandwidth and the hyperparameters alpha and lambda used in this rerun.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 5's central claim explicitly includes 'power-law repulsion-attraction potential' as a successfully reconstructed kernel. The evidence for that case is internally inconsistent. The dictionary D_poly in Section 3.1.1 is defined with integer exponents p, so r^{-1.5} is not in the dictionary. In Case B of Section 4.2 the true kernel is r - r^{-1.5}, but Table 1 reports a 'True' coefficient of -1 in the x^{-p} column and a 'Learned' coefficient of -0.815. If p=1, the 'True' entry is impossible because r^{-1.5} cannot be represented by r^{-1}; if p=1.5, the learned coefficient has an 18.5% error. Either way, the reported experiment does not demonstrate accurate reconstruction of the singular power-law kernel. The surrounding plots of density agreement and loss convergence only show that a misspecified model can fit the data in a coarse sense; they do not show that the correct kernel was recovered. Since this advertised example is one of the paper's four headline demonstrations, the conclusion that the method works 'in an accurate and robust way' across singular kernels is not supported by the provided evidence.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a two-phase regression method for extracting interaction kernels of stochastic many-particle systems from full particle trajectories. The trajectories are converted into a density field by kernel density estimation, and the residual of the mean-field PDE (Eq. 13) is minimized in Phase I with importance sampling and an adaptive threshold sparsification, after which Phase II refines the non-masked coefficients using all time steps. The method is demonstrated on one-dimensional systems with cubic, power-law, and piecewise-discontinuous kernels and on a two-dimensional radially symmetric cubic kernel, with comparisons based on learned coefficients, Wasserstein distance, free energy, and convergence plots.","tokens_in":10997,"tokens_out":6628,"duration_ms":66525,"significance":"If the results are valid, the proposal is a useful addition to sparse-identification methods for interaction kernels: it works directly from raw trajectories, uses a principled mean-field residual, and includes a sensible sparsification/refinement split. The clean recovery of the cubic coefficient (3.003 vs 3) in the one-dimensional case, the extension to a two-dimensional radially symmetric potential, and the system-level comparisons via Wasserstein distance and free energy are strong points, as is the explicit discussion of the method's limitations in high dimensions and for small data sets. The main reservations concern the consistency of the power-law validation in Table 1, missing hyperparameter and bandwidth disclosure, and a flaw in the displayed weak-form derivation; none of these is necessarily fatal, but they must be fixed before the advertised claims are fully supported.","major_comments":[{"comment":"The displayed weak formulation is not correct. The time integral must act on the whole expression: the correct term is ∫_0^t ∫ g'(x) · [∫ K(x,y)u(y,s) dy] u(x,s) dx ds, followed by the analogous diffusion term. As written, Eq. (10) has u(x,t) outside the ds integral, so the subsequent step of differentiating with respect to t is not valid. Since Eq. (13) is the basis of the regression loss in Eq. (17), this derivation should be corrected.","section":"Section 2, Eq. (10)"},{"comment":"The headline 'power-law repulsion-attraction' experiment is internally inconsistent. The true kernel is r - r^{-1.5}, but the dictionary D_poly in Section 3.1.1 contains only integer powers; moreover, the modified potential with cutoff radius r_c is constant on r ≤ r_c, so the exact kernel is not in the dictionary. Table 1 lists a 'True' coefficient of -1 for the x^{-p} term without stating p: if p=1 the target is not r^{-1.5}, and if p=1.5 the term is outside the declared dictionary. The learned coefficient -0.815 therefore does not demonstrate accurate reconstruction of the singular power-law kernel. Please either extend the dictionary to include p=1.5, rerun the experiment and report the fitted coefficient with its error, or explicitly analyze the projection error of the true kernel onto the dictionary and temper the corresponding claim in Section 5.","section":"Section 4.2 and Table 1"},{"comment":"The reproducibility of the method is blocked by missing numerical details. The KDE bandwidth, the finite-difference grid parameters, the regularization parameter λ, the sparsity intensity α, the optimization algorithm and stopping criterion for Phases I and II, and the exact dictionary used for the power-law case are not reported. Since the loss in Eq. (17) uses finite differences of the KDE density, the results can depend strongly on these choices; Figure 3 shows order-of-magnitude variation in accuracy with N and Δt. A complete table of hyperparameters for all four experiments should be provided.","section":"Sections 3 and 4"}],"minor_comments":[{"comment":"The text 'power-low repulsion-attraction potential' contains a typo and should read 'power-law'.","section":"Section 5"},{"comment":"The citation 'Kac[?]' is an unresolved placeholder and should be completed.","section":"Section 2, Theorem 2.3"},{"comment":"The row labeled 'Case D' in Table 1 is not linked to any 'Case D' in Section 4.4; please harmonize the case labels between the table and the text.","section":"Table 1"},{"comment":"The per-time-step error E_n used for importance sampling is never defined; please state its exact formula.","section":"Section 3.1.2, Eq. (18)"},{"comment":"The DNN baseline is described only in a few sentences; please provide architecture, training details, or a reference so that the comparison is meaningful.","section":"Section 4.3"},{"comment":"The paper does not include a data/code availability statement; for a purely numerical methodology, releasing code and parameter settings would substantially aid reproducibility.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The manuscript fits the scope of the journal. The main risk is that Table 1, Case B, will be cited as validation of singular-kernel recovery even though the dictionary and the reported true/learned coefficients are inconsistent; the authors should either rerun the experiment with a matching dictionary or explicitly quantify the projection error and soften the claim. No concerns about citation ethics were identified beyond the unresolved '[?]' placeholder."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: the two-phase algorithm (importance-sampled adaptive thresholding, then frozen-mask refinement on all time steps) is a real, modest improvement over existing SINDy-style sparse regression for mean-field kernel extraction, and the clean recovery of the cubic and 2D radial cases is credible. But the power-law case is not demonstrated as claimed, and the missing code, hyperparameters, and baseline comparison make the paper hard to evaluate. It deserves a serious referee, not a desk reject, but it needs major revision.\n\nWhat is new: the two-phase design is not in Lang & Lu (2022) or SINDy. Phase I using importance sampling on time steps with largest residual, plus an adaptive threshold to lock the sparsity mask, then Phase II refitting on all data with the mask frozen, is a sensible way to balance sparsity and accuracy. The discontinuity-location pre-step via DNN gradients is also a reasonable add-on for piecewise kernels. The numerical results for the cubic (3.003 vs 3) and 2D radial (2.916 vs 3) cases, and the piecewise linear opinion dynamics coefficients, are decent evidence the pipeline works when the kernel is in the dictionary.\n\nSoft spots, in order of severity. First, the power-law validation is internally inconsistent. The dictionary is defined with integer exponents p, but the test kernel is r - r^{-1.5}. Table 1 lists a \"True\" coefficient of -1 in the x^{-p} column and learned -0.815. If p=1, the \"True\" entry is impossible; if p=1.5, the learned coefficient has an 18.5% error. Either way the singular-kernel claim is not supported. The paper should either use a dictionary that actually contains r^{-1.5}, or report the approximation error honestly. Second, hyperparameters (lambda, alpha, KDE bandwidth, the exponent p) are undisclosed, and no code or data is provided, so the experiments are not reproducible. Third, there is no baseline comparison against Lang & Lu 2022, the most directly relevant prior method, so the claimed speedup or accuracy gain is not quantified against the state of the art. Minor: Eq. (10) has a time-index slip (u(x,t) outside the time integral should be u(x,s)), and the Kac reference is missing in Theorem 2.3. These are fixable.\n\nWho it is for: people working on data-driven discovery of interaction rules from particle trajectories or mean-field PDEs. The central idea is useful enough to warrant referee time, but the paper as written overclaims \"accurate and robust\" reconstruction across singular kernels. I would recommend sending to peer review with a request for major revision: fix the power-law experiment, add code/data and hyperparameters, and run a baseline comparison.","headline":"A plausible speedup-and-sparsity trick for mean-field kernel inversion, but the advertised power-law reconstruction is not supported as reported.","tokens_in":11540,"tokens_out":2579,"would_cite":false,"duration_ms":23282,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that unknown pairwise interaction rules in stochastic many-particle systems can be recovered directly from trajectory data by converting snapshots into a density and fitting the mean-field equation with a sparse…","keywords":["interaction kernels","many-particle systems","stochastic dynamics","mean-field equation","sparse regression","importance sampling","kernel density estimation","machine learning"],"falsifier":"Simulate a system with an interaction kernel that is not in the chosen dictionary, for example $\\varphi(r)=r^{-1.5}$ using only integer-power basis functions, and monitor the mean-field residual after Phase II. The learned coefficient $-0.815$ instead of $-1$ shows that no coefficient vector can represent the true kernel; if the residual floor stays well above machine precision, the method is recovering the best dictionary fit rather than the true kernel.","tokens_in":1693,"feed_emoji":"🧮","tokens_out":2766,"duration_ms":90083,"temperature":0.7,"pith_summary":"The paper aims to establish that the interaction kernel of a stochastic many-particle system can be extracted from raw trajectory data without knowing the force law in advance. Its pipeline turns particle trajectories into a density via kernel density estimation, treats that density as a solution of the McKean–Vlasov mean-field equation, and solves a sparse regression problem for the kernel coefficients. A first phase uses importance-weighted time sampling and an adaptive threshold to select which dictionary terms matter, and a second phase refines those coefficients using the full dataset. If the method works as claimed, experimental or simulated trajectories alone would yield explicit force laws such as cubic potentials, singular repulsion–attraction forces, and piecewise opinion dynamics, with coefficients matching the true values to within a few percent.","feed_headline":"Two-phase method recovers interaction kernels from trajectories","feed_subtitle":"Fitting the mean-field equation to estimated densities recovers cubic, singular, and piecewise potentials.","key_machinery":"The load-bearing object is the constant-diffusion McKean–Vlasov mean-field equation $\\partial_t u + \\nabla_x\\cdot[u(K*u)] = \\frac{\\sigma_c^2}{2}\\Delta u$, which becomes linear in the coefficients once the kernel is written as $K(x,y)=\\varphi(\\|x-y\\|_2)\\frac{x-y}{\\|x-y\\|_2}$ and $\\varphi$ is expanded in a dictionary. Kernel density estimation supplies the density $u$; finite differences supply $\\partial_t u$ and $\\Delta u$; discrete convolution supplies $K*u$. Phase I samples time steps with probability proportional to their current residual error, updates coefficients, and applies the threshold $\\tau_k=\\alpha\\max_i|\\zeta_i^{(k)}|$ to build a binary mask; Phase II fixes that mask and solves the full-data least-squares problem with an $\\ell^1$ penalty.","core_discovery":"The central claim is that the residual of the mean-field equation, computed from kernel-density-estimated densities, is linear in the kernel coefficients, so identifying the interaction kernel becomes a sparse linear regression problem. Using a polynomial or physics-informed dictionary, Phase I selects active terms with importance sampling and adaptive thresholding, and Phase II refines the masked coefficients over all time steps. The reported results include a cubic kernel $\\varphi(r)=3r^2$ recovered as $3.003$, a power-law repulsion–attraction kernel with coefficients $0.986$ and $-0.815$ (true values $1$ and $-1$), piecewise opinion kernels recovered to within about $0.01$, and a two-dimensional radial cubic kernel recovered as $2.916$ (true $3$). The same pipeline also reproduces free-energy curves, with Wasserstein distances between true and learned densities staying below $10^{-2}$ in the tested cases.","pith_inferences":["The method's accuracy is bounded by dictionary coverage: a kernel outside the chosen basis can only be approximated, as the power-law case's $-0.815$ coefficient shows when the true term is $r^{-1.5}$; moving to adaptive dictionaries or nonparametric bases is a natural next step.","The same residual-regression idea should transfer to weak-form or sample-based formulations for high-dimensional systems, a direction the authors themselves flag as future work.","If trajectories of interacting agents can be obtained, this pipeline offers a way to infer effective interaction rules for agent-based systems without knowing the rules a priori.","A testable extension would be to use the Phase I mask as a model-selection device across physically motivated dictionaries, such as Coulomb, Morse, or Lennard-Jones forms, on real experimental tracks."],"forward_implications":["Explicit interaction laws can be inferred from trajectory data alone, so downstream simulations can use the learned coefficients rather than a fitted black-box model.","Importance sampling reduces the Phase I computation time by about an order of magnitude, making large candidate dictionaries feasible.","The method handles smooth, singular, discontinuous, and two-dimensional radially symmetric kernels, so it applies to a broad class of pairwise interaction models.","Accuracy improves sharply with particle number, with Wasserstein error dropping by more than two orders of magnitude from $N=5000$ to $N=20000$, confirming the mean-field limit as the right bridge for large systems.","Macroscopic quantities such as free energy are also recoverable, allowing thermodynamic or collective-behavior comparisons beyond the kernel itself."],"supporting_citations":[{"why":"Establishes the mean-field-equation-based regression formulation for learning interaction kernels that this paper adapts and extends to a two-phase sparse scheme.","marker":"[5]"},{"why":"Derives the McKean–Vlasov equation used as the residual target for kernel recovery.","marker":"[12]"},{"why":"Provides propagation-of-chaos convergence results justifying the replacement of the N-particle system by its mean-field density.","marker":"[13]"},{"why":"Supplies the sparse-identification and thresholding paradigm on which Phase I's adaptive sparsification is built.","marker":"[2]"},{"why":"Defines the opinion-dynamics kernels with discontinuities used as numerical test cases.","marker":"[11]"},{"why":"Gives the Coulomb $r^{-2}$ motivation for the power-law repulsion-attraction test kernel.","marker":"[14]"},{"why":"Points to weak-form alternatives the paper cites for future high-dimensional extensions.","marker":"[16]"}],"fun_headline_variants":["Two-phase sparse regression learns particle interaction kernels","Sparse mean-field fit recovers interaction kernels from trajectories","Two-phase method extracts kernels from particle trajectories","Mean-field regression identifies interaction kernels in many-particle systems","Sparse regression on mean-field residual identifies interaction kernels"],"cache_read_input_tokens":13568,"weakest_assumption_plain":"Everything rests on the assumption that the density estimated from a finite number of particle snapshots is close enough to a smooth solution of the mean-field equation that its numerical derivatives and convolutions have small error, which requires many particles, a fine grid, and a small time step.","fun_headline_variants_meta":{"raw":{"variants":["Two-phase sparse regression learns particle interaction kernels","Sparse mean-field fit recovers interaction kernels from trajectories","Two-phase method extracts kernels from particle trajectories","Mean-field regression identifies interaction kernels in many-particle systems","Sparse regression on mean-field residual identifies interaction kernels"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000575,"raw_usage":{"total_tokens":2668,"prompt_tokens":852,"completion_tokens":1816,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":468,"completion_tokens_details":{"reasoning_tokens":1743}},"tokens_in":468,"tokens_out":1816,"duration_ms":13969,"temperature":1.0,"reasoning_tokens":1743,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T22:44:07.285691+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate a system with an interaction kernel that is not in the chosen dictionary, for example $\\varphi(r)=r^{-1.5}$ using only integer-power basis functions, and monitor the mean-field residual after Phase II. The learned coefficient $-0.815$ instead of $-1$ shows that no coefficient vector can represent the true kernel; if the residual floor stays well above machine precision, the method is recovering the best dictionary fit rather than the true kernel.","supporting_citations":[{"cited_title":"Learning interaction kernels in mean-field equations of first-order systems of interacting particles","cited_arxiv_id":null,"evidence_quote":"Establishes the mean-field-equation-based regression formulation for learning interaction kernels that this paper adapts and extends to a two-phase sparse scheme."},{"cited_title":"A class of markov processes associated with nonlinear parabolic equations","cited_arxiv_id":null,"evidence_quote":"Derives the McKean–Vlasov equation used as the residual target for kernel recovery."},{"cited_title":"Topics in propagation of chaos","cited_arxiv_id":null,"evidence_quote":"Provides propagation-of-chaos convergence results justifying the replacement of the N-particle system by its mean-field density."},{"cited_title":"Heterophilious dynamics enhances consensus","cited_arxiv_id":null,"evidence_quote":"Defines the opinion-dynamics kernels with discontinuities used as numerical test cases."},{"cited_title":"Maxwell’s equations","cited_arxiv_id":null,"evidence_quote":"Gives the Coulomb $r^{-2}$ motivation for the power-law repulsion-attraction test kernel."},{"cited_title":"Weak collocation regression method: Fast reveal hidden stochastic dynamics from high-dimensional aggregate data","cited_arxiv_id":null,"evidence_quote":"Points to weak-form alternatives the paper cites for future high-dimensional extensions."}],"review_version":1}