{"id":"0d449949-7bb0-4f2b-8208-a379adfee77a","arxiv_id":"2506.15657","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Machine learning models trained on DEM simulation data can interpolate and, within narrow ranges, extrapolate the gas-cluster transition boundary in a vibrated granular gas.","lead":"This paper trains machine learning models on DEM simulations of granular gases to predict whether particles cluster under vibration. The models reproduce simulation-derived clustering criteria on held-out data and could help plan microgravity experiments without running expensive simulations.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Sec. 3.2 says only θ=0 and θ=π were simulated, while Sec. 4 builds the dataset on θ/π=0,0.25,…,1; this direct contradiction means the reported three-parameter surrogate may not be supported by the data as described.","rationale":"I agree with the reader that the physical validity of the clustering criteria is an important caveat, but it is explicitly conceded in Sec. 5, making it a stated limitation rather than a hidden defect. The more immediate problem is the unacknowledged contradiction in the data-generation protocol: Sec. 3.2 restricts the piston phase shift to θ=0 and θ=π, while Sec. 4 uses five or nine θ levels in the dataset and reports interpolation/extrapolation performance over that θ grid. The reported sample sizes (385 and 693) are only consistent with multi-level θ grids, not with a two-level θ grid. Unless the intermediate-θ simulations exist but are misdescribed, the training labels for those points have no stated source, and the central claim of a three-parameter surrogate over N, A, θ is not supported. This is a concrete, checkable issue rather than a matter of consensus or external validation. I therefore recommend UNVERDICTED: the paper should not be accepted, rejected, or conditionally accepted until the parameter grid used for training is confirmed and the contradiction is resolved. If the authors can show that intermediate θ values were indeed simulated, the verdict would return to CONDITIONAL with the reader's external-criterion caveat as the main remaining condition.","tokens_in":16571,"tokens_out":5711,"duration_ms":74580,"concrete_test":"Request the DEM input parameter files or simulation logs and enumerate the unique θ values actually run. If no simulation used θ/π=0.25, 0.5, or 0.75, then the reported dataset does not match the stated protocol and the θ dimension of the surrogate is unsupported. If such runs do exist, correct the Sec. 3.2 sentence and verify the sample counts against the unique θ levels.","verdict_should_be":"UNVERDICTED","load_bearing_attack":"The paper's central claim is that ML models trained on DEM data predict the gas-cluster state as a function of N, A, and θ without further DEM runs. For that claim, the training set must actually contain the parameter values described. This premise is contradicted internally. Sec. 3.2 states: 'The choice of phases between the wall vibrations was restricted to in-phase (θ = 0) and anti-phase (θ = π) piston movement.' Yet Table 2 and Sec. 4.1.1 report θ/π from 0 to 1 with step 0.25, and the interpolation datasets have sizes 385 and 693, which factor as 11×7×5 and 11×7×9, respectively. Such full-factorial grids require 5 and 9 θ levels, not 2. If only θ=0 and θ=π were simulated, the reported θ-dependent training points (and any interpolation/extrapolation metrics over θ) could not have been generated as described. This is not a physics disagreement; it is a data-provenance inconsistency that directly affects the surrogate's claimed input space. The reader's external-validation caveat about the clustering criteria is genuine but explicitly acknowledged in Sec. 5, so it is not hidden; the θ contradiction is unacknowledged and should be resolved before the central claim can be assessed.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript trains supervised machine-learning regressors and classifiers on Discrete Element Method (DEM) simulations of a vibrated granular gas in the VIP-Gran geometry. The target quantities are two established clustering indicators derived from the simulations: the minimum KS-test margin δ_min and the maximum caged-particle fraction φ_caged. The input parameters are particle number N, excitation amplitude A, and phase shift θ between the two vibrating walls. The authors report interpolation and extrapolation experiments for five regression and five classification algorithms, selecting a best model in each task. The central claim is that machine learning can predict the gas-cluster state for a given set of system parameters without running new DEM simulations.","tokens_in":16847,"tokens_out":6584,"duration_ms":74292,"significance":"If the reported results are reproducible, the paper would provide a useful proof-of-concept for using ML surrogates to accelerate parameter-space exploration in microgravity granular-gas experiments. The study has several genuine strengths: the learning targets come from DEM simulations external to the ML models, so there is no circularity; the parameter grid is systematic; five regression and five classification methods are compared; and the extrapolation tests go beyond simple in-sample fitting. The authors also explicitly acknowledge in Sec. 5 that the physical appropriateness of the two clustering criteria was not verified, which is an honest limitation. However, the manuscript currently contains a load-bearing data-provenance contradiction concerning the phase-shift parameter, as well as internal inconsistencies in the reported performance metrics, so the central claim cannot yet be assessed as stated.","major_comments":[{"comment":"There is a direct contradiction about the θ grid. Sec. 3.2 states that 'the choice of phases between the wall vibrations was restricted to in-phase (θ = 0) and anti-phase (θ = π) piston movement,' but Sec. 4.1.1 builds the interpolation dataset on 'θ/π from 0 to 1 with step size of 0.25,' and the reported dataset size 385 = 11×7×5 requires five θ levels. Similarly, Sec. 4.2.1 reports a caging-interpolation dataset of 693 = 11×7×9 samples, which requires nine θ levels. The extrapolation dataset sizes in Sec. 4.1.1 (315 = 9×7×5 and 70 = 2×7×5) and Sec. 4.2.1 (567 = 9×7×9 and 126 = 2×7×9) also require multiple intermediate θ values. These statements cannot both be true. Because every interpolation and extrapolation result and the abstract's three-parameter claim depend on θ-resolved training data, this is a load-bearing issue: either Sec. 3.2 is wrong and intermediate θ values were actually simulated, or the dataset descriptions and the surrogate's claimed input space are unsupported. Please state the actual θ grid explicitly, reconcile the sample counts, and if intermediate θ values were not simulated, remove or explicitly restrict all claims about θ dependence.","section":"Sec. 3.2 vs Secs. 4.1.1 and 4.2.1"},{"comment":"The reported metrics in Table 3 are internally inconsistent. The Random Forest Regression row lists RMSE = 0.001 and R² = 0.894, while the Neural Network Regression row lists RMSE = 0.007 and R² = 0.967. For the same target dataset, a lower RMSE must correspond to a higher R², so these entries cannot both be correct as printed. The text then states that the ANN 'has the lowest RMSE and MAE values,' which is not supported by the table because the Random Forest RMSE is an order of magnitude smaller. Since the ANN is selected as the best interpolation model on this basis, this inconsistency directly affects the main regression claim. Please verify all entries in Tables 3–5 and 10–12 and correct any typos.","section":"Sec. 4.1.1, Table 3"},{"comment":"All performance metrics are computed on a single random split, and the 'best' model is selected by comparing test-set metrics across algorithms. With a single split, the reported RMSE, R², and AUC values have no uncertainty estimates, and the model ranking may be an artifact of that particular split. In addition, using the test set for model selection means the reported test performance is optimistically biased. To support the quantitative 'best model' claims, the authors should report repeated k-fold or bootstrap confidence intervals, or clearly separate a validation set used for model selection from a hold-out test set used exactly once for final evaluation.","section":"Secs. 3.6, 4.1.1, 4.2.1"},{"comment":"The abstract and conclusions claim prediction of the 'gas-cluster transition' without qualification, but Sec. 5 acknowledges that the appropriateness of the two clustering criteria (KS test with α = 0.05 and caging with T_caged = 0.285, φ_caged > 0.05) was not verified. The ML prediction target is therefore the state as defined by these two criteria applied to DEM data, not necessarily the physical transition in a real microgravity experiment. This caveat is acknowledged, but it should also be reflected in the abstract and in the phrasing of the central claim.","section":"Sec. 5 and Abstract"}],"minor_comments":[{"comment":"Typo: 'howest' should be 'lowest' in the sentence introducing the ANN as the best interpolation model.","section":"Sec. 4.1.1"},{"comment":"Typo: 'Of cause' should be 'Of course' in the summary paragraph of the caging-regression results.","section":"Sec. 4.2.1"},{"comment":"The caption contains a duplicated article: 'a cross-sectional view of the of the VIP-Gran experiment.' Please edit.","section":"Fig. 1 caption"},{"comment":"Reference [28] (Detectron2) and reference [29] (a general article on p-values) do not appear to support the statements they are attached to: the transient-period discussion in Sec. 3.2 and the φ_caged > 0.05 threshold in Sec. 3.4, respectively. Please replace them with appropriate citations.","section":"References"},{"comment":"Several figure captions state that θ = π was fixed, while the surrounding text describes datasets with θ varying. Please clarify in each caption whether the displayed data are restricted to θ = π or are aggregated across θ.","section":"Figures 4, 6, 8, 12, 14, 16"},{"comment":"Typo: 'comprized' should be 'comprised' in the description of the simulation duration.","section":"Sec. 3.2"}],"recommendation":"major_revision","confidential_remarks":"The θ contradiction is the chief risk to the paper's central claim. If the simulations truly used only θ = 0 and θ = π, then the reported interpolation and extrapolation results involving θ cannot be generated as described, and the paper's main conclusion would not be supported. I would ask the authors to provide the actual simulation parameter grid or the raw dataset, and to clarify whether intermediate θ values were simulated. The manuscript also lacks a data availability statement; given the ML focus, making the dataset and code available would substantially strengthen the reproducibility of the claims."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know two things before spending time on this one. It is a clearly written, well-scoped benchmark of off-the-shelf ML models for emulating DEM-derived clustering in the VIP-Gran geometry. It is also, as written, not a trustworthy surrogate paper, because the data description contradicts itself on a load-bearing point.\n\nThe genuinely new piece is the systematic comparison of five regression and five classification models for two clustering criteria (KS-test delta_min and caging phi_caged), with separate interpolation and extrapolation tests, and with the phase shift theta treated as a third input. That is a legitimate engineering contribution, though a modest one. The methods are standard; the paper does not oversell the physics. Section 5 explicitly states that the two clustering criteria were not verified here, which is refreshing honesty. The reporting of metrics, confusion matrices, and ROC curves is thorough, and the extrapolation tables show failures as well as successes, which is more than many such papers do.\n\nNow the soft spots, in proportion. The load-bearing problem is an internal contradiction about the phase shift. Section 3.2 says the wall vibrations were restricted to in-phase (theta=0) and anti-phase (theta=pi). But Section 4 describes datasets with theta/pi = 0, 0.25, 0.5, 0.75, 1 (and in one place nine theta levels), and the stated sample sizes (385, 693, etc.) only work out if those intermediate theta values were actually simulated. If only two phases were simulated, the reported three-parameter interpolation and extrapolation results cannot have been produced as described. This is not a typo in a formula; it is a contradiction about what data exist. The paper needs to correct one of those statements before the central claim can be assessed.\n\nSecond, the model selection is done on the test set. The paper uses the test-split metrics to pick the \"best\" model, which is test-set peeking. A single random split with no confidence intervals compounds the problem. Third, the extrapolation results are modest: polynomial regression gets R^2 about 0.49 for delta_min when extrapolating in N, and the paper's summary says the models \"effectively interpolate and extrapolate\" — that overstates what the tables show, even though the limited-extension caveat is present.\n\nWho is this for? Someone building ML surrogates for DEM-based granular gas studies, or planning microgravity experiments with VIP-Gran. It is a useful case study in benchmarking and a cautionary example of data-provenance reporting. My bottom line: the underlying work may be salvageable, but the current version should not be accepted. If I were the editor, I would ask the authors to resolve the theta contradiction and resubmit, then send it to a serious referee. The topic and the effort deserve that much; the current manuscript does not.","headline":"Clean ML benchmark for DEM-derived clustering, but a direct contradiction about the phase-shift grid guts the data provenance as written.","tokens_in":17340,"tokens_out":4740,"would_cite":false,"duration_ms":51877,"reading_group":"maybe","serious_thinker":"no","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Machine learning models trained on discrete-element simulations predict the gas–cluster transition in vibrated granular gases from particle number, excitation amplitude, and wall phase shift, without running new simulations.","keywords":["granular gases","dynamical clustering","gas-cluster transition","machine learning","discrete element method","Kolmogorov-Smirnov test","caging criterion","microgravity"],"falsifier":"A microgravity experiment that measures particle positions for parameter sets inside the trained range, for example $N \\approx 4000$, $A \\approx 4.5$ mm, and $\\theta \\approx \\pi/2$, and compares the observed onset of dense clustering with the predicted gas/cluster label would settle the matter. If the predicted boundary disagrees with observation by more than the model's interpolation error, the central claim fails. A cheaper simulation-level check is to test whether the KS-based and caging-based labels agree with each other on the full parameter grid; systematic disagreement would mean the models are fitting an ambiguous target rather than a single transition.","tokens_in":16406,"feed_emoji":"📊","tokens_out":8011,"duration_ms":79508,"temperature":0.7,"pith_summary":"The paper argues that the gas–cluster transition in a microgravity granular gas can be predicted by machine learning once it has been labeled by two established statistical criteria. The labels come from discrete element method (DEM) simulations of frictional spheres in a cuboid container with two vibrating walls, where clustering is detected by a Kolmogorov–Smirnov test on local packing-fraction profiles and by a caging criterion based on Voronoi cell volumes. Across the tested range of particle number, excitation amplitude, and wall phase shift, several regression and classification models reproduce the DEM-derived transition parameters and the gas/cluster labels, with the best performers differing between interpolation and extrapolation tasks. If correct, this gives a fast surrogate for DEM simulations that currently cost roughly ten hours per run, and it can map the phase space of a preparative microgravity experiment before flight.","feed_headline":"ML predicts granular gas clustering without costly simulations","feed_subtitle":"Trained on discrete-element simulations, the models map particle number, amplitude, and wall phase to the gas or cluster state.","key_machinery":"The machinery is a two-stage pipeline. Stage one reduces the high-dimensional DEM state—positions, velocities, and contacts of thousands of spheres—to two scalar order parameters: $\\delta_{\\min}$, the minimum over the excitation cycle of the gap between the Kolmogorov–Smirnov statistic and its threshold, and $\\phi_{\\mathrm{caged}}$, the maximum fraction of particles whose Voronoi cell volume ratio exceeds the caging threshold. Stage two trains standard regression and classification models on the map from particle number, amplitude, and phase shift to these scalars and to the resulting binary gas/cluster label. The load-bearing identity is the correspondence between the sign of $\\delta_{\\min}$ (or the $\\phi_{\\mathrm{caged}} > 0.05$ threshold) and the physically observed clustered state; the machine learning only learns that correspondence from DEM-generated labels.","core_discovery":"The central claim is that for a specific experimental geometry, the gas–cluster transition of frictional spheres is a predictable function of just three parameters—number of particles, piston amplitude, and phase shift between the two vibrating walls—and that a small set of standard machine learning models can learn this function from DEM data. With the KS-derived label $\\delta_{\\min}$, a neural-network regressor reaches $R^2=0.967$ on interpolation, while polynomial and XGBoost regressors are the only ones that extrapolate to larger particle numbers and amplitudes without collapsing; support vector and neural-network regressors fail on these extrapolations. With the caging criterion $\\phi_{\\mathrm{caged}}$, XGBoost is best at interpolation and polynomial regression is best at both extrapolation tasks. For classification, random forest is best for KS-based interpolation and neural network is best for extrapolation, while XGBoost and neural network share top performance for caging interpolation. The paper does not claim a universal model; it claims that this low-dimensional phase space is simple enough that a few hundred DEM samples suffice to train reliable predictors, and that the same workflow can be adapted to other clustering metrics and geometries.","pith_inferences":["The best-model hierarchy is dataset-dependent: the neural network interpolates $\\delta_{\\min}$ best but fails on extrapolation, suggesting its smoothness prior fits the data interior but not the boundary; a model that respects physical constraints at low fill fractions might generalize better.","Treating the phase shift $\\theta$ as a periodic feature (for example feeding $\\cos\\theta$ and $\\sin\\theta$ instead of a linear input) would exploit the symmetry of the two-wall excitation and could improve extrapolation in $\\theta$, a direction the paper does not explore.","Because both criteria ignore particle dynamics, a surrogate trained on either captures only static density structure; adding velocity-correlation or collision-rate features could target the dynamical onset the paper acknowledges is missing.","The extrapolation claims hold only within the tested ranges and for the fixed geometry and particle diameter; they do not suggest validity across container size, particle size, or material."],"forward_implications":["A trained surrogate maps the full tested N–A–θ grid in seconds, replacing roughly ten-hour DEM runs for each point.","Training data can be taken on a regular grid or randomly distributed over phase space, so the sample budget is small and flexible.","If a better clustering metric is found, the same ML pipeline can be retrained on it with little adaptation, as the paper notes.","Experiment planning for microgravity campaigns can use the predicted phase diagram to pre-select parameter ranges before flight.","The same approach should transfer to other dynamic systems with many degrees of freedom that are statistically classified by a few parameters."],"supporting_citations":[{"why":"Defines dynamical clustering and supplies both the KS-test and caging criteria used to label states.","marker":"[1]"},{"why":"Establishes the experimental phase diagram in low gravity and fixes the KS significance level and transient length.","marker":"[5]"},{"why":"Parametric study of the clustering transition for frictional spheres whose friction and restitution choices are adopted.","marker":"[8]"},{"why":"Provides the GPU-based DEM software that generated all simulation trajectories.","marker":"[11]"},{"why":"Shows that reduced Young's modulus does not alter statistical results, justifying the fast DEM setup.","marker":"[19]"},{"why":"Source of the phi_caged greater than 0.05 threshold used to label clustered states.","marker":"[29]"},{"why":"Supplies the machine learning implementations for regression and classification used in the study.","marker":"[30]"}],"fun_headline_variants":["ML predicts granular gas clusters from just three parameters","Three parameters, one geometry: ML nails gas clustering","Machine learning forecasts gas-cluster transition in granular gases","ML model maps gas vs cluster state with only three inputs","Predicting granular gas clustering: ML needs only three numbers"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The argument rests on the assumption that the two DEM-derived criteria—the KS test at $\\alpha=0.05$ and the caging threshold $\\phi_{\\mathrm{caged}} > 0.05$—actually mark the real gas–cluster transition in a microgravity granular gas, and the paper states explicitly that this was not verified.","fun_headline_variants_meta":{"raw":{"variants":["ML predicts granular gas clusters from just three parameters","Three parameters, one geometry: ML nails gas clustering","Machine learning forecasts gas-cluster transition in granular gases","ML model maps gas vs cluster state with only three inputs","Predicting granular gas clustering: ML needs only three numbers"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000679,"raw_usage":{"total_tokens":3086,"prompt_tokens":945,"completion_tokens":2141,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":561,"completion_tokens_details":{"reasoning_tokens":2064}},"tokens_in":561,"tokens_out":2141,"duration_ms":15379,"temperature":1.0,"reasoning_tokens":2064,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T23:51:46.432712+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A microgravity experiment that measures particle positions for parameter sets inside the trained range, for example $N \\approx 4000$, $A \\approx 4.5$ mm, and $\\theta \\approx \\pi/2$, and compares the observed onset of dense clustering with the predicted gas/cluster label would settle the matter. If the predicted boundary disagrees with observation by more than the model's interpolation error, the central claim fails. A cheaper simulation-level check is to test whether the KS-based and caging-based labels agree with each other on the full parameter grid; systematic disagreement would mean the models are fitting an ambiguous target rather than a single transition.","supporting_citations":[{"cited_title":"Europhys","cited_arxiv_id":null,"evidence_quote":"Establishes the experimental phase diagram in low gravity and fixes the KS significance level and transient length."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Parametric study of the clustering transition for frictional spheres whose friction and restitution choices are adopted."},{"cited_title":"AIP Conference Proceedings 1542, 169–172 (2013) https://doi.org/10","cited_arxiv_id":null,"evidence_quote":"Provides the GPU-based DEM software that generated all simulation trajectories."},{"cited_title":"Scientific Reports 11(1), 10621 (2021) https://doi.org/10.1038/ s41598-021-89949-z","cited_arxiv_id":null,"evidence_quote":"Shows that reduced Young's modulus does not alter statistical results, justifying the fast DEM setup."},{"cited_title":"https://rdcu.be/dv7LL","cited_arxiv_id":null,"evidence_quote":"Source of the phi_caged greater than 0.05 threshold used to label clustered states."}],"review_version":1}