{"id":"0c734a24-c532-45d7-b6f6-f0292e5eb800","arxiv_id":"2508.05308","paper_version":2,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"On a DEM packing-fraction dataset, a histogram-based gradient boosting model beat 15 other regression models in a k-fold cross-validation benchmark.","lead":"The abstract reports a cross-validation benchmark that selects a histogram-based gradient boosting model as the best of 16 regression surrogates for DEM particle-simulation data. The manuscript body supplied is a different paper (a vision transformer), so this result could not be verified and the review rests on the abstract alone.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Abstract claims HGBM optimal for DEM regression, but supplied full text is a different paper (arXiv:2508.05307, CoCA ViT); no benchmark evidence is present.","rationale":"The reader's verdict is UNVERDICTED with LOW confidence, and I agree. The supplied full text is internally inconsistent with the abstract: the body is a completely different paper on vision transformers, as evidenced by its own arXiv ID footer. Under the review rule that every part of the manuscript is in-scope evidence, this mismatch is treated as an explicit lack of support for the abstract's claims. The abstract's central claim about HGBM optimality cannot be verified from any data, methods, or results in the document. The reader's weakest_assumption focuses on generality across datasets, but the more fundamental issue is that the benchmark itself is absent. Before assessing external validity, one must first verify internal existence and reporting. My concrete test is a direct lookup of the official arXiv record to settle whether the supplied text is a packaging error and whether the real paper contains the claimed benchmark. Thus the verdict should remain UNVERDICTED; no change from the reader's verdict is needed.","tokens_in":13078,"tokens_out":2571,"duration_ms":25936,"concrete_test":"Retrieve the full text associated with arXiv ID 2508.05308 (from arXiv's official PDF or HTML source) and verify that (i) the abstract matches the body, (ii) the body contains a benchmark of 16 regression models on the described packing-fraction dataset, (iii) k-fold cross-validation settings and hyperparameters are specified, and (iv) performance differences include error bars or statistical significance. If the official full text is actually the CoCA ViT paper rather than the DEM benchmark, the central claim is unverifiable and the verdict remains UNVERDICTED.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central assertion is that, on a DEM packing-fraction dataset with five varying inter-particle properties, a histogram-based gradient boosting model is optimal among 16 regression models selected via k-fold cross-validation. The manuscript body supplied in the review packet is not that paper. It carries the arXiv footer 'arXiv:2508.05307v1 [cs.CV]' and presents a compact vision transformer (CoCA ViT) with ImageNet classification results. There is no description of the DEM dataset, no list of the 16 models, no k-fold protocol, no hyperparameter choices, no error bars, and no code or reproducibility artifacts. Therefore the central claim is unsupported by the available evidence. This is not a disagreement with conventional tabular-ML expectations; it is an internal evidentiary gap. The abstract's claim could be true, but nothing in the supplied full text demonstrates it. The reader's concern about dataset generality is secondary: before asking whether the result transfers to other DEM tasks, one must establish that the experiment exists and is reported correctly.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript as supplied is internally inconsistent. The abstract (arXiv:2508.05308, physics.comp-ph) announces a study of regression surrogates for discrete-element-method (DEM) datasets: a k-fold cross-validation framework, a DEM packing-fraction dataset with five varying particle properties, a comparison of 16 regression models, and the conclusion that a histogram-based gradient boosting model is optimal. However, the full text of the submission is a different paper (arXiv:2508.05307v1, cs.CV), titled \"CoCA ViT: Compact Vision Transformer with Robust Global Coordination,\" which describes a vision transformer architecture, ImageNet classification, COCO detection, and ADE20K segmentation. No DEM dataset, no list of the 16 models, no k-fold protocol, no hyperparameters, no error bars, no code, and no reproducibility artifacts appear anywhere in the submitted text. The central claim of the abstract is therefore entirely unsupported by the submitted manuscript.","tokens_in":13216,"tokens_out":1396,"duration_ms":15793,"significance":"If the DEM-regression study described in the abstract were actually presented and correct, the practical contribution would be useful: a reusable model-selection workflow for tabular DEM surrogate modeling, with a concrete recommendation (histogram-based gradient boosting) on a packing-fraction dataset. However, none of that content is present in the manuscript text under review. The supplied body instead reports a computer-vision architecture with ImageNet results. Consequently, the claimed significance cannot be evaluated, and the manuscript in its current form contains no falsifiable evidence for its stated conclusions. It is also not reproducible from the submitted materials.","major_comments":[{"comment":"The abstract and the full text are different papers. The abstract (arXiv:2508.05308, physics.comp-ph) claims a DEM regression benchmark with 16 models, k-fold cross-validation, and an optimal histogram-based gradient boosting model; the full text is \"CoCA ViT: Compact Vision Transformer with Robust Global Coordination,\" an image-classification paper. None of the claimed DEM content appears. This is a load-bearing mismatch: the manuscript provides no methods, dataset description, model list, results tables, or code for its stated central claim. The central assertion is therefore unsupported in the submitted text.","section":"Abstract vs. Full Text (arXiv footer line: arXiv:2508.05307v1 [cs.CV])"},{"comment":"The full text contains only vision experiments: ImageNet-1K top-1 accuracy, ImageNet-Real, ImageNet-V2, ImageNet-A/R, COCO mAP, and ADE20K mIOU. There is no packing-fraction dataset, no list of the 16 regression models, no k-fold cross-validation description, no hyperparameter settings, no training/inference timings for tabular surrogates, and no error bars or variance estimates. The reader's report and the abstract cannot be cross-checked against any supporting evidence in the manuscript.","section":"Full Text (Sections 1-5, Tables 1-2)"},{"comment":"Even under a generous reading where the abstract is the intended contribution, the submission is not reviewable as a physics.comp-ph paper: it contains no reproducibility artifacts, no dataset description (size, feature ranges, simulation parameters), no model-selection protocol details beyond the phrase \"k-fold cross-validation,\" and no definition of \"optimal\" (accuracy vs. training/inference time trade-off). The absence of all of these is not a presentation issue but a complete evidentiary gap for the central claim.","section":"Entire submission"}],"minor_comments":[{"comment":"The title, abstract, and arXiv identifier in the header conflict with the full text's title and arXiv identifier. This needs resolution before any substantive review can occur.","section":"Title and metadata"}],"recommendation":"reject","confidential_remarks":"The submitted manuscript appears to be a different paper from the one described in its abstract. This may be a submission or file-conversion mix-up rather than a scientific problem with the DEM work itself. The appropriate action may be for the authors to resubmit the correct manuscript; however, as supplied, the paper is not reviewable because the claimed subject matter is absent. I recommend rejection of the current submission, with the possibility of resubmission of the correct manuscript."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick read. The full text we were sent is not the DEM paper; it's a CoCA ViT vision paper with a different arXiv ID. So I can only react to the abstract, and that's a real problem: the central claim—that a histogram-based gradient boosting model beat 15 others on a DEM packing-fraction dataset—is unsupported by anything in front of us. This isn't a scientific error I can point to; it's an evidentiary gap. If the actual manuscript matches the abstract, the topic is worthwhile. Applying k-fold CV to choose surrogate models for DEM tabular data is a practical need, and the result is consistent with what tabular ML would predict: gradient boosting usually wins on structured data. The proposed framework, simple as it is, could help practitioners stop picking models by habit.\n\nSoft spots, in proportion. The single-dataset generality is a minor-to-moderate concern: packing fractions in a beaker with five varied properties is a narrow slice of DEM regression tasks. More importantly, there's a winner's-curse risk—if the 'optimal' model is declared by the same CV metric used to pick among 16 candidates, the margins could be noise unless there are repeated splits or nested CV. The abstract also doesn't report dataset size, feature ranges, or how much better HGBM was than the runner-up. Those are exactly what a referee should check. The mismatch in the review packet is the load-bearing issue, not the paper's own content.\n\nWho's this for? People doing DEM surrogate modeling who want a quick, defensible model-selection workflow. They'd get value from the protocol even if the specific winner doesn't transfer.\n\nMy recommendation: yes, send it to peer review if the real manuscript exists and matches the abstract. It's empirical, checkable, and addresses a real gap in the particle-tech literature. But the editor should confirm the manuscript identity first—this copy would be desk-rejected on consistency alone.\n\nReading group: maybe, if we can get the actual PDF. I wouldn't cite the result yet.","headline":"The review copy is the wrong paper—judged on the abstract alone, this is a plausible and useful benchmark, but there is no evidence to verify the central claim.","tokens_in":13821,"tokens_out":3082,"would_cite":false,"duration_ms":32450,"reading_group":"maybe","serious_thinker":"unclear","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Histogram-based gradient boosting is the best surrogate among 16 regression models for tabular DEM data, selected by k-fold cross-validation.","keywords":["DEM","surrogate models","regression","k-fold cross-validation","histogram-based gradient boosting","packing fraction","tabular data","metamodel"],"falsifier":"Run the same 16-model comparison with k-fold cross-validation on an independent DEM dataset targeting a different quantity (e.g., discharge time, stress, or mixing index) and with a different geometry; if histogram-based gradient boosting is not among the top models, the claimed optimality and framework transferability would be contradicted.","tokens_in":12885,"feed_emoji":"📊","tokens_out":2160,"duration_ms":26513,"temperature":0.7,"pith_summary":"The paper argues that surrogate model selection for discrete element method (DEM) simulations is often overlooked, leading to suboptimal predictive accuracy and slow evaluations. It proposes a simple k-fold cross-validation framework that practitioners can follow to choose the best regression model for their own tabular DEM datasets. Demonstrating the framework on a dataset of packing fractions measured in a measuring beaker with five varying particle properties, the paper tests 16 models and finds a histogram-based gradient boosting model to be optimal, offering a good fit with acceptable training and inference times. If correct, this gives DEM practitioners a reproducible protocol for picking fast, accurate surrogates for real-time process optimization.","feed_headline":"Histogram boosting tops 16 models for DEM surrogate fitting","feed_subtitle":"A k-fold cross-validation recipe makes surrogate model choice reproducible for particle simulation datasets.","key_machinery":"The key machinery is the k-fold cross-validation evaluation protocol applied uniformly across 16 regression model families, with the histogram-based gradient boosting model emerging as the best trade-off between fit quality and computational cost. The framework's transferability rests on this protocol being a reliable arbiter of model quality for tabular DEM data.","core_discovery":"The central discovery is that, on the studied DEM dataset of packing fractions, a histogram-based gradient boosting model outperforms the other 15 candidate regression models when assessed through k-fold cross-validation, balancing predictive accuracy with practical training and inference speed. The paper frames this as a concrete instance of a broader claim: that k-fold cross-validation is an appropriate and effective way to select regression surrogates for tabular DEM data, and that a simple, generalizable framework built around it can help readers identify the optimal model for their own simulation datasets.","pith_inferences":["I infer that the success of histogram-based gradient boosting on this dataset likely reflects its ability to capture interactions among the five varied particle properties without heavy tuning, but the paper does not isolate this mechanism directly.","I infer that the framework's generality would be strengthened by demonstrations on additional DEM datasets with different target quantities and geometries; the current single-dataset demonstration leaves open how far the recommendation transfers.","I infer that the protocol could be extended to classification tasks or to DEM datasets with spatial or time-series structure, though the paper only addresses tabular regression.","I infer that the optimal model choice may shift with dataset size, feature count, or the cost ratio of training vs. inference, so practitioners should treat the HGBM result as a starting point, not a universal answer."],"forward_implications":["Practitioners can use the proposed k-fold cross-validation protocol to reproducibly select surrogate regression models for their own DEM datasets, avoiding defaulting to suboptimal choices.","For tabular DEM regression tasks similar to the packing-fraction example, histogram-based gradient boosting appears to be a strong default candidate, offering good accuracy without excessive training or inference cost.","Adopting such a selection framework could make real-time evaluations of DEM-based processes more feasible by ensuring the surrogate model used is actually well-suited to the data.","The comparison of 16 models highlights that model choice materially affects predictive performance, reinforcing that model selection deserves explicit attention in particle technology research."],"supporting_citations":[],"fun_headline_variants":["K-fold CV picks histogram boosting for DEM data","Best DEM surrogate: histogram boosting wins out","16 models, one winner: histogram boosting for DEM","Cross-validation framework finds top DEM surrogate","Histogram boosting beats 15 rivals for DEM data"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The benchmark's generality rests on a single dataset of packing fractions measured in one simple beaker geometry, so the optimal model choice may not transfer to other DEM regression tasks.","fun_headline_variants_meta":{"raw":{"variants":["K-fold CV picks histogram boosting for DEM data","Best DEM surrogate: histogram boosting wins out","16 models, one winner: histogram boosting for DEM","Cross-validation framework finds top DEM surrogate","Histogram boosting beats 15 rivals for DEM data"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000497,"raw_usage":{"total_tokens":2233,"prompt_tokens":662,"completion_tokens":1571,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":406,"completion_tokens_details":{"reasoning_tokens":1501}},"tokens_in":406,"tokens_out":1571,"duration_ms":14090,"temperature":1.0,"reasoning_tokens":1501,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T23:25:42.572079+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same 16-model comparison with k-fold cross-validation on an independent DEM dataset targeting a different quantity (e.g., discharge time, stress, or mixing index) and with a different geometry; if histogram-based gradient boosting is not among the top models, the claimed optimality and framework transferability would be contradicted.","supporting_citations":[],"review_version":1}