{"id":"f50940cb-79bc-4710-a7f2-270b929d7d90","arxiv_id":"2605.05234","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Benchmark study finds quantile and z-score marking strategies most robust for adaptive mesh refinement in steady mechanics problems, with Dörfler effective at large parameters and Isolation Forest competitive only under generous settings.","lead":"This paper benchmarks marking strategies for deciding which elements to refine in adaptive mesh refinement for finite element simulations of steady solid and fluid problems. It identifies quantile and z-score approaches as the most robust options for balancing refinement quality against computational cost.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.3","headline":"Robustness rankings hinge on untested generalizability of the specific test problems and Kelly estimator","rationale":"The reader's weakest assumption correctly isolates the single load-bearing risk for an empirical benchmark study. Because the work is purely numerical and offers no analytic proof of superiority, the validity of the practical recommendations rests entirely on whether the test corpus is representative; the proposed concrete_test directly probes that assumption.","tokens_in":1651,"tokens_out":300,"duration_ms":19895,"concrete_test":"Re-run the full benchmark suite on two additional problems outside the original set (one transient Navier-Stokes case and one 3D linear elasticity case with a goal-oriented estimator) using identical marking implementations and metrics; if the top-two strategies change or the sensitivity ordering of maximum marking reverses, the headline robustness conclusions do not generalize.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim ranks marking strategies (quantile/z-score most robust, Dörfler effective at large bulk, maximum sensitive, Isolation Forest conditional on contamination) based solely on steady solid/fluid mechanics problems driven by the residual-based Kelly estimator. For these rankings to support the stated practical guidance, the chosen problems must adequately sample the variability of real engineering workloads (different geometries, singularities, nonlinearities, and error-indicator behaviors). No evidence is provided that the results are insensitive to problem selection or estimator choice; an atypical test suite could invert the observed ordering without contradicting the reported data.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript presents a benchmark study comparing classical marking strategies (maximum, Dörfler bulk-chasing, quantile) with statistical approaches (z-score, Isolation Forest) for adaptive mesh refinement in finite element methods. All strategies are driven by the residual-based Kelly error estimator and tested on steady solid and fluid mechanics problems. The central claim is that quantile and z-score markings are the most robust, Dörfler is effective for large bulk parameters, maximum marking is sensitive to irregular fields, and Isolation Forest can rival the top classical methods when the contamination level is set generously but may fail under aggressive settings. These findings are positioned as practical guidance for balancing refinement aggressiveness and computational cost in AMR workflows.","tokens_in":1759,"tokens_out":567,"duration_ms":47093,"significance":"If the reported performance orderings hold, the work supplies useful empirical data on marking strategy selection for AMR, extending prior comparisons by including non-classical statistical methods. The efficiency-focused framing and use of a standard residual estimator are strengths for practitioners in solid and fluid mechanics FEM. However, the significance is constrained by the narrow problem class and the qualitative presentation of robustness rankings, limiting immediate applicability to broader engineering workloads.","major_comments":[{"comment":"The robustness rankings (quantile/z-score most robust; maximum sensitive to irregular fields) are load-bearing for the practical guidance claim yet rest on qualitative summaries of performance across the chosen test problems. No quantitative metrics such as error-vs-DOF curves with error bars, wall-clock timings, or mesh statistics are referenced in the abstract or results, and no statistical tests of variability are reported, making it impossible to assess the magnitude or significance of differences between strategies.","section":"Abstract and Results"},{"comment":"The claim that the observed orderings provide general practical guidance assumes the selected steady solid- and fluid-mechanics problems together with the Kelly estimator adequately sample real engineering variability. No sensitivity study to problem selection, geometry, singularities, nonlinearity, or alternative error indicators is presented; an atypical test suite could invert the rankings without contradicting the reported data.","section":"Discussion and Conclusions"}],"minor_comments":[{"comment":"The methods section should explicitly tabulate the exact parameter values used for each strategy (e.g., bulk parameter range for Dörfler, contamination levels for Isolation Forest, z-score threshold) so that the experiments are fully reproducible.","section":"Methods"},{"comment":"Figure captions and legends would benefit from clearer indication of which curves correspond to which marking strategy and parameter setting, especially when multiple contamination levels are shown for Isolation Forest.","section":"Figures"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive feedback on our benchmark study of marking strategies for adaptive mesh refinement. We address each major comment below, proposing targeted revisions to improve clarity and balance while preserving the manuscript's focus and scope.","responses":[{"response":"We acknowledge that the abstract and results summaries are primarily descriptive. The manuscript does contain error-versus-DOF curves, mesh statistics, and comparative performance data in the results section and figures, but these are presented visually without explicit numerical call-outs or variability measures in the text. To strengthen the presentation, we will revise the abstract to reference key quantitative trends (such as typical DOF counts at target error levels) and enhance the results section with additional quantitative summaries, error bars on relevant plots where feasible, and notes on the magnitude of observed differences. Wall-clock timings for the marking step itself can be added as supplementary data, though the primary efficiency metric remains error reduction per degree of freedom. Full statistical hypothesis testing across all problems would require additional analysis but can be noted as a limitation.","revision_made":"partial","referee_comment":"[Abstract and Results] The robustness rankings (quantile/z-score most robust; maximum sensitive to irregular fields) are load-bearing for the practical guidance claim yet rest on qualitative summaries of performance across the chosen test problems. No quantitative metrics such as error-vs-DOF curves with error bars, wall-clock timings, or mesh statistics are referenced in the abstract or results, and no statistical tests of variability are reported, making it impossible to assess the magnitude or significance of differences between strategies."},{"response":"We agree that the reported orderings are tied to the specific steady problems and residual-based Kelly estimator used. The manuscript frames the work as a focused benchmark study rather than a universal claim, but we will strengthen the discussion and conclusions by explicitly qualifying the practical guidance as applicable to the tested class of problems. A dedicated limitations section will be added to note that rankings could vary with different geometries, singularities, nonlinearities, or error indicators, and to recommend case-specific validation. A comprehensive sensitivity study across all such variations lies beyond the scope of this efficiency-focused comparison of marking strategies.","revision_made":"partial","referee_comment":"[Discussion and Conclusions] The claim that the observed orderings provide general practical guidance assumes the selected steady solid- and fluid-mechanics problems together with the Kelly estimator adequately sample real engineering variability. No sensitivity study to problem selection, geometry, singularities, nonlinearity, or alternative error indicators is presented; an atypical test suite could invert the rankings without contradicting the reported data."}],"tokens_in":1379,"tokens_out":547,"duration_ms":30233,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main thing to know is that this paper compares maximum, Dörfler, quantile, z-score, and Isolation Forest marking strategies, all driven by the Kelly residual estimator, on a set of steady solid and fluid mechanics benchmarks. It reports that quantile and z-score are the most robust across the tests, Dörfler performs well at large bulk values, maximum marking is sensitive to irregular fields, and Isolation Forest can match the top performers only with a generous contamination parameter.","headline":"The paper runs a clean empirical comparison of classical and statistical marking strategies for AMR on steady mechanics problems and gives usable practical advice, though the rankings rest on a narrow test suite.","tokens_in":2225,"tokens_out":177,"would_cite":false,"duration_ms":19397,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Quantile and z-score marking strategies are the most robust for adaptive mesh refinement in steady solid and fluid mechanics problems.","keywords":["adaptive mesh refinement","marking strategies","finite element method","error estimation","quantile","z-score","Isolation Forest","solid and fluid mechanics"],"falsifier":"Repeating the benchmark study using transient problems or a different error estimator such as the Zienkiewicz-Zhu method and finding that the robustness order of the marking strategies changes would falsify the generalizability of these results.","tokens_in":2556,"feed_emoji":"📊","tokens_out":679,"duration_ms":36203,"temperature":0.7,"pith_summary":"This paper benchmarks classical and statistical marking strategies for adaptive mesh refinement in finite element models of steady solid and fluid mechanics problems. Driven by the residual-based Kelly error estimator, the study identifies quantile and z-score markings as the most robust options overall. Dörfler marking performs effectively with large bulk parameters, whereas maximum marking is sensitive to irregular fields. Isolation Forest can compete with the best methods if its contamination level is set high enough, but it risks underperforming with more aggressive settings. This offers practical advice for engineers seeking to improve the efficiency of their adaptive simulations by selecting appropriate marking approaches.","feed_headline":"Quantile and z-score marking strategies lead AMR benchmarks","feed_subtitle":"Benchmark study on solid and fluid mechanics problems identifies them as most robust for balancing refinement and computational cost.","key_machinery":"Marking strategies applied to elements selected by the residual-based Kelly error estimator, including maximum, Dörfler, quantile, z-score, and Isolation Forest methods.","core_discovery":"The study benchmarks marking strategies for adaptive mesh refinement driven by the Kelly residual estimator on steady mechanics problems. It concludes that quantile and z-score markings are the most robust, Dörfler marking is effective when using large bulk parameters, maximum marking is sensitive to irregular fields, and Isolation Forest can match top performers only with generous contamination settings but risks failure under aggressive parameters.","pith_inferences":["The preference for quantile and z-score methods might apply to time-dependent or multiphysics simulations if the error estimator behaves similarly.","Alternative error estimators could alter which marking strategy ranks highest, suggesting the need for problem-specific validation.","In very large-scale computations, the efficiency gains from robust marking could translate to significant savings in time and resources.","Hybrid strategies that combine quantile marking with elements of Isolation Forest might offer further improvements."],"forward_implications":["Quantile and z-score markings maintain consistent performance across different problem types and field characteristics.","Dörfler marking achieves good results when the bulk parameter is chosen sufficiently large.","Maximum marking can lead to overly sensitive or irregular refinement patterns in solutions with varying smoothness.","Isolation Forest requires a sufficiently high contamination parameter to match classical methods but becomes unreliable if set too aggressively.","Selecting robust marking strategies can help reduce overall computational costs in adaptive finite element workflows."],"fun_headline_variants":["Quantile z-score markings robust in AMR benchmarks","Dörfler marking effective with large bulk params","Maximum marking sensitive to irregular AMR fields","Isolation Forest rivals with generous contamination"],"cache_read_input_tokens":64,"weakest_assumption_plain":"The steady solid and fluid mechanics test problems paired with the residual-based Kelly error estimator are representative of broader engineering applications.","fun_headline_variants_meta":{"raw":{"variants":["Quantile z-score markings robust in AMR benchmarks","Dörfler marking effective with large bulk params","Maximum marking sensitive to irregular AMR fields","Isolation Forest rivals with generous contamination"]},"model":"grok-4.3","cost_usd":0.006477,"raw_usage":{"total_tokens":2913,"prompt_tokens":591,"num_sources_used":0,"completion_tokens":53,"cost_in_usd_ticks":64765500,"prompt_tokens_details":{"text_tokens":591,"audio_tokens":0,"image_tokens":0,"cached_tokens":64},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":2269,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":591,"tokens_out":53,"duration_ms":18673,"temperature":1.0,"reasoning_tokens":2269,"cache_read_input_tokens":64,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-05-09T20:30:09.561468+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Repeating the benchmark study using transient problems or a different error estimator such as the Zienkiewicz-Zhu method and finding that the robustness order of the marking strategies changes would falsify the generalizability of these results.","supporting_citations":[],"review_version":1}