{"id":"47f22366-28c1-4bfa-b938-6a6b3ab2f82a","arxiv_id":"2605.30278","paper_version":2,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":4.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"modelimportance is an R package that quantifies the contribution of individual models to ensemble forecast accuracy for point and probabilistic predictions, supporting multiple methods, metrics, missing-data handling, and hubverse integration.","lead":"This paper describes an R package called modelimportance that calculates how much each model adds to the accuracy of a multi-model ensemble forecast. Researchers who combine forecasts for decisions in fields like public health or policy may use it to inspect which models matter most.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"No significant objection identified","rationale":"The reader's weakest_assumption targets metric reliability, but that is an open question for downstream users rather than a premise the paper itself relies upon. The paper's argument is descriptive and holds if the package exists and performs the advertised computations; the absence of validation studies is therefore not a load-bearing concern for the claim as written.","tokens_in":1684,"tokens_out":262,"duration_ms":14568,"concrete_test":"Install the package (CRAN/GitHub) and execute its core functions on the vignette example data; confirm that the returned importance scores match the documented output format and that no runtime errors occur for both point and probabilistic cases.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The manuscript is a software description paper whose central claim is that the modelimportance package supplies implemented tools for quantifying component-model contributions to ensemble accuracy (point and probabilistic) under multiple methods and within the hubverse framework. No novel theoretical result, empirical performance claim, or assertion that the chosen metrics are optimal or guaranteed to be reliable is advanced; the text limits itself to describing supported functionality, missing-value handling, and integration. Consequently there is no load-bearing assumption whose falsity would invalidate the paper's stated purpose.","agreement_with_reader":"disagree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript describes the R package modelimportance, which supplies implemented tools to quantify the contribution of each component model to ensemble accuracy for both point and probabilistic forecasts. The package supports multiple ensemble methods and model-importance metrics, includes options for handling missing values, and integrates with the hubverse framework for collaborative modeling.","tokens_in":1743,"tokens_out":255,"duration_ms":14743,"significance":"If the described functionality is correctly implemented, the package would serve as a practical, reusable tool for researchers analyzing ensembles in forecasting applications. Its hubverse compatibility could streamline workflows in collaborative hubs and support post-hoc interpretation of model value without requiring users to code custom importance calculations.","major_comments":[],"minor_comments":[{"comment":"The manuscript would benefit from a short table or enumerated list in the main text that explicitly names the supported ensemble methods and importance metrics (currently only alluded to in the abstract).","section":null},{"comment":"A minimal worked example (code snippet plus output) illustrating a typical call and interpretation of results would clarify usage for readers who have not yet installed the package.","section":null}],"recommendation":"minor_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for their supportive review of the modelimportance package manuscript and for recommending minor revision. No major comments were provided in the report.","responses":[],"tokens_in":1156,"tokens_out":48,"duration_ms":7933,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"This paper describes the modelimportance R package, which gives users functions to measure the contribution of each model to an ensemble's performance on both point and probabilistic forecasts. It supports several ensemble methods and importance metrics, plus options for missing values, all within the hubverse setup.\n\nThe new part is the specific implementation and its seamless tie-in to hubverse tools, which could make it easy for people already working in those collaborative forecasting projects to add importance calculations to their workflow. The package itself looks like a useful utility for practitioners who want to understand which models are driving the ensemble results.\n\nThe main shortcoming is the lack of any empirical check. The manuscript explains what the package can do but does not show results from applying the metrics to real or simulated data, nor does it compare the different metrics to see if they agree or which one is more reliable. Without that, it's difficult to know if the outputs are stable or interpretable in the way the abstract suggests.\n\nThe citation pattern is fine for a software paper, mostly pointing to the hubverse and related forecasting work.\n\nThis is for R users in statistical forecasting who need a ready-made tool rather than a new theoretical approach. A reader interested in methodological innovation or evidence on ensemble construction will not find it here.\n\nI think it deserves peer review in a venue that publishes software descriptions, as the implementation could be worth verifying for bugs or usability issues, even though the paper does not make strong scientific claims.","headline":"The paper describes an R package for computing model contributions in ensembles but supplies no validation or examples of the metrics.","tokens_in":2211,"tokens_out":366,"would_cite":false,"duration_ms":18544,"reading_group":"no","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"The modelimportance R package quantifies each model's contribution to ensemble forecast accuracy.","keywords":["ensemble forecasting","model importance","R package","forecast accuracy","point forecasts","probabilistic forecasts","multi-model ensembles"],"falsifier":"A test in which models the package ranks as low importance are removed from the ensemble and the resulting accuracy does not improve or declines.","tokens_in":2577,"feed_emoji":"📊","tokens_out":544,"duration_ms":21315,"temperature":0.7,"pith_summary":"The paper presents an R package that measures how much each individual model adds to the accuracy of a combined forecast made from multiple models. Ensembles are widely used because they tend to be more accurate and stable than any single model, yet without ways to score each component's value it is difficult to decide which models to retain or how to refine the group. The package computes importance scores for both point forecasts and probabilistic forecasts, supports several methods for combining models, and includes several different importance metrics. It also provides options to manage cases where some models have no prediction for a given time point. These functions let users build stronger ensembles while learning which models drive the gains.","feed_headline":"R package scores each model's value in forecast ensembles","feed_subtitle":"Metrics for point and probability predictions help identify which models boost combined accuracy.","key_machinery":"The modelimportance package, which applies chosen importance metrics to score each model's effect on ensemble accuracy.","core_discovery":"The modelimportance package quantifies how each component model contributes to the accuracy of ensemble performance for both point and probabilistic forecasts. The package supports multiple ensemble methods and multiple model importance metrics, and it supplies customizable options for handling missing values.","pith_inferences":["Importance scores could be used to assign different weights to models when forming new ensembles.","The same metrics might be applied to combined systems outside forecasting, such as stacked predictors in other domains.","Direct comparisons of the package's different metrics on the same data could show which one best matches actual accuracy gains."],"forward_implications":["Users can identify which models contribute most to ensemble accuracy.","The tools support construction of more effective ensembles across forecasting tasks.","Analysis yields insights into the role played by each model inside the ensemble.","Options for missing values allow the methods to work with incomplete real data."],"fun_headline_variants":["R package quantifies model importance in forecast ensembles","Evaluate each model's role in ensemble forecast accuracy","R package assesses model value for forecast ensembles","Score model contributions to multi-model forecast performance"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The chosen importance metrics and ensemble methods produce reliable and interpretable measures of each model's value.","fun_headline_variants_meta":{"raw":{"variants":["R package quantifies model importance in forecast ensembles","Evaluate each model's role in ensemble forecast accuracy","R package assesses model value for forecast ensembles","Score model contributions to multi-model forecast performance"]},"model":"grok-4.3","cost_usd":0.007016,"raw_usage":{"total_tokens":3217,"prompt_tokens":606,"num_sources_used":0,"completion_tokens":54,"cost_in_usd_ticks":70162000,"prompt_tokens_details":{"text_tokens":606,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":2557,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":606,"tokens_out":54,"duration_ms":16784,"temperature":1.0,"reasoning_tokens":2557,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-28T23:25:25.398542+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A test in which models the package ranks as low importance are removed from the ensemble and the resulting accuracy does not improve or declines.","supporting_citations":[],"review_version":1}