{"id":"15dcc676-cfe5-425a-96a3-0d7029588fb0","arxiv_id":"2608.08069","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"An explainable multi-criteria system ranks 71,274 Hugging Face models using metadata, functional features, and community-perceived quality, matching commercial LLM recommenders on coverage while exposing criterion-level scores.","lead":"HugSelect is a decision-support system that scores and ranks AI foundation models on Hugging Face using explicit criteria, such as task, license, and community-reported quality. It recommends models with a transparent breakdown of why each one ranks where it does, offering an auditable alternative to asking a chatbot.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central competitive claim depends on the 44 literature-derived proxy labels and on the functional-feature pipeline; with only point estimates at n=44 and no confidence intervals, the headline 'comparable' differences (0.61 vs 0.64, 0.91 vs 0.93) are within ordinary sampling noise.","rationale":"The reader's weakest_assumption — that the 44 literature-derived cases are a valid proxy for recommendation quality — is exactly the load-bearing concern I identify. The argument for comparability with commercial LLM recommenders depends on these labels being representative and unbiased; the authors themselves acknowledge this in Section 5.6. My stress-test adds two concrete technical points not fully developed in the reader's assessment: (a) the absence of confidence intervals for the headline metrics, which matters at n=44; and (b) the ablation results showing extreme sensitivity to the functional-feature pipeline, which was validated with LLM-assisted ground truth, so the pipeline's quality is not fully independent of the very kind of LLM reasoning the baselines use. The paper is generally honest about limitations, provides a replication package, and the internal logic of the WSM ranking is sound. The concern is therefore not fatal but does justify a CONDITIONAL verdict: the central comparative claim is plausible but not yet statistically pinned down relative to the proxy labels and sampling noise. My recommendation is CONDITIONAL, not REJECT, because the reported evidence is substantial and the identified gaps are addressable with additional analysis of the existing replication data.","tokens_in":26626,"tokens_out":1855,"duration_ms":21267,"concrete_test":"Use the published replication package to recompute the 44-case Comparison with two changes: (1) report bootstrap 95% confidence intervals for all Coverage@10 and NDCG@10 values in Table 11; if the intervals for HugSelect and Claude overlap substantially, the claim of 'comparable' should be softened to 'not distinguishable at this sample size.' (2) Independently re-derive the literature ground truth for a random subset of 10 cases by having two new annotators, blind to HugSelect outputs, select the model they would choose from the same source papers; measure inter-annotator agreement and recompute Coverage@10 against the original labels and against each annotator's labels. If agreement is low or the revised Coverage@10 drops materially relative to commercial baselines, the proxy-label assumption is not secure.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim is that HugSelect achieves recommendation quality comparable to commercial LLM recommenders, as measured by Coverage@10 and NDCG@10 on 44 literature-derived scenarios. The most load-bearing condition is that those proxy labels actually measure selection quality. The authors acknowledge in Sections 5.3.1 and 5.6 that literature-reported model choices are proxy ground truth and that 'other suitable alternatives may exist'; they mitigate this with family-level metrics. However, the strength of the entire comparative claim depends on the quality and representativeness of these 44 labels. If the labels are biased toward popular or conveniently documented models, then high coverage may partly reflect popularity signals that are already embedded in the knowledge base, rather than genuine multi-criteria selection quality. This concern is reinforced by the ablation in Table 13: removing functional features drops model-level Coverage@10 from 0.61 to 0.23 (p<0.001), while removing all functional and quality features leaves only 0.09. That shows the ranking is heavily driven by the functional-feature pipeline, whose validation ground truth was itself constructed with LLM assistance (Section 5.2.1). Although the human-in-the-loop protocol reduces circularity, the functional features are still extracted from the same model-card text that the commercial LLM baselines also see, so the comparison may partially reflect different ways of reading the same documentation rather than independent evidence of model quality. Additionally, Table 11 reports only point estimates with no confidence intervals; with n=44, a difference between 0.61 and 0.64 is well within sampling variability, and the omnibus Friedman test is borderline (p=0.0667). The paper's own threats-to-validity section is honest about these limits, but the headline abstract claims 'comparable' without quantifying the uncertainty.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes HugSelect, an explainable multi-criteria decision-support framework for selecting foundation models from Hugging Face. It builds a knowledge base of 71,274 models by combining repository metadata, functional features extracted from model cards, and quality attributes derived from community feedback via LLM-based mapping. Rankings are produced with a weighted sum model with criterion-level score decomposition. The evaluation includes extraction pipeline validation, comparison with four commercial LLM-based recommenders on 44 literature-derived scenarios, ablation, and a 10-participant user study. The central claim is that HugSelect achieves recommendation quality comparable to commercial LLM recommenders while providing transparent reasoning.","tokens_in":26882,"tokens_out":6795,"duration_ms":63299,"significance":"If the results hold, the paper demonstrates that a deterministic, explainable system with an automatically constructed knowledge base can be competitive with LLM chat recommenders on a realistic selection-retrieval task. The strengths include a publicly available replication package, a repository-scale knowledge base, an external literature-based evaluation benchmark, and an honest and explicit treatment of threats to validity. The functional-feature pipeline validation uses a human-in-the-loop protocol that partially avoids circularity, and the ablation study provides clear evidence of the contribution of functional features. However, the quality-attribute mapping validation is partially circular, and the comparative evaluation at n=44 lacks confidence intervals, so the strength of the central claim is limited.","major_comments":[{"comment":"The reported 0.84 accuracy for quality-attribute mapping is validated against a reference set formed by three LLMs (labels appearing in at least two outputs, refined by a human annotator), while the mapping pipeline itself uses LLM prompting. This is a circular evaluation for the mapping step: the pipeline is measured against ground truth generated by the same kind of model it uses. The abstract and Section 5.2.2 should either report a human-only validation or explicitly qualify the 0.84 figure as an agreement rate among LLMs, not as accuracy against independent human labels.","section":"Section 5.2.2, Table 7"},{"comment":"The comparative claim of 'no significant overall differences' rests on 44 cases with point estimates and no confidence intervals. The observed differences (model Coverage@10 0.61 vs. Claude's 0.64; family-level 0.91 vs. 0.93) are well within sampling noise, and the only significant McNemar comparison is against Gemini. The manuscript should report confidence intervals for Coverage@10 and NDCG@10, and rephrase 'comparable' as 'no statistically significant difference was detected', with an explicit discussion of the low statistical power at n=44. Without this, the reader cannot gauge the strength of the evidence for equivalence.","section":"Section 5.3, Table 11"},{"comment":"The 44 literature-derived proxy labels are a load-bearing condition for the entire comparative evaluation. The paper acknowledges that 'other suitable alternatives may exist' (Section 5.6), but it does not examine whether the ground-truth models are systematically biased toward popular or well-documented models, which would inflate Coverage@10 for any system that reads model-card text. I recommend adding an analysis of the popularity and documentation-quality distribution of the ground-truth models, and a sensitivity check of Coverage@10 when the top-k most popular models are removed from the evaluation.","section":"Section 5.3.1, Ground Truth Construction"},{"comment":"The framework's sensitivity to the MoSCoW criterion weights is a core design feature but is never evaluated. Section 3.6 mentions a sensitivity analysis under ±20% weight perturbations, yet no results of such an analysis appear in Section 5. The abstract's claim of 'stable' reasoning is thus unsupported. Either report the sensitivity results or explicitly defer them to future work in the evaluation section.","section":"Section 3.6 and Section 5.5"}],"minor_comments":[{"comment":"The HugSelect row appears malformed ('0.61-0.91-0.37 0.74') and does not align with the six-column header; please format the row to match the header, using an em dash or 'N/A' for Overlap@10.","section":"Table 11"},{"comment":"Duplicate word in the text: 'perceived perceived quality attributes'.","section":"Section 2.2"},{"comment":"The two paragraphs describing functional-feature ground-truth construction are inconsistent: the first says features appearing in at least two LLM outputs were retained and validated by a human expert, while the second says a feature was included only when identified by the human annotator and confirmed by at least one LLM. Please clarify which protocol was actually used.","section":"Section 5.2.1"},{"comment":"Several sentences start with lowercase 'foundation-models' (e.g., the first sentence of Section 1 and Section 2.2); please fix capitalization at sentence starts.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The paper is unusually transparent about its limitations, which is commendable, but the headline claims in the abstract ('0.84 accuracy', 'comparable') are not fully qualified by those limitations. The circularity in the quality-mapping validation is the most serious issue; a re-validation with human-only labels or a reworded claim is needed. The statistical reporting at n=44 should be strengthened with confidence intervals and effect sizes. If these points are addressed, the paper could be a solid contribution to the software-engineering literature on model selection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Thanks for the pass. I read HugSelect carefully, and I think it is more solid than the stress-test headline suggests. The real contribution is a running, reproducible system: a 71k-model knowledge base combining metadata, functional features extracted from model cards, and community-derived quality attributes, all feeding a WSM/SAW ranking that decomposes scores for inspection. The authors ship code and data on Mendeley, which is more than most of this literature does. The 44-case evaluation against four commercial LLM recommenders uses literature-reported model choices as external ground truth, a genuine attempt to avoid circularity, and the paper is honest about the proxy nature of those labels. The human-in-the-loop validation of the extraction pipelines also softens the circularity concern: ground truth is not pure LLM agreement.\n\nThe soft spots are real but manageable. First, the headline claim of comparability rests on n=44 point estimates with no confidence intervals. A difference between 0.61 and 0.64 is within sampling noise, and the omnibus Friedman test is borderline (p=0.0667). The abstract should either report uncertainty or soften the wording. Second, the metadata/filtering baseline used in the McNemar tests is never defined, so the reader cannot tell what the comparison actually removes. Third, the explainability claim is slightly oversold: the arithmetic is transparent, but the criterion weights come from an LLM-based intent extractor, so the audit trail stops at the weighting step. The user study is n=10 and explicitly exploratory, which the paper acknowledges, so I do not count that as a flaw.\n\nThe circularity concern in the stress test is overstated. The recommendation ground truth is literature-derived, not LLM-generated, and the family-level metrics blunt the label-bias problem. The quality-attribute ground truth does use LLM labels, but the human-in-the-loop protocol and the explicit threat-to-validity discussion make that a moderate caveat, not a load-bearing flaw.\n\nWho should read this: researchers working on model selection, model hubs, or decision support for ML components, and practitioners who need an auditable selection process. It will not change how models are built, but it is a solid, honest systems paper. I would send it to serious peer review, with a request for confidence intervals, a definition of the metadata-only baseline, and a toned-down abstract.","headline":"An honest, reproducible systems paper whose 'comparable to commercial LLMs' claim rests on small-sample point estimates, but whose integrated pipeline is a real and useful contribution.","tokens_in":27555,"tokens_out":2805,"would_cite":true,"duration_ms":29657,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that foundation-model selection can be treated as an auditable multi-criteria decision problem, and that a deterministic weighted-sum recommender over an automatically built knowledge base matches commercial LLM-based…","keywords":["foundation-model selection","multi-criteria decision-making","explainable recommender","weighted sum model","Hugging Face","knowledge base","software component selection","perceived quality attributes"],"falsifier":"Build a ground-truth set from benchmark leaderboards and deployment logs instead of literature-reported choices, run the same or an expanded set of scenarios, and check whether model-level Coverage@10 stays near 0.61; if it falls toward the metadata-only ablation level of 0.09 or the no-functional-features level of 0.23, the literature ground truth was carrying the result.","tokens_in":26448,"feed_emoji":"🤖","tokens_out":7167,"duration_ms":85691,"temperature":0.7,"pith_summary":"This paper argues that choosing a foundation model for a software project is a software-component decision, not a search or popularity problem, and that the choice should be auditable. It presents HugSelect, which combines Hugging Face metadata, capability phrases extracted from model cards, and community-discussion quality signals into a knowledge base of 71,274 models, then ranks candidates with a weighted-sum model whose per-criterion scores can be inspected. Across 44 realistic selection scenarios, HugSelect's top-10 lists contained the literature-reported model in 61% of cases at the model level and 91% at the family level, statistically comparable to ChatGPT, Claude, Gemini, and Perplexity while remaining deterministic and traceable. The paper also reports an ablation showing that functional features are the main driver of retrieval, and a small user study in which ten practitioners rated the system useful and easy to use.","feed_headline":"Explainable ranker ties LLM chatbots at model retrieval","feed_subtitle":"A deterministic weighted-sum recommender reaches family-level coverage of 0.91 and explains every score.","key_machinery":"The engine is a Weighted Sum Model (WSM/SAW): each candidate model $m_i$ receives a score $R(m_i,q)=\\sum_j w_j \\cdot s_{ij}$, where $w_j$ are user-derived criterion weights that sum to 1 and $s_{ij}\\in[0,1]$ are normalized criterion scores. Perceived quality scores are aggregated from community review snippets as $S_{\\text{fuzzy}}=H/(L+H)$, the ratio of positive over directional reviews, and a quality attribute is scored only when at least three evidence-bearing reviews exist. A graph-structured knowledge base stores models, features, quality attributes, and source snippets, but the ranking logic itself is the additive model, whose per-term products $w_j\\cdot s_{ij}$ are what make the recommendations inspectable and tunable.","core_discovery":"The central claim is that repository-scale foundation-model selection can be operationalized as an explicit multi-criteria decision problem without sacrificing recommendation quality. HugSelect builds a unified knowledge base from three evidence families: structured repository metadata, functional capabilities extracted from model-card text, and ISO/IEC 25010-inspired perceived quality attributes derived from community reviews. It ranks candidates with a Weighted Sum Model (WSM/SAW), computing $R(m_i,q)=\\sum_j w_j \\cdot s_{ij}$ and decomposing each score into criterion-level contributions for inspection. The evaluation shows model-level Coverage@10 of 0.61 and family-level Coverage@10 of 0.91, with no significant overall difference in ranking quality against four commercial zero-shot LLM recommenders, and an ablation shows that removing functional features drops model-level coverage from 0.61 to 0.23 while removing both functional and quality features drops it to 0.09. The value is not that HugSelect beats chatbots, but that competitive retrieval comes with score decomposition, feature-level traceability, and comparative trade-off views that opaque conversational systems do not provide.","pith_inferences":["An implication the authors leave implicit is that the same WSM/SAW shell could port to other model hubs or private registries, since the framework is repository-agnostic in principle; the open question is whether the extraction pipelines transfer without retraining.","The measured parity with LLM chatbots is time-bound: the baselines were zero-shot with no prompt optimization, so a conversational recommender augmented with retrieval or fine-tuning could plausibly close or reverse the gap.","A natural testable extension is to replace the fuzzy sentiment ratio with benchmark-based quality evidence and see whether family-level discrimination improves; the paper lists richer evidence as future work, and the modular design makes this straightforward.","The jump from 0.61 model-level coverage to 0.91 family-level coverage reflects the paper's choice to treat family members as substitutes; a user who needs an exact checkpoint would see lower effective performance, which suggests a UI that separates exact matches from family matches."],"forward_implications":["A deterministic, auditable recommender can substitute for opaque conversational advice in foundation-model selection without a measurable drop in top-10 retrieval quality.","Functional capability extraction from model cards is the main driver of exact-model retrieval, so teams building such systems should invest in model-card and README parsing before community sentiment.","Family-level coverage near 0.91 means the system is particularly useful for surfacing substitute variants such as base, quantized, and fine-tuned checkpoints rather than a single canonical model.","Removing both functional and quality features drops model-level coverage to 0.09, showing that metadata-only or popularity-based ranking is far weaker than multi-source evidence.","Because weights are explicit and adjustable, practitioners can audit and adapt rankings to project constraints such as license, format, and reliability without re-prompting a black-box system."],"supporting_citations":[{"why":"Supplies the WSM/SAW additive ranking model that HugSelect uses as its decision engine.","marker":"Triantaphyllou (2000)"},{"why":"Provides the multi-layer multi-criteria decision-making framework architecture that HugSelect adapts.","marker":"Farshidi (2020)"},{"why":"Supplies the dependency-parsing noun-phrase extraction method used in the functional-feature pipeline.","marker":"Honnibal and Montani (2017)"},{"why":"The Hub API that supplies the 71,274-model metadata collection.","marker":"Hugging Face (2023)"},{"why":"The closest large-scale Hugging Face knowledge-graph work, used as a positioning and comparison baseline.","marker":"Chen et al. (2025)"},{"why":"Supplies the Technology Acceptance Model constructs used in the user study.","marker":"Davis (1989)"},{"why":"Supplies the TAM2 external variables used in the user study.","marker":"Venkatesh and Davis (2000)"},{"why":"The dataset and replication package that back the evaluation results and reproducibility claims.","marker":"Adalı et al. (2026)"}],"fun_headline_variants":["HugSelect matches LLM chatbots with explainable ranking","Deterministic model picker ties AI recommenders, explains every score","Transparent scoring matches LLM bots with 91% family coverage","Explainable multi-criteria ranking ties commercial AI recommenders","HugSelect: auditable model selection with chatbot-level recall"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the model or model family chosen in a peer-reviewed paper, verified through its linked GitHub repository, is a valid proxy for what a good recommendation should retrieve; if those literature-derived labels are biased toward popular or convenient models, the Coverage@10 and NDCG@10 comparisons against commercial LLMs measure agreement with proxy labels rather than true selection quality.","fun_headline_variants_meta":{"raw":{"variants":["HugSelect matches LLM chatbots with explainable ranking","Deterministic model picker ties AI recommenders, explains every score","Transparent scoring matches LLM bots with 91% family coverage","Explainable multi-criteria ranking ties commercial AI recommenders","HugSelect: auditable model selection with chatbot-level recall"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000704,"raw_usage":{"total_tokens":3224,"prompt_tokens":1041,"completion_tokens":2183,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":657,"completion_tokens_details":{"reasoning_tokens":2096}},"tokens_in":657,"tokens_out":2183,"duration_ms":16524,"temperature":1.0,"reasoning_tokens":2096,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T00:28:05.744748+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Build a ground-truth set from benchmark leaderboards and deployment logs instead of literature-reported choices, run the same or an expanded set of scenarios, and check whether model-level Coverage@10 stays near 0.61; if it falls toward the metadata-only ablation level of 0.09 or the no-functional-features level of 0.23, the literature ground truth was carrying the result.","supporting_citations":[],"review_version":1}