{"id":"89d1d2d7-4789-45d8-b1ad-e41a94839082","arxiv_id":"2607.27421","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"Across eight zero-shot intent datasets, instruction-tuned ~3B open-weight models can match or beat larger base models, top systems are statistically tied on MASSIVE, and SNIPS is saturated.","lead":"A head-to-head zero-shot test of 41 open-weight language models (135M–9B) on intent classification finds that a tuned 3B model can beat several 7B base models, while classic benchmarks like SNIPS no longer separate today’s models. Practitioners picking deployable chat models for voice and e-commerce bots get concrete accuracy, latency, robustness, and calibration tradeoffs.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.5","headline":"Finding-2’s IT-over-scale claim confounds format compliance with intent discrimination under free-text exact-match scoring.","rationale":"The reader’s strongest claim is the protocol-bound accuracy package (IT vs scale, SNIPS saturation, MASSIVE McNemar non-significance). That package is internally coherent as measured: Table 2 and the six-pair statement are clear, SNIPS 32/41 >80% is decisive, and the ten MASSIVE McNemar p-values are reported. The reader’s weakest assumption (MASSIVE free-text logprobs as ECE/Brier proxy) correctly flags a real limitation but attaches to the secondary calibration finding (Finding-8), not to the load-bearing accuracy/IT claim. The softer joint in the strongest claim is the free-text exact-match protocol without invalid-rate reporting: it can inflate instruct-over-base gaps via format following, which the paper never quantifies. That does not justify REJECT—the work is explicitly a deployment-oriented selection study under free-text vLLM settings, exclusions are transparent, and weighting sensitivity (Spearman ρ=0.995) is checked—but it keeps the right posture at CONDITIONAL pending code/data release and either invalid-rate disclosure or a constrained-decoding control. No change to the reader’s verdict label; only a sharper primary soft spot than calibration alone.","tokens_in":21061,"tokens_out":682,"duration_ms":76855,"concrete_test":"For the six matched base–instruct pairs, report EMPTY_PRED/invalid rates on the eight zero-shot sets and re-score accuracy conditional on valid label-shaped outputs (or re-run with constrained decoding / allowed-token sets to the label list). If pair deltas shrink below the reported +0.012–0.095 band or reverse on valid-only accuracy, Finding-2 is largely a format-compliance effect.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central accuracy claim—that instruction tuning outweighs moderate scale (Finding-2; Qwen2.5-3B-Instruct 0.632 beating multiple 7B bases; all six matched pairs improve)—is measured only under free-text greedy generation, max 20 tokens, exact-match after normalization, with invalid outputs scored wrong and per-model invalid/empty rates not reported (Evaluation Framework; Supplementary Answer-parsing logic; EMPTY_PRED in the error taxonomy). Instruct models are trained to obey “Reply with the label only”; base models are not. A material share of the observed IT advantage may therefore be format- and budget-compliance rather than better intent separation. The paper’s own DeepSeek-R1-Distill-Qwen-1.5B near-zero score is already attributed to trace tokens exhausting the 20-token budget, showing the protocol can zero out models for non-IC reasons. Without invalid-rate tables or a constrained-decoding control, the scientific reading of “IT outweighs scale for zero-shot IC” is under-supported even though the same numbers remain usable as menu-style ranks under this exact free-text deployment protocol. Calibration-proxy scope (reader’s weakest assumption) is real but secondary; it does not underwrite Finding-1/2.","agreement_with_reader":"partial"},"referee_report":{"model":"grok-4.5","summary":"The paper presents a systematic zero-shot evaluation of 41 open-weight LLMs (15 families, 135M–9B) on eight English single-label intent-classification datasets, with ATIS reported separately as a five-shot auxiliary result. Beyond exact-match accuracy it reports confidence calibration (scoped to MASSIVE), robustness to input perturbations, McNemar tests and CIs on MASSIVE rankings, Pareto deployment efficiency, and benchmark saturation. Main claims are that instruction-tuned 3B models can outperform several 7B base models under the stated protocol, that top models on MASSIVE are pairwise statistically indistinguishable, that SNIPS is saturated (32/41 models >80%), and that instruction tuning’s effect on calibration is inconsistent rather than uniformly harmful. The work is framed as practical selection guidance for compute- and latency-constrained deployment.","tokens_in":21411,"tokens_out":1503,"duration_ms":40479,"significance":"If the reported rankings and tradeoffs hold under the stated free-text deployment protocol, the paper fills a genuine gap: prior intent benchmarks either target frontier proprietary models, use multiple-choice reformulations, or omit calibration, robustness, efficiency, and ranking reliability for the sub-9B open-weight range practitioners actually deploy. Strengths include breadth (41 models, production HINT3 sets), explicit statistical tests on MASSIVE, Pareto filtering with measured latency/VRAM, saturation counts that justify de-emphasizing SNIPS, and transparent limitations (ATIS subset skew, single perturbation seed, free-text calibration proxy). These make the resource useful as a menu of protocol-conditioned ranks even where causal claims about instruction tuning need tighter qualification.","major_comments":[{"comment":"Finding-2 (Instruction Tuning vs. Parameter Scale) and the abstract claim that instruction-tuned 3B models outperform 7B bases are measured only under free-text greedy generation, max 20 tokens, exact-match after normalization, with invalid outputs scored wrong (Evaluation Framework; Supplementary Answer-parsing logic). Instruct models are trained to obey “Reply with the label only”; base models are not. The paper’s own DeepSeek-R1-Distill-Qwen-1.5B near-zero score is attributed to reasoning tokens exhausting the budget, showing the protocol can zero models for non-IC reasons. Without per-model invalid/EMPTY_PRED rates (or a constrained-decoding / forced-choice control), a material share of the IT advantage may be format- and budget-compliance rather than intent separation. Please report invalid/empty rates by model (base vs instruct) and either (a) qualify Finding-2 as protocol-conditio","section":"Results: Instruction Tuning vs. Parameter Scale; Finding-2; Evaluation Framework"},{"comment":"Calibration (RQ-8 / Finding-8) rests on sequence-level free-text logprobabilities on MASSIVE as a proxy for label-level confidence (Why Calibration is Restricted to MASSIVE; Limitations). The authors correctly flag length sensitivity and that low ECE can arise from consistent underconfidence. Given that the abstract and Finding-8 still state that instruction tuning’s calibration effect is “inconsistent rather than uniformly harmful,” the claim should be explicitly scoped in the abstract and finding box to “under approximate free-text sequence logprob confidence on MASSIVE,” and any base–instruct ECE comparison should note whether instruct models also change output-length distributions (which mechanically affect sequence logprobs).","section":"Confidence Calibration; Finding-8; Why Calibration is Restricted to MASSIVE"},{"comment":"Main accuracy uses deterministic first-500 index slices (complete HINT3 tests). Limitations documents a real ATIS majority-class skew (+3.8 pp for flight) that likely inflates ATIS figures, and states representativeness was not checked for the other five non-HINT3 sets. Because aggregate ranks and Finding-1 rest on these slices, either (i) report full-test-set scores for at least the top-10 models on CLINC150/Banking77/MASSIVE/MTOP, or (ii) quantify subset-vs-full agreement (e.g., rank correlation / accuracy delta) so readers can bound how much the 40-model ordering could move.","section":"Datasets; Limitations; Overall Model Rankings"}],"minor_comments":[{"comment":"Robustness (Finding-7) uses a single perturbation seed (seed 0). The text already cautions against over-generalizing Qwen2.5-7B-Instruct’s 1.3 pp typo drop; please also report at least 3 seeds or bootstrap CIs for the typo condition on the competitive subset so the “lowest observed degradation” claim is not a one-draw artifact.","section":"Robustness to Input Perturbations; Finding-7"},{"comment":"Saturation is defined as “>50% of models exceed 80% accuracy.” The threshold is reasonable but free; a one-sentence sensitivity note (e.g., 70%/90% cutoffs) would show SNIPS remains saturated under nearby definitions.","section":"Benchmark Saturation Analysis; Finding-3"},{"comment":"Figure 1 and Table 2 report three-decimal aggregates; Qwen2.5-3B-Instruct and Llama-3.1-8B both show 0.632. State how ties are broken in rank order, or mark them as tied.","section":"Overall Model Rankings; Figure 1; Supplementary Table 2"},{"comment":"Prompt templates differ slightly across datasets (spaces vs underscores; Utterance/Intent vs Query/Label; uppercase for HINT3). A short note on whether normalization fully absorbs these surface differences would help reproducibility.","section":"Supplementary Prompt Templates; Scoring"},{"comment":"Related Work cites IntentGrasp and Park et al. appropriately; when discussing saturation, a brief pointer to Larson & Leach (2022) already in the text is good—ensure the survey’s specific saturation observations are not overstated beyond what they claimed.","section":"Related Work; Extended Related Work"},{"comment":"Minor clarity: “group-balanced aggregate” is defined clearly in Evaluation Framework; consider repeating the six-group definition once in the Figure 1 caption so the figure is self-contained.","section":"Figure 1 caption"}],"recommendation":"major_revision","confidential_remarks":"Solid evaluation paper for an empirical venue; the format-compliance confound on Finding-2 is the main reason I chose major_revision rather than minor. If the authors add invalid-rate tables and reframe IT-vs-scale as protocol-conditioned (or add a small constrained-decoding control), this is close to acceptable. Scope fits cs.CL evaluation / dialogue systems tracks; not a theory paper."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"Punchline: this is a careful, deployment-oriented bake-off—41 open-weight models, eight zero-shot IC datasets plus auxiliary 5-shot ATIS—not a new method. If you pick 3B–8B models for task-oriented dialogue, you will actually use the tables.\n\nWhat is new is the scope and the extras, not the task. Nobody had systematically covered the sub-9B open-weight single-label zero-shot setting with calibration, typo robustness, McNemar on MASSIVE, Pareto efficiency, and an explicit SNIPS saturation call. They document exclusions cleanly (ATIS out of aggregates, Qwen2.5-1.5B-Instruct protocol mismatch, DeepSeek-1.5B token-budget artifact). Matched base–instruct pairs all move the same way; top-5 MASSIVE differences are non-significant; SNIPS is cooked (32/41 >80%). That is honest measurement work.\n\nSoft spots, in proportion. The stress-test on Finding-2 lands: free-text greedy, max 20 tokens, exact-match, invalids scored wrong, no invalid-rate table and no constrained-decoding control. Instruct models are trained to emit “label only”; base models are not. Some of the IT-over-scale gap is almost certainly format/budget compliance rather than pure intent separation—the DeepSeek trace eating the budget already shows the protocol can zero a model for non-IC reasons. The numbers remain valid as ranks under this exact serving recipe; the scientific gloss “IT outweighs scale for zero-shot IC” is thinner than the abstract suggests. Calibration only on MASSIVE via sequence logprobs is a second, smaller caveat they already flag. First-500 index subsets and single-seed typos are minor if labeled as such.\n\nMath and stats are ordinary and fine (Wald CIs, McNemar n=500, ECE bins). Citations look appropriate; IntentGrasp and Park et al. are positioned correctly as complementary. No circular fitting.\n\nWho it is for: NLU/dialogue engineers and anyone writing an IC eval suite. Not for people hunting a new architecture or theory. I would bring it to a systems/NLU reading group, cite the saturation and Pareto bits, and send it to peer review with a request for invalid-rate breakdowns or a constrained-decoding ablation before leaning hard on Finding-2.","headline":"Useful practitioner bake-off of 41 sub-9B open-weight models on zero-shot IC; the IT-over-scale headline is real under their free-text protocol but partly confounded with format compliance.","tokens_in":22017,"tokens_out":606,"would_cite":true,"duration_ms":18405,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Instruction-tuned 3B open-weight models can beat larger 7B base models on zero-shot intent classification, while SNIPS no longer ranks them.","keywords":["intent classification","zero-shot learning","open-weight language models","instruction tuning","confidence calibration","benchmark saturation","deployment efficiency","robustness"],"falsifier":"Re-run the same 41 models with constrained decoding or full-test-set sampling and multi-seed typo perturbations: if matched base–instruct pairs no longer show consistent aggregate gains, top-five MASSIVE differences become significant, or SNIPS stops looking saturated relative to production sets, the selection guidance collapses.","tokens_in":21934,"feed_emoji":"🎯","tokens_out":1058,"duration_ms":19888,"temperature":0.7,"pith_summary":"Practitioners who ship task-oriented dialogue systems need to pick an open-weight language model that fits real compute, latency, and noise constraints, but most existing evaluations either chase frontier APIs or stop at accuracy on legacy benchmarks. This paper runs a uniform zero-shot evaluation of 41 open-weight models from 135M to 9B parameters across eight English single-label intent datasets, plus an auxiliary five-shot ATIS result, covering academic suites, a large voice-assistant corpus, and production e-commerce data. Beyond exact-match accuracy it measures calibration, typo and formatting robustness, ranking significance, Pareto efficiency, and benchmark saturation. The central practical claim is that instruction tuning outweighs moderate size jumps in this range: a 3B instruct model matches or beats several 7B base models, leading models are statistically hard to separate on MASSIVE alone, SNIPS is saturated and should be de-emphasized, and calibration effects of instruction tuning are mixed rather than uniformly bad. The point is actionable model-selection guidance for deployable open-weight intent classifiers.","feed_headline":"3B instruct models beat bigger 7B bases on intent classification","feed_subtitle":"A 41-model zero-shot study shows SNIPS is saturated and top ranks often tie on MASSIVE","key_machinery":"A group-balanced eight-dataset zero-shot aggregate (standard IC composite of CLINC150/Banking77/SNIPS, plus MASSIVE, MTOP, and three HINT3 sets), scored by free-text exact-match accuracy under fixed greedy decoding, then extended with McNemar ranking tests, MASSIVE-scoped ECE/Brier calibration, single-seed input perturbations, and a Pareto front over accuracy, latency, and parameters.","core_discovery":"Within the evaluated open-weight sub-9B population, instruction tuning is a stronger signal than moderate parameter scale for zero-shot single-label intent classification: Qwen2.5-3B-Instruct reaches the same reported aggregate as a top-five 8B model and outperforms multiple 7B base models, all six cleanly matched base–instruct pairs improve with instruction tuning, differences among the top five on MASSIVE are pairwise non-significant under McNemar tests, and SNIPS is saturated (most models exceed 80% accuracy) while production-style sets still discriminate.","pith_inferences":["Label-space design may matter as much as model choice: the cross-model play-music/music-query confusion suggests many “model errors” are actually ambiguous taxonomies that few-shot definitions or merged-label metrics could fix.","Reasoning-distilled models need deployment protocols with larger output budgets; fixed short-generation limits can zero out otherwise capable systems and skew comparative tables.","A natural next test is whether the same instruct-over-scale pattern holds for multi-label or multilingual intent under the same efficiency and robustness axes."],"forward_implications":["For maximum zero-shot accuracy in this range, prefer a strong 7B instruct model such as Mistral-7B-Instruct-v0.3 over larger base models alone.","Under tight memory or latency budgets, a 3B instruct model can be Pareto-rational and still competitive with much larger bases.","SNIPS should be demoted in modern IC leaderboards; MASSIVE, MTOP, Banking77, and HINT3-style production sets carry more ranking signal.","Single-benchmark rankings of leading open-weight models are unreliable without multi-dataset aggregates and significance tests.","Typo robustness and calibration must be checked per model; clean accuracy and instruction-tuning status do not reliably predict either."],"fun_headline_variants":["3B instruct beats multiple 7B bases in zero-shot intent tests","Instruction tuning tops moderate scale across 41 open-weight models","SNIPS saturated; top MASSIVE ranks tie under McNemar tests","Matched base-instruct pairs all gain; 3B instruct matches top 8B","Production intent sets still discriminate where SNIPS no longer does"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"The paper treats sequence-level free-text log-probabilities on MASSIVE’s short labels as a usable stand-in for true label confidence, which is the basis for saying instruction tuning’s calibration effect is inconsistent rather than uniformly harmful.","fun_headline_variants_meta":{"raw":{"variants":["3B instruct beats multiple 7B bases in zero-shot intent tests","Instruction tuning tops moderate scale across 41 open-weight models","SNIPS saturated; top MASSIVE ranks tie under McNemar tests","Matched base-instruct pairs all gain; 3B instruct matches top 8B","Production intent sets still discriminate where SNIPS no longer does"]},"model":"grok-4.5","effort":"low","cost_usd":0.003734,"raw_usage":{"total_tokens":1210,"prompt_tokens":827,"num_sources_used":0,"completion_tokens":78,"cost_in_usd_ticks":37344000,"prompt_tokens_details":{"text_tokens":827,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":305,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":827,"tokens_out":78,"duration_ms":6356,"temperature":1.0,"reasoning_tokens":305,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-31T01:24:14.258592+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Re-run the same 41 models with constrained decoding or full-test-set sampling and multi-seed typo perturbations: if matched base–instruct pairs no longer show consistent aggregate gains, top-five MASSIVE differences become significant, or SNIPS stops looking saturated relative to production sets, the selection guidance collapses.","supporting_citations":[],"review_version":1}