{"id":"2a9942e3-a35b-4a25-9fdc-b578b0dffd8a","arxiv_id":"2501.17195","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"The paper presents Selene Mini, an 8B open-weights judge model that reports state-of-the-art average scores across 11 LLM evaluation benchmarks, with gains on medical and financial expert agreement.","lead":"Atla Selene Mini is an open 8B language model trained to grade other AI models' responses, and the paper reports it outperforms similar small judges and GPT-4o-mini on an average of 11 evaluation benchmarks. The claim matters because reliable automated judging could replace costly human review of LLM outputs.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 0.007 overall margin over SFR-LLaMA-3.1-8B-Judge lacks error bars and mixes externally reported baseline numbers; the 'outperforms' claim is not established.","rationale":"The reader's weakest assumption correctly identifies the OOD/contamination risk and the small margin without error bars. I agree with the conditional verdict, but the most load-bearing issue for the central claim is the absence of variance estimates in the comparison: a 0.007 difference on a noisy average of 11 heterogeneous metrics is not enough to assert superiority, especially when several baseline numbers are quoted from other papers rather than measured in the same harness. The OOD concern is also serious and independently testable, but it does not fully capture the fragility of the quantitative claim. My proposed test—re-running the closest baseline under identical conditions with multiple seeds and bootstrapping—would directly settle whether the headline margin is real. The RewardBench contradiction in Table 1 (Selene Mini 0.584 vs. GPT-4o-mini 0.615) further supports caution but is a narrower issue than the overall comparison. Therefore the reader's conditional acceptance remains appropriate, pending the missing statistical and contamination evidence.","tokens_in":13559,"tokens_out":9772,"duration_ms":76452,"concrete_test":"Run Selene Mini and SFR-LLaMA-3.1-8B-Judge on the same 11 benchmarks with identical prompts, seeds, and parsing, for at least 3 independent repetitions; compute bootstrap 95% CIs on the overall average. If the 0.007 difference is not significant (e.g., CI overlaps zero or the effect is <0.005), the superiority claim is unsupported. Additionally, publish the training dataset list and run an n-gram overlap check between training and eval sets to validate the OOD claim.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central comparative claim rests on an equal-weight average of 11 benchmarks where Selene Mini scores 0.756 vs. SFR-LLaMA-3.1-8B-Judge at 0.749, a 0.007 difference. The paper reports no confidence intervals, significance tests, or repeated runs, and several baseline rows (†) are taken from external technical reports rather than reproduced in the same prompt/harness/parsing. Because the overall average combines Pearson correlations (absolute scoring) with accuracies (classification/pairwise) and weighs task types equally, small perturbations in any single benchmark could flip the ranking. Without variance estimates, the headline 'outperforms the best SLMJs' is a point estimate within noise. Additionally, the 'out-of-distribution' guarantee is uncheckable because the 16 training datasets are never enumerated; Appendix A only shows an embedding map, so contamination of MT-Bench, RewardBench, or others cannot be ruled out.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript introduces Atla Selene Mini, an 8B-parameter Llama-3.1-based language-model judge trained on a curated mixture of 16 public datasets augmented with synthetically generated chosen/rejected chain-of-thought critiques, filtered by a reward model and consistency checks, and optimized with a DPO+NLL objective. The central claim is that the resulting model is a general-purpose evaluator that outperforms existing small language-model judges and GPT-4o-mini on an equal-weight average over 11 benchmarks covering absolute scoring, pairwise preference, and classification, and is the highest-scoring 8B generative model on RewardBench. Additional claims concern zero-shot agreement on finance and medical expert labels, robustness to prompt formatting, and top ranking in the authors' community Judge Arena.","tokens_in":13776,"tokens_out":5054,"duration_ms":42783,"significance":"If fully substantiated, an open-weights 8B judge that matches or beats GPT-4o-mini and specialized judges on a broad benchmark suite would be practically significant, lowering the cost of automated evaluation and enabling local deployment. The paper's strengths are its reproducible release of weights, explicit ablations of reward-model filtering and dataset inclusion, and a reasonably detailed curation pipeline. The main weakness is that the headline comparative claim rests on a 0.007 average margin with no statistical uncertainty, and some abstract-level factual claims are contradicted by the paper's own tables. The contribution is primarily empirical; the method itself is a combination of known techniques.","major_comments":[{"comment":"The abstract's RewardBench claim is contradicted by the paper's own data. Table 1 reports RewardBench scores of 0.688 for Selene Mini and 0.689 for SFR-LLaMA-3.1-8B-Judge, so Selene Mini is not the highest-scoring 8B generative model on RewardBench among the models listed. Similarly, §3.1 states that Selene Mini beats GPT-4o on RewardBench, EvalBiasBench, and Auto-J, but Table 3 shows GPT-4o at 0.765, 0.932, and 0.769 versus Selene Mini's 0.688, 0.900, and 0.732, respectively. These claims need to be corrected or removed; as written, the headline superiority statement is not supported.","section":"Abstract; §3.1, Tables 1 and 3"},{"comment":"The overall comparison is a single-run point estimate without error bars, confidence intervals, or significance tests, and the margin over SFR-LLaMA-3.1-8B-Judge is 0.007 (0.756 vs 0.749). Moreover, several baseline rows marked with a dagger are taken from external technical reports, so the comparison mixes evaluation harnesses, prompts, and parsing conventions. Since a small perturbation in one benchmark could flip the ranking, the claim that Selene Mini outperforms the best SLMJs is not established. Please report per-run variance or repeated evaluations with different seeds, and either reproduce all baselines under identical conditions or restrict the comparative claim to the subset that was run in-house.","section":"§3.1, Table 1"},{"comment":"The out-of-distribution status of the 11 evaluation benchmarks is asserted but not verifiable. The training mixture is described only as 16 public datasets inspired by FLAMe, and Appendix A contains an embedding visualization rather than a dataset list. Without enumerating the training datasets and, ideally, reporting overlap or contamination checks against MT-Bench, RewardBench, FLASK, HHH, and the other evaluation benchmarks, the OOD premise cannot be assessed. Please disclose the composition of the training mix and any decontamination procedure.","section":"§2.1, §3.1, Appendix A"},{"comment":"The Judge Arena evidence is self-referential: the authors developed the arena and report an early snapshot of their own model as top-ranking. This is not independent validation. At minimum, state the number of votes, the evaluation procedure, and the relationship between the authors and the platform, or treat the arena result as anecdotal rather than as part of the headline evidence.","section":"§3.3, Abstract"}],"minor_comments":[{"comment":"The text says 'six different prompt formats' but lists only five (original, markdown, JSON, PrePair, and simplified instructions); please correct the count or add the missing format.","section":"§3.2.2"},{"comment":"The column label 'RewardB' is ambiguous; consider writing 'RewardBench' in full for readability.","section":"Tables 1 and 3"},{"comment":"Reference [25] contains a typo in 'Foundation'; please fix it.","section":"References"},{"comment":"The filtering ablation results in Figure 6 are reported without error bars or sample sizes; adding them would strengthen the conclusion that the effects are dataset-dependent.","section":"Appendix C, Figure 6"},{"comment":"The rejected-judgment sampling for absolute-scoring tasks is described only for a 1–5 scale; please clarify how it generalizes to other numeric scales used in the evaluation suite.","section":"§2.2"}],"recommendation":"major_revision","confidential_remarks":"This is an empirical model report. The factual inconsistencies in the abstract and §3.1 are the main barrier: they suggest the headline claims were not checked against Tables 1 and 3. If the authors correct the claims, add uncertainty quantification, and disclose the training data composition, the manuscript could be acceptable as a systems or benchmark report. If the 0.007 margin remains the only evidence for superiority, the paper would not support the claimed outperformance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my take. The paper's real contribution is the released artifact: an open-weights 8B judge trained on a curated mix of public preference data with synthetic critiques and a DPO+NLL objective. That's useful, and the ablation on reward-model filtering is a legitimately informative piece of engineering. The industry-dataset results (CRAFT-MD, FinanceBench) are also a sensible sanity check that the training transfers to domains outside the academic benchmarks.\n\nThe soft spots are all about the evidence for the headline. The abstract says Selene Mini is 'the highest-scoring 8B generative model on RewardBench,' but Table 1 shows SFR-LLaMA-3.1-8B-Judge at 0.689 vs. Selene's 0.688. That's a direct internal contradiction and it needs to be fixed. Second, the overall average advantage over the same SFR model is 0.007 (0.756 vs 0.749). No error bars, no significance test, and some baseline rows are taken from external reports rather than reproduced in the same harness. With that margin, 'outperforms' is a point estimate within noise. Third, the 11 benchmarks are called out-of-distribution, but the 16 training datasets are never enumerated; Appendix A is just an embedding map. So no one can check contamination.\n\nNone of these are fatal to the engineering value. The model weights are public, so anyone can run their own evaluation and see if the ranking holds. But the paper as written overstates what its own numbers support. I'd want the RewardBench claim corrected, the training data list and contamination analysis added, and at least a few repeated runs or an error bar on the average before taking the superiority claim at face value.\n\nThis paper deserves a serious referee, not a desk reject. The artifact is real and the curation details are worth scrutiny. But the referee should push for the evidential fixes. I'd cite it as a resource if I need an open judge model, though not for the comparative claim.","headline":"Useful open judge model, but the headline benchmark superiority is a point estimate within noise and the abstract's RewardBench claim contradicts Table 1.","tokens_in":14337,"tokens_out":2776,"would_cite":true,"duration_ms":23464,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper introduces an 8-billion-parameter open-weights judge model that outperforms existing small judges and GPT-4o-mini on the average of 11 evaluation benchmarks, and claims the top spot among 8B generative models on RewardBench.","keywords":["LLM-as-a-judge","small language model","direct preference optimization","synthetic data curation","evaluation benchmarks","RewardBench","prompt robustness","open-weight model"],"falsifier":"Check the 11 benchmark test sets for exact or near-duplicate overlap with the 16 training datasets; if any of MT-Bench, RewardBench, FLASK, or the others appear in training, the out-of-distribution premise fails. Then rerun the full 11-benchmark comparison several times with different seeds or checkpoints to see whether the roughly 0.007 average margin over the closest baseline survives run-to-run noise.","tokens_in":13335,"feed_emoji":"⚖️","tokens_out":7230,"duration_ms":63739,"temperature":0.7,"pith_summary":"The paper sets out to show that a capable, general-purpose LLM-as-a-judge can be built at 8 billion parameters rather than requiring frontier-scale models. It presents Selene Mini, a judge fine-tuned from Llama 3.1 8B Instruct, claims it beats the best existing small-language-model judges and GPT-4o-mini on the average of 11 evaluation benchmarks, and claims the top score among 8B generative models on RewardBench. The wider importance is practical: if true, reliable automated evaluation can be run locally, cheaply, and with open weights, and the main lever is data curation and a hybrid training objective rather than model size. The paper also reports gains in zero-shot agreement with human expert labels on finance and medical datasets and stable performance when prompt formatting changes.","feed_headline":"An 8B open judge beats larger models on 11 evaluation benchmarks","feed_subtitle":"Fine-tuned with DPO on synthetic critiques, Selene Mini tops GPT-4o-mini on average.","key_machinery":"The load-bearing mechanism is the data-and-loss recipe, not a new architecture. Each training point pairs a chosen evaluation, which argues for the ground-truth label or a correct score, with a rejected evaluation arguing for a wrong label or a score two points off; both are written as chain-of-thought critiques with a final judgment. The model is trained with $L_{\\text{DPO+NLL}} = L_{\\text{DPO}} + \\alpha\\,L_{\\text{NLL}}$, where the extra negative log-likelihood term is applied only to chosen responses. Before training, a reward model filters low-quality raw examples and a prompted consistency checker removes synthetic critiques whose reasoning contradicts their assigned judgment. Around 70 percent of training pairs use the critique-plus-judgment format and 30 percent use judgments only, following the baseline judge recipe.","core_discovery":"Selene Mini is a fine-tuned Llama 3.1 8B Instruct model trained on 16 public datasets that were augmented with synthetic chosen and rejected critiques, then filtered for quality. On the paper's headline comparison, it scores 0.756 on the unweighted average across 11 benchmarks, ahead of SFR-LLaMA-3.1-8B-Judge (0.749) and GPT-4o-mini (0.743). It also reports the highest score among 8B generative models on RewardBench, surpassing GPT-4o and specialized judges in that comparison. The authors attribute the improvement to their curation pipeline: reward-model filtering of raw data, a critique-consistency checker, and a DPO objective that adds a negative log-likelihood term on chosen responses so the margin over rejected critiques is widened while good critiques become more probable.","pith_inferences":["The reported 'best overall' claim is an equal-weight average over the 11 benchmarks; a user who weights absolute scoring, classification, and pairwise tasks differently could see a different ordering, and the paper itself notes that practitioners prefer absolute scoring.","The 0.007 gap in overall average over the closest baseline is small enough that run-to-run variance or a different benchmark mix could plausibly change the rank order, so the stability of the headline result is not yet established.","Because the weights are open, the community can rerun the evaluations and extend them to new domains; that independent evidence, rather than the paper's own runs, may settle how general the capability really is."],"forward_implications":["An 8B open-weights judge can match or exceed proprietary mini judges on general evaluation tasks, lowering the cost and latency of automated evaluation.","The same model can plausibly serve as a reward signal for preference optimization, given its top ranking among 8B generative models on RewardBench.","Practitioners should be able to vary prompt templates in production without retraining, since Selene Mini's score stays roughly stable across the tested formats.","The zero-shot gains on finance and medical expert-labeled data suggest the training recipe transfers beyond academic benchmarks into regulated, domain-specific settings."],"supporting_citations":[{"why":"Defines the judge training-pair format used here (70% chain-of-thought, 30% judgment-only) and provides SFR-LLaMA-3.1-8B-Judge, the closest baseline Selene Mini is compared against.","marker":"[10]"},{"why":"Supplies the FLAMe dataset collection across pairwise, absolute-scoring, and classification tasks that inspired the training mix.","marker":"[7]"},{"why":"Introduces the DPO variant with an added negative log-likelihood term that the training loss is built on.","marker":"[13]"},{"why":"Provides the ArmoRM reward model used to filter low-quality raw training data before synthetic augmentation.","marker":"[14]"},{"why":"Is the RewardBench leaderboard where Selene Mini claims to be the highest-scoring 8B generative model.","marker":"[18]"},{"why":"Is CRAFT-MD, the medical expert-annotated dataset used to measure zero-shot agreement.","marker":"[25]"},{"why":"Is FinanceBench, the finance expert-annotated dataset used to measure zero-shot agreement.","marker":"[26]"},{"why":"Is the Judge Arena platform whose early leaderboard snapshot ranks Selene Mini first among judge models.","marker":"[12]"}],"fun_headline_variants":["8B judge tops larger models on 11 benchmarks","Atla Selene Mini: small judge, huge win rate","Open 8B model beats GPT-4o-mini as evaluator","Tiny judge: 8B Selene Mini wins 11 eval tests"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The central claim depends on treating an unweighted average of scores on 11 chosen benchmarks as the definition of general-purpose evaluation quality, and on those benchmarks being genuinely outside the 16 training datasets.","fun_headline_variants_meta":{"raw":{"variants":["8B judge tops larger models on 11 benchmarks","Atla Selene Mini: small judge, huge win rate","Open 8B model beats GPT-4o-mini as evaluator","Tiny judge: 8B Selene Mini wins 11 eval tests"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001023,"raw_usage":{"total_tokens":4323,"prompt_tokens":962,"completion_tokens":3361,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":578,"completion_tokens_details":{"reasoning_tokens":3286}},"tokens_in":578,"tokens_out":3361,"duration_ms":20573,"temperature":1.0,"reasoning_tokens":3286,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T13:43:11.024553+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Check the 11 benchmark test sets for exact or near-duplicate overlap with the 16 training datasets; if any of MT-Bench, RewardBench, FLASK, or the others appear in training, the out-of-distribution premise fails. Then rerun the full 11-benchmark comparison several times with different seeds or checkpoints to see whether the roughly 0.007 average margin over the closest baseline survives run-to-run noise.","supporting_citations":[{"cited_title":"Craft-md: A conversational evaluation framework for comprehensive assessment of clinical llms","cited_arxiv_id":null,"evidence_quote":"Is CRAFT-MD, the medical expert-annotated dataset used to measure zero-shot agreement."},{"cited_title":"Judge arena: Benchmarking llms as evaluators","cited_arxiv_id":null,"evidence_quote":"Is the Judge Arena platform whose early leaderboard snapshot ranks Selene Mini first among judge models."}],"review_version":1}