{"id":"d6f6f90c-aff3-4c23-b310-dfeb1741dbfa","arxiv_id":"2504.13439","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A fine-tuned Llama model generates multiple-choice distractors that keep model rankings nearly unchanged (Spearman 0.99) and matched confidence entropy, while human scores are only reported on a separate set of tasks.","lead":"The paper introduces D-GEN, a fine-tuned Llama model that writes plausible wrong answers to turn open-ended questions into multiple-choice tests. A smart generalist might use it because it offers a cheap way to build and audit MC benchmarks for comparing language models.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Rank alignment masks a systematic ~5% accuracy drop (Table 3) that is never checked by human evaluation on MMLU-DGEN, so the central replacement claim lacks evidence of difficulty equivalence.","rationale":"The reader's weakest assumption identifies the validity of proxy metrics as proof of distractor quality, specifically noting that rank correlation is insensitive to a uniform difficulty shift and that Table 3 shows MMLU-DGEN is slightly harder on average. My analysis converges on that same point as the most load-bearing concern: the paper's headline quantitative support (rank correlation and entropy) cannot detect a systematic accuracy drop, and the paper's own Table 3 provides direct evidence of such a drop. The additional issue that human evaluation never covers MMLU-DGEN reinforces the concern but is secondary; the primary gap is the missing difficulty-equivalence test. I agree with the reader that the central claim is conditionally supported: the released models and code, the high rank correlations, and the entropy comparisons are real evidence, but the 'closely matches' and 'reliable assessment' conclusions would require either demonstrating that the ~5% mean accuracy drop is not systematic (e.g., not significantly different from zero) or reframing the claim to be explicitly about ranking only. The paper itself includes a Limitations section acknowledging the dependency on ground-truth distractors and the unreliability of LLM-based evaluation, which is honest but does not resolve the difficulty-shift issue. A conditional verdict remains appropriate: the work is useful and mostly well-executed, but the central generalizability claim needs one additional verification step. Therefore I do not change the reader's verdict.","tokens_in":40287,"tokens_out":4962,"duration_ms":47946,"concrete_test":"Recompute from the released MMLU and MMLU-DGEN results the 42 paired per-configuration accuracy differences (original minus D-GEN) and run a paired Wilcoxon signed-rank test (or paired t-test on the differences). If the median difference is significantly below zero at p < 0.05, then D-GEN systematically increases benchmark difficulty, directly contradicting the 'closely matches' wording in the central claim. This single check uses data already reported in Table 1/Table 3 and settles whether the ranking-alignment result is compatible with a meaningful absolute-score shift.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that D-GEN can replace human-written MMLU distractors without changing the measured relative performance of models, supported by rank correlation (Spearman 0.99, Kendall 0.94) and entropy matching. The load-bearing weakness is that these proxy metrics are insensitive to exactly the kind of systematic shift the paper's own data reveal. Section 4.4, Table 3 reports accuracy differences (MMLU-DGEN minus MMLU) with mean -0.05, minimum -0.09/-0.10, and maximum +0.01/+0.04 across 42 configurations; a mean drop of 5 percentage points is a substantial change in benchmark difficulty. Spearman and Kendall correlations are invariant to monotone transformations, so a uniform difficulty shift produces perfect rank correlation while changing every absolute score. The entropy analysis (Section 5) only compares three large models (Llama-3.3-70B-Instruct, Qwen2.5-72B-Instruct, Mixtral-8x7B-Instruct-v0.1) at the domain-aggregate level and does not establish item-level or model-level difficulty equivalence; indeed, the paper notes slightly higher entropy in MMLU-DGEN. Furthermore, the human evaluation (fluency, coherence, distracting ability, incorrectness) is conducted only on 700 FLAN examples (Section 6), never on MMLU-DGEN, so the abstract's 'Human evaluation further confirms' does not apply to the product for which the replacement claim is made. The paper honestly acknowledges MMLU-DGEN is 'slightly more challenging' (Section 4.4), but the central assertion of 'closely matches' requires difficulty preservation, not merely ranking preservation. Without a direct check for systematic difficulty shift, the evidence does not rule out a benchmark that is uniformly harder and yields different absolute scores, undermining the 'reliable assessment' claim.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces D-GEN, an open-source distractor generation model obtained by fine-tuning Llama-3.1-8B-Instruct and Llama-3.3-70B-Instruct on the MMLU auxiliary training set. It proposes two automated evaluation metrics for distractor quality: ranking alignment (Spearman/Kendall correlation of model accuracies between original MMLU and MMLU-DGEN) and entropy analysis (comparison of model confidence distributions). Experiments show Spearman rho 0.99 and Kendall tau 0.94 over 42 configurations, and entropy differences that are mostly not statistically significant. Human evaluation on FLAN tasks reports high fluency, coherence, distractiveness, and incorrectness scores. The paper claims that D-GEN can replace human-written distractors in multiple-choice benchmarks without altering relative model rankings.","tokens_in":40599,"tokens_out":5836,"duration_ms":48517,"significance":"If the central claim is established, D-GEN offers a practical, low-cost way to convert open-ended tasks into multiple-choice format and to enlarge or refresh MC benchmarks with automatically generated distractors. The paper's strengths include the release of the models and datasets, the use of a broad set of 21 models with 0- and 5-shot settings, and the honest acknowledgment in Section 4.4 and the Limitations that MMLU-DGEN is slightly harder and that the evaluation methods depend on ground-truth distractors. However, as detailed below, the evidence for difficulty equivalence is incomplete: the headline rank-correlation metric is insensitive to the systematic accuracy drop reported in Table 3, and the human evaluation does not cover the MMLU-DGEN set. With additional targeted analyses, the work could become a solid contribution to reliable MC evaluation.","major_comments":[{"comment":"The paper reports mean accuracy differences (MMLU-DGEN minus MMLU) of -0.05 overall, with per-domain means of -0.03 to -0.07 and minimums of -0.09/-0.10. This is a systematic downward shift in scores. Spearman's rho and Kendall's tau are invariant to monotone transformations of accuracy, so the high rank correlations in Table 2 cannot detect such a uniform difficulty shift. Since the central claim is that D-GEN can replace human-written MMLU distractors in benchmark evaluation, the paper needs additional evidence of difficulty equivalence, e.g., item-level accuracy distributions, calibration plots, or human scoring on a sample of MMLU-DGEN items.","section":"§4.4, Table 3"},{"comment":"Human evaluation is conducted only on 700 FLAN examples across seven tasks; no human evaluation is performed on the MMLU-DGEN set that is the primary product of the paper. The abstract's statement that \"Human evaluation further confirms the fluency, coherence, distractiveness, and incorrectness\" therefore does not apply to MMLU-DGEN, where the replacement claim actually stands. Without human judgment on MMLU-DGEN, the quality of these distractors rests entirely on proxy metrics, which the paper itself shows are imperfect (Table 3 difficulty shift).","section":"§6.2–6.3"},{"comment":"The entropy analysis finds no statistically significant difference in 11 of 12 model-domain pairs, but the direction is highly consistent: D-GEN entropy is higher in every domain for Llama and Qwen, and in 2 of 4 domains for Mixtral. With domain-level sample sizes of 12–19 subcategories, the Wilcoxon tests have limited power, and no multiple-comparison correction is applied. The paper's conclusion that the entropy distributions \"closely match\" is therefore not strongly supported; the data are consistent with a small but systematic increase in uncertainty, matching the difficulty shift in Table 3. I recommend reporting effect sizes, confidence intervals, and item-level analyses.","section":"§5.3, Table 4"},{"comment":"The human evaluation section does not report the number of annotators, their qualifications, or inter-annotator agreement. As a result, the reliability of the reported average scores (Table 7) cannot be assessed. Since human evaluation is used to validate distractor quality, this information is essential for the claims.","section":"§6.2, Appendix F.3"}],"minor_comments":[{"comment":"The abstract says \"closely matches the entropy distribution,\" but Section 5 only reports mean entropy and p-values; the full distributions are not shown.","section":"Abstract, §5"},{"comment":"The text states \"Most of these p-values exceed 0.05,\" but Table 65 lists rank-correlation p-values that are all far below 0.05; this appears to be a copy-paste error from the entropy analysis.","section":"Appendix D.2"},{"comment":"Please clarify whether \"preserving the original distribution of correct answer indices\" means preserving the sequence of correct positions or just the marginal distribution.","section":"§4.2"},{"comment":"The one significant p-value (0.0342) is described as \"the only exception,\" but with 12 tests, one significant result is expected by chance; no multiple-comparison correction is discussed.","section":"§5.3"},{"comment":"Consider reporting the raw accuracy values in addition to the differences, since a mean difference of -0.05 may be small for some models but substantial for others.","section":"Table 3"}],"recommendation":"major_revision","confidential_remarks":"The paper is a plausible and useful contribution to MC-format evaluation, but the central replacement claim needs to be either strengthened with direct difficulty-equivalence evidence on MMLU-DGEN or softened to a ranking-preservation claim. The authors are transparent about the difficulty shift in Section 4.4, which is commendable, but the abstract and conclusion currently overstate what the evidence supports."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe useful core is real: D-GEN fine-tunes Llama-3.1-8B and 3.3-70B to generate distractors, releases both sets of weights, and shows that swapping MMLU's human-written distractors for theirs preserves the ordering of 21 models across 42 configurations (Spearman ~0.99). That is a practical result for anyone building MC benchmarks. The two proposed metrics—ranking alignment and entropy analysis—are sensible, and the entropy comparison across three models is mostly clean. The human evaluation on FLAN tasks shows decent quality, and the baseline comparison against the unfine-tuned base model makes the case for fine-tuning.\n\nNow the soft spots. The stress-test concern is legitimate: the paper claims to preserve difficulty, but Table 3 reports a mean accuracy drop of about 5 percentage points on MMLU-DGEN. Rank correlations are blind to uniform difficulty shifts, so the strong Spearman/Tau numbers do not address that. The paper acknowledges the shift but calls it 'slightly more challenging'; a 5% drop is not trivial for a benchmark intended to 'closely match' the original. And the human evaluation was done on FLAN examples, not on MMLU-DGEN, so we never get direct human judgment of the actual product. That is a real gap.\n\nA minor point: the entropy analysis uses Llama-3.3-70B-Instruct, the base model of D-GEN, which introduces some circularity; including Qwen and Mixtral partially controls for it. The human evaluation also lacks annotator counts and agreement metrics, which should be added.\n\nNone of this sinks the paper. The central claim—that D-GEN preserves model rankings—holds up. The difficulty-equivalence claim is weaker, and the abstract's 'new standard' phrasing overreaches. A revision that adds human scores on MMLU-DGEN, or at least a direct difficulty analysis, would solidify it.\n\nThis is for researchers building or modifying MC benchmarks and people working on automatic distractor generation. It deserves peer review; a serious referee would ask for those additions, but the core contribution is worth engaging with.","headline":"Useful distractor-generation model with honest experiments, but the claim of difficulty preservation is overstated given the ~5% accuracy drop in Table 3.","tokens_in":41159,"tokens_out":3162,"would_cite":true,"duration_ms":28215,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"D-GEN, a fine-tuned language model, generates multiple-choice distractors that preserve the ranking of 42 model configurations (Spearman's $\\rho$ 0.99) and match the confidence distribution of human-written distractors.","keywords":["distractor generation","multiple-choice evaluation","MMLU","ranking alignment","entropy analysis","large language models","human evaluation","benchmark construction"],"falsifier":"Run the paper's own human-evaluation protocol (fluency, coherence, distractiveness, incorrectness) on a random sample of MMLU-DGEN questions: if the distractiveness or incorrectness ratings fall clearly below the original MMLU distractors', or if naive human test-takers score significantly lower on MMLU-DGEN than on MMLU, the claim that the generated distractors are interchangeable with the originals fails.","tokens_in":40083,"feed_emoji":"📝","tokens_out":8190,"duration_ms":66897,"temperature":0.7,"pith_summary":"D-GEN is a fine-tuned large language model that turns open-ended questions into multiple-choice format by generating three plausible but incorrect distractors for each question. The paper argues that these machine-written distractors are interchangeable with human-written ones: across 21 models and two few-shot settings, model rankings on the rewritten MMLU benchmark match the original with Spearman's $\\rho$ 0.99 and Kendall's $\\tau$ 0.94, and the entropy of models' answer-choice distributions is statistically indistinguishable from the original for two of three judge models. Human ratings on seven other tasks confirm the distractors are fluent, coherent, distracting, and incorrect. If the claim holds, expensive human distractor writing can be replaced by an automated pipeline, making multiple-choice evaluation fast, scalable, and reliable.","feed_headline":"Machine-written test distractors preserve model rankings","feed_subtitle":"A fine-tuned LLM writes multiple-choice wrong answers that align model rankings at 0.99, enabling automated benchmarks.","key_machinery":"The central object is the D-GEN model itself: a LLaMA-3.3-70B-Instruct (and an 8B variant) fine-tuned on the MMLU auxiliary training set to output three semantically close but incorrect distractors for a question and its correct answer. An automatic correction loop filters and regenerates distractors until they are unique and non-overlapping with the correct answer, and the correct answer is randomly placed among the four options to cancel position bias. The evaluation rests on two quantitative instruments: ranking alignment (Spearman's $\\rho$ and Kendall's $\\tau$ between model performance ranks on original MMLU and MMLU-DGEN) and entropy analysis (the Shannon entropy $H(p)=-\\sum_i p_i \\log p_i$ of the softmax probability distribution over A/B/C/D, compared with Wilcoxon signed-rank tests). These instruments are what carry the argument that the generated distractors have the same discriminatory power and plausibility as human-written ones.","core_discovery":"The paper's central claim is that D-GEN-generated distractors preserve the measurement properties of the original MMLU test set. Concretely, replacing the human-written wrong answers with D-GEN's distractors does not change the relative ordering of 42 model configurations (Spearman's $\\rho$ 0.99, Kendall's $\\tau$ 0.94), and the distributions of model confidence, measured as entropy over the four choices, are close enough that Wilcoxon signed-rank tests find no significant difference for two of the three judge models and only one domain (Social Sciences on Llama-3.3-70B-Instruct) shows $p < 0.05$. The paper also demonstrates on seven FLAN tasks that the generated distractors receive high average scores (mostly 4-5 on a 1-5 scale) for fluency, coherence, distractiveness, and incorrectness. The authors intend this as evidence that D-GEN is the first open-source distractor generator reliable enough for automated multiple-choice evaluation.","pith_inferences":["The rank-correlation evidence would survive a uniform difficulty shift, so the paper's strongest statistics do not by themselves prove that absolute difficulty is preserved; a direct human or behavioral comparison of MMLU versus MMLU-DGEN would settle this.","The same entropy-matching methodology could be adapted as a screening test for adversarial-robustness benchmarks: distractors that maximize model uncertainty are exactly the options that stress-test calibrated confidence.","A natural extension is to train D-GEN on the distractors generated by itself, bootstrapping new MC benchmarks from open-ended data with no human-written gold options, which would remove the ground-truth dependency the paper names as its main scalability limitation.","The paper's finding that general-purpose LLM judges penalize correct-but-intentionally-incorrect distractors suggests automated distractor evaluation by such judges needs task-specific calibration, a caveat that applies beyond this paper."],"forward_implications":["Benchmarks like MMLU can be regenerated or extended with fresh distractors without re-running human annotation, at a fraction of the cost.","Model evaluation on open-ended tasks can be converted to multiple-choice format automatically, reducing false negatives caused by format inconsistencies in generation.","The ranking-alignment and entropy tests give future distractor generators a quantitative, scalable acceptance criterion that does not require expert human judgment.","Because the D-GEN distractors slightly increase difficulty (negative mean accuracy differences across all domains), benchmarks built this way may be marginally harder than their originals, a shift the headline rank-correlation statistic does not reveal.","The 8B variant and open-source release allow other researchers to generate distractors for new datasets without access to proprietary models."],"supporting_citations":[{"why":"Supplies the MMLU auxiliary training set and the original test questions and distractors used for ranking alignment.","marker":"Hendrycks et al., 2021a"},{"why":"Provides the Llama-3.1-8B-Instruct and Llama-3.3-70B-Instruct base models that D-GEN fine-tunes.","marker":"Grattafiori et al., 2024"},{"why":"Provides the evaluation harness used to compute accuracies across the 21 models.","marker":"Gao et al., 2024"},{"why":"Contributes the FLAN collection that supplies the seven tasks for the task-applicability and human-evaluation experiments.","marker":"Wei et al., 2022"},{"why":"Extends the FLAN collection, giving the consolidated instruction datasets used for task applicability.","marker":"Longpre et al., 2023"},{"why":"Supplies the four human-evaluation metrics (fluency, coherence, distracting ability, incorrectness) used to rate D-GEN distractors.","marker":"Zhou et al., 2019"},{"why":"Provides LoRA, the efficient fine-tuning method used for the 70B variant of D-GEN.","marker":"Hu et al., 2021"},{"why":"Informs the decision to randomly assign the correct answer among the four options to avoid position bias.","marker":"Zheng et al., 2023"}],"fun_headline_variants":["AI-generated wrong answers keep model rankings intact","Benchmark distractor generator preserves model order at 0.99","Automated distractors match human quality for MC tests","Open-source model writes wrong answers, rankings don't budge","Spearman rho 0.99: AI distractors keep benchmark rankings"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The argument assumes that two aggregate statistics, rank correlation across 42 model configurations and matched entropy distributions on three judge models, prove that D-GEN's distractors are as good as human-written ones, even though no human ever scores the MMLU-DGEN distractors themselves; human scores come only from different FLAN tasks.","fun_headline_variants_meta":{"raw":{"variants":["AI-generated wrong answers keep model rankings intact","Benchmark distractor generator preserves model order at 0.99","Automated distractors match human quality for MC tests","Open-source model writes wrong answers, rankings don't budge","Spearman rho 0.99: AI distractors keep benchmark rankings"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000604,"raw_usage":{"total_tokens":2801,"prompt_tokens":913,"completion_tokens":1888,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":529,"completion_tokens_details":{"reasoning_tokens":1804}},"tokens_in":529,"tokens_out":1888,"duration_ms":14498,"temperature":1.0,"reasoning_tokens":1804,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T12:08:06.477048+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the paper's own human-evaluation protocol (fluency, coherence, distractiveness, incorrectness) on a random sample of MMLU-DGEN questions: if the distractiveness or incorrectness ratings fall clearly below the original MMLU distractors', or if naive human test-takers score significantly lower on MMLU-DGEN than on MMLU, the claim that the generated distractors are interchangeable with the originals fails.","supporting_citations":[{"cited_title":"Co-Attention Hierarchical Network: Generating Coherent Long Distractors for Reading Comprehension","cited_arxiv_id":"1911.08648","evidence_quote":"Supplies the four human-evaluation metrics (fluency, coherence, distracting ability, incorrectness) used to rate D-GEN distractors."}],"review_version":1}