{"id":"50ccca0c-dc27-44cc-86a2-636fb576a16a","arxiv_id":"2506.20274","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"A 14-task enterprise LLM benchmark built mostly from GPT-4o-generated labels and scored by GPT-4o-as-judge shows open-source models closing the reasoning gap, but the dataset is not public and the evaluation is partly self-referential.","lead":"This paper presents a 14-task benchmark for assessing large language models on enterprise workflows, including summarization, bias detection, and Jira query conversion, organized by Bloom's Taxonomy. It reports that open-source models such as DeepSeek R1 approach proprietary models on reasoning tasks, though the benchmark data are not released and the labels come from the same model family that scores them.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"GPT-4o serves as labeler, judge, and evaluated model; without blinded human re-scoring, reported rankings may reflect judge self-preference rather than true model capability.","rationale":"The reader's weakest assumption—that GPT-4o labels and G-Eval scores are valid ground truth—is precisely the load-bearing point. I agree with that assessment. A benchmark is only robust if the measurement is independent of the thing measured. Here the measurement instrument and one competitor share the same model family and version (GPT-4o-2024-11-20). Prior work has documented LLM self-preference in evaluation, and the paper itself acknowledges judge limitations but does not implement standard mitigations such as blinding or order swapping. Therefore the central ranking claims are not supportable without additional validation. The proposed test would settle the question by measuring judge bias directly. Assuming the reader's REJECT is based on this unaddressed circularity, I recommend no change in verdict. If the test revealed no bias, the model rankings would be more credible; until then, REJECT is the correct disposition. Additional internal inconsistencies, such as the 'seven models' wording in Section 4.3 while Table 2 lists six, and the sample-size sum of 9,261 versus the stated approximate 9,700, further reduce confidence in the reported quantities, but they are secondary to the measurement-validity concern.","tokens_in":15484,"tokens_out":4637,"duration_ms":45021,"concrete_test":"Select 200 responses per open-ended task stratified across the six models, anonymize model identity, and have human experts and a different LLM judge (e.g., Claude 3.5 Sonnet) score them with the same G-Eval rubric. Compare per-model mean scores and inter-judge agreement with the original GPT-4o scores. If GPT-4o scores its own outputs significantly higher than either independent judge does, or human agreement falls below a pre-registered threshold, the reported ranking is not supported. Separately, re-annotate 100 random items per LLM-labeled task by human experts and measure agreement on the full label set, not just low-confidence cases.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that the benchmark is robust and that rankings such as 'DeepSeek R1 rivals proprietary models in reasoning but lags in judgment-based scenarios' are informative—depends on the validity of the labels and scores. Section 3.2 uses GPT-4o as LLM-as-a-Labeler for most tasks, with human review only for 'a subset of the annotated data that received low confidence scores.' Section 3.3 uses G-Eval with GPT-4o as judge for correctness, relevance, and coherence. Section 4.1 states that GPT-4o-2024-11-20 is both one of the evaluated models and the judge version. Thus the measurement instrument and a measured subject are the same system. Known LLM-judge biases, including self-preference, are cited in the paper (refs [6], [58], [61]), but no mitigation is reported: model identity is not blinded, response order is not swapped, and G-Eval scores are not validated against human ratings on a random sample. If GPT-4o systematically assigns higher G-Eval scores to its own outputs, every cross-model comparison on open-ended tasks (1-1, 1-2, 2-5, 3-1, 3-2, 3-3, 6-1) is confounded, and the 'overthinking' explanation for DeepSeek R1 in Task 5-1 is unsupported. The label-side problem is similar: ground truth for tasks such as 2-3, 2-4, 3-2, and 3-3 is LLM-generated, and only low-confidence labels received human review, leaving most labels unverified. These issues are not peripheral; they bear directly on whether the reported numbers measure anything stable.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a 14-task enterprise LLM benchmark organized by Bloom's Taxonomy, drawing on roughly 9,700 samples from Atlassian Confluence, Rovo chat, customer feedback, Slack queries, and developer documentation. Labels for most tasks are generated by GPT-4o via an LLM-as-a-Labeler pipeline with CRAG, and evaluation uses G-Eval with GPT-4o as judge. Six models are evaluated: Llama 3.2 3B, Llama 3.3 70B, Llama 4 Scout, DeepSeek R1, DeepSeek Distilled Llama 3.3 70B, and GPT-4o-2024-11-20. The central claims are that the benchmark is robust, that open-source models rival proprietary ones in reasoning tasks, and that DeepSeek R1 lags in judgment-based scenarios, attributed to overthinking.","tokens_in":15842,"tokens_out":3873,"duration_ms":40735,"significance":"If the benchmark and its results were valid, the paper would provide a reusable enterprise evaluation instrument and concrete guidance for model selection, which is a practically important contribution. The task taxonomy based on Bloom's Taxonomy, the use of real enterprise data, and the incorporation of CRAG for labeling are commendable design choices. However, the current evidence does not support the robustness claim because the measurement and the measured subject are the same system: GPT-4o generates most ground truth, serves as the judge, and is also one of the evaluated models. The paper also reports no dataset release, no error bars, and no human validation for the majority of labels. These limitations directly undermine the trustworthiness of the reported rankings and the 'overthinking' explanation, so the significance is conditional on major rework.","major_comments":[{"comment":"The evaluation loop is circular: GPT-4o is used as the LLM-as-a-Labeler for most tasks (Section 3.2), as the judge in G-Eval for correctness, relevance, and coherence (Section 3.3), and it is also one of the evaluated models (Section 4.1). The paper acknowledges in Section 2.3 that LLM judges exhibit self-preference and positional bias, but reports no mitigation such as blinding the model identity or swapping response order. Consequently, GPT-4o's higher scores on open-ended tasks (1-1, 1-2, 2-5, 3-1, 3-2, 3-3, 6-1) may reflect agreement with its own labels and judging criteria rather than true capability. This confounds every cross-model comparison, so the reported rankings are not trustworthy as evidence for the paper's conclusions.","section":"§3.2, §3.3, §4.1"},{"comment":"The paper states that 'human experts reviewed a subset of the annotated data that received low confidence scores,' but provides no details about the size of this subset, the selection criteria, or inter-annotator agreement. Because the majority of labels were never human-checked, the claim of a 'robust 9,700-sample benchmark' is unsupported. The authors should report the proportion of labels that received low confidence, the human correction rate, and ideally a random sample of high-confidence labels to verify that the LLM-generated ground truth is accurate beyond the low-confidence tail.","section":"§3.2 Human Validation"},{"comment":"G-Eval scores are reported without any validation against human judgments for these specific tasks. Citing the general G-Eval paper (ref [63]) is not sufficient; the authors need to show that GPT-4o's relevance, coherence, and correctness scores correlate with human ratings on a sample of the benchmark data. Without such calibration, differences like 0.88 vs. 0.87 in relevance or 0.97 vs. 0.91 in coherence in Table 2 cannot be interpreted as meaningful performance gaps, and the derived conclusions about model superiority are not supported.","section":"§3.3 and Table 2"},{"comment":"The paper claims that 'a sample size of approximately 600 provides stable evaluation results,' but no ablation study, confidence intervals, or error bars are presented anywhere. This claim is load-bearing because the benchmark's robustness argument rests on it, and two tasks (3-4 with 218 samples and 5-1 with 265 samples) fall well below that threshold. The authors should provide the ablation data, report variance estimates, and justify why the smaller manual-label tasks still yield stable comparisons.","section":"§3.1 Sample Size Claim"},{"comment":"The explanation that DeepSeek R1 lags in judgment-based scenarios 'likely due to overthinking' is speculative and unsupported by the reported data. No analysis of reasoning token counts, no comparison of R1 with its non-reasoning counterpart on task 5-1, and no ablation are provided; the cited reference [93] concerns agentic tasks and does not directly support this conclusion. The claim should be either supported with direct evidence or removed from the abstract and conclusions.","section":"§4.3 and Abstract"}],"minor_comments":[{"comment":"In the Data Curation section, 'similarly in Task 4-1 (\"NL2JQL\")' should reference Task 3-4, since NL2JQL is listed as 3-4 in Table 1.","section":"§3.2"},{"comment":"The text says 'seven popular LLMs' but only six models are listed in Section 4.1 and Table 2; the count should be corrected.","section":"§4.3"},{"comment":"For tasks 3-1 and 6-1, the labeling method is marked '-', and the text later says these tasks do not require ground truth; please clarify in the table caption or in Section 3.2 that 'no ground truth' is intended.","section":"Table 1"},{"comment":"The paper does not include a data availability statement or a link to the benchmark dataset, which is a significant omission for a benchmark paper aiming to provide a 'blueprint for enterprises'; consider adding a repository or an explicit statement about proprietary data restrictions.","section":"General"},{"comment":"The experimental setup reports temperature and top-p but does not specify the exact prompt templates or the G-Eval rubric details; providing these in an appendix or supplementary material would aid reproducibility.","section":"§4.2"}],"recommendation":"reject","confidential_remarks":"The circular evaluation design (GPT-4o as labeler, judge, and evaluated model) is a fundamental validity threat that cannot be fixed with minor edits. The paper would need a substantially different evaluation setup—such as using an independent judge with blinded model identities, full human validation of labels, error bars, and a released dataset—before it could be considered for publication. As submitted, the reported rankings and the 'overthinking' claim are not well supported, and the paper offers limited reproducibility due to the absence of data and prompts."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Hi,\n\nShort version: this is an honest applied-benchmark paper with a real design idea, but the reported numbers are not trustworthy because GPT-4o is simultaneously labeler, judge, and contestant. That is the one thing to know before reading the tables.\n\nWhat's new: a 14-task enterprise benchmark grouped by Bloom's taxonomy, with tasks drawn from actual Confluence/Rovo/Slack/developer-doc data. The sample-size ablation (about 600 per task) is a practical plus, and manually labeled NL2JQL and LLM-as-a-Judge tasks show the authors know when automatic labeling won't cut it. The related work is broad and appropriately cites the known limitations of LLM-as-a-judge, including position bias and leniency.\n\nThe soft spot is not minor. Most ground-truth labels come from GPT-4o, G-Eval scores come from GPT-4o, and GPT-4o is one of the six models. The paper does not blind model identity, swap response order, or validate G-Eval against a random sample of human ratings. Human review only touched low-confidence labels, so the bulk of the dataset is unverified. That makes the cross-model comparisons on open-ended tasks partly a measure of agreement with the labeler/judge. The 'overthinking' explanation for DeepSeek R1 in Task 5-1 is an unsupported hypothesis, not a finding. There are also small inconsistencies: Section 4.3 says seven models though six are listed, and the Table 1 sample sizes sum to about 9.3k, not 9,700. None of these are fatal to the benchmark concept, but they are fatal to the current version of the claimed results.\n\nI would believe this paper if the dataset were public, the judge were independent (or at least validated against a held-out human set), and the numbers were corrected. As is, the results are not independently checkable. I would not cite the rankings in their current form.\n\nRecommendation: desk-reject this version and invite a resubmission that includes the released benchmark and a blinded human validation set. The idea is worth a serious look; the evidence is not there yet.","headline":"Useful enterprise benchmark blueprint, but GPT-4o serving as labeler, judge, and contestant plus no released data makes the rankings unreliable as-is.","tokens_in":16339,"tokens_out":3348,"would_cite":false,"duration_ms":35749,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that a 14-task, Bloom's Taxonomy-based benchmark reveals enterprise-specific LLM gaps that more general benchmarks miss.","keywords":["Large Language Models (LLMs)","Evaluation Benchmark","Bloom's Taxonomy","LLM-as-a-Labeler","LLM-as-a-Judge","Corrective Retrieval-Augmented Generation (CRAG)","enterprise applications","model evaluation"],"falsifier":"Re-label a random sample of the 9,700 benchmark items entirely by human experts and recompute every model's score on that sample; if the rankings diverge from the GPT-4o-judged rankings, the reported comparisons collapse. A cheaper first check is to rerun the G-Eval scoring with labels and judge both produced by a different model family.","tokens_in":15297,"feed_emoji":"📊","tokens_out":5154,"duration_ms":50773,"temperature":0.7,"pith_summary":"The paper proposes a 14-task benchmark for evaluating LLMs in enterprise settings, organized by Bloom's six cognitive levels from Remember to Create. It argues that general-purpose benchmarks such as MMLU do not capture the proprietary knowledge and judgment demands of real enterprise work, and that its roughly 9,700-sample dataset, built with LLM-generated labels, corrective retrieval-augmented generation, and G-Eval judging, exposes those gaps. Evaluating six models, it reports that open-source DeepSeek R1 is competitive with proprietary models on reasoning tasks but lags on judgment-based ones, which it attributes to overthinking. If the benchmark is sound, it gives enterprises a reusable instrument for model selection and a blueprint for tailoring evaluations to their own data.","feed_headline":"14-task benchmark claims to reveal where enterprise LLMs actually fail","feed_subtitle":"Open-source DeepSeek R1 rivals proprietary models on reasoning but lags on judgment-based tasks.","key_machinery":"The central machinery is the task taxonomy paired with a data curation pipeline. The taxonomy maps 14 tasks onto six cognitive levels — Remember, Understand, Apply, Analyze, Evaluate, Create — each with a specific metric, such as G-Eval correctness for question-answering, exact match for named entity recognition, and Spearman's r for the judge-alignment task. The pipeline combines LLM-as-a-Labeler (GPT-4o) with corrective retrieval-augmented generation to reduce hallucination during annotation, then uses LLM-as-a-Judge through G-Eval in the DeepEval package to score outputs, with human review applied only to low-confidence labels. This machinery is what allows the benchmark to reach about 9,700 samples while claiming quality, and it also underpins the overthinking explanation, since reasoning models receive extra tokens for chain-of-thought yet still score lower on judgment tasks.","core_discovery":"The central claim is that an enterprise evaluation benchmark grounded in Bloom's Taxonomy can differentiate LLM capabilities in ways existing benchmarks do not. The paper reports that all six evaluated models score below 0.30 in G-Eval correctness on acronym memorization and factual question-answering over internal Atlassian data, indicating a systematic lack of proprietary enterprise knowledge; that open-source models win eight, tie one, and lose five comparisons against proprietary models overall; and that DeepSeek R1 leads in summarization and content generation while GPT-4o achieves the highest Spearman's r (0.47) on the LLM-as-a-Judge task, with reasoning models lagging there, likely due to overthinking. These results are presented as evidence that the benchmark surfaces actionable performance gaps and offers a practical blueprint for enterprise LLM evaluation and post-training decisions.","pith_inferences":["A direct test of judge bias would rerun the benchmark with labels and G-Eval scores produced by a non-GPT-4o judge; if GPT-4o's relative ranking drops, part of the reported judgment gap is an artifact of self-preference.","The paper validates only low-confidence labels by human review, so the accuracy of the majority of GPT-4o-generated labels remains unverified; a random fully human-labeled sample would quantify label reliability across all tasks.","Because the data come from a single company, the benchmark's transferability to other enterprises is open; porting the task templates to another organization's documents would test whether the Bloom's Taxonomy structure generalizes.","The overthinking explanation can be tested more directly by running DeepSeek R1 on the Evaluate-level task with chain-of-thought disabled or token-limited and checking whether its Spearman's r increases."],"forward_implications":["Enterprises should base model selection on task-level scores rather than overall averages, since no single model leads across all 14 tasks and the paper recommends against further post-training of Llama-3.2-3B-Instruct, the weakest performer.","Proprietary enterprise knowledge is a core bottleneck: all models score below 0.30 on acronym and factual question-answering tasks, supporting continuous pre-training on company-specific data.","The overthinking hypothesis implies that giving reasoning models more chain-of-thought tokens can hurt judgment-based performance, so deployments may need to constrain or calibrate reasoning budgets.","Open-source models are competitive enough on several tasks that enterprises can reduce reliance on proprietary APIs for those workloads, which the paper frames as a cost-saving opportunity."],"supporting_citations":[{"why":"Supplies the corrective retrieval-augmented generation method used in the data curation pipeline to reduce hallucination during labeling.","marker":"[52]"},{"why":"Introduces G-Eval, the chain-of-thought based LLM-as-a-Judge method the paper uses to score model outputs.","marker":"[63]"},{"why":"Provides the what/where/how evaluation framework that structures the benchmark's task construction, data construction, and metric calculation.","marker":"[65]"},{"why":"Defines the revised Bloom's Taxonomy that organizes the 14 tasks into six cognitive levels.","marker":"[66]"},{"why":"Documents the GPT-4 model family, the basis for GPT-4o, which serves as the LLM-as-Labeler and as the judge in G-Eval scoring.","marker":"[71]"},{"why":"The DeepEval package implements the G-Eval and other metrics the paper relies on for evaluation.","marker":"[72]"},{"why":"Supplies the overthinking concept used to explain why reasoning models underperform on judgment-based tasks.","marker":"[93]"},{"why":"Describes DeepSeek-R1, the open-source reasoning model that the paper compares against proprietary models.","marker":"[35]"}],"fun_headline_variants":["Open-source LLMs win reasoning, falter on judgment in enterprise test","DeepSeek R1 shines on reasoning but overthinks judge tasks","14-task benchmark finds enterprise LLMs lack internal knowledge","Bloom's Taxonomy benchmark: six models, gap in enterprise QA"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the labels produced by GPT-4o and the G-Eval scores it computes are accurate enough to serve as ground truth, even though GPT-4o is itself one of the evaluated models and only low-confidence labels receive human review.","fun_headline_variants_meta":{"raw":{"variants":["Open-source LLMs win reasoning, falter on judgment in enterprise test","DeepSeek R1 shines on reasoning but overthinks judge tasks","14-task benchmark finds enterprise LLMs lack internal knowledge","Bloom's Taxonomy benchmark: six models, gap in enterprise QA"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000737,"raw_usage":{"total_tokens":3250,"prompt_tokens":856,"completion_tokens":2394,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":472,"completion_tokens_details":{"reasoning_tokens":2322}},"tokens_in":472,"tokens_out":2394,"duration_ms":19574,"temperature":1.0,"reasoning_tokens":2322,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T22:52:09.988243+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-label a random sample of the 9,700 benchmark items entirely by human experts and recompute every model's score on that sample; if the rankings diverge from the GPT-4o-judged rankings, the reported comparisons collapse. A cheaper first check is to rerun the G-Eval scoring with labels and judge both produced by a different model family.","supporting_citations":[],"review_version":1}