{"id":"82ab2dea-75b7-49b6-92a2-900705e946f3","arxiv_id":"2506.13651","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"xbench is a new evaluation suite that scores AI agents on real-world Recruitment and Marketing tasks and uses Item Response Theory to track their capability growth over time.","lead":"xbench introduces two profession-aligned benchmarks, Recruitment and Marketing, that score AI agents on real headhunting and influencer-matching tasks defined with industry experts. The goal is to measure agent productivity and track capability growth over time, rather than isolated technical skills.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central claim that xbench scores strongly correlate with productivity value rests entirely on an unvalidated LLM judge; the paper defers the required human-alignment study, leaving all reported scores scientifically uninterpretable as productivity measures.","rationale":"The paper's central claim is that xbench metrics 'strongly correlate with productivity value' and that the reported scores establish baselines for professional domains. For this claim to hold, the LLM judge scores must faithfully reflect the quality and economic value of agent outputs. The paper provides no direct evidence for this: no human-expert agreement study, no correlation with business outcomes, and no cross-judge consistency analysis. Instead, Section 4.1 explicitly defers these validations. The reader's weakest-assumption analysis correctly identifies this unvalidated judge as the load-bearing point, and the paper's own limitation statement confirms it. My review does not find a different, more fundamental flaw: the task-collection methodology (live expert demands, anonymized real business cases) is a reasonable construction approach, and the transparency about future-work limitations is commendable. The conditional verdict remains appropriate because the benchmark infrastructure and baseline results are potentially valuable, but the productivity-value claim cannot be accepted without the deferred validation. The proposed test directly targets the missing evidence: a human-judge alignment study with a preregistered threshold, plus a judge-model swap to detect judge dependence. If the correlation with human experts is low or rankings shift across judges, the central claim would need to be substantially weakened. If the test passes, the benchmark would have much stronger support for its productivity-alignment claim.","tokens_in":18439,"tokens_out":3345,"duration_ms":33739,"concrete_test":"Run a human-expert validation study on a stratified random sample of 20 recruitment and 20 marketing tasks. Have 3–5 professional headhunters or marketing operators independently score agent outputs using the same rubrics (or rank them) and compute inter-rater agreement (e.g., quadratic weighted Cohen's kappa) and correlation with Gemini-2.5-Flash scores. Preregister a threshold (e.g., Spearman rho >= 0.7 and agreement on top/bottom agent ranks). Also swap the judge to a different model family (e.g., Claude or GPT-4o) on the same responses; if rankings change materially, the scores are judge-dependent and cannot support the productivity correlation claim.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract asserts that xbench 'creates metrics that strongly correlate with productivity value,' and Section 4.2 presents agent rankings as baselines for Recruitment and Marketing. The sole scoring mechanism is Gemini-2.5-Flash as LLM judge (Section 4.1), with marketing rubrics generated by LLM summarization of client-selected influencers (Section 3.3). The paper explicitly defers 'metric stability, alignment with human judgment, and consistency across different judge models' (Section 4.1). No data are presented linking xbench scores to any external measure of productivity or economic value, such as client re-selection rates, hiring outcomes, or time saved. Therefore the central claim is currently unfalsified: if the judge is biased toward certain response styles or model families, the rankings could misrepresent real-world value. The self-referential structure, where the same model family generates the ideal-influencer persona and scores candidates, compounds this risk. This is a correctness risk, not merely a missing nicety, because the paper's headline contribution is the productivity correlation, not the task collection itself.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces xbench, a profession-aligned evaluation suite for AI agents in two commercially significant domains: recruitment and marketing. For recruitment, 50 tasks were collected from headhunting business scenarios and divided into Company Mapping, People-to-Info, and Info-to-People categories. For marketing, 50 advertiser requirements are paired with a pool of 836 influencers, and agents recommend influencers for each campaign. All open-ended responses are scored by Gemini-2.5-Flash as an LLM judge, with recruitment tasks scored against human-annotated verifier answers and marketing tasks scored against an LLM-generated 'ideal influencer persona' derived from client-selected influencers. The paper reports leaderboards for nine agents (Tables 8 and 9), proposes an IRT-based xbench-index for tracking capability growth over time, and sketches technology-market-fit (TMF) curves. The intended contribution is a dynamic, profession-aligned benchmark whose metrics 'strongly correlate with productivity value' and can track real-world agent progress.","tokens_in":18689,"tokens_out":4454,"duration_ms":43327,"significance":"If the validity concerns are resolved, xbench would be a useful complement to capability-centric benchmarks: the tasks are grounded in real business workflows, co-constructed with domain professionals, run in live web environments, and the authors commit to continuous updates, which addresses the saturation and contamination problems of static benchmarks. The paper is also commendably transparent in acknowledging the potential bias of using Gemini-2.5-Flash as its judge. However, the central claim that the metrics 'strongly correlate with productivity value' is not demonstrated in the current manuscript. All reported scores come from a single LLM judge with no human validation, no inter-rater agreement, no repeated runs, and no linkage to any external productivity outcome such as client re-selection rates, hiring outcomes, or time saved. The marketing evaluation is additionally self-referential: the ideal influencer persona is generated by an LLM and then scored by the same model family. Consequently, the benchmark is currently best interpreted as an internally consistent LLM-opinion leaderboard, not as a validated measure of economic value.","major_comments":[{"comment":"The abstract claims that xbench 'creates metrics that strongly correlate with productivity value,' but the manuscript provides no evidence for this correlation. Section 4.1 states that 'alignment with human judgment' is deferred to future work, and all scores in Tables 8 and 9 are produced by a single LLM judge (Gemini-2.5-Flash) without human validation, inter-annotator agreement, or an external productivity criterion. As written, the reported scores are self-consistent LLM opinions, not validated measures of productivity. This is load-bearing because the paper's headline contribution is the productivity-value claim, not merely the task collection. I request either (a) a human validation study on a sample of tasks demonstrating that judge scores agree with expert ratings and correlate with at least one external productivity proxy, or (b) a revised abstract, title, and discussion that remove or substantially weaken the correlation claim.","section":"Abstract; §4.1"},{"comment":"The marketing evaluation is self-referential: the 'ideal influencer persona' is generated by an LLM from client-selected influencers, and the same model family (Gemini-2.5-Flash) is then used to score candidate influencers against that persona. If the LLM-generated persona is inaccurate, incomplete, or style-biased, the scores will reward conformity to the LLM's stereotype rather than to the client's actual preferences. The paper does not validate the LLM-generated rubric against the client's own criteria, nor does it measure agreement between the LLM persona and human experts. Because the marketing metric is described as an 'estimated re-selection rate,' this loop is load-bearing. A concrete remedy is to have human experts independently produce personas for a subset of campaigns from the same client selections and report agreement of the LLM judge with human judgments on the final influencer lists, together with any actual client re-selection data that can be disclosed.","section":"§3.3, Figure 5"},{"comment":"All reported leaderboard scores are point estimates from a single evaluation run, with no error bars, no repeated runs, and no significance tests. Several adjacent ranks are separated by very small margins (e.g., marketing ranks 2–4: 47.6, 46.5, 45.9; recruitment ranks 3–4: 61.4 for both), so the ordering may not be robust. Since the paper's stated purpose is to establish baselines for Recruitment and Marketing and to track progress over time, the absence of uncertainty quantification makes the ranking claims unsupported. Please report per-task score distributions, bootstrap confidence intervals, or repeated runs with different random seeds or judge calls.","section":"§4.2, Tables 8 and 9"},{"comment":"The IRT-based xbench-index and the technology-market-fit analysis are presented as deliverables in the abstract, but the manuscript contains no demonstration on xbench data. The OpenCompass validation in Figure 8 concerns standard LLM benchmark scores, not agent task scores in dynamic environments, and the text does not explain how IRT parameters (θ, a, b) are estimated from an incomplete score matrix, how identifiability is handled, or how changes in tasks and environments over time are modeled. Figure 9 is a schematic with no fitted curves or data. These proposals are reasonable future directions, but the abstract and summary should not present 'prediction of TMF' and 'tracking of product capabilities over time' as achieved results unless a concrete demonstration is added.","section":"§5.1, Figure 8, Figure 9"}],"minor_comments":[{"comment":"The section heading 'EVALUTIONS' contains a typo; it should read 'EVALUATIONS'.","section":"§4 heading"},{"comment":"Figure 6 is in the Marketing section and describes marketing task distribution, but its caption reads 'Task distribution across recruitment tasks’ categories and human time cost.' This appears to be a copy-paste error from Figure 4 and should be corrected.","section":"Figure 6 caption"},{"comment":"The list of evaluated agents in §4.1 does not include Grok3-Search, yet Grok3-Search appears in both Tables 8 and 9. Please clarify the evaluation setup and version used for Grok3-Search.","section":"§4.1 vs. Tables 8–9"},{"comment":"The text states that 'Items with a higher discrimination index a typically exhibit a gentler slope in relation to ability θ,' but in the logistic IRT model a higher a produces a steeper slope. Please correct this description.","section":"§5.1, Eq. (1)"},{"comment":"Table 10 appears to be a placeholder with colored cells but no actual data or legend explaining the colors. Either provide the actual available results or remove the table until data exist.","section":"Table 10"}],"recommendation":"major_revision","confidential_remarks":"The benchmark resource itself has potential value: the tasks are expert-grounded, live, and domain-specific, and the authors are transparent about the judge bias. However, the abstract's 'strongly correlate with productivity value' claim is currently unsupported by any evidence in the paper, and the marketing pipeline has a genuine self-referential structure. These are fixable in a revision that adds human validation, uncertainty quantification, and a more careful framing of the TMF/IRT portions. If the authors are unwilling to add such validation, the paper would need to be repositioned as an LLM-judged leaderboard rather than a productivity-value benchmark."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The real contribution here is the benchmark construction: 50 recruitment tasks (company mapping, people-to-info, info-to-people) and 50 marketing influencer-search tasks with 836 curated influencers, co-designed with practitioners and drawn from live business cases. That is concrete and valuable, and the baselines across ten agents are a useful starting point even if the rankings are predictable. The IRT-based tracking idea is also sensible, and the OpenCompass illustration shows it works on existing leaderboards.\n\nThe soft spots are where the abstract oversells. The claim that xbench scores 'strongly correlate with productivity value' is not demonstrated anywhere in the paper. All scoring uses Gemini-2.5-Flash as judge, with no human alignment, no error bars, and no cross-judge consistency checks. The authors explicitly defer that work, and I take them at their word, but the abstract should not outrun the evidence. The marketing loop is genuinely circular: the ideal-influencer persona is generated by an LLM from client-selected influencers, and then an LLM scores candidates against that persona. Same-family judging is a known risk, and they flag it, but it remains unmitigated. Task data and code are not released yet, which makes independent verification impossible.\n\nThe IRT section is more proposal than result: it is validated on OpenCompass, not on xbench itself, so it does not yet support the tracking claims. The TMF curves in Section 5.2 are entirely conceptual.\n\nDespite these issues, this is a serious paper. The benchmark design is thoughtful, the tasks are real, the authors are transparent about the judge limitation, and the failure modes are the standard ones for LLM-judged open-ended evaluations. It deserves a referee, and the likely path to acceptance is a major revision that adds human judgment alignment, judge-consistency checks, error bars, and a release of at least a sample of the task data. I believe the authors are positioned to supply that. The paper is not ready in its current form because the central productivity claim is unsupported, but the underlying asset is sound.","headline":"A genuinely useful profession-aligned benchmark pair, but the productivity-value claim in the abstract outruns the evidence.","tokens_in":19271,"tokens_out":1799,"would_cite":false,"duration_ms":20459,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A new benchmark scores AI agents on real recruitment and marketing work, aiming to predict their economic value.","keywords":["AI agent evaluation","profession-aligned benchmark","recruitment","influencer marketing","LLM-as-a-judge","item response theory","technology-market fit","productivity measurement"],"falsifier":"Run a field study where a group of professional headhunters and marketing operations staff independently score the same agent outputs using the same rubrics, then compare their scores to the LLM judge's scores; if agreement is low (e.g., correlation near zero) or if the judge consistently favors a particular model family, the claim that scores track productivity value is weakened.","tokens_in":18275,"feed_emoji":"📊","tokens_out":3492,"duration_ms":38350,"temperature":0.7,"pith_summary":"This paper introduces xbench, an evaluation suite designed to measure how much real-world productivity AI agents deliver in professional settings, rather than how well they perform isolated technical tasks. It argues that existing benchmarks fail to capture the economic value agents create, and that profession-aligned tasks defined by industry experts can close this gap. The first two implementations cover Recruitment, with 50 headhunting tasks, and Marketing, where agents match influencers to 50 advertiser briefs. The paper claims its metrics strongly correlate with productivity value, can predict technology-market fit, and can track capability growth over time. If correct, xbench could let businesses price agent services by benchmark score and help predict which agent products are commercially viable.","feed_headline":"Benchmark ties agent scores to real hiring and marketing value","feed_subtitle":"xbench scores agents on live headhunter and influencer-campaign tasks, then tracks progress with item response theory.","key_machinery":"The core machinery is the profession-aligned evaluation pipeline: tasks are collected live from expert business operations, categorized by feasibility and evaluability, and scored by an LLM judge that follows a chain-of-thought rubric (co-developed with domain experts) to produce a 1-5 score linearly mapped to 0-100. For Marketing, a rubric generator first summarizes the client-selected influencers into an 'ideal persona', then each agent-recommended influencer is scored against that persona. For tracking capabilities over time, the paper uses item response theory (IRT), modeling the probability of a correct response as $p(\\theta) = 1/(1+e^{-a(\\theta-b)})$, where $\\theta$ is agent ability, $b$ is item difficulty, and $a$ is item discrimination; this allows estimation of capability from incomplete evaluation results across product versions.","core_discovery":"The central claim is that profession-aligned evaluation, built from live business demands and scored with LLM-based judges using professional rubrics, produces benchmark scores that track the economic value agents deliver. For Recruitment, agents are tested on company mapping, people-to-info (completing a person's professional history), and info-to-people (finding people from constraints); for Marketing, agents recommend influencers from a curated pool of 836 candidates and are scored against the persona of influencers clients actually selected. The paper reports baseline results for leading contemporary agents, with the top-ranked agent scoring 78.5 on Recruitment and 50.8 on Marketing, and shows that scores vary by task theme in ways that highlight specific capability gaps. It also proposes item response theory to estimate underlying agent capability from an incomplete score matrix, so that capability growth can be tracked as both agents and evaluation sets evolve.","pith_inferences":["An implicit testable extension is to validate the claimed correlation by running a field study where agents with high xbench scores are used in real client engagements and measuring actual client re-selection rates or hiring outcomes; the paper defers this validation.","The same profession-aligned construction could be applied to other labor-intensive domains (e.g., legal research, financial analysis) where expert tasks are information-heavy and scorability is feasible, potentially broadening the suite into a general productivity index.","The IRT approach could be extended to predict future capability growth from early score trajectories, turning the benchmark from a measurement tool into a forecasting tool for agent development, though the paper does not yet demonstrate such forecasting.","Because the marketing benchmark scores agents on matching to an ideal persona derived from client selections, it implicitly assumes client selections are stable and rational; shifts in client preferences or influencer availability would require re-normalization of the persona."],"forward_implications":["If benchmark scores track productivity value, businesses can use xbench scores to price agent services and decide which agent products to adopt in recruitment and marketing workflows.","The IRT-based tracking method would let developers and investors observe capability growth of agent products over time, even as the underlying evaluation tasks and environments change.","Cost-performance analysis of benchmark scores against human labor cost could indicate when a domain reaches technology-market fit, meaning agents can deliver value cheaper and faster than human experts.","The baseline results suggest that end-to-end trained agents with strong search capabilities currently lead in these professional domains, while models with weaker search or shorter responses lag.","Continuous updates to the evaluation set are designed to reduce test-case leakage and keep scores aligned with a dynamic internet environment."],"supporting_citations":[{"why":"Supplies the observation that AI task-completion time approximately doubles every seven months, motivating the need to track productivity scaling.","marker":"(METR, 2025)"},{"why":"Provides the item response theory model used to estimate agent capability from an incomplete score matrix.","marker":"(Hambleton et al., 1991)"},{"why":"Provides the dynamically updated leaderboard evaluation data used to validate the IRT-based capability tracking method.","marker":"(Contributors, 2023)"},{"why":"Identifies the test-case leakage problem that motivates the design of dynamic, continuously updated evaluation sets.","marker":"(Xu et al., 2024)"},{"why":"Precedent for a dynamically updated benchmark with contamination-free test cases, used to motivate xbench's live evaluation approach.","marker":"(White et al., 2024)"},{"why":"Another precedent for dynamic evaluation sets that mitigate contamination, supporting the design of continuously updated evalsets.","marker":"(Jain et al., 2024)"},{"why":"Precedent for reporting demand curves, human capability curves, and technology supply curves in a performance-cost graph, used for the technology-market-fit analysis.","marker":"(Chollet et al., 2024)"}],"fun_headline_variants":["Benchmark scores agents on real hiring and marketing tasks","Agent eval mirrors real job value, not just tech chops","Profession-crafted tasks link agent scores to economic value","xbench tracks agent productivity with live business scenarios","From headhunting to influencers: benchmark measures real worth"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The scores produced by the LLM judge are assumed to accurately reflect the quality and economic value of an agent's work in recruitment and marketing, an assumption the paper explicitly defers validating against human judgment.","fun_headline_variants_meta":{"raw":{"variants":["Benchmark scores agents on real hiring and marketing tasks","Agent eval mirrors real job value, not just tech chops","Profession-crafted tasks link agent scores to economic value","xbench tracks agent productivity with live business scenarios","From headhunting to influencers: benchmark measures real worth"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000286,"raw_usage":{"total_tokens":1661,"prompt_tokens":905,"completion_tokens":756,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":521,"completion_tokens_details":{"reasoning_tokens":678}},"tokens_in":521,"tokens_out":756,"duration_ms":9254,"temperature":1.0,"reasoning_tokens":678,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T00:27:24.308433+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a field study where a group of professional headhunters and marketing operations staff independently score the same agent outputs using the same rubrics, then compare their scores to the LLM judge's scores; if agreement is low (e.g., correlation near zero) or if the judge consistently favors a particular model family, the claim that scores track productivity value is weakened.","supporting_citations":[],"review_version":1}