{"id":"788c26cc-c394-4708-ac3a-7d87cef0d545","arxiv_id":"2504.16778","paper_version":2,"verdict":"UNVERDICTED","confidence":"MODERATE","novelty_score":2.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"A white paper recommending a shift from static benchmark-based evaluation to holistic, dynamic, outcome-oriented evaluation of generative AI systems in real-world contexts.","lead":"This white paper argues that static benchmarks do not capture how generative AI systems perform in real-world use, and proposes a framework for continuous, human-in-the-loop evaluation. It is a position piece for practitioners and policymakers, illustrated with two case studies.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Central claim rests on an unvalidated premise: the paper asserts that dynamic, outcome-oriented, human-in-the-loop evaluation tracks real-world performance better than static benchmarks, but provides no empirical evidence and its own caveats concede human evaluation is slow, subjective, and biased.","rationale":"The paper is a white paper, so I read it as a set of recommendations rather than a falsifiable empirical claim. The strongest claim is about evaluation practice: static benchmarks leave a lab-to-real gap, and a holistic, dynamic, human-in-the-loop framework closes it. For that claim to be true, two empirical conditions must hold: benchmark performance must be a poor predictor of real-world outcomes in at least some important deployment contexts, and the proposed alternative must be a better predictor at acceptable cost. The manuscript argues the first condition with conceptual examples and illustrates the second with two case studies, but it provides no measurements of either condition. I do not treat disagreement with current consensus as a concern by itself; the concern is internal support. The text itself flags the key weakness: human-centered evaluation is 'slow and subjective' and 'influenced by individual biases,' and 'Evaluating the Evaluations' is listed as an open need rather than a solved validation step. Without a validation protocol, a policymaker or practitioner cannot know whether the proposed framework improves on benchmarks or simply replaces a known bias with an unmeasured one. No machine-checked proofs, code, or data accompany the paper; the case studies are illustrative. The reader's UNVERDICTED verdict is therefore appropriate, and my concern does not move it. A concrete shadow-study comparison in one case-study domain would settle whether the premise holds in at least one representative setting.","tokens_in":17658,"tokens_out":3460,"duration_ms":32384,"concrete_test":"Run a small shadow study in one of the two case-study domains: deploy two candidate LLMs for clinical note summarization, record static benchmark scores (e.g., ROUGE or LLM-as-judge) alongside outcome-oriented metrics (clinician time saved, action-item completeness, patient comprehension), and compare model rankings. If benchmark ranking matches outcome ranking, the central motivation gap is unsupported in that domain; if rankings diverge and the outcome metrics show acceptable inter-rater reliability and predictive validity, the framework gains preliminary support.","verdict_should_be":"UNCHANGED","load_bearing_attack":"To make the central claim hold, the paper would need to show that (1) static benchmark performance diverges from real-world outcomes in a way that matters, and (2) the proposed dynamic, outcome-oriented, human-in-the-loop framework has better criterion validity than benchmarks, at acceptable cost and subjectivity. Neither is established. The 'In-the-wild evaluation' section motivates the gap via a validity mismatch and incomplete coverage, but offers no empirical demonstration; the clinical case study asserts that 'time saved by clinicians, improvement in patient outcomes, and cost savings' are not captured by summary quality evaluations, but does not measure them. The 'Human-centered ML Evaluation' section simultaneously states that human-centered evaluations 'tend to be slow and subjective, relying on human intervention and influenced by individual biases,' and the 'LLM as a Judge' section lists biases of automated judges. The 'Evaluating the Evaluations' section says benchmarks should correlate with human judgment of usefulness, but the paper never applies this validity check to its own proposed methods. Thus the load-bearing assumption is that the new framework's outcome metrics are measurable and better aligned with real-world priorities; the text acknowledges measurement difficulties but supplies no validation protocol or data. This is a correctness risk for the prescription, not merely a stylistic gap.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This white paper argues that existing static benchmark-based evaluation is inadequate for generative AI (GenAI) systems deployed in real-world settings, and proposes a framework based on holistic, dynamic, continuous, and human-in-the-loop evaluation. It organizes the discussion around three questions—what is being evaluated, who evaluates and how, and how long evaluation remains relevant—and supplements this with two case studies (clinical note summarization and content moderation) plus audience-specific recommendations for practitioners, policymakers, business leaders, evaluation designers, and funding agencies.","tokens_in":18009,"tokens_out":5523,"duration_ms":49281,"significance":"If the central claim were established, the paper would provide a useful synthesis of current evaluation concerns and a plausible template for outcome-oriented evaluation. Its strengths are its coherent taxonomy, its candid discussion of tradeoffs (e.g., human evaluation being slow and subjective, LLM-as-judge biases, estimation error in sampling), and its recognition that evaluation must be maintained over time to avoid saturation and leakage. However, the paper is a position statement rather than an empirical or methodological contribution: it presents no data, no validation protocol, and no reproducible artifact, and it does not apply its own stated validity criteria to the framework it recommends. The paper is therefore best viewed as a starting point for a research agenda, not as a demonstrated evaluation method.","major_comments":[{"comment":"The central premise, stated in the Executive Summary ('these static evaluations often fail to capture how models perform in real-world scenarios') and elaborated in the 'In-the-wild evaluation' section, is asserted rather than demonstrated. The paper provides no empirical example, prior study, or dataset showing a divergence between benchmark scores and consequential real-world outcomes for GenAI models, nor does it provide evidence that the proposed continuous, outcome-oriented metrics track real-world performance better than static benchmarks. This is load-bearing: if the gap is not real or if the proposed metrics do not close it, the framework loses its rationale. Please add either a systematic review of documented benchmark-deployment divergence or an original measurement in one of the case studies, and state explicitly what evidence would falsify the central claim.","section":"Executive Summary; In-the-wild evaluation"},{"comment":"The paper's recommended shift to human-in-the-loop and LLM-as-judge evaluation is not supported by evidence of criterion validity, and the paper's own caveats cut against it. 'Human-centered ML Evaluation' concedes that such evaluations 'tend to be slow and subjective, relying on human intervention and influenced by individual biases,' and 'LLM as a Judge' documents self-preference and style biases in automated judges. The paper does not explain how the proposed framework mitigates these problems (e.g., inter-rater reliability targets, bias audits, cost-effectiveness thresholds) or why the residual subjectivity is acceptable in high-stakes settings. Without such a validity and cost argument, the recommendation is an appeal to best practice rather than an evidenced conclusion.","section":"Human-centered ML Evaluation; LLM as a Judge"},{"comment":"This section correctly says that benchmarks should 'correlate with human judgment of the models' usefulness in real-world applications,' and that metrics must be reliable and reproducible. The paper never applies these same criteria to its own outcome-oriented, human-in-the-loop metrics: no correlation analysis, reliability estimate, or validation protocol is given for 'time saved by clinicians,' 'improvement in patient outcomes,' or 'moderator well-being.' The framework needs an explicit meta-evaluation plan that would verify the proposed metrics against independently measured real-world outcomes, and the manuscript should report at least one such check or clearly mark it as a required future step.","section":"Evaluating the Evaluations"},{"comment":"Both case studies illustrate the framework's themes but do not provide evidence that the approach works. In 'Health AI: Clinical Note Summarization,' the paper asserts that hospital priorities (time saved, patient outcomes, cost savings) are 'not captured by summary quality evaluations,' but it does not measure these outcomes, describe how they would be operationalized, or discuss how to separate model effect from confounding workflow changes. In 'Content Moderation,' the call for evaluators from 'different types of backgrounds' lacks a concrete aggregation procedure for conflicting judgments and a scaling strategy. Case studies should be expanded into worked examples with realistic measurement protocols, or reframed as hypotheses to be tested rather than demonstrations.","section":"Case Studies: Health AI; Content Moderation"}],"minor_comments":[{"comment":"The sentence 'it is important to ensure the same model is used to produce the assessment' appears to contradict the previous section's advice to use a different LLM from the generator; this is likely a typo and should be corrected.","section":"Evaluating the Evaluations"},{"comment":"The claim that existing literature 'often falls short by providing coarse estimations of energy consumption' and the example of a 200B-parameter model using ~11.9 GWh are presented without citations; these empirical assertions need references.","section":"Power, energy, and sustainability considerations"},{"comment":"The term 'temporal incongruence' is used without definition or example; please define it or provide a reference.","section":"Dynamic Evaluation + Benchmark Automation"},{"comment":"The bulleted list in this section uses the symbol 'ο' for each item, which renders as a Greek letter rather than a standard bullet; this formatting issue should be fixed.","section":"In-the-wild evaluation"},{"comment":"The recommendations list separate actions for five audiences but do not address how to reconcile conflicts among them, such as the tension between regulatory transparency and proprietary model secrecy, or between continuous evaluation and cost.","section":"Summary of Recommendations"}],"recommendation":"major_revision","confidential_remarks":"This is a position/white paper rather than a standard experimental contribution. The authors are candid about tradeoffs, and the synthesis is coherent. For a peer-reviewed CS journal, the lack of evidentiary support for the main claim is the central obstacle; a major revision that adds a validation section, literature evidence, and worked case studies would make it publishable. If the journal does not normally accept position papers of this scope, the editor may consider whether a magazine or opinion venue is more appropriate."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a well-organized position paper that says most of the right things about why static benchmarks are insufficient for deployed GenAI, but it never establishes that its proposed alternative actually tracks real-world performance better. So treat it as a policy document, not a research contribution.\n\nWhat it does well: the taxonomy (in-the-lab, human-capability, in-the-wild) is clear and useful. The two case studies—clinical note summarization and content moderation—are concrete and illustrate real tradeoffs, like the hospital procurement officer who cares about time saved and patient outcomes, not ROUGE scores. The paper also deserves credit for acknowledging the costs and biases of human evaluation, the limits of LLM-as-judge, and the risk of data leakage. The “Evaluating the Evaluations” section is the most useful part, because it points to validity checks like whether benchmark scores correlate with human-perceived utility.\n\nNow the soft spots. The central premise is exactly what the stress-test says: the paper asserts that dynamic, outcome-oriented, human-in-the-loop evaluation will track real-world performance better than static benchmarks, but gives no evidence that this is true, measurable, or affordable. In fact, its own caveats—human evaluation is slow, subjective, biased; LLM judges have their own biases—undercut the prescription. The case studies mention metrics like cost savings and moderator well-being but do not measure them or show how they would be validated. Oddly, the paper says benchmarks should be checked for correlation with human judgment of usefulness, but it never applies that same validity check to its own proposed framework. That is a real gap, and it is load-bearing for a paper that tells people to change how they evaluate.\n\nAlso, the manuscript has no reference list. For a white paper that synthesizes many well-known critiques, that is unusual and makes it harder to track which claims come from where. Minor edit: the clinical case study’s “How long should we evaluate?” section repeats a paragraph verbatim.\n\nWho is this for? Practitioners and policymakers who want a checklist of considerations before designing an evaluation, not researchers looking for a new method or result. I would bring it to a reading group as a discussion piece, but I would not cite it in my own work because it offers no citable evidence.\n\nRecommendation for peer review: if the venue is a research journal, desk-reject—this is not a research paper. But for an outlet that accepts position papers, send it to referees with instructions that the key claim needs at least a validation sketch or one empirical illustration; otherwise the recommendations are just good intentions.","headline":"A coherent white paper that synthesizes known evaluation critiques, but the central claim that dynamic, outcome-oriented evaluation works better is asserted, not shown.","tokens_in":18438,"tokens_out":2821,"would_cite":false,"duration_ms":28668,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Static benchmarks systematically fail to capture how generative AI behaves in real deployments, this white paper argues; evaluation must become continuous, outcome-oriented, and human-in-the-loop.","keywords":["generative AI evaluation","in-the-wild evaluation","human-in-the-loop","dynamic benchmarks","LLM as a judge","outcome-oriented metrics","AI policy","AI safety and fairness"],"falsifier":"Take one deployed GenAI application, record its static benchmark scores alongside real-world outcome metrics such as clinician time saved, content-moderation error costs, or user trust, and test whether a continuous human-in-the-loop evaluation predicts those outcomes and triggers corrective action better than the benchmark scores; if benchmark scores predict outcomes at least as well, the paper's central claim would be wrong.","tokens_in":17480,"feed_emoji":"🧭","tokens_out":3411,"duration_ms":33546,"temperature":0.7,"pith_summary":"This white paper argues that current model evaluation, built on standardized benchmarks and fixed datasets, systematically under-measures generative AI in real deployment and can even mislead. It proposes a framework organized around three questions: what is being evaluated, who evaluates and how, and how long the evaluation remains valid. The central claim is that meaningful evaluation must be holistic, dynamic, continuous, and human-in-the-loop, integrating performance, fairness, ethics, efficiency, and societal impact. Practitioners are directed to design context-specific metrics and workflow-aware studies, while policymakers are directed to regulate outcomes and societal impacts rather than model parameters. A fair reader takes away a checklist for building evaluation plans that track evolving, real-world use rather than a single reproducible score.","feed_headline":"Static benchmarks miss how AI behaves in the wild","feed_subtitle":"A white paper argues for continuous, outcome-focused, human-in-the-loop evaluation of deployed generative AI.","key_machinery":"The organizing framework is a three-question structure: 'What is being evaluated?', 'Who evaluates and how?', and 'How long is the evaluation relevant?' Each question carries design principles: choose metrics from the deployment context, include human and domain expertise alongside automated scaling, watch for LLM-as-judge bias, treat benchmarks as rolling artifacts that must be refreshed, validate scores against human judgment, and check for data leakage. This structure does the argumentative work by turning 'evaluate in the wild' from a slogan into a sequence of decisions that practitioners, policymakers, and funders can make.","core_discovery":"The paper's core claim is that there is a structural gap between lab-tested performance and real-world outcomes for generative AI, caused by static benchmarks' limited coverage, saturation, contamination, and mismatch with deployment contexts. To close this gap, the paper defines in-the-wild evaluation as evaluation tailored to a specific practical use case and stakeholder priorities, with three desiderata: contextually appropriate metrics, capture of unintended impacts, and attention to workflow effects. It then maps the evaluation space along two axes: who evaluates (automated benchmarks, human stakeholders, LLM-as-judge, and combinations) and how long evaluation remains relevant (dynamic benchmarks, continuous monitoring, and evaluating the evaluations). The discovery is not a new metric but a re-framing: evaluation should be an ongoing, outcome-oriented, stakeholder-inclusive process rather than a one-time benchmark.","pith_inferences":["The paper leaves open how continuous evaluation is funded; I infer that evaluation will increasingly become a service-layer activity, similar to auditing, with independent evaluators certifying deployed systems.","A testable extension of the framework is that the divergence between static benchmark rankings and in-the-wild rankings grows with task complexity and subjective disagreement; this could be measured directly across high-stakes and low-stakes applications.","I infer that benchmark scores will be reinterpreted as calibration signals rather than endpoints, so a model's raw score matters less than how the model changes outcomes in the specific workflow where it operates."],"forward_implications":["Practitioners will need continuous monitoring plans with thresholds tied to outcome metrics, not just accuracy scores.","Policymakers can regulate outcomes such as fairness, transparency, and environmental impact instead of model size or architecture, keeping rules relevant as AI changes rapidly.","Benchmark builders should treat benchmarks as rolling artifacts and use automated refreshment to prevent saturation, leakage, and obsolescence.","Evaluation budgets must include human expert time because automated methods alone cannot supply context-aware judgment about values, workflows, and unintended impacts.","Procurement decisions should shift from reading model leaderboards to reviewing workflow-specific evidence about how a model changes real-world outcomes."],"supporting_citations":[],"fun_headline_variants":["Benchmarks don't predict real-world AI behavior","End static benchmarks: AI evaluation must be continuous","In-the-wild AI needs ongoing, human-checked evaluation","Why fixed benchmarks fail deployed AI models","Evaluate AI like it's deployed: dynamic and outcome-focused"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"Outcome-oriented, dynamic, human-in-the-loop evaluation measures real-world performance more accurately than static benchmarks, and the extra cost and subjectivity are acceptable trade-offs.","fun_headline_variants_meta":{"raw":{"variants":["Benchmarks don't predict real-world AI behavior","End static benchmarks: AI evaluation must be continuous","In-the-wild AI needs ongoing, human-checked evaluation","Why fixed benchmarks fail deployed AI models","Evaluate AI like it's deployed: dynamic and outcome-focused"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000323,"raw_usage":{"total_tokens":1773,"prompt_tokens":863,"completion_tokens":910,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":479,"completion_tokens_details":{"reasoning_tokens":836}},"tokens_in":479,"tokens_out":910,"duration_ms":7793,"temperature":1.0,"reasoning_tokens":836,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T10:54:59.825632+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take one deployed GenAI application, record its static benchmark scores alongside real-world outcome metrics such as clinician time saved, content-moderation error costs, or user trust, and test whether a continuous human-in-the-loop evaluation predicts those outcomes and triggers corrective action better than the benchmark scores; if benchmark scores predict outcomes at least as well, the paper's central claim would be wrong.","supporting_citations":[],"review_version":1}