{"id":"336d6d41-97dd-4539-a614-72a9378e8b6d","arxiv_id":"2608.00298","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"WM-Cov defines testing adequacy for world-model-based driving simulation by separating requested, realized, and valid evidence and stopping when valid coverage saturates.","lead":"This paper proposes WM-Cov, a scoring layer that checks whether simulated driving tests from generative world-model simulators count as valid evidence of a failure. It reports coverage, valid failures, duplicates, artifacts, and stopping criteria instead of counting raw dangerous events.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Validity-label reliability is the load-bearing assumption; the author-verified audit's Cohen's kappa of 0.281 undermines the central 'valid evidence' claim.","rationale":"The reader identified the same weakest assumption, and the manuscript's own limitation statement confirms it. I agree this is the most load-bearing point. The framework design is otherwise coherent; the DriveArena realization counts (304/360 fully realized) are objective and do not depend on validity labels, but the selection and stopping conclusions do. The low kappa is not fatal by itself — some fine-grained taxonomies have low inter-rater agreement — but the paper makes no attempt to quantify label uncertainty in downstream metrics (e.g., by bootstrapping over label noise). Until an independent audit is done, the central comparison of WM-Cov against baselines in Tables VII and VIII is essentially an evaluation of the authors' own labels. The public code is a strength because it makes the proposed audit test feasible. I would keep the CONDITIONAL verdict and require the independent audit and parameter disclosure.","tokens_in":15379,"tokens_out":3663,"duration_ms":36047,"concrete_test":"Run an independent-labeling study: have ≥2 raters who are not authors apply Table II's validity ledger to the 141-row stratified sample (and ideally the full 1219-event TeraSim/SUMO pool and 1500-trace WM-like pool) using the public replication package. Compute exact agreement and Cohen's kappa per category. Then recompute the budget-100 selection (Eq. 14 / Table VII / Table VIII) and the Fig. 6 stopping point separately for each rater's labels. If kappa stays below ~0.6 or the selected valid-failure count/precision changes by >10%, the validity-ledger foundation for the central claim is not established. Also require the paper to report the exact ε_c, ε_f, ε_p, w values used for Fig. 6.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that WM-Cov's adequacy counts rest on 'valid interactive evidence' (Abstract, Eq. 3, Table II). The paper's own Section V-F audit reports exact agreement 0.411 and Cohen's kappa 0.281 on a 141-row sample, and the Limitations section states validity labels 'should be evaluated across ... independently audited cases' — i.e., no independent audit exists. Since the labels are assigned by the authors using a rubric and adjudicated by the authors, and the reported inter-rater agreement is low (especially at the duplicate boundary), the headline results (Table VII: 99/100 valid; Table VIII: 76/100; Fig. 6 stopping at 66 traces) are not reproducible by an independent tester. The whole value of WM-Cov over raw failure counts is that it filters 'artifacts' and 'duplicates'; if those categories are subjective, the filtering is subjective. This is not a mere calibration issue: duplicate-boundary disagreements can flip a trace from 'valid ADS failure' to 'duplicate failure', changing VF, FM, and Π_v. The stopping rule (Eqs. 8–10) is also under-specified — no values are given for ε_c, ε_f, ε_p, or w — compounding the reproducibility problem.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces WM-Cov, a provider-agnostic evaluation layer for interactive world-model-style autonomous-driving simulation. A scenario is formalized as an interactive family F(q) = {m, ρ, x0, e, π_bg, p, Z}, and each closed-loop trace is classified along requested, realized, and valid coverage dimensions. The framework defines a validity ledger, a budgeted trace-selection objective, and a stopping rule based on coverage growth, valid-failure discovery, and uncertainty. Experiments use an executed TeraSim/SUMO event pool, a WM-like mixed trace pool, and a real DriveArena TrafficManager–WorldDreamer matrix with UniAD and VAD planners. The headline results are that WM-Cov selects 99/100 valid failures on the executed pool and 76/100 on the WM-like pool with zero artifacts, and that on the DriveArena matrix 304/360 requested attempts are fully realized and 56 are partial. The central claim is that world-model-style testing adequacy should be judged by convergence of valid interactive evidence under budget rather than by raw generated failures or prompt coverage alone.","tokens_in":15941,"tokens_out":3107,"duration_ms":34085,"significance":"If the validity-labeling layer is reliable, WM-Cov addresses a real gap: modern generative simulators produce dangerous-looking traces whose evidential status is uncertain, and the paper provides a structured accounting mechanism (requested/realized/valid coverage, duplicate/artifact ledger, budgeted selection, stopping rule) that is provider-agnostic and reusable. The paper includes a public replication package and concrete provider-chain evidence from DriveArena, and its sensitivity and contamination stress tests are useful checks. However, the central claim depends entirely on the trustworthiness of the validity labels, and the paper's own audit reports low inter-rater agreement (Cohen's kappa 0.281). The stopping rule is also under-specified, with no concrete values for the tolerances or window size. These issues are load-bearing because every headline number—99, 76, 66—is a count of 'valid' evidence.","major_comments":[{"comment":"The central results (Tables VII and VIII, Figure 6) count 'valid ADS failures' and 'artifact/invalid' cases, but the validity labels are assigned by the authors using a rubric and audited only by the authors. The reported 141-row audit gives exact agreement 0.411 and Cohen's kappa 0.281, with the main disagreements at the duplicate boundary. A duplicate-boundary disagreement can flip a trace between 'valid ADS failure' and 'duplicate failure', directly changing VF, FM, and valid-evidence precision Π_v. The Limitations section itself states that labels 'should be evaluated across ... independently audited cases,' which confirms that no independent audit exists. Without an independent audit or a conservative duplicate-label variant, the headline precision-of-1.00 claims are not reproducible by an external tester. Please provide an independent audit (or at least a sensitivity analysis under","section":"Section V-F and Limitations (Section VIII)"},{"comment":"The stopping-oriented adequacy rule is formally defined but never instantiated. No values are given for the tolerances ε_c, ε_f, ε_p or the window size w, and Figure 6 merely says 'under a windowed rule.' This makes the stopping claim (66 traces) non-reproducible and untestable. Please specify the exact parameter values used for the figure, justify them, and report sensitivity of the stopping point and precision to these parameters. Without this, the paper's stopping-oriented adequacy contribution is only a schema, not an evaluated procedure.","section":"Section III-A, Eqs. (8)–(10), and Figure 6"},{"comment":"There is a circularity concern in the validation strategy. WM-Cov is validated on pools where the 'valid' / 'artifact' / 'duplicate' ground truth is itself assigned by the authors using the same ledger categories that WM-Cov is supposed to audit. The reported artifact precision/recall of 1.00 in the author-verified audit is therefore partly self-confirmation. The executed TeraSim/SUMO pool and the WM-like pool have no external ground truth. Please either use independently established labels, or clearly re-frame the evaluation as a consistency check of the accounting layer rather than as evidence that WM-Cov can distinguish valid from invalid evidence in an objective sense.","section":"Section V-F and Section VI (Tables VII–X)"}],"minor_comments":[{"comment":"The greedy selection weights β_r, β_q, β_c, β_m, β_a are not given concrete values. The sensitivity grid in Table IX varies only a subset of weights and only on the WM-like pool. Please report the exact default weights and include at least one sensitivity row on the executed TeraSim/SUMO pool.","section":"Section IV-D, Eq. (14)"},{"comment":"The heatmap in Figure 4 is not visually self-contained; the numeric realization rates are printed only in the table. Adding values inside the heatmap cells would improve readability. Also, the color scale is not defined in the caption.","section":"Table IV and Figure 4"},{"comment":"The feature vector ϕ(τ_i) = (o_i, u_i, r_i, b_i, f_i, v_i) is introduced before the ODD/interaction/risk/behavior taxonomy is described in Section IV-B. A forward reference or a brief definition of each component at Eq. (3) would help the reader.","section":"Section III, Eq. (3)"},{"comment":"The description says the UniAD default-cloudy 7.5-s cell produces 'partial artifacts but no UniAD planner responses.' It is not clear whether this is a provider-chain failure, a timeout, or a semantic mismatch. Please clarify the concrete failure mechanism, since this cell is used as a key example of requested-to-realized accounting.","section":"Section V-B"},{"comment":"Several references have future dates (e.g., [6], [12], [13], [14], [15]) relative to the arXiv submission. If these are accepted or online-first papers, please add DOIs or publication status; if they are preprints, mark them accordingly.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The core adequacy-accounting idea is reasonable and timely, and the DriveArena requested/realized counts are a concrete contribution. The main risk is that the paper's quantitative claims are built on author-assigned validity labels with low inter-rater reliability (kappa 0.281). I would like to see either an independent audit or a substantial weakening of the headline 'valid failure' and 'precision 1.00' claims. If the authors can provide independent labels or a conservative duplicate-boundary sensitivity analysis, the paper could be suitable for acceptance after revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe short version: this paper has a good idea — count valid interactive evidence instead of raw dangerous rollouts — but its own numbers show the validity labels aren't reliable enough to support the conclusions. Worth reading, worth citing with caution, and worth sending back for a serious revision rather than desk rejecting.\n\nWhat's actually new: the requested/realized/valid evidence split, the provider-agnostic adapter schema, and the stopping-oriented adequacy report. Scenario coverage work talks about coverage; world-model papers talk about generation; this is the first to define when a trace counts as evidence for a given planner and intent. The DriveArena evaluation is genuine provider-chain accounting: 360 requests, 304 fully realized, 56 partial, with a disjoint 80-request slice giving similar numbers. That part is concrete and reproducible. Code is public.\n\nThe soft spots are real. The validity ledger is the load-bearing piece, and the 141-row author-verified audit reports exact agreement 0.411 and Cohen's kappa 0.281. Those are low. If independent raters can't agree on what counts as a valid ADS failure versus a duplicate or artifact, then every downstream number — 99/100, 76/100, stop at 66 — inherits that subjectivity. The paper itself says independent audit is needed, but that's a limitation, not a fix.\n\nSecond, the stopping rule is under-specified. No values are given for epsilon_c, epsilon_f, epsilon_p, or the window w. So the stopping at 66 traces in Figure 6 isn't reproducible as stated. Same for the selection weights — they report sensitivity over 16 settings, but the actual parameter values aren't pinned down.\n\nThird, part of the validation is circular. The selection results are measured against labels assigned using WM-Cov's own ledger categories. That doesn't sink the paper, because the DriveArena realization counts are independent and support the accounting layer. But it means the headline claims about 'valid evidence precision' are not independently confirmed.\n\nOverall, the central argument — that adequacy should be judged by convergence of valid interactive evidence under budget — is sound and important for this subfield. The measurement instrument needs work. I'd send it to peer review and ask the authors for an independent validity audit and exact stopping-rule parameters. It's a useful paper for anyone building or evaluating world-model-based AV testbeds.","headline":"Valid evidence is the right thing to count, but with kappa 0.281 on the validity labels, the counting instrument needs independent calibration before the headline numbers can be trusted.","tokens_in":16336,"tokens_out":2180,"would_cite":true,"duration_ms":19729,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"World-model-based autonomous-driving testing should be judged by convergence of valid interactive evidence under budget, not by raw generated failures or prompt coverage.","keywords":["autonomous driving","world models","test adequacy","scenario coverage","validity audit","budgeted selection","simulation testing","stopping rule"],"falsifier":"An independent blind audit of the same traces using the paper's rubric; if inter-rater agreement remains low (e.g., kappa below ~0.4) and the set of traces labeled as valid failures changes materially across raters, then WM-Cov's reported valid-failure counts and stopping points are not reproducible. Concretely: have two raters label the 141-row stratified audit sample and compute the overlap of the resulting valid-evidence sets—if overlap is low, the framework's adequacy conclusions are labeler-dependent.","tokens_in":15325,"feed_emoji":"✅","tokens_out":4913,"duration_ms":45005,"temperature":0.7,"pith_summary":"This paper argues that when world-model simulators generate interactive, planner-conditioned driving scenarios, testers cannot count raw collisions or near-misses as evidence. It introduces WM-Cov, a provider-agnostic evaluation layer that converts raw traces into requested, realized, and valid realized coverage, along with a validity ledger that filters artifacts and duplicates. The authors show in evaluations that dangerous-looking events often include artifacts, duplicates, or partial realizations, so adequacy should be measured by convergence of audited valid evidence, not by raw failure counts. If right, this gives the AV testing community a stopping rule: keep generating until new valid evidence saturates rather than chasing a fixed number of failures.","feed_headline":"World-model driving tests need validity ledgers, not crash counts","feed_subtitle":"WM-Cov separates requested, realized, and valid evidence so testers know when they have enough coverage to stop.","key_machinery":"The key machinery is WM-Cov's requested–realized–valid coverage ledger: three counters that separate what the test asked for, what the provider actually produced, and what passes an audit as plausible, non-duplicate, non-artifact evidence. The ledger feeds a budgeted greedy selector score that rewards marginal valid coverage gain, risk, and realism while penalizing redundancy and artifact risk, and a windowed stopping rule that halts when coverage growth, new valid failure modes, and uncertainty-width thresholds saturate.","core_discovery":"The paper's central claim is that test adequacy for interactive world-model-style simulators is not a property of the generator's output volume or the number of prompts that were answered, but of a triple: requested coverage (what the test intended), realized coverage (what actually happened in closed loop), and valid realized coverage (what is plausible, replayable, non-duplicate, and relevant to the ADS under test). WM-Cov is the accounting layer that tracks this triple and reports it via coverage growth, valid-failure discovery, failure-mode diversity, realism, artifact suppression, duplicate accounting, and valid-evidence precision. The empirical evidence—from an executed simulator event","pith_inferences":["The framework implies that the quality of a generative simulator as a test oracle can be measured by its valid evidence yield per rollout, a metric the authors do not name but that follows directly from their validity ledger.","If independent raters reproduce the low agreement seen in the author audit (exact 0.411, Cohen's kappa 0.281), then WM-Cov's outputs depend on who performs the audit; the natural extension is an automated validity classifier trained on the rubric.","The requested-to-realized gap could itself become a quality signal for prompt conditioning: a drop in realization rate at longer horizons, as in the paper's 7.5-second cells, indicates the provider chain loses conditionality over time.","A testable extension: apply WM-Cov's ledger to other closed-loop simulators (e.g., rule-based traffic simulators) to confirm the validity categories transfer beyond the providers studied."],"forward_implications":["Test campaigns should report valid-evidence precision and artifact counts, not just collision rates.","Stopping decisions can be based on evidence-saturation windows, reducing wasted simulation budget.","Requested-to-realized tracking exposes provider-chain failures that raw request counts hide, such as horizon- and planner-conditioned partial realizations.","Generators that produce many dangerous-looking traces but low valid yield are identifiable and can be penalized as test infrastructure.","Adequacy reports become a standard contract: coverage + validity + budget, enabling cross-provider comparison."],"fun_headline_variants":["Validity beats volume in world-model driving tests","WM-Cov: adequacy from valid evidence, not crash volume","Interactive sim tests: count valid evidence, not failures","Driving sim adequacy: requested vs realized vs valid","World-model testing: track valid evidence to know when to stop"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that the validity labels—which determine what counts as valid evidence—are reliable and reproducible; the paper's own author-verified audit reports exact agreement of 0.411 and Cohen's kappa of 0.281, so if independent raters cannot reproduce these labels, every downstream coverage and stopping conclusion collapses.","fun_headline_variants_meta":{"raw":{"variants":["Validity beats volume in world-model driving tests","WM-Cov: adequacy from valid evidence, not crash volume","Interactive sim tests: count valid evidence, not failures","Driving sim adequacy: requested vs realized vs valid","World-model testing: track valid evidence to know when to stop"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000203,"raw_usage":{"total_tokens":1250,"prompt_tokens":797,"completion_tokens":453,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":541,"completion_tokens_details":{"reasoning_tokens":388}},"tokens_in":541,"tokens_out":453,"duration_ms":4325,"temperature":1.0,"reasoning_tokens":388,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T00:46:00.360141+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"An independent blind audit of the same traces using the paper's rubric; if inter-rater agreement remains low (e.g., kappa below ~0.4) and the set of traces labeled as valid failures changes materially across raters, then WM-Cov's reported valid-failure counts and stopping points are not reproducible. Concretely: have two raters label the 141-row stratified audit sample and compute the overlap of the resulting valid-evidence sets—if overlap is low, the framework's adequacy conclusions are labeler-dependent.","supporting_citations":[],"review_version":1}