{"id":"a8b0f0c8-da68-449f-937c-aaf9d23329cc","arxiv_id":"2512.10958","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A five-aspect, 24-metric benchmark, a 26K human-annotated dataset, and an AI evaluator show that today's driving world models cannot simultaneously look real, respect geometry, and behave safely.","lead":"WorldLens is a new testing suite for AI driving world models — systems that generate realistic driving videos. It scores models on 24 measures across five areas (visual quality, 3D reconstruction, reaction to driving actions, real perception tasks, human opinion) and finds that no current model wins everywhere: the most realistic-looking ones often violate physics or drive unsafely.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Downstream-task scores may measure detector domain brittleness rather than world-model fidelity; the claim that OpenDWM's realism coincides with poor usability is not yet controlled.","rationale":"The reader's weakest_assumption already identifies the core issue: evaluation aspects may measure the external toolchain rather than the world model itself. My stress-test converges on the same point, with the downstream-task aspect as the most load-bearing instance because it directly supports the paper's headline trade-off between appearance and behavior. If BEVFusion et al. are brittle to OpenDWM's synthetic distribution, then the NDS/AMOTA rankings in Table 3 do not establish that OpenDWM is less usable; they establish only that one perception stack was not adapted to that model's outputs. This would not destroy the benchmark's framework—the protocols remain coherent and the appendix is unusually detailed—but it would weaken the paper's central empirical generalization and the specific 'usability' conclusion. The proposed test is decisive because it either reproduces the gap under a controlled domain shift or shows the ranking is detector-dependent. I agree with the reader's conditional verdict: the concern is substantive but addressable, so no verdict change is needed.","tokens_in":53029,"tokens_out":5382,"duration_ms":60362,"concrete_test":"Re-run the D.2 downstream evaluation under two conditions: (i) replace BEVFusion with a different frozen camera-only detector (e.g., BEVFormer) and re-rank the six world models; (ii) apply BEVFusion to real nuScenes frames perturbed to match each world model's low-level statistics (color, contrast, resolution, artifact profile). If OpenDWM's ~11-point NDS deficit vs DiST-4D disappears with the alternate detector, or if the same deficit appears on perturbed real frames, then §5.3's 'perceptual quality does not imply usability' conclusion is a toolchain artifact rather than a property of the world models.","verdict_should_be":"UNCHANGED","load_bearing_attack":"WorldLens's central empirical claim—'Perceptual Quality Does Not Imply Usability' (§5.3)—rests on Aspect 4 (Downstream Task, §3.4). OpenDWM attains the best Subject Fidelity (G.1 = 36.30, Table 1) but NDS = 21.96% vs DiST-4D's 33.22% (Table 3 / Table 21), interpreted as 'large-scale multi-domain training can hinder adaptation.' However, every D.x metric uses a single frozen perception toolchain: BEVFusion for D.1/D.2, ADA-Track for D.3, SparseOcc for D.4 (§10.1–10.4). A detector trained on real nuScenes frames can degrade on OpenDWM's particular color/contrast/resolution statistics regardless of scene fidelity; no control perturbs real frames to match each world model's synthetic distribution. The same toolchain confound affects Action-Following (LimSim and the DriveArena protocol co-determine A.3/A.4) and Reconstruction (the OmniRe optimizer co-determines R.1/R.2, §8.1.3). Since the headline finding—'no existing world model excels universally'—depends on comparing Generation scores against these toolchain-mediated behavior scores, the measurement validity of the central claim is not established. The manuscript's limitation section (§13.3) does not flag this confound.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"WorldLens proposes a five-aspect benchmark (Generation, Reconstruction, Action-Following, Downstream Task, Human Preference) with 24 metrics for evaluating driving world models. On six recent models (MagicDrive, DreamForge, DriveDreamer-2, OpenDWM, DiST-4D, X-Scene), the paper reports that no model dominates across all aspects: OpenDWM leads subject fidelity while DiST-4D tends to lead geometry, reconstruction, and downstream perception; open-loop action-following is moderate, closed-loop route completion is uniformly low (6.89–13.51%), and human ratings cluster around 2–3/10. To align automated evaluation with human judgment, the authors collect WorldLens-26K, a dataset of 26,808 human-annotated score/rationale records, and train WorldLens-Agent, a Qwen3-VL-based critic that predicts scores and generates textual rationales. The central claims are that visual realism does not imply behavioral usability and that geometry-aware, temporally conditioned generation yields more physically coherent worlds.","tokens_in":53282,"tokens_out":6254,"duration_ms":68929,"significance":"The benchmark is timely and the empirical breadth is substantial. Strengths include the detailed per-dimension appendices, the transparent use of fixed pretrained evaluators, the unusually large human annotation dataset with structured rationales, and the stated commitment to release toolkit, dataset, and model. If the measurement-validity issues are addressed, WorldLens could become a useful standardized evaluation ecosystem for driving world models. At present, however, the headline conclusion 'Perceptual Quality Does Not Imply Usability' is not yet established, because the downstream metrics conflate world-model fidelity with the domain robustness of the frozen perception models. Likewise, the claim that WorldLens-Agent shows 'strong alignment with human annotations' is supported only by qualitative examples. Both issues are fixable with additional analyses, but they are load-bearing for the paper's main contributions.","major_comments":[{"comment":"The downstream-task aspect uses a single frozen perception toolchain (BEVFusion for D.1/D.2, ADA-Track for D.3, SparseOcc for D.4) pretrained on real nuScenes frames. The paper interprets OpenDWM's low NDS (21.96% vs. DiST-4D's 33.22%) and the statement 'large-scale multi-domain training can hinder adaptation' as evidence about the world model. An equally plausible reading is that these detectors are brittle to OpenDWM's particular synthetic distribution (color, contrast, resolution, object appearance), independent of scene fidelity. No control is provided that perturbs real frames to match each world model's distribution, or that uses multiple perception backbones, or that measures a domain-gap baseline. This confound bears directly on the headline 'Perceptual Quality Does Not Imply Usability' in §5.3. The limitation section (§13.3) does not flag this issue. I recommend adding a control","section":"§3.4, §5.3, Tables 3 and 21"},{"comment":"The paper claims that WorldLens-Agent's predicted scores 'exhibit strong alignment with human annotations across all evaluated dimensions,' but no quantitative evidence is reported. Section 5.2 and Appendix 12.4 show only qualitative examples on Gen3C videos. There is no correlation coefficient (e.g., Spearman or Pearson), no per-dimension agreement statistics, no sample size for the zero-shot test, and no comparison with the base Qwen3-VL model without LoRA fine-tuning. Since WorldLens-Agent and WorldLens-26K are billed as a core contribution that enables 'scalable, explainable scoring,' the absence of a quantitative validation is a load-bearing gap. Please report agreement metrics on a held-out or OOD set, ideally with confidence intervals and a base-model baseline.","section":"§4.3, §5.2, Figure 8, §12.4"},{"comment":"The Action-Following scores are co-determined by the LimSim traffic engine and the DriveArena closed-loop protocol, so the uniformly low Route Completion rates (6.89–13.51%) may partly reflect simulator and planner limitations rather than world-model deficiencies. More importantly, the two models featured in the paper's main trade-off narrative — OpenDWM and DiST-4D — are absent from Table 2, so the claim that 'geometry-stable ones lack behavioral fidelity' is not directly tested for the models that are central to the headline. Either include these models in the closed-loop evaluation or restrict the claim to the models actually evaluated. I also suggest a real-data closed-loop baseline (e.g., same planner and simulator on real recorded frames) to calibrate the absolute route-completion numbers.","section":"§3.3, Table 2, §9.2.3"},{"comment":"The human preference scores are heavily concentrated at the low end: for most dimensions the median and quartiles are all 2.0, with mean differences between models often around 0.2–0.3 points. The paper nevertheless makes comparative claims such as DiST-4D 'achieves the most balanced scores' and 'leads in physical plausibility' and 'behavioral safety.' No inter-annotator agreement, significance testing, or confidence intervals are reported, and the number of unique videos per model is not stated. Given the small apparent effect sizes, these comparative human-preference claims are not statistically supported. Please add agreement metrics and a statistical analysis of the model-level differences.","section":"§4.1, §5.2, Tables 24–29"}],"minor_comments":[{"comment":"The OpenDWM mAP value reads '0.944'; from the NDS value and surrounding rows this appears to be a typo for '0.0944'. Please correct.","section":"Table 21"},{"comment":"The 'Empirical Max' row is not defined in the main text. State whether it is a per-sample maximum, a video-level upper bound, or an oracle value, and clarify why some cells are omitted (e.g., Perceptual Discrepancy).","section":"Tables 1–3"},{"comment":"The phrase 'the first benchmark that measures both the appearance and behavior' is too strong given existing closed-loop and behavior-oriented evaluations such as DriveArena and NAVSIM. Please qualify the novelty claim.","section":"§2"},{"comment":"The statement that 'World Realism and Consistency scores correlate strongly' is made without reporting a correlation coefficient or scatter plot. Please add the quantitative value.","section":"§5.2"},{"comment":"The section title says 'Physical Plausibility' but the content and rubric describe 3D & 4D Consistency; the heading appears to be a copy-paste error.","section":"§11.5"}],"recommendation":"major_revision","confidential_remarks":"This is a substantial and potentially influential benchmark paper, but the toolchain confound in Aspect 4 is the main risk. The fix — a controlled perturbation of real frames matched to each model's distribution, plus quantitative validation of the evaluation agent — is within the scope of a major revision and does not require re-collecting the benchmark. I would not reject on the current evidence, but I would not accept without the controls."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The useful thing here is the integrated protocol: five aspects, 24 dimensions, all specified with formulas and appendix detail, run on six driving world models. The 26K human-preference dataset with rationales and the distilled evaluator are genuinely new resources, and the main empirical pattern—no model dominates across generation, reconstruction, action-following, downstream tasks, and human judgment—holds up internally. The closed-loop collapse (route completion 6.9–13.5% for every model) is a striking and believable finding. This is a benchmark paper that could plausibly become a standard reference if adopted.\n\nWhat it does well: the authors are unusually concrete about metric definitions and implementation choices, the tables are mostly coherent, and the qualitative examples match the numbers. The human ratings averaging 2–3/10 across all models is honestly reported and consistent with the low closed-loop scores.\n\nThe soft spots are real but not fatal. The stress-test concern is correct: the downstream-task scores and the 'perceptual quality does not imply usability' conclusion treat degradation of frozen detectors/trackers as a property of the world model, but no control perturbs real frames to match each synthetic distribution. OpenDWM's low NDS could just mean BEVFusion is brittle to OpenDWM's particular color and contrast statistics. The same toolchain confound affects the reconstruction and action-following aspects. That doesn't sink the whole framework—the generation and human-preference aspects are less confounded—but it means the headline trade-off claim is not yet fully controlled, and the limitations section does not flag it.\n\nAlso worth fixing: the WorldLens-Agent 'strong alignment with human annotations' claim is supported only by qualitative examples, with no quantitative agreement numbers; the human-preference dataset reports no inter-annotator agreement; and Table 21 has an impossible mAP of 0.944 for OpenDWM, inconsistent with its NDS and with the other columns. These are addressable without changing the framework.\n\nWho this is for: anyone building or evaluating driving world models, and the broader video-generation evaluation community. It deserves a serious referee and probably acceptance after revision. I'd send it to peer review, with a specific request to add a synthetic-distribution control or soften the claim.","headline":"A serious, mostly well-built benchmark for driving world models, with one real confound in the downstream and action metrics that should be controlled before the headline claim is taken at face value.","tokens_in":53964,"tokens_out":1125,"would_cite":true,"duration_ms":14115,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"WorldLens claims no current driving world model excels universally: perceptual realism and functional usability are decoupled, so a full-spectrum benchmark must measure both appearance and behavior.","keywords":["world model evaluation","driving world models","benchmark","generative video","human preference","closed-loop simulation","4D reconstruction","evaluation agent"],"falsifier":"Take real driving frames, apply distribution shifts that mimic each generated model's visual style, and run the same frozen perception stack; if detection and tracking scores fall as much as they do on generated videos, the downstream aspect mainly measures perception-model brittleness rather than world-model fidelity.","tokens_in":52793,"feed_emoji":"🚗","tokens_out":4808,"duration_ms":51137,"temperature":0.7,"pith_summary":"WorldLens is a benchmark for driving world models that claims to be the first to measure both appearance and behavior. It scores six models across five aspects—generation, reconstruction, action-following, downstream tasks, and human preference—spanning 24 dimensions. Its central finding is that no model excels everywhere: the best-looking model ranks second-lowest in 3D detection and lowest in tracking, while all models finish only 6.89–13.51% of closed-loop routes, and human raters give overall realism an average of 2–3 out of 10. The paper argues that this decoupling makes unified, human-aligned evaluation essential, and it provides one: a 26,808-annotation dataset and a distilled, explainable evaluation agent.","feed_headline":"No driving world model both looks real and drives safely","feed_subtitle":"Five-aspect benchmark finds top realism model ranks second-lowest in detection; closed-loop route completion stays under 14%.","key_machinery":"The carrying mechanism is a five-aspect, 24-dimension protocol that pairs objective signals—monocular depth stability, semantic label stability, 4D Gaussian-splatting reconstructability, frozen-planner trajectory adherence, closed-loop route completion, and frozen-perception downstream scores—with a large human-annotated preference dataset of 26,808 scored videos with textual rationales. A vision-language critic distilled from these annotations outputs both 1–10 scores and evidence-based explanations, enabling scalable, explainable evaluation. The key move is testing each model both as an appearance generator and as an environment that a planner can operate in, which exposes the appearance–b","core_discovery":"The paper's central claim is that perceptual quality and functional usability are decoupled in current driving world models. Empirically, the model with the highest subject fidelity scores poorly on downstream detection and tracking, while the geometrically most stable model is also the most balanced overall yet still fails to complete more than 13.51% of closed-loop routes. Human ratings of world realism, physical plausibility, and behavioral safety cluster around 2–3 out of 10 for every model, with a strong correlation between perceived realism and geometric consistency. WorldLens makes these trade-offs visible and reproducible through a standardized five-aspect protocol.","pith_inferences":["The toolchain confound implies that downstream scores may partly reflect how brittle the frozen perception models are to each model's synthetic distribution; a calibration run on real frames perturbed to match each style would separate the two effects.","The strong correlation between human realism ratings and geometric consistency suggests depth and novel-view metrics could serve as a cheap proxy for human preference, lowering annotation cost.","WorldLens-26K could be repurposed as a reward model for reinforcement fine-tuning of world models, turning the benchmark from a measurement stick into an optimization target.","The closed-loop collapse below 14% route completion indicates a ceiling not visible in open-loop metrics; testing whether self-forcing or streaming-diffusion training raises that ceiling is a direct next step."],"forward_implications":["If perceptual and functional decoupling holds, optimizing only appearance metrics will not produce safe driving simulators; geometry and temporal conditioning must be explicit objectives.","A common five-aspect protocol makes results across models and datasets comparable, standardizing world-model evaluation.","The distilled evaluation agent can replace costly human annotation, returning both scores and reasons at scale.","The uniformly low human ratings (2–3/10) quantify a large headroom for improvement, not just incremental gains.","Geometry-aware supervision consistently improves reconstruction, novel-view, and downstream scores, pointing to a concrete design direction."],"fun_headline_variants":["Realism doesn't equal safety in driving world models","Top realism world model fails detection and route completion","Driving world models: pretty but not reliable, benchmark shows","No world model balances realism and driving safety yet","Perceptual quality and driving reliability remain decoupled"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"Each aspect score is interpreted as a property of the world model, but no control isolates the external toolchain—the 4D reconstruction optimizer, traffic simulator, and frozen perception models—that co-determines every score.","fun_headline_variants_meta":{"raw":{"variants":["Realism doesn't equal safety in driving world models","Top realism world model fails detection and route completion","Driving world models: pretty but not reliable, benchmark shows","No world model balances realism and driving safety yet","Perceptual quality and driving reliability remain decoupled"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000141,"raw_usage":{"total_tokens":998,"prompt_tokens":736,"completion_tokens":262,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":480,"completion_tokens_details":{"reasoning_tokens":187}},"tokens_in":480,"tokens_out":262,"duration_ms":3226,"temperature":1.0,"reasoning_tokens":187,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T16:58:06.674845+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take real driving frames, apply distribution shifts that mimic each generated model's visual style, and run the same frozen perception stack; if detection and tracking scores fall as much as they do on generated videos, the downstream aspect mainly measures perception-model brittleness rather than world-model fidelity.","supporting_citations":[],"review_version":1}