{"id":"efa3b40e-6dbe-4534-b300-7d2d6681beb6","arxiv_id":"2412.05893","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"doScenes adds retroactive natural language instructions and referentiality tags to 1,000 nuScenes clips so autonomous driving systems can learn instruction-following behavior.","lead":"The paper introduces doScenes, a dataset built on 1,000 nuScenes driving clips with natural language instructions and static/dynamic referentiality tags. The goal is to help autonomous vehicles learn to follow short-term human directives in real-world scenes.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'first public real-world dataset' claim is falsified by the uncited Talk2Car dataset, which already provides natural-language driving commands, referred-object referentiality annotations, and ego trajectories on nuScenes; the central novelty claim needs reframing or removal.","rationale":"The paper's most consequential assertion is in Section II-C: first public real-world dataset with driving instructions and referentiality information. That claim is externally checkable and appears to fail because Talk2Car predates doScenes and already provides natural-language commands, referred-object annotations, and ego trajectories on nuScenes. This is more decisive than the annotation-reliability concern because even perfect annotation reliability cannot rescue a false priority claim. I agree with the reader's overall CONDITIONAL disposition: the appropriate remedy is to rewrite the novelty claim, add a Talk2Car comparison, and release the data. I do not think the paper must be rejected outright, because its static/dynamic referentiality taxonomy and multi-annotation-per-scene format may be useful extensions, but the current claim is not maintainable as stated. The reader's weakest assumption focused on retroactive annotation; the Talk2Car omission was flagged in the rationale but not as the primary weakness, hence partial agreement.","tokens_in":8359,"tokens_out":5518,"duration_ms":58523,"concrete_test":"Construct a comparison row for Talk2Car: check (1) the repository is public and downloadable, (2) annotations are from real-world nuScenes scenes, (3) commands are natural-language directives, (4) each command has a referred-object annotation, and (5) commands are associated with ego trajectories. If all five hold, Section II-C's 'first public real-world dataset' claim is false as stated; the paper must either cite Talk2Car, state what doScenes adds (e.g., explicit static/dynamic referentiality tags, multiple instructions per scene), or drop the firstness claim.","verdict_should_be":"UNCHANGED","load_bearing_attack":"In Section II-C, the paper claims that doScenes is 'the first public real-world dataset to provide driving instructions and referentiality information as the natural language annotation' and uses this to distinguish it from nuScenes-QA, DRAMA, and simulation-based frameworks. This claim ignores Talk2Car (Deruyttere et al., 2019), a public real-world dataset constructed from nuScenes that contains natural-language commands issued to a driver, referred-object annotations for each command, and the ego vehicle's resulting trajectory. Talk2Car thereby already supplies driving instructions, referentiality information, and the instruction-to-motion link that doScenes claims to introduce. The doScenes contribution may still be valuable as a re-annotation with explicit static/dynamic referentiality tags and multiple plausible instructions per scene, but the 'unlike existing datasets' framing overstates novelty. The annotation-reliability caveat raised by the reader is also valid, but it is secondary: even with reliable annotations, the central claim as written is not defensible without a comparison to Talk2Car.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces doScenes, a dataset built by retroactively annotating all 1,000 nuScenes 12-second driving scenes with natural language instructions and referentiality tags (non-referential, static-referential, dynamic-referential, and both). Five annotators used a 'taxi test' heuristic to infer what passenger instruction, if any, would have caused the observed ego motion. The paper reports aggregate statistics on referentiality, discusses the retroactive-annotation limitation, and proposes applications in vision-language navigation and vision-language-action models. The dataset is made public through a GitHub link.","tokens_in":8565,"tokens_out":3071,"duration_ms":28621,"significance":"If the annotations are reliable, doScenes could provide a useful resource for linking imperative language to real-world driving motion, complementing simulation-based instruction datasets and enabling research on referential instructions that require closed-loop observation. The explicit static/dynamic referentiality tags are a useful and relatively novel annotation axis. However, the paper does not yet provide the evidence needed to assess annotation reliability, and its central novelty claim is overstated by the omission of Talk2Car, a public real-world dataset that already supplies natural-language driving commands, referred-object annotations, and ego trajectories on nuScenes. The retroactive-annotation approach is honestly acknowledged, but its validity for establishing an instruction-to-motion link is left unvalidated.","major_comments":[{"comment":"The claim that doScenes is 'the first public real-world dataset to provide driving instructions and referentiality information as the natural language annotation' is not defensible, because Talk2Car (Deruyttere et al., 2019) already provides natural-language driving commands, referred-object annotations, and ego trajectories on nuScenes data. The absence of any citation or discussion of Talk2Car in Section II is a substantive omission that directly affects the paper's central contribution claim.","section":"II-C"},{"comment":"The annotation statistics in Table II sum to 1,001 annotations for a stated 1,000 scenes, and the paper does not explain how five independent annotators relate to this total (e.g., number of scenes annotated per annotator, agreement rates, or handling of multiple annotations per scene). The paper also provides no example annotation rows, no annotation guideline text, and no inter-annotator agreement metric. Without this information, readers cannot assess the stability or quality of the labels, which is load-bearing for a dataset paper.","section":"III"},{"comment":"The paper's own limitation statement in Section V that the retroactive instruction is 'only a proxy for a true signal' is candid, but it means the central premise -- that the instruction can serve as the cause of the observed motion -- is an unverified assumption. The authors should either validate the taxi-test heuristic (e.g., by comparing annotator-recovered instructions with actual human instructions in a small naturalistic study, or by showing that the annotated instructions predict the recorded trajectory better than a baseline) or explicitly scope the dataset as containing plausible instructions rather than causal instructions.","section":"V"},{"comment":"No data file, schema description, or sample of the annotation table is included in the manuscript, making it impossible to verify the format of the instructions, the handling of blank instruction fields, or the exact content of the referentiality tags. The GitHub link alone is insufficient for review; a few example rows should be included in the paper or an accessible appendix.","section":"III"}],"minor_comments":[{"comment":"In the abstract, 'A Vs' should be 'AVs' or 'autonomous vehicles.'","section":"Abstract"},{"comment":"In Section III, 'thestatic reference tag' is missing a space; it should read 'the static reference tag.'","section":"III"},{"comment":"The citation formatting is inconsistent for references [8], [9], and [14]; for example, nuScenes-MQA and DRAMA use different styles for venue and year.","section":"II-B"},{"comment":"Section IV mentions 'the first t seconds' but does not suggest a method for choosing t; a brief discussion of the temporal validity of instructions would be helpful.","section":"IV"},{"comment":"Figure 2 is referenced without axis labels or a caption; adding a concise caption and labeled axes would improve readability.","section":"III"}],"recommendation":"major_revision","confidential_remarks":"The omission of Talk2Car is a significant oversight, as it is a well-known dataset in the driving-instruction community. This may indicate a narrow literature review, and the authors should be asked to compare against it explicitly. The manuscript is honest about the retroactive annotation limitation, which counts in its favor, but the lack of data-quality evidence is a major barrier to acceptance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: doScenes is a plausible and potentially useful re-annotation of nuScenes with free-form driving instructions and static/dynamic referentiality tags, but its headline novelty claim is overstated because Talk2Car already covers natural-language commands and referred objects on nuScenes, and the paper neither cites nor engages with it. The dataset itself is honestly described as a retroactive proxy, but the paper gives no inter-annotator agreement, no example annotations, and its own Table II sums to 1,001 for 1,000 scenes.\n\nWhat is actually new here is the specific package: the taxi-test heuristic, multiple annotations per scene, and the explicit static/dynamic referentiality tags. That is a legitimate extension over prior datasets, and the writing is refreshingly clear about limitations—the proxy nature of retroactive instructions, the 12-second clip mismatch, and the absence of speed/style directives. Those are good signs.\n\nThe soft spots are real. First, the central claim in Section II-C that this is \"the first public real-world dataset to provide driving instructions and referentiality information as the natural language annotation\" is not defensible. Talk2Car (Deruyttere et al., 2019) is a public real-world dataset built on nuScenes with natural-language commands, referred-object annotations, and ego trajectories. The doScenes static/dynamic tags are a refinement of the referentiality idea, not its first appearance. That claim needs reframing or removal. Second, the dataset's utility is undercut by missing evidence: no example annotation rows, no inter-annotator agreement, no clear account of how the five annotators map to the 1,001 annotations in Table II, and no data file in the manuscript. These are all fixable, but as presented the annotation reliability is unmeasured.\n\nWho is this for? Researchers who want a quick real-world label layer on nuScenes for instruction-conditioned motion planning. It could be a handy resource, but I would not cite it until the repository actually ships, the protocol is documented, and the numbers add up.\n\nRecommendation: send it to peer review as a major revision. The idea is useful, the authors are honest about the proxy caveat, but the novelty claim must be corrected and the annotation quality demonstrated with concrete artifacts.","headline":"Useful nuScenes instruction layer, but the 'first real-world' claim ignores Talk2Car and the annotation stats are unverified.","tokens_in":9106,"tokens_out":1743,"would_cite":false,"duration_ms":17334,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that doScenes is the first public real-world dataset to pair autonomous-driving sensor clips with natural-language driving instructions and referentiality tags, linking imperative language to observed vehicle motion.","keywords":["autonomous driving","natural language instruction","vision-language navigation","dataset annotation","referentiality","motion planning","nuScenes","human-vehicle interaction"],"falsifier":"Collect real in-car passenger instructions synchronized with ego trajectories and compare their content and referentiality distribution to the taxi-test annotations; systematic differences would undermine the proxy. A simpler check is measuring inter-annotator agreement on the taxi test for a sample of scenes, since near-chance agreement would show the causal instruction is not reliably recoverable from playback.","tokens_in":8200,"feed_emoji":"🚗","tokens_out":2969,"duration_ms":30237,"temperature":0.7,"pith_summary":"This paper introduces doScenes, a dataset built by retroactively annotating 1,000 real-world driving clips from nuScenes with short, passenger-style natural-language instructions. Each clip receives the instruction that would have caused the observed motion, plus a tag indicating whether the instruction refers to static objects, dynamic objects, both, or neither. The authors argue this is the first public real-world dataset to connect imperative language directly to vehicle motion, in contrast to simulated instruction datasets and real-world datasets that only caption risks or scenes. If the annotations are reliable, doScenes would let vision-language-action models learn instruction-conditioned motion planning from real sensor data rather than from fixed command sets.","feed_headline":"1,000 real driving scenes paired with the instructions that caused them","feed_subtitle":"The doScenes dataset tags each nuScenes clip with a passenger-style command and whether it refers to static or moving objects.","key_machinery":"The central mechanism is the retroactive taxi-test annotation procedure: for each nuScenes clip, five independent annotators imagine riding as a passenger and write the short instruction that would initiate the observed motion, optionally with multiple phrasings per scene. The referentiality tags (none, static, dynamic, both) operationalize whether the instruction requires continued observation of scene objects, which distinguishes closed-loop evaluation from open-loop prediction. This procedure turns an existing large multimodal driving corpus into instruction-conditioned training data without instrumented in-vehicle recording.","core_discovery":"doScenes augments all 1,000 nuScenes scenes with retroactive driving-instruction annotations and referentiality tags. Annotators watching sensor playback apply the taxi test, asking what instruction a passenger would give to trigger the observed maneuver, and mark whether the instruction is non-referential, static-referential, dynamic-referential, or both. The dataset statistics report 535 non-referential, 214 static-referential, 159 dynamic-referential, and 93 both-referential annotations. The paper claims this makes doScenes the first public real-world dataset to provide driving instructions and referentiality information as the natural-language annotation, creating a link between imperative language and motion for autonomous vehicles, and that it supports nuanced, flexible responses beyond the predefined action sets used in simulated instruction datasets.","pith_inferences":["If the taxi-test proxy holds, doScenes could become a standard benchmark for instruction-conditioned trajectory prediction, letting researchers measure how much language grounding improves forecasting beyond scene-only models.","A testable extension is to compare model performance on static-referential versus dynamic-referential instructions, predicting that dynamic ones are harder because the referred object's future state is what drives the motion plan.","The multiple annotations per scene capture paraphrased instructions with similar intent, which could be used to evaluate how robust language-grounding models are to varied phrasings of the same directive.","If retroactive annotation proves reliable, it offers a low-cost route to convert other existing driving corpora into instruction datasets, accelerating data collection without instrumented vehicles."],"forward_implications":["Vision-language-action models trained on doScenes could respond to open-vocabulary natural-language commands rather than a fixed set of predefined directives.","The static versus dynamic referentiality tags allow training and evaluation on subsets where models must either ground instructions in map-visible static objects or track moving objects that require closed-loop observation.","The annotations support the reverse task of generating a natural-language description of a trajectory, contributing to interpretable and interactive autonomous driving.","Multi-stage motion planning can be studied, since some 12-second clips contain maneuvers longer than a single instruction, and accurate response may appear only in the first portion of a path.","Models using only rasterized maps may handle non-referential instructions, but referential instructions will require LiDAR or front-view images, giving a natural test bed for sensor-modality choices."],"supporting_citations":[{"why":"Supplies the 1,000 multimodal nuScenes scenes with camera, LiDAR, radar, maps, and trajectories that doScenes re-annotates with instructions.","marker":"[6]"},{"why":"Provides the do-calculus inspiration behind the dataset name and the causal framing that instructions should emulate commands causing the observed action sequence.","marker":"[2]"},{"why":"Represents the motion-planning-as-language paradigm that doScenes extends by adding egocentric views and human instructions beyond generic goals.","marker":"[10]"},{"why":"Shows a simulated instruction-following driving framework whose fixed action set and synthetic data doScenes contrasts with real-world flexible instructions.","marker":"[11]"},{"why":"Another simulated instruction-following framework, LMDrive, that doScenes brings to real-world data with open-vocabulary instructions.","marker":"[12]"},{"why":"The closest real-world natural-language driving dataset, but focused on risk captioning rather than directives; doScenes distinguishes itself from this baseline.","marker":"[14]"}],"fun_headline_variants":["doScenes: 1,000 real scenes annotated with passenger-style instructions","First real-world driving dataset with instruction referentiality","doScenes pairs 1,000 driving scenes with the commands that triggered them","New dataset links passenger commands to nuScenes scene objects","doScenes: the first driving dataset with instruction tags for AVs"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The dataset's value rests on the assumption that instructions written after the fact by annotators watching sensor playback faithfully capture the instruction a real passenger would have given to produce the observed driving.","fun_headline_variants_meta":{"raw":{"variants":["doScenes: 1,000 real scenes annotated with passenger-style instructions","First real-world driving dataset with instruction referentiality","doScenes pairs 1,000 driving scenes with the commands that triggered them","New dataset links passenger commands to nuScenes scene objects","doScenes: the first driving dataset with instruction tags for AVs"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001248,"raw_usage":{"total_tokens":5086,"prompt_tokens":883,"completion_tokens":4203,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":499,"completion_tokens_details":{"reasoning_tokens":4111}},"tokens_in":499,"tokens_out":4203,"duration_ms":31352,"temperature":1.0,"reasoning_tokens":4111,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T20:13:37.467151+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Collect real in-car passenger instructions synchronized with ego trajectories and compare their content and referentiality distribution to the taxi-test annotations; systematic differences would undermine the proxy. A simpler check is measuring inter-annotator agreement on the taxi test for a sample of scenes, since near-chance agreement would show the causal instruction is not reliably recoverable from playback.","supporting_citations":[{"cited_title":"nuscenes: A multimodal dataset for autonomous driving,","cited_arxiv_id":null,"evidence_quote":"Supplies the 1,000 multimodal nuScenes scenes with camera, LiDAR, radar, maps, and trajectories that doScenes re-annotates with instructions."},{"cited_title":"Lmdrive: Closed-loop end-to-end driving with large language models,","cited_arxiv_id":null,"evidence_quote":"Another simulated instruction-following framework, LMDrive, that doScenes brings to real-world data with open-vocabulary instructions."},{"cited_title":"Drama: Joint risk localization and captioning in driving,","cited_arxiv_id":null,"evidence_quote":"The closest real-world natural-language driving dataset, but focused on risk captioning rather than directives; doScenes distinguishes itself from this baseline."}],"review_version":1}