{"id":"5ff71731-5867-460f-a868-7b4ab584ffeb","arxiv_id":"2411.17912","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"In all six real-world path-planning scenarios, GPT-4, Gemini, and Mistral produced routes with major errors, leading the authors to judge all three unreliable for navigation.","lead":"This study asked three large language models to produce driving and walking directions for six real routes, and all three returned directions with major mistakes, including gaps in the route and wrong destinations. The authors conclude that such models should not be used to direct vehicle navigation until they gain reality-check and transparency mechanisms.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"One-shot prompts conflate missing world knowledge with path-planning inability; the 'fundamentally unsuitable' conclusion overreaches.","rationale":"The reader's weakest assumption was that six local, hand-chosen routes and one prompt style are representative. My concern is adjacent but distinct: even on the routes actually tested, the task design does not isolate path-planning ability from external-knowledge retrieval and one-shot stochasticity. Both concerns target the gap between the evidence and the strong 'not capable / fundamentally unsuitable' wording. I agree with the reader that the narrow finding—these three models made numerous errors on these six routes—is supported by concrete, sometimes striking examples (e.g., Mistral's 2,500-mile route to Nevada; the 8.8-mile and 83.3-mile discontinuities in Section 4.2). Those examples earn credit for the paper as a negative demonstration of standalone zero-shot LLM navigation. However, the categorical conclusion requires more than one failed draw per model-scenario; a single sample cannot establish incapability, and several failures are transparently due to lack of real-time schedule data or place-name disambiguation rather than spatial reasoning. The proposed controlled replication with map-augmented prompts and repeated sampling would settle whether the bottleneck is path planning per se or missing knowledge. I keep the verdict at CONDITIONAL/UNCHANGED because the paper's useful contribution is a set of documented failure cases and a clear limitation section, and the overreach can be corrected by narrowing the claim and adding such a control. No ad hominem is intended; this is an argument-design critique.","tokens_in":16370,"tokens_out":6298,"duration_ms":62233,"concrete_test":"Run a controlled replication with 20-50 start-destination pairs (mix of well-known landmarks and ambiguous names), each repeated 10 times at temperature 0 and 0.7, in two conditions: (A) the original standalone prompt, and (B) the same prompt supplemented with a road-network graph or explicit permission to query a mapping API. Have blind raters score route validity against multiple independent reference routes, not a single Waze route. If condition (B) yields high accuracy (e.g., >90% valid) while (A) remains low, the bottleneck is world knowledge/tool access, not path planning, and the categorical conclusion must be narrowed. If (B) is also poor, the planning deficit is confirmed.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim—'our tested LLMs are not capable of path planning in the real world' and 'fundamentally unsuitable for reliable path-planning applications' (Sections 6, 4.4)—requires that the observed failures be failures of spatial planning. But the prompts bundle at least three distinct competencies: (1) retrieving current facts (MLB first-weekend schedule, federal holiday, operating hours), (2) disambiguating place names (Hogsback Mountain Paintball Center vs. Hogsback Rd; Shot Tower State Park in Virginia vs. Austin, NV), and (3) constructing a valid route. Appendix A shows that several failures are in (1)/(2), not (3): Mistral cannot browse the web and routes to Nevada; Gemini initially returns Google Maps links and misidentifies Friday as a weekend game before correcting. Section 4.3 separately concedes that Mistral 'could not search the internet in real-time.' Because each model-scenario cell was run once, with no temperature control or repeated sampling, the data cannot separate 'cannot plan' from 'lacked a fact on this draw.' The paper itself limits the sample in Section 5.4, but the conclusion is stated unconditionally. If the same models were given a map or mapping API, route generation might be substantially better; the evidence as presented supports 'unreliable standalone zero-shot LLMs,' not 'fundamentally unsuitable for path planning.' This is the load-bearing weakness: the strong categorical claim depends on an attribution that the experiment does not isolate.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports an empirical evaluation of three LLMs (GPT-4, Gemini, and Mistral) on six real-world path-planning scenarios around George Mason University: three turn-by-turn driving scenarios (urban, suburban, rural) with time constraints and three vision-and-language navigation scenarios on campus. For each scenario, the model was prompted once (with follow-ups when the initial response was unhelpful), and the outputs were compared against a Waze-derived ground-truth route and in-person landmark checks. The authors classify errors as major or minor, report discontinuity distances and turn counts, assign severity weights to error categories, and rank the models. They conclude that the tested LLMs are 'not capable of path planning in the real world' and are 'fundamentally unsuitable for reliable path-planning applications' (Sections 4.4 and 6). They also discuss directions for future work: reality checks, in-context transparency, and smaller models.","tokens_in":16661,"tokens_out":4067,"duration_ms":35938,"significance":"If the paper's sweeping conclusion were fully supported, it would be an important cautionary result for the automotive industry's recent integration of LLMs into vehicle voice assistants. The work has clear strengths: it moves beyond simulated environments to real roads and buildings, verifies landmark claims in person, and documents specific, concrete route errors (e.g., the 8.8-mile discontinuity in Gemini's suburban route, Mistral routing to Nevada instead of Virginia). These documented failures are a valuable case study of LLM limitations in high-stakes spatial tasks. However, the breadth of the conclusion exceeds what the experimental design can support: with one trial per model-scenario, no control for stochasticity, and prompts that mix factual retrieval with spatial reasoning, the evidence supports a more limited claim about unreliable zero-shot performance in a small set of instances. The significance is therefore conditional on the authors reframing the claim and strengthening the analysis.","major_comments":[{"comment":"The conclusion that the LLMs are 'not capable of path planning in the real world' and 'fundamentally unsuitable for reliable path-planning applications' is not supported by the experimental design, which does not isolate spatial planning from other competencies. Several documented failures are attributable to missing knowledge rather than route-construction ability. For example, in the rural TbT scenario (Appendix A, Mistral), the model routes to Austin, NV because it misidentified the state; given the erroneous destination, the route sequence is internally plausible. Gemini's failure to identify the correct game date (Section 4.3) is a scheduling-information error, not a path-existence error. The authors themselves concede in Section 5.4 that the routes were local and may not generalize, and that real-time mapping access could be 'game-changing.' To make the categorical claim, the authors need either to disentangle knowledge failures from planning failures (e.g., by supplying the correct destination and schedule in the prompt), to compare against a model with browsing or mapping tools, or to temper the conclusion to 'unreliable as standalone zero-shot planners without tool access.'","section":"Sections 4.4 and 6"},{"comment":"The performance rankings in Table 2 depend entirely on hand-assigned severity weights: TbT discontinuities and wrong directions receive 0.4, wrong exits 0.2; VLN discontinuities and failure to reach the destination receive 0.8, landmark errors 0.2. No justification is given for these weights, and no sensitivity analysis is reported. Because the rankings flip under alternative plausible weights (for example, weighting failure-to-reach as more severe than a discontinuity, or weighting landmark fabrication as more severe than a wrong exit), the comparative claims (e.g., 'Gemini performed the best' in all VLN, 'ChatGPT-4 had the weakest performance in urban') are not robust. I recommend either deriving the weights from a user-impact survey or rationale, testing a range of weight vectors, or reporting raw error counts and discontinuities without a single aggregate ranking.","section":"Section 4.4"},{"comment":"Each model-scenario cell was run once, with no repeated sampling and no temperature control. The paper itself reports that approximately 77.8% of TbT scenarios required at least one follow-up question, and the follow-up wording varies across models and scenarios (e.g., Gemini was asked to avoid Google Map links, Mistral was asked to find federal holidays). Under these conditions, the differences in Table 2 (e.g., Gemini first vs. GPT-4 last in the short VLN scenario) cannot be distinguished from sampling noise or prompt-variation effects. This is load-bearing because Section 4.4 makes model-to-model comparisons. At minimum, the authors should either present the results as individual case studies without ranking claims, or add repeated trials with fixed prompts and report variance and inter-rater reliability for the error classification.","section":"Sections 3 and 4.3"},{"comment":"The error taxonomy is defined qualitatively, but the quantitative discontinuity measurements (e.g., 8.8 miles, 83.3 miles) require a precise gap-filling protocol. The text states that 'we identify the shortest path to fill these substantial gaps,' but does not specify the algorithm, road network, or treatment of the 'shortest path' when no direct connection exists. Similarly, the threshold for a major VLN discontinuity ('three or more buildings') is not justified. Without a reproducible protocol, independent verification of the central quantitative results is impossible. I recommend including the full route transcripts, the gap-filling procedure, and the major/minor decision rules as supplemental material.","section":"Section 4.1"}],"minor_comments":[{"comment":"The statement that Mistral 'could not search the internet in real-time' is presented without a citation or a description of the model interface; it is unclear whether this is a property of the model itself or of the experimental setup (e.g., no browsing tool enabled).","section":"Section 4.3"},{"comment":"The figures are referenced but not embedded in the text, and the color scheme (red vs. green) is described only in captions. Please ensure the final version includes clearly legible figures with a legend explaining the route colors, and consider adding scale bars.","section":"Figures 1 and 2"},{"comment":"The categories 'Beginner,' 'Intermediate,' and 'Expert' are defined qualitatively, but the assignment of a particular route to a category uses the phrase 'we believe,' which is too subjective. Please provide an explicit rubric (e.g., number of discontinuous miles, complexity of the local area) for assigning these ratings.","section":"Section 4.2"},{"comment":"The prompt labeled 'Suburban VLN' for Mistral appears to be a copy-paste of the TbT prompt: it asks for 'turn-by-turn driving instructions' to Garage C and mentions the paintball center, not visual landmarks. This inconsistency undermines reproducibility of the VLN results for Mistral and should be corrected.","section":"Appendix A, Mistral Suburban VLN prompt"},{"comment":"The claim that 'None of the models accurately planned any of the paths' lacks a definition of 'accurately planned.' Is a path with a single missing turn 'inaccurate'? Please specify the accuracy threshold (e.g., all turns correct, zero discontinuities, arrival at destination) before making the claim.","section":"Section 4.4"}],"recommendation":"major_revision","confidential_remarks":"The paper is better framed as a case study documenting specific failures of three LLMs on six hand-crafted route-planning tasks than as a general proof that LLMs cannot plan paths. The main risk is overclaiming from a very small, non-randomized sample. A major revision that narrows the conclusions, adds sensitivity analysis for the weights, and reports raw trial data would make the contribution solid and publishable as a negative-result or evaluation study. I would also encourage the authors to pin down the exact model versions and access dates, since LLM behavior changes rapidly, and to consider archiving their prompts and responses as a supplementary dataset."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth reading for the concrete failure cases. The narrow empirical claim—that GPT-4, Gemini, and Mistral each made major errors on six real routes around George Mason University—is supported by the documentation: the 50-mile phantom address, the 8.8-mile discontinuity, the misrouting to Nevada. The in-person landmark checks and Waze ground truth are real evidence, and the authors are upfront about local-area limits. If I needed an example of why a standalone LLM should not be handed the navigation stack, I would point here.\n\nThe soft spot is the conclusion, not the data. \"Not capable of path planning in the real world\" and \"fundamentally unsuitable for reliable path-planning applications\" overreach. The prompts mix at least three skills: retrieving current facts (game times, federal holidays), disambiguating names (Hogsback Mountain Paintball Center vs. Hogsback Rd.; Shot Tower State Park in Virginia vs. Austin, NV), and generating a continuous route. Several documented failures are fact-retrieval or name-matching failures, not spatial-planning failures. Mistral's Nevada route is a world-knowledge error; Gemini's Friday-as-weekend error is a schedule lookup error. The paper even concedes Mistral cannot search the internet in real time. With one run per model-scenario cell and no temperature control, the data cannot separate \"cannot plan\" from \"lacked a fact on this draw.\" The stress-test concern lands: the strong categorical claim depends on an attribution the experiment does not isolate. The right conclusion is that these three models, used zero-shot without tools, are unreliable navigation assistants—which is itself a useful safety finding for car companies.\n\nThe scoring weights in Section 4.4 are arbitrary, and the rankings are the weakest part; I would ignore them. The six hand-picked routes are a real limitation, and the paper says so. Those are addressable, not fatal. The citation pattern is fine—it engages both the optimistic and skeptical prior work.\n\nThis paper is for people building LLM-based voice assistants in vehicles and for anyone studying tool-augmented versus standalone planning. It deserves a serious referee; the core observation is sound even if the categorical language needs to be dialed back.","headline":"Real failure cases worth documenting; the 'fundamentally unsuitable' conclusion overstates what six zero-shot runs can show.","tokens_in":17155,"tokens_out":2790,"would_cite":true,"duration_ms":26112,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"All three LLMs failed every real-world route test in this study","keywords":["LLM path planning","real-world navigation","turn-by-turn navigation","vision-and-language navigation","GPT-4","Gemini","Mistral","navigation safety"],"falsifier":"Re-run the exact six scenarios with the same three models ten times each, using varied prompt phrasings, and map every output onto a road or campus network; if any model produces a fully connected route that reaches the destination with no wrong turns or invented landmarks on any trial, the paper's claim that none of the tested models can accurately plan any of the paths would be falsified.","tokens_in":16199,"feed_emoji":"🗺️","tokens_out":7368,"duration_ms":58662,"temperature":0.7,"pith_summary":"This paper sets out to test whether large language models can produce usable path-planning directions in real-world settings, not just in simulated environments. The authors gave GPT-4, Gemini, and Mistral six route-planning tasks—three driving routes of rising difficulty and three pedestrian routes using visual landmarks—and compared every output against a commercial navigation app's ground truth. Every model produced numerous major errors on every route, including gaps of up to 83 miles between generated segments, turns in the wrong direction, missed exits, invented destinations, and landmarks that do not exist. The paper concludes that the tested LLMs are not capable of real-world path planning and are fundamentally unsuitable for reliable navigation, and urges caution for car makers integrating LLM voice assistants.","feed_headline":"All three LLMs failed every real-world route test","feed_subtitle":"Route gaps up to 83 miles and invented destinations make standalone LLM navigation unsafe.","key_machinery":"The load-bearing machinery is the six-scenario evaluation protocol paired with an explicit error taxonomy. Three turn-by-turn scenarios (urban, suburban, rural) ask the model for step-by-step GPS-style directions; three vision-and-language scenarios ask for landmark-based walking directions. Every LLM output is mapped onto a road or campus network and compared with the ground-truth route from the Waze navigation app. Errors are classified as major (route discontinuities, opposite-direction turns, wrong or missed exits, failure to reach the destination, nonexistent or wrongly described landmarks) or minor (small local misdirections), and the paper also records total discontinuous miles and number of turns to rank the models. This comparison is what supports the conclusion that all three models are unreliable.","core_discovery":"The central claim is that none of the tested LLMs—GPT-4, Gemini, and Mistral—could accurately plan any of the six paths. The paper demonstrates this by classifying errors into major and minor categories: discontinuities where the generated route skips road segments (the largest was 83.3 miles), directions that send drivers onto a highway in the opposite direction, wrong or missed exits, routes that never reach the destination, and visual landmarks that either do not exist or are described incorrectly. Even the best-performing model varied by scenario, and no model met the time-constraint requirement except GPT-4 on two of the three driving tasks. The authors conclude that the tested LLMs are 'fundamentally unsuitable for reliable path-planning applications' and should not be used to direct vehicle navigation.","pith_inferences":["The pattern of plausible-sounding but disconnected routes suggests these LLMs generate route-like text from associations rather than querying a consistent spatial model of the road network; if true, the same failure would appear in any unfamiliar geography, not just the tested region.","A natural benchmark extension would sample dozens of routes across multiple cities, road types, and languages; if error rates drop sharply for fresh models, the 'fundamentally unsuitable' verdict may be limited to the specific models and prompt style tested here.","The authors' 'reality check' suggestion could be operationalized as a separate route-verification module that validates start, end, and each turn against a map API before the LLM's answer is displayed; this would be a concrete testable extension of their proposal.","Because all six routes start from the same university campus, the study does not test origin-point diversity; adding varied origins and destinations would test whether route-planning errors concentrate near the model's knowledge of a locale."],"forward_implications":["Car makers should not deploy current LLMs as standalone navigation planners; a route with an 83-mile discontinuity or an invented destination would strand or misdirect a driver.","LLM-based assistants that touch navigation need an external verification layer—for example, checking every turn and waypoint against a map database—before any route is shown to a user.","Model scale does not predict path-planning skill: the 7-billion-parameter Mistral sometimes outperformed the much larger GPT-4, and no model was consistently best.","Prompt follow-ups did not fix the underlying failures; roughly 78% of driving prompts required follow-up questions, and major errors still appeared in the final answers.","Future work should focus on mechanisms for reality checks, in-context transparency about limitations, and fine-tuned smaller models rather than ever-larger general-purpose models."],"supporting_citations":[{"why":"Claimed GPT-3.5-turbo outperforms A* and RRT in path planning; this optimistic baseline is what the paper tests against in real-world settings.","marker":"(Latif 2024)"},{"why":"Proposed LLM-A*, using an LLM to generate waypoints with A* connecting them; represents the hybrid approach that the paper's failures motivate.","marker":"(Meng et al. 2024)"},{"why":"Introduced NavGPT, a purely LLM-based navigation agent; the paper's vision-and-language scenarios directly test this line of work.","marker":"(Zhou, Hong, and Wu 2024)"},{"why":"Benchmark showing LLMs struggle to generalize to larger or more obstacle-heavy environments; supports the observed failure pattern.","marker":"(Aghzal, Plaku, and Yao 2023)"},{"why":"Critical benchmark finding a very low success rate for LLM planning; the paper's real-world results align with this.","marker":"(Valmeekam et al. 2023)"},{"why":"Argues LLMs rely on fine-tuning or human-in-the-loop prompting to plan; the paper's follow-up prompting procedure is relevant to this claim.","marker":"(Kambhampati 2024)"},{"why":"VELMA, an embodied LLM agent for vision-and-language navigation; provides context for the experimental design of the VLN scenarios.","marker":"(Schumann et al. 2024)"}],"fun_headline_variants":["LLMs flunk real-world route planning with 83-mile gaps","GPT-4, Gemini, and Mistral all fail every route test","Path-planning LLMs deemed fundamentally unsuitable for navigation","Real-world routes expose LLM navigation errors, missing exits","LLM route errors: 83-mile gaps, wrong exits, phantom landmarks"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The conclusion that the tested LLMs are fundamentally unsuitable for path planning rests on the assumption that six routes in one U.S. region and one prompt style represent real-world path planning as a whole; the paper itself concedes this in its limitations section.","fun_headline_variants_meta":{"raw":{"variants":["LLMs flunk real-world route planning with 83-mile gaps","GPT-4, Gemini, and Mistral all fail every route test","Path-planning LLMs deemed fundamentally unsuitable for navigation","Real-world routes expose LLM navigation errors, missing exits","LLM route errors: 83-mile gaps, wrong exits, phantom landmarks"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000572,"raw_usage":{"total_tokens":2608,"prompt_tokens":752,"completion_tokens":1856,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":368,"completion_tokens_details":{"reasoning_tokens":1767}},"tokens_in":368,"tokens_out":1856,"duration_ms":12955,"temperature":1.0,"reasoning_tokens":1767,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T11:41:45.414515+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the exact six scenarios with the same three models ten times each, using varied prompt phrasings, and map every output onto a road or campus network; if any model produces a fully connected route that reaches the destination with no wrong turns or invented landmarks on any trial, the paper's claim that none of the tested models can accurately plan any of the paths would be falsified.","supporting_citations":[],"review_version":1}