{"id":"be689546-da0f-44d4-89a3-975abeccd3b4","arxiv_id":"2501.04437","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":2.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The paper surveys recent work, models, applications, and challenges of using LLMs in intelligent transportation systems, without presenting new experimental results.","lead":"This paper reviews how large language models are being used in intelligent transportation systems, covering traffic prediction, signal control, autonomous driving, and public transit. It is a survey of existing work rather than a new experiment, intended to guide researchers and practitioners.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central claim is supported by evidence that conflates LLMs with generic transformer models; a systematic audit of cited systems would determine whether LLMs specifically, rather than transformers generally, drive the reported ITS gains.","rationale":"The reader's verdict is CONDITIONAL with medium correctness risk, citing factual errors and an incoherent passage. My concern is more specific: the review systematically conflates LLMs with generic transformer architectures and with non-LLM frameworks, so the evidence cited for the central claim does not isolate the effect of large language models. This is a load-bearing weakness because the paper's conclusion that 'LLMs can revolutionize ITS' is only as strong as the evidence that LLMs specifically—rather than transformers generally—enable the reported gains. I do not think this changes the verdict to REJECT: the paper is a narrative review, and the core possibility that LLMs have useful applications in ITS remains plausible even after removing the conflated examples. However, the manuscript needs major revisions to accurately characterize its evidence and to remove the incoherent 'Sora's LLMs' passage and the GPT-3/GPT-4 inconsistency. The conditional verdict already implies such revisions, so I keep the verdict UNCHANGED. My agreement with the reader is partial because the reader identified real-time deployment as the weakest assumption, whereas I find the conflation of evidence to be the most load-bearing issue for the central claim; both point to the need for substantial revision.","tokens_in":45910,"tokens_out":4964,"duration_ms":50653,"concrete_test":"Create a table classifying every system cited in Sections IV and V (e.g., TrafficBERT, LA-Light/LLMLight, STPT, TransTTE, BERT4ITS/BERT-ITS, GGT, STLLM, LG-Traj, VTPM, LiMeda, OpenFlamingo, GPT-NeoX) as: (a) a true LLM (large-scale language model pretrained primarily on text), (b) a non-LLM transformer or domain-adapted BERT variant, or (c) an LLM-based prototype evaluated only in simulation or synthetic settings. If categories (b) and (c) contain the majority of claimed successes, the review's central claim that LLMs specifically can revolutionize ITS is not supported by its own cited evidence.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that LLMs can 'potentially revolutionize' ITS. The load-bearing support is the set of cited applications in Sections IV and V. However, many of these cited systems are not large language models at all. TrafficBERT is a BERT-like time-series transformer pretrained on traffic data, not a language model; STPT and TransTTE are general transformer architectures for spatiotemporal data; BERT4ITS and BERT-ITS adapt BERT for time-series and sensor data, not for language understanding; and the 'decentralized LLMs' in Section III include frameworks and tools such as Colossal-AI, Mesh TensorFlow, Petals, and GPT-NeoX, which are not language models themselves. The case studies in Section V (LLM-Light, STransformer, TransTTE, BERT4ITS) are largely simulation-based or use non-LLM transformers. If these are excluded, the remaining direct evidence for LLM-specific benefits is thin: LA-Light and related LLM-based signal control are shown in synthetic or limited real-world scenarios, and several other cited works are research prototypes. The review therefore does not establish that LLMs as such, rather than transformer architectures broadly, produce the reported ITS improvements. This is compounded by an incoherent inserted passage in Section II-E about 'Sora's LLMs' and the 'torrent protocol,' which appears to be an artifact and undermines the reliability of the synthesis, and by an internal inconsistency in Section II-A that calls GPT-3 the 'latest iteration' while later describing GPT-4. Because the central claim rests on this unreliable and conflated evidence base, it is not currently supported in the way the paper presents it.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript is a survey of the role of large language models in intelligent transportation systems. It motivates the topic with market projections, reviews centralized LLM families (GPT, T5, BERT, LLaMA, Falcon) and decentralized training/inference frameworks (GPT-NeoX, OpenFlamingo, BLOOM, Colossal-AI, Mesh TensorFlow, Petals), and then surveys uses in traffic prediction, signal control, route planning, autonomous driving, public transport, V2X, ADAS, traffic control centers, smart cities, pedestrian flow, and multimodal transport. It also presents five case studies and a discussion of challenges and future directions including data, computation, ethics, integration, latency, scalability, edge computing, and 6G/quantum computing. The paper's central claim is that LLMs can potentially revolutionize user interaction with transportation systems and substantially improve ITS functions.","tokens_in":46196,"tokens_out":6962,"duration_ms":63283,"significance":"The topic is timely and the reference list is broad, covering 2023-2024 developments, and the paper offers a useful high-level taxonomy of challenges. The strongest contribution is the organization of applications and the enumeration of open problems rather than a quantitative comparison. However, several load-bearing examples in Sections IV and V are not language models, and the text contains internal inconsistencies; as a result, the manuscript does not currently establish the claimed LLM-specific advantages.","major_comments":[{"comment":"Several systems presented as evidence of LLM benefits are not large language models. TrafficBERT (V-A) is a BERT-style transformer pretrained on traffic time-series data; STransformer/STPT (V-C), TransTTE (V-D), and BERT4ITS/BERT-ITS (V-E) are general transformer architectures for spatiotemporal or time-series data without natural-language training. Including these as 'case studies of LLMs' conflates LLMs with transformer architectures generally. Since the central claim that LLMs improve ITS depends on these examples, the authors must reclassify each cited system as an LLM, a non-LLM transformer baseline, or a framework, and then reassess which remaining systems (e.g., LA-Light, TF-LLM) support the LLM-specific claim.","section":"Sections IV-A, V-A, V-C, V-D, V-E"},{"comment":"The paragraph beginning 'Distributed architectures distribute...' introduces 'Sora's LLMs' without definition, describes data being stored 'within currently consigned data-protecting devices,' refers to a 'torrent protocol' for decentralized model updates, and ends with the incomplete sentence 'would not be just like that.' This passage is not coherent and cannot be verified against the cited literature. It should be removed or rewritten with proper technical definitions and citations before the survey can be considered reliable.","section":"Section II-E"},{"comment":"Table II contains factual errors and category confusions. LLaMA is described as an 'Encoder-decoder Transformer,' but LLaMA models are decoder-only transformers. Colossal-AI, Mesh TensorFlow, and Petals are listed as 'LLM models' with architecture and training entries, although they are distributed training/inference frameworks. GPT-NeoX is also a library rather than a pretrained model, and the 'multi-query attention' attribution appears unsupported. The table should be corrected and reorganized to distinguish models from frameworks.","section":"Table II"},{"comment":"This section repeatedly labels Colossal-AI, Mesh TensorFlow, Petals, and GPT-NeoX as 'decentralized LLMs' and says 'we discuss emerging DLLMs,' but these are tools for distributed computation rather than language models. This conflation affects the technical discussion of communication complexity and model updates, since the complexity of training infrastructure is not the same as the complexity of an LLM itself. The section should separate model architectures from systems/frameworks.","section":"Section III"}],"minor_comments":[{"comment":"GPT-3 is called 'the latest iteration' in the first paragraph of Section II-A, but GPT-4 is discussed later in the same subsection; the wording should be corrected to reflect the chronology.","section":"Section II-A"},{"comment":"There are typographical errors such as 'Colossoal-AI,' 'LlaMa,' 'pedestrain,' and 'logO(n log n)'; these should be corrected throughout.","section":"Sections I-C and III"},{"comment":"The row for reference [47] describes 'generative LLM models in traffic simulation' and mentions DriveGAN, but the cited work is a survey of vision-language models in autonomous driving and ITS; the table entry should match the cited source.","section":"Table I"},{"comment":"The model name is inconsistently written as 'LLMLight,' 'LLM-Light,' and 'LA-Light'; pick one notation and use it consistently.","section":"Sections I and V"},{"comment":"The reference metadata in [109] and [110] appears inconsistent with the cited papers in author names and titles; check and update these entries.","section":"References [109]-[110]"}],"recommendation":"major_revision","confidential_remarks":"The paper would benefit from a systematic technical revision rather than a surface proofread. The reclassification of non-LLM transformers (TrafficBERT, STPT/STransformer, TransTTE, BERT4ITS) and the removal of the incoherent Sora/torrent passage are prerequisites for relying on this survey. If those are addressed, the manuscript could be a useful roadmap; as it stands, the LLM-specific contribution is not established."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First: this is a review, not a research contribution, and it reads like one. Its value is as a wide map of LLM-related work in ITS. For a newcomer, it does a decent job of assembling model names, application areas, and open problems. The challenge and future-directions sections (data, compute, privacy, edge, 6G, quantum) are broad and would help a student scope a thesis. That is the honest upside.\n\nThe soft spots are not minor. The central claim that LLMs can 'revolutionize' ITS is carried by examples that mostly are not LLMs. TrafficBERT is a BERT-style time-series transformer pre-trained on traffic data, not a language model; STPT and TransTTE are general transformer architectures; BERT4ITS adapts BERT for sensor and time-series data. Those are transformers, not LLMs. The 'decentralized LLMs' section lists GPT-NeoX, Colossal-AI, Mesh TensorFlow, and Petals, which are training and distribution frameworks, not LLMs themselves. Once you remove those, the direct evidence for LLM-specific gains is thin, mostly prototype results like LA-Light in synthetic settings. The survey needs to either reframe its claim as covering transformer architectures broadly or tighten the evidence to actual LLMs.\n\nThere are also internal errors a referee would catch immediately. Section II-A calls GPT-3 the 'latest iteration' while later describing GPT-4. Section II-E contains an incoherent passage about 'Sora's LLMs' and the 'torrent protocol' that belongs nowhere and looks like a paste artifact. Table II labels LLaMA as encoder-decoder when it is decoder-only. These undermine trust in the synthesis.\n\nWhat is genuinely new: not much. The paper explicitly positions itself against several earlier surveys and acknowledges them in Table I. That is honest, but the contribution is organizational rather than novel.\n\nWho is this for? Graduate students and practitioners wanting a broad, non-systematic tour of the area. A corrected version would be genuinely useful. In its current form, I would not cite it, and I would not trust its specific claims without checking the original sources.\n\nRecommendation: a serious editor could send this to peer review, because the topic is timely and a corrected survey would be useful. But it needs major revision first: fix the factual errors, remove the incoherent passage, and audit whether the cited systems are LLMs or general transformers. I would ask for that audit before accepting.","headline":"A broad but uneven survey whose central LLM-specific claim is undercut by its own evidence, much of which is about transformer architectures rather than language models.","tokens_in":46760,"tokens_out":2420,"would_cite":false,"duration_ms":25429,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This review argues that large language models can serve as the reasoning layer of intelligent transportation systems, improving prediction, signal control, and driving assistance while facing real-time and ethical barriers.","keywords":["intelligent transportation systems","large language models","traffic prediction","traffic signal control","autonomous driving","V2X communication","decentralized LLMs","edge computing"],"falsifier":"A field deployment that measures end-to-end latency of an LLM-driven traffic signal controller against the signal phase cycle would settle the practical claim: if median decision latency during peak traffic exceeds the phase-change interval, or if the controller causes a safety-critical failure, the real-time viability thesis is falsified. A second observation would be a benchmark where an LLM-based traffic forecaster, trained on the same historical data, fails to beat a standard graph neural network on held-out incident days.","tokens_in":1485,"feed_emoji":"🚦","tokens_out":2265,"duration_ms":79681,"temperature":0.7,"pith_summary":"This review sets out to establish that large language models, though built for text, can serve as the reasoning layer of intelligent transportation systems: predicting traffic, timing signals, planning routes, assisting autonomous vehicles, scheduling transit, and enabling V2X communication. It treats the transformer's self-attention mechanism as the bridge, since the same machinery that captures word dependencies can capture spatiotemporal dependencies in traffic data, and the same model that generates language can generate alerts, explanations, and control suggestions. The paper's own evidence is a set of research prototypes in which LLM-based or transformer-based systems match or beat classical baselines for traffic forecasting and signal control. If the claim holds, transportation agencies would integrate LLMs into traffic management rather than treating them as isolated chatbots, while meeting data quality, latency, privacy, and network reliability constraints.","feed_headline":"Language models could run the traffic control room","feed_subtitle":"A review links LLMs to signal timing, autonomous driving, and transit, with latency and data limits as the barriers.","key_machinery":"The central object is the transformer's multi-head self-attention mechanism, generalized beyond language so it can encode both temporal and spatial dependencies in traffic data; the paper also relies on pre-training and fine-tuning, and on decentralized or edge deployment, as the machinery for making LLMs practical in ITS. The attention mechanism lets a model weigh every relevant input at once, which is what makes it suited to forecasting traffic from many heterogeneous streams and to explaining its own control choices. Pre-training supplies general world knowledge that can be adapted to traffic with limited in-domain data, while decentralization distributes inference across vehicles, roadside units, and edge nodes to cut latency and preserve data locality.","core_discovery":"The paper argues that LLMs are broadly applicable to transportation management and should be integrated into future ITS: they can interpret, predict, and respond to complex scenarios within transportation networks, improving decision-making and traffic management. Its central claim is that the transformer's attention mechanism can be transferred from language to traffic data, letting the same models that understand text also model spatiotemporal traffic patterns and generate human-readable explanations for control actions. The paper also claims that decentralized and edge-deployed LLM architectures are needed to make this practical in connected vehicle environments, and that data quality, computational cost, bias, privacy, and network latency are the barriers that future research must remove.","pith_inferences":["If LLMs enter live traffic systems, their first safe role is likely advisory: producing explainable signal recommendations, incident summaries, and passenger alerts, with a human or a verified controller retaining final authority, because current latency and verification gaps make direct control risky.","The case studies reviewed are mostly research prototypes and simulation results, so the natural next step is shared, standardized benchmarks that measure not only forecast accuracy but also inference latency, bandwidth, and safety outcomes across LLM-based and conventional ITS controllers.","The paper's emphasis on edge computing suggests a concrete testable route: small fine-tuned LLMs running at roadside units could meet real-time constraints while large cloud models handle non-critical planning, a hybrid architecture the review points toward but does not itself test."],"forward_implications":["Traffic prediction can move from task-specific deep networks to pre-trained LLMs that integrate text, weather, and sensor data with limited labeled data, reducing the dependence on large historical traffic datasets.","Traffic signal control can become human-mimetic and explainable: LLM-based agents can justify each signal phase change, which helps operators trust and audit decisions in complex urban intersections.","Autonomous driving and V2X communication can gain a linguistic layer that interprets instructions, road signs, and alerts, and generates context-aware responses, improving situational awareness and hazard detection.","Public transit can use LLM-driven analysis of traffic patterns, sensor data, and rider information to adjust schedules in real time and provide passengers with useful, up-to-date travel information.","Realizing these applications at scale requires decentralized and edge deployment, model compression, and new network architectures, because cloud-only LLMs are too slow and bandwidth-hungry for safety-critical traffic loops."],"supporting_citations":[{"why":"It supplies the survey's baseline map of frontier AI, foundation models, and LLM applications in intelligent transportation systems.","marker":"[23]"},{"why":"It describes the LA-Light human-mimetic traffic signal control framework, the central example in the signal-optimization and case-study sections.","marker":"[28]"},{"why":"It argues that large language models could support intelligent transportation and provides the paper's framing of model-based opportunities and risks.","marker":"[30]"},{"why":"It introduces TrafficBERT, a pre-trained transformer for long-range traffic flow forecasting, which the paper uses as case-study evidence.","marker":"[132]"},{"why":"It presents the LA-Light hybrid decision approach that combines perception and decision tools with LLM reasoning for traffic signal control.","marker":"[143]"},{"why":"It provides evidence on using LLMs as traffic signal control agents, including capacity and data limitations that the paper cites.","marker":"[144]"},{"why":"It proposes using LLMs as the decision-making brain of autonomous vehicles, which is the core idea of the paper's autonomous-driving section.","marker":"[156]"},{"why":"It introduces BERT4ITS, a BERT-based deep learning framework for traffic prediction and incident detection, used as a case study.","marker":"[199]"}],"fun_headline_variants":["LLMs for traffic: promises, pitfalls, and the road ahead","How language models could shape the future of traffic management","Integrating LLMs into smart transportation: a critical review","Language models in ITS: from prediction to autonomous driving","The potential and limits of LLMs in intelligent transport systems"],"cache_read_input_tokens":48896,"weakest_assumption_plain":"The load-bearing premise is that LLM-based systems can be run in real time in safety-critical transportation settings with acceptable latency, reliability, and resource use, and that decentralized LLM architectures can meet the communication and computational constraints of vehicular networks.","fun_headline_variants_meta":{"raw":{"variants":["LLMs for traffic: promises, pitfalls, and the road ahead","How language models could shape the future of traffic management","Integrating LLMs into smart transportation: a critical review","Language models in ITS: from prediction to autonomous driving","The potential and limits of LLMs in intelligent transport systems"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00098,"raw_usage":{"total_tokens":4126,"prompt_tokens":879,"completion_tokens":3247,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":495,"completion_tokens_details":{"reasoning_tokens":3166}},"tokens_in":495,"tokens_out":3247,"duration_ms":23586,"temperature":1.0,"reasoning_tokens":3166,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T21:32:28.961202+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A field deployment that measures end-to-end latency of an LLM-driven traffic signal controller against the signal phase cycle would settle the practical claim: if median decision latency during peak traffic exceeds the phase-change interval, or if the controller causes a safety-critical failure, the real-time viability thesis is falsified. A second observation would be a benchmark where an LLM-based traffic forecaster, trained on the same historical data, fails to beat a standard graph neural network on held-out incident days.","supporting_citations":[],"review_version":1}