Pith. sign in

REVIEW 2 cited by

Advancing Object Detection in Transportation with Multimodal Large Language Models (MLLMs): A Comprehensive Review and Empirical Testing

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.18286 v1 pith:LTZRHBKX submitted 2024-09-26 cs.CV cs.AIcs.CL

classification cs.CVcs.AIcs.CL
keywords detectiontransportationobjectmllmslargemodelsreviewapplications
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This study aims to comprehensively review and empirically evaluate the application of multimodal large language models (MLLMs) and Large Vision Models (VLMs) in object detection for transportation systems. In the first fold, we provide a background about the potential benefits of MLLMs in transportation applications and conduct a comprehensive review of current MLLM technologies in previous studies. We highlight their effectiveness and limitations in object detection within various transportation scenarios. The second fold involves providing an overview of the taxonomy of end-to-end object detection in transportation applications and future directions. Building on this, we proposed empirical analysis for testing MLLMs on three real-world transportation problems that include object detection tasks namely, road safety attributes extraction, safety-critical event detection, and visual reasoning of thermal images. Our findings provide a detailed assessment of MLLM performance, uncovering both strengths and areas for improvement. Finally, we discuss practical limitations and challenges of MLLMs in enhancing object detection in transportation, thereby offering a roadmap for future research and development in this critical area.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Large Language Models (LLMs) as Traffic Control Systems at Urban Intersections: A New Paradigm

    cs.CL 2024-11 conditional novelty 5.0 of 10

    A fine-tuned GPT-4o-mini detects conflicts in synthetic four-leg intersection scenarios with 83% accuracy and produces traffic-management text with high ROUGE-L scores against the simulator's templated references.

  2. Detecting AI-Generated Text in Educational Content: Leveraging Machine Learning and Explainable AI for Academic Integrity

    cs.CL 2025-01 conditional novelty 4.0 of 10

    XGBoost and Random Forest can distinguish ChatGPT-written cybersecurity paragraphs from human Wikipedia paragraphs with 81 to 83% accuracy, and a narrow XGBoost model beat GPTZero in a three-class test, yet the benchm...

Pith tools