RailVQA-bench supplies 21,168 QA pairs for ATO visual cognition while RailVQA-CoM combines large-model reasoning with small-model efficiency via transparent modules and temporal sampling.
Retrieval-based interleaved visual chain-of-thought in real-world driving scenarios
5 Pith papers cite this work. Polarity classification is still indexing.
representative citing papers
OmniDrive-R1 boosts VLM reasoning score from 51.77% to 80.35% and answer accuracy from 37.81% to 73.62% on DriveLMM-o1 via reinforcement-driven interleaved multi-modal chain-of-thought with annotation-free grounding.
Alpamayo-R1 introduces a VLA model with a Chain of Causation dataset and multi-stage SFT-plus-RL training that reports 12% better planning accuracy and 35% fewer close encounters versus trajectory-only baselines in driving tasks.
CaVe-VLM-CoT is a closed-loop agentic-RAG framework with Extractor, Retriever, Solver, Citation Injector and Verifier stages plus 23 metrics anchored by CaVeScore that reports 87.1% accuracy on ScienceQA and 55.2% on MMMU without model changes.
SliceScorer combines an exposure-based coverage prior and a neighbor-failure prior into a simple deterministic score for recommending coverage gaps in driving VLMs, embedded in the LLM-orchestrated SliceNav pipeline.
citing papers explorer
-
RailVQA: A Benchmark and Framework for Efficient Interpretable Visual Cognition in Automatic Train Operation
RailVQA-bench supplies 21,168 QA pairs for ATO visual cognition while RailVQA-CoM combines large-model reasoning with small-model efficiency via transparent modules and temporal sampling.
-
OmniDrive-R1: Reinforcement-driven Interleaved Multi-modal Chain-of-Thought for Trustworthy Vision-Language Autonomous Driving
OmniDrive-R1 boosts VLM reasoning score from 51.77% to 80.35% and answer accuracy from 37.81% to 73.62% on DriveLMM-o1 via reinforcement-driven interleaved multi-modal chain-of-thought with annotation-free grounding.
-
Alpamayo-R1: Bridging Reasoning and Action Prediction for Generalizable Autonomous Driving in the Long Tail
Alpamayo-R1 introduces a VLA model with a Chain of Causation dataset and multi-stage SFT-plus-RL training that reports 12% better planning accuracy and 35% fewer close encounters versus trajectory-only baselines in driving tasks.
-
CaVe-VLM-CoT: An Interpretable Vision-Language Model Framework
CaVe-VLM-CoT is a closed-loop agentic-RAG framework with Extractor, Retriever, Solver, Citation Injector and Verifier stages plus 23 metrics anchored by CaVeScore that reports 87.1% accuracy on ScienceQA and 55.2% on MMMU without model changes.
-
What to Test Next: Interpretable Coverage Gap Discovery in Driving VLMs
SliceScorer combines an exposure-based coverage prior and a neighbor-failure prior into a simple deterministic score for recommending coverage gaps in driving VLMs, embedded in the LLM-orchestrated SliceNav pipeline.