A 3-billion-parameter vision-language model distilled from GPT-4o and o3-mini pseudo-annotations matches its teachers on traffic-video captioning metrics.
A review of the effect of traffic and weather characteristics on road safety
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
UNVERDICTED 1representative citing papers
citing papers explorer
-
Structured Prompting and Multi-Agent Knowledge Distillation for Traffic Video Interpretation and Risk Inference
A 3-billion-parameter vision-language model distilled from GPT-4o and o3-mini pseudo-annotations matches its teachers on traffic-video captioning metrics.