Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T00:42:44.346158Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 2 inbound Pith citation observations for arXiv:2412.19406.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T00:42:44.346158Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T19:15:52.844188Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-19T11:13:02.891653Z
42 of 42 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 87c0b55f-b40d-4b62-8e3f-8a007d717d1f · outbound
MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios Traffic sign interpretation via natural language description,
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation e6711e8b-336b-4441-ac12-aca8caee8ad9 · outbound
MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios Transcrib3D: 3D Referring Expression Resolution through Large Language Models
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 8f18536d-d3bf-49be-b7fa-2f57fbce27ea · outbound
MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios Large Language Models Powered Context-aware Motion Prediction in Autonomous Driving
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff33035c-5fdd-4795-8676-1349756d0fa6 · outbound
MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios Rlingua: Improving reinforcement learning sample efficiency in robotic manipulations with large language models,
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 4cf745ae-1036-4eda-85b0-812ffbaa0170 · outbound
MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios The Llama 3 Herd of Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7be3f154-15f2-45e0-b790-5e2a06f9fe14 · outbound
MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios GPT-4 Technical Report
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 014d3312-670b-4b8f-bf0a-2f4624944513 · outbound
MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios Drive like a human: Rethinking autonomous driving with large language models,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 029cb633-85fa-4179-8147-3a39251a4529 · outbound
MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios Trafficgpt: Viewing, processing and interacting with traffic foundation models,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 41efa04d-4b54-4b72-90e6-cfd3171a36da · outbound
MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios EMMA: End-to-End Multimodal Model for Autonomous Driving
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 293799c0-70cc-4cb8-bb7a-7cdd9a0f7b76 · outbound
MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06bc1c7a-60c4-4779-8cb7-51ac5d3d867a · outbound
MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios IDD-X: A Multi-View Dataset for Ego-relative Important Object Localization and Explanation in Dense and Unstructured Traffic
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 4a4f0d62-5085-467e-9e39-741b8095069f · outbound
MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios Rank2tell: A multimodal driving dataset for joint importance ranking and reasoning,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation fce9ec8e-7fbf-4051-a382-fc628571f259 · outbound
MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios Drama: Joint risk localization and captioning in driving,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation e42e428e-9a98-44b8-af07-48b9f5ae7d38 · outbound
MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7bca6ded-722b-4fb3-a332-78bbb5938f42 · outbound
MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios Instructblip: Towards general-purpose vision- language models with instruction tuning,
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6aad559a-12cf-4186-a68f-b45a408e7178 · outbound
MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 429c9b32-8a8f-4f05-939c-d7e54e0d693b · outbound
MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios A hybrid cnn-lstm approach for image caption generation,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 29f315b2-b4f3-47dd-b46b-ab66ce1b7520 · outbound
MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios Improving pre-trained cnn-lstm models for image captioning with hyper-parameter optimization,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation f2371168-6553-4872-85e2-e634d7785124 · outbound
MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios Benet: bi-directional enhanced network for image captioning,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 49245009-b56e-4ab6-8525-46a9d9d14be6 · outbound
MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios Regular constrained multi- modal fusion for image captioning,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 8d2a82f4-8da4-43ed-81d2-a8d1c1f07106 · outbound
MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios A dual-feature-based adaptive shared transformer network for image captioning,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 3bea342d-20f6-4e7c-af8e-c01205514a6f · outbound
MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios Dual-adaptive interactive transformer with textual and visual context for image captioning,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 2b6b9a87-47c4-4cb7-9307-c8d73441dbf8 · outbound
MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios Visual instruction tuning,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation b99e2775-e418-488e-9782-f5f7355c7736 · outbound
MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b280f1a-f7ee-4211-8a88-65c6165f75af · outbound
MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios eP-ALM: Efficient Perceptual Augmentation of Language Models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation fefbb950-5ba8-4c67-b5cc-67237a7d2e80 · outbound
MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios Vlaad: Vision and language assistant for autonomous driving,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 645adf7e-4c5c-4b01-bc9e-f59824a70e34 · outbound
MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios Dolphins: Multimodal language model for driving,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 2ec9fcc0-44ed-4f91-bd7e-a78245fe51dc · outbound
MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios Drivegpt4: Interpretable end-to-end autonomous driving via large language model,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 060ea510-2359-4eac-bbf6-b29869458082 · outbound
MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios Holistic Autonomous Driving Understanding by Bird's-Eye-View Injected Multi-Modal Large Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49be8927-e62b-4a1c-bd5d-947d15130ca9 · outbound
MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios Rag-driver: Generalisable driving explanations with retrieval-augmented in-context learning in multi-modal large language model,
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ee5c6f2-2ec7-4ffd-8f8f-90ff05d76dc3 · outbound
MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios HiLM-D: Enhancing MLLMs with Multi-Scale High-Resolution Details for Autonomous Driving
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec1e870d-a8d1-4e66-bb84-cd6cef2588ae · outbound
MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios Deep residual learning for image recognition,
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 124f003b-3f36-4157-8831-2915137583bc · outbound
MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios Faster r-cnn: Towards real- time object detection with region proposal networks,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 6463d84b-a642-465a-9215-59d3d326ab2b · outbound
MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d355e3be-536c-4d40-a7fe-1da4164d81a5 · outbound
MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios Swin transformer: Hierarchical vision transformer using shifted windows,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 2d2e1e52-d195-4943-afde-07ab9425001f · outbound
MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c52151b-2b49-4d90-8f75-8cafadc65744 · outbound
MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios Generalized intersection over union: A metric and a loss for bounding box regression,
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation cbf66b7c-001d-4d95-84a8-d59996592f6d · outbound
MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios Bleu: a method for automatic evaluation of machine translation,
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 4cc20059-aae8-43cd-b748-352095ae956f · outbound
MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios Meteor: An automatic metric for mt evaluation with improved correlation with human judgments,
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 69c83753-9638-4893-9e35-dd51af1bd1d2 · outbound
MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios Cider: Consensus- based image description evaluation,
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 01cc0114-9ffa-4273-a8ca-0730e4cdff10 · outbound
MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios Swin transformer v2: Scaling up capacity and resolution,
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 1a2ffc71-991a-466a-bb94-01e9ee4edce7 · outbound
MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14ee273e-135e-47aa-9a9f-7a009c3d6e88 · inbound
AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 75cccf7c-c386-4424-b487-f15c47ba9a51 · inbound
MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.