Pith. sign in

Paper Citation Record · LEDGER

MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios

As of 17 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 2 inbound Pith citation observations for arXiv:2412.19406.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.19406 v1

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T00:42:44.346158Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:15:52.844188Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-19T11:13:02.891653Z

Reference resolution

42 of 42 outbound references displayed

  • verified exact3
  • verified fuzzy23
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 87c0b55f-b40d-4b62-8e3f-8a007d717d1f · outbound

This paper cites Traffic sign interpretation via natural language description,.

MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios Traffic sign interpretation via natural language description,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:42:45.104195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T00:42:44.143058Z digest=sha256:91a2972524d3812d9d38bf1f1fc4aab046ca60b88d4ce7f9268837fd8977e911

Observation e6711e8b-336b-4441-ac12-aca8caee8ad9 · outbound

This paper cites Transcrib3D: 3D Referring Expression Resolution through Large Language Models.

MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios Transcrib3D: 3D Referring Expression Resolution through Large Language Models

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-11T00:42:44.731141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T00:42:44.148383Z digest=sha256:e2e8f7269491da9008da5cbb12e7269e6d022167ff18438438eb7c4d6b297293

Observation 8f18536d-d3bf-49be-b7fa-2f57fbce27ea · outbound

This paper cites Large Language Models Powered Context-aware Motion Prediction in Autonomous Driving.

MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios Large Language Models Powered Context-aware Motion Prediction in Autonomous Driving

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T00:42:44.154073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:42:44.154073Z digest=sha256:7cacc9764ef2945796f0c89112dcc3c4b054414956b7e84f8e32d2154a4b0ad1

Observation ff33035c-5fdd-4795-8676-1349756d0fa6 · outbound

This paper cites Rlingua: Improving reinforcement learning sample efficiency in robotic manipulations with large language models,.

MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios Rlingua: Improving reinforcement learning sample efficiency in robotic manipulations with large language models,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:42:45.086704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T00:42:44.159597Z digest=sha256:31cef195ed9dc8df8cd1bcb2af07bbbe3be350167481f1b686b88a853a842c81

Observation 4cf745ae-1036-4eda-85b0-812ffbaa0170 · outbound

This paper cites The Llama 3 Herd of Models.

MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios The Llama 3 Herd of Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T00:42:44.165165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:42:44.165165Z digest=sha256:a01f5d7a5a3baa4ed8a2b1596aae7538c527bc2b235d2a2ec29638e536cde627

Observation 7be3f154-15f2-45e0-b790-5e2a06f9fe14 · outbound

This paper cites GPT-4 Technical Report.

MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios GPT-4 Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T00:42:44.170821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:42:44.170821Z digest=sha256:3d7faef9a44d7202f680a33c81f8606197204c9ff49a90cae015976ff0d4708e

Observation 014d3312-670b-4b8f-bf0a-2f4624944513 · outbound

This paper cites Drive like a human: Rethinking autonomous driving with large language models,.

MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios Drive like a human: Rethinking autonomous driving with large language models,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:42:45.071072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T00:42:44.176903Z digest=sha256:2686f01c586fd61a77680809ee9d56a005cf44902ffc2593a4962ecece218796

Observation 029cb633-85fa-4179-8147-3a39251a4529 · outbound

This paper cites Trafficgpt: Viewing, processing and interacting with traffic foundation models,.

MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios Trafficgpt: Viewing, processing and interacting with traffic foundation models,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:42:45.055856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T00:42:44.181581Z digest=sha256:aaf4df2c000441a855a6fc82bc324186f4eea926fa6f1b6b8d230d5484369492

Observation 41efa04d-4b54-4b72-90e6-cfd3171a36da · outbound

This paper cites EMMA: End-to-End Multimodal Model for Autonomous Driving.

MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios EMMA: End-to-End Multimodal Model for Autonomous Driving

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T00:42:44.186156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:42:44.186156Z digest=sha256:ac5a6f9e42518b186da0c66c98a6bfadc3deedc0a19a4f9fbd2eed55049343be

Observation 293799c0-70cc-4cb8-bb7a-7cdd9a0f7b76 · outbound

This paper cites DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models.

MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T00:42:44.191189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:42:44.191189Z digest=sha256:e62656b0730423681162b8053109fc6ecaf10d818bfd12b4e45da684f3f87a8d

Observation 06bc1c7a-60c4-4779-8cb7-51ac5d3d867a · outbound

This paper cites IDD-X: A Multi-View Dataset for Ego-relative Important Object Localization and Explanation in Dense and Unstructured Traffic.

MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios IDD-X: A Multi-View Dataset for Ego-relative Important Object Localization and Explanation in Dense and Unstructured Traffic

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-11T00:42:44.637871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T00:42:44.196127Z digest=sha256:52c010596587a750d46b770cb2c3a23de1323c4b38df189148381f38c691fe26

Observation 4a4f0d62-5085-467e-9e39-741b8095069f · outbound

This paper cites Rank2tell: A multimodal driving dataset for joint importance ranking and reasoning,.

MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios Rank2tell: A multimodal driving dataset for joint importance ranking and reasoning,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:42:45.040505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T00:42:44.201247Z digest=sha256:27ae1b3bac2deee9bd79e3dbb85d0116bc89727670f2b7e5d55c2fab2e51dd83

Observation fce9ec8e-7fbf-4051-a382-fc628571f259 · outbound

This paper cites Drama: Joint risk localization and captioning in driving,.

MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios Drama: Joint risk localization and captioning in driving,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:42:45.025211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T00:42:44.205636Z digest=sha256:e8e0de84fc7bc893c75dbef9389c0bfdad7c5dda707f6e7edc29c2b4f2483ba2

Observation e42e428e-9a98-44b8-af07-48b9f5ae7d38 · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T00:42:44.210191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:42:44.210191Z digest=sha256:7afc174d366a311244d6bdacf64d2b6e485cf22a69ffbc41c335108242c7db3a

Observation 7bca6ded-722b-4fb3-a332-78bbb5938f42 · outbound

This paper cites Instructblip: Towards general-purpose vision- language models with instruction tuning,.

MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios Instructblip: Towards general-purpose vision- language models with instruction tuning,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T00:42:44.215475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:42:44.215475Z digest=sha256:23c743b3a6da89a6377c2d613eb9cb17029b56ffe19e97ecf5675d2bc708307c

Observation 6aad559a-12cf-4186-a68f-b45a408e7178 · outbound

This paper cites LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention.

MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T00:42:44.220696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:42:44.220696Z digest=sha256:4cbe0251ae18fc6a8a8e609e6eee423c37aa223079a74637c6c6176b846b43d3

Observation 429c9b32-8a8f-4f05-939c-d7e54e0d693b · outbound

This paper cites A hybrid cnn-lstm approach for image caption generation,.

MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios A hybrid cnn-lstm approach for image caption generation,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:42:44.999001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T00:42:44.226296Z digest=sha256:bb506c6df9c6fd1c0dcd372cc02c28717c54404dedc2807e05b20cbe7b8fa8bb

Observation 29f315b2-b4f3-47dd-b46b-ab66ce1b7520 · outbound

This paper cites Improving pre-trained cnn-lstm models for image captioning with hyper-parameter optimization,.

MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios Improving pre-trained cnn-lstm models for image captioning with hyper-parameter optimization,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:42:44.984581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T00:42:44.231820Z digest=sha256:f6c1382a1d4c2f4808adb65261703ba57d61a2218a8f7038578cab265a72716e

Observation f2371168-6553-4872-85e2-e634d7785124 · outbound

This paper cites Benet: bi-directional enhanced network for image captioning,.

MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios Benet: bi-directional enhanced network for image captioning,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:42:44.970286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T00:42:44.237336Z digest=sha256:ba7a6fc31c0be82a0739f2280e4b8ada4ec66976e8ef39ffd6ff5aa6e6889ac8

Observation 49245009-b56e-4ab6-8525-46a9d9d14be6 · outbound

This paper cites Regular constrained multi- modal fusion for image captioning,.

MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios Regular constrained multi- modal fusion for image captioning,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:42:44.956207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T00:42:44.242072Z digest=sha256:92b47b290babdd004a768179f252e85126ba3902420b35764e3639d78ce62b85

Observation 8d2a82f4-8da4-43ed-81d2-a8d1c1f07106 · outbound

This paper cites A dual-feature-based adaptive shared transformer network for image captioning,.

MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios A dual-feature-based adaptive shared transformer network for image captioning,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:42:44.941678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T00:42:44.246567Z digest=sha256:d70bfed943ae9ca1ba3b18e0e0007ea26cffa13f62140cfddaf593845856744b

Observation 3bea342d-20f6-4e7c-af8e-c01205514a6f · outbound

This paper cites Dual-adaptive interactive transformer with textual and visual context for image captioning,.

MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios Dual-adaptive interactive transformer with textual and visual context for image captioning,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:42:44.925408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T00:42:44.250928Z digest=sha256:e92c886c72d26e1cf31f68f9ea241c08738712796b605d433cc1992e65f893f1

Observation 2b6b9a87-47c4-4cb7-9307-c8d73441dbf8 · outbound

This paper cites Visual instruction tuning,.

MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios Visual instruction tuning,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:42:44.910910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T00:42:44.255358Z digest=sha256:dbd87ed4c291c935ed2698c69b9e900e5d2c40b74deb05c45078c5e4da54fc3a

Observation b99e2775-e418-488e-9782-f5f7355c7736 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T00:42:44.259899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:42:44.259899Z digest=sha256:9503147baf3edd57b91dd2d3522be6a6610b041883101e6e21a04aa3ef9b50d0

Observation 6b280f1a-f7ee-4211-8a88-65c6165f75af · outbound

This paper cites eP-ALM: Efficient Perceptual Augmentation of Language Models.

MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios eP-ALM: Efficient Perceptual Augmentation of Language Models

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-08-11T00:42:44.572557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T00:42:44.264746Z digest=sha256:8d50498ec6eb1c49f3b2599dfefebac050739eb783495b77098b8944adf72df2

Observation fefbb950-5ba8-4c67-b5cc-67237a7d2e80 · outbound

This paper cites Vlaad: Vision and language assistant for autonomous driving,.

MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios Vlaad: Vision and language assistant for autonomous driving,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:42:44.896740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T00:42:44.269806Z digest=sha256:531bf5653d7d4a1f48baa2891bdc53fd972fa3f6f478bfa00fd4cb9d88900a9b

Observation 645adf7e-4c5c-4b01-bc9e-f59824a70e34 · outbound

This paper cites Dolphins: Multimodal language model for driving,.

MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios Dolphins: Multimodal language model for driving,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:42:44.881778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T00:42:44.274317Z digest=sha256:e57bc2eb3a858e50eb58903c25991a1033e5c33da5c13cbf4c5e497262f16d23

Observation 2ec9fcc0-44ed-4f91-bd7e-a78245fe51dc · outbound

This paper cites Drivegpt4: Interpretable end-to-end autonomous driving via large language model,.

MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios Drivegpt4: Interpretable end-to-end autonomous driving via large language model,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:42:44.865695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T00:42:44.279054Z digest=sha256:b1c7f42c6f0f68b05864556caa28459d3cccab0ac730ef587c4c616c97e2d13b

Observation 060ea510-2359-4eac-bbf6-b29869458082 · outbound

This paper cites Holistic Autonomous Driving Understanding by Bird's-Eye-View Injected Multi-Modal Large Models.

MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios Holistic Autonomous Driving Understanding by Bird's-Eye-View Injected Multi-Modal Large Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T00:42:44.283698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:42:44.283698Z digest=sha256:cc0d86eb45ac31af8bd6c2a2bedbe6704bb416ff72ec41dcf708e0a9f50ce429

Observation 49be8927-e62b-4a1c-bd5d-947d15130ca9 · outbound

This paper cites Rag-driver: Generalisable driving explanations with retrieval-augmented in-context learning in multi-modal large language model,.

MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios Rag-driver: Generalisable driving explanations with retrieval-augmented in-context learning in multi-modal large language model,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T00:42:44.288850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:42:44.288850Z digest=sha256:f3341d04eef6f78f5a11c49048f54c7156ca636cff111232ad9951a57430e133

Observation 2ee5c6f2-2ec7-4ffd-8f8f-90ff05d76dc3 · outbound

This paper cites HiLM-D: Enhancing MLLMs with Multi-Scale High-Resolution Details for Autonomous Driving.

MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios HiLM-D: Enhancing MLLMs with Multi-Scale High-Resolution Details for Autonomous Driving

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T00:42:44.293660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:42:44.293660Z digest=sha256:cfafc26c39747cbf37c1a63c9d4097b4450a9d71e0d5ce514bdac4f4bb398ae9

Observation ec1e870d-a8d1-4e66-bb84-cd6cef2588ae · outbound

This paper cites Deep residual learning for image recognition,.

MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios Deep residual learning for image recognition,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T00:42:44.298567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:42:44.298567Z digest=sha256:139a68d699554bdf4a73789941ed3a12b3b5289ade16bdfe5088483b3dd06bc5

Observation 124f003b-3f36-4157-8831-2915137583bc · outbound

This paper cites Faster r-cnn: Towards real- time object detection with region proposal networks,.

MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios Faster r-cnn: Towards real- time object detection with region proposal networks,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:42:44.839808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T00:42:44.303123Z digest=sha256:568719bb2731d34faadd4d7f6eced34e2b8b2fb70728cdbb7bc396207fe0496c

Observation 6463d84b-a642-465a-9215-59d3d326ab2b · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T00:42:44.307520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:42:44.307520Z digest=sha256:d7e73f47b2d4a45a1f5488d6522f9c02fba8a3f7baacc63e63b685903fa3944c

Observation d355e3be-536c-4d40-a7fe-1da4164d81a5 · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows,.

MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios Swin transformer: Hierarchical vision transformer using shifted windows,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:42:44.825053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T00:42:44.312055Z digest=sha256:42803d1dad60812a1df625e21f844ea8e2b9f2316a6e7fc4ba50dc41fc47ea15

Observation 2d2e1e52-d195-4943-afde-07ab9425001f · outbound

This paper cites SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing.

MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T00:42:44.317027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:42:44.317027Z digest=sha256:62c77f270bdde054575f73ca3fb895e501361b889430fc62c8fe80e248228af2

Observation 2c52151b-2b49-4d90-8f75-8cafadc65744 · outbound

This paper cites Generalized intersection over union: A metric and a loss for bounding box regression,.

MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios Generalized intersection over union: A metric and a loss for bounding box regression,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:42:44.809432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T00:42:44.322245Z digest=sha256:dc66ebc2fcbc081bcf00ef68e8857088268c031373677bed87b1468018dd885a

Observation cbf66b7c-001d-4d95-84a8-d59996592f6d · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation,.

MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios Bleu: a method for automatic evaluation of machine translation,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:42:44.794099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T00:42:44.326751Z digest=sha256:a2731eb8d92c22fece3191dc67b42e9c73b47090923b5e348b6d23c6cfe2066a

Observation 4cc20059-aae8-43cd-b748-352095ae956f · outbound

This paper cites Meteor: An automatic metric for mt evaluation with improved correlation with human judgments,.

MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios Meteor: An automatic metric for mt evaluation with improved correlation with human judgments,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:42:44.779090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T00:42:44.331519Z digest=sha256:8aa1a744fbf11d5eecb739aed95654dbf8f420a30f6dce6be744c583fc0f5d3f

Observation 69c83753-9638-4893-9e35-dd51af1bd1d2 · outbound

This paper cites Cider: Consensus- based image description evaluation,.

MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios Cider: Consensus- based image description evaluation,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:42:44.762900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T00:42:44.336391Z digest=sha256:24f53c279aa16bbe5c995bdeb020383ff97d7d7a89dba9fc6e2a798a64a54788

Observation 01cc0114-9ffa-4273-a8ca-0730e4cdff10 · outbound

This paper cites Swin transformer v2: Scaling up capacity and resolution,.

MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios Swin transformer v2: Scaling up capacity and resolution,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:42:44.747378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T00:42:44.341007Z digest=sha256:99775cc09720b71dddc51a316fadd7a68468940cf27eeaa5a40959ef0a4a694b

Observation 1a2ffc71-991a-466a-bb94-01e9ee4edce7 · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T00:42:44.346158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:42:44.346158Z digest=sha256:5d625e051c2cea75df8857c52b323e40f371d33a4d7c20e1365bdf48d8b75425

Pith citing papers

Observation 14ee273e-135e-47aa-9a9f-7a009c3d6e88 · inbound

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models cites this paper.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:13:02.893274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:541531cd931ef10faf3630f36b57298a0d1c103f053847214dab577d463faaa0

Observation 75cccf7c-c386-4424-b487-f15c47ba9a51 · inbound

MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding cites this paper.

MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:52.844188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:52.844188Z digest=sha256:064e7ce6b91f8f54aceb03ce630192657222d29dd1729a5a9cd73c645cd3b09a