Pith. sign in

Paper Citation Record · LEDGER

Dolphins: Multimodal Language Model for Driving

As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 15 inbound Pith citation observations for arXiv:2312.00438.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.00438 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T04:34:34.313636Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-23T16:35:42.142836Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a3f43326-7900-45c2-bdc4-146e51600f16 · inbound

Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models cites this paper.

Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models Dolphins: Multimodal Language Model for Driving

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:06.872427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:06.872427Z digest=sha256:f4fd0f4398c06d908a2952e3b42469a8fcbce95811bb4d4e91fc22bdb9a2e321

Observation 2df64d50-fdae-4ad5-9a76-5bf2c79e2cf9 · inbound

Visual Adversarial Attack on Vision-Language Models for Autonomous Driving cites this paper.

Visual Adversarial Attack on Vision-Language Models for Autonomous Driving Dolphins: Multimodal Language Model for Driving

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-23T16:35:42.147144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-23T16:35:24.063578Z digest=sha256:d6788210abd537a803fe20822aa441cbc6ce556cf4dd25096c3136b85435dceb

Observation 8e7de841-eaaf-4e62-a729-15887c84e069 · inbound

World knowledge-enhanced Reasoning Using Instruction-guided Interactor in Autonomous Driving cites this paper.

World knowledge-enhanced Reasoning Using Instruction-guided Interactor in Autonomous Driving Dolphins: Multimodal Language Model for Driving

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T19:52:15.211566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:52:15.211566Z digest=sha256:3811a1e2a8e309a4191751677625aeb8696b83b6cf12c0425f1b7ba143e82c2a

Observation 5b039ff6-a4d9-47d5-99f3-f1ff7c772e5f · inbound

A Review of Multimodal Explainable Artificial Intelligence: Past, Present and Future cites this paper.

A Review of Multimodal Explainable Artificial Intelligence: Past, Present and Future Dolphins: Multimodal Language Model for Driving

Reference 274

Resolution
unresolved
no resolver link, observed 2026-08-11T12:33:41.797744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:33:41.797744Z digest=sha256:d55cb91f9f61a4bcc5eea0b9e47386ede6458bd61d48d9836e745ca1bcf00efa

Observation 8d2eb5e0-e746-41d3-ae7c-91c0612868ce · inbound

Visual Large Language Models for Generalized and Specialized Applications cites this paper.

Visual Large Language Models for Generalized and Specialized Applications Dolphins: Multimodal Language Model for Driving

Reference 164

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:09.527124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:09.527124Z digest=sha256:f70e765253ce5ae27817041d519fb23a567c72aa32521d25c29f7cd63504cdab

Observation e66cdf29-19c8-4b39-8f1e-a7e7e353c8f8 · inbound

Are VLMs Ready for Autonomous Driving? An Empirical Study from the Reliability, Data, and Metric Perspectives cites this paper.

Are VLMs Ready for Autonomous Driving? An Empirical Study from the Reliability, Data, and Metric Perspectives Dolphins: Multimodal Language Model for Driving

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T21:44:55.880123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:44:55.880123Z digest=sha256:5661de4bacce277458147fb3cb671a2d0817c7a1f543c0e030122e32d03c0ed7

Observation 91d70c6a-5fcb-43cd-b09f-25ba7e4b8510 · inbound

LeapVAD: A Leap in Autonomous Driving via Cognitive Perception and Dual-Process Thinking cites this paper.

LeapVAD: A Leap in Autonomous Driving via Cognitive Perception and Dual-Process Thinking Dolphins: Multimodal Language Model for Driving

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-10T20:34:11.911195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:34:11.911195Z digest=sha256:0387b20821d0cfabfe78f9f1c7c8fd7e45eea1bea34b4145c1b9f043300a2cbc

Observation 68bb0ec4-245a-4f04-8f36-e3f0ec66f53d · inbound

Explainability for Vision Foundation Models: A Survey cites this paper.

Explainability for Vision Foundation Models: A Survey Dolphins: Multimodal Language Model for Driving

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-10T17:26:35.310664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:26:35.310664Z digest=sha256:245cbf850a81ebe94107c0b470540fa7d4eceddb3418f4560087305e26fc8425

Observation 9999d067-60d3-4ceb-ad95-aa859421de92 · inbound

Black-Box Adversarial Attack on Vision Language Models for Autonomous Driving cites this paper.

Black-Box Adversarial Attack on Vision Language Models for Autonomous Driving Dolphins: Multimodal Language Model for Driving

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T15:55:07.852163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:55:07.852163Z digest=sha256:9ca79db1deb389af2c90f8572f6b68115687d7ea15c73817796fae71b004e9bd

Observation 68622cfc-ae15-4c87-9263-e6bd30a6981e · inbound

Transferable Adversarial Attacks on Black-Box Vision-Language Models cites this paper.

Transferable Adversarial Attacks on Black-Box Vision-Language Models Dolphins: Multimodal Language Model for Driving

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T04:34:34.313636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:34:34.313636Z digest=sha256:eae07968930374d7ab5e478df47c1f1bf7cc209b8063eb0b58aec0681becc3e8

Observation 96dabfef-6f05-4c3e-9f54-f162ed66ddf5 · inbound

Natural Reflection Backdoor Attack on Vision Language Model for Autonomous Driving cites this paper.

Natural Reflection Backdoor Attack on Vision Language Model for Autonomous Driving Dolphins: Multimodal Language Model for Driving

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T22:46:49.589275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:46:49.589275Z digest=sha256:f3247284ee37f1de7d2f8ba77f369f657b8e2090c94e63fc9822fac2cd052873

Observation a5409517-a9f0-48ae-a5f2-665bc24bc018 · inbound

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation cites this paper.

S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation Dolphins: Multimodal Language Model for Driving

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:43.273602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:43.273602Z digest=sha256:226075451ec25081af14f6df730c1b3564d92c6167e5e65763b518c345fba227

Observation 16892d25-c216-4f37-903a-92aefed9ae5b · inbound

AD^2-Bench: A Hierarchical CoT Benchmark for MLLM in Autonomous Driving under Adverse Conditions cites this paper.

AD^2-Bench: A Hierarchical CoT Benchmark for MLLM in Autonomous Driving under Adverse Conditions Dolphins: Multimodal Language Model for Driving

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T04:49:28.374098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:49:28.374098Z digest=sha256:9a4f1051b0cc96981facb632f4a4b47364730d681b631e277408b840aa1ba842

Observation 6dd443bf-64e6-47c3-be1f-8389e6fb831e · inbound

UrbanLLaVA: A Multi-modal Large Language Model for Urban Intelligence with Spatial Reasoning and Understanding cites this paper.

UrbanLLaVA: A Multi-modal Large Language Model for Urban Intelligence with Spatial Reasoning and Understanding Dolphins: Multimodal Language Model for Driving

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:23.026538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:23.026538Z digest=sha256:52f335d392e79b7ae9f375ea0a5013bfbcf0948081b5940606efb2ea06a8a153

Observation 1baa118f-46e3-49ca-8045-ca67ee201dbf · inbound

CosmosAlign: Adapting a World Foundation Model for Generative Traffic Video Forecasting cites this paper.

CosmosAlign: Adapting a World Foundation Model for Generative Traffic Video Forecasting Dolphins: Multimodal Language Model for Driving

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T00:28:52.246338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:28:52.246338Z digest=sha256:ecea367b2664e9c9664e92844f2a84a318ab6230d3d11356b557af4e38fc253d