Pith. sign in

Paper Citation Record · LEDGER

DriveSOTIF: Advancing Perception SOTIF Through Multimodal Large Language Models

As of 17 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 0 inbound Pith citation observations for arXiv:2505.07084.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.07084 v3

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T22:29:54.820604Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

54 of 54 outbound references displayed

  • verified exact0
  • verified fuzzy27
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 46b64b86-d5e1-43e5-94b0-7b51ac3b73d4 · outbound

This paper cites A holistic robust motion control framework for autonomous platooning,.

DriveSOTIF: Advancing Perception SOTIF Through Multimodal Large Language Models A holistic robust motion control framework for autonomous platooning,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:29:55.328863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T22:29:54.642437Z digest=sha256:07ad4fc5a13d8a2396601264795f3d0500f694ef550f9841d82be80a2830fc50

Observation 7cc2ee29-394f-40d7-b4b9-6e5e3edb318f · outbound

This paper cites A survey on an emerging safety challenge for autonomous vehicles: Safety of the intended functionality,.

DriveSOTIF: Advancing Perception SOTIF Through Multimodal Large Language Models A survey on an emerging safety challenge for autonomous vehicles: Safety of the intended functionality,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:29:55.319248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T22:29:54.646944Z digest=sha256:210d4dced7817e43950631dd201bb830813d807ce4e9f757b5a30be6002f187d

Observation d4b8a711-22db-4e15-b3e6-fd2dfe32d7ea · outbound

This paper cites Sotif entropy: Online sotif risk quantification and mitigation for autonomous driving,.

DriveSOTIF: Advancing Perception SOTIF Through Multimodal Large Language Models Sotif entropy: Online sotif risk quantification and mitigation for autonomous driving,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:29:55.309427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T22:29:54.650419Z digest=sha256:51bda3fc6485cad89020d1456b5c839d543505a134e0c12fd8a95d532acd4152

Observation 0cc07648-4951-4ff0-bb73-9781a85c92b5 · outbound

This paper cites Pesotif: A challenging visual dataset for perception sotif problems in long-tail traffic scenarios,.

DriveSOTIF: Advancing Perception SOTIF Through Multimodal Large Language Models Pesotif: A challenging visual dataset for perception sotif problems in long-tail traffic scenarios,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:29:55.299979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T22:29:54.654024Z digest=sha256:c2f9374ca49475d64deb3e5be492bd2207472e7fd54fba7898339940e565d028

Observation 96f132e6-bfe6-43e2-b5fb-baa52b4601ff · outbound

This paper cites Multi-modal answer validation for knowledge-based vqa,.

DriveSOTIF: Advancing Perception SOTIF Through Multimodal Large Language Models Multi-modal answer validation for knowledge-based vqa,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:29:55.289494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T22:29:54.657603Z digest=sha256:1c64070399f68290e9168d28c25f611677c51c4f65c3051745e4cdbc7c3e791e

Observation 5c50d009-e0c0-4f99-9bb8-7ce3852b80d5 · outbound

This paper cites Context-vqa: Towards context- aware and purposeful visual question answering,.

DriveSOTIF: Advancing Perception SOTIF Through Multimodal Large Language Models Context-vqa: Towards context- aware and purposeful visual question answering,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:29:55.279321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T22:29:54.661203Z digest=sha256:332cb8d484253279a39687116d1a923a167ea3a794f98d4cdaa86f8f2f5c1a67

Observation 73245253-0d0d-44d8-9e8c-07e3d677d042 · outbound

This paper cites CODA: A Real-World Road Corner Case Dataset for Object Detection in Autonomous Driving.

DriveSOTIF: Advancing Perception SOTIF Through Multimodal Large Language Models CODA: A Real-World Road Corner Case Dataset for Object Detection in Autonomous Driving

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T22:29:54.664755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:29:54.664755Z digest=sha256:85588f024ebbd751fd6429bba4910b5ed9cb57e44f455aa7fff872936de97323

Observation 852d2156-8d11-4cd0-8905-5566c6d68739 · outbound

This paper cites Simulation-based performance evaluation of 3d object detection methods with deep learning for a lidar point cloud dataset in a sotif-related use case,.

DriveSOTIF: Advancing Perception SOTIF Through Multimodal Large Language Models Simulation-based performance evaluation of 3d object detection methods with deep learning for a lidar point cloud dataset in a sotif-related use case,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:29:55.269004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T22:29:54.668454Z digest=sha256:e2e3d94f0e3fc454b0d29e99e462567f84b3da10d8b33e8ed18d03b876423b27

Observation f88aa260-008f-4e0b-9ee8-99087bfb7a06 · outbound

This paper cites Automated Evaluation of Large Vision-Language Models on Self-driving Corner Cases.

DriveSOTIF: Advancing Perception SOTIF Through Multimodal Large Language Models Automated Evaluation of Large Vision-Language Models on Self-driving Corner Cases

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T22:29:54.671697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:29:54.671697Z digest=sha256:dbd3f388b25ddb12e2d310370be37c9b8199ca9e448b40d9f21ed7cd8e8a5272

Observation a924ca6b-d498-4d65-938a-59aa82a2395d · outbound

This paper cites Enhancing autonomous vehicle safety based on operational design domain defi- nition, monitoring, and functional degradation: A case study on lane keeping system,.

DriveSOTIF: Advancing Perception SOTIF Through Multimodal Large Language Models Enhancing autonomous vehicle safety based on operational design domain defi- nition, monitoring, and functional degradation: A case study on lane keeping system,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:29:55.260103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T22:29:54.675246Z digest=sha256:9ea75b4e7e6d8543d3636a7ee57722c55416f424cb783e1408e579132c14c8ef

Observation e3b557f6-f83f-43f0-9521-eb9712307c92 · outbound

This paper cites Sotif-oriented percep- tion evaluation method for forward obstacle detection of autonomous vehicles,.

DriveSOTIF: Advancing Perception SOTIF Through Multimodal Large Language Models Sotif-oriented percep- tion evaluation method for forward obstacle detection of autonomous vehicles,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:29:55.250743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T22:29:54.678423Z digest=sha256:008c702dc9a89e00f9ee7eb5e9aa7895b3d42570ba10384a355a528023668ccd

Observation 05235559-8731-4a8c-9456-f672812c62ca · outbound

This paper cites The sotif meta-algorithm: Quantitative analyses of the safety of autonomous behaviors,.

DriveSOTIF: Advancing Perception SOTIF Through Multimodal Large Language Models The sotif meta-algorithm: Quantitative analyses of the safety of autonomous behaviors,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:29:55.241685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T22:29:54.682201Z digest=sha256:c3cabafa6939eb7f1b5c97eeed9015b28a96e240575c33a8c650e864b6d08a97

Observation 00407ee8-a828-44bf-9e15-14a62d08664f · outbound

This paper cites A hazard analysis approach for the sotif in intelligent railway driving assistance systems using stpa and complex network,.

DriveSOTIF: Advancing Perception SOTIF Through Multimodal Large Language Models A hazard analysis approach for the sotif in intelligent railway driving assistance systems using stpa and complex network,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:29:55.231610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T22:29:54.685310Z digest=sha256:9880b3b409a8f384c22ecf73208297b86538067efdc1a743c47ac3a4aeb16f47

Observation 19804e08-fff1-4471-8ace-f86518bc4018 · outbound

This paper cites Safety of the intended functionality (sotif) based on system theoretic process analysis (stpa): Study for specific control action in blind spot detection (bsd),.

DriveSOTIF: Advancing Perception SOTIF Through Multimodal Large Language Models Safety of the intended functionality (sotif) based on system theoretic process analysis (stpa): Study for specific control action in blind spot detection (bsd),

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:29:55.222175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T22:29:54.688551Z digest=sha256:36558c2aea3b8fc4b3e65814145df58b2954ceaa68f30fbcc03516861655d3c4

Observation 71aab8f8-35b1-499f-95fa-c5f944899136 · outbound

This paper cites Formal cer- tification methods for automated vehicle safety assessment,.

DriveSOTIF: Advancing Perception SOTIF Through Multimodal Large Language Models Formal cer- tification methods for automated vehicle safety assessment,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:29:55.212108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T22:29:54.691865Z digest=sha256:7387e278fec1133f9500b5df3f87ad17eb355ea6160aa48ca4a203f4a979abc7

Observation 8a562cba-ab94-49f0-9a37-e6317d4be6a6 · outbound

This paper cites Systematic modeling ap- proach for environmental perception limitations in automated driving,.

DriveSOTIF: Advancing Perception SOTIF Through Multimodal Large Language Models Systematic modeling ap- proach for environmental perception limitations in automated driving,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:29:55.201684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T22:29:54.695079Z digest=sha256:f305144175ba31f14851a4d81d45f3e1431ab922157733b5bbf87f91019166af

Observation c8a783b0-c624-4284-a950-697bfb4d9e07 · outbound

This paper cites Decomposition and quan- tification of sotif requirements for perception systems of autonomous vehicles,.

DriveSOTIF: Advancing Perception SOTIF Through Multimodal Large Language Models Decomposition and quan- tification of sotif requirements for perception systems of autonomous vehicles,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:29:55.192369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T22:29:54.698250Z digest=sha256:1b8d45a01cdd12bd9b1851a238a3c9dd5020f4e7c73cccac37e0b2cb16b8735d

Observation 76c3cb82-d18d-42af-a172-2380f7f728fd · outbound

This paper cites Online quantitative analysis of perception uncertainty based on high- definition map,.

DriveSOTIF: Advancing Perception SOTIF Through Multimodal Large Language Models Online quantitative analysis of perception uncertainty based on high- definition map,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:29:55.182003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T22:29:54.701499Z digest=sha256:939e9d4ae4280fef7f36004b2119a46ee0ee137240120e21c2ac890df4fa162a

Observation 4219b357-3f28-4387-a7a4-672ddf0ed6c1 · outbound

This paper cites Towards causal vqa: Revealing and reducing spurious correlations by invariant and covariant semantic editing,.

DriveSOTIF: Advancing Perception SOTIF Through Multimodal Large Language Models Towards causal vqa: Revealing and reducing spurious correlations by invariant and covariant semantic editing,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:29:55.171141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T22:29:54.704597Z digest=sha256:af968677ea0a0adeaef089b609a541e4afd3ad5956a53893a431a3ad94ee7ed3

Observation 3c0229bb-9ef0-484d-a319-f94e6d8c4689 · outbound

This paper cites Drivellm: Charting the path toward full autonomous driving with large language models,.

DriveSOTIF: Advancing Perception SOTIF Through Multimodal Large Language Models Drivellm: Charting the path toward full autonomous driving with large language models,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:29:55.160212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T22:29:54.708051Z digest=sha256:ea9a686919785a65225c6921096608e06d72e661441d39bd26586af64aa69acc

Observation c9f1ba13-06d1-4dc6-832e-89847479b820 · outbound

This paper cites Driving with llms: Fusing object- level vector modality for explainable autonomous driving,.

DriveSOTIF: Advancing Perception SOTIF Through Multimodal Large Language Models Driving with llms: Fusing object- level vector modality for explainable autonomous driving,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T22:29:54.711204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:29:54.711204Z digest=sha256:66ba0aee5e3c3509a90900ee8ea5af0b30053f50b83fde94ce3e8e8b925d36a4

Observation a0e3d80c-b9ed-4c79-90c7-fc6be35d14cc · outbound

This paper cites VLM-MPC: Vision Language Foundation Model (VLM)-Guided Model Predictive Controller (MPC) for Autonomous Driving.

DriveSOTIF: Advancing Perception SOTIF Through Multimodal Large Language Models VLM-MPC: Vision Language Foundation Model (VLM)-Guided Model Predictive Controller (MPC) for Autonomous Driving

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T22:29:54.714301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:29:54.714301Z digest=sha256:55fb0f5316c00e3890cad1c66ca74e07d2078273bbf5e9e33d6e88a44a578c1f

Observation 63bca8eb-f201-4ec8-8110-aea7246f8399 · outbound

This paper cites Scene understanding for autonomous driving using visual question answering,.

DriveSOTIF: Advancing Perception SOTIF Through Multimodal Large Language Models Scene understanding for autonomous driving using visual question answering,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:29:55.144045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T22:29:54.717707Z digest=sha256:c364dc808c2d4d2295488fef15f0ad34fba61819623df2f9c2e4961a1a6ab004

Observation ee408494-4133-4f7c-8f51-c8ece088a1d1 · outbound

This paper cites Nuscenes-qa: A multi-modal visual question answering benchmark for autonomous driving scenario,.

DriveSOTIF: Advancing Perception SOTIF Through Multimodal Large Language Models Nuscenes-qa: A multi-modal visual question answering benchmark for autonomous driving scenario,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T22:29:54.720784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:29:54.720784Z digest=sha256:e5590229e44bc8263833eb7c69b5fe063961787c63d5373764ca928020351b4f

Observation 77deb080-5531-476b-a481-9c0a3fd65c76 · outbound

This paper cites Explaining autonomous driving actions with visual question answering,.

DriveSOTIF: Advancing Perception SOTIF Through Multimodal Large Language Models Explaining autonomous driving actions with visual question answering,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:29:55.126937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T22:29:54.724099Z digest=sha256:ed1ac8f7e061aeb3f761f75cd6ff2213a30873a6cd98f67f1981d439b18285cd

Observation 82ff2c98-48a3-4da4-935d-06f6c562cf89 · outbound

This paper cites GPT-4V(ision) is a Generalist Web Agent, if Grounded.

DriveSOTIF: Advancing Perception SOTIF Through Multimodal Large Language Models GPT-4V(ision) is a Generalist Web Agent, if Grounded

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T22:29:54.727447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:29:54.727447Z digest=sha256:cca17da322fec11a2d2de42bb8242f4a481ecbf99c63a3893baf5c89f6bed0d5

Observation e4ea5224-34b7-485b-bf84-20fb8a0f16e0 · outbound

This paper cites Multi-Frame, Lightweight & Efficient Vision-Language Models for Question Answering in Autonomous Driving.

DriveSOTIF: Advancing Perception SOTIF Through Multimodal Large Language Models Multi-Frame, Lightweight & Efficient Vision-Language Models for Question Answering in Autonomous Driving

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T22:29:54.731003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:29:54.731003Z digest=sha256:4462f59b4590968839f123e5da4687c7a66e13d4ca7ca67fc9cd134c364a4aac

Observation 451be2cc-3810-430e-a1df-cd1b4f64d60a · outbound

This paper cites Drama: Joint risk localization and captioning in driving,.

DriveSOTIF: Advancing Perception SOTIF Through Multimodal Large Language Models Drama: Joint risk localization and captioning in driving,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:29:55.115916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T22:29:54.734596Z digest=sha256:436f1866c776583b7663a4186c66dc56f50dcf387ed8a7b6ce3f31c069e64a3d

Observation 3723ccbd-3c21-457a-8421-c908c6fae8d0 · outbound

This paper cites Referring multi-object tracking,.

DriveSOTIF: Advancing Perception SOTIF Through Multimodal Large Language Models Referring multi-object tracking,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:29:55.106127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T22:29:54.737923Z digest=sha256:06c6f799dc4e382e7036540ade435767125a21d5e1f81d71d856bad2a974a36e

Observation 2088789f-d601-4812-857e-17ad26a5502d · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

DriveSOTIF: Advancing Perception SOTIF Through Multimodal Large Language Models Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T22:29:54.741051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:29:54.741051Z digest=sha256:8cd8f36f19fdf1aa1f4a27180091ef647dd71fb1e6286e7268febfe3f955943b

Observation e580f922-9d48-4390-81ef-de0b8f662e33 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answering,.

DriveSOTIF: Advancing Perception SOTIF Through Multimodal Large Language Models Making the v in vqa matter: Elevating the role of image understanding in visual question answering,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:29:55.096080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T22:29:54.744672Z digest=sha256:de50fe65359a7ca1746ea65b4d3bf5fb3eb9b87180f6e51135a0ac858bbd07f2

Observation 00262dd7-76b9-4726-831d-df780c21d896 · outbound

This paper cites Cider: Consensus- based image description evaluation,.

DriveSOTIF: Advancing Perception SOTIF Through Multimodal Large Language Models Cider: Consensus- based image description evaluation,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T22:29:54.747867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:29:54.747867Z digest=sha256:edcbad583a1c141e9fbb92197bb9940f9f509ca8e11c7b86497e9131e056bdb2

Observation e25fa1a8-aa61-42b0-a797-2018c82bbc48 · outbound

This paper cites Spice: Semantic propositional image caption evaluation,.

DriveSOTIF: Advancing Perception SOTIF Through Multimodal Large Language Models Spice: Semantic propositional image caption evaluation,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T22:29:54.751096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:29:54.751096Z digest=sha256:91a58fe0cecb2182f15dad6642c722f5a225b2dccab75c7de4df90f5d9c2b400

Observation 01237a3f-619e-40df-8396-a0b309e4bfc6 · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena,.

DriveSOTIF: Advancing Perception SOTIF Through Multimodal Large Language Models Judging llm-as-a-judge with mt-bench and chatbot arena,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T22:29:54.754213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:29:54.754213Z digest=sha256:4dccb62a08a0c2fcd28b3ba653004fd417eb437efb90d22cd26c03b154d10cc5

Observation 9e3e05d4-7447-44bb-84dd-44b438aad7c3 · outbound

This paper cites Judge Anything: MLLM as a Judge Across Any Modality.

DriveSOTIF: Advancing Perception SOTIF Through Multimodal Large Language Models Judge Anything: MLLM as a Judge Across Any Modality

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T22:29:54.757189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:29:54.757189Z digest=sha256:b7fe0c259a8998303920e4254b7513f07ae3a6ce410135b67ca4473f797c12f1

Observation 1792280e-73aa-4912-a9af-55f00774f848 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,.

DriveSOTIF: Advancing Perception SOTIF Through Multimodal Large Language Models Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T22:29:54.760786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:29:54.760786Z digest=sha256:bce2efceddeefcdf02c90089af56c2d5820aebf3c01d7d3e1cbac472e832579d

Observation f63bb5fa-8cf4-49b2-a13a-9885ec1cc7c7 · outbound

This paper cites Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,.

DriveSOTIF: Advancing Perception SOTIF Through Multimodal Large Language Models Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T22:29:54.763942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:29:54.763942Z digest=sha256:14160ee3c24baa2d9ab87b1cd43100963d4c3155558868b6647cc5b4a8629217

Observation 79343155-fe7a-4cdb-8231-6f28bf60a842 · outbound

This paper cites Visual instruction tuning,.

DriveSOTIF: Advancing Perception SOTIF Through Multimodal Large Language Models Visual instruction tuning,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T22:29:54.767003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:29:54.767003Z digest=sha256:c0b5331aa29de3aa840bb33aa8221bef91581a603db0c8d5914310078a7d281f

Observation 49174d0a-cfd0-465f-b4bd-94e858efdc42 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

DriveSOTIF: Advancing Perception SOTIF Through Multimodal Large Language Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T22:29:54.769693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:29:54.769693Z digest=sha256:25ae7a3cfdf960facbc9b0010153173290bc5b6746ee9e5eb87655490115f86a

Observation 0bc66d7f-f9e9-4a12-920b-7d3bc587a902 · outbound

This paper cites Qwen2.5-VL Technical Report.

DriveSOTIF: Advancing Perception SOTIF Through Multimodal Large Language Models Qwen2.5-VL Technical Report

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T22:29:54.772797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:29:54.772797Z digest=sha256:614ad87defbc06ec292e2db56afb38426166f98dbfda14f9dbc6027d4615f4ce

Observation 01893d7a-a86d-4c02-bd13-f0c42186a738 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

DriveSOTIF: Advancing Perception SOTIF Through Multimodal Large Language Models InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T22:29:54.776155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:29:54.776155Z digest=sha256:52456ba3c2c80b741e23de2c954ffa4c27d111b653715baa8dcf3282f60464b3

Observation 06ef0d92-e2ae-44f7-b28e-7cfa5e9d2e12 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

DriveSOTIF: Advancing Perception SOTIF Through Multimodal Large Language Models LoRA: Low-Rank Adaptation of Large Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T22:29:54.779436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:29:54.779436Z digest=sha256:5a2cf0fa2fd697105649178d6666891a118d602be4f159bc2ba8829f2f7cc77b

Observation c7397c6c-1fb6-4475-a94f-57ad15c34ca9 · outbound

This paper cites LA VIS: A one-stop library for language-vision intelligence,.

DriveSOTIF: Advancing Perception SOTIF Through Multimodal Large Language Models LA VIS: A one-stop library for language-vision intelligence,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:29:55.054754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T22:29:54.782476Z digest=sha256:2e7356e4060f232f4003e344e4997e8c0efee67f6a26900932f3d3e3df694c5d

Observation 12b0af8c-3221-4f41-867e-6e4052e0bfa6 · outbound

This paper cites LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models.

DriveSOTIF: Advancing Perception SOTIF Through Multimodal Large Language Models LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T22:29:54.785776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:29:54.785776Z digest=sha256:03037a04764ab993960b04cc66c60807aff4702da246e102b1feb3ffe4ad1645

Observation 7ef106db-afe3-4eca-b312-f51ccc38a32a · outbound

This paper cites TensorRT-LLM,.

DriveSOTIF: Advancing Perception SOTIF Through Multimodal Large Language Models TensorRT-LLM,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:29:55.044107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T22:29:54.789211Z digest=sha256:94d0f15e97abcb3bf21864d55641f6eb74bfaa4007589704e6d482cc41e0a446

Observation b4220204-30f7-4b63-b043-700143b1478a · outbound

This paper cites Efficient memory management for large language model serving with pagedattention,.

DriveSOTIF: Advancing Perception SOTIF Through Multimodal Large Language Models Efficient memory management for large language model serving with pagedattention,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T22:29:54.792756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:29:54.792756Z digest=sha256:9010693384dfa68149519f105e607d9a60bf6b41a99f41297d9f87bbf510d167

Observation bbd5893c-327a-4e9a-9797-d751232f4e45 · outbound

This paper cites Towards Human-Centric Autonomous Driving: A Fast-Slow Architecture Integrating Large Language Model Guidance with Reinforcement Learning.

DriveSOTIF: Advancing Perception SOTIF Through Multimodal Large Language Models Towards Human-Centric Autonomous Driving: A Fast-Slow Architecture Integrating Large Language Model Guidance with Reinforcement Learning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T22:29:54.795961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:29:54.795961Z digest=sha256:0d8b2097c900673093968f3ec9075198947ffca3797a13923c22aa671c791335

Observation 430ff2b3-3d49-4c12-9e90-9f78dc6224ac · outbound

This paper cites VLMPlanner: Integrating Visual Language Models with Motion Planning.

DriveSOTIF: Advancing Perception SOTIF Through Multimodal Large Language Models VLMPlanner: Integrating Visual Language Models with Motion Planning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T22:29:54.799574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:29:54.799574Z digest=sha256:7529933a846d795b31c00349cb866001451e688ae29d223b2783900ef77f1066

Observation 0fd600ec-0e51-427f-a690-266130e510a2 · outbound

This paper cites Canadian adverse driving conditions dataset,.

DriveSOTIF: Advancing Perception SOTIF Through Multimodal Large Language Models Canadian adverse driving conditions dataset,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:29:55.028217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T22:29:54.803567Z digest=sha256:aa9602749a776cf1b92b30cd7d9a4b956b71882eb2d30a7c39e4bfa18d3c95b7

Observation 9822037e-2cb3-41c5-bdf7-b701dac82b8e · outbound

This paper cites Safe RLHF: Safe Reinforcement Learning from Human Feedback.

DriveSOTIF: Advancing Perception SOTIF Through Multimodal Large Language Models Safe RLHF: Safe Reinforcement Learning from Human Feedback

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T22:29:54.807314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:29:54.807314Z digest=sha256:c9f639b6b21cee87e0f707ed8d311adea5050c42aba0a1c68eb3985849e366c9

Observation 47041eae-e0ac-474b-a444-8aa83f236321 · outbound

This paper cites Siren's Song in the AI Ocean: A Survey on Hallucination in Large Language Models.

DriveSOTIF: Advancing Perception SOTIF Through Multimodal Large Language Models Siren's Song in the AI Ocean: A Survey on Hallucination in Large Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T22:29:54.810696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:29:54.810696Z digest=sha256:5ea3c02995b20dfdf2fd57246fa938a32d0dd87209d3966fe2ccfc087561436c

Observation 4f88fff4-c8b0-4fb3-bf11-6c00db24f296 · outbound

This paper cites Retrieval- augmented generation for knowledge-intensive nlp tasks,.

DriveSOTIF: Advancing Perception SOTIF Through Multimodal Large Language Models Retrieval- augmented generation for knowledge-intensive nlp tasks,

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T22:29:54.814475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:29:54.814475Z digest=sha256:a6010dcd158c49e3cb28b69ac9f85d4c2a834dbb84ef1a070011a598264dd77d

Observation 0bca7d1b-2853-42e1-9ace-42286c56d3d6 · outbound

This paper cites Rlaif vs. rlhf: Scaling reinforcement learning from human feedback with ai feedback,.

DriveSOTIF: Advancing Perception SOTIF Through Multimodal Large Language Models Rlaif vs. rlhf: Scaling reinforcement learning from human feedback with ai feedback,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:29:55.013397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T22:29:54.817524Z digest=sha256:8014ee1f28c6c09397bc2188855ca108a0c0cc245307992e3a480b96d95a6340

Observation 55fb241e-8f23-4ed5-86ef-d71ce61e03de · outbound

This paper cites Reducing Hallucinations in Vision-Language Models via Latent Space Steering.

DriveSOTIF: Advancing Perception SOTIF Through Multimodal Large Language Models Reducing Hallucinations in Vision-Language Models via Latent Space Steering

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T22:29:54.820604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:29:54.820604Z digest=sha256:981035467aac7815fa0c3d854dda337ef79106de188feb1059c0b46ffb5d4bd9

Pith citing papers

No inbound Pith citation observations are available.