Pith. sign in

Paper Citation Record · LEDGER

Towards a Multi-Agent Vision-Language System for Zero-Shot Novel Hazardous Object Detection for Autonomous Driving Safety

As of 17 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 3 inbound Pith citation observations for arXiv:2504.13399.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.13399 v1

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:13:55.176948Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:42:41.687023Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-13T17:33:02.317711Z

Reference resolution

23 of 23 outbound references displayed

  • verified exact0
  • verified fuzzy10
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 910ba7a2-a3f7-47f7-8858-99a86654c316 · outbound

This paper cites A parametric top- view representation of complex road scenes,.

Towards a Multi-Agent Vision-Language System for Zero-Shot Novel Hazardous Object Detection for Autonomous Driving Safety A parametric top- view representation of complex road scenes,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:13:55.529043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T12:13:55.067398Z digest=sha256:2a74612339498c74dbc9475d3f6e3fae8d9208f8533ca934e1460186723185f1

Observation daa3c045-5e9d-4c0b-8295-668d9d6771db · outbound

This paper cites Understanding road layout from videos as a whole,.

Towards a Multi-Agent Vision-Language System for Zero-Shot Novel Hazardous Object Detection for Autonomous Driving Safety Understanding road layout from videos as a whole,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:13:55.514349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T12:13:55.072935Z digest=sha256:14a141887e1860445dc2e3278ff2bb2759c4bae7b2fd503cb55ec8807d719833

Observation a0eb17f4-d0c4-4b85-913e-97f5f9e60350 · outbound

This paper cites Aide: An automatic data engine for object detection in autonomous driving,.

Towards a Multi-Agent Vision-Language System for Zero-Shot Novel Hazardous Object Detection for Autonomous Driving Safety Aide: An automatic data engine for object detection in autonomous driving,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:13:55.498155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T12:13:55.078045Z digest=sha256:5593fffe2deb0212a4d4d5e276ff5195430d773c0d1de0c990cf484a28f1febc

Observation c8275e5d-fb50-4c95-9014-3a576b1fa7c5 · outbound

This paper cites Vila: On pre-training for visual language models,.

Towards a Multi-Agent Vision-Language System for Zero-Shot Novel Hazardous Object Detection for Autonomous Driving Safety Vila: On pre-training for visual language models,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T12:13:55.082978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:13:55.082978Z digest=sha256:9cd2852293e40dc18479b861529c1f454463960694eea8db8f8dcce14dc4f3ed

Observation ebfc814a-2b63-4886-b3b3-af9e57caf3f8 · outbound

This paper cites Vila: On pre-training for visual language models,.

Towards a Multi-Agent Vision-Language System for Zero-Shot Novel Hazardous Object Detection for Autonomous Driving Safety Vila: On pre-training for visual language models,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T12:13:55.087924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:13:55.087924Z digest=sha256:73027e2ebbf45515913a7990164ff4d90837931d6a56b4ce804a69bd492804ce

Observation 5b8bfaac-31c1-45de-bc8f-6812f8096ec4 · outbound

This paper cites NVILA: Efficient Frontier Visual Language Models.

Towards a Multi-Agent Vision-Language System for Zero-Shot Novel Hazardous Object Detection for Autonomous Driving Safety NVILA: Efficient Frontier Visual Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T12:13:55.094551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:13:55.094551Z digest=sha256:657ec7fc4c2d16f46776bb9de530b97238a73dcc4798212f143a2e4b683b0b1b

Observation 55649044-730d-41c5-bbe2-e7f7f061fbba · outbound

This paper cites OmniVLM: A Token-Compressed, Sub-Billion-Parameter Vision-Language Model for Efficient On-Device Inference.

Towards a Multi-Agent Vision-Language System for Zero-Shot Novel Hazardous Object Detection for Autonomous Driving Safety OmniVLM: A Token-Compressed, Sub-Billion-Parameter Vision-Language Model for Efficient On-Device Inference

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T12:13:55.100419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:13:55.100419Z digest=sha256:e10aa9c99c73fb91c540097c8c3222b499ea0702b5e81d002927e020fc96cd2e

Observation 164811e9-a918-4582-be2c-b6e7deed9304 · outbound

This paper cites COOOL: Challenge Of Out-Of-Label A Novel Benchmark for Autonomous Driving.

Towards a Multi-Agent Vision-Language System for Zero-Shot Novel Hazardous Object Detection for Autonomous Driving Safety COOOL: Challenge Of Out-Of-Label A Novel Benchmark for Autonomous Driving

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T12:13:55.106250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:13:55.106250Z digest=sha256:0839e480361a8576def2a07364a521bded0e668884c64a0c7e90bcbe703f0650

Observation 19fd4cde-bdea-4854-8a4a-fb246583ffce · outbound

This paper cites Simple baselines for image restoration,.

Towards a Multi-Agent Vision-Language System for Zero-Shot Novel Hazardous Object Detection for Autonomous Driving Safety Simple baselines for image restoration,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T12:13:55.111572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:13:55.111572Z digest=sha256:8b93eeb6f3b2f3c26fcdfef3df33bb05b3b211a2dcd3f5092f55972ea06ca557

Observation 7f553563-90b8-464a-86c1-8c2fb5468a55 · outbound

This paper cites The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision).

Towards a Multi-Agent Vision-Language System for Zero-Shot Novel Hazardous Object Detection for Autonomous Driving Safety The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision)

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T12:13:55.116277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:13:55.116277Z digest=sha256:835e65b6526adf7993f4738e872d55cf5213840fdd4d5d8266a42885a70b7dd6

Observation 679a1826-2836-44af-972d-0cc6c1f109f2 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Towards a Multi-Agent Vision-Language System for Zero-Shot Novel Hazardous Object Detection for Autonomous Driving Safety Learning transferable visual models from natural language supervision,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T12:13:55.121358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:13:55.121358Z digest=sha256:b66622920b091a87c979d3629cd4e2c7090ae26e6eafdef9ad0ffbbb539ca85e

Observation 63cc6115-07b0-48a9-ba26-3eea1851aa94 · outbound

This paper cites nuscenes: A multimodal dataset for autonomous driving,.

Towards a Multi-Agent Vision-Language System for Zero-Shot Novel Hazardous Object Detection for Autonomous Driving Safety nuscenes: A multimodal dataset for autonomous driving,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T12:13:55.125840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:13:55.125840Z digest=sha256:f29f25a18df96c7ac86def539360042b3ebac411b739be45e18e4394f9fcefec

Observation 61577742-bfe3-4626-8486-a5da7b5fe439 · outbound

This paper cites The cityscapes dataset for semantic urban scene understanding,.

Towards a Multi-Agent Vision-Language System for Zero-Shot Novel Hazardous Object Detection for Autonomous Driving Safety The cityscapes dataset for semantic urban scene understanding,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T12:13:55.130554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:13:55.130554Z digest=sha256:a0d060a7848432584aa8053961fd5f4786abea78b81e19f8bde387f16234ead2

Observation 46e35478-0971-4b51-8d86-79363269d75a · outbound

This paper cites Vision meets robotics: The kitti dataset,.

Towards a Multi-Agent Vision-Language System for Zero-Shot Novel Hazardous Object Detection for Autonomous Driving Safety Vision meets robotics: The kitti dataset,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T12:13:55.135258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:13:55.135258Z digest=sha256:4596d7a9f8326b57577a6e1deefb705406e5ff997cc5205ce0598fd610bddef4

Observation 0d8902a6-cfcb-40eb-af98-33ddbd515da7 · outbound

This paper cites Exploring the potential of multi-modal ai for driving hazard prediction,.

Towards a Multi-Agent Vision-Language System for Zero-Shot Novel Hazardous Object Detection for Autonomous Driving Safety Exploring the potential of multi-modal ai for driving hazard prediction,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T12:13:55.139599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:13:55.139599Z digest=sha256:458d4dfb0be5beb121feb0e3071131ec30a74859a37d346544e82bdb03a6d469

Observation 16407ca4-3b67-411e-9863-eab3d00c32d1 · outbound

This paper cites Driver Assistance System Based on Multimodal Data Hazard Detection.

Towards a Multi-Agent Vision-Language System for Zero-Shot Novel Hazardous Object Detection for Autonomous Driving Safety Driver Assistance System Based on Multimodal Data Hazard Detection

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T12:13:55.144129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:13:55.144129Z digest=sha256:61d9bfe46016eb8080de9ecacdf1d22b8e7e4ff5c0fdcec986417271ab9661c9

Observation 75cd17fc-71b9-451b-a0f1-b5b87425844f · outbound

This paper cites Insight: Enhancing autonomous driving safety through vision-language models on context- aware hazard detection and edge case evaluation,.

Towards a Multi-Agent Vision-Language System for Zero-Shot Novel Hazardous Object Detection for Autonomous Driving Safety Insight: Enhancing autonomous driving safety through vision-language models on context- aware hazard detection and edge case evaluation,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:13:55.392238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T12:13:55.148964Z digest=sha256:e148fcafaf979e3bbe4f801217d345db427f46aa08f5209c4f761c6b498f8192

Observation c00ff6fe-745d-42db-bc5a-bd25719862a8 · outbound

This paper cites Detecting hazardous events: A framework for automated vehicle safety systems,.

Towards a Multi-Agent Vision-Language System for Zero-Shot Novel Hazardous Object Detection for Autonomous Driving Safety Detecting hazardous events: A framework for automated vehicle safety systems,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:13:55.375549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T12:13:55.153454Z digest=sha256:013d4b133fe97fcd4b74b7d9575e1f84aa4d81041413505bb4ad6dc7d104219b

Observation 551e7acd-49c1-4628-a5b6-1cb2a04d47b7 · outbound

This paper cites Lost and found: detecting small road hazards for self-driving vehicles. in 2016 ieee,.

Towards a Multi-Agent Vision-Language System for Zero-Shot Novel Hazardous Object Detection for Autonomous Driving Safety Lost and found: detecting small road hazards for self-driving vehicles. in 2016 ieee,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:13:55.358412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T12:13:55.158224Z digest=sha256:5046448b917c69612d7f10b03369e370c17ff332b8fa3dd14b1b479ae2ff800f

Observation 7e5b8858-6c6c-45c2-94c9-83a4bd994aea · outbound

This paper cites spaCy 2: Natural language understanding with Bloom embeddings, convolutional neural networks and incremental parsing,.

Towards a Multi-Agent Vision-Language System for Zero-Shot Novel Hazardous Object Detection for Autonomous Driving Safety spaCy 2: Natural language understanding with Bloom embeddings, convolutional neural networks and incremental parsing,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:13:55.342465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T12:13:55.163312Z digest=sha256:b3bba0bdb911dff3af305894b1dbdf1cb0a9abc116dfa8aded1f476cf04e5e8d

Observation 2343b412-5755-4c51-9e5b-0cb60b1a9f34 · outbound

This paper cites Open-world hazard detection and captioning for autonomous driving with a unified multimodal pipeline,.

Towards a Multi-Agent Vision-Language System for Zero-Shot Novel Hazardous Object Detection for Autonomous Driving Safety Open-world hazard detection and captioning for autonomous driving with a unified multimodal pipeline,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:13:55.326387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T12:13:55.167942Z digest=sha256:8d53e680744de53f5f7490d19e8224a92148fc55d7b8173d777cbd90f392d318

Observation ef713320-1a11-46d8-93e6-7f419c03c5dc · outbound

This paper cites COOL-W ACV25 Competition - Leaderboard,.

Towards a Multi-Agent Vision-Language System for Zero-Shot Novel Hazardous Object Detection for Autonomous Driving Safety COOL-W ACV25 Competition - Leaderboard,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:13:55.310153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T12:13:55.172619Z digest=sha256:ef6073224ac2e4713c885347db78e8e03d43825e8386387c39ae2b72141a21a7

Observation f05e4332-adfc-415e-9091-756432f689e3 · outbound

This paper cites Addressing out-of-label hazard detection in dashcam videos: Insights from the coool challenge,.

Towards a Multi-Agent Vision-Language System for Zero-Shot Novel Hazardous Object Detection for Autonomous Driving Safety Addressing out-of-label hazard detection in dashcam videos: Insights from the coool challenge,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:13:55.293324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T12:13:55.176948Z digest=sha256:6fbe43850d9a0eac400008040a88cf2cb2d2ba538b1f5322bc6a784e77f2183e

Pith citing papers

Observation 55e5b0ed-2023-4475-8c7c-239c54ad6b01 · inbound

Beyond General Prompts: Automated Prompt Refinement using Contrastive Class Alignment Scores for Disambiguating Objects in Vision-Language Models cites this paper.

Beyond General Prompts: Automated Prompt Refinement using Contrastive Class Alignment Scores for Disambiguating Objects in Vision-Language Models Towards a Multi-Agent Vision-Language System for Zero-Shot Novel Hazardous Object Detection for Autonomous Driving Safety

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T21:42:41.687023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:42:41.687023Z digest=sha256:d86f8a4d0fb49d801cd9c42c9aeffc8bb97395b02b77a4a031ae726ec7189529

Observation 390999ee-fdfb-4292-ab6d-2bdeb7d413c4 · inbound

A Multimodal Foundation Model of Spatial Transcriptomics and Histology for Biological Discovery and Clinical Prediction cites this paper.

A Multimodal Foundation Model of Spatial Transcriptomics and Histology for Biological Discovery and Clinical Prediction Towards a Multi-Agent Vision-Language System for Zero-Shot Novel Hazardous Object Detection for Autonomous Driving Safety

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-13T12:44:39.273341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T12:44:39.273341Z digest=sha256:932a101f846b86d4be7e00b34fcf13d429f8f9e6002a5bcfe36aea26ae0681ba

Observation 57d2ec93-e7b5-4c1f-8582-32b23747eb6d · inbound

Single-agent vs. Multi-agents for Automated Video Analysis of On-Screen Collaborative Learning Behaviors cites this paper.

Single-agent vs. Multi-agents for Automated Video Analysis of On-Screen Collaborative Learning Behaviors Towards a Multi-Agent Vision-Language System for Zero-Shot Novel Hazardous Object Detection for Autonomous Driving Safety

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-13T17:33:02.319063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T17:32:06.256142Z digest=sha256:73fa49dab44e9ce3ab910182da855e452901ae8e747f34c7c4aea11a0daeff2b