Pith. sign in

Paper Citation Record · LEDGER

HMVLM: Multistage Reasoning-Enhanced Vision-Language Model for Long-Tailed Driving Scenarios

As of 8 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 6 inbound Pith citation observations for arXiv:2506.05883.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.05883 v1

Coverage vector

measured 14 of 14 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:17:37.155533Z

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T21:12:28.565090Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-28T23:02:46.934447Z

Reference resolution

14 of 14 outbound references displayed

  • verified exact0
  • verified fuzzy5
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7aaf111c-0015-400e-95b6-195e1789067d · outbound

This paper cites Qwen2.5-VL Technical Report.

HMVLM: Multistage Reasoning-Enhanced Vision-Language Model for Long-Tailed Driving Scenarios Qwen2.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:35.974514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:35.974514Z digest=sha256:13faae58d596c1b17c1948a8e990e64fcc40b4617ec9bf70ec4fefdcbc1e0764

Observation 234b758d-13b7-4279-bb4f-69ab2e2b4c1b · outbound

This paper cites Vlmevalkit: An open-source toolkit for evaluating large multi-modality models.

HMVLM: Multistage Reasoning-Enhanced Vision-Language Model for Long-Tailed Driving Scenarios Vlmevalkit: An open-source toolkit for evaluating large multi-modality models

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:17:38.126413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:17:36.029742Z digest=sha256:b2d36ce3430ed05dedf0171920c809e5afefd380b7ad2e55ebf4333ae97d8f01

Observation 498bbcf9-55e1-4e33-a5b3-86479a254546 · outbound

This paper cites EMMA: End-to-End Multimodal Model for Autonomous Driving.

HMVLM: Multistage Reasoning-Enhanced Vision-Language Model for Long-Tailed Driving Scenarios EMMA: End-to-End Multimodal Model for Autonomous Driving

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:36.133717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:36.133717Z digest=sha256:ade0a9dae1a4ff24b780eb90c23cc1fc62d9ed7c769af11debeab5963a59e477

Observation 4348ceac-57e5-4328-9fa5-966bd156d144 · outbound

This paper cites Vad: Vectorized scene representa- tion for efficient autonomous driving.

HMVLM: Multistage Reasoning-Enhanced Vision-Language Model for Long-Tailed Driving Scenarios Vad: Vectorized scene representa- tion for efficient autonomous driving

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:36.227403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:36.227403Z digest=sha256:8a586febdcf28a1c130e991601aa41a1d7c03fe991c36d32bac78284fdb3d974

Observation 19b53aff-7b85-46c1-a293-1a2b827d9b71 · outbound

This paper cites Senna: Bridging Large Vision-Language Models and End-to-End Autonomous Driving.

HMVLM: Multistage Reasoning-Enhanced Vision-Language Model for Long-Tailed Driving Scenarios Senna: Bridging Large Vision-Language Models and End-to-End Autonomous Driving

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:36.360513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:36.360513Z digest=sha256:83fc99955340cac516c598cd74d15666ec98f77770f4575b8cd263bbfcc3824e

Observation a91321a4-5de0-4db9-a954-38ba7e17d20d · outbound

This paper cites AlphaDrive: Unleashing the Power of VLMs in Autonomous Driving via Reinforcement Learning and Reasoning.

HMVLM: Multistage Reasoning-Enhanced Vision-Language Model for Long-Tailed Driving Scenarios AlphaDrive: Unleashing the Power of VLMs in Autonomous Driving via Reinforcement Learning and Reasoning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:36.438890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:36.438890Z digest=sha256:b9dc465ae7d70619c6b2e22c797e63238a3b1d03044fee7a7b4635ba14e2436b

Observation 2304d855-f9e5-40ac-b922-144a0fd49286 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention.

HMVLM: Multistage Reasoning-Enhanced Vision-Language Model for Long-Tailed Driving Scenarios Efficient memory management for large language model serving with pagedattention

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:17:37.929342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:17:36.520854Z digest=sha256:dfe795a834991a902f7ddd8f02e73c5264620ff3dead3fe56e1485c71c5c30ee

Observation c22b8cf3-ddea-49d7-9c4f-f521de0ff169 · outbound

This paper cites Devel- oping and validating an adaptive multi-layer vehicle trajec- tory reconstruction method for outlier removal.

HMVLM: Multistage Reasoning-Enhanced Vision-Language Model for Long-Tailed Driving Scenarios Devel- oping and validating an adaptive multi-layer vehicle trajec- tory reconstruction method for outlier removal

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:17:37.766009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:17:36.597305Z digest=sha256:97aaf159c4ed5fcf569c35f9adc985d3ba9e5076cc93622adba174aab2ef5443

Observation 01b998b6-db5c-46e2-ba5c-f471d07c37a1 · outbound

This paper cites A Survey on Hallucination in Large Vision-Language Models.

HMVLM: Multistage Reasoning-Enhanced Vision-Language Model for Long-Tailed Driving Scenarios A Survey on Hallucination in Large Vision-Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:36.697142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:36.697142Z digest=sha256:df57c7fa915d616a4c01fada6e62881f649eef1144d9e516c6c5647c87f0e749

Observation e7f115ff-35e1-4d31-8147-1e959cd0d3c4 · outbound

This paper cites Deepspeed: System optimizations enable train- ing deep learning models with over 100 billion parameters.

HMVLM: Multistage Reasoning-Enhanced Vision-Language Model for Long-Tailed Driving Scenarios Deepspeed: System optimizations enable train- ing deep learning models with over 100 billion parameters

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:17:37.581196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:17:36.774176Z digest=sha256:7bd005a8cf835b71f0047983365ca45cc309d15e5ee1935fa33f3073dc662823

Observation e3817ae6-5dd3-46b2-b250-d6665932e4ec · outbound

This paper cites What is a savitzky-golay filter?[lecture notes].

HMVLM: Multistage Reasoning-Enhanced Vision-Language Model for Long-Tailed Driving Scenarios What is a savitzky-golay filter?[lecture notes]

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:17:37.383036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:17:36.857015Z digest=sha256:dbd0165571d5f3248db0261418e1c25f4f9536e86feb107e691c95482789364a

Observation 52cd9536-bc1e-4d43-9145-a5a67e8ea63b · outbound

This paper cites DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models.

HMVLM: Multistage Reasoning-Enhanced Vision-Language Model for Long-Tailed Driving Scenarios DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:36.953338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:36.953338Z digest=sha256:ce07a52104a988b9d73c728d231ce3025aabd81845f657841d199191531cbff2

Observation c65c632d-b867-4fd9-908f-a239a630fecd · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large lan- guage models.

HMVLM: Multistage Reasoning-Enhanced Vision-Language Model for Long-Tailed Driving Scenarios Chain-of-thought prompting elicits reasoning in large lan- guage models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:37.060777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:37.060777Z digest=sha256:81480075fa1a69a49a9f6508891d8761af1ea0d00f1b4ec83837854f8833df93

Observation 57b0c53d-3b16-49f2-b38a-c0d8564911ac · outbound

This paper cites LLaVA-CoT: Let Vision Language Models Reason Step-by-Step.

HMVLM: Multistage Reasoning-Enhanced Vision-Language Model for Long-Tailed Driving Scenarios LLaVA-CoT: Let Vision Language Models Reason Step-by-Step

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:37.155533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:37.155533Z digest=sha256:52a127e9c39940da3051386b0fe0f887f3cf0b1fa29ba39e61e3bcdaa18fe138

Pith citing papers

Observation f4793115-5db0-4ddb-a436-315f54075697 · inbound

NoRD: A Data-Efficient Vision-Language-Action Model that Drives without Reasoning cites this paper.

NoRD: A Data-Efficient Vision-Language-Action Model that Drives without Reasoning HMVLM: Multistage Reasoning-Enhanced Vision-Language Model for Long-Tailed Driving Scenarios

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T21:12:28.565090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:12:28.565090Z digest=sha256:212b5c7c705478283f53bd468aa3cd616093418628839a024cb2f5be3d4925f6

Observation e26871c1-1573-49e7-9ba9-76229a56de1f · inbound

Action Emergence from Streaming Intent cites this paper.

Action Emergence from Streaming Intent HMVLM: Multistage Reasoning-Enhanced Vision-Language Model for Long-Tailed Driving Scenarios

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:02:59.109706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-14T20:59:42.456958Z digest=sha256:8d4143eee6b2aca2ffe8435b3d8bb912941d4fba21e5a6adb93310902fd4bb79

Observation 1da9637a-cbdc-42c7-8af0-cd541d13d649 · inbound

Action Emergence from Streaming Intent cites this paper.

Action Emergence from Streaming Intent HMVLM: Multistage Reasoning-Enhanced Vision-Language Model for Long-Tailed Driving Scenarios

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:15:02.933927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-15T05:13:44.795483Z digest=sha256:17d61d5e7c5077869bb37a0c36bfdbdf14490f376da5a07090f3707db8af5207

Observation 885f1345-7842-4f9b-8277-6f1ed710cde2 · inbound

MindVLA-U1: VLA Beats VA with Unified Streaming Architecture for Autonomous Driving cites this paper.

MindVLA-U1: VLA Beats VA with Unified Streaming Architecture for Autonomous Driving HMVLM: Multistage Reasoning-Enhanced Vision-Language Model for Long-Tailed Driving Scenarios

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:59:28.338032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-14T20:54:50.887381Z digest=sha256:1749390974c78d5fe294d2ed9c18e88fcac5163cad05c2517d4b2b1e0143764e

Observation 37a9b975-654b-4b0f-a209-1fc136e6f53a · inbound

MindVLA-U1: VLA Beats VA with Unified Streaming Architecture for Autonomous Driving cites this paper.

MindVLA-U1: VLA Beats VA with Unified Streaming Architecture for Autonomous Driving HMVLM: Multistage Reasoning-Enhanced Vision-Language Model for Long-Tailed Driving Scenarios

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:09:45.390320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T05:07:58.953866Z digest=sha256:9dcabb6dbaf45ccea59e169ebdd534143bf955b391a4335d83ccfddcfee73e8e

Observation e0dada4d-c031-46c6-8470-896fd16e5ef0 · inbound

NTR: Neural Token Reconstruction for Scene Token Bottleneck in End-to-End Driving cites this paper.

NTR: Neural Token Reconstruction for Scene Token Bottleneck in End-to-End Driving HMVLM: Multistage Reasoning-Enhanced Vision-Language Model for Long-Tailed Driving Scenarios

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-06-28T23:02:46.935701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T22:57:01.742536Z digest=sha256:77b129bf9783a2217877f7fa19c2ad493b2caf0d33fa72e5d6758dba1cc3440e