Pith. sign in

Paper Citation Record · LEDGER

LaVida Drive: Vision-Text Interaction VLM for Autonomous Driving with Token Selection, Recovery and Enhancement

As of 15 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2411.12980.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.12980 v3

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T17:04:12.534638Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

31 of 31 outbound references displayed

  • verified exact0
  • verified fuzzy19
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 86041218-8aff-4b6b-84bc-689843ff16d2 · outbound

This paper cites Meteor: An auto- matic metric for mt evaluation with improved correla- tion with human judgments.

LaVida Drive: Vision-Text Interaction VLM for Autonomous Driving with Token Selection, Recovery and Enhancement Meteor: An auto- matic metric for mt evaluation with improved correla- tion with human judgments

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:04:13.000461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:04:12.403437Z digest=sha256:04a849d159390adfd6416f125e2d35506f599d298f66b31eee52c7e2782dcc62

Observation f011736c-f8a9-4ef8-a26c-fe26925ceb21 · outbound

This paper cites Is space-time attention all you need for video under- standing? In Proceedings of the International Confer- ence on Machine Learning (ICML) , 2021.

LaVida Drive: Vision-Text Interaction VLM for Autonomous Driving with Token Selection, Recovery and Enhancement Is space-time attention all you need for video under- standing? In Proceedings of the International Confer- ence on Machine Learning (ICML) , 2021

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:04:12.986694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:04:12.408133Z digest=sha256:6ad91c37d7f1b23e4e07c229bde92e20282d416f30cd27ca32445ff78c422e69

Observation f84fc7a9-acd8-41d7-8f0b-201a3dc32b79 · outbound

This paper cites Driving with llms: Fusing object-level vector modality for explainable au- tonomous driving.

LaVida Drive: Vision-Text Interaction VLM for Autonomous Driving with Token Selection, Recovery and Enhancement Driving with llms: Fusing object-level vector modality for explainable au- tonomous driving

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:04:12.972759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:04:12.413197Z digest=sha256:e087f49f6847e83df2857e5fd62d623d9b1198d516678d827b3a5ac4ec48e98b

Observation 6aafe554-1c1a-4b4f-abac-76eda3a1acb6 · outbound

This paper cites an unresolved cited work.

LaVida Drive: Vision-Text Interaction VLM for Autonomous Driving with Token Selection, Recovery and Enhancement Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-12T17:04:12.959236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:04:12.418422Z digest=sha256:cf0e2d99ca56516d0932c02468415c65616017708296c79468138d12b7c4e9fa

Observation d8aea3a9-ae68-4007-bee3-80b756e9160e · outbound

This paper cites an unresolved cited work.

LaVida Drive: Vision-Text Interaction VLM for Autonomous Driving with Token Selection, Recovery and Enhancement Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-12T17:04:12.943907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:04:12.423188Z digest=sha256:720442db214f36de0a425093582fdd060b2a5d8f914f54b47a0890eb934f071f

Observation 538aa569-23ce-4222-b5e1-06dda7eca6ca · outbound

This paper cites Multi-frame, lightweight & efficient vision- language models for question answering in autonomous driving, 2024.

LaVida Drive: Vision-Text Interaction VLM for Autonomous Driving with Token Selection, Recovery and Enhancement Multi-frame, lightweight & efficient vision- language models for question answering in autonomous driving, 2024

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:04:12.929941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:04:12.427829Z digest=sha256:9c5d785c890ba43ae142d110c85a6b028e52e7e66e878808aaaf2f84fd9f59d4

Observation a4625a06-954b-4d22-b348-91195ef325f5 · outbound

This paper cites Lan- guage is not all you need: Aligning perception with language models.

LaVida Drive: Vision-Text Interaction VLM for Autonomous Driving with Token Selection, Recovery and Enhancement Lan- guage is not all you need: Aligning perception with language models

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:04:12.916476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:04:12.433051Z digest=sha256:2eac97a871b7d4897150650e8e52123e1429f6915130754d031eabe290589cf1

Observation ba3795a1-fcc4-443f-8bc1-29b4958c04c7 · outbound

This paper cites an unresolved cited work.

LaVida Drive: Vision-Text Interaction VLM for Autonomous Driving with Token Selection, Recovery and Enhancement Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-12T17:04:12.903365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:04:12.437309Z digest=sha256:0ed80936a286df23e1dbde6636d51774d63ae6a3b033d18b87dc328a70c0ea34

Observation 6d7d7217-284e-4943-b22a-6c26ee1f3184 · outbound

This paper cites Vlm2scene: Self-supervised image-text-lidar learning with foundation models for autonomous driving scene understanding.

LaVida Drive: Vision-Text Interaction VLM for Autonomous Driving with Token Selection, Recovery and Enhancement Vlm2scene: Self-supervised image-text-lidar learning with foundation models for autonomous driving scene understanding

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:04:12.890378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:04:12.441331Z digest=sha256:cd037a2e0c48e4fc445f71768de99fdc166f26bc7c0568f4ca3ebbabbd6d2e06

Observation bad55b0c-c5c0-42a3-aaf7-2de3c733c1a4 · outbound

This paper cites Rouge: A package for automatic eval- uation of summaries.

LaVida Drive: Vision-Text Interaction VLM for Autonomous Driving with Token Selection, Recovery and Enhancement Rouge: A package for automatic eval- uation of summaries

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:04:12.876859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:04:12.446091Z digest=sha256:96b9eb5baab9f7619da136007959e211d48d786841ff2955fe1cab8f39fecc75

Observation b553a8a7-9463-4a09-90ec-85fa53af4f8d · outbound

This paper cites Reason2drive: Towards interpretable and chain-based reasoning for autonomous driving.

LaVida Drive: Vision-Text Interaction VLM for Autonomous Driving with Token Selection, Recovery and Enhancement Reason2drive: Towards interpretable and chain-based reasoning for autonomous driving

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:04:12.862162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:04:12.450285Z digest=sha256:2b929e7721d80e8144d87a09c64bbc904c2bf665200036d9f81d0587eb0a5070

Observation c1ef6a27-0967-4f2a-9925-6cf2dd410e98 · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation.

LaVida Drive: Vision-Text Interaction VLM for Autonomous Driving with Token Selection, Recovery and Enhancement Bleu: a method for automatic evaluation of machine translation

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:04:12.849459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:04:12.454302Z digest=sha256:e55103817a30ff2deabb4a1967ff40e048cd83586b4b4f0046277f64a276f1a5

Observation ecb3508d-e648-4b9c-8038-f9a54b2c54bc · outbound

This paper cites Nuscenes-qa: A multi-modal vi- sual question answering benchmark for autonomous driving scenario.

LaVida Drive: Vision-Text Interaction VLM for Autonomous Driving with Token Selection, Recovery and Enhancement Nuscenes-qa: A multi-modal vi- sual question answering benchmark for autonomous driving scenario

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:04:12.835795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:04:12.458461Z digest=sha256:1ba88e29d75d803e6d235f8d1d80dcd45c31c984d06e49227f80cf1343b4eb70

Observation f224662a-2fc6-4579-8a4b-aafadeff4360 · outbound

This paper cites Nuscenes-qa: A multi-modal vi- sual question answering benchmark for autonomous driving scenario, 2024.

LaVida Drive: Vision-Text Interaction VLM for Autonomous Driving with Token Selection, Recovery and Enhancement Nuscenes-qa: A multi-modal vi- sual question answering benchmark for autonomous driving scenario, 2024

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:04:12.820295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:04:12.462299Z digest=sha256:69462d10e324e880035e1750590a7942ba52aff43171fa482360aaa58993c2ae

Observation 61fe9706-91d6-4d97-b1d3-bd4be4186a4e · outbound

This paper cites Learning trans- ferable visual models from natural language supervi- sion.

LaVida Drive: Vision-Text Interaction VLM for Autonomous Driving with Token Selection, Recovery and Enhancement Learning trans- ferable visual models from natural language supervi- sion

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:04:12.806538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:04:12.466285Z digest=sha256:c7244550ad6f5ca85cf55eecea9e173372fe915f56efd69afa63c8735ec949c3

Observation 6c8d36c6-10d2-4da4-a0f2-dc81b6105dc9 · outbound

This paper cites Learning transferable visual models from natural lan- guage supervision.

LaVida Drive: Vision-Text Interaction VLM for Autonomous Driving with Token Selection, Recovery and Enhancement Learning transferable visual models from natural lan- guage supervision

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:04:12.792941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:04:12.470510Z digest=sha256:a167b42516c1da0c6aabb987b65ef40d1375053bc003c3f576dd753e98e79af8

Observation 01bb169f-3f79-4e84-95bf-473735016ed3 · outbound

This paper cites an unresolved cited work.

LaVida Drive: Vision-Text Interaction VLM for Autonomous Driving with Token Selection, Recovery and Enhancement Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-12T17:04:12.778559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:04:12.474717Z digest=sha256:830fb7896a8a42c4a90332f982cad188fb2cc6a940742e6cae6db362cae765eb

Observation 2dfbf586-d531-4699-bb65-c4f8b611ec18 · outbound

This paper cites LanguageMPC: Large Language Models as Decision Makers for Autonomous Driving.

LaVida Drive: Vision-Text Interaction VLM for Autonomous Driving with Token Selection, Recovery and Enhancement LanguageMPC: Large Language Models as Decision Makers for Autonomous Driving

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T17:04:12.479080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:04:12.479080Z digest=sha256:fe1803caf988658f76d1552b31947b7ebfa7f6f83a80d14475271712c83f42ca

Observation 84653957-d64b-4755-a4ba-b2c4cc630430 · outbound

This paper cites DriveLM: Driving with Graph Visual Question Answering.

LaVida Drive: Vision-Text Interaction VLM for Autonomous Driving with Token Selection, Recovery and Enhancement DriveLM: Driving with Graph Visual Question Answering

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T17:04:12.484511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:04:12.484511Z digest=sha256:0a46f62a1523cc3851b41a14df0bec2db4c7ad35dc06c47ea395abc8d756707a

Observation 5b51f4e3-4620-47fb-8dae-1566368f4a0c · outbound

This paper cites DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models.

LaVida Drive: Vision-Text Interaction VLM for Autonomous Driving with Token Selection, Recovery and Enhancement DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T17:04:12.489019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:04:12.489019Z digest=sha256:975f9f971b6bd3755317dbbebf89c0ad84bc02481d48c9a1b0f77e9051608438

Observation 35c8512b-0fc0-4eab-bb37-bd93adca5750 · outbound

This paper cites Cider: Consensus-based image descrip- tion evaluation.

LaVida Drive: Vision-Text Interaction VLM for Autonomous Driving with Token Selection, Recovery and Enhancement Cider: Consensus-based image descrip- tion evaluation

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:04:12.765032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:04:12.493514Z digest=sha256:45cf6bb66afedb40dec1904995c892088f44e1cc1761a76cd05ee5e2b89918fa

Observation f69c3665-b71d-42e9-a2c5-7c6f890bdb49 · outbound

This paper cites Driving into the future: Multiview visual forecasting and planning with world model for autonomous driving.

LaVida Drive: Vision-Text Interaction VLM for Autonomous Driving with Token Selection, Recovery and Enhancement Driving into the future: Multiview visual forecasting and planning with world model for autonomous driving

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:04:12.751668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:04:12.497165Z digest=sha256:f486f4c60641432e8913c2283adfa8378017c38f2aa0af9a96d70ae4f75ba258

Observation 15cdc746-7f21-4404-b411-0f9b966323bc · outbound

This paper cites an unresolved cited work.

LaVida Drive: Vision-Text Interaction VLM for Autonomous Driving with Token Selection, Recovery and Enhancement Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-12T17:04:12.738054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:04:12.500769Z digest=sha256:d09f330638b3aea45114ba684b2a6d2abdca0cd0ad728405a60cb8dd0e92f8ce

Observation c36ce008-a22d-42ed-92d7-eeac749c2f44 · outbound

This paper cites On the Road with GPT-4V(ision): Early Explorations of Visual-Language Model on Autonomous Driving.

LaVida Drive: Vision-Text Interaction VLM for Autonomous Driving with Token Selection, Recovery and Enhancement On the Road with GPT-4V(ision): Early Explorations of Visual-Language Model on Autonomous Driving

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T17:04:12.504683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:04:12.504683Z digest=sha256:094450a23cc2d634cf98f8a88548e2248215aa7fdd4e7994165787acf447c6d4

Observation 886f1fca-9726-41e7-ba3f-309ae6077c0f · outbound

This paper cites Drivegpt4: Interpretable end-to-end au- tonomous driving via large language model.

LaVida Drive: Vision-Text Interaction VLM for Autonomous Driving with Token Selection, Recovery and Enhancement Drivegpt4: Interpretable end-to-end au- tonomous driving via large language model

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:04:12.724266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:04:12.508759Z digest=sha256:71468472ce8b41d13b8b68d2a3e77e516b3d06fda522660c55d89451da6e55d5

Observation ec8d9c15-47d0-4e7d-a638-83c0f32f72fd · outbound

This paper cites Generalized predictive model for autonomous driving.

LaVida Drive: Vision-Text Interaction VLM for Autonomous Driving with Token Selection, Recovery and Enhancement Generalized predictive model for autonomous driving

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:04:12.708896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:04:12.512334Z digest=sha256:f3f172a80e67b1c2cf22e6b29d28378a01e53ab91f4fa449084b748232b162ca

Observation c18a6a17-0799-481b-a108-076ff5cd00e6 · outbound

This paper cites an unresolved cited work.

LaVida Drive: Vision-Text Interaction VLM for Autonomous Driving with Token Selection, Recovery and Enhancement Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-12T17:04:12.691591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:04:12.516332Z digest=sha256:56dfa9cc26786068a8bde5bd67784be8788640530108bcd30ed019f600bcc2f0

Observation 89310149-e210-4e64-9f4a-f30d081609b5 · outbound

This paper cites an unresolved cited work.

LaVida Drive: Vision-Text Interaction VLM for Autonomous Driving with Token Selection, Recovery and Enhancement Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-12T17:04:12.675307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:04:12.521039Z digest=sha256:5610d2103e1dc212deda42b4f0a0419e3e9a3e94d7be06bf21599d642d7e9611

Observation 0f57dc8e-2d87-49ac-805a-5f9b85c6f755 · outbound

This paper cites Zhang, L.

LaVida Drive: Vision-Text Interaction VLM for Autonomous Driving with Token Selection, Recovery and Enhancement Zhang, L

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:04:12.659776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:04:12.526011Z digest=sha256:70288bbb428ac8f34d09496d4f1af960753b3d50045fe7e0ac24a2744573f2d1

Observation 6517acc2-d355-421e-87a7-9ab5d979b4c1 · outbound

This paper cites Vinvl: Revisiting visual representa- tions in vision-language models.

LaVida Drive: Vision-Text Interaction VLM for Autonomous Driving with Token Selection, Recovery and Enhancement Vinvl: Revisiting visual representa- tions in vision-language models

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:04:12.643933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:04:12.530750Z digest=sha256:35bfcfc2704718f09c36d5a8a9456e71f7ec513c735ea996ec2e8daa37fe8891

Observation 3794ec17-d05f-4a00-9956-3b7668594d7e · outbound

This paper cites an unresolved cited work.

LaVida Drive: Vision-Text Interaction VLM for Autonomous Driving with Token Selection, Recovery and Enhancement Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-12T17:04:12.626385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T17:04:12.534638Z digest=sha256:eb555c102c11ade183ea90e55b111e983a23da6959928894b4327124f0fc230b

Pith citing papers

No inbound Pith citation observations are available.