Pith. sign in

Paper Citation Record · LEDGER

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios

As of 15 August 2026, this Paper Citation Record lists 70 of 70 outbound references and 6 inbound Pith citation observations for arXiv:2501.04671.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.04671 v2

Coverage vector

measured 70 of 70 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T21:30:36.752852Z

measured 76 of 76 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:03:02.400451Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T21:28:58.962434Z

Reference resolution

70 of 70 outbound references displayed

  • verified exact0
  • verified fuzzy40
  • unresolved30
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6eac0e54-24dc-4f11-a550-b4dfa5e8b48a · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Flamingo: a visual language model for few-shot learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T21:30:36.149229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:30:36.149229Z digest=sha256:9786d28bedcf1d62dd1dbbb501faec0c83395c260466573de4b128c6a3906dfc

Observation 2e26796c-e549-427a-a216-75b008ededc2 · outbound

This paper cites Lawrence Zitnick, and Devi Parikh.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Lawrence Zitnick, and Devi Parikh

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:39.197797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T21:30:36.162752Z digest=sha256:7aaee31c7886b83583ee50a8eb065a09f3b03c80529772a1bf74b91249f4fa1b

Observation d8725a61-6142-439d-8372-47b3fd01cea6 · outbound

This paper cites Response Wide Shut: Surprising Observations in Basic Vision Language Model Capabilities.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Response Wide Shut: Surprising Observations in Basic Vision Language Model Capabilities

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T21:30:36.172684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:30:36.172684Z digest=sha256:8dcc1280074d5851fd149470bcf887c2abb898e78ec15741ad6e83e8a42d5486

Observation 170c53ea-a194-4a52-8804-b90bc786f8a9 · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T21:30:36.183356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:30:36.183356Z digest=sha256:fca12e1f6821f93dc7d5364a010f0eeee1802d3a9fa7366a1e159cd5c9699a38

Observation a2bde1ed-0c76-42ea-86fe-afe2792a6806 · outbound

This paper cites PuzzleVQA: Diagnosing Multimodal Reasoning Challenges of Language Models with Abstract Visual Patterns.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios PuzzleVQA: Diagnosing Multimodal Reasoning Challenges of Language Models with Abstract Visual Patterns

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T21:30:36.202132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:30:36.202132Z digest=sha256:002c34fa9613a7c0dc95b4a78030e1890ab8274f24a60b6b370230c46e879952

Observation e9b4910e-bd5c-4a08-8208-9d3b0941154b · outbound

This paper cites Gpt-4 technical report, 2024.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Gpt-4 technical report, 2024

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:39.167769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T21:30:36.212947Z digest=sha256:6e23aec7d51a15c4775ffd550a24d6c233bd4f0e0bf06866509948438dcc797a

Observation f619a843-7cc5-4e5b-8946-22802fc1eb0f · outbound

This paper cites Making the V in VQA matter: El- evating the role of image understanding in visual question answering.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Making the V in VQA matter: El- evating the role of image understanding in visual question answering

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:39.130806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T21:30:36.223655Z digest=sha256:d9e55c4131aea51735283631ff983c8d5c51dbaeb6fa973c96fc58701ec02299

Observation 1983a855-2c2c-4077-b4d2-5c3557fb5f16 · outbound

This paper cites A Survey on LLM-as-a-Judge.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios A Survey on LLM-as-a-Judge

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T21:30:36.240289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:30:36.240289Z digest=sha256:c7b3f81dd79a60b716213a913dd7b4d055c2ae706d23c7392c0528adb545dff8

Observation 6debd138-9dad-49e2-b3b7-33b1df524c3e · outbound

This paper cites Hal- lusionbench: An advanced diagnostic suite for entangled language hallucination and visual illusion in large vision- language models.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Hal- lusionbench: An advanced diagnostic suite for entangled language hallucination and visual illusion in large vision- language models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:39.103518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T21:30:36.247879Z digest=sha256:7fb2e6b2d0788c6e5f2d7a97d999c5b81dedd6d3c14951b14d989b196f785841

Observation 2481ea86-57d9-4fee-ba29-f746ea242c48 · outbound

This paper cites Visual program- ming: Compositional visual reasoning without training.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Visual program- ming: Compositional visual reasoning without training

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:39.065647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T21:30:36.255730Z digest=sha256:34423b498c45de38aef64aaff51d9a54114300050f317b87aae92371056dee79

Observation 7545632b-f99b-434d-b33e-617ce985ee69 · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and composi- tional question answering.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Gqa: A new dataset for real-world visual reasoning and composi- tional question answering

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:39.017237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T21:30:36.262399Z digest=sha256:22322487024ae5bba220d4f35d9da57b8ad6f53e09907ef06eca20f8eb6549e3

Observation be5f94ff-42c1-44b5-90b0-58647365c8c8 · outbound

This paper cites EMMA: End-to-End Multimodal Model for Autonomous Driving.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios EMMA: End-to-End Multimodal Model for Autonomous Driving

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T21:30:36.269691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:30:36.269691Z digest=sha256:fd88e4d49ff37000f3a1573dc129b3ab4a1caea30ed1b709550ffee280ae9200

Observation 1c722649-ff0e-479a-b36d-4a84de4c5db3 · outbound

This paper cites an unresolved cited work.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:30:38.988473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T21:30:36.278634Z digest=sha256:08875f7f55e857d5654b8c1032c7a84a913a720e8777c283f1b10152edafc2ba

Observation a992410b-0699-4e80-bf10-ebca3b119452 · outbound

This paper cites Hallucination augmented contrastive learn- ing for multimodal large language model.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Hallucination augmented contrastive learn- ing for multimodal large language model

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:38.954899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T21:30:36.288146Z digest=sha256:d8df83601a02dc32879ce8eecc4990d4e8ddd33355952946318b06df29d5969f

Observation 720b1843-4ba2-4ccf-8c86-785e32544d45 · outbound

This paper cites Textual explanations for self-driving vehicles.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Textual explanations for self-driving vehicles

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:38.924396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T21:30:36.296054Z digest=sha256:d886ba1447d2495a9c1ec71d6d949ea423b7081937d8f0355ba5306ff0db55ee

Observation 188272a2-6265-4d64-b66a-783612a9a630 · outbound

This paper cites Scaffolding Coordinates to Promote Vision-Language Coordination in Large Multi-Modal Models.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Scaffolding Coordinates to Promote Vision-Language Coordination in Large Multi-Modal Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T21:30:36.312300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:30:36.312300Z digest=sha256:2571d29396b61247c173ceacdafc4ecd8c2589d9b6b9b1d0388b8173df63ceb4

Observation 7bb1aab0-1a22-477b-876e-037b4f056aca · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios LLaVA-OneVision: Easy Visual Task Transfer

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T21:30:36.324818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:30:36.324818Z digest=sha256:013c743744b740076f3f7ede77620007c83f4a95f2f3eda9ec844e7b962373db

Observation 3be75811-bf98-4772-be0c-09b8a9447a60 · outbound

This paper cites From representa- tion to reasoning: Towards both evidence and commonsense reasoning for video question-answering.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios From representa- tion to reasoning: Towards both evidence and commonsense reasoning for video question-answering

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:38.888105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T21:30:36.335361Z digest=sha256:8b556e563a9e71e978e03496ead039054f738c6a5468017a3bc7dffc66395e34

Observation 7cbf36dd-cf39-448b-a6b4-9df9eb0d374b · outbound

This paper cites Chain-of-region: Visual language models need details for diagram analysis.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Chain-of-region: Visual language models need details for diagram analysis

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:38.861322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T21:30:36.343783Z digest=sha256:47101615da106bf95646ed7d7da5b5236e27dba18793eb31c99e2e5eb69a8025

Observation 3accbeb3-2ebb-4ce3-b660-730752a9e7f4 · outbound

This paper cites Enhancing advanced visual reasoning ability of large language models.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Enhancing advanced visual reasoning ability of large language models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:38.834923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T21:30:36.350843Z digest=sha256:db41775aaa707bae5feddca64e3bbc7aa3eff49d786ff5e4921b02a2c883a423

Observation 1607bf4c-20b8-4e3d-91fd-0c2b3270945f · outbound

This paper cites Vila: On pre-training for vi- sual language models.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Vila: On pre-training for vi- sual language models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:38.800965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T21:30:36.358581Z digest=sha256:c63090d7a541fbadff6c1474e57e28d79c2c6dce8b98686311d37fef12b8c3a7

Observation ab371db2-69cb-4c8a-9f81-cb3562882ca2 · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T21:30:36.365886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:30:36.365886Z digest=sha256:a233f0c932cbbbc468be803d58e60d06237ccd72c9ef84206559cd9ec7051a57

Observation 93eee8a4-0a34-466c-840f-62cbf698f9cf · outbound

This paper cites Visual instruction tuning.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Visual instruction tuning

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:38.743639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T21:30:36.371144Z digest=sha256:d43d21b5933ea4435b972d39d578217e70c09211aa786423ecedd8eae30db1aa

Observation 816f331d-2bf9-433b-8222-4abc381eb137 · outbound

This paper cites A Survey on Hallucination in Large Vision-Language Models.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios A Survey on Hallucination in Large Vision-Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T21:30:36.377405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:30:36.377405Z digest=sha256:ddd1cb1565cbe3b92d2ba27c123abd5df012c9dc6514617c0d96bfec21572e9c

Observation c476e120-ccb2-419d-8774-b1239cbad4c0 · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T21:30:36.385263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:30:36.385263Z digest=sha256:d1e3796c4adfee9faea0d6bb05220b5c445a076397709711b725ecc3775bb88b

Observation 44dc0183-0c48-4173-9e4e-b63d19ed9b6d · outbound

This paper cites Aligning with Human Judgement: The Role of Pairwise Preference in Large Language Model Evaluators.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Aligning with Human Judgement: The Role of Pairwise Preference in Large Language Model Evaluators

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T21:30:36.390831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:30:36.390831Z digest=sha256:9298e785cf394437bf279aa7d7059c9e95b472cda65d310552c984dda94f5864

Observation 18877e83-5800-48e1-ae17-c318030e4a98 · outbound

This paper cites SpatialCoT: Advancing Spatial Reasoning through Coordinate Alignment and Chain-of-Thought for Embodied Task Planning.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios SpatialCoT: Advancing Spatial Reasoning through Coordinate Alignment and Chain-of-Thought for Embodied Task Planning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T21:30:36.398003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:30:36.398003Z digest=sha256:ce2ea491fb008fe5549aecccc6bfc33818a0638612f51c97fd0be33081df7950

Observation 3c0eeb34-a7e3-4ab0-b455-8a067e927ebc · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? In European Conference on Computer Vision, pages 216–233.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Mmbench: Is your multi-modal model an all-around player? In European Conference on Computer Vision, pages 216–233

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:38.717175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T21:30:36.414293Z digest=sha256:d580064366c38253da145026950242437fbeaffd7ff29532f9a7a3afb825434f

Observation c598d8a0-c409-402c-8336-af120426c895 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T21:30:36.420260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:30:36.420260Z digest=sha256:34aab81e3054c9061626650179eda503a3460d26b36ba20dfab62164fafba6a0

Observation a306cbfd-7391-4d0b-a444-94c2384f3c63 · outbound

This paper cites Lingoqa: Visual question an- swering for autonomous driving.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Lingoqa: Visual question an- swering for autonomous driving

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:38.658998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T21:30:36.430131Z digest=sha256:a068b1a6b99b3e642bec3014fe542c501543b1c45b70ceb00e33bd97ceedda18

Observation 7e0d9d59-297e-4d3a-acc9-8af6ae3c31b9 · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Ok-vqa: A visual question answering benchmark requiring external knowledge

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:38.628535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T21:30:36.437404Z digest=sha256:2d7bf0a5e4eb28b7023ae823953dc87f469448ac3b520a6e2ddf77f9c0ae0a74

Observation e5d10df3-bf1c-42e5-8782-8b15ee4477ce · outbound

This paper cites Compositional chain-of-thought prompting for large multimodal models.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Compositional chain-of-thought prompting for large multimodal models

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:38.604162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T21:30:36.448556Z digest=sha256:681318596501e5678620412ac2fbff358b35a69e93afb7d36ca838cb499d3615

Observation b5901ceb-6a78-4511-bfbb-ebf7cd7bcf2e · outbound

This paper cites Gpt-4o system card, 2024.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Gpt-4o system card, 2024

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:38.571257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T21:30:36.455451Z digest=sha256:da8a1772b99076f404126e774ad0c48bf5a5ecc9346495c561146f7a0922219f

Observation 2cdf85ba-8a17-45f4-a2ed-e9b400298c5a · outbound

This paper cites Openai o1 system card, 2024.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Openai o1 system card, 2024

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:38.536737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T21:30:36.461161Z digest=sha256:6b6aad469c59df061181f545eb5133ad60751e1bbacb358bbf743a8ffdb73e84

Observation 315e8c15-ac7a-4028-aea9-04a7aaa22793 · outbound

This paper cites Enhancing visual question answering through question-driven image captions as prompts.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Enhancing visual question answering through question-driven image captions as prompts

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:38.496766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T21:30:36.466714Z digest=sha256:930608d18449780ccbf11fab865385fd85afc5a482992623b36d65beefa6e214

Observation ab18df49-45f2-4dd3-b3a7-a91f34b91cc5 · outbound

This paper cites Piaget’s theory of intelligence.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Piaget’s theory of intelligence

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:38.459157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T21:30:36.472476Z digest=sha256:f56c31b25783d4244332baff38fdcb2324016cf4d59901e56793f17b6ab708b1

Observation cbbb9c9f-49d9-4292-95c3-dbc310f9ad4d · outbound

This paper cites CogCoM: A Visual Language Model with Chain-of-Manipulations Reasoning.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios CogCoM: A Visual Language Model with Chain-of-Manipulations Reasoning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T21:30:36.478381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:30:36.478381Z digest=sha256:debee164a57fb8bb48dd07fae3bd165c8058feea4c1a0103dd36c3a048f76435

Observation afd6f303-18b7-408f-b4ec-c257d5eab6d8 · outbound

This paper cites Nuscenes-qa: A multi-modal visual ques- tion answering benchmark for autonomous driving scenario,.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Nuscenes-qa: A multi-modal visual ques- tion answering benchmark for autonomous driving scenario,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:38.435527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T21:30:36.485056Z digest=sha256:9022b939c9cedc316e98f484d39b1f87c72898e9ce0449e7712648486784047e

Observation e49f37cf-2b07-4896-aaa2-02cc98e11da3 · outbound

This paper cites Prism: A framework for decoupling and assessing the capabilities of vlms.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Prism: A framework for decoupling and assessing the capabilities of vlms

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:38.393214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T21:30:36.491409Z digest=sha256:2e114e1e32986d8de88580a017706bef1c4970172fca2ed4bc118b5f4ab78d06

Observation 93b07525-29fd-4455-a296-13be0a3f35d3 · outbound

This paper cites A-okvqa: A benchmark for visual question answering using world knowl- edge.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios A-okvqa: A benchmark for visual question answering using world knowl- edge

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:38.365510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T21:30:36.497058Z digest=sha256:53259c0a881377e8f93d032f73a8be50b4d182655444f67975455b38660abc50

Observation 4e7aab82-e0b8-4f80-9015-05064d369046 · outbound

This paper cites Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T21:30:36.506782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:30:36.506782Z digest=sha256:fcc44ca6ab8ba6524c02dcbc5e3e6feec89cbb7d0cfa2a1bc3be50d09c0ad985

Observation 90af283d-834c-48a2-a291-677c09ccb2a9 · outbound

This paper cites Drivelm: Driving with graph visual ques- tion answering.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Drivelm: Driving with graph visual ques- tion answering

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:38.336952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T21:30:36.514338Z digest=sha256:a56ef514ab553733b446e06c105d33e5f807b181ab826a77f0711d3524f62aaa

Observation 7f0ea722-c5fd-4956-939f-032f56899a7d · outbound

This paper cites Towards vqa models that can read.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Towards vqa models that can read

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:38.314568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T21:30:36.522235Z digest=sha256:c7ce6feb9b70bfc06dbd6aeb5ecb9b899e067ff6976a3cc2b874d8d51bbecc26

Observation 5b5c25b4-0b85-440f-b652-98f777adcca8 · outbound

This paper cites To CoT or not to CoT? Chain-of-thought helps mainly on math and symbolic reasoning.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios To CoT or not to CoT? Chain-of-thought helps mainly on math and symbolic reasoning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T21:30:36.527626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:30:36.527626Z digest=sha256:a7b43d242c861ee86f49fd589d4319080b3b9b05a1a73f9b60abfc504386921e

Observation e14b3747-d2fb-430b-b096-0899bfa22de4 · outbound

This paper cites Vipergpt: Visual inference via python execution for reasoning.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Vipergpt: Visual inference via python execution for reasoning

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:38.289634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T21:30:36.534132Z digest=sha256:5376731ac6699d11fea4c0c1ee0a54987c9a0163441e01844e93b0a8ad793e68

Observation 15e39f8e-855c-481f-9923-0c43491f3ba6 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios LLaMA: Open and Efficient Foundation Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T21:30:36.540535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:30:36.540535Z digest=sha256:25d9cf55a10992e1476d1275fb1f548572d7f9df01dc11e1afe415096d2e3c46

Observation 49cf1156-a5f3-4770-952d-aafffb4ef24c · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T21:30:36.556236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:30:36.556236Z digest=sha256:c2300e1eaff433bf190b0735afe0bccab8a3ea948cdaae1954586eb60777a952

Observation 6f5ddf51-ea5d-4951-8760-992ffa7f679f · outbound

This paper cites an unresolved cited work.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:30:38.251058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T21:30:36.561374Z digest=sha256:629b4d004529349955c0ede65fa156a46136e01e30e160da62ce6327e2fdf0d6

Observation 98c8e57b-12de-4f7f-9ad7-3b42b7c86340 · outbound

This paper cites Drive anywhere: Generalizable end-to-end autonomous driving with multi-modal foundation models,.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Drive anywhere: Generalizable end-to-end autonomous driving with multi-modal foundation models,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:38.222761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T21:30:36.570857Z digest=sha256:f394ce4eb95ab57ba65393e0a6c12eaf64cda534d22d885d92cf90eebc336c6d

Observation d7181fe8-ce77-4cc8-9b33-66d1e145248b · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Chain-of-thought prompting elicits reasoning in large language models

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:38.181322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T21:30:36.577033Z digest=sha256:4118ce9d6ddc536b9807ebadb62ee91d417316f96b9610190239117b3f87971b

Observation d63dfdb6-12da-46c9-a446-3686f6de3944 · outbound

This paper cites Drivegpt4: Interpretable end-to-end autonomous driving via large language model.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Drivegpt4: Interpretable end-to-end autonomous driving via large language model

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-10T21:30:36.583053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:30:36.583053Z digest=sha256:a957ac2e18376cdaf32ca9a1977ec08a5ca46db108baea53fe03325d62c04890

Observation 25ccc62b-e148-4520-b3d1-7a631ec14ada · outbound

This paper cites List Items One by One: A New Data Source and Learning Paradigm for Multimodal LLMs.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios List Items One by One: A New Data Source and Learning Paradigm for Multimodal LLMs

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T21:30:36.589188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:30:36.589188Z digest=sha256:2fab233fe5f0bb14c98b0dc1094820f83203ceb8e3ba1c37159bec5e43ef3a59

Observation 16276f81-c2a7-4957-a3b5-1284f8ccddc9 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understand- ing and reasoning benchmark for expert agi.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Mmmu: A massive multi-discipline multimodal understand- ing and reasoning benchmark for expert agi

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:38.118245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T21:30:36.597676Z digest=sha256:83728f76497ed33390f2e4ed046e49a397176504f3fbb1b38eb39bfbe955d8ea

Observation 107f1902-3405-4a7f-80e0-d0f480279289 · outbound

This paper cites Sigmoid loss for language image pre-training.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Sigmoid loss for language image pre-training

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-10T21:30:36.606307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:30:36.606307Z digest=sha256:a44e66bbe81b70687eab716ea2534420d8de4f6479f4b16c7ce50965d409dcb6

Observation 9105b1a2-9d67-4191-9bda-b93fcbc9e225 · outbound

This paper cites Multimodal chain-of-thought rea- soning in language models.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Multimodal chain-of-thought rea- soning in language models

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:38.056246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T21:30:36.617223Z digest=sha256:4e31082f55afc11b4f8dfee000d6effa2f46c27357de7668c0ec53c2889e8795

Observation a497be7e-228a-40db-8662-6f16761e7bd8 · outbound

This paper cites Multimodal chain-of-thought reasoning in language models.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Multimodal chain-of-thought reasoning in language models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-10T21:30:36.626364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:30:36.626364Z digest=sha256:30ebab3641e4c580599302ec044d10d9d73e01ffcb75060c79528ade78f9f283

Observation d8dfc508-a7b9-4b03-b700-9c1c3c72fc18 · outbound

This paper cites Ddcot: Duty-distinct chain-of-thought prompting for multimodal reasoning in language models.Advances in Neu- ral Information Processing Systems, 36:5168–5191, 2023.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Ddcot: Duty-distinct chain-of-thought prompting for multimodal reasoning in language models.Advances in Neu- ral Information Processing Systems, 36:5168–5191, 2023

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:38.011547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T21:30:36.633857Z digest=sha256:b5ffd57a30871d2ad118cc6cbecdeb2d55e4faac0bfff1a64f9c66b74869c979

Observation 4f8b75b4-414a-4196-8d22-946366919bff · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in Neural Information Processing Systems, 36:46595–46623, 2023.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in Neural Information Processing Systems, 36:46595–46623, 2023

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:37.979560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T21:30:36.640347Z digest=sha256:587052ded1e97ede8c27ae730a4dc8e264f1f3da3d6a03eae40d912f01e738d7

Observation 88c8e4ed-9e2b-4904-ae87-bafb6fde5c37 · outbound

This paper cites an unresolved cited work.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:30:37.929860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T21:30:36.647145Z digest=sha256:6fe80b34d8a6009f383426ef82327a35d48f559d02bef179ef7021d8a76f4abc

Observation 03099f83-4d89-4e61-b2e7-621853fab1f8 · outbound

This paper cites an unresolved cited work.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Unresolved cited work

Reference 61

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:30:37.896710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T21:30:36.653178Z digest=sha256:a1958fa205cf463924c826e02c88ed18a0d7c86451260cbff29946277964928b

Observation dbf4994c-883c-449e-a3eb-f488b0e35fe9 · outbound

This paper cites Replicate bound- ing box coordinates exactly as provided in the list.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Replicate bound- ing box coordinates exactly as provided in the list

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:37.861424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T21:30:36.663540Z digest=sha256:0ab3b5550fed5d30e9e1237c45ccbcd0c7a725eb8f640763ea2dfbf142c3cc3b

Observation b6dc4b4f-c5ef-44d4-b990-29549955602d · outbound

This paper cites an unresolved cited work.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:30:37.822666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T21:30:36.672684Z digest=sha256:ec7ff8312d687fbfe107672727c839f35c270314baa3ae0f7d0472f2aab85626

Observation e89bb2bf-b40d-4a7a-a170-d73304b334e2 · outbound

This paper cites 1 Demonstration 1 Question: [”I am turning right at the next intersec- tion.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios 1 Demonstration 1 Question: [”I am turning right at the next intersec- tion

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:37.786060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T21:30:36.684686Z digest=sha256:7e97511c0984905e7aae5b9df6618ac63c0fd3d3318c559692a1bc483e33267d

Observation a29777f7-d99d-46cd-b5e7-68a6d40df5fc · outbound

This paper cites an unresolved cited work.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Unresolved cited work

Reference 65

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:30:37.748045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T21:30:36.692496Z digest=sha256:20f8ac2b5a82741ccf8e06b298024b37ea8602db942ef5722fe94e795ef1ba20

Observation ef6e2bd7-dc18-4447-82e3-2d448fb9565f · outbound

This paper cites an unresolved cited work.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Unresolved cited work

Reference 66

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:30:37.711537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T21:30:36.699972Z digest=sha256:a3ae8485866ea018657e633cc0191bf23bcea878de28985b864ad6e805a2f762

Observation f9faba4c-7921-4df5-8c31-85da6192b6f1 · outbound

This paper cites correct reasoning.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios correct reasoning

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:37.661846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T21:30:36.710086Z digest=sha256:9c8e05279074afb81e239df49a96a43e28167b0c9eb781ee4464cb7750077188

Observation a42d84ca-e754-4ff3-aa01-48bbdf57abf3 · outbound

This paper cites Step-by-Step Instructions:.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios Step-by-Step Instructions:

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:37.634427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T21:30:36.718007Z digest=sha256:1d8e939756fccb346f231d5ec49605a2f42f411454dbc862cc0115e39a41ef24

Observation 0a4c0857-cd4a-4e09-9741-40eccefa8d5c · outbound

This paper cites • For each argument, briefly state whether it is correct or not, given the provided correct reasoning.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios • For each argument, briefly state whether it is correct or not, given the provided correct reasoning

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:37.595726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T21:30:36.733599Z digest=sha256:679278bbd7b39458c0c386b1dbc049c9cde49133c2e2ca82ef96eabd627e91ae

Observation 64ee6a68-f966-40df-82dc-1caaf70048b5 · outbound

This paper cites • List important points or steps from the correct reasoning that the student omits or directly contradicts.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios • List important points or steps from the correct reasoning that the student omits or directly contradicts

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:37.569194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T21:30:36.741666Z digest=sha256:0750666bc6c53d29efacadc5a43588f042b2457a5ef98bdcee70f45867bff7e7

Observation 7e0f1922-d574-4a0d-b64e-7aa12ba5d9f5 · outbound

This paper cites 1” if you judge the student’s reasoning is overall correct, “0.

Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios 1” if you judge the student’s reasoning is overall correct, “0

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:30:37.528655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T21:30:36.752852Z digest=sha256:abfed24e03b59855a0adfda8485412c82a14898d40a1bfa847a64c569994998e

Pith citing papers

Observation 9c7da779-1d85-4c52-8184-f79c7db28ea5 · inbound

Chain-of-Thought for Autonomous Driving: A Comprehensive Survey and Future Prospects cites this paper.

Chain-of-Thought for Autonomous Driving: A Comprehensive Survey and Future Prospects Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios

Reference 127

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:02.400451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:02.400451Z digest=sha256:f94da181aa0f570609a15fab627b93548a1706ec37cc954a379b5ce79343cb14

Observation 20675e45-9e4a-438a-9623-f38d744f27ab · inbound

Alpamayo-R1: Bridging Reasoning and Action Prediction for Generalizable Autonomous Driving in the Long Tail cites this paper.

Alpamayo-R1: Bridging Reasoning and Action Prediction for Generalizable Autonomous Driving in the Long Tail Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-18T02:35:13.306854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T02:35:13.126171Z digest=sha256:c03596c69e0220cbe29b44dafd0e0d6a0d38f799dd657eb993321bf03582b14b

Observation f7b6ac57-9991-47a6-a0a8-920044a279fd · inbound

OmniDrive-R1: Reinforcement-driven Interleaved Multi-modal Chain-of-Thought for Trustworthy Vision-Language Autonomous Driving cites this paper.

OmniDrive-R1: Reinforcement-driven Interleaved Multi-modal Chain-of-Thought for Trustworthy Vision-Language Autonomous Driving Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:38:37.815554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T22:34:00.895252Z digest=sha256:140b9dd4137c9ddc390e051a8d183bcd899b2b8e540df1e65a187f3f6e29fac7

Observation 9e10aa20-2566-4a03-94a5-b7daf20abbb7 · inbound

RailVQA: A Benchmark and Framework for Efficient Interpretable Visual Cognition in Automatic Train Operation cites this paper.

RailVQA: A Benchmark and Framework for Efficient Interpretable Visual Cognition in Automatic Train Operation Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:28:04.771507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-14T22:24:13.897439Z digest=sha256:aba8d98edb7e3aee4b86b1f308daa7098509e7b579ef5272fe96d512ef88e078

Observation bec3ebc2-c327-4c0a-97e8-589271f2df65 · inbound

What to Test Next: Interpretable Coverage Gap Discovery in Driving VLMs cites this paper.

What to Test Next: Interpretable Coverage Gap Discovery in Driving VLMs Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:16:16.242167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-28T15:36:25.219607Z digest=sha256:5db05ed449e1e58d4b0edfaf4e78b17c37991ca1130cd167d5c5a14fe4ab1a02

Observation 2ae9e646-ecb7-41cd-a9d2-b45ce7697733 · inbound

CaVe-VLM-CoT: An Interpretable Vision-Language Model Framework cites this paper.

CaVe-VLM-CoT: An Interpretable Vision-Language Model Framework Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T21:28:58.963898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-27T00:31:01.692743Z digest=sha256:df5f62e2242952745a0c0f0aa919efcbbdbbd1a6f6ae87fa1c46ffcc6060e07e