Pith. sign in

Paper Citation Record · LEDGER

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation

As of 10 August 2026, this Paper Citation Record lists 100 of 121 outbound references and 12 inbound Pith citation observations for arXiv:2604.18486.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.18486 v3

Coverage vector

measured 100 of 121 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-12T00:54:23.508845Z

measured 112 of 112 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T19:56:57.482174Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-02T16:57:10.319204Z

Reference resolution

100 of 121 outbound references displayed

  • verified exact49
  • verified fuzzy50
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation afe21470-a494-48dc-97f9-e6a16805d14c · outbound

This paper cites Claude 3.7 Sonnet and Claude Code.https://www.anthropic.com/news/claude-3-7-sonnet.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation Claude 3.7 Sonnet and Claude Code.https://www.anthropic.com/news/claude-3-7-sonnet

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T09:28:26.956782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:ba033c39992cc65dc83e3ee1414525c92b1f358a5945b0077f3bf1f22d7cf201

Observation ff616f81-e00f-4a4e-8134-d1475698d7bd · outbound

This paper cites Self-supervised learning from images with a joint-embedding predictive architecture.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation Self-supervised learning from images with a joint-embedding predictive architecture

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T09:28:27.146836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:62c2ee09cfcf2b022a0772c829731145bf82e207350e733b923b5fda12ccb37e

Observation 3627fa32-91d8-45b1-8876-bba15b85d603 · outbound

This paper cites Cosmos-Reason1: From Physical Common Sense To Embodied Reasoning.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation Cosmos-Reason1: From Physical Common Sense To Embodied Reasoning

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:47:10.348602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:0a408ba8ffd84038e0e145d8d3047673dca0adb68521eda09a7987a16cc327c6

Observation cb490e54-c2fb-4f2b-921b-d4dbcdf5a3da · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:36:25.958151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:cd17b2b91bcf4ca821bf4af9361d9f8cece32277241d2b34c46d510e350b6673

Observation 53e59826-aefe-43c2-809c-72c9aa545c43 · outbound

This paper cites Qwen3-VL Technical Report.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation Qwen3-VL Technical Report

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:36:26.145301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:16526a9f9c7dec643e1e76d2906b375e9f28d72be0d806847267369eff9c3c79

Observation 885fe00b-5b3f-4ad9-8230-102eb2512c51 · outbound

This paper cites Qwen2.5-VL Technical Report.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation Qwen2.5-VL Technical Report

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:36:26.005487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:e1c45cc2fe0524b3c2379968694e4255bd5dc2e6134166fe019a7243f037f465

Observation 64fbd631-b097-441b-ac7b-26de69cc7a57 · outbound

This paper cites Dynamiccity: Large-scale 4d oc- cupancy generation from dynamic scenes.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation Dynamiccity: Large-scale 4d oc- cupancy generation from dynamic scenes

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:36:26.114624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:22cb276e573729d1c9ab5645e42a53e5080ba3d8776a00ee3aac0db3eaeea7c2

Observation b816d1b3-3b7e-4592-9b40-00e4523b8b45 · outbound

This paper cites nuscenes: A multimodal dataset for autonomous driving.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation nuscenes: A multimodal dataset for autonomous driving

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T09:28:26.948035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:94177a1af30bf42291d1ab66e49e28529cbed8f8f9c9f8135cec7eac058896e1

Observation 9dc12a58-83f3-49b9-adcf-dbafa8bd9fad · outbound

This paper cites Maplm: A real-world large-scale vision-language benchmark for map and traffic scene understanding.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation Maplm: A real-world large-scale vision-language benchmark for map and traffic scene understanding

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T09:28:27.102138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:84a34f9b9dfb44fdc132acb6354510f920fe1b5ed93e464fbbaf46f8fd3bbf9a

Observation 7ea35755-2bd2-4cd7-9e3b-2514ce54b21d · outbound

This paper cites arXiv preprint arXiv:2510.25122 (2025).

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation arXiv preprint arXiv:2510.25122 (2025)

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:36:25.952483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:262a2fb33aa5aefa75f3b200d6f7027a1d74954b740e9f8afc267a0aca3f29a6

Observation 1938e40b-b0f9-4977-a83d-a2cf3c406a3e · outbound

This paper cites M3-Embedding: Multi-Linguality, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation M3-Embedding: Multi-Linguality, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:36:26.135572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:c43f9c45d3860de4a84e14bf479db5936e45d294476ca8524f2652b53c2d745f

Observation 6f6069d0-72d4-4e34-aaa0-6e99044706f0 · outbound

This paper cites Automated evaluation of large vision-language models on self-driving corner cases.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation Automated evaluation of large vision-language models on self-driving corner cases

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T09:28:27.130709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:634dd2d9b50ec92d952498bbc8500b83fcc44aeb86d9e87a9ad235e72aa7f94b

Observation 41071d63-b645-49a3-9130-b66678e30519 · outbound

This paper cites Driving with llms: Fusing object-level vector modality for explainable autonomous driving.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation Driving with llms: Fusing object-level vector modality for explainable autonomous driving

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T09:28:27.081174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:823142c2d272fecf4967f4c5f6ec05717005867add44cf4f8e0812daab7d09d6

Observation 3a5e2368-0cbf-46a5-ac7b-06088f13084f · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation Evaluating Large Language Models Trained on Code

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:36:26.082150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:07d3ac67e3374fa83b4b93d8e1ffc99df3928715ceae935056b2eb00fe73811d

Observation 112f04bb-3a77-45e8-b544-4996b1204386 · outbound

This paper cites Vilta: A vlm-in-the-loop adversary for enhancing driving policy robustness.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation Vilta: A vlm-in-the-loop adversary for enhancing driving policy robustness

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:36:25.994615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:7e3809a684fa700edfee74316d5441ab19cb3b5d25a13c1feca06d4085bd17b2

Observation 47737d8e-5862-4011-bebb-f6469cdb3ff9 · outbound

This paper cites Compressed Chain of Thought: Efficient Reasoning Through Dense Representations.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation Compressed Chain of Thought: Efficient Reasoning Through Dense Representations

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-17T04:47:40.470022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:3654a47f88bbb675755317b6d28fccd19d43190edbaeb452b307e82197767c19

Observation e7693050-a7ed-4eb9-a224-00b35724fa7f · outbound

This paper cites Impromptu VLA: Open weights and open data for driving vision-language-action models.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation Impromptu VLA: Open weights and open data for driving vision-language-action models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T09:28:27.122045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:5da0a453faa4990a0695a911df1590cbaf79692ff775a8c7e9f3c7e4b7304e49

Observation 19177e61-0370-4e00-b196-c64d4bda4c92 · outbound

This paper cites Emu3.5: Native Multimodal Models are World Learners.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation Emu3.5: Native Multimodal Models are World Learners

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:12:13.931083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:39bd33d8c9f4e6584e022fea46b4a59fffcb2c807798c2705506b3b83ce565f3

Observation b46d9a9d-4e92-4705-9491-c54b7881db37 · outbound

This paper cites Navsim: Data-driven non-reactive autonomous vehicle simulation and benchmarking.Advancesin Neural Information Processing Systems, 37:28706–28719.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation Navsim: Data-driven non-reactive autonomous vehicle simulation and benchmarking.Advancesin Neural Information Processing Systems, 37:28706–28719

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T09:28:26.993180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:fbaeebc7b817e9866fbe2bddf0da8f6b6dc8d71359f66078f4fe56fc6b025a91

Observation 8ac4bf0a-0f8b-4089-a5d3-767c35939e5c · outbound

This paper cites Language Modeling Is Compression.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation Language Modeling Is Compression

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-17T22:36:13.497932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:d206090d45452321df55389833469be73b020de9883ace9c7d0cb64b404543da

Observation 1b9bba79-0104-4460-a111-fe205e9fbb16 · outbound

This paper cites From Explicit CoT to Implicit CoT: Learning to Internalize CoT Step by Step.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation From Explicit CoT to Implicit CoT: Learning to Internalize CoT Step by Step

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-16T11:44:32.634190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:287a0a467a43f947d36ff2397ff01d6c8ad0e63984648e53facb276708514284

Observation c8166422-dbb3-4821-a7d7-ed67d279e2bf · outbound

This paper cites Holistic autonomous driving understanding by bird’s-eye-view injected multi-modal large models.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation Holistic autonomous driving understanding by bird’s-eye-view injected multi-modal large models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T09:28:26.933883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:283e01d27deb4255ca54f0aa68774dc75c74d11cf744c84b7f227d13995fdac5

Observation f61d2ea3-f73d-4c45-acf4-5c9aec3bc003 · outbound

This paper cites Hauptmann, and Zhi-Qi Cheng.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation Hauptmann, and Zhi-Qi Cheng

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:36:26.067131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:e7a0b257912aa3b087407cc1c9e6594f574d3842230a5f52e2fed0145b13ecea

Observation 545c5e5d-ef61-4ce5-8580-1ecc29d9dd02 · outbound

This paper cites Language-conditioned world modeling for visual navigation.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation Language-conditioned world modeling for visual navigation

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:36:25.942926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:9947ae58ee400180efea58e7a7e3cea3ea782387746ff50a95f90d25ab9383fe

Observation 2780ccca-bd98-4ccc-8173-eb91a5d33c86 · outbound

This paper cites Large scale interactive motion forecasting for autonomous driving: The waymo open motion dataset.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation Large scale interactive motion forecasting for autonomous driving: The waymo open motion dataset

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T09:28:27.000828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:019372b007ff19d1eeae4ad52e07305cf9272c2ab978550d66d8bf3c9cf77084

Observation 54e96c37-84ae-46f7-a65d-ae63c24a258d · outbound

This paper cites Advancing sequential numerical prediction in autoregressive models.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation Advancing sequential numerical prediction in autoregressive models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T09:28:26.980942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:c197df84e969c5981eddfed3ee889f28dcf875951f300cbff129a5d877b89a7e

Observation f1ba7fa9-8fde-4ebd-a3d6-2d641ca8e39e · outbound

This paper cites Dolphin: Document image parsing via heterogeneous anchor prompting.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation Dolphin: Document image parsing via heterogeneous anchor prompting

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T09:28:26.977003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:5b14b477166c0c50738b1f42ffb52502979e30d3d6b3da0c213297a8d86e189d

Observation cc7d6ea5-f360-4445-afc6-a084108778b6 · outbound

This paper cites A Survey of World Models for Autonomous Driving.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation A Survey of World Models for Autonomous Driving

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:36:25.917007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:0a9c4d743b490c40087667fab14159a80acb45f151aa25407a0642298a030c22

Observation 236f4343-fad7-49f7-87d5-8d519b6f4cde · outbound

This paper cites Orion: A holistic end-to-end autonomous driving framework by vision-language instructed action generation.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation Orion: A holistic end-to-end autonomous driving framework by vision-language instructed action generation

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T09:28:27.031896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:b9342b304d61fe35b5e42ad5ce79b154ca869f974b2b918d36a2a30a196bc680

Observation e5a188c7-178f-4979-96f9-2b79aac3b7fc · outbound

This paper cites MindDrive: A Vision-Language-Action Model for Autonomous Driving via Online Reinforcement Learning.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation MindDrive: A Vision-Language-Action Model for Autonomous Driving via Online Reinforcement Learning

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-07-20T02:18:22.862599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:60a3898ca2745aff73ac3c25eafecbb28ed9acc58dd686a9c9f230974cbbe275

Observation 8198289c-3a26-4a1f-b16b-00bf42971b9e · outbound

This paper cites Vision meets robotics: The KITTI dataset.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation Vision meets robotics: The KITTI dataset

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T09:28:27.004571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:cc5c60e9e7122172309d056f5f9b77d870eaf02bba4713df0d3ec25a46dfba51

Observation 085ff846-a7ce-4fad-933e-735946f9a7ea · outbound

This paper cites Driving in Corner Case: A Real-World Adversarial Closed-Loop Evaluation Platform for End-to-End Autonomous Driving.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation Driving in Corner Case: A Real-World Adversarial Closed-Loop Evaluation Platform for End-to-End Autonomous Driving

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:36:26.025230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:98e8eae8288ab0729c99923f9bc102924a0af9c05e896e7d9549790e45338782

Observation ba9f214d-e395-417a-9c21-2f59262bfd37 · outbound

This paper cites Narasimhan.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation Narasimhan

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T09:28:27.035456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:dbee42c89ebc89aa99018206458f5c9b262f9c4dc75ad91a6d9978526c62b8ae

Observation 0deda81c-23b1-4b65-92e2-67a27e86957d · outbound

This paper cites Gemini 2.5 Pro preview: even better coding performance.https://developers.googleblog.com/en/ gemini-2-5-pro-io-improved-coding-performance.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation Gemini 2.5 Pro preview: even better coding performance.https://developers.googleblog.com/en/ gemini-2-5-pro-io-improved-coding-performance

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T09:28:26.972718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:689e7534b51c109a995c2d746123079b132cba334acb0ba1205a6247f4e8246d

Observation dcb3504d-5c72-4a43-9cf2-69b471620150 · outbound

This paper cites World models for autonomous driving: An initial survey.IEEE Transactionson Intelligent Vehicles, pages 1–17.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation World models for autonomous driving: An initial survey.IEEE Transactionson Intelligent Vehicles, pages 1–17

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T09:28:27.117978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:1661ec56458b4ece6b33b8a340f856c19c7afca76e03857796cd263b648cdfaa

Observation 982ffee2-ddec-4146-8e38-ea7f991d2168 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:36:26.045917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:6694c33eca2400efca20027c6ec8d9eaa84b0b469ae1aef7316107971ba15127

Observation 1298f9e2-02e3-458a-af07-8c4ef8c99a56 · outbound

This paper cites Seed1.5-VL Technical Report.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation Seed1.5-VL Technical Report

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:36:25.937007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:95c94ab8508f6cca98cc92c5389d5e9ba3f453bc0a7c6d46f3e72c019c02dfa3

Observation d06f0d02-d1db-45ea-b128-a8d00e9bcfcd · outbound

This paper cites World Models.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation World Models

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:36:25.974962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:c624147346a34f277930ee5b8aca3f3845ac9f0c8e50b9ef1b3fb23495f48d3e

Observation d69e091e-0f6d-4ec1-bd80-73ae172fc2ba · outbound

This paper cites Dream to Control: Learning Behaviors by Latent Imagination.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation Dream to Control: Learning Behaviors by Latent Imagination

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:36:26.020489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:019e7ff44d0a1d583fd4ae064452f2f2314f8be55a736e202d210989395a449a

Observation 6b8688fd-4d23-4e64-9556-f587ef618d05 · outbound

This paper cites Training Large Language Models to Reason in a Continuous Latent Space.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation Training Large Language Models to Reason in a Continuous Latent Space

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:36:26.108770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:6057b4974a7c1abf17add314e971c3c9ce1bec55c9912b078d9e41f6ec10d04f

Observation e8132dd7-4299-41ac-8c75-432621e774e7 · outbound

This paper cites MiMo-Embodied: X-Embodied Foundation Model Technical Report.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation MiMo-Embodied: X-Embodied Foundation Model Technical Report

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:36:26.119384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:8d69a3b33ec692f79948d0a307b37f5beed4e5902113b6bf37faf4e7d548e765

Observation 90ab0c2e-c23b-4c66-bf27-d0b1e8da0585 · outbound

This paper cites DriveMRP: Enhancing Vision-Language Models with Synthetic Motion Data for Motion Risk Prediction.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation DriveMRP: Enhancing Vision-Language Models with Synthetic Motion Data for Motion Risk Prediction

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:36:26.150807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:dc3c9dc24673dd5803541bc505da7cb146f7d73a83ffadf687545eaf8d57fe48

Observation f0455a59-ab08-4d7a-8771-7bc65ef14683 · outbound

This paper cites GAIA-1: A Generative World Model for Autonomous Driving.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation GAIA-1: A Generative World Model for Autonomous Driving

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:36:26.160483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:663d80fd9c3db7b2dd52be41b6f26e75ff8de8c382f6e243ccc7f7035841fb00

Observation add2c6a1-1e1d-426a-8654-d2e3493891c2 · outbound

This paper cites Vision-language-action models for autonomous driving: Past, present, and future.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation Vision-language-action models for autonomous driving: Past, present, and future

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:36:25.979896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:d329bbd2cf94ad44051af1911f4f045a6997f40a6be4ae27d48386c30551a8b2

Observation b47ebd11-2c33-49f0-b3c9-80650a543b56 · outbound

This paper cites NavThinker: Action-conditioned world models for coupled prediction and planning in social navigation.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation NavThinker: Action-conditioned world models for coupled prediction and planning in social navigation

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:36:25.989888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:23f4b2cffd30853d80e4f2afae9c16c110df41a67bf546aa4834e20baf8e5dad

Observation e9d6a343-9e18-47e9-9d52-25a48fc86a22 · outbound

This paper cites Fuller: Unified multi-modality multi-task 3D perception via multi-level gradient calibration.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation Fuller: Unified multi-modality multi-task 3D perception via multi-level gradient calibration

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T09:28:27.012473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:87f5cf9d240580d711660add40d5d59f69c8fe02e8e209b00d61cd2f48ff2f61

Observation 5e3ee279-4a83-439c-bbd2-d9bdaa1a2ab2 · outbound

This paper cites Making large language models better planners with reasoning-decision alignment.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation Making large language models better planners with reasoning-decision alignment

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T09:28:27.052531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:7141f7e66088930789058e8db6b8c3ab5e4ac0e5e17c77e3f823c134c960c93d

Observation cbeb9ec0-be71-4d1d-8835-4b1049b6e960 · outbound

This paper cites RoboTron-Drive: All-in-one large multimodal model for autonomous driving.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation RoboTron-Drive: All-in-one large multimodal model for autonomous driving

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T09:28:27.026950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:6af81894ac2fd87f229b4fd045fdd800d2f6fafec9bc53b5435c420fa59bb1ea

Observation 6927d97a-b8b8-4ca1-932c-ca107d61b1e0 · outbound

This paper cites GPT-4o System Card.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation GPT-4o System Card

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:36:25.984468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:d6480984da68558f07ce9ad8edcccab4031dd9d5d61fa462e15c6a43a4943229

Observation 722e832f-126e-4e72-896f-a079bb27268f · outbound

This paper cites DriveLMM-o1: A Step-by-Step Reasoning Dataset and Large Multimodal Model for Driving Scenario Understanding.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation DriveLMM-o1: A Step-by-Step Reasoning Dataset and Large Multimodal Model for Driving Scenario Understanding

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:36:26.000557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:7522fcd44d2d9036e7439f224c4eb424630e31d4e5e7c3ec8725e331aae85f3f

Observation 50bed8ca-bb2c-49d5-86e7-7465238af746 · outbound

This paper cites OpenAI o1 System Card.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation OpenAI o1 System Card

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:36:26.154997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:f9cc7c166ec1beacab4c7246b960933f00dc61fd1fa4854e43533d32f4efd11e

Observation f9ab1206-fe37-469a-b28a-96803dcd9022 · outbound

This paper cites Meml-grpo: Heterogeneous multi-expert mutual learning for rlvr advancement.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation Meml-grpo: Heterogeneous multi-expert mutual learning for rlvr advancement

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T09:28:27.068710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:85bb1dba325ba1b0e5ac4c5aaebff88df57697aa780c297ccb43855d169648a6

Observation 7f35206e-1553-408e-b54e-36717c420eb4 · outbound

This paper cites Towards learning- based planning: The nuPlan benchmark for real-world autonomous driving.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation Towards learning- based planning: The nuPlan benchmark for real-world autonomous driving

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T09:28:26.922400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:843b73253e82975dda6569ab8e6028e1643fe937d0cf807a549594f17ff7b45d

Observation 8f759bc3-d470-441e-9ab4-9a5b191998a6 · outbound

This paper cites The RoboDrive Challenge: Drive Anytime Anywhere in Any Condition.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation The RoboDrive Challenge: Drive Anytime Anywhere in Any Condition

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:36:26.010977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:da5c4a417671b1309677a724c1f043338ce438d318238f200ada24345ed2c736

Observation e9739e53-cc52-4c1f-8a49-a9d343ed6b9d · outbound

This paper cites Multi-modal data-efficient 3D scene understanding for autonomous driving.IEEE Transactions on Pattern Analysis and Machine Intelligence, 47(5):3748–3765.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation Multi-modal data-efficient 3D scene understanding for autonomous driving.IEEE Transactions on Pattern Analysis and Machine Intelligence, 47(5):3748–3765

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T09:28:26.918393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:3e55266297342a2f6fa3988281fea020e1e27c011bd5e30eb602e68f8efe458d

Observation 244fed89-f186-451a-a330-46bdeb84d43b · outbound

This paper cites 3D and 4D World Modeling: A Survey.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation 3D and 4D World Modeling: A Survey

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-07-21T03:22:29.733321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:75be77c3799755f2d272b5775e6403ea12594ecb6efbec328691456d9e62296f

Observation 9534465d-35cf-4059-a05c-1760ba4fe1f5 · outbound

This paper cites LargeAD: Large-scale cross-sensor data pretraining for autonomous driving.IEEE Transactionson Pattern Analysis and Machine Intelligence, 48(2):1291–1308.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation LargeAD: Large-scale cross-sensor data pretraining for autonomous driving.IEEE Transactionson Pattern Analysis and Machine Intelligence, 48(2):1291–1308

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T09:28:27.155138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:312c1ac624b088dceab4c9c94bcf5b00fcd04073dd3eeeae41b668cf9410ce7b

Observation 671dee57-d1f4-4b40-a601-7272c70ae337 · outbound

This paper cites Universal intelligence: A definition of machine intelligence.Minds and Machines, 17(4):391–444.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation Universal intelligence: A definition of machine intelligence.Minds and Machines, 17(4):391–444

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T09:28:27.048217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:1d66c94405845dbd9a7aa91b49dd06a8d0973a9fd506a6cef694469bbe77e01f

Observation 113837ec-0886-4220-9ede-9672b07a86b7 · outbound

This paper cites Enhancing end- to-end autonomous driving with latent world model.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation Enhancing end- to-end autonomous driving with latent world model

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T09:28:27.039632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:0cc4dc30fbc081d1b73b8c1ab641f0c913a0dc69828da7a2e78ec5b543b2c84c

Observation 59a1dff2-a1c9-4ce8-aa54-1f5aba3c953a · outbound

This paper cites End-to-end driving with online trajectory evaluation via BEV world model.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation End-to-end driving with online trajectory evaluation via BEV world model

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T09:28:27.023303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:0a89f5250453dd78838aee8823c58fa90d90d31d8eef013f0d60750b9b20c36a

Observation 99fe55d7-c1ef-4ccb-b1c4-456f6bc51bca · outbound

This paper cites DriveVLA-W0: World models amplify data scaling law in autonomous driving.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation DriveVLA-W0: World models amplify data scaling law in autonomous driving

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T09:28:27.139153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:954cbfa3dac6e21073557ad73c3531bf22b6053478506b12c99476442c4fdd77

Observation 1030d577-5b2a-4a88-b885-20bb80d20d79 · outbound

This paper cites ReCogDrive: A Reinforced Cognitive Framework for End-to-End Autonomous Driving.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation ReCogDrive: A Reinforced Cognitive Framework for End-to-End Autonomous Driving

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-15T07:36:24.555133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:a50cb47cda2db91113bb7aeaa936584fcd4d8f38738c6bed7e7f03b3f85933f2

Observation b9fde2de-5229-4987-b210-c34141b1e6a4 · outbound

This paper cites Lidarcrafter: Dynamic 4d world modeling from lidar sequences.Proceedings of the AAAI Conference on Artificial Intelligence, 40(22):18406–18414, Mar.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation Lidarcrafter: Dynamic 4d world modeling from lidar sequences.Proceedings of the AAAI Conference on Artificial Intelligence, 40(22):18406–18414, Mar

Reference 63

Resolution
verified exact
doi, observed 2026-05-12T00:56:13.429159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:8f5ecae197daebc272855ab018cfabb35768ddcfc787f612f34f20d0e3cca1c9

Observation 9b11874e-ae5d-42d1-8860-e69e1c4f8220 · outbound

This paper cites Cottereau, Changxin Gao, Liang Pan, Wei Tsang Ooi, and Ziwei Liu.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation Cottereau, Changxin Gao, Liang Pan, Wei Tsang Ooi, and Ziwei Liu

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T09:28:27.113849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:1735012ed40558f27ecf8b0f322729f95836dd47f76627dd8ed1bb8714d74e56

Observation 8abcd7b0-e6d1-45fa-93ae-13363e14d86e · outbound

This paper cites Let’s verify step by step.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation Let’s verify step by step

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T09:28:26.913878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:c87243f3675812ce8fa5870875924fafe9590d4fbdd269b1ab4cb92a6beb3bcb

Observation 23119ec1-5e0a-42ea-a692-29b98f360b55 · outbound

This paper cites Visual instruction tuning.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation Visual instruction tuning

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T09:28:26.988446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:24576eee6643731bbf563b2d9c9b0deb53933966c5a795584a1df240b9e06237

Observation 5018f948-4b49-487f-978e-7e0315496a61 · outbound

This paper cites Improved baselines with visual instruction tuning.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation Improved baselines with visual instruction tuning

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T09:28:26.937629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:74b730d77c8d9a95da54522dc3ba6201e8c012f073ac2b6dfdb9e76ed13fa888

Observation 5520fd27-c576-483a-b886-bd67562cff5a · outbound

This paper cites Guideflow: Constraint-guided flow matching for planning in end-to-end autonomous driving.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation Guideflow: Constraint-guided flow matching for planning in end-to-end autonomous driving

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:36:26.056267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:c26e9885958ab40f8d7d4f3270afb95c566bf3633e58574e28d50bf9ecfde910

Observation 460b551d-9731-412b-8413-a7f793b171f6 · outbound

This paper cites Driveworld-vla: Unified latent-space world modeling with vision-language-action for au- tonomous driving.ArXiv, abs/2602.06521.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation Driveworld-vla: Unified latent-space world modeling with vision-language-action for au- tonomous driving.ArXiv, abs/2602.06521

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:36:25.923214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:c9970495e5a6279201997fee8e8de9572fa231ff2fe160a4c18d07ba232d7212

Observation 71739495-92fe-4d76-926f-37b578189c94 · outbound

This paper cites ReasonPlan: Unified scene prediction and decision reasoning for closed-loop autonomous driving.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation ReasonPlan: Unified scene prediction and decision reasoning for closed-loop autonomous driving

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T09:28:26.926123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:d6e16a9e92948ab0f03b8e0a92c77c8a6cc6491209bb0de5dc9db918f597129c

Observation 16f167f0-0d6f-4e38-bdbf-8ec01025f4af · outbound

This paper cites A rationale-centric framework for human-in-the-loop machine learning.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation A rationale-centric framework for human-in-the-loop machine learning

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T09:28:27.105989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:5b70417dde3327c355a91a22707ea65341e6e73c493ec8d25a3115145e8f966d

Observation b6af7b92-b15e-4d71-9c87-0214aa6f91fe · outbound

This paper cites Punifiedner: a prompting-based unified ner system for diverse datasets.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation Punifiedner: a prompting-based unified ner system for diverse datasets

Reference 72

Resolution
verified exact
doi, observed 2026-05-12T00:56:13.432442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:fae5d67f7f97d27e0ad915b77bdc8ddd60e842fd977a9a4bf490f5b24acb8322

Observation 3005b26f-07ac-4b60-ae7c-c22cf75f3a93 · outbound

This paper cites What makes pre-trained language models better zero-shot learners? InAnnual Meeting of the Association for Computational Linguistics, pages 2288–2303.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation What makes pre-trained language models better zero-shot learners? InAnnual Meeting of the Association for Computational Linguistics, pages 2288–2303

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T09:28:27.097798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:587c88d3dd9f23ff40790103b2c7178fbcd3496e446d443b8f0d1b3ec89f44bf

Observation 1cf21f5a-3921-4256-b9a6-abd6b2426cb8 · outbound

This paper cites PaDeLLM-NER: Parallel decoding in large language models for named entity recognition.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation PaDeLLM-NER: Parallel decoding in large language models for named entity recognition

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T09:28:27.093495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:70a424e7891da8b93eb8856de05af6763f9f2811d2c3216d75f6c2bd3a86079d

Observation 5be4e1d5-51ab-4b3a-b53e-481f8ef7ab7e · outbound

This paper cites A bounding box is worth one token - interleaving layout and text in a large language model for document understanding.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation A bounding box is worth one token - interleaving layout and text in a large language model for document understanding

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T09:28:27.085499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:86864bb4826952ba2511f8a3400762b82b151562cd5332ae4ba812bedfc3ef43

Observation 6fb73279-136c-4dbe-acc9-dcd487444f67 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T09:28:26.984854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:7e64f30afe0a4b5cc61e5e4a960e9a3f15abff9c3585de6c6fb4e415da976ae9

Observation 3565fb29-e396-4de8-9929-aa45e3dad758 · outbound

This paper cites Last-vla: Thinking in latent spatio-temporal space for vision-language-action in autonomous driving.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation Last-vla: Thinking in latent spatio-temporal space for vision-language-action in autonomous driving

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:36:25.903632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:18f383a42c97736f3cd87632df5cffa7f0a9471197478932978d1422ea1f6eff

Observation 391d1782-8dd5-487c-adfc-baa6fd312bca · outbound

This paper cites Adathinkdrive: Adaptive thinking via reinforcement learning for autonomous driving.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation Adathinkdrive: Adaptive thinking via reinforcement learning for autonomous driving

Reference 78

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:36:25.910801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:10f6c6908752a46ba33243542f0e9d03a5ac70c6319f58db63e535b5123e8092

Observation 291a81bb-1f12-4e16-b88e-edb540f16845 · outbound

This paper cites Unleashing vla potentials in autonomous driving via explicit learning from failures.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation Unleashing vla potentials in autonomous driving via explicit learning from failures

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:36:26.051200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:bee02f818ed6f4d1c687e757e3669c2725bf8f4f9c3d4277dc474d2af590378a

Observation 98cf6bc8-3244-4b8e-8c2d-f27aecb19e8f · outbound

This paper cites MTRDrive: Memory-tool synergistic reasoning for robust autonomous driving in corner cases.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation MTRDrive: Memory-tool synergistic reasoning for robust autonomous driving in corner cases

Reference 80

Resolution
verified exact
arxiv_id, observed 2026-05-12T00:56:14.430355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:a22ad01948e9f18cb5a6f6786f1b33f144a5a7dcef123c8632351c95592c77eb

Observation 733ad1f3-1b83-4b2f-9dfe-252b148f3a1f · outbound

This paper cites DRAMA: Joint risk localization and captioning in driving.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation DRAMA: Joint risk localization and captioning in driving

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T09:28:27.076819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:cc7c5057cf4a6f3969d43281d6803a5cafce5efa315c55c40c53699e191f0143

Observation b47adb1f-d950-4cc1-928a-956fd29128e1 · outbound

This paper cites One Million Scenes for Autonomous Driving: ONCE Dataset.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation One Million Scenes for Autonomous Driving: ONCE Dataset

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:36:25.931211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:3f7a8752dacb658167b603b766886f98fa41d4764a229f325eda654f601d34f1

Observation 600053ac-8981-4a6c-a56d-15215dc29816 · outbound

This paper cites LingoQA: Visual question answering for autonomous driving.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation LingoQA: Visual question answering for autonomous driving

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T09:28:27.008416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:24c96df61bc55de8b40adcb2958bd5b778c448b8739607c02242205ad7c6853f

Observation 214dbcf6-5d3f-462c-a6a9-3e4b4f8afef3 · outbound

This paper cites The Mapillary Vistas dataset for semantic understanding of street scenes.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation The Mapillary Vistas dataset for semantic understanding of street scenes

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T09:28:27.150957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:ef63ccb1bf0a8f4913326d22de8210aa8b6a9cf40cc2d77dd6d40b7fc36136df

Observation 3663ab44-5b14-4ae8-810a-d62c23301983 · outbound

This paper cites DynVLA: Learning world dynamics for action reasoning in autonomous driving.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation DynVLA: Learning world dynamics for action reasoning in autonomous driving

Reference 85

Resolution
verified exact
arxiv_id, observed 2026-05-12T00:56:14.434052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:2c404bccaa8373e968dd3b6f5d64918bc1261395e08b1df1ec442c1a231cb495

Observation d1d00094-37b6-426c-b03e-a1f5e829b847 · outbound

This paper cites CODI: Compressing chain-of- thought into continuous space via self-distillation.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation CODI: Compressing chain-of- thought into continuous space via self-distillation

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T09:28:27.109975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:4ff267370511c176f4f936a9e36d4a349e627576c348315870ab8ce1063ddf1b

Observation f8304efd-5119-48ac-b378-b8122ecca706 · outbound

This paper cites Scalableimagetokenizationwith index backpropagation quantization.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation Scalableimagetokenizationwith index backpropagation quantization

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T09:28:27.064985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:543266e88c07d759ff96a3170aa1a2f24a63dfe827101f22e7fc473605af0358

Observation 6f0abdb4-c447-443b-aab4-e40daedde74a · outbound

This paper cites DriveLM: Driving with graph visual question answering.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation DriveLM: Driving with graph visual question answering

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T09:28:27.043786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:0f21dd935621fe35dbe15700c9032cc74dfce29b1c1f5752522fd7f8ebadc329

Observation 0a79c90d-afd3-4cab-ba58-fcd97b15a62b · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 89

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:36:26.093692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:8b5c088a6e57264143648c31d5cd2730b1656766bd503ac02847378eee3e2f56

Observation 39c8e822-17f8-415c-a770-5d479634be57 · outbound

This paper cites Token Assorted: Mixing Latent and Text Tokens for Improved Language Model Reasoning.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation Token Assorted: Mixing Latent and Text Tokens for Improved Language Model Reasoning

Reference 90

Resolution
verified exact
arxiv_id, observed 2026-05-12T00:56:14.437815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:2ceef8843aac2268482b17850cc2b84d17eb4c4a5baea97dbec9fdb468d66b57

Observation 5f8407aa-2f0d-434f-98f2-8281dd1c6aaf · outbound

This paper cites MTVQA: Benchmarking multilingual text-centric visual question answering.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation MTVQA: Benchmarking multilingual text-centric visual question answering

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T09:28:27.072718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:e13dfb7b29c0c27d34b5c81a593d6f956bc61f6669bfa7b646dd31ac263ac306

Observation 1bbfb0b5-56c6-4d27-9c4a-d245529bf937 · outbound

This paper cites What Drives Success in Physical Planning with Joint-Embedding Predictive World Models?.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation What Drives Success in Physical Planning with Joint-Embedding Predictive World Models?

Reference 92

Resolution
verified exact
arxiv_id, observed 2026-05-20T00:04:20.071030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:a2e4bda6beb039c0fded40b6f1b51924619d5b3b1cfc426166a49bf684b01d7e

Observation 17a0933a-0ce5-4828-ae43-65e7eeeef114 · outbound

This paper cites SimScale: Learning to Drive via Real-World Simulation at Scale.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation SimScale: Learning to Drive via Real-World Simulation at Scale

Reference 93

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:36:26.062231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:47e238d769e3f3b22d19b3641450d4a086229925fba81fe517fbf4355fe25e42

Observation 067b5e6b-a2b7-4d3d-874a-9372da15e125 · outbound

This paper cites Deep learning and the information bottleneck principle.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation Deep learning and the information bottleneck principle

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T09:28:27.056894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:d9cefc9ac52dd58f1a2152e9710f08df7ac8fb713f96454ea0d06d08accdadff

Observation e8a6840c-f5d1-4add-9ac5-4920faf1c7bb · outbound

This paper cites The information bottleneck method.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation The information bottleneck method

Reference 95

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T08:36:25.947345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:74d946ab3790f23da5a7c97c61b7e7cb161ad881d8587816422913d1b76a926d

Observation 951ca6c9-148f-41f2-8ae5-843ccaf433da · outbound

This paper cites Cambrian-1: A fully open, vision-centric exploration of multimodal LLMs.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation Cambrian-1: A fully open, vision-centric exploration of multimodal LLMs

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T09:28:27.143052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:09def81865bf0f026152ce534cf20d7ed5faad1b869b6a9d845ac52150a6902d

Observation b755d325-b7f4-46b6-9a6e-32772f023871 · outbound

This paper cites IDD: A dataset for exploring problems of autonomous navigation in unconstrained environments.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation IDD: A dataset for exploring problems of autonomous navigation in unconstrained environments

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T09:28:27.089472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:11eeee8a17b22fdabb92d64c416dcd2d0b99977fe39c932e1dbe68da1e0f4122

Observation ed3c0551-9ba4-47ff-b39b-92fc18c2c410 · outbound

This paper cites Attention is all you need.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation Attention is all you need

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T09:28:26.940888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:eaf85860491fb30c6635acf6989e7b8936cd810ee6d7ade3fa8c5238f3e889f8

Observation 900348eb-cf57-4efa-927b-84fc5448bae2 · outbound

This paper cites Vision as LoRA.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation Vision as LoRA

Reference 99

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:36:26.103976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:970ad728b016409ee5fe57e551137922b481b3029a9e4d8191bd7514638a338d

Observation 81cfff75-923f-4a2c-914e-a7cdeca2dc19 · outbound

This paper cites VGGT: Visual geometry grounded transformer.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation VGGT: Visual geometry grounded transformer

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T09:28:27.016154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:44ae80ef4c06bc0566074f3508b4583c54745169295d36e1b78997a74ce245ad

Pith citing papers

Observation 331997d7-08e6-4ca7-b2bb-4b05796870be · inbound

Is Your Driving World Model an All-Around Player? cites this paper.

Is Your Driving World Model an All-Around Player? Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-12T06:01:24.008567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:41:50.292305Z digest=sha256:c4ac501bea8433a3e3c9db9bc750ce1c98990c2a7da2655b6b290b184299c6a1

Observation 4e6f403d-8549-477d-ac75-a963e1c621a3 · inbound

OmniLiDAR: A Unified Diffusion Framework for Multi-Domain 3D LiDAR Generation cites this paper.

OmniLiDAR: A Unified Diffusion Framework for Multi-Domain 3D LiDAR Generation Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-14T19:22:50.653700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T19:20:15.333859Z digest=sha256:cf48a50f77f40477d2967a07989cfe1800a595358e49085d26863cfee15a3b7e

Observation 86600751-4c2f-497e-9562-608bea60a07d · inbound

Beyond Imitation: Learning Safe End-to-End Autonomous Driving from Hard Negatives cites this paper.

Beyond Imitation: Learning Safe End-to-End Autonomous Driving from Hard Negatives Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation

Reference 31

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T05:33:04.352441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T05:28:21.056792Z digest=sha256:af274a0b2672a6b1e1812c63b4c95f38d0bfbaf71283de3d664b0340d960b445

Observation 0f008b2a-8686-49db-9bf1-fe950c51d477 · inbound

OneVLA: A Unified Framework for Embodied Tasks cites this paper.

OneVLA: A Unified Framework for Embodied Tasks Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-07-01T21:26:14.122300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T17:05:08.124096Z digest=sha256:e32bf69357d779c0b8f939e38ba66c2dd04d76579937e41e103d5f64a7e9cd5c

Observation 9eab05e9-a088-4b20-8431-84405cf6c45c · inbound

WALL-WM: Carving World Action Modeling at the Event Joints cites this paper.

WALL-WM: Carving World Action Modeling at the Event Joints Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-07-01T23:36:22.326219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T14:15:14.454649Z digest=sha256:ef801b89755b834b4927f492506e7829c1aa4ff7a0df41383e9c7556c17621c8

Observation 8ac20d96-80fe-467f-a21d-417dbc8bd2eb · inbound

Imagine Before You Predict: Interleaved Latent Visual Reasoning for Video Event Prediction cites this paper.

Imagine Before You Predict: Interleaved Latent Visual Reasoning for Video Event Prediction Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation

Reference 69

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T11:56:55.429542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-28T02:46:50.373450Z digest=sha256:052c72714708750fb9ecb432909c0a9583ebb058556aaf6cd5908ed0753d397c

Observation e1259f2f-b0cd-40ce-9d3d-f39e2bf1ad4b · inbound

Dash2Sim: Closed-Loop Driving Simulation from in-the-wild Dashcam Videos cites this paper.

Dash2Sim: Closed-Loop Driving Simulation from in-the-wild Dashcam Videos Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-07-02T16:57:10.321628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T22:15:08.063308Z digest=sha256:858131149715df1ccbc145c1fc8f2401253c67c9361c43b39ff70e1686764ec0

Observation ddd72e8b-ce8b-4d34-a859-27ce4e0d5e0e · inbound

FutureNav: Unified World-Action Modeling for Vision-and-Language Navigation cites this paper.

FutureNav: Unified World-Action Modeling for Vision-and-Language Navigation Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-06-30T15:44:49.434536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T05:05:56.065261Z digest=sha256:80e58fb9caed76bf052a009aaedd4d353aa7d70e77d75487360b11681fc19cb1

Observation 7eb48b8a-c4c7-4953-925b-e620c9234984 · inbound

DriveVer: Lightweight Trajectory Evaluator as Test-Time Verifier for Autonomous Driving cites this paper.

DriveVer: Lightweight Trajectory Evaluator as Test-Time Verifier for Autonomous Driving Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-02T15:07:04.053954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-02T14:59:35.226245Z digest=sha256:b779693dd297ad7865ba4478d10a01cb9f4652bb01835e5e20796857beb9ddf1

Observation 09405c93-1edd-42bc-b8c0-98cd78597ea7 · inbound

Token-Wise Latent Streaming from Slow Reasoners to Fast Planners for Dynamic Vision Language Navigation cites this paper.

Token-Wise Latent Streaming from Slow Reasoners to Fast Planners for Dynamic Vision Language Navigation Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-01T19:56:57.482174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T19:56:57.482174Z digest=sha256:6f0b4323c892b8ec70ad81a4226cccd337cb3a2daeb85b24640fdc3b98005b61

Observation 90980c78-0455-40f7-b57c-a88828014828 · inbound

HyWorldVLA: A Vision-Language-Action Model with Hybrid World Modeling for Autonomous Driving cites this paper.

HyWorldVLA: A Vision-Language-Action Model with Hybrid World Modeling for Autonomous Driving Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T08:49:48.309380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:49:48.309380Z digest=sha256:0b8d9d745ad92a58da552607ea9b749e0dc4b7d70363ca80d2994f561b350a5f

Observation 23dbe812-28cd-4331-b7eb-82aea1231cb7 · inbound

Data Pyramid for Embodied Manipulation cites this paper.

Data Pyramid for Embodied Manipulation Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation

Reference 255

Resolution
unresolved
no resolver link, observed 2026-07-31T06:18:55.846788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:18:55.846788Z digest=sha256:aaaeac46b6ff1bbd04bdc16ce0d1d2693ffe305f34923cbc89057ea392488094