Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T18:37:46.441294Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 1 inbound Pith citation observation for arXiv:2507.07818.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T18:37:46.441294Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-28T06:02:52.244884Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-02T08:26:48.315943Z
44 of 44 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation dd9b177f-34be-43b8-89c5-4f0e9559916c · outbound
MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41e8ba27-e195-4415-a321-bae0fcf75039 · outbound
MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 226d1cf1-0197-4c11-93c2-7c9e1ece8f0c · outbound
MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Stable LM 2 1.6B Technical Report
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 035bce93-1f69-4de4-9c7f-0245ceb4ee3c · outbound
MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines nuScenes: A multimodal dataset for autonomous driving
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f46c68ad-b279-4d63-8b96-9937854540a5 · outbound
MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f75bcb01-631b-40ea-904d-0508db5d7e0d · outbound
MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Automated Evaluation of Large Vision-Language Models on Self-driving Corner Cases
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37f2feec-768b-4037-bf41-91debb223555 · outbound
MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Driving with llms: Fusing object-level vector modality for explainable autonomous driving
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 17a83839-8f01-40a8-ac31-e33322cd34e4 · outbound
MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines AdaMV-MoE : Adaptive multi-task vision mixture-of-experts
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3387ea54-5524-4ebf-be5e-7f6bcbaaea40 · outbound
MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bcd5ea30-12e8-45de-8fbc-928b722b9736 · outbound
MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Qwen-vl-max: A high-performance vision-language model, 2024
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ceb72f98-e359-465e-8e4e-5268d0b94383 · outbound
MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Unresolved cited work
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13687786-8b02-4c1c-9bed-2131bd069e15 · outbound
MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1817b2e-acbc-4aea-97be-fdec8af77a8f · outbound
MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1fbd3099-efba-4e77-aba1-8223580de194 · outbound
MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37dff330-5c94-4656-9c29-e46c03b36b6c · outbound
MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines From regional to general: A vision-language model-based framework for corner cases comprehension in autonomous driving
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c39b43f7-553d-4fa8-b882-a8ce5af221a1 · outbound
MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Planning-oriented autonomous driving
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 208520fc-ef92-4241-8d61-1fd602e8cc42 · outbound
MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines RoboTron-Drive: All-in-One Large Multimodal Model for Autonomous Driving
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a50447ba-ee14-44df-8cae-aa3de68d0a6a · outbound
MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Making large language models better planners with reasoning-decision alignment, 2024 b
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 705c16b4-2634-4fa7-834e-4c878261629f · outbound
MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Unresolved cited work
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 08425c3b-cea7-4a62-89f6-060f4486ea08 · outbound
MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Med-MoE: Mixture of Domain-Specific Experts for Lightweight Medical Vision-Language Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15c83ebb-9597-477f-9ab8-140bf0a5e143 · outbound
MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3024295-eb79-4046-bdc1-e521673d8a7e · outbound
MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines LLaVA-OneVision: Easy Visual Task Transfer
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08b1fab0-830a-4e8d-b47b-88599f0f41cb · outbound
MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a49f957b-7b22-46b0-8cdf-7bf838e74df8 · outbound
MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines MoE-LLaVA: Mixture of Experts for Large Vision-Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c2e9f44-4840-4713-a710-4540a1339876 · outbound
MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Moma: Efficient early-fusion pre-training with mixture of modality-aware experts, 2024 b
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e38584df-0703-4f92-bb3a-649120cbda23 · outbound
MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Improved baselines with visual instruction tuning, 2023 a
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8613ffa9-7c0a-45e5-856e-b692bbb3df22 · outbound
MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Visual instruction tuning, 2023 b
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54b403ca-5d3b-4914-8e15-9e01f8d2f9bd · outbound
MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Llava-next: Improved reasoning, ocr, and world knowledge, 2024
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7525c862-a7e6-4949-b811-0db14d56237f · outbound
MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines GPT-Driver: Learning to Drive with GPT
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8f031f5-16fd-4c46-9e18-a20906993e08 · outbound
MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Gpt-4v system card
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b4a901a6-6a8b-4ad1-896b-39bd0ebdaf35 · outbound
MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines LanguageMPC: Large Language Models as Decision Makers for Autonomous Driving
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e9f60ed-40dd-46b7-a0a6-c7f113b717c8 · outbound
MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Lmdrive: Closed-loop end-to-end driving with large language models
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation bf62d979-6ff7-4c40-8fc3-0f5d9bc1b76f · outbound
MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21e183e7-d9bf-463b-b53d-feccc26bf41e · outbound
MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines DriveLM: Driving with Graph Visual Question Answering
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23b76a4e-41e6-4102-909b-d05e0d7b7417 · outbound
MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b07d8080-d6a9-4313-9b60-0fe51ef14d6f · outbound
MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines OmniDrive: A Holistic Vision-Language Dataset for Autonomous Driving with Counterfactual Reasoning
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab119678-ae2b-4583-be42-ab3add7e1a52 · outbound
MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines On the Road with GPT-4V(ision): Early Explorations of Visual-Language Model on Autonomous Driving
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d789d401-3624-4b30-a015-00980ab687e3 · outbound
MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Two-stage lvlm system: 1st place solution for eccv 2024 corner case scene understanding challenge
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6d778d01-a12e-4015-a18a-b2fe06d7fc4e · outbound
MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2be8034-3639-4772-9c79-34d983dd8a63 · outbound
MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Flex-moe: Modeling arbitrary modality combination via the flexible mixture-of-experts
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e0d118b6-0f14-48c4-b74d-3d9dad7c8798 · outbound
MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Sigmoid loss for language image pre-training
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f813ce0-2054-427d-b11b-70dab2f5dd13 · outbound
MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines MiniDrive: More Efficient Vision-Language Models with Multi-Level 2D Features as Text Tokens for Autonomous Driving
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e04c8bf4-71a5-4603-8a24-c4349f920eec · outbound
MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Diversifying the expert knowledge for task-agnostic pruning in sparse mixture-of-experts, 2024 b
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c3dd1e9b-0d51-4115-8e98-174d7f77f9a8 · outbound
MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines ST-MoE: Designing Stable and Transferable Sparse Expert Models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c0db8a3-3f58-49a1-8d04-192b6c2c2481 · inbound
D$^3$-MoE:Dual Disentangled Diffusion Mixture-of-Experts for Style-Controllable End-to-End Autonomous Driving MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.