Pith. sign in

Paper Citation Record · LEDGER

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines

As of 7 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 1 inbound Pith citation observation for arXiv:2507.07818.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.07818 v2

Coverage vector

measured 44 of 44 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:37:46.441294Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-28T06:02:52.244884Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T08:26:48.315943Z

Reference resolution

44 of 44 outbound references displayed

  • verified exact0
  • verified fuzzy13
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation dd9b177f-34be-43b8-89c5-4f0e9559916c · outbound

This paper cites write newline.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:43.006424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:43.006424Z digest=sha256:597f33081f7effd2c307d2f3327e617b2ccfdd22f09ffea68c37af73606b4829

Observation 41e8ba27-e195-4415-a321-bae0fcf75039 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:43.131103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:43.131103Z digest=sha256:38e2680ca30a802f12d534d4060d1354e125fdc5957ffc5603f3aa9d39a7a978

Observation 226d1cf1-0197-4c11-93c2-7c9e1ece8f0c · outbound

This paper cites Stable LM 2 1.6B Technical Report.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Stable LM 2 1.6B Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:43.188523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:43.188523Z digest=sha256:80ed9537de6bfe63b10cc7c6f67cee764a4419de3f7b20d0c670cd29a5db91e6

Observation 035bce93-1f69-4de4-9c7f-0245ceb4ee3c · outbound

This paper cites nuScenes: A multimodal dataset for autonomous driving.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines nuScenes: A multimodal dataset for autonomous driving

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:43.254429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:43.254429Z digest=sha256:4a8833f2b918554629816558d665856b9726c11c9da4a9bb996b3bf27dab2579

Observation f46c68ad-b279-4d63-8b96-9937854540a5 · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:43.370613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:43.370613Z digest=sha256:3261d936b9c6bb473a00b14055bf7e931280f512dcfaf24bd372caa39d322814

Observation f75bcb01-631b-40ea-904d-0508db5d7e0d · outbound

This paper cites Automated Evaluation of Large Vision-Language Models on Self-driving Corner Cases.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Automated Evaluation of Large Vision-Language Models on Self-driving Corner Cases

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:43.421592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:43.421592Z digest=sha256:835932e6abffc64f016325fb9a409f755872ae0a35feab814426f04a1d0866ec

Observation 37f2feec-768b-4037-bf41-91debb223555 · outbound

This paper cites Driving with llms: Fusing object-level vector modality for explainable autonomous driving.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Driving with llms: Fusing object-level vector modality for explainable autonomous driving

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:37:50.409725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:37:43.493213Z digest=sha256:6a94af3ae754649398ccc4c1d71688d0df2d10ce88f16309da70ad9d56d5cd52

Observation 17a83839-8f01-40a8-ac31-e33322cd34e4 · outbound

This paper cites AdaMV-MoE : Adaptive multi-task vision mixture-of-experts.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines AdaMV-MoE : Adaptive multi-task vision mixture-of-experts

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:37:50.091903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:37:43.577182Z digest=sha256:fadce79328f61a0e62db39b4b6fdb67628209e98e92d8b1685207fbb4b259bcd

Observation 3387ea54-5524-4ebf-be5e-7f6bcbaaea40 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:43.654779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:43.654779Z digest=sha256:e4dffa2091b61dc2a2d55e55d6df7101762a72b5c5d9019cb6095430254080ea

Observation bcd5ea30-12e8-45de-8fbc-928b722b9736 · outbound

This paper cites Qwen-vl-max: A high-performance vision-language model, 2024.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Qwen-vl-max: A high-performance vision-language model, 2024

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:37:49.760319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:37:43.746929Z digest=sha256:80d22625a685041d7831bcf49b1171423c212947f1e8db66b0c3d4b1dec38052

Observation ceb72f98-e359-465e-8e4e-5268d0b94383 · outbound

This paper cites an unresolved cited work.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:43.824536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:43.824536Z digest=sha256:9bb16c629ebb34782eac7d12814fe0689e758ed9e39984276cac6cc6081f260d

Observation 13687786-8b02-4c1c-9bed-2131bd069e15 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:43.912623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:43.912623Z digest=sha256:819a62e9bf4e830955711b4f99a26af39a77433e7cb9d278c0b4621693f84bac

Observation f1817b2e-acbc-4aea-97be-fdec8af77a8f · outbound

This paper cites Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:44.005943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:44.005943Z digest=sha256:caf6deb8a85ac06988e483663d6020869e83e249821d03aab5de01d600373966

Observation 1fbd3099-efba-4e77-aba1-8223580de194 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:44.073982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:44.073982Z digest=sha256:c26e5f933f7f93695912d21b15c185c8beeb9b88ed9e47aa30068c7957426fb2

Observation 37dff330-5c94-4656-9c29-e46c03b36b6c · outbound

This paper cites From regional to general: A vision-language model-based framework for corner cases comprehension in autonomous driving.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines From regional to general: A vision-language model-based framework for corner cases comprehension in autonomous driving

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:37:49.466464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:37:44.175293Z digest=sha256:7982f2205eb227adfd6d174d9876ba6c6168927bab1a7df0036436edb642eae4

Observation c39b43f7-553d-4fa8-b882-a8ce5af221a1 · outbound

This paper cites Planning-oriented autonomous driving.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Planning-oriented autonomous driving

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:37:48.849077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:37:44.227270Z digest=sha256:1b793d369a61643044c4d0392484a0a5d015ca19e184d1872f38e6278396c182

Observation 208520fc-ef92-4241-8d61-1fd602e8cc42 · outbound

This paper cites RoboTron-Drive: All-in-One Large Multimodal Model for Autonomous Driving.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines RoboTron-Drive: All-in-One Large Multimodal Model for Autonomous Driving

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:44.330575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:44.330575Z digest=sha256:97e1e7bdc1e04df739af2587b35b68eaf81a12fe0a23c45f3055003d6ae2af62

Observation a50447ba-ee14-44df-8cae-aa3de68d0a6a · outbound

This paper cites Making large language models better planners with reasoning-decision alignment, 2024 b.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Making large language models better planners with reasoning-decision alignment, 2024 b

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:37:48.360150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:37:44.391895Z digest=sha256:d982116c8613d486291440c1e92e89923f59e6db14195f7f23d1c2eaf3c056e8

Observation 705c16b4-2634-4fa7-834e-4c878261629f · outbound

This paper cites an unresolved cited work.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:37:48.176300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:37:44.429974Z digest=sha256:61171d3aca6e64ddc76ca9722904693bc7bca34dfed9f76179d2ba5f84a131ff

Observation 08425c3b-cea7-4a62-89f6-060f4486ea08 · outbound

This paper cites Med-MoE: Mixture of Domain-Specific Experts for Lightweight Medical Vision-Language Models.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Med-MoE: Mixture of Domain-Specific Experts for Lightweight Medical Vision-Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:44.507398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:44.507398Z digest=sha256:be3cc1419cc9609ff09d1ad4216aa41de3512a5c410ce3c00b970af5278ef938

Observation 15c83ebb-9597-477f-9ab8-140bf0a5e143 · outbound

This paper cites GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:44.567764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:44.567764Z digest=sha256:5e9b2221e2deab7cfd22fe8d5c6845ef01916a3c152f834d9714b53c20e66ada

Observation f3024295-eb79-4046-bdc1-e521673d8a7e · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines LLaVA-OneVision: Easy Visual Task Transfer

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:44.668987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:44.668987Z digest=sha256:c6b023c7b2e38d24b31982ed89c68753f946cbb2c7d7c90ddc4cfdb1545a9599

Observation 08b1fab0-830a-4e8d-b47b-88599f0f41cb · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:37:48.030031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:37:44.737427Z digest=sha256:001977ed3c06782d4c0e89ba3cfd172a4f96728575e98437c3613823266f16ca

Observation a49f957b-7b22-46b0-8cdf-7bf838e74df8 · outbound

This paper cites MoE-LLaVA: Mixture of Experts for Large Vision-Language Models.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines MoE-LLaVA: Mixture of Experts for Large Vision-Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:44.808524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:44.808524Z digest=sha256:38f474d601861b52ba60078f79ec99547eb5fbfdb0e1a839b511e28310ab312d

Observation 2c2e9f44-4840-4713-a710-4540a1339876 · outbound

This paper cites Moma: Efficient early-fusion pre-training with mixture of modality-aware experts, 2024 b.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Moma: Efficient early-fusion pre-training with mixture of modality-aware experts, 2024 b

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:37:47.808005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:37:44.865125Z digest=sha256:ea20fa582588084975a4e5c26f6f42e2dc195456434710a88ed98827098adc4f

Observation e38584df-0703-4f92-bb3a-649120cbda23 · outbound

This paper cites Improved baselines with visual instruction tuning, 2023 a.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Improved baselines with visual instruction tuning, 2023 a

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:44.972759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:44.972759Z digest=sha256:cdad8ed813853cd1482d38c53de751c259aa6170dd368e4ee3e34bb19ee56e5c

Observation 8613ffa9-7c0a-45e5-856e-b692bbb3df22 · outbound

This paper cites Visual instruction tuning, 2023 b.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Visual instruction tuning, 2023 b

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:45.037569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:45.037569Z digest=sha256:6e79cf31b7e7a461796604b3f5fd1fcca529affde2e94eaa4180e9dc6b4955bb

Observation 54b403ca-5d3b-4914-8e15-9e01f8d2f9bd · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge, 2024.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Llava-next: Improved reasoning, ocr, and world knowledge, 2024

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:45.119217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:45.119217Z digest=sha256:76f7787870633f30c74a9175ff6061e4b6cb39266b7b0dd2500b1f90fed1de1e

Observation 7525c862-a7e6-4949-b811-0db14d56237f · outbound

This paper cites GPT-Driver: Learning to Drive with GPT.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines GPT-Driver: Learning to Drive with GPT

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:45.201566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:45.201566Z digest=sha256:15f9112b70b6da535499f88fe09d06ddbaa8ada64a6137193fa2de191fc776cd

Observation c8f031f5-16fd-4c46-9e18-a20906993e08 · outbound

This paper cites Gpt-4v system card.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Gpt-4v system card

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:37:47.608696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:37:45.275342Z digest=sha256:8bf2b94ddae8665b4e909af5797b3e529e0321c4e3ad14191975d67547c75378

Observation b4a901a6-6a8b-4ad1-896b-39bd0ebdaf35 · outbound

This paper cites LanguageMPC: Large Language Models as Decision Makers for Autonomous Driving.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines LanguageMPC: Large Language Models as Decision Makers for Autonomous Driving

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:45.330042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:45.330042Z digest=sha256:4da050420a2d99dbf496a08440b73add774fc1127318e4279bf3c064f0d780d3

Observation 8e9f60ed-40dd-46b7-a0a6-c7f113b717c8 · outbound

This paper cites Lmdrive: Closed-loop end-to-end driving with large language models.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Lmdrive: Closed-loop end-to-end driving with large language models

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:37:47.428519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:37:45.416140Z digest=sha256:e8c09041712878048265b6a40bda043a0d2e5f17abe83a366c2bb1e75c5f5353

Observation bf62d979-6ff7-4c40-8fc3-0f5d9bc1b76f · outbound

This paper cites Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:45.491881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:45.491881Z digest=sha256:23f186e479bd394a35831295a30980afaf1938486f6715ac36d6e678b69afcd6

Observation 21e183e7-d9bf-463b-b53d-feccc26bf41e · outbound

This paper cites DriveLM: Driving with Graph Visual Question Answering.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines DriveLM: Driving with Graph Visual Question Answering

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:45.568121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:45.568121Z digest=sha256:e3de5e6876168c31fb179ccd189915f8ebda9576be0fdd5893e34b883959eaba

Observation 23b76a4e-41e6-4102-909b-d05e0d7b7417 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:45.673966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:45.673966Z digest=sha256:792a559d51528e807fdf5407d4c95b57f76c0a19f915ce059d0d9b6d3e5698bf

Observation b07d8080-d6a9-4313-9b60-0fe51ef14d6f · outbound

This paper cites OmniDrive: A Holistic Vision-Language Dataset for Autonomous Driving with Counterfactual Reasoning.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines OmniDrive: A Holistic Vision-Language Dataset for Autonomous Driving with Counterfactual Reasoning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:45.752864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:45.752864Z digest=sha256:e60d5bbc56720cb87a557593af7c61dafadb4fa0b2b4b7654bac55f525af1aca

Observation ab119678-ae2b-4583-be42-ab3add7e1a52 · outbound

This paper cites On the Road with GPT-4V(ision): Early Explorations of Visual-Language Model on Autonomous Driving.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines On the Road with GPT-4V(ision): Early Explorations of Visual-Language Model on Autonomous Driving

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:45.837334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:45.837334Z digest=sha256:d90f975bc456943a19537559476cf31d4f8b6e9f65f660797c0e2049dbac06e1

Observation d789d401-3624-4b30-a015-00980ab687e3 · outbound

This paper cites Two-stage lvlm system: 1st place solution for eccv 2024 corner case scene understanding challenge.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Two-stage lvlm system: 1st place solution for eccv 2024 corner case scene understanding challenge

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:37:47.244661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:37:45.918065Z digest=sha256:86fa9449d535fd03c4cd57697434e47930132a55a31fcee7ad9ef4c8496d49f9

Observation 6d778d01-a12e-4015-a18a-b2fe06d7fc4e · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:45.972426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:45.972426Z digest=sha256:c9a25e95b833e3f528133bbee849bc17e22a201df7463e3c3f37d31c01ce151a

Observation b2be8034-3639-4772-9c79-34d983dd8a63 · outbound

This paper cites Flex-moe: Modeling arbitrary modality combination via the flexible mixture-of-experts.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Flex-moe: Modeling arbitrary modality combination via the flexible mixture-of-experts

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:37:47.083633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:37:46.073957Z digest=sha256:f6eaec83bfdd10937affa50968f72c58a267da0041b75c0c73458ac978bdef88

Observation e0d118b6-0f14-48c4-b74d-3d9dad7c8798 · outbound

This paper cites Sigmoid loss for language image pre-training.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Sigmoid loss for language image pre-training

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:46.195794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:46.195794Z digest=sha256:01a703496be3846d68e25c192f1364080401195e77d64a7e49f94f688bf2faf6

Observation 5f813ce0-2054-427d-b11b-70dab2f5dd13 · outbound

This paper cites MiniDrive: More Efficient Vision-Language Models with Multi-Level 2D Features as Text Tokens for Autonomous Driving.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines MiniDrive: More Efficient Vision-Language Models with Multi-Level 2D Features as Text Tokens for Autonomous Driving

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:46.269978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:46.269978Z digest=sha256:211099266f45f0ff9cc3157a61e035ff5c42d3c577ebbd99f2d67aadd84b7217

Observation e04c8bf4-71a5-4603-8a24-c4349f920eec · outbound

This paper cites Diversifying the expert knowledge for task-agnostic pruning in sparse mixture-of-experts, 2024 b.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Diversifying the expert knowledge for task-agnostic pruning in sparse mixture-of-experts, 2024 b

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:37:46.894735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:37:46.338934Z digest=sha256:7fd63da41da2fbdad88902785760a7d857af3bfa77e024917f9c06981a3b7946

Observation c3dd1e9b-0d51-4115-8e98-174d7f77f9a8 · outbound

This paper cites ST-MoE: Designing Stable and Transferable Sparse Expert Models.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines ST-MoE: Designing Stable and Transferable Sparse Expert Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:46.441294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:46.441294Z digest=sha256:b0bee13799f20acb0882f3d033a211bd473872c078b6cfc2fa8c059e1aa28e8e

Pith citing papers

Observation 6c0db8a3-3f58-49a1-8d04-192b6c2c2481 · inbound

D$^3$-MoE:Dual Disentangled Diffusion Mixture-of-Experts for Style-Controllable End-to-End Autonomous Driving cites this paper.

D$^3$-MoE:Dual Disentangled Diffusion Mixture-of-Experts for Style-Controllable End-to-End Autonomous Driving MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-02T08:26:48.317451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T06:02:52.244884Z digest=sha256:f31c1fe4afa5adbdb3b795c1d8624ea5c1db2884ba33d2a943c86e52818e2ac5