Pith. sign in

REVIEW 3 major objections 6 minor 33 references

Robot policies route experts by motion type learned from actions, then act from vision and language alone.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-30 20:28 UTC pith:NBR3ZBAR

load-bearing objection Solid systems paper: kinematic clustering as explicit MoE-router labels with observation-only inference, plus a cheap real dual-arm platform; gains look real, stats and N=K coupling are the soft spots. the 3 major comments →

arxiv 2607.26807 v1 pith:NBR3ZBAR submitted 2026-07-29 cs.RO

Route by Kinematics, Act by Observation: Kinematics-Supervised Expert Routing in MoE-Augmented VLA

classification cs.RO
keywords vision-language-actionmixture of expertskinematic routingrobot manipulationprivileged informationflow matchingcross-embodiment
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Vision-language-action models struggle when many manipulation tasks share similar scenes but demand very different motions, or look different but move the same way. Mixture-of-experts helps only if the router assigns the right specialist; ordinary routers, trained only from observations, often assign badly because true motion structure is invisible at test time. This paper shows that diverse tasks collapse into a few kinematic archetypes, clusters demonstration trajectories into those archetypes offline, and trains the router to predict the archetype ID from vision and language. At deployment the router never sees actions—only images and instructions—yet still dispatches the matching experts. On a standard simulation suite and on a cheap 3D-printed dual-arm platform the authors built, the method beats both dense and implicitly routed mixture models by large margins, especially on rare bimanual and contact-precise skills.

Core claim

Most semantically distinct manipulation tasks reduce to a small set of kinematic archetypes. Supervising a mixture-of-experts router with offline cluster IDs from action-velocity trajectories, then running that router on vision-language observations only, bridges the train-time privilege of kinematics into inference and yields large gains over dense VLAs and over MoE VLAs that route implicitly from observations.

What carries the argument

KinRT (kinematics-supervised explicit routing): offline K-means on concatenated action-chunk and velocity features produces prototype labels; a global router is trained with cross-entropy on those labels from pooled vision-language prefix features; at inference Top-K experts are chosen from observations alone, with a residual shared FFN always on.

Load-bearing premise

A handful of motion clusters found offline on training trajectories are both the right expert split and reliably predictable from vision and language alone, including on a new robot body.

What would settle it

If clustering action-velocity chunks into the same number of groups and supervising the router with those IDs does not beat matched dense and implicitly routed MoE baselines on RoboTwin and on the DIYRobot tasks—especially the rare bimanual and fine-contact ones—the central claim fails.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Expert specialization in robot VLAs should be driven by kinematic isomorphism, not visual or language similarity.
  • Privileged action trajectories can supervise routing without being needed at deployment.
  • Minority motion regimes (large bimanual coordination, fine positioning) get dedicated capacity instead of being washed out by dominant patterns.
  • A sub-$2000 3D-printed dual-arm platform and its five-task benchmark become a reusable real-world testbed for cross-embodiment claims.
  • The same asymmetric bridge can be dropped into other VLA backbones as a plug-in router, not only one architecture.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If kinematic archetypes transfer across embodiments, a shared offline motion taxonomy could standardize MoE design for multi-robot fleets without per-robot re-clustering.
  • Failure modes should concentrate where vision-language embeddings cannot separate two motions that look alike but move differently; those pairs are the natural stress tests for the bridge.
  • Adaptive expert counts or online reclustering would be the natural next control if fixed K under-covers long-tail skills as task libraries grow.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper proposes KinRT, a kinematics-supervised explicit routing scheme for MoE-augmented vision-language-action (VLA) policies. Offline K-means on concatenated action-velocity chunks yields kinematic archetype labels that supervise a global router via cross-entropy; at inference the router dispatches Top-K experts from pooled visual-language prefix features alone, realizing an asymmetric (privileged-training / observation-only-inference) bridge. The method is instantiated on a residual shared+routed MoE action head and evaluated against dense and implicitly routed MoE VLAs on RoboTwin and on a new low-cost 14-DoF DIYRobot platform and five-task real-world benchmark built by the authors. Reported gains exceed 23% relative on RoboTwin (KinRT-LoRA vs π0.5-LoRA) and 20% on DIYRobot (KinRT-Full vs π0.5-Full), with ablations on balanced sampling (α) and clustering source.

Significance. If the result holds, KinRT supplies a concrete, LUPI-style recipe for expert routing in embodied MoE-VLAs: use privileged action kinematics only to label the router, then run observation-only at deployment. That addresses a genuine failure mode of gradient-only routers under kinematic heterogeneity and is architecture-agnostic in the reported plug-in experiments. The DIYRobot platform and benchmark, if released as promised, are a useful community resource for small labs. Strengths that should be credited include the clustering-source ablation (Table 3), the VLM-vs-kinematics similarity analysis (Figure 3), residual MoE design (Eq. 1), and consistent gains across multiple backbones and both sim and real settings.

major comments (3)
  1. [§Kinematics-Supervised Global Router; Eq. (4); Table 1] The central asymmetric-bridge claim is that the router recovers kinematic archetype identity from observation alone (L_sup, §Kinematics-Supervised Global Router). Downstream task success (Table 1) is necessary but not sufficient evidence: please report router classification accuracy (overall and per-cluster), a confusion matrix, and optionally calibration, on held-out frames for both RoboTwin and DIYRobot. Without these metrics it remains possible that gains come from extra MoE capacity or residual design rather than successful kinematics distillation.
  2. [Table 1; Abstract; RQ1–RQ2] Table 1 reports raw success counts with no seeds, error bars, or significance tests, yet the abstract and RQ1 quote precise relative improvements (23.26%, 20.27%). For load-bearing SOTA claims, rerun with ≥3 seeds (or bootstrap over evaluation episodes) and report mean±std; mark pairwise comparisons that remain significant. This is especially important for DIYRobot (50 tests per task) and for minority-archetype tasks (e.g., Handover Block, Press Button) where variance is likely high.
  3. [§Kinematic Archetype Clustering; Implementation] N is set equal to the number of K-means clusters discovered on the training set (N=K=4 on RoboTwin), coupling the expert inventory to a single offline partition. Please (i) state explicitly whether DIYRobot is re-clustered independently or reuses RoboTwin prototypes, (ii) ablate K∈{2,4,8} (or equivalent) on at least one benchmark, and (iii) discuss failure modes when test kinematics fall outside the training prototype set. The current N=K choice is a free parameter that affects the main claim’s generality.
minor comments (6)
  1. [Full manuscript] Throughout the PDF, many word boundaries are missing (e.g., “suffersfromineffective”, “kinematicarchetypes”). This appears to be a compilation/encoding issue and should be fixed for readability.
  2. [Eq. (1)] Eq. (1) fixes shared/routed mix weights at 1/2–1/2 with no ablation; a one-line sensitivity check (or justification from pretraining stability) would help.
  3. [Table 1] KinRT-AdaMoE rows are blank on RoboTwin (Table 1); either complete the numbers or explain why that hybrid was evaluated only on DIYRobot.
  4. [Figure 3] Figure 3 caption and axes should state the exact PCA dimension and cosine definition so the density plots are reproducible from the text alone.
  5. [Implementation; Eq. (3)] Hyperparameters σ, τ for router noise/temperature and the supervised-loss coefficient 0.05 are stated but not ablated; a short appendix note on stability would suffice.
  6. [§DIYRobot Platform and Benchmark; Experimental Setups] Clarify data split protocol on DIYRobot (train/test episode separation, whether evaluation uses held-out objects/poses) to support the cross-platform claim.

Circularity Check

0 steps flagged

No significant circularity: router labels come from offline action clustering, while claimed gains are external task-success rates under observation-only control.

full rationale

KinRT’s chain is a standard privileged-information training setup, not a closed definitional loop. Offline K-means on action–velocity chunks yields prototype IDs y_i used only as supervised targets for the router (L_sup); at inference the router sees only pooled vision–language context c and never the kinematics used to build the labels. The headline metrics are per-task and average success counts on RoboTwin and DIYRobot under observation-only deployment—quantities that are not algebraic restatements of the cluster objective, the cross-entropy L_sup, or the choice N=K. Setting the expert count equal to the number of clusters is a design hyperparameter, not a construction that forces high success. Ablations (clustering source, balanced-sampling α) and backbone plug-ins further treat the method as an empirical hypothesis tested against independent baselines (dense VLAs, implicitly routed MoEs). There is no self-definitional equation, no fitted constant renamed as a prediction, no load-bearing uniqueness theorem imported from overlapping authors, and no renaming of a known closed-form result. The paper is self-contained against external benchmarks; circularity score is therefore 0.

Axiom & Free-Parameter Ledger

6 free parameters · 5 axioms · 3 invented entities

The central claim rests on standard MoE/VLA machinery, the LUPI-style privilege of training-only kinematics, and several hand-chosen clustering and routing hyperparameters. No new physical entities; DIYRobot and 'kinematic archetypes as router GT' are methodological constructs. Free parameters (K/N, α, loss weight, feature pipeline) are fit or selected on the same benchmarks used for claims.

free parameters (6)
  • number of kinematic clusters / experts K=N = 4
    Set to 4 to match K-means archetypes on RoboTwin frames; directly defines expert capacity and supervision cardinality.
  • balanced sampling coefficient α = 0.5
    Interpolates empirical vs uniform prototype sampling; chosen by ablation (best at 0.5 on RoboTwin).
  • supervised routing loss coefficient = 0.05
    Weight on L_sup relative to flow-matching action loss; set to 0.05 without broad sensitivity study.
  • action chunk horizon H and PCA dimension = H=50, PCA=64
    H=50 and PCA to 64 define the kinematic descriptor before K-means; hand-set design choices that shape clusters.
  • shared/routed FFN mix weights and Top-K = 1/2+1/2, Top-1
    Fixed 1/2–1/2 residual mix and Top-1 routing; architectural knobs not derived from theory.
  • router noise σ and temperature τ = unspecified numeric values
    Exploration noise and softmax temperature during training; paper says set 'accordingly' without unique values.
axioms (5)
  • domain assumption Kinematic prototype collapse: heterogeneous manipulation tasks reduce to a small number of coherent motion archetypes discoverable by clustering action-velocity features.
    Stated as an empirical observation (Fig. 1, RoboTwin 4 clusters) and used as the justification for expert count and supervision.
  • domain assumption Privileged training-only signals can improve a student that lacks them at inference (LUPI / privileged teacher paradigm).
    Related Work and method frame KinRT as distilling action-space structure into observation-space routing.
  • ad hoc to paper Pooled prefix VLM tokens are sufficiently discriminative for kinematic archetype classification even when raw VLM cosine similarities collapse (~0.95).
    Authors reject raw embeddings for routing and substitute masked-mean prefix context c∈R^2048; load-bearing for observation-only dispatch.
  • domain assumption Standard sparsely-gated MoE training and flow-matching action heads are valid policy learners for continuous robot control.
    Architecture builds on cited MoE and π0-style flow-matching without re-deriving those foundations.
  • ad hoc to paper K-means in PCA-reduced action-velocity space yields semantically meaningful expert partitions for routing labels.
    Clustering pipeline (standardize, PCA 64, frame-level K-means) is a methodological choice treated as ground truth for L_sup.
invented entities (3)
  • KinRT asymmetric bridging router no independent evidence
    purpose: Map observation-only inputs to experts whose specialization was defined by privileged kinematic clusters.
    Core proposed mechanism; not a physical entity but a new training/inference contract for MoE-VLA.
  • DIYRobot 14-DoF platform and 5-task benchmark no independent evidence
    purpose: Provide a cheap real-world testbed for cross-platform generalization claims.
    Author-constructed hardware and dataset; evidence quality depends on future public release and third-party use.
  • Kinematic archetype labels as router ground truth no independent evidence
    purpose: Replace implicit gradient-only routing targets with explicit motion-family IDs.
    Labels are defined by the paper's clustering, not an external ontology; validated only via downstream success.

pith-pipeline@v1.2.0-daily-grok45 · 19447 in / 3952 out tokens · 84547 ms · 2026-07-30T20:28:45.335911+00:00 · methodology

0 comments
read the original abstract

While MoE augments VLA via expert specialization, router suffers from ineffective expert routing owing to the kinematic heterogeneity of actions across manipulation tasks and, even worse, the unavailability of the kinematic signals at inference time. In this work, we first observe that most semantically distinct manipulation tasks reduce to multiple kinematic archetypes. Motivated by this finding, we propose Kinematics-supervised explicit routing (KinRT), a new paradigm that shifts from implicit, observation-driven expert routing to explicit, kinematics-guided expert dispatching. Specifically, we perform kinematic clustering on action trajectories into multiple kinematically coherent groups, whose IDs serve as ground truth to supervise the training of the router; at inference time, the router dispatches experts only using visual-language observations, without any reliance on action kinematics. KinRT actually introduces an asymmetric bridging mechanism that distills the task kinematics from the action space in training into the observation space at inference. In addition, to assess KinRT's cross-platform generalization, we build an economical, Do-It-Yourself robot (DIYRobot) platform from scratch using 3D-print technology ($<$ 2,000USD). Extensive experiments demonstrate KinRT's superiority over both dense and MoE-featured VLAs by more than 23.26% on RoboTwin benchmark and 20.27% on our introduced DIYRobot platform. Our code and DIYRobot platform will be open-sourced.

Figures

Figures reproduced from arXiv: 2607.26807 by Junjie Wang, Ruotong Li, Tianhang Yang, Wei-Bin Kou, Yanze Zheng, Yujiu Yang.

Figure 1
Figure 1. Figure 1: Illustration of kinematic archetype collapse. [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Overview of the proposed KinRT paradigm and our newly introduced DIYRobot platform. [PITH_FULL_IMAGE:figures/full_fig_p002_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Demonstrations of the relationship between action [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Demonstrations of the five manipulation tasks performed on our DIYRobot platform, where the left-to-right sequence [PITH_FULL_IMAGE:figures/full_fig_p007_4.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

33 extracted references · 8 linked inside Pith

  1. [1]

    X.; Tanner, J.; Vuong, Q.; Walling, A.; Wang, H.; and Zhilinsky, U

    Black, K.; Brown, N.; Driess, D.; Esmail, A.; Equi, M.; Finn, C.; Fusai, N.; Groom, L.; Hausman, K.; Ichter, B.; Jakubczak, S.; Jones, T.; Ke, L.; Levine, S.; Li-Bell, A.; Mothukuri, M.; Nair, S.; Pertsch, K.; Shi, L. X.; Tanner, J.; Vuong, Q.; Walling, A.; Wang, H.; and Zhilinsky, U. 2026. _0 : A Vision-Language-Action Flow Model for General Robot Control

  2. [2]

    G.; Gopalakrishnan, K.; Han, K.; Hausman, K.; Herzog, A.; Hsu, J.; Ichter, B.; Irpan, A.; Joshi, N.; Julian, R.; Kalashnikov, D.; Kuang, Y.; Leal, I.; Lee, L.; Lee, T.-W

    Brohan, A.; Brown, N.; Carbajal, J.; Chebotar, Y.; Chen, X.; Choromanski, K.; Ding, T.; Driess, D.; Dubey, A.; Finn, C.; Florence, P.; Fu, C.; Arenas, M. G.; Gopalakrishnan, K.; Han, K.; Hausman, K.; Herzog, A.; Hsu, J.; Ichter, B.; Irpan, A.; Joshi, N.; Julian, R.; Kalashnikov, D.; Kuang, Y.; Leal, I.; Lee, L.; Lee, T.-W. E.; Levine, S.; Lu, Y.; Michalew...

  3. [3]

    Brohan, A.; Brown, N.; Carbajal, J.; Chebotar, Y.; Dabis, J.; Finn, C.; Gopalakrishnan, K.; Hausman, K.; Herzog, A.; Hsu, J.; Ibarz, J.; Ichter, B.; Irpan, A.; Jackson, T.; Jesmonth, S.; Joshi, N. J.; Julian, R.; Kalashnikov, D.; Kuang, Y.; Leal, I.; Lee, K.-H.; Levine, S.; Lu, Y.; Malla, U.; Manjunath, D.; Mordatch, I.; Nachum, O.; Parada, C.; Peralta, J...

  4. [4]

    X.; Gao, H.; Chen, D.; Li, J.; Zeng, W.; Yu, X.; Wu, Y.; Xie, Z.; Li, Y

    Dai, D.; Deng, C.; Zhao, C.; Xu, R. X.; Gao, H.; Chen, D.; Li, J.; Zeng, W.; Yu, X.; Wu, Y.; Xie, Z.; Li, Y. K.; Huang, P.; Luo, F.; Ruan, C.; Sui, Z.; and Liang, W. 2024. DeepSeekMoE : Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models. arXiv preprint arXiv:2401.06066

  5. [5]

    M.; Tong, S.; Lepikhin, D.; Xu, Y.; Krikun, M.; Zhou, Y.; Yu, A

    Du, N.; Huang, Y.; Dai, A. M.; Tong, S.; Lepikhin, D.; Xu, Y.; Krikun, M.; Zhou, Y.; Yu, A. W.; Firat, O.; Zoph, B.; Fedus, L.; Bosma, M.; Zhou, Z.; Wang, T.; Wang, Y. E.; Webster, K.; Pellat, M.; Robinson, K.; Meier-Hellstern, K.; Duke, T.; Dixon, L.; Zhang, K.; Le, Q. V.; Wu, Y.; Chen, Z.; and Cui, C. 2022. GLaM : Efficient Scaling of Language Models wi...

  6. [6]

    Du, Z.; Liu, B.; Liang, Y.; Shen, Y.; Cao, H.; Zheng, X.; Feng, Z.; Wu, Z.; Yang, J.; and Jiang, Y.-G. 2025. HiMoE-VLA: Hierarchical Mixture-of-Experts for Generalist Vision-Language-Action Policies. arXiv preprint arXiv:2512.05693

  7. [7]

    Fedus, W.; Zoph, B.; and Shazeer, N. 2022. Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity. Journal of Machine Learning Research, 23(120): 1--39

  8. [8]

    Y.; Ghosh, D.; Groom, L.; Hausman, K.; Ichter, B.; Jakubczak, S.; Jones, T.; Ke, L.; LeBlanc, D.; Levine, S.; Li-Bell, A.; Mothukuri, M.; Nair, S.; Pertsch, K.; Ren, A

    Intelligence, P.; Black, K.; Brown, N.; Darpinian, J.; Dhabalia, K.; Driess, D.; Esmail, A.; Equi, M.; Finn, C.; Fusai, N.; Galliker, M. Y.; Ghosh, D.; Groom, L.; Hausman, K.; Ichter, B.; Jakubczak, S.; Jones, T.; Ke, L.; LeBlanc, D.; Levine, S.; Li-Bell, A.; Mothukuri, M.; Nair, S.; Pertsch, K.; Ren, A. Z.; Shi, L. X.; Smith, L.; Springenberg, J. T.; Sta...

  9. [9]

    Q.; Sablayrolles, A.; Roux, A.; Mensch, A.; Savary, B.; Bamford, C.; Chaplot, D

    Jiang, A. Q.; Sablayrolles, A.; Roux, A.; Mensch, A.; Savary, B.; Bamford, C.; Chaplot, D. S.; de las Casas, D.; Hanna, E. B.; Bressand, F.; Lengyel, G.; Bour, G.; Lample, G.; Lavaud, L. R.; Saulnier, L.; Lachaux, M.-A.; Stock, P.; Subramanian, S.; Yang, S.; Antoniak, S.; Le Scao, T.; Gervet, T.; Lavril, T.; Wang, T.; Lacroix, T.; and El Sayed, W. 2024. M...

  10. [10]

    Jiang, Y.; Gupta, A.; Zhang, Z.; Wang, G.; Dou, Y.; Chen, Y.; Fei-Fei, L.; Anandkumar, A.; Zhu, Y.; and Fan, L. 2023. VIMA: General Robot Manipulation with Multimodal Prompts. In International Conference on Machine Learning (ICML)

  11. [11]

    Kim, M. J.; Pertsch, K.; Karamcheti, S.; Xiao, T.; Balakrishna, A.; Nair, S.; Rafailov, R.; Foster, E.; Lam, G.; Sanketi, P.; Vuong, Q.; Kollar, T.; Burchfiel, B.; Tedrake, R.; Sadigh, D.; Levine, S.; Liang, P.; and Finn, C. 2024. OpenVLA: An Open-Source Vision-Language-Action Model. arXiv preprint arXiv:2406.09246

  12. [12]

    Lee, J.; Hwangbo, J.; Wellhausen, L.; Koltun, V.; and Hutter, M. 2020. Learning quadrupedal locomotion over challenging terrain. Science robotics, 5(47): eabc5986

  13. [13]

    Lepikhin, D.; Lee, H.; Xu, Y.; Chen, D.; Firat, O.; Huang, Y.; Krikun, M.; Shazeer, N.; and Chen, Z. 2021. GShard : Scaling Giant Models with Conditional Computation and Automatic Sharding. In Proc. Int. Conf. Learning Representations (ICLR)

  14. [14]

    Lewis, M.; Bhosale, S.; Dettmers, T.; Goyal, N.; and Zettlemoyer, L. 2021. BASE Layers: Simplifying Training of Large, Sparse Models. In Proc. Int. Conf. Machine Learning (ICML), 6265--6274

  15. [15]

    Liang, H.; Fan, Z.; Sarkar, R.; Jiang, Z.; Chen, T.; Zou, K.; Cheng, Y.; Hao, C.; and Wang, Z. 2022. M ^3 ViT: Mixture-of-Experts Vision Transformer for Efficient Multi-task Learning with Model-Accelerator Co-design. In Koyejo, S.; Mohamed, S.; Agarwal, A.; Belgrave, D.; Cho, K.; and Oh, A., eds., Advances in Neural Information Processing Systems, volume ...

  16. [16]

    Liang, Y.; Ellis, K.; and Henriques, J. 2024. Rapid motor adaptation for robotic manipulator arms. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 16404--16413

  17. [17]

    Liu, B.; Liu, X.; Jin, X.; et al. 2021. Conflict-Averse Gradient Descent for Multi-Task Learning. In Advances in Neural Information Processing Systems (NeurIPS)

  18. [18]

    Liu, S.; Wu, L.; Li, B.; Tan, H.; Chen, H.; Wang, Z.; Xu, K.; Su, H.; and Zhu, J. 2025. Rdt-1b: a diffusion foundation model for bimanual manipulation. In International Conference on Learning Representations, volume 2025, 29982--30009

  19. [19]

    Mu, S.; and Lin, S. 2025. A Comprehensive Survey of Mixture-of-Experts: Algorithms, Theory, and Applications. arXiv preprint arXiv:2503.07137

  20. [20]

    Mustafa, B.; Riquelme, C.; Puigcerver, J.; Jenatton, R.; and Houlsby, N. 2022. Multimodal contrastive learning with limoe: the language-image mixture of experts. Advances in Neural Information Processing Systems, 35: 9564--9576

  21. [21]

    Octo Model Team . 2024. Octo: An Open-Source Generalist Robot Policy. In Robotics: Science and Systems (RSS)

  22. [22]

    Riquelme, C.; Puigcerver, J.; Mustafa, B.; Neumann, M.; Jenatton, R.; Susano Pinto, A.; Keysers, D.; and Houlsby, N. 2021. Scaling Vision with Sparse Mixture of Experts. Advances in Neural Information Processing Systems (NeurIPS), 34: 8583--8595

  23. [23]

    Shazeer, N.; Mirhoseini, A.; Maziarz, K.; Davis, A.; Le, Q.; Hinton, G.; and Dean, J. 2017. Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer. In International Conference on Learning Representations (ICLR)

  24. [24]

    Shen, W.; Liu, Y.; Wu, Y.; Liang, Z.; Gu, S.; Wang, D.; Nian, T.; Xu, L.; Qin, Y.; Pang, J.; Guan, X.; Yang, X.; and Mu, Y. 2025. Expertise need not monopolize: Action-Specialized Mixture of Experts for Vision-Language-Action Learning. arXiv preprint arXiv:2510.14300

  25. [25]

    Shridhar, M.; Manuelli, L.; and Fox, D. 2023. Perceiver-Actor: A Multi-Task Transformer for Robotic Manipulation. In Conference on Robot Learning (CoRL)

  26. [26]

    Vapnik, V.; and Vashist, A. 2009. A New Learning Paradigm: Learning Using Privileged Information. Neural Networks, 22(5-6): 544--557

  27. [27]

    Wu, W.; Lu, F.; Wang, Y.; Yang, S.; Liu, S.; Wang, F.; Ma, S.; Sun, H.; Wang, Y.; Qiu, Z.; Xiong, H.; Wang, Z.; Zhou, S.; Ren, Y.; Zhang, K.; Yu, H.; Zhao, J.; Zhu, Q.; Cheng, R.; Li, Y.-L.; Huang, Y.; Zhu, X.; Shen, Y.; and Zheng, K. 2026 a . A Pragmatic VLA Foundation Model. arXiv preprint arXiv:2601.18692v1

  28. [28]

    Wu, W.; Wang, F.; Lu, F.; Sun, H.; Liu, S.; Wang, Y.; Yan, Y.; Wang, Y.; Ma, S.; Wang, X.; Liu, Y.; Yang, S.; Zhou, T.; Zhang, K.; Zhou, L.; Su, C.; Xue, N.; Tan, B.; Zhang, H.; Zhang, Y.; Liao, F.; Zhu, X.; Shen, Y.; and Zheng, K. 2026 b . From Foundation to Application: Improving VLA Models in Practice. arXiv preprint arXiv:2607.06403

  29. [29]

    Yu, T.; Kumar, S.; Gupta, A.; Levine, S.; Hausman, K.; and Finn, C. 2020. Gradient surgery for multi-task learning. Advances in neural information processing systems, 33: 5824--5836

  30. [30]

    Zhang, L.; Tang, T.; Zhan, Z.; Chen, X.; Chen, Z.; Han, J.; Zhu, J.; Xu, P.; Xu, H.; Wu, H.; Lin, L.; and Liang, X. 2026. Atomicvla: Unlocking the potential of atomic skill learning in robots. arXiv preprint arXiv:2603.07648

  31. [31]

    Y.; Dai, A

    Zhou, Y.; Lei, T.; Liu, H.; Du, N.; Huang, Y.; Zhao, V. Y.; Dai, A. M.; Chen, Z.; Le, Q. V.; and Laudon, J. 2022. Mixture-of-Experts with Expert Choice Routing. In Advances in Neural Information Processing Systems (NeurIPS), volume 35, 7103--7114

  32. [32]

    Zoph, B.; Bello, I.; Kumar, S.; Du, N.; Huang, Y.; Dean, J.; Shazeer, N.; and Fedus, W. 2022. ST-MoE: Designing Stable and Transferable Sparse Expert Models. arXiv preprint arXiv:2202.08906

  33. [33]

    J.; Hassan, H.; Zhang, R.; Zhao, T.; and Gao, J

    Zuo, S.; Liu, X.; Jiao, J.; Kim, Y. J.; Hassan, H.; Zhang, R.; Zhao, T.; and Gao, J. 2022. Taming Sparsely Activated Transformer with Stochastic Experts. In International Conference on Learning Representations (ICLR)