REVIEW 3 major objections 52 references
Different action interfaces can share one video world model when experts separate common world dynamics from control-specific details.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-11 22:41 UTC pith:ZITXFZHK
load-bearing objection Solid systems paper that unifies camera/robot/hand control in one DiT-MoE world model; the multi-benchmark wins are real, but the MoE-vs-dense ablation does not cleanly prove specialization over capacity. the 3 major comments →
Worldscape-MoE: A Unified Mixture-of-Experts World Model for Scalable Heterogeneous Action Control
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Heterogeneous action supervision improves individual control capabilities when a single Diffusion Transformer world model factors computation into a shared expert for cross-control physical regularities and modality-specific experts for each action interface, instead of forcing all controls through one dense pathway.
What carries the argument
Worldscape-MoE: modality-aware control injection into a shared DiT backbone plus control-aware MoE feed-forwards that always activate a shared expert and only eligible control-specific experts, trained with progressive expert expansion and grouped learning rates.
Load-bearing premise
The claim rests on the premise that masking experts by control type, always updating a shared expert at a lower learning rate, and cloning new experts from that shared expert is enough to separate common world physics from control-specific residuals—so measured gains are specialization, not just extra capacity or a lucky data mix.
What would settle it
Train a capacity-matched dense multi-branch baseline on the same mixed data, budget, and injection paths without modality-masked MoE routing; if it matches or beats Worldscape-MoE on locomotion, manipulation, and hand metrics while expert-loading no longer shows modality-dependent dedicated use, the specialization claim is falsified.
If this is right
- Fragmented camera, robot, and hand datasets can be pooled as one growing training resource instead of training siloed single-control models.
- New action modalities can be added by extending experts and injection branches while retaining previously learned controls.
- Shared experts transfer physical and scene regularities to out-of-distribution environments and to new single-arm manipulation settings with limited fine-tuning.
- Once multiple controls share one world representation, the model can begin to compose previously isolated skills such as loco-manipulation.
- Controllable world models can scale by absorbing heterogeneous supervision rather than by repeatedly partitioning capacity along control boundaries.
Where Pith is reading between the lines
- If the gains are truly from factorization, equal-parameter dense multi-branch models should still underperform even when data mix and injection paths are held fixed—a stricter capacity control the paper only partly isolates.
- The same shared-plus-specific expert pattern may transfer to other action-conditioned predictors that currently keep separate stacks per robot or sensor suite.
- Progressive cloning of experts from a shared world prior could become a practical recipe for lifelong addition of embodiments without full retraining.
- Reliable loco-manipulation composition would reduce reliance on hand-designed skill libraries when planning over mixed action types in one simulator.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes Worldscape-MoE, a DiT-based video world model that unifies heterogeneous action controls (camera locomotion, dual-arm robot actions, and dense hand-joint action maps) via modality-aware injection, a control-aware MoE-FFN with a always-on shared expert plus eligibility-masked modality experts (Eqs. 4–7), and progressive Worldscape-MoE Tuning with shared-expert cloning and grouped learning rates. It argues that different controls are interfaces to shared world dynamics, so joint training should improve rather than interfere with individual controls (RQ1–RQ4). Empirically, the model reports leading aggregate scores on iWorld-Bench locomotion (Table 1), WorldArena EWMScore for dual-arm manipulation (Table 2; 62.84), and EgoDex-subset hand metrics (Table 3), with routing load analysis (Fig. 5), dense mixed-training ablations, modality-extension notes, OOD qualitative cases, and loco-manipulation composition demos.
Significance. If the factorization claim holds, the work is a useful systems contribution for embodied world models: it reframes fragmented control-specific generators as a single extensible training resource and shows competitive multi-control performance on standard-ish benchmarks (WorldArena, iWorld-Bench-style locomotion, hand FID/FVD). Strengths include a clear architectural recipe (asymmetric injection + shared/dedicated experts), progressive expansion protocol, multi-regime evaluation, and expert-loading diagnostics. The significance is primarily empirical and engineering rather than theoretical; the paper would matter most if it cleanly shows that MoE specialization—not just capacity, data mix, or staged multi-branch conditioning—is what turns heterogeneous supervision into a scaling resource.
major comments (3)
- The central factorization claim (shared expert = world dynamics; dedicated experts = control residuals; heterogeneous training improves rather than interferes) is under-identified by the decisive ablation. §3.4 and Tables 1–3 compare Worldscape-MoE to a dense “w/o MoE” mixed-training baseline under a claimed same budget, but MoE multiplies FFN capacity (Eq. 7), uses progressive expert cloning and grouped LRs (§2.4), and routes by deterministic modality eligibility. Without a capacity-matched multi-expert dense control, multi-branch dense conditioner, or parameter-matched shared-only MoE, gains may reflect extra parameters and better-conditioned adapters rather than the claimed shared/specific factorization (RQ2).
- Aggregate leadership overstates cross-control fidelity. In Table 1, Worldscape-MoE’s overall locomotion score (0.7556) leads, but Trajectory Accuracy (0.6300) trails VideoX-Fun-Wan (0.7645) and RealCam-I2V (0.7050); motion smoothness is near ceiling for many methods. The paper should report primary control-faithfulness metrics with equal weight to aggregates, and state clearly where unified training trades off pure trajectory following.
- RQ3 scalability and “improves rather than interferes” are only partially evidenced. §3.4 and Appendix D.3 describe transient locomotion degradation after adding a modality and later recovery, plus expert L2 update tables, but there is no quantitative before/after table of all primary metrics when experts/data are added, no multi-seed statistics, and no held-out joint-control quantitative protocol for loco-manipulation (Fig. 32 is qualitative). The continual-extension claim needs a compact scaling table and uncertainty estimates.
Circularity Check
No significant circularity: empirical MoE world-model claims are tested on held-out benchmarks, not derived by redefining targets as fitted inputs.
full rationale
Worldscape-MoE is an empirical systems/ML paper. Its load-bearing claims—that a DiT with shared plus modality-specific experts can absorb heterogeneous action supervision without cross-control interference, and that joint training improves locomotion, manipulation, and hand-control metrics—are evaluated by training and scoring on external or held-out protocols (iWorld-Bench locomotion set, WorldArena EWMScore, EgoDex hand subset), with ablations (w/o MoE dense mixed training, routing workload, progressive extension). There is no first-principles derivation in which a predicted quantity is algebraically identical to a fitted constant, no uniqueness theorem imported from the authors that forces the architecture, and no ansatz smuggled in as a mathematical necessity. Self-citations (e.g., WorldArena, iWorld-Bench, Matrix-Game) supply evaluation protocols and related-work context; they do not close a definitional loop with the reported scores. Design choices (deterministic modality eligibility, shared expert always updated, progressive expert cloning, grouped LRs) are architectural hypotheses tested empirically, not circular identities. Concerns that the MoE-vs-dense comparison under-isolates capacity versus specialization are experimental-identification issues, not circularity of the derivation chain.
Axiom & Free-Parameter Ledger
free parameters (5)
- Router temperature τ
- Grouped learning rates (shared vs modality-specific)
- Progressive stage schedule and expert-init cloning
- Action tensor shape Ta=17, dact=14
- Number of DiT MoE layers / expert set size M+1
axioms (5)
- domain assumption Heterogeneous controls (camera, robot actions, hand maps) are different interfaces to one shared latent world with reusable physical and temporal regularities.
- domain assumption A pretrained video DiT/VAE/text encoder provides transferable spatiotemporal priors that fine-tuning should preserve.
- ad hoc to paper Deterministic modality eligibility (shared always on; modality expert only if control present) induces useful specialization without auxiliary load-balancing loss.
- ad hoc to paper Asymmetric injection—dense controls via visual tokens, compact actions via timestep modulation—is an appropriate unified conditioning form.
- domain assumption Standard diffusion training and video-quality / EWM metrics are adequate proxies for world-model utility under action control.
invented entities (3)
-
Worldscape-MoE control-aware MoE-FFN (shared + modality experts with eligibility-masked router)
no independent evidence
-
Worldscape-MoE Tuning (progressive control expansion with shared-expert cloning and grouped LRs)
no independent evidence
-
Modality-aware heterogeneous control injection pathways (Control Adapter / VAE action maps / Action MLP)
no independent evidence
read the original abstract
World models are rapidly becoming a core infrastructure for embodied intelligence and interactive agents: they provide controllable simulators in which agents can perceive, act, forecast, and acquire scalable experience. Yet current video generation world models are still organized around isolated control interfaces, such as camera trajectories, robot actions, or hand-joint signals. This fragmentation is increasingly a scaling bottleneck. The central challenge is not the absence of controllable generators, but the lack of a unified and extensible learning framework that can absorb heterogeneous action supervision while preserving a shared model of world dynamics. In this work, we introduce Worldscape-MoE, a Mixture-of-Experts world model built on Diffusion Transformers for scalable heterogeneous action control. Our key observation is that different controls specify different interfaces to the same underlying world: although their representations differ, they constrain shared physical regularities, scene dynamics, and interaction semantics. Worldscape-MoE operationalizes this observation through modality-aware control injection, shared and control-specific experts, and a progressive MoE tuning strategy that supports continual extension to new action modalities. Experiments across locomotion, robotic manipulation, and egocentric hand control show that heterogeneous supervision improves rather than interferes with individual control capabilities. Worldscape-MoE achieves strong results on WorldArena, improves locomotion and hand-control metrics, exhibits robust out-of-distribution generalization, and demonstrates scaling behavior as additional control data and experts are integrated.
Figures
Reference graph
Works this paper leans on
-
[1]
Wan: Open and advanced large-scale video generative models.arXiv preprint arXiv:2503.20314, 2025
Team Wan, Ang Wang, Baole Ai, Bin Wen, Chaojie Mao, Chen-Wei Xie, Di Chen, Feiwu Yu, Haiming Zhao, Jianxiao Yang, et al. Wan: Open and advanced large-scale video generative models.arXiv preprint arXiv:2503.20314, 2025
Pith/arXiv arXiv 2025
-
[2]
Hunyuanvideo 1.5 technical report.arXiv preprint arXiv:2511.18870, 2025
Bing Wu, Chang Zou, Changlin Li, Duojun Huang, Fang Yang, Hao Tan, Jack Peng, Jianbing 13 Wu, Jiangfeng Xiong, Jie Jiang, et al. Hunyuanvideo 1.5 technical report.arXiv preprint arXiv:2511.18870, 2025
Pith/arXiv arXiv 2025
-
[3]
Zhuoyi Yang, Jiayan Teng, Wendi Zheng, Ming Ding, Shiyu Huang, Jiazheng Xu, Yuanming Yang, Wenyi Hong, Xiaohan Zhang, Guanyu Feng, et al. Cogvideox: Text-to-video diffusion models with an expert transformer.arXiv preprint arXiv:2408.06072, 2024
Pith/arXiv arXiv 2024
-
[4]
Understanding world or predicting future? a comprehensive survey of world models.ACM Computing Surveys, 58(3):1–38, 2025
Jingtao Ding, Yunke Zhang, Yu Shang, Yuheng Zhang, Zefang Zong, Jie Feng, Yuan Yuan, Hongyuan Su, Nian Li, Nicholas Sukiennik, et al. Understanding world or predicting future? a comprehensive survey of world models.ACM Computing Surveys, 58(3):1–38, 2025
2025
-
[5]
GigaWorld Team, Angen Ye, Boyuan Wang, Chaojun Ni, Guan Huang, Guosheng Zhao, Haoyun Li, Jiagang Zhu, Kerui Li, Mengyuan Xu, et al. Gigaworld-0: World models as data engine to empower embodied ai.arXiv preprint arXiv:2511.19861, 2025
arXiv 2025
-
[6]
Genie: Generative interactive environments
Jake Bruce, Michael D Dennis, Ashley Edwards, Jack Parker-Holder, Yuge Shi, Edward Hughes, Matthew Lai, Aditi Mavalankar, Richie Steigerwald, Chris Apps, et al. Genie: Generative interactive environments. InForty-first International Conference on Machine Learning, 2024
2024
-
[7]
Xianglong He, Chunli Peng, Zexiang Liu, Boyang Wang, Yifan Zhang, Qi Cui, Fei Kang, Biao Jiang, Mengyin An, Yangyang Ren, et al. Matrix-game 2.0: An open-source real-time and streaming interactive world model.arXiv preprint arXiv:2508.13009, 2025
Pith/arXiv arXiv 2025
-
[8]
Zile Wang, Zexiang Liu, Jaixing Li, Kaichen Huang, Baixin Xu, Fei Kang, Mengyin An, Peiyu Wang, Biao Jiang, Yichen Wei, et al. Matrix-game 3.0: Real-time and streaming interactive world model with long-horizon memory.arXiv preprint arXiv:2604.08995, 2026
Pith/arXiv arXiv 2026
-
[9]
Matrix-game: Interactive world foundation model.arXiv preprint arXiv:2506.18701, 2025
Yifan Zhang, Chunli Peng, Boyang Wang, Puyi Wang, Qingcheng Zhu, Fei Kang, Biao Jiang, Zedong Gao, Eric Li, Yang Liu, et al. Matrix-game: Interactive world foundation model.arXiv preprint arXiv:2506.18701, 2025
Pith/arXiv arXiv 2025
-
[10]
Wenqiang Sun, Haiyu Zhang, Haoyuan Wang, Junta Wu, Zehan Wang, Zhenwei Wang, Yunhong Wang, Jun Zhang, Tengfei Wang, and Chunchao Guo. Worldplay: Towards long-term geometric consistency for real-time interactive world modeling.arXiv preprint arXiv:2512.14614, 2025
Pith/arXiv arXiv 2025
-
[11]
Team HY-World, Chenjie Cao, Xuhui Zuo, Zhenwei Wang, Yisu Zhang, Junta Wu, Zhenyang Liu, Yuning Gong, Yang Liu, Bo Yuan, et al. Hy-world 2.0: A multi-modal world model for reconstructing, generating, and simulating 3d worlds.arXiv preprint arXiv:2604.14268, 2026
Pith/arXiv arXiv 2026
-
[12]
Yanjiang Guo, Lucy Xiaoyang Shi, Jianyu Chen, and Chelsea Finn. Ctrl-world: A controllable generative world model for robot manipulation.arXiv preprint arXiv:2510.10125, 2025
Pith/arXiv arXiv 2025
-
[13]
World simulation with video foundation models for physical ai.arXiv preprint arXiv:2511.00062, 2025
Arslan Ali, Junjie Bai, Maciej Bala, Yogesh Balaji, Aaron Blakeman, Tiffany Cai, Jiaxin Cao, Tianshi Cao, Elizabeth Cha, Yu-Wei Chao, et al. World simulation with video foundation models for physical ai.arXiv preprint arXiv:2511.00062, 2025
Pith/arXiv arXiv 2025
-
[14]
Yu Shang, Zhuohang Li, Yiding Ma, Weikang Su, Xin Jin, Yinzhou Tang, Haisheng Su, Chen Gao, Wei Wu, Xihui Liu, Zhibo Chen, Jun Zhu, Yonghong Tian, Tat-Seng Chua, Ziyou Wang, Dhruv Shah, Wenwu Zhu, Lei Jin, Xin Zhang, Zhaoxiang Zhang, and Yong Li. Worldarena: A unified benchmark for evaluating perception and functional utility of embodied world models. arX...
arXiv 2026
-
[15]
Yuxi Wang, Wenqi Ouyang, Tianyi Wei, Yi Dong, Zhiqi Shen, and Xingang Pan. Hand2world: Autoregressive egocentric interaction generation via free-space hand gestures.arXiv preprint arXiv:2602.09600, 2026. 14
arXiv 2026
-
[16]
Linxi Xie, Lisong C Sun, Ashley Neall, Tong Wu, Shengqu Cai, and Gordon Wetzstein. Generated reality: Human-centric world simulation using interactive video generation with hand and camera control.arXiv preprint arXiv:2602.18422, 2026
arXiv 2026
-
[17]
Tinghui Zhou, Richard Tucker, John Flynn, Graham Fyffe, and Noah Snavely. Stereo magni- fication: Learning view synthesis using multiplane images.arXiv preprint arXiv:1805.09817, 2018
Pith/arXiv arXiv 2018
-
[18]
Dl3dv-10k: A large-scale scene dataset for deep learning-based 3d vision
Lu Ling, Yichen Sheng, Zhi Tu, Wentian Zhao, Cheng Xin, Kun Wan, Lantao Yu, Qianyu Guo, Zixun Yu, Yawen Lu, et al. Dl3dv-10k: A large-scale scene dataset for deep learning-based 3d vision. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 22160–22169, 2024
2024
-
[19]
Jianjie Fang, Yingshan Lei, Qin Wan, Ziyou Wang, Yuchao Huang, Yongyan Xu, Baining Zhao, Weichen Zhang, Chen Gao, Xinlei Chen, and Yong Li. iworld-bench: A benchmark for interactive world models with a unified action generation framework.arXiv preprint arXiv:2605.03941, 2026
Pith/arXiv arXiv 2026
-
[20]
Sekai: A video dataset towards world exploration.arXiv preprint arXiv:2506.15675, 2025
Zhen Li, Chuanhao Li, Xiaofeng Mao, Shaoheng Lin, Ming Li, Shitian Zhao, Zhaopan Xu, Xinyue Li, Yukang Feng, Jianwen Sun, et al. Sekai: A video dataset towards world exploration.arXiv preprint arXiv:2506.15675, 2025
arXiv 2025
-
[21]
Jiahao Wang, Yufeng Yuan, Rujie Zheng, Youtian Lin, Jian Gao, Lin-Zhuo Chen, Yajie Bao, Yi Zhang, Chang Zeng, Yanxi Zhou, et al. Spatialvid: A large-scale video dataset with spatial annotations.arXiv preprint arXiv:2509.09676, 2025
arXiv 2025
-
[22]
Vipe: Video pose engine for 3d geometric perception.arXiv preprint arXiv:2508.10934, 2025
Jiahui Huang, Qunjie Zhou, Hesam Rabeti, Aleksandr Korovko, Huan Ling, Xuanchi Ren, Tianchang Shen, Jun Gao, Dmitry Slepichev, Chen-Hsuan Lin, et al. Vipe: Video pose engine for 3d geometric perception.arXiv preprint arXiv:2508.10934, 2025
Pith/arXiv arXiv 2025
-
[23]
Robotwin: Dual-arm robot benchmark with generative digital twins
Yao Mu, Tianxing Chen, Zanxin Chen, Shijia Peng, Zhiqian Lan, Zeyu Gao, Zhixuan Liang, Qiaojun Yu, Yude Zou, Mingkun Xu, et al. Robotwin: Dual-arm robot benchmark with generative digital twins. InProceedings of the computer vision and pattern recognition conference, pages 27649–27660, 2025
2025
-
[24]
Tianxing Chen, Zanxin Chen, Baijun Chen, Zijian Cai, Yibin Liu, Zixuan Li, Qiwei Liang, Xianliang Lin, Yiheng Ge, Zhenyu Gu, et al. Robotwin 2.0: A scalable data generator and benchmark with strong domain randomization for robust bimanual robotic manipulation.arXiv preprint arXiv:2506.18088, 2025
Pith/arXiv arXiv 2025
-
[25]
Seedance 2.0: Advancing video generation for world complexity.arXiv preprint arXiv:2604.14148, 2026
Team Seedance, De Chen, Liyang Chen, Xin Chen, Ying Chen, Zhuo Chen, Zhuowei Chen, Feng Cheng, Tianheng Cheng, Yufeng Cheng, et al. Seedance 2.0: Advancing video generation for world complexity.arXiv preprint arXiv:2604.14148, 2026
Pith/arXiv arXiv 2026
-
[26]
Qwen3.6-Plus: Towards real world agents, April 2026
Qwen Team. Qwen3.6-Plus: Towards real world agents, April 2026. URLhttps://qwen.ai/ blog?id=qwen3.6
2026
-
[27]
Bo Liu, Yifeng Zhu, Chongkai Gao, Yihao Feng, Qiang Liu, Yuke Zhu, and Peter Stone. Libero: Benchmarking knowledge transfer for lifelong robot learning.arXiv preprint arXiv:2306.03310, 2023
Pith/arXiv arXiv 2023
-
[28]
Ryan Hoque, Peide Huang, David J Yoon, Mouli Sivapurapu, and Jian Zhang. Egodex: Learning dexterous manipulation from large-scale egocentric video.arXiv preprint arXiv:2505.11709, 2025. 15
Pith/arXiv arXiv 2025
-
[29]
Ego4d: Around the world in 3,000 hours of egocentric video
Kristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang, Miao Liu, Xingyu Liu, et al. Ego4d: Around the world in 3,000 hours of egocentric video. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 18995–19012, 2022
2022
-
[30]
Reconstructing hands in 3d with transformers
Georgios Pavlakos, Dandan Shan, Ilija Radosavovic, Angjoo Kanazawa, David Fouhey, and Jitendra Malik. Reconstructing hands in 3d with transformers. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9826–9836, 2024
2024
-
[31]
Hassan Abu Alhaija, Jose Alvarez, Maciej Bala, Tiffany Cai, Tianshi Cao, Liz Cha, Joshua Chen, Mike Chen, Francesco Ferroni, Sanja Fidler, et al. Cosmos-transfer1: Conditional world generation with adaptive multimodal control.arXiv preprint arXiv:2503.14492, 2025
Pith/arXiv arXiv 2025
-
[32]
Cameractrl: Enabling camera control for video diffusion models
Hao He, Yinghao Xu, Yuwei Guo, Gordon Wetzstein, Bo Dai, Hongsheng Li, and Ceyuan Yang. Cameractrl: Enabling camera control for video diffusion models. InThe Thirteenth International Conference on Learning Representations, 2025
2025
-
[33]
Motionctrl: A unified and flexible motion controller for video generation
Zhouxia Wang, Ziyang Yuan, Xintao Wang, Yaowei Li, Tianshui Chen, Menghan Xia, Ping Luo, and Ying Shan. Motionctrl: A unified and flexible motion controller for video generation. In ACM SIGGRAPH 2024 Conference Papers, pages 1–11, 2024
2024
-
[34]
Cami2v: Camera-controlled image-to-video diffusion model.arXiv preprint arXiv:2410.15957, 2024
Guangcong Zheng, Teng Li, Rui Jiang, Yehao Lu, Tao Wu, and Xi Li. Cami2v: Camera-controlled image-to-video diffusion model.arXiv preprint arXiv:2410.15957, 2024
Pith/arXiv arXiv 2024
-
[35]
Realcam-i2v: Real-world image-to-video generation with interactive complex camera control
Teng Li, Guangcong Zheng, Rui Jiang, Shuigen Zhan, Tao Wu, Yehao Lu, Yining Lin, Chuanyun Deng, Yepan Xiong, Min Chen, et al. Realcam-i2v: Real-world image-to-video generation with interactive complex camera control. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 28785–28796, 2025
2025
-
[36]
Videox-fun: A flexible framework for video generation at any resolution
AIGC-Apps and Alibaba PAI Team. Videox-fun: A flexible framework for video generation at any resolution. https://github.com/aigc-apps/VideoX-Fun, 2024. Open-source project / technical framework
2024
-
[37]
Ac3d: Analyzing and improving 3d camera control in video diffusion transformers
Sherwin Bahmani, Ivan Skorokhodov, Guocheng Qian, Aliaksandr Siarohin, Willi Menapace, Andrea Tagliasacchi, David B Lindell, and Sergey Tulyakov. Ac3d: Analyzing and improving 3d camera control in video diffusion transformers. InProceedings of the Computer Vision and Pattern Recognition Conference, pages 22875–22889, 2025
2025
-
[38]
Yuang Zhang, Jiaxi Gu, Li-Wen Wang, Han Wang, Junqi Cheng, Yuefeng Zhu, and Fangyuan Zou. Mimicmotion: High-quality human motion video generation with confidence-aware pose guidance.arXiv preprint arXiv:2406.19680, 2024
Pith/arXiv arXiv 2024
-
[39]
Di Chang, Yichun Shi, Quankai Gao, Jessica Fu, Hongyi Xu, Guoxian Song, Qing Yan, Yizhe Zhu, Xiao Yang, and Mohammad Soleymani. Magicpose: Realistic human poses and facial expressions retargeting with identity-aware diffusion.arXiv preprint arXiv:2311.12052, 2023
Pith/arXiv arXiv 2023
-
[40]
QuankaiGao, JiaweiYang, QiangengXu, LeChen, andYueWang. Lome: Learninghuman-object manipulation with action-conditioned egocentric world model.arXiv preprint arXiv:2603.27449, 2026
arXiv 2026
-
[41]
Adaptive mixtures of local experts.Neural computation, 3(1):79–87, 1991
Robert A Jacobs, Michael I Jordan, Steven J Nowlan, and Geoffrey E Hinton. Adaptive mixtures of local experts.Neural computation, 3(1):79–87, 1991. 16
1991
-
[42]
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean. Outrageously large neural networks: The sparsely-gated mixture-of-experts layer. InInternational Conference on Learning Representations, 2017. doi: 1701.06538. URL https://arxiv.org/abs/1701.06538
Pith/arXiv arXiv 2017
-
[43]
Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen, Orhan Firat, Yanping Huang, Maxim Krikun, Noam Shazeer, and Zhifeng Chen. Gshard: Scaling giant models with conditional computation and automatic sharding.arXiv preprint arXiv:2006.16668, 2020
Pith/arXiv arXiv 2006
-
[44]
William Fedus, Barret Zoph, and Noam Shazeer. Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity.arXiv preprint arXiv:2101.03961, 2022
Pith/arXiv arXiv 2022
-
[45]
Damai Dai, Chengqi Deng, Chenggang Zhao, R. X. Xu, Huazuo Gao, Deli Chen, Jiashi Li, Wangding Zeng, Xingkai Yu, Y. Wu, Zhenda Xie, Y. K. Li, Panpan Huang, Fuli Luo, Chong Ruan, Zhifang Sui, and Wenfeng Liang. Deepseekmoe: Towards ultimate expert specialization in mixture-of-experts language models.arXiv preprint arXiv:2401.06066, 2024
Pith/arXiv arXiv 2024
-
[46]
Moe-llava: Mixture of experts for large vision-language models.arXiv preprint arXiv:2401.15947, 2024
Bin Lin, Zhenyu Tang, Yang Ye, Jinfa Huang, Junwu Zhang, Yatian Pang, Peng Jin, Munan Ning, Jiebo Luo, and Li Yuan. Moe-llava: Mixture of experts for large vision-language models.arXiv preprint arXiv:2401.15947, 2024
Pith/arXiv arXiv 2024
-
[47]
Zhiyu Wu, Xiaokang Chen, Zizheng Pan, Xingchao Liu, Wen Liu, Damai Dai, Huazuo Gao, Yiyang Ma, Chengyue Wu, Bingxuan Wang, Zhenda Xie, Yu Wu, Kai Hu, Jiawei Wang, Yaofeng Sun, Yukun Li, Yishi Piao, Kang Guan, Aixin Liu, Xin Xie, Yuxiang You, Kai Dong, Xingkai Yu, Haowei Zhang, Liang Zhao, Yisong Wang, and Chong Ruan. Deepseek-vl2: Mixture-of-experts visio...
Pith/arXiv arXiv 2024
-
[48]
Nucleus-image: Sparse moe for image generation.arXiv preprint arXiv:2604.12163, 2026
Chandan Akiti, Ajay Modukuri, Murali Nandan Nagarapu, Gunavardhan Akiti, and Haozhe Liu. Nucleus-image: Sparse moe for image generation.arXiv preprint arXiv:2604.12163, 2026
Pith/arXiv arXiv 2026
-
[49]
Yixuan Zhu, Jiaqi Feng, Wenzhao Zheng, Yuan Gao, Xin Tao, Pengfei Wan, Jie Zhou, and Jiwen Lu. Astra: General interactive world model with autoregressive denoising.arXiv preprint arXiv:2512.08931, 2026. URLhttps://arxiv.org/abs/2512.08931
arXiv 2026
-
[50]
Gang Cheng, Xin Gao, Li Hu, Siqi Hu, Mingyang Huang, Chaonan Ji, Ju Li, Dechao Meng, Jinwei Qi, Penchong Qiao, Zhen Shen, Yafei Song, Ke Sun, Linrui Tian, Feng Wang, Guangyuan Wang, Qi Wang, Zhongjian Wang, Jiayu Xiao, Sheng Xu, Bang Zhang, Peng Zhang, Xindi Zhang, Zhe Zhang, Jingren Zhou, and Lian Zhuo. Wan-animate: Unified character animation and replac...
arXiv 2025
-
[51]
Javier Romero, Dimitrios Tzionas, and Michael J. Black. Embodied hands: Modeling and capturing hands and bodies together.ACM Transactions on Graphics, 36(6), 2017. 17 A Broader Impacts This work may have several positive broader impacts. By unifying heterogeneous control modalities within a single world model, Worldscape-MoE provides a scalable framework ...
2017
-
[52]
is a fully Transformer-based method for hand mesh recovery. It uses a large-scale Vision 21 Transformer backbone and a Transformer decoder to regress hand parameters from a monocular RGB image. The method adopts the MANO parametric hand model [51] to represent hand geometry. Given a hand crop𝐼𝑡, HaMeR predicts the MANO pose parameters, shape parameters, a...
arXiv 2083
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.