Pith. sign in

Paper Citation Record · LEDGER

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence

As of 5 August 2026, this Paper Citation Record lists 100 of 131 outbound references and 4 inbound Pith citation observations for arXiv:2607.07675.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.07675 v1

Coverage vector

measured 100 of 131 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-09T02:55:58.018234Z

measured 104 of 104 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T07:53:39.262642Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T04:16:48.701434Z

Reference resolution

100 of 131 outbound references displayed

  • verified exact42
  • verified fuzzy54
  • unresolved0
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 70853329-5a65-4852-a42a-03772111d8dd · outbound

This paper cites GPT-4 Technical Report.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence GPT-4 Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-09T03:05:55.225604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:eab1c88c1af67c075e9b5022d5b6330778406fa63b1a1f3997b7fc2fc6b31dff

Observation ad2ace2e-2a21-4fd2-b5d8-120e307a1492 · outbound

This paper cites Cosmos 3: Omnimodal World Models for Physical AI.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Cosmos 3: Omnimodal World Models for Physical AI

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-09T03:05:55.157813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:b2242df9e3320e408ca3ac9f1a9db36f6f50b5282d47dd845adca7b9ed4940b1

Observation 00c84fef-d6df-4237-8229-0a35e6d9ae50 · outbound

This paper cites Pytorch 2: Faster machine learning through dynamic python bytecode transformation and graph compilation.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Pytorch 2: Faster machine learning through dynamic python bytecode transformation and graph compilation

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T03:05:55.727547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:76e14350f19ef207878f995474cdc79239c08a33baae1dca1e69f9c6b2066a6a

Observation 96a0c4da-5074-4812-965a-8512c139310b · outbound

This paper cites V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-09T03:05:55.138662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:9be30aa13cda53ecb73d0a82fa158d7797255eeb599e9006cb6cb47d157a31a3

Observation 790734dc-2f3f-40e6-9758-20a24e3e9b82 · outbound

This paper cites Qwen Technical Report.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Qwen Technical Report

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-09T03:05:55.345416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:b97017604afe99723d4ba73278d2c9d2bac5cbfc3f68c61afe0d8422c5adb076

Observation b2073d22-0f88-49c7-9e87-21da66fe9cdd · outbound

This paper cites Qwen3-VL Technical Report.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Qwen3-VL Technical Report

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-09T03:05:55.250167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:3fe4823518823fbd9065f840d364e32a13329b445c9ac4f221bd41611706a764

Observation 81243553-b743-49d7-bb9a-3133304947e3 · outbound

This paper cites Lumiere: A space-time diffusion model for video generation.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Lumiere: A space-time diffusion model for video generation

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T03:05:55.764181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:9ebc34f382741630fa010543457e5aad8143b567360082be9f982f8d17239a4a

Observation 78e0aa2b-8e95-4a3e-af79-48401eafa4b4 · outbound

This paper cites DeepSeek LLM: Scaling Open-Source Language Models with Longtermism.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence DeepSeek LLM: Scaling Open-Source Language Models with Longtermism

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-09T03:05:55.314557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:9ac2902152c8cad7fdee92ad4ae444f2f69bf2bd4f541835c4d2553e76c6cb8f

Observation a2a0244b-024c-42ad-a3ed-ae2f800a8981 · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-09T03:05:55.206327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:271e33aa83dade4e813dd922325b62cf46d2693b6894b3548c0ca33e7e7139af

Observation 8f3ac1ac-3e91-4896-8cb6-3417118b3c05 · outbound

This paper cites Genie: Generative interactive environments.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Genie: Generative interactive environments

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T03:05:55.690954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:874a2a9f69c6d4698f61eca8714a9998a235373cb060491b79bf03cf1b15784f

Observation 95ef248b-fb5e-48f7-91f7-c950704d2ed8 · outbound

This paper cites Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-09T03:05:55.311053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:bc132c1c64cd7166609d58d76300c4921c0228647988954f8717cc61dc909dd3

Observation d4c12d5c-89ba-4ff0-a00e-4c9cee1c9fce · outbound

This paper cites Longcat-video technical report.arXiv preprint arXiv:2510.22200.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Longcat-video technical report.arXiv preprint arXiv:2510.22200

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-09T03:05:55.255582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:64fc64820c924d668fbdb69a088638ac4f327e91b57327e4d6958b95079a71dc

Observation b94d3e9f-2e63-40d2-827b-f1cfcb88fb3b · outbound

This paper cites Pixart- alpha: Fast training of diffusion transformer for photorealistic text-to-image synthesis.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Pixart- alpha: Fast training of diffusion transformer for photorealistic text-to-image synthesis

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T03:05:55.869506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:4ff78cfaac038691328f892b420fa41f4047cc7d1a29c17be39b73de5d8db507

Observation 31900569-fe83-41c8-9184-883a84f9f380 · outbound

This paper cites Training Deep Nets with Sublinear Memory Cost.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Training Deep Nets with Sublinear Memory Cost

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-09T03:05:55.263611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:1d661fe5dbfa92bc8524b0afb3ee6ec1ff848d258d1461a05ef0dc7b8e2c9618

Observation 3a7e8a4e-fd38-4800-9a60-48a79854cdf1 · outbound

This paper cites Realdpo: Real or not real, that is the preference.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Realdpo: Real or not real, that is the preference

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-09T03:05:55.255294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:33eca352d67fd182fc229e61a4b28d6cf6f658ebfed391fe4444d4ca52345c28

Observation 3d6dec75-acb6-4329-bfd9-ab782929c6f7 · outbound

This paper cites Local all-pair correspondence for point tracking.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Local all-pair correspondence for point tracking

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T03:05:55.871387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:8182d6485b26b00579759729818eccba133bc34c59057bce71562e0b76275023

Observation 39e2ad2d-15b6-4305-b0b0-c84e6e515996 · outbound

This paper cites Palm: Scaling language modeling with pathways.Journal of machine learning research, 24(240):1–113.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Palm: Scaling language modeling with pathways.Journal of machine learning research, 24(240):1–113

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T03:05:55.859303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:3b94b5613c4c2d7f5a3a88c5ce2b8ec715593a9dd65d753dc868fa37cb146fd6

Observation b7ae041d-9c39-407a-a9fe-696665c95d96 · outbound

This paper cites Deepseekmoe: Towards ultimate expert specialization in mixture-of-experts language models.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Deepseekmoe: Towards ultimate expert specialization in mixture-of-experts language models

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T03:05:55.854259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:f20dfc11c90e45f9407845369b73d300de021342355fd0501c335f02694ada2c

Observation 2cff44b4-ed21-4aaa-8474-9e300a5d59af · outbound

This paper cites Fu, Stefano Ermon, Atri Rudra, and Christopher Ré.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Fu, Stefano Ermon, Atri Rudra, and Christopher Ré

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T03:05:55.857277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:f42dc0438ce9c13a0957392a4521e3bd2ee17efc4b9765002268119c16888815

Observation 361bc574-79af-4a30-a5ca-63a06608670f · outbound

This paper cites Deepep: An efficient expert-parallel communication library.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Deepep: An efficient expert-parallel communication library

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T03:05:55.861339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:b10a00e21c8f6c2f23693ff36426fb1401bc2b017a85d80df8357dec27dc1c2f

Observation 84b6211c-a79b-4471-a909-5d5e15156e8e · outbound

This paper cites an unresolved cited work.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Unresolved cited work

Reference 21

Resolution
parse uncertain
raw_fallback, observed 2026-07-09T03:05:55.873149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:ac0d80a711b44161fe82d276e68b5fba04975b82769470a4da3253163afd506d

Observation 3f10346e-3d51-4593-a139-84fc97169474 · outbound

This paper cites Scaling vision transformers to 22 billion parameters.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Scaling vision transformers to 22 billion parameters

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T03:05:55.876791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:86840f75a1a1c6d00a4ffd3dda51152db7fba5d5a9d5aa8ab2adf329820d44ec

Observation ff20bb95-561d-4a7e-8693-9672bdf59099 · outbound

This paper cites Rethinking video generation model for the embodied world.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Rethinking video generation model for the embodied world

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T03:05:55.849007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:c5eac21c379ce2756444748106ac11d2eeabe8d38c5c1b8d37850d71f7f9bced

Observation 751852de-e464-4e8b-ac62-1b7b76841d70 · outbound

This paper cites Structure and content- guided video synthesis with diffusion models.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Structure and content- guided video synthesis with diffusion models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T03:05:55.839376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:929c9f0c16315c1597eb405478df50c82f698795b4242cacbfaeb30f02dcaaef

Observation 0b9319d0-9ca9-4227-8d00-6f52d6408626 · outbound

This paper cites Scaling rectified flow transformers for high-resolution image synthesis.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Scaling rectified flow transformers for high-resolution image synthesis

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T03:05:55.836466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:040b204c7aa7e3d6d788efc526f93781fe8768dedf7365c0bfe7ad2e2e9a8456

Observation 24c48c8d-ee24-4d31-9fe7-7aee347baa19 · outbound

This paper cites Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity.Journal of Machine Learning Research, 23(120):1–39.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity.Journal of Machine Learning Research, 23(120):1–39

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T03:05:55.836894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:2526f7f11767db742dd9c88dae2b5f9cadc179d872e4ff53bd0492b1e94bc42b

Observation e55f2f81-05ef-4fe7-8e12-5e194e8ba58e · outbound

This paper cites DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-07-09T03:05:55.287932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:47696c28afb3e975d631a8bff7a85a8e3dba566e65c80a51c52722859137706a

Observation 8d650eef-2cf5-4fd0-812c-1f61edeaa5a8 · outbound

This paper cites The pulse of motion: Measuring physical frame rate from visual dynamics.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence The pulse of motion: Measuring physical frame rate from visual dynamics

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-07-09T03:05:55.274549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:f2eb1dbaea476481a53842a1f8ece57533559b9052951784be0bd8d42d8261e3

Observation 9fb154c2-6d77-4b74-b654-2aa890816496 · outbound

This paper cites Vlaw: Iterative co-improvement of vision-language-action policy and world model.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Vlaw: Iterative co-improvement of vision-language-action policy and world model

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-09T03:05:55.348812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:055ba0725a6f2b7cd8e4643b97dafb2219d4c0058a45f802409d0d05f6a3deef

Observation 0bdb0a86-22fd-478f-96a3-5836865a55e5 · outbound

This paper cites Ctrl-World: A Controllable Generative World Model for Robot Manipulation.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Ctrl-World: A Controllable Generative World Model for Robot Manipulation

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-07-09T03:05:55.295756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:2fb53e0644672dd32596cbc154417168f647207a78e4379e961785e733f7ace6

Observation 5acb3f5d-73af-468a-af59-ce45d6b02848 · outbound

This paper cites OmniAID: Decoupling Semantics and Artifacts for Universal AI-Generated Image Detection in the Wild.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence OmniAID: Decoupling Semantics and Artifacts for Universal AI-Generated Image Detection in the Wild

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-07-09T03:05:55.361324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:4d1bd372f7be8b632bbfb3cd635d0a5a8401f13d8d3c2d34f454bf7b5ca05616

Observation 41247c9b-e748-4939-9afe-8a5c203df4ef · outbound

This paper cites Animatediff: Animate your personalized text-to-image diffusion models without specific tuning.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Animatediff: Animate your personalized text-to-image diffusion models without specific tuning

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T03:05:55.807674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:e79b5eedbef97dfe62fdaaaa67ad93fdf2bf17f22b568c9a06177947f488d95a

Observation b4dbe36b-8ed4-4402-bbd1-d37a5fccfdce · outbound

This paper cites Photorealistic video generation with diffusion models.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Photorealistic video generation with diffusion models

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T03:05:55.804135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:4e8a843716e15dd89533d7c510c1239b96d04385c3b3c2d0360984066cd25f31

Observation 75f559d0-8341-45aa-a9a8-0e1b63c88c6a · outbound

This paper cites Generating an image from 1,000 words: Enhancing text-to-image with structured captions.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Generating an image from 1,000 words: Enhancing text-to-image with structured captions

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-07-09T03:05:55.270798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:4678241a61d83415107618353a00da766c8dde7af1ea18c0684088a568ae3a89

Observation 81e1d7e4-80b8-425c-88b0-5c90c330cae9 · outbound

This paper cites Flash-GRPO: Efficient Alignment for Video Diffusion via One-Step Policy Optimization.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Flash-GRPO: Efficient Alignment for Video Diffusion via One-Step Policy Optimization

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-07-09T03:05:55.295529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:839d50f4e291e31b1af689027bde05224beeaf0111fe903b52ed23545c97f775

Observation 88fd76ff-7fe0-4edb-839d-93ac2586b840 · outbound

This paper cites TempFlow-GRPO: When Timing Matters for GRPO in Flow Models.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence TempFlow-GRPO: When Timing Matters for GRPO in Flow Models

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-07-09T03:05:55.314045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:020a441e1a848bce36721718625b37da9dbeb6e38892b5d9ef6076dbc252e648

Observation bc549c17-ce48-4bf8-b2aa-111d831738f1 · outbound

This paper cites Imagen Video: High Definition Video Generation with Diffusion Models.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Imagen Video: High Definition Video Generation with Diffusion Models

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-07-09T03:05:55.336407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:d2fb0c7a7f6727c1f4724fa8e45d35a53ccaa67d613aa91293d3b57ce82e15a5

Observation 689ba7c5-0767-467f-a07d-7ea3db82e036 · outbound

This paper cites Video diffusion models.Advances in neural information processing systems, 35:8633–8646.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Video diffusion models.Advances in neural information processing systems, 35:8633–8646

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T03:05:55.831622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:546aa46bcd36014b4113194ae7e0f9fa63d974f5f6741ac0e0627c0d4472272a

Observation 57e26a9b-e1f4-4288-9a4f-621ca83404ab · outbound

This paper cites Rae, Oriol Vinyals, and Laurent Sifre.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Rae, Oriol Vinyals, and Laurent Sifre

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T03:05:55.851486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:34a2caaecd19c0d148928625d4499a6135c9fb5847ec400b0cde08c789a83d98

Observation 9dd81e01-b8af-4197-81d7-3f8557cb4add · outbound

This paper cites Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T03:05:55.844301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:8b4f7559d1c14b7a9e80d1c11705a59806da323aa51c9ce6b6c522dd7a20e70c

Observation 56cc5ead-2d39-44be-9093-e4e9708f1749 · outbound

This paper cites OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-07-09T03:05:55.340810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:113c7ba9dc6ddff74974e99dcb93600d1a60d2913d22c0105b042226d14c7ac1

Observation 11127d1f-bc22-4b80-9969-5c438155834d · outbound

This paper cites Tutel: Adaptive mixture-of-experts at scale.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Tutel: Adaptive mixture-of-experts at scale

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T03:05:55.880891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:85b8609c373db6a4752a56dacceb9d993f28b06aea92a566254a5a33eaa6b637

Observation 9f6d8b17-da7d-412c-80d6-a3ab8ef9d57a · outbound

This paper cites Jacobs, Michael I.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Jacobs, Michael I

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T03:05:55.784513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:f7192f2c5d245be3c67d248b345e1e08242bcf70f3436b793dcc2bbaaf9dfaa6

Observation 2cbc8996-b7ae-432a-acea-528bc3fff068 · outbound

This paper cites DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-07-09T03:05:55.170422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:e626266ca0e2de4522901eb3d8eea40242d712f70ffa67ccd4b6251de2070939

Observation 4fc0cfcb-86c5-456e-805f-1171b3c1e9a9 · outbound

This paper cites WoVR: World Models as Reliable Simulators for Post-Training VLA Policies with RL.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence WoVR: World Models as Reliable Simulators for Post-Training VLA Policies with RL

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-07-09T03:05:55.347531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:b620b85ddf0695d65794646885f28791ebbbccd7c72b330ac1b0b6e23f3c162b

Observation 0ce3b01e-01ee-406c-8090-fcf3a2469eac · outbound

This paper cites Hierarchical mixtures of experts and the em algorithm.Neural computation, 6(2):181– 214, 1994.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Hierarchical mixtures of experts and the em algorithm.Neural computation, 6(2):181– 214, 1994

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T03:05:55.756750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:e37e780ca7aeb2ef27acd0e99b9fa20fafd17eaf1c8b5a6b69e1568cf71365e2

Observation 92a09cda-2931-4fa5-9a7e-aaa0f977994c · outbound

This paper cites Scaling Laws for Neural Language Models.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Scaling Laws for Neural Language Models

Reference 47

Resolution
metadata mismatch
local_arxiv, observed 2026-07-09T03:05:55.200980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:d579fffff3d14f667552f7b7cec49828908eecfa64597174a625171236544aee

Observation 394d93db-e316-430d-b94a-dc8a9f3e1955 · outbound

This paper cites Pick-a-pic: An open dataset of user preferences for text-to-image generation.Advances in neural information processing systems, 36:36652–36663.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Pick-a-pic: An open dataset of user preferences for text-to-image generation.Advances in neural information processing systems, 36:36652–36663

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T03:05:55.761087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:dc3215ec77553d163fb4747029075ca1a5412983dc30c078af24736a67a79623

Observation fb7231a6-0d13-4315-b6a0-535134fd160d · outbound

This paper cites Reducing activation recomputation in large transformer models.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Reducing activation recomputation in large transformer models

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T03:05:55.775761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:12893be9377716a76f3d8a2638cad8a818f07ec8726123813acddee9c98e7732

Observation 31adb489-bf50-4f65-b85c-c432ba8c3717 · outbound

This paper cites Efficient Sequence Packing without Cross-contamination: Accelerating Large Language Models without Impacting Performance.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Efficient Sequence Packing without Cross-contamination: Accelerating Large Language Models without Impacting Performance

Reference 50

Resolution
metadata mismatch
local_arxiv, observed 2026-07-09T03:05:55.225351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:62d4222b705ca7fe45451061540ba52092ed133d3c6bb6cd6b8c1df63459a8b0

Observation 6a38081c-64f9-42e9-939b-50f0f509036c · outbound

This paper cites Gshard: Scaling giant models with conditional computation and automatic sharding.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Gshard: Scaling giant models with conditional computation and automatic sharding

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T03:05:55.791386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:1c725881656c94f74d6334a4f1b366ccff3b3a32fe0ac43b660300e0406814bc

Observation eeef1473-3eaf-4837-88a0-1ef8999344ec · outbound

This paper cites MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-07-09T03:05:55.325140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:a950cbd48d34a474b125e7a681237f46881358a244b1b47b270e7510906b5a98

Observation c7c95d25-a18b-41a9-9f97-59a0d7a6139b · outbound

This paper cites Causal World Modeling for Robot Control.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Causal World Modeling for Robot Control

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-07-09T03:05:55.339484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:7edadb366a9f300cfd93708159dc4a5bac510386af0c442d47e00985b7f4fbce

Observation c4cbe177-a852-4d45-9eca-d3a27456b067 · outbound

This paper cites Pytorch distributed: Experiences on accelerating data parallel training.Proceedings of the VLDB Endowment, 13(12):3005–3018, 2020.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Pytorch distributed: Experiences on accelerating data parallel training.Proceedings of the VLDB Endowment, 13(12):3005–3018, 2020

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T03:05:55.794858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:07152c263bae50dc50dfa078944d9248d348c0ce7513b53d02ee1a8b729c5b43

Observation 40d7770b-4240-4ce4-a229-93bb23617050 · outbound

This paper cites Torchtitan: One-stop pytorch native solution for production ready LLM pretraining.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Torchtitan: One-stop pytorch native solution for production ready LLM pretraining

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T03:05:55.738988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:9af6f6d11f1f426b81af2e3e96dea5d5b49c0dc5dec7bb8a3fabe7af7bd1ffb8

Observation 473c2d7a-4c49-488f-80d6-8945664f9613 · outbound

This paper cites Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-07-09T03:05:55.333017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:0c3a27bd86e018c4e33755cbefd9072b40846c4866aae826d65007129910a830

Observation 464c14f7-12a4-4a2f-9696-f95fba431445 · outbound

This paper cites Flow matching for generative modeling.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Flow matching for generative modeling

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T03:05:55.736148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:278d04680d12acddca020997fc8aff3dbc35834d6b550490dba1c344532e81d7

Observation d6ff8eda-bb2f-446a-ac7b-fe0605904ecf · outbound

This paper cites DeepSeek-V3 Technical Report.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence DeepSeek-V3 Technical Report

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-07-09T03:05:55.318622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:0aeceeca4f28ec82bcc5d1f94a2e371bc32f8020e5768e55e5d9f178c3436134

Observation fcb022fb-6e5d-4feb-9d85-acc5e2a9d5d4 · outbound

This paper cites Flow-grpo: Training flow matching models via online rl.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Flow-grpo: Training flow matching models via online rl

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T03:05:55.732991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:3b280c01fa1da9117a4e01846ea13872876460a0669bd0e1049b31966d5dc3c3

Observation 11334e03-e355-4ae6-924e-e86d06c2c4c8 · outbound

This paper cites GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-07-09T03:05:55.271067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:7f20b253291c10b812e591571e8cae89c1faeb35e4c6c2038c8e612006622c90

Observation 52239194-9700-4306-9bd4-5cc0c812d290 · outbound

This paper cites Flow straight and fast: Learning to generate and transfer data with rectified flow.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Flow straight and fast: Learning to generate and transfer data with rectified flow

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T03:05:55.889150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:5a8628d1cf630c6a40c584e930e1fb3e8daa52fc68c10d1c475b7f4a1b80c262

Observation e232313d-ae74-41ba-85ca-1aa9ac3cfe1e · outbound

This paper cites VeOmni: Scaling Any Modality Model Training with Model-Centric Distributed Recipe Zoo.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence VeOmni: Scaling Any Modality Model Training with Model-Centric Distributed Recipe Zoo

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-07-09T03:05:55.350703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:575c93092dfe2cd3cf91bd3cb00fd9a89550a9accf18df8a9aeec1d4975ac80b

Observation 9c5523d9-31bb-45bc-a636-bb086c6a3f6a · outbound

This paper cites Hpsv3: Towards wide-spectrum human preference score.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Hpsv3: Towards wide-spectrum human preference score

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T03:05:55.724491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:e24748aa08cee887251c607ba877c1022d52473a574d8d47bbb99b688ec29ca3

Observation d4824ef0-bd1d-40e3-849c-37499dde4a93 · outbound

This paper cites Real: Efficient rlhf training of large language models with parameter reallocation.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Real: Efficient rlhf training of large language models with parameter reallocation

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T03:05:55.727250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:e0308cfd1e0147a332c584a2cddafab9de5bf59612805417abec21ec08c783d9

Observation f58a3575-7986-45ba-91bf-c53d208e01e2 · outbound

This paper cites Ray: A distributed framework for emerging {AI} applications.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Ray: A distributed framework for emerging {AI} applications

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T03:05:55.750194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:81ee1331f560f63444c5e7ef5f175ac01da6fb0aec794dd700de3abbadea4ed0

Observation bc37b517-4919-4e5d-9f5f-d77b863bd67a · outbound

This paper cites Do generative video models understand physical principles? InProceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 2026.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Do generative video models understand physical principles? InProceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 2026

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T03:05:55.717244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:c08f200380b90daab9ecafe4d279daa543435dada896f97397571bae3935ca84

Observation 968dffcc-9434-4bda-8baf-a5f669c945b4 · outbound

This paper cites Efficient large-scale language model training on gpu clusters using megatron-lm.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Efficient large-scale language model training on gpu clusters using megatron-lm

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T03:05:55.851883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:1c1cb1e9fd41079e3ee5056091a8877088aa10651ade49ded9394586bda8d40c

Observation 09dac93e-8f76-4794-8acf-32225fee767f · outbound

This paper cites Nowlan and Geoffrey E.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Nowlan and Geoffrey E

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T03:05:55.816957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:4a70984f2ae58b9e7f4f7dc3450ccc83377ebafb36a3d500b46b551a00d4f889

Observation aec767cb-602e-4f70-90bc-7685a2c7220f · outbound

This paper cites The pagerank citation ranking : Bringing order to the web.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence The pagerank citation ranking : Bringing order to the web

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T03:05:55.741486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:f80101b48ef6dcbf526dce983697530dc70343b2459efec9723afce02ff9f4d1

Observation d1773c59-9cad-413b-9689-f6e09ad96be3 · outbound

This paper cites Switch diffusion transformer: Synergizing denoising tasks with sparse mixture-of-experts.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Switch diffusion transformer: Synergizing denoising tasks with sparse mixture-of-experts

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T03:05:55.792609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:ec4f60b111120f3c81fe35f9bff220c5d0ce51b4576493a469673fcb437a717c

Observation 2be8db7f-9845-4cdf-9b24-72d192a2f82e · outbound

This paper cites Scalable diffusion models with transformers.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Scalable diffusion models with transformers

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T03:05:55.882836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:b72635352d915b8de9dccf9378f7b642073f3e315dd36143d9cb3d2f2db1fb11

Observation 84511f17-e38c-46c3-a1ff-6e671e40ce22 · outbound

This paper cites Lumina- image 2.0: A unified and efficient image generative framework.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Lumina- image 2.0: A unified and efficient image generative framework

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T03:05:55.846711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:5491008db52cdfcd16559103bee47f6d4100b2291057ad3077b19abd2d4b5128

Observation 6ad04e81-d91d-4b06-9c47-5dca10302cd9 · outbound

This paper cites Qwen3.6-27B: Flagship-level coding in a 27B dense model.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Qwen3.6-27B: Flagship-level coding in a 27B dense model

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T03:05:55.830705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:a3dfadb7079651dffd607b65c36f9c8151cd620873e45b15583b7a9f2faff57b

Observation 4bc46d9b-1a6d-4bcd-bc14-32014c801094 · outbound

This paper cites Physics-IQ Verified.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Physics-IQ Verified

Reference 74

Resolution
verified exact
local_arxiv, observed 2026-07-09T03:05:55.354852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:a3dd3a1d6a5fab0265557dd9e5dd9f875ed4ea4c292d5d2cedf131753ac5f2bd

Observation da621f34-97d6-409b-b817-23bb2e5dcb2e · outbound

This paper cites Manning, and Chelsea Finn.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Manning, and Chelsea Finn

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T03:05:55.810864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:c73b0016c9f27fd6400c14479fb76fce99949c0c8ab7fda6a45709550e668319

Observation 038b75f1-3de7-44e6-8e8c-3777eaf8cbf2 · outbound

This paper cites Deepspeed-moe: Advancing mixture-of-experts inference and training to power next-generation ai scale.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Deepspeed-moe: Advancing mixture-of-experts inference and training to power next-generation ai scale

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T03:05:55.814087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:52e37ac4f22ad5b99c573923699aea1f5a1912628c2ff7bc4ca67ceab560296c

Observation 919fc223-ef0e-4f77-af59-e31f86ed8a71 · outbound

This paper cites Zero: Memory optimizations toward training trillion parameter models.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Zero: Memory optimizations toward training trillion parameter models

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T03:05:55.817163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:768e946fa98e53f5661780ef8ccf61e88043af7f3460fcd2b48d6655856d1adc

Observation dae434a5-f7d6-44e8-9344-81cd38555028 · outbound

This paper cites Cosmos-Drive-Dreams: Scalable Synthetic Driving Data Generation with World Foundation Models.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Cosmos-Drive-Dreams: Scalable Synthetic Driving Data Generation with World Foundation Models

Reference 78

Resolution
metadata mismatch
local_arxiv, observed 2026-07-09T03:05:55.168574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:d5aaabc6532bae03f1a8fa7d27010266516c93ad2683e7d9284de5ce9a758444

Observation a650b96b-0a8a-4a9f-8ab7-24366bf6f934 · outbound

This paper cites GAIA-2: A Controllable Multi-View Generative World Model for Autonomous Driving.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence GAIA-2: A Controllable Multi-View Generative World Model for Autonomous Driving

Reference 79

Resolution
verified exact
local_arxiv, observed 2026-07-09T03:05:55.259747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:b9ffadb74f56b470909833eaa0263912c654fbafb83f73cda689222232a8a991

Observation 97ea944d-6de3-4aa4-9a62-14759002b823 · outbound

This paper cites Image super-resolution via iterative refinement.IEEE transactions on pattern analysis and machine intelligence, 45(4):4713–4726, 2022.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Image super-resolution via iterative refinement.IEEE transactions on pattern analysis and machine intelligence, 45(4):4713–4726, 2022

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T03:05:55.806491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:80323edce37aea367048e2e9e5fcb03611e9fbfc35b25938d8d8b064d8719315

Observation 38a07966-514c-4155-a6e7-fe155d515dfa · outbound

This paper cites Proximal Policy Optimization Algorithms.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Proximal Policy Optimization Algorithms

Reference 81

Resolution
verified exact
local_arxiv, observed 2026-07-09T03:05:55.322189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:9447bacb15ea39cb889889805b48683a4650310a7704a65142cfab1aa7cceb86

Observation c47deaf5-7d9f-4e98-973e-5430389a9be2 · outbound

This paper cites Seedance 2.0: Advancing Video Generation for World Complexity.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Seedance 2.0: Advancing Video Generation for World Complexity

Reference 82

Resolution
verified exact
local_arxiv, observed 2026-07-09T03:05:55.356921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:7026a539bc856a64a7b7b39d321523d928a240436115586a6abbe2da835581e1

Observation 5d982b03-7d9b-4c05-8e04-0c63c9ced34f · outbound

This paper cites Sparse Mixture-of-Experts Routing in Visual Diffusion Transformers:Diagnosis, Boundary Calibration and Evolutionary Roadmap from Routing Collapse to Selective Deadlock.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Sparse Mixture-of-Experts Routing in Visual Diffusion Transformers:Diagnosis, Boundary Calibration and Evolutionary Roadmap from Routing Collapse to Selective Deadlock

Reference 83

Resolution
verified exact
local_arxiv, observed 2026-07-09T03:05:55.352039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:b17779b18a34e1a06f315d84b8b464a60c8d974558c044c4dfd3392f2e98f360

Observation 835da1a1-d149-49c7-a65f-865026b1d31b · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 84

Resolution
verified exact
local_arxiv, observed 2026-07-09T03:05:55.353715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:6433e4fcc79d6e4535e40547ba71d40fb37b3c76bd4c061f33480718edc9dab1

Observation 2a2a4251-64d1-4448-869e-7ab78b2c0bc5 · outbound

This paper cites GLU Variants Improve Transformer.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence GLU Variants Improve Transformer

Reference 85

Resolution
verified exact
local_arxiv, observed 2026-07-09T03:05:55.184581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:49a80cfdf3ff6ca557f3c0f19d3208ab5dfc5e55698af7b58fbd20e84bd7f4f3

Observation fb08d893-8931-4f43-89d8-362c3c877578 · outbound

This paper cites Mesh-tensorflow: Deep learning for supercomputers.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Mesh-tensorflow: Deep learning for supercomputers

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T03:05:55.870274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:78bbb71094bd724891d8e200b17d13c7573f2462b41bfe2b6c9ebf68c781caba

Observation 32ae3d71-85b3-4adf-a9a1-8d7689e95057 · outbound

This paper cites Outrageously large neural networks: The sparsely-gated mixture-of-experts layer.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Outrageously large neural networks: The sparsely-gated mixture-of-experts layer

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T03:05:55.864664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:e3d6b12fee3592f034af31a7a90f25cb0306feeb000308a5e98a0660f0f2f545

Observation 1e44ba84-e90d-490b-a7b2-88690518180c · outbound

This paper cites Hybridflow: A flexible and efficient rlhf framework.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Hybridflow: A flexible and efficient rlhf framework

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T03:05:55.884832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:667fb5b5d4d7082bf0cae424a8d5c7ec33fdabb9b97e4f439b147cc7d2710906

Observation 6d976b50-22dc-484d-a0af-becea7fd26c9 · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 89

Resolution
verified exact
local_arxiv, observed 2026-07-09T03:05:55.328143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:da6fa262cdd61ea725f1015ff6852372cd6e777c7a2f8f18b16eacb8ac7def85

Observation 1d797439-c349-460b-898b-43d24b56e346 · outbound

This paper cites Make-a-video: Text-to-video generation without text-video data.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Make-a-video: Text-to-video generation without text-video data

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T03:05:55.868532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:203df7ace5edfb6334bfa9192dbd5d49a5e098254ce4aba0933a413d654284e8

Observation 8cf65560-843c-4178-95a5-a23029179f14 · outbound

This paper cites Denoising diffusion implicit models.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Denoising diffusion implicit models

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T03:05:55.875666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:1d2e029acddf33171bd9dd310c20b9cad8695562cf5efde7731521f4ce285f96

Observation d6d45c22-70df-4efe-bd4e-c9ed7ebb22cf · outbound

This paper cites Transnet v2: An effective deep network architecture for fast shot transition detection.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Transnet v2: An effective deep network architecture for fast shot transition detection

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T03:05:55.862914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:062dd76dba43feadd923f450e36ddb60ca5aeda7fd3bf1348f1db6782ff2c75f

Observation 0755ce10-6287-42c9-857e-670fedd53d33 · outbound

This paper cites WorldPlay: Towards Long-Term Geometric Consistency for Real-Time Interactive World Modeling.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence WorldPlay: Towards Long-Term Geometric Consistency for Real-Time Interactive World Modeling

Reference 93

Resolution
verified exact
local_arxiv, observed 2026-07-09T03:05:55.236141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:df9ba19e6f6f5f4e1cd8b31097d613ec3235d1ebd2b1f40325ca381b687cc89f

Observation 160d8aa8-d12f-43e4-a359-51c79c75ad8d · outbound

This paper cites Enhancing spatial understanding in image generation via reward modeling.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Enhancing spatial understanding in image generation via reward modeling

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T03:05:55.857468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:5bcc85b37d0e750dfc0f2e71626d5f6c36ad9cc98ea036dcd3e62729b73c0ed3

Observation fcc3de0e-a0fe-440d-96e2-fab7e393983c · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Gemini: A Family of Highly Capable Multimodal Models

Reference 95

Resolution
verified exact
local_arxiv, observed 2026-07-09T03:05:55.328674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:ecb0a83595634cf1cfab5c209d48795b29bc312dcf2288806222bcfe9dc7181d

Observation f74cdbbd-8267-4731-a1ee-2b41c2990fad · outbound

This paper cites Evaluating Gemini Robotics Policies in a Veo World Simulator.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Evaluating Gemini Robotics Policies in a Veo World Simulator

Reference 96

Resolution
verified exact
arxiv_id, observed 2026-07-09T03:05:55.344354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:0974e557bef97427aca7967cc19d495f9d5076a0863a79965be2a769da5bfc7e

Observation 29b8e9b5-aa8a-4167-ad5f-f5db53fff0b4 · outbound

This paper cites Advancing Open-source World Models.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Advancing Open-source World Models

Reference 97

Resolution
verified exact
local_arxiv, observed 2026-07-09T03:05:55.240536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:23d3e4749ddef40692ce3d09cefda3e30f1143bf4da1d4052b205141586b844c

Observation 5eef6e33-39c3-4207-9b22-36a51d5beee7 · outbound

This paper cites slime: An LLM post-training framework for reinforcement learning at scale.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence slime: An LLM post-training framework for reinforcement learning at scale

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T03:05:55.859462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:85b9b02eb8d684f6d8a88a3bf5b45cefbe5f338772bcbba9cecd668ee9d61438

Observation 6ee41ac9-c8d5-49ce-bebb-de466474b3b1 · outbound

This paper cites Improving and generalizing flow-based generative models with minibatch optimal transport.Transactions on Machine Learning Research, 2024.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Improving and generalizing flow-based generative models with minibatch optimal transport.Transactions on Machine Learning Research, 2024

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T03:05:55.847540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:9ed34cbcf3f7f10e314ac09088ff6331c8d77da729acd33248444ec0234391dc

Observation f36e0406-9dc6-46d4-9573-787d6d1d8b00 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence LLaMA: Open and Efficient Foundation Language Models

Reference 100

Resolution
verified exact
local_arxiv, observed 2026-07-09T03:05:55.285108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:ef1a9cc5fcd571b9012138cce50301c33d6914eadd40833407130a6d7409e5ea

Pith citing papers

Observation 8aa0e333-3d98-4862-baff-6907f5db724c · inbound

Native Video-Action Pretraining for Generalizable Robot Control cites this paper.

Native Video-Action Pretraining for Generalizable Robot Control Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence

Reference 70

Resolution
verified exact
local_arxiv, observed 2026-07-10T04:16:48.702844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-10T04:12:29.153764Z digest=sha256:cdca6caa331a7cd7407f739674578535f8d293ec29b450bc8061cf93398e2131

Observation a760df59-1c6b-4b7a-b322-088f2b59a3c9 · inbound

Native Video-Action Pretraining for Generalizable Robot Control cites this paper.

Native Video-Action Pretraining for Generalizable Robot Control Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:39.262642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:39.262642Z digest=sha256:fe8024aaf57e8dc5aff4327c7f341e54214385c8a3cbecffd286c7b58817359a

Observation 2119e194-4fbd-4a60-b3f8-fc2cb5f98de2 · inbound

SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation cites this paper.

SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T07:09:07.343135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:09:07.343135Z digest=sha256:633b9e8a189ce3cf666c033ba896b6769f460983031dc8e85c21b6175cc9ffee

Observation 1c80363e-6ece-424e-95f9-5f98349c3f9f · inbound

$N_0$-TWAM: Scaling Tactile-Native World-Action Model for Contact-Rich Manipulation cites this paper.

$N_0$-TWAM: Scaling Tactile-Native World-Action Model for Contact-Rich Manipulation Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence

Reference 55

Resolution
unresolved
no resolver link, observed 2026-07-30T12:42:18.533960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:42:18.533960Z digest=sha256:21f6aa5e624fe0895c970f6a1b3c9add1081062ad644739cb6e79b7bbf37a2c2