Pith. sign in

Paper Citation Record · LEDGER

Cosmos 3: Omnimodal World Models for Physical AI

As of 23 July 2026, this Paper Citation Record lists 15 of 15 outbound references and 27 inbound Pith citation observations for arXiv:2606.02800.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.02800 v4

Coverage vector

measured 15 of 15 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-28T15:08:33.957835Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-07-22T06:31:00.163083+00:00

measured 27 of 27 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-15T10:20:59.147440Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T12:15:01.137692Z

Reference resolution

15 of 15 outbound references displayed

  • verified exact4
  • verified fuzzy0
  • unresolved2
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch9

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 90ab2315-cd92-4d4c-b723-4373ad1efe21 · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

Cosmos 3: Omnimodal World Models for Physical AI PaliGemma: A versatile 3B VLM for transfer

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T22:46:18.409103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-06-28T15:08:33.957835Z digest=sha256:a21196b5a751be67134e62cd55da7a510ce0ad8a28c7b859833c3dfdb31f565f

Observation 92d4c903-e3f6-41ab-b4b9-d406b0253b4e · outbound

This paper cites Internvla-a1: Unifying understanding, generation and action for robotic manipulation.

Cosmos 3: Omnimodal World Models for Physical AI Internvla-a1: Unifying understanding, generation and action for robotic manipulation

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T22:46:18.364550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-06-28T15:08:33.957835Z digest=sha256:319906d30c8113cd2734829648635118e13319e317f13e75cad95da8edfbf715

Observation e3df3b5b-4aa8-4deb-9ede-690dc6f95bab · outbound

This paper cites GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation.

Cosmos 3: Omnimodal World Models for Physical AI GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:46:18.382387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-06-28T15:08:33.957835Z digest=sha256:1d94e68e45d18a55283cd37428bb0cf0f59e60d791e93ecf8cad48a692689ef7

Observation 7fc89fb4-9072-4458-b829-3c7d59a981c5 · outbound

This paper cites Out of time: Automated lip sync in the wild.

Cosmos 3: Omnimodal World Models for Physical AI Out of time: Automated lip sync in the wild

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-28T15:08:33.957835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T15:08:33.957835Z digest=sha256:b40ea2895bae78d2f54a3dc1467de25abe34a75975a91c0f43acea39cb79bb0d

Observation 82577fa1-45ae-4aa8-ba34-435ca39c7be7 · outbound

This paper cites NVLM: Open Frontier-Class Multimodal LLMs.

Cosmos 3: Omnimodal World Models for Physical AI NVLM: Open Frontier-Class Multimodal LLMs

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T22:46:18.387412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-06-28T15:08:33.957835Z digest=sha256:dca7dc99193a15457a1dc7f7fa99a0abe27fe196994dae7744ba0ab6f545b19b

Observation b7c293f8-d6ef-4633-9cef-b54d65c05d07 · outbound

This paper cites Emerging Properties in Unified Multimodal Pretraining.

Cosmos 3: Omnimodal World Models for Physical AI Emerging Properties in Unified Multimodal Pretraining

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:46:18.392456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-06-28T15:08:33.957835Z digest=sha256:fee87da075e27ea6288d41c42eeed510fc476665807e068938370983b220f9dd

Observation 26399458-3208-4e5f-ad80-20b554940130 · outbound

This paper cites VLMEvalKit: An open-source toolkit for evaluating large multi-modality models.

Cosmos 3: Omnimodal World Models for Physical AI VLMEvalKit: An open-source toolkit for evaluating large multi-modality models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-28T15:08:33.957835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T15:08:33.957835Z digest=sha256:9c89fd8a63dd74494e3a19e36a62e0fe19f5014f17bc53c7be6ddd81a5e685d4

Observation a1e8a19a-7040-462b-b9c1-5a55972fb968 · outbound

This paper cites CausalVQA: A Physically Grounded Causal Reasoning Benchmark for Video Models.

Cosmos 3: Omnimodal World Models for Physical AI CausalVQA: A Physically Grounded Causal Reasoning Benchmark for Video Models

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T22:36:18.020144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-06-28T15:08:33.957835Z digest=sha256:b3db818e13d3846a8e01809b6c4d53cab1ff03a1eaa95c5d271e6fc868e83050

Observation b47562f0-b06b-412d-a181-48b19ee21455 · outbound

This paper cites CameraCtrl: Enabling Camera Control for Text-to-Video Generation.

Cosmos 3: Omnimodal World Models for Physical AI CameraCtrl: Enabling Camera Control for Text-to-Video Generation

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T22:36:18.034862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-06-28T15:08:33.957835Z digest=sha256:483178608a83cc04c810fcb0afb3b483b2578365031a76f9c0d7508d08fc6c66

Observation 3131c469-4361-4749-a281-04a1258d019f · outbound

This paper cites MolmoAct: Action Reasoning Models that can Reason in Space.

Cosmos 3: Omnimodal World Models for Physical AI MolmoAct: Action Reasoning Models that can Reason in Space

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T22:46:18.358146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-06-28T15:08:33.957835Z digest=sha256:2ea0dfef487e7914a661738bebca173be8a4bc7bb5bf9172266aa1ff6cbbef83

Observation 56a83cb6-2c8d-42b5-86c2-232b9e46d46c · outbound

This paper cites SceneSmith: Agentic Generation of Simulation-Ready Indoor Scenes.

Cosmos 3: Omnimodal World Models for Physical AI SceneSmith: Agentic Generation of Simulation-Ready Indoor Scenes

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T22:46:18.396613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-06-28T15:08:33.957835Z digest=sha256:7f6edbcb5070bebb1ba282a4957b67215a295548786409660e34be8e104485bc

Observation 3b55fcdc-a366-48ec-bb80-095cba4609f7 · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

Cosmos 3: Omnimodal World Models for Physical AI GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:46:18.405483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-06-28T15:08:33.957835Z digest=sha256:3bd4bc07e9ce21323f8a7788243060e2b3ca3b724d25c9d5cfef5aa4806731f0

Observation 29b28eea-60da-4b26-8ade-f90d4a326b3a · outbound

This paper cites Learning to Act without Actions.

Cosmos 3: Omnimodal World Models for Physical AI Learning to Act without Actions

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T22:46:18.377940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-06-28T15:08:33.957835Z digest=sha256:8b1dffbe34b1aa061b8aee192fa75d84254b3c49e2197dc1938c25bf40c1f678

Observation 0ed2161f-86dc-48bf-80ec-fcf968de0120 · outbound

This paper cites Video models are zero-shot learners and reasoners.

Cosmos 3: Omnimodal World Models for Physical AI Video models are zero-shot learners and reasoners

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:46:18.400971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-06-28T15:08:33.957835Z digest=sha256:b8bc9fb0c9d1915d774e054e3423faffafc70133ec7056c79dafccf09fe7d6ad

Observation 61e6f52a-62a4-41f7-bfa8-7662d99a7d2d · outbound

This paper cites Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation.

Cosmos 3: Omnimodal World Models for Physical AI Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T22:46:18.371943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-06-28T15:08:33.957835Z digest=sha256:a34339abfcaa708c810c56170bfa7b2ec9331e4a612c871318587304def56957

Pith citing papers

Observation 29cba8cc-5629-4e9f-953f-62913f2fb8f9 · inbound

What Spatial Memory Must Store: Occlusion as the Test for Language-Agent Memory cites this paper.

What Spatial Memory Must Store: Occlusion as the Test for Language-Agent Memory Cosmos 3: Omnimodal World Models for Physical AI

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-03T04:37:37.213686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-06-27T13:46:29.718269Z digest=sha256:0142386244b39b410981a3854b6618e3b3d7f4fca692a2c2e33b9f3ed98a6964

Observation 91adc934-0bf2-40d7-944c-8c6ddbb0fb3a · inbound

ActWorld: From Explorable to Interactive World Model via Action-Aware Memory cites this paper.

ActWorld: From Explorable to Interactive World Model via Action-Aware Memory Cosmos 3: Omnimodal World Models for Physical AI

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:28:55.397948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-06-27T01:22:12.771098Z digest=sha256:b97a0d08a7e5c8b1d32c2c2bdfb13e98680c67a585316413163c43644424c3d8

Observation 6ca8f406-2a4b-422f-9777-76ce08f88f85 · inbound

PAIWorld: A 3D-Consistent World Foundation Model for Robotic Manipulation cites this paper.

PAIWorld: A 3D-Consistent World Foundation Model for Robotic Manipulation Cosmos 3: Omnimodal World Models for Physical AI

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-03T21:38:58.754179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-06-27T00:19:33.170645Z digest=sha256:b86c0fb0bacd1a26c75334322e6c1add7f6305751bbec91c90779f4ff7510527

Observation fa8e1f51-2122-4199-8976-eac1433f015c · inbound

SC3-Eval: Evaluating Robot Foundation Models via Self-Consistent Video Generation cites this paper.

SC3-Eval: Evaluating Robot Foundation Models via Self-Consistent Video Generation Cosmos 3: Omnimodal World Models for Physical AI

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:19:13.362817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-06-26T21:18:07.494794Z digest=sha256:25d4a4d250f3064c759cac5a13e1a76ad9371a4355d4b833b8dc2948527027b5

Observation b6c19015-8863-419a-a00d-ae261b110023 · inbound

SC3-Eval: Evaluating Robot Foundation Models via Self-Consistent Video Generation cites this paper.

SC3-Eval: Evaluating Robot Foundation Models via Self-Consistent Video Generation Cosmos 3: Omnimodal World Models for Physical AI

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-06-29T18:23:51.389791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-06-29T05:06:30.594399Z digest=sha256:cddcba400137922905928e16cd9c0a37ff8ec9146d3bc68b28c1f9ca78c759b0

Observation 4a382069-ca46-4061-8322-4b9cea202198 · inbound

Physics-IQ Verified cites this paper.

Physics-IQ Verified Cosmos 3: Omnimodal World Models for Physical AI

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:29:15.556417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-06-26T21:13:36.568700Z digest=sha256:372234a894ace03858e8937a5a2afabd870a449c3e6c8baa7d704afa281be7be

Observation 891ebd45-3d91-4e59-b077-7546724672b3 · inbound

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? cites this paper.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Cosmos 3: Omnimodal World Models for Physical AI

Reference 84

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:39:17.503977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:523cee8ef20899a989563ee5066efb2847618dc18b719ed3ff364865c6711c6c

Observation dc43e339-80bc-457f-acb7-8a04d4428f56 · inbound

Sol Video Inference Engine: Agent-Native Full-Stack Acceleration Framework for Efficient Video Generation cites this paper.

Sol Video Inference Engine: Agent-Native Full-Stack Acceleration Framework for Efficient Video Generation Cosmos 3: Omnimodal World Models for Physical AI

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-04T08:49:41.558859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-06-26T11:05:44.972128Z digest=sha256:09bcf477574213eb4c0b60d75c899ab348d514c35483ebe6c6b2c351dbd93b38

Observation 4fb51621-8fce-4a31-9f5a-9e337858f7a3 · inbound

Critique of Agent Model cites this paper.

Critique of Agent Model Cosmos 3: Omnimodal World Models for Physical AI

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-07-04T11:29:51.076095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-06-26T07:57:20.830605Z digest=sha256:c25ccdc43ca1a409a89a392651ecfee567e0d9db5acaef3f0c72b7a363fe4464

Observation 2d8ece93-2c98-486b-b0a7-44a037d7cca0 · inbound

DiffusionBench: On Holistic Evaluation of Diffusion Transformers cites this paper.

DiffusionBench: On Holistic Evaluation of Diffusion Transformers Cosmos 3: Omnimodal World Models for Physical AI

Reference 212

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T16:59:58.195744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=arxiv_source observed=2026-06-26T00:06:11.951205Z digest=sha256:5dc71584a49f7a4033e654b784cbf52992094da6339aab21fa2f4c0e43d39da3

Observation 65459639-2a72-4e6e-b9b7-6ba7fbd8258e · inbound

Causal-rCM: A Unified Teacher-Forcing and Self-Forcing Open Recipe for Autoregressive Diffusion Distillation in Streaming Video Generation and Interactive World Models cites this paper.

Causal-rCM: A Unified Teacher-Forcing and Self-Forcing Open Recipe for Autoregressive Diffusion Distillation in Streaming Video Generation and Interactive World Models Cosmos 3: Omnimodal World Models for Physical AI

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-07-04T19:50:11.445672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-06-25T20:57:30.765802Z digest=sha256:c84089d910cd5898233259ae1fd17eb47179ef9f291f034506441eb44ca2dd09

Observation e1402191-1b5d-44fd-97dc-807539f2af53 · inbound

Learning Action Priors for Cross-embodiment Robot Manipulation cites this paper.

Learning Action Priors for Cross-embodiment Robot Manipulation Cosmos 3: Omnimodal World Models for Physical AI

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-04T21:00:09.820221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-06-25T19:09:56.409766Z digest=sha256:8abd0c45dc8754f543b6809a3bcef2c1d633087586387c6b59c0b07277dbf4da

Observation 0707065f-fc4d-47f5-9996-73e35bf64de4 · inbound

PhysisForcing: Physics Reinforced World Simulator for Robotic Manipulation cites this paper.

PhysisForcing: Physics Reinforced World Simulator for Robotic Manipulation Cosmos 3: Omnimodal World Models for Physical AI

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-06-29T20:03:56.948658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-06-29T04:34:38.286863Z digest=sha256:17c00a6c62ecb72dae720a7f40d64ec01893d4a7075de00e65d95ff774349c21

Observation 42699cb3-ed34-4566-a74e-45d7db3f8368 · inbound

RoboGaze: Evaluating Robot World Models via Structured Vision-Language Analysis cites this paper.

RoboGaze: Evaluating Robot World Models via Structured Vision-Language Analysis Cosmos 3: Omnimodal World Models for Physical AI

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-06-30T10:54:35.850307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-06-30T10:49:58.019238Z digest=sha256:aabb825c6ef2aa94187d3433f2d0ac8620f0975a8efc38284c132739b0bfe32f

Observation 8df2960d-1060-4c18-84e1-97ccce87eeca · inbound

Mural: Transferring LLM knowledge to image generation via Mixture-of-Transformers cites this paper.

Mural: Transferring LLM knowledge to image generation via Mixture-of-Transformers Cosmos 3: Omnimodal World Models for Physical AI

Reference 20

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T09:24:32.131834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-06-30T09:23:59.752793Z digest=sha256:3701393522418fbdee39e0efb209d32101d96825c3aaa5960dc83ea3455899e5

Observation 8df1cee3-7063-430e-9ec5-4fda6ccc8145 · inbound

World Narrative Model for Highly Controllable Video Generation: A Paradigm Shift from Pixel Sampling to Physical World Orchestration cites this paper.

World Narrative Model for Highly Controllable Video Generation: A Paradigm Shift from Pixel Sampling to Physical World Orchestration Cosmos 3: Omnimodal World Models for Physical AI

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-07-01T10:25:41.706604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-07-01T05:31:43.678762Z digest=sha256:69741ce577302c0f696240a550567e8b5a731dbd5d48b2ad03ab9bcb41ad7ab7

Observation 5aae7d9a-c867-495b-95f1-9df375d77f2f · inbound

World Narrative Model for Highly Controllable Video Generation: A Paradigm Shift from Pixel Sampling to Physical World Orchestration cites this paper.

World Narrative Model for Highly Controllable Video Generation: A Paradigm Shift from Pixel Sampling to Physical World Orchestration Cosmos 3: Omnimodal World Models for Physical AI

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-15T10:20:59.147440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T10:20:59.147440Z digest=sha256:083e10bc76949e95194d099c48798291a9a35dbad8b784128a2fc11ed80c6ad2

Observation 595f0b64-ad7b-46dc-a55b-0de06bfff587 · inbound

ROSA: A Robotics Foundation Model Serving System for Robot Factories cites this paper.

ROSA: A Robotics Foundation Model Serving System for Robot Factories Cosmos 3: Omnimodal World Models for Physical AI

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-02T11:16:52.666121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-07-02T11:11:39.584755Z digest=sha256:ed9f7e1cc7482ff5f468c545d392ccc13695c1207c41f85f35fff2547794cb2e

Observation 6e6bf47c-766d-4292-bb0c-2aeecf7e3122 · inbound

DynaWM: A Base-VLA-Guided World Foundation Model for Moving-Object Manipulation cites this paper.

DynaWM: A Base-VLA-Guided World Foundation Model for Moving-Object Manipulation Cosmos 3: Omnimodal World Models for Physical AI

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:40.344868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:40.344868Z digest=sha256:06bd32ef291cbebbb9390fcb753e74f991aea70a68760b97d146835564dd6718

Observation e815add8-2379-4d1f-97f4-1b2bbe946ad3 · inbound

GigaWorld-1: A Roadmap to Build World Models for Robot Policy Evaluation cites this paper.

GigaWorld-1: A Roadmap to Build World Models for Robot Policy Evaluation Cosmos 3: Omnimodal World Models for Physical AI

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-12T08:04:48.963890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T08:04:48.963890Z digest=sha256:fb98fbd075d7613e18f44bb6626ab535a97122c6b814a2a845396aa987ee3180

Observation 2e3ebbcd-57e3-4dbe-8636-e49b96634871 · inbound

ACE-Brain-0.5: A Unified Embodied Foundational Model for Physical Agentic AI cites this paper.

ACE-Brain-0.5: A Unified Embodied Foundational Model for Physical Agentic AI Cosmos 3: Omnimodal World Models for Physical AI

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-11T19:16:57.396710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:16:57.396710Z digest=sha256:efce88316aaed4d18cbddab622db5d28de1fc0ecfc554a7612387b5148b56478

Observation 340b9926-bd9a-4491-b57a-94fc9fb807a6 · inbound

A Definition and Roadmap for World Models cites this paper.

A Definition and Roadmap for World Models Cosmos 3: Omnimodal World Models for Physical AI

Reference 128

Resolution
verified exact
local_arxiv, observed 2026-07-08T07:14:44.336323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=arxiv_source observed=2026-07-08T07:10:33.826140Z digest=sha256:6e5febf872f8d27b2f0c86dc3cd2205cacb0487502f58bd2998a51e75991d43d

Observation 5cfc9951-4d25-4d42-a0ba-7eb2e0803ad2 · inbound

From Foundation to Application: Improving VLA Models in Practice cites this paper.

From Foundation to Application: Improving VLA Models in Practice Cosmos 3: Omnimodal World Models for Physical AI

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-08T07:14:45.485926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:a21544f13ad0d439396b688534b691c78423bfb5964ff228a9372590d1817ce7

Observation fea50311-ff49-4e5f-9c6d-776b88c0059e · inbound

WAM-TTT: Steering World-Action Models by Watching Human Play at Test Time cites this paper.

WAM-TTT: Steering World-Action Models by Watching Human Play at Test Time Cosmos 3: Omnimodal World Models for Physical AI

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-13T06:48:14.554799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T06:48:14.554799Z digest=sha256:e3002f4c324a127fd53973d7c3286c6d75c067725db883906ec9c2e13db2c215

Observation d70e5411-8c99-4851-8678-3fdf7bf4fa3f · inbound

Infinite Worlds with Versatile Interactions cites this paper.

Infinite Worlds with Versatile Interactions Cosmos 3: Omnimodal World Models for Physical AI

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-09T07:56:04.717628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-07-09T07:51:36.802801Z digest=sha256:f2b242af1e5f1dad24b4ae593a4cec0339b6e7fa39200018016b9f09a90339ab

Observation ad2ace2e-2a21-4fd2-b5d8-120e307a1492 · inbound

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence cites this paper.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Cosmos 3: Omnimodal World Models for Physical AI

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-09T03:05:55.157813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:88209db451e0ccbd3776624aa2d147aa2c775f51d3b16712fb672d8c5ca95a77

Observation 25d1cb50-f7ab-4983-8cba-96e8942e7e39 · inbound

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model cites this paper.

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model Cosmos 3: Omnimodal World Models for Physical AI

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-14T04:10:14.360463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:10:14.360463Z digest=sha256:85466e3bdecc3deb91c0621e1465ecee05e8a9155794d148c7154c1919502d08