Pith. sign in

Paper Citation Record · LEDGER

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control

As of 22 August 2026, this Paper Citation Record lists 92 of 92 outbound references and 6 inbound Pith citation observations for arXiv:2412.11198.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.11198 v1

Coverage vector

measured 92 of 92 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T15:16:34.345618Z

measured 98 of 98 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:13:42.902690Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T21:06:13.797390Z

Reference resolution

92 of 92 outbound references displayed

  • verified exact1
  • verified fuzzy42
  • unresolved49
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f0c943d2-5c7b-4af8-9697-38696965940a · outbound

This paper cites Stereo vision and laser odom- etry for autonomous helicopters in gps-denied indoor envi- ronments.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Stereo vision and laser odom- etry for autonomous helicopters in gps-denied indoor envi- ronments

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:33.963013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:33.963013Z digest=sha256:3056b203d5554d32dba25b1fdb3617caf6b803a4db1d60aa6fa2700795aec307

Observation 79fcb6d7-3857-4729-8c18-155ef65d53a7 · outbound

This paper cites LIMT: Language-Informed Multi-Task Visual World Models.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control LIMT: Language-Informed Multi-Task Visual World Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:33.967601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:33.967601Z digest=sha256:f60ce9444a5c1f98792e97f26ab39fd56f33ed45f3fa6fbb8b60f94e9c7865d8

Observation b2bacd06-2342-4a9b-9d1a-fd0330842872 · outbound

This paper cites Uncertainty-based traffic accident anticipation with spatio-temporal relational learn- ing.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Uncertainty-based traffic accident anticipation with spatio-temporal relational learn- ing

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:33.972862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:33.972862Z digest=sha256:aba3ad159b5ab19c9d6839610ccedb61f8332144f12e26da0175ecaa5c974c33

Observation f23d4975-f0c0-4ee0-985f-580052c2e3e0 · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:33.977882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:33.977882Z digest=sha256:632230105fbab09d4ae8af6b3048c57aefc8068220ad389f8d03b1ff1507a706

Observation f0217454-faf5-41fb-b2d2-e04426a28ed7 · outbound

This paper cites Align your latents: High-resolution video synthesis with la- tent diffusion models.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Align your latents: High-resolution video synthesis with la- tent diffusion models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:33.982431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:33.982431Z digest=sha256:6a7953093591fd9658d238cd3da60de8ebf0434aec7b8a139bdb950ac8579ddc

Observation f2693969-09fd-4680-adab-7b9a39fcff14 · outbound

This paper cites Marius Z¨ollner.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Marius Z¨ollner

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:33.986497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:33.986497Z digest=sha256:d9fa7077c9605717007491320d1861f2f2d687214f054c98ac3553366c8aa53a

Observation e531c33e-6bcb-4925-b59a-77063474a104 · outbound

This paper cites Generating long videos of dynamic scenes.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Generating long videos of dynamic scenes

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:33.990515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:33.990515Z digest=sha256:debe066400c3051146ee4ef92476a2e4d1ba39785f0e96ece032ad54482abbc2

Observation aee29c92-7710-4524-88e1-966e84c13c27 · outbound

This paper cites nuscenes: A multi- modal dataset for autonomous driving.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control nuscenes: A multi- modal dataset for autonomous driving

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:33.994832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:33.994832Z digest=sha256:72375f107eb45cb63f3dc3ecbbb4c20cd2381250aec828b48d04936bcab41d5a

Observation 85f2f802-717c-483b-bc90-7eeb4d159fe9 · outbound

This paper cites D$^2$-City: A Large-Scale Dashcam Video Dataset of Diverse Traffic Scenarios.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control D$^2$-City: A Large-Scale Dashcam Video Dataset of Diverse Traffic Scenarios

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:33.998408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:33.998408Z digest=sha256:36f4071c7516d8b126e403b1d23ac638e7bd2507f26f0373161a73d51ebded1b

Observation 30438a74-a2f1-4cea-9b4a-a1235963287b · outbound

This paper cites Diffusion forcing: Next-token prediction meets full-sequence diffu- sion.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Diffusion forcing: Next-token prediction meets full-sequence diffu- sion

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:34.002308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:34.002308Z digest=sha256:1367b81e3319923c0134d7641b367bdd5f7dffe736ab424884b0c09d6497a697

Observation 19b80778-3157-41ac-843d-6a5a877a932b · outbound

This paper cites Videocrafter1: Open diffusion models for high-quality video generation, 2023.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Videocrafter1: Open diffusion models for high-quality video generation, 2023

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:34.005983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:34.005983Z digest=sha256:9e3e8b0c0fbfac5ed3d2a96be3b1384d0f05ce497f8543c6c00e4a748cd69f66

Observation 07125f32-80f0-4fcc-8a86-801405cd1789 · outbound

This paper cites Videocrafter2: Overcoming data limitations for high-quality video diffu- sion models.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Videocrafter2: Overcoming data limitations for high-quality video diffu- sion models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:34.009937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:34.009937Z digest=sha256:b2198a48597c0e56e899af7784fdc431962559951876963c020be4b355c4ecb3

Observation 0b802405-eea4-4dab-b622-4ba0539a3f11 · outbound

This paper cites Seine: Short-to-long video diffu- sion model for generative transition and prediction.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Seine: Short-to-long video diffu- sion model for generative transition and prediction

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:34.014775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:34.014775Z digest=sha256:d2a020017a3081a7612d4f403549888edda992d3eccdc4e37e52db77cab77aa4

Observation d08aa199-e8c3-4b98-be7a-75e46d61ce8e · outbound

This paper cites CAGE: Unsupervised Visual Composition and Animation for Controllable Video Generation.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control CAGE: Unsupervised Visual Composition and Animation for Controllable Video Generation

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-08-11T15:16:34.794898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:16:34.018680Z digest=sha256:bb44830d2e721821f84f6f04334e1054372b289d1a373621ea7debfcf7930ef3

Observation dbc556a1-dcb3-47da-a210-f934efbc4c48 · outbound

This paper cites Diffusion models beat gans on image synthesis, 2021.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Diffusion models beat gans on image synthesis, 2021

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:34.023037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:34.023037Z digest=sha256:7ced07c39fa49e576fa98f7529c476fcf718c7f52b74c927c15abb60f7a0abf4

Observation d57f5800-a0d9-4f50-a31a-47072ce02c1e · outbound

This paper cites Vista: A generalizable driving world model with high fidelity and versatile controllability.Advances in Neural Information Processing Systems, 2024.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Vista: A generalizable driving world model with high fidelity and versatile controllability.Advances in Neural Information Processing Systems, 2024

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:34.026799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:34.026799Z digest=sha256:3cc3837459b3fe70ff7bfa1ee6e97b7f0c4490d13c98b35b345b7ce49158ec79

Observation ed0f69db-665f-483b-b7db-d425581e2d33 · outbound

This paper cites Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:34.031763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:34.031763Z digest=sha256:5e46e4728d79d82162e41e3f665a82751e91137ad3ef91c173c5e21b13fb2c51

Observation 44ca4930-10cc-4699-8354-70ae585e082a · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Ego4d: Around the world in 3,000 hours of egocentric video

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:34.036409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:34.036409Z digest=sha256:d027fcfd9407af24b2f35d8d5853abe4a4e206e9b861eebed8ff32af1d651fc5

Observation aaed36c6-6714-4100-810c-2e2c87a59148 · outbound

This paper cites Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:16:35.475087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:16:34.044063Z digest=sha256:22cd0181fa07163586d989bd2a9d6df7c37998697eb1f2e64cdcae5d05382e41

Observation 9833566a-808b-4bd1-808e-f9b3832588fb · outbound

This paper cites Animatediff: Animate your personalized text-to- image diffusion models without specific tuning, 2024.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Animatediff: Animate your personalized text-to- image diffusion models without specific tuning, 2024

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:16:35.463197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:16:34.048122Z digest=sha256:46c5c503e032d30d73e4e108e6479beba6e851712da6ee138cd95ea9c9f5cfaf

Observation ac20f99a-15e1-42bc-995b-2212db177b20 · outbound

This paper cites Dream to control: Learning behaviors by la- tent imagination.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Dream to control: Learning behaviors by la- tent imagination

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:16:35.451926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:16:34.051812Z digest=sha256:8e14f55d329ef9071d9f24e93e6a7db120bab7641d560dbfd63599dffa00d396

Observation b1e5a28c-32c9-47c9-a508-1219fb203d8a · outbound

This paper cites Hierarchical World Models as Visual Whole-Body Humanoid Controllers.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Hierarchical World Models as Visual Whole-Body Humanoid Controllers

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:34.056278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:34.056278Z digest=sha256:e9d5f0fcfa4e9c33ce1fd9ab801482c0ba42114d688259b29ad86074c1e7794a

Observation 59608956-1ebe-4e6c-92a3-75466b10a5ca · outbound

This paper cites Temporal difference learning for model predictive control.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Temporal difference learning for model predictive control

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:16:35.439374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:16:34.061360Z digest=sha256:549a0b42229b4f56538aa3a2a2a54186f144c44e847c287158f22a006f6b9c19

Observation 5a9a4bd3-dd76-421f-bb02-714fcd14764a · outbound

This paper cites Reasoning with language model is planning with world model.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Reasoning with language model is planning with world model

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:16:35.426209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:16:34.065233Z digest=sha256:c289a9bc6b7a7712d585431a3ba56ce371d076c0a26495303b21ef2f8f2779ed

Observation 0ff444b0-d8d2-4735-9284-aaef416c4d73 · outbound

This paper cites Large-scale actionless video pre-training via discrete diffusion for efficient policy learning, 2024.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Large-scale actionless video pre-training via discrete diffusion for efficient policy learning, 2024

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:16:35.412200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:16:34.069400Z digest=sha256:f88c314889a49a372da3d4ec0da69ecb48c9af41e52ef6927464cd6da577d2a1

Observation 11b00641-8c68-4e8f-ba92-0f25a5466ec7 · outbound

This paper cites End-to-end learning of driving models with surround-view cameras and route planners.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control End-to-end learning of driving models with surround-view cameras and route planners

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:16:35.399476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:16:34.073357Z digest=sha256:e8d45afbecbfb1bcf9980a927f3c66cdab6a73f655cae2b87bd423495ae5415c

Observation 7112a02b-1522-488f-b883-3aac346ad990 · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilib- rium.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Gans trained by a two time-scale update rule converge to a local nash equilib- rium

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:16:35.384026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:16:34.076953Z digest=sha256:8485c42f690d163583e6e2d125d4e21d2f648c7ea75bb716c863247459f0006a

Observation 73e31547-eb73-4850-b0d8-e0f86d1c6325 · outbound

This paper cites Imagen Video: High Definition Video Generation with Diffusion Models.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Imagen Video: High Definition Video Generation with Diffusion Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:34.080962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:34.080962Z digest=sha256:e1fd8f9b0c3227489f7260cc7d0cb3ab0f136be7285f78f0a5f92196158a8a5b

Observation 32fe1254-b292-4730-a47b-7fce52e7a0a8 · outbound

This paper cites an unresolved cited work.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:16:35.369461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:16:34.085256Z digest=sha256:041b61d1727f0f9ef0bfbcbd9fe4740ea8fcdc21af78da909918b6f2f8886525

Observation caf35113-5329-4041-8711-ff2636902fa8 · outbound

This paper cites Gaia-1: A generative world model for au- tonomous driving, 2023.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Gaia-1: A generative world model for au- tonomous driving, 2023

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:16:35.357703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:16:34.088769Z digest=sha256:dbf4fb0cab3e09ed3dbd3d6d2b921ed369689bddb8d0f08b76ff8402790c9ec7

Observation 7976e2c1-5eb1-4f6e-9b48-e033cf95f423 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control LoRA: Low-Rank Adaptation of Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:34.092126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:34.092126Z digest=sha256:1a74e701fd57dc8ce406406f6ddefd7ebb04c98475cbe2186629e2119f6edbfd

Observation c61068d7-457c-4447-b1df-d8417595ccd2 · outbound

This paper cites DepthCrafter: Generating Consistent Long Depth Sequences for Open-world Videos.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control DepthCrafter: Generating Consistent Long Depth Sequences for Open-world Videos

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:34.095535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:34.095535Z digest=sha256:10cd526f35ed73a175433a46f9cdd50cfcf1ff1e8b90dfc907d57b2a57ee10e3

Observation 478f36dc-6428-4fbc-8083-19fb125c2532 · outbound

This paper cites Toward general-purpose robots via foundation mod- els: A survey and meta-analysis, 2023.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Toward general-purpose robots via foundation mod- els: A survey and meta-analysis, 2023

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:16:35.343294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:16:34.099513Z digest=sha256:95c9336340fee257e4dfe969b4b0e7e55acaadd5322cff5b20ea4a470a43fe44

Observation 275aa98d-c1ca-40be-af15-33ecba4d8c64 · outbound

This paper cites Language Models, Agent Models, and World Models: The LAW for Machine Reasoning and Planning.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Language Models, Agent Models, and World Models: The LAW for Machine Reasoning and Planning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:34.103629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:34.103629Z digest=sha256:da9caab793d9a1e09ec5103db56cd11ace5856f232fbbf9040d80faa233669ff

Observation 6b18fa19-9d5b-43e0-83d6-f325614c7e97 · outbound

This paper cites ADriver-I: A General World Model for Autonomous Driving.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control ADriver-I: A General World Model for Autonomous Driving

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:34.107971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:34.107971Z digest=sha256:c46e3a9264ec7ae3ddc5f327456ed42c4adf661d23f309ad51ae7cfd20e7aad1

Observation 1a1d81bf-ee41-4b76-a794-9413ea73231f · outbound

This paper cites Elucidating the Design Space of Diffusion-Based Generative Models.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Elucidating the Design Space of Diffusion-Based Generative Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:34.112333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:34.112333Z digest=sha256:b5692614d0e571dacd45fc20b135bab93b8c1de77113504ef8bd4341318da930

Observation e6c70fd4-e4ea-4e51-8b7c-d7c0c53b3390 · outbound

This paper cites YOLOv11: An Overview of the Key Architectural Enhancements.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control YOLOv11: An Overview of the Key Architectural Enhancements

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:34.116587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:34.116587Z digest=sha256:97d47b6b041af33ca4d76ecdcc60f87a0eb2b745ffe438b0d5556728c7e3db82

Observation 679ea084-840c-4aee-9d03-ff668f46e032 · outbound

This paper cites Grounding human-to-vehicle advice for self-driving vehicles.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Grounding human-to-vehicle advice for self-driving vehicles

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:16:35.330645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:16:34.120449Z digest=sha256:60ab5552081b27563558336881ff0806412ee6441af5662640861cd72b582452

Observation 2354599e-12f3-4135-bc64-c2d3479d4128 · outbound

This paper cites Drivegan: Towards a controllable high-quality neural simulation.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Drivegan: Towards a controllable high-quality neural simulation

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:16:35.318280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:16:34.124548Z digest=sha256:76747a29f05f9eb46226c061a28edd9da8925f8d9cac54f2bf2f7a363c41d0b6

Observation 1918ead4-40ed-4bb2-b153-551dcded7567 · outbound

This paper cites A path towards autonomous machine intelli- gence.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control A path towards autonomous machine intelli- gence

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:16:35.305571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:16:34.128703Z digest=sha256:f5ac8258b1b7494c06e99dad5a42fe43c82fa65dfc526555669e23f95134761f

Observation 0747f4d1-d52a-4336-a492-a60e81d6f402 · outbound

This paper cites WoVoGen: World Volume-aware Diffusion for Controllable Multi-camera Driving Scene Generation.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control WoVoGen: World Volume-aware Diffusion for Controllable Multi-camera Driving Scene Generation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:34.132700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:34.132700Z digest=sha256:2e0e1820b6f33efd5703ee2a3c5ec9304bdb3d05c58736931a5f5099dc35ba95

Observation 493923a2-6d7d-42d9-8f01-2d96a63a5e76 · outbound

This paper cites Fit: Flexible vision trans- former for diffusion model, 2024.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Fit: Flexible vision trans- former for diffusion model, 2024

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:16:35.294159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:16:34.137416Z digest=sha256:81420c6d75a3d3b445c77e2d8a52d98b00a1224a5c275ab5b39071fa472a8c27

Observation 684dd68f-f55a-47f5-8d85-95fb7dee1e0f · outbound

This paper cites Struc- tured world models from human videos, 2023.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Struc- tured world models from human videos, 2023

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:16:35.280112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:16:34.141266Z digest=sha256:220aa62974b561214c980b1eae654351643c819acc60011750f7fa9a5ae3a828

Observation 17adda47-40aa-46b9-be64-5169017f8d44 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control DINOv2: Learning Robust Visual Features without Supervision

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:34.144911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:34.144911Z digest=sha256:bff2698bdaa5fd0ae9f69784239dc9ee7fad0a79bdcc91ed7f390a5df9529493

Observation 329a02ad-df82-49a8-b8a0-690db601d10d · outbound

This paper cites an unresolved cited work.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:16:35.267574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:16:34.149083Z digest=sha256:39ca70296b38770afb603bb6a476f450b67a1daceb4db8109a350b830d0572bc

Observation 4c3230d9-a7a7-4085-b26a-398a15278514 · outbound

This paper cites Dutta Roy, Sugosh Nagavara Ravindra, Priya Goyal, and Matthijs Douze.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Dutta Roy, Sugosh Nagavara Ravindra, Priya Goyal, and Matthijs Douze

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:16:35.254388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:16:34.153447Z digest=sha256:a1c8adeeb7d7673347f41b935b622fdc095cfd8fece4e18b2e14baf22a747b22

Observation 2d1fb9e7-fd5f-4803-9c3d-864686cf3de9 · outbound

This paper cites Reasoning with large lan- guage models, a survey.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Reasoning with large lan- guage models, a survey

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:34.158058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:34.158058Z digest=sha256:f4d0cfab813366d5c7c99c4a22c462904cb37fcec19c8a1668d9337bf27dd3d6

Observation c348d7e8-32fb-4413-a947-a301fc99734f · outbound

This paper cites Sampson, Shikai Li, Simone Parmeggiani, Steve Fine, Tara Fowler, Vladan Petro- vic, and Yuming Du.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Sampson, Shikai Li, Simone Parmeggiani, Steve Fine, Tara Fowler, Vladan Petro- vic, and Yuming Du

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:16:35.241195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:16:34.162913Z digest=sha256:348e098389f6a79eeca13b70258b4e68cddc462597adc6bf3b1584f0d5dc4d54

Observation aed6bf9a-9447-4a3c-a717-905e8fb42de0 · outbound

This paper cites Toward driving scene understanding: A dataset for learning driver behavior and causal reasoning.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Toward driving scene understanding: A dataset for learning driver behavior and causal reasoning

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:16:35.227930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:16:34.166894Z digest=sha256:52237b6d2c040763f3abd8a40847dc34fc01d0385ad6d844b49585d929164b27

Observation e79a5712-0d83-43ff-8a45-278b8934ed9f · outbound

This paper cites Deepspeed: System optimizations enable train- ing deep learning models with over 100 billion parameters.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Deepspeed: System optimizations enable train- ing deep learning models with over 100 billion parameters

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:16:35.215751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:16:34.170813Z digest=sha256:66654dc805f4637ebd1d6b8b94ead4fc059cfaa338ebf47ad8c4ba902dd1983a

Observation c1b0ef6c-bb61-42a2-b009-546f21d7e47b · outbound

This paper cites High-resolution image syn- thesis with latent diffusion models, 2022.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control High-resolution image syn- thesis with latent diffusion models, 2022

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:34.174473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:34.174473Z digest=sha256:3c10dc41ef3e9e9723ec388b34fdeadec85c39f53934eeb02699e996e8caff97

Observation cbf2978f-858f-4d31-a3cb-6a31081f7d9b · outbound

This paper cites LAION-5B: An open large-scale dataset for training next generation image-text models.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control LAION-5B: An open large-scale dataset for training next generation image-text models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:34.177833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:34.177833Z digest=sha256:401f60dce27519d774c7f67ccb3466dd0344e79f24bf5887a5f998a98a17be1f

Observation 20b04af5-cd00-40ac-b420-7d398a99b6ab · outbound

This paper cites Make-a-video: Text-to-video generation without text-video data, 2022.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Make-a-video: Text-to-video generation without text-video data, 2022

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:34.181414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:34.181414Z digest=sha256:1b74360b6962ff2f1a731a1e38dbe4619e948c1c45ffa1c24262b73b93e36e5a

Observation 27127ad6-60d7-44c7-b12e-e618d5be4876 · outbound

This paper cites Fourier features let networks learn high frequency functions in low dimen- sional domains.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Fourier features let networks learn high frequency functions in low dimen- sional domains

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:34.185082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:34.185082Z digest=sha256:37604efd26dca56e0a38440f307464969e7576182de3cc1099ae5ba4c8ee0b68

Observation 00c901e5-fe80-4d84-b48a-c7b542f2d00a · outbound

This paper cites Raft: Recurrent all-pairs field transforms for optical flow.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Raft: Recurrent all-pairs field transforms for optical flow

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:16:35.176086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:16:34.189407Z digest=sha256:6ac648d579cb30cef902d9d6131035b46e049663f7696c9b9a0511de3d0e5d1d

Observation a53140e5-8058-4024-a7e7-e4ba50a7a0c9 · outbound

This paper cites Droid-slam: Deep visual slam for monocular, stereo, and rgb-d cameras.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Droid-slam: Deep visual slam for monocular, stereo, and rgb-d cameras

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:16:35.163278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:16:34.193274Z digest=sha256:3aab89fe12a9e5990140508f6c592a392190148826775926a110a8220d29dacd

Observation a48b5b46-1a54-420a-ae8c-8b2aacef3454 · outbound

This paper cites Towards Accurate Generative Models of Video: A New Metric & Challenges.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Towards Accurate Generative Models of Video: A New Metric & Challenges

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:34.197022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:34.197022Z digest=sha256:6229874ad5796f4463c601ef3832d58c8b53946141a12d1228ef051cf2fd9652

Observation cf9e70ef-3750-4a86-8499-d6debda17c84 · outbound

This paper cites GeoCalib: Single-image Cali- bration with Geometric Optimization.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control GeoCalib: Single-image Cali- bration with Geometric Optimization

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:16:35.150407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:16:34.201145Z digest=sha256:5e22fcaa7fb9a666c24757bc8e9c7c5c10474dc56dc1ecc00b1736631197a943

Observation 4880cfa8-55d8-4724-a32f-079a73ef61bf · outbound

This paper cites Channappayya, and Swarup S.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Channappayya, and Swarup S

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:16:35.137510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:16:34.205641Z digest=sha256:fc82c5c9c2f9affd0555912219f388543832352fd3e51f7d5a517b071b948a9f

Observation 67d38f81-1555-4636-8539-526f298f566f · outbound

This paper cites Mcvd: Masked conditional video diffusion for predic- tion, generation, and interpolation, 2022.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Mcvd: Masked conditional video diffusion for predic- tion, generation, and interpolation, 2022

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:16:35.124740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:16:34.209464Z digest=sha256:df926058a7d17326c6a7cc0d5eb44f594d6d4c227d70e024317ca14a57e949f1

Observation 36e8d615-636d-41e5-901a-a1c0dbcbc765 · outbound

This paper cites DriveDreamer: Towards Real-world-driven World Models for Autonomous Driving.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control DriveDreamer: Towards Real-world-driven World Models for Autonomous Driving

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:34.213408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:34.213408Z digest=sha256:5d35b4ef0f5d4a27547cbd4f6d6d964f495e4c97206d5fd95cceb7226f7e79ab

Observation 2f2105f8-9434-4efa-8cbf-f44c4fcb92b8 · outbound

This paper cites Videocomposer: Compositional video synthesis with motion controllability.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Videocomposer: Compositional video synthesis with motion controllability

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:16:35.110970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:16:34.218155Z digest=sha256:5e98dd102b1587f3f296428bd9b0ea6062dd3f4a2289dc4aa4ffd84353f45a6b

Observation f33f9adc-e1a7-4b3b-934e-64ccb307eff8 · outbound

This paper cites Pseudo- lidar from visual depth estimation: Bridging the gap in 3d object detection for autonomous driving.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Pseudo- lidar from visual depth estimation: Bridging the gap in 3d object detection for autonomous driving

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:16:35.096364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:16:34.221857Z digest=sha256:79e80b8f3dae8f968082140562c819661d1fb7a68bfed900a2e1c7def87b5a60

Observation f83f2773-7020-4dc1-ad3f-c3b402b3fa66 · outbound

This paper cites DrivingDojo Dataset: Advancing Interactive and Knowledge-Enriched Driving World Model.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control DrivingDojo Dataset: Advancing Interactive and Knowledge-Enriched Driving World Model

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:34.225388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:34.225388Z digest=sha256:e258dc29b5817929495177dcbe1d9322b8e82bbd8f9e4c78b8f15589ef080915

Observation dfc508c9-9db5-42da-8e58-274abbee091b · outbound

This paper cites Driving into the future: Multiview visual forecasting and planning with world model for au- tonomous driving.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Driving into the future: Multiview visual forecasting and planning with world model for au- tonomous driving

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:16:35.082482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:16:34.229463Z digest=sha256:7058985e9bec932aaae2e71eb79e604681d73ebdd1032135d08181247bbeb57e

Observation 2332a368-0687-437c-a674-3262e5c98836 · outbound

This paper cites Pre-training contextualized world models with in-the-wild videos for reinforcement learning.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Pre-training contextualized world models with in-the-wild videos for reinforcement learning

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:16:35.070390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:16:34.233575Z digest=sha256:3f17551317531b8f4598ff6e209d3826eabde123e39afeb1b9cf270c8cbc8f12

Observation 65c3c2cd-1dbf-4b97-94fb-5428114a0bf2 · outbound

This paper cites Pandora: Towards General World Model with Natural Language Actions and Video States.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Pandora: Towards General World Model with Natural Language Actions and Video States

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:34.237524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:34.237524Z digest=sha256:50b066bb01d749a47ae230a84f03245f1adf4b36e2a49c7fd71653ed5bab00e7

Observation 2a78fe71-635e-461f-8730-054a20671f52 · outbound

This paper cites Progressive Autoregressive Video Diffusion Models.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Progressive Autoregressive Video Diffusion Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:34.242395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:34.242395Z digest=sha256:d705b8425b5671c0e9fb271307fb382cb40054dad64358dbecc297b6d8347a26

Observation 8d42c684-6d69-41f9-b521-548132ccc25e · outbound

This paper cites End- to-end learning of driving models from large-scale video datasets.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control End- to-end learning of driving models from large-scale video datasets

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:16:35.057052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:16:34.246288Z digest=sha256:3969e5012bcb2e37331fff23a161a40d4db203b882d0de9ede21b5423736fab8

Observation a111aac7-e4e0-403b-9ffe-e55922e446ac · outbound

This paper cites Videogpt: Video generation using vq-vae and trans- formers, 2021.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Videogpt: Video generation using vq-vae and trans- formers, 2021

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:34.250365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:34.250365Z digest=sha256:b9f12614bb5448fdf142807d8d797dfd652a7f1daa3e5d56f20000ecb1d506db

Observation 2c6cebde-293c-4dda-ba97-e344c47b0158 · outbound

This paper cites Generalized Predictive Model for Autonomous Driving.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Generalized Predictive Model for Autonomous Driving

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:16:35.035670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:16:34.254190Z digest=sha256:f17567c0c7f7dc4531ed6bc82ebcccee66036af3959515937fc73c1d0b32d20c

Observation 411a0b95-e99a-4c74-94ca-a56661d28e62 · outbound

This paper cites Depth anything: Unleashing the power of large-scale unlabeled data.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Depth anything: Unleashing the power of large-scale unlabeled data

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:16:35.022034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:16:34.257651Z digest=sha256:3eb47ef6282f39ce8bc982466389f9e4909d0c9c26c015af223f6f0df65acc2c

Observation 931b43e4-f357-451e-b850-3b1b9b6680f0 · outbound

This paper cites Depth Anything V2.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Depth Anything V2

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:34.261320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:34.261320Z digest=sha256:6a236d9153e10e629772ace81b3d4e7465c8873ba06b1f82dd63072ff3061fff

Observation cfc1a449-bff5-4400-b8ee-792964ca4c90 · outbound

This paper cites Learning interactive real-world simulators.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Learning interactive real-world simulators

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:16:35.008514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:16:34.265441Z digest=sha256:10b5976ce54c625e702209e9e8af3883db5f804bd638d883eef53ced99baee62

Observation dc4080eb-e9cd-427b-9ec4-7642e08da771 · outbound

This paper cites Effec- tive whole-body pose estimation with two-stages distillation.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Effec- tive whole-body pose estimation with two-stages distillation

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:16:34.995986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:16:34.268915Z digest=sha256:73043f87464028cc1edf87cafaf4120c1174a830122b467d2df08d66a6ac3c93

Observation 6a08682e-23b1-4848-bcf6-63cb486c9467 · outbound

This paper cites Visual point cloud forecasting enables scalable autonomous driving.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Visual point cloud forecasting enables scalable autonomous driving

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:16:34.983668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:16:34.272580Z digest=sha256:735654a1958bb66f375d98eb7e43238292f767baf693581f6bc8adefabfb83c9

Observation 5723f6bc-9fb7-4f25-930a-a69cfcad6f2d · outbound

This paper cites When, Where, and What? A New Dataset for Anomaly Detection in Driving Videos.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control When, Where, and What? A New Dataset for Anomaly Detection in Driving Videos

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:34.277111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:34.277111Z digest=sha256:8b53112658c8ef15a7c79b69b9b27d042a3d5425a45c01e009055bc19c2c36d6

Observation e924538b-f16b-4c7c-9aa8-3546111a436e · outbound

This paper cites Adding conditional control to text-to-image diffusion models.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Adding conditional control to text-to-image diffusion models

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:34.281670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:34.281670Z digest=sha256:b3dd97badd91bb06c5387f4a4d1c7569306f01785c6862cc3218ee8a2170d4bc

Observation 7986282e-c3f4-44e4-ba87-18961c7c0a15 · outbound

This paper cites Copilot4d: Learning unsupervised world models for autonomous driving via discrete diffusion,.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Copilot4d: Learning unsupervised world models for autonomous driving via discrete diffusion,

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:16:34.963543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:16:34.285298Z digest=sha256:bacdb87b49fe0d8a9f5b0418106a63ab49dc5bad23a43aed41cbb549e96f1b17

Observation 57d2e743-2eb4-4cff-bbad-77f4a5a42864 · outbound

This paper cites I2vgen-xl: High-quality image-to-video synthe- sis via cascaded diffusion models, 2023.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control I2vgen-xl: High-quality image-to-video synthe- sis via cascaded diffusion models, 2023

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:16:34.949449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:16:34.289172Z digest=sha256:e86185d191a0f57a29ea8f9d6b2249488ce8657b5c9059cf4b28390166e47d30

Observation a8b80298-4f0e-4f77-9a8e-3000b71d3179 · outbound

This paper cites MimicMotion: High-Quality Human Motion Video Generation with Confidence-aware Pose Guidance.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control MimicMotion: High-Quality Human Motion Video Generation with Confidence-aware Pose Guidance

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:34.298271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:34.298271Z digest=sha256:a44e1ff652d2b3a99a2aec11f37c21198efe96af4f8799ee648a6a167da00e7a

Observation 0eb6e488-ea6f-44ef-bea7-0d1b4f35d57f · outbound

This paper cites Can LLM Graph Reasoning Generalize beyond Pattern Memorization?.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Can LLM Graph Reasoning Generalize beyond Pattern Memorization?

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:34.303797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:34.303797Z digest=sha256:de3361769e5cc8d3bab479455157fe84ed46ca5a7ece2d5c320a93878beca431

Observation d9d0647d-6d8c-43fe-aa1f-9b74889d98b2 · outbound

This paper cites DriveDreamer-2: LLM-Enhanced World Models for Diverse Driving Video Generation.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control DriveDreamer-2: LLM-Enhanced World Models for Diverse Driving Video Generation

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:34.308100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:34.308100Z digest=sha256:b68198a0ff50d26a70bd575bf2e09a1f0cf7d76b7ee97c0f2fd042f8d885df0d

Observation 2bb1fa72-7ae8-404e-916b-bd23865f7562 · outbound

This paper cites OccWorld: Learning a 3D Occupancy World Model for Autonomous Driving.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control OccWorld: Learning a 3D Occupancy World Model for Autonomous Driving

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:34.312572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:34.312572Z digest=sha256:d929873cde793586029109eb79a4815a93e2b8e07a6f3142d0dad3ef70068271

Observation 97f7abb0-83a3-48d6-bf22-62d330648a67 · outbound

This paper cites On the continuity of rotation representations in neural networks.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control On the continuity of rotation representations in neural networks

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:34.317611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:34.317611Z digest=sha256:79c7db85af672430393291dedec841df78e27d3931e22dddef99c82ca61b0391

Observation 747f885b-9498-4c5e-b4c3-d3e0fc0699aa · outbound

This paper cites Is sora a world simulator? a comprehensive survey on general world models and beyond, 2024.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Is sora a world simulator? a comprehensive survey on general world models and beyond, 2024

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:16:34.927880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:16:34.321791Z digest=sha256:a568f66282f507309e800fd88d9abae50f0f8da308f8dba723eab9c7bf1afcee

Observation 8d07ec75-c5e7-4d2b-9bef-8cc8a9a0b194 · outbound

This paper cites Pseudo-labeling Depth.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Pseudo-labeling Depth

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:16:34.912699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:16:34.325405Z digest=sha256:bf5819083e6263426fda7b8a26410aedc81f60e7af7bd15a26a907742e56ed42

Observation a9d5013e-3dd9-4d7b-a38b-1b7e91e5efda · outbound

This paper cites Due to the increased size of our network, we in- corporate activation checkpointing and optimizer sharding to mitigate memory constraints, utilizing the DeepSpeed li- brary [50].

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Due to the increased size of our network, we in- corporate activation checkpointing and optimizer sharding to mitigate memory constraints, utilizing the DeepSpeed li- brary [50]

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:16:34.898427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:16:34.329547Z digest=sha256:9e3a283748f9004ed0a48ed49531570f6a7ad375ef99df91483f245f0bf2c7ad

Observation 629adf3f-3746-4c34-8de7-6350a4bfadfb · outbound

This paper cites To achieve fine- grained, high-quality control, we employ a two-stage train- ing regime, detailed as follows: 8.1.1.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control To achieve fine- grained, high-quality control, we employ a two-stage train- ing regime, detailed as follows: 8.1.1

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:16:34.884683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:16:34.333665Z digest=sha256:a3ef61d440fa82f6ebe96c1ad9b20ed7d513a94bd78061f5f613b61010afae7f

Observation 9d2cc4b5-3be1-482b-8a1c-22cbd71ad899 · outbound

This paper cites an unresolved cited work.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Unresolved cited work

Reference 91

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:16:34.870716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:16:34.337488Z digest=sha256:f901022eaa5988fefd3d8341ebf23f76c54e44a863d19614efd6f729495030c1

Observation a63f93af-8fd5-4512-a3a9-74f8bd13825a · outbound

This paper cites Depth generation quality comparison.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Depth generation quality comparison

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:16:34.855888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:16:34.341458Z digest=sha256:b769fcb01b36c181614c4dc90cf4e18d9e057dfb5444a83fd8aa583ae7bf13eb

Observation 04e3547b-2e58-49ad-b45e-cfdceeed581d · outbound

This paper cites 11 to 14 show qualitative examples of our generations, our controls, long generation and multimodal outputs.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control 11 to 14 show qualitative examples of our generations, our controls, long generation and multimodal outputs

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:16:34.841585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T15:16:34.345618Z digest=sha256:ca100957b416a7640ea447ed882b059d49523d3716ddd7a242646180d2a01b4e

Pith citing papers

Observation ac5b25be-eb94-4d8b-a35f-de8e19f1f6cc · inbound

GAIA-2: A Controllable Multi-View Generative World Model for Autonomous Driving cites this paper.

GAIA-2: A Controllable Multi-View Generative World Model for Autonomous Driving GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-15T13:48:22.339230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T13:48:22.279405Z digest=sha256:fda526d60be7fbf5983ab79fda654a109bc2ed99e73019bb34fd13059034f199

Observation e64e1309-a665-4a08-8dff-aa00159e5cea · inbound

GeoDrive: 3D Geometry-Informed Driving World Model with Precise Action Control cites this paper.

GeoDrive: 3D Geometry-Informed Driving World Model with Precise Action Control GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:13:42.902690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:13:42.902690Z digest=sha256:f17f88193ea1b0c7d21afc6b06f3777c64422d2419f542cd479743940c9508d9

Observation da65dc62-cd75-4d42-b702-276a4d3cc2e8 · inbound

Seeing Clearly, Forgetting Deeply: Revisiting Fine-Tuned Video Generators for Driving Simulation cites this paper.

Seeing Clearly, Forgetting Deeply: Revisiting Fine-Tuned Video Generators for Driving Simulation GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T17:19:11.673125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:19:11.673125Z digest=sha256:d47eeb7ac845e85dd5932ce7576c89b3c04b02627b9ae8557b6c31fb2892fd3c

Observation 8d9971b4-40b9-40d2-a10e-e34611b53554 · inbound

How You Move Tells What You'll Do: Trajectory-Conditioned Egocentric Prediction cites this paper.

How You Move Tells What You'll Do: Trajectory-Conditioned Egocentric Prediction GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:19:47.034828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:14:59.175822Z digest=sha256:72c7c27c04acaa639994b081141762a92c37ed11c0ecb2f358c71da4b3931d0b

Observation 645d106c-2757-4d39-905c-4faf702b1677 · inbound

Towards Interactive Video World Modeling: Frontiers, Challenges, Benchmarks, and Future Trends cites this paper.

Towards Interactive Video World Modeling: Frontiers, Challenges, Benchmarks, and Future Trends GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control

Reference 211

Resolution
verified exact
arxiv_id, observed 2026-07-01T21:06:13.799031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-28T17:29:18.513507Z digest=sha256:4b54a800d894f7853ec8e72aeedc11b9e8ff0c68de0c149a6a67c4c0a9d4a8b0

Observation e209b637-182d-4001-a185-e1286e9b4871 · inbound

Instant NuRec: Feed-Forward 3D Gaussian Reconstruction for Driving Scene Simulation cites this paper.

Instant NuRec: Feed-Forward 3D Gaussian Reconstruction for Driving Scene Simulation GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-02T02:55:45.953730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:55:45.953730Z digest=sha256:e82b540244608a11981a466fa159f1a1d30e445f580d5f4ea46659a1daa73c9a