Pith. sign in

Paper Citation Record · LEDGER

Back to the Features: DINO as a Foundation for Video World Models

As of 17 August 2026, this Paper Citation Record lists 70 of 70 outbound references and 19 inbound Pith citation observations for arXiv:2507.19468.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.19468 v1

Coverage vector

measured 70 of 70 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T14:24:27.413833Z

measured 89 of 89 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 19 of 19 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T00:13:32.849963Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T20:38:55.747199Z

Reference resolution

70 of 70 outbound references displayed

  • verified exact0
  • verified fuzzy50
  • unresolved19
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b1793104-a74f-4e87-964c-91983d040163 · outbound

This paper cites Recurrent world models facilitate policy evolution.

Back to the Features: DINO as a Foundation for Video World Models Recurrent world models facilitate policy evolution

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:24:31.055775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T14:24:19.077215Z digest=sha256:6b5b78423b27ae2917e82a991edce7791a2a736c1f37a2b381e04ba60860e44d

Observation 0179cea5-6424-4d8f-ac0e-f1d2fef54403 · outbound

This paper cites GAIA-1: A Generative World Model for Autonomous Driving.

Back to the Features: DINO as a Foundation for Video World Models GAIA-1: A Generative World Model for Autonomous Driving

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T14:24:19.230074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:24:19.230074Z digest=sha256:432f461b64a1006753944580d6230996b83692b04f59f262291d8f6c3dacb8c5

Observation 4b154396-d440-430a-8111-04af2298a4ba · outbound

This paper cites Learning interactive real-world simulators.

Back to the Features: DINO as a Foundation for Video World Models Learning interactive real-world simulators

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:24:31.041389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T14:24:19.337762Z digest=sha256:88de0b3c1a991509a026e8e4b819446a084a4a4b7558b14b02e3d160a0ae182b

Observation 7fb5ea1f-3d80-4c4e-8b2a-92e7f5e8334c · outbound

This paper cites Video gen- eration models as world simulators.

Back to the Features: DINO as a Foundation for Video World Models Video gen- eration models as world simulators

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:24:31.026043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T14:24:19.507700Z digest=sha256:8dad054cc8e5c382cae27accf7c7c3fdab64bf5287dd47efa35026f4121ba5ff

Observation 2d701c55-b1d1-4af9-8117-9b6a962f88d6 · outbound

This paper cites Genie: Generative interactive environments.

Back to the Features: DINO as a Foundation for Video World Models Genie: Generative interactive environments

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:24:31.011696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T14:24:19.584629Z digest=sha256:94be1345a45ecb1899d42e6af91983f8902b081bd572c381d9375798a3666b77

Observation 73d6bbeb-9f8c-491a-a2a3-78dd21eb5f6d · outbound

This paper cites Genie 2: A large- scale foundation world model.

Back to the Features: DINO as a Foundation for Video World Models Genie 2: A large- scale foundation world model

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:24:30.997168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T14:24:19.683448Z digest=sha256:c0178df3c8cf92a9a3d7e15c7b2fb8970834bea1da60fcdf1e162831aac6af6b

Observation 4adec688-ce17-42ab-b3d9-c6d341f9820c · outbound

This paper cites VaViM and VaVAM: Autonomous Driving through Video Generative Modeling.

Back to the Features: DINO as a Foundation for Video World Models VaViM and VaVAM: Autonomous Driving through Video Generative Modeling

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T14:24:19.834310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:24:19.834310Z digest=sha256:2b802c485c6d908b440cbe2ed872972ac60ed773d7b85c159f4b421fdc783df7

Observation ffa81ed9-aa70-491f-b8cb-3ff58afef1fb · outbound

This paper cites Cosmos World Foundation Model Platform for Physical AI.

Back to the Features: DINO as a Foundation for Video World Models Cosmos World Foundation Model Platform for Physical AI

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T14:24:19.930972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:24:19.930972Z digest=sha256:abefa9e6b3bf986f381f9b3fa555d07001bf4773dc267372c61ba9f3ab007ce2

Observation fc51d744-670f-40f3-9f11-8694ce637b23 · outbound

This paper cites GAIA-2: A Controllable Multi-View Generative World Model for Autonomous Driving.

Back to the Features: DINO as a Foundation for Video World Models GAIA-2: A Controllable Multi-View Generative World Model for Autonomous Driving

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T14:24:20.014729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:24:20.014729Z digest=sha256:fa6d317fc5d678ea442fa3a7a985bb55aff4186e4b1a12e94f88d0a264bb3b14

Observation bb315865-5db7-4819-bc0d-baafc365e7c2 · outbound

This paper cites Movie Gen: A Cast of Media Foundation Models.

Back to the Features: DINO as a Foundation for Video World Models Movie Gen: A Cast of Media Foundation Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T14:24:20.142280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:24:20.142280Z digest=sha256:036305dd2a5abecfa58b1bfb62db4f25488ce87a9038857e5212e3d72d2ff362

Observation bcbccda0-410e-4d95-a9a9-cb3ab5d0cd5d · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

Back to the Features: DINO as a Foundation for Video World Models Wan: Open and Advanced Large-Scale Video Generative Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T14:24:20.265832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:24:20.265832Z digest=sha256:9d9b2becad692f0d73cd851a41fd49ddb376c9ae3cec86fffee2f1c8daa378c1

Observation 0eccd809-ad50-47f5-9447-5aeef1810379 · outbound

This paper cites A path towards autonomous machine intelligence, 2022.

Back to the Features: DINO as a Foundation for Video World Models A path towards autonomous machine intelligence, 2022

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:24:30.983100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T14:24:20.456825Z digest=sha256:e2a6ee5934be3f71a325cb75b25c12cafe34366df579083ab0849ca9acc79777

Observation d0c234e6-1a2b-4618-9ab2-719c06c4d710 · outbound

This paper cites OpenEQA: Embodied question answering in the era of foundation models.

Back to the Features: DINO as a Foundation for Video World Models OpenEQA: Embodied question answering in the era of foundation models

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:24:30.967819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T14:24:20.586542Z digest=sha256:c3810d46bb0eb68d648ca766598a95228e0121f2bcfd7a4444b655879d18488a

Observation e02f184d-7cac-4cfc-81d6-28773b0476c5 · outbound

This paper cites Eyes wide shut? exploring the visual shortcomings of multimodal llms.

Back to the Features: DINO as a Foundation for Video World Models Eyes wide shut? exploring the visual shortcomings of multimodal llms

Reference 14

Resolution
malformed identifier
no resolver link, observed 2026-08-06T14:24:20.719121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:24:20.719121Z digest=sha256:c5bc1fa24ddb5a7866125901388db0c4ed6be44fd90d8b12af36fda0d3390aee

Observation 1f3285bb-c00b-4e96-95c8-4a43c1316cd9 · outbound

This paper cites Mastering diverse control tasks through world models.

Back to the Features: DINO as a Foundation for Video World Models Mastering diverse control tasks through world models

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:24:30.953943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T14:24:20.834970Z digest=sha256:8a9c06ec87a9c4d387eb9d1bfdabe27d91d889905b395a164aa9aeffc85c5b36

Observation ea5e4040-e89c-4eaf-ac5c-274cf314330b · outbound

This paper cites Diffusion for world modeling: Visual details matter in atari.

Back to the Features: DINO as a Foundation for Video World Models Diffusion for world modeling: Visual details matter in atari

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:24:30.939126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T14:24:20.915469Z digest=sha256:dcfe0d0ce5d4b0e740f1b215383f1fbb5702d7c2ac8abdc053fc1a69ffa05196

Observation ba063c0b-b3c1-41c5-96aa-68cefb1d2837 · outbound

This paper cites TD-MPC2: Scalable, robust world models for continuous control.

Back to the Features: DINO as a Foundation for Video World Models TD-MPC2: Scalable, robust world models for continuous control

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:24:30.925431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T14:24:21.018373Z digest=sha256:9c5159429b83d794b2736ffe06713d229fadc0197e9084d1900831491e1f3019

Observation 4db01b73-96d7-4aeb-a5fa-acb017067f80 · outbound

This paper cites DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning.

Back to the Features: DINO as a Foundation for Video World Models DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T14:24:21.154250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:24:21.154250Z digest=sha256:c617ced61c5f9b42ab02c98e860656455dc26d6f4b8fb48ff5b32ea0a012bebe

Observation 3c0dde91-6f37-42c5-a384-a85f000ee1a8 · outbound

This paper cites an unresolved cited work.

Back to the Features: DINO as a Foundation for Video World Models Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T14:24:21.241933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:24:21.241933Z digest=sha256:e062a30e2432b390cf3916fb45a088baff3d5e41dc593e3bdca90743d70255c9

Observation 53a54ef5-62be-437b-9dcd-efb08d7d9e91 · outbound

This paper cites DINO-Foresight: Looking into the future with dino.

Back to the Features: DINO as a Foundation for Video World Models DINO-Foresight: Looking into the future with dino

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T14:24:21.304379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:24:21.304379Z digest=sha256:742bd32e71705adf7df008ab2ad3078515082e4e7fd8035e9b961c9fe74bd2a4

Observation 016f5174-ad6c-4bcf-a506-bd1506d81b97 · outbound

This paper cites Intuitive physics understanding emerges from self-supervised pretraining on natural videos.

Back to the Features: DINO as a Foundation for Video World Models Intuitive physics understanding emerges from self-supervised pretraining on natural videos

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T14:24:21.361886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:24:21.361886Z digest=sha256:7e1221e1b9576b7d65a54626f1efefee101c8291f39547c5c8f8f1a145304672

Observation 654ba3bc-8375-471f-b089-b03401cd4432 · outbound

This paper cites Do generative video models understand physical principles?.

Back to the Features: DINO as a Foundation for Video World Models Do generative video models understand physical principles?

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T14:24:21.470319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:24:21.470319Z digest=sha256:8a794779f122641d5d3bd4a140a9a4f8fd3b1513cd196b22731934520cc4a983

Observation 36e28ed2-1d9e-4a14-9032-6cbdc86660a2 · outbound

This paper cites Dinov2: Learning robust visual features without supervision.

Back to the Features: DINO as a Foundation for Video World Models Dinov2: Learning robust visual features without supervision

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:24:30.911680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T14:24:21.647569Z digest=sha256:e5154a517b47840034068caeebbe0e4e64fcb8350c52b7a140bfac9b79cf2918

Observation 3b3052f3-e33d-4d7b-a8ae-1e252b6e5418 · outbound

This paper cites Revisiting feature prediction for learning visual representations from video.

Back to the Features: DINO as a Foundation for Video World Models Revisiting feature prediction for learning visual representations from video

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:24:30.896830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T14:24:21.744798Z digest=sha256:8bc4dae314bb302c76ea6d2722ebdabccef15952eed877595f113bfacb5dced6

Observation 4c73bb99-2c26-4c60-b9ed-c9a6c6e69886 · outbound

This paper cites an unresolved cited work.

Back to the Features: DINO as a Foundation for Video World Models Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:24:30.883309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T14:24:21.880961Z digest=sha256:761436b55ecc10f3fe36b10a734e7f4a7a4239cd34459a4aecadf2cadc63c845

Observation ef1d1eb6-988e-4a84-af0c-3f1db23db1bb · outbound

This paper cites Making the world differentiable: on using self supervised fully recurrent neural networks for dynamic reinforcement learning and planning in non-stationary environments.

Back to the Features: DINO as a Foundation for Video World Models Making the world differentiable: on using self supervised fully recurrent neural networks for dynamic reinforcement learning and planning in non-stationary environments

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:24:30.869882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T14:24:22.040712Z digest=sha256:dfb92454e50fde82a641f8a31a809b90a5dc96f0f2dff0b118ef8d0e9d27439c

Observation 26f586b3-8529-437f-89b5-dada133df858 · outbound

This paper cites Goodwin and K.S.

Back to the Features: DINO as a Foundation for Video World Models Goodwin and K.S

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:24:30.855677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T14:24:22.182653Z digest=sha256:687448e14829428bf98fd6e987f93107a9f1a4039e9a08fe7eb408f00dcc248b

Observation 19fc8569-9ab7-4751-8b8e-5d8ac6a8e45d · outbound

This paper cites Lillicrap, Jimmy Ba, and Mohammad Norouzi.

Back to the Features: DINO as a Foundation for Video World Models Lillicrap, Jimmy Ba, and Mohammad Norouzi

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:24:30.842316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T14:24:22.266415Z digest=sha256:f4f311376baf1fc70c8ac21a156f469850abc114e81d7144edbf7193af8ed5e1

Observation 3b706e76-2b7a-4f7d-a657-83b5d9e96710 · outbound

This paper cites Transformers are sample-efficient world models.

Back to the Features: DINO as a Foundation for Video World Models Transformers are sample-efficient world models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:24:30.827733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T14:24:22.432112Z digest=sha256:cc16ae9d81b256c218a3ec509bc786af446a672a280d7cae21ff3f50b5ff9572

Observation 8e661078-69cf-4222-95ea-870e109aac45 · outbound

This paper cites Lillicrap, Ian S.

Back to the Features: DINO as a Foundation for Video World Models Lillicrap, Ian S

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:24:30.809530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T14:24:22.521742Z digest=sha256:c09d6b33d44fb64b10fc2933823e467f839867df1268eb582124154289f12035

Observation 52f6a0c8-1e60-40e8-9650-88639314e742 · outbound

This paper cites Navigation World Models.

Back to the Features: DINO as a Foundation for Video World Models Navigation World Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T14:24:22.613110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:24:22.613110Z digest=sha256:fd2b04d0b450ce54de0d2ff7d04ba5a31a76049a23152df0c681b4d5c1a7ce09

Observation 0384afb7-0a13-4177-9083-d338d19ed56d · outbound

This paper cites Video pixel networks.

Back to the Features: DINO as a Foundation for Video World Models Video pixel networks

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:24:30.795669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T14:24:22.763667Z digest=sha256:eebfc44f02926897f0854d7e66be80955c8435578b372d04ae0380ea2fb837d9

Observation 804d5eee-2073-439a-ba4e-d462db24c9ab · outbound

This paper cites Campbell, and Sergey Levine.

Back to the Features: DINO as a Foundation for Video World Models Campbell, and Sergey Levine

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:24:30.781330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T14:24:22.874245Z digest=sha256:9653b0af541f30147efc03e9c5e43b9a6f4f1f5bb86b2c446be037cebf61294c

Observation 44a791d4-8383-43bc-a26a-aa3708502244 · outbound

This paper cites Stochastic video generation with a learned prior.

Back to the Features: DINO as a Foundation for Video World Models Stochastic video generation with a learned prior

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:24:30.767071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T14:24:22.976795Z digest=sha256:53780480580f2730628f529e7a878c89f7bd270a11e7e6f03c15e1b9f5d5e90d

Observation 565226b3-6805-4451-b262-4fc06a239848 · outbound

This paper cites VideoGPT: Video Generation using VQ-VAE and Transformers.

Back to the Features: DINO as a Foundation for Video World Models VideoGPT: Video Generation using VQ-VAE and Transformers

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T14:24:23.088610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:24:23.088610Z digest=sha256:3967047955de16362f079c6870ebf6411fdad55f30443b95f20dfef69aa8b0a0

Observation 29563b80-7240-4290-ae59-85480841aacb · outbound

This paper cites Hauptmann, Ming-Hsuan Yang, Yuan Hao, Irfan Essa, and Lu Jiang.

Back to the Features: DINO as a Foundation for Video World Models Hauptmann, Ming-Hsuan Yang, Yuan Hao, Irfan Essa, and Lu Jiang

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:24:30.750767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T14:24:23.203680Z digest=sha256:39fa44470221621e338a284fb2047f522b4b5dee3d2e291e828bcde0797bf540

Observation bb5284a2-4701-4b36-879f-9bd3700b14a7 · outbound

This paper cites Photorealistic video generation with diffusion models.

Back to the Features: DINO as a Foundation for Video World Models Photorealistic video generation with diffusion models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:24:30.735696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T14:24:23.318467Z digest=sha256:546a2a8a9ef9e2ba799206d5880b140db10daa36c9e8acbc9750c416d03be0f1

Observation dce58889-429a-411d-a136-2d6f2ed8005b · outbound

This paper cites Phenaki: Variable length video generation from open domain textual descriptions.

Back to the Features: DINO as a Foundation for Video World Models Phenaki: Variable length video generation from open domain textual descriptions

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:24:30.721118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T14:24:23.443562Z digest=sha256:5bada34d53142243d1b6d93c7591f8a0e5b85de490f9a5be6566ce6326ff2e3e

Observation 01cadf25-56c1-4351-8f1d-db96fb6ef124 · outbound

This paper cites an unresolved cited work.

Back to the Features: DINO as a Foundation for Video World Models Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:24:30.706950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T14:24:23.584776Z digest=sha256:cdd559ab5c1d976b4abad6cc07e4dd8a85795f28bb4d2c7678915d6376b57a72

Observation 3c46f61a-997a-47da-8a7d-3af9cf1aa378 · outbound

This paper cites Action-conditional video prediction using deep networks in atari games.

Back to the Features: DINO as a Foundation for Video World Models Action-conditional video prediction using deep networks in atari games

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:24:30.693050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T14:24:23.693050Z digest=sha256:c151ade440db10719f9353cfbbcfc402ec16d07cc8dd5ec5edf25a952950319e

Observation b44315c2-8172-4e68-b4ba-9daa4ac01abd · outbound

This paper cites Recurrent environment simulators.

Back to the Features: DINO as a Foundation for Video World Models Recurrent environment simulators

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:24:30.679684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T14:24:23.796136Z digest=sha256:340df3b3ffa1150f1f3d949fea18b66fbd021db86318c6452270cc0d3ba4ae2f

Observation c172c475-0d52-44f7-8ffa-e3ec1e963f03 · outbound

This paper cites How Far is Video Generation from World Model: A Physical Law Perspective.

Back to the Features: DINO as a Foundation for Video World Models How Far is Video Generation from World Model: A Physical Law Perspective

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T14:24:23.915976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:24:23.915976Z digest=sha256:eb8ccb195fb68a6e40369f74faaa37f7651bb1fec5d70fc52bb401729b309d14

Observation 9fe1ac1d-86aa-4c3b-81bb-e7701028d066 · outbound

This paper cites Predicting deeper into the future of semantic segmentation.

Back to the Features: DINO as a Foundation for Video World Models Predicting deeper into the future of semantic segmentation

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:24:30.665757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T14:24:24.025940Z digest=sha256:d9f3efed96267e57f83359fd02911a3991c169c0cd12644b374d8a138b5ea234

Observation 1eb57702-7e8d-4226-b2fe-eb621e331a5e · outbound

This paper cites Segmenting the future.

Back to the Features: DINO as a Foundation for Video World Models Segmenting the future

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:24:30.652445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T14:24:24.189747Z digest=sha256:eb13df97577e924f8f64434a8e5493866956ec270e49c6b91ac3b0926b295ab7

Observation 68260170-b9e6-4000-b060-9342b2c4ff86 · outbound

This paper cites Predicting future instance segmentation by forecasting convolutional features.

Back to the Features: DINO as a Foundation for Video World Models Predicting future instance segmentation by forecasting convolutional features

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:24:30.638948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T14:24:24.324980Z digest=sha256:034d39f8f1f1c51423fda1a33c14c064891529aed6b27a5d1f7d1b933c1677c1

Observation 2628495f-951c-4f04-a059-b058736068da · outbound

This paper cites Anticipating visual representations from unlabeled video.

Back to the Features: DINO as a Foundation for Video World Models Anticipating visual representations from unlabeled video

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:24:30.624286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T14:24:24.422184Z digest=sha256:6ff14d720f394362b11436ac07ec81591d780852d9c03197f568d713855e85d5

Observation a016a948-361f-4659-a82a-eeae9b4a274c · outbound

This paper cites Anticipative video transformer.

Back to the Features: DINO as a Foundation for Video World Models Anticipative video transformer

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:24:30.609549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T14:24:24.592648Z digest=sha256:fb1d9e6befe4d00832884bc26164422bf6872391ec43dced656123df90c04d13

Observation 9148fbf3-2de1-40ca-889d-cee7904f7a5c · outbound

This paper cites Anticipative feature fusion transformer for multi-modal action anticipation.

Back to the Features: DINO as a Foundation for Video World Models Anticipative feature fusion transformer for multi-modal action anticipation

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:24:30.594484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T14:24:24.712355Z digest=sha256:8c8c594941d976e1477b90ee29095217f65af48b8a646273f50529a816acf4a3

Observation eaf442db-a459-40d0-89ac-26a65676b1cd · outbound

This paper cites Auto-encoding variational bayes.

Back to the Features: DINO as a Foundation for Video World Models Auto-encoding variational bayes

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:24:30.580519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T14:24:24.837044Z digest=sha256:682decc5f5ecee170b2ba16216d87539ce7b10bc506ee5e7607776dd9a370c42

Observation c738cca5-7455-441b-87b1-54eefcf11fe9 · outbound

This paper cites Neural discrete representation learning.

Back to the Features: DINO as a Foundation for Video World Models Neural discrete representation learning

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:24:30.566076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T14:24:24.973247Z digest=sha256:cc81e06cfb1ced711e1160eed88f1baefc541b415afeb644653af283acee4217

Observation a6acc1d8-ba77-4c9d-8a94-b5d7dc414d3a · outbound

This paper cites Siglip 2: Multilingual vision-language encoders with improved semantic understanding, localization, and dense features, 2025.

Back to the Features: DINO as a Foundation for Video World Models Siglip 2: Multilingual vision-language encoders with improved semantic understanding, localization, and dense features, 2025

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:24:30.551500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T14:24:25.107537Z digest=sha256:8a115a3f5b8d87e43900089b56a39cfc788bed55e9fe271296f67fb22b3a7e94

Observation b8227ae9-e0dd-4a69-8b53-caeaff8f0025 · outbound

This paper cites Is sora a world simulator? a comprehensive survey on general world models and beyond.

Back to the Features: DINO as a Foundation for Video World Models Is sora a world simulator? a comprehensive survey on general world models and beyond

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T14:24:25.211077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:24:25.211077Z digest=sha256:9537d41a63bff37d4e5d5cf8d572d2c9aa3a1ea0cc4383534f43fd2a05ba6595

Observation 39511b93-9ee4-443d-897b-b0109e6079c5 · outbound

This paper cites Attention is all you need.

Back to the Features: DINO as a Foundation for Video World Models Attention is all you need

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:24:30.537053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T14:24:25.350254Z digest=sha256:cda380c30d81674f54c41171cd90bde79112d19d696902518dd2b2c503183dae

Observation eed14a7d-ab82-4edb-9878-e52b80edcad6 · outbound

This paper cites Rethinking patch dependence for masked autoencoders.

Back to the Features: DINO as a Foundation for Video World Models Rethinking patch dependence for masked autoencoders

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:24:30.523486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T14:24:25.415133Z digest=sha256:0093f45ded2a103de0077cda218e636d175efa57938fee1e14b315c4168ef76f

Observation 4ac9ebb9-dc8c-4d04-b230-e64a4fec428b · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding.

Back to the Features: DINO as a Foundation for Video World Models Roformer: Enhanced transformer with rotary position embedding

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:24:30.509534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T14:24:25.556611Z digest=sha256:5f72382eae20197d7f4acda4365d9139f1a035cb07cdd8b0c9351411168826e0

Observation 68af77d5-c39d-42b9-b83f-16f835a1cf8b · outbound

This paper cites Vision transformers need registers.

Back to the Features: DINO as a Foundation for Video World Models Vision transformers need registers

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:24:30.495017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T14:24:25.695128Z digest=sha256:42c7171cec5ecd3545af80d2bb1f38219b80a6de23cbb2228532695d3fe2d758

Observation e4f855f6-375b-4026-a3b6-823c7d425163 · outbound

This paper cites Decoupled weight decay regularization.

Back to the Features: DINO as a Foundation for Video World Models Decoupled weight decay regularization

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:24:30.480014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T14:24:25.818978Z digest=sha256:12f09a6e76b8f6b3fa32cc48e2449448fc11a7571c6ff3d359b180f6d3c85b55

Observation a06719da-5e2a-4bf3-a51a-6c7ebfa765ac · outbound

This paper cites The cityscapes dataset for semantic urban scene understanding.

Back to the Features: DINO as a Foundation for Video World Models The cityscapes dataset for semantic urban scene understanding

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:24:30.464341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T14:24:25.921907Z digest=sha256:cc0610c037e144f43098021edf4cb6de2cd99974fd6c20e1e95f1a7a905b7e1d

Observation 976a6f12-4538-489b-b861-bb5e2580d606 · outbound

This paper cites something something.

Back to the Features: DINO as a Foundation for Video World Models something something

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:24:30.405943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T14:24:26.044296Z digest=sha256:c600baeda9517ef1f4315da4053a428513047c64d9a8858d6405236a1e072034

Observation 286aa182-e8f9-43dd-a393-f26013f487e0 · outbound

This paper cites HowTo100M: Learning a text-video embedding by watching hundred million narrated video clips.

Back to the Features: DINO as a Foundation for Video World Models HowTo100M: Learning a text-video embedding by watching hundred million narrated video clips

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:24:30.143267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T14:24:26.154041Z digest=sha256:be7bcd3d10b94ef4528ac352f36dfbf4cd4f89c908b8aa87250b5d382d742334

Observation 4333be0b-3730-4a2a-83f7-8fb509f8e3f5 · outbound

This paper cites The Kinetics Human Action Video Dataset.

Back to the Features: DINO as a Foundation for Video World Models The Kinetics Human Action Video Dataset

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T14:24:26.303259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:24:26.303259Z digest=sha256:ef6b9d513dc8ac521fc35cd435d8dbf1db6e3a62c07086927468c652702300b6

Observation 6f9f5bbd-1ed7-4973-98be-02026336fcce · outbound

This paper cites Vspw: A large-scale dataset for video scene parsing in the wild.

Back to the Features: DINO as a Foundation for Video World Models Vspw: A large-scale dataset for video scene parsing in the wild

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:24:29.977033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T14:24:26.392949Z digest=sha256:5ab0e70521f267e8ddd56a8d4622f8e8a5e2c4951facb43da57454f8caf7118e

Observation c5d6f065-3eca-452f-8ff1-22224bcd48fb · outbound

This paper cites Vision meets robotics: The kitti dataset.

Back to the Features: DINO as a Foundation for Video World Models Vision meets robotics: The kitti dataset

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:24:29.765314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T14:24:26.501283Z digest=sha256:a2f881feafa7d33088024b4d9441592d414e6e5846b67ed7c16b017527c61123

Observation f433dafc-05c4-462c-8ddb-ff829a673852 · outbound

This paper cites Intphys 2019: A benchmark for visual intuitive physics understanding.

Back to the Features: DINO as a Foundation for Video World Models Intphys 2019: A benchmark for visual intuitive physics understanding

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:24:29.528079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T14:24:26.639434Z digest=sha256:3f949988f8fce2c42986327a457f11da4a05a86fdfec696cd32ac4b31adfc04f

Observation 2ac9a11d-82c4-4829-bb97-74c685065b4c · outbound

This paper cites Grasp: A novel benchmark for evaluating language grounding and situated physics understanding in multimodal language models.

Back to the Features: DINO as a Foundation for Video World Models Grasp: A novel benchmark for evaluating language grounding and situated physics understanding in multimodal language models

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:24:29.215945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T14:24:26.806966Z digest=sha256:8045bba164d40f731e23961270a3e1d799257a880fccdc3bdfd1692a1a7440f5

Observation c268f0f6-7bb6-43a1-9e2d-9435c87e56a4 · outbound

This paper cites Benchmarking progress to infant-level physical reasoning in AI.

Back to the Features: DINO as a Foundation for Video World Models Benchmarking progress to infant-level physical reasoning in AI

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:24:28.996247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T14:24:26.935176Z digest=sha256:d64b630808c5edc1af6180e4870f980a5b8bab620d1d7716ab9801264c9412ce

Observation 1740f303-981d-4def-b921-67142f76e457 · outbound

This paper cites Scaling rectified flow transformers for high-resolution image synthesis.

Back to the Features: DINO as a Foundation for Video World Models Scaling rectified flow transformers for high-resolution image synthesis

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:24:28.842843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T14:24:27.014609Z digest=sha256:9cc046cc95ba7129ae1b8bcc6e27e1cab67a39f723c76637ffecaa229482902a

Observation d1c76065-4ee6-4099-a90b-120c29683baf · outbound

This paper cites Diffusion policy: Visuomotor policy learning via action diffusion.

Back to the Features: DINO as a Foundation for Video World Models Diffusion policy: Visuomotor policy learning via action diffusion

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:24:28.524596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T14:24:27.155767Z digest=sha256:899741207bbee5d4e07fd03db2be7697de8e0e813c5954812ad68a89ad58a3af

Observation d2fa7a46-05c6-4f23-94ad-e2e3f5929667 · outbound

This paper cites D4RL: Datasets for Deep Data-Driven Reinforcement Learning.

Back to the Features: DINO as a Foundation for Video World Models D4RL: Datasets for Deep Data-Driven Reinforcement Learning

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T14:24:27.292378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:24:27.292378Z digest=sha256:a607615d4377c9abbad96650b84f0422d4c6b0d2db9b8a4e001c304d567d6711

Observation 90779470-eb7c-4758-bf85-4ec3483610a5 · outbound

This paper cites frameskip.

Back to the Features: DINO as a Foundation for Video World Models frameskip

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:24:28.242547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T14:24:27.413833Z digest=sha256:bdd8f41ef7dbf2b93023db3f7df74b04b8ad9ae8091be3855cd09b18c08b0f90

Pith citing papers

Observation 45352b30-b791-45f5-99dd-c580d9001316 · inbound

A Comprehensive Survey on World Models for Embodied AI cites this paper.

A Comprehensive Survey on World Models for Embodied AI Back to the Features: DINO as a Foundation for Video World Models

Reference 182

Resolution
unresolved
no resolver link, observed 2026-08-04T09:12:50.711030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:12:50.711030Z digest=sha256:8851952c382cfe5b6475b4661169c93576217f805b1fc06ce9f8973f9bd31495

Observation 6990b1c1-0b58-4d1d-bea8-3cf3c3357260 · inbound

What Drives Success in Physical Planning with Joint-Embedding Predictive World Models? cites this paper.

What Drives Success in Physical Planning with Joint-Embedding Predictive World Models? Back to the Features: DINO as a Foundation for Video World Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-21T15:34:15.048528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-21T15:33:24.616338Z digest=sha256:49e733978b5733e92f40df5bbf7d341d06b64d6a881d974f5abe0c8818915178

Observation 2d0f6416-88ae-45fd-90ad-13e7e65f9fc5 · inbound

Representations Before Pixels: Semantics-Guided Hierarchical Video Prediction cites this paper.

Representations Before Pixels: Semantics-Guided Hierarchical Video Prediction Back to the Features: DINO as a Foundation for Video World Models

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T08:45:58.662898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T16:30:53.578491Z digest=sha256:7b53f644ce2ccff2c245d039ce8f4d0cf52257a5fe56183664c5bd3780821ee5

Observation c622365f-1f1e-4c0b-86fa-1a3eabadefd5 · inbound

Learning Long-term Motion Embeddings for Efficient Kinematics Generation cites this paper.

Learning Long-term Motion Embeddings for Efficient Kinematics Generation Back to the Features: DINO as a Foundation for Video World Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:36:05.038995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T15:55:26.222900Z digest=sha256:f159c58dd085505cae447f282432b22c01bedb1ab6d7efc141d571731c3a5ca5

Observation ff25bd88-1c4c-4954-a78c-2a65cc27c661 · inbound

Video Generation with Predictive Latents cites this paper.

Video Generation with Predictive Latents Back to the Features: DINO as a Foundation for Video World Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:26:09.276917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-09T16:55:54.709581Z digest=sha256:36e6f071ecbdf4d6622aae63fc701dd7d4887511c73709f049fa2c61340238d3

Observation 3d292a4f-e3b0-4ca5-8273-791a8fff967f · inbound

Text-Conditional JEPA for Learning Semantically Rich Visual Representations cites this paper.

Text-Conditional JEPA for Learning Semantically Rich Visual Representations Back to the Features: DINO as a Foundation for Video World Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:51:30.617299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-07T17:58:43.650308Z digest=sha256:d234612fa2209726050410827df61cb0ac54d8fd15e81d56ae05eaf9be2b7d93

Observation fbfa1ba5-dd6c-4464-80fa-3e540fd130ff · inbound

Reconstruction or Semantics? What Makes a Latent Space Useful for Robotic World Models cites this paper.

Reconstruction or Semantics? What Makes a Latent Space Useful for Robotic World Models Back to the Features: DINO as a Foundation for Video World Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:56:07.075706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-08T13:16:27.868477Z digest=sha256:27c25f319c2669e74bfe0f88132969943bc40a3211c2826cb1f58fb34c42d30e

Observation 4db30d88-b06a-4425-b9c8-8c25c42dd8a0 · inbound

Learning Visual Feature-Based World Models via Residual Latent Action cites this paper.

Learning Visual Feature-Based World Models via Residual Latent Action Back to the Features: DINO as a Foundation for Video World Models

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:40:51.976743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-11T01:37:27.658078Z digest=sha256:0682db26dd47ca940e609c018c77eb8f5874830458e939e546ccae9b54ebf1a6

Observation daa4d663-5748-4cd1-bb00-e70f719cbccc · inbound

Back to Parsimonious Latents: Learning Task-Centric World Models from Visual Foundations cites this paper.

Back to Parsimonious Latents: Learning Task-Centric World Models from Visual Foundations Back to the Features: DINO as a Foundation for Video World Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:04:00.631958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-29T21:43:59.953809Z digest=sha256:574c9f1cb9346e5953c99b2eb4ed8873edff7d9f07a11e8ea7a116b8023147d9

Observation ae265a73-b4f7-4c7b-b72d-bfb4a17d3ea0 · inbound

Envision4D: Envisioning Visual Futures via Feed-forward 4D Gaussian Splatting for Autonomous Driving cites this paper.

Envision4D: Envisioning Visual Futures via Feed-forward 4D Gaussian Splatting for Autonomous Driving Back to the Features: DINO as a Foundation for Video World Models

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-07-03T04:37:36.945597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T13:48:15.379724Z digest=sha256:6f732c2ef126dc3b2e3129b99754e6e8be51721bb10cb2e1f2bc8b600aa15efd

Observation bb059efd-1c82-45a5-bb56-67e876b71806 · inbound

Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application cites this paper.

Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application Back to the Features: DINO as a Foundation for Video World Models

Reference 214

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:58:02.812991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T09:46:30.702256Z digest=sha256:c6dcfdf7fd0162000b5d3d7dfd0010e09ac7e115844f1fd44eb9e8333e3bdcf7

Observation c07cf121-1c4a-42a8-8068-55d7b3522e2e · inbound

Future Dynamic 3D Reconstruction: A 3D World Model with Disentangled Ego-Motion cites this paper.

Future Dynamic 3D Reconstruction: A 3D World Model with Disentangled Ego-Motion Back to the Features: DINO as a Foundation for Video World Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:38:55.748766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-27T01:14:19.195054Z digest=sha256:57a5b6262a83765173385d1cf8095b1d8e1e6b7c446d322ff38f1e66d73b0b89

Observation 1c8089d9-01b6-4a7d-bfb3-f439e3855bb7 · inbound

Learning Task-Sufficient World Models by Synergizing Agentic Exploration and Structured Modeling cites this paper.

Learning Task-Sufficient World Models by Synergizing Agentic Exploration and Structured Modeling Back to the Features: DINO as a Foundation for Video World Models

Reference 117

Resolution
unresolved
no resolver link, observed 2026-07-11T19:24:48.899301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T19:24:48.899301Z digest=sha256:94521d3efa5fcf454c7103aa6626ec5eb33463dc4c837799c6b86faf2039b723

Observation 645df39c-c401-4de0-81ef-3b53daab1ac7 · inbound

Self-Supervised Learning of Structured Dynamics from Videos cites this paper.

Self-Supervised Learning of Structured Dynamics from Videos Back to the Features: DINO as a Foundation for Video World Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:42.787898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:42.787898Z digest=sha256:7fc7bf200e8678547172f167d3ad48444a0171c045eaf3868be340e9fa65d30f

Observation f1a7fb7f-ec3d-455c-afbc-2c92c56a6548 · inbound

Failure Detection for Surgical Robot Imitation Policies via Flow-Matching World Modeling cites this paper.

Failure Detection for Surgical Robot Imitation Policies via Flow-Matching World Modeling Back to the Features: DINO as a Foundation for Video World Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T06:34:04.062862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:34:04.062862Z digest=sha256:1a6b069abe7ec223a6d6a54e46a9b4f20014c488000db6cb7107460f8c64828e

Observation 291abc56-8acb-4515-949a-231de4f18a6e · inbound

DF$^3$: World Modeling via Decoder-Free Feature Forecasting in Autonomous Navigation cites this paper.

DF$^3$: World Modeling via Decoder-Free Feature Forecasting in Autonomous Navigation Back to the Features: DINO as a Foundation for Video World Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T07:37:52.354911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:37:52.354911Z digest=sha256:661302c41bf1d3ddbd835a27952bff8ba9a05134910457c1c17f30e3b756e22d

Observation 4e666f9c-e272-4158-ab78-ad6e6b9fe3e9 · inbound

Uncertainty-Aware World Model for Aerial Image-Goal Navigation cites this paper.

Uncertainty-Aware World Model for Aerial Image-Goal Navigation Back to the Features: DINO as a Foundation for Video World Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T05:52:17.229060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:52:17.229060Z digest=sha256:4358df818aaf2a9b12128ac322d37df503368c9d3e21932616c3c29f7eb10c5b

Observation 352b986b-b00f-4e1c-950a-b0af7760ac81 · inbound

UniJEPA: A Unified Joint-Embedding Predictive Architecture for Task-Agnostic Visual World Modeling cites this paper.

UniJEPA: A Unified Joint-Embedding Predictive Architecture for Task-Agnostic Visual World Modeling Back to the Features: DINO as a Foundation for Video World Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-10T05:06:02.500245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:06:02.500245Z digest=sha256:b2f60d5dd35c5410e2df67b5c715e641baf8c362ed26398207b7268ea92e0710

Observation e49d4b57-609f-4f4a-bc96-a628a10630a8 · inbound

Latent World Models with Monotone Planning Costs for Image-Goal Navigation cites this paper.

Latent World Models with Monotone Planning Costs for Image-Goal Navigation Back to the Features: DINO as a Foundation for Video World Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T00:13:32.849963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:13:32.849963Z digest=sha256:46d29b24ab6dd8fba6a89d2cc17b5dec8ad52c4336a7a0ef34e30370b3a8bd2a