Pith. sign in

Paper Citation Record · LEDGER

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

As of 5 August 2026, this Paper Citation Record lists 69 of 69 outbound references and 100 inbound Pith citation observations for arXiv:2506.09985.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.09985 v1

Coverage vector

measured 69 of 69 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-11T00:33:50.471804Z

measured 169 of 169 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 100 of 376 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T06:04:03.420693Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

69 of 69 outbound references displayed

  • verified exact48
  • verified fuzzy9
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch11

External citation measurements

3
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 201bc43f-a55d-4a40-abd1-5a4ab1bd16a4 · outbound

This paper cites Cosmos World Foundation Model Platform for Physical AI.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning Cosmos World Foundation Model Platform for Physical AI

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-11T00:33:50.706975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:1fe5fa619906397ee2b043321b67f7d1f25fe85ee5b20ca90939637b8536e7ce

Observation 84124e4f-7ec8-46d3-9e7c-2f185788a459 · outbound

This paper cites The Hidden Uniform Cluster Prior in Self-Supervised Learning.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning The Hidden Uniform Cluster Prior in Self-Supervised Learning

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:33:50.615165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:69d3e870f370266edb84b7ae7775f5e0ed19be7b9e9267dbb3ac1d1d6e58ad6c

Observation 0c4c3889-4ac4-43e6-b0b7-542a2be6494a · outbound

This paper cites Revisiting Feature Prediction for Learning Visual Representations from Video.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning Revisiting Feature Prediction for Learning Visual Representations from Video

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-12T12:40:24.452580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:0ddc7636b76c192221e0a102b9216e58cedf8817c42cb5660f2361c1d4f517d0

Observation 1cac71f1-58df-4132-bd8c-e6bfe8e9cf69 · outbound

This paper cites GR00T N1: An Open Foundation Model for Generalist Humanoid Robots.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-11T00:33:50.658349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:3b55615d18977244f0f65c2229f049de190c708bfd547127bd3a09309c6ec63f

Observation 5b6f6463-1cb4-4704-8729-e5b45bc57d45 · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-11T00:33:50.663856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:266cf1c64064d9fee2da5c67b7baacb9d6396c2d8d08ad798c814135ab0ee12e

Observation 72dc4682-8985-4280-81fd-d359e9130445 · outbound

This paper cites $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-11T00:33:50.677015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:b69bb721d7b2ef7432b0f5503ca431625526d18bcb27018d20f5a3f3b07af5c3

Observation a641c187-947a-4c2a-a8d5-7f5e6f296f4c · outbound

This paper cites Perception Encoder: The best visual embeddings are not at the output of the network.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning Perception Encoder: The best visual embeddings are not at the output of the network

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:21:16.084670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:827957a966a1fdd42cb7057869703286ed2e8500354c56df3b138410af7a8d03

Observation c6b04c45-8f01-4280-97ac-3488f0dd72d2 · outbound

This paper cites RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-11T00:33:50.699817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:cad3a1d17f1bc3053f181e249458a40df817f4774c670d032885e341c674cc58

Observation b3e7157c-7b90-4817-989b-f5200612b389 · outbound

This paper cites TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:33:50.948202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:f35a8837bbaa3c870a9cc524ae35e5eb75c758cf1d8ff51a4ff5eceba676a0ed

Observation 44f9af9e-5d1a-4739-96ae-26e3ee04a784 · outbound

This paper cites A Short Note about Kinetics-600.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning A Short Note about Kinetics-600

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:33:50.714339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:83f8d391895bce76293ddd4f6edf3a77a6b54a5d797610a2f69bef1008d1fae4

Observation cb071916-8124-410a-98cb-d137010fa565 · outbound

This paper cites A Short Note on the Kinetics-700 Human Action Dataset.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning A Short Note on the Kinetics-700 Human Action Dataset

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:33:50.725871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:51aa877af41933c337b84e0a3da1e7a5ff0ec4fd05a611de5bc1be2e3c8b4a1f

Observation 3f95764d-646c-414b-9d2c-358331e7599a · outbound

This paper cites Scaling 4D Representations.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning Scaling 4D Representations

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:33:50.730654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:6b7e2dc4d17d29631de84da7bf92475401c3044900dbccc7cb2fc1408aec5e3d

Observation cb939337-9d3a-4bff-99dd-4dab5c33a910 · outbound

This paper cites Actionable Models: Unsupervised Offline Reinforcement Learning of Robotic Skills.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning Actionable Models: Unsupervised Offline Reinforcement Learning of Robotic Skills

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:33:50.741002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:f471539f4b10a9e66b01bb1e38d8a14fb9da757545830e0ba4f39d509d58850a

Observation 3cd5d2a7-0d85-490e-a597-fadc8b1949d0 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-11T00:33:50.745932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:84272e84ffc6ef6466d688c5105e97d849f8891c522dd46c46c573d39b1f1dab

Observation 5481e072-022d-4fb6-974c-a1f73ccbf903 · outbound

This paper cites PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:33:50.751197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:91de967dec5b23b1b44f485e70b6273cf9636ca4164e41263494d8ae09e9f89e

Observation 2ddcf280-f912-4b04-a12e-5e84b7c485e2 · outbound

This paper cites Lost in Time: A New Temporal Benchmark for VideoLLMs.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning Lost in Time: A New Temporal Benchmark for VideoLLMs

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T00:33:50.755434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:6d3a702da9722782060f1c0533f713db1fd1e1363696224d8b609f83b9f624d3

Observation 1605c1b9-48e0-4470-8876-c126e7c7da52 · outbound

This paper cites Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens-100.International Journal of Computer Vision (IJCV).

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens-100.International Journal of Computer Vision (IJCV)

Reference 17

Resolution
metadata mismatch
doi, observed 2026-05-11T00:33:50.598343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:16be632426408dfaea6658a56060ea0e1f3fd5fabca74a4b00e65fb9cdc2f046

Observation 26548820-34cc-46ec-b905-9e3d7a804b1f · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-11T00:33:50.759248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:2314788362aa49cd30c5f44d5d0c975e1d89145fd6d02d3cda16451db0c2ae9f

Observation 57437d9d-66ca-4e5a-87c3-adf9bf9e453e · outbound

This paper cites Driess, F.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning Driess, F

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:33:50.763515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:8f94ca580d5e53265ad8de7136b0697f1418b9d69e09a9f8b1dffee2ee835627

Observation fcb8aa3a-2c89-4015-82af-6b32beb1c621 · outbound

This paper cites Visual Foresight: Model-Based Deep Reinforcement Learning for Vision-Based Robotic Control.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning Visual Foresight: Model-Based Deep Reinforcement Learning for Vision-Based Robotic Control

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:33:50.779946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:051e192bb16e79dcdc19d645e0c8183555a0c11b9f95ce1bea6daf9d4ce8a43b

Observation 0b4fc876-9190-4d61-8583-94c60ded441c · outbound

This paper cites Scaling Language-Free Visual Representation Learning.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning Scaling Language-Free Visual Representation Learning

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:33:50.790701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:fb2c9145ad0758bd6660547e6949a355db02b3b29642050215c0268172664849

Observation 0d1cdbad-61e0-4aa1-8757-c176ebea21e6 · outbound

This paper cites Multimodal Autoregressive Pre-training of Large Vision Encoders.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning Multimodal Autoregressive Pre-training of Large Vision Encoders

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T00:33:50.796334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:9b98c1850cab15f705b5ed759af3d4514a42d5ac0e47fdfa549749b7e80ffa48

Observation 1df28384-6207-4a77-8bf9-e715e6d7bd4f · outbound

This paper cites Learning Visual Predictive Models of Physics for Playing Billiards.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning Learning Visual Predictive Models of Physics for Playing Billiards

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:33:50.811345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:56c3c71da8221eb66e121d70f0741ec07dd8607e60a1170037a4d44cea0dab0b

Observation da9481c3-0afc-4f49-993a-4ed25268cb13 · outbound

This paper cites He, B., Yin, L., Zhen, H.-L., Liu, S., Wu, H., Zhang, X., Yuan, M., and Ma, C.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning He, B., Yin, L., Zhen, H.-L., Liu, S., Wu, H., Zhang, X., Yuan, M., and Ma, C

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T00:33:50.821430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:84d3689744c2826006526134e951f7779214fb2f043000ced0a091b2e434e41f

Observation 3cb7acec-43f4-4ac7-8ef5-9fbac5a594c6 · outbound

This paper cites The Llama 3 Herd of Models.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning The Llama 3 Herd of Models

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-11T00:33:50.829441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:ed6aedb71f1d8fbe1465b38d36961305b924ff6d4d00b96c3063bb83d32edcf4

Observation c7a600d1-840b-4ff7-86ef-153bff29820d · outbound

This paper cites MaskViT: Masked Visual Pre-Training for Video Prediction.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning MaskViT: Masked Visual Pre-Training for Video Prediction

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T00:33:50.835310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:2a5b8089094ae0c63c67c091024374840c0a227ea069e6e2eeff7f67a348d614

Observation cbbef07d-29b2-47c4-95d4-fddd0f75e9ac · outbound

This paper cites World Models.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning World Models

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:08:36.665942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:3cec764e6ecdf803a3d0d36d5a77278abbd6fca4adfe7ca5dc139c917ce0fddd

Observation 19a1361b-43c3-48e1-aff6-8e92026358e7 · outbound

This paper cites Dream to Control: Learning Behaviors by Latent Imagination.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning Dream to Control: Learning Behaviors by Latent Imagination

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-12T01:16:36.618198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:65cd555154492e116853cc6495d29a41ddd3854a260c600950a7b98e660dd740

Observation 5754e964-70a6-4abd-a996-3b33939c6594 · outbound

This paper cites Temporal Difference Learning for Model Predictive Control.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning Temporal Difference Learning for Model Predictive Control

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:33:50.866977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:fd756909baba1fdea83db00592faf16e4291b6bde56d7f0de1bd3d0de4039336

Observation b35811ad-c69e-4a7d-8415-fa3bcaf5986e · outbound

This paper cites TD-MPC2: Scalable, Robust World Models for Continuous Control.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning TD-MPC2: Scalable, Robust World Models for Continuous Control

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-14T17:27:36.461279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:b3c6a329cb50067db0811bfd47842a9eae36eb99a7953fa2f14323b3d434fea5

Observation c4ba759a-8199-4962-aa8e-2a77bd90c9d6 · outbound

This paper cites GAIA-1: A Generative World Model for Autonomous Driving.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning GAIA-1: A Generative World Model for Autonomous Driving

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:15:10.779550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:81e5e917268fcdda6297ec8d308b9838bf2dedc1e285e0746905e71da98876b3

Observation 05a596de-a1e0-4a57-8668-a476c90b7169 · outbound

This paper cites Huang, P., Liu, S., Liu, Z., Yan, Y ., Wang, S., Chen, Z., and Xiao, T.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning Huang, P., Liu, S., Liu, Z., Yan, Y ., Wang, S., Chen, Z., and Xiao, T

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T00:33:50.897096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:8b61ab8d044b41c089083c1a7c52d086028cbff8b034681e4648ccd2ac69e7d5

Observation a353802c-bc41-43c8-8ec1-ee2124de08e8 · outbound

This paper cites Perceiver IO: A General Architecture for Structured Inputs & Outputs.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning Perceiver IO: A General Architecture for Structured Inputs & Outputs

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T19:47:14.368858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:9f5830e1c03aa1409f5f58cd9103d54d6f2173216c97a82e4e752bdf960bd82b

Observation 512192e1-5c2f-421b-8717-7545025e78eb · outbound

This paper cites The Kinetics Human Action Video Dataset.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning The Kinetics Human Action Video Dataset

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:13:45.682066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:c05a5ae6923e253476bc172f105342de73a06c5e41e8ae3afc2e56e0b6db46ed

Observation d9daf3d7-f670-43fe-b71d-216708b2a2c3 · outbound

This paper cites DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:51:19.504346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:82721373e2a5b4d4ef6afe5db20d27f97f9ec004f2a8c8d7a83bc0929754a133

Observation dee87fdd-3f34-4d1e-919a-a6fce3010e77 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning OpenVLA: An Open-Source Vision-Language-Action Model

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-11T00:33:50.935341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:dfd06b948ab696bf1caecf99dec161d04359336dd28d9b15dbf860015b7ad088

Observation 95317c8e-06fc-4c1f-bef4-3d1bc77394f1 · outbound

This paper cites A path towards autonomous machine intelligence version 0.9.2, 2022-06-27.Open Review, 62(1):1–62.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning A path towards autonomous machine intelligence version 0.9.2, 2022-06-27.Open Review, 62(1):1–62

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T00:33:51.213849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:b420434c77f9e7877c45fea38ba44e43b267d121c4224eb71b9fb6b901d8bf01

Observation 1a49219d-35ae-4aec-a482-1252cd056abb · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning LLaVA-OneVision: Easy Visual Task Transfer

Reference 38

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T00:33:51.102812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:69ab5463469eb3baab6a8a3fae689b017d95c1e96bcb303ac8b0db9f575cad99

Observation 8359af7d-9ab3-4be8-9d3c-1fd8d95608af · outbound

This paper cites TempCompass: Do Video LLMs Really Understand Videos?.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning TempCompass: Do Video LLMs Really Understand Videos?

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:46:17.144047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:349fc1275fad5e351922260fc8cdd3500225799f02e02bedb012021edd357bfe

Observation 45cc6c29-ea7d-4df9-993c-173a43c8c513 · outbound

This paper cites Decoupled Weight Decay Regularization.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning Decoupled Weight Decay Regularization

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-11T00:33:51.112985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:facc8808324250b3ae02ee0ef1759a3d1b880941dc3bb500477ee3aa5e02474d

Observation 06953f0c-1aeb-41ad-a154-2c2e310e8590 · outbound

This paper cites Keypoints into the Future: Self-Supervised Correspondence in Model-Based Reinforcement Learning.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning Keypoints into the Future: Self-Supervised Correspondence in Model-Based Reinforcement Learning

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T00:33:51.120897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:c7bb41b0d2b9e773f4d2472f351eda7c79ebaf09ecd5e75f715b8a29243abb38

Observation 9d623b42-e6a5-42be-9ef5-29b84742618d · outbound

This paper cites Octo: An Open-Source Generalist Robot Policy.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning Octo: An Open-Source Generalist Robot Policy

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-05-11T00:33:51.126566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:b6ac28ba293f9a5e7e4dc93852e545d4fb775453881bf2358cd79304544ebb6b

Observation 57365b87-d7e6-4990-aca8-c82a929eeb0a · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning DINOv2: Learning Robust Visual Features without Supervision

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-05-11T00:33:51.136489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:f9e0769509d1c93ddd47204f50e8ac7e27cfec9e06a854a22a4319eb7a15f4d9

Observation 06ea8c53-e396-447e-a2e7-7f7852559e51 · outbound

This paper cites Qwen2.5-VL Technical Report.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning Qwen2.5-VL Technical Report

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-11T00:33:51.149057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:624d96815ecf76757a338f08b6dc15beb135c778caefc0278f0aff39613678d8

Observation 93043d74-8061-46d9-9ec8-2118051a49ef · outbound

This paper cites An Empirical Study of Autoregressive Pre-training from Videos.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning An Empirical Study of Autoregressive Pre-training from Videos

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:33:51.155917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:9a0669ebf6f20508c7ffb37249699be39769eda256ae838576553005cce55734

Observation d57d4997-31f2-4682-a732-3cee5148dacf · outbound

This paper cites GAIA-2: A Controllable Multi-View Generative World Model for Autonomous Driving.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning GAIA-2: A Controllable Multi-View Generative World Model for Autonomous Driving

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-15T13:48:22.447793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:35f6c5865e1bbb7d8cc8cd045525feaec2dfd57a35c7ce17d9e156e8bdb67b49

Observation dd1c37b8-704b-4224-888d-0992d7313db4 · outbound

This paper cites TOMATO: Assessing Visual Temporal Reasoning Capabilities in Multimodal Foundation Models.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning TOMATO: Assessing Visual Temporal Reasoning Capabilities in Multimodal Foundation Models

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:33:51.178602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:e5cdd292866cc5779c3268ea8127b3bde6dec325e9ae0250a6933a14df4d18af

Observation f501c319-0753-418f-8e25-5100c409fa1e · outbound

This paper cites Learning from reward-free offline data: A case for planning with latent dynamics models.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning Learning from reward-free offline data: A case for planning with latent dynamics models

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:33:51.190282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:02f1093d9683b1a36ea3e7fbe336c97866997183c6eafae23a4f4c0182c1a0e0

Observation ed8e984d-368c-4623-95b6-256643006d18 · outbound

This paper cites Video Occupancy Models.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning Video Occupancy Models

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T00:33:51.198027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:b2035a1d65f897bbda2eb62b8ab1dc1288b7ae58f5c30d7c1dc6ae7abf5b7c2e

Observation eec793d8-d0c9-4b3f-b870-7fcde67106e0 · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-05-11T00:33:51.202552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:9c75810a1b225d07d61e4b95757f63bf2c2ba99828e56922ce0afd609dd42f59

Observation 03b7c11b-16a0-4b97-8a9c-c4fa9e3062d9 · outbound

This paper cites The Evolution of Multimodal Model Architectures.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning The Evolution of Multimodal Model Architectures

Reference 51

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T00:33:50.954145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:e160f8be04cf22714f2b7ca139b74897852825ffe20489ddb295bba0e858f305

Observation efc3332f-1f50-491c-85bb-3c960ecc534a · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-05-11T00:33:50.959202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:89458b96fe22d9476383756abb61cf649b0689a715b0a3e642cc0bc6bc4cf84f

Observation 0ea17cc3-a4f7-4b31-8377-a22e27bd17ec · outbound

This paper cites Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-13T16:32:05.974462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:36fb475526ecadd738d0b5c916c67acf9ac64819826dca8cd3612594f4ff3008

Observation 506152f2-e699-4975-852f-f94d86af54a0 · outbound

This paper cites Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:33:50.994322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:b6882a7ca50e20eaf850424ba732afc0cb8e0070f7a17aeb99b57cdb8a7794d8

Observation 5f53b7b9-dd40-42ca-9226-7639e797b152 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-13T15:02:01.218694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:ce96c80525d177d379d31f1ac6a539053bef10d29ede79acb1dc3648a17cfcf2

Observation 15bd45bc-a1b8-4c77-8dd8-83626e44a128 · outbound

This paper cites LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:19:22.513527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:e743e94f9f5e84510193d65402bd602ce6e7b621b02b66ee270013ef8971b586

Observation f937889c-ac7f-4d16-b967-bec5f6d8c96a · outbound

This paper cites CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:21:45.399848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:faba95e8a364e7aea734ef2c39d4bad3ed9e71675d2d3147cccdf58399daadbf

Observation 5ad1e2d9-6c37-401d-bb2c-b32321d5c289 · outbound

This paper cites FLARE: Robot Learning with Implicit World Modeling.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning FLARE: Robot Learning with Implicit World Modeling

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-17T15:59:09.111136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:4cb12c923a73cbffd46d707c963c075031c4da435e57ae72a3e670ec1d16cfc1

Observation 339aa09a-d664-4b98-98dd-6f5afc436f7f · outbound

This paper cites DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-17T16:06:10.055251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:0dc1e21c62427a82a0e42f60fd1bde7e9d2a4d7e4c9e7c8ccfad16d8cdab6e04

Observation 93e6b606-7cf2-4cff-936d-0abc15437426 · outbound

This paper cites Unified World Models: Coupling Video and Action Diffusion for Pretraining on Large Robotic Datasets.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning Unified World Models: Coupling Video and Action Diffusion for Pretraining on Large Robotic Datasets

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-13T16:25:00.606134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:aa53832de460522b1b4d266d14c8d773074db2644067eab083a8862883eac395

Observation a1188a44-03fc-4259-831f-f1f75f816862 · outbound

This paper cites abbreviated.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning abbreviated

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T00:33:51.234814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:8ff25124526bb578659f9291a1528a8230315aa9baa89c44745bce0c066c9397

Observation 9652bef8-da87-4409-b4ff-cbec3f69077c · outbound

This paper cites (2020), using the standard16 × 16 patch size.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning (2020), using the standard16 × 16 patch size

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T00:33:51.250197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:8f4d6ecd036360d5f1260028e35b240bb29b89aab710cb88a5276e73e7d47240

Observation 1df189b3-4176-4803-8df5-004026c7a7ba · outbound

This paper cites For the pick-and-place tasks we present two sub-goal images to the model in addition to the final goal.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning For the pick-and-place tasks we present two sub-goal images to the model in addition to the final goal

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T00:33:51.255541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:6af9b8307ef065ffd715efe26856a64e8c8437408e94b537d5d4a7b4cba71aa1

Observation 3cca2a2f-f3c5-47cb-bfcf-cb6bbd4d06a8 · outbound

This paper cites The MLLM ingests the output embeddings of the vision encoder, which are projected to the hidden dimension of the LLM backbone using aprojector module.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning The MLLM ingests the output embeddings of the vision encoder, which are projected to the hidden dimension of the LLM backbone using aprojector module

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T00:33:51.261444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:78fcc749b7c540983fb871a4b0ae41d568f29470c70ba277fd54508dc28582da

Observation 763b4da0-03cc-4965-9bad-c241fb7b9ac3 · outbound

This paper cites We describe the training details in the following sections.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning We describe the training details in the following sections

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T00:33:51.217765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:e03387e122fbaef7b2e5caf70f9db369bd3d1328e76f18096c5d66125c686f0b

Observation 98a591e2-65dc-4d12-aa70-daabf6b9780d · outbound

This paper cites To assess the ability of V-JEPA 2 to capture spatiotemporal details for VidQA, we compare to leading off-shelf image encoders.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning To assess the ability of V-JEPA 2 to capture spatiotemporal details for VidQA, we compare to leading off-shelf image encoders

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T00:33:51.221451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:63a6eb0ce97feb1b43902d89678d05ec2a5a363c5d86f5e63bca42fa0377776b

Observation 689f5c30-4102-4141-a14d-e6b9b166681a · outbound

This paper cites Unlike Cho et al.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning Unlike Cho et al

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T00:33:51.224273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:c2502161774d7d33c42ea4d78c03608ffd82c82b46d752b3c6da7a527c992a24

Observation 8c6206a5-0347-4ae4-a5f0-65cce4f07999 · outbound

This paper cites We scale up the data size to 88.5 million samples.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning We scale up the data size to 88.5 million samples

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T00:33:51.227970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:f90579b0ff66a29914817762809a38b5e396bcac734e2ce17cc9aa4dce61a958

Observation 5766f544-a493-431f-bb50-22e375a8bae9 · outbound

This paper cites an unresolved cited work.

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning Unresolved cited work

Reference 69

Resolution
unresolved
raw_fallback, observed 2026-05-11T00:33:51.230791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T00:33:50.471804Z digest=sha256:4bad41871c1c27fec25ccb2ecbbf174c82cf61258f7ae339eea87e604773618a

Pith citing papers

Observation 21a9cdcc-18e3-4368-8bc6-7f934b17a38c · inbound

3D Foundation Model for Generalizable Disease Detection in Head Computed Tomography cites this paper.

3D Foundation Model for Generalizable Disease Detection in Head Computed Tomography V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-23T03:27:27.024372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T03:26:46.665351Z digest=sha256:c64f4c99870ad265496663248e0a289ac309351324ff6127a06bcc3bc35eea68

Observation 3f2e97c4-f009-4f96-bf8a-1d1219edbc22 · inbound

A Survey on Vision-Language-Action Models: An Action Tokenization Perspective cites this paper.

A Survey on Vision-Language-Action Models: An Action Tokenization Perspective V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 298

Resolution
verified exact
local_arxiv, observed 2026-05-17T14:08:35.462325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T14:08:34.893876Z digest=sha256:323f96ceec19d9d4e629181e21e00175772364274d29e1dab47fe9e2f75bca4c

Observation 61102fb8-a0dc-4cd7-b391-66e53c861e88 · inbound

Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey cites this paper.

Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 201

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:28:16.240454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T20:28:15.818016Z digest=sha256:5cbaf706da1c7441712981802b0005bfd6556fbb8ad2f4432e65649f362940e3

Observation d30e72f5-c6d1-4ed0-9cef-06c758ff7752 · inbound

3D and 4D World Modeling: A Survey cites this paper.

3D and 4D World Modeling: A Survey V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T06:04:03.420693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T06:04:03.420693Z digest=sha256:0e1b0e915c60b1c3524e6e566681acbd65825af3462d3afad6e477b749929b37

Observation d4b14d17-b51d-4562-b0c4-b0b79f98482b · inbound

A Survey of Reinforcement Learning for Large Reasoning Models cites this paper.

A Survey of Reinforcement Learning for Large Reasoning Models V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-18T00:02:24.869240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T00:02:24.352947Z digest=sha256:aeb5a334dfcb3b0e9b76cb97951ba377a214eca01b6e8f777a01d5ba5e19efba

Observation 96208e04-2720-4f38-9ca4-5006ef5ab42c · inbound

Video models are zero-shot learners and reasoners cites this paper.

Video models are zero-shot learners and reasoners V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-05-14T02:16:45.787940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:8fcf0527deafc06ed831d0955c4f3e0bf851a4a5745ca3e98613d747db188f8a

Observation 7d0e224a-43b3-4916-99fb-1379ff8754b7 · inbound

World-Env: Leveraging World Model as a Virtual Environment for VLA Post-Training cites this paper.

World-Env: Leveraging World Model as a Virtual Environment for VLA Post-Training V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-18T12:51:23.537135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T12:48:32.123998Z digest=sha256:a05013adeec24e3e0a3f44bafb08c17211be6a8a90bbe7d79685935982f4b146

Observation c0a10ac1-b6f2-4aeb-9670-ff93d76bbf30 · inbound

Inferring Dynamic Physical Properties from Video Foundation Models cites this paper.

Inferring Dynamic Physical Properties from Video Foundation Models V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-18T10:11:14.124897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T10:08:11.191706Z digest=sha256:d2818e6494d4d14f119c2137591e2038b95b1613dfd835ed2f36f05e910940af

Observation 0ce7c558-6d52-4d25-ad06-07093d0bdb01 · inbound

AtomWorld: A Benchmark for Evaluating Spatial Reasoning in Large Language Models on Crystalline Materials cites this paper.

AtomWorld: A Benchmark for Evaluating Spatial Reasoning in Large Language Models on Crystalline Materials V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T11:26:04.215319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:26:04.215319Z digest=sha256:5cdfbb025ba47eb211e766708677e28058415ba3257860d90ed551a341588224

Observation f2381dac-5508-4b30-95f3-13fea060946d · inbound

Hybrid Architectures for Language Models: Systematic Analysis and Design Insights cites this paper.

Hybrid Architectures for Language Models: Systematic Analysis and Design Insights V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-18T10:21:15.006979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T10:18:04.431436Z digest=sha256:16b1a67d60b6cfaac7683625ecb162b2d28388a9df24f49868552b067cd2904b

Observation 38d2428e-0744-43b7-88d2-6041cc7a3d29 · inbound

A Comprehensive Survey on World Models for Embodied AI cites this paper.

A Comprehensive Survey on World Models for Embodied AI V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T09:12:30.341116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:12:30.341116Z digest=sha256:bd1056483a35c701b06b4d3ae7b2d97bef9edd60c1df8ed1635eb4b74a56cfdf

Observation ab3cb743-68a0-4ace-b948-d3a1d9d89180 · inbound

World Simulation with Video Foundation Models for Physical AI cites this paper.

World Simulation with Video Foundation Models for Physical AI V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-12T23:01:13.764797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:c6f4f55e9763d982ee21b6a2c4204bbd448131db71783d4db06ae11a8de6f3c1

Observation 4e4b2129-5193-46d3-a387-fc376242db98 · inbound

Cambrian-S: Towards Spatial Supersensing in Video cites this paper.

Cambrian-S: Towards Spatial Supersensing in Video V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-18T03:46:04.539471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T03:46:04.363500Z digest=sha256:0d7264086b26948a34b9fa327506fdba6d78dacb33f7ec18aa60bc2f24cedcad

Observation d7ad238f-d48f-47bf-b56d-6e0de984978b · inbound

IPR-1: Interactive Physical Reasoner cites this paper.

IPR-1: Interactive Physical Reasoner V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:55:14.957837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T20:55:10.819298Z digest=sha256:d7cf6d1a2260dd5f71ad16412a5416c28b3232e9b0b4effc08a859cb8bb8f068

Observation dda7fd4f-7d7d-4d7e-a0bc-1b85c5188586 · inbound

IPR-1: Interactive Physical Reasoner cites this paper.

IPR-1: Interactive Physical Reasoner V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T21:29:31.124396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:29:31.124396Z digest=sha256:e21f2e8f1a897669ccf5dba762e9e9d44ccea405952e929c5dd6bbde1462c6bb

Observation 9d703c97-a3c9-4ed9-8e12-712545cabcff · inbound

POMA-3D: The Point Map Way to 3D Scene Understanding cites this paper.

POMA-3D: The Point Map Way to 3D Scene Understanding V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:30:11.562918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T20:27:27.347592Z digest=sha256:cc67d53bf0947b23ac30b33cc8bc3d3f4066c6e0d5af98e6b867baadb89b0a65

Observation 7a26b75e-ccdb-4541-94ad-957a160d7f01 · inbound

Rethinking Reward Signals in Video GRPO: When Scores Become Targets cites this paper.

Rethinking Reward Signals in Video GRPO: When Scores Become Targets V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T20:36:05.488757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:36:05.488757Z digest=sha256:2844d2b122e97733d8104d69d41e4f71ae69ee02a53236261de5e5d06c7e7e8b

Observation c58b4f50-b54f-4fa5-b31c-2ea5c5e1e5d0 · inbound

MIND-V: Hierarchical World Model for Long-Horizon Robotic Manipulation with RL-based Physical Alignment cites this paper.

MIND-V: Hierarchical World Model for Long-Horizon Robotic Manipulation with RL-based Physical Alignment V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T18:09:47.097419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:09:47.097419Z digest=sha256:3397b8d2ec1dd85e59b2c2d71936112c82c6b673cb7dbf1d809128187abde113

Observation adac12f7-08b5-4027-be17-8ce31abf9a29 · inbound

CLARITY: Medical World Model for Guiding Treatment Decisions by Modeling Context-Aware Disease Trajectories in Latent Space cites this paper.

CLARITY: Medical World Model for Guiding Treatment Decisions by Modeling Context-Aware Disease Trajectories in Latent Space V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T17:52:06.917547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:52:06.917547Z digest=sha256:2dc3a6ed12aa87b58e73c9f41147c63dc37e2adc696145789133e0bbf6f87141

Observation aef58132-70af-47e3-981d-a8f1dd0a24c8 · inbound

Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation cites this paper.

Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:18.919548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:18.919548Z digest=sha256:642ef8a4304eb03b68fc0b93114c8c395012e4394e7c9f400266d3b5da62ad6f

Observation 31f0bd24-cdaf-4924-a56e-f53b0ba3020d · inbound

A Geometric Theory of Cognition for Machine Intelligence cites this paper.

A Geometric Theory of Cognition for Machine Intelligence V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T16:46:58.559497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:46:58.559497Z digest=sha256:5bf69be7a8b1e84f0f59d57ac3b4d3c46e877e9817a58ade3dff85f3bf455853

Observation af32d96c-5b12-46c1-b227-b72cec58107a · inbound

Recurrent Video Masked Autoencoders cites this paper.

Recurrent Video Masked Autoencoders V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-16T21:58:35.830030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T21:55:42.555679Z digest=sha256:6a8813f6aa7a41e74653f14ba6a251c1b78031f10bed62122e9a4a7bc5763a5b

Observation 3f6f8a27-8e83-43a3-9a5b-74f822c77d77 · inbound

mimic-video: Video-Action Models for Generalizable Robot Control Beyond VLAs cites this paper.

mimic-video: Video-Action Models for Generalizable Robot Control Beyond VLAs V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-15T10:41:00.299216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T10:41:00.142543Z digest=sha256:32acac6c1759f5354983438e4e23e4dfc4f3a7a56e3040777c52784693f7bea1

Observation ae34eebe-d67b-4c98-ad3d-87c153291f73 · inbound

Probing and Leveraging Video Diffusion Transformer Features for Robust Point Tracking cites this paper.

Probing and Leveraging Video Diffusion Transformer Features for Robust Point Tracking V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T14:24:49.389952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:24:49.389952Z digest=sha256:dfa57531b800a00b98361875887ad475d80aba7150eef81965e2153b040d0694

Observation f9499361-b3e8-4730-a19b-f05269409e1f · inbound

DriveLaW:Unifying Planning and Video Generation in a Latent Driving World cites this paper.

DriveLaW:Unifying Planning and Video Generation in a Latent Driving World V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-16T19:38:21.072718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T19:34:39.518649Z digest=sha256:5987587f44d734835e726725bd22a94673e044886345afcb00a6c1e02888979d

Observation 0fd2bcdf-49c0-47e7-b972-a559bf8f87f7 · inbound

Learning to Feel the Future: DreamTacVLA for Contact-Rich Manipulation cites this paper.

Learning to Feel the Future: DreamTacVLA for Contact-Rich Manipulation V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-16T18:41:10.949440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T18:39:59.449746Z digest=sha256:23686132caeae7192392fb039fa34096c74d0e573a98d666b74ea69afb58446f

Observation ad22814e-df61-40a5-9c25-ad49bd2911d4 · inbound

Learning to Feel the Future: DreamTacVLA for Contact-Rich Manipulation cites this paper.

Learning to Feel the Future: DreamTacVLA for Contact-Rich Manipulation V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T13:35:54.025798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:35:54.025798Z digest=sha256:aaaa145255600c49a58fcf42b06859c9fc6abcb89c07ed75ac4f5a909691a5f9

Observation baa17504-3b25-4f8f-9df5-f436e1ccdb58 · inbound

Flow Equivariant World Models: Memory for Partially Observed Dynamic Environments cites this paper.

Flow Equivariant World Models: Memory for Partially Observed Dynamic Environments V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T13:00:51.146564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T13:00:51.146564Z digest=sha256:f311482ee666bd253b63edf6d932459522fccf3b738f0e433c19902ae1ca0267

Observation 8bc65591-6b88-4c89-bb64-14a49b14b01e · inbound

CycleVLA: Proactive Self-Correcting Vision-Language-Action Models via Subtask Backtracking and Minimum Bayes Risk Decoding cites this paper.

CycleVLA: Proactive Self-Correcting Vision-Language-Action Models via Subtask Backtracking and Minimum Bayes Risk Decoding V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-03T12:37:27.943411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:37:27.943411Z digest=sha256:ea525b671a145e981923cc76cb04039bb9ba2cac4f9f52a5e75ec7eba0f25cb9

Observation 59019341-dc3a-42cf-9eb0-5eb63e765887 · inbound

Advancing Open-source World Models cites this paper.

Advancing Open-source World Models V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-16T09:07:01.064580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T09:07:00.904794Z digest=sha256:75dfd5b6c5bd5a8f7a0d233c42cced7b2a0ab43ed406d8f66f7f97db188f11a1

Observation ccbc70af-4af1-4c23-95a0-32fde3d2353b · inbound

Drive-JEPA: Video JEPA Meets Multimodal Trajectory Distillation for End-to-End Driving cites this paper.

Drive-JEPA: Video JEPA Meets Multimodal Trajectory Distillation for End-to-End Driving V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T06:48:13.911798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:48:13.911798Z digest=sha256:fd3e3b83382bf35bd606249b5fa148196217ec18de1728081c3ac0a6d3ff5e14

Observation 10d87c3c-b641-4255-8841-35c48e2c3b4b · inbound

PEPR: Privileged Event-based Predictive Regularization for Domain Generalization cites this paper.

PEPR: Privileged Event-based Predictive Regularization for Domain Generalization V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-16T07:40:44.111305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T07:37:38.263234Z digest=sha256:e40e3d786752bc8d160a60dbabd0bf7bcdbb630b56c9c28ef086eee2142b1d85

Observation b5c27669-8051-4f81-9fca-a33b4e8e9d6e · inbound

SurgMotion: A Video-Native Foundation Model for Universal Understanding of Surgical Videos cites this paper.

SurgMotion: A Video-Native Foundation Model for Universal Understanding of Surgical Videos V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-16T07:17:30.309716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T07:16:29.588452Z digest=sha256:2c9c894fd7faee0254205bb8d412683e697c41862ed5b0fd9a27fa123b31ad7b

Observation d47e3b7b-3f15-4065-a7fd-c0029e97b160 · inbound

DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos cites this paper.

DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-16T17:02:34.074516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T17:02:33.997887Z digest=sha256:db671867843136cd1c6526c78ee9f735e848b1b02203733a9086d1ee59a6a865

Observation 53dafe14-8e19-4c33-bf29-1a261ef3d584 · inbound

OmniFysics: Towards Physical Intelligence Evolution via Omni-Modal Signal Processing and Network Optimization cites this paper.

OmniFysics: Towards Physical Intelligence Evolution via Omni-Modal Signal Processing and Network Optimization V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-16T07:10:43.176327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T07:09:46.254851Z digest=sha256:2f677144f46d347ba95ee7a92fe472c6f12cdbeb0e6f51db8dadd3c68d144cf6

Observation 97ca2aad-0555-4096-b31d-63295199f0b3 · inbound

Going with the Flow: Koopman Behavioral Models as Pseudo Planners for Visuo-Motor Dexterity cites this paper.

Going with the Flow: Koopman Behavioral Models as Pseudo Planners for Visuo-Motor Dexterity V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T03:41:53.258454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:41:53.258454Z digest=sha256:074db4ead680acc26fa53f6554862f857885e93602f4bbd59ed9f92382d87418

Observation 8574663c-49b5-495e-ad2d-e4e5c55c2037 · inbound

Rolling Sink: Bridging Limited-Horizon Training and Open-Ended Testing in Autoregressive Video Diffusion cites this paper.

Rolling Sink: Bridging Limited-Horizon Training and Open-Ended Testing in Autoregressive Video Diffusion V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T07:07:30.029514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T07:02:38.876518Z digest=sha256:155e363eff7b60cf65fb75dc92b4368b4457cc16029c4b81aa1f20af80e48ee1

Observation 612e5dc3-2012-409f-9649-c32baa742f19 · inbound

Olaf-World: Orienting Latent Actions for Video World Modeling cites this paper.

Olaf-World: Orienting Latent Actions for Video World Modeling V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-03T01:20:04.620117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:20:04.620117Z digest=sha256:97dbc49e5eccda46cff3019d6f5f95f9005ff94f8fc3210a69618004220f0fad

Observation 3cde03a6-daf4-4890-863d-8934ac6bd406 · inbound

RISE: Self-Improving Robot Policy with Compositional World Model cites this paper.

RISE: Self-Improving Robot Policy with Compositional World Model V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-16T02:30:32.159613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T02:28:37.997148Z digest=sha256:46293a8172627db9e3cacbfef09900e229f4e3739c0077f39093f685b7d6841b

Observation edea824e-f551-4587-8b00-c2326bdfcfe1 · inbound

World Action Models are Zero-shot Policies cites this paper.

World Action Models are Zero-shot Policies V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-11T16:18:15.244552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T16:18:15.003371Z digest=sha256:3306f871b11d8e9b53c871e834964e0f6a6d414a97da38293d7feef8d1ee8aa9

Observation 8ca2d4af-bf68-4a7a-b8bf-f51fae9b828e · inbound

Xray-Visual Models: Scaling Vision models on Industry Scale Data cites this paper.

Xray-Visual Models: Scaling Vision models on Industry Scale Data V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-02T22:25:57.827119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:25:57.827119Z digest=sha256:37c2529ff639f702a82a20f0fad8f785eaa9ccebac1cb499e12cffb9f5b57baa

Observation 8c0c1ec0-bcc9-43cf-814f-8a6d3b71c92b · inbound

Space-Time Forecasting of Dynamic Scenes with Motion-aware Gaussian Grouping cites this paper.

Space-Time Forecasting of Dynamic Scenes with Motion-aware Gaussian Grouping V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-15T19:56:33.994859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T19:51:58.756605Z digest=sha256:fb2bbd5dfbf4f81b9869cae91bf38eacc08c7966cd616b5b0607bed4beee7847

Observation 6276ba32-2cb3-4495-86f7-5f86d7344504 · inbound

TrajTok: Learning Trajectory Tokens enables better Video Understanding cites this paper.

TrajTok: Learning Trajectory Tokens enables better Video Understanding V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-15T19:16:31.936451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T19:11:52.694778Z digest=sha256:4eadef44c3606c109b3f04830e4c3a52285b3728c2194cc7314eae202517d0a6

Observation abbaeea2-1fbe-4f5e-a560-ab79d31205d2 · inbound

TrajTok: Learning Trajectory Tokens enables better Video Understanding cites this paper.

TrajTok: Learning Trajectory Tokens enables better Video Understanding V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T20:38:35.343597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:38:35.343597Z digest=sha256:7d1d5e6495a677dbfb214bfc5f6fd3aaa5fc8e0b97a263cf176d79751a8e9e29

Observation ed6458fc-c39e-49d3-a453-7e7473b89bd0 · inbound

GeoWorld: Geometric World Models cites this paper.

GeoWorld: Geometric World Models V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-21T11:40:03.357822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T11:39:15.308355Z digest=sha256:aadbd25a0f5718f4596a541218e382a14c627fc9c5d28cfb04953a070279a020

Observation e85ab659-0887-475e-a06d-6ba5ae6d989d · inbound

EgoMoD: Predicting Global Maps of Dynamics from Local Egocentric Observations cites this paper.

EgoMoD: Predicting Global Maps of Dynamics from Local Egocentric Observations V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T20:38:16.567411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:38:16.567411Z digest=sha256:928134b4ffb6fda02346886553d135d8cab574964ddda58808c8874e0ea30fbc

Observation 1bd60c73-fe88-451c-88f1-86c1b145c81f · inbound

ReMoT: Reinforcement Learning with Motion Contrast Triplets cites this paper.

ReMoT: Reinforcement Learning with Motion Contrast Triplets V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T19:58:12.904007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:58:12.904007Z digest=sha256:3f50b743446c2ec876529864c00831356c2833a1782013fe73bc78944ff05bfd

Observation 89965aa2-5d0d-493d-8d98-55fb15729978 · inbound

RoboLight: A Dataset with Linearly Composable Illumination for Robotic Manipulation cites this paper.

RoboLight: A Dataset with Linearly Composable Illumination for Robotic Manipulation V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T18:56:44.652978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:56:44.652978Z digest=sha256:523b041d255b2b6cad525e41d1af8fe32bbf9468a306bbba4c7d2aa10ea3aa08

Observation 06261469-2999-4665-af3d-125c049f3567 · inbound

Dreamer-CDP: Improving Reconstruction-free World Models Via Continuous Deterministic Representation Prediction cites this paper.

Dreamer-CDP: Improving Reconstruction-free World Models Via Continuous Deterministic Representation Prediction V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-15T14:50:05.131996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T14:46:18.645009Z digest=sha256:b6b0bdcd04b3af8d0c86647cb1bf2bdc5bb2f0e5b28bdb305bea26bc50fdd4f5

Observation 1d328c51-3b8a-4d96-bf3b-8ec7053f88a0 · inbound

PlayWorld: Learning Robot World Models from Autonomous Play cites this paper.

PlayWorld: Learning Robot World Models from Autonomous Play V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-05-15T14:05:54.592109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T14:05:32.112432Z digest=sha256:197e476670efb92b5f8bbf88e9b5eda673c5af39dcf2a9b07eaf20d330af1e76

Observation d247fa89-3dd1-4fcb-8121-5ef947af325a · inbound

Self-Distillation of Hidden Layers for Self-Supervised Representation Learning cites this paper.

Self-Distillation of Hidden Layers for Self-Supervised Representation Learning V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T18:11:29.452431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T18:11:29.452431Z digest=sha256:f01f564ed5b5a759bb9ed686b6a25eae0b7cd12a679a74c576c431461eb79891

Observation 16560a48-022d-4352-950b-e8246d764ca0 · inbound

Generation Models Know Space: Unleashing Implicit 3D Priors for Scene Understanding cites this paper.

Generation Models Know Space: Unleashing Implicit 3D Priors for Scene Understanding V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T17:54:04.043468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:54:04.043468Z digest=sha256:73b85cb41d06e0b219768780a8863271927fb97d80a75a2e9708bf48d4b9e0c3

Observation 23af3160-2c36-400f-84c3-89ff5668f910 · inbound

LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixels cites this paper.

LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixels V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-15T04:09:22.421581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T04:09:22.328844Z digest=sha256:fb6fa497546a79d954157fb3ddf90084ce5835a2332ec788c158eec7c4823e22

Observation ceb99782-4f30-4b07-9610-178fd1f13b3b · inbound

ThinkJEPA: Empowering Latent World Models with Large Vision-Language Reasoning Model cites this paper.

ThinkJEPA: Empowering Latent World Models with Large Vision-Language Reasoning Model V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-13T20:17:45.587021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:17:45.587021Z digest=sha256:3f9d284c49543178b191486d1af8b7faa961a4ac6273cdabd76555fdd3bb0304

Observation 7f95f2fe-9259-479e-b25b-5e566ce335f8 · inbound

Factorization Regret mediates compositional generalization in latent space cites this paper.

Factorization Regret mediates compositional generalization in latent space V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-14T23:18:15.909718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T23:15:08.716714Z digest=sha256:be66830b14f7f14a8454b268b363ff7a8da4b35fb2be400e6b91251a1d0742fe

Observation 100bf105-1541-4c8b-8aa8-2d51137a3a71 · inbound

Factorization Regret mediates compositional generalization in latent space cites this paper.

Factorization Regret mediates compositional generalization in latent space V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T05:46:17.900933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:46:17.900933Z digest=sha256:1caf1bea85930921968280d7d16bf0863a6334fb45fde794b202c04582397854

Observation c3bbbd9e-94f2-4f4b-8571-300692b94b8b · inbound

Topological sum rule for geometric phases of quantum gates cites this paper.

Topological sum rule for geometric phases of quantum gates V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-13T15:33:59.448352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T15:33:59.448352Z digest=sha256:3cd2c3eafa46909fa68c837c59e4dd692780d848f4ada1faf0df29e2b50f081c

Observation 80ecc837-2a2f-4b48-be64-ab3373eb5029 · inbound

World Action Verifier: Self-Improving World Models via Forward-Inverse Asymmetry cites this paper.

World Action Verifier: Self-Improving World Models via Forward-Inverse Asymmetry V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-13T14:05:26.303000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T14:05:26.303000Z digest=sha256:db0deae93a198b211913d383b757dafd62d7cd1bede7f21bf9717fab2285acc3

Observation d094bea9-e17b-486c-9f17-a1f06f2bd0b3 · inbound

Hierarchical Planning with Latent World Models cites this paper.

Hierarchical Planning with Latent World Models V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-13T20:18:13.819302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T20:13:58.991298Z digest=sha256:9355abd1cef9f1e6c9cf18d9b71b8f372b0f7e5ba57e03c029815d63fda04085

Observation cf57247f-acab-491e-bf07-88a05c31f1ce · inbound

Emergent Compositional Communication for Latent World Properties cites this paper.

Emergent Compositional Communication for Latent World Properties V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-15T08:29:52.183924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T08:28:10.971822Z digest=sha256:b3ff7bd60efb230e9d2d9fc82a9ae15f51d0b298ff7d88b9bee402eb5ce868d0

Observation b0e9a11d-3fbc-4e1a-a690-7786c981094e · inbound

Veo-Act: How Far Can Frontier Video Models Advance Generalizable Robot Manipulation? cites this paper.

Veo-Act: How Far Can Frontier Video Models Advance Generalizable Robot Manipulation? V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:33:51.269344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T20:10:54.362107Z digest=sha256:6a9887eb53dc0b030712dd946b63943d0d00dccbb60d6641590f384e88302961

Observation 7323336e-7f29-4016-b2f9-33980d4733dd · inbound

OpenWorldLib: A Unified Codebase and Definition of Advanced World Models cites this paper.

OpenWorldLib: A Unified Codebase and Definition of Advanced World Models V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:33:51.269344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:36:42.100191Z digest=sha256:63f7f7dc13c5378973145dc1a0f2b48b21b0453a5b111447cd04790696772d5f

Observation ef4339cd-0695-4a61-8b2a-e82da2a66563 · inbound

OpenWorldLib: A Unified Codebase and Definition of Advanced World Models cites this paper.

OpenWorldLib: A Unified Codebase and Definition of Advanced World Models V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-13T09:42:23.808691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T09:42:23.808691Z digest=sha256:3affb11608b9220c754f9d78ba68e69841791209e4f6bd779058f34b4ceb14b7

Observation 6169bb95-c826-4c1f-af34-866a793f05be · inbound

StarVLA: A Lego-like Codebase for Vision-Language-Action Model Developing cites this paper.

StarVLA: A Lego-like Codebase for Vision-Language-Action Model Developing V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:33:51.269344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:01:02.641396Z digest=sha256:95c7a5381d8e3e91833735028bee69205fadb8a0cd485787f3411419d36019c3

Observation 8d1c0a46-fbe9-4a36-9baf-565ffe1c7911 · inbound

Action Images: End-to-End Policy Learning via Multiview Video Generation cites this paper.

Action Images: End-to-End Policy Learning via Multiview Video Generation V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T00:33:51.269344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:51:05.206602Z digest=sha256:98cf816ee165babce4dc5fc67c2a766a11041d9ad7e7d8088cae9822fd60e3bf

Observation 2e3c5bc1-0b32-489a-bc99-9233bafaa622 · inbound

The Cartesian Cut in Agentic AI cites this paper.

The Cartesian Cut in Agentic AI V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-11T05:51:10.828754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T17:54:45.333038Z digest=sha256:395b90934974cb8fd3e7e252afad1068ffc7232ea6e81a0e8cbd0d7a393d2677

Observation c4dbef46-0207-4e20-b674-434e937dacc5 · inbound

A Machine Learning Framework for Turbofan Health Estimation via Inverse Problem Formulation cites this paper.

A Machine Learning Framework for Turbofan Health Estimation via Inverse Problem Formulation V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:33:51.269344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:26:48.918721Z digest=sha256:36134748548625f4a906cdb3c9d26f8a6994b0ee387d13d8cbc353bbab4be06b

Observation ff9a67a8-604d-4515-ae69-06bea3cbb133 · inbound

Phantom: Physics-Infused Video Generation via Joint Modeling of Visual and Latent Physical Dynamics cites this paper.

Phantom: Physics-Infused Video Generation via Joint Modeling of Visual and Latent Physical Dynamics V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-11T05:10:55.621528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:15:37.338442Z digest=sha256:df1d22bd4f0b2bafeb47abd9fb9cfdd054415b1bb1f284b527804581398aa45c

Observation f334355b-f8ce-4e64-914c-c9c4e2cb2118 · inbound

Phantom: Physics-Infused Video Generation via Joint Modeling of Visual and Latent Physical Dynamics cites this paper.

Phantom: Physics-Infused Video Generation via Joint Modeling of Visual and Latent Physical Dynamics V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-21T09:19:56.619333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T09:17:53.731701Z digest=sha256:079b3fc3e5096fe673524683bfb5c62a12461323df3e8022453a9e034aed6368

Observation f6276d56-bbf8-4642-bfac-e566e0fd01c8 · inbound

Matrix-Game 3.0: Real-Time and Streaming Interactive World Model with Long-Horizon Memory cites this paper.

Matrix-Game 3.0: Real-Time and Streaming Interactive World Model with Long-Horizon Memory V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-11T00:41:04.710352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:24:14.349469Z digest=sha256:8d67bb1f60d85147adc47eaeed05275b78cac1f5aaf11d9ccebe2de9a3c1940c

Observation 9b339fed-d2bf-4fa8-b583-4f83f8c6f0b8 · inbound

Zero-shot World Models Are Developmentally Efficient Learners cites this paper.

Zero-shot World Models Are Developmentally Efficient Learners V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-11T10:16:08.390658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:33:39.342672Z digest=sha256:09f1fb2f04dd81bf172190b8c06263ef4e82869dd0a852f0f22190dfb35a9be7

Observation 9b94b971-c780-4c1d-a3f7-406da2b4b39a · inbound

GTASA: Ground Truth Annotations for Spatiotemporal Analysis, Evaluation and Training of Video Models cites this paper.

GTASA: Ground Truth Annotations for Spatiotemporal Analysis, Evaluation and Training of Video Models V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T08:26:01.705672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T16:38:10.956726Z digest=sha256:5a3420698efbf07a7d54cfc7a0a56e77d2a2029c1b664d16a0a45ed65421bf77

Observation 50996131-99c6-4bb5-8670-9375b8c28b1c · inbound

Observe Less, Understand More: Cost-aware Cross-scale Observation for Remote Sensing Understanding cites this paper.

Observe Less, Understand More: Cost-aware Cross-scale Observation for Remote Sensing Understanding V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-11T10:51:02.844753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:18:11.837448Z digest=sha256:ab7a6b0a9b098758395cf1ae33c3f0d4dfe455a6b6e1976e8068d664da3611cb

Observation d5400338-cc49-433f-bb25-0d875da0e26f · inbound

Representations Before Pixels: Semantics-Guided Hierarchical Video Prediction cites this paper.

Representations Before Pixels: Semantics-Guided Hierarchical Video Prediction V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T08:45:58.539499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T16:30:53.578491Z digest=sha256:2ff2cc8e3a5932b05831162a61f87dbcee196205becd3e5a3461511700780233

Observation e0490e0e-8127-4f85-9983-874e5fea1123 · inbound

Grounded World Model for Semantically Generalizable Planning cites this paper.

Grounded World Model for Semantically Generalizable Planning V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:33:51.269344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:05:29.465402Z digest=sha256:f2960afd5132f40edb14471939959e990fcd159d4bdcd302cc0d2c8b06de52e2

Observation 57faeb5d-c3c7-4e0a-a986-b4f43b48116d · inbound

Learning Versatile Humanoid Manipulation with Touch Dreaming cites this paper.

Learning Versatile Humanoid Manipulation with Touch Dreaming V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-11T11:31:03.269322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T14:48:33.942119Z digest=sha256:fdd37e68a9a5311290e059df98dcc87f40fd56f74ab7ba3503785793e99ba2ca

Observation bc09f644-da8c-42d7-86ed-5f21a588671c · inbound

NTIRE 2026 Challenge on Video Saliency Prediction: Methods and Results cites this paper.

NTIRE 2026 Challenge on Video Saliency Prediction: Methods and Results V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:33:51.269344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T12:23:07.074182Z digest=sha256:a1d07f03879008c7370a2a1f8b6afa13087c88967ff744e21df78b01cb78773e

Observation c071330a-ffda-485a-9964-3a03dad4a69f · inbound

AnimationBench: Are Video Models Good at Character-Centric Animation? cites this paper.

AnimationBench: Are Video Models Good at Character-Centric Animation? V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:33:51.269344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T11:47:27.131584Z digest=sha256:98c964fb20d422b5abdb063baa753127af2d69c9574f64c90893c36b43fcc2b3

Observation 4a1bb3cd-6d4e-48ba-813c-d47acc0a6cd3 · inbound

Stylistic-STORM (ST-STORM) : Perceiving the Semantic Nature of Appearance cites this paper.

Stylistic-STORM (ST-STORM) : Perceiving the Semantic Nature of Appearance V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:33:51.269344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T08:59:40.341143Z digest=sha256:ae8e6b815a48e687806e06d769e3daa4ec02c45991f8ac968cec8584d9204056

Observation aa9b2aa5-4bde-4283-888b-d13bc903e180 · inbound

Human Cognition in Machines: A Unified Perspective of World Models cites this paper.

Human Cognition in Machines: A Unified Perspective of World Models V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T00:33:51.269344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T08:12:15.663761Z digest=sha256:c25952b435d8872c4069dfa8d11927b61dd36fb141b088708c4f996f010a6f40

Observation f054379c-ff58-47a3-ab08-0affcc7c8d33 · inbound

Active World-Model with 4D-informed Retrieval for Exploration and Awareness cites this paper.

Active World-Model with 4D-informed Retrieval for Exploration and Awareness V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:33:51.269344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T08:20:40.569482Z digest=sha256:06dc7723b2e2df35253c0a6df1d3e6ebf182509a4cbbe5313e2e690fd0009219

Observation 98c68955-945f-49bb-9590-6929e3521b14 · inbound

Watching Physics: the Generative Science of Matter and Motion cites this paper.

Watching Physics: the Generative Science of Matter and Motion V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:33:51.269344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T07:32:27.962931Z digest=sha256:6c2214f418c963e4dad7b8d360dc5f9da4f41e1d0976788c0c081f16a309c537

Observation 7b640478-bbcf-40c1-9021-ef2fd37152b5 · inbound

RoboWM-Bench: A Benchmark for Evaluating World Models in Robotic Manipulation cites this paper.

RoboWM-Bench: A Benchmark for Evaluating World Models in Robotic Manipulation V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T00:33:51.269344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T03:02:26.084185Z digest=sha256:b588542f2d21786680c173031f7a8333099df809cf4e3940b8f7f86c9d9d951c

Observation b8663770-f9a9-4c91-b695-9640e57bf87d · inbound

RoboWM-Bench: A Benchmark for Evaluating World Models in Robotic Manipulation cites this paper.

RoboWM-Bench: A Benchmark for Evaluating World Models in Robotic Manipulation V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T06:19:50.267503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T06:15:32.881140Z digest=sha256:f1a8072b875275b458aac8a4534f000eafecfcc8c28aa247d834d339b3d0ef74

Observation 3a9e2724-55be-464a-840f-c73f6113e49c · inbound

Mask World Model: Predicting What Matters for Robust Robot Policy Learning cites this paper.

Mask World Model: Predicting What Matters for Robust Robot Policy Learning V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T13:11:06.331616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T02:14:17.676675Z digest=sha256:10b374b147dc8e05b6d8131745a0cba1873a484d24de3041eb1acebbef75f6e2

Observation 2f9ddda3-64b1-4c17-a97c-023df2cd2804 · inbound

Cortex 2.0: Grounding World Models in Real-World Industrial Deployment cites this paper.

Cortex 2.0: Grounding World Models in Real-World Industrial Deployment V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:33:51.269344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T00:52:05.348704Z digest=sha256:b53d0df55cb73da8ee672609a0dae3cbb08a596c9616c9d8da7ef1e1ef218194

Observation 70f8e283-5974-4e01-9be8-c535914a03ee · inbound

Exploring High-Order Self-Similarity for Video Understanding cites this paper.

Exploring High-Order Self-Similarity for Video Understanding V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T00:33:51.269344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T00:59:03.890135Z digest=sha256:1a3980807757e7b8cfbfe19bcc1ccca7a0f6408f4a63047c847c8c76c8c21d2d

Observation c8f847a4-e859-4099-9380-7d3e2a2d9cf7 · inbound

Open-H-Embodiment: A Large-Scale Dataset for Enabling Foundation Models in Medical Robotics cites this paper.

Open-H-Embodiment: A Large-Scale Dataset for Enabling Foundation Models in Medical Robotics V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-11T13:56:05.568412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-09T23:53:09.614994Z digest=sha256:162b986942b620434d286f06ef6cf513837349d2f6d877056225562e75b40a23

Observation 6e70a9cc-382c-40c9-8444-cc4252b97c92 · inbound

Sapiens2 cites this paper.

Sapiens2 V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-11T14:21:06.834481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-09T21:59:43.755956Z digest=sha256:156eb06458dab7670dd1356d066df4f02a36992aa0c4e66a3039a7e22bb0a607

Observation 32102979-e1a1-4f54-b2ea-65e91e5a982a · inbound

Only Brains Align with Brains: Cross-Region Alignment Patterns Expose Limits of Normative Models cites this paper.

Only Brains Align with Brains: Cross-Region Alignment Patterns Expose Limits of Normative Models V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-11T18:56:07.566890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-08T13:13:03.129195Z digest=sha256:ce15e3a2cb6979036cb3fd2e4d47f42c8822ba66fd3fe62fc3570adbf421cc36

Observation 33747b10-1494-4a32-a753-d5d9f3fc89ac · inbound

A Co-Evolutionary Theory of Human-AI Coexistence: Mutualism, Governance, and Dynamics in Complex Societies cites this paper.

A Co-Evolutionary Theory of Human-AI Coexistence: Mutualism, Governance, and Dynamics in Complex Societies V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-11T20:16:08.977561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T09:51:33.783535Z digest=sha256:1b8b678219c3f78d4d3c87d4680aef552a56dc3bc296cd0d58a4b7492010807b

Observation bd7e36e3-8c87-49ea-940a-dfe051c0c54e · inbound

SS3D: End2End Self-Supervised 3D from Web Videos cites this paper.

SS3D: End2End Self-Supervised 3D from Web Videos V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T19:06:12.301942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T12:32:40.901242Z digest=sha256:8598d0bf56acdeea81f1fc0e4169e60ada207f148cb5fcfad354016ff0ff56ba

Observation 5dc7b409-eda9-4c7f-9e76-954eb13f19f5 · inbound

SS3D: End2End Self-Supervised 3D from Web Videos cites this paper.

SS3D: End2End Self-Supervised 3D from Web Videos V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T07:46:30.201939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:57:09.197785Z digest=sha256:9d690eeb5eab632ed4784fb3c07a43d974747b8d8a5d360c429b4ad8790f382d

Observation a7c581a3-6796-411e-ba2a-f53831bca7df · inbound

SS3D: End2End Self-Supervised 3D from Web Videos cites this paper.

SS3D: End2End Self-Supervised 3D from Web Videos V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-05-14T21:27:59.609043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T21:26:38.410299Z digest=sha256:3ec1ed4181738f0f97b888001fd071aced70669edf4670f0758528f000bde28f

Observation bfc17488-b348-4be3-9279-508eedc1a782 · inbound

Agentic World Modeling: Foundations, Capabilities, Laws, and Beyond cites this paper.

Agentic World Modeling: Foundations, Capabilities, Laws, and Beyond V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-11T19:26:07.829429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-08T12:02:07.027775Z digest=sha256:402806e3f820acd6f19dae05d4c7ac5ea33d41f1a9ca09f8ddbe573abb0b339f

Observation 3483b91e-7ff2-4218-b80f-a014b115de23 · inbound

Agentic World Modeling: Foundations, Capabilities, Laws, and Beyond cites this paper.

Agentic World Modeling: Foundations, Capabilities, Laws, and Beyond V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-04T17:29:59.538944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-04T17:29:43.764085Z digest=sha256:4770f5de0bb4e9f7ad471972eab3a7106100eaa09225de46daca20cb9f386e15

Observation 77f8cd15-cd30-45c4-8b22-2ecc5e7e580f · inbound

Learning Human-Intention Priors from Large-Scale Human Demonstrations for Robotic Manipulation cites this paper.

Learning Human-Intention Priors from Large-Scale Human Demonstrations for Robotic Manipulation V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-11T22:26:11.646080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T02:51:27.662262Z digest=sha256:dce686cc420c031e0f0bff3ddb3c48f07420838f103c3b11262639e4987ac73e

Observation 59f04218-6cf5-4c70-bc8d-1b8089e884b0 · inbound

Learning Human-Intention Priors from Large-Scale Human Demonstrations for Robotic Manipulation cites this paper.

Learning Human-Intention Priors from Large-Scale Human Demonstrations for Robotic Manipulation V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-22T11:21:29.203425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T11:16:58.104663Z digest=sha256:5199edbff63531a34f6785f7e096102e7baf0c775a29303995b0c1aee8115201

Observation 7198b9f8-18c0-452e-843c-25599a2684a7 · inbound

Lifting Embodied World Models for Planning and Control cites this paper.

Lifting Embodied World Models for Planning and Control V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T23:41:15.767575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:32:46.059831Z digest=sha256:74f558f16c0c985d2b6319ac1c9bee2375441c8a08939cfc63438a90761e1d4e

Observation bd2d379d-f512-49d7-8a73-019fc8909cf0 · inbound

Lifting Embodied World Models for Planning and Control cites this paper.

Lifting Embodied World Models for Planning and Control V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T15:20:24.269257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T15:20:24.269257Z digest=sha256:5f7a597d6b6765542faf8881cc50f6a280de7d0124051918d2e4811d0c36adc9