Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-11T00:33:50.471804Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 69 of 69 outbound references and 100 inbound Pith citation observations for arXiv:2506.09985.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-11T00:33:50.471804Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-05T06:04:03.420693Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
69 of 69 outbound references displayed
External citation measurements
3
pith, observed 2026-08-05T02:28:24.338817Z
Observation 201bc43f-a55d-4a40-abd1-5a4ab1bd16a4 · outbound
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning Cosmos World Foundation Model Platform for Physical AI
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 84124e4f-7ec8-46d3-9e7c-2f185788a459 · outbound
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning The Hidden Uniform Cluster Prior in Self-Supervised Learning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0c4c3889-4ac4-43e6-b0b7-542a2be6494a · outbound
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning Revisiting Feature Prediction for Learning Visual Representations from Video
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1cac71f1-58df-4132-bd8c-e6bfe8e9cf69 · outbound
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning GR00T N1: An Open Foundation Model for Generalist Humanoid Robots
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5b6f6463-1cb4-4704-8729-e5b45bc57d45 · outbound
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 72dc4682-8985-4280-81fd-d359e9130445 · outbound
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a641c187-947a-4c2a-a8d5-7f5e6f296f4c · outbound
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning Perception Encoder: The best visual embeddings are not at the output of the network
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c6b04c45-8f01-4280-97ac-3488f0dd72d2 · outbound
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b3e7157c-7b90-4817-989b-f5200612b389 · outbound
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 44f9af9e-5d1a-4739-96ae-26e3ee04a784 · outbound
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning A Short Note about Kinetics-600
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cb071916-8124-410a-98cb-d137010fa565 · outbound
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning A Short Note on the Kinetics-700 Human Action Dataset
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3f95764d-646c-414b-9d2c-358331e7599a · outbound
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning Scaling 4D Representations
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cb939337-9d3a-4bff-99dd-4dab5c33a910 · outbound
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning Actionable Models: Unsupervised Offline Reinforcement Learning of Robotic Skills
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3cd5d2a7-0d85-490e-a597-fadc8b1949d0 · outbound
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5481e072-022d-4fb6-974c-a1f73ccbf903 · outbound
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2ddcf280-f912-4b04-a12e-5e84b7c485e2 · outbound
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning Lost in Time: A New Temporal Benchmark for VideoLLMs
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1605c1b9-48e0-4470-8876-c126e7c7da52 · outbound
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens-100.International Journal of Computer Vision (IJCV)
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 26548820-34cc-46ec-b905-9e3d7a804b1f · outbound
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 57437d9d-66ca-4e5a-87c3-adf9bf9e453e · outbound
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning Driess, F
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation fcb8aa3a-2c89-4015-82af-6b32beb1c621 · outbound
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning Visual Foresight: Model-Based Deep Reinforcement Learning for Vision-Based Robotic Control
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0b4fc876-9190-4d61-8583-94c60ded441c · outbound
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning Scaling Language-Free Visual Representation Learning
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0d1cdbad-61e0-4aa1-8757-c176ebea21e6 · outbound
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning Multimodal Autoregressive Pre-training of Large Vision Encoders
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1df28384-6207-4a77-8bf9-e715e6d7bd4f · outbound
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning Learning Visual Predictive Models of Physics for Playing Billiards
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation da9481c3-0afc-4f49-993a-4ed25268cb13 · outbound
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning He, B., Yin, L., Zhen, H.-L., Liu, S., Wu, H., Zhang, X., Yuan, M., and Ma, C
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3cb7acec-43f4-4ac7-8ef5-9fbac5a594c6 · outbound
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning The Llama 3 Herd of Models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c7a600d1-840b-4ff7-86ef-153bff29820d · outbound
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning MaskViT: Masked Visual Pre-Training for Video Prediction
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cbbef07d-29b2-47c4-95d4-fddd0f75e9ac · outbound
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning World Models
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 19a1361b-43c3-48e1-aff6-8e92026358e7 · outbound
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning Dream to Control: Learning Behaviors by Latent Imagination
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5754e964-70a6-4abd-a996-3b33939c6594 · outbound
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning Temporal Difference Learning for Model Predictive Control
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b35811ad-c69e-4a7d-8415-fa3bcaf5986e · outbound
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning TD-MPC2: Scalable, Robust World Models for Continuous Control
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c4ba759a-8199-4962-aa8e-2a77bd90c9d6 · outbound
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning GAIA-1: A Generative World Model for Autonomous Driving
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 05a596de-a1e0-4a57-8668-a476c90b7169 · outbound
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning Huang, P., Liu, S., Liu, Z., Yan, Y ., Wang, S., Chen, Z., and Xiao, T
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a353802c-bc41-43c8-8ec1-ee2124de08e8 · outbound
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning Perceiver IO: A General Architecture for Structured Inputs & Outputs
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 512192e1-5c2f-421b-8717-7545025e78eb · outbound
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning The Kinetics Human Action Video Dataset
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d9daf3d7-f670-43fe-b71d-216708b2a2c3 · outbound
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation dee87fdd-3f34-4d1e-919a-a6fce3010e77 · outbound
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning OpenVLA: An Open-Source Vision-Language-Action Model
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 95317c8e-06fc-4c1f-bef4-3d1bc77394f1 · outbound
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning A path towards autonomous machine intelligence version 0.9.2, 2022-06-27.Open Review, 62(1):1–62
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1a49219d-35ae-4aec-a482-1252cd056abb · outbound
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning LLaVA-OneVision: Easy Visual Task Transfer
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8359af7d-9ab3-4be8-9d3c-1fd8d95608af · outbound
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning TempCompass: Do Video LLMs Really Understand Videos?
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 45cc6c29-ea7d-4df9-993c-173a43c8c513 · outbound
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning Decoupled Weight Decay Regularization
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 06953f0c-1aeb-41ad-a154-2c2e310e8590 · outbound
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning Keypoints into the Future: Self-Supervised Correspondence in Model-Based Reinforcement Learning
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9d623b42-e6a5-42be-9ef5-29b84742618d · outbound
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning Octo: An Open-Source Generalist Robot Policy
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 57365b87-d7e6-4990-aca8-c82a929eeb0a · outbound
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning DINOv2: Learning Robust Visual Features without Supervision
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 06ea8c53-e396-447e-a2e7-7f7852559e51 · outbound
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning Qwen2.5-VL Technical Report
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 93043d74-8061-46d9-9ec8-2118051a49ef · outbound
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning An Empirical Study of Autoregressive Pre-training from Videos
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d57d4997-31f2-4682-a732-3cee5148dacf · outbound
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning GAIA-2: A Controllable Multi-View Generative World Model for Autonomous Driving
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation dd1c37b8-704b-4224-888d-0992d7313db4 · outbound
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning TOMATO: Assessing Visual Temporal Reasoning Capabilities in Multimodal Foundation Models
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f501c319-0753-418f-8e25-5100c409fa1e · outbound
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning Learning from reward-free offline data: A case for planning with latent dynamics models
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ed8e984d-368c-4623-95b6-256643006d18 · outbound
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning Video Occupancy Models
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation eec793d8-d0c9-4b3f-b870-7fcde67106e0 · outbound
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 03b7c11b-16a0-4b97-8a9c-c4fa9e3062d9 · outbound
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning The Evolution of Multimodal Model Architectures
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation efc3332f-1f50-491c-85bb-3c960ecc534a · outbound
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0ea17cc3-a4f7-4b31-8377-a22e27bd17ec · outbound
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 506152f2-e699-4975-852f-f94d86af54a0 · outbound
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5f53b7b9-dd40-42ca-9226-7639e797b152 · outbound
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 15bd45bc-a1b8-4c77-8dd8-83626e44a128 · outbound
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f937889c-ac7f-4d16-b967-bec5f6d8c96a · outbound
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5ad1e2d9-6c37-401d-bb2c-b32321d5c289 · outbound
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning FLARE: Robot Learning with Implicit World Modeling
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 339aa09a-d664-4b98-98dd-6f5afc436f7f · outbound
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 93e6b606-7cf2-4cff-936d-0abc15437426 · outbound
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning Unified World Models: Coupling Video and Action Diffusion for Pretraining on Large Robotic Datasets
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a1188a44-03fc-4259-831f-f1f75f816862 · outbound
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning abbreviated
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9652bef8-da87-4409-b4ff-cbec3f69077c · outbound
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning (2020), using the standard16 × 16 patch size
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1df189b3-4176-4803-8df5-004026c7a7ba · outbound
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning For the pick-and-place tasks we present two sub-goal images to the model in addition to the final goal
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3cca2a2f-f3c5-47cb-bfcf-cb6bbd4d06a8 · outbound
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning The MLLM ingests the output embeddings of the vision encoder, which are projected to the hidden dimension of the LLM backbone using aprojector module
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 763b4da0-03cc-4965-9bad-c241fb7b9ac3 · outbound
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning We describe the training details in the following sections
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 98a591e2-65dc-4d12-aa70-daabf6b9780d · outbound
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning To assess the ability of V-JEPA 2 to capture spatiotemporal details for VidQA, we compare to leading off-shelf image encoders
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 689f5c30-4102-4141-a14d-e6b9b166681a · outbound
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning Unlike Cho et al
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8c6206a5-0347-4ae4-a5f0-65cce4f07999 · outbound
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning We scale up the data size to 88.5 million samples
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5766f544-a493-431f-bb50-22e375a8bae9 · outbound
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning Unresolved cited work
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 21a9cdcc-18e3-4368-8bc6-7f934b17a38c · inbound
3D Foundation Model for Generalizable Disease Detection in Head Computed Tomography V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3f2e97c4-f009-4f96-bf8a-1d1219edbc22 · inbound
A Survey on Vision-Language-Action Models: An Action Tokenization Perspective V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 298
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 61102fb8-a0dc-4cd7-b391-66e53c861e88 · inbound
Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 201
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d30e72f5-c6d1-4ed0-9cef-06c758ff7752 · inbound
3D and 4D World Modeling: A Survey V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4b14d17-b51d-4562-b0c4-b0b79f98482b · inbound
A Survey of Reinforcement Learning for Large Reasoning Models V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 96208e04-2720-4f38-9ca4-5006ef5ab42c · inbound
Video models are zero-shot learners and reasoners V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7d0e224a-43b3-4916-99fb-1379ff8754b7 · inbound
World-Env: Leveraging World Model as a Virtual Environment for VLA Post-Training V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c0a10ac1-b6f2-4aeb-9670-ff93d76bbf30 · inbound
Inferring Dynamic Physical Properties from Video Foundation Models V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0ce7c558-6d52-4d25-ad06-07093d0bdb01 · inbound
AtomWorld: A Benchmark for Evaluating Spatial Reasoning in Large Language Models on Crystalline Materials V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2381dac-5508-4b30-95f3-13fea060946d · inbound
Hybrid Architectures for Language Models: Systematic Analysis and Design Insights V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 38d2428e-0744-43b7-88d2-6041cc7a3d29 · inbound
A Comprehensive Survey on World Models for Embodied AI V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab3cb743-68a0-4ace-b948-d3a1d9d89180 · inbound
World Simulation with Video Foundation Models for Physical AI V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4e4b2129-5193-46d3-a387-fc376242db98 · inbound
Cambrian-S: Towards Spatial Supersensing in Video V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d7ad238f-d48f-47bf-b56d-6e0de984978b · inbound
IPR-1: Interactive Physical Reasoner V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation dda7fd4f-7d7d-4d7e-a0bc-1b85c5188586 · inbound
IPR-1: Interactive Physical Reasoner V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d703c97-a3c9-4ed9-8e12-712545cabcff · inbound
POMA-3D: The Point Map Way to 3D Scene Understanding V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7a26b75e-ccdb-4541-94ad-957a160d7f01 · inbound
Rethinking Reward Signals in Video GRPO: When Scores Become Targets V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c58b4f50-b54f-4fa5-b31c-2ea5c5e1e5d0 · inbound
MIND-V: Hierarchical World Model for Long-Horizon Robotic Manipulation with RL-based Physical Alignment V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation adac12f7-08b5-4027-be17-8ce31abf9a29 · inbound
CLARITY: Medical World Model for Guiding Treatment Decisions by Modeling Context-Aware Disease Trajectories in Latent Space V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aef58132-70af-47e3-981d-a8f1dd0a24c8 · inbound
Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31f0bd24-cdaf-4924-a56e-f53b0ba3020d · inbound
A Geometric Theory of Cognition for Machine Intelligence V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af32d96c-5b12-46c1-b227-b72cec58107a · inbound
Recurrent Video Masked Autoencoders V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3f6f8a27-8e83-43a3-9a5b-74f822c77d77 · inbound
mimic-video: Video-Action Models for Generalizable Robot Control Beyond VLAs V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ae34eebe-d67b-4c98-ad3d-87c153291f73 · inbound
Probing and Leveraging Video Diffusion Transformer Features for Robust Point Tracking V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9499361-b3e8-4730-a19b-f05269409e1f · inbound
DriveLaW:Unifying Planning and Video Generation in a Latent Driving World V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0fd2bcdf-49c0-47e7-b972-a559bf8f87f7 · inbound
Learning to Feel the Future: DreamTacVLA for Contact-Rich Manipulation V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ad22814e-df61-40a5-9c25-ad49bd2911d4 · inbound
Learning to Feel the Future: DreamTacVLA for Contact-Rich Manipulation V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation baa17504-3b25-4f8f-9df5-f436e1ccdb58 · inbound
Flow Equivariant World Models: Memory for Partially Observed Dynamic Environments V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8bc65591-6b88-4c89-bb64-14a49b14b01e · inbound
CycleVLA: Proactive Self-Correcting Vision-Language-Action Models via Subtask Backtracking and Minimum Bayes Risk Decoding V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59019341-dc3a-42cf-9eb0-5eb63e765887 · inbound
Advancing Open-source World Models V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ccbc70af-4af1-4c23-95a0-32fde3d2353b · inbound
Drive-JEPA: Video JEPA Meets Multimodal Trajectory Distillation for End-to-End Driving V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10d87c3c-b641-4255-8841-35c48e2c3b4b · inbound
PEPR: Privileged Event-based Predictive Regularization for Domain Generalization V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b5c27669-8051-4f81-9fca-a33b4e8e9d6e · inbound
SurgMotion: A Video-Native Foundation Model for Universal Understanding of Surgical Videos V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d47e3b7b-3f15-4065-a7fd-c0029e97b160 · inbound
DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 53dafe14-8e19-4c33-bf29-1a261ef3d584 · inbound
OmniFysics: Towards Physical Intelligence Evolution via Omni-Modal Signal Processing and Network Optimization V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 97ca2aad-0555-4096-b31d-63295199f0b3 · inbound
Going with the Flow: Koopman Behavioral Models as Pseudo Planners for Visuo-Motor Dexterity V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8574663c-49b5-495e-ad2d-e4e5c55c2037 · inbound
Rolling Sink: Bridging Limited-Horizon Training and Open-Ended Testing in Autoregressive Video Diffusion V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 612e5dc3-2012-409f-9649-c32baa742f19 · inbound
Olaf-World: Orienting Latent Actions for Video World Modeling V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3cde03a6-daf4-4890-863d-8934ac6bd406 · inbound
RISE: Self-Improving Robot Policy with Compositional World Model V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation edea824e-f551-4587-8b00-c2326bdfcfe1 · inbound
World Action Models are Zero-shot Policies V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8ca2d4af-bf68-4a7a-b8bf-f51fae9b828e · inbound
Xray-Visual Models: Scaling Vision models on Industry Scale Data V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c0c1ec0-bcc9-43cf-814f-8a6d3b71c92b · inbound
Space-Time Forecasting of Dynamic Scenes with Motion-aware Gaussian Grouping V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6276ba32-2cb3-4495-86f7-5f86d7344504 · inbound
TrajTok: Learning Trajectory Tokens enables better Video Understanding V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation abbaeea2-1fbe-4f5e-a560-ab79d31205d2 · inbound
TrajTok: Learning Trajectory Tokens enables better Video Understanding V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed6458fc-c39e-49d3-a453-7e7473b89bd0 · inbound
GeoWorld: Geometric World Models V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e85ab659-0887-475e-a06d-6ba5ae6d989d · inbound
EgoMoD: Predicting Global Maps of Dynamics from Local Egocentric Observations V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1bd60c73-fe88-451c-88f1-86c1b145c81f · inbound
ReMoT: Reinforcement Learning with Motion Contrast Triplets V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89965aa2-5d0d-493d-8d98-55fb15729978 · inbound
RoboLight: A Dataset with Linearly Composable Illumination for Robotic Manipulation V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06261469-2999-4665-af3d-125c049f3567 · inbound
Dreamer-CDP: Improving Reconstruction-free World Models Via Continuous Deterministic Representation Prediction V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1d328c51-3b8a-4d96-bf3b-8ec7053f88a0 · inbound
PlayWorld: Learning Robot World Models from Autonomous Play V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d247fa89-3dd1-4fcb-8121-5ef947af325a · inbound
Self-Distillation of Hidden Layers for Self-Supervised Representation Learning V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16560a48-022d-4352-950b-e8246d764ca0 · inbound
Generation Models Know Space: Unleashing Implicit 3D Priors for Scene Understanding V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23af3160-2c36-400f-84c3-89ff5668f910 · inbound
LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixels V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ceb99782-4f30-4b07-9610-178fd1f13b3b · inbound
ThinkJEPA: Empowering Latent World Models with Large Vision-Language Reasoning Model V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f95f2fe-9259-479e-b25b-5e566ce335f8 · inbound
Factorization Regret mediates compositional generalization in latent space V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 100bf105-1541-4c8b-8aa8-2d51137a3a71 · inbound
Factorization Regret mediates compositional generalization in latent space V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3bbbd9e-94f2-4f4b-8571-300692b94b8b · inbound
Topological sum rule for geometric phases of quantum gates V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80ecc837-2a2f-4b48-be64-ab3373eb5029 · inbound
World Action Verifier: Self-Improving World Models via Forward-Inverse Asymmetry V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d094bea9-e17b-486c-9f17-a1f06f2bd0b3 · inbound
Hierarchical Planning with Latent World Models V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cf57247f-acab-491e-bf07-88a05c31f1ce · inbound
Emergent Compositional Communication for Latent World Properties V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b0e9a11d-3fbc-4e1a-a690-7786c981094e · inbound
Veo-Act: How Far Can Frontier Video Models Advance Generalizable Robot Manipulation? V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7323336e-7f29-4016-b2f9-33980d4733dd · inbound
OpenWorldLib: A Unified Codebase and Definition of Advanced World Models V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ef4339cd-0695-4a61-8b2a-e82da2a66563 · inbound
OpenWorldLib: A Unified Codebase and Definition of Advanced World Models V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6169bb95-c826-4c1f-af34-866a793f05be · inbound
StarVLA: A Lego-like Codebase for Vision-Language-Action Model Developing V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8d1c0a46-fbe9-4a36-9baf-565ffe1c7911 · inbound
Action Images: End-to-End Policy Learning via Multiview Video Generation V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2e3c5bc1-0b32-489a-bc99-9233bafaa622 · inbound
The Cartesian Cut in Agentic AI V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c4dbef46-0207-4e20-b674-434e937dacc5 · inbound
A Machine Learning Framework for Turbofan Health Estimation via Inverse Problem Formulation V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ff9a67a8-604d-4515-ae69-06bea3cbb133 · inbound
Phantom: Physics-Infused Video Generation via Joint Modeling of Visual and Latent Physical Dynamics V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f334355b-f8ce-4e64-914c-c9c4e2cb2118 · inbound
Phantom: Physics-Infused Video Generation via Joint Modeling of Visual and Latent Physical Dynamics V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f6276d56-bbf8-4642-bfac-e566e0fd01c8 · inbound
Matrix-Game 3.0: Real-Time and Streaming Interactive World Model with Long-Horizon Memory V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9b339fed-d2bf-4fa8-b583-4f83f8c6f0b8 · inbound
Zero-shot World Models Are Developmentally Efficient Learners V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9b94b971-c780-4c1d-a3f7-406da2b4b39a · inbound
GTASA: Ground Truth Annotations for Spatiotemporal Analysis, Evaluation and Training of Video Models V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 50996131-99c6-4bb5-8670-9375b8c28b1c · inbound
Observe Less, Understand More: Cost-aware Cross-scale Observation for Remote Sensing Understanding V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d5400338-cc49-433f-bb25-0d875da0e26f · inbound
Representations Before Pixels: Semantics-Guided Hierarchical Video Prediction V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e0490e0e-8127-4f85-9983-874e5fea1123 · inbound
Grounded World Model for Semantically Generalizable Planning V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 57faeb5d-c3c7-4e0a-a986-b4f43b48116d · inbound
Learning Versatile Humanoid Manipulation with Touch Dreaming V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation bc09f644-da8c-42d7-86ed-5f21a588671c · inbound
NTIRE 2026 Challenge on Video Saliency Prediction: Methods and Results V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c071330a-ffda-485a-9964-3a03dad4a69f · inbound
AnimationBench: Are Video Models Good at Character-Centric Animation? V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4a1bb3cd-6d4e-48ba-813c-d47acc0a6cd3 · inbound
Stylistic-STORM (ST-STORM) : Perceiving the Semantic Nature of Appearance V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation aa9b2aa5-4bde-4283-888b-d13bc903e180 · inbound
Human Cognition in Machines: A Unified Perspective of World Models V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f054379c-ff58-47a3-ab08-0affcc7c8d33 · inbound
Active World-Model with 4D-informed Retrieval for Exploration and Awareness V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 98c68955-945f-49bb-9590-6929e3521b14 · inbound
Watching Physics: the Generative Science of Matter and Motion V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7b640478-bbcf-40c1-9021-ef2fd37152b5 · inbound
RoboWM-Bench: A Benchmark for Evaluating World Models in Robotic Manipulation V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b8663770-f9a9-4c91-b695-9640e57bf87d · inbound
RoboWM-Bench: A Benchmark for Evaluating World Models in Robotic Manipulation V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3a9e2724-55be-464a-840f-c73f6113e49c · inbound
Mask World Model: Predicting What Matters for Robust Robot Policy Learning V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2f9ddda3-64b1-4c17-a97c-023df2cd2804 · inbound
Cortex 2.0: Grounding World Models in Real-World Industrial Deployment V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 70f8e283-5974-4e01-9be8-c535914a03ee · inbound
Exploring High-Order Self-Similarity for Video Understanding V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c8f847a4-e859-4099-9380-7d3e2a2d9cf7 · inbound
Open-H-Embodiment: A Large-Scale Dataset for Enabling Foundation Models in Medical Robotics V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6e70a9cc-382c-40c9-8444-cc4252b97c92 · inbound
Sapiens2 V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 32102979-e1a1-4f54-b2ea-65e91e5a982a · inbound
Only Brains Align with Brains: Cross-Region Alignment Patterns Expose Limits of Normative Models V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 33747b10-1494-4a32-a753-d5d9f3fc89ac · inbound
A Co-Evolutionary Theory of Human-AI Coexistence: Mutualism, Governance, and Dynamics in Complex Societies V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation bd7e36e3-8c87-49ea-940a-dfe051c0c54e · inbound
SS3D: End2End Self-Supervised 3D from Web Videos V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5dc7b409-eda9-4c7f-9e76-954eb13f19f5 · inbound
SS3D: End2End Self-Supervised 3D from Web Videos V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a7c581a3-6796-411e-ba2a-f53831bca7df · inbound
SS3D: End2End Self-Supervised 3D from Web Videos V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation bfc17488-b348-4be3-9279-508eedc1a782 · inbound
Agentic World Modeling: Foundations, Capabilities, Laws, and Beyond V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3483b91e-7ff2-4218-b80f-a014b115de23 · inbound
Agentic World Modeling: Foundations, Capabilities, Laws, and Beyond V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 77f8cd15-cd30-45c4-8b22-2ecc5e7e580f · inbound
Learning Human-Intention Priors from Large-Scale Human Demonstrations for Robotic Manipulation V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 59f04218-6cf5-4c70-bc8d-1b8089e884b0 · inbound
Learning Human-Intention Priors from Large-Scale Human Demonstrations for Robotic Manipulation V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7198b9f8-18c0-452e-843c-25599a2684a7 · inbound
Lifting Embodied World Models for Planning and Control V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation bd2d379d-f512-49d7-8a73-019fc8909cf0 · inbound
Lifting Embodied World Models for Planning and Control V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.