Pith. sign in

Paper Citation Record · LEDGER

UniJEPA: A Unified Joint-Embedding Predictive Architecture for Task-Agnostic Visual World Modeling

As of 10 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 0 inbound Pith citation observations for arXiv:2608.07409.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.07409 v1

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T05:06:02.556817Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

13 of 13 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 24cbaa79-2054-4393-8f54-e63325486322 · outbound

This paper cites RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.

UniJEPA: A Unified Joint-Embedding Predictive Architecture for Task-Agnostic Visual World Modeling RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T05:06:02.510016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:06:02.510016Z digest=sha256:d019a9a57fdba83699d1ea4fa66ef5e38742acbbc09459bfe09ff3dd1f31070b

Observation 8cc7a71d-ba13-4532-8ecf-be3043235c94 · outbound

This paper cites Learning and Leveraging World Models in Visual Representation Learning.

UniJEPA: A Unified Joint-Embedding Predictive Architecture for Task-Agnostic Visual World Modeling Learning and Leveraging World Models in Visual Representation Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T05:06:02.521473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:06:02.521473Z digest=sha256:df19f57c1d0cc746b2cbf932c5e29bb665e2e21742ed97928b268b310ccbd70a

Observation 30bd77c7-f4ca-4cbb-a6b0-8ed961ac26bf · outbound

This paper cites LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixels.

UniJEPA: A Unified Joint-Embedding Predictive Architecture for Task-Agnostic Visual World Modeling LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixels

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T05:06:02.537117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:06:02.537117Z digest=sha256:04dacfa7bb716b4eba945a9fa3acc991014c53b159daeb49b55efc204eb41272

Observation 060cbd3e-4b45-4538-a678-032f4d4a006c · outbound

This paper cites MUSE: Resolving Manifold Misalignment in Visual Tokenization via Topological Orthogonality.

UniJEPA: A Unified Joint-Embedding Predictive Architecture for Task-Agnostic Visual World Modeling MUSE: Resolving Manifold Misalignment in Visual Tokenization via Topological Orthogonality

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T05:06:02.546806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:06:02.546806Z digest=sha256:d4b6cf01867a937dbf777bf7057e802e98da269b087d9cc37630267c0c224705

Observation f78f59b3-f56f-4c3c-bb46-2df38e4b4c4f · outbound

This paper cites DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning.

UniJEPA: A Unified Joint-Embedding Predictive Architecture for Task-Agnostic Visual World Modeling DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T05:06:02.551652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:06:02.551652Z digest=sha256:c0dcbe5afca5aaf4255859a5d511e9b1231c2ea93b5c36894a640705c93a7ab4

Observation c6a2148f-a133-4997-aaa7-aa526c13740e · outbound

This paper cites Implementation Details We use a ViT-Small/16 encoder (15M parameters) for control and ViT-Large for large-scale image/video representation.

UniJEPA: A Unified Joint-Embedding Predictive Architecture for Task-Agnostic Visual World Modeling Implementation Details We use a ViT-Small/16 encoder (15M parameters) for control and ViT-Large for large-scale image/video representation

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T05:06:02.767745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T05:06:02.556817Z digest=sha256:548aeba611c4591df6aa5f77947450c09fdeaff1c42706058f616c15ad7a79ff

Observation 5865df4c-7a22-4d94-84a9-4a385d0569b1 · outbound

This paper cites World Model on Million-Length Video And Language With Blockwise RingAttention.

UniJEPA: A Unified Joint-Embedding Predictive Architecture for Task-Agnostic Visual World Modeling World Model on Million-Length Video And Language With Blockwise RingAttention

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-10T05:06:02.541971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:06:02.541971Z digest=sha256:96b20e0c6fda878ac1cfba46918edb476c64b51c7c1c6dca1b4b8f36316ee09f

Observation fda6f588-d0ca-437e-99fd-4408a6368685 · outbound

This paper cites Mastering Diverse Domains through World Models.

UniJEPA: A Unified Joint-Embedding Predictive Architecture for Task-Agnostic Visual World Modeling Mastering Diverse Domains through World Models

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-10T05:06:02.532154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:06:02.532154Z digest=sha256:ed4412785044bc1deae20840ca3f11ccacca5497d442e6e96733e6bd02f42ad8

Observation 0fb21eff-7dbb-464a-8316-875e9d85aca7 · outbound

This paper cites World Models.

UniJEPA: A Unified Joint-Embedding Predictive Architecture for Task-Agnostic Visual World Modeling World Models

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-10T05:06:02.527212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:06:02.527212Z digest=sha256:945a7f9f2b153ac4a0d81968ab620ca9dfc95d55af8e287f29bad8ca73b7bd99

Observation 5f1d49ef-b69e-4614-9827-9b04cfb27af0 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

UniJEPA: A Unified Joint-Embedding Predictive Architecture for Task-Agnostic Visual World Modeling OpenVLA: An Open-Source Vision-Language-Action Model

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-10T05:06:02.516123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:06:02.516123Z digest=sha256:622bae1c40ec03ce1f6defb7e0ca0b702311dfd5e2efff55a4755329bd939786

Observation 352b986b-b00f-4e1c-950a-b0af7760ac81 · outbound

This paper cites Back to the Features: DINO as a Foundation for Video World Models.

UniJEPA: A Unified Joint-Embedding Predictive Architecture for Task-Agnostic Visual World Modeling Back to the Features: DINO as a Foundation for Video World Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-10T05:06:02.500245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:06:02.500245Z digest=sha256:754c26472920092092403e215201b764ee05cba5318677c3fdde87549cf1bad3

Observation 1d969f65-d119-40c4-8b92-833dea55a436 · outbound

This paper cites V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning.

UniJEPA: A Unified Joint-Embedding Predictive Architecture for Task-Agnostic Visual World Modeling V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-10T05:06:02.494586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:06:02.494586Z digest=sha256:c14f34746fbd4a24d0c08dcfcd3b809170c367fbf718dbb180e6eca1893a5529

Observation da6c6179-14b6-45d5-a77f-5e42090c050b · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

UniJEPA: A Unified Joint-Embedding Predictive Architecture for Task-Agnostic Visual World Modeling Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-10T05:06:02.505056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:06:02.505056Z digest=sha256:0197bbcb7b4045895244df5a0d3f78b442993751caa5fc86be4f7758c6499f95

Pith citing papers

No inbound Pith citation observations are available.