Pith. sign in

Paper Citation Record · LEDGER

Video Representation Learning with Joint-Embedding Predictive Architectures

As of 15 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 4 inbound Pith citation observations for arXiv:2412.10925.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.10925 v1

Coverage vector

measured 16 of 16 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T15:33:35.034268Z

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:37:35.343073Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T15:49:58.016388Z

Reference resolution

16 of 16 outbound references displayed

  • verified exact2
  • verified fuzzy5
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 43a3a51e-1fea-4970-8d14-b650079c9fd7 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Video Representation Learning with Joint-Embedding Predictive Architectures Adam: A Method for Stochastic Optimization

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T15:33:34.746792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:33:34.746792Z digest=sha256:ba1dee61e555bd4486f0df7d6576cc908b6371b47a6e4db8c65192186be157dd

Observation e6bc7076-ee9e-4933-a28b-511a6fced05b · outbound

This paper cites Evolving Losses for Unlabeled Video Representation Learning.

Video Representation Learning with Joint-Embedding Predictive Architectures Evolving Losses for Unlabeled Video Representation Learning

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-11T15:33:35.253458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:33:34.938748Z digest=sha256:33fd45112abefde918b0414e0936c9efa82431f141929fdd10d27838c4bc6b91

Observation c11d5329-9947-4222-be01-4c3990ebc822 · outbound

This paper cites An Information-Theoretic Perspective on Variance-Invariance-Covariance Regularization.

Video Representation Learning with Joint-Embedding Predictive Architectures An Information-Theoretic Perspective on Variance-Invariance-Covariance Regularization

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T15:33:34.964758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:33:34.964758Z digest=sha256:ab256f95c47e7615ced48769c14859a32b3c4fc773bfa1fbf90422cc7937998c

Observation 5c63c0fe-e37a-4732-a3e2-af880b1cc40e · outbound

This paper cites CLEVRER: CoLlision Events for Video REpresentation and Reasoning.

Video Representation Learning with Joint-Embedding Predictive Architectures CLEVRER: CoLlision Events for Video REpresentation and Reasoning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T15:33:34.998477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:33:34.998477Z digest=sha256:2df143c160f06fba390c09dc80e5b0fe06a6ddd7004bdb9bad41c5b81198c768

Observation 511f9f14-6c56-43fe-9fb1-0790c3881300 · outbound

This paper cites Variance-Covariance Regularization Improves Representation Learning.

Video Representation Learning with Joint-Embedding Predictive Architectures Variance-Covariance Regularization Improves Representation Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T15:33:35.020514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:33:35.020514Z digest=sha256:661c5025540022290845bd90e70b806dbff90cc13b657574fc04fddaa93fa290

Observation e6c97888-4f75-4dde-8177-38c0eb58ea32 · outbound

This paper cites an unresolved cited work.

Video Representation Learning with Joint-Embedding Predictive Architectures Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:33:35.577831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:33:35.026041Z digest=sha256:cfe09acdc66e8fa07ee0952c0afb4d87f68627ab08ff8b340babbd76437d62b9

Observation 3ae96479-0343-4154-8efe-6ad4cac5a881 · outbound

This paper cites Each video contains 300 frames at 24 frames per second at 320x240 resolution.

Video Representation Learning with Joint-Embedding Predictive Architectures Each video contains 300 frames at 24 frames per second at 320x240 resolution

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:33:35.564661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:33:35.030323Z digest=sha256:0ba7eb89cddd89c96581add6ae897a523c5d2a65847a835d97a84f5f217168e0

Observation 9c4a4b5b-e4ea-438f-b450-8482bcfe85ec · outbound

This paper cites The model on the left has PSNR of 22.8 and the one on the right has PSNR of 21.2.

Video Representation Learning with Joint-Embedding Predictive Architectures The model on the left has PSNR of 22.8 and the one on the right has PSNR of 21.2

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:33:35.533152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:33:35.034268Z digest=sha256:1109ba835d3161031838f8b4bd0facc583951c3cee9653dec58a70a11f5e699a

Observation 2b1bb1f0-92fb-4a67-9fbd-d791367c1a04 · outbound

This paper cites neurips.cc/paper_files/paper/1993/file/288cc0ff022877bd3df94bc9360b9c5d-Paper.pdf.

Video Representation Learning with Joint-Embedding Predictive Architectures neurips.cc/paper_files/paper/1993/file/288cc0ff022877bd3df94bc9360b9c5d-Paper.pdf

Reference 1993

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:33:35.901120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:33:34.633727Z digest=sha256:50e533ec312a010e10003258d1612987bf8e13d2715ce20d39413c995ecebf92

Observation cd786fd0-65b6-47bf-8cb6-dbda02d415cb · outbound

This paper cites A path towards autonomous machine intelligence version 0.9.

Video Representation Learning with Joint-Embedding Predictive Architectures A path towards autonomous machine intelligence version 0.9

Reference 2014

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:33:35.590337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:33:34.822235Z digest=sha256:ae13a0ab699908889146d6061b67610fc5c3a715a2a5df123dc42c0c23267b9a

Observation 4512778e-7f4d-4259-a0ba-9404a5e88499 · outbound

This paper cites A Comprehensive Review on Autonomous Navigation.

Video Representation Learning with Joint-Embedding Predictive Architectures A Comprehensive Review on Autonomous Navigation

Reference 2015

Resolution
verified exact
local_arxiv, observed 2026-08-11T15:33:35.311554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:33:34.872349Z digest=sha256:e5ab53a16e1212971d105744c2c9df15d9abc75de9e32c9782149c02878060c4

Observation 59569077-ed74-432d-bad6-9d9b22838424 · outbound

This paper cites Understanding Dimensional Collapse in Contrastive Self-supervised Learning.

Video Representation Learning with Joint-Embedding Predictive Architectures Understanding Dimensional Collapse in Contrastive Self-supervised Learning

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-11T15:33:34.723038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:33:34.723038Z digest=sha256:8161c7b1aea3b6cc876da980d0126992ecd9ee2dc9dd45d8a83cf215d62a10db

Observation 2aa3a84e-aebd-407a-b62d-46d5d3986f52 · outbound

This paper cites Prediction Under Uncertainty with Error-Encoding Networks.

Video Representation Learning with Joint-Embedding Predictive Architectures Prediction Under Uncertainty with Error-Encoding Networks

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-11T15:33:34.718900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:33:34.718900Z digest=sha256:58a5fe5ff3a46dc062524f1848bf3a803f6d81495eaff9bbc005ad7492fd509a

Observation cb4c7050-19b8-4581-8495-d89385d36385 · outbound

This paper cites Dimensionality reduction by learning an invariant mapping.

Video Representation Learning with Joint-Embedding Predictive Architectures Dimensionality reduction by learning an invariant mapping

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:33:35.800813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:33:34.714554Z digest=sha256:76daea7eb537a06a1ae12d762b345d00c5e009d9b9ca8c7eddee57abfb5b7de0

Observation 1249f408-1edf-44fa-a101-ed43c70cdf8f · outbound

This paper cites Deep multi-scale video prediction beyond mean square error.

Video Representation Learning with Joint-Embedding Predictive Architectures Deep multi-scale video prediction beyond mean square error

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-11T15:33:34.828546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:33:34.828546Z digest=sha256:5799f8050608d38a03a420f9df37d16231b333ef50f7e184394f3abb7aa3ab5d

Observation 13911945-1afd-4b63-80e6-a897e2f06dfb · outbound

This paper cites CATER: A diagnostic dataset for Compositional Actions and TEmporal Reasoning.

Video Representation Learning with Joint-Embedding Predictive Architectures CATER: A diagnostic dataset for Compositional Actions and TEmporal Reasoning

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T15:33:34.688128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:33:34.688128Z digest=sha256:b083cc6e7c8ac61376d6b216666fd66046c298fce74d5f600e12e385187896ff

Pith citing papers

Observation 99be13f5-de93-4de8-aef3-531f3a746e39 · inbound

Time to Embed: Unlocking Foundation Models for Time Series with Channel Descriptions cites this paper.

Time to Embed: Unlocking Foundation Models for Time Series with Channel Descriptions Video Representation Learning with Joint-Embedding Predictive Architectures

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T15:37:35.343073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:37:35.343073Z digest=sha256:7c409cd7c69a5bf3132c0d31830524900d56f46fde07931c505c5fe6eabaad06

Observation 4119aa35-b59e-48b4-999e-0424ccc17b56 · inbound

A Survey: Learning Embodied Intelligence from Physical Simulators and World Models cites this paper.

A Survey: Learning Embodied Intelligence from Physical Simulators and World Models Video Representation Learning with Joint-Embedding Predictive Architectures

Reference 266

Resolution
unresolved
no resolver link, observed 2026-08-06T21:09:16.043159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:09:16.043159Z digest=sha256:e0cbedbdb4c56c4f3e191b03478b93417852b87dbc15be7ddcaeb868502d86dd

Observation 8f9a0b6f-ffd7-4153-9a00-be14658ebe90 · inbound

AeroJEPA: Learning Semantic Latent Representations for Scalable 3D Aerodynamic Field Modeling cites this paper.

AeroJEPA: Learning Semantic Latent Representations for Scalable 3D Aerodynamic Field Modeling Video Representation Learning with Joint-Embedding Predictive Architectures

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:36:09.367752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-09T15:53:39.074115Z digest=sha256:e4d6a139347f3fc1cc9741fd56357afdddd745e0b099043419c5f66055c2831c

Observation 46b3717d-678d-4112-bd29-144f9f242670 · inbound

Neural Voxel Dynamics: Learning Implicit 3D Physics via Volumetric Feature Advection cites this paper.

Neural Voxel Dynamics: Learning Implicit 3D Physics via Volumetric Feature Advection Video Representation Learning with Joint-Embedding Predictive Architectures

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:49:58.017770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-26T01:17:06.587807Z digest=sha256:1d21464496310abbbe19e26c4d656eb9c957dbb5ad0e589d915110ae6b1bb522