Pith. sign in

Paper Citation Record · LEDGER

JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling

As of 21 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 0 inbound Pith citation observations for arXiv:2608.09381.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.09381 v1

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T18:27:58.510808Z

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

20 of 20 outbound references displayed

  • verified exact1
  • verified fuzzy4
  • unresolved13
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2303cbd8-4dc4-4888-9817-23ff0b784e44 · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T18:27:58.440915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:27:58.440915Z digest=sha256:16103a56532ced922e4610f4fe91ea63a82d3336de1a7f8ca51f82bdc3190973

Observation 293745fc-0811-4b86-b9b6-1d60f3e39454 · outbound

This paper cites LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models.

JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T18:27:58.444976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:27:58.444976Z digest=sha256:179f06255327a7764934593c69ce7dd7b7c1ac036566b62cdceaef58b7bd14ee

Observation f5b30d59-8018-437d-9fb8-9752196624bb · outbound

This paper cites 0, .25, .50, .75,1 Three-block placement Place all three target blocks on the plate, with partial credit for completed placements.

JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling 0, .25, .50, .75,1 Three-block placement Place all three target blocks on the plate, with partial credit for completed placements

Reference 5

Resolution
malformed identifier
raw_fallback, observed 2026-08-11T18:27:58.812634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T18:27:58.510808Z digest=sha256:3ca79d58abae6c3ee863e1dc23e78d7251fd5796e5d5f0200a93822a37b27c8c

Observation 816fe6d9-7ca7-44ac-a613-8886548cfac5 · outbound

This paper cites Training Vision-Language-Action Models with Dense Embodied Chain-of-Thought Supervision.

JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling Training Vision-Language-Action Models with Dense Embodied Chain-of-Thought Supervision

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-08-11T18:27:58.757020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T18:27:58.458145Z digest=sha256:4f572d030e219ca7c4473057fdcf62f023af9a255010ea8b269794ad62050fd2

Observation 926b2cab-789c-4ed3-8301-018a6cbbb54f · outbound

This paper cites RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models.

JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T18:27:58.462125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:27:58.462125Z digest=sha256:bb7f67a4daff3dda2dc49db0049def87f852c372eb108cc36d50eba8aabfe82d

Observation 698618c8-f15b-4b4d-9065-bfe62e7954d2 · outbound

This paper cites Miao, S.; Feng, N.; Wu, J.; Lin, Y.; He, X.; Li, D.; and Long,M.2026.

JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling Miao, S.; Feng, N.; Wu, J.; Lin, Y.; He, X.; Li, D.; and Long,M.2026

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T18:27:58.465794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:27:58.465794Z digest=sha256:c9e61124179de01d4eba05c2db7c70e1bd1121f3f3d7d4dd1504f98523da70b2

Observation 96d463c9-6265-4e8d-88b3-1f03d7dd53f8 · outbound

This paper cites Physical Intelligence; Black, K.; Brown, N.; Darpinian, J.; Dhabalia,K.;Driess,D.;Esmail,A.;Equi,M.;Finn,C.;Fu- sai, N.; et al.

JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling Physical Intelligence; Black, K.; Brown, N.; Darpinian, J.; Dhabalia,K.;Driess,D.;Esmail,A.;Equi,M.;Finn,C.;Fu- sai, N.; et al

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:27:58.889066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T18:27:58.469594Z digest=sha256:731169bd0ec61ad19f204345c2ecae7df4787c1887779bc918d465f5c74e77ad

Observation e514e95f-37c6-440b-8713-498185037704 · outbound

This paper cites StarVLA: A Lego-like Codebase for Vision-Language-Action Model Developing.

JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling StarVLA: A Lego-like Codebase for Vision-Language-Action Model Developing

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T18:27:58.473548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:27:58.473548Z digest=sha256:d251c0a3db5194d2dd58c444057d7f12f1b82caee81c3600e62a19f4986dca64

Observation f0de75d1-b698-4252-8ff9-aa133464bf2d · outbound

This paper cites arXiv:2602.10098.

JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling arXiv:2602.10098

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T18:27:58.477033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:27:58.477033Z digest=sha256:a44c5f5d885e85035f9ff38539643e4ecef0206b76abfa4b43f76f56ad40db95

Observation 9693e49e-6508-49b1-8203-b676b420398d · outbound

This paper cites World Action Models are Zero-shot Policies.

JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling World Action Models are Zero-shot Policies

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T18:27:58.480327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:27:58.480327Z digest=sha256:868de15d2bf3a7e3e0f28493bfdabef2ba16a05fc54ef1b70968f4726992d0fe

Observation c9b4c20a-af3b-4a10-8003-c2fa25b596fa · outbound

This paper cites PokeVLA: Empowering Pocket-Sized Vision-Language-Action Model with Comprehensive World Knowledge Guidance.

JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling PokeVLA: Empowering Pocket-Sized Vision-Language-Action Model with Comprehensive World Knowledge Guidance

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T18:27:58.485559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:27:58.485559Z digest=sha256:4e67ccaee4058045a0e2acfaa7c5e26937948fbdb0ee082ef4a71674ad71bd29

Observation dcd39c70-e222-48b7-812c-3bea6e98989d · outbound

This paper cites From Noise to Intent: Anchoring Generative VLA Policies with Residual Bridges.

JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling From Noise to Intent: Anchoring Generative VLA Policies with Residual Bridges

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-08-11T18:27:58.542580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T18:27:58.489065Z digest=sha256:c25efe5c9888efd8722c9818fda01a7aca68147191a122786eb33cfae9a75e69

Observation 253f5c72-3d5f-4866-b74a-b7ffe55ee3df · outbound

This paper cites We report the average task success rate over all 20 tasks.

JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling We report the average task success rate over all 20 tasks

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:27:58.824845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T18:27:58.507689Z digest=sha256:8dac395a49418b34f3686eb7d76c1efff78d4f6e15eda6bbb4bcfa623ae983ea

Observation ce4913ae-baa9-4306-87c9-193129b725ea · outbound

This paper cites an unresolved cited work.

JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-11T18:27:58.863555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T18:27:58.496584Z digest=sha256:10b88b6ead94d4fd5a43cce8450351864cface33a39507e96c6e441fbf6c7137

Observation a941725f-048f-4c70-8607-e400892a8cc1 · outbound

This paper cites Only the Qwen LoRA adapters, transition prediction head, and action expert are optimized.

JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling Only the Qwen LoRA adapters, transition prediction head, and action expert are optimized

Reference 128

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:27:58.850605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T18:27:58.499962Z digest=sha256:4c32c3dd675f85d9dc3abba54a8e8e5714b7c7a12daeb627899c76d10d573b87

Observation ab67598c-f27b-4f18-8114-ca5b555e49bd · outbound

This paper cites an unresolved cited work.

JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling Unresolved cited work

Reference 896

Resolution
unresolved
raw_fallback, observed 2026-08-11T18:27:58.876459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T18:27:58.492463Z digest=sha256:5a9f3988456f78c74f1b7326e177e92bb14a4dd7ac7c1ed48b169304eec3110f

Observation b1dd2530-887d-4ce4-97d4-053f5f09c2c9 · outbound

This paper cites an unresolved cited work.

JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling Unresolved cited work

Reference 2023

Resolution
unresolved
raw_fallback, observed 2026-08-11T18:27:58.838440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T18:27:58.503962Z digest=sha256:8c907d965e84b01da8e2df8b03f548db29b74b14cf7291ab937abbe31098c0c9

Observation 0079f681-f4c2-4610-a6ec-b52073f66b85 · outbound

This paper cites InForty-first International Conference on Machine Learn- ing,ICML2024,Vienna,Austria,July21–27,2024,volume 235 ofProceedings of Machine Learning Research, 23123– 23144.

JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling InForty-first International Conference on Machine Learn- ing,ICML2024,Vienna,Austria,July21–27,2024,volume 235 ofProceedings of Machine Learning Research, 23123– 23144

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:27:58.900366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T18:27:58.454400Z digest=sha256:03d3ee2abcd29bd1d6d079714090d6f6b662fade75d79d92e825632e8e948100

Observation fed3a5c9-3654-4bff-911b-8d7a96e6b72f · outbound

This paper cites V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning.

JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-11T18:27:58.435802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:27:58.435802Z digest=sha256:96370dd76cdbffb82295fd9242dcab6d0adfd00bf0c4eeaf418732c105f3e4ad

Observation ff31df74-baf2-480e-901b-07ac4f7cb4a7 · outbound

This paper cites Vision-aligned Latent Reasoning for Multi-modal Large Language Model.

JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling Vision-aligned Latent Reasoning for Multi-modal Large Language Model

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-11T18:27:58.449889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:27:58.449889Z digest=sha256:b5084e39119bef0f5ce5bf4fb7e8d6e11846ed90cc1c0c4e7cbd6fa5faf57d86

Pith citing papers

No inbound Pith citation observations are available.