Pith. sign in

Paper Citation Record · LEDGER

JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling

As of 14 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 0 inbound Pith citation observations for arXiv:2608.09381.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.09381 v1

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T18:27:58.510808Z

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

20 of 20 outbound references displayed

  • verified exact1
  • verified fuzzy4
  • unresolved13
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2303cbd8-4dc4-4888-9817-23ff0b784e44 · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T18:27:58.440915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:27:58.440915Z digest=sha256:12a7bc41e092a3372e32b8a5475204c20b6d28a5f832ab3bba0d60935a945bd5

Observation 293745fc-0811-4b86-b9b6-1d60f3e39454 · outbound

This paper cites LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models.

JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T18:27:58.444976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:27:58.444976Z digest=sha256:3ac4dce3240f6596dba9c9f35b01e89bb72e8fbb7744fb9f3adbb2834a20b779

Observation f5b30d59-8018-437d-9fb8-9752196624bb · outbound

This paper cites 0, .25, .50, .75,1 Three-block placement Place all three target blocks on the plate, with partial credit for completed placements.

JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling 0, .25, .50, .75,1 Three-block placement Place all three target blocks on the plate, with partial credit for completed placements

Reference 5

Resolution
malformed identifier
raw_fallback, observed 2026-08-11T18:27:58.812634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T18:27:58.510808Z digest=sha256:4267256b148c52863da210e654e023779e6fc04ff1084ada35b532c0cff31641

Observation 816fe6d9-7ca7-44ac-a613-8886548cfac5 · outbound

This paper cites Training Vision-Language-Action Models with Dense Embodied Chain-of-Thought Supervision.

JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling Training Vision-Language-Action Models with Dense Embodied Chain-of-Thought Supervision

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-08-11T18:27:58.757020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T18:27:58.458145Z digest=sha256:e48994c5d7054461e7f749f067d07bd0c267df096b59e8c93dc6d922f5be7f01

Observation 926b2cab-789c-4ed3-8301-018a6cbbb54f · outbound

This paper cites RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models.

JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T18:27:58.462125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:27:58.462125Z digest=sha256:22c02063e90d3da48e87b3e011d220754183421c1230bcf78d720d50daebc639

Observation 698618c8-f15b-4b4d-9065-bfe62e7954d2 · outbound

This paper cites Miao, S.; Feng, N.; Wu, J.; Lin, Y.; He, X.; Li, D.; and Long,M.2026.

JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling Miao, S.; Feng, N.; Wu, J.; Lin, Y.; He, X.; Li, D.; and Long,M.2026

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T18:27:58.465794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:27:58.465794Z digest=sha256:5a0832b57c270e49c2f920523559c33caaaef524b6bca6ec5c8b6373c7d5f32e

Observation 96d463c9-6265-4e8d-88b3-1f03d7dd53f8 · outbound

This paper cites Physical Intelligence; Black, K.; Brown, N.; Darpinian, J.; Dhabalia,K.;Driess,D.;Esmail,A.;Equi,M.;Finn,C.;Fu- sai, N.; et al.

JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling Physical Intelligence; Black, K.; Brown, N.; Darpinian, J.; Dhabalia,K.;Driess,D.;Esmail,A.;Equi,M.;Finn,C.;Fu- sai, N.; et al

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:27:58.889066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T18:27:58.469594Z digest=sha256:db47dbb714bd183d6a335e824f17705a848e6dc072fa226c38f020b8e508f3a4

Observation e514e95f-37c6-440b-8713-498185037704 · outbound

This paper cites StarVLA: A Lego-like Codebase for Vision-Language-Action Model Developing.

JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling StarVLA: A Lego-like Codebase for Vision-Language-Action Model Developing

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T18:27:58.473548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:27:58.473548Z digest=sha256:435a838002eee0abc8aef603d8fe648113debeef367368e267a7cc6047cbd221

Observation f0de75d1-b698-4252-8ff9-aa133464bf2d · outbound

This paper cites arXiv:2602.10098.

JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling arXiv:2602.10098

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T18:27:58.477033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:27:58.477033Z digest=sha256:eeea1c4565749d9d84c70705307eef2a81ab3cee5945a8c7524f04f36effe10e

Observation 9693e49e-6508-49b1-8203-b676b420398d · outbound

This paper cites World Action Models are Zero-shot Policies.

JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling World Action Models are Zero-shot Policies

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T18:27:58.480327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:27:58.480327Z digest=sha256:21f9e8fde0678c9968a88b13e3c5a0d659af109c93188fd85d260c37ca2cacc7

Observation c9b4c20a-af3b-4a10-8003-c2fa25b596fa · outbound

This paper cites PokeVLA: Empowering Pocket-Sized Vision-Language-Action Model with Comprehensive World Knowledge Guidance.

JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling PokeVLA: Empowering Pocket-Sized Vision-Language-Action Model with Comprehensive World Knowledge Guidance

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T18:27:58.485559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:27:58.485559Z digest=sha256:a17b2fddfba061f934dab259fefa92901adb617448558232cd88ff5a56a8edd5

Observation dcd39c70-e222-48b7-812c-3bea6e98989d · outbound

This paper cites From Noise to Intent: Anchoring Generative VLA Policies with Residual Bridges.

JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling From Noise to Intent: Anchoring Generative VLA Policies with Residual Bridges

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-08-11T18:27:58.542580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T18:27:58.489065Z digest=sha256:96a25a4c77ea54abf4fdb9329513b9a39ebf57568a99235e6c4796fda4461f1a

Observation 253f5c72-3d5f-4866-b74a-b7ffe55ee3df · outbound

This paper cites We report the average task success rate over all 20 tasks.

JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling We report the average task success rate over all 20 tasks

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:27:58.824845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T18:27:58.507689Z digest=sha256:d86adacfaff798729a03ef51c40bcc473bb40d6f2bb8490ac5ca8452b9963a20

Observation ce4913ae-baa9-4306-87c9-193129b725ea · outbound

This paper cites an unresolved cited work.

JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-11T18:27:58.863555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T18:27:58.496584Z digest=sha256:3b5bee1b1bbc2c33d41ae72db0ea2adf336271aaaf812445f13db951a2cab1e2

Observation a941725f-048f-4c70-8607-e400892a8cc1 · outbound

This paper cites Only the Qwen LoRA adapters, transition prediction head, and action expert are optimized.

JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling Only the Qwen LoRA adapters, transition prediction head, and action expert are optimized

Reference 128

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:27:58.850605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T18:27:58.499962Z digest=sha256:4f236b0e91b9efc49be19027f39be9ad864b32626d9518dee5b932c3b7ad2ed7

Observation ab67598c-f27b-4f18-8114-ca5b555e49bd · outbound

This paper cites an unresolved cited work.

JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling Unresolved cited work

Reference 896

Resolution
unresolved
raw_fallback, observed 2026-08-11T18:27:58.876459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T18:27:58.492463Z digest=sha256:37e0283d20e37e8c6d4fb7e9811dbf74d4f402b6e19ee825f097d64f8501239c

Observation b1dd2530-887d-4ce4-97d4-053f5f09c2c9 · outbound

This paper cites an unresolved cited work.

JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling Unresolved cited work

Reference 2023

Resolution
unresolved
raw_fallback, observed 2026-08-11T18:27:58.838440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T18:27:58.503962Z digest=sha256:aa4519332d395f1dbf86bf6b965e46095087808df2a8161326bd5cc9c9cc62d9

Observation 0079f681-f4c2-4610-a6ec-b52073f66b85 · outbound

This paper cites InForty-first International Conference on Machine Learn- ing,ICML2024,Vienna,Austria,July21–27,2024,volume 235 ofProceedings of Machine Learning Research, 23123– 23144.

JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling InForty-first International Conference on Machine Learn- ing,ICML2024,Vienna,Austria,July21–27,2024,volume 235 ofProceedings of Machine Learning Research, 23123– 23144

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:27:58.900366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T18:27:58.454400Z digest=sha256:c5734da65b728656e310c3fbcb03af2152fb250cc0b8ce45421f492a828dfcb2

Observation fed3a5c9-3654-4bff-911b-8d7a96e6b72f · outbound

This paper cites V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning.

JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-11T18:27:58.435802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:27:58.435802Z digest=sha256:7fe1505c4c04a711e04aa7c1cf145906bbd38d1bb70172590ec330381134efc1

Observation ff31df74-baf2-480e-901b-07ac4f7cb4a7 · outbound

This paper cites Vision-aligned Latent Reasoning for Multi-modal Large Language Model.

JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling Vision-aligned Latent Reasoning for Multi-modal Large Language Model

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-11T18:27:58.449889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:27:58.449889Z digest=sha256:9ed6128cdfdf76adb2b87292a22d0f96dde7e8dc820a4c98b43e04f2051c338b

Pith citing papers

No inbound Pith citation observations are available.