Pith. sign in

Paper Citation Record · LEDGER

World Tokens: Enhancing Embodied Policies with Training-Time World Modeling

As of 16 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 0 inbound Pith citation observations for arXiv:2608.09730.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.09730 v1

Coverage vector

measured 14 of 14 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T11:52:55.674025Z

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

14 of 14 outbound references displayed

  • verified exact1
  • verified fuzzy1
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f6db37f8-df69-4a5b-a9bc-1bec9f8e7372 · outbound

This paper cites Qwen3-VL Technical Report.

World Tokens: Enhancing Embodied Policies with Training-Time World Modeling Qwen3-VL Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:55.636885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:55.636885Z digest=sha256:f45ee3b1af21cafe9907f5be952c12f8925c84dff8fb739909d3951c591c9124

Observation 0dc1c50f-f5d6-4f63-afd5-926905ff0386 · outbound

This paper cites VLA-JEPA: Enhancing vision-language-action model with latent world model.arXiv preprint arXiv:2602.10098,.

World Tokens: Enhancing Embodied Policies with Training-Time World Modeling VLA-JEPA: Enhancing vision-language-action model with latent world model.arXiv preprint arXiv:2602.10098,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:55.642353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:55.642353Z digest=sha256:7adb023cb6007a9ccb8bf1e2da84920667191f63b0dd697e2b45c78db34a24b7

Observation a21cfbb2-8175-4df4-9e22-40192f6ccc35 · outbound

This paper cites One Token Per Frame: Reconsidering Visual Bandwidth in World Models for VLA Policy.

World Tokens: Enhancing Embodied Policies with Training-Time World Modeling One Token Per Frame: Reconsidering Visual Bandwidth in World Models for VLA Policy

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:55.646630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:55.646630Z digest=sha256:d5c89185529d698dbc5ba0486c15c97f018f236aad8ed621dea67e5ba39ee8fb

Observation b75afe70-f9f2-464d-80c5-046d6dc53446 · outbound

This paper cites World2Act: Latent Action Post-Training from World Model Dynamics.

World Tokens: Enhancing Embodied Policies with Training-Time World Modeling World2Act: Latent Action Post-Training from World Model Dynamics

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-11T11:52:55.836408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T11:52:55.651035Z digest=sha256:469dc157394370f9f9b98bc092a7246ca1beca100faca602ecc9c5f02695077f

Observation f6c4bcfb-0b1a-4d54-9a97-a942ecadba9a · outbound

This paper cites BridgeData V2: A Dataset for Robot Learning at Scale.

World Tokens: Enhancing Embodied Policies with Training-Time World Modeling BridgeData V2: A Dataset for Robot Learning at Scale

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:55.655345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:55.655345Z digest=sha256:6d0fcb15cdbbcdc623286ef89046eff58acebe6ab057562f59921bfbf3f6756e

Observation 95c30bd7-31d3-4bf8-b858-34080c05bac9 · outbound

This paper cites StarVLA-$\alpha$: Reducing Complexity in Vision-Language-Action Systems.

World Tokens: Enhancing Embodied Policies with Training-Time World Modeling StarVLA-$\alpha$: Reducing Complexity in Vision-Language-Action Systems

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:55.659762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:55.659762Z digest=sha256:7d4930c40647c324c7463a55ed297e8377a3fcde09a9a1998e05e7421f54aeb4

Observation 75e865f1-f1c4-4b87-837a-817bd924df4e · outbound

This paper cites 𝛿VLA:Prior-guidedvision-language-actionmodelsviaworldknowledgevariation.arXiv preprint arXiv:2603.08361,.

World Tokens: Enhancing Embodied Policies with Training-Time World Modeling 𝛿VLA:Prior-guidedvision-language-actionmodelsviaworldknowledgevariation.arXiv preprint arXiv:2603.08361,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:55.664989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:55.664989Z digest=sha256:612a128675d3cd78d17f10cc73c260ea12d3ece2d4553e5852413aa763951b6c

Observation dabfba20-dc68-409a-b791-160f285c5780 · outbound

This paper cites We use AdamW with𝛽= (0.9,0.95) ,𝜖=10 −8, and weight decay10−8, under a cosine schedule with linear warmup.

World Tokens: Enhancing Embodied Policies with Training-Time World Modeling We use AdamW with𝛽= (0.9,0.95) ,𝜖=10 −8, and weight decay10−8, under a cosine schedule with linear warmup

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:52:56.158365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T11:52:55.669408Z digest=sha256:7e37efa9155717eaeeea3ce92325eef55ef0d5263b2497ae145c16b60e6dabd2

Observation 101c7be5-a322-4587-99dc-11055849eb3c · outbound

This paper cites an unresolved cited work.

World Tokens: Enhancing Embodied Policies with Training-Time World Modeling Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-11T11:52:56.142199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T11:52:55.674025Z digest=sha256:5c252ccd806c2c3e70d508593c203276b8731dfa2855101a6ff7f0becaf112f8

Observation e2b5cf19-cc19-471b-90f2-8f11b8fcc6a7 · outbound

This paper cites WorldVLA: Towards Autoregressive Action World Model.

World Tokens: Enhancing Embodied Policies with Training-Time World Modeling WorldVLA: Towards Autoregressive Action World Model

Reference 1986

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:55.613388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:55.613388Z digest=sha256:ef1ec168226b4ef1cbf7a476cb815543748bc58df3e4b4a6daa287d366133c60

Observation 99eccc8d-ad0a-41e6-8d94-65296ad4ed31 · outbound

This paper cites Robotic VLA benefits from joint learning with motion image diffusion.arXiv preprint arXiv:2512.18007,.

World Tokens: Enhancing Embodied Policies with Training-Time World Modeling Robotic VLA benefits from joint learning with motion image diffusion.arXiv preprint arXiv:2512.18007,

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:55.618677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:55.618677Z digest=sha256:ddfe663eaca8b38be0d442ec1d23bfc4f535a019f44efe9c8a9a79fded5a7a8f

Observation 3cfe59c5-a037-4b0e-95d7-ee56f79def5b · outbound

This paper cites Causal World Modeling for Robot Control.

World Tokens: Enhancing Embodied Policies with Training-Time World Modeling Causal World Modeling for Robot Control

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:55.623005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:55.623005Z digest=sha256:88dfff8d927f13565f1a273c8f86a746f26593444ff3919858a2a4def7649e7b

Observation 70551851-8387-412a-b9da-29d0052a4284 · outbound

This paper cites DiT4DiT: Jointly modeling video dynamics and actions for generalizable robot control.

World Tokens: Enhancing Embodied Policies with Training-Time World Modeling DiT4DiT: Jointly modeling video dynamics and actions for generalizable robot control

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:55.628035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:55.628035Z digest=sha256:114af473bcf049f678bb4fe1581298454ad15f44dd03db537de6653617e2a8f5

Observation bd1375fc-bb65-4083-8e81-a160f95fffbd · outbound

This paper cites GR00T N1: An Open Foundation Model for Generalist Humanoid Robots.

World Tokens: Enhancing Embodied Policies with Training-Time World Modeling GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:55.632302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:55.632302Z digest=sha256:069137125ce4549e1f67ae806adb9f9fdd5dac11d1a180377d8820971c783dd7

Pith citing papers

No inbound Pith citation observations are available.