Pith. sign in

Paper Citation Record · LEDGER

HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models

As of 4 August 2026, this Paper Citation Record lists 66 of 66 outbound references and 7 inbound Pith citation observations for arXiv:2605.10942.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.10942 v1

Coverage vector

measured 66 of 66 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-12T03:28:18.730751Z

measured 73 of 73 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-03T06:30:56.289259+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-15T11:04:29.593725Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T12:47:05.504971Z

Reference resolution

66 of 66 outbound references displayed

  • verified exact38
  • verified fuzzy27
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7a78616a-8a50-4b71-b9e0-301ddb72e4d6 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.Advances in Neural Information Processing Systems, 35:23716–23736.

HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models Flamingo: a visual language model for few-shot learning.Advances in Neural Information Processing Systems, 35:23716–23736

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T19:36:49.041453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T03:28:18.730751Z digest=sha256:22a19e7ca3685420a09fd21be04b897484e72d4de3bd13288dbeca4812b4c977

Observation 2d52ce14-1e7a-4ea7-928a-98f2bd4e3fd4 · outbound

This paper cites Qwen3-VL Technical Report.

HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models Qwen3-VL Technical Report

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:21:24.011339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T03:28:18.730751Z digest=sha256:270ee589de1a0049a6194c13c10cf536f99585e20bf9b83804f149fe99d2f9fa

Observation b1e1f89d-d7d3-4f70-a098-bf0644911523 · outbound

This paper cites Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation.

HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:17:01.597427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T03:28:18.730751Z digest=sha256:311cde40a8db4d99acf7535e6439b807f881e200a4b89670dc83bb7f6d5c6b66

Observation 59d926a2-ddc3-4be6-b89d-93f915d96b73 · outbound

This paper cites Motus: A Unified Latent Action World Model.

HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models Motus: A Unified Latent Action World Model

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:44:37.047333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T03:28:18.730751Z digest=sha256:f746a156c1d33daeb7bf45553273a205b2989b4534c8e48a899fce44a7bdcc2d

Observation 58ecaf62-69fd-460e-b81b-5b719f27ec0f · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:21:24.144417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T03:28:18.730751Z digest=sha256:4aebeecc2a5eee894038b5197f4cd263d078caaee881d645ad17443ab72722ea

Observation a671914b-7405-4624-bd33-af65b676b82f · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models RT-1: Robotics Transformer for Real-World Control at Scale

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:21:24.231863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T03:28:18.730751Z digest=sha256:d5302eb3438fa11f24608f92c190ff7c9abc92231c7417be8c1d617a0f8c34ba

Observation 9a1c4ce2-3f16-49d6-92d7-8af53a6e7720 · outbound

This paper cites AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems.

HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:21:24.034211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T03:28:18.730751Z digest=sha256:87b3d662f3fec9ce1f3eb7facc492af972d9d28f9aa2629ea441ca293e7d078d

Observation bc57b6fd-992c-4786-a242-48351ca85ba3 · outbound

This paper cites RynnVLA-002: A Unified Vision-Language-Action and World Model.

HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models RynnVLA-002: A Unified Vision-Language-Action and World Model

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-02T02:03:36.254972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T03:28:18.730751Z digest=sha256:93ef9104abd5e3e23778c4a67520e986eea739dced9a2634a4e7b6e035ee5abc

Observation f85c4c3f-3820-421e-a56d-db42ef4ff9fe · outbound

This paper cites WorldVLA: Towards Autoregressive Action World Model.

HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models WorldVLA: Towards Autoregressive Action World Model

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:21:24.271985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T03:28:18.730751Z digest=sha256:2d931554cdbe3f3c30e476510f3b13e589a500821e12a7952182ac6253d7e9bc

Observation 81873f46-217e-4e51-8bf9-2bae8b4a7f69 · outbound

This paper cites GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation.

HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:21:24.169467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T03:28:18.730751Z digest=sha256:abf1501275045f66be6fd6e7704d775230e8b723fc950da62ffe023300cf72d9

Observation 812085f3-9b33-4416-8906-59e9eaf2051b · outbound

This paper cites Spatialvlm: Endowing vision-language models with spatial reasoning capabilities.

HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models Spatialvlm: Endowing vision-language models with spatial reasoning capabilities

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T19:36:48.943686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T03:28:18.730751Z digest=sha256:0d8502cf22a6b8ce1a2d69322f6dba3d38d507dfac21f40c3c57da48a3460363

Observation f7308767-1780-4984-b5d1-f43cc746db46 · outbound

This paper cites Fast-in-Slow: A Dual-System Foundation Model Unifying Fast Manipulation within Slow Reasoning.

HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models Fast-in-Slow: A Dual-System Foundation Model Unifying Fast Manipulation within Slow Reasoning

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:21:24.290792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T03:28:18.730751Z digest=sha256:f793708ead4cd9abf0e06e7012b92cea58dfaed62a9dbffec805aee970fe699e

Observation 0270b260-09e3-413a-9ba3-b4a6db734e92 · outbound

This paper cites Wow: Towards a world omniscient world model through embodied interaction.

HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models Wow: Towards a world omniscient world model through embodied interaction

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:21:24.109420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T03:28:18.730751Z digest=sha256:2feddaed3d45af0ac6ad4e1029a45c48a840762bb0eb87cf02e1b7a92c2e1b2a

Observation 32c2c10c-d6d3-4a31-b45e-227de2f88070 · outbound

This paper cites Lightewm: Light embodied world model.

HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models Lightewm: Light embodied world model

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T19:36:48.938116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T03:28:18.730751Z digest=sha256:37576f8b5016bbc6769018d1c352f13fbb7552e5ef3b6d1707ed4c9be0be8d66

Observation c7f680d8-8427-4931-af30-a394a60f2ce3 · outbound

This paper cites Learning universal policies via text-guided video generation.Advances in neural information processing systems, 36:9156–9172.

HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models Learning universal policies via text-guided video generation.Advances in neural information processing systems, 36:9156–9172

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T19:36:49.067151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T03:28:18.730751Z digest=sha256:8423506f412be13bde8cb79092e2541df98d1ab7d37596165a3eb2a36b396466

Observation d3e9c807-bf4c-4659-b8cf-c1fb6090487a · outbound

This paper cites Vidar: Embodied Video Diffusion Model for Generalist Manipulation.

HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models Vidar: Embodied Video Diffusion Model for Generalist Manipulation

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:54:28.431425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T03:28:18.730751Z digest=sha256:00c5733416d19b69d78f046dca43b40849fff56d1bf6773491a76344893bede1

Observation f86d8ef4-4c46-4be0-b242-731255207f25 · outbound

This paper cites Manualvla: A unified vla model for chain-of-thought manual generation and robotic manipulation.

HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models Manualvla: A unified vla model for chain-of-thought manual generation and robotic manipulation

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:21:24.257010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T03:28:18.730751Z digest=sha256:ff19b58f5231952c502f7dff5c6340a0d140bc745c48410ed2e6f3abd90ff4a9

Observation 26af7d2e-a15c-4337-b056-fe8d4c5c1d2f · outbound

This paper cites Video prediction policy: A generalist robot policy with predictive visual representations.

HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models Video prediction policy: A generalist robot policy with predictive visual representations

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T19:36:49.026579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T03:28:18.730751Z digest=sha256:a58e3e5076844883bd3c989d8268a5807b7d2a9ba6ed1139844ec9b1373ca8ba

Observation f2eebd49-ca7a-4385-b735-ea6654994c30 · outbound

This paper cites ${\pi}_{0.7}$: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities.

HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models ${\pi}_{0.7}$: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:21:24.085346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T03:28:18.730751Z digest=sha256:8e8931897ff4f0f2752cb4ef61d0edce114be9adedf5c1434ce751a97a067f88

Observation 8fc11e20-a9da-4943-9955-8f5e11447eeb · outbound

This paper cites π0.5: a vision-language-action model with open-world generalization.

HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models π0.5: a vision-language-action model with open-world generalization

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T19:36:49.032334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T03:28:18.730751Z digest=sha256:19b38ae24fcba33a88ba5ad0726499ac0384993d1fda9d457874c2fe6706093a

Observation a65a6360-47b1-44ef-b929-9e1f4ddd3525 · outbound

This paper cites Dreamgen: Unlocking generalization in robot learning through video world models.

HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models Dreamgen: Unlocking generalization in robot learning through video world models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T19:36:49.036984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T03:28:18.730751Z digest=sha256:601b5337abdd9d61885b85df64a9f90ce722abf753315482255b5f66312f25af

Observation 9534572a-51e2-4204-a582-d8c864b30921 · outbound

This paper cites Video2act: A dual-system video diffusion policy with robotic spatio-motional modeling.

HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models Video2act: A dual-system video diffusion policy with robotic spatio-motional modeling

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T19:36:49.011287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T03:28:18.730751Z digest=sha256:bd646f2491107a43d30d7afd21e9c81ec2b32acd75040d1a81067a326eda8756

Observation 46f0efc8-a03e-4377-9661-218729592cbf · outbound

This paper cites Prismatic vlms: Investigating the design space of visually-conditioned language models.

HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models Prismatic vlms: Investigating the design space of visually-conditioned language models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T19:36:49.016248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T03:28:18.730751Z digest=sha256:49fe7e5e0d176505d1edc62682db1619eb65b03dc1cbedc93dd7484a53c0287b

Observation e93507b2-c079-423b-9062-44f770e58242 · outbound

This paper cites DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset.

HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:21:24.261599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T03:28:18.730751Z digest=sha256:f2e799ffda2c0e23186ba2c45c3a983c99968dd3bf7e543de90e0108a70d8001

Observation 88f4f6b9-9dff-4237-8768-663e2041dee1 · outbound

This paper cites Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success.

HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:21:24.277136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T03:28:18.730751Z digest=sha256:3fd3c5af1af11bbbdc4e81d273ad469e71e1db4801999a6e24e26bd26b0eb29f

Observation 63c70900-4d56-4002-914b-e54e1959a94b · outbound

This paper cites Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning.

HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-12T14:50:13.054714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T03:28:18.730751Z digest=sha256:36f4bba49968c360b3b4d71b6d7a880f9870683cbaf2b70b5544a312299b3faa

Observation f25e8460-3d19-4ca4-844e-93b679213c42 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models OpenVLA: An Open-Source Vision-Language-Action Model

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:21:24.026208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T03:28:18.730751Z digest=sha256:526994cb7d8c7416aaa813cd486274f61c493c572c37466c45570d942f73bd35

Observation 22bb5909-ba3d-466e-a224-a431777dda70 · outbound

This paper cites Causal World Modeling for Robot Control.

HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models Causal World Modeling for Robot Control

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-12T13:53:52.791726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T03:28:18.730751Z digest=sha256:eae2e6c07eda6eefb3204054c322bc3265775a5af19c5dd0301c173bcb708a59

Observation 0f748f54-4a10-4a34-b3cb-b7e1a97ddc5e · outbound

This paper cites CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation.

HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:21:24.239747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T03:28:18.730751Z digest=sha256:9a1a50c194c08aa9b9a656a5096e62f88e440dee0c24fd19710d050dac62a7ef

Observation bc87e695-01d9-4559-9bec-5b31adba7018 · outbound

This paper cites Unified video action model.

HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models Unified video action model

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T19:36:49.005730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T03:28:18.730751Z digest=sha256:9409993be48343d2cf94950328b9280a27e758fb12ad0deaa12ce9f4663c63c6

Observation 6e0139f9-82cf-4edd-ab91-21cb44394166 · outbound

This paper cites Manipllm: Embodied multimodal large language model for object-centric robotic manipulation.

HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models Manipllm: Embodied multimodal large language model for object-centric robotic manipulation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T19:36:49.061347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T03:28:18.730751Z digest=sha256:9dcd2f37d7adf63ae108d105df6b0171efd2a9623fc612b04aa87236ffdd1320

Observation 132fbe1c-5663-470f-94b7-0dd526ea9c99 · outbound

This paper cites Video generators are robot policies.

HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models Video generators are robot policies

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T19:36:49.055737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T03:28:18.730751Z digest=sha256:ea158889eda8a62c3bf98dc617532f2339e111e5a0a8775ea4d47325a3d085c5

Observation eaa14da8-c3c6-4ea9-896f-ff1c91baa5db · outbound

This paper cites Genie envisioner: A unified world foundation platform for robotic manipulation.

HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models Genie envisioner: A unified world foundation platform for robotic manipulation

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T19:36:48.996746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T03:28:18.730751Z digest=sha256:04ad54cbedf5d2198191c4419d40933297bc83fa8599d99672c521ae2a6f989b

Observation 6228396d-8c21-405b-92b6-9ff98a26ed8c · outbound

This paper cites Onetwovla: A unified vision-language-action model with adaptive reasoning.

HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models Onetwovla: A unified vision-language-action model with adaptive reasoning

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:21:24.299110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T03:28:18.730751Z digest=sha256:b4e2d0b3cd8eda061ce173c0b94423c16cedd6e0ef7ce356957e6d039cd57d92

Observation f1ada0cc-a523-4315-807a-3941bc6871cb · outbound

This paper cites HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model.

HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:00:49.507028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T03:28:18.730751Z digest=sha256:c9399520b629af81224d5f692461fd640c8d8a57c37be29787467041d558f174

Observation ff7024c0-ef4e-4e5b-97b1-628e585df75f · outbound

This paper cites RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation.

HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:21:24.063345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T03:28:18.730751Z digest=sha256:d700dddda719994100fe1a07daee974becee6cea897582716c69e137b0405138

Observation e2a9b8af-7fad-4ef0-be8d-32cc27d017c2 · outbound

This paper cites Last {0}: Latent spatio-temporal chain-of- thought for robotic vision-language-action model.

HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models Last {0}: Latent spatio-temporal chain-of- thought for robotic vision-language-action model

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:21:24.018915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T03:28:18.730751Z digest=sha256:56a4f353475e234084051b778b257e060414941772aa0204eab11cce4bbec8b3

Observation 379316af-aa50-43b0-8e7a-217c01a96006 · outbound

This paper cites Tc-idm: Grounding video generation for executable zero-shot robot motion.

HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models Tc-idm: Grounding video generation for executable zero-shot robot motion

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:21:24.223249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T03:28:18.730751Z digest=sha256:ed536169207a22dca30a55a7678d908e08f3d388f8fca2b2c60b7cbd1942af07

Observation 86f621a1-888b-4ff4-a6d4-01a0fd4c4881 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models DINOv2: Learning Robust Visual Features without Supervision

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:21:24.174165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T03:28:18.730751Z digest=sha256:2b8cf97441e7681ba4ec7c3ef0b926c863438090c653d7e1be11b08870a17191

Observation b259c3ba-5090-4645-adb2-dea4611b6150 · outbound

This paper cites mimic-video: Video-Action Models for Generalizable Robot Control Beyond VLAs.

HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models mimic-video: Video-Action Models for Generalizable Robot Control Beyond VLAs

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-15T10:41:00.438292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T03:28:18.730751Z digest=sha256:a0831e4b4528dd58c3209fbb6c0830456a46f39e0b20808ab9511a29e0e8556a

Observation 6413e0b0-1e2a-4b6b-8ef9-fc3217469303 · outbound

This paper cites FAST: Efficient Action Tokenization for Vision-Language-Action Models.

HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:21:24.215121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T03:28:18.730751Z digest=sha256:10e64cd8b418acbaa0be016365cd18b15cca02ff8293c7b1d5f6601fcc6afcdb

Observation 252def1a-8466-48c4-ae65-35434ed142ac · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.Journal of machine learning research, 21(140):1–67.

HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models Exploring the limits of transfer learning with a unified text-to-text transformer.Journal of machine learning research, 21(140):1–67

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T19:36:49.001463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T03:28:18.730751Z digest=sha256:35ea8768a0a72c5064fa4782ac2a3cd65ac30979dc64989a8193fc49ce19d41e

Observation 1737f6ec-d680-4f29-833e-b426c0b25273 · outbound

This paper cites Perceiver-actor: A multi-task transformer for robotic manipulation.

HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models Perceiver-actor: A multi-task transformer for robotic manipulation

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T19:36:49.021630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T03:28:18.730751Z digest=sha256:235c30b3c1a83df5399c77ee896fc887946787bbcd685d6fc0a5915a4cbd3e20

Observation e2e91c98-d5c3-4fc5-905f-6062760e7578 · outbound

This paper cites AnyPos: Automated Task-Agnostic Actions for Bimanual Manipulation.

HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models AnyPos: Automated Task-Agnostic Actions for Bimanual Manipulation

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:21:23.995672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T03:28:18.730751Z digest=sha256:de85a61e64ef9740669fb5949177d2b914b6ac6dc39cfc5fe9bcc114c1de9df5

Observation c772bb04-cb80-4414-ac4e-e587a0ac5b50 · outbound

This paper cites Octo: An Open-Source Generalist Robot Policy.

HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models Octo: An Open-Source Generalist Robot Policy

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:21:24.307773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T03:28:18.730751Z digest=sha256:cc2b8238e3e9b50972357bdfd3ae205f4f8de00eb223bac98614cd195427afac

Observation f5b50ab9-9f3e-4d29-a9ca-16ca75ff7421 · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models Wan: Open and Advanced Large-Scale Video Generative Models

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:21:24.283470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T03:28:18.730751Z digest=sha256:68df6e89cc558ef857d73ee4891c87d4bab1f85eeed3e124d292b9ef060d1e37

Observation 975ceea2-8c42-4770-8a5e-878814d0c7ff · outbound

This paper cites Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning.

HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:21:24.315613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T03:28:18.730751Z digest=sha256:fb3d6d7cba68dc5d02cc64b07cba7e7ce6e5f6b070166a2a1fa8c8f811806461

Observation 474f2a2f-8dac-4e48-b495-6b2fe069137a · outbound

This paper cites Tinyvla: Towards fast, data-efficient vision-language-action models for robotic manipulation.IEEE Robotics and Automation Letters.

HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models Tinyvla: Towards fast, data-efficient vision-language-action models for robotic manipulation.IEEE Robotics and Automation Letters

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T19:36:48.948400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T03:28:18.730751Z digest=sha256:2bd6a41e9f1fd4728ca69ef1556cf823effa40d74931ce2fae163221af00e558

Observation c7321da2-609e-46e7-91df-c5fbc8ddaa54 · outbound

This paper cites Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model.

HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-29T02:04:58.522203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T03:28:18.730751Z digest=sha256:0b1140d7923696ba3bed4ed1d8bee6ec637a25c1db57a028d1022200cadeff44

Observation bea135fc-734f-4da2-ab08-090489c27de8 · outbound

This paper cites Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation.

HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-13T16:32:05.974462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T03:28:18.730751Z digest=sha256:7d1cd4a54580f6756323eb30354f7a94202e532df62d440c2426ee7b7303d828

Observation 68307ac7-d89b-4093-b2d2-39759037960a · outbound

This paper cites RoboMIND: Benchmark on Multi-embodiment Intelligence Normative Data for Robot Manipulation.

HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models RoboMIND: Benchmark on Multi-embodiment Intelligence Normative Data for Robot Manipulation

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:14:18.432154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T03:28:18.730751Z digest=sha256:62721e10e9f2794570ef12f1a23becf394a2544d498d286ee1acbcba3795f959

Observation 84d21b0c-ec04-49f7-8f7d-0bdbe02092cb · outbound

This paper cites arXiv preprint arXiv:2603.17240 , year=.

HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models arXiv preprint arXiv:2603.17240 , year=

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:21:24.248123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T03:28:18.730751Z digest=sha256:15911f183a5dd9a2a89388808648c700026bffcefdd0bdae6293fb5d1e4e096f

Observation 08983b88-2532-477b-9d28-63ee7e310093 · outbound

This paper cites World action models are zero-shot policies.

HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models World action models are zero-shot policies

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T19:36:48.982161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T03:28:18.730751Z digest=sha256:18914401032c6d5d44db9889089f35df21b90b9ebaef85ed111cb151e6c0074d

Observation a93511a7-520d-4c31-b425-839c1c891dc4 · outbound

This paper cites Fast-wam: Do world action models need test-time future imagination?.

HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models Fast-wam: Do world action models need test-time future imagination?

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T19:36:49.051193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T03:28:18.730751Z digest=sha256:62495dd477a537cc9aa19b182cdace25f5872d89defee264745e6f3586b63dd2

Observation 092511af-6622-4857-a400-4fa8e7e5bc38 · outbound

This paper cites Sigmoid loss for language image pre-training.

HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models Sigmoid loss for language image pre-training

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T19:36:48.972716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T03:28:18.730751Z digest=sha256:3d0796b57874028b506ce1c6f5fe6fad486727a08d312f8022849c3de5631b55

Observation 8253379b-d0ba-4646-93da-53dca8e7c747 · outbound

This paper cites FLARE: Robot Learning with Implicit World Modeling.

HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models FLARE: Robot Learning with Implicit World Modeling

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-17T15:59:09.111136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T03:28:18.730751Z digest=sha256:402564ce07cb11f4c79f9cc8fca7cec2f918c143fd2f4474dbf37ebd54cdf434

Observation d75f1b64-a43c-481c-9299-72dbc2b2f529 · outbound

This paper cites Act2goal: From world model to general goal-conditioned policy.

HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models Act2goal: From world model to general goal-conditioned policy

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:21:24.046062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T03:28:18.730751Z digest=sha256:9682ba1a7eda3eefef2398e04f95be893ae1cf81a5f7a4b097245219f52535cb

Observation 42b9092e-4591-4fb4-97ea-b2295fcd4974 · outbound

This paper cites RoboDreamer: Learning Compositional World Models for Robot Imagination.

HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models RoboDreamer: Learning Compositional World Models for Robot Imagination

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:47:30.400020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T03:28:18.730751Z digest=sha256:60f9945b86e90b658005fe59bba56a757f05a6db9fb5e9ffcb4750d143294baf

Observation 590cdcf1-ad3b-4f6c-a3a3-aa452731ad23 · outbound

This paper cites Unified world models: Coupling video and action diffusion for pretraining on large robotic datasets.

HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models Unified world models: Coupling video and action diffusion for pretraining on large robotic datasets

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T19:36:48.977599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T03:28:18.730751Z digest=sha256:4b227479c38d5637b5609aa5e6d13a40cf674f14966780575511eb94da53f033

Observation a63ee307-0755-4b94-82f7-652465223746 · outbound

This paper cites Rt-2: Vision-language-action models transfer web knowledge to robotic control.

HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models Rt-2: Vision-language-action models transfer web knowledge to robotic control

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T19:36:48.987396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T03:28:18.730751Z digest=sha256:f95983374c64e548033665dcfd8a77aa42f41ff898c96ce9acaed4a3be8aa937

Observation 95a20d0d-dba4-4390-a9ef-670968298c56 · outbound

This paper cites S1: grasp and place banana; S2: grasp and place carrot.

HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models S1: grasp and place banana; S2: grasp and place carrot

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T19:36:48.992159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T03:28:18.730751Z digest=sha256:d76e85f08c0d66dc485488cf24f14cf16f63d63c17345d55ade05f475ec47858

Observation eda433fc-8e92-42f0-a7b2-0e0ff677063f · outbound

This paper cites S1: place the second can beside the first; S2: place the third can on top.

HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models S1: place the second can beside the first; S2: place the third can on top

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T19:36:49.046436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T03:28:18.730751Z digest=sha256:3615eb9032fa0345d1f4cf6a63a7243ac9dce6ead91e06541c737e60349560b1

Observation 34a7fbad-6712-45bd-a66f-3840f532bb9e · outbound

This paper cites S1: grasp bottle; S2: pour into beaker.

HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models S1: grasp bottle; S2: pour into beaker

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T19:36:48.967927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T03:28:18.730751Z digest=sha256:c079864a086951139e29fe1af6a2f65c8c94fcb7ba260813552448a5b546877a

Observation 5af47688-cf5f-4073-9508-2cdb082820ae · outbound

This paper cites Yes”.The robot picks up a marker and writes “Y.

HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models Yes”.The robot picks up a marker and writes “Y

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T19:36:48.953887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T03:28:18.730751Z digest=sha256:812f0050cd499b0cf1ec251f45c75377e43a32a5ee00a2c29719b2b3d499c3a8

Observation ade77d6a-1f2a-4606-9db5-2b009ad35dde · outbound

This paper cites S1: pick flower; S2: bimanual handover; S3: insert into vase.

HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models S1: pick flower; S2: bimanual handover; S3: insert into vase

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T19:36:48.958974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T03:28:18.730751Z digest=sha256:6bcd3a76ce95bd9ff6ad9c1aa5872b5806e46098ab2e8a61460d8f150e700ee4

Observation 3f66ea93-702b-43e3-b452-f1ce583e6912 · outbound

This paper cites S1 →S 2: pick up item and place into bag; S3 →S 4 →S 5: one arm grips the bag to hold it steady, the other grips and pulls the zipper to close.

HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models S1 →S 2: pick up item and place into bag; S3 →S 4 →S 5: one arm grips the bag to hold it steady, the other grips and pulls the zipper to close

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T19:36:48.963721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T03:28:18.730751Z digest=sha256:23467402ce184b9b492ccc948aa4ce4d67832ebc1387d4f933d3f23e7e3b905f

Pith citing papers

Observation 714f4102-baea-4211-b3c5-74e6371c648e · inbound

SANTS: A State-Adaptive Scheduler for World Action Models cites this paper.

SANTS: A State-Adaptive Scheduler for World Action Models HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-06-29T12:03:24.059813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-29T11:57:36.200662Z digest=sha256:4d07762cfe68569d81abbe585ad708ea0a2fb988d795ac22cacfd5653351287c

Observation e29e8131-fd96-4c5e-95ac-7846f1ec905b · inbound

SANTS: A State-Adaptive Scheduler for World Action Models cites this paper.

SANTS: A State-Adaptive Scheduler for World Action Models HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-15T11:04:29.593725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T11:04:29.593725Z digest=sha256:c14e0a614b66fd515b10eb8a638b59f806208e800235fb6cbbb6d84b4adc00ad

Observation 7afae3ed-3b57-435d-a119-a9676f480ff1 · inbound

RepWAM: World Action Modeling with Representation Visual-Action Tokenizers cites this paper.

RepWAM: World Action Modeling with Representation Visual-Action Tokenizers HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-03T14:58:33.448356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-27T06:47:05.236028Z digest=sha256:5c654afd08a7ebe2e48850dd7a0f06bc3154d41febcec4da032babe007928f7b

Observation 0c35a945-8d13-4c05-ac41-a0a0dd632641 · inbound

World Action Models: A Survey cites this paper.

World Action Models: A Survey HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-07-04T04:09:35.556529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-26T17:11:12.686936Z digest=sha256:56a3871d5dadb4009e174ee93d69dab6aa7a16a75faf66f13bf14d9986a5940c

Observation 5c4fab1a-8822-4d4d-8f9e-dbc9cd90cbf8 · inbound

TACO: TActile World Model as a Self-COrrector forScalable VLA Post-Training cites this paper.

TACO: TActile World Model as a Self-COrrector forScalable VLA Post-Training HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-12T06:41:57.276146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:41:57.276146Z digest=sha256:3cc67d3f28eee8aa76ffda9333e715d8f4aa33e8cac9a12895e0a2bd2d49622d

Observation 984341a5-1e83-49d5-a77c-e7c2f02d8c17 · inbound

HALO-WA: Hybrid-Attention Latent-Guided Online Reinforcement Learning for World-Action Models cites this paper.

HALO-WA: Hybrid-Attention Latent-Guided Online Reinforcement Learning for World-Action Models HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-11T20:32:22.412216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:32:22.412216Z digest=sha256:afdea9ad2aef86d0b62ca0294e61071beb7df7273284ccf46cc0c1a9021f619d

Observation 119570ee-d95f-4a2c-8c80-a09a675dd0c9 · inbound

Understanding and Mitigating the Video-Action Generalization Gap via Temporal Ratio cites this paper.

Understanding and Mitigating the Video-Action Generalization Gap via Temporal Ratio HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-07-10T12:47:05.506373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-07-10T12:39:02.545780Z digest=sha256:d1a15654ec5dc2306e8379c40a55779e3dcaa563b3f91bba14d10188ff9e4967