Pith. sign in

Paper Citation Record Β· LEDGER

WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning

As of 5 August 2026, this Paper Citation Record lists 66 of 66 outbound references and 0 inbound Pith citation observations for arXiv:2607.29613.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.29613 v1

Coverage vector

measured 66 of 66 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T03:34:07.067380Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

66 of 66 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved66
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7fa90ab6-857f-4b8d-a7d3-59b97102d7d5 Β· outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T03:34:03.242685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:34:03.242685Z digest=sha256:3356d8a33bccf797d7c153358891cf9dd8fda2dfb28461e5580cd2e125efac97

Observation f0f35e92-5023-4295-bcfa-3bf76c62f8a2 Β· outbound

This paper cites πœ‹0.5: Avision-language-actionmodelwithopen-worldgeneralization.

WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning πœ‹0.5: Avision-language-actionmodelwithopen-worldgeneralization

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T03:34:03.288984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:34:03.288984Z digest=sha256:443a7545ac7246b0aff096e81de50d1c7502214d90a4b148cbd26ab78efbfc15

Observation 83ab3ffb-07f9-4034-bfe5-34ba90136c3c Β· outbound

This paper cites ${\pi}_{0.7}$: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities.

WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning ${\pi}_{0.7}$: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T03:34:03.301615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:34:03.301615Z digest=sha256:4522a79a5439abff7ff83b7eb20ae2e91d352164df2240b0052d678841d6c872

Observation bd1eba51-00ff-4a53-aff5-652d4dd5198c Β· outbound

This paper cites Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success.

WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T03:34:03.333897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:34:03.333897Z digest=sha256:a0b11ffd197e8666169e57f995de18b470be1494646ab2aff97421f06aa7520f

Observation 343dbda9-a696-4e7f-93d9-645eeaf20350 Β· outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning OpenVLA: An Open-Source Vision-Language-Action Model

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T03:34:03.377367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:34:03.377367Z digest=sha256:c8808ecf2808870826bb963240ce84170092933df6409435165b77c97c67feec

Observation 1a4001d8-33fd-4efa-985f-ccd29084a647 Β· outbound

This paper cites What can rl bring to vla generalization? an empirical study.arXiv preprint arXiv:2505.19789, 2025.

WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning What can rl bring to vla generalization? an empirical study.arXiv preprint arXiv:2505.19789, 2025

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T03:34:03.430509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:34:03.430509Z digest=sha256:d96678734b526e4dcd2c0d6d7d8ad746d67789b4266383943aca354af26a437e

Observation d08c0812-ab4b-4173-aef3-12c0d56b7712 Β· outbound

This paper cites Srpo: Self-referential policy optimization for vision-language-action models.arXiv preprint arXiv:2511.15605, 2025.

WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning Srpo: Self-referential policy optimization for vision-language-action models.arXiv preprint arXiv:2511.15605, 2025

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T03:34:03.462501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:34:03.462501Z digest=sha256:313f69e2a22c940d98569094051ec8180725422f9e5bfff9be37b3006da116b8

Observation 6eca5fab-20cf-4c76-a10f-91fa5b50df44 Β· outbound

This paper cites SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning.

WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T03:34:03.478972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:34:03.478972Z digest=sha256:b1a883668940add4ed79db254847fc5218a77e4928720400e5d09f3de9706e36

Observation 73cdb8fc-c7af-487f-8777-ccc53fcbe341 Β· outbound

This paper cites Rlinf-vla: A unified and efficient framework for vla+ rl training.arXiv preprint arXiv:2510.06710, 2025.

WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning Rlinf-vla: A unified and efficient framework for vla+ rl training.arXiv preprint arXiv:2510.06710, 2025

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T03:34:03.506647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:34:03.506647Z digest=sha256:074a3fdaee9bbefaa86f24c9a4f63766294786c7daf5da88325c096196da56b9

Observation b647665a-7334-4173-b611-d5ef02bce27f Β· outbound

This paper cites $\pi^{*}_{0.6}$: a VLA That Learns From Experience.

WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T03:34:03.544151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:34:03.544151Z digest=sha256:d6f09a25fc4d42122f44aef3a09ab814dee4de2b510940a57aba2e1f60d2ddc1

Observation 242eec4c-a43e-4b75-a5b1-bfd812172088 Β· outbound

This paper cites an unresolved cited work.

WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T03:34:03.573555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:34:03.573555Z digest=sha256:a5cf9b720f80071669209ef7513be5021a9ff2c83a07124c98c22fb1a73701c6

Observation b83d68a0-1171-4bd1-86d9-1894af254b05 Β· outbound

This paper cites Rlinf-user: A unified and extensible system for real-world online policy learning in embodied ai.

WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning Rlinf-user: A unified and extensible system for real-world online policy learning in embodied ai

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T03:34:03.622824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:34:03.622824Z digest=sha256:692c64bacfbd3e4ae4adb6f1a69934a1757762d57fefe322413aec76e60f7007

Observation df350158-23c2-4575-b6b6-4ea84dd51391 Β· outbound

This paper cites Predictive representations of state.Advances in neural information processing systems, 14, 2001.

WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning Predictive representations of state.Advances in neural information processing systems, 14, 2001

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T03:34:03.653698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:34:03.653698Z digest=sha256:706e8678f650f54d5aa0fbd13d6ce7c88305868286f8ce01e03ad6a52895cac5

Observation 53cfd8ac-add9-43aa-8fa0-aa4a4fe6a9e1 Β· outbound

This paper cites Learning predictive state representations.

WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning Learning predictive state representations

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T03:34:03.671286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:34:03.671286Z digest=sha256:f3759e3d56a0ab81c5dd3e7257e0d56dfa646a4bb0a50eb9d93e1b0561583a4a

Observation bdb30716-95e7-4ae3-b169-de0b687edc1d Β· outbound

This paper cites Predictive State Representations: A New Theory for Modeling Dynamical Systems.

WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning Predictive State Representations: A New Theory for Modeling Dynamical Systems

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T03:34:03.701039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:34:03.701039Z digest=sha256:c27e78d79293544be9e02e5286bbbde7d11f6c1044dff38ff05290e773f8c974

Observation 1ebfbe89-9d2b-40b7-acaa-08e4cdbe4917 Β· outbound

This paper cites When is partially observable reinforcement learning not scary? InConference on Learning Theory, pages 5175–5220.

WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning When is partially observable reinforcement learning not scary? InConference on Learning Theory, pages 5175–5220

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T03:34:03.774921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:34:03.774921Z digest=sha256:83515a2717df2d6c743d29e357f70626eba015878d758c0b4b83cef1f32bfbd0

Observation 6106fe5c-1350-4222-b9dd-df70d277639f Β· outbound

This paper cites Approximate information state for approximateplanningandreinforcementlearninginpartiallyobservedsystems.JournalofMachineLearningResearch, 23(12):1–83, 2022.

WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning Approximate information state for approximateplanningandreinforcementlearninginpartiallyobservedsystems.JournalofMachineLearningResearch, 23(12):1–83, 2022

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T03:34:03.790811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:34:03.790811Z digest=sha256:181cc36b2d07f0ff45258d6f1e07eb21062d19014141500493924da9c59ce715

Observation 123f041f-05e5-42a0-9235-954ec290d52b Β· outbound

This paper cites VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning.

WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T03:34:03.842090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:34:03.842090Z digest=sha256:7e2f8e134f0a6c9d0cbb49914c404904769aa4ff32ca4051c6f1a5dddea01693

Observation 8a9ee29c-22c2-413d-90fd-6ce154bb1a8b Β· outbound

This paper cites Reinforcement learning with latent flow.Advances in Neural Information Processing Systems, 34:22171–22183, 2021.

WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning Reinforcement learning with latent flow.Advances in Neural Information Processing Systems, 34:22171–22183, 2021

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T03:34:03.872169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:34:03.872169Z digest=sha256:986345915cbb2539054277d395d800efcbdc3cadc77d60e62d4d3b4acab87ba1

Observation 8c54494c-aa1f-4e5f-bf04-52f172b5b90d Β· outbound

This paper cites Provable reinforcement learning with a short-term memory.

WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning Provable reinforcement learning with a short-term memory

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T03:34:03.885055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:34:03.885055Z digest=sha256:d6cf5e57f1f5feba28cd25850911781880a31c89acbdca6384886612ce56e0a2

Observation 19972567-022c-4080-8240-65516e2fcbbc Β· outbound

This paper cites Improving sample efficiencyinmodel-freereinforcementlearningfromimages.

WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning Improving sample efficiencyinmodel-freereinforcementlearningfromimages

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T03:34:03.919296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:34:03.919296Z digest=sha256:832444d40d888dbbd3609f5c8fd0136559f1910730829aa16d9afadc77ad91de

Observation 4e14a391-0959-47d7-9eab-7f60f8093f18 Β· outbound

This paper cites Weakly supervised representation learning with sparse perturbations.Advances in Neural Information Processing Systems, 35:15516–15528, 2022.

WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning Weakly supervised representation learning with sparse perturbations.Advances in Neural Information Processing Systems, 35:15516–15528, 2022

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T03:34:03.948783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:34:03.948783Z digest=sha256:c1d9e92f86ad4b0f30a28ee9dbfc4b4274f4b6daf24c03e98034555911f73b49

Observation f2ecac8e-9a05-4392-a829-95c0f40aa3da Β· outbound

This paper cites GPT-4 Technical Report.

WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning GPT-4 Technical Report

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T03:34:03.975909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:34:03.975909Z digest=sha256:9d0b641e4081df36a910cec846a0daba9eec65b74990d1048fbba67a897df67c

Observation 02b0f225-4330-464f-90c6-73748e5f0158 Β· outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning Gemini: A Family of Highly Capable Multimodal Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T03:34:04.050849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:34:04.050849Z digest=sha256:47f80152b63bd1c6296df719de22d99630052aec34de24603cf25a70e5df898d

Observation 87d8efb6-e8b5-418f-9a6b-46290389fc05 Β· outbound

This paper cites Data-efficient reinforcement learning with self-predictive representations.International Conference on Learning Representations, 2020.

WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning Data-efficient reinforcement learning with self-predictive representations.International Conference on Learning Representations, 2020

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T03:34:04.075614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:34:04.075614Z digest=sha256:09654673214b8b5ec4a74816b3f972d6ceabb3c4b272e804c7eeb3d054d8ec76

Observation 7003de70-2833-4330-8149-940cfaadb01c Β· outbound

This paper cites PaLM-E: An Embodied Multimodal Language Model.

WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning PaLM-E: An Embodied Multimodal Language Model

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T03:34:04.107819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:34:04.107819Z digest=sha256:d376e99e60163b0902d85f4e2441dffbac202e7a7efe81b12866e4849645dc62

Observation 88171681-e18f-4375-9572-f08ab7c8fa8d Β· outbound

This paper cites Rt-2: Vision-language-action models transfer web knowledge to robotic control.

WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning Rt-2: Vision-language-action models transfer web knowledge to robotic control

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T03:34:04.139042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:34:04.139042Z digest=sha256:5a7b32f76a245bbf09186531a966bdadc8800cfda3e59b1cb6b83ee77926cfd7

Observation 708ea319-aba6-410d-8201-79efc96041d7 Β· outbound

This paper cites FAST: Efficient Action Tokenization for Vision-Language-Action Models.

WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T03:34:04.174749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:34:04.174749Z digest=sha256:8abc49ae228d6fd598d1fdcc9803725fceebcb920c9a83fcf38649c8e6727b65

Observation 592d2881-d740-4607-8194-2e85d338ba83 Β· outbound

This paper cites Beyond Imitation: Reinforcement Learning-Based Sim-Real Co-Training for VLA Models.

WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning Beyond Imitation: Reinforcement Learning-Based Sim-Real Co-Training for VLA Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T03:34:04.267745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:34:04.267745Z digest=sha256:9be1495dd40964282d9935abe6b243fcab38ee93ab61f99e614bfd00f6e1630b

Observation 822de1b3-429a-4177-ad8f-3c28a816cdcb Β· outbound

This paper cites Interactive Post-Training for Vision-Language-Action Models.

WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning Interactive Post-Training for Vision-Language-Action Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T03:34:04.350940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:34:04.350940Z digest=sha256:2427eef2973e79995fb586dcd575e6d952dfec3f58160c732d0be42dade35a36

Observation 853ca24b-1724-4fcd-b5b7-e97e81336a3a Β· outbound

This paper cites Proximal Policy Optimization Algorithms.

WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T03:34:04.400722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:34:04.400722Z digest=sha256:b8836efd14cf8c95c724297cef25726d79238e1d1a511dc70de132735aafecaf

Observation 1b2310d6-92ae-40b1-b520-4c407fb0f638 Β· outbound

This paper cites DeepSeek-V3 Technical Report.

WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning DeepSeek-V3 Technical Report

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T03:34:04.436544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:34:04.436544Z digest=sha256:08f4ce201142b5868a52df2be4723794d5c8db251bb53be0e289c747eb3f709b

Observation 44dfb27f-1f49-4687-b600-a0f442de68c8 Β· outbound

This paper cites Soft Actor-Critic Algorithms and Applications.

WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning Soft Actor-Critic Algorithms and Applications

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T03:34:04.492471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:34:04.492471Z digest=sha256:3ca301a132c0c1d1e29cd6619e6ebcf16d4cbce770601e9bc366debc6ef2f397

Observation 57438874-8622-461f-8331-e7470c642835 Β· outbound

This paper cites Offline Reinforcement Learning with Implicit Q-Learning.

WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning Offline Reinforcement Learning with Implicit Q-Learning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T03:34:04.603280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:34:04.603280Z digest=sha256:4e84e655e25db57c5acdaceac575364aba1fc941aff35f631e1c60f4a4bd791a

Observation 8aab59ed-a513-42f3-bd7a-1a6b6c3595f0 Β· outbound

This paper cites Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning.

WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-03T03:34:04.663606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:34:04.663606Z digest=sha256:241bf98a24dc3373b002273ff9f9efe131deb0505a46d5815374fc290f53f02f

Observation a11aba10-313d-4f63-905d-e8f4b6034c33 Β· outbound

This paper cites Reinflow: Fine-tuning flow matching policy with online reinforcement learning.arXiv preprint arXiv:2505.22094, 2025.

WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning Reinflow: Fine-tuning flow matching policy with online reinforcement learning.arXiv preprint arXiv:2505.22094, 2025

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T03:34:04.698327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:34:04.698327Z digest=sha256:2013d3689891004021744103727fc889578c39401fff0adcce2418130cdf0b6a

Observation 68c8c9e2-c66e-4544-ac17-6f0fce99ed93 Β· outbound

This paper cites Flow-GRPO: Training Flow Matching Models via Online RL.

WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning Flow-GRPO: Training Flow Matching Models via Online RL

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-03T03:34:04.727487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:34:04.727487Z digest=sha256:50709d068e8d92ce8826e1c8bcb14872b0e47375f38ad24893cb87a9efca3f7f

Observation 9f13788d-107f-4628-ba61-208e54ae897a Β· outbound

This paper cites Diffusion Policy Policy Optimization.

WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning Diffusion Policy Policy Optimization

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T03:34:04.794445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:34:04.794445Z digest=sha256:fd0198a0394fe497451364c2f9b71cc0fe9f996ba7016e6b418474e43eadbcb1

Observation f698fb0e-85e1-4c77-9a51-3c1495ac8224 Β· outbound

This paper cites Steering Your Diffusion Policy with Latent Space Reinforcement Learning.

WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning Steering Your Diffusion Policy with Latent Space Reinforcement Learning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T03:34:04.852016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:34:04.852016Z digest=sha256:c5a713d54ab00eb926ca2927404066dd1b7ed9e00fc901da64de4072df0dd675

Observation 73ca62b2-2d62-4f7c-a049-4a8f05474fff Β· outbound

This paper cites Precise and dexterous robotic manipulation via human-in- the-loop reinforcement learning.Science Robotics, 10(105):eads5033, 2025.

WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning Precise and dexterous robotic manipulation via human-in- the-loop reinforcement learning.Science Robotics, 10(105):eads5033, 2025

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-03T03:34:04.926298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:34:04.926298Z digest=sha256:b02b7e45c0b6a0e74b689f3556ca09f7ecbcb7da713e1bfb5076067ef3d1f70d

Observation cfaad6db-cb59-4d64-aaf1-de56b7e983b3 Β· outbound

This paper cites Gigabrain-0.5 m*: a vla that learns from world model-based reinforcement learning.arXiv preprint arXiv:2602.12099, 2026.

WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning Gigabrain-0.5 m*: a vla that learns from world model-based reinforcement learning.arXiv preprint arXiv:2602.12099, 2026

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-03T03:34:04.975813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:34:04.975813Z digest=sha256:445d0a12b396692d7af7b5a251baa96580e7f0047536f8330b68fce66e1f318e

Observation e929bafd-2a44-4e76-a65b-9c98b9284e70 Β· outbound

This paper cites Optimal control of markov decision processes with incomplete state estimation.J.

WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning Optimal control of markov decision processes with incomplete state estimation.J

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-03T03:34:04.998876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:34:04.998876Z digest=sha256:4dc86e1f658691227ab8ba8dad660e7cded3766d8ce758089594fb0fcaf846c4

Observation 84aed31f-fe45-4273-9cab-e6d23b77c2b1 Β· outbound

This paper cites The optimal control of partially observable markov processes over a finite horizon.Operations research, 21(5):1071–1088, 1973.

WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning The optimal control of partially observable markov processes over a finite horizon.Operations research, 21(5):1071–1088, 1973

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-03T03:34:05.048458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:34:05.048458Z digest=sha256:9f4b3497d3723dd39c3e60fe3b94bbca483f76bef771e80dd1670df26985d03c

Observation 98d935dc-c828-43c3-aff9-317e556c7864 Β· outbound

This paper cites Reinforcement learning with augmented data.Advances in neural information processing systems, 33:19884–19895, 2020.

WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning Reinforcement learning with augmented data.Advances in neural information processing systems, 33:19884–19895, 2020

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-03T03:34:05.103198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:34:05.103198Z digest=sha256:27d835374b5bb0988a2c7d7205404fb34b7f131b62ce41c6ce43d14f43c26cc4

Observation 78cd7873-4be1-4ea0-82f1-b37fe1bcb0c7 Β· outbound

This paper cites Contextual decision processes with low bellman rank are pac-learnable.

WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning Contextual decision processes with low bellman rank are pac-learnable

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-03T03:34:05.185713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:34:05.185713Z digest=sha256:0d52d6b9c633e43deae45e95e1ac015f61fbf2bf797adaa448d811f3cc739204

Observation 4e701487-ded9-4a2a-8fde-615fa793c6e4 Β· outbound

This paper cites Human-level control through deep reinforcement learning.nature, 518(7540):529–533, 2015.

WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning Human-level control through deep reinforcement learning.nature, 518(7540):529–533, 2015

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-03T03:34:05.242212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:34:05.242212Z digest=sha256:a3cafc8559c256f73aedc5ff434dc1898077abab003ad65d59901d4910ba3600

Observation 0dadea8c-6c53-4ffa-a2f7-8b220aa2ae01 Β· outbound

This paper cites Deep recurrent q-learning for partially observable mdps.

WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning Deep recurrent q-learning for partially observable mdps

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-03T03:34:05.288843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:34:05.288843Z digest=sha256:30187c255eb5d7e66b3f0d539999b2e8b1a7319d7bb1e50181ff413694f0d705

Observation 3dca8b78-4ee0-49d2-a690-58ae25b7be3c Β· outbound

This paper cites Robust Reinforcement Learning in POMDPs with Incomplete and Noisy Observations.

WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning Robust Reinforcement Learning in POMDPs with Incomplete and Noisy Observations

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-03T03:34:05.347655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:34:05.347655Z digest=sha256:ca3c20d5f676e29aee0d2398e82ca000d07a2f9137f7b57b27d7aad4cf3dff44

Observation 79d01f92-8e09-4a1d-a044-a73562782aeb Β· outbound

This paper cites Robust Finite-Memory Policy Gradients for Hidden-Model POMDPs.

WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning Robust Finite-Memory Policy Gradients for Hidden-Model POMDPs

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-03T03:34:05.377359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:34:05.377359Z digest=sha256:32527c70e2cc0b96248e586bcd48b1ba1b90776af7e98df0b38a7bd0aadfc955

Observation 5a65ab53-f0d4-4fc7-8650-380116ace64a Β· outbound

This paper cites MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation.

WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-03T03:34:05.408344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:34:05.408344Z digest=sha256:c828146ded8566c482aa358c665468062bf52f49ca1a43f5ebaeb2afd5721d82

Observation 6b3b49e7-5eb6-40f0-80b9-2abc4f9ac89e Β· outbound

This paper cites HAMLET: Switch your Vision-Language-Action Model into a History-Aware Policy.

WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning HAMLET: Switch your Vision-Language-Action Model into a History-Aware Policy

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-03T03:34:05.444734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:34:05.444734Z digest=sha256:036d57b4797aa990dcf68c2d07084a257c22e868c8b87bb8f78c207a522486ae

Observation 5a1ba548-5787-4be7-8c7e-704f6aa7dac7 Β· outbound

This paper cites Cronusvla: Transferring latent motion across time for multi-frame prediction in manipulation.arXiv e-prints, pages arXiv–2506, 2025.

WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning Cronusvla: Transferring latent motion across time for multi-frame prediction in manipulation.arXiv e-prints, pages arXiv–2506, 2025

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-03T03:34:05.487924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:34:05.487924Z digest=sha256:7190808aa458af429a56225a2396cf366ea18b2f0b89ff11b4bf5bf41256370b

Observation 6a9c4f3a-696b-4556-816b-9316f18d4e5b Β· outbound

This paper cites Decision transformer: Reinforcement learning via sequence modeling.Advances in neural information processing systems, 34:15084–15097, 2021.

WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning Decision transformer: Reinforcement learning via sequence modeling.Advances in neural information processing systems, 34:15084–15097, 2021

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-03T03:34:05.503119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:34:05.503119Z digest=sha256:8072cdf6308745434077e31e23ba9fa195eadb4b8fd083412b1ac5c03e05bcdb

Observation 6ad8aa0f-39e6-4432-abc4-c809ee450e54 Β· outbound

This paper cites World Action Models are Zero-shot Policies.

WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning World Action Models are Zero-shot Policies

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-03T03:34:05.569104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:34:05.569104Z digest=sha256:edb81ccd993bf6ad6258a25cb1cdea3dfcb3bfcea5e3013f7731f2fa7b4ec2bc

Observation e4186373-6e16-4fe1-8f3b-b2c0d989aa43 Β· outbound

This paper cites Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning.

WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-03T03:34:05.699196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:34:05.699196Z digest=sha256:53644aca5c730f3d7ac6dda24298ab1691645e23f8f7198132492e6f20badfb6

Observation 252b9028-a7b9-48d0-a9bb-9921dc501027 Β· outbound

This paper cites Motus: A Unified Latent Action World Model.

WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning Motus: A Unified Latent Action World Model

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-03T03:34:05.960639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:34:05.960639Z digest=sha256:5128a5ad9eab070c135b49dc7b7a32cb1bf0212c157350098e607d63a0f68b12

Observation 6ab9c4a7-d9ce-4ee2-9fab-67980c921be9 Β· outbound

This paper cites Causal World Modeling for Robot Control.

WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning Causal World Modeling for Robot Control

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-03T03:34:06.048945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:34:06.048945Z digest=sha256:6db9500277dad482fa95c72bbce4322fd5cf9f1fd4ab5e1626fe729136def05b

Observation 7024ebcd-9741-4e84-95a0-7e1585df37cb Β· outbound

This paper cites Fast-WAM: Do World Action Models Need Test-time Future Imagination?.

WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning Fast-WAM: Do World Action Models Need Test-time Future Imagination?

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-03T03:34:06.116648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:34:06.116648Z digest=sha256:ce652f72e4465be916f02f4dbc9d9d93d671cab02feade3449be15e25618ec36

Observation 439b7114-9c0f-4163-9f03-e08b3947a626 Β· outbound

This paper cites LeJEPA: Provable and Scalable Self-Supervised Learning Without the Heuristics.

WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning LeJEPA: Provable and Scalable Self-Supervised Learning Without the Heuristics

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-03T03:34:06.268673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:34:06.268673Z digest=sha256:2aa87f1b6b2f378e15832923c509c497868e23836536729ef402a310ede4b05b

Observation a3693436-16f0-403b-b9e0-2b1fc073cd44 Β· outbound

This paper cites LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixels.

WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixels

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-03T03:34:06.397273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:34:06.397273Z digest=sha256:1457a09f106a89e62a5421440ef39447f191735592ff9617e7fd86ed666ba59b

Observation 1e559c64-4cae-4de1-9a9a-0406d32435bc Β· outbound

This paper cites Learning transferable visual models from natural language supervision.

WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning Learning transferable visual models from natural language supervision

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-03T03:34:06.517745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:34:06.517745Z digest=sha256:3a22c1ad70a0e50330e5ed62fe19edfe10564b9bc1eb8d9bbf68f2e697d1c154

Observation 484980e3-2d16-41dc-96b1-c0e3964ba11c Β· outbound

This paper cites Film: Visual reasoning with a general conditioning layer.

WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning Film: Visual reasoning with a general conditioning layer

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-03T03:34:06.583766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:34:06.583766Z digest=sha256:5bdd9bfec633774c1ce3e03a2409fd131ccfa519d5e5c1629c11a1efc7d0fbd8

Observation 99773263-4905-41ea-a982-3fde8f212d8e Β· outbound

This paper cites ManiSkill: Generalizable Manipulation Skill Benchmark with Large-Scale Demonstrations.

WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning ManiSkill: Generalizable Manipulation Skill Benchmark with Large-Scale Demonstrations

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-03T03:34:06.739901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:34:06.739901Z digest=sha256:a30fd30d195c514e4473386b3c9a4cbbf9f0ff375eab892077e714bba84f5244

Observation 59d15a0a-12a6-4016-81f4-17d904b8a8dc Β· outbound

This paper cites Meta- world: A benchmark and evaluation for multi-task and meta reinforcement learning.

WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning Meta- world: A benchmark and evaluation for multi-task and meta reinforcement learning

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-03T03:34:06.865518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:34:06.865518Z digest=sha256:d42d9c0532c6fdf6887eed3f12f29e4d65cee66adcdd1e559f32941ae8702eff

Observation 079bb149-e56d-498b-a474-c4f90377d6ef Β· outbound

This paper cites an unresolved cited work.

WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning Unresolved cited work

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-03T03:34:07.005364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:34:07.005364Z digest=sha256:1e77b4e7246689de66dbd3138356f524819a3c047e3b5cb3b2c7ca90fadcd231

Observation d3d21f88-ce07-42eb-963b-2567fafcbccf Β· outbound

This paper cites LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models.

WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-03T03:34:07.067380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:34:07.067380Z digest=sha256:ff03f2db7e544251590b44bd79470e940685b49a1c6fb140d27c9082f4e108a8

Pith citing papers

No inbound Pith citation observations are available.