Pith. sign in

Paper Citation Record · LEDGER

MolmoAct2: Action Reasoning Models for Real-world Deployment

As of 6 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 31 inbound Pith citation observations for arXiv:2605.02881.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.02881 v2

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-11T00:59:54.787472Z

measured 87 of 87 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 31 of 31 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T05:05:10.384913Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

56 of 56 outbound references displayed

  • verified exact50
  • verified fuzzy3
  • unresolved2
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

0
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 7b49e2f3-a5cc-4d08-8360-badc491287dc · outbound

This paper cites Qwen3-VL Technical Report.

MolmoAct2: Action Reasoning Models for Real-world Deployment Qwen3-VL Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-11T04:55:57.306199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T00:59:54.787472Z digest=sha256:26adfedb353a43ee75d1712403ac189e1e1fb37cac26f6fafd966fcf0e4f36be

Observation 3e5d8cb7-63a3-4773-b41b-4fc86aa6d1fb · outbound

This paper cites Berscheid, P.

MolmoAct2: Action Reasoning Models for Real-world Deployment Berscheid, P

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T00:40:27.359456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T00:59:54.787472Z digest=sha256:20be352f722771f3bb07fd8b991b333ed670d495a67a01068909315160c424d4

Observation 1ac003eb-44b0-4763-8ce3-3239a1d058a7 · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

MolmoAct2: Action Reasoning Models for Real-world Deployment $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-11T04:55:57.212355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T00:59:54.787472Z digest=sha256:25ef5712fe3b200cce25791415094e821c8ff47695872ce7a6fc8d4d935c59db

Observation aab02a0c-3713-4454-9e9a-c38b6153cc66 · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

MolmoAct2: Action Reasoning Models for Real-world Deployment RT-1: Robotics Transformer for Real-World Control at Scale

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-11T04:55:57.351606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T00:59:54.787472Z digest=sha256:2b6184dd41bd8ef961b9fdae46287898cc665391413e55c606e33de7ec8678bb

Observation 66f523f6-efba-43bf-90e5-a5ed53676da5 · outbound

This paper cites Sims-v: Simulated instruction-tuning for spatial video understanding.

MolmoAct2: Action Reasoning Models for Real-world Deployment Sims-v: Simulated instruction-tuning for spatial video understanding

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:55:57.081169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T00:59:54.787472Z digest=sha256:0ef5c750e2545b4e8ff11ad153c8eb02ec36a5afeaf09c25995e2f0345ea9bea

Observation 4d45beaf-1cd7-46ee-966a-70f787df3caf · outbound

This paper cites Scaling spatial intelligence with multimodal foundation models.

MolmoAct2: Action Reasoning Models for Real-world Deployment Scaling spatial intelligence with multimodal foundation models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:55:57.329449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T00:59:54.787472Z digest=sha256:8e4077188ca1e0e6e00793e9b49e41ee77dedecbfb43226eb803359d6936b84d

Observation 9fc55664-f5a3-49a2-9a06-8e511421e035 · outbound

This paper cites PointArena: Probing Multimodal Grounding Through Language-Guided Pointing.

MolmoAct2: Action Reasoning Models for Real-world Deployment PointArena: Probing Multimodal Grounding Through Language-Guided Pointing

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:55:57.197919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T00:59:54.787472Z digest=sha256:441013bfc92af6160bd4866d1fc07ec16b7222662d64f00d0d8f60c45b5ba89b

Observation b4f01cc1-e0ce-40b3-9f98-d2eee21ca948 · outbound

This paper cites Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding.

MolmoAct2: Action Reasoning Models for Real-world Deployment Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-16T04:21:30.454351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T00:59:54.787472Z digest=sha256:c4e208de620357be35e62c7f5df1d9464f7baf5a4c042971ab7e873cd836b72d

Observation b37d6612-e929-4a24-96ca-b04ff8b74e75 · outbound

This paper cites RoboNet: Large-Scale Multi-Robot Learning.

MolmoAct2: Action Reasoning Models for Real-world Deployment RoboNet: Large-Scale Multi-Robot Learning

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-17T19:08:29.717161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T00:59:54.787472Z digest=sha256:9647b34290c8fe025fd1b0b4fc20449574a5e784b0e9768d8c681d7910573023

Observation 5f5dc25b-ad35-4b5c-8263-6d52775a873b · outbound

This paper cites StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision.

MolmoAct2: Action Reasoning Models for Real-world Deployment StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-06-29T02:14:23.857212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T00:59:54.787472Z digest=sha256:0762c7a8d1dd4019a0511c0ac08f459935e334dc6df15747ca62d8901a321122

Observation 3d8b7f54-3007-4915-a627-c67e87743aed · outbound

This paper cites Deshpande, M.

MolmoAct2: Action Reasoning Models for Real-world Deployment Deshpande, M

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:55:57.045565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T00:59:54.787472Z digest=sha256:e95a93e0209e394b23e35dfffb5ba5baf2b2bcac29e955ff473b5d1c96951910

Observation 8dbdd4d7-1fdc-4c73-b31d-af0e838aaf27 · outbound

This paper cites Knowledge Insulating Vision-Language-Action Models: Train Fast, Run Fast, Generalize Better.

MolmoAct2: Action Reasoning Models for Real-world Deployment Knowledge Insulating Vision-Language-Action Models: Train Fast, Run Fast, Generalize Better

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:55:57.430704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T00:59:54.787472Z digest=sha256:192426321fed9c4846d141362ed754a4bc9e26d19a38e3f121fc20390beb7a1d

Observation 7e396091-4020-48fc-86c0-e580b6466b72 · outbound

This paper cites Bridge Data: Boosting Generalization of Robotic Skills with Cross-Domain Datasets.

MolmoAct2: Action Reasoning Models for Real-world Deployment Bridge Data: Boosting Generalization of Robotic Skills with Cross-Domain Datasets

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-13T19:55:55.512010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T00:59:54.787472Z digest=sha256:9dd775b8d28f7e95d103a18f158333861515a0b0002eb0958824179149986159

Observation c0097d0f-158d-4a9b-ba98-9559a4982cde · outbound

This paper cites RH20T: A Comprehensive Robotic Dataset for Learning Diverse Skills in One-Shot.

MolmoAct2: Action Reasoning Models for Real-world Deployment RH20T: A Comprehensive Robotic Dataset for Learning Diverse Skills in One-Shot

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:55:56.986582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T00:59:54.787472Z digest=sha256:3f737d55797db781a5741d21cf88472e1d915717b04fd79c8233ff50761a76bf

Observation 0a135132-adaf-410e-84db-cd80d96e06ac · outbound

This paper cites Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives.

MolmoAct2: Action Reasoning Models for Real-world Deployment Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:55:57.151671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T00:59:54.787472Z digest=sha256:c5f48f04295e3f7f1d56b02586c148d64c5501693d6be841bd9e5193cfbe68cb

Observation e1723710-3fb4-474a-a5bb-d2cdcf4fa735 · outbound

This paper cites Accessed: 2024-12-31.

MolmoAct2: Action Reasoning Models for Real-world Deployment Accessed: 2024-12-31

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:55:57.249924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T00:59:54.787472Z digest=sha256:088bab78760cdd6bedb2664842d1da9935a9a88159aad8d038ec5bfad53b6126

Observation 06149cdf-1a51-4c86-a7a6-27d23e0645ac · outbound

This paper cites NORA: A Small Open-Sourced Generalist Vision Language Action Model for Embodied Tasks.

MolmoAct2: Action Reasoning Models for Real-world Deployment NORA: A Small Open-Sourced Generalist Vision Language Action Model for Embodied Tasks

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-16T15:53:29.501619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T00:59:54.787472Z digest=sha256:2da342cda6736de098ea6faf7f9c7b4dd2d08ec53ae2e26cefb477aa152afb35

Observation 3ca969a2-2062-4c90-b9a8-ed251b7bb76d · outbound

This paper cites $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization.

MolmoAct2: Action Reasoning Models for Real-world Deployment $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-11T04:55:57.006466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T00:59:54.787472Z digest=sha256:7a6526b1ae8dc78bf5714877fbb0319903556b17e10a8aadbaf9578085d3ffef

Observation 253cf7e7-ce59-40b3-bac6-93a5e57c4cb6 · outbound

This paper cites CLEVR: A Diagnostic Dataset for Compositional Language and Elementary Visual Reasoning.

MolmoAct2: Action Reasoning Models for Real-world Deployment CLEVR: A Diagnostic Dataset for Compositional Language and Elementary Visual Reasoning

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:55:57.465368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T00:59:54.787472Z digest=sha256:e9b44bc2b0db58b80607375ebcc1ae623f4b015d2cce596f64ba7eef358f8383

Observation 9884ad99-e5e0-40e4-abf6-3b92aa236e63 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

MolmoAct2: Action Reasoning Models for Real-world Deployment OpenVLA: An Open-Source Vision-Language-Action Model

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-11T04:55:57.339764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T00:59:54.787472Z digest=sha256:3e1681bd5ff097db83b960571b96276374aec8f3a63a736ff4c67cdc2dbd2f93

Observation b43b9662-87bb-4766-a465-f370660abc5f · outbound

This paper cites Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success.

MolmoAct2: Action Reasoning Models for Real-world Deployment Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-11T04:55:57.173053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T00:59:54.787472Z digest=sha256:86d7aa5482ddb60bc53e046a4ebb1e50c9c4bb42acd110082663c74f13559e9d

Observation 9f7c52cd-7bea-44f3-a88b-12673aa95206 · outbound

This paper cites Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning.

MolmoAct2: Action Reasoning Models for Real-world Deployment Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-12T14:50:13.054714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T00:59:54.787472Z digest=sha256:226d081aaee4b35b9f6b1760d8f6b029bb556be75a1662342c2742b9b8ccf60d

Observation cf981437-b674-4b8d-be86-395cb9a906e7 · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

MolmoAct2: Action Reasoning Models for Real-world Deployment Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:06:39.239829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T00:59:54.787472Z digest=sha256:5c4a0438435b99588f44db340422d81ed17afc93bbbfb5ec322801d468d80016

Observation 6c8e35ff-79cd-4217-8463-ca23deed3a71 · outbound

This paper cites MolmoAct: Action Reasoning Models that can Reason in Space.

MolmoAct2: Action Reasoning Models for Real-world Deployment MolmoAct: Action Reasoning Models that can Reason in Space

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-14T23:35:23.409560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T00:59:54.787472Z digest=sha256:00133930b422dc7181543e5026517f37ae9d9edae54946dd5e77a832e757b6e9

Observation 311a71b0-a972-4418-98f2-2958e89f1a43 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

MolmoAct2: Action Reasoning Models for Real-world Deployment LLaVA-OneVision: Easy Visual Task Transfer

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-11T04:55:57.014057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T00:59:54.787472Z digest=sha256:d442ca7982b7ea7518f3910469f35a095de38facf30e246b30c30582230f5d26

Observation 4832b423-1abc-4ad6-be4c-223d11943244 · outbound

This paper cites Flow Matching for Generative Modeling.

MolmoAct2: Action Reasoning Models for Real-world Deployment Flow Matching for Generative Modeling

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-11T04:55:57.565000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T00:59:54.787472Z digest=sha256:e2f7e8454d2328ace423a5f6f2cf3f7d072e9a80d05b4c068e72f5731ac71711

Observation d5fe3bc5-f1d4-40f9-9864-95067fe8dcc7 · outbound

This paper cites RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation.

MolmoAct2: Action Reasoning Models for Real-world Deployment RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:46:30.715521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T00:59:54.787472Z digest=sha256:1913b0d8e19a4fa589a3f011b1be95f4499cb9c995ce19c110ceab74a3604590

Observation b5198592-196f-4a01-bd40-075bf75d892f · outbound

This paper cites All in Tokens: Unifying Output Space of Visual Tasks via Soft Token.

MolmoAct2: Action Reasoning Models for Real-world Deployment All in Tokens: Unifying Output Space of Visual Tasks via Soft Token

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:55:57.159781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T00:59:54.787472Z digest=sha256:b38a1836f4a4c13a419635cab8b45227f1cf57504ceafe024a05597c4159685d

Observation 1efaa6d5-e8ea-4234-b4b4-192f69377328 · outbound

This paper cites an unresolved cited work.

MolmoAct2: Action Reasoning Models for Real-world Deployment Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-05-16T00:40:27.347841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T00:59:54.787472Z digest=sha256:16c98caae101ff0c3e5d4637630eed26b5e837f0677e27d628051822b013d48d

Observation 68850316-85f1-4886-90d4-7fbf7a35f772 · outbound

This paper cites FAST: Efficient Action Tokenization for Vision-Language-Action Models.

MolmoAct2: Action Reasoning Models for Real-world Deployment FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:52:32.433811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T00:59:54.787472Z digest=sha256:3f9c3204e1c9a54f3260512ea677b8027e9b7d9992ec7e108bd886b6a04aa0a6

Observation d69b634a-61f3-4932-af09-ecda30f8abe3 · outbound

This paper cites SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model.

MolmoAct2: Action Reasoning Models for Real-world Deployment SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:12:22.898827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T00:59:54.787472Z digest=sha256:843b1eb391a3c5108dda9ea0bc0e19e100708d0ad3c8ff63f04434bc854bfd05

Observation 0e126d8d-3b9a-4827-86f8-dc0f0d96dca6 · outbound

This paper cites Sat: Spa- tial aptitude training for multimodal language models.

MolmoAct2: Action Reasoning Models for Real-world Deployment Sat: Spa- tial aptitude training for multimodal language models

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:55:57.497058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T00:59:54.787472Z digest=sha256:1b102d2be172e14b9a5931fbe5290753607ec93b5e41ea813568eb30578e9383

Observation 4907ffc0-347b-4fac-9e9c-20fc9e54a0dc · outbound

This paper cites RoboVQA: Multimodal Long-Horizon Reasoning for Robotics.

MolmoAct2: Action Reasoning Models for Real-world Deployment RoboVQA: Multimodal Long-Horizon Reasoning for Robotics

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:55:57.182333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T00:59:54.787472Z digest=sha256:67b05a1b07de11cbe1a998a20b3e44eb9c86bf60be226623cf76fab7dd845e24

Observation 1761c014-ccf4-4622-9f5f-142a8353ac6d · outbound

This paper cites SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics.

MolmoAct2: Action Reasoning Models for Real-world Deployment SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:22:37.721356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T00:59:54.787472Z digest=sha256:8197d9ce4836fabe7645a9ed41f082925b34021443c26214034833be91d01352

Observation e8c2aec0-2777-49a3-871b-a3f5f472a521 · outbound

This paper cites Emma-X: An Embodied Multimodal Action Model with Grounded Chain of Thought and Look-ahead Spatial Reasoning.

MolmoAct2: Action Reasoning Models for Real-world Deployment Emma-X: An Embodied Multimodal Action Model with Grounded Chain of Thought and Look-ahead Spatial Reasoning

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:55:57.288841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T00:59:54.787472Z digest=sha256:14553232e9d2deac3a9104fffb2dbdbfc14dddd5d77dafd74c4ae2a5a45ec97b

Observation c1ba78f0-1713-4945-8bc1-7ae4d1fd107e · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

MolmoAct2: Action Reasoning Models for Real-world Deployment Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-11T04:55:57.517964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T00:59:54.787472Z digest=sha256:d14a683077e337e627af9671b8836833e3bbbec0d32bc465199ef5d51a1c9261

Observation 9cd6854a-fb94-4f8e-9ddf-54c6e423b1f1 · outbound

This paper cites Gemini Robotics 1.5: Pushing the Frontier of Generalist Robots with Advanced Embodied Reasoning, Thinking, and Motion Transfer.

MolmoAct2: Action Reasoning Models for Real-world Deployment Gemini Robotics 1.5: Pushing the Frontier of Generalist Robots with Advanced Embodied Reasoning, Thinking, and Motion Transfer

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:37:16.476937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T00:59:54.787472Z digest=sha256:67056d836195f98211fd583df4ffa31aaee9f1de280352737acf50aad894ca05

Observation 9a277eb4-2dab-43cd-a824-73d204ffdaf6 · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

MolmoAct2: Action Reasoning Models for Real-world Deployment SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-11T04:55:57.541532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T00:59:54.787472Z digest=sha256:297e94fa1782f198848efe891d6be21995eb9376ad74c9d5c7183fdc97c42ee1

Observation cc4791da-7040-494c-924f-7e43e55e401a · outbound

This paper cites Recurrent-depth vla: Implicit test-time compute scaling of vision-language-action models via latent iterative reasoning.

MolmoAct2: Action Reasoning Models for Real-world Deployment Recurrent-depth vla: Implicit test-time compute scaling of vision-language-action models via latent iterative reasoning

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:55:57.320331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T00:59:54.787472Z digest=sha256:914de1fc30127a4f816165a6e76f9fdd92cd851b79e7bbebcdee9f321d527de8

Observation 2070972d-def5-46e4-8069-1ffaf03e9bd6 · outbound

This paper cites an unresolved cited work.

MolmoAct2: Action Reasoning Models for Real-world Deployment Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-05-16T00:40:27.351697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T00:59:54.787472Z digest=sha256:382a8003c3af32ef23528f8256066adc76cf2e52249c94a3e2729b0dba4db0d3

Observation 1e30d036-7643-4495-9378-b6e5dc7db8ff · outbound

This paper cites InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.

MolmoAct2: Action Reasoning Models for Real-world Deployment InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-05-11T04:55:57.119774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T00:59:54.787472Z digest=sha256:345df134cae73bb1cc0fe49fdc2480678ecf8c7ddeb1a524299bf1375e46ffcd

Observation 39649cf4-4c0f-4cde-a1d6-ab093dd36295 · outbound

This paper cites RoboEval: Where Robotic Manipulation Meets Structured and Scalable Evaluation.

MolmoAct2: Action Reasoning Models for Real-world Deployment RoboEval: Where Robotic Manipulation Meets Structured and Scalable Evaluation

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-05-11T04:55:57.471549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T00:59:54.787472Z digest=sha256:123ba8499658153c4732a8a3e9213205cecb90ee3729bdb2509796633db5d24a

Observation ec0386b6-485f-48d7-8de1-52b225dbff63 · outbound

This paper cites DexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot Control.

MolmoAct2: Action Reasoning Models for Real-world Deployment DexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot Control

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:55:57.052538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T00:59:54.787472Z digest=sha256:aa7b807edde5828e5f4d58ea4721078dc54afeef33879a9dbc50cf528fdba4db

Observation 3fb9f65f-1f04-4626-864e-0611401c9829 · outbound

This paper cites Qwen3 Technical Report.

MolmoAct2: Action Reasoning Models for Real-world Deployment Qwen3 Technical Report

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-11T04:55:57.231212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T00:59:54.787472Z digest=sha256:bfc08b11d702f7cfa85edcfb5acfa7ce56921640223947bea5e8870677fd2ceb

Observation 1c2d934b-4331-44e2-97de-e7e4253f0fe0 · outbound

This paper cites Thinking in space: How multimodal large language models see, remember, and recall spaces.

MolmoAct2: Action Reasoning Models for Real-world Deployment Thinking in space: How multimodal large language models see, remember, and recall spaces

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:55:57.137498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T00:59:54.787472Z digest=sha256:aaa5a18208f47906a052715cb41f20a6a18b07ed085869ce616991a4858f8323

Observation e397e0c9-1eca-4ee1-a70d-57ad984b25a5 · outbound

This paper cites Robotic Control via Embodied Chain-of-Thought Reasoning.

MolmoAct2: Action Reasoning Models for Real-world Deployment Robotic Control via Embodied Chain-of-Thought Reasoning

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-15T08:14:32.442267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T00:59:54.787472Z digest=sha256:7cbee55eb4b49bcd90c1013004c57e53aa2fe2b6dd0527ef7d93ba5f57f6f6e2

Observation c382cbc3-8316-41ae-ba84-780f1f191de2 · outbound

This paper cites Lap: Language-action pre-training enables zero-shot cross-embodiment transfer.

MolmoAct2: Action Reasoning Models for Real-world Deployment Lap: Language-action pre-training enables zero-shot cross-embodiment transfer

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:55:57.386378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T00:59:54.787472Z digest=sha256:2534a799ba052a105cfbd0842dae0333d8c26fa4a7eb4c303cc263202342613e

Observation bb6ff770-3cbe-4fed-91d1-3a1be755a76b · outbound

This paper cites Zhang, M.

MolmoAct2: Action Reasoning Models for Real-world Deployment Zhang, M

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:55:57.070216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T00:59:54.787472Z digest=sha256:e43106c34a900c739d8cabbabe6c35a3e0b1fd010fd93dc802f32add3eda095c

Observation b28737d2-3197-4476-a7d0-40be74a274db · outbound

This paper cites ALOHA Unleashed: A Simple Recipe for Robot Dexterity.

MolmoAct2: Action Reasoning Models for Real-world Deployment ALOHA Unleashed: A Simple Recipe for Robot Dexterity

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:55:57.279060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T00:59:54.787472Z digest=sha256:85285511493d443e9ae0187c35eee9de2f17ee2ecfed4140748f15133a72c121

Observation c98ff74f-5701-493b-ade9-a2a257ccdb3b · outbound

This paper cites X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model.

MolmoAct2: Action Reasoning Models for Real-world Deployment X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-12T14:57:47.957311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T00:59:54.787472Z digest=sha256:a4f8ff22c280415ef88c3fcb87218d8ec6bc67c9bea79ec5f023cc305db35f67

Observation d7134601-83cb-4ead-b4a8-600efcea92f5 · outbound

This paper cites TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies.

MolmoAct2: Action Reasoning Models for Real-world Deployment TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:27:23.210785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T00:59:54.787472Z digest=sha256:1e818b988044f016e702aa47008c3fa61a67271c3b0e26dcd9b21c45121c757c

Observation e40096fe-ae0d-4a55-8aa2-ca2a545bd9c0 · outbound

This paper cites RoboRefer: Towards spatial referring with rea- soning in vision-language models for robotics.

MolmoAct2: Action Reasoning Models for Real-world Deployment RoboRefer: Towards spatial referring with rea- soning in vision-language models for robotics

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:55:57.510897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T00:59:54.787472Z digest=sha256:36d7e071098d18d3075e7d67cfa08c4da5c0e826c9dd4eafda277745bce08a6b

Observation ae8b9d07-5882-4ca3-a015-b5796b8f2888 · outbound

This paper cites Unified World Models: Coupling Video and Action Diffusion for Pretraining on Large Robotic Datasets.

MolmoAct2: Action Reasoning Models for Real-world Deployment Unified World Models: Coupling Video and Action Diffusion for Pretraining on Large Robotic Datasets

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-13T16:25:00.606134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T00:59:54.787472Z digest=sha256:af91734cb2cd13756b23dd31625e717fecf8ac2e599c792657fe949038fa0a47

Observation 50db70d3-7d83-4e4d-8c07-daf3aa314e8b · outbound

This paper cites 4.1 and Sec.

MolmoAct2: Action Reasoning Models for Real-world Deployment 4.1 and Sec

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T00:40:27.345018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T00:59:54.787472Z digest=sha256:319e73c6c1d2f7270675901d5e2b5d04b7ccdede00742706e0a7910c40dbf74e

Observation b7b33432-b2bc-4d6b-9460-e278eef8bf02 · outbound

This paper cites dialects.

MolmoAct2: Action Reasoning Models for Real-world Deployment dialects

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T00:40:27.362757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T00:59:54.787472Z digest=sha256:179154ac34b58cf6d11462466e23df607c01eee388e828dc95c4c1f6e042013c

Observation f08d8320-2dcd-4cb8-95b8-666c8cfa5884 · outbound

This paper cites Camera positions are held fixed across all three models.

MolmoAct2: Action Reasoning Models for Real-world Deployment Camera positions are held fixed across all three models

Reference 56

Resolution
malformed identifier
raw_fallback, observed 2026-05-16T00:40:27.356182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T00:59:54.787472Z digest=sha256:ce63f8cf8e8fc36de56e650e42eb7c410b626431fa69ed55c12346d8aadceb54

Pith citing papers

Observation 04e7d929-150b-4cb2-a099-c465f3a27b73 · inbound

Coarse-to-Control: Action-Token Planning for Vision-Language-Action Models cites this paper.

Coarse-to-Control: Action-Token Planning for Vision-Language-Action Models MolmoAct2: Action Reasoning Models for Real-world Deployment

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-06-27T22:01:20.760224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T21:57:19.195413Z digest=sha256:41d1b4bd5539f92d516566525b31cb2ec7b57f93d37402c7526b390437a43874

Observation eb9ddbb9-46b3-4f54-a794-b2b49074cac9 · inbound

VoLo: A Physical Orchestrator for Open-Vocabulary Long-Horizon Manipulation cites this paper.

VoLo: A Physical Orchestrator for Open-Vocabulary Long-Horizon Manipulation MolmoAct2: Action Reasoning Models for Real-world Deployment

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-07-02T19:07:18.160599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T21:40:00.330510Z digest=sha256:3c03d071ab962fe14fdb7da83ef970fd119f69134d4276cdebb3a6e5b683dbaf

Observation 2892ac8a-b5f1-41be-b870-9b9f73669512 · inbound

Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models cites this paper.

Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models MolmoAct2: Action Reasoning Models for Real-world Deployment

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-14T18:07:09.018997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T18:07:09.018997Z digest=sha256:a5c72e8dfb9ac3a16c676939240fbd7c94a26de77bf20ca5bb803afe716738cc

Observation 46c47a8b-3c3e-4d08-9f05-852227d9b745 · inbound

Learning What to Say to Your VLA: Mostly Harmless Vision Language Action Model Steering cites this paper.

Learning What to Say to Your VLA: Mostly Harmless Vision Language Action Model Steering MolmoAct2: Action Reasoning Models for Real-world Deployment

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-03T10:27:56.320924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T10:02:35.406713Z digest=sha256:07239d23e59cf106d84acd96344a50bbb001248ee2d3d2d03aebe72be25fe44c

Observation 4f1a819b-a456-4ef7-a2c2-edd894d80e2c · inbound

SPARC: Reliable Spatial Annotations from Robot Demonstrations at Scale cites this paper.

SPARC: Reliable Spatial Annotations from Robot Demonstrations at Scale MolmoAct2: Action Reasoning Models for Real-world Deployment

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-03T15:28:34.327902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T06:28:14.694168Z digest=sha256:f021afa55a1e884e9d3015cb20a5740ba8432d82cb144d1ad771d184e02e1e22

Observation 0a03aa81-afba-469c-af01-20ee83e1281b · inbound

Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models cites this paper.

Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models MolmoAct2: Action Reasoning Models for Real-world Deployment

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:48:55.821613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T01:08:59.969040Z digest=sha256:e6b7b2ab729811bda542fd6ed4133ae9becc609898ab0c982ff0f460a78228b3

Observation 21f80482-41de-4e7a-87a9-2f98742c3f5b · inbound

Guava: An Effective and Universal Harness for Embodied Manipulation cites this paper.

Guava: An Effective and Universal Harness for Embodied Manipulation MolmoAct2: Action Reasoning Models for Real-world Deployment

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-03T21:38:58.342197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T00:22:53.323832Z digest=sha256:3eddd8b71dccf5ef9c3285fef04fb3fd66e683b075304708ee84bf38dadd5b12

Observation 18a2dc62-579a-4ebf-8a19-2a4bb1094065 · inbound

MolmoMotion: Forecasting Point Trajectories in 3D with Language Instruction cites this paper.

MolmoMotion: Forecasting Point Trajectories in 3D with Language Instruction MolmoAct2: Action Reasoning Models for Real-world Deployment

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:09:14.640541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T21:27:47.702578Z digest=sha256:c2ac1ca4b767a66c7d5986349c4e14c69f95bc2da245d331c2045204c556daca

Observation a4c14bdc-b8a5-41df-b0b4-a0beb3aeae1e · inbound

HumanScale: Egocentric Human Video Can Outperform Real-Robot Data for Embodied Pretraining cites this paper.

HumanScale: Egocentric Human Video Can Outperform Real-Robot Data for Embodied Pretraining MolmoAct2: Action Reasoning Models for Real-world Deployment

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-04T03:39:29.615529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T17:53:24.431287Z digest=sha256:93180e62a2f12b27ea8107c3293ed7e5a8279b37f382d5eb2070b30d55595b10

Observation 4e46e123-6b38-46d8-817f-1ae7e82c2f32 · inbound

NAC: Neural Action Codec for Vision-Language-Action Models cites this paper.

NAC: Neural Action Codec for Vision-Language-Action Models MolmoAct2: Action Reasoning Models for Real-world Deployment

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-07-04T06:59:37.867290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T13:59:53.484305Z digest=sha256:4ae5d24c7db971f7ff7abf3718649125094e2ee0d13d8e3530a68197f5612b97

Observation 847fe876-6ec4-4871-af86-8f0437a3323a · inbound

VLA-FAIL: Efficient Task Failure Detection for Finetuned Vision-Language-Action Models cites this paper.

VLA-FAIL: Efficient Task Failure Detection for Finetuned Vision-Language-Action Models MolmoAct2: Action Reasoning Models for Real-world Deployment

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-04T05:59:38.225358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T14:57:41.592627Z digest=sha256:d90304194ca5013801cc14728f74ed11d7487c4bc095c05b1f47c346f22d5328

Observation acee10e4-2651-4036-9024-c4ddc7a00c95 · inbound

LA4VLA: Learning to Act without Seeing via Language-Action Pretraining cites this paper.

LA4VLA: Learning to Act without Seeing via Language-Action Pretraining MolmoAct2: Action Reasoning Models for Real-world Deployment

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-04T13:59:52.636525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T04:41:13.473022Z digest=sha256:5cc6a77050c8c7f33b81d0dd3bc056ab69c66b321705db50358ad75ef3a4e341

Observation a2f593a0-0818-4c5b-987c-d47e1f91cd89 · inbound

LA4VLA: Learning to Act without Seeing via Language-Action Pretraining cites this paper.

LA4VLA: Learning to Act without Seeing via Language-Action Pretraining MolmoAct2: Action Reasoning Models for Real-world Deployment

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-06-29T19:53:56.090438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T04:38:02.797595Z digest=sha256:d7b8209325c9b11bb1be07d727622d9dc70e4a6709e596ce2f49933c762f4598

Observation c76d0277-9ff3-4e85-b551-db3b24e3c7ad · inbound

Scalable Behavior Cloning with Open Data, Training, and Evaluation cites this paper.

Scalable Behavior Cloning with Open Data, Training, and Evaluation MolmoAct2: Action Reasoning Models for Real-world Deployment

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-04T14:19:53.796200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T04:15:34.832025Z digest=sha256:1db1068e0596fe97669e09e7cf8b0465095c465b11f0eef1f90c3ff7c378e594

Observation 48bca585-2f79-4aaa-8761-df9354ff7d9e · inbound

Training Vision-Language-Action Models with Dense Embodied Chain-of-Thought Supervision cites this paper.

Training Vision-Language-Action Models with Dense Embodied Chain-of-Thought Supervision MolmoAct2: Action Reasoning Models for Real-world Deployment

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-06-30T15:44:48.929931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T05:10:51.005004Z digest=sha256:d3baf8abcbf231f6d426c2feab09979a5377eceac18c1248551cf583d9c0fde7

Observation 56d5dedb-cbf1-44d7-805c-d7d0fc0dd38d · inbound

Training Vision-Language-Action Models with Dense Embodied Chain-of-Thought Supervision cites this paper.

Training Vision-Language-Action Models with Dense Embodied Chain-of-Thought Supervision MolmoAct2: Action Reasoning Models for Real-world Deployment

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-07-02T20:37:22.343593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-02T20:29:13.282030Z digest=sha256:4b730ec816a7740bf71eda69b0139ba6468f469ad4564af771abb941c9c39586

Observation 8b2c9acf-80ab-4367-af94-76d5f2f519c1 · inbound

Sequential Planning via Anchored Robotic Keypoints cites this paper.

Sequential Planning via Anchored Robotic Keypoints MolmoAct2: Action Reasoning Models for Real-world Deployment

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-06-30T05:04:20.439809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T04:59:12.363425Z digest=sha256:3b099a975a1ed0027eb6d69ad646362ced39a995c4c337da34eee2cc9352c0bc

Observation 7093daec-0540-48bc-a710-2b00f05698e1 · inbound

ACE-Brain-0.5: A Unified Embodied Foundational Model for Physical Agentic AI cites this paper.

ACE-Brain-0.5: A Unified Embodied Foundational Model for Physical Agentic AI MolmoAct2: Action Reasoning Models for Real-world Deployment

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-11T19:16:57.396710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:16:57.396710Z digest=sha256:64d78d9c57ee4ede5fc9c7eaf19b756420a0e747cfb9843418ab89d6822dbbdb

Observation b679f0cd-85c1-4ed9-b183-4dcea1c56464 · inbound

RoboDojo: A Unified Sim-and-Real Benchmark for Comprehensive Evaluation of Generalist Robot Manipulation Policies cites this paper.

RoboDojo: A Unified Sim-and-Real Benchmark for Comprehensive Evaluation of Generalist Robot Manipulation Policies MolmoAct2: Action Reasoning Models for Real-world Deployment

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-11T19:13:23.494763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:13:23.494763Z digest=sha256:864db2ded37ba1e08705dc64617e017a92ba810ad77b8af3cbba3165d837a291

Observation c4b2beab-3ba5-4477-83bc-c6220315f131 · inbound

GaP: A Graph-as-Policy Multi-Agent Self-Learning Harness For Variational Automation Tasks cites this paper.

GaP: A Graph-as-Policy Multi-Agent Self-Learning Harness For Variational Automation Tasks MolmoAct2: Action Reasoning Models for Real-world Deployment

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-11T07:12:02.239732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:12:02.239732Z digest=sha256:67907aa240f9ce0725fc1b2a5c8344272cf9f2c6f86107234de09855b31fb2f2

Observation 88039e0d-e590-4bef-9e81-6e2a5f164489 · inbound

Understanding and Mitigating the Video-Action Generalization Gap via Temporal Ratio cites this paper.

Understanding and Mitigating the Video-Action Generalization Gap via Temporal Ratio MolmoAct2: Action Reasoning Models for Real-world Deployment

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-07-10T12:47:05.529649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-10T12:39:02.545780Z digest=sha256:f49aa67dfae7efebe2c764df9ab3b08f171be068bc00be8e1b403af5c569fa3c

Observation 3b8ab96d-0154-4c86-92ed-fc610b9b7ac7 · inbound

Worlds in One Demo: A Synthetic Data Engine for Learning Open-World Mobile Manipulation cites this paper.

Worlds in One Demo: A Synthetic Data Engine for Learning Open-World Mobile Manipulation MolmoAct2: Action Reasoning Models for Real-world Deployment

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T06:07:00.047773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T06:07:00.047773Z digest=sha256:27c00221cb9e8574c5348878d156ea65d1d5e898160f1052b1be306e351715ac

Observation ae0e4585-db95-4289-9304-68fa09bea061 · inbound

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination cites this paper.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination MolmoAct2: Action Reasoning Models for Real-world Deployment

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:53.016209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:53.016209Z digest=sha256:912b4b325b0ac9db204ef4c82efe8ab40aaefcac33c4f27817ff884d73e265a7

Observation cf7ae672-17fb-434a-a3b0-84017c78795e · inbound

DiMaS: Distribution Matching for Steering Vision-Language-Action Models cites this paper.

DiMaS: Distribution Matching for Steering Vision-Language-Action Models MolmoAct2: Action Reasoning Models for Real-world Deployment

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T02:40:36.286530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:40:36.286530Z digest=sha256:4e3200cb72f0a3d723b3bdbe53f64a2dfe73b2d0226e323622ec0c5ebde1dc02

Observation 3ce97011-8962-433a-b591-c49bd1ccf417 · inbound

RoboTTT: Context Scaling for Robot Policies cites this paper.

RoboTTT: Context Scaling for Robot Policies MolmoAct2: Action Reasoning Models for Real-world Deployment

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T23:43:17.812765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:43:17.812765Z digest=sha256:e9f483db9bfe5cce1cc635856507a0acb590f2196475730ddcf5276620f52117

Observation b0517684-f747-44da-b962-a8232099456e · inbound

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories cites this paper.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories MolmoAct2: Action Reasoning Models for Real-world Deployment

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:01.196872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:01.196872Z digest=sha256:94039f85893174cadfd7e30785d65d8626589084e5563d56049a99ad52196b39

Observation 0fd56c30-abf1-46c4-9068-942911e35c7f · inbound

RoboHarness: Memory-Driven Orchestration of Heterogeneous Robot Policies for Long-Horizon Planning cites this paper.

RoboHarness: Memory-Driven Orchestration of Heterogeneous Robot Policies for Long-Horizon Planning MolmoAct2: Action Reasoning Models for Real-world Deployment

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-01T16:19:09.418294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T16:19:09.418294Z digest=sha256:b59c0723de3582ca54761ac694a2a65ab14db4de98db8084d29f28f522e863d9

Observation 2412d0e3-8b04-46d3-91ac-a787f5a05d83 · inbound

A Few Words Go a Long Way: Language Guided Robot Policy Synthesis cites this paper.

A Few Words Go a Long Way: Language Guided Robot Policy Synthesis MolmoAct2: Action Reasoning Models for Real-world Deployment

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-30T12:41:24.818011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:41:24.818011Z digest=sha256:5595889f0bc6c8a92026e5e0471822801a151b59607b8c1f4a4431e55907b2ab

Observation 2e971b77-d2cb-479d-9117-adfd9f3c2613 · inbound

ArmnetBench v0.1: Parallel Real-World Evaluation of Manipulation Policies on a Low-Cost Arm Farm cites this paper.

ArmnetBench v0.1: Parallel Real-World Evaluation of Manipulation Policies on a Low-Cost Arm Farm MolmoAct2: Action Reasoning Models for Real-world Deployment

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-31T13:44:33.034018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T13:44:33.034018Z digest=sha256:dfae79bbfc59bf8e3edc5a467aa24e059386345cd89f18441cf894b5ae98a312

Observation a058621a-cf7a-4176-9494-957a4831b69f · inbound

Data Pyramid for Embodied Manipulation cites this paper.

Data Pyramid for Embodied Manipulation MolmoAct2: Action Reasoning Models for Real-world Deployment

Reference 97

Resolution
unresolved
no resolver link, observed 2026-07-31T06:18:55.466819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:18:55.466819Z digest=sha256:f70a0a7cd80fa601dbe2a8ed81a1b1199c9b89d497c0ac9a6862940b67c3a2a2

Observation 4cbbffd6-21e5-4856-8f36-8bfb00622caa · inbound

Why Does Action Chunking Improve Behavioral Cloning Performance in Robotic Control? cites this paper.

Why Does Action Chunking Improve Behavioral Cloning Performance in Robotic Control? MolmoAct2: Action Reasoning Models for Real-world Deployment

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T05:05:10.384913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:05:10.384913Z digest=sha256:3f506d38bf7702b49ffdb0f4acecd34c6c41f27f34b47a45a25c7702bff01271