Pith. sign in

Paper Citation Record · LEDGER

Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models

As of 7 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 18 inbound Pith citation observations for arXiv:2606.17846.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.17846 v2

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-27T01:08:59.969040Z

measured 64 of 64 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:11:17.686150Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T04:16:48.688281Z

Reference resolution

46 of 46 outbound references displayed

  • verified exact37
  • verified fuzzy0
  • unresolved5
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch4

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7792fe7d-f2c6-4bcd-b9be-5f62b467dd3a · outbound

This paper cites AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems.

Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:48:55.860510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T01:08:59.969040Z digest=sha256:667f4508661b7b8dacf033592c417e5b51451acdaff2341e32e945515bacf07e

Observation 7cbfd898-25f3-471f-bc6e-f5255aee8a0a · outbound

This paper cites Qwen3-VL Technical Report.

Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models Qwen3-VL Technical Report

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:48:55.810032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T01:08:59.969040Z digest=sha256:a525248b0b55c51b4a40c6df5ce6e42e736a65a84387491d39aad4b95fa14321

Observation 5e50a8d8-7a0e-48f3-b799-89afcd1cba2d · outbound

This paper cites GR00T N1: An Open Foundation Model for Generalist Humanoid Robots.

Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:48:55.850773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T01:08:59.969040Z digest=sha256:2fb6808e1ce03d523f284d76a999843a958d2a0de48ef36139e831b35249f6a2

Observation 46e7a4a5-c9ed-4367-adcd-4c35c4da402c · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:48:55.832463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T01:08:59.969040Z digest=sha256:308e558b67d998674cd23c8227f4af73871d8b69f651623b3ab86f6f79651e51

Observation d82e847d-53d6-46b0-8d0e-450f4db6e77d · outbound

This paper cites Internvla-a1: Unifying understanding, generation and action for robotic manipulation.

Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models Internvla-a1: Unifying understanding, generation and action for robotic manipulation

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:48:55.840266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T01:08:59.969040Z digest=sha256:84f80e4de85f036f6667f08cf16a3a64a7fc238f729604bfd58ad5c059d27977

Observation a5867eb9-0fea-4218-949c-a9408c03a0ca · outbound

This paper cites SAM 3: Segment Anything with Concepts.

Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models SAM 3: Segment Anything with Concepts

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T20:48:55.827390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T01:08:59.969040Z digest=sha256:c7ed7f69fbb2c2cb8e6f7d00f1dbf8a237c3e953e1d628e733906ff6313f7912

Observation e1d42c77-e512-45ce-8395-22fd1f38ee48 · outbound

This paper cites Toward embodiment equivariant vision-language-action policy.arXiv preprint arXiv:2509.14630, 2025a.

Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models Toward embodiment equivariant vision-language-action policy.arXiv preprint arXiv:2509.14630, 2025a

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:48:55.865127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T01:08:59.969040Z digest=sha256:2cc4f47da2a7cc29a568e0d0cd2a8dc67c0977b1866780b403cb9b26dc42d073

Observation fe281978-d163-46d2-8c9a-5c58ea138587 · outbound

This paper cites StarVLA: A Lego-like Codebase for Vision-Language-Action Model Developing.

Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models StarVLA: A Lego-like Codebase for Vision-Language-Action Model Developing

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:48:55.793734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T01:08:59.969040Z digest=sha256:65dd04bf8ea835dd57cb824d092e21767e0a7860d291ab484215c8ed3a3337ab

Observation acc3dd4e-35c2-4f98-880c-3d66b142c9a2 · outbound

This paper cites Vision transformers need registers.

Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models Vision transformers need registers

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-27T01:08:59.969040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:08:59.969040Z digest=sha256:fde982d036638b237b9dc5c10c565ac53a5d6135643f5c758dcd24914c7eefcd

Observation e9920926-cd7b-4b55-8d32-cab090d1d162 · outbound

This paper cites an unresolved cited work.

Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-27T01:08:59.969040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:08:59.969040Z digest=sha256:9249b87dc5b7523f5aaaad3073712fbc3ba517b9a84990475883a9f27e09c603

Observation a05d9af2-a934-414e-999b-babc534a4e73 · outbound

This paper cites Knowledge Insulating Vision-Language-Action Models: Train Fast, Run Fast, Generalize Better.

Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models Knowledge Insulating Vision-Language-Action Models: Train Fast, Run Fast, Generalize Better

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:48:55.801115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T01:08:59.969040Z digest=sha256:754f9df6bf607be7821b1ab3d42760a84d3995f74587abaa73729a8a5749fd5a

Observation 1471d4e4-8c9e-4168-83bf-189c86a4d29b · outbound

This paper cites The Llama 3 Herd of Models.

Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models The Llama 3 Herd of Models

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:48:55.812451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T01:08:59.969040Z digest=sha256:c915e1054202658a5de6eea4d679ef06d5df07ffa74d4d6deb7b0c94f7e7d975

Observation 0a03aa81-afba-469c-af01-20ee83e1281b · outbound

This paper cites MolmoAct2: Action Reasoning Models for Real-world Deployment.

Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models MolmoAct2: Action Reasoning Models for Real-world Deployment

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:48:55.821613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T01:08:59.969040Z digest=sha256:5689b004e170daad272da599a53cdf1b0e215b78a006ef18d7ea1b23c3c6a587

Observation 14c7e704-ab7b-4910-ada2-45f5f76dc3ae · outbound

This paper cites LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models.

Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:48:55.858265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T01:08:59.969040Z digest=sha256:84c6ee3b32388a62b2ee4d1bcc347a6633e836a9b19bc272cde901c56491375b

Observation 5ef8fc13-dd01-472f-bb95-1df712d9f320 · outbound

This paper cites ProcVLM: Learning Procedure-Grounded Progress Rewards for Robotic Manipulation.

Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models ProcVLM: Learning Procedure-Grounded Progress Rewards for Robotic Manipulation

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:48:55.862671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T01:08:59.969040Z digest=sha256:630549791cd1ac1b4067ec39f818f26154a4a602ba901a502fbef72b11877d94

Observation 20910669-4e40-4c8f-bb64-caa6adac6337 · outbound

This paper cites Galaxea Open-World Dataset and G0 Dual-System VLA Model.

Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models Galaxea Open-World Dataset and G0 Dual-System VLA Model

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:48:55.778096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T01:08:59.969040Z digest=sha256:6231b3c697eb6700fa2c17c19afebd3bafb88e1ad8779846b6309cea1bb03575

Observation c882aede-3073-44b2-aaf5-9fa7aaf667d3 · outbound

This paper cites EgoDex: Learning Dexterous Manipulation from Large-Scale Egocentric Video.

Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models EgoDex: Learning Dexterous Manipulation from Large-Scale Egocentric Video

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:48:55.824335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T01:08:59.969040Z digest=sha256:e7a56db19ac5e2e8d950f3b9ba2971bff225bdf860ba86712507a3b77d9eb064

Observation c8e98220-7d5b-4d7b-b6cd-78a3caba9bc7 · outbound

This paper cites RLDX-1 Technical Report.

Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models RLDX-1 Technical Report

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:48:55.837778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T01:08:59.969040Z digest=sha256:05647c86f793f2c41b63598b55527f49f539a7bf4e6edaa5c88f7a6ebb146a7e

Observation 07db1c2f-5e27-4aa5-9445-0a971de18bc0 · outbound

This paper cites Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success.

Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:48:55.801162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T01:08:59.969040Z digest=sha256:6ecd50571a538619a2ee2ad94f986a55d2f57ef8a75447d97b0e7a7e59763049

Observation f219d875-d3ef-4889-9246-b68bc7198e2f · outbound

This paper cites Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning.

Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:48:55.853373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T01:08:59.969040Z digest=sha256:99ae49c9e2633c4b2293ccae2916be5e5db99521bedf1bf5fc8cd864a0fe58b6

Observation b72d442f-4909-4e37-8a82-91272b6dc6e3 · outbound

This paper cites Towards Long-horizon Embodied Agents with Tool-Aligned Vision-Language-Action Models.

Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models Towards Long-horizon Embodied Agents with Tool-Aligned Vision-Language-Action Models

Reference 21

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T20:48:55.766802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T01:08:59.969040Z digest=sha256:9ce341e745f1c8b68ae0718e5d53f271acd35fd3c68f1902f227c4bb85ae3c63

Observation 2329b5ee-e71c-44a5-b576-14dfaad6aac8 · outbound

This paper cites Masquerade: Learning from In-the-wild Human Videos using Data-Editing.

Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models Masquerade: Learning from In-the-wild Human Videos using Data-Editing

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:48:55.848413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T01:08:59.969040Z digest=sha256:a589c27b6dbdd653247170f339a7c42c994ebe40b069b593bb6cacb498eeb1a8

Observation f009df3c-07cd-4322-9226-8ef7509c6927 · outbound

This paper cites Scalable vision-language-action model pretraining for robotic manipulation with real-life human activity videos.

Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models Scalable vision-language-action model pretraining for robotic manipulation with real-life human activity videos

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:48:55.830107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T01:08:59.969040Z digest=sha256:7447d203f03e03a6d5c66f10dff3971c9d02f287371176a08064eaa3c6ef73f9

Observation e800d41b-0a48-47f4-bdaa-c2e97936b2f5 · outbound

This paper cites Depth Anything 3: Recovering the Visual Space from Any Views.

Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models Depth Anything 3: Recovering the Visual Space from Any Views

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:48:55.861204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T01:08:59.969040Z digest=sha256:a2e93d9f7e9a34ef0230f49269b1de4f07021b922b17a98b975b03e9b13aba2d

Observation 06360092-ba71-4869-be2f-5a08b0eb493b · outbound

This paper cites LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning.

Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:48:55.842798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T01:08:59.969040Z digest=sha256:a9facec6cadb4f3d35491b62bbfdc7ab56388e619f34f23774e13dfbdfe8d567

Observation 67082a9d-78e1-4134-b360-31a571b052fa · outbound

This paper cites Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos.

Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:48:55.855989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T01:08:59.969040Z digest=sha256:a4abb0616c9f872be7d0343eda9b035550ec89f259687c240c76e5c1b51d147f

Observation 27c9e34c-c76c-4053-b6f7-1f9927fa29fa · outbound

This paper cites Being-h0.

Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models Being-h0

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T20:48:55.832301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T01:08:59.969040Z digest=sha256:26b07b62812b1b6f65a9cb589172b4a6ec76004b504d2d95208ff5527a3a8015

Observation 462f815d-d991-41d7-9e4c-1e90737327c1 · outbound

This paper cites ELECTRA and GPT-4o: Cost-Effective Partners for Sentiment Analysis.

Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models ELECTRA and GPT-4o: Cost-Effective Partners for Sentiment Analysis

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:48:55.824979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T01:08:59.969040Z digest=sha256:ef7382487dc0f3fa5871fd4ca46fa4ef6ff063e64242c5534d31b44987203557

Observation 10043a7a-ce3a-4c84-bd68-1479378f7573 · outbound

This paper cites GPT-4 Technical Report.

Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models GPT-4 Technical Report

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:48:55.870919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T01:08:59.969040Z digest=sha256:bf11fef4967c242a1bee521bf9a5129d54e31c74d3acd69d8e40838789e559a8

Observation 4e516733-8adf-417a-9794-37c9d4e8b41a · outbound

This paper cites EgoVerse: An Egocentric Human Dataset for Robot Learning from Around the World.

Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models EgoVerse: An Egocentric Human Dataset for Robot Learning from Around the World

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:48:55.868731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T01:08:59.969040Z digest=sha256:57d99f0c074c1035b1867bc6fb62234233f32f8e330d4d7d3a1e437916d37e26

Observation da47f46e-8796-4b19-af27-26bb046beca3 · outbound

This paper cites Yoon, Ryan Hoque, Lars Paulsen, Ge Yang, Jian Zhang, Sha Yi, Guanya Shi, and Xiaolong Wang.

Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models Yoon, Ryan Hoque, Lars Paulsen, Ge Yang, Jian Zhang, Sha Yi, Guanya Shi, and Xiaolong Wang

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T20:48:55.866126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T01:08:59.969040Z digest=sha256:3030e39793edf9f764b043d8996bd11d50e48770d6a384cfb087c6a347252c6d

Observation 73cd3371-c47e-4107-b97c-c180e41acadc · outbound

This paper cites Embodied Hands: Modeling and Capturing Hands and Bodies Together.

Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models Embodied Hands: Modeling and Capturing Hands and Bodies Together

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:48:55.827067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T01:08:59.969040Z digest=sha256:6e7082aa6c3ab4a82183511234bca159685add22a6ab8cc7d56dc99eb39fc2c0

Observation b13d6c66-42aa-4d7b-8b3f-2795f03d2acf · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models Gemini: A Family of Highly Capable Multimodal Models

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:48:55.848138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T01:08:59.969040Z digest=sha256:e3894fe8f6194ca58d5972d91e77ae54ad546a38c8e3979ed45d332e1393a7d0

Observation 92bcecc6-8c39-4998-94cd-47f86775475d · outbound

This paper cites Interndata-a1: Pioneering high-fidelity synthetic data for pre-training generalist policy.

Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models Interndata-a1: Pioneering high-fidelity synthetic data for pre-training generalist policy

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:48:55.850882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T01:08:59.969040Z digest=sha256:8bf06920e73c2616839f98d9bde0d238dbafa0effd94b8540a1e985e64cc7fe3

Observation 5c42effb-06ae-4930-aaeb-c8219bbfa5f1 · outbound

This paper cites Mujoco: A physics engine for model-based control.

Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models Mujoco: A physics engine for model-based control

Reference 35

Resolution
unresolved
no resolver link, observed 2026-06-27T01:08:59.969040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:08:59.969040Z digest=sha256:03f2d235761f7d64f37f5a3c3d33de1e1bfe93812a974c6dabdc115747ea0ed5

Observation 9d26fa16-95f4-47ca-b689-7fb45dd06aa5 · outbound

This paper cites Rethinking visual-language-action model scaling: Alignment, mixture, and regularization.

Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models Rethinking visual-language-action model scaling: Alignment, mixture, and regularization

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:48:55.853776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T01:08:59.969040Z digest=sha256:f6c6bee2ee594cb21dc956e6a9c4e293173550412676f0e275eb9eb0edfc1f8b

Observation 554b48c2-6268-444e-b161-c81e069b0f16 · outbound

This paper cites Robomind 2.0: A multimodal, bimanual mobile manipulation dataset for generalizable embodied intelligence.

Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models Robomind 2.0: A multimodal, bimanual mobile manipulation dataset for generalizable embodied intelligence

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:48:55.845703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T01:08:59.969040Z digest=sha256:f83f1296f78021e7ed589430a0ad42ef2b38d1ae16ec8fe3ea0c8253a475df20

Observation f79a3620-9d27-4921-a0b7-269941123534 · outbound

This paper cites Qwen3 Technical Report.

Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models Qwen3 Technical Report

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:48:55.856125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T01:08:59.969040Z digest=sha256:a73366d4ce9da4c1dea1d2e360e5229988e7d1fb3b624468c183ca4432a102a6

Observation 6f824e62-24a6-437b-abca-7a171ad580db · outbound

This paper cites ABot-M0: VLA Foundation Model for Robotic Manipulation with Action Manifold Learning.

Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models ABot-M0: VLA Foundation Model for Robotic Manipulation with Action Manifold Learning

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:48:55.858489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T01:08:59.969040Z digest=sha256:d1b7cc0a96ab775c922d47dc11a96b65846585044a9430922977d81a7483f0cc

Observation 705d79dd-fe30-4408-8330-4406864d2882 · outbound

This paper cites Capsfusion: Rethinking image-text data at scale.

Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models Capsfusion: Rethinking image-text data at scale

Reference 40

Resolution
unresolved
no resolver link, observed 2026-06-27T01:08:59.969040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:08:59.969040Z digest=sha256:a016be9d6d4f1d1ed719d9d06e9c0df92704aeee5c52052ab235945a3c277869

Observation e546d2f2-5309-4788-a033-18cece2ba1b2 · outbound

This paper cites Robopoint: A vision-language model for spatial affordance prediction in robotics.

Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models Robopoint: A vision-language model for spatial affordance prediction in robotics

Reference 41

Resolution
unresolved
no resolver link, observed 2026-06-27T01:08:59.969040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:08:59.969040Z digest=sha256:554a19aa492c5ec77d84484964121e03dd13bffd1b902bc5b2406016c5a1f263

Observation a6bc78f5-cf00-4d47-9b85-3c5ed4508ff6 · outbound

This paper cites Robotic Control via Embodied Chain-of-Thought Reasoning.

Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models Robotic Control via Embodied Chain-of-Thought Reasoning

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:48:55.863414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T01:08:59.969040Z digest=sha256:6878cee6a2a9c5c588ba36af939573e8067358014a4196451531c4975c48896f

Observation 787c745d-5760-40ba-b7b2-97acbd09e2ce · outbound

This paper cites Easymimic: A low-cost framework for robot imitation learning from human videos.arXiv preprint arXiv:2602.11464, 2026a.

Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models Easymimic: A low-cost framework for robot imitation learning from human videos.arXiv preprint arXiv:2602.11464, 2026a

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:48:55.842891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T01:08:59.969040Z digest=sha256:e82dcd3e0dfc9445803464c5b93cde99411de049854c67fa0093c353a8a31896

Observation 8b7b98dd-8a28-43ee-9e37-8d4b113603b6 · outbound

This paper cites X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model.

Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:48:55.840223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T01:08:59.969040Z digest=sha256:2c4f318b71acbfd96ee0347fee6c72ed0514b107ae5fe0802525b32e33156639

Observation 143e8bc2-e92c-4986-b4c6-98cb65063b14 · outbound

This paper cites EgoScale: Scaling Dexterous Manipulation with Diverse Ego- centric Human Data.

Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models EgoScale: Scaling Dexterous Manipulation with Diverse Ego- centric Human Data

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:48:55.837878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T01:08:59.969040Z digest=sha256:cbab92321f92e839e3aa56ca025f076cc2cf92d08136d31edc5c4ae3a617448c

Observation 6088ca18-0988-413f-9897-498e9fdba47a · outbound

This paper cites RoboRefer: Towards spatial referring with rea- soning in vision-language models for robotics.

Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models RoboRefer: Towards spatial referring with rea- soning in vision-language models for robotics

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:48:55.835028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T01:08:59.969040Z digest=sha256:d3c6c938a679696f96355569c3042350e23451b4753776982bb1e59ece642eb4

Pith citing papers

Observation 03a13d1f-f04a-4cb9-984e-4ac1f0b7413f · inbound

Supervise What Survives: Geometry-Guided VLA Adaptation from Synthetic Robot Videos cites this paper.

Supervise What Survives: Geometry-Guided VLA Adaptation from Synthetic Robot Videos Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-04T17:09:58.456768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-25T23:57:48.395172Z digest=sha256:a30fdb38add1ddcfa864b55ca5adbb18da9567220a93524ae6ea63b6e681b0db

Observation 220d2683-6c36-45a2-89e0-91623ddd73f7 · inbound

Learning Action Priors for Cross-embodiment Robot Manipulation cites this paper.

Learning Action Priors for Cross-embodiment Robot Manipulation Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-07-04T21:00:09.758743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-25T19:09:56.409766Z digest=sha256:f9096d9b475a04d73a27a096934c5f78962b43b0ad5e7b196c6245b36237fef7

Observation 1709861a-067b-4608-811b-f1ac78b595c8 · inbound

Training Vision-Language-Action Models with Dense Embodied Chain-of-Thought Supervision cites this paper.

Training Vision-Language-Action Models with Dense Embodied Chain-of-Thought Supervision Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-06-30T15:44:48.960169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T05:10:51.005004Z digest=sha256:794bd286f8b6905eb155363330d38e1cec1aea9d82817b881177e3114de2a8cd

Observation 782069d7-fac3-4d16-81dc-468a3389e979 · inbound

Training Vision-Language-Action Models with Dense Embodied Chain-of-Thought Supervision cites this paper.

Training Vision-Language-Action Models with Dense Embodied Chain-of-Thought Supervision Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-07-02T20:37:22.363902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T20:29:13.282030Z digest=sha256:cefafb5d67e139c3eb32ab0a1fb4fd8ff699444d6c49563fbb81b1efd0f4a2c5

Observation f19c76ee-d3b4-475a-bda8-5f7a4bee8757 · inbound

ABot-M0.5: Unified Mobility-and-Manipulation World Action Model cites this paper.

ABot-M0.5: Unified Mobility-and-Manipulation World Action Model Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models

Reference 75

Resolution
verified exact
local_arxiv, observed 2026-07-02T14:27:03.141017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T14:24:23.187164Z digest=sha256:956a1fc7547db334aad6a4cd3212695b718202dd9fd9dc407f66c0e569ce952a

Observation d731f79d-8218-4634-9ea8-125c7ab87526 · inbound

ABot-M0.5: Unified Mobility-and-Manipulation World Action Model cites this paper.

ABot-M0.5: Unified Mobility-and-Manipulation World Action Model Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models

Reference 75

Resolution
unresolved
no resolver link, observed 2026-07-12T09:22:08.000379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:22:08.000379Z digest=sha256:c2542de763f74e6809bd028301c6668c8ee63fc270c616a8b887f4bc05a2f88c

Observation a1a78ed6-52c2-4e16-9127-fcb0dc4d66a5 · inbound

ACE-Brain-0.5: A Unified Embodied Foundational Model for Physical Agentic AI cites this paper.

ACE-Brain-0.5: A Unified Embodied Foundational Model for Physical Agentic AI Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-11T19:16:57.396710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:16:57.396710Z digest=sha256:0cc0cbf4be57940ef883e3edb22a3f278d091644b4c665370212112f861475ce

Observation 34bc4026-6426-4ed8-a6b6-040b5f9e289b · inbound

RoboDojo: A Unified Sim-and-Real Benchmark for Comprehensive Evaluation of Generalist Robot Manipulation Policies cites this paper.

RoboDojo: A Unified Sim-and-Real Benchmark for Comprehensive Evaluation of Generalist Robot Manipulation Policies Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models

Reference 69

Resolution
unresolved
no resolver link, observed 2026-07-11T19:13:23.494763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:13:23.494763Z digest=sha256:62d86193f3949e06e471af4190fc1066da8dfc3d7b67b437efee50c005088d4f

Observation 4fb16920-1e39-415e-98b1-a5c900566fc6 · inbound

Native Video-Action Pretraining for Generalizable Robot Control cites this paper.

Native Video-Action Pretraining for Generalizable Robot Control Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models

Reference 118

Resolution
verified exact
local_arxiv, observed 2026-07-10T04:16:48.689527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-10T04:12:29.153764Z digest=sha256:0c3385f2f1c9fba0653d470011883b979543b08fea27014a075cb6e90c56e6e6

Observation 48af4ce6-da46-471c-8966-6fb1a7fd7102 · inbound

Native Video-Action Pretraining for Generalizable Robot Control cites this paper.

Native Video-Action Pretraining for Generalizable Robot Control Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models

Reference 119

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:45.461220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:45.461220Z digest=sha256:5223dd8b48c07d73de8c70b2de8e51fe84098580e88acccec805e7b3e36e0635

Observation 1eafb56b-242d-48e1-a279-b236400e7160 · inbound

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination cites this paper.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:53.165132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:53.165132Z digest=sha256:55a99fbee661ebb9bd671712908d5c214c04a5156ddf418cd682e73d952bf6e2

Observation 720b3c77-7128-468e-8bed-864145797870 · inbound

LeapBot-WA: World-Anchor Action Models via Predictive Latent Alignments cites this paper.

LeapBot-WA: World-Anchor Action Models via Predictive Latent Alignments Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models

Reference 96

Resolution
unresolved
no resolver link, observed 2026-07-31T23:29:15.462624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T23:29:15.462624Z digest=sha256:3c95ff2f567ebe1e8b1f1cde70e3c9ef4b949ab72cc7ae7665d747eb6177e595

Observation 6b6e4223-43cf-4315-afdc-a7dad6cc1d00 · inbound

HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone cites this paper.

HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T01:15:07.952361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T01:15:07.952361Z digest=sha256:28044a4f5bed93043910da90907c74d62834dcba0dae64fbe6e1981d1b726fba

Observation d4e33866-fd90-48c8-b426-5fd62dbf524d · inbound

CheckVLA: Execution-Time Verification with Action-Conditioned World Model for Long-Horizon Mobile Manipulation cites this paper.

CheckVLA: Execution-Time Verification with Action-Conditioned World Model for Long-Horizon Mobile Manipulation Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-07-30T21:05:52.762823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T21:05:52.762823Z digest=sha256:15fb54a12b9a95abc29f9cef485399e374fe64b328f8402f902463eca8fb9f41

Observation a4ca67ac-a508-44b8-8353-3541ad5e4084 · inbound

DLAM: Distributional Latent Actions with Temporal Constraints cites this paper.

DLAM: Distributional Latent Actions with Temporal Constraints Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-30T11:05:16.112572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T11:05:16.112572Z digest=sha256:d851cc11bf0eafee3bdc4b0faa56b5701e8a0dde8c14f7bb0fe7ea9262fcc13b

Observation e691e4ee-2e52-431d-8b73-fe3809035e73 · inbound

Hand-Object Interaction in the Age of Large Foundation Models:Reconstruction, Generation, and Embodied Transfer cites this paper.

Hand-Object Interaction in the Age of Large Foundation Models:Reconstruction, Generation, and Embodied Transfer Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models

Reference 260

Resolution
unresolved
no resolver link, observed 2026-07-31T08:51:32.436517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T08:51:32.436517Z digest=sha256:aa7a2770ae7530a4a2bef1a566692377fb559bf29372431692ee3e02a83b3591

Observation 28185618-4ba0-4082-8094-ac01891329aa · inbound

Hand-Object Interaction in the Age of Large Foundation Models:Reconstruction, Generation, and Embodied Transfer cites this paper.

Hand-Object Interaction in the Age of Large Foundation Models:Reconstruction, Generation, and Embodied Transfer Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models

Reference 241

Resolution
unresolved
no resolver link, observed 2026-08-04T01:23:12.520335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:23:12.520335Z digest=sha256:1db29ae0749a7e5d5ea04648db5dbbff1e1f2a656a170ea75862fb656498c9d8

Observation ca2363f2-ee95-4de9-a150-5722fb4947f1 · inbound

GeniWorld: A Generalizable Interactive World Model for Robotic Manipulation via Visual Actions cites this paper.

GeniWorld: A Generalizable Interactive World Model for Robotic Manipulation via Visual Actions Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:11:17.686150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:11:17.686150Z digest=sha256:34631397ddef789add16fe859ce98948146e7121022481aca7975309fb2a359f