Pith. sign in

Paper Citation Record · LEDGER

G0.5: One Autoregressive Stream for Robot Reasoning and Action

As of 21 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 0 inbound Pith citation observations for arXiv:2608.11739.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.11739 v1

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:36:56.172379Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

61 of 61 outbound references displayed

  • verified exact0
  • verified fuzzy14
  • unresolved47
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e6a4538b-0d66-4493-b259-442553d7553a · outbound

This paper cites Rt-2: Vision-language-action models transfer web knowledge to robotic control.

G0.5: One Autoregressive Stream for Robot Reasoning and Action Rt-2: Vision-language-action models transfer web knowledge to robotic control

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:55.912331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:55.912331Z digest=sha256:3ce808097088dba91568e4688d1a9613580e03358051868021ec0bc896da1976

Observation ca02fd4b-c510-4933-be4d-b64467a05af3 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

G0.5: One Autoregressive Stream for Robot Reasoning and Action OpenVLA: An Open-Source Vision-Language-Action Model

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:55.917278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:55.917278Z digest=sha256:5f986fa6816a94fbd900dc6aa70ce5a682763de08cc2bf52433c6908fda94f52

Observation 51933ffc-6a2d-4d64-b315-b1ec127ea0d9 · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

G0.5: One Autoregressive Stream for Robot Reasoning and Action $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:55.922396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:55.922396Z digest=sha256:02e3e7cd87dcd37569ec8865015b13090ec2e83b8af16fa3ecdf8fde3c1c78c0

Observation 02353279-4265-4c9e-b1e8-d512588509bd · outbound

This paper cites $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization.

G0.5: One Autoregressive Stream for Robot Reasoning and Action $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:55.926962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:55.926962Z digest=sha256:1409ecab4f06b110dd1b80c3e195ed60c9c3e3fd9d1683e8ede754722b209643

Observation d23df188-45d1-43f4-a669-643c936e91e4 · outbound

This paper cites GR00T N1: An Open Foundation Model for Generalist Humanoid Robots.

G0.5: One Autoregressive Stream for Robot Reasoning and Action GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:55.931882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:55.931882Z digest=sha256:ed87ce545beadfe03885f72b39e1216a626b74eda9f1a66910623a4f983d1e50

Observation 7ee44f63-a2a0-42c6-bc90-a411cc0b2f02 · outbound

This paper cites SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics.

G0.5: One Autoregressive Stream for Robot Reasoning and Action SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:55.936631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:55.936631Z digest=sha256:4b9cff368cc65d689c5a3672153f0f1019d553ec20fa7f22178f5b233f242c51

Observation 6e441feb-240b-4ce5-bf4e-1c0d1232238b · outbound

This paper cites Cot-vla: Visualchain-of-thoughtreasoningforvision-language-action models.

G0.5: One Autoregressive Stream for Robot Reasoning and Action Cot-vla: Visualchain-of-thoughtreasoningforvision-language-action models

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:57.749609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:36:55.941670Z digest=sha256:832c81c859e0e9d12066a77662d4e2080d4b410ca8662b923451f89caf5dd35a

Observation 81937156-f8c8-4258-9fd5-d8b916e73a54 · outbound

This paper cites Dualcot-vla: Visual-linguistic chain of thought via parallel reasoning for vision-language-action models.arXiv preprint arXiv:2603.22280, 2026.

G0.5: One Autoregressive Stream for Robot Reasoning and Action Dualcot-vla: Visual-linguistic chain of thought via parallel reasoning for vision-language-action models.arXiv preprint arXiv:2603.22280, 2026

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:55.946552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:55.946552Z digest=sha256:9f2c0860dc17e8cca6378ca025198f6def0cbc4ed5449fbc34514c0495353d47

Observation 0df9b212-c39b-443a-99c4-191cd45f70da · outbound

This paper cites Halo: A unified vision-language-action model for embodied multimodal chain-of-thought reasoning.arXiv preprint arXiv:2602.21157, 2026.

G0.5: One Autoregressive Stream for Robot Reasoning and Action Halo: A unified vision-language-action model for embodied multimodal chain-of-thought reasoning.arXiv preprint arXiv:2602.21157, 2026

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:55.950855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:55.950855Z digest=sha256:494b45efe485f91c97a55095bc8db090827afff7d181000c1c2baa1d2a75e5b2

Observation 71786589-ee15-4fe5-82c8-6a46adf63cc0 · outbound

This paper cites Mem: Multi-scale embodied memory for vision language action models.arXiv preprint arXiv:2603.03596, 2026.

G0.5: One Autoregressive Stream for Robot Reasoning and Action Mem: Multi-scale embodied memory for vision language action models.arXiv preprint arXiv:2603.03596, 2026

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:55.955170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:55.955170Z digest=sha256:11e4b9f5ffb294559750d650b1f7ae990660e2d36f66e0ade9682d6896e21efc

Observation d176f7ec-a79f-454b-9802-63139c1eb7ad · outbound

This paper cites FAST: Efficient Action Tokenization for Vision-Language-Action Models.

G0.5: One Autoregressive Stream for Robot Reasoning and Action FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:55.959246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:55.959246Z digest=sha256:34b84d83b44422bf44c98217f9641636db009e0d2889afa5718caf9787906601

Observation 7c8e958b-89b0-421e-bc2c-72cdb841c485 · outbound

This paper cites Flowvla: Visual chain of thought-based motion reasoning for vision-language-action models.arXiv preprint arXiv:2508.18269, 2025.

G0.5: One Autoregressive Stream for Robot Reasoning and Action Flowvla: Visual chain of thought-based motion reasoning for vision-language-action models.arXiv preprint arXiv:2508.18269, 2025

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:55.963973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:55.963973Z digest=sha256:29a122765b63b80dcde26ba5d3d21bee1c6479498c44762bd242f01bfcb1ff45

Observation c40c6bab-be37-45e5-b986-223bac6ea4aa · outbound

This paper cites BEHAVIOR-1K: A benchmark for embodied ai with 1,000 everyday activities and realistic simulation.

G0.5: One Autoregressive Stream for Robot Reasoning and Action BEHAVIOR-1K: A benchmark for embodied ai with 1,000 everyday activities and realistic simulation

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:57.734827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:36:55.968316Z digest=sha256:1d625adce864a042a4fbfaaf59521fed3398a8ff45556469c488f69234bbd066

Observation e1d83742-33f5-4f8d-ae2a-aad23cb452f5 · outbound

This paper cites LIBERO: Benchmarking knowledge transfer for lifelong robot learning.Advances in Neural Information Processing Systems, 36:44776–44791, 2023.

G0.5: One Autoregressive Stream for Robot Reasoning and Action LIBERO: Benchmarking knowledge transfer for lifelong robot learning.Advances in Neural Information Processing Systems, 36:44776–44791, 2023

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:57.720481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:36:55.972164Z digest=sha256:4892e4609f697bfabf500b024b3287ebf42386667e99bd43d1d294dfa0799655

Observation 7b33b106-c09e-4051-8f90-173be0b8b1c0 · outbound

This paper cites Evaluating Real-World Robot Manipulation Policies in Simulation.

G0.5: One Autoregressive Stream for Robot Reasoning and Action Evaluating Real-World Robot Manipulation Policies in Simulation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:55.976306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:55.976306Z digest=sha256:3d64b1a120b62fb48f03f5aaab136c46861a6d8e5fead32d461d189d89a2c2b2

Observation dc35eb31-beca-464b-93dc-5e949eb3f3b5 · outbound

This paper cites Fast-WAM: Do World Action Models Need Test-time Future Imagination?.

G0.5: One Autoregressive Stream for Robot Reasoning and Action Fast-WAM: Do World Action Models Need Test-time Future Imagination?

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:55.980689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:55.980689Z digest=sha256:efb9da72604f81678bb1dd608f0e73d3daa48b65382598cf484dec6f98fd61d5

Observation 5ab3b970-d16d-4385-9d8c-40280cfee4cf · outbound

This paper cites Knowledge insulating vision-language-action models: Trainfast, runfast, generalizebetter.AdvancesinNeuralInformationProcessingSystems, 38:102867–102888, 2026.

G0.5: One Autoregressive Stream for Robot Reasoning and Action Knowledge insulating vision-language-action models: Trainfast, runfast, generalizebetter.AdvancesinNeuralInformationProcessingSystems, 38:102867–102888, 2026

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:57.705781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:36:55.984853Z digest=sha256:48dfbe7145bafeef89a0256ecc8e1fcdd9b43d8b3b720d1db07f25690731288b

Observation 118d3919-840b-41ad-bf6d-a962e1b4a4c7 · outbound

This paper cites an unresolved cited work.

G0.5: One Autoregressive Stream for Robot Reasoning and Action Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:55.989259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:55.989259Z digest=sha256:44c5495d143781b53b92d535e4b7b2ea254623d1ab28e2e5354c6caac54b8df4

Observation 7ff129e5-8f3a-414f-ae96-3b0998c99f15 · outbound

This paper cites Vq-vla: Improving vision-language-action models via scaling vector-quantized action tokenizers.

G0.5: One Autoregressive Stream for Robot Reasoning and Action Vq-vla: Improving vision-language-action models via scaling vector-quantized action tokenizers

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:55.993408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:55.993408Z digest=sha256:4b1145027786e4eb69ffa1fb22a0cca018c7cd89f318652f5392e8b1c6e830a9

Observation ea4e75db-a95c-4866-8bb0-420c4973051e · outbound

This paper cites Beast: Efficient tokenization of b-splines encoded action sequences for imitation learning.Advances in Neural Information Processing Systems, 38:172934–172959, 2026.

G0.5: One Autoregressive Stream for Robot Reasoning and Action Beast: Efficient tokenization of b-splines encoded action sequences for imitation learning.Advances in Neural Information Processing Systems, 38:172934–172959, 2026

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:57.681457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:36:55.997542Z digest=sha256:66d76ce9f4120b1fcdc6d4e9824c2ad4236dcc608d675301d19fc636d16a0379

Observation e2732b29-346a-459a-899c-54589328e9bb · outbound

This paper cites Behavior Generation with Latent Actions.

G0.5: One Autoregressive Stream for Robot Reasoning and Action Behavior Generation with Latent Actions

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:56.001612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:56.001612Z digest=sha256:9db897f8dcfe7696063e163f726d1a17c05ce6c886da1889f99431e4ee9e8d45

Observation 3b081125-96cd-41cd-b48d-7c7014fc943a · outbound

This paper cites SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model.

G0.5: One Autoregressive Stream for Robot Reasoning and Action SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:56.006016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:56.006016Z digest=sha256:73857e81cd671f26293cc16a4f6cd7afe83335dac68137d73f094aeb2bfebaec

Observation 4770f8a3-0545-4b1f-9cef-9f5944c71d0d · outbound

This paper cites Being-h0.

G0.5: One Autoregressive Stream for Robot Reasoning and Action Being-h0

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:56.010632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:56.010632Z digest=sha256:02c1f2de632e78fa643ca70eb52fce60a098bf764e145595d672097b36548805

Observation b97e535f-9489-41da-a876-6ee918418af3 · outbound

This paper cites Green-vla: Staged vision-language-action model for generalist robots.

G0.5: One Autoregressive Stream for Robot Reasoning and Action Green-vla: Staged vision-language-action model for generalist robots

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:56.014811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:56.014811Z digest=sha256:a79b1bfc978470d258de7cc221229b5a3b9e7b51b353e2a4307def962f80ce21

Observation 21573125-87f0-453f-80c9-339104c92188 · outbound

This paper cites HEX: Humanoid-Aligned Experts for Cross-Embodiment Whole-Body Manipulation.

G0.5: One Autoregressive Stream for Robot Reasoning and Action HEX: Humanoid-Aligned Experts for Cross-Embodiment Whole-Body Manipulation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:56.019095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:56.019095Z digest=sha256:5e9eebce739059ba8b282b1d63683a136eab3aa679d3ca54d5513afaff843fca

Observation c28e83a9-1918-4345-be6d-1c57232d5224 · outbound

This paper cites Hamster: Hierarchicalactionmodelsforopen-worldrobotmanipulation.

G0.5: One Autoregressive Stream for Robot Reasoning and Action Hamster: Hierarchicalactionmodelsforopen-worldrobotmanipulation

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:57.667542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:36:56.024803Z digest=sha256:c6244b4da2f83672a156c1f157b51bb084446f8f5de882b8eaccdeeb1844e453

Observation e64ba385-8d49-4c8f-bff4-e58458c96197 · outbound

This paper cites Robotic Control via Embodied Chain-of-Thought Reasoning.

G0.5: One Autoregressive Stream for Robot Reasoning and Action Robotic Control via Embodied Chain-of-Thought Reasoning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:56.029284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:56.029284Z digest=sha256:a35a4ab3be6122bca956cd3a8595037dba8349539aac8779a164ef433587da91

Observation c1ca599c-a6eb-43b1-b760-3a5440ec15a2 · outbound

This paper cites Emma-x: An embodied multimodal action model with grounded chain of thought and look-ahead spatial reasoning.

G0.5: One Autoregressive Stream for Robot Reasoning and Action Emma-x: An embodied multimodal action model with grounded chain of thought and look-ahead spatial reasoning

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:57.650531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:36:56.033759Z digest=sha256:0dc29d40f3a2fbb519a08ceb2fd85fe4e1210f6c9da4b57329270d5ff281609c

Observation dfcc9841-b045-48b1-b8ac-c5ba5ddc5af5 · outbound

This paper cites Tracevla: Visual trace prompting enhances spatial-temporal awareness for generalist robotic policies.

G0.5: One Autoregressive Stream for Robot Reasoning and Action Tracevla: Visual trace prompting enhances spatial-temporal awareness for generalist robotic policies

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:57.636018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:36:56.037632Z digest=sha256:d9c494568b0b4c35c7a8f5f0c85d754c6dcbd055641c2b2072041cfe4289fc1a

Observation 0c10f5de-1a2d-4f14-ad6d-f924c74e88c5 · outbound

This paper cites an unresolved cited work.

G0.5: One Autoregressive Stream for Robot Reasoning and Action Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-16T00:36:57.621745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:36:56.041973Z digest=sha256:bf9352cf4afc5769d14ebd3ee8fa4088eb621f461d691c288c3bcc7c67703f5f

Observation 18d078ad-820f-431c-8f5e-11f50a2c9449 · outbound

This paper cites Minivla: A better vla with a smaller footprint, 2024.

G0.5: One Autoregressive Stream for Robot Reasoning and Action Minivla: A better vla with a smaller footprint, 2024

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:56.046109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:56.046109Z digest=sha256:2cb372c850e165a98358ebc9d6a5ca4b2f74bcff098d65172452eca9ba1e0951

Observation 55842ded-b04b-4160-9501-ab1b003a48a8 · outbound

This paper cites Actioncodec: What makes for good action tokenizers.arXiv preprint arXiv:2602.15397, 2026.

G0.5: One Autoregressive Stream for Robot Reasoning and Action Actioncodec: What makes for good action tokenizers.arXiv preprint arXiv:2602.15397, 2026

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:56.050256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:56.050256Z digest=sha256:2004e629a475543e1fad4f85a81c3b8ea017b07b17cbfbb158b70ecef8e2c742

Observation def8cd19-5178-4684-a2ad-24f0a36b3f3e · outbound

This paper cites FASTer: Toward powerful and efficient autoregressive vision–language–action models with learnableactiontokenizerandblock-wisedecoding.

G0.5: One Autoregressive Stream for Robot Reasoning and Action FASTer: Toward powerful and efficient autoregressive vision–language–action models with learnableactiontokenizerandblock-wisedecoding

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:57.598968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:36:56.054258Z digest=sha256:97a62db49a89f707fea79ebb3e50459e8974d6ca24dc766b556cc2030b28c63e

Observation 6330763b-1a08-4b67-98cf-93100ef78cdf · outbound

This paper cites ${\pi}_{0.7}$: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities.

G0.5: One Autoregressive Stream for Robot Reasoning and Action ${\pi}_{0.7}$: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:56.058688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:56.058688Z digest=sha256:0b390e417c9471992f83fb3203f2bbfebd3171017178f9e83e8400545a67e51c

Observation c61b6bc6-ffad-4425-8e66-20d8d76ae290 · outbound

This paper cites Gemini 3 pro model card, 2026.

G0.5: One Autoregressive Stream for Robot Reasoning and Action Gemini 3 pro model card, 2026

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:57.584095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:36:56.062853Z digest=sha256:91aa89371c35474e3649a6f4f14f4b5e41b2ced24be0580c450e8a636ae9964d

Observation c5d5b08e-8d74-4bff-8c67-d665754cc486 · outbound

This paper cites Seed 2.0 official launch, 2026.

G0.5: One Autoregressive Stream for Robot Reasoning and Action Seed 2.0 official launch, 2026

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:57.570693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:36:56.066778Z digest=sha256:9f3902781b9bdd1afe288c4fe9278dcad2800c9cd2a8cd0f1c78d4e107480636

Observation fb8ed6d8-9ab3-453f-be36-0c5c6dc4add6 · outbound

This paper cites SAM 3: Segment Anything with Concepts.

G0.5: One Autoregressive Stream for Robot Reasoning and Action SAM 3: Segment Anything with Concepts

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:56.070976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:56.070976Z digest=sha256:533577c493d161d4a58d27213f126151fec92c7b09b72b0f66823325923698b5

Observation 08764ba2-cd26-4b8a-8a8e-086090018625 · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge, January 2024.

G0.5: One Autoregressive Stream for Robot Reasoning and Action Llava-next: Improved reasoning, ocr, and world knowledge, January 2024

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:57.556683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:36:56.075505Z digest=sha256:0098d08e44e3da6c329ca5367648af1f9f6fbf813e79a7776bff8e25854fe414

Observation a864ae24-7f0f-460f-9d79-cc2202f82167 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

G0.5: One Autoregressive Stream for Robot Reasoning and Action LLaVA-OneVision: Easy Visual Task Transfer

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:56.079605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:56.079605Z digest=sha256:7421e8ae22b5e314dbd273cad0c7d04cb707765090718f0b0b4f265b0be9db6c

Observation b12fdf99-bef0-46b6-98e9-272d5d50a9c5 · outbound

This paper cites RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics.

G0.5: One Autoregressive Stream for Robot Reasoning and Action RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:56.083603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:56.083603Z digest=sha256:45a74f0cca8b49ee3aab5ac4f3adfa2c13e19c62d15dd782127d997d6cb0c4d0

Observation fcfeb2e4-5c44-42e2-a26e-8f9abcd78d03 · outbound

This paper cites MolmoAct: Action Reasoning Models that can Reason in Space.

G0.5: One Autoregressive Stream for Robot Reasoning and Action MolmoAct: Action Reasoning Models that can Reason in Space

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:56.087930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:56.087930Z digest=sha256:a0ed71e01b54c2cbb3b5dc6ef7658fe82f1ebf217f6c92e42745c79a5622717d

Observation 711b78ad-9685-4331-b31d-4b7f8ba580e5 · outbound

This paper cites RoboBrain: A Unified Brain Model for Robotic Manipulation from Abstract to Concrete.

G0.5: One Autoregressive Stream for Robot Reasoning and Action RoboBrain: A Unified Brain Model for Robotic Manipulation from Abstract to Concrete

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:56.092518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:56.092518Z digest=sha256:53b19143f986e6745feb22592b99d4ce993a42db42abf3c60c8ffc2694163a8b

Observation 20e85a7a-93e7-4229-a702-78c68aee629f · outbound

This paper cites DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset.

G0.5: One Autoregressive Stream for Robot Reasoning and Action DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:56.096943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:56.096943Z digest=sha256:09c38db7cfc4e2813a5ed2e9f4aa5622b079255b8ca3460bd85b566b2b227a67

Observation cbd6ad81-2b84-43fb-998e-1a6019684428 · outbound

This paper cites MolmoAct2: Action Reasoning Models for Real-world Deployment.

G0.5: One Autoregressive Stream for Robot Reasoning and Action MolmoAct2: Action Reasoning Models for Real-world Deployment

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:56.101107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:56.101107Z digest=sha256:7ec836b10bc50ccd46ff7937386acb4fdffa32c614198769a5e87abb42749c02

Observation 07313b5c-7309-4835-b187-d4095547961d · outbound

This paper cites Bridgedata v2: A dataset for robot learning at scale.

G0.5: One Autoregressive Stream for Robot Reasoning and Action Bridgedata v2: A dataset for robot learning at scale

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:56.105340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:56.105340Z digest=sha256:ba7801262ff90f9b99b392b877cb9a8d2568620bb6256f0747f917e91bc74c34

Observation f25edd1a-9048-4098-b340-bc88846a725d · outbound

This paper cites Starvla: A lego-like codebase for vision- language-action model developing, 2026.

G0.5: One Autoregressive Stream for Robot Reasoning and Action Starvla: A lego-like codebase for vision- language-action model developing, 2026

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:57.533657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:36:56.109557Z digest=sha256:b917b1a1ef52e54a54f150485ce850e6be3583fca74fe38e3334319d3eca69a0

Observation b14255fa-8191-4c86-aebe-383f6ba8736e · outbound

This paper cites Memoryvla: Perceptual-cognitive memory in vision-language-action models for robotic manipulation, 2025.

G0.5: One Autoregressive Stream for Robot Reasoning and Action Memoryvla: Perceptual-cognitive memory in vision-language-action models for robotic manipulation, 2025

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:57.518633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:36:56.113432Z digest=sha256:7e9abd5e939525828cd38a1e2e03f7ad462fddc6ca86d6800b86711242c759bd

Observation 52e3020b-d5c2-4547-a907-2bb516040913 · outbound

This paper cites Eo-1: An open unified embodied foundation model for general robot control.arXiv preprint arXiv:2508.21112, 2026.

G0.5: One Autoregressive Stream for Robot Reasoning and Action Eo-1: An open unified embodied foundation model for general robot control.arXiv preprint arXiv:2508.21112, 2026

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:56.117394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:56.117394Z digest=sha256:68377be490cb8c7fdc6d33d6a2b36bf6b0284bef1b4292dd518c5726353fa672

Observation 82b131a1-4540-4122-8476-084a12e47f8d · outbound

This paper cites Xiaomi-robotics-0: An open-sourced vision-language-action model with real-time execution.arXiv preprint arXiv:2602.12684, 2026.

G0.5: One Autoregressive Stream for Robot Reasoning and Action Xiaomi-robotics-0: An open-sourced vision-language-action model with real-time execution.arXiv preprint arXiv:2602.12684, 2026

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:56.121460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:56.121460Z digest=sha256:e6d721ce575a84de10fc656b7abced726aeee8b101e715c926b762c20ff12786

Observation 57c37a23-dd03-495c-bf76-86495884294d · outbound

This paper cites Motus: A Unified Latent Action World Model.

G0.5: One Autoregressive Stream for Robot Reasoning and Action Motus: A Unified Latent Action World Model

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:56.125563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:56.125563Z digest=sha256:9c2e0f8c49cca02bcef60b9c84dcb029ca870479cf041189f3ea43f4b8fecd4c

Observation 3d70e131-f2a7-4f5c-b0b1-c165da572692 · outbound

This paper cites Causal World Modeling for Robot Control.

G0.5: One Autoregressive Stream for Robot Reasoning and Action Causal World Modeling for Robot Control

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:56.129732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:56.129732Z digest=sha256:9808621d56df71368b4c7561c3d8424a51280ed58a00f4e3e7395f6fdf97369e

Observation 4fdca3a8-c7db-45a3-8229-e71a6c246cf1 · outbound

This paper cites A Pragmatic VLA Foundation Model.

G0.5: One Autoregressive Stream for Robot Reasoning and Action A Pragmatic VLA Foundation Model

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:56.133832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:56.133832Z digest=sha256:50d7018029d2e7ba6d38380a2ab8019c2014a3c2b47ba7100bc8ad3279cc9c1a

Observation 13325ed3-0e75-47b8-8ef1-5e0fbc3eaa2f · outbound

This paper cites Qwen-VLA: Unifying Vision-Language-Action Modeling across Tasks, Environments, and Robot Embodiments.

G0.5: One Autoregressive Stream for Robot Reasoning and Action Qwen-VLA: Unifying Vision-Language-Action Modeling across Tasks, Environments, and Robot Embodiments

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:56.138129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:56.138129Z digest=sha256:df2667152e5cc44be7e56e9848de72417aa1c6756ccf1fb513a62f932a924c07

Observation 99d173db-18b3-4ae3-bbd6-b2c9c362a1e6 · outbound

This paper cites InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy.

G0.5: One Autoregressive Stream for Robot Reasoning and Action InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:56.142437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:56.142437Z digest=sha256:bae794fa53da83e514ed2f3652faf053b3684b9838f8e23742d7372cb3fd1d89

Observation 7140d6cc-f9cc-433a-99f4-3cec222604c5 · outbound

This paper cites Igniting vlms toward the embodied space.arXiv preprint arXiv:2509.11766, 2025.

G0.5: One Autoregressive Stream for Robot Reasoning and Action Igniting vlms toward the embodied space.arXiv preprint arXiv:2509.11766, 2025

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:56.146862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:56.146862Z digest=sha256:759474514a8f0ff71ebc149648da5177af0a2230ca48397f5b765830cf5af1cc

Observation 82e07ac9-34f3-4fcf-9b67-e54122a22b80 · outbound

This paper cites Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success.

G0.5: One Autoregressive Stream for Robot Reasoning and Action Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:56.151108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:56.151108Z digest=sha256:7cfeb3ccd019debec2eb5e1f286c0f3dbfa5831c62a096a25deeec43eed9e428

Observation 29de1143-8027-4771-b45c-c47939b9b0d3 · outbound

This paper cites Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning.

G0.5: One Autoregressive Stream for Robot Reasoning and Action Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:56.155511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:56.155511Z digest=sha256:34f64d46db67fd27dbf8591d143c74ae299d81786da2d0f3e3e8da310db8bd3d

Observation b08c06af-78f7-4153-ab1d-7dcb50220005 · outbound

This paper cites Task adaptation of vision-language-action model: 1st place solution for the 2025 behavior challenge.arXiv preprint arXiv:2512.06951, 2025.

G0.5: One Autoregressive Stream for Robot Reasoning and Action Task adaptation of vision-language-action model: 1st place solution for the 2025 behavior challenge.arXiv preprint arXiv:2512.06951, 2025

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:56.160221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:56.160221Z digest=sha256:3049a8dc6eed9899d6324771d0b79cb250a4ca088b661769d847e8b064f4d825

Observation e2785799-22a6-40bd-99c0-33d023b9dab1 · outbound

This paper cites Openpi comet: Competition solution for 2025 behavior challenge.arXiv preprint arXiv:2512.10071, 2025.

G0.5: One Autoregressive Stream for Robot Reasoning and Action Openpi comet: Competition solution for 2025 behavior challenge.arXiv preprint arXiv:2512.10071, 2025

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:56.164151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:56.164151Z digest=sha256:f887b4c2442013f6ab4498fee1683a844f2eba1efdfa733329d88fd0197640a4

Observation 86aa10a3-011e-440d-8c3d-553e90849c7e · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

G0.5: One Autoregressive Stream for Robot Reasoning and Action DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:56.168485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:56.168485Z digest=sha256:8f767d154d81289da074d11aadf60b819708eb6b8e6fe2457e4a7440e349c89a

Observation b55cee17-9784-4f27-9e01-bc25f179af50 · outbound

This paper cites RLinf: Flexible and efficient large-scale reinforcement learning via macro-to-micro flow transformation.arXiv preprint arXiv:2509.15965, 2025.

G0.5: One Autoregressive Stream for Robot Reasoning and Action RLinf: Flexible and efficient large-scale reinforcement learning via macro-to-micro flow transformation.arXiv preprint arXiv:2509.15965, 2025

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:56.172379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:56.172379Z digest=sha256:7c13ff2d3650183453b2abb2aa629f1407b60104ec221a34b4ab6700b33bc3f0

Pith citing papers

No inbound Pith citation observations are available.