{"total":27,"items":[{"citing_arxiv_id":"2607.06706","ref_index":16,"ref_count":1,"confidence":0.88,"is_internal_anchor":false,"paper_title":"Vision Language Action (VLA) Models for Unmanned Aerial Robotics and Bimanual Manipulation: A Review","primary_cat":"cs.RO","submitted_at":"2026-07-07T18:24:09+00:00","verdict":"ACCEPT","verdict_confidence":"HIGH","novelty_score":5.5,"formal_verification":"none","one_line_summary":"Bimanual VLA coordination strategies, training recipes, and continuous action chunking transfer to unmanned aerial systems; the survey maps 183 works and lists fourteen shared research directions.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.01111","ref_index":4,"ref_count":1,"confidence":0.88,"is_internal_anchor":false,"paper_title":"FAR: Failure-Aware Retry for Test-Time Recovery and Continual Policy Improvement","primary_cat":"cs.RO","submitted_at":"2026-07-01T16:01:54+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":4.0,"formal_verification":"none","one_line_summary":"FAR combines failure-contrastive preference adaptation with action perturbations for test-time recovery and continual policy improvement, reporting 17.6% and 11.7% success gains over diffusion policies in simulation and real-world manipulation tasks.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.29948","ref_index":44,"ref_count":1,"confidence":0.88,"is_internal_anchor":false,"paper_title":"Heterogeneous Tactile Transformer","primary_cat":"cs.RO","submitted_at":"2026-06-29T08:24:38+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"HTT learns shared representations across heterogeneous tactile sensors using a new paired dataset and pretraining objectives, enabling transfer to unseen sensors and tasks.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.25215","ref_index":17,"ref_count":1,"confidence":0.88,"is_internal_anchor":false,"paper_title":"Reflective VLA: In-Context Action Consequences Make VLAs Generalize","primary_cat":"cs.CV","submitted_at":"2026-06-23T22:23:35+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"Reflective VLA improves VLA generalization on LIBERO-Plus and LIBERO-Plus-Hard by 5.4 and 4.2 percentage points by conditioning on action consequences instead of reactive single-frame inputs.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.22142","ref_index":1,"ref_count":1,"confidence":0.88,"is_internal_anchor":false,"paper_title":"RoboLineage: Agent-Native Data Lifecycle Governance Across Robot Policy Iterations","primary_cat":"cs.RO","submitted_at":"2026-06-20T16:48:51+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":5.0,"formal_verification":"none","one_line_summary":"RoboLineage introduces an agent-native data lifecycle governance system that represents robot policy iteration steps as typed lineage artifacts to improve speed and auditability in real-robot workflows.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.20990","ref_index":11,"ref_count":1,"confidence":0.88,"is_internal_anchor":false,"paper_title":"Duet: Dual-Robot Understanding via Efficient Teaching","primary_cat":"cs.RO","submitted_at":"2026-06-18T23:47:27+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":5.0,"formal_verification":"none","one_line_summary":"DUET pretrains collaborative policies on human-human VR demonstrations then fine-tunes on minimal robot teleoperation data, achieving equal or better performance than robot-only baselines with 5.4x faster collection across four tasks.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.13355","ref_index":14,"ref_count":1,"confidence":0.88,"is_internal_anchor":false,"paper_title":"Real-Time Execution with Autoregressive Policies","primary_cat":"cs.RO","submitted_at":"2026-06-11T13:43:01+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":5.0,"formal_verification":"none","one_line_summary":"Autoregressive VLA policies achieve real-time execution via tokenization horizon adjustment and constrained decoding, outperforming flow-matching policies in speed and performance across simulated and real environments.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.13279","ref_index":4,"ref_count":1,"confidence":0.88,"is_internal_anchor":false,"paper_title":"See Selectively, Act Adaptively: Dual-Level Structural Decomposition for Bimanual Robot Manipulation","primary_cat":"cs.RO","submitted_at":"2026-06-11T12:33:55+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"A VLA policy using view-selective visual routing and interaction-aware action MoE improves average success by 27.7% in simulation and 43.3% in real-world bimanual tasks over monolithic baselines.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.12406","ref_index":51,"ref_count":1,"confidence":0.88,"is_internal_anchor":false,"paper_title":"FACTR 2: Learning External Force Sensing for Commodity Robot Arms Improves Policy Learning","primary_cat":"cs.RO","submitted_at":"2026-06-10T17:59:35+00:00","verdict":null,"verdict_confidence":null,"novelty_score":null,"formal_verification":null,"one_line_summary":null,"context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.12403","ref_index":11,"ref_count":1,"confidence":0.88,"is_internal_anchor":false,"paper_title":"World Pilot: Steering Vision-Language-Action Models with World-Action Priors","primary_cat":"cs.RO","submitted_at":"2026-06-10T17:59:08+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":5.0,"formal_verification":"none","one_line_summary":"World Pilot augments VLA policies with world-action priors through latent and action steering pathways, reporting 84.7% success on LIBERO-Plus zero-shot OOD and top real-robot results across four tasks.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.12334","ref_index":58,"ref_count":1,"confidence":0.88,"is_internal_anchor":false,"paper_title":"Fourier Features Let Agents Learn High Precision Policies with Imitation Learning","primary_cat":"cs.LG","submitted_at":"2026-06-10T17:05:50+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"Mapping point clouds to Fourier features improves high-precision imitation learning policies on RoboCasa, ManiSkill3, and real-robot tasks compared with Cartesian inputs.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.12497","ref_index":76,"ref_count":1,"confidence":0.88,"is_internal_anchor":false,"paper_title":"$\\mu$VLA: On Recurrent Memory for Partially Observable Manipulation in VLA Models","primary_cat":"cs.LG","submitted_at":"2026-06-10T13:26:40+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"Adding recurrent memory tokens to VLA models raises success rates on partially observable manipulation tasks from 0.42 to 0.84 on training and 0.07 to 0.23 on held-out tasks while preserving performance under full observability.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.09056","ref_index":10,"ref_count":1,"confidence":0.88,"is_internal_anchor":false,"paper_title":"MilliVid: Hierarchical Latents for Long-Range Consistency in Video Generation","primary_cat":"cs.CV","submitted_at":"2026-06-08T05:46:01+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":5.0,"formal_verification":"none","one_line_summary":"MilliVid compresses video frames into multi-scale token hierarchies and uses coarse-to-fine rollout in a diffusion model to maintain long-range geometric and object consistency on Minecraft videos.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.02274","ref_index":33,"ref_count":1,"confidence":0.88,"is_internal_anchor":false,"paper_title":"Dexterity-BEV: Aligning 3D World and Actions for Generalizable Robot Policies Learning","primary_cat":"cs.RO","submitted_at":"2026-06-01T14:01:11+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"Dexterity-BEV creates 3D vertex-based inputs and BEV-aligned outputs to reduce spatial-temporal misalignments in end-to-end robot policies trained on diverse datasets and embodiments.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.01238","ref_index":4,"ref_count":1,"confidence":0.88,"is_internal_anchor":false,"paper_title":"Training-Free Imitation Learning with Closed-Form Diffusion Policies","primary_cat":"cs.RO","submitted_at":"2026-05-31T13:40:47+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"Closed-Form Diffusion Policies enable training-free imitation learning by using closed-form scores derived from demonstration data, achieving competitive benchmark performance with millisecond inference and composable editing of pre-trained policies.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.00515","ref_index":10,"ref_count":1,"confidence":0.88,"is_internal_anchor":false,"paper_title":"PaCo-VLA: Passivity-Shielded Compliance Prior for Contact-Rich Vision-Language-Action Manipulation","primary_cat":"cs.RO","submitted_at":"2026-05-30T04:06:39+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":5.0,"formal_verification":"none","one_line_summary":"PaCo-VLA adds an independent passivity shield to VLA outputs so that semantic proposals for compliance and admittance can be used in contact-rich tasks without violating passivity or causing damage.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.24152","ref_index":43,"ref_count":1,"confidence":0.88,"is_internal_anchor":false,"paper_title":"Neuro-Inspired Inverse Learning for Planning and Control","primary_cat":"cs.AI","submitted_at":"2026-05-22T19:19:32+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"The Inverter framework formalizes inverse learning to generate coherent multi-step trajectories, outperforming offline RL and diffusion baselines on D4RL maze tasks by 24% on average with 10-100x less inference time while also matching GRAPE fidelity on single-qubit gates at >1000x speed.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.21811","ref_index":58,"ref_count":1,"confidence":0.88,"is_internal_anchor":false,"paper_title":"Safe and Steerable Geometric Motion Policies for Robotic Dexterous Manipulation","primary_cat":"cs.RO","submitted_at":"2026-05-20T23:16:01+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"SafePBDS uses pullback control barrier functions and a task manifold action interface to generate certifiably safe, steerable motions on high-DOF robots from objectives defined on arbitrary geometric spaces.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.19924","ref_index":21,"ref_count":1,"confidence":0.88,"is_internal_anchor":false,"paper_title":"RoHIL: Robust Human-in-the-Loop Robotic Reinforcement Learning Against Illumination Variations","primary_cat":"cs.RO","submitted_at":"2026-05-19T14:47:38+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"RoHIL adapts human-in-the-loop RL policies to new illumination conditions offline by combining world-model image relighting, illumination-retention replay, and anchored Bellman regularisation, improving shifted-light performance while preserving source performance on four real-robot tasks.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.18287","ref_index":45,"ref_count":1,"confidence":0.88,"is_internal_anchor":false,"paper_title":"StableVLA: Towards Robust Vision-Language-Action Models without Extra Data","primary_cat":"cs.CV","submitted_at":"2026-05-18T12:15:16+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":5.0,"formal_verification":"none","one_line_summary":"StableVLA adds an Information Bottleneck Adapter to VLA models that improves robustness to visual corruptions by 30% on average with under 10M extra parameters and no extra data, even when using a much smaller backbone.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.15536","ref_index":40,"ref_count":1,"confidence":0.88,"is_internal_anchor":false,"paper_title":"SkiP: When to Skip and When to Refine for Efficient Robot Manipulation","primary_cat":"cs.RO","submitted_at":"2026-05-15T02:16:34+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"SkiP introduces action relabeling and Motion Spectrum Keying to skip redundant steps in robot trajectories, cutting executed steps by 15-40% while maintaining success rates across 72 simulated and 3 real tasks.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.09989","ref_index":10,"ref_count":2,"confidence":0.88,"is_internal_anchor":false,"paper_title":"StereoPolicy: Improving Robotic Manipulation Policies via Stereo Perception","primary_cat":"cs.RO","submitted_at":"2026-05-11T05:06:12+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":4.0,"formal_verification":"none","one_line_summary":"StereoPolicy fuses left-right image features via cross-attention to deliver consistent gains over RGB, RGB-D, point cloud, and multi-view baselines in simulation and real-robot manipulation tasks.","context_count":1,"top_context_role":"background","top_context_polarity":"background","context_text":"Research, 2024. 10 [9] T. Z. Zhao, V . Kumar, S. Levine, and C. Finn. Learning fine-grained bimanual manipulation with low-cost hardware. In K. E. Bekris, K. Hauser, S. L. Herbert, and J. Yu, editors,Robotics: Science and Systems XIX, Daegu, Republic of Korea, July 10-14, 2023, 2023. doi:10.15607/ RSS.2023.XIX.016. URLhttps://doi.org/10.15607/RSS.2023.XIX.016. [10] M. J. Kim, K. Pertsch, S. Karamcheti, T. Xiao, A. Balakrishna, S. Nair, R. Rafailov, E. P. Foster, P. R. Sanketi, Q. Vuong, T. Kollar, B. Burchfiel, R. Tedrake, D. Sadigh, S. Levine, P. Liang, and C. Finn. OpenVLA: An open-source vision-language-action model. In8th Annual Conference on Robot Learning, 2024. URLhttps://openreview.net/forum?id=ZMnD6QZAE6."},{"citing_arxiv_id":"2604.24018","ref_index":67,"ref_count":1,"confidence":0.88,"is_internal_anchor":false,"paper_title":"Betting for Sim-to-Real Performance Evaluation","primary_cat":"cs.RO","submitted_at":"2026-04-27T03:58:50+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"Betting mechanisms can yield provably more accurate and efficient estimates of real-world robot behavior than Monte Carlo sampling under specified conditions, with practical approximations demonstrated on synthetic data and a robotic manipulator task.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2604.23570","ref_index":50,"ref_count":1,"confidence":0.88,"is_internal_anchor":false,"paper_title":"EgoLive: A Large-Scale Egocentric Dataset from Real-World Human Tasks","primary_cat":"cs.RO","submitted_at":"2026-04-26T07:21:15+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":4.0,"formal_verification":"none","one_line_summary":"EgoLive is presented as the largest open-source annotated egocentric dataset for real-world task-oriented human routines, captured with a custom head-mounted device and multi-modal annotations exclusively in unconstrained environments.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2602.08167","ref_index":103,"ref_count":1,"confidence":0.88,"is_internal_anchor":false,"paper_title":"Self-Supervised Bootstrapping of Action-Predictive Embodied Reasoning","primary_cat":"cs.RO","submitted_at":"2026-02-09T00:10:17+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"R&B-EnCoRe uses self-supervised importance-weighted variational inference to distill action-predictive reasoning datasets that improve VLA performance on manipulation, navigation, and driving tasks without external verifiers.","context_count":1,"top_context_role":"method","top_context_polarity":"use_method","context_text":"a reasoning dropout rate ofd= 0.2, with the same7primitives as in Section V-A, where the reasoning content was generated in [11] by the Gemini 1.0 [75] model. The actionAconsists of seven discrete tokens based on the tokenization of [5, 37]. To support high-throughput repeated sampling when generating our synthetic reasoning candidate for refinement, we extended SGLang-VLA [103, 104] to support rapid inferencing of reasoning-based [11, 12] Prismatic Models [5, 28]. Q3How effective canR&B-EnCoReenable manipulation VLAs to generalize when test-time reasoning is suppressed for speed, compared to reasoning with all primitives? Since generating textual reasoning traces at test-time can dras- tically increase latency (often taking several seconds per step),"},{"citing_arxiv_id":"2601.21971","ref_index":38,"ref_count":1,"confidence":0.88,"is_internal_anchor":false,"paper_title":"Supervised Mixture-of-Experts for Surgical Grasping and Retraction","primary_cat":"cs.RO","submitted_at":"2026-01-29T16:50:14+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"Supervised MoE on top of ACT achieves higher success in bowel grasping/retraction from <150 demos than standard ACT or generalist VLAs, with OOD robustness, unseen viewpoint generalization, and zero-shot ex vivo porcine transfer.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2505.03233","ref_index":88,"ref_count":1,"confidence":0.88,"is_internal_anchor":false,"paper_title":"GraspVLA: a Grasping Foundation Model Pre-trained on Billion-scale Synthetic Action Data","primary_cat":"cs.RO","submitted_at":"2025-05-06T06:59:28+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"GraspVLA shows that pretraining a grasping model on a billion synthetic action frames enables zero-shot open-vocabulary performance and sim-to-real transfer.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null}],"limit":50,"offset":0}