{"total":13,"items":[{"citing_arxiv_id":"2607.06706","ref_index":6,"ref_count":1,"confidence":0.88,"is_internal_anchor":false,"paper_title":"Vision Language Action (VLA) Models for Unmanned Aerial Robotics and Bimanual Manipulation: A Review","primary_cat":"cs.RO","submitted_at":"2026-07-07T18:24:09+00:00","verdict":"ACCEPT","verdict_confidence":"HIGH","novelty_score":5.5,"formal_verification":"none","one_line_summary":"Bimanual VLA coordination strategies, training recipes, and continuous action chunking transfer to unmanned aerial systems; the survey maps 183 works and lists fourteen shared research directions.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.27355","ref_index":24,"ref_count":1,"confidence":0.88,"is_internal_anchor":false,"paper_title":"RouterVLA: Turning Smoke Tests into Supervision for Heterogeneous VLA Selection","primary_cat":"cs.RO","submitted_at":"2026-06-25T17:56:33+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":4.0,"formal_verification":"none","one_line_summary":"RouterVLA reports that a simple probe-success rule from outcome-separated smoke tests raises held-out VLA success by 14.64pp on 34,752 LIBERO-Plus records, with learned scorers adding no further gain.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.27251","ref_index":29,"ref_count":1,"confidence":0.88,"is_internal_anchor":false,"paper_title":"Advancing Omnimodal Embodied Agents from Isolated Skills to Everyday Physical Autonomy","primary_cat":"cs.RO","submitted_at":"2026-06-25T16:36:35+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":5.0,"formal_verification":"none","one_line_summary":"OmniAct framework integrates planning, memory, and verification to enable persistent autonomy in omnimodal embodied agents, showing improved success and stable context in 40 real-world tasks.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.25215","ref_index":14,"ref_count":1,"confidence":0.88,"is_internal_anchor":false,"paper_title":"Reflective VLA: In-Context Action Consequences Make VLAs Generalize","primary_cat":"cs.CV","submitted_at":"2026-06-23T22:23:35+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"Reflective VLA improves VLA generalization on LIBERO-Plus and LIBERO-Plus-Hard by 5.4 and 4.2 percentage points by conditioning on action consequences instead of reactive single-frame inputs.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.22142","ref_index":13,"ref_count":1,"confidence":0.88,"is_internal_anchor":false,"paper_title":"RoboLineage: Agent-Native Data Lifecycle Governance Across Robot Policy Iterations","primary_cat":"cs.RO","submitted_at":"2026-06-20T16:48:51+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":5.0,"formal_verification":"none","one_line_summary":"RoboLineage introduces an agent-native data lifecycle governance system that represents robot policy iteration steps as typed lineage artifacts to improve speed and auditability in real-robot workflows.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.13279","ref_index":15,"ref_count":1,"confidence":0.88,"is_internal_anchor":false,"paper_title":"See Selectively, Act Adaptively: Dual-Level Structural Decomposition for Bimanual Robot Manipulation","primary_cat":"cs.RO","submitted_at":"2026-06-11T12:33:55+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"A VLA policy using view-selective visual routing and interaction-aware action MoE improves average success by 27.7% in simulation and 43.3% in real-world bimanual tasks over monolithic baselines.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.12497","ref_index":19,"ref_count":1,"confidence":0.88,"is_internal_anchor":false,"paper_title":"$\\mu$VLA: On Recurrent Memory for Partially Observable Manipulation in VLA Models","primary_cat":"cs.LG","submitted_at":"2026-06-10T13:26:40+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"Adding recurrent memory tokens to VLA models raises success rates on partially observable manipulation tasks from 0.42 to 0.84 on training and 0.07 to 0.23 on held-out tasks while preserving performance under full observability.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.07895","ref_index":39,"ref_count":1,"confidence":0.88,"is_internal_anchor":false,"paper_title":"TBD-VLA: Temporal Block Diffusion Vision Language Action Model","primary_cat":"cs.CV","submitted_at":"2026-06-05T23:10:43+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":5.0,"formal_verification":"none","one_line_summary":"TBD-VLA partitions action sequences into temporal blocks, performs masked discrete diffusion within blocks, and autoregressive generation across blocks to unify parallel decoding with temporal coherence in discrete VLA models.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.05737","ref_index":11,"ref_count":1,"confidence":0.88,"is_internal_anchor":false,"paper_title":"Let It Be Simple: One-Step Action Generation for Vision-Language-Action Models","primary_cat":"cs.CV","submitted_at":"2026-06-04T05:58:30+00:00","verdict":"CONDITIONAL","verdict_confidence":"HIGH","novelty_score":5.0,"formal_verification":"none","one_line_summary":"High-noise flow-matching training makes one-step VLA action decoding competitive with multi-step decoding because actions are compact targets under rich observations.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.02274","ref_index":38,"ref_count":1,"confidence":0.88,"is_internal_anchor":false,"paper_title":"Dexterity-BEV: Aligning 3D World and Actions for Generalizable Robot Policies Learning","primary_cat":"cs.RO","submitted_at":"2026-06-01T14:01:11+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"Dexterity-BEV creates 3D vertex-based inputs and BEV-aligned outputs to reduce spatial-temporal misalignments in end-to-end robot policies trained on diverse datasets and embodiments.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.00515","ref_index":7,"ref_count":1,"confidence":0.88,"is_internal_anchor":false,"paper_title":"PaCo-VLA: Passivity-Shielded Compliance Prior for Contact-Rich Vision-Language-Action Manipulation","primary_cat":"cs.RO","submitted_at":"2026-05-30T04:06:39+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":5.0,"formal_verification":"none","one_line_summary":"PaCo-VLA adds an independent passivity shield to VLA outputs so that semantic proposals for compliance and admittance can be used in contact-rich tasks without violating passivity or causing damage.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.20299","ref_index":17,"ref_count":1,"confidence":0.88,"is_internal_anchor":false,"paper_title":"Mechanisms of Misgeneralization in Physical Sequence Modeling","primary_cat":"cs.LG","submitted_at":"2026-05-19T12:34:16+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"Generative sequence models for physical tasks exhibit physical misgeneralization where local prediction errors propagate through physical measurements to distort aggregate distributions over quantities like distance or energy; a data deviation kernel explains and predicts the shifts and supports a内核","context_count":1,"top_context_role":"background","top_context_polarity":"background","context_text":"The adaptive bandwidth prevents a near-exact match to a reference trajectory from collapsing the posterior to a degenerate point mass. When we report a single recovered quantity value rather than a posterior overr, we use the posterior mode. A.2.2 Double-pendulum Energy For a double-pendulum angle trajectory sampled with ∆t= 0.01 , we recover energy by first estimating angular velocities with central differences, ˙qt = qt+1 −q t−1 2∆t ,(17) and compute Et = 1 2 ˙q⊤ t M(q t) ˙qt +V(q t)−V min,(18) at each interior timestep, where Vmin is the potential energy at the downward equilibrium. We then take the median over time. The median suppresses occasional spikes from finite differences caused by local irregularities in generated angles. 14 A.2.3 Maze2D Path Length For a Maze2D position trajectory x= (q 0, ."},{"citing_arxiv_id":"2506.10137","ref_index":17,"ref_count":1,"confidence":0.88,"is_internal_anchor":false,"paper_title":"Self-Predictive Representations for Combinatorial Generalization in Behavioral Cloning","primary_cat":"cs.LG","submitted_at":"2025-06-11T19:32:41+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":5.0,"formal_verification":"none","one_line_summary":"BYOL-γ uses self-predictive representations to approximate successor representations, improving zero-shot combinatorial generalization in goal-conditioned behavioral cloning.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null}],"limit":50,"offset":0}