{"total":16,"items":[{"citing_arxiv_id":"2606.23686","ref_index":12,"ref_count":2,"confidence":0.9,"is_internal_anchor":false,"paper_title":"LIBERO-Safety: A Comprehensive Benchmark for Physical and Semantic Safety in Vision-Language-Action Models","primary_cat":"cs.RO","submitted_at":"2026-06-22T17:59:53+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"LIBERO-Safety supplies a scalable benchmark, data-generation pipeline, and 19,664-demonstration dataset that exposes a generalization-safety tension in current VLA models where diverse training improves collision avoidance but task success stays limited by trajectory quality and semantic understandi","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.20905","ref_index":19,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Vesta: A Generalist Embodied Reasoning Model","primary_cat":"cs.RO","submitted_at":"2026-06-18T20:01:32+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"Vesta is a unified embodied generalist model that outperforms specialist baselines by over 20% on average and improves real-world robotic task success by over 35%.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.19965","ref_index":21,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"ROSE: Benchmarking the Perception-to-Action Gap in Multimodal Models","primary_cat":"cs.CV","submitted_at":"2026-06-18T09:05:48+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"ROSE benchmark shows MLLMs drop up to 44.5 percentage points from counting tasks to region-conditioned action on identical scenes, with the gap persisting even when counts are correct.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.17539","ref_index":9,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Reinforcing Dual-Path Reasoning in Spatial Vision Language Models","primary_cat":"cs.CV","submitted_at":"2026-06-16T05:32:39+00:00","verdict":"UNVERDICTED","verdict_confidence":"UNKNOWN","novelty_score":6.0,"formal_verification":"none","one_line_summary":"SR-REAL equips spatial VLMs with dual LOR and DTR reasoning paths trained via RL, achieving better benchmark performance through mutual reinforcement and generalization without per-task tuning.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.15476","ref_index":25,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"FARM: Find Anything using Relational Spatial Memory","primary_cat":"cs.RO","submitted_at":"2026-06-13T21:21:24+00:00","verdict":"CONDITIONAL","verdict_confidence":"MODERATE","novelty_score":6.0,"formal_verification":"none","one_line_summary":"A real-time relational spatial memory that parses object queries into spatial predicates, scores them against per-object 3D Gaussians, and retrieves object instances with substantially higher top-K recall than prior scene-graph or video-VLM baselines.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.13497","ref_index":9,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"SPARC: Reliable Spatial Annotations from Robot Demonstrations at Scale","primary_cat":"cs.RO","submitted_at":"2026-06-11T15:46:28+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":5.0,"formal_verification":"none","one_line_summary":"SPARC generates reliable spatial annotations for robot demonstrations by leveraging spatio-temporal task structure, outperforming detection baselines on localization accuracy while retaining more samples and enabling competitive model performance without manual annotations.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.12402","ref_index":55,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"DIRECT: When and Where Should You Allocate Test-Time Compute in Embodied Planners?","primary_cat":"cs.RO","submitted_at":"2026-06-10T17:58:49+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"DIRECT is a multimodal-context router that allocates test-time compute across chain-of-thought depth, model size, and memory history for VLM embodied planners, improving the success-cost Pareto frontier and matching stronger models at up to 65% lower latency on benchmarks and a physical Franka arm.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.05979","ref_index":17,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"World-Language-Action Model for Unified World Modeling, Language Reasoning, and Action Synthesis","primary_cat":"cs.RO","submitted_at":"2026-06-04T10:23:01+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":5.0,"formal_verification":"none","one_line_summary":"WLA models use an autoregressive Transformer to jointly predict textual subtasks, subgoal images, and robot actions from instructions, images, and states, reporting SOTA success rates on RoboTwin2.0 and RMBench.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.03890","ref_index":8,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"OVO-S-Bench: A Hierarchical Benchmark for Streaming Spatial Intelligence in Multimodal LLMs","primary_cat":"cs.CV","submitted_at":"2026-06-02T16:51:32+00:00","verdict":"ACCEPT","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"OVO-S-Bench provides 1680 human-annotated questions on 348 videos to measure streaming spatial intelligence in MLLMs across instantaneous perception, spatiotemporal tracking, spatial simulation, and allocentric mapping.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.01810","ref_index":9,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Token Predictors Are Not Planners: Building Physically Grounded Causal Reasoners","primary_cat":"cs.AI","submitted_at":"2026-06-01T07:27:39+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"Introduces a new diagnostic benchmark and million-scale reasoning corpus showing that training on explicit causal traces improves next-state prediction in embodied planning, with reported gains from data scaling.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.28548","ref_index":17,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"GEM: Generative Supervision Helps Embodied Intelligence","primary_cat":"cs.CV","submitted_at":"2026-05-27T14:39:42+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":5.0,"formal_verification":"none","one_line_summary":"GEM adds generative depth supervision to VLM pre-training and reports improved results on embodied benchmarks plus real-world robot execution.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.22536","ref_index":83,"ref_count":2,"confidence":0.9,"is_internal_anchor":false,"paper_title":"SpaceDG: Benchmarking Spatial Intelligence under Visual Degradation","primary_cat":"cs.CV","submitted_at":"2026-05-21T14:25:15+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"SpaceDG is the first large-scale benchmark dataset (~1M QA pairs) simulating nine visual degradations in 3DGS-rendered scenes to measure and improve spatial intelligence robustness in MLLMs.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.08747","ref_index":43,"ref_count":3,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Done, But Not Sure: Disentangling World Completion from Self-Termination in Embodied Agents","primary_cat":"cs.AI","submitted_at":"2026-05-09T07:24:07+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"VIGIL decouples world-state completion from terminal commitment in embodied agents, exposing up to 19.7 pp gaps in benchmark success despite comparable execution across 20 models.","context_count":1,"top_context_role":"background","top_context_polarity":"background","context_text":"[41] Kimi Team, Angang Du, et al. Kimi-VL Technical Report.arXiv preprint arXiv:2504.07491, 2025. [42] Yuheng Ji, Huajie Tan, Jiayu Shi, Xiaoshuai Hao, Yuan Zhang, Hengyuan Zhang, Pengwei Wang, Mengdi Zhao, Yao Mu, Pengju An, et al. RoboBrain: A Unified Brain Model for Robotic Manipula- tion from Abstract to Concrete.arXiv preprint arXiv:2502.21257, 2025. [43] Ronghao Dang, Jiayan Guo, Bohan Hou, Sicong Leng, Kehan Li, Xin Li, Jiangping Liu, Yunxuan Mao, Zhikai Wang, Yuqian Yuan, Minghao Zhu, Xiao Lin, Yang Bai, Qian Jiang, Yaxi Zhao, Minghua Zeng, Junlong Gao, Yuming Jiang, Jun Cen, Siteng Huang, Liuyi Wang, Wenqiao Zhang, Chengju Liu, Jianfei Yang, Shijian Lu, and Deli Zhao. RynnBrain: Open Embodied Foundation"},{"citing_arxiv_id":"2605.00080","ref_index":13,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"World Model for Robot Learning: A Comprehensive Survey","primary_cat":"cs.RO","submitted_at":"2026-04-30T14:35:31+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":3.0,"formal_verification":"none","one_line_summary":"A comprehensive survey that organizes the literature on world models in robot learning, their roles in policy learning, planning, simulation, and video-based generation, with connections to navigation, driving, datasets, and benchmarks.","context_count":1,"top_context_role":"background","top_context_polarity":"background","context_text":"architectural conclusion. At a high level, this family replaces the two-stage factorization of \"predict first, then act\" with a unified multimodal generative objective. Letx= [ zv; za]denote the concatenation of future visual and action representations. A shared backbonefθ is trained on corrupted inputs˜xτ under conditioning(o t, l) ˆy=fθ(˜xτ , ot, l, τ),x= [z v;z a],(13) whereτmeans denoising steps, with a generic unified objective Lunified =E \u0002 ℓ(ˆy, y) \u0003 ,(14) where the exact target depends on the specific instantiation:y may correspond to diffusion noise in continuous denoising models, a velocity field in flow-matching variants, or masked tokens in discrete denoising formulations. Representative early unified designs such as UVA (Li et al."},{"citing_arxiv_id":"2604.21924","ref_index":16,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Long-Horizon Manipulation via Trace-Conditioned VLA Planning","primary_cat":"cs.RO","submitted_at":"2026-04-23T17:59:04+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"LoHo-Manip enables robust long-horizon robot manipulation by using a receding-horizon VLM manager to output progress-aware subtask sequences and 2D visual traces that condition a VLA executor for automatic replanning.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2604.11789","ref_index":36,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation","primary_cat":"cs.CV","submitted_at":"2026-04-13T17:55:02+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":3.0,"formal_verification":"none","one_line_summary":"This review organizes literature on large multimodal models and object-centric vision into four themes—understanding, referring segmentation, editing, and generation—while summarizing paradigms, strategies, and challenges like instance permanence and consistent interaction.","context_count":1,"top_context_role":"background","top_context_polarity":"background","context_text":"When integrated with LMMs, this perspective shifts multimodal systems from passive scene description toward more actionable and controllable capabilities: identifying objects, segmenting them, editing them, and generating them under explicit object-level constraints. Such capabilities are central to a wide range of applications, including interactive agents [44, 199], embodied intelligence [36, 37, 62, 168, 200, 218], 4D reasoning [66, 209], visual design [13, 176] and medical analysis [93], where object faithfulness, spatial accuracy, and controllability are essential. In this paper, we organize the emerging landscape ofobject-centric vision in the era of large multimodal modelsinto four closely related research themes. Object-Centric Visual Understandingforms the foundation."}],"limit":50,"offset":0}