VLATIM benchmark reveals large VLMs excel at high-level planning in physics puzzles but struggle with precise visual grounding and mouse control, so they lack human-like problem-solving capabilities.
ABot-Claw: A Foundation for Persistent, Cooperative, and Self-Evolving Robotic Agents
4 Pith papers cite this work. Polarity classification is still indexing.
abstract
Current embodied intelligent systems still face a substantial gap between high-level reasoning and low-level physical execution in open-world environments. Although Vision-Language-Action (VLA) models provide strong perception and intuitive responses, their open-loop nature limits long-horizon performance. Agents incorporating System 2 cognitive mechanisms improve planning, but usually operate in closed sandboxes with predefined toolkits and limited real-system control. OpenClaw provides a localized runtime with full system privileges, but lacks the embodied control architecture required for long-duration, multi-robot execution. We therefore propose ABot-Claw, an embodied extension of OpenClaw that integrates: 1) a unified embodiment interface with capability-driven scheduling for heterogeneous robot coordination; 2) a visual-centric cross-embodiment multimodal memory for persistent context retention and grounded retrieval; and 3) a critic-based closed-loop feedback mechanism with a generalist reward model for online progress evaluation, local correction, and replanning. With a decoupled architecture spanning the OpenClaw layer, shared service layer, and robot embodiment layer, ABot-Claw enables real-world interaction, closes the loop from natural language intent to physical action, and supports progressively self-evolving robotic agents in open, dynamic environments.
years
2026 4representative citing papers
Aligning temporal granularity, action subspaces, and train-test conditioning yields SOTA long-horizon mobile and fine-grained manipulation success for a unified world-action model.
MindClaw adds belief memory and a skill-guided cognitive trigger so an embodied agent intervenes only when false beliefs or hidden goals block progress, beating direct VLMs on MindPower intervention metrics.
AerialClaw provides a modular open-source framework for LLM-driven UAV agents using a brain-skill-runtime architecture with hard and soft skills, memory reflection, and simulation support.
citing papers explorer
-
Do Vision-Language-Models show human-like logical problem-solving capability in point and click puzzle games?
VLATIM benchmark reveals large VLMs excel at high-level planning in physics puzzles but struggle with precise visual grounding and mouse control, so they lack human-like problem-solving capabilities.
-
ABot-M0.5: Unified Mobility-and-Manipulation World Action Model
Aligning temporal granularity, action subspaces, and train-test conditioning yields SOTA long-horizon mobile and fine-grained manipulation success for a unified world-action model.
-
MindClaw: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
MindClaw adds belief memory and a skill-guided cognitive trigger so an embodied agent intervenes only when false beliefs or hidden goals block progress, beating direct VLMs on MindPower intervention metrics.
-
AerialClaw: An Open-Source Framework for LLM-Driven Autonomous Aerial Agents
AerialClaw provides a modular open-source framework for LLM-driven UAV agents using a brain-skill-runtime architecture with hard and soft skills, memory reflection, and simulation support.