Pith. sign in

REVIEW 4 cited by

Safety Aware Task Planning via Large Language Models in Robotics

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.15707 v1 pith:H7JIDRSO submitted 2025-03-19 cs.RO cs.AI

classification cs.ROcs.AI
keywords safetytaskframeworkplanningsafermodelsplannerrobotic
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

The integration of large language models (LLMs) into robotic task planning has unlocked better reasoning capabilities for complex, long-horizon workflows. However, ensuring safety in LLM-driven plans remains a critical challenge, as these models often prioritize task completion over risk mitigation. This paper introduces SAFER (Safety-Aware Framework for Execution in Robotics), a multi-LLM framework designed to embed safety awareness into robotic task planning. SAFER employs a Safety Agent that operates alongside the primary task planner, providing safety feedback. Additionally, we introduce LLM-as-a-Judge, a novel metric leveraging LLMs as evaluators to quantify safety violations within generated task plans. Our framework integrates safety feedback at multiple stages of execution, enabling real-time risk assessment, proactive error correction, and transparent safety evaluation. We also integrate a control framework using Control Barrier Functions (CBFs) to ensure safety guarantees within SAFER's task planning. We evaluated SAFER against state-of-the-art LLM planners on complex long-horizon tasks involving heterogeneous robotic agents, demonstrating its effectiveness in reducing safety violations while maintaining task efficiency. We also verify the task planner and safety planner through actual hardware experiments involving multiple robots and a human.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LVLM-MPC Collaboration for Autonomous Driving: A Safety-Aware and Task-Scalable Control Architecture

    cs.RO 2025-05 conditional novelty 5.0 of 10

    A system that lets a large vision-language model propose driving maneuvers while a model predictive controller verifies and safely executes them, rejecting or assisting unsafe lane changes.

  2. HyCodePolicy: Hybrid Language Controllers for Multimodal Monitoring and Decision in Embodied Agents

    cs.RO 2025-08 conditional novelty 4.0 of 10

    HyCodePolicy closes the loop between generated robot code, visual checkpoint monitoring, and iterative repair, raising average success rates on 10 simulated manipulation tasks by up to 16.5 points over one-shot code-a...

  3. Grounding Language Models with Semantic Digital Twins for Robotic Planning

    cs.RO 2025-06 reject novelty 4.0 of 10

    The system grounds an LLM's action plans in hand-built semantic rules about a simulated home and reports success on all 14 selected ALFRED tasks.

  4. Semantic Intelligence: Integrating GPT-4 with A Planning in Low-Cost Robotics

    cs.RO 2025-05 conditional novelty 4.0 of 10

    A hybrid system where GPT-4 selects and adjusts A* paths lets a cheap quadruped follow semantic instructions like avoiding a toxic spill or collecting a resource first, with 90-100% success on the authors' tests.

Pith tools