REVIEW 4 cited by
Back to Newton's Laws: Learning Vision-based Agile Flight via Differentiable Physics
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
abstract
Swarm navigation in cluttered environments is a grand challenge in robotics. This work combines deep learning with first-principle physics through differentiable simulation to enable autonomous navigation of multiple aerial robots through complex environments at high speed. Our approach optimizes a neural network control policy directly by backpropagating loss gradients through the robot simulation using a simple point-mass physics model and a depth rendering engine. Despite this simplicity, our method excels in challenging tasks for both multi-agent and single-agent applications with zero-shot sim-to-real transfer. In multi-agent scenarios, our system demonstrates self-organized behavior, enabling autonomous coordination without communication or centralized planning - an achievement not seen in existing traditional or learning-based methods. In single-agent scenarios, our system achieves a 90% success rate in navigating through complex environments, significantly surpassing the 60% success rate of the previous state-of-the-art approach. Our system can operate without state estimation and adapt to dynamic obstacles. In real-world forest environments, it navigates at speeds up to 20 m/s, doubling the speed of previous imitation learning-based solutions. Notably, all these capabilities are deployed on a budget-friendly $21 computer, costing less than 5% of a GPU-equipped board used in existing systems. Video demonstrations are available at https://youtu.be/LKg9hJqc2cc.
Forward citations
Cited by 4 Pith papers
-
CORB-Planner: Corridor as Observations for RL Planning in High-Speed Flight
CORB-Planner uses safe flight corridors as low-dimensional observations for an RL policy that generates B-spline control points, enabling real-time cross-platform UAV planning after about ten minutes of training.
-
HEPP: Hyper-efficient Perception and Planning for High-speed Obstacle Avoidance of UAVs
A three-module mapping, search, and trajectory optimization system lets a LiDAR drone avoid obstacles at up to 15 m/s in dense simulated forests and over 11 m/s outdoors.
-
Reactive Aerobatic Flight via Reinforcement Learning
An RL policy trained with an automated curriculum and domain randomization lets a quadrotor perform continuous inverted flight while reactively navigating a moving gate.
-
ABPT: Amended Backpropagation through Time with Partially Differentiable Rewards
ABPT averages a zero-step value gradient with an N-step backpropagation gradient so that non-differentiable reward components do not fully block policy learning.
Discussion (0). Continue with ORCID to comment.