Pith. sign in

REVIEW 4 cited by

Demonstrating Agile Flight from Pixels without State Estimation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.12505 v1 pith:MH4JCTUW submitted 2024-06-18 cs.RO

classification cs.RO
keywords agiletrainingcontroldroneduringestimationflightstate
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Quadrotors are among the most agile flying robots. Despite recent advances in learning-based control and computer vision, autonomous drones still rely on explicit state estimation. On the other hand, human pilots only rely on a first-person-view video stream from the drone onboard camera to push the platform to its limits and fly robustly in unseen environments. To the best of our knowledge, we present the first vision-based quadrotor system that autonomously navigates through a sequence of gates at high speeds while directly mapping pixels to control commands. Like professional drone-racing pilots, our system does not use explicit state estimation and leverages the same control commands humans use (collective thrust and body rates). We demonstrate agile flight at speeds up to 40km/h with accelerations up to 2g. This is achieved by training vision-based policies with reinforcement learning (RL). The training is facilitated using an asymmetric actor-critic with access to privileged information. To overcome the computational complexity during image-based RL training, we use the inner edges of the gates as a sensor abstraction. This simple yet robust, task-relevant representation can be simulated during training without rendering images. During deployment, a Swin-transformer-based gate detector is used. Our approach enables autonomous agile flight with standard, off-the-shelf hardware. Although our demonstration focuses on drone racing, we believe that our method has an impact beyond drone racing and can serve as a foundation for future research into real-world applications in structured environments.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Vision-Based Agile Landing on Turbulent Waters

    cs.RO 2026-05 unverdicted novelty 6.0 of 10

    Reinforcement learning policy trained on synthetic visual features in simulation enables zero-shot real-world agile multirotor landing on turbulent maritime platforms without explicit platform-state estimation.

  2. YOPOv2-Tracker: An End-to-End Agile Tracking and Navigation Framework from Perception to Action

    cs.RO 2025-05 conditional novelty 6.0 of 10

    A single lightweight network replaces the cascade of detection, mapping, planning, and control, and tracks an uncooperative human target at up to 6 m/s in cluttered real-world environments.

  3. One Net to Rule Them All: Domain Randomization in Quadcopter Racing Across Different Platforms

    cs.RO 2025-04 conditional novelty 6.0 of 10

    A domain-randomized neural network policy trained in simulation races both a 3-inch and a 5-inch quadcopter in the real world, and randomization level trades speed for sim-to-real robustness.

  4. Self-Supervised Monocular Visual Drone Model Identification through Improved Occlusion Handling

    cs.RO 2025-04 conditional novelty 5.0 of 10

    A self-supervised scheme uses a vision model as a teacher to train a neural drone model from onboard data, improving velocity estimates and VIO accuracy at high speeds, with a proposed occlusion-handling loss that cut...

Pith tools