Pith. sign in

REVIEW 1 cited by

Learning to View: Decision Transformers for Active Object Detection

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2301.09544 v1 pith:RPDWUYBI submitted 2023-01-23 cs.RO cs.CV

classification cs.ROcs.CV
keywords detectionrobotperceptionplanningpolicyresultsactivedataset
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Active perception describes a broad class of techniques that couple planning and perception systems to move the robot in a way to give the robot more information about the environment. In most robotic systems, perception is typically independent of motion planning. For example, traditional object detection is passive: it operates only on the images it receives. However, we have a chance to improve the results if we allow planning to consume detection signals and move the robot to collect views that maximize the quality of the results. In this paper, we use reinforcement learning (RL) methods to control the robot in order to obtain images that maximize the detection quality. Specifically, we propose using a Decision Transformer with online fine-tuning, which first optimizes the policy with a pre-collected expert dataset and then improves the learned policy by exploring better solutions in the environment. We evaluate the performance of proposed method on an interactive dataset collected from an indoor scenario simulator. Experimental results demonstrate that our method outperforms all baselines, including expert policy and pure offline RL methods. We also provide exhaustive analyses of the reward distribution and observation space.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. GoStop: Reinforcement Learning for Adaptive Temporal Aggregation in Event-Based Feature Tracking

    cs.CV 2026-07 conditional novelty 6.0 of 10

    An RL agent that adaptively decides when to accumulate events and when to run tracking inference improves event-based feature tracking on a new dynamic benchmark, but the gains are less consistent on an existing benchmark.

Pith tools