Pith. sign in

REVIEW 1 cited by

Optimizing Algorithms From Pairwise User Preferences

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2308.04571 v1 pith:GTA2XHGD submitted 2023-08-08 cs.RO cs.CVcs.HC

classification cs.ROcs.CVcs.HC
keywords userpreferencesrewardrobotapproachesbehaviorgroundhowever
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Typical black-box optimization approaches in robotics focus on learning from metric scores. However, that is not always possible, as not all developers have ground truth available. Learning appropriate robot behavior in human-centric contexts often requires querying users, who typically cannot provide precise metric scores. Existing approaches leverage human feedback in an attempt to model an implicit reward function; however, this reward may be difficult or impossible to effectively capture. In this work, we introduce SortCMA to optimize algorithm parameter configurations in high dimensions based on pairwise user preferences. SortCMA efficiently and robustly leverages user input to find parameter sets without directly modeling a reward. We apply this method to tuning a commercial depth sensor without ground truth, and to robot social navigation, which involves highly complex preferences over robot behavior. We show that our method succeeds in optimizing for the user's goals and perform a user study to evaluate social navigation results.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Generative Representational Learning of Foundation Models for Recommendation

    cs.IR 2025-06 conditional novelty 6.0 of 10

    A single recommendation model with task-aware Mixture of Low-rank Experts and convergence-based sample scheduling beats baselines on a new 13-task benchmark.

Pith tools