REVIEW 4 cited by
Surprise Potential as a Measure of Interactivity in Driving Scenarios
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Validating the safety and performance of an autonomous vehicle (AV) requires benchmarking on real-world driving logs. However, typical driving logs contain mostly uneventful scenarios with minimal interactions between road users. Identifying interactive scenarios in real-world driving logs enables the curation of datasets that amplify critical signals and provide a more accurate assessment of an AV's performance. In this paper, we present a novel metric that identifies interactive scenarios by measuring an AV's surprise potential on others. First, we identify three dimensions of the design space to describe a family of surprise potential measures. Second, we exhaustively evaluate and compare different instantiations of the surprise potential measure within this design space on the nuScenes dataset. To determine how well a surprise potential measure correctly identifies an interactive scenario, we use a reward model learned from human preferences to assess alignment with human intuition. Our proposed surprise potential, arising from this exhaustive comparative study, achieves a correlation of more than 0.82 with the human-aligned reward function, outperforming existing approaches. Lastly, we validate motion planners on curated interactive scenarios to demonstrate downstream applications.
Forward citations
Cited by 4 Pith papers
-
Safety Evaluation of Motion Plans Using Trajectory Predictors as Forward Reachable Set Estimators
FORCE-OPT extracts calibrated, multi-modal reachable sets from GMM trajectory predictors using convex optimization and conformal prediction, achieving the lowest balanced error rate in safety evaluation on nuScenes.
-
CrashAgent: Crash Scenario Generation via Multi-modal Reasoning
A multi-agent vision-language framework converts NHTSA crash reports into simulation-ready road layouts and collision scenarios, with modest accuracy gains over direct VLM baselines.
-
Test Automation for Interactive Scenarios via Promptable Traffic Simulation
A goal-prompt search with Bayesian optimization over a data-driven traffic simulator automatically finds safety-critical scenarios for testing autonomous vehicle planners.
-
Sim2Val: Leveraging Correlation Across Test Platforms for Variance-Reduced Metric Estimation
Sim2Val adapts control variates and prediction-powered inference to robot validation, using correlated simulator outputs to reduce the real-world sample count needed for a given confidence interval.
Discussion (0). Continue with ORCID to comment.