SIEVES improves selective prediction coverage by up to 3x on OOD VQA benchmarks by training a selector to score the quality of visual evidence produced by reasoner models, generalizing across benchmarks and proprietary models without internal access or per-task retraining.
Pooled” evaluates C@5 on the set of question-answer pairs aggregated over all repetitions. “Avg. per-rep
2 Pith papers cite this work, alongside 32 external citations. Polarity classification is still indexing.
2
Pith papers citing it
32
external citations · external index
years
2026 2representative citing papers
citing papers explorer
-
SIEVES: Selective Prediction Generalizes through Visual Evidence Scoring
SIEVES improves selective prediction coverage by up to 3x on OOD VQA benchmarks by training a selector to score the quality of visual evidence produced by reasoner models, generalizing across benchmarks and proprietary models without internal access or per-task retraining.
- MCMit: Hardware-Software Co-Design for Mid-Circuit Measurement Error Mitigation