A new benchmark and clean-room harness show frontier AI agents reach only 0.337 factual F1 when synthesizing conclusions from scientific evidence.
best- effort
4 Pith papers cite this work. Polarity classification is still indexing.
years
2026 4verdicts
UNVERDICTED 4representative citing papers
LLM self-reports predict behavior selectively: TPB reaches human-level coherence within shared conversations but collapses across sessions for primed behaviors, unlike Big 5, with persona prompting stabilizing reports but not actions.
CaliPPer introduces a distance-based framework that quantifies generalizability, predicts performance metrics like AUROC with low error, and improves predictions on unseen binding data across multiple models and domains.
LLMs exhibit Bayesian-like hypothesis updating with strong-sampling bias and an evaluation-generation gap but generalize poorly outside observed data.
citing papers explorer
-
CaliPPer: quantifying, predicting and improving AI model performance for binding prediction
CaliPPer introduces a distance-based framework that quantifies generalizability, predicts performance metrics like AUROC with low error, and improves predictions on unseen binding data across multiple models and domains.