A new benchmark and clean-room harness show frontier AI agents reach only 0.337 factual F1 when synthesizing conclusions from scientific evidence.
Richard Landis and Gary G
5 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
verdicts
UNVERDICTED 5roles
background 1polarities
background 1representative citing papers
ImproBR combines a hybrid detector with GPT-4o mini and RAG to raise bug report structural completeness from 7.9% to 96.4% and executable steps from 28.8% to 67.6% on 139 Mojira reports.
ValueBlindBench is a preregistered agreement-gated stress-test protocol for deciding when LLM-judged investment-rationale claims are stable enough to report, using 1,100 trajectories and 5,500 judge calls to gate claims by weighted kappa agreement.
API misuses in data-centric libraries share key characteristics with deep learning misuses and occur regardless of whether documentation directives are present.
Four new Reddit-derived datasets for mental health detection tasks are presented with inter-annotator agreement above 0.8 and reported model F1 scores of 93-99%.
citing papers explorer
-
Can AI Agents Synthesize Scientific Conclusions?
A new benchmark and clean-room harness show frontier AI agents reach only 0.337 factual F1 when synthesizing conclusions from scientific evidence.
-
ImproBR: Bug Report Improver Using LLMs
ImproBR combines a hybrid detector with GPT-4o mini and RAG to raise bug report structural completeness from 7.9% to 96.4% and executable steps from 28.8% to 67.6% on 139 Mojira reports.
-
ValueBlindBench: Agreement-Gated Stress Testing of LLM-Judged Investment Rationales Before Returns Are Observable
ValueBlindBench is a preregistered agreement-gated stress-test protocol for deciding when LLM-judged investment-rationale claims are stable enough to report, using 1,100 trajectories and 5,500 judge calls to gate claims by weighted kappa agreement.
-
An Empirical Study of API Misuses of Data-Centric Libraries
API misuses in data-centric libraries share key characteristics with deep learning misuses and occur regardless of whether documentation directives are present.
-
A Benchmark Suite of Reddit-Derived Datasets for Mental Health Detection
Four new Reddit-derived datasets for mental health detection tasks are presented with inter-annotator agreement above 0.8 and reported model F1 scores of 93-99%.