A new benchmark and clean-room harness show frontier AI agents reach only 0.337 factual F1 when synthesizing conclusions from scientific evidence.
ArXivabs/2512.10791 (2025),https://api.semanticscholar.org/CorpusID:283737312
7 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
years
2026 7roles
background 1polarities
unclear 1representative citing papers
Agreeableness in AI personas reliably predicts sycophantic behavior in 9 of 13 tested language models.
PROBE pipeline with deterministic PCAP normalization, verdict-aware evidence ensembles, and composite reliability scoring raises weighted evidence F1 to 0.957 on 87 Wi-Fi captures while avoiding LLM self-confidence and evaluation bias issues.
LLMs need metacognition to align expressed uncertainty with their actual knowledge boundaries, moving beyond knowledge expansion to reduce confident errors.
An Explicit Logic Channel of LLM, VFM and probabilistic inference validates and improves zero-shot MLLMs via Consistency Rate without ground-truth labels.
A new harness and benchmark, VeRO, lets coding agents be scored on how much they improve target agents under a fixed evaluation budget.
The paper introduces a 10-faculty Cognitive Taxonomy and a held-out task protocol to generate cognitive profiles for measuring AI progress toward AGI.
citing papers explorer
-
Can AI Agents Synthesize Scientific Conclusions?
A new benchmark and clean-room harness show frontier AI agents reach only 0.337 factual F1 when synthesizing conclusions from scientific evidence.
-
Too Nice to Tell the Truth: Quantifying Agreeableness-Driven Sycophancy in Role-Playing Language Models
Agreeableness in AI personas reliably predicts sycophantic behavior in 9 of 13 tested language models.
-
Evidence-Grounded Ensemble Diagnosis of 802.11 Packet Captures: A Multi-Stage Pipeline with Deterministic Reliability Scoring
PROBE pipeline with deterministic PCAP normalization, verdict-aware evidence ensembles, and composite reliability scoring raises weighted evidence F1 to 0.957 on 87 Wi-Fi captures while avoiding LLM self-confidence and evaluation bias issues.
-
Hallucinations Undermine Trust; Metacognition is a Way Forward
LLMs need metacognition to align expressed uncertainty with their actual knowledge boundaries, moving beyond knowledge expansion to reduce confident errors.
-
Explicit Logic Channel for Validation and Enhancement of MLLMs on Zero-Shot Tasks
An Explicit Logic Channel of LLM, VFM and probabilistic inference validates and improves zero-shot MLLMs via Consistency Rate without ground-truth labels.
-
VeRO: A Harness for Agents to Optimize Agents
A new harness and benchmark, VeRO, lets coding agents be scored on how much they improve target agents under a fixed evaluation budget.
-
Measuring Progress Toward AGI: A Cognitive Framework
The paper introduces a 10-faculty Cognitive Taxonomy and a held-out task protocol to generate cognitive profiles for measuring AI progress toward AGI.