Pith. sign in

Can large language models match the conclusions of systematic reviews?

3 Pith papers cite this work. Polarity classification is still indexing.

3 Pith papers citing it

citation-role summary

other 1

citation-polarity summary

years

2026 2 2025 1

verdicts

UNVERDICTED 3

roles

other 1

polarities

unclear 1

representative citing papers

Can AI Agents Synthesize Scientific Conclusions?

cs.AI · 2026-06-09 · unverdicted · novelty 7.0

A new benchmark and clean-room harness show frontier AI agents reach only 0.337 factual F1 when synthesizing conclusions from scientific evidence.

Treatment, evidence, imitation, and chat

stat.OT · 2025-06-29 · unverdicted · novelty 4.0

LLMs cannot solve the medical treatment problem through imitation alone because it requires evidence from experiments or observations, posing ethical challenges for training such systems.

citing papers explorer

Showing 3 of 3 citing papers.