REVIEW 11 cited by
DeepReview: Improving LLM-based Paper Review with Human-like Deep Thinking Process
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Large Language Models (LLMs) are increasingly utilized in scientific research assessment, particularly in automated paper review. However, existing LLM-based review systems face significant challenges, including limited domain expertise, hallucinated reasoning, and a lack of structured evaluation. To address these limitations, we introduce DeepReview, a multi-stage framework designed to emulate expert reviewers by incorporating structured analysis, literature retrieval, and evidence-based argumentation. Using DeepReview-13K, a curated dataset with structured annotations, we train DeepReviewer-14B, which outperforms CycleReviewer-70B with fewer tokens. In its best mode, DeepReviewer-14B achieves win rates of 88.21\% and 80.20\% against GPT-o1 and DeepSeek-R1 in evaluations. Our work sets a new benchmark for LLM-based paper review, with all resources publicly available. The code, model, dataset and demo have be released in http://ai-researcher.net.
Forward citations
Cited by 11 Pith papers
-
The Ideation-Execution Gap: Execution Outcomes of LLM-Generated versus Human Research Ideas
A randomized execution study with 43 experts shows that LLM-generated research ideas lose more of their appeal than human ideas when actually implemented, reversing part of their ideation-stage advantage.
-
OmniPresent: Generating Coherent Presentation Suites from Scientific Papers
A multi-agent HTML pipeline with shared knowledge and cross-artifact verify-and-repair generates coherent poster/slides/video/page suites from papers and beats specialized baselines on OmniPreBench.
-
Paper2Rebuttal: A Multi-Agent Framework for Transparent Author Response Assistance
A multi-agent 'verify-then-write' system for writing peer-review rebuttals beats direct LLM prompting on a new benchmark, but the gains are measured by an LLM judge, not by the original reviewers.
-
OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning
A tool-augmented reward model trained with GRPO on 27K synthetic pairs beats existing reward models on long-form QA judgment and improves downstream alignment.
-
ReviewRL: Towards Automated Scientific Review with RL
ReviewRL combines arXiv retrieval, supervised fine-tuning, and reinforcement learning with a composite reward to generate paper reviews that better match human ratings and judged quality.
-
THE-Tree: Can Tracing Historical Evolution Enhance Scientific Verification and Reasoning?
THE-Tree constructs causally-linked semantic evolution trees from surveys and literature, and the authors report improved graph completion, future prediction, and LLM-based paper evaluation.
-
AI for Auto-Research: Roadmap & User Guide
The paper delivers a stage-by-stage roadmap for AI in research, showing reliable assistance in retrieval and tool tasks but fragility in novelty and judgment, advocating human-governed collaboration.
-
When Your Reviewer is an LLM: Biases, Divergence, and Prompt Injection Risks in Peer Review
GPT-5-mini gives weaker papers systematically higher scores than human reviewers, and hidden field-specific prompts in PDFs can force it to assign perfect scores or suppress weaknesses.
-
How Far Are AI Scientists from Changing the World?
This survey proposes a four-level capability framework for AI Scientist systems and, using an AI reviewer, finds that current systems produce papers rated well below normal scientific standards.
-
Deep Research Agents: A Systematic Examination And Roadmap
A survey that organizes LLM-powered deep research agents into static versus dynamic workflows and single versus multi agent architectures, and reviews their benchmarks and open challenges.
-
AI Scientists Fail Without Strong Implementation Capability
AI scientist systems can propose ideas but cannot reliably implement and verify experiments, making the implementation gap, not idea generation, the current bottleneck.
Discussion (0). Sign in to comment.