Pith. sign in

REVIEW 14 cited by

Reviewer2: Optimizing Review Generation Through Prompt Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.10886 v2 pith:ORD5NZVN submitted 2024-02-16 cs.CL

Reviewer2: Optimizing Review Generation Through Prompt Generation

classification cs.CL
keywords reviewgenerationreviewsaddressaspectsauthorscoverdraft
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Recent developments in LLMs offer new opportunities for assisting authors in improving their work. In this paper, we envision a use case where authors can receive LLM-generated reviews that uncover weak points in the current draft. While initial methods for automated review generation already exist, these methods tend to produce reviews that lack detail, and they do not cover the range of opinions that human reviewers produce. To address this shortcoming, we propose an efficient two-stage review generation framework called Reviewer2. Unlike prior work, this approach explicitly models the distribution of possible aspects that the review may address. We show that this leads to more detailed reviews that better cover the range of aspects that human reviewers identify in the draft. As part of the research, we generate a large-scale review dataset of 27k papers and 99k reviews that we annotate with aspect prompts, which we make available as a resource for future research.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 14 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. NovBench: Evaluating Large Language Models on Academic Paper Novelty Assessment

    cs.CL 2026-04 unverdicted novelty 8.0

    NovBench is the first large-scale benchmark with 1,684 expert-annotated pairs to evaluate LLMs on assessing academic paper novelty via a four-dimensional framework of Relevance, Correctness, Coverage, and Clarity.

  2. ReviewGrounder: Improving Review Substantiveness with Rubric-Guided, Tool-Integrated Agents

    cs.CL 2026-04 unverdicted novelty 7.0

    ReviewGrounder decomposes review generation into rubric-guided drafting and tool-integrated grounding stages, outperforming larger baseline models on a new benchmark measuring alignment with human judgments and review...

  3. Judgment-Grounded Expansion for Peer Review Generation

    cs.CL 2026-06 unverdicted novelty 6.0

    Formalizes judgment-grounded expansion as a human-AI collaborative task for peer review generation, supported by a user study and conformal prediction methods for scalable evaluation.

  4. From Passive Generation to Investigation: A Proactive Scientific Peer Review Agent

    cs.CL 2026-06 unverdicted novelty 6.0

    ProReviewer is an MDP-formulated proactive peer review agent trained with SFT and RL on an 8B model that outperforms larger frontier LLMs on review quality metrics.

  5. PRISM: A Multi-Dimensional Benchmark for Evaluating LLM Peer Reviewers

    cs.CL 2026-05 unverdicted novelty 6.0

    PRISM benchmark finds LLMs match or exceed humans on isolated review dimensions like novelty verification but none achieve the balanced performance of human reviewers across depth, flaw prioritization, and constructiveness.

  6. HiRAS: A Hierarchical Multi-Agent Framework for Paper-to-Code Generation and Execution

    cs.CL 2026-04 unverdicted novelty 6.0

    HiRAS introduces hierarchical multi-agent coordination for paper-to-code generation and experiment reproduction, claiming over 10% relative gains over prior state-of-the-art on a refined benchmark with reduced hallucination.

  7. EGTR-Review: Efficient Evidence-Grounded Scientific Peer Review Generation via Multi-Agent Teacher Distillation

    cs.CL 2026-06 unverdicted novelty 5.0

    EGTR-Review distills a multi-agent evidence-grounded review generator into an efficient student model that outperforms baselines on quality, grounding, and traceability while using fewer tokens.

  8. LLM-as-a-Reviewer: Benchmarking Their Ability, Divergence, and Prompt Injection Resistance as Paper Reviewers

    cs.CL 2026-05 unverdicted novelty 5.0

    LLMs overrate weak papers, diverge from humans on criteria like clarity and reproducibility, write longer less diverse reviews, and remain vulnerable to prompt injection attacks that can boost low-scoring papers to ac...

  9. SafeReview: Defending LLM-based Review Systems Against Adversarial Hidden Prompts

    cs.CL 2026-04 unverdicted novelty 5.0

    SafeReview trains a Generator to create adversarial prompts and a Defender to detect them via co-evolution with an IR-GAN-inspired loss, claiming better resilience than static defenses for LLM-based peer review.

  10. Impact of large language models on peer review opinions from a fine-grained perspective: Evidence from top conference proceedings in AI

    cs.CL 2026-04 unverdicted novelty 5.0

    Peer review reports in AI conferences have grown longer and more standardized after LLMs, with increased emphasis on surface-level clarity and summaries at the expense of deeper critiques on originality and replicability.

  11. LLM-Based Scientific Peer Review: Methods, Benchmarks, and Reliability Challenges

    cs.CL 2026-06 unverdicted novelty 4.0

    A survey synthesizing LLM methods for peer review critique generation and score prediction, including taxonomies, benchmark limitations, domain biases, and robustness risks such as prompt injection.

  12. SciAtlas: A Large-Scale Knowledge Graph for Automated Scientific Research

    cs.AI 2026-05 unverdicted novelty 4.0

    SciAtlas builds a large-scale multi-disciplinary academic knowledge graph and a neuro-symbolic retrieval system to support automated scientific research tasks such as literature review and idea positioning.

  13. AI for Auto-Research: Roadmap & User Guide

    cs.AI 2026-05 unverdicted novelty 4.0

    The paper delivers a stage-by-stage roadmap for AI in research, showing reliable assistance in retrieval and tool tasks but fragility in novelty and judgment, advocating human-governed collaboration.

  14. Evolving Roles of LLMs in Scientific Innovation: Assistant, Collaborator, Scientist, and Evaluator

    cs.DL 2025-07 unverdicted novelty 4.0

    The paper proposes a four-role framework for LLMs in scientific innovation and reviews methods, benchmarks, and limitations across Assistant, Collaborator, Scientist, and Evaluator roles.