Pith. sign in

REVIEW 6 cited by

Automated Peer Reviewing in Paper SEA: Standardization, Evaluation, and Analysis

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.12857 v2 pith:YORJ6NG3 submitted 2024-07-09 cs.CL cs.DLcs.IR

Automated Peer Reviewing in Paper SEA: Standardization, Evaluation, and Analysis

classification cs.CL cs.DLcs.IR
keywords automatedevaluationreviewingreviewsstandardizationanalysiscapabilitiesconsistency
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

In recent years, the rapid increase in scientific papers has overwhelmed traditional review mechanisms, resulting in varying quality of publications. Although existing methods have explored the capabilities of Large Language Models (LLMs) for automated scientific reviewing, their generated contents are often generic or partial. To address the issues above, we introduce an automated paper reviewing framework SEA. It comprises of three modules: Standardization, Evaluation, and Analysis, which are represented by models SEA-S, SEA-E, and SEA-A, respectively. Initially, SEA-S distills data standardization capabilities of GPT-4 for integrating multiple reviews for a paper. Then, SEA-E utilizes standardized data for fine-tuning, enabling it to generate constructive reviews. Finally, SEA-A introduces a new evaluation metric called mismatch score to assess the consistency between paper contents and reviews. Moreover, we design a self-correction strategy to enhance the consistency. Extensive experimental results on datasets collected from eight venues show that SEA can generate valuable insights for authors to improve their papers.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. FARS: A Fully Automated Research System Deployed at Scale

    cs.AI 2026-06 unverdicted novelty 7.0

    FARS deployed at scale produced 166 AI/ML papers across 67 topics that received 282 structured human reviews indicating some review-worthy outputs alongside recurring failure modes.

  2. FARS: A Fully Automated Research System Deployed at Scale

    cs.AI 2026-06 conditional novelty 6.0

    A fully automated AI-for-AI research system produced 166 papers across 67 topics; human reviews of 140 papers show occasional review-worthy work but mostly low scores and recurring integrity and scope failures.

  3. Beyond Rating: A Comprehensive Evaluation and Benchmark for AI Reviews

    cs.CL 2026-04 unverdicted novelty 6.0

    Beyond Rating proposes five text-centric metrics for AI reviewers and demonstrates that aligning AI focus on paper weaknesses with human experts is required for reliable automated review scoring.

  4. MEIC-DT: Memory-Efficient Incremental Clustering for Long-Text Coreference Resolution with Dual-Threshold Constraints

    cs.IR 2025-12 unverdicted novelty 6.0

    MEIC-DT delivers competitive coreference resolution on long texts via a memory-bounded dual-threshold incremental clustering scheme built on a lightweight Transformer.

  5. SafeReview: Defending LLM-based Review Systems Against Adversarial Hidden Prompts

    cs.CL 2026-04 unverdicted novelty 5.0

    SafeReview trains a Generator to create adversarial prompts and a Defender to detect them via co-evolution with an IR-GAN-inspired loss, claiming better resilience than static defenses for LLM-based peer review.

  6. LLM-Based Scientific Peer Review: Methods, Benchmarks, and Reliability Challenges

    cs.CL 2026-06 unverdicted novelty 4.0

    A survey synthesizing LLM methods for peer review critique generation and score prediction, including taxonomies, benchmark limitations, domain biases, and robustness risks such as prompt injection.