Pith. sign in

REVIEW 8 cited by

LLMs Assist NLP Researchers: Critique Paper (Meta-)Reviewing

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.16253 v3 pith:J2ZYQEPT submitted 2024-06-24 cs.CL

classification cs.CL
keywords llmsresearchersassistreviewingreviewsworkhandindividual
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

This work is motivated by two key trends. On one hand, large language models (LLMs) have shown remarkable versatility in various generative tasks such as writing, drawing, and question answering, significantly reducing the time required for many routine tasks. On the other hand, researchers, whose work is not only time-consuming but also highly expertise-demanding, face increasing challenges as they have to spend more time reading, writing, and reviewing papers. This raises the question: how can LLMs potentially assist researchers in alleviating their heavy workload? This study focuses on the topic of LLMs assist NLP Researchers, particularly examining the effectiveness of LLM in assisting paper (meta-)reviewing and its recognizability. To address this, we constructed the ReviewCritique dataset, which includes two types of information: (i) NLP papers (initial submissions rather than camera-ready) with both human-written and LLM-generated reviews, and (ii) each review comes with "deficiency" labels and corresponding explanations for individual segments, annotated by experts. Using ReviewCritique, this study explores two threads of research questions: (i) "LLMs as Reviewers", how do reviews generated by LLMs compare with those written by humans in terms of quality and distinguishability? (ii) "LLMs as Metareviewers", how effectively can LLMs identify potential issues, such as Deficient or unprofessional review segments, within individual paper reviews? To our knowledge, this is the first work to provide such a comprehensive analysis.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. AbGen: Evaluating Large Language Models in Ablation Study Design and Evaluation for Scientific Research

    cs.CL 2025-07 conditional novelty 7.0 of 10

    A benchmark of 1,500 expert-annotated ablation study designs from 807 NLP papers shows frontier LLMs underperform human experts and that LLM-as-a-judge evaluations correlate weakly with human judgments.

  2. I Can Find You in Seconds! Leveraging Large Language Models for Code Authorship Attribution

    cs.SE 2025-01 conditional novelty 6.0 of 10

    Zero-shot and few-shot LLM prompting can attribute source code authorship across C++ and Java, and a tournament prompting scheme scales this to 500 to 686 candidate authors with 65 to 69 percent top-1 accuracy.

  3. LLM4SR: A Survey on Large Language Models for Scientific Research

    cs.CL 2025-01 conditional novelty 5.0 of 10

    A systematic review of LLM-based systems for hypothesis discovery, experiment planning, scientific writing, and peer review, including benchmarks, evaluation methods, and open challenges.

  4. Exploring the Implicit Semantic Ability of Multimodal Large Language Models: A Pilot Study on Entity Set Expansion

    cs.CL 2024-12 conditional novelty 5.0 of 10

    LUSAR applies listwise sampling and ranking to multimodal LLMs for entity set expansion and reports improved MESED scores, though the gains are confounded with supervised fine-tuning.

  5. Can Large Language Models Be Trusted Paper Reviewers? A Feasibility Study

    cs.CY 2025-06 conditional novelty 4.0 of 10

    An LLM-based review system processing 290 real submissions was far faster and cheaper than human review, yet agreed with the conference's acceptance decisions only 38.6% of the time, indicating LLMs should assist rath...

  6. OpenReview Should be Protected and Leveraged as a Community Asset for Research in the Era of Large Language Models

    cs.CY 2025-05 conditional novelty 4.0 of 10

    The paper advocates protecting and leveraging OpenReview's peer review corpus as a community asset for LLM-based review assistance, benchmarks, and alignment.

  7. AI-Driven Scholarly Peer Review via Persistent Workflow Prompting, Meta-Prompting, and Meta-Reasoning

    cs.AI 2025-05 conditional novelty 4.0 of 10

    A persistent, structured prompt loaded into an LLM chat session can guide reasoning models through critical analysis of experimental chemistry papers, but the evidence is a single qualitative case study.

  8. Evaluating the Predictive Capacity of ChatGPT for Academic Peer Review Outcomes Across Multiple Platforms

    cs.DL 2024-11 conditional novelty 4.0 of 10

    Averaging 30 ChatGPT scores yields weak to moderate correlations with peer review scores on ICLR and SciPost Physics, but no correlation on F1000Research.

Pith tools