REVIEW 8 cited by
LLMs Assist NLP Researchers: Critique Paper (Meta-)Reviewing
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
This work is motivated by two key trends. On one hand, large language models (LLMs) have shown remarkable versatility in various generative tasks such as writing, drawing, and question answering, significantly reducing the time required for many routine tasks. On the other hand, researchers, whose work is not only time-consuming but also highly expertise-demanding, face increasing challenges as they have to spend more time reading, writing, and reviewing papers. This raises the question: how can LLMs potentially assist researchers in alleviating their heavy workload? This study focuses on the topic of LLMs assist NLP Researchers, particularly examining the effectiveness of LLM in assisting paper (meta-)reviewing and its recognizability. To address this, we constructed the ReviewCritique dataset, which includes two types of information: (i) NLP papers (initial submissions rather than camera-ready) with both human-written and LLM-generated reviews, and (ii) each review comes with "deficiency" labels and corresponding explanations for individual segments, annotated by experts. Using ReviewCritique, this study explores two threads of research questions: (i) "LLMs as Reviewers", how do reviews generated by LLMs compare with those written by humans in terms of quality and distinguishability? (ii) "LLMs as Metareviewers", how effectively can LLMs identify potential issues, such as Deficient or unprofessional review segments, within individual paper reviews? To our knowledge, this is the first work to provide such a comprehensive analysis.
Forward citations
Cited by 8 Pith papers
-
AbGen: Evaluating Large Language Models in Ablation Study Design and Evaluation for Scientific Research
A benchmark of 1,500 expert-annotated ablation study designs from 807 NLP papers shows frontier LLMs underperform human experts and that LLM-as-a-judge evaluations correlate weakly with human judgments.
-
I Can Find You in Seconds! Leveraging Large Language Models for Code Authorship Attribution
Zero-shot and few-shot LLM prompting can attribute source code authorship across C++ and Java, and a tournament prompting scheme scales this to 500 to 686 candidate authors with 65 to 69 percent top-1 accuracy.
-
LLM4SR: A Survey on Large Language Models for Scientific Research
A systematic review of LLM-based systems for hypothesis discovery, experiment planning, scientific writing, and peer review, including benchmarks, evaluation methods, and open challenges.
-
Exploring the Implicit Semantic Ability of Multimodal Large Language Models: A Pilot Study on Entity Set Expansion
LUSAR applies listwise sampling and ranking to multimodal LLMs for entity set expansion and reports improved MESED scores, though the gains are confounded with supervised fine-tuning.
-
Can Large Language Models Be Trusted Paper Reviewers? A Feasibility Study
An LLM-based review system processing 290 real submissions was far faster and cheaper than human review, yet agreed with the conference's acceptance decisions only 38.6% of the time, indicating LLMs should assist rath...
-
OpenReview Should be Protected and Leveraged as a Community Asset for Research in the Era of Large Language Models
The paper advocates protecting and leveraging OpenReview's peer review corpus as a community asset for LLM-based review assistance, benchmarks, and alignment.
-
AI-Driven Scholarly Peer Review via Persistent Workflow Prompting, Meta-Prompting, and Meta-Reasoning
A persistent, structured prompt loaded into an LLM chat session can guide reasoning models through critical analysis of experimental chemistry papers, but the evidence is a single qualitative case study.
-
Evaluating the Predictive Capacity of ChatGPT for Academic Peer Review Outcomes Across Multiple Platforms
Averaging 30 ChatGPT scores yields weak to moderate correlations with peer review scores on ICLR and SciPost Physics, but no correlation on F1000Research.
Discussion (0). Continue with ORCID to comment.