Pith. sign in

REVIEW 2 cited by

Pushing the Limits of ChatGPT on NLP Tasks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.09719 v2 pith:YH27JDC5 submitted 2023-06-16 cs.CL cs.AI

classification cs.CLcs.AI
keywords taskschatgptperformancesstrategysupervisedaddressbaselinesbetter
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Despite the success of ChatGPT, its performances on most NLP tasks are still well below the supervised baselines. In this work, we looked into the causes, and discovered that its subpar performance was caused by the following factors: (1) token limit in the prompt does not allow for the full utilization of the supervised datasets; (2) mismatch between the generation nature of ChatGPT and NLP tasks; (3) intrinsic pitfalls of LLMs models, e.g., hallucination, overly focus on certain keywords, etc. In this work, we propose a collection of general modules to address these issues, in an attempt to push the limits of ChatGPT on NLP tasks. Our proposed modules include (1) a one-input-multiple-prompts strategy that employs multiple prompts for one input to accommodate more demonstrations; (2) using fine-tuned models for better demonstration retrieval; (3) transforming tasks to formats that are more tailored to the generation nature; (4) employing reasoning strategies that are tailored to addressing the task-specific complexity; (5) the self-verification strategy to address the hallucination issue of LLMs; (6) the paraphrase strategy to improve the robustness of model predictions. We conduct experiments on 21 datasets of 10 representative NLP tasks, including question answering, commonsense reasoning, natural language inference, sentiment analysis, named entity recognition, entity-relation extraction, event extraction, dependency parsing, semantic role labeling, and part-of-speech tagging. Using the proposed assemble of techniques, we are able to significantly boost the performance of ChatGPT on the selected NLP tasks, achieving performances comparable to or better than supervised baselines, or even existing SOTA performances.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MEKiT: Multi-source Heterogeneous Knowledge Injection Method via Instruction Tuning for Emotion-Cause Pair Extraction

    cs.CL 2025-07 conditional novelty 5.0 of 10

    MEKiT improves LLM emotion-cause pair extraction by adding emotional label knowledge to instruction prompts and mixing causal examples into training data, achieving 61.49 F1 on NTCIR-13.

  2. Instruction Learning Paradigms: A Dual Perspective on White-box and Black-box LLMs

    cs.CL 2025-06 reject novelty 5.0 of 10

    A hybrid black-box and white-box instruction optimizer, built on InstructZero and INSTINCT, reports the highest mean score on 30 tasks but with small margins, missing error bars, and unreleased code.

Pith tools