REVIEW 20 cited by
On the Robustness of ChatGPT: An Adversarial and Out-of-distribution Perspective
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
ChatGPT is a recent chatbot service released by OpenAI and is receiving increasing attention over the past few months. While evaluations of various aspects of ChatGPT have been done, its robustness, i.e., the performance to unexpected inputs, is still unclear to the public. Robustness is of particular concern in responsible AI, especially for safety-critical applications. In this paper, we conduct a thorough evaluation of the robustness of ChatGPT from the adversarial and out-of-distribution (OOD) perspective. To do so, we employ the AdvGLUE and ANLI benchmarks to assess adversarial robustness and the Flipkart review and DDXPlus medical diagnosis datasets for OOD evaluation. We select several popular foundation models as baselines. Results show that ChatGPT shows consistent advantages on most adversarial and OOD classification and translation tasks. However, the absolute performance is far from perfection, which suggests that adversarial and OOD robustness remains a significant threat to foundation models. Moreover, ChatGPT shows astounding performance in understanding dialogue-related texts and we find that it tends to provide informal suggestions for medical tasks instead of definitive answers. Finally, we present in-depth discussions of possible research directions.
Forward citations
Cited by 20 Pith papers
-
Fisher Random Walk: Automatic Debiasing Contextual Preference Inference for Large Language Model Evaluation
A Fisher random walk weighted residual estimator achieves semiparametric efficient confidence intervals for contextual Bradley-Terry-Luce preference comparisons with flexible score estimators.
-
What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations)
Chat UI and API access to the same chatbot produce different accuracy, consistency, citation, and refusal behaviors on safety benchmarks, and web search changes these patterns further.
-
Would You Walk to the Car Wash? Revealing the Salience Bias of Large Language Models in Commonsense Reasoning
LLMs' failures on physically impossible, number-laden tasks are mostly knowledge suppression by salient distractors, not missing commons sense, and light prompting largely fixes them.
-
ASSURE: Metamorphic Testing for AI-powered Browser Extensions
A modular metamorphic testing framework for LLM-based browser extensions reports 531 automatically detected issues across six real-world extensions.
-
Jailbreak Defense in a Narrow Domain: Limitations of Existing Methods and a New Transcript-Classifier Approach
No tested defense, including the paper's own transcript classifier, can fully stop an LLM from giving competent bomb-making instructions under a grey-box attacker.
-
Evaluating LLMs Robustness in Less Resourced Languages with Proxy Models
Important words chosen by a small proxy model, when perturbed with typos or spacing errors, push Bielik, Mistral-7B, and Llama-3.1-8B to wrong answers on Polish classification tasks more often than random edits.
-
Bridging Distribution Shift and AI Safety: Conceptual and Methodological Synergies
The paper proposes a one-to-one mapping between six causes of distribution shift and several AI safety issues, arguing for mutual method transfer through aligned definitions.
-
Normative Evaluation of Large Language Models with Everyday Moral Dilemmas
Seven LLMs give different moral verdicts on AITA dilemmas, differ from Redditors, and only in an ensemble approximate human consensus.
-
Multi-Granularity Tibetan Textual Adversarial Attack Method Based on Masked Language Model
TSTricker uses masked Tibetan language models to generate syllable- and word-level substitutions that flip over 90% of fine-tuned Tibetan and multilingual classifiers' predictions.
-
Pay Attention to the Robustness of Chinese Minority Language Models! Syllable-level Textual Adversarial Attack on Tibetan Script
A Tibetan syllable-level black-box attack using syllable embeddings and a probability-based scoring mechanism successfully fools fine-tuned CINO models, with attack success rates up to 76%.
-
Hallucinations in medical devices
AI hallucinations in medical devices are defined as plausible errors, either impactful or benign, to guide device evaluation.
-
Look Within or Look Beyond? A Theoretical Comparison Between Parameter-Efficient and Full Fine-Tuning
The paper claims PEFT is a strict, less robust, lower-capacity subset of full fine-tuning, but the mathematical proofs contain load-bearing errors and the experiments, while suggestive, cannot repair them.
-
AI Ethics and Social Norms: Exploring ChatGPT's Capabilities From What to How
A cross-country survey and interview study finds that ChatGPT users and experts perceive transparency, bias, and data collection as the main ethical concerns, and that perceptions differ across groups for trust, secur...
-
Human-Readable Adversarial Prompts: An Investigation into LLM Vulnerabilities Using Situational Context
Human-readable adversarial insertions placed inside movie-summary prompts can jailbreak several open and closed LLMs, but the paper's measured attack rates are not statistically supported.
-
On Adversarial Robustness and Out-of-Distribution Robustness of Large Language Models
The paper claims model-specific correlations between adversarial and OOD robustness in LLMs, but these are based on a tiny number of strategies and are not statistically reliable.
-
Handling Out-of-Distribution Data: A Survey
A survey that organizes covariate and semantic shift handling methods into one taxonomy and argues for unified models, while contributing no new experiments or benchmark evaluation.
-
Evaluating and Improving Robustness in Large Language Models: A Survey and Future Directions
LLM robustness research is organized into adversarial robustness, out-of-distribution robustness, and evaluation, with an accompanying GitHub collection of papers.
-
Challenges in Guardrailing Large Language Models for Science
A position paper proposing a guardrail framework with four dimensions (trustworthiness, ethics & bias, safety, legal) and implementation strategies for scientific LLM use.
-
Building Task Bots with Self-learning for Enhanced Adaptability, Extensibility, and Factuality
A thesis that combines self-learning from dialog logs, schema-guided prompting, and self-aligned factuality to build task bots with minimal human intervention.
-
A Survey on Privacy Risks and Protection in Large Language Models
The paper surveys LLM privacy leaks and attacks, organizes them into a taxonomy, and reviews defenses without adding new empirical results.
Discussion (0). Continue with ORCID to comment.