REVIEW 32 cited by
Better Zero-Shot Reasoning with Role-Play Prompting
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Modern large language models (LLMs) exhibit a remarkable capacity for role-playing, enabling them to embody not only human characters but also non-human entities. This versatility allows them to simulate complex human-like interactions and behaviors within various contexts, as well as to emulate specific objects or systems. While these capabilities have enhanced user engagement and introduced novel modes of interaction, the influence of role-playing on LLMs' reasoning abilities remains underexplored. In this study, we introduce a strategically designed role-play prompting methodology and assess its performance under the zero-shot setting across twelve diverse reasoning benchmarks. Our empirical results illustrate that role-play prompting consistently surpasses the standard zero-shot approach across most datasets. Notably, in experiments conducted using ChatGPT, accuracy on AQuA rises from 53.5% to 63.8%, and on Last Letter from 23.8% to 84.2%.Upon further comparison with the Zero-Shot-CoT technique, which prompts the model to "think step by step", our study demonstrates that role-play prompting acts as a more effective trigger for the CoT process. This highlights its potential to augment the reasoning capabilities of LLMs. We release our code at https://github.com/NKU-HLT/Role-Play-Prompting.
Forward citations
Cited by 32 Pith papers
-
Prompt engineering using order-of-addition experiments: An application to generating two-level fractional factorial designs
Order-of-addition designs and logistic pairwise-ordering models measure and optimize prompt-element order, lifting LLM success on 16-run fractional factorial design tasks from low teens or mid-thirties to near 100%.
-
Toward Adaptive Reasoning in Large Language Models with Thought Rollback
Thought Rollback enables LLMs to revise prior reasoning steps through rollback and accumulated error analysis, improving solve rates on math and multi-task benchmarks while increasing token usage dramatically.
-
Quality-Aware Personalized AI Service Provisioning in UAV-Assisted 6G Networks
HyPE integrates DRL mobility prediction, LLM-based UAV trajectory and inference assignment, and greedy heuristics for service placement and routing to jointly optimize latency, output fidelity, and personalization con...
-
Orchid: Orchestrating Context Across Creative Workflows with Generative AI
A notebook-style GenAI tool that supports specifying, referencing, and monitoring context produced more novel, feasible, and valuable creative outcomes than a fragmented toolbelt in a within-subjects study of 12 participants.
-
LLMs on Trial: Evaluating Judicial Fairness for Large Language Models
A new 177,100-case benchmark shows that 16 LLMs systematically vary criminal sentences based on extra-legal demographic and procedural details, revealing pervasive judicial unfairness.
-
Stabilizing Black-Box Prompt Optimization with Textual Regularization and Signal Aggregation
TRAS adds success-based textual regularization and Monte Carlo signal aggregation to black-box prompt optimization, improving accuracy and reducing instruction loss when moving prompts across models.
-
To Protect the LLM Agent Against the Prompt Injection Attack with Polymorphic Prompt
Polymorphic Prompt Assembling randomizes per-request system-prompt separators, cutting prompt-injection attack success to as low as 1.83% on GPT-3.5 with 0.06 ms runtime overhead.
-
ORPP: Self-Optimizing Role-playing Prompts to Enhance Language Model Capabilities
ORPP generates task-specific role-playing prompts through iterative reward-model-guided optimization on a small sample, then uses few-shot transfer to create prompts for new questions.
-
Toward Structured Knowledge Reasoning: Contrastive Retrieval-Augmented Generation on Experience
CoRE improves structured knowledge reasoning by retrieving both correct and incorrect past examples into the prompt, using MCTS-generated experience memory.
-
ViDDAR: Vision Language Model-Based Task-Detrimental Content Detection for Augmented Reality
ViDDAR is a VLM-based system that detects task-detrimental AR content, achieving 92.15% obstruction detection and 82.46% information manipulation detection accuracy on its own dataset.
-
CKGFuzzer: LLM-Based Fuzz Driver Generation Enhanced By Code Knowledge Graph
CKGFuzzer uses a code knowledge graph to guide LLM agents in generating, repairing, and mutating fuzz drivers, reporting a pooled 8.73% relative coverage gain over PromptFuzz on eight libraries and 9 new bugs.
-
Agentic Harnesses: LLM-Driven Verification Layers for Robot Autonomy
An LLM-driven gate between robot planning and execution labels plans accept, reject, or escalate, reporting 81 percent accuracy and no direct accept/reject errors on small test sets.
-
RESBev: Making BEV Perception More Robust
A latent world model predicts clean BEV semantic features from sequential observations to recover existing Lift-Splat-Shoot pipelines under natural and adversarial corruption with few-shot fine-tuning.
-
REFLEX: Self-Refining Explainable Fact-Checking via Verdict-Anchored Style Control
REFLEX improves explainable fact-checking by using verdict-anchored style control and self-disagreement signals to disentangle fact from style in LLM outputs, achieving SOTA results with minimal self-refined samples.
-
A Role-Aware Multi-Agent Framework for Financial Education Question Answering with LLMs
A role-aware multi-agent pipeline with retrieval and expert critique raises financial multiple-choice accuracy by 6.6-8.3 percentage points over zero-shot CoT across four LLMs.
-
CS-Agent: LLM-based Community Search via Dual-agent Collaboration
CS-Agent, a Solver-Validator two-agent dialogue with a Decider selector, improves LLM community search on synthetic graphs, and GraphCS is a new benchmark for measuring it.
-
Statistical Hypothesis Testing for Auditing Robustness in Language Models
A permutation-based hypothesis test on pairwise semantic similarities detects whether LLM outputs shift under arbitrary input or model perturbations.
-
AgentMisalignment: Measuring the Propensity for Misaligned Behaviour in LLM-Based Agents
A new nine-task benchmark measures LLM agents' propensity for misalignment and finds more capable models misalign more on average, with persona effects sometimes exceeding model effects.
-
Exploring the Impact of Occupational Personas on Domain-Specific QA
Profession-based personas slightly improve LLM accuracy on science QA, while occupational personality personas often reduce it, even when semantically related.
-
The Eye of Sherlock Holmes: Uncovering User Private Attribute Profiling via Vision-Language Model Agentic Framework
A VLM-LLM agentic pipeline and a new 251-person benchmark show that ordinary personal photo sets can reveal private attributes, including abstract traits like income and MBTI, at rates above human evaluators.
-
BinMetric: A Comprehensive Binary Analysis Benchmark for Large Language Models
BinMetric is a new 1,000-question, six-task benchmark that measures LLM ability across the binary analysis lifecycle, with an empirical study of 12 models showing strong semantic understanding but weak low-level lifti...
-
Exploring Zero-Shot App Review Classification with ChatGPT: Challenges and Potential
Zero-shot GPT-4o mini classifies app reviews into functional, non-functional, both, or neither with F1 0.84, outperforming classical ML models on a 1,880-review benchmark.
-
LLM-Assisted Automated Deductive Coding of Dialogue Data: Leveraging Dialogue-Specific Characteristics to Enhance Contextual Understanding
An LLM-assisted pipeline that separates communicative acts from events, uses multi-model voting, and applies a consistency check reaches Cohen's kappa above 0.80 with human coders on a small student-dialogue corpus.
-
Edge Agentic AI Framework for Autonomous Network Optimisation in O-RAN
A simulated edge agentic AI framework with LSTM traffic prediction and tiered Tx power control reports zero network outages in high-stress 5G scenarios.
-
Pedestrian Intention Prediction via Vision-Language Foundation Models
Time-aware vehicle-speed prompts improve vision-language model accuracy for pedestrian crossing intent, but the claimed edge over specialized vision models is not consistent across the paper's own benchmarks.
-
Generating Automotive Code: Large Language Models for Software Development and Verification in Safety-Critical Systems
An LLM-based, feedback-driven code generation pipeline produced an ISO-inspired ACC implementation that passed static and CARLA simulation checks in all three test runs.
-
Automating Security Audit Using Large Language Model based Agent: An Exploration Experiment
A GPT-4 LangChain agent successfully performed basic Windows password policy compliance checks, with acknowledged limitations in ambiguous situations.
-
AI-Driven Scholarly Peer Review via Persistent Workflow Prompting, Meta-Prompting, and Meta-Reasoning
A persistent, structured prompt loaded into an LLM chat session can guide reasoning models through critical analysis of experimental chemistry papers, but the evidence is a single qualitative case study.
-
Learning-by-teaching with ChatGPT: The effect of teachable ChatGPT agent on programming education
Students who taught a ChatGPT agent to solve the Eight Queens puzzle learned more and wrote clearer pseudocode than a video-only control group, but showed no extra gain in code correctness.
-
PKRD-CoT: A Unified Chain-of-thought Prompting for Multi-Modal Large Language Models in Autonomous Driving
PKRD-CoT structures multimodal LLM prompts into perception, knowledge, reasoning, and decision steps, and the authors report improved driving decision accuracy for GPT-4.0 and several other models.
-
Survey of GenAI for Automotive Software Development: From Requirements to Executable Code
A review of roughly 60 papers and 9 industry respondents finds GPT-family models dominate automotive code generation while requirements handling lags due to confidentiality constraints.
-
Leveraging the Potential of Prompt Engineering for Hate Speech Detection in Low-Resource Languages
Relabeling hate speech as metaphor pairs (red/green, summer/winter) in prompts raises Llama2's F1 on a 500-item Bengali subsample to 95.89, though the gain is reported without matched test-set comparisons or error bars.
Discussion (0). Continue with ORCID to comment.