REVIEW 18 cited by
ExpertPrompting: Instructing Large Language Models to be Distinguished Experts
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The answering quality of an aligned large language model (LLM) can be drastically improved if treated with proper crafting of prompts. In this paper, we propose ExpertPrompting to elicit the potential of LLMs to answer as distinguished experts. We first utilize In-Context Learning to automatically synthesize detailed and customized descriptions of the expert identity for each specific instruction, and then ask LLMs to provide answer conditioned on such agent background. Based on this augmented prompting strategy, we produce a new set of instruction-following data using GPT-3.5, and train a competitive open-source chat assistant called ExpertLLaMA. We employ GPT4-based evaluation to show that 1) the expert data is of significantly higher quality than vanilla answers, and 2) ExpertLLaMA outperforms existing open-source opponents and achieves 96\% of the original ChatGPT's capability. All data and the ExpertLLaMA model will be made publicly available at https://github.com/OFA-Sys/ExpertLLaMA.
Forward citations
Cited by 18 Pith papers
-
XCR-Bench: Benchmarking Cross-Cultural Reasoning in LLMs via Culture-Specific Items and Hall's Triad
XCR-Bench provides 4,100+ parallel sentences with 1,098 culture-specific items mapped to Hall's Triad, and shows LLMs struggle most with deeper, semi-visible cultural norms.
-
Exploring Advanced LLM Multi-Agent Systems Based on Blackboard Architecture
A blackboard-based LLM multi-agent system with controller-selected agents achieves competitive benchmark accuracy at lower token cost than several static and dynamically optimized baselines.
-
LLMs Can Teach Themselves to Better Predict the Future
DPO fine-tuning on outcome-ranked self-play forecasts improves LLM Brier scores by 7 to 10 percent over base and randomized-label controls.
-
Leveraging Multimodal LLM for Inspirational User Interface Search
A GPT-4o-based system extracts UI semantics from screenshots and provides semantic search for mobile UI design inspiration, beating CLIP-based retrieval in designer ratings.
-
Are Human Interactions Replicable by Generative Agents? A Case Study on Pronoun Usage in Hierarchical Interactions
In a leader/subordinate discussion simulation, most LLM agents do not reproduce the human pronoun pattern, and knowing the pattern does not help them show it.
-
What Makes Cryptic Crosswords Challenging for LLMs?
LLMs solve cryptic crossword clues with at most 11.4% accuracy, far below human experts, and fail mainly at definition extraction and wordplay type identification.
-
TransitGPT: A Generative AI-based framework for interacting with GTFS data using Large Language Models
A prompt-only LLM framework that generates and executes Python code answers 90 to 93 percent of 100 GTFS transit-data queries without fine-tuning.
-
Integrating gender inclusivity into large language models via instruction tuning
Abstract promises gender-inclusive Polish LLM tuning with the IPIS dataset, while the full text is an unrelated quantum transformer paper; no evidence for the declared claims is present.
-
Explaining GitHub Actions Failures with Large Language Models: Challenges, Insights, and Limitations
A 31-developer survey found that LLM-generated explanations of GitHub Actions failures are perceived as correct and clear for simple logs, but less useful for complex CI/CD failures.
-
Perspective Transition of Large Language Models for Solving Subjective Tasks
Reasoning through Perspective Transition (RPT) improves LLM performance on subjective NLP tasks by ranking direct, role, and third-person perspectives by self-reported confidence and answering from the top-ranked perspective.
-
SuperCode: Sustainability PER AI-driven CO-DEsign
The paper proposes an AI-driven hardware-software-science co-design methodology for radio astronomy, using sustainability as the key performance indicator, with no empirical results yet.
-
Improving Physics Reasoning in Large Language Models Using Mixture of Refinement Agents
MoRA uses GPT-4o to detect miscomprehension, wrong-concept, and computational errors in open-source LLM solutions, then routes specialized agents to fix them, improving multiple-choice physics accuracy by up to 16 per...
-
Leveraging Large Language Models for Institutional Portfolio Management: Persona-Based Ensembles
Persona-ensembled GPT-4 predictions improve Sharpe ratio over buy-and-hold during rising-CPI months in a 26-month stock-bond backtest, but underperform during falling-CPI months.
-
Evaluating Large Language Models on Business Process Modeling: Framework, Benchmark, and Self-Improvement Analysis
A benchmark of 20 business processes and 16 large language models finds Claude-3.5-Sonnet produces the highest-quality process models, and suggests that output optimization improves weaker models.
-
An Integrated Framework of Prompt Engineering and Multidimensional Knowledge Graphs for Legal Dispute Analysis
A prompt-plus-knowledge-graph framework for legal dispute analysis reports improved LLM sensitivity and citation accuracy on a 100-pair test set, but with limited statistical support.
-
Enhancing LLM Reasoning with Multi-Path Collaborative Reactive and Reflection agents
A multi-path, reactive-plus-reflection agent framework improves gpt-3.5-turbo accuracy on MMLU physics, math, and moral reasoning subsets compared with CoT, self-consistency, and self-refine baselines.
-
Practical Considerations for Agentic LLM Systems
This paper is a practical survey that organizes research on LLM-based agents into design considerations for planning, memory, tools, and control flow.
-
Generalist Virtual Agents: A Survey on Autonomous Agents Across Digital Platforms
A survey that proposes the Generalist Virtual Agent concept and taxonomies for agent environments, tasks, perceptions, actions, models, and evaluation, concluding that real-world-like environments favor human-like int...
Discussion (0). Continue with ORCID to comment.