Pith. sign in

REVIEW 18 cited by

ExpertPrompting: Instructing Large Language Models to be Distinguished Experts

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.14688 v2 pith:7LIH3W5C submitted 2023-05-24 cs.CL cs.AI

classification cs.CLcs.AI
keywords expertllamadataanswerdistinguishedexpertexpertpromptingexpertslanguage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The answering quality of an aligned large language model (LLM) can be drastically improved if treated with proper crafting of prompts. In this paper, we propose ExpertPrompting to elicit the potential of LLMs to answer as distinguished experts. We first utilize In-Context Learning to automatically synthesize detailed and customized descriptions of the expert identity for each specific instruction, and then ask LLMs to provide answer conditioned on such agent background. Based on this augmented prompting strategy, we produce a new set of instruction-following data using GPT-3.5, and train a competitive open-source chat assistant called ExpertLLaMA. We employ GPT4-based evaluation to show that 1) the expert data is of significantly higher quality than vanilla answers, and 2) ExpertLLaMA outperforms existing open-source opponents and achieves 96\% of the original ChatGPT's capability. All data and the ExpertLLaMA model will be made publicly available at https://github.com/OFA-Sys/ExpertLLaMA.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 18 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 53 citations worldwide. Full citation record

  1. XCR-Bench: Benchmarking Cross-Cultural Reasoning in LLMs via Culture-Specific Items and Hall's Triad

    cs.CL 2026-01 conditional novelty 6.0 of 10

    XCR-Bench provides 4,100+ parallel sentences with 1,098 culture-specific items mapped to Hall's Triad, and shows LLMs struggle most with deeper, semi-visible cultural norms.

  2. Exploring Advanced LLM Multi-Agent Systems Based on Blackboard Architecture

    cs.MA 2025-07 conditional novelty 6.0 of 10

    A blackboard-based LLM multi-agent system with controller-selected agents achieves competitive benchmark accuracy at lower token cost than several static and dynamically optimized baselines.

  3. LLMs Can Teach Themselves to Better Predict the Future

    cs.CL 2025-02 conditional novelty 6.0 of 10

    DPO fine-tuning on outcome-ranked self-play forecasts improves LLM Brier scores by 7 to 10 percent over base and randomized-label controls.

  4. Leveraging Multimodal LLM for Inspirational User Interface Search

    cs.HC 2025-01 conditional novelty 6.0 of 10

    A GPT-4o-based system extracts UI semantics from screenshots and provides semantic search for mobile UI design inspiration, beating CLIP-based retrieval in designer ratings.

  5. Are Human Interactions Replicable by Generative Agents? A Case Study on Pronoun Usage in Hierarchical Interactions

    cs.CL 2025-01 conditional novelty 6.0 of 10

    In a leader/subordinate discussion simulation, most LLM agents do not reproduce the human pronoun pattern, and knowing the pattern does not help them show it.

  6. What Makes Cryptic Crosswords Challenging for LLMs?

    cs.CL 2024-12 conditional novelty 6.0 of 10

    LLMs solve cryptic crossword clues with at most 11.4% accuracy, far below human experts, and fail mainly at definition extraction and wordplay type identification.

  7. TransitGPT: A Generative AI-based framework for interacting with GTFS data using Large Language Models

    cs.CL 2024-12 conditional novelty 6.0 of 10

    A prompt-only LLM framework that generates and executes Python code answers 90 to 93 percent of 100 GTFS transit-data queries without fine-tuning.

  8. Integrating gender inclusivity into large language models via instruction tuning

    cs.CL 2025-08 reject novelty 5.0 of 10

    Abstract promises gender-inclusive Polish LLM tuning with the IPIS dataset, while the full text is an unrelated quantum transformer paper; no evidence for the declared claims is present.

  9. Explaining GitHub Actions Failures with Large Language Models: Challenges, Insights, and Limitations

    cs.SE 2025-01 conditional novelty 5.0 of 10

    A 31-developer survey found that LLM-generated explanations of GitHub Actions failures are perceived as correct and clear for simple logs, but less useful for complex CI/CD failures.

  10. Perspective Transition of Large Language Models for Solving Subjective Tasks

    cs.CL 2025-01 conditional novelty 5.0 of 10

    Reasoning through Perspective Transition (RPT) improves LLM performance on subjective NLP tasks by ranking direct, role, and third-person perspectives by self-reported confidence and answering from the top-ranked perspective.

  11. SuperCode: Sustainability PER AI-driven CO-DEsign

    astro-ph.IM 2024-12 unverdicted novelty 5.0 of 10

    The paper proposes an AI-driven hardware-software-science co-design methodology for radio astronomy, using sustainability as the key performance indicator, with no empirical results yet.

  12. Improving Physics Reasoning in Large Language Models Using Mixture of Refinement Agents

    cs.AI 2024-12 conditional novelty 5.0 of 10

    MoRA uses GPT-4o to detect miscomprehension, wrong-concept, and computational errors in open-source LLM solutions, then routes specialized agents to fix them, improving multiple-choice physics accuracy by up to 16 per...

  13. Leveraging Large Language Models for Institutional Portfolio Management: Persona-Based Ensembles

    cs.CE 2024-11 conditional novelty 5.0 of 10

    Persona-ensembled GPT-4 predictions improve Sharpe ratio over buy-and-hold during rising-CPI months in a 26-month stock-bond backtest, but underperform during falling-CPI months.

  14. Evaluating Large Language Models on Business Process Modeling: Framework, Benchmark, and Self-Improvement Analysis

    cs.DB 2024-11 conditional novelty 5.0 of 10

    A benchmark of 20 business processes and 16 large language models finds Claude-3.5-Sonnet produces the highest-quality process models, and suggests that output optimization improves weaker models.

  15. An Integrated Framework of Prompt Engineering and Multidimensional Knowledge Graphs for Legal Dispute Analysis

    cs.AI 2025-07 conditional novelty 4.0 of 10

    A prompt-plus-knowledge-graph framework for legal dispute analysis reports improved LLM sensitivity and citation accuracy on a 100-pair test set, but with limited statistical support.

  16. Enhancing LLM Reasoning with Multi-Path Collaborative Reactive and Reflection agents

    cs.CL 2024-12 conditional novelty 4.0 of 10

    A multi-path, reactive-plus-reflection agent framework improves gpt-3.5-turbo accuracy on MMLU physics, math, and moral reasoning subsets compared with CoT, self-consistency, and self-refine baselines.

  17. Practical Considerations for Agentic LLM Systems

    cs.AI 2024-12 conditional novelty 3.0 of 10

    This paper is a practical survey that organizes research on LLM-based agents into design considerations for planning, memory, tools, and control flow.

  18. Generalist Virtual Agents: A Survey on Autonomous Agents Across Digital Platforms

    cs.MA 2024-11 conditional novelty 3.0 of 10

    A survey that proposes the Generalist Virtual Agent concept and taxonomies for agent environments, tasks, perceptions, actions, models, and evaluation, concluding that real-world-like environments favor human-like int...

Pith tools