REVIEW 5 cited by
An Empirical Study on Challenges for LLM Application Developers
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
In recent years, large language models (LLMs) have seen rapid advancements, significantly impacting various fields such as computer vision, natural language processing, and software engineering. These LLMs, exemplified by OpenAI's ChatGPT, have revolutionized the way we approach language understanding and generation tasks. However, in contrast to traditional software development practices, LLM development introduces new challenges for AI developers in design, implementation, and deployment. These challenges span different areas (such as prompts, APIs, and plugins), requiring developers to navigate unique methodologies and considerations specific to LLM application development. Despite the profound influence of LLMs, to the best of our knowledge, these challenges have not been thoroughly investigated in previous empirical studies. To fill this gap, we present the first comprehensive study on understanding the challenges faced by LLM developers. Specifically, we crawl and analyze 29,057 relevant questions from a popular OpenAI developer forum. We first examine their popularity and difficulty. After manually analyzing 2,364 sampled questions, we construct a taxonomy of challenges faced by LLM developers. Based on this taxonomy, we summarize a set of findings and actionable implications for LLM-related stakeholders, including developers and providers (especially the OpenAI organization).
Forward citations
Cited by 5 Pith papers
-
Small Models, Big Tasks: An Exploratory Empirical Study on Small Language Models for Function Calling
Fine-tuned small language models can reach high accuracy on function calling, but zero-shot and few-shot performance is poor, and the study's few-shot results are compromised by using test-set examples in the prompt.
-
Rule-ATT&CK Mapper (RAM): Mapping SIEM Rules to TTPs Using LLMs
A multi-stage LLM agent pipeline (RAM) maps structured SIEM rules to MITRE ATT&CK techniques, achieving AR 0.75 and AP 0.52 with GPT-4-Turbo on recent Splunk rules.
-
Adaptive Testing for LLM-Based Applications: A Diversity-based Approach
A farthest-first diversity-based selection method for prompt templates finds LLM failures faster than random selection, with compression distance giving the strongest average gains.
-
Defining and Detecting the Defects of the Large Language Model-based Autonomous Agents
This study defines eight defect types for LLM-based agents and presents Agentable, a CPG-plus-LLM static analysis tool that detects them with reported precision of 88.79% and recall of 91.03%.
-
Developer Challenges on Large Language Models: A Study of Stack Overflow and OpenAI Developer Forum Posts
A BERTopic analysis of 8,593 Stack Overflow posts and 26,474 OpenAI Developer Forum posts yields 9 and 17 LLM developer challenge topics, with API usage dominant and high unresolved rates.
Discussion (0). Continue with ORCID to comment.