Pith. sign in

REVIEW 20 cited by

Domain Specialization as the Key to Make Large Language Models Disruptive: A Comprehensive Survey

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.18703 v7 pith:LWCBQGVZ submitted 2023-05-30 cs.CL cs.AI

classification cs.CLcs.AI
keywords domainlanguagelargellmsapplicationsmodelscomprehensivetechniques
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Large language models (LLMs) have significantly advanced the field of natural language processing (NLP), providing a highly useful, task-agnostic foundation for a wide range of applications. However, directly applying LLMs to solve sophisticated problems in specific domains meets many hurdles, caused by the heterogeneity of domain data, the sophistication of domain knowledge, the uniqueness of domain objectives, and the diversity of the constraints (e.g., various social norms, cultural conformity, religious beliefs, and ethical standards in the domain applications). Domain specification techniques are key to make large language models disruptive in many applications. Specifically, to solve these hurdles, there has been a notable increase in research and practices conducted in recent years on the domain specialization of LLMs. This emerging field of study, with its substantial potential for impact, necessitates a comprehensive and systematic review to better summarize and guide ongoing work in this area. In this article, we present a comprehensive survey on domain specification techniques for large language models, an emerging direction critical for large language model applications. First, we propose a systematic taxonomy that categorizes the LLM domain-specialization techniques based on the accessibility to LLMs and summarizes the framework for all the subcategories as well as their relations and differences to each other. Second, we present an extensive taxonomy of critical application domains that can benefit dramatically from specialized LLMs, discussing their practical significance and open challenges. Last, we offer our insights into the current research status and future trends in this area.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 20 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. When LLM Meets Time Series: Can LLMs Perform Multi-Step Time Series Reasoning and Inference

    cs.LG 2025-09 conditional novelty 6.0 of 10

    TSAIA, a new benchmark, tests eight LLMs on 1,054 multi-step time series tasks and finds they cannot reliably complete the required workflows.

  2. ReCatcher: Towards LLMs Regression Testing for Code Generation

    cs.SE 2025-07 conditional novelty 6.0 of 10

    ReCatcher systematically measures regressions in LLM code generation across correctness, static quality, and performance, and its evaluation shows fine-tuning, merging, and new releases each introduce specific regressions.

  3. Toward Structured Knowledge Reasoning: Contrastive Retrieval-Augmented Generation on Experience

    cs.CL 2025-06 conditional novelty 6.0 of 10

    CoRE improves structured knowledge reasoning by retrieving both correct and incorrect past examples into the prompt, using MCTS-generated experience memory.

  4. QualBench: Benchmarking Chinese LLMs with Localized Professional Qualifications for Vertical Domain Evaluation

    cs.CL 2025-05 conditional novelty 6.0 of 10

    QualBench offers a 17,316-question Chinese evaluation from professional qualification exams, and Chinese models outperform non-Chinese models on it.

  5. All-in-One Tuning and Structural Pruning for Domain-Specific LLMs

    cs.CL 2024-12 conditional novelty 6.0 of 10

    ATP jointly searches for pruning decisions and fine-tunes LLaMA models with LoRA in one stage, outperforming two-stage pruning on domain-specific tasks.

  6. Large Action Models: From Inception to Implementation

    cs.AI 2024-12 conditional novelty 6.0 of 10

    A four-phase training pipeline converts a 7B language model into a Windows GUI action model that reaches 81.2% offline and 71.0% online task success on the authors' Word test set, beating text-only GPT-4o.

  7. KBAlign: Efficient Self Adaptation on Specific Knowledge Bases

    cs.CL 2024-11 conditional novelty 6.0 of 10

    KBAlign is a self-supervised method that generates multi-grained QA pairs from a small text knowledge base and iteratively self-verifies to adapt a RAG model, reaching about 90% of GPT-4-supervised gains on LooGLE F1.

  8. CEQuest: Benchmarking Large Language Models for Construction Estimation

    cs.CL 2025-08 conditional novelty 5.0 of 10

    A new 164-question benchmark shows that current LLMs score between 62% and 75% on construction drawing interpretation and estimation tasks.

  9. Climate-Eval: A Comprehensive Benchmark for NLP Tasks Related to Climate Change

    cs.CL 2025-05 conditional novelty 5.0 of 10

    ClimateEval unifies 25 climate-related NLP tasks into one benchmark and shows that open-source LLMs gain from few-shot examples but lag on misinformation and fine-grained entity recognition.

  10. Keep the General, Inject the Specific: Structured Dialogue Fine-Tuning for Knowledge Injection without Catastrophic Forgetting

    cs.CL 2025-04 conditional novelty 5.0 of 10

    Structured Dialogue Fine-Tuning (SDFT) injects domain knowledge into vision-language models through caption, contrastive, and specialization dialogue turns, reporting improved specialization with modest general-capabi...

  11. TelcoLM: collecting data, adapting, and benchmarking language models for the telecommunication domain

    cs.CL 2024-12 conditional novelty 5.0 of 10

    For Llama-2-7B on telecommunications tasks, instruction tuning on a telco-generated dataset suffices; continuing pretraining on raw telco text adds little (max +0.03 accuracy).

  12. GEE-OPs: An Operator Knowledge Base for Geospatial Code Generation on the Google Earth Engine Platform Powered by Large Language Models

    cs.SE 2024-12 conditional novelty 5.0 of 10

    A structured operator knowledge base mined from Google Earth Engine scripts improves LLM-generated geospatial code by 20-30 percentage points in executability and correctness when used with retrieval-augmented generation.

  13. Improving Topic Modeling of Social Media Short Texts with Rephrasing: A Case Study of COVID-19 Related Tweets

    cs.CL 2025-10 conditional novelty 4.0 of 10

    LLM rephrasing of tweets before topic modeling raises Wikipedia-measured coherence (LDA 0.31→0.50) but the abstract's claim of broad improvements is contradicted by the paper's own table for LDA and by the metric choice.

  14. LLMREI: Automating Requirements Elicitation Interviews with LLMs

    cs.SE 2025-07 conditional novelty 4.0 of 10

    A GPT-4o chatbot using zero-shot or iteratively refined prompts can conduct requirements elicitation interviews with error rates comparable to trained human interviewers in a simulated student study.

  15. Toward Edge General Intelligence with Multiple-Large Language Model (Multi-LLM): Architecture, Trust, and Orchestration

    cs.NI 2025-07 conditional novelty 4.0 of 10

    A survey of multi-LLM systems in edge computing, covering architectures, enabling technologies, trust mechanisms, applications, and open datasets for edge general intelligence.

  16. A Domain Adaptation of Large Language Models for Classifying Mechanical Assembly Components

    cs.LG 2025-05 conditional novelty 4.0 of 10

    Fine-tuning GPT-3.5 Turbo on 681 OSDR parts yields about 89% accuracy on held-out OSDR data and produces function labels for ABC parts, although the ABC labels are never checked against ground truth.

  17. A Comparative Study of DSPy Teleprompter Algorithms for Aligning Large Language Models Evaluation Metrics to Human Evaluation

    cs.CL 2024-12 conditional novelty 4.0 of 10

    On a 375-sample HaluBench test set, MIPROv2 and Bootstrap Few Shot with Random Search achieve the highest weighted and macro F1 scores for LLM-based hallucination detection, though without statistical significance tests.

  18. LP Data Pipeline: Lightweight, Purpose-driven Data Pipeline for Large Language Models

    cs.CL 2024-11 reject novelty 4.0 of 10

    This paper presents a CPU-only data curation pipeline for LLMs, but the central claim of high-quality output is not supported by any training or quality evaluation.

  19. Augmenting Large Language Models with Static Code Analysis for Automated Code Quality Improvements

    cs.SE 2025-06 conditional novelty 3.0 of 10

    Generating fixes with GPT-3.5 Turbo and GPT-4o, prompted with SonarQube findings and web-retrieved examples, removed most flagged bugs, vulnerabilities, and code smells from one codebase, with success judged solely by...

  20. SKETCH: Structured Knowledge Enhanced Text Comprehension for Holistic Retrieval

    cs.CL 2024-12 reject novelty 3.0 of 10

    SKETCH combines semantic chunking and a knowledge graph retriever, and the paper claims it tops Naive RAG, RAPTOR, semantic-only, and KG-only baselines on RAGAS metrics, though the reported results are internally inco...

Pith tools