Pith. sign in

REVIEW 12 cited by

Fine-tuning Large Language Models for Domain-specific Machine Translation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.15061 v2 pith:ZXQJ6753 submitted 2024-02-23 cs.CL cs.LG

classification cs.CLcs.LG
keywords domain-specificdragftllmsperformanceexamplesfine-tuningmodelsthree
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Large language models (LLMs) have shown great potential in domain-specific machine translation (MT). However, one major issue is that LLMs pre-trained on general domain corpus might not generalize well to specific domains due to the lack of domain-specific knowledge. To address this issue, this paper focuses on enhancing the domain-specific MT capability of LLMs, by providing high-quality training datasets and proposing a novel fine-tuning framework denoted by DragFT. DragFT augments LLMs via three techniques: (i) Dictionary-enhanced prompting integrates dictionary information into prompts to improve the translation of domain-specific terminology.; (ii) RAG-based few-shot example selection provides high-quality examples that simulate both the domain and style characteristics; (iii) Fine-tuning with few-shot examples further enhances performance when using in-domain examples. We deploy DragFT on three well-known LLM backbones with 13B training parameters to validate its effectiveness. The results on three domain-specific datasets show that DragFT achieves a significant performance boost and shows superior performance compared to advanced models such as GPT-3.5 and GPT-4o. The drastic performance improvement of DragFT over existing LLMs can be attributed to incorporating relevant knowledge while mitigating noise.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Training-Free Token-Level Steering for LLM Personalized Co-Writing

    cs.CL 2026-08 conditional novelty 6.0 of 10

    A token-level, training-free steering framework that improves LLM personalized co-writing by mixing the base model's posterior with a kernel-density estimate from a small user dataset.

  2. TAT-R1: Terminology-Aware Translation with Reinforcement Learning and Word Alignment

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Word-alignment rewards for RL-trained translation raise terminology accuracy on RTT from 54.42 to 56.42 TA without hurting general translation quality.

  3. Robot Operation of Home Appliances by Reading User Manuals

    cs.RO 2025-05 conditional novelty 6.0 of 10

    A robot system that constructs a symbolic appliance model from a user manual and uses it to reliably execute natural language appliance operation tasks, outperforming direct VLM-based policies.

  4. All-in-One Tuning and Structural Pruning for Domain-Specific LLMs

    cs.CL 2024-12 conditional novelty 6.0 of 10

    ATP jointly searches for pruning decisions and fine-tunes LLaMA models with LoRA in one stage, outperforming two-stage pruning on domain-specific tasks.

  5. Position: Restructuring of Categories and Implementation of Guidelines Essential for VLM Adoption in Healthcare

    cs.CY 2025-05 conditional novelty 5.0 of 10

    The authors propose a four-category taxonomy of healthcare vision-language model studies with category-specific reporting standards and a peer-review checklist, arguing existing ML reporting guidelines are unfit for VLMs.

  6. Collaborative Editable Model

    cs.AI 2025-06 reject novelty 4.0 of 10

    CoEM scores user-contributed knowledge fragments using user ratings and LLM attribution, keeps the high scorers in a prompt-level knowledge pool, and reports 76% agreement with FinGPT on fragment value.

  7. Privacy-Preserving Transformers: SwiftKey's Differential Privacy Implementation

    cs.CL 2025-05 reject novelty 4.0 of 10

    Microsoft's small DP-finetuned transformer for keyboard prediction beats an older GRU in offline tests but shows no aggregate live gain, and its privacy guarantee for the modified sampling is unproven.

  8. Optimizing GPT for Video Understanding: Zero-Shot Performance and Prompt Engineering

    cs.CV 2025-02 reject novelty 4.0 of 10

    Zero-shot GPT-4o comes close to production classifiers on several video-moderation tasks, and prompt simplification plus decomposition-aggregation improves accuracy, though key gains are partly fitted to the test set.

  9. Aligning Knowledge Graphs and Language Models for Factual Accuracy

    cs.CL 2025-07 conditional novelty 3.0 of 10

    ALIGNed-LLM aligns knowledge graph entity embeddings with language model text embeddings through a trainable projection layer, improving question answering accuracy on KG-derived datasets.

  10. FuxiMT: Sparsifying Large Language Models for Chinese-Centric Multilingual Machine Translation

    cs.CL 2025-05 reject novelty 3.0 of 10

    FuxiMT combines a frozen BLOOMz model with sparse mixture-of-experts layers, Chinese-first pretraining, and curriculum learning to translate into Chinese from 65 languages, with claimed low-resource gains that the pap...

  11. Exploring Variability in Fine-Tuned Models for Text Classification with DistilBERT

    cs.CL 2024-12 reject novelty 3.0 of 10

    A regression study of 55 Hugging Face model cards reports correlations between hyperparameters and metrics, but does not establish comparability or causality.

  12. LIMBA: An Open-Source Framework for the Preservation and Valorization of Low-Resource Languages using Generative Models

    cs.CL 2024-11 conditional novelty 3.0 of 10

    LIMBA is a proposed pipeline that combines collection, grammatical tagging, translation, speech, and generative modules to build language models for low-resource languages, with preliminary Sardinian experiments.

Pith tools