Pith. sign in

REVIEW 2 cited by

Fine-Tuning and Evaluating Open-Source Large Language Models for the Army Domain

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.20297 v1 pith:F6CT34US submitted 2024-10-27 cs.CL cs.AIcs.LGcs.NE

classification cs.CLcs.AIcs.LGcs.NE
keywords armyllmsdomainfine-tuningmodelstraclmaddresscases
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In recent years, the widespread adoption of Large Language Models (LLMs) has sparked interest in their potential for application within the military domain. However, the current generation of LLMs demonstrate sub-optimal performance on Army use cases, due to the prevalence of domain-specific vocabulary and jargon. In order to fully leverage LLMs in-domain, many organizations have turned to fine-tuning to circumvent the prohibitive costs involved in training new LLMs from scratch. In light of this trend, we explore the viability of adapting open-source LLMs for usage in the Army domain in order to address their existing lack of domain-specificity. Our investigations have resulted in the creation of three distinct generations of TRACLM, a family of LLMs fine-tuned by The Research and Analysis Center (TRAC), Army Futures Command (AFC). Through continuous refinement of our training pipeline, each successive iteration of TRACLM displayed improved capabilities when applied to Army tasks and use cases. Furthermore, throughout our fine-tuning experiments, we recognized the need for an evaluation framework that objectively quantifies the Army domain-specific knowledge of LLMs. To address this, we developed MilBench, an extensible software framework that efficiently evaluates the Army knowledge of a given LLM using tasks derived from doctrine and assessments. We share preliminary results, models, methods, and recommendations on the creation of TRACLM and MilBench. Our work significantly informs the development of LLM technology across the DoD and augments senior leader decisions with respect to artificial intelligence integration.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. RRS-10K: A Multitask Vision-Language Model Benchmark for Rare Remote Sensing Image Interpretation

    cs.AI 2026-07 conditional novelty 6.0 of 10

    Current vision-language models score only moderately on a new 10,738-image military remote-sensing benchmark and collapse on visual grounding and referring segmentation.

  2. Clone What You Can't Steal: Black-Box LLM Replication via Logit Leakage and Distillation

    cs.CR 2025-08 reject novelty 4.0 of 10

    A two-stage attack claims to recover a black-box LLM's output projection from under 10k top-k logit queries and distill a compact clone, but the core matrix-completion step is not justified.

Pith tools