Pith. sign in

REVIEW 3 major objections 5 minor 28 references

The Use of Artificial Intelligence in Military Intelligence: An Experimental Investigation of Added Value in the Analysis Process

T0 review · 3 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read AI assistance measurably improves military analysts' assessments under time pressure.

desk verdict A real experiment on AI in military analysis with a plausible positive effect, but the unvalidated expert scoring key means 'clearly superior' is not yet established. read the letter →

arxiv 2412.03610 v1 pith:OBAIV6FV submitted 2024-12-04 cs.AI cs.HC

classification cs.AIcs.HC
keywords militaryintelligenceartificiallargelanguagemodelopensourcesemanticsearchnamedentityrecognitiontextsummarizationexperimentalevaluation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper sets out to test whether artificial intelligence actually adds value in the military intelligence analysis process, rather than just promising it. It reports a controlled experiment in which 29 soldiers analyzed a realistic open-source intelligence scenario, the April 2017 Khan Shaykhun poison gas attack, under a 30-minute time limit. Half the participants used the deepCOM demonstrator with three AI functions, semantic search, automatic text summarization, and named entity recognition, while the control group used keyword search and no AI support. The authors claim the AI-assisted group produced assessments clearly closer to the judgments of seven expert military analysts, both on factual questions and on probability estimates, without reporting higher confidence in their own work. For a reader, the significance is concrete: this is some of the first empirical evidence about where AI helps and does not help in intelligence analysis.

What carries the argument

The load-bearing object is the deepCOM demonstrator, a German-language analysis tool built on a large language model, which provides three functions: semantic search that answers whole questions and cites source passages; automatic paragraph-level summarization that reduces texts to one-third to one-half of their length; and a named entity recognition module that tags time, place, organization, and person mentions. The experimental machinery is a randomized comparison: 29 soldiers were split into an AI-supported group and a control group, given the same 50-report database and 30 minutes, and their answers were scored by proximity to the average judgment of seven experts who worked without a time limit. The argument hangs on that scoring key: 'superior' means closer to the experts' consensus.

What would settle it

Re-run the same scenario with the expert baseline checked against independently verified facts of the Khan Shaykhun attack, such as casualty counts and responsible actors per declassified records, or add a no-time-limit condition: if the control group then matches the AI group's accuracy, the claimed advantage is limited to time-pressured settings. A simpler sign: if the seven experts' individual judgments diverge substantially on the scoring items, the averaged baseline used here may not be stable enough to support the result.

Watch

Extended reading notes

Core claim

The central claim is that, under time pressure, the combined use of AI-based text search, automatic summarization, and named entity recognition improves the quality of military intelligence analysis. The experimental group scored more than six and a half points higher on the factual task (M = 18.2 vs 11.5, p = 0.007) and showed significantly smaller deviations from expert probability judgments than the control group (difference 0.851 vs 1.039, p = 0.047). The paper is careful to limit the claim: the benefit appeared mainly on direct, factual questions and faded on more complex or argumentative items, and confidence in one's own assessment did not rise alongside the objective improvement. The authors also report that this advantage was measured against a baseline of seven experienced military analysts who completed the same task with unlimited time.

Load-bearing premise

The paper's measure of 'correct' analysis is the average judgment of seven military intelligence experts who did the same task without a time limit; if those experts are biased or unrepresentative, the AI group's higher score only means it moved closer to those seven people's opinions.

Editorial extensions

If this is right

  • If the claim holds, AI support can be positioned at the analysis and production stage of the intelligence cycle, not just in collection.
  • The benefit appears largest for questions with short factual answers; for argumentative or ambiguous questions, AI adds little measurable value.
  • Because confidence did not rise with accuracy, AI assistance in this setting does not seem to produce overconfidence, a key worry for military use.
  • The perceived speed gain reported by participants, combined with the experts needing 3 hours 49 minutes versus 30 minutes in the experiment, suggests time savings as well as accuracy gains.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The expert-judgment baseline is an assumption, not a ground truth; the result would be more decisive if the experts' own answers were validated against independently established facts about the Khan Shaykhun event.
  • Because the three AI functions were tested only in combination, the experiment cannot tell which function carries the effect; a factorial design could isolate whether NER, summarization, or search alone accounts for the gains.
  • A natural extension is to vary the time pressure: if the control group matches the AI group when given unlimited time, the AI's contribution is specifically about compressing the analysis timeline rather than raising the ceiling of accuracy.
  • The four-day reporting window of the scenario means the finding may not transfer to long-term monitoring, where contradictory information accumulates and source reliability varies.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper reports an experimental study (n=29) in which active-duty soldiers analyzed a realistic open-source intelligence scenario (the 2017 Khan Shaykhun poison gas attack) under a 30-minute time limit, either with or without three AI functions in a demonstrator called deepCOM: LLM-based semantic search, automatic summarization, and named entity recognition. Performance was measured as closeness to the judgments of seven experienced military intelligence experts who completed the same tasks without a time limit. The experimental group scored higher overall on factual analysis (18.2 vs. 11.5, p=0.007) and on probability estimates (0.85 vs. 1.04 deviation, p=0.047), with no difference in self-reported confidence. The paper interprets these results as evidence that the combined AI functions provide added value for factual and probability judgments under time pressure, while also listing limitations including the single scenario and the inability to attribute value to individual functions.

Significance. If the central result is robust, the study is a valuable contribution to the empirical literature on AI support in military intelligence, a domain where controlled experiments are rare. The use of a realistic, documented scenario, random assignment, and a pre-specified task structure (albeit with post-hoc task blocking) are strengths, as is the candid discussion of limitations in Section 6.4. The paper also makes the useful observation that AI assistance improved factual outcome accuracy without inflating confidence, which has operational implications. The main obstacles to accepting the claims as stated are the unvalidated expert baseline and the statistical reporting around the task-level tests. A revision that addresses these issues, or qualifies the conclusions accordingly, would substantially strengthen the paper.

major comments (3)
  1. [§4.2 and §5] The outcome measure for Part 1 is the distance between each participant's answers and the judgments of seven experts, but the manuscript does not report (i) how expert answers were converted into item scores, (ii) whether a scoring rubric was fixed before participant responses were read, (iii) whether scoring was blind to treatment group, or (iv) any inter-rater reliability statistic (e.g., percent agreement or Krippendorff's alpha) for the seven experts. The full items in Annex B contain factual questions (e.g., 3a, 3b, 6c) that are checkable against the public record of the Khan Shaykhun attack and the US strike on Al Shayrat; the paper does not use this external record to validate the expert baseline. The abstract's claim of 'clearly superior' assessments is therefore underdetermined: it may simply reflect convergence to one of several expert response patterns. The authors should either provide reliability and validity evidence or substantially qualify the conclusion.
  2. [§5, Tables 2 and 3] The tables label the test statistic as χ², but the text describes 'mean differences in independent samples' without stating the test used (e.g., Welch's t, Mann-Whitney U, or Kruskal-Wallis), whether assumptions were verified, or effect sizes. The 21 items are split into seven task blocks (and six probability blocks), and each block is tested separately with no correction for multiple comparisons. The split of Task 6 into 6/1 and 6/2 appears to be made during analysis, and the grouping of tasks by significance in Section 6.1 is derived from the same pattern of p-values, making the 'complexity' interpretation post-hoc. This does not invalidate the overall comparison (p=0.007), but it leaves the task-level claims in Table 2 and Table 3 with an inflated Type I error rate and a risk of circular interpretation.
  3. [§4.2, §6.4] With n=14 and n=15, a single scenario, and no pre-registration, the study is a small convenience-sample experiment. The authors acknowledge the single-scenario limitation and the combination-only design in Section 6.4, which is commendable, but the title and abstract go beyond what the design can support ('clearly superior'), and the claim in Section 1 that this is 'the first study to empirically analyze the added value of AI in the context of intelligence' is too strong without a more systematic literature review. The authors should report effect sizes with confidence intervals, discuss statistical power, and frame the conclusion as a proof-of-concept for this specific scenario and AI combination.
minor comments (5)
  1. [§4.2 and §5] The text gives 30 minutes for the analysis task in Section 4.2 but later refers to a 25-minute time limit for the first part of the analysis task in Section 5; please reconcile the timeline.
  2. [Annex A] Several source titles contain typos or formatting artifacts (e.g., 'Spigel Online' for 'Spiegel Online', 'Refueat' for 'Ruefat?'), and some URLs contain line-break spaces; please check and standardize.
  3. [§6.2] The sentence 'In the context of the labeling of aviation accident documents...' appears without a citation or connection to the study; it should be tied to the cited prior work (likely Perboli et al. [22]) or removed.
  4. [§5] The correlations of performance with age and gender are reported with χ² values; please specify the statistical test and whether these were pre-specified or exploratory.
  5. [§6.1] The division of tasks into Groups 1-3 based on observed significance levels is presented without any sensitivity analysis or correction; at minimum, the exploratory nature should be flagged in the results section, not only in the limitations.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the AI-added-value claim is an empirical group comparison against an independently collected expert scoring key.

full rationale

The paper's central claim—that the AI-supported experimental group outperformed the control group under time pressure—rests on comparing participant answers to a scoring key produced by seven military intelligence experts who completed the same task without time pressure (Sections 4.2 and 5). This is an external benchmark, not a quantity fitted from the participants' responses, and it is not defined in terms of the experimental manipulation. The paper's own limitation section acknowledges the benchmark's ceiling: "It cannot be excluded that the experimental group performed better than the experts due to the support of the AI functions, but that this could not be measured" (Section 6.4), which shows the authors treat expert performance as a measurement ceiling rather than as a constructed target. The self-citations ([3], [21], [27]) appear in background or related-work contexts (keyword extraction, NER description, intelligence-cycle diagram) and are not load-bearing for the experimental result; no uniqueness theorem is imported, and no fitted parameter is relabeled as a prediction. The skeptical concern about inter-rater reliability or validity of the expert baseline is a correctness and measurement limitation, not circularity, because the paper does not claim to derive the superiority result from the baseline by construction. Accordingly, the circularity score is 0.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

This is an empirical study with no fitted parameters or invented theoretical entities. It relies on domain assumptions about the representativeness of the scenario, the validity of expert judgments as ground truth, the use of soldiers as proxy analysts, and the effectiveness of random assignment in a small sample.

assumptions (4)
  • domain assumption Active-duty soldiers aged 20-33 are a valid proxy for military intelligence analysts.
    Participants are soldiers, not necessarily trained analysts; the task requires general reading and source-citation skills, but analytical experience may differ. Invoked in Section 4.2.
  • domain assumption The seven experts' untimed average judgments define the correct answers for all scoring.
    All performance scores are computed as deviations from expert judgment, so the experts' accuracy is assumed rather than independently verified. Section 4.2 and Section 5.
  • domain assumption The Khan Shaykhun scenario and the 50 selected sources are representative of real military intelligence analysis tasks.
    The authors note in Section 6.4 that only one scenario was used and that generalization to other time windows or source types is untested.
  • domain assumption Random assignment balanced all unmeasured participant characteristics despite the small sample.
    With n=29, chance imbalances are possible; the authors report no significant age or gender correlations with performance, but statistical power is low. Section 4.2.

how reviews work

0 comments
Cite this review

Pith. "Pith review of The Use of Artificial Intelligence in Military Intelligence: An Experimental Investigation of Added Value in the Analysis Process." pith.science (2026). https://pith.science/paper/OBAIV6FV

@misc{pith2026241203610,
  author       = {Pith},
  title        = {Pith review of: The Use of Artificial Intelligence in Military Intelligence: An Experimental Investigation of Added Value in the Analysis Process},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/OBAIV6FV}},
  note         = {Machine review of arXiv:2412.03610}
}
read the original abstract

It is beyond dispute that the potential benefits of artificial intelligence (AI) in military intelligence are considerable. Nevertheless, it remains uncertain precisely how AI can enhance the analysis of military data. The aim of this study is to address this issue. To this end, the AI demonstrator deepCOM was developed in collaboration with the start-up Aleph Alpha. The AI functions include text search, automatic text summarization and Named Entity Recognition (NER). These are evaluated for their added value in military analysis. It is demonstrated that under time pressure, the utilization of AI functions results in assessments clearly superior to that of the control group. Nevertheless, despite the demonstrably superior analysis outcome in the experimental group, no increase in confidence in the accuracy of their own analyses was observed. Finally, the paper identifies the limitations of employing AI in military intelligence, particularly in the context of analyzing ambiguous and contradictory information.

Figures

Figures reproduced from arXiv: 2412.03610 by the authors.

Figure 1
Figure 1. illustrates the typical intelligence cycle [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Sample AI search answers to the questions ’How was the US air strike on [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗
Figure 3
Figure 3. Top image: NER automatically extracts time, place, organization, and [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figures from the paper (5 more)
Figure 4
Figure 4. Figure 4: Example of an automated text summary: The text ’On Tuesday morning [PITH_FULL_IMAGE:figures/full_fig_p008_4.png]
Figure 5
Figure 5. Figure 5: Examples of tasks for the two parts of the analysis task. [PITH_FULL_IMAGE:figures/full_fig_p010_5.png]
Figure 6
Figure 6. Figure 6: Percentage mean value per item separately for experimental and control [PITH_FULL_IMAGE:figures/full_fig_p011_6.png]
Figure 7
Figure 7. Figure 7: Absolute value of the difference between the experimental group and the [PITH_FULL_IMAGE:figures/full_fig_p013_7.png]
Figure 8
Figure 8. Figure 8: Examples of items for each of the three groups of items, in order of [PITH_FULL_IMAGE:figures/full_fig_p016_8.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

28 extracted references · 24 canonical work pages

  1. [1]

    In: Computer Aided Systems Theory - EUROCAST 2022 - 18th International Conference, Las Palmas de Gran Canaria, Spain, Feb 20–25, 2022, Revised Selected Papers

    Albrecht, K., Nitzl, C., Borghoff, U.M.: Transdisciplinary software development for early crisis detection. In: Computer Aided Systems Theory - EUROCAST 2022 - 18th International Conference, Las Palmas de Gran Canaria, Spain, Feb 20–25, 2022, Revised Selected Papers. pp. 3–10. Lecture Notes in Computer Science 13789, Springer (2022). https://doi.org/10.10...

  2. [2]

    Digital Society 2(1), 12 (2023)

    Blanchard, A., Taddeo, M.: The ethics of artificial intelligence for intelligence anal- ysis: a review of the key challenges with recommendations. Digital Society 2(1), 12 (2023)

  3. [3]

    In: Hardy, M.R.B., Tompa, F.W

    Bohne, T., R¨ onnau, S., Borghoff, U.M.: Efficient keyword extraction for meaningful document perception. In: Hardy, M.R.B., Tompa, F.W. (eds.) Proceedings of the 2011 ACM Symposium on Document Engineering, Mountain View, CA, USA, Sep 19-22, 2011. pp. 185–194. ACM (2011). https://doi.org/10.1145/2034691.2034732

  4. [4]

    IEEE Access (2023)

    Chamola, V., Hassija, V., Sulthana, A.R., Ghosh, D., Dhingra, D., Sikdar, B.: A review of trustworthy and explainable artificial intelligence (xai). IEEE Access (2023)

  5. [5]

    Electronics 9(12), 2187 (2020)

    Cho, S., Shin, W., Kim, N., Jeong, J., In, H.P.: Priority determination to apply artificial intelligence technology in military intelligence areas. Electronics 9(12), 2187 (2020)

  6. [6]

    California, USA: CQ Press, SAGE Publications (2013)

    Clark, R.M.: Intelligence collection. California, USA: CQ Press, SAGE Publications (2013)

  7. [7]

    California, USA: CQ Press, SAGE Publications (2019)

    Clark, R.M.: Intelligence analysis: a target-centric approach. California, USA: CQ Press, SAGE Publications (2019)

  8. [8]

    arXiv preprint arXiv:1810.04805 (2018)

    Devlin, J.: Bert: Pre-training of deep bidirectional transformers for language un- derstanding. arXiv preprint arXiv:1810.04805 (2018)

Show all 28 references
  1. [9]

    ACM Computing Surveys 55(9), 1–33 (2023)

    Dwivedi, R., et al.: Explainable ai (xai): Core ideas, techniques, and solutions. ACM Computing Surveys 55(9), 1–33 (2023)

  2. [10]

    Studies in Intelligence 63(2), 1–5 (2019)

    Gartin, J.W.: The future of analysis. Studies in Intelligence 63(2), 1–5 (2019)

  3. [11]

    Political Analysis 31(4), 481–499 (2023)

    H¨ affner, S., Hofer, M., Nagl, M., Walterskirchen, J.: Introducing an interpretable deep learning approach to domain-specific dictionary creation: A use case for con- flict prediction. Political Analysis 31(4), 481–499 (2023)

  4. [12]

    Intelligence and National Security 31(6), 858–870 (2016)

    Hare, N., Coghill, P.: The future of the intelligence analysis task. Intelligence and National Security 31(6), 858–870 (2016)

  5. [13]

    Intelligence and National Security 38(3), 447–469 (2023)

    Horlings, T.: Dealing with data: coming to grips with the information age in intel- ligence studies journals. Intelligence and National Security 38(3), 447–469 (2023)

  6. [14]

    Intelligence and national Security 21(6), 959–979 (2006) 28 C

    Hulnick, A.S.: What’s wrong with the intelligence cycle. Intelligence and national Security 21(6), 959–979 (2006) 28 C. Nitzl, A. Cyran, S. Krstanovic & U. M. Borghoff

  7. [15]

    International Journal of In- telligence and Counter Intelligence 1(4), 1–23 (1986)

    Johnson, L.K.: Making the intelligence “cycle” work. International Journal of In- telligence and Counter Intelligence 1(4), 1–23 (1986)

  8. [16]

    arXiv preprint arXiv:1603.01360 (2016)

    Lample, G., et al.: Neural architectures for named entity recognition. arXiv preprint arXiv:1603.01360 (2016)

  9. [17]

    Expert Systems with Applications 167, 114152 (2021)

    Lamsiyah, S., El Mahdaouy, A., Espinasse, B., Ouatik, S.E.A.: An unsupervised method for extractive multi-document summarization based on centroid approach and sentence embeddings. Expert Systems with Applications 167, 114152 (2021)

  10. [18]

    International Journal of Human–Computer Interaction 34(7), 577–590 (2018)

    Lewis, J.R.: The system usability scale: past, present, and future. International Journal of Human–Computer Interaction 34(7), 577–590 (2018)

  11. [19]

    California, USA: CQ Press, SAGE Publications (2022)

    Lowenthal, M.M.: Intelligence: From secrets to policy. California, USA: CQ Press, SAGE Publications (2022)

  12. [20]

    https://jadl.act.nato.int (2016)

    NATO: AJP-2.1 Allied joint doctrine for intelligence procedures. https://jadl.act.nato.int (2016)

  13. [21]

    In: Computer Aided Systems Theory - EUROCAST 2024 - 19th International Conference, Las Palmas de Gran Canaria, Spain, Feb 25 – Mar 1, 2024, Revised Selected Papers

    Nitzl, C., Cyran, A., Krstanovic, S., Borghoff, U.M.: The application of named entity recognition in military intelligence. In: Computer Aided Systems Theory - EUROCAST 2024 - 19th International Conference, Las Palmas de Gran Canaria, Spain, Feb 25 – Mar 1, 2024, Revised Selec...

  14. [22]

    Expert Systems with Applications 186, 115694 (2021)

    Perboli, G., Gajetti, M., Fedorov, S., Giudice, S.L.: Natural language processing for the identification of human factors in aviation accidents causes: An application to the shel methodology. Expert Systems with Applications 186, 115694 (2021)

  15. [23]

    Routledge London (2013)

    Phythian, M., et al.: Understanding the intelligence cycle. Routledge London (2013)

  16. [24]

    In: IEEE International Engi- neering Conference (IEC)

    Qader, W.A., Ameen, M.M., Ahmed, B.I.: An overview of bag of words; impor- tance, implementation, applications, and challenges. In: IEEE International Engi- neering Conference (IEC). pp. 200–204 (2019)

  17. [25]

    A Primer on Multiple Intelli- gences pp

    Sadiku, M.N.O., Musa, S.M.: Military intelligence. A Primer on Multiple Intelli- gences pp. 249–262 (2021)

  18. [26]

    Intelligence and National Security 36(6), 827–848 (2021)

    Vogel, K.M., Reid, G., Kampe, C., Jones, P.: The impact of ai on intelligence analysis: tackling issues of collaboration, algorithmic transparency, accountability, and management. Intelligence and National Security 36(6), 827–848 (2021)

  19. [27]

    arXiv preprint arXiv:2405.06957 (2024)

    Werro, A., Nitzl, C., Borghoff, U.M.: On the role of intelligence and business wargaming in developing foresight. arXiv preprint arXiv:2405.06957 (2024)

  20. [28]

    arXiv preprint arXiv:1910.11470 (2019)

    Yadav, V., Bethard, S.: A survey on recent advances in named entity recognition from deep learning models. arXiv preprint arXiv:1910.11470 (2019)

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.