Pith. sign in

REVIEW 2 major objections 2 minor 10 references

Glossary-augmented prompting improves terminology accuracy in rock art translation to 81.4 percent while preserving overall quality.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

Glossary-augmented LLM prompting reaches 81.4% terminology accuracy on rock art documents versus 69.1% for basic LLM and 64.4% for DeepL, with no loss in overall quality scores.

T0 review reviewed 2026-06-30 challenge →

load-bearing objection Glossary-augmented Gemini prompting lifts terminology accuracy from 69% to 81% on one Spanish rock art text while holding overall DA scores steady, but the single-document design caps how much practical advice follows. the 2 major comments →

arxiv 2605.14679 v1 pith:LYPT5XYY submitted 2026-05-14 cs.CL cs.AI

AI-assisted cultural heritage dissemination: Comparing NMT and glossary-augmented LLM translation in rock art documents

classification cs.CL cs.AI
keywords machine translationcultural heritageterminology accuracyLLM promptingrock artNMTglossary augmentation
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper compares three translation approaches for a Spanish academic text on rock art: a standard neural machine translation system, a basic large language model prompt, and the same model with glossary-augmented prompting that retrieves term pairs. It reports that the glossary method raises exact-match terminology accuracy from 69.1 percent and 64.4 percent to 81.4 percent. Overall quality scores remain essentially unchanged at around 85 on direct assessment, exceeding the baseline neural system at 80.3. The work positions this as a low-overhead intervention that cultural heritage institutions can adopt if they maintain small term lists and simple evaluation steps.

Core claim

Gemini-RAG with glossary-augmented prompting via term-pair retrieval yields the highest exact-match terminology accuracy at 81.4 percent, versus 69.1 percent for Gemini-Simple and 64.4 percent for DeepL, while preserving overall quality with mean direct assessment scores of 85.3 versus 85.2 and outperforming DeepL at 80.3.

What carries the argument

Glossary-augmented prompting via term-pair retrieval, which supplies relevant specialized term pairs to the model prompt to enforce consistent vocabulary in terminology-dense text.

Load-bearing premise

The single Spanish rock art academic text and the chosen human evaluation protocol using direct assessment plus restricted MQM are representative of the broader cultural heritage domain.

What would settle it

A follow-up evaluation on a second rock art document or a different heritage subdomain that shows no gain in exact-match terminology accuracy from glossary augmentation would falsify the reported benefit.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • Cultural heritage institutions can achieve better term consistency in multilingual outputs by maintaining minimal terminology resources.
  • Lightweight human evaluation procedures are sufficient to confirm terminology gains without complex model changes.
  • Glossary-augmented methods offer a practical alternative to standard neural machine translation for specialized domains.
  • Quality scores remain comparable, allowing adoption without sacrificing overall translation fidelity.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The same retrieval-augmented approach could be tested in other term-heavy fields such as legal or scientific translation to check transferability.
  • Small maintained glossaries might lower long-term costs for institutions that produce repeated multilingual materials.
  • Expanding the test set beyond one document would clarify whether the accuracy lift holds across varying text lengths and term densities.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 2 minor

Summary. The paper compares three English MT setups for one Spanish academic rock art text: DeepL (NMT baseline), Gemini-Simple (basic LLM prompt), and Gemini-RAG (glossary-augmented LLM prompting). Human evaluation via Direct Assessment (0-100) and restricted MQM terminology auditing shows Gemini-RAG at 81.4% exact-match terminology accuracy (vs. 69.1% and 64.4%), with mean DA scores of 85.3 (RAG), 85.2 (Simple), and 80.3 (DeepL). The authors conclude that glossary-augmented prompting offers a low-overhead method for terminology control in cultural-heritage translation when institutions maintain minimal terminology resources and lightweight evaluation procedures.

Significance. If the results hold beyond the single document, the work supplies concrete human-evaluation evidence that a simple, retrieval-based prompting intervention can measurably improve specialized-term accuracy without degrading overall quality or requiring model-side changes. The direct numeric comparison of terminology accuracy and DA scores on a terminology-dense domain provides a practical data point for low-resource cultural-heritage institutions.

major comments (2)
  1. [Abstract] Abstract and evaluation description: the study reports results from only a single Spanish rock art academic text. This single-instance design is load-bearing for the claim that glossary-augmented prompting constitutes a reliable, low-overhead solution for cultural-heritage institutions, because no evidence is given that the 12.3-point terminology-accuracy lift survives changes in terminology density, document length, author style, or sub-domain.
  2. [Abstract] Evaluation protocol (implied in Abstract): no inter-annotator agreement statistics or statistical significance tests are referenced for the DA scores or the exact-match terminology accuracy figures. Without these, it is not possible to assess whether the reported differences (e.g., 81.4% vs. 69.1%) are reliable or could be affected by rater variability or small sample size.
minor comments (2)
  1. The abstract mentions the PEARMUT evaluation framework but provides no details on the number of annotators, annotation guidelines, or how the restricted MQM taxonomy was operationalized for terminology auditing.
  2. Clarify the precise definition of 'exact-match terminology accuracy' (e.g., whether it requires identical surface forms or allows morphological variants) and how glossary terms were selected and retrieved.

Simulated Author's Rebuttal

2 responses · 0 unresolved

We thank the referee for the constructive comments, which help clarify the scope and limitations of our work. We respond to each major comment below.

read point-by-point responses
  1. Referee: [Abstract] Abstract and evaluation description: the study reports results from only a single Spanish rock art academic text. This single-instance design is load-bearing for the claim that glossary-augmented prompting constitutes a reliable, low-overhead solution for cultural-heritage institutions, because no evidence is given that the 12.3-point terminology-accuracy lift survives changes in terminology density, document length, author style, or sub-domain.

    Authors: We agree that the single-document design limits the strength of general claims about reliability across varying conditions. The manuscript presents the work as a targeted case study on a terminology-dense rock art text rather than a broad validation. We will revise the abstract, introduction, and conclusion to explicitly describe the study as an initial case study, qualify the scope of the findings, and note that additional documents would be required to assess robustness to changes in terminology density, length, style, or sub-domain. revision: yes

  2. Referee: [Abstract] Evaluation protocol (implied in Abstract): no inter-annotator agreement statistics or statistical significance tests are referenced for the DA scores or the exact-match terminology accuracy figures. Without these, it is not possible to assess whether the reported differences (e.g., 81.4% vs. 69.1%) are reliable or could be affected by rater variability or small sample size.

    Authors: The evaluation was conducted by a single domain expert to ensure consistency in terminology auditing, which is why inter-annotator agreement is not reported and no statistical significance tests appear in the manuscript. We accept that this constitutes a limitation for assessing rater variability. We will add a limitations paragraph in the evaluation section to describe the single-annotator protocol, its rationale, and the implications for interpreting the observed differences. revision: yes

Circularity Check

0 steps flagged

No significant circularity; paper reports direct empirical measurements

full rationale

The paper is an empirical comparison of three MT systems on one Spanish rock art text, reporting human-evaluated terminology accuracy (exact-match percentages) and Direct Assessment scores. No equations, fitted parameters, predictions derived from models, or derivation chains exist. Claims rest on observed human ratings rather than any reduction to inputs by construction. No self-citation load-bearing steps or ansatz smuggling are present. This is a standard honest non-finding for an experimental study without mathematical derivations.

Axiom & Free-Parameter Ledger

0 free parameters · 1 axioms · 0 invented entities

The paper is a comparative empirical study; it introduces no mathematical derivations, new physical entities, or fitted constants beyond the reported human evaluation scores.

axioms (1)
  • domain assumption Human raters using Direct Assessment and a restricted MQM taxonomy produce reliable and relevant quality judgments for non-specialist readers of rock art translations.
    The central claim that Gemini-RAG improves terminology control rests on these evaluation results being valid.

reviewed 2026-06-30 · how reviews work

0 comments
Cite this review

Pith. "Pith review of AI-assisted cultural heritage dissemination: Comparing NMT and glossary-augmented LLM translation in rock art documents." pith.science (2026). https://pith.science/paper/LYPT5XYY

@misc{pith2026260514679,
  author       = {Pith},
  title        = {Pith review of: AI-assisted cultural heritage dissemination: Comparing NMT and glossary-augmented LLM translation in rock art documents},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/LYPT5XYY}},
  note         = {Machine review of arXiv:2605.14679}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Cultural heritage institutions increasingly disseminate research and interpretive materials globally, but multilingual dissemination is constrained by limited budgets and staffing. In terminology-dense domains such as rock art, translation quality depends on accurate, consistent specialised terms, and small lexical errors can mislead non-specialists and reduce reuse. We compare three English MT setups for a Spanish academic rock art text, focusing on simple, operationally feasible interventions rather than complex model-side modifications: (1) DeepL as a strong NMT baseline, (2) Gemini-Simple (LLM with a basic prompt), and (3) Gemini-RAG (the same LLM with glossary-augmented prompting via term-pair retrieval). Using PEARMUT, we conduct a human evaluation via (i) multi-way Direct Assessment (0--100) and (ii) targeted terminology auditing with a restricted MQM taxonomy. Gemini-RAG yields the highest exact-match terminology accuracy (81.4\%), versus Gemini-Simple (69.1\%) and DeepL (64.4\%), while preserving overall quality (mean DA 85.3 Gemini-RAG vs. 85.2 Gemini-Simple), outperforming DeepL (80.3). These results show that glossary-augmented prompting is a low-overhead way to improve terminology control in cultural-heritage translation if institutions maintain minimal terminology resources and lightweight evaluation procedures.

Figures

Figures reproduced from arXiv: 2605.14679 by Mar\'ia Ferre-Fern\'andez, Vicent Briva-Iglesias.

Figure 1
Figure 1. Figure 1: PEARMUT interface for Task 1 (multi-way DA-style quality rating). For each Spanish source segment, three anonymised system outputs are shown side by side and scored on a 0–100 scale [PITH_FULL_IMAGE:figures/full_fig_p005_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: PEARMUT interface for Task 2 (targeted terminology audit). Each item shows the Spanish source segment, the expected glossary term, and three anonymised system outputs for labeling terminology errors as wrong, missing, or inconsistent. segment-system ratings. For analysis, the paper reports mean DA scores by system, paired per-segment score differences, 95% bootstrap confidence intervals based on 5,000 resa… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

10 extracted references · 10 canonical work pages · 1 internal anchor

  1. [1]

    Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

    Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference. https://arxiv.org/abs/2403.04132v1, March. Chippindale, Christopher. 2001. What are the right words for rock-art in Australia?Australian Archae- ology, 53:12–15. Colace, Francesco, Rosario Gaeta, Angelo Lorusso, Michele Pellegrino, and Domenico Santaniello

  2. [2]

    Core, MQM

    New AI challenges for cultural heritage pro- tection: A general overview.Journal of Cultural Heritage, 75:168–193, September. Core, MQM. 2025. MQM (Multidimensional Quality Metrics). De Lara L´opez, Hugo, Mart´ı Mas Cornell`a, and M´onica Sol´ıs Delgado. 2025. Chronocultural proposal for the Atlanterra Cave (Cadiz, Spain).Rock Art Re- search, 42(2):213–23...

  3. [3]

    Faber Ben´ıtez, Pamela and Clara In´es Lopez Rodriguez

    Latest developments in rock art record- ing: Towards an integral documentation of Levan- tine rock art sites combining 2D and 3D record- ing techniques.Journal of Archaeological Science, 40(4):1879–1889, April. Faber Ben´ıtez, Pamela and Clara In´es Lopez Rodriguez

  4. [4]

    InA Cognitive Linguistics View of Terminology and Spe- cialized Language, pages 9–31

    Terminology and Specialized Language. InA Cognitive Linguistics View of Terminology and Spe- cialized Language, pages 9–31. Mouton de Gruyter, July. F´oris, ´Agota and Andrea Faludi. 2021. The role of documentation and document management in trans- lation and terminology. InLinguistic Research in the Fields of Content Development and Documentation, pages ...

  5. [5]

    https://heritage- standards.museologi.st, April

    Terminologies. https://heritage- standards.museologi.st, April. Forum on Information Standards in Heritage

  6. [6]

    https://heritage- standards.museologi.st, December

    FISH Terminologies. https://heritage- standards.museologi.st, December. Gao, Yuan, Ruili Wang, and Feng Hou. 2023. How to Design Translation Prompts for ChatGPT: An Em- pirical Study, April. Getty Research Institute. 2017. Cultural Ob- jects Name Authority (CONA).https: //www.getty.edu/research/tools/ vocabularies/cona/, November. Getty Research Institute...

  7. [7]

    In Moniz, He- lena, Lieve Macken, Andrew Rufener, Lo¨ıc Barrault, Marta R

    Europeana Translate: Providing multilingual access to digital cultural heritage. In Moniz, He- lena, Lieve Macken, Andrew Rufener, Lo¨ıc Barrault, Marta R. Costa-juss `a, Christophe Declercq, Maarit Koponen, Ellie Kemp, Spyridon Pilos, Mikel L. For- cada, Carolina Scarton, Joachim Van den Bogaert, Joke Daems, Arda Tezcan, Bram Vanroy, and Margot Fonteyne,...

  8. [8]

    In Mazel, Aron, George Nash, and Clive Waddington, editors,Art as Metaphor: The Prehistoric Rock-Art of Britain

    Rock art and dating. In Mazel, Aron, George Nash, and Clive Waddington, editors,Art as Metaphor: The Prehistoric Rock-Art of Britain. Ar- chaeopress. Mellinger, Christopher D., Nicoletta Spinolo, Maureen Ehrensberger-Dow, and Sharon O’Brien. 2025. De- signing studies with naturalistic tasks. InResearch Methods in Cognitive Translation and Interpreting Stu...

  9. [9]

    Terminology

    Terminological competence in translation. Terminology. International Journal of Theoretical and Applied Issues in Specialized Communication, 15(1):88–104, January. Petti, Luigi, Claudia Trillo, and Chiko Ncube. 2020. Cultural Heritage and Sustainable Development Tar- gets: A Possible Harmonisation? Insights from the European Perspective.Sustainability, 12...

  10. [10]

    F1000Research, 12:523, May

    Digital Rock Art: Beyond ’pretty pictures’. F1000Research, 12:523, May. Whitley, David S. 2005.Introduction to Rock Art Re- search. Left Coast Press. Zouhar, Vil´em and Tom Kocmi. 2026. Pearmut: Hu- man Evaluation of Translation Made Trivial, Jan- uary

This paper was first reviewed by grok-4.3 on June 30, 2026.