Pith. sign in

REVIEW 3 major objections 5 minor 15 references

LLMs match a trained linguist better than student annotators on subjective Appraisal Attitude labels in TED talks.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-31 17:17 UTC pith:EB247F7Y

load-bearing objection Useful three-way Appraisal annotation study; the binary F1 and single-expert gold slightly oversell the multiclass “aid” claim, but the empirical work is honest and worth refereeing. the 3 major comments →

arxiv 2607.28119 v1 pith:EB247F7Y submitted 2026-07-30 cs.CL cs.SI

Challenges in annotations by humans and LLMs: A case study of evaluative language

classification cs.CL cs.SI
keywords evaluative languageAppraisal theoryAttitudelarge language modelscomplex annotationmachine-supported annotationTED talksinter-annotator agreement
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper asks whether large language models hit the same walls as people when labeling highly subjective evaluative language. Using Appraisal theory’s Attitude categories—Affect, Judgement, and Appreciation—on English TED-talk transcripts, the authors compare twenty-four linguists in training, one senior researcher treated as gold standard, and three prompted LLMs. Students show low agreement with each other and with the senior researcher, especially on Judgement and on implicit or multi-clause sentences. With a carefully chosen prompt, the models reach higher average agreement with the senior researcher than the students do on the Attitude classes; fine-tuning further lifts binary detection of evaluative sentences to an F1 of 0.77. The authors conclude that LLMs can usefully assist complex annotation in digital humanities, but only after case-by-case validation against a human gold standard, because the same categories remain hard for both humans and machines.

Core claim

On sentence-level Appraisal Attitude labeling of TED talks, prompted LLMs achieve higher average agreement with a trained linguist’s gold labels than trained students do, and fine-tuning raises binary evaluativeness detection to F1 0.77; both humans and models still struggle most with Judgement and with nested explicit/implicit meanings.

What carries the argument

A two-stage sentence-level pipeline (binary evaluativeness, then multi-label Affect/Judgement/Appreciation plus Ambiguous) driven by three compared prompts, the best of which is few-shot and theory-close, evaluated against a senior researcher’s labels and optionally QLoRA-finetuned.

Load-bearing premise

That one senior researcher’s sentence-level labels form a stable gold standard for a subjective, multi-label scheme when students are non-native, no annotation iterations occur, and explicit versus implicit spans are not marked.

What would settle it

Re-annotate the same 440-sentence sample with two or more independent trained Appraisal experts who mark spans and explicit/implicit status; if student–expert and model–expert agreements then reverse or collapse relative to the current single-gold results, the central claim fails.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • LLM prompting can scale Attitude annotation for digital-humanities corpora faster and cheaper than student-only pipelines once a gold sample exists.
  • Judgement remains the hardest Attitude class for both humans and models and needs tighter operationalization or span-level schemes.
  • Binary evaluativeness is already strong enough (F1 ~0.77 after fine-tune) to serve as a pre-filter before finer Attitude labeling.
  • Case-by-case validation against human gold remains mandatory; LLMs are not a general panacea for complex pragmatic schemes.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Span- or clause-level annotation with forced explicit/implicit flags would likely raise both human IAA and model F1 and should be the next controlled experiment.
  • A perspectivist multi-label gold set (several plausible labels per sentence) may better match Appraisal’s genuine ambiguity than a single senior label.
  • The same prompt-plus-gold protocol could transfer to other subjective DH schemes such as Engagement, Graduation, or solidarity coding.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper compares Attitude (Affect, Judgement, Appreciation) annotations under Appraisal theory on English TED-talk transcripts by linguists-in-training, one senior researcher treated as gold, and three LLMs. Students annotate 22 talks (sentence-level, multi-label allowed); the senior labels a 440-sentence sample. Three prompts are compared on a held-out subsample; the best is used for Llama, Qwen and Mistral, then Qwen is QLoRA-finetuned. Binary evaluativeness reaches prompted F1 around 0.68 and finetuned F1 0.77; multiclass Attitude is weaker (prompted macro F1 ~0.40; finetuned multi-label micro/macro 0.34/0.30). Models agree more with the senior researcher than trainees do (Table 5), especially on Affect; Judgement is hardest for both. The authors conclude LLMs can aid complex annotation in DH.

Significance. If the comparison is valid, the work supplies a concrete, reproducible case study of LLM vs human performance on a canonical subjective pragmatic scheme, with transparent guidelines, prompt variants, error analysis of target-of-evaluation and explicit/implicit confusions, and public code. That is useful for DH/CSS annotation practice even if absolute IAA remains modest. Strengths include the staged binary-then-multiclass design, domain-wise IAA reporting, and the candid discussion (§8) of sentence-level limits and perspectivism. The headline “aid” claim and abstract F1 0.77, however, rest heavily on binary evaluativeness and a single-expert gold, so significance for fine-grained Attitude resolution is more limited than the abstract suggests.

major comments (3)
  1. [§4.1, Table 5, §7] §4.1 and Table 5: The central claim that models “perform best compared to annotations conducted by the trained linguist” and “aid in complex annotation task resolution” treats one senior researcher’s 440 sentence-level labels as gold for a task the paper itself calls highly subjective and multi-plausible (§1, §8; Cabitza et al. 2023 cited). There is no second expert, no adjudication, and no annotation iterations (contra Fuoli 2018, which the paper cites). Higher model–expert α than trainee–expert α may show style-matching to one annotator rather than superior resolution of Attitude. A multi-expert or perspectivist gold (or explicit framing as agreement with one expert, not accuracy) is needed before the aid claim can carry the weight given in the abstract and §7.
  2. [Abstract, §5.2–5.3, §7] Abstract and §5.3: The reported F1-score of 0.77 is for finetuned binary evaluativeness only. Multiclass Attitude—the theoretically interesting subsystem—remains weak (prompted macro F1 ~0.40; finetuned multi-label micro/macro F1 0.34/0.30, exact-match ~0.04; Judgement F1 0.26–0.32). Presenting 0.77 in the abstract without immediately scoping it to binary overstates performance on the complex task the paper set out to study. Clarify scope in abstract, results, and conclusion, and do not let binary gains stand in for Attitude classification.
  3. [§4.1–4.3, §8] §4.1–4.3 and §8: Sentence-level labelling without forced span marking or explicit/implicit flags produces the very confusions the error analysis documents (target of evaluation, nested clauses, explicit vs invoked Attitude). The authors acknowledge this in §8 but still use the resulting labels as gold for LLM scoring. Either re-annotate a subset at span level with explicit/implicit tags, or restrict claims to “sentence-level overall Attitude as operationalized here” and treat low multiclass performance as partly an artefact of the scheme.
minor comments (5)
  1. [§5.1, Appendix 5] Table 2 / Appendices: Prompt version 3 is said to follow theory closely and to use indicative adjectives/n-grams, but Appendix 5 also shows inconsistent field names (appraisal vs appreciation; judgment spelling) and a thinner definition block than v1. Align the prose description of v3 with the actual appendix text.
  2. [§5.3, Fig. 9] Fig. 9 and §5.3: Finetuning improves binary F1 but sharply drops Affect F1 (0.74→0.17). Briefly discuss why staged multi-label fine-tuning underperforms direct prompting (label co-occurrence, imbalance, supervision only on label tokens).
  3. [§3, §5] §3 / footnote: EmotionalizTED GitHub is given as github.com/anonymous/; replace with the real repository or a stable archival link before publication. Code link in §5 is clearer; keep corpus and code consistent.
  4. [§4.1] Non-native trainee pool and lack of a no-training control are noted (§4.1) but not quantified (e.g., by CEFR or by agreement stratified by proficiency). A short sensitivity note would help readers weigh IAA.
  5. [Abstract, Appendices] Minor typos and wording: “finetune the model, reaching an F1-score of 0.77” (abstract) needs binary scope; “appraisal” vs “Appreciation” in prompt v3; occasional “Judgment”/“Judgement” inconsistency.

Circularity Check

0 steps flagged

Empirical human–LLM comparison; no derivation that equates outputs to inputs by construction

full rationale

This paper is an empirical annotation study: student and senior-linguist labels on TED sentences are compared to prompted and QLoRA-finetuned LLM labels under Appraisal Attitude. Performance numbers (binary F1≈0.77 after finetuning; multiclass macro F1≈0.40 prompted; Krippendorff α in Table 5) are measured against held-out or gold labels, not obtained by renaming a fitted parameter as a prediction. Self-citations (Imamovic 2026; Imamovic et al. 2024, 2025) appear only as related-work background on prior Appraisal–LLM prompting; they do not supply a uniqueness theorem, ansatz, or load-bearing premise that forces the present metrics. Choice of a single senior annotator as gold is a validity threat, not circular math. No self-definitional loop, fitted-input-as-prediction, or renaming of a known result is present. Score 0 is therefore appropriate.

Axiom & Free-Parameter Ledger

3 free parameters · 5 axioms · 1 invented entities

Load-bearing premises are methodological domain assumptions about Appraisal operationalization and evaluation design, not free physical parameters. The central comparative claim rests on treating one expert’s labels as reference, sentence-level multi-label Attitude as the unit, and standard IAA/F1 as success measures.

free parameters (3)
  • Prompt version selection (v3 chosen on subsample F1) = v3 selected at F1=0.49
    Best prompt picked by macro F1 on a small excluded subsample (v3=0.49 vs 0.43/0.24); selection criterion and split size are researcher choices that condition all later LLM rankings.
  • QLoRA fine-tune hyperparameters and 60/40 split = 60/40 split; 4-bit NF4 QLoRA (details partial)
    Train/test split and adapter/quantization settings govern the reported 0.77 evaluative F1 and the multiclass collapse; not uniquely determined by theory.
  • Decision threshold / conservative labeling bias in prompts
    Prompts explicitly encourage conservative decisions, shaping high precision / low recall and Affect-heavy confusions.
axioms (5)
  • domain assumption Martin & White (2005) Attitude tripartition (Affect, Judgement, Appreciation) plus optional Ambiguous is an adequate operational scheme at sentence level.
    Entire annotation pipeline and model targets presuppose this taxonomy and sentence-level aggregation (§2, §4.1).
  • ad hoc to paper A single trained linguist’s labels are the gold standard for scoring both students and LLMs.
    Stated in §4.1; critical because the paper simultaneously argues the task is highly subjective and multi-interpretable (§8).
  • domain assumption Krippendorff’s alpha / Cohen’s kappa and macro F1 are appropriate primary success measures for multi-label subjective categories.
    Used throughout §4–6 without perspectivist or soft-label alternatives until discussion.
  • ad hoc to paper PREVIOUS/NEXT sentence context plus persona constraints suffice for models to recover speaker Attitude without span annotation.
    Prompt design §5.1; later error analysis admits span/explicitness underspecification harms humans and likely models.
  • standard math Standard classification metrics and IAA comparisons are computed under conventional independence assumptions on TED transcript sentences.
    Routine NLP evaluation; not re-derived.
invented entities (1)
  • EmotionalizTED corpus (190 TED transcripts; 22-text annotation subset) independent evidence
    purpose: Domain-balanced spoken popular-science data for Attitude annotation experiments.
    Compiled by authors; full release deferred to another publication—dataset is a contribution but not an ontological postulate.

pith-pipeline@v1.2.0-daily-grok45 · 26472 in / 3575 out tokens · 78566 ms · 2026-07-31T17:17:29.058576+00:00 · methodology

0 comments
read the original abstract

In this paper, we draw a comparison between linguists in training, a trained linguist, and annotations generated by large language models (LLMs) to find out if they struggle with complex linguistic phenomena in a similar way. For this purpose, we analyse evaluative language in spoken popular science discourse, with the example of a corpus of English TED talk transcripts. We focus on the Appraisal theory and its Attitude subsystem, including the categories (classes) of Affect, Judgement, and Appreciation. In this context, Appraisal theory is an example of a highly subjective annotation task, making it a suitable example for the study of complex annotation challenges. First, we assess human annotations on a sentence level in specific scientific domains. Then, we develop three prompts and compare them for model performance for the automatic classification of Appraisal classes. We assess the performance of three LLMs using the best-performing prompt and finetune the model, reaching an F1-score of 0.77. We find that models perform best compared to annotations conducted by the trained linguist, while linguists in training do not reach high agreement scores. We conclude that LLMs can aid in complex annotation task resolution, opening new pathways for the complex theories annotated and analyzed in digital humanities studies.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

15 extracted references · 1 linked inside Pith

  1. [1]

    However, annotator 2 classified it as non-evaluative. Even though the sentence looks purely factual, due to the subjective interpretation on the part of the given annotator, it might be deemed as implicitly offering an Appreciation, as evidenced by at least two annotators agreeing on it. Both linguists in training have an excellent command of English, whi...

  2. [2]

    For version 1 prompt, ambiguity is avoided

    In version 3, we provide indicative adjectives and n-grams instead of real examples. For version 1 prompt, ambiguity is avoided. Next to this, version 1 is a Chain-of-thought prompting approach For testing all prompts, the model qwen3-30b-a3b-instruct-25075 (Qwen) is employed, as it performs particularly well on reasoning tasks (Qwen Team, 2025). All mode...

  3. [3]

    Section 7 concludes, and Section 8 contains a discussion of problematic issues and limitations, as well as the following steps

    While RQ1 is addressed in Section 4, we follow with RQ2 and RQ3 in Sections 5 and 6, respectively. Section 7 concludes, and Section 8 contains a discussion of problematic issues and limitations, as well as the following steps. 2 Appraisal Theory for the Annotation of Evaluative Language Previous studies annotated evaluative language using Appraisal theory...

  4. [4]

    Uncertain

    if the sentence expresses multiple Attitudes (e.g. Affect and Judgement). 22 Notes Use only when you select "Uncertain." Explain briefly what you're unsure about (e.g. “Not sure if it's emotion or factual description”). Use the reasons 1-4 under Uncertain to explain why. Table 6 Annotation instructions for labelling Attitude in text Appendix 2 Prompt for ...

  5. [5]

    Affect was supposed to be used if the sentence expresses feelings or emotions (e.g

    For the annotation of Attitude, we used the following classes: Affect (emotions), Judgement (ethics), and Appreciation (aesthetics). Affect was supposed to be used if the sentence expresses feelings or emotions (e.g. happy, scared, grateful), Judgement if it evaluates behaviour or character of people (e.g. kind, unfair, brave) and Appreciation if things, ...

  6. [6]

    The assistance of ChatGPT 5.0

    All models mentioned were used via the SAIA API. The assistance of ChatGPT 5.0. was used to perform preliminary data coding for the creation of graphs. All content generated by the tool was checked and approved by the authors. All code and output data underlying this article are available on GitHub (https://github.com/happy522/Challenges-in-annotations-by...

  7. [9]

    Do NOT invent extra information [...]

    We craft the first version of the prompt (Appendix 3), building on the learnings from Hamilton et al. (2024), Hou et al. (2024) and Cruickshank et al. (2025), who worked with complex annotation schemes in the fields of linguistics or social science coding. In the end, the prompt involves a persona description and chain-of-thought prompting, as it elicits ...

  8. [10]

    Instead, we focus on agreement between humans and LLMs

    We do not conduct domain-wise interpretation of the results, as this is beyond the scope of this paper. Instead, we focus on agreement between humans and LLMs. Evaluative Affect Judgement Appreciation Linguists in training 0.38 0.23 0.19 0.20 Linguists in training and senior researcher 0.30 0.25 0.16 0.25 Qwen and senior researcher 0.35 0.40 0.23 0.33 Lla...

  9. [12]

    Without knowing if annotators see a sentence as evaluative because of explicit or implicit Attitude, we are unfortunately not able to differentiate this

    that the lowest agreement is normally achieved for segments containing implicit Attitude, which the authors attributed to the inherently subjective and context-dependent nature of such evaluations. Without knowing if annotators see a sentence as evaluative because of explicit or implicit Attitude, we are unfortunately not able to differentiate this. Our q...

  10. [15]

    Despite everything, she stood tall and smiled

    – multi-label annotation is allowed (when you are confident that different Attitude types are expressed, as in e.g. “Despite everything, she stood tall and smiled.”) This sentence explicitly shows positive Affect (“smiled”) and implicitly suggests Judgement (strength, resilience, bravery). - If you genuinely cannot decide which category is primary, use "a...

  11. [49]

    https://doi.org/10.1515/9783110223972 Dong, M., & Fang, A

    Berlin: De Gruyter Mouton. https://doi.org/10.1515/9783110223972 Dong, M., & Fang, A. (2023). Appraisal theory and the annotation of speaker-writer engagement. In Proceedings of the 19th Joint ACL-ISO Workshop on Interoperable Semantics (ISA-19), pp. 18–26, Nancy, France. Association for Computational Linguistics. Eggins, S., & Slade, D. (1997). Analysing...

  12. [1960]

    and Krippendorff’s alpha (Krippendorff, 2011). Next, we conduct an error analysis to answer our first research question: Which Attitude categories are challenging for human annotators and what are the underlying problems? 4.1 Annotation Process and Guidelines The annotations for this study were done in a class throughout the course of one semester. The st...

  13. [2015]

    Still, we consider some aspects in the design of the qualitative study as reasons for low agreement

    and collective discussion (Muller et al., 2021). Still, we consider some aspects in the design of the qualitative study as reasons for low agreement. First, we find that the document level of annotation could be selected differently. While annotators operated on a sentence level in this study, our evidence suggests that a span level of annotation could be...

  14. [2018]

    However, difficult examples were discussed in class together

    in terms of discussions between annotators were not conducted. However, difficult examples were discussed in class together. The time spent on the annotation practice in class was considered sufficient for the annotators to be equipped with enough knowledge and perform the classification of evaluative meanings. Since training and recruiting annotators can...

  15. [2022]

    They reported a mean PABAK agreement score of 0.69

    contributed new annotation guidelines for creative metaphors using Appraisal theory in film reviews. They reported a mean PABAK agreement score of 0.69. However, among the labels used, the involved (implicit) evaluation showed the lowest agreement. The authors attributed this to the inherently subjective and context-dependent nature of such evaluations. P...