Pith. sign in

REVIEW 4 major objections 3 minor 3 references

Automated scoring of the Ambiguous Intentions Hostility Questionnaire using fine-tuned large language models

T0 review · 4 major / 3 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read Fine-tuned large language models can automate AIHQ scoring, matching trained human raters.

desk verdict The abstract promises an AIHQ scoring study, but the supplied manuscript is a completely different paper on educational question generation—the claimed results are simply not in the text. read the letter →

arxiv 2508.10007 v1 pith:WTOFGQI6 submitted 2025-08-05 cs.CL stat.ME

classification cs.CLstat.ME
keywords hostileattributionbiasAmbiguousIntentionsHostilityQuestionnairelargelanguagemodelsfine-tuningautomatedscoringtraumaticbraininjuryaggressionclinicalassessment
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tries to establish that large language models, after fine-tuning on human-rated examples, can score the open-ended portions of the Ambiguous Intentions Hostility Questionnaire (AIHQ) as reliably as trained human raters. The AIHQ is a standard measure of hostile attribution bias, but its open-ended items on perceived intentions and likely responses demand labor-intensive manual coding. If the claim holds, automated scoring could remove a bottleneck in research and clinical work with the AIHQ, including studies of populations such as people with traumatic brain injury. The reported results indicate that fine-tuned model ratings align with human ratings for both hostility and aggression items, do so across ambiguous, intentional, and accidental scenarios, replicate known group differences, and generalize to an independent nonclinical dataset.

What carries the argument

The load-bearing mechanism is supervised fine-tuning on human-scored AIHQ responses. Human raters' scores on the questionnaire's hostility-attribution item and aggression-response item serve as training labels; a fine-tuned language model maps participant free text to the same ratings. The fine-tuning step is what the paper credits for the alignment, and the comparison against held-out human ratings is what carries the validity argument.

What would settle it

Score a fresh set of AIHQ responses with both the fine-tuned models and an independent panel of trained human raters; the central claim fails if model-human agreement is no better than chance on the ambiguous scenarios, or if the model scores do not reproduce the known TBI-versus-control difference in that fresh sample.

Watch

Extended reading notes

Core claim

The central claim is that fine-tuned LLMs produce AIHQ hostility and aggression ratings aligned with trained human raters, and that this alignment is stable across ambiguous, intentional, and accidental scenario types. The support is an experiment: half of a previously collected, human-rated AIHQ dataset from individuals with traumatic brain injury and healthy controls was used to fine-tune two LLMs; the other half was scored by the models, and the model-generated ratings were compared with the existing human ratings. The authors report that model-generated ratings aligned with human ratings for both hostility and aggression items, that fine-tuning improved alignment relative to the base mod

Load-bearing premise

The whole pipeline rests on the human ratings being reliable enough to serve as ground truth; if those labels are noisy or biased, the model learns the noise and alignment with those same ratings is not evidence of true measurement validity.

Editorial extensions

If this is right

  • AIHQ scoring can be offloaded to fine-tuned LLMs, reducing the time and cost of studies using this instrument.
  • Model scores preserve scenario-type distinctions and reproduce known TBI-versus-control differences, so existing research designs could be run at larger scale.
  • Because the fine-tuned models generalized to an independent nonclinical dataset, the approach is not confined to one clinical sample.
  • The reported local and cloud scoring interface could move automated scoring from a research procedure into routine clinical use.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the alignment is genuine, the same fine-tuning recipe could be transferred to other open-ended clinical instruments with structured rating rubrics, but each transfer would need separate validation.
  • Editorial caveat: the full text supplied with this manuscript is a different paper on educational question generation, not the AIHQ study; the abstract's claims therefore rest on the title and abstract alone in the material reviewed here.
  • Automated scoring can inherit the biases and blind spots of the human ratings it is trained on, so agreement with human raters should not be read as ground truth about another person's intentions.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 3 minor

Summary. The paper under review, arXiv:2508.10007, is claimed in the abstract to evaluate fine-tuned LLMs for scoring open-ended AIHQ responses, using a prior TBI/HC dataset, held-out evaluation, scenario-type consistency, group-difference replication, and independent generalization. The supplied full text, however, is arXiv:2508.10005v1 ('From Answers to Questions: EQGBench'), a Chinese educational question-generation benchmark. The body contains no mention of AIHQ, hostile attribution bias, traumatic brain injury, human hostility/aggression ratings, fine-tuning, or the scoring interface. Consequently, every empirical claim in the abstract is unsupported by the submitted manuscript.

Significance. Automated AIHQ scoring, if validated, would be valuable for clinical research and practice. The abstract's design—training on half and testing on a held-out half—is methodologically sensible and non-circular in principle. However, the submitted artifact provides no data, no numeric agreement metrics, no model names, no sample sizes, no fine-tuning protocol, and no generalization analysis. The EQGBench content is a different paper and does not bear on the claimed result. Thus the significance of the claimed result cannot be assessed from the submitted manuscript. Because the supporting study is absent, I cannot credit the paper with reproducible code, machine-checked proofs, or parameter-free derivations; no such elements appear.

major comments (4)
  1. [Full text / Abstract mismatch] The full text is entirely the EQGBench paper. Sections 1–6, Tables 1–3, and the references concern educational question generation for Chinese middle-school STEM and never mention AIHQ, TBI, hostility, aggression, or human ratings of open-ended responses. This is not a missing appendix or a local error; the claimed AIHQ scoring study is absent from the manuscript. Any verdict on the abstract's central claims cannot be grounded in the provided text.
  2. [Abstract (quantitative support)] The abstract reports that 'model-generated ratings aligned with human ratings' and that fine-tuned models showed 'higher alignment,' but it gives no agreement coefficients (e.g., ICC, Pearson r, Cohen's kappa), no sample sizes, no model names, no error bars, and no significance tests. The body also contains no such numbers for the AIHQ task. Even if the correct full text were intended, the abstract alone is insufficient evidence for the consistency, replication, and generalization claims.
  3. [Methodology (missing details)] The described train/holdout protocol is non-circular in principle, but operational details are absent: how responses were split, whether the held-out set was used for any prompt/threshold selection, which models were fine-tuned, how ratings were aggregated, and whether human-rater reliability was reported. The abstract's weakest assumption—that the human ratings used as labels and ground truth are reliable—is not addressed by any inter-rater reliability statistic in the supplied text. Without these details, leakage or label-noise effects cannot be ruled out.
  4. [Results (group-difference and generalization claims)] The claimed replication of TBI-versus-HC group differences and generalization to an independent nonclinical dataset require group-level analyses (means, standard deviations, inferential tests, dataset descriptions). None of these appear. The EQGBench tables contain LLM scores for question-generation dimensions, not hostility or aggression ratings, so they cannot serve as evidence for the abstract's replication or generalization claims.
minor comments (3)
  1. [Title/author consistency] The body's title, author list, and affiliation (Beijing Normal University, EQGBench) do not match the abstract's claimed AIHQ automated-scoring study. If this is an upload error, the intended manuscript must be resubmitted; the current body is a different paper.
  2. [References] The reference list contains no citations to the AIHQ instrument or its validation, no description of the hostility/aggression scoring rubric, and no references to the prior TBI/HC dataset. These would be essential even in a correct submission.
  3. [Scoring interface] The abstract promises an accessible scoring interface with local and cloud-based options, but no link, screenshot, code repository, or usage description appears anywhere in the manuscript.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity identifiable: supplied full text is arXiv:2508.10005 (EQGBench), not the claimed AIHQ automated-scoring paper, so the central derivation chain is absent rather than circular.

full rationale

The abstract describes a held-out evaluation design: half of previously collected, human-rated AIHQ responses are used to fine-tune two LLMs, and the remaining half are used for testing. As described, this is structurally non-circular because the test responses are not used for fine-tuning. However, the supplied full text is not the AIHQ paper at all: it is arXiv:2508.10005, 'From Answers to Questions: EQGBench for Evaluating LLMs’ Educational Question Generation.' The full text contains no AIHQ dataset, no fine-tuning protocol, no hostile-attribution or aggression ratings, no TBI/HC group comparison, and no independent nonclinical generalization dataset. Therefore, the claimed derivation chain cannot be walked, and there is no quotable equation, fitted parameter, or self-citation chain that exhibits a reduction of the claimed result to its inputs. Under the hard rule that circularity may only be claimed when the specific reduction can be exhibited from the paper text, the correct finding is no circularity (score 0), with the caveat that the provided manuscript does not contain the evidence needed to assess the AIHQ paper's validity.

Assumptions & free parameters 0 free parameters · 2 assumptions · 0 invented entities

The abstract alone does not expose any explicit free parameters or invented entities. The main external inputs are the human ratings and the dataset itself.

assumptions (2)
  • domain assumption Human ratings of AIHQ open-ended responses are a reliable and valid ground truth for hostility and aggression.
    The abstract treats trained human rater scores as the reference standard for both fine-tuning labels and evaluation; if these ratings are noisy or biased, the model alignment is not meaningful.
  • domain assumption The previously collected TBI/HC dataset is representative and free of sampling or item-level bias.
    The abstract relies on this dataset for both training and testing; its representativeness underpins the reported generalization.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Automated scoring of the Ambiguous Intentions Hostility Questionnaire using fine-tuned large language models." pith.science (2026). https://pith.science/paper/WTOFGQI6

@misc{pith2026250810007,
  author       = {Pith},
  title        = {Pith review of: Automated scoring of the Ambiguous Intentions Hostility Questionnaire using fine-tuned large language models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/WTOFGQI6}},
  note         = {Machine review of arXiv:2508.10007}
}
read the original abstract

Hostile attribution bias is the tendency to interpret social interactions as intentionally hostile. The Ambiguous Intentions Hostility Questionnaire (AIHQ) is commonly used to measure hostile attribution bias, and includes open-ended questions where participants describe the perceived intentions behind a negative social situation and how they would respond. While these questions provide insights into the contents of hostile attributions, they require time-intensive scoring by human raters. In this study, we assessed whether large language models can automate the scoring of AIHQ open-ended responses. We used a previously collected dataset in which individuals with traumatic brain injury (TBI) and healthy controls (HC) completed the AIHQ and had their open-ended responses rated by trained human raters. We used half of these responses to fine-tune the two models on human-generated ratings, and tested the fine-tuned models on the remaining half of AIHQ responses. Results showed that model-generated ratings aligned with human ratings for both attributions of hostility and aggression responses, with fine-tuned models showing higher alignment. This alignment was consistent across ambiguous, intentional, and accidental scenario types, and replicated previous findings on group differences in attributions of hostility and aggression responses between TBI and HC groups. The fine-tuned models also generalized well to an independent nonclinical dataset. To support broader adoption, we provide an accessible scoring interface that includes both local and cloud-based options. Together, our findings suggest that large language models can streamline AIHQ scoring in both research and clinical contexts, revealing their potential to facilitate psychological assessments across different populations.

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

3 extracted references · 3 linked inside Pith

  1. [2019]

    ��������, arXiv:1909.01953

    Mixture content selection for diverse sequence generation. ��������, arXiv:1909.01953. Bryan R Christ, Jonathan Kropko, and Thomas Hartvigsen. 2024. Mathwell: Generating educa- tional math word problems using teacher annotations. ��������, arXiv:2402.15861. Li Dong, Nan Yang, Wenhui Wang, Furu Wei, Xi- aodong Liu, Yu Wang, Jianfeng Gao, Ming Zhou, and Hsi...

  2. [2024]

    ��������, arXiv:2309.12546

    Automatic answerability evaluation for ques- tion generation. ��������, arXiv:2309.12546. Bin Xu, Yu Bai, Huashan Sun, Yiguan Lin, Sim- ing Liu, Xinyue Liang, Yaolin Li, Yang Gao, and Heyan Huang. 2025. Edubench: A comprehensive benchmarking dataset for evaluating large language models in diverse educational scenarios. ��������, arXiv:2505.16160. Diyi Yan...

  3. [2025]

    Ruslan Mitkov and 1 others

    Can large language models meet the challenge of generating school-level questions? ��������� ��� ���������� ��������� ������������, 8:100370. Ruslan Mitkov and 1 others. 2003. Computer-aided generation of multiple-choice tests. In ����������� �� ��� ��������� �� �������� �� �������� ������ ������ ������������ ����� ������� �������� �������� ���, pages 17–...

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.