Pith. sign in

REVIEW 4 major objections 4 minor 8 references

Unveiling tortured phrases in Humanities and Social Sciences

T0 review · 4 major / 4 minor · reviewed 2026-08-08 · deepseek-v4-flash

Pith's one-line read The paper claims that paraphrased 'tortured' phrases, well-known in STEM literature, also appear in humanities and social sciences, and that an extended screening tool has flagged 32 documents in education, psychology, and economics.

desk verdict A modest, honest extension of the tortured-phrase toolkit into the social sciences; the fingerprint set is useful, but precision is unproven and the 32-document count is a screening output, not a prevalence estimate. read the letter →

arxiv 2502.04944 v1 pith:IAZTKSTX submitted 2025-02-07 cs.DL

classification cs.DL
keywords torturedphrasesabbreviationspapermillsscientificintegrityhumanitiesandsocialsciencestextparaphrasingProblematicScreenerliteraturedecontamination
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper sets out to establish that the paraphrased nonsense phrases known as 'tortured phrases' — which have been widely detected in science, technology, engineering, and mathematics papers — also appear in humanities and social sciences (HSS) literature. The authors generate 121 new 'tortured abbreviations' by paraphrasing terms from two social-science thesauri with a text-spinning tool, add them to the existing Problematic Paper Screener list, and scan bibliometric records plus two open-access corpora. They report 32 flagged documents in education, psychology, and economics, and make four public review comments inviting expert assessment. If the claim holds, it extends a literature-integrity screening method to a domain where vocabulary is less standardized and where paper-mill contamination had been less visible.

What carries the argument

The central object is the 'tortured abbreviation' fingerprint: a paraphrased expanded form that no longer matches its attached abbreviation, such as 'non-administrative associations (NGOs)' instead of 'non-governmental organizations'. The method works by taking abbreviations from two social-science thesauri, paraphrasing them with a text-spinning tool to produce candidate fingerprints, adding the unaltered ones to the Problematic Paper Screener, and querying full-text indexes for documents containing them. Because HSS vocabulary is less standardized, a manual curation step is essential; the paper documents a high false-positive rate and derives additional filtering rules from it.

What would settle it

Have social-science experts conduct a blind review of the 32 flagged documents, looking for independent paper-mill indicators such as fabricated references, duplicate images, or nonsensical methodology; if the majority show no such indicators, the fingerprints would be exposed as unreliable markers of fraud in HSS.

Watch

Extended reading notes

Core claim

The central claim, stated in the abstract and detailed in the results, is that tortured abbreviations occur in HSS publications and that a screening approach built for STEM can be adapted to find them. Using 121 fingerprints generated from thesaurus terms, the authors matched 543 bibliometric documents, and after filtering and manual checking arrived at 32 problematic documents across education, psychology, and economics. Mining one open-access repository produced 9,322 candidate 'tortured' abbreviations, but manual review of 5,048 showed the majority to be false positives such as foreign institution names and reversed word orders; five additional problematic documents came out of that review. The paper also proposes new filtering rules to reduce such false positives in future automation.

Load-bearing premise

The load-bearing premise is that paraphrasing thesaurus terms with an automated text-spinning tool produces fingerprints that match the actual paraphrased text of paper-mill articles, and that a match reliably marks a document as problematic; the authors' own manual review, which found most automated matches in one repository to be false positives, shows this premise requires human verification.

Editorial extensions

If this is right

  • The fingerprint list of the Problematic Paper Screener grows by 121 entries, so future scans will catch these paraphrased HSS abbreviations.
  • Four of the flagged documents now carry public review comments, making them available for domain experts to reassess.
  • The manual review of thousands of candidate matches shows that automated screening of HSS literature must account for legitimate foreign institution names and word-order variants to avoid false positives.
  • The discovery of flagged documents in education, psychology, and economics suggests that paper-mill output has reached social-science venues and warrants expert scrutiny in those fields.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The same fingerprint-generation strategy could be applied to other controlled vocabularies (for example, legal or medical taxonomies) to expand screening coverage faster, since the method only requires a thesaurus and a paraphrasing tool.
  • A detector that incorporates a whitelist of legitimate institution names and cross-language expansions would likely cut the high false-positive rate seen in repository mining.
  • If paper-mill paraphrasing uses generic synonym substitution, other detectable artifacts such as unnatural collocations should co-occur with tortured abbreviations; this could lead to machine-learning models that flag documents rather than relying solely on exact fingerprint matches.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The paper extends the Problematic Paper Screener (PPS) to the humanities and social sciences (HSS). The authors generate 121 "tortured abbreviations" by spinning ELSST and THESOZ thesaurus terms with SpinBot, add them to the PPS fingerprint list, and screen Dimensions, the Hindawi EDRI journal, and the GESIS SSOAR repository. They report 32 flagged documents (26 from PPS screening, 1 from EDRI, 5 from SSOAR), four new PubPeer comments, and several filtering rules for the TPTK detector. The HSS-specific fingerprint dataset is deposited on Zenodo.

Significance. If the 32 documents are genuine paper-mill products, the result is a useful extension of PPS coverage into HSS and a concrete demonstration that tortured abbreviations appear beyond STEM. The paper is commendably transparent about its workflow and reports a large manual validation effort (5,048 SSOAR matches), with an honest acknowledgment that most flagged matches are false positives. The release of 121 fingerprints and the explicit invitation for HSS domain experts to review the flagged documents are practical contributions. However, the validity of the central numeric claim depends on precision information that is not yet provided.

major comments (4)
  1. [Materials and method] The 121 fingerprints are produced by spinning ELSST and THESOZ abbreviations with SpinBot, but the manuscript offers no evidence that SpinBot-generated variants are representative of the paraphrasing found in real paper-mill texts in HSS; the later SSOAR results in Table 2 (9,322 flagged, 5,048 checked, majority false positives, only 5 problematic) suggest that abbreviation-dissimilarity alone has very low positive predictive value in this domain, and the PPS stage that produced 26 of the 32 documents reports no precision figure at all.
  2. [Results and discussion, Table 2] The manual validation covered only 5,048 of the 9,322 SSOAR documents flagged as featuring 'tortured' abbreviations, leaving 4,274 matches unexamined; the abstract nevertheless states 'a total of 32' without an 'at least' qualifier, so the headline count is a lower bound, not a complete enumeration.
  3. [Results and discussion] The classification of a document as 'problematic' is not defined operationally; no decision rule (e.g., number of distinct fingerprints, context verification, domain-expert adjudication) is given, and no inter-annotator agreement or independent confirmation is reported for the 26 documents found via PPS screening, making the 32-document claim non-reproducible.
  4. [Results and discussion and Conclusion] The paper states that the case studies 'highlighted new filtering rules to be implemented through the TPTK tortured abbreviations detector,' and that these rules improve precision, but none of the rules are described; without specifying them, this methodological contribution cannot be evaluated or adopted by other TPTK users.
minor comments (4)
  1. [Abstract] The phrase 'A small amount of unscrupulous people' should read 'A small number of unscrupulous people'; the current wording is non-standard.
  2. [Materials and method] The acronym ANZRC is incorrect; the standard is ANZSRC (Australian and New Zealand Standard Research Classification).
  3. [Introduction] The name 'Texeira da Silva' is misspelled in the text in two places; the reference list correctly spells 'Teixeira da Silva.'
  4. [Table 1] Table 1 would benefit from a column indicating whether each tortured phrase originates from the thesauri-spinning procedure or from the existing PPS list.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: fingerprints are generated from external thesauri before screening, and the flagged documents are manually validated against the definition of a tortured abbreviation.

full rationale

The paper's derivation chain is not circular. The 121 new fingerprints are produced from the ELSST and THESOZ thesauri by SpinBot spinning before any document screening ('we extracted all the abbreviations contained in the ELSST (n=60) and THESOZ (n=75) thesauri, we spun them using SpinBot, then we filtered out the unaltered abbreviations'), so they are not fitted to the documents they later flag. The 42 existing HSS fingerprints come from the prior PPS list and are independent of the screened corpora. The 26 PPS-derived documents, the 1 Hindawi EDRI article, and the 5 SSOAR documents were all subject to manual assessment or validation; the SSOAR screen even shows the authors discarding the majority of 5,048 checked automated matches as false positives. A tortured abbreviation is defined as a mismatch between an abbreviation and its expanded form, so manually finding such a mismatch in a document is direct evidence, not a prediction forced by the fingerprint's origin. The main caveat is precision, not circularity: the authors acknowledge that 'some of the matched abbreviations may still be false positives as they may have different meanings given the HSS field of research.' The one near-circular element is that the case-study screens produced new filtering rules from the same documents later recommended for future screening, but the central claims (32 documents and 121 fingerprints) do not depend on those rules. Self-citations to Clausse et al. (2023, 2025) and Cabanac et al. (2021, 2022) supply the PPS tool and dataset, but the generation method is described in-paper and the tortured-abbreviation concept has independent coverage (O'Grady, 2024), so these citations are not load-bearing circular justification.

Assumptions & free parameters 1 free parameters · 4 assumptions · 0 invented entities

The central screening rests on the assumptions listed: SpinBot-generated fingerprints correspond to real paper-mill language, PPS and Dimensions indexing is reliable, and manual review can separate genuine concepts from tortured phrases. The paper itself provides partial evidence for the last point but does not validate the first. No new theoretical entities are introduced.

free parameters (1)
  • Selection threshold for 'problematic' = at least one fingerprint match per document, followed by manual filtering
    The paper flags documents containing at least one generated tortured abbreviation, then manually filters; this threshold is chosen by the authors and affects the 32-document count.
assumptions (4)
  • domain assumption Tortured abbreviations, defined as phrases that do not match their expansions, are reliable indicators of paper-mill paraphrasing.
    Used throughout Materials and method and Results; the authors concede many matches were false positives, so this assumption only holds partially.
  • ad hoc to paper SpinBot-generated variants of ELSST and THESOZ abbreviations approximate the paraphrasing found in real paper-mill HSS texts.
    Introduced in Materials and method: the 121 fingerprints are produced by spinning thesaurus abbreviations with SpinBot; no external evidence links these specific machine outputs to actual paper-mill products.
  • domain assumption The Dimensions database full-text search and ANZRC 2020 classification correctly identify HSS documents and their textual content.
    Used in Materials and method to screen and filter documents; the paper does not validate recall or precision of this indexing step.
  • domain assumption Genuine HSS concepts such as 'civil war' and 'internal war' can be distinguished from tortured variants by manual review.
    Noted in Motivation; manual review is the final arbiter, but no inter-rater reliability or formal criteria are provided.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Unveiling tortured phrases in Humanities and Social Sciences." pith.science (2026). https://pith.science/paper/IAZTKSTX

@misc{pith2026250204944,
  author       = {Pith},
  title        = {Pith review of: Unveiling tortured phrases in Humanities and Social Sciences},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/IAZTKSTX}},
  note         = {Machine review of arXiv:2502.04944}
}
read the original abstract

A small amount of unscrupulous people, concerned by their career prospects, resort to paper mill services to publish articles in renowned journals and conference proceedings. These include patchworks of synonymized contents using paraphrasing tools, featuring tortured phrases, increasingly polluting the scientific literature. The Problematic Paper Screener (PPS) has been developed to allow articles (re)assessment on PubPeer. Since most of the known tortured phrases are found in publications in science, technology, engineering, and mathematics (STEM), we extend this work by exploring their presence in the humanities and social sciences (HSS). To do so, we used the PPS to look for tortured abbreviations, generated from the two social science thesauri ELSST and THESOZ. We also used two case studies to find new tortured abbreviations, by screening the Hindawi EDRI journal and the GESIS SSOAR repository. We found a total of 32 multidisciplinary problematic documents, related to Education, Psychology, and Economics. We also generated 121 new fingerprints to be added to the PPS. These articles and future screening have to be investigated by social scientists, as most of it is currently done by STEM domain experts.

Figures

Figures reproduced from arXiv: 2502.04944 by the authors.

Figure 1
Figure 1. The research quality insurance workflow, describing how an analyst looks for tortured [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Word cloud of terms and phrases from the fingerprints list. [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

8 extracted references · 6 canonical work pages

  1. [1]

    &Wise,N.(2025).‘Stampoutpapermills’–sciencesleuthsonhowtofightfakeresearch

    Abalkina,A.,Aquarius,R.,Bik,E.,Bimler,D.,Bishop,D.,Byrne,J.,Cabanac,G.,Day,A.,Labbé,C. &Wise,N.(2025).‘Stampoutpapermills’–sciencesleuthsonhowtofightfakeresearch. Nature, 637(8048)(pp.1047-1050).DOI: https://doi.org/10.1038/d41586-025-00212-1

  2. [2]

    & Stell, B.-M

    Barbour, B. & Stell, B.-M. (2020). PubPeer: scientific assessment without metrics. In Biagioli, M. & Lippman,A.(Eds.), Gaming the Metrics: Misconduct and Manipulation in Academic Research(pp. 149-155).TheMITPress.DOI: https://doi.org/10.7551/mitpress/11087.003.0015

  3. [3]

    & Lippman, A

    Biagioli, M. & Lippman, A. (Eds.). (2020).Gaming the Metrics: Misconduct and Manipulation in Academic Research.TheMITPress.DOI: https://doi.org/10.7551/mitpress/11087.001.0001

  4. [4]

    Cabanac, G. (2022). Decontamination of the scientific literature. arXiv preprint: https://arxiv.org/abs/2210.15912

  5. [5]

    Cabanac, G. (2024). Chain retraction: how to stop bad science propagating through the literature. Nature,632(8027)(pp.977-979).DOI: https://doi.org/10.1038/d41586-024-02747-1 Cabanac,G.,Labbé,C.&Magazinov,A.(2021).Torturedphrases:adubiouswritingstyleemergingin science. Evidence of Critical issues affecting established journals. arXiv preprint: https://arx...

  6. [6]

    & Mayr, P

    Clausse, A., Badalova, F., Cabanac, G. & Mayr, P. (2025). Tortured Phrases from the Humanities and Social Sciences – A fingerprints dataset [Data set]. Zenodo. DOI: https://zenodo.org/records/14753785

  7. [7]

    Nazarovets, S. (2024). Dealing with research paper mills, tortured phrases, and data fabrication and falsificationinscientificpapers.InP.-B.Joshi,P.-P.Churi&M.Pandey(Eds.), ScientificPublishing Ecosystem: An Author-Editor-Reviewer Axis (pp. 233-254). Springer Nature. DOI: https://doi.org/10.1007/978-981-97-4060-4_14

  8. [8]

    Tortured phrases

    O'Grady, C. (2024). Software that detects 'tortured acronyms' in research papers could help root out misconduct. Science.DOI: https://doi.org/10.1126/science.znqe1aq TeixeiradaSilva,J.-A.(2021).Atorturedphraseclaimsheterosexualityofthecarbonstructure. Results in Physics,30.DOI: https://doi.org/10.1016/j.rinp.2021.104842 Teixeira da Silva, J.-A. (2023). “T...

Pith tools

Reviewed August 8, 2026 · model on record in the stance chip above.