Pith. sign in

REVIEW 3 major objections 4 minor 36 references

Authoring and Management of Transparent Research Integrity Assessments of Randomised Clinical Trial Publications Using LLM-assisted Tools and Provenance Knowledge Graphs

T0 review · 3 major / 4 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read This paper shows that an LLM-assisted, human-in-the-loop workflow can produce transparent research integrity assessments of RCT publications, with provenance captured in a reusable ontology and knowledge graph.

desk verdict Useful open infrastructure for research-integrity screening, but the 86.4% human-AI agreement is anchored by the human-in-the-loop design and should be read as post-exposure concordance, not independent validation. read the letter →

arxiv 2608.07202 v1 pith:QBDMOKAG submitted 2026-08-07 cs.AI

classification cs.AI
keywords ResearchIntegrityLargeLanguageModelOntologyProvenanceKnowledgeGraphRandomisedControlledTrialsINSPECT-SRHuman-in-the-loop
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper claims that an end-to-end pipeline for semi-automated research integrity assessment of published randomised controlled trials now exists and is publicly reusable. The pipeline pairs an LLM-based tool, INSPECT-AI, with a provenance ontology, RIPE-O, and a published knowledge graph, RIPE-KG, containing 140 expert assessments of 95 trial publications. The paper's central empirical finding is that automated and human-reviewed outcomes agree on 86.4% of 514 question pairs, with disagreement concentrated on trial registration checks, where human reviewers themselves also disagree most. A sympathetic reader would take away that transparent, auditable AI-assisted integrity screening is feasible at low cost and that provenance graphs make the reasoning behind each verdict inspectable.

What carries the argument

The load-bearing object is RIPE-O, a provenance ontology that models a research integrity assessment as a collection of investigated questions, evidence pieces, evaluation activities, and hypotheses, with human and automated agents attributed to their respective outputs. RIPE-O's competency questions, framed by the seven W's of provenance, keep the pattern generic so that new integrity questions can be added without changing the model. RIPE-KG is the materialisation of that ontology, currently holding 1,221 hypotheses across 140 assessments, linked to author identities in an external scholarly knowledge graph and exposed through SPARQL queries that can compare automated versus human outcomes, trace rationales, and federate with other scholarly graphs.

What would settle it

Conduct a blinded crossover study: have the same set of publications assessed twice by the same or equivalent reviewers, once with INSPECT-AI suggestions shown and once with them hidden. If the agreement rate between automated and human outcomes remains near 86.4% under blinding, the concordance claim is solid; if it falls substantially, the reported agreement is largely an anchoring artifact.

Watch

Extended reading notes

Core claim

On its own terms, the paper shows that LLM assistance can be embedded in a human-in-the-loop integrity assessment workflow without handing over the decision. INSPECT-AI extracts evidence from a publication's PDF, queries external registries and databases, evaluates each piece of evidence against selected INSPECT-SR checks using conditional rules, and presents suggested yes/no/unclear outcomes that a reviewer confirms or overrides. Every accepted, modified, or overridden decision is logged together with the evidence and rationale, and the log is transformed through YARRRML mappings into RIPE-KG, where each assessment, hypothesis, and evidence item is connected by provenance relations. The reported 86.4% agreement between automated and human-reviewed outcomes, alongside the uneven disagreement across the four implemented checks, supports the paper's argument that documenting provenance is necessary because human assessors themselves disagree, particularly on registration timing.

Load-bearing premise

The paper treats the recorded human review outcomes as independent expert judgments, even though every reviewer saw the automated suggestion before finalising an answer and some reviewers accepted the suggestion without carrying out the extra publisher-website checks.

Editorial extensions

If this is right

  • Evidence synthesis teams can deploy INSPECT-AI to screen candidate RCTs for integrity concerns, with the pilot reporting a marginal cost below $0.10 per paper.
  • Because RIPE-KG links each verdict to its evidence and rationale, systematic reviewers can audit why a publication received a particular integrity outcome rather than treating the verdict as a black box.
  • The disagreement pattern, with 26.2% of automated-human pairs differing on registration checks, identifies exactly where decision support tools need better external data or clearer guidance.
  • RIPE-O's generic provenance pattern can be reused by other integrity assessment tools, letting their outputs be merged into RIPE-KG or comparable knowledge graphs without rebuilding the model.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If a follow-up study has human reviewers record their own answers before seeing INSPECT-AI's suggestions, the 86.4% agreement figure will very likely drop, because the current design lets reviewers anchor on the automated answer and some admitted to doing so unreflectively.
  • The registration-check disagreement probably reflects the rule-based comparison of registration and recruitment dates being too blunt for cases where the reported timeline allows acceptable prospective registration; a more nuanced model of clinical trial practice would reduce noise.
  • As RIPE-KG grows, its author-pair co-authorship counts for serious-concerns publications could be read as a public reputational metric, raising fairness considerations that the paper does not address.
  • The low cost of the pilot suggests that large-scale integrity screening of entire systematic review candidate sets is feasible, which would let evidence synthesists prioritise human review effort rather than expand it.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. This paper presents INSPECT-AI, an LLM-assisted tool that guides human reviewers through research integrity assessments of randomised controlled trials using the INSPECT-SR framework, together with RIPE-O, an ontology for representing the provenance of such assessments, and RIPE-KG, a knowledge graph of 140 assessments of 95 publications with a SPARQL endpoint and web GUI. The core contribution is an end-to-end pipeline: PDF upload, automated evidence aggregation using GROBID and Gemini 2.0 Flash, rule-based suggested outcomes, human confirmation or override, and RDF conversion via YARRRML. The paper reports a pilot deployment with 13 volunteers, ontology validation via OOPS! and SPARQL competency queries, and an analysis of 514 question pairs with automated and human-reviewed outcomes, finding 86.4% agreement, with the lowest agreement for study-registration checks.

Significance. If the infrastructure claims hold, this is a timely and useful contribution to evidence synthesis and metascience. The public availability of the ontology, SPARQL endpoint, mappings, and an explicit LLM guide makes the pipeline inspectable and reusable, and the cost data suggest scalability is plausible. The provenance model's separation of automated and human contributions is a genuine design strength. However, the paper's empirical validation is weaker than the abstract implies: the headline agreement figure is collected in a workflow where reviewers always see the automated suggestion first, and no accuracy metrics against labelled ground truth are reported. The infrastructure and the empirical claim should be judged separately; the former is largely supported, while the latter needs rework.

major comments (3)
  1. [Section 6.1; Figure 1 (Section 4)] The headline agreement statistic of 86.4% (514 pairs, p. 12) is not a measure of independent AI-human concordance. In the workflow shown in Figure 1, the tool presents suggested outcomes to the reviewer (steps 3–4), and the human outcome is recorded only after the reviewer confirms or overrides that suggestion. Every 'human' outcome in RIPE-KG is therefore posterior to exposure to the automated suggestion. The paper's own statement in Section 6.1 that 'some human reviewers sided with the automated INSPECT-AI suggestions without following the additional guidance to check publishers' websites' indicates that anchoring occurred. The 13.6% disagreement rate is consequently a lower bound on true disagreement, not a measured rate. To support the empirical sub-claim, the authors need either a blinded validation study in which reviewers record their answer before seeing the suggestion, or a clear reframing of the statistic as human-in-the-loop workflow agreement rather than concordance.
  2. [Abstract; Section 4.1] The abstract's label '140 expert research integrity assessments' is not supported by the pilot description: Section 4.1 reports 104 traces produced by 13 volunteers, of whom only 61.5% had previously undertaken integrity assessments, with additional assessments contributed by core team members and research sleuths. The manuscript should state how many assessments came from each group and define what qualifies the contributors as 'expert.'
  3. [Section 6.1; Section 3] No accuracy or extraction-quality metric is reported for the automated pipeline. The paper mentions a set of 50 known problematic publications used to guide reviewers (Section 3), but it does not use this or any labelled set to report precision/recall for the automated outcomes, nor does it report error rates for LLM-extracted metadata such as registration IDs and trial dates. Without such benchmarks, the agreement rates cannot be interpreted as evidence that the tool produces correct assessments; they only show that human reviewers often accept the suggestions.
minor comments (4)
  1. [Section 4.1] Section 4.1 reports percentages with small denominators (e.g., '31% reporting n=2 or very frequently n=2' out of 13 participants); please give absolute counts alongside percentages or avoid unnecessary precision.
  2. [Listing 1.3; Section 6] The federated query in Listing 1.3 relies on the SemOpenAlex SPARQL endpoint, which Section 6 notes can be incomplete or unreliable; the paper should state that the example result is illustrative and that reproducibility depends on endpoint availability.
  3. [Section 6.1] The paper should explicitly state that the 86.4% agreement is measured under the human-in-the-loop protocol and is not a blinded concordance rate, so that readers do not overinterpret the figure.
  4. [Figure 3; Section 5] The ontology diagram includes many classes and properties, but the accompanying text does not define every property shown (e.g., ripe:concerns used on multiple classes); consider listing the intended domains and ranges in the ontology documentation and in the paper.

Circularity Check

1 steps flagged · score 5.0 of 10

The 86.4% agreement statistic is anchored by the human-in-the-loop workflow, so the paper's headline agreement does not establish independent AI-human concordance.

  1. self definitional [Section 4 (INSPECT-AI Tool workflow) and Section 6.1 (Analysing Assessment Outputs)]
    "Human reviewers accept, modify, or override these suggestions before submitting their final assessment of the publication... We have analysed 514 assessments×question pairs for which both automated and human-reviewed outcomes are available. Of these, 444 (86.4%) show agreement... some human reviewers sided with the automated INSPECT-AI suggestions without following the additional guidance to check publishers' websites."

    The comparison standard (human-reviewed outcome) is recorded only after the reviewer has seen and may simply accept the automated suggestion, so the automated outcome is an input to the human outcome rather than an independent check of it. The paper explicitly concedes that some reviewers sided with the automated suggestions instead of performing the extra publisher-website checks. The 86.4% agreement is therefore not a measured concordance rate between independent expert judgments and the AI; it is inflated by the workflow's design, and the 13.6% disagreement is a lower bound on true disagreement. The agreement statistic thus cannot support the conclusion that automated outcomes reliably reproduce expert integrity judgments.

full rationale

The paper's infrastructure contributions (RIPE-O ontology, RIPE-KG, SPARQL endpoint, provenance mappings, federated queries) are self-contained and do not reduce to their own inputs; they are demonstrated by the artefacts themselves. However, the empirical sub-claim in Section 6.1, the 86.4% automated/human agreement, is methodologically anchored: the human final answers are produced after the reviewer is shown the automated suggestion and can accept it without further verification, and the paper admits some reviewers did exactly that. Because RIPE-KG also labels as 'expert' assessments produced by pilot volunteers with limited integrity-assessment experience and by the paper's own research-sleuth team, the reference labels are not independent of the system being evaluated. The INSPECT-SR framework [35] has substantial author overlap with the present paper, though it is community-approved and Cochrane-endorsed, so that overlap is not itself load-bearing. The deficiency is confined to the agreement analysis; the provenance and knowledge-graph claims remain supported. Score 5 reflects one prominent self-referential evaluation loop rather than a fully circular derivation.

Assumptions & free parameters 0 free parameters · 4 assumptions · 1 invented entities

This is a systems and data paper, so the ledger contains no fitted numerical parameters. The rule thresholds in the automated checks (for example, registration date after recruitment start date triggers a concern) are hand-set design decisions, documented in the check logic, rather than values fit to data. The load-bearing assumptions are domain assumptions. First, the Cochrane-endorsed INSPECT-SR checklist is treated as the criterion for integrity concerns. Second, external databases (Retraction Watch, PubPeer, ClinicalTrials.gov, OpenAlex) are treated as sufficiently complete, with incompleteness acknowledged by the paper itself. Third, the LLM extraction of dates and registration identifiers is assumed accurate, with no reported precision or recall. Fourth, human outcomes are assumed to be independent judgments, even though reviewers recorded them after seeing the tool's suggestion. The only invented entity, RIPE-O, is the artifact under test and carries independent public evidence.

assumptions (4)
  • domain assumption The INSPECT-SR checklist is a valid and sufficient operationalization of RCT research integrity for evidence synthesis inclusion decisions.
    The paper's automated checks and question set are built directly on INSPECT-SR (Sections 4 and 6.1) without independent validation of the checklist itself; INSPECT-SR was authored by a group that overlaps with the present paper's authors (reference [35]).
  • domain assumption External evidence sources (Retraction Watch Database, PubPeer, ClinicalTrials.gov, WHO ICTRP, OpenAlex) are sufficiently complete and accurate for the checks to be reliable.
    Section 4 relies on these for retraction, comment, and registration evidence; Section 6.1 and Section 7 acknowledge that the Retraction Watch database 'may be incomplete' and that RIPE-KG 'will inherit any insufficiencies' from OpenAlex.
  • domain assumption The LLM's structured extraction of dates, registration identifiers, and trial timeline values is accurate enough for the rule-based suggestions to be meaningful.
    Section 4 states GEMINI 2.0 Flash extracts registration IDs and the study timeline, but no precision or recall figures are reported anywhere; the prospective-registration check compares the extracted recruitment start date with the registry registration date, so an extraction error becomes a wrong suggested outcome.
  • domain assumption The recorded human outcomes in RIPE-KG represent the reviewers' own considered judgment rather than endorsement of the tool's suggestion.
    Section 4's workflow shows the reviewer accepts or overrides the displayed suggestion, and Section 6.1 speculates that some reviewers 'sided with the automated INSPECT-AI suggestions.' The agreement statistics in Section 6.1 require this premise to be interpretable.
invented entities (1)
  • RIPE-O ontology classes (ripe:ResearchIntegrityAssessment, ripe:IntegrityAssessmentQuestion, ripe:IntegrityAssessmentHypothesis, ripe:EvidenceAggregation, and evidence subclasses) independent evidence
    purpose: A semantic vocabulary for publishing and querying provenance traces of research integrity assessments, integrating PROV-O, TIDO, FABIO, and CITO.
    The ontology is the paper's deliverable rather than an explanatory postulate. It is publicly documented at https://w3id.org/ripe/ripe-o, checked with OOPS!, evaluated through SPARQL competency queries, and instantiated in the queryable RIPE-KG, so third parties can inspect and falsify it directly.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Authoring and Management of Transparent Research Integrity Assessments of Randomised Clinical Trial Publications Using LLM-assisted Tools and Provenance Knowledge Graphs." pith.science (2026). https://pith.science/paper/QBDMOKAG

@misc{pith2026260807202,
  author       = {Pith},
  title        = {Pith review of: Authoring and Management of Transparent Research Integrity Assessments of Randomised Clinical Trial Publications Using LLM-assisted Tools and Provenance Knowledge Graphs},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/QBDMOKAG}},
  note         = {Machine review of arXiv:2608.07202}
}
read the original abstract

Systematic reviews of Randomised Controlled Trials (RCTs) are routinely used as evidence for clinical care guidelines. Such evidence has to meet high research integrity standards to prevent low quality or false research outputs influencing the clinical care. However, assessing research integrity of published RCTs is a complex process requiring manual effort, and potentially resulting in diverse opinions of the human assessors. This paper describes INSPECT-AI, an LLM-based interactive tool that assists human reviewers with research integrity assessments of published RCTs based on the community approved INSPECT-SR framework, and the Research Integrity Provenance and Evidence ontology (RIPE-O) for documenting the provenance of the assessment process. In addition, we present the Research Integrity Provenance and Evidence knowledge graph (RIPE-KG), an initial set of 140 expert research integrity assessments of 95 RCT publications generated by INSPECT-AI and described using RIPE-O.

Figures

Figures reproduced from arXiv: 2608.07202 by the authors.

Figure 1
Figure 1. The INSPECT-AI workflow detailing two automated steps for evidence aggre￾gation and initial evaluation, human review step, and finally storage and publishing of provenance traces documenting the assessment process. 4 INSPECT-AI Tool To produce a systematic review, reviewers have to assess whether an RCT pub￾lication should be included. One of the criteria preventing inclusion of otherwise relevant publication is the… view at source ↗
Figure 2
Figure 2. Web interface of INSPECT-AI application. The assessed publication is pre￾sented on the left. Evidence and the results corresponding to individual checks are displayed on the right side of the UI. also contributed additional assessments included in RIPE-KG. Participants used INSPECT-AI independently to assess 69 publications and generated 104 assess￾ment traces (i.e., some publications were assessed by more than one … view at source ↗
Figure 3
Figure 3. The RIPE-O Ontology. relating to the timing or absence of study registration. These questions are an￾swered by hypotheses (ripe:IntegrityAssessmentHypothesis) which are re￾sults of the evaluation activity (tido:Evaluation). The hypotheses are gener￾ated following the evaluation of evidence (ripe:IntegrityAssessmentEvidence) which is generated by the INSPECT-AI’s automated evidence aggregation activ￾ity (ripe:Evidenc… view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: An example partial provenance trace documenting considered evidence and assessment results for question about concerns related to the timing of the trial regis￾tration. 6 Research Integrity Provenance and Evidence Knowledge Graph (RIPE-KG) RIPE-KG contains provenance t…
Figure 5
Figure 5. Figure 5: A partial view of the RIPE-KG web interface. Basic bibliographic information about the publication is displayed alongside three research integrity assessments. The view can be expanded to allow inspection of the individual pieces of evidence considered for each check. …

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

36 extracted references · 27 canonical work pages

  1. [1]

    https://doi.org/10.1016/j.jclinepi.2025.111672

    Au, L.S., Qu, L., Nielsen, J., Ge, Z., Gurrin, L.C., Mol, B.W., Wang, R.: Using artificial intelligence to semi-automate trustworthiness assessment of randomized controlledtrials:acasestudy.JournalofClinicalEpidemiology180,111672(2025). https://doi.org/10.1016/j.jclinepi.2025.111672

  2. [2]

    Accountability in Research31(1), 14–37 (2024).https://doi.org/10.1080/08989621.2022.2082290

    Avenell, A., Bolland, M.J., Gamble, G.D., Grey, A.: A randomized trial alerting authors,withorwithoutcoauthorsoreditors,thatresearchtheycitedinsystematic reviews and guidelines has been retracted. Accountability in Research31(1), 14–37 (2024).https://doi.org/10.1080/08989621.2022.2082290

  3. [3]

    BMJ Evidence-Based Medicine 29(2), 121–126 (2024).https://doi.org/10.1136/bmjebm- 2022- 111921, https://ebm.bmj.com/content/29/2/121

    Bakker, C., Boughton, S., Faggion, C.M., Fanelli, D., Kaiser, K., Schneider, J.: Reducing the residue of retractions in evidence synthesis: ways to minimise in- appropriate citation and use of retracted data. BMJ Evidence-Based Medicine 29(2), 121–126 (2024).https://doi.org/10.1136/bmjebm- 2022- 111921, https://ebm.bmj.com/content/29/2/121

  4. [4]

    ClinicalTrials.gov API,https://clinicaltrials.gov/data-api/api, last ac- cessed 2026/04/30

  5. [5]

    Cochrane: Methods in cochrane.https://www.cochrane.org/authors/methods -cochrane(nd), accessed: 2026-04-27

  6. [6]

    In: The Semantic Web – ISWC 2023

    Färber, M., Lamprecht, D., Krause, J., Aung, L., Haase, P.: Semopenalex: The scientific landscape in 26 billion rdf triples. In: The Semantic Web – ISWC 2023. pp. 94–112. Springer, Cham (2023).https://doi.org/10.1007/978-3-031-472 43-5_6

  7. [7]

    In: International Seman- tic Web Conference

    Garijo, D.: Widoco: a wizard for documenting ontologies. In: International Seman- tic Web Conference. pp. 94–102. Springer, Cham (2017).https://doi.org/10.1 007/978-3-319-68204-4_9,http://dgarijo.com/papers/widoco-iswc2017.pdf

  8. [8]

    Accountability in Research32(4), 488–508 (2025).https://doi.org/10.1080/08989621.2023

    Grey, A., Avenell, A., Bolland, M.J.: Ten years later: Assessments of the integrity of publications from one research group with multiple retractions. Accountability in Research32(4), 488–508 (2025).https://doi.org/10.1080/08989621.2023. 2295996

Show all 36 references
  1. [9]

    evidence and suggested improvements

    Grey, A., Avenell, A., Gaby, A., Bolland, M.J.: Inconsistency in publishers’ re- sponses to integrity concerns about published research. evidence and suggested improvements. Journal of Clinical Epidemiology186(2025).https://doi.org/ 10.1016/j.jclinepi.2025.111918

  2. [10]

    Nature577(7789), 167–169 (2020).https: //doi.org/10.1038/d41586-019-03959-6

    Grey, A., Bolland, M.J., Avenell, A., Klein, A.A., Gunsalus, C.K.: Check for pub- lication integrity before misconduct. Nature577(7789), 167–169 (2020).https: //doi.org/10.1038/d41586-019-03959-6

  3. [11]

    (eds.) European Semantic Web Conference

    Heyvaert, P., De Meester, B., Dimou, A., Verborgh, R.: Declarative rules for linked data generation at your fingertips! In: Gangemi, A., Gentile, A.L., Nuzzolese, A.G., Rudolph, S., Maleshkova, M., Paulheim, H., Pan, J.Z., Alam, M. (eds.) European Semantic Web Conference. pp. ...

  4. [12]

    In: Proceedings of the 10th International Conference on Knowledge Capture

    Jaradeh, M.Y., Oelen, A., Farfar, K.E., Prinz, M., D’Souza, J., Kismihók, G., Stocker, M., Auer, S.: Open research knowledge graph: Next generation infrastruc- ture for semantic scholarly knowledge. In: Proceedings of the 10th International Conference on Knowledge Capture. p. ...

  5. [13]

    ArXiv abs/2301.10140(2023),https://api.semanticscholar.org/CorpusID:256194 545

    Kinney, R.M., Anastasiades, C., Authur, R., Beltagy, I., Bragg, J., Buraczynski, A., Cachola, I., Candra, S., Chandrasekhar, Y., Cohan, A., Crawford, M., Downey, D., Dunkelberger, J., Etzioni, O., Evans, R., Feldman, S., Gorney, J., Graham, D.W., Hu, F., Huff, R., King, D., Ko...

  6. [14]

    W3C rec- ommendation, W3C (Apr 2013),http://www.w3.org/TR/2013/REC-prov-o-201 30430/

    Lebo, T., Sahoo, S., McGuinness, D.: PROV-O: The PROV ontology. W3C rec- ommendation, W3C (Apr 2013),http://www.w3.org/TR/2013/REC-prov-o-201 30430/

  7. [15]

    In: Agosti, M., Borbinha, J., Kapidakis, S., Papatheodorou, C., Tsakonas, G

    Lopez, P.: Grobid: Combining automatic bibliographic data recognition and term extraction for scholarship publications. In: Agosti, M., Borbinha, J., Kapidakis, S., Papatheodorou, C., Tsakonas, G. (eds.) Research and Advanced Technology for Digital Libraries. pp. 473–474. Spri...

  8. [16]

    Manghi, P., Bardi, A., Atzori, C., Baglioni, M., Manola, N., Schirrwagen, J., Principe, P.: The openaire research graph data model (Apr 2019).https://do i.org/10.5281/zenodo.2643199

  9. [17]

    Research Integrity and Peer Review8(1), 6 (2023).https://doi.org/10.1186/ s41073-023-00130-8

    Mol,B.W.,Lai,S.,Rahim,A.,Bordewijk,E.M.,Wang,R.,vanEekelen,R.,Gurrin, L.C., Thornton, J.G., van Wely, M., Li, W.: Checklist to assess trustworthiness in randomised controlled trials (TRACT checklist): concept proposal and pilot. Research Integrity and Peer Review8(1), 6 (2023)...

  10. [18]

    AI Matters 1(4), 4–12 (2015).https://doi.org/10.1145/2757001.2757003

    Musen, M.A.: The protégé project: a look back and a look forward. AI Matters 1(4), 4–12 (2015).https://doi.org/10.1145/2757001.2757003

  11. [19]

    In: Blomqvist, E., Hose, K., Paulheim, H., Ławrynowicz, A., Ciravegna, F., Hartig, O

    Nielsen, F.Å., Mietchen, D., Willighagen, E.: Scholia, scientometrics and wikidata. In: Blomqvist, E., Hose, K., Paulheim, H., Ławrynowicz, A., Ciravegna, F., Hartig, O. (eds.) The Semantic Web: ESWC 2017 Satellite Events. pp. 237–259. Springer, Cham (2017).https://doi.org/10....

  12. [20]

    OpenAlex,https://openalex.org/, last accessed 2026/04/27

  13. [21]

    In: The Semantic Web – ISWC 2018: 17th International Semantic Web Conference, Monterey, CA, USA, October 8–12, 2018, Proceedings, Part II

    Peroni, S., Shotton, D.: The spar ontologies. In: The Semantic Web – ISWC 2018: 17th International Semantic Web Conference, Monterey, CA, USA, October 8–12, 2018, Proceedings, Part II. p. 119–136. Springer-Verlag, Berlin, Heidelberg (2018). https://doi.org/10.1007/978-3-030-00668-6_8

  14. [22]

    Minerva 61(2), 147–174 (Jun 2023).https://doi.org/10.1007/s11024-023-09490-3

    Peterson, D., Panofsky, A.: Metascience as a scientific social movement. Minerva 61(2), 147–174 (Jun 2023).https://doi.org/10.1007/s11024-023-09490-3

  15. [23]

    International Journal on Semantic Web and Information Systems (IJSWIS)10(2), 7–34 (2014).https: //doi.org/10.4018/ijswis.2014040102 Title Suppressed Due to Excessive Length 19

    Poveda-Villalón, M., Gómez-Pérez, A., Suárez-Figueroa, M.C.: Oops! (ontology pitfall scanner!): An on-line tool for ontology evaluation. International Journal on Semantic Web and Information Systems (IJSWIS)10(2), 7–34 (2014).https: //doi.org/10.4018/ijswis.2014040102 Title Su...

  16. [24]

    Engineer- ing Applications of Artificial Intelligence111, 104755 (2022).https://doi.org/ 10.1016/j.engappai.2022.104755

    Poveda-Villalón, M., Fernández-Izquierdo, A., Fernández-López, M., García- Castro, R.: Lot: An industrial oriented ontology engineering framework. Engineer- ing Applications of Artificial Intelligence111, 104755 (2022).https://doi.org/ 10.1016/j.engappai.2022.104755

  17. [25]

    PubPeer,https://pubpeer.com/, last accessed 2026/04/27

  18. [26]

    In: Pro- ceedings of the First International Conference on Semantic Web in Provenance Management - Volume 526

    Ram, S., Liu, J.: A new perspective on semantics of data provenance. In: Pro- ceedings of the First International Conference on Semantic Web in Provenance Management - Volume 526. p. 35–40. SWPM’09, CEUR-WS.org, Aachen, DEU (2009)

  19. [27]

    Retraction Watch: Meet the scientific sleuths: Ten who’ve had an impact on the scientific literature (Jun 2018),https://retractionwatch.com/2018/06/17/mee t-the-scientific-sleuths-ten-whove-had-an-impact-on-the-scientific-l iterature/, accessed: 2026-03-03

  20. [28]

    Retraction Watch Database,https://retractiondatabase.org/RetractionSea rch.aspx?, last accessed 2026/04/27

  21. [29]

    In: Proceedings of the 13th Knowledge Capture Conference 2025

    Roothaert, R., Schlobach, S., Massacci, F., Stork, L.: Tido: The threat intelligence decision ontology. In: Proceedings of the 13th Knowledge Capture Conference 2025. p. 82–89. K-CAP ’25, Association for Computing Machinery, New York, NY, USA (2025).https://doi.org/10.1145/373...

  22. [30]

    UK Research Integrity Office (UKRIO): Introductory guide to the concordat to support research integrity (2025),https://ukrio.org/research-integrity/the -concordat-to-support-research-integrity/introductory-guide-to-the-c oncordat-to-support-research-integrity/, accessed: 14 April 2026

  23. [31]

    Complex & Intelligent Systems 9(1), 1059–1095 (Feb 2023).https://doi.org/10.1007/s40747-022-00806-6

    Verma, S., Bhatia, R., Harit, S., Batish, S.: Scholarly knowledge graphs through structuring scholarly communication: a review. Complex & Intelligent Systems 9(1), 1059–1095 (Feb 2023).https://doi.org/10.1007/s40747-022-00806-6

  24. [32]

    Weeks, J., Cuthbert, A., Alfirevic, Z.: Trustworthiness assessment as an inclusion criterion for systematic reviews—what is the impact on results? Cochrane Evidence Synthesis and Methods1(10), e12037 (2023).https://doi.org/10.1002/cesm.1 2037

  25. [33]

    Research Synthesis Methods14(3), 357–369 (2023).https://doi.org/10.1002/jrsm.1599

    Weibel, S., Popp, M., Reis, S., Skoetz, N., Garner, P., Sydenham, E.: Identifying and managing problematic trials: A research integrity assessment tool for random- ized controlled trials in evidence synthesis. Research Synthesis Methods14(3), 357–369 (2023).https://doi.org/10....

  26. [34]

    WHO International Clinical Trials Registry Platform (ICTRP) ,https://trials earch.who.int/, last accessed 2026/04/30

  27. [35]

    medRxiv (2025).https: //doi.org/10.1101/2025.09.03.25334905,https://www.medrxiv.org/conten t/early/2025/10/21/2025.09.03.25334905

    Wilkinson, J., Heal, C., Flemyng, E., Antoniou, G.A., Aburrow, T., Alfirevic, Z., Avenell, A., Barbour, V., Berghella, V., Bishop, D.V.M., Bordewijk, E.M., Brown, N.J.L., Christopher, J., Clarke, M., Dahly, D., Dennis, J., Dicker, P., Dumville, J., Frankish, H., Grey, A., Groh...

  28. [36]

    Wilkinson, M.D., Dumontier, M., Aalbersberg, I.J., Appleton, G., Axton, M., Baak, A., Blomberg, N., Boiten, J.W., da Silva Santos, L.B., Bourne, P.E., Bouw- man, J., Brookes, A.J., Clark, T., Crosas, M., Dillo, I., Dumon, O., Edmunds, S., 20 Markovic et al. Evelo, C.T., Finker...

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.