Pith. sign in

REVIEW 5 major objections 6 minor 3 cited by

Generative AI in Academic Writing: A Comparison of DeepSeek, Qwen, ChatGPT, Gemini, Llama, Mistral, and Gemma

T0 review · 5 major / 6 minor · reviewed 2026-08-08 · deepseek-v4-flash

Pith's one-line read None of the twelve leading large language models tested can produce academic text that simultaneously passes plagiarism, AI-detection, and readability checks, even though all of them preserve the source's meaning.

desk verdict Useful but thin comparison of new LLMs; the semantic-similarity claim needs calibration before it can carry the paper's conclusions. read the letter →

arxiv 2503.04765 v2 pith:2SIRMS6P submitted 2025-02-11 cs.CY cs.HC

classification cs.CYcs.HC
keywords generativeAIacademicwritinglargelanguagemodelsplagiarismdetectionsemanticsimilarityreadabilitydigitaltwinhealthcare
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper sets out to measure whether current large language models can produce original, readable, undetectable academic text. It compares twelve models, including DeepSeek v3, Qwen 2.5 Max, and Qwen3 235B, against ChatGPT, Gemini, Llama, Mistral, and Gemma on two tasks: answering a question about digital twins and paraphrasing forty healthcare-related abstracts. The central finding is that no model passes the full set of checks: paraphrased abstracts trip plagiarism detectors at rates of 9–57%, all outputs are flagged as AI-generated, and readability is rated poor, even though semantic similarity to the source stays high. If correct, the study maps which models are safer or riskier for different academic uses and shows that detection tools currently outperform attempts to evade them.

What carries the argument

The measuring instrument is a five-part evaluation protocol applied to two text-generation tasks. Each model answered "What is Digital Twin?" and wrote a section on "Digital Twin in Healthcare," and each paraphrased the abstracts of forty papers on digital twin healthcare from 2020 to 2023. The outputs were then scored with iThenticate for originality, Quillbot and StealthWriter for AI detectability, Hemingway Editor, Grammarly, and WebFX for readability, and four LLMs (ChatGPT 4o, DeepSeek v3, Qwen 2.5 Max, and Qwen 3 235B) as semantic-similarity judges against the original abstracts. The semantic-similarity step is the piece that produces the paper's positive result that meaning is preserved, while the other four steps produce the cautions about plagiarism, detectability, and readability.

What would settle it

Re-run the paraphrase evaluation with human raters and with evaluator models that were not the generator; if the high similarity scores fall below roughly 90% or diverge sharply across judges, the claim that all models preserve semantic integrity would not stand.

Watch

Extended reading notes

Core claim

The paper's central claim is that twelve current large language models, when asked to produce academic text about digital twins in healthcare, generate content that is semantically faithful to its source but fails the other checks that scholarly publishing relies on. On iThenticate, paraphrased abstracts matched existing text at 9–57%, with ChatGPT 4o mini worst at 57%; question-answer outputs matched at 1–39%, with only Gemini 2.5 Pro (1%) and Qwen 3 235B (7%) inside commonly accepted limits. Two AI detectors flagged essentially all outputs as machine-written, with the least-detected paraphrase still rated 62% AI on one detector. Semantic similarity between paraphrases and originals stayed at or near 90% across all four LLM judge tools. Readability, however, was uniformly poor: every model's output scored "Poor" on Hemingway Editor and low on Grammarly and WebFX, with WebFX text scores ranging from 3.4% to 25.2%.

Load-bearing premise

The study assumes that similarity scores given by LLM evaluators reflect true semantic preservation even when the evaluator is the same model that produced the paraphrase.

Editorial extensions

If this is right

  • Paraphrase tasks are riskier than open-answer tasks: for most models the iThenticate match rate for paraphrased abstracts (up to 57%) exceeds the 10–20% range many institutions tolerate, while some question-answer outputs fall inside it.
  • AI-generated text remains detectable in practice: even the least-detected paraphrase, Llama 2 7B at 62% on one detector, was flagged as mostly machine-written.
  • Semantic fidelity is not the bottleneck: with all models scoring above roughly 85% similarity, meaning preservation is achieved even when the wording is not.
  • Readability is the uniform weakness: every model's output scored "Poor" on Hemingway Editor and low on Grammarly and WebFX, so generated academic text will need human rewriting to be accessible.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The semantic-similarity scores in Table 7 may overstate fidelity because each paraphrase is judged partly by models that include the generator itself; a held-out judge design would test whether the roughly 90% overlap is genuine or a self-preference artifact.
  • The low plagiarism rate of Llama 3.1 8B (9%) is ambiguous: it could mean the model genuinely rephrases, or that its shorter, less detailed output (2,615 words versus 7,341 in the source) simply contains fewer matchable strings; comparing normalized match rates per 100 words would separate these possibilities.
  • A natural extension is to include a human-written paraphrase baseline; if human paraphrases of the same abstracts also score poorly on readability and show moderate match rates, then some of the reported weaknesses are properties of the paraphrase genre rather than of AI.
  • Because the study uses one domain, digital twin healthcare, and a fixed prompt set, the ranking of models could shift with topic and prompt; testing on other disciplines would show whether the cross-model pattern generalizes.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 6 minor

Summary. This manuscript reports an empirical comparison of twelve large language models (ChatGPT 4o, ChatGPT 4o mini, Gemini 2.5 Pro, Gemini 1.5 Flash, Qwen 3 235B, Qwen 2.5 Max, DeepSeek v3, DeepSeek-Coder-v2 16B, Llama 3.1 8B, Llama 2 7B, Gemma 27B, and Mistral 7B) on two academic-writing tasks: answering the question "What is Digital Twin?" and composing a section on "Digital Twin in Healthcare," plus paraphrasing the abstracts of 40 digital twin/healthcare papers. Outputs were assessed with iThenticate for text similarity, Quillbot and StealthWriter for AI detectability, Hemingway, Grammarly, and WebFX for readability, and four LLMs (ChatGPT 4o, DeepSeek v3, Qwen 2.5 Max, Qwen3 235B) as judges for semantic similarity. The main reported findings are that paraphrased abstracts have high similarity rates, question-based answers also often exceed common acceptance thresholds, AI detectors label nearly all outputs as AI-generated, word counts are generally sufficient, semantic similarity is high, and readability is low. The authors conclude that the newer Chinese models are competitive but that plagiarism, detectability, and readability concerns must be addressed before such tools are used in scholarly writing.

Significance. If the methodological gaps were closed, this study would provide a useful broad snapshot of current LLM performance on concrete academic-writing tasks. The longitudinal comparison with the authors' earlier ChatGPT- and Bard-based studies is a genuine strength, and the explicit limitation section shows awareness of several threats to validity. The study also covers a wider set of models than most prior comparisons. However, the semantic-similarity instrument is unvalidated, no raw data or statistical inferential tests are provided, and some abstract-level claims are inconsistent with the paper's own tables. The significance is therefore conditional on substantial methodological revision.

major comments (5)
  1. [Section 3.3; Table 7] The semantic-similarity scores are produced by ChatGPT 4o, DeepSeek v3, Qwen 2.5 Max, and Qwen3 235B acting as judges, but the manuscript never specifies the judge prompt, the response scale, the aggregation over the 40 abstracts, or any validation of this instrument. No negative controls (e.g., unrelated abstract pairs), no independent metric (e.g., BERTScore or human ratings), and no per-pair scores are reported. Because the raw percentages in Table 7 are treated as interval measurements, the Discussion's conclusion that the models "maintain semantic integrity ... and are reliable for this process" is unsupported. This is load-bearing because the abstract's claim that outputs are "semantically accurate content" rests on this table.
  2. [Section 4; Tables 3-6] All quantitative results are reported as aggregate values without per-paper data, standard deviations, confidence intervals, or statistical tests. For n=40 abstracts, differences such as Llama 3.1 8B at 9% versus ChatGPT 4o mini at 57% iThenticate similarity (Table 4) are presented as definitive model rankings, but the reader cannot tell whether these differences are stable or within measurement noise. The authors should provide the underlying data and dispersion measures, or explicitly reframe the claims as descriptive observations rather than comparative findings.
  3. [Abstract; Section 5] The abstract states that "question-based responses also exceeded acceptable levels," but Table 4 reports Gemini 2.5 Pro at 1% and Qwen3 235B at 7% similarity for question-answer outputs, values the Discussion itself calls "acceptable rates for academic world." The blanket statement in the abstract is therefore inconsistent with the paper's own data and should be revised to reflect the range of results.
  4. [Section 3.2; Table 5] The AI-detection results differ strongly across the two tools for the same text (e.g., Qwen 3 235B paraphrase: 54.33% from Quillbot versus 94.54% from StealthWriter; Mistral 7B paraphrase: 80% versus 62%). The paper nonetheless concludes that "all outputs" are identified as AI-generated without calibrating either tool or reporting thresholds. The Limitations section acknowledges possible false positives, but the abstract and Findings do not qualify the claim. The authors should give threshold justifications or report detector-specific scores without overgeneralizing.
  5. [Section 3.2; Section 3.3] The manuscript does not report the exact prompts beyond the two question texts, the generation parameters (temperature, max tokens, sampling), or the dates of access for the cloud-based models. Because LLM outputs vary with these settings and the paper explicitly claims a point-in-time comparison, the absence of this information prevents replication and weakens the comparative rankings. A protocol appendix with prompts, settings, and access dates should be added.
minor comments (6)
  1. [Title/running header] The running title and author line contain spacing errors ("Comparison ofDeepSeek," "Gemm a"); these should be corrected throughout.
  2. [Figure 2 caption] The caption for Figure 2 describes a "scatter plot," but the figure content is a benchmark table; the caption should match the displayed content.
  3. [Table 7] Table 7 uses inconsistent numeric precision (e.g., "91.2%" alongside "89%") and misspells "DeepSeek" as "DeepsSeek" in the header column; use a consistent format and correct the spelling.
  4. [Section 3.3] The semantic-similarity description does not state whether each judge saw the original and paraphrased abstracts side by side or independently, nor whether each judge produced a single score per pair or per sentence; specify the procedure.
  5. [Section 4; Table 4] The term "plagiarism" is used interchangeably with iThenticate's "matching rate"; consider using "text similarity rate" to avoid conflating a match percentage with intentional plagiarism.
  6. [Section 3.3; Table 1] The text says the abstracts were "originally published between 2020 and 2022," but Table 1 includes two 2023 papers (references [69] and [70]); reconcile the statement or the table.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity; the paper is an empirical measurement study, and the LLM-judged semantic-similarity issue is a validity concern rather than a construction-level circularity.

full rationale

The paper's results are measurements, not derivations: plagiarism percentages come from iThenticate, AI-detection rates from QuillBot and StealthWriter, readability from Hemingway, Grammarly, and WebFX, and semantic similarity from LLM judges applied to separately generated texts. No fitted parameter is renamed as a prediction, no result is defined in terms of another result, and no claim reduces to an input by construction. The semantic-similarity evaluation does use LLMs, including some models being evaluated, as judges, and the paper provides no calibration, control pairs, or independent metric; this is a real methodological limitation that weakens the inference that outputs are 'semantically accurate' or 'reliable for this process.' However, that is an epistemic and validity problem, not circularity: the similarity scores are not logically entailed by the conclusion, and the paper does not define semantic preservation as the judges' output. Self-citations to prior work by the same authors describe methodological lineage and the reuse of the same 40 abstracts for consistency, but they are not load-bearing evidence for the current models' performance, and no uniqueness theorem or ansatz is imported from those papers. The plagiarism, AI-detection, word-count, and readability findings are independently supported by external tools and would stand even if the semantic-similarity instrument were invalid. Thus no circular step meets the required standard of quoteable reduction-by-construction, and the appropriate score is 0.

Assumptions & free parameters 0 free parameters · 5 assumptions · 0 invented entities

The paper does not introduce a mathematical model or fitted parameters. Its conclusions rest on the validity of the selected evaluation tools and on the representativeness of single LLM runs.

assumptions (5)
  • domain assumption iThenticate matching rates are a valid indicator of plagiarism.
    Section 3.3 states 'Originality was assessed using the plagiarism detection tool iThenticate' and Table 4 interprets match rates as plagiarism rates.
  • domain assumption AI detection scores from Quillbot and StealthWriter accurately identify AI-generated text.
    Section 3.3 uses these tools to measure 'human-likeness'; Table 5 and the abstract conclude all outputs are AI-generated based on these scores.
  • domain assumption Readability scores from Hemingway, Grammarly, and WebFX reflect the clarity and accessibility of academic text.
    Section 3.3 uses these tools to evaluate 'complexity and clarity'; Table 6 and discussion interpret low scores as insufficient readability.
  • domain assumption LLM-based semantic similarity scores are unbiased measures of semantic preservation.
    Section 3.3 states semantic similarity was analyzed using ChatGPT 4o, DeepSeek v3, Qwen 2.5 Max, and Qwen3 235B; Table 7 builds all conclusions on these scores, including when the evaluator generated the text.
  • domain assumption A single generation per task with default settings is representative of each model's typical output.
    Sections 3.1 and 3.2 describe one generation pass per model without replication; the study treats these single runs as stable measurements.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Generative AI in Academic Writing: A Comparison of DeepSeek, Qwen, ChatGPT, Gemini, Llama, Mistral, and Gemma." pith.science (2026). https://pith.science/paper/2SIRMS6P

@misc{pith2026250304765,
  author       = {Pith},
  title        = {Pith review of: Generative AI in Academic Writing: A Comparison of DeepSeek, Qwen, ChatGPT, Gemini, Llama, Mistral, and Gemma},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/2SIRMS6P}},
  note         = {Machine review of arXiv:2503.04765}
}
read the original abstract

DeepSeek v3, developed in China, was released in December 2024, followed by Alibaba's Qwen 2.5 Max in January 2025 and Qwen3 235B in April 2025. These free and open-source models offer significant potential for academic writing and content creation. This study evaluates their academic writing performance by comparing them with ChatGPT, Gemini, Llama, Mistral, and Gemma. There is a critical gap in the literature concerning how extensively these tools can be utilized and their potential to generate original content in terms of quality, readability, and effectiveness. Using 40 papers on Digital Twin and Healthcare, texts were generated through AI tools based on posed questions and paraphrased abstracts. The generated content was analyzed using plagiarism detection, AI detection, word count comparisons, semantic similarity, and readability assessments. Results indicate that paraphrased abstracts showed higher plagiarism rates, while question-based responses also exceeded acceptable levels. AI detection tools consistently identified all outputs as AI-generated. Word count analysis revealed that all chatbots produced a sufficient volume of content. Semantic similarity tests showed a strong overlap between generated and original texts. However, readability assessments indicated that the texts were insufficient in terms of clarity and accessibility. This study comparatively highlights the potential and limitations of popular and latest large language models for academic writing. While these models generate substantial and semantically accurate content, concerns regarding plagiarism, AI detection, and readability must be addressed for their effective use in scholarly work.

Figures

Figures reproduced from arXiv: 2503.04765 by the authors.

Figure 1
Figure 1. is a scatter plot comparing different AI models based on their MMLU Redux ZeroEval Score (y-axis) and Input API Price per 1M tokens (x-axis) using a logarithmic scale. The chart highlights the performance-to-price ratio, with an optimal range indicated in a shaded area. DeepSeek-V3 is positioned at the top-left corner, marked with a red star, indicating it has a high￾performance score (~89) while maintaining a low i… view at source ↗
Figure 2
Figure 2. Scatter plot comparing different AI models based on their MMLU Redux ZeroEval Score and Input API Price [19]. 2.2 QWEN Models Qwen is trained on a large, diverse dataset that covers both general language usage and specialized domains such as science, mathematics, and programming. This carefully curated dataset selection, combined with domain-specific fine-tuning techniques, enhances Qwen’s capabilities in areas requ… view at source ↗
Figure 3
Figure 3. presents the benchmark comparison across leading models such as Qwen2.5-Max, DeepSeek-V3, Llama-3.1-405B-Inst, GPT-4o-0806, and Claude-3.5-Sonnet-1022. Qwen2.5-Max consistently outperforms other models across Arena-Hard (89.4) and LiveBench (62.2), among other benchmarks [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: Comparison of performance results of Qwen2 and other models[104] 3. MATERIALS and METHOD This study adopts a structured methodology to expand upon previous works [31], incorporating a broader range of large language models (LLMs) and more detailed evaluation metrics. T…

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Let's Get You Hired: A Job Seeker's Perspective on Multi-Agent Recruitment Systems for Explaining Hiring Decisions

    cs.CY 2025-05 conditional novelty 6.0 of 10

    A multi-agent LLM chatbot for job seekers was perceived by 20 interviewed participants as more actionable, trustworthy, and fair than their recalled experiences with traditional hiring methods.

  2. Beyond Traditional Algorithms: Leveraging LLMs for Accurate Cross-Border Entity Identification

    cs.CL 2025-07 reject novelty 3.0 of 10

    A 65-case comparison claims commercial chatbot LLMs are the most accurate for Portuguese entity matching, but the reported false-positive rates contradict the claim.

  3. DeepSeek in Healthcare: A Survey of Capabilities, Risks, and Clinical Applications of Open-Source Large Language Models

    cs.CL 2025-06 conditional

    A narrative review of DeepSeek-R1's healthcare capabilities, risks, and applications, without new experiments.

Reference graph

Works this paper leans on

105 extracted references · 53 canonical work pages · cited by 3 Pith papers

  1. [1]

    Hello GPT-4o

    OpenAI. (2024). "Hello GPT-4o" Retrieved from https://openai.com/index/hello-gpt-4o/

  2. [2]

    GPT -4o mini: advancing cost -efficient intelligence

    OpenAI. (2024). “GPT -4o mini: advancing cost -efficient intelligence”. https://openai.com/index/gpt-4o-mini-advancing-cost-efficient-intelligence/

  3. [3]

    Introducing Gemini 1.5, Google's next -generation AI model

    Google AI. (2024). "Introducing Gemini 1.5, Google's next -generation AI model" Retrieved from https://blog.google/technology/ai/google-gemini-next-generation-model-february-2024/

  4. [4]

    Qwen2.5-Max: Exploring the Intelligence of Large -scale MoE Model

    Qwen Team. (2025). "Qwen2.5-Max: Exploring the Intelligence of Large -scale MoE Model" Retrieved from https://qwenlm.github.io/blog/qwen2.5-max/

  5. [5]

    Introducing DeepSeek -V3

    DeepSeek AI. (2024). "Introducing DeepSeek -V3" Retrieved from https://api- docs.deepseek.com/news/news1226

  6. [6]

    DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence

    DeepSeek AI. (2024). "DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence" Retrieved from https://github.com/deepseek-ai/DeepSeek-Coder-V2

  7. [7]

    Introducing Llama 3.1: Our most capable models to date

    Meta AI. (2024). "Introducing Llama 3.1: Our most capable models to date." Retrieved from https://ai.meta.com/blog/meta-llama-3-1/

  8. [8]

    Llama-2-7b

    Meta AI. (2024). "Llama-2-7b" Retrieved from https://huggingface.co/meta-llama/Llama-2-7b

Show all 105 references
  1. [9]

    Gemma 2 27B

    Gemma Research Team. (2024). "Gemma 2 27B" Retrieved from https://huggingface.co/google/gemma-2-27b

  2. [10]

    Mistral 7B: The best model to date, Apache 2.0

    Mistral AI. (2023). "Mistral 7B: The best model to date, Apache 2.0." Retri eved from https://mistral.ai/news/announcing-mistral-7b/

  3. [11]

    Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., & Steinhardt, J. (2020). Measuring massive multitask language understanding. arXiv preprint arXiv:2009.03300. https://doi.org/10.48550/arXiv.2009.03300

  4. [12]

    Srivastava, A., Rastogi, A., Rao, A., Shoeb, A. A. M., Abid, A., Fisch, A., ... & Wang, G. (2022). Beyond the imitation game: Quantifying and extrapolating the capabilities of language models. arXiv preprint arXiv:2206.04615. https://doi.org/10.48550/arXiv.2206.04615

  5. [13]

    & Guo, D

    Shao, Z., Wang, P., Zhu, Q., Xu, R., Song, J., Zhang, M., ... & Guo, D. (2024). Deepseekmath: Pushing the limits of mathematical reasoning in open language models. arXiv preprint arXiv:2402.03300. https://doi.org/10.48550/arXiv.2402.03300

  6. [14]

    Liu, A., Feng, B., Wang, B., Wang, B., Liu, B., Zhao, C., ... & Xu, Z. (2024). Deepseek-v2: A strong, economical, and efficient mixture -of-experts language model. arXiv preprint arXiv:2405.04434. https://doi.org/10.48550/arXiv.2405.04434 Cite (APA): Aydin, O., Karaarslan, E.,...

  7. [15]

    & Piao, Y

    Liu, A., Feng, B., Xue, B., Wang, B., Wu, B., Lu, C., ... & Piao, Y. (2024). Deepseek -v3 technical report. arXiv preprint arXiv:2412.19437. https://doi.org/10.48550/arXiv.2412.19437

  8. [16]

    Guo, D., Yang, D., Zhang, H., Song, J., Zhang, R., Xu, R., ... & He, Y. (2025). DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning. arXiv preprint arXiv:2501.12948. https://doi.org/10.48550/arXiv.2501.12948

  9. [17]

    & Zhu, T

    Bai, J., Bai, S., Chu, Y., Cui, Z., Dan g, K., Deng, X., ... & Zhu, T. (2023). Qwen technical report. arXiv preprint arXiv:2309.16609. https://doi.org/10.48550/arXiv.2309.16609

  10. [18]

    & Qiu, Z

    Yang, A., Yang, B., Zhang, B., Hui, B., Zheng, B., Yu, B., ... & Qiu, Z. (2024). Qwen2. 5 technical report. arXiv preprint arXiv:2412.15115. https://doi.org/10.48550/arXiv.2412.15115

  11. [19]

    DeepSeek API Docs

    DeepSeek-V3 (2024) Introducing DeepSeek -V3. DeepSeek API Docs. https://api - docs.deepseek.com/news/news1226

  12. [20]

    Hugging Face Web Site

    Hugging Face (2025). Hugging Face Web Site. https://huggingface.co/

  13. [21]

    Olteanu, A. (2025 ). Qwen 2.5 Max: Features, DeepSeek V3 Comparison & More. Datacamp.com https://www.datacamp.com/blog/qwen-2-5-max

  14. [22]

    Github.com OpenAI Simple -evals

    Simple-evals (2025). Github.com OpenAI Simple -evals. https://github.com/openai/simple- evals?tab=readme-ov-file#benchmark-results

  15. [23]

    & Goldblum, M

    White, C., Dooley, S., Roberts, M., Pal, A., Feuer, B., Jain, S., ... & Goldblum, M. (2024). Livebench: A challenging, contamination -free LLM benchmark. arXiv preprint arXiv:2406.19314 . https://doi.org/10.48550/arXiv.2406.19314

  16. [24]

    DeepSeek-V3

    DeepSeek-AI (2024). DeepSeek-V3. https://github.com/deepseek-ai/DeepSeek-V3

  17. [25]

    Qwen2.5 -Max: Exploring the Intelligence of Large -scale MoE Model

    Qwen Team (2025). Qwen2.5 -Max: Exploring the Intelligence of Large -scale MoE Model. https://qwenlm.github.io/blog/qwen2.5-max/

  18. [26]

    Gurney, M. (2024). Llama Model Details. https://github.com/meta - llama/llama/blob/main/MODEL_CARD.md

  19. [27]

    & Scialom, T

    Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., ... & Scialom, T. (2023). Llama 2: Open foundation and fine -tuned chat models. arXiv preprint arXiv:2307.09288. https://doi.org/10.48550/arXiv.2307.09288

  20. [28]

    Llama 2: open source, free for research and commercial use

    Llama (2023). Llama 2: open source, free for research and commercial use. Meta. https://www.llama.com/llama2/

  21. [29]

    G., Hardin, C., Bhupatiraju, S.,

    Team, G., Riviere, M., Pathak, S., Sessa, P. G., Hardin, C., Bhupatiraju, S., ... & Garg, S. (2024). Gemma 2: Improving open language models at a prac tical size. arXiv preprint arXiv:2408.00118. https://doi.org/10.48550/arXiv.2408.00118

  22. [30]

    Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D

    Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., Casas, D. D. L., ... & Sayed, W. E. (2023). Mistral 7B. arXiv preprint arXiv:2310.06825. https://doi.org/10.48550/arXiv.2310.06825

  23. [31]

    (2020, June)

    Aydın, Ö., & Karaarslan, E. (2020, June). Covid-19 Belirtilerinin Tespiti İçin Dijital İkiz Tabanlı Bir Sağlık Bilgi Sistemi. Online International Conference of COVID-19 (CONCOVID)

  24. [32]

    A., Fletcher, D

    Coorey, G., Figtree, G. A., Fletcher, D . F., & Redfern, J. (2021). The health digital twin: advancing precision cardiovascular medicine. Nature Reviews Cardiology, 18(12), 803 -804. https://doi.org/10.1038/s41569-021-00630-4

  25. [33]

    Volkov, I., Radchenko, G., & Tchernykh, A. (2021). Digital twins, internet of things and mobile medicine: a review of current platforms to support smart healthcare. Programming and Computer Software, 47, 578-590. https://doi.org/10.1134/S0361768821080284

  26. [34]

    Shengli, W. (2021). Is human digital twin possible?. Computer Methods an d Programs in Biomedicine Update, 1, 100014. https://doi.org/10.1016/j.cmpbup.2021.100014 Cite (APA): Aydin, O., Karaarslan, E., Erenay, F.S., & Bacanin, N. (2025). Generative AI in Academic Writing :A Co...

  27. [35]

    Garg, H. (2020). Digital twin technology: revolutionary to improve personalized healthcare. Science Progress and Research (SPR), 1(1), 31-34. https://doi.org/10.52152/spr.2020.01.104

  28. [36]

    O., van Hilten, M., Oosterkamp, E., & Bogaardt, M

    Popa, E. O., van Hilten, M., Oosterkamp, E., & Bogaardt, M. J. (2021). The use of digital twins in healthcare: socio-ethical benefits and socio-ethical risks. Life sciences, society and policy, 17, 1-25. https://doi.org/10.1186/s40504-021-00113-x

  29. [37]

    Elayan, H., Aloqaily, M., & Guizani, M. (2021). Digital twin for intelligent context -aware IoT healthcare systems. IEEE Internet of Things Journal, 8(23), 16749 -16757. https://doi.org/10.1109/JIOT.2021.3051158

  30. [38]

    Gupta, D., Kayode, O., Bhatt, S., Gupta, M., & Tosun, A. S. (2021, December). Hierarchical federated learning based anomaly detection using digital twins for smart healthcare. In 2021 IEEE 7th international conference on collaboration and internet computing (CIC) (pp. 16 -25)....

  31. [39]

    (2021, June)

    Zheng, Y., Lu, R., Guan, Y., Zhang, S., & Shao, J. (2021, June). Towards private similarity query based healthcare monitoring over digital twin cloud platform. In 2021 IEEE/ACM 29th International Symposium on Quality of Ser vice (IWQOS) (pp. 1 -10). IEEE. https://doi.org/10.11...

  32. [40]

    Benson, M. (2021). Digital twins will revolutionise healthcare. Engineering & technology, 16(2), 50-53. https://doi.org/10.1049/et.2021.0210

  33. [41]

    R., Kaynak, O., & Yin, S

    Yang, D., Karimi, H. R., Kaynak, O., & Yin, S. (2021). Developments of digital twin technologies in industrial, smart city and healthcare sectors: a survey. Complex Engineering Systems, 1(1), 1-21. https://doi.org/10.20517/ces.2021.06

  34. [42]

    EL Azzaoui, A., Kim, T.W., Loia, V., Park, J.H. (2021). Blockchain-Based Secure Digital Twin Framework for Smart Healthy City. In: Park, J.J., Loia, V., Pan, Y., Sung, Y. (eds) Advanced Multimedia and Ubiquitous Engineering. Lecture Notes in Electrical Engineering, vol 716. Sp...

  35. [43]

    De Maeyer, C., & Markopoulos, P. (2021). Experts’ View on the Future Outlook on the Materialization, Expectations and Implementation of Digital Twins in Healthcare . Interacting with Computers, 33(4), 380-394. https://doi.org/10.1093/iwc/iwac010

  36. [44]

    Angulo, C., Gonzalez-Abril, L., Raya, C., Ortega, J.A. (2020). A Proposal to Evolving Towards Digital Twins in Healthcare. In: Rojas, I., Valenzuela, O., Rojas, F., Herrera, L., Ortuño, F. (eds) Bioinformatics and Biomedical Engineering. I WBBIO 2020. Lecture Notes in Computer...

  37. [45]

    Boată, A., Angelescu, R., & Dobrescu, R. (2021). Using digital twins in health care. UPB Scientific Bulletin, Series C: Electrical Engineering and Computer Science, 83(4), 53-62

  38. [46]

    C., & Anumba, C

    Madubuike, O. C., & Anumba, C. J. (2021). Digital twin application in healthcare facilities management. In Computing in Civil Engineering 2021 (pp. 366 -373). https://doi.org/10.1061/9780784483893.046

  39. [47]

    Voigt, I., Inojosa, H., Dillenseger, A., Haase, R., Akgün, K., & Ziemssen, T. (2021). Digital Twins for Multiple Sclerosis. Frontiers in Immunology, 12. https://doi.org/10.3389/fimmu.2021.669811

  40. [48]

    N., & Zhang, P

    Kamel Boulos, M. N., & Zhang, P. (2021). Digital Twins: From Persona lised Medicine to Precision Public Health. Journal of Personalized Medicine, 11(8), 745. https://doi.org/10.3390/jpm11080745

  41. [49]

    A., & Park, S

    Hussain, I., Hossain, M. A., & Park, S. J. (2021, December). A healthcare digital twin for diagnosis of stroke. In 2021 IEEE International conference on biomedical engineering, computer and information technology for health (BECITHCON) (pp. 18 -21). IEEE. https://doi.org/10.11...

  42. [50]

    M., & Rutka, J

    Thiong’o, G. M., & Rutka, J. T. (2022). Digital Twin Technology: The Future of P redicting Neurological Complications of Pediatric Cancers and Their Treatment. Frontiers in Oncology, 11. https://doi.org/10.3389/fonc.2021.781499

  43. [51]

    U., Koppu, S., Ramu, S

    Alazab, M., Khan, L. U., Koppu, S., Ramu, S. P., M, I., Boobalan, P., Baker, T., Maddikunta, P. K. R., Gadekallu, T. R., & Aljuhani, A. (2023). Digital Twins for Healthcare 4.0—Recent Advances, Architecture, and Open Challenges. IEEE Consumer Electronics Magazine, 12(6), 29 –3...

  44. [52]

    Yu, Z., Wang, K., Wan, Z., Xie, S., & Lv, Z. (2023). FMCPNN in Digital Twins Smart Healthcare. IEEE Consumer Electronics Magazine, 12(4), 66 –73. https://doi.org/10.1109/mce.2022.3184441

  45. [53]

    D., Cai, J., Niyato, D., & Yi, C

    Okegbile, S. D., Cai, J., Niyato, D., & Yi, C. (2023). Human Digital Twin for Personalized Healthcare: Vision, Arch itecture and Future Directions. IEEE Network, 37(2), 262 –269. https://doi.org/10.1109/mnet.118.2200071

  46. [54]

    Subramanian, B., Kim, J., Maray, M., & Paul, A. (2022). Digital Twin Model: A Real -Time Emotion Recognition System for Personalized Healthcare. IEEE Acce ss, 10, 81155 –81165. https://doi.org/10.1109/access.2022.3193941

  47. [55]

    Hassani, H., Huang, X., & MacFeely, S. (2022). Impactful Digital Twin in the Healthcare Revolution. Big Data and Cognitive Computing, 6(3), 83. https://doi.org/10.3390/bdcc6030083

  48. [56]

    Huang, P., Kim, K., & Schermer, M. (2022). Ethical Issues of Digital Twins for Personalized Health Care Service: Preliminary Mapping Study. Journal of Medical Internet Research, 24(1), e33081. https://doi.org/10.2196/33081

  49. [57]

    H., & Brown, K

    Sahal, R., Alsamhi, S. H., & Brown, K. N. (2 022). Personal Digital Twin: A Close Look into the Present and a Step towards the Future of Personalised Healthcare Industry. Sensors, 22(15),

  50. [58]

    Wickramasinghe, N. (2022). The Case for Digital Twins in Healthcare. Digi tal Disruption in Health Care, 59 –65. Cham: Springer International Publishing. https://doi.org/10.1007/978-3-030- 95675-2_5

  51. [59]

    S., & Ferdous, M

    Akash, S. S., & Ferdous, M. S. (2022). A Blockchain Based System for Healthcare Digital Twin. IEEE Access, 10, 50523–50547. https://doi.org/10.1109/access.2022.3173617

  52. [60]

    A., Yang, C., & El Saddik, A

    Ferdousi, R., Laamarti, F., Hossain, M. A., Yang, C., & El Saddik, A. (2022). Digital twins for well-being: an overview. Digital Twin, 1, 7. https://doi.org/10.12688/digitaltwin.17475.2

  53. [61]

    E., Jones, R

    Khan, A., Milne-Ives, M., Meinert, E., Iyawa, G. E., Jones, R. B., & Josephraj, A. N. (2022). A Scoping Review of Digital Twins in the Context of the Covid-19 Pandemic. Biomedical Engineering and Computational Biology, 13. https://doi.org/10.1177/11795972221102115

  54. [62]

    Kleftakis, S., Mavrogiorgou, A., Mavrogiorgos, K., Kiourtis, A., & Kyriazis, D. (2022). Digital Twin in Healthcare Through the Eyes of the Vitruvian Man. Innovation in Medicine and Healthcare, 75–85. https://doi.org/10.1007/978-981-19-3440-7_7

  55. [63]

    T., Omidvari, A

    Mulder, S. T., Omidvari, A. -H., Rueten-Budde, A. J., Huang, P. -H., Kim, K. -H., Bais, B., Rousian, M., Hai, R., Akgun, C., van Lennep, J. R., Willemsen, S., Rijnbeek, P. R., Tax, D. M., Reinders, M., Boersma, E., Rizopoulos, D., Visch, V., & Steegers -Theunissen, R. (2022). ...

  56. [64]

    P., Boopalan, P., Pham, Q

    Ramu, S. P., Boopalan, P., Pham, Q. -V., Maddikunta, P. K. R., Huynh -The, T., Alazab, M., Nguyen, T. T., & Gadekallu, T. R. (2022). Federated learning enabled digital twins for smart cities: Concepts, recent advances, and future directions. Sustainable Cities and Society, 79,...

  57. [65]

    Song, Y., & Li, Y. (2022). Digital Twin Aided Healthcare Facility Management: A Case Study of Shanghai Tongji Hospital. Construction Research Congress 2022, 1145 –1155. https://doi.org/10.1061/9780784483961.120 Cite (APA): Aydin, O., Karaarslan, E., Erenay, F.S., & Bacanin, N....

  58. [66]

    Khan, S., Arslan, T., & Ratnarajah, T. (2022). Digita l Twin Perspective of Fourth Industrial and Healthcare Revolution. IEEE Access, 10, 25732 –25754. https://doi.org/10.1109/access.2022.3156062

  59. [67]

    Pesapane, F., Rotili, A., Penco, S., Nicosia, L., & Cassano, E. (2022). Digital Twins in Radiology. Journal of Clinical Medicine, 11(21), 6553. https://doi.org/10.3390/jcm11216553

  60. [68]

    Ricci, A., Croatti, A., & Montagna, S. (2022). Pervasive and Connected Digital Twins —A Vision for Digital Health. IEEE Internet Computing, 26(5), 26 –32. https://doi.org/10.1109/mic.2021.3052039

  61. [69]

    Sun, T., He, X., & Li, Z. (2023). Digital twin in healthcare: Recent updates and challenges. Digital Health, 9. https://doi.org/10.1177/20552076221149651

  62. [70]

    M., & Berssaneti, F

    Machado, T. M., & Berssaneti, F. T. (2023). Literature review of digital twin in healthcare. Heliyon, 9(9), e19390. https://doi.org/10.1016/j.heliyon.2023.e19390

  63. [71]

    Aydın, Ö., & Karaarslan, E. (2022). OpenAI ChatGPT Generated Literature Review: Digital Twin in Healthcare. In Ö. Aydın (Ed.), Emerging Computer Technologies 2 (pp. 22 -31). İzmir Akademi Dernegi. https://doi.org/10.2139/ssrn.4308687

  64. [72]

    Aydin, Ö., & Karaarslan, E. (2023). Is ChatGPT Leading Generative AI? What is Beyond Expectations? Academic Platform Journal of Engineering and Smart Systems, 11(3), 118 –134. https://doi.org/10.21541/apjess.1293702

  65. [73]

    Dergaa, I., Chamari, K., Zmijewski, P., & Ben Saad, H. (2023). From human writing to artificial intelligence generated text: examining the prospects and potential threats of ChatGPT in academic writing. Biology of Sport, 40(2), 615–622. https://doi.org/10.5114/biolsport.2023.125623

  66. [74]

    Bom, H.-S. H. (2023). Exploring the Opportunities and Challenges of ChatGPT in Academic Writing: a Roundtable Discussion. Nuclear Medicine and Molecular Imaging, 57(4), 165 –167. https://doi.org/10.1007/s13139-023-00809-2

  67. [75]

    Mondal, H., & Mondal, S. (2023). ChatGPT in academic writing: Maximizing its benefits and minimizing the risks. Indian Journal of Ophthalmology, 71(12), 3600 –3606. https://doi.org/10.4103/ijo.ijo_718_23

  68. [76]

    Lingard, L. (2023). Writing with ChatGPT: An Illustration of i ts Capacity, Limitations & Implications for Academic Writers. Perspectives on Medical Education, 12(1), 261 –270. https://doi.org/10.5334/pme.1072

  69. [77]

    M., Wardat, Y., & Fidalgo, P

    Jarrah, A. M., Wardat, Y., & Fidalgo, P. (2023). Using ChatGPT in academic writing is (not) a form of plag iarism: What does the literature say? Online Journal of Communication and Media Technologies, 13(4), e202346. https://doi.org/10.30935/ojcmt/13572

  70. [78]

    P., & Botchu, R

    Ariyaratne, S., Iyengar, K. P., & Botchu, R. (2023). Will collaborative publishing with ChatGPT drive academi c writing in the future? British Journal of Surgery, 110(9), 1213 –1214. https://doi.org/10.1093/bjs/znad198

  71. [79]

    Gruda, D. (2024). Three ways ChatGPT helps me in my academic writing. Nature. https://doi.org/10.1038/d41586-024-01042-3

  72. [80]

    Zohery, M. (2024). ChatGPT in Academic Writing and Publishing: A Comprehensive Guide. Içinde Artificial Intelligence in Academia, Research and Science: ChatGPT as a Case Study (ss. 10- 61). Novabret Publishing. https://doi.org/10.5281/zenodo.12703273

  73. [81]

    X., Zhou, K., Li, J., Ta ng, T., Wang, X., Hou, Y.,

    Zhao, W. X., Zhou, K., Li, J., Ta ng, T., Wang, X., Hou, Y., ... & Wen, J. R. (2023). A survey of large language models. arXiv preprint arXiv:2303.18223. https://doi.org/10.48550/arXiv.2303.18223

  74. [82]

    Minaee, S., Mikolov, T., Nikzad, N., Chenaghlu, M., Socher, R., Amatriain, X., & Gao, J. (2024). Large language models: A survey. arXiv preprint arXiv:2402.06196. https://doi.org/10.48550/arXiv.2402.06196 Cite (APA): Aydin, O., Karaarslan, E., Erenay, F.S., & Bacanin, N. (2025...

  75. [83]

    Blevins, T., Gonen, H., & Zettlemoyer, L. (2022). Prompting language models for linguistic structure. arXiv preprint arXiv:2211.07830. https://doi.org/10.48550/arXiv. 2211.07830

  76. [84]

    C., Wang, B., & Kuo, C

    Wei, C., Wang, Y. C., Wang, B., & Kuo, C. C. J. (2023). An overview on language models: Recent developments and outlook. arXiv preprint arXiv:2303.05759. https://doi.org/10.48550/arXiv. 2303.05759

  77. [85]

    A., Howard, F

    Gao, C. A., Howard, F. M., Markov, N. S., Dyer, E. C., Ramesh, S., Luo, Y., & Pearson, A. T. (2023). Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewers. Npj Digital Medicine, 6(1). https://doi.org/10.1038/s41746-023-00819-6

  78. [86]

    Stokel-Walker, C. (2023). ChatGPT listed as author on research papers: many scientists disapprove. Nature, 613(7945), 620–621. https://doi.org/10.1038/d41586-023-00107-z

  79. [87]

    Hu, G. (2023). Challenges for enforcing editorial policies on AI -generated papers. Accountability in Research, 31(7), 978–980. https://doi.org/10.1080/08989621.2023.2184262

  80. [88]

    R., & Rachid, L

    Elali, F. R., & Rachid, L. N. (2023). AI-generated research paper fabrication and plagiarism in the scientific community. Patterns, 4(3), 100706. https://doi.org/10.1016/j.patter.2023.100706

  81. [89]

    L., Perle, S

    Anderson, N., Belavy, D. L., Perle, S. M., Hendricks, S., Hespanhol, L., Verhagen, E., & Memon, A. R. (2023). AI did not write this manuscript, or did it? Can we trick the AI text detector into generated texts? The potential future of ChatGPT and AI in Sports & Exercise Me...

  82. [90]

    Artificial Intelligence (AI)

    Springer (2025). Artificial Intelligence (AI). Springer Editor ial Policies. https://www.springer.com/gp/editorial-policies/artificial-intelligence--ai-/25428500

  83. [91]

    E., Magnus, D

    Kaebnick, G. E., Magnus, D. C., Kao, A., Hosseini, M., Resnik, D., Dubljević, V., Rentmeester, C., Gordijn, B., & Cherry, M. J. (2023). Editors’ statement on the responsible use of generative AI technologies in scholarly journal publishing. Medicine, Health Care a nd Philosoph...

  84. [92]

    AI Policy

    Taylor & Francis (2025). AI Policy. Taylor & Francis Policies. https://taylorandfrancis.com/our-policies/ai-policy/

  85. [93]

    Generative AI policies for journals

    Elsevier (2025). Generative AI policies for journals. Elsevier Policies.https://www.elsevier.com/about/policies-and-standards/generative-ai-policies-for-journals

  86. [94]

    APA Journals policy on generative AI: Additional guidance

    American Psychological Association (2023). APA Journals policy on generative AI: Additional guidance. APA Journals Publishing Resource Center https://www.apa.org/pubs/journals/resources/publishing-tips/policy-generative-ai

  87. [95]

    Artificial Intelligence (AI)

    Natura Portfolio(2025). Artificial Intelligence (AI). Editorial Policies. https://www.nature.com/nature-portfolio/editorial-policies/ai

  88. [96]

    Artificial Intelligence Policy

    Sage (2025). Artificial Intelligence Policy. Sage Editorial Policies. https://us.sagepub.com/en- us/nam/artificial-intelligence-policy

  89. [97]

    Misra, D. P. (2022). A Practical Guide for Plagiarism Detection Through the Coordinated Use of Software and (Human) Hardware. Journal of Gastrointestinal Infections, 12(01), 047 –050. CLOCKSS. https://doi.org/10.1055/s-0042-1757486

  90. [98]

    P., & Ravindran, V

    Misra, D. P., & Ravindran, V. (2021). Detecting and Handling Suspected Plagiarism in Submitted Manuscripts. Journal of the Royal College of Physicians of Edinburgh, 51(2), 115 –117. https://doi.org/10.4997/jrcpe.2021.201

  91. [99]

    (2024) How Much Plagiarism Is Allowed? Originality.ai https://originality.ai/blog/how-much-plagiarism-is-allowed

    Jacob, S. (2024) How Much Plagiarism Is Allowed? Originality.ai https://originality.ai/blog/how-much-plagiarism-is-allowed

  92. [100]

    N., Sarcinella, L., Gabriele, L., & Bonifazi, M

    Palmisano, T., Convertini, V. N., Sarcinella, L., Gabriele, L., & Bonifazi, M. (2021). Notarization and Anti -Plagiarism: A New Blockchain Approach. Applied Sciences, 12(1), 243. https://doi.org/10.3390/app12010243 Cite (APA): Aydin, O., Karaarslan, E., Erenay, F.S., & Bacanin...

  93. [101]

    How do I interpret the similarity score? https://support.paperpal.com/support/solutions/articles/3000124505-how-do-i-interpret-the- similarity-score-

    Paperpal Help Center (2024). How do I interpret the similarity score? https://support.paperpal.com/support/solutions/articles/3000124505-how-do-i-interpret-the- similarity-score-

  94. [102]

    How to Use Readability Scores in Your Writing

    Grammarly (2020). How to Use Readability Scores in Your Writing. Grammarly Spotlight. https://www.grammarly.com/blog/product/readability-scores/

  95. [103]

    WebFX (2025) What is Flesch -Kincaid readability? https://www.webfx.com/tools/read- able/flesch-kincaid/

  96. [104]

    https://qwenlm.github.io/blog/qwen3/

    Qwen Team (2025) Qwen3: Think Deeper, Act Faster. https://qwenlm.github.io/blog/qwen3/

  97. [5918]

    https://doi.org/10.3390/s22155918

Pith tools

Reviewed August 8, 2026 · model on record in the stance chip above.