{"id":"0f8cf354-78d3-42e7-93d8-5d9f73ed804f","arxiv_id":"2505.24681","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A qualitative content analysis of 53 academic influencer videos identifies an input-process-output pipeline for ChatGPT-assisted academic writing, and the study generalizes this into a proposed generative knowledge production framework.","lead":"This study analyzes 53 YouTube tutorials by academic influencers, watched by 5.3 million viewers, to describe how ChatGPT is taught as a step-by-step workflow for academic publishing. It proposes a three-phase human-AI pipeline and policy suggestions for universities and publishers navigating generative AI.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The three-phase pipeline may be an artifact of a sampling frame and coding prompt that already assume step-by-step structure; no reliability or exclusion data are reported.","rationale":"The reader's weakest assumption was sampling representativeness, and I agree that the sampling logic is a serious problem. However, the more load-bearing issue is the combination of a sampling criterion that selects for step-by-step tutorials and a coding prompt that explicitly instructs the model to decompose content into input, process, and output. This makes the central empirical finding vulnerable to circularity: the 'discovered' pipeline is baked into both the inclusion criteria and the coding instrument. The absence of any reliability check, exclusion counts, or audit trail means the paper cannot currently distinguish a genuine bottom-up norm from a projection of the authors' analytic frame. The viewership discrepancy between the abstract (5.3 million) and Section 7 (nearly 80 million) is a concrete red flag that the reported empirical basis is not yet stable. These problems are addressable: a released corpus, preregistered coding, independent blind coding, and transparent screening counts would let a reader assess whether the pipeline is real. Because the central claim may still survive such a test, the appropriate outcome remains conditional rather than rejection. The reader's verdict already captures this, so I recommend no change to the verdict, while sharpening the reason.","tokens_in":10784,"tokens_out":2735,"duration_ms":33895,"concrete_test":"Release the full search and screening log (platform, query terms, date, raw hit count, and number of videos excluded at each criterion), and have two independent coders blind to the study hypothesis open-code a random sample of 20 included and 20 excluded videos without the input/process/output frame. If the trichotomy does not emerge unprompted, or if structured step-by-step tutorials are a small minority of the initially retrieved set, the pipeline is an artifact of selection and coding. In the same revision, reconcile the 5.3 million versus 80 million viewership figures with the raw view counts for each of the 53 videos.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that academic influencer videos exhibit a standardized input-process-output pipeline is not independently established: the method appears to pre-impose the structure it then discovers. Section 4 ('Generating a Code System and Coding for Analysis') states the authors first 'manually tested video structures' and 'identified three primary components in the step-by-step guides: input, process, and output,' and then prompted GPT-4o with exactly those three categories. The sampling frame is also circular in effect: the corpus was restricted to videos that already provided 'step-by-step tutorials,' excluding 'a simple list of instructions, opinion or promotion videos.' Thus the finding that tutorials present a structured pipeline may reflect how videos were selected and coded rather than a property of academic influencer content. The paper reports no inter-coder reliability metric, no count of videos excluded at each screening stage, and no comparison with rejected videos, so it is impossible to estimate what fraction of influencer content actually follows this structure. Additionally, the empirical anchor is internally inconsistent: the abstract reports 53 videos reaching 5.3 million viewers, while Section 7 says these videos reached 'nearly 80 million users.' This unresolved discrepancy further weakens the claim that a detectable, influential, consistent workflow exists.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper analyzes 53 YouTube videos by academic influencers offering step-by-step guidance on using ChatGPT in academic publication. Through qualitative content analysis with GPT-4o-assisted coding, it identifies a three-phase input-process-output 'Generative Knowledge Production Pipeline' and argues that this bottom-up, influencer-driven workflow balances productivity with ethical compliance, challenging top-down institutional policies. The paper proposes policy implications for universities, publishers, and platforms, and presents the pipeline as evidence of a paradigm shift in academic knowledge production.","tokens_in":11076,"tokens_out":2557,"duration_ms":29339,"significance":"If the empirical claims were fully supported, the paper would make a useful contribution by documenting an informal, bottom-up channel through which AI-use norms in academia are being formed, an area that is under-researched. The authors are transparent about using GPT-4o for initial coding and about the manual refinement and validation process, and the use of a real corpus of videos with verbatim quotes provides a tangible empirical anchor. However, the central finding—that academic influencer content exhibits a stable, three-phase pipeline—is currently undermined by methodological gaps in sampling, coding, and reliability reporting, so the significance of the contribution cannot yet be assessed with confidence.","major_comments":[{"comment":"The sampling frame appears to pre-determine the main finding: the corpus was restricted to videos that 'provided step-by-step tutorials,' excluding 'a simple list of instructions, opinion or promotion videos,' so the discovery that the selected videos present a structured input-process-output workflow may be a consequence of the inclusion criteria rather than a property of academic influencer content in general. The manuscript does not report how many videos were screened, how many were excluded at each stage, or any comparison with rejected videos, making it impossible to estimate what fraction of influencer content actually follows the proposed pipeline structure; a detailed screening flow diagram and exclusion counts are needed.","section":"Section 4, 'Sampling and Corpus'"},{"comment":"The coding process is potentially circular: the authors first 'identified three primary components in the step-by-step guides: input, process, and output' and then prompted GPT-4o to break the content into exactly those three categories, so the subsequent finding that videos exhibit this three-phase structure is at least partly an artifact of the coding prompt. The paper should show that these categories emerged from an open, bottom-up coding pass or otherwise demonstrate that the structure is independently present in the data, for example by reporting a separate human open-coding of a subsample without pre-specified categories.","section":"Section 4, 'Generating a Code System and Coding for Analysis'"},{"comment":"No intercoder reliability or independent validation of the AI-generated codes is reported; with two authors and an acknowledged data analyst, a reliability metric such as Cohen's kappa on a subsample of videos, together with the full codebook including category definitions and inclusion/exclusion rules, should be provided to establish that the coding is consistent and reproducible.","section":"Section 4, 'Generating a Code System and Coding for Analysis' and Section 5"},{"comment":"The abstract reports that the 53 videos reached 5.3 million viewers, while Section 7 states the same videos collectively reached 'nearly 80 million users'; this discrepancy spans an order of magnitude and is not reconciled anywhere in the manuscript, and it weakens the credibility of the reported audience-impact figures that form part of the paper's significance argument.","section":"Abstract and Section 7, 'Limitations'"},{"comment":"The claim that 'no negative feedback or criticism indicates that the video audience largely accepts the content' is not supported by any described analysis of the comment sections; if comment sentiment is being used as evidence, the sampling and coding procedure for comments should be specified, and the absence of criticism should not be inferred without systematic analysis.","section":"Section 5, 'Answering RQs'"}],"minor_comments":[{"comment":"The keyword list contains 'ChatPGT,' which appears to be a typo for 'ChatGPT'; this should be corrected.","section":"Keywords"},{"comment":"The references to 'DARWIN19, Bloom20, and KOSMOS21' appear as superscript numbers but no corresponding footnotes or reference entries are visible in the manuscript; please provide proper citations or clarify the notation.","section":"Section 2"},{"comment":"The statement that the corpus is 'appropriately sized' is asserted without a formal saturation justification; please explain how saturation or adequacy of the sample was determined.","section":"Section 4, 'Sampling and Corpus'"},{"comment":"The claim that 'the manual and automated analyses yielded almost identical results' is not quantified; please specify what 'almost identical' means, for example by reporting agreement rates or overlap statistics.","section":"Section 4, 'Generating a Code System and Coding for Analysis'"},{"comment":"Figure 1, the Generative Knowledge Production Pipeline, is referenced but its content is not described in the text and the figure itself is not visible in the manuscript; please ensure the figure is included and its elements are explained in the body.","section":"Figure 1"}],"recommendation":"major_revision","confidential_remarks":"The unresolved mismatch between the 5.3 million and 'nearly 80 million' viewer figures in the abstract versus the limitations section is a red flag for careful fact-checking; even if it is a simple data-entry error, the manuscript should be checked for other numeric inconsistencies. The paper's scope fits a venue that welcomes qualitative studies of academic practice and AI policy, but the current methodological transparency is below the standard needed to support the strong paradigmatic claims."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know two things about this one. First, the raw material is genuinely new: 53 English-language YouTube videos by academic influencers teaching ChatGPT workflows, with transcripts, metadata, and a qualitative analysis. That corpus is worth having. Second, the central claim—that these videos promote a structured input-process-output pipeline—is not independently established; it is baked into the sampling and coding procedure. The paper restricts the corpus to \"step-by-step tutorials,\" manually identifies three components in those guides, then instructs GPT-4o to code with exactly those categories. The pipeline is therefore a property of the selection, not a discovery about influencer content as a whole. That does not make the analysis useless, but it changes what the paper can claim.\n\nWhat the paper does well: the descriptive material on how influencers balance productivity talk with ethical-compliance talk is real. The quotes on rewriting, fact-checking, and originality tools are informative. The reactive-vs-proactive policy table is a reasonable organizer for thinking about institutional responses. The authors also acknowledge obvious limitations—English-only, Global North, small creators excluded—and they note the absence of feedback loops in the workflow, which is a small check on the paradigm-shift rhetoric.\n\nThe soft spots are more than cosmetic. No inter-coder reliability metric, no full coding scheme, no count of excluded videos, no comparison with rejected videos. The viewer discrepancy is a genuine red flag: the abstract says 53 videos reached 5.3 million viewers, while the limitations section says \"nearly 80 million users.\" That is a factor of fifteen, and since the paper's significance argument leans on reach, it needs an explanation. The \"paradigm shift\" framing also outruns the evidence; a descriptive study of a small, curated corpus cannot support that language.\n\nWho is this for? People working on AI policy in higher education, academic integrity, or social-media influence in scholarly practice. They will find the corpus and the three-phase framing a useful starting point, not a finished result.\n\nMy recommendation: send it to peer review, but brace the authors for a major revision. The empirical material deserves referee time, and the methodological flaws are addressable. Require the coding scheme, reliability checks, screening counts, and a consistent viewer figure. Tone down the paradigm-shift claims. Then it could be a solid qualitative contribution.","headline":"A useful empirical corpus undermined by a sampling and coding design that largely pre-ordains the three-phase pipeline it claims to discover.","tokens_in":11483,"tokens_out":1625,"would_cite":false,"duration_ms":20482,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that academic influencer videos teach a consistent three-phase ChatGPT workflow (input, process, output) that functions as an emerging bottom-up standard for AI-assisted publication, and that institutions should…","keywords":["generative AI","ChatGPT","academic influencers","academic integrity","knowledge production pipeline","human-AI co-intelligence","academic policy","YouTube tutorials"],"falsifier":"Take an unfiltered or randomly sampled set of academic ChatGPT tutorials — including videos with under 500 followers, non-English languages, and videos the current study would have excluded as opinions or promotions — code them with the same input/process/output scheme, and compare the prevalence of the three-phase structure. If the structure is substantially weaker or absent in the broader sample, the pipeline is an artifact of how the 53 videos were selected rather than a property of influencer teaching.","tokens_in":10542,"feed_emoji":"🎓","tokens_out":9513,"duration_ms":103794,"temperature":0.7,"pith_summary":"This paper argues that academic influencers on YouTube are teaching a structured, repeatable way to use ChatGPT for scholarly publication, and that this informal instruction is becoming a de facto standard for AI-assisted research. Examining 53 tutorial videos with 5.3 million combined views, the authors identify a three-phase workflow — input, process, and output — in which prompt design, rewriting, plagiarism checks, and human accountability are built into each stage. The paper calls this the Generative Knowledge Production Pipeline and contends that it balances productivity with ethical compliance, making it a proof of concept for proactive university and publisher policies rather than a threat to integrity. If the claim holds, the real evolution of academic AI norms is happening bottom-up through social-media teaching, and formal policy is currently lagging behind it.","feed_headline":"53 influencer videos reveal a structured ChatGPT publishing pipeline","feed_subtitle":"Influencer-driven workflows may be setting academic AI norms before institutions do.","key_machinery":"The Generative Knowledge Production Pipeline (GKPP) is the central object: a three-stage model — input (prompts, tasks, data), process (writing, editing, ethics, workflow efficiency), and output (quality, originality, publication) — derived from coding 53 video transcripts totaling 120,667 words. The pipeline carries the argument by showing that the same structure recurs across independent influencers, with each stage containing safeguards (human oversight, rewriting, plagiarism checks, fact-checking) that the authors read as evidence of a bottom-up ethical standard. The coding itself was produced with ChatGPT (GPT-4o) and then manually validated, which the authors present as a demonstration of the human-AI co-intelligence the pipeline describes.","core_discovery":"The paper's central discovery is that academic influencer videos about ChatGPT for research are not a scattered collection of tips but a consistent, three-phase structure: input (prompts, tasks, and data), process (writing, editing, workflow efficiency, and ethical checks), and output (quality, originality, citation, and publication). The authors argue that this pipeline enforces human oversight at every stage — influencers instruct viewers to refine prompts, rewrite AI-generated text, fact-check, and run plagiarism or originality checks — and therefore represents a proactive, credibility-centered alternative to top-down bans. They further claim that the pipeline is a bottom-up norm-setting mechanism that should be incorporated into institutional policy, and that it exemplifies a broader shift toward human-AI co-intelligence in academic knowledge production.","pith_inferences":["If the pipeline is real, influencer ethics is largely procedural — rewriting and originality checks — which suggests a testable extension: does following the pipeline reduce actual plagiarism, or mainly reduce detection by AI-text classifiers?","The same bottom-up standardization dynamic likely applies to other generative tools and other professions; a comparative study of coding, legal, or medical tutorial channels could reveal whether the input-process-output structure is a general pattern of informal AI guidance.","A direct test of the sampling assumption would be to code a random, broader set of ChatGPT academic tutorials — including small creators and non-English content — to see whether the clean three-phase structure survives outside the curated corpus."],"forward_implications":["Universities and publishers can use the input-process-output pipeline as a proof of concept for proactive AI policy, training modules, and ethics-by-design workflows.","Academic influencers become legitimate intermediaries: institutions can collaborate with them to translate bottom-up practices into official guidelines rather than treating their content as unregulated noise.","The pipeline's next version needs explicit feedback loops and quality-control checkpoints; the authors argue these are underrepresented in the videos but necessary for a scalable framework.","Because the pipeline works across academic writing tasks, it can be adapted beyond academia to professional knowledge work that uses generative AI.","If the pipeline is adopted formally, authorship and evaluation standards will have to shift from banning AI to auditing how humans and AI collaborated at each phase."],"supporting_citations":[{"why":"Establishes AI-assisted literature review as a workflow that generative tools now support, grounding the study's premise of transformed knowledge production.","marker":"Bolanos et al. 2024"},{"why":"Supplies the multidisciplinary framing of ChatGPT's opportunities and challenges in academic research and publishing.","marker":"Dwivedi et al., 2023"},{"why":"Provides the COPE-based, top-down publisher guideline that the influencer-driven bottom-up pipeline is contrasted against.","marker":"Cacciamani et al., 2023"},{"why":"Documents university and publisher positions on AI and academic integrity, defining the reactive policy landscape the paper wants to move beyond.","marker":"Gulumbe, 2024"},{"why":"Frames the paradox of AI as both banned and popular, motivating attention to informal bottom-up adoption.","marker":"Lim et al., 2023"},{"why":"Supplies the co-intelligence concept the paper uses to characterize human-AI collaboration in the pipeline.","marker":"Mollick, 2024"},{"why":"Provides the benefits, challenges, and future-research agenda for ChatGPT in higher education, including human agency and ethics.","marker":"Rasul et al. 2023"}],"fun_headline_variants":["53 influencer videos reveal a 3-phase ChatGPT research pipeline","Academic influencers are setting AI research norms from below","ChatGPT workflow pipeline built by influencers relies on human oversight","Influencer-led pipeline proposes co-intelligence for academic publishing"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the 53 English-language videos chosen through keyword search and manual filtering faithfully represent academic influencer content about ChatGPT, despite the authors' own note that smaller creators, non-English videos, and content excluded as opinion or promotion were left out.","fun_headline_variants_meta":{"raw":{"variants":["53 influencer videos reveal a 3-phase ChatGPT research pipeline","Academic influencers are setting AI research norms from below","ChatGPT workflow pipeline built by influencers relies on human oversight","Influencer-led pipeline proposes co-intelligence for academic publishing"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000197,"raw_usage":{"total_tokens":1308,"prompt_tokens":834,"completion_tokens":474,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":450,"completion_tokens_details":{"reasoning_tokens":408}},"tokens_in":450,"tokens_out":474,"duration_ms":5855,"temperature":1.0,"reasoning_tokens":408,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T12:13:39.659302+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take an unfiltered or randomly sampled set of academic ChatGPT tutorials — including videos with under 500 followers, non-English languages, and videos the current study would have excluded as opinions or promotions — code them with the same input/process/output scheme, and compare the prevalence of the three-phase structure. If the structure is substantially weaker or absent in the broader sample, the pipeline is an artifact of how the 53 videos were selected rather than a property of influencer teaching.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the multidisciplinary framing of ChatGPT's opportunities and challenges in academic research and publishing."}],"review_version":1}